[["Objectives: The ability to perform a context-free 3-dimensional multiple object tracking (3D-MOT) task has been highly related to athletic performance. In the present study, we assessed the transferability of a perceptual-cognitive 3D-MOT training from a laboratory setting to a soccer field, a sport in which the capacity to correctly read the dynamic visual scene is a prerequisite to performance. Design: Throughout pre- and post-training sessions, we looked at three essential skills (passing, dribbling, shooting) that are used to gain the upper hand over the opponent. Method: We recorded decision-making accuracy during small-sided games in university-level soccer players (n = 23) before and after a training protocol. Experimental (n = 9) and active control (n=7) groups were respectively trained during 10 sessions of 3D-MOT or 3D soccer videos. A passive control group (n = 7) did not received any particular training or instructions. Results: Decision-making accuracy in passing, but not in dribbling and shooting, between pre- and post-sessions was superior for the 3D-MOT trained group compared to control groups. This result was correlated with the players' subjective decision-making accuracy, rated after pre- and post-sessions through a visual analogue scale questionnaire. Conclusions: To our knowledge, this study represents the first evidence in which a non-contextual, perceptual-cognitive training exercise has a transfer effect onto the field in athletes. --------------------------------------------------------------------------------","In dynamic sports such as soccer (Association Football), the ability to ‘read the game’ distinguishes skilled from less skilled players (Williams, 2000). However, athletes are not characterized by superior vision (Abernethy, 1987; Helsen & Starkes, 1999) and visual training programs have not shown any evidence of transfer to the field (Wood & Abernethy, 1997). Rather, sport scientists identified a number of abilities that are tightly related to superior anticipation and decision-making to better explain the ability to ‘read the game’. Anticipation and decision-making represents the human brain's ability to extract meaningful contextual information from the visual scene and are essential for high-level performance in sports (Casanova, Oliveira, Williams, & Garganta, 2009). They are typically referred to as perceptual-cognitive skills, illustrating the role played by both perceptual and cognitive processes. Two main approaches have been proposed to identify athletes' perceptual-cognitive superiority. The first and most common theory that supports athletic expertise relies on the expert performance approach. It reflects comparisons between elite, sub-elite and/or novice performers in tasks that are domain specific and, in some cases representative of the behavioral requirements of the competitive setting. In essence, experts have been shown to be superior to sub-elite and/or novices in sports- specific tasks including advance visual cue utilization (Abernethy, Gill, Parks, & Packer, 2001; Ward, Williams, & Bennett, 2002), pattern recall and recognition (Abernethy, Baker, & Côté, 2005; Smeeton, Ward, & Williams, 2004), visual search strategies (Vaeyens, Lenoir, Williams, & Philippaerts, 2007; Williams, 2000) and the knowledge of situational probabilities (North & Williams, 2008; Williams, Hodges, North, & Barton, 2006). Those abilities have been linked to game intelligence. On the other hand, the cognitive component skill approach examines whether sport expertise influences fundamental cognitive and perceptual functions outside the sport-specific domain (Nougier, Stein, & Bonnel, 1991). It relies on more fundamental, sport context-free, paradigms that confer a cognitive fidelity rather than a physical fidelity with the sport environment. In fact, it is well accepted that physical activity enhances brain plasticity and improves cognitive and executive functions (for recent reviews see Erickson, Gildengers, & Butters, 2013; Vivar, Potter, & van Praag, 2013). For example, a significant correlation has been demonstrated between the results from the executive functions tests (neuropsychological assessment tool) versus the number of goals and assists the players had scored two seasons later (Vestberg, Gustafson, Maurex, Ingvar, & Petrovic, 2012). The authors suggested that results in cognitive function tests predict the success of top-soccer players. Furthermore, higher order cognitive function has been suggested to be relevant for talent identification and development in youth soccer players (Verburgh, Scherder, van Lange, & Oosterlaan, 2014). In a recent meta-analysis, Voss and colleagues showed that expertise in sport was related to high levels of performance on measures of processing speed and visual attention (Voss, Kramer, Basak, Prakash, & Roberts, 2010). Moreover, Alves and colleagues found that volleyball players differed from non-athlete controls on two executive control tasks and one visuo-spatial attentional processing task (Alves et al., 2013). Furthermore, interesting significant differences have recently been found between athletes that outperformed non-athletes in socially realistic multitasking crowd scenes involving pedestrians crossing streets (Chaddock, Neider, Voss, Gaspar, & Kramer, 2011) or in learning complex and neutral dynamic visual scenes through a three dimensional multiple object tracking (3D-MOT) task (Faubert, 2013). These studies support the claim that the cognitive component approach captures a fundamental cognitive skill associated with competitive sport training (Voss et al., 2010). Given the emerging evidence of brain plasticity following learning or injury (Draganski & May, 2008; Ptito, Kupers, Lomber, & Pietrini, 2012), Faubert and Sidebottom (2012) introduced a perceptual-cognitive training methodology for athletes (Faubert & Sidebottom, 2012). The technique used is a “highly leveled” 3D-MOT perceptual-cognitive task because it stimulates a high number of brain networks that have to work together during the exercise including complex motion integration, dynamic, sustained and distributed attention processing and working memory. In an earlier publication by Faubert (2013), the 3D-MOT training technique revealed striking superior skills in professional athletes compared to sub-elites and novices when rapidly learning complex and neutral dynamic visual scenes (Faubert, 2013). The results showed a clear distinction between the level of athletic performance and corresponding fundamental mental capacities for learning an abstract and demanding dynamic scene task. The author suggested that rapid learning in complex and unpredictable dynamic contexts is one of the critical components required for elite performance. Lately, a study revealed that 3D-MOT performance was most likely related to the athletes' ability to see and respond to various stimuli on the basketball court, however the simple visuo-motor reaction time that was not related to any of the basketball specific performance measures (Mangine et al., 2014). Furthermore, recent neurological evidence has demonstrated the role of 3D-MOT in enhancing cognitive function in healthy young adults (Parsons et al., 2014). In fact, 10 sessions of 3D-MOT training improved attention, visual information processing speed and working memory recorded through neuropsychological tests and quantitative electroencephalography. In addition, other evidence has demonstrated that 3D-MOT training can show transfer to socially relevant tasks such as biological motion perception in the elderly (Legault & Faubert, 2012). In sport science and especially in perceptual-cognitive training studies, focus on transfer measures is essential to determine whether any improvements observed in the laboratory may transfer back to a live game situation. A common hypothesis suggests that transfer can occur if the trained and transfer tasks engage specific overlapping cognitive processes and brain networks (Dahlin, Neely, Larsson, Backman, & Nyberg, 2008). Studies supporting the expert performance approach have raised the question of perceptual-cognitive transfer in athletes (Caserta, Young, & Janelle, 2007; Gabbett, Carius, & Mulvey, 2008; Hopwood, Mann, Farrow, & Nielsen, 2011; Williams, Ward, & Chapman, 2003), however; to our knowledge, studies from the cognitive component approach have yet to demonstrate this transfer. For instance, Gabbett et al. (2008) investigated the effects of video-based perceptual training on decision- making skills during small-sided games (SSG) in elite women soccer players (Gabbett et al., 2008). Video-based training yielded on-field improvements in passing, dribbling and shooting decision-making skills. In addition, the use of SSG as a measure of transfer seemed to be an efficient strategy in capturing the dynamic and strategic components of soccer. In fact, a major challenge with transfer settings is to develop objective and sensitive measures of transfer. Soccer is an invasion game where the main goal is to invade an opponent's territory (offensive scenario) to score and/or to contain space and regain ball possession (defensive scenario) to avoid conceding goals (Mitchell, Oslin, & Griffin, 2013). Players have to make different decisions whether they have to defend or to attack. During the attack, the offensive aspect is to score a goal (e.g. shooting) while conservation of the ball (e.g. passing, dribbling) is the defensive aspect (Gréhaigne, Richard, & Griffin, 2012). Passing and shooting the ball have been recognized as important factors that contribute to the success or improvement of performance in invasion games (Hughes & Bartlett, 2002). On the other hand, recovering the ball or putting pressure (e.g. pressing) on the opposing team to regain possession of the ball is the offensive aspect of the defense. Defending one's goal (e.g. tackling) consists in the defensive aspect of the defense. Therefore, soccer presents a complex and rapidly changing environment where different decisions have to be made under pressure and time constraints. The quality of response to those decision-making situations is crucial in the success of the team and to improve performance. SSG allows players to experience similar situations that they encounter in competitive matches while optimizing training duration. It also reduces space and therefore the amount of time to respond. This process increases the number of opportunities for decision-making and thus increases the ratio of players' participation in decision-making (Aguiar, Botelho, Lago, Maças, & Sampaio, 2012). In the present study, we assessed the transfer capability of perceptual-cognitive 3D-MOT training void of sports context on offensive decision-making with soccer players. In a dynamic sport environment such as soccer, players must be able to correctly read the key information from a visual scene to make accurate decisions. They are usually confronted with multiple choices. To score goals, players have to select the best options especially during an offensive scenario. For example, within a split second, players need to decide whether they protect the ball, when and where to pass or whether they should take a shot on net. The decision-making skill refers to the capability of individuals to make a choice and achieve a specific task goal from a set of possibilities (Bar-Eli, Plessner, & Raab, 2011). Becoming an expert in decision-making is thought to be acquired following sport- specific deliberate practice (10 000 hour rule) (Ericsson, Krampe, & Tesch-Römer, 1993) even if there is some speculation about the benefits of the involvement in non-sport- specific activities during the first years of practice (Baker, Cote, & Abernethy, 2003). This refers to the general benefits of involvement in any physical activities explained earlier (cognitive component approach). Decision-making in sport relies on three major cognitive components such as perception, knowledge and decision strategies (Bar-Eli et al., 2011). Expert decision-makers depend on advanced perceptual and memory processes to better execute short term decisions. In particular, they rely on advanced visual search strategies in their central and peripheral vision (Vaeyens et al., 2007) as well as selective, focused and divided attention (Bar-Eli et al., 2011). While general visual training programs hardly improve decision-making in sports, 3D-MOT includes dynamic visual information that has to be processed actively, which is a crucial part of the perceptual component involved in decision-making. The 3D-MOT selective attention and processing speed of multiple moving targets task may be a crucial skill to help execute decision-making. This is supported by studies showing that performance on multiple target tracking tasks is superior in different kinds of experts such as professional radar operators, video-gamers and athletes (Allen, McGeorge, Pearson, & Milne, 2004; Green & Bavelier, 2003; Zhang, Yan, & Yangang, 2009). To train perceptual and attentional processes involved in decision- making, we used a highly leveled perceptual-cognitive training task that includes complex motion integration, dynamic, sustained and distributed attention processing and working memory. Three-dimensional MOT training has already shown evidence in enhancing cognitive function by improving attention, visual information processing speed and working memory. Additionally, this is a task in which has previously highlighted athletes' impressive learning capabilities in complex and dynamic visual scenes. Recently, the technique has shown training transfer onto a perceptual-cognitive task, such as biological motion perception, within a laboratory setting. To assess the role of 3D-MOT in decision-making accuracy on the field, we looked at three essential skills that are used to gain the upper hand over the opponent throughout soccer offensive scenarios. We suggested that passing, dribbling and shooting could be improved in the 3D-MOT trained group compared to control groups between pre- and post-training sessions. Ethics statement ~~~~~~~~~~~~~~~~ The experimental protocol and related ethical issues were evaluated and approved by the Comité d’Éthique de la Recherche en Santé of Université de Montréal. All subjects were given verbal and written information about the study and gave their verbal and written informed consent to participate.","Twenty-three young males from the Carabins soccer team of Université de Montréal participated (Table 1). All subjects reported normal or corrected-to-normal vision (6/6 or better) with normal stereoacuity (50 s of arc or better). None of the subjects had ever taken part in any previous 3D-MOT or perceptual-cognitive experiment. Laboratory tests The 3D-MOT experiment and 3D soccer videos (active control) were conducted using a fully immersive virtual environment thanks to a head-mounted display (Sony HMZ-T2) in a room with controlled lighting. The head-mounted display is a 3D-ready system that allows image projection on two OLED panels with a resolution of 1280 × 720 pixels and covering a maximal visual field of 45°. Inter-pupillary distance was adjusted for each subject. The 3D-MOT experiment was supported by a Hewlett- Packard ProBook 4530s with a Core i5 processor and an Intel HD Graphics 3000 graphic card. The 3D soccer videos, from the official 2010 FIFA world cup™ blu- ray, were played on a Sony PlayStation 3™ system. Both the computer (for 3D-MOT) and the Sony PlayStation 3™ (for 3D soccer videos) were connected to the head- mounted display to offer an immersive experience. Field test Decision-making assessment was conducted during standardized SSG before and after the training period. SSG consisted of standard 5 × 5 soccer matches on a 30 m × 40 m interior turf soccer field to avoid weather influence. Coaches were positioned on the side of the pitch to give their instructions as in a real game situation. Players were randomly distributed in five different teams composed of five players each including a keeper. The five teams were randomly facing each other two times during ten games of 5 min each. Every player was then taking part in eight games of 5 min for a total of 40 min during both pre- and post-sessions. Players who were waiting for the start of the next game were stretching or exercising with the ball. SSG were recorded using two video cameras (Sony, HDR-CX260VW). Cameras were positioned in the bleachers of the stadium, approximately 10 m above the field of play to cover the entire playing area. Players were identified by jerseys and numbers. The video recordings were analyzed using Dartfish Connect v6.0. Decision-making coding On-field decision making ability during SSG was coded using standardized coding criteria adapted from previous studies (French & Thomas, 1987; Gabbett et al., 2008). Passing, dribbling and shooting were the skills assessed (see Table 2). The coding instrument made it possible to separate the cognitive decision-making component of performance from the motor skill execution component of performance. When initially used by French and Thomas (1987) in basketball players, the coding instrument was built to evaluate three aspects of performance: control (e.g. a player catches the ball), decision (e.g. a player decides which action is appropriate), and execution (e.g. a player then executes the skill). In a recent study, Gabbett et al. (2008) adapted the instrument for coding decision-making in soccer. They assessed one aspect of performance, the decision, which is central in the context of our study. The decision component involves selection of the skill (e.g. pass, dribble, shoot), as well as which teammate to pass to, what direction to dribble, when to shoot, when to stop dribbling, and so on. The quality of each decision was coded as 1 for an appropriate decision and 0 for an inappropriate decision according to the criteria (Table 2). Decisions that were neither appropriate nor inappropriate were not coded. For instance, when the player made a pass that did not: a) directly or indirectly created a shot attempt; b) went to a teammate who was in a better position than the passer; c) went to a player who was closely guarded; d) being intercepted or cause a turn over; e) reached an area of the field where no teammates was positioned or out of the field of play. Moreover, where the player did not have time to assess the options (e.g. player was tackled as soon as he received the ball) the disposal was not considered for assessment. Decision-making coding was assessed by an experienced soccer coach blinded to the experimental protocol and trained to use the instrument for coding. Then, the total score of each player by session was converted to percentage for analysis. Percentage accuracy values were established for each participant by dividing the number of points awarded by the total number available and then multiplying by 100. Assessment of subjective judgements To assess whether perceptual-cognitive learning was directly perceived or related to unconscious processes, we used participants' judgments in on-field decision- making. Players' confidence levels in decision-making accuracy were assessed promptly after pre- and post-sessions using the Sport Performance Scale application developed in our laboratory (http://vision.opto.umontreal.ca/english/technologies/apps_en.html). In the present study, we used a simple visual analog scale (rated from min [0%] to max [100%]) to assess players' confidence levels in decision-making. Measurements were performed on a ‘10′ Samsung Galaxy Tab II™’ tablet. The following instructions were individually given to the players: 1) to rate their ‘decision-making accuracy during the last play’ and to do so by 2) scrolling their finger on the visual analog scale of the touchscreen tablet until they had reached the appropriate score. Decision-making accuracy was described and contextualized on the scale as follows: ‘Your level of accuracy in anticipating teammate or opponent movements and to deliver a correct response (e.g. assist, shot)’. No time constraint was imposed and each player took approximately 15 s to both receive the instructions and give an answer. 3D-MOT The 3D-MOT task (Fig. 1) was working under the NeuroTracker™ system licenced by the Université de Montréal to CogniSens Athletics, Inc. (Montreal, Canada). The CORE mode of the NeuroTracker™ system was used. During the exercise, four of eight projected spheres had to be tracked within a 3D virtual volumetric cube space with virtual light grey walls, subtending a visual angle of 42°. The spheres followed a linear trajectory in the 3D virtual space. Deviation occurred only when the balls collided against each other or the walls. In order to support an effective distribution of attention, a fixation spot was presented in the center of the cube throughout the experimentation. An instructed part of the training task was to focus on this green fixation square throughout the tracking phase which serves as an anchor point from which to extract information from the visual periphery (Ripoll, 1991). In other words, the anchor point is localized in a central position where the fovea could be directed while the relative movement of the spheres could be monitored using the peripheral visual field which is also able to detect movement (McKee & Nakayama, 1984). The effective use of such a strategy has already been demonstrated in experts through a variety of sports (e.g. Ripoll, Kerlirzin, Stein, & Reine, 1995; Savelsbergh, Williams, Van der Kamp, & Ward, 2002). Each session, based on a staircase procedure, lasted about 8 min. The staircase procedure consists of increasing speed if the subject got all the indexed targets or decreasing speed if at least one target was missed. Speed thresholds were then evaluated using a 1-up 1-down staircase procedure (Levitt, 1971). After each correct response, the dependent variable (speed ball displacement) was increased by 0.05 log and decreased by the same proportion after each incorrect response, resulting in a threshold criterion of 50%. The staircase was interrupted after eight inversions and the threshold was estimated by the mean of the speeds at the last four inversions. 3D soccer videos The 3D soccer videos consisted of game replay from group and knock-out stages of the FIFA world cup 2010. Participants went through the whole blu-ray video once during the first four sessions of training. They were asked to watch soccer actions during approximately 25 min/session. During the last six sessions of training, participants would watch the 3D soccer videos for a second time (20 min/session) and were challenged by 5 min interviews during which questions about decision-making accuracy occurring throughout the soccer videos were raised (e.g. Did the scorer make the right decision on the first goal of Brazil against North Korea according to the situation? What would you have done in his place: pass, dribble, shoot?). The questions only ensured that the players were focusing on the videos and reinforced their belief that the training was effective. Importantly, no feedback was given to the players after they answered.","We used a computer randomization script to allocate players into three separate groups including an experimental (n = 9), active (n = 7) and passive (n = 7) control group. All of them completed a pre- and post-on-field session during standardized SSG. Participants were constrained to not take part in any other training research activities throughout the duration of the testing period. In addition, the university athletes maintained a similar weekly routine, which was limited to their academic class schedule, training and soccer practice. Experimental group Players were actively trained ten times; twice a week for five consecutive weeks. During each evaluation, they participated in three CORE sessions of 3D-MOT. All of the observers reached a total of thirty sessions at the end of the training. Each participant followed the same standard procedure and completed each task while seated. Active control Participants focused on 3D soccer videos from the official 2010 FIFA world cup™ blu-ray twice a week for a five week period and a total of 10 sessions. Players were informed that this training was expected to have a positive effect on their decision-making performance during soccer games. This procedure was undertaken to provide an expectancy set for training benefits comparable to that of the perceptual-cognitive training group (cf Williams et al., 2003). Passive control No instruction or training was provided for this group. Analysis ~~~~~~~~ Due to injuries, two players from the experimental and the passive control groups were removed from the decision-making analysis. As the two (active–passive) control groups showed no statistical differences and small sample size, they were analyzed as a single control group. On-field decision-making accuracy in passing, dribbling and shooting was compared between the experimental and the active–passive control group. To find out whether the players' performance was different between pre- and post-sessions in the 3D-MOT group compared to the other group, we used a mixed-design analysis of variance (ANOVA; group × sessions). A Levene test yielded no significant differences in homogeneity between groups (p > 0.05). We conducted paired Student t-tests to compare pre- and post- evaluation of players' subjective decision-making accuracy, rated by the Sport Performance Scale, in the 3D-MOT and the active–passive control group. Athletes' 3D-MOT speed threshold means analysis between pre and post-training showed a comparable improvement as previously demonstrated in athletes (Faubert, 2013; Faubert & Sidebottom, 2012). Objective decision-making assessment On-field decision-making analysis revealed a significant improvement in passing accuracy only for the 3D-MOT trained group between pre- and post-sessions compared to the other groups (F(1, 17) = 4.708, p = 0.044, ɳ2 = 0.162) (Fig. 2A). No significant difference was observed in decision-making accuracy for dribbling (F(1, 14) = 3.628, p = 0.078, ɳ2 = 0.200) and shooting (F(1, 13) = 0.210, p = 0.654, ɳ2 = 0.015) between pre- and post-sessions for the experimental compared to the other groups. However, there is a clear tendency for improvement in dribbling that did not reach the significance threshold (p = 0.078). We suggest that a high variance could explain those results. In fact, mean number of dribbles (5.1 ± 0.63 SEM) and shots (4.0 ± 0.52 SEM) during games were far too small compared to mean number of passes (15.5 ± 0.93 SEM) to reach a decisive conclusion on the impact of 3D-MOT training on the on-field performance for these abilities. Subjective decision-making assessment A general improvement in subjective confidence levels of decision-making accuracy between pre- and post-sessions was observed in the 3D-MOT group following a paired student t-test analysis (t[6] = −3.547, p = 0.012) while no difference was observed in the active–passive control group (t[11] = −1.515, p = 0.158) (Fig. 2B).","This study was designed to assess the transferability of perceptual-cognitive training void of sports context on decision-making accuracy in soccer players. The main result demonstrates a significant 15% improvement in passing decision-making accuracy in soccer players trained with the 3D-MOT technique compared to active and passive controls. Furthermore, this result was corroborated with a proportional quantitative increase in subjective decision-making accuracy for the experimental group. Finally, players' 3D-MOT speed thresholds confirmed their superior capacity for processing a complex and dynamic visual scene task. It should be noted that even if we observed a trend (p = 0.078) in favor of an improvement in dribbling decision-making for the 3D-MOT group, no significant differences were found in dribbling and shooting abilities between groups. However, the low number of dribbles and shots attempted by the players restrict our ability to firmly conclude that the 3D-MOT training does not affect these decision-making abilities. For this reason, we will center the discussion on passing decision-making. Decision-making improvement in passing ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the present study, we trained soccer players to track multiple elements through virtual, dynamic, complex and neutral (sport environment free) visual scenes using a perceptual-cognitive 3D-MOT technique. Ten training sessions in laboratory were sufficient to improve players' passing decision-making accuracy by 15% between pre- and post-sessions of SSG. Meanwhile, the active (10 sessions of 3D soccer matches) and passive control groups did not show any improvement. Considering 3D-MOT transferability has already been shown in a biological motion perception laboratory task (Legault & Faubert, 2012), this study represents the first evidence of 3D-MOT transfer from the laboratory to the field. Moreover, we argued in the introduction that expertise in athletes has been contextualized according to both the (sport specific) expert performance skills approach and the (sport environment free) cognitive component skills approach. While few studies have reported transfer following perceptual-cognitive training in specific sport contexts, the present study is, to our knowledge, the first evidence in favor of a non-contextual perceptual- cognitive training transfer on the sport-field. According to the most accepted theory, it is suggested that specific overlapping cognitive processes and brain networks were engaged for the laboratory 3D-MOT task and on-field decision-making process during SSG (Dahlin et al., 2008). Potential implications of 3D-MOT in decision-making improvement are discussed below. Attentional tracking of multiple elements One of the critical requirements to allow decision-making is to accurately extract meaningful information from the visual scene which is possible by using both perceptual and attentional processes. Attention and concentration are crucial abilities that affect the decision-making of athletes (Bar-Eli et al., 2011). During a soccer action, an athlete has to divide attention on the field (e.g. teammates, opponents, ball), to use selective attention (e.g. which player to give the ball to) and to focus attention (e.g. staring at the net to score). To this purpose, many benefits may arise from the highly leveled 3D-MOT technique. A core feature of the 3D-MOT technique relies on distributed attention on a number of separated dynamic elements (Cavanagh & Alvarez, 2005). The ability to track multiple elements has been reported to be superior with the involvement in sport activities in adults and youngsters during laboratory MOT tasks (Barker, Allen, & McGeorge, 2010; Trick, Jaspers-Fayer, & Sethi, 2005; Zhang et al., 2009). This result is not surprising knowing that the ability to maintain attention on multiple stimuli or locations for quite a prolonged period of time is important for sport (Memmert, 2009). Importantly the 3D-MOT technique includes speed thresholds as a dependent variable which is considered as a crucial part of MOT performance by requiring more attentional resources to track at higher speeds (Feria, 2012). Recently, neurological evidence has demonstrated the role of 3D-MOT in improving attention, visual information processing speed and working memory (Parsons et al., 2014). From other imagery studies, the MOT technique has reported activation of higher-level brain areas involved in attentional processes (Culham et al., 1998; Howe, Horowitz, Morocz, Wolfe, & Livingstone, 2009). These areas include parietal and frontal regions of the cortex and are believed to be responsible for attention shifts and eye movement. As well, the middle temporal complex has, not surprisingly, been implicated during MOT processing for motion perception (Culham et al., 1998). These brain pathways could potentially be involved during the action of reading the play in soccer players; a perception-in- action process which is known to activate both dorsal and ventral streams (Goodale & Milner, 1992). Training the brain to simultaneously activate those networks may possibly help to enhance perceptual-cognitive execution in athletes. In this sense, the attentional tracking of multiple elements during 3D-MOT could overlap brain networks required during decision-making. Training on 3D-MOT could serve as a tool to help automatize those networks and could lead to superior decision- makings abilities. Future imagery or electroencephalography study will help us to explain the neural process behind 3D-MOT improvements. Engagement of visual search strategies in peripheral vision Beyond the attentional tracking of multiple elements, the 3D-MOT technique engages wide visual field stimulation especially because peripheral vision has been suggested to play an important role in the performance of sports teams (Knudson & Kluka, 1997). With players spread all along a field of about 60 m in width, soccer utilizes a large amount of peripheral vision. Peripheral vision refers to the ability to detect and react to stimuli outside of foveal vision. In soccer referees, peripheral vision has been showed to be useful in decision-making accuracy (De Oliveira, Orbetelli, & De Barros Neto, 2011). To extract meaningful information from the visual scene, including the periphery, expert athletes rely on advanced visual search strategies (Vaeyens et al., 2007; Williams, 2000). There is evidence to support that relative motion information is picked up effectively via peripheral vision (Williams, Davids, & Williams, 1999). A common occurrence during a game is to use foveal and peripheral vision simultaneously. For instance, in ‘time-constrained’ situations (e.g. 5 vs 5 situation), skilled soccer players fixate the ball in foveal vision while using peripheral vision to monitor the positions of teammates and opponents in the periphery (Williams & Davids, 1997). A study by Vaeyens et al. (2007) showed that successful soccer decision-makers use the player in possession of the ball as the central point on which to fixate gaze to explore and pick up the key information underpinning decision making in offensive situations (Vaeyens et al., 2007). Researchers have reported the use of those ‘visual pivots’ in other sports (Ripoll et al., 1995; Savelsbergh et al., 2002) and is a reason why 3D-MOT includes such an anchor point. On the other hand, it has been proposed that athletes from visual demanding sports (e.g. netball) can generate more frequent (and shorter) eye movements therefore enabling information to be extracted for the visual field more rapidly (Morgan & Patterson, 2009). Research on eye movement during MOT experiments have identified viewing strategies exercised by observers. When multiple targets (e.g. 3 spheres) were presented, participants usually adopted a ‘center-looking’ strategy as if they were grouping the targets into a single object (e.g. triangle) and were looking closer to the center of the object formed by the targets (Fehd & Seiffert, 2008). This strategy is in contrast to a ‘target-looking’ strategy where participants would saccade from target to target. However, another study by Fehd and Seiffert (2010) demonstrated that participants often engaged in both ‘target-looking’ and ‘center- looking’ strategies by switching their gaze from the center to the targets and so on (Fehd & Seiffert, 2010). Visual search strategies involved during MOT could be closely linked to those engaged by sport experts during the process of extracting visual information from the action. By training those strategies, which are part of the perceptual component involved in decision-making, could help to improve in game decision-making. Virtual reality (3D vision) Another major asset of the 3D-MOT methodology is the involvement of virtual reality, a technology that is recognized as an important tool to potentially improve sport performance (Bideau et al., 2010; Carling, Reilly, & Williams, 2009). Immersive environments, such as those afforded by virtual reality, engender automaticity and therefore implicit learning – which yields decision-making that is robust under pressure (Patterson, Pierce, Bell, Andrews, & Winterbottom, 2009). Moreover, virtual reality involves stereoscopy (binocular disparity) which is required in situations where fast, complex and dynamic elements collide or overlap. For instance, stereoscopy has been shown to help in disambiguating object occlusions when processing dynamic visual scenes (Faubert & Allard, 2013). This is typically the kind of critical situation that can usually be found in soccer when players are close to each other (e.g. an attacker wants to make a deep run from behind the defender but needs to stay close until the last moment to avoid being called offside). Whether it is fast attentional multiple element tracking, wide visual field stimulation or stereoscopy, all those components are involved during sport actions. Therefore, previous results revealing athletes' extraordinary skills for rapidly learning complex and neutral dynamic visual scenes using the 3D-MOT technique appear logical (Faubert, 2013). This non-contextual perceptual- cognitive training seems to involve higher-level cognitive abilities subserved by the central nervous system. Presumably, 3D-MOT may capture the dynamic components of soccer actions where players have to maintain focus and attention on teammates, opponents or the ball to make the best decisions. In summary, this paradigm, even if void of sports context, is in keeping with the complex and dynamic nature of strategic sport such as soccer. From a sport performance point of view, the result is of particular interest in regards to the implication of decision-making and passing in modern soccer. Today's elite soccer players require faster decision- making than ever mainly because as a competitive sport, the game of soccer is increasing in ball speed (15%) as well as the intensity of play and passing rate (35%) according to an analysis on world cup soccer final games from the last 40 years (Wallace & Norton, 2014). Whereas accurately passing the ball to a teammate is an essential and fundamental ability required by soccer players (Ali, 2011), it also represents the essence of keeping the ball and the source of scoring opportunities (Chassy, 2013). The 3D-MOT technique could play a crucial role in improving passing accuracy in elite soccer players and could be implemented in training centers. Subjective decision-making assessment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Results of subjective decision-making assessment, collected from the Sport Performance Scale, support the on-field improvement observed in the 3D-MOT training group. One possible explanation is that it confirms that players who received the perceptual- cognitive training were conscious of their on-field improvement in decision-making after the training. Furthermore, confidence level improvement in decision-making accuracy of trained players was quantitatively proportional to the improvement in decision-making accuracy rated during video analysis. These results seem to demonstrate that passing decision-making accuracy improvement in the 3D-MOT group represents a meaningful training effect rather than the result of increased familiarity with the test environment or expectancy set for the training benefits. Athlete skills for learning complex and neutral dynamic visual scenes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The 3D-MOT speed thresholds of soccer players were qualitatively similar to those previously obtained in professional and elite-amateurs athletes showing superior capacity for processing a complex and dynamic visual scene task (Faubert, 2013; Faubert & Sidebottom, 2012). This ability has been argued to be one of the critical components for elite performance and 3D-MOT performance has been shown to be highly associated with athletes' performance level (Faubert, 2013; Mangine et al., 2014). Future requirements and limitations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To resolve the issue of whether this training can also transfer to better decision making in regards to dribbles and shooting decisions, we may have to use a different standardized SSG situation. For instance, we could reduce pitch size to favor ‘one-on-one’ situations (increase dribbling ratio) and direct ‘goal-to-goal’ actions (increase shooting ratio). To underline the potential of non-contextual 3D-MOT training, it will also be interesting to address the degree of transfer of the technique in other invasion (e.g. hockey) or net (e.g. tennis) games. Another important aspect is to evaluate perceptual-cognitive skills in youngsters with the 3D-MOT technique. Using memory recall and structured pattern of play, Ward and Williams (2003) have previously demonstrated superior perceptual-cognitive skills in elite compared to sub-elites soccer players as early as 9 years old (Ward & Williams, 2003). However, little or no studies have compared perceptual-cognitive skills of elite and novice youngsters using a perceptual-cognitive method void of sports context. One study has revealed better MOT performance in children involved in physical activity compared to more sedentary children (Trick et al., 2005). Knowing that rapid learning in complex and dynamic visual scenes is a critical component for sport performance (Faubert, 2013) and that 3D-MOT performance has been linked to sport specific performance measures (Mangine et al., 2014), 3D-MOT speed thresholds could serve as a tool to determine a player's ability. Non-contextual perceptual-cognitive techniques may also have implications in the screening or detection of new talent.","Expertise in athletes has been well characterized using specific as well as non-contextual perceptual-cognitive paradigms. However, the present study represents the first evidence of an on-field improvement (transfer) following a laboratory perceptual-cognitive training void of sports context. In fact, training to process complex and dynamic visual scenes has not only revealed superior learning ability in soccer players but has led to improvements in passing decision-making accuracy in the field as well. Future laboratory and in-field studies will be needed to evaluate the degree of transferability of such training on other dynamic sports.","One of the authors is director of Visual Psychophysics and Perception Laboratory at the University of Montreal and he is the Chief Science Officer of Cognisens Athletics Inc. who produces the commercial version of the NeuroTracker used in this study. In this capacity, he holds shares in the company. This does not alter our adherence to your journal policies on sharing data and materials."],["This paper considers communication in terms of inference about the behaviour of others (and our own behaviour). It is based on the premise that our sensations are largely generated by other agents like ourselves. This means, we are trying to infer how our sensations are caused by others, while they are trying to infer our behaviour: for example, in the dialogue between two speakers. We suggest that the infinite regress induced by modelling another agent - who is modelling you - can be finessed if you both possess the same model. In other words, the sensations caused by others and oneself are generated by the same process. This leads to a view of communication based upon a narrative that is shared by agents who are exchanging sensory signals. Crucially, this narrative transcends agency - and simply involves intermittently attending to and attenuating sensory input. Attending to sensations enables the shared narrative to predict the sensations generated by another (i.e. to listen), while attenuating sensory input enables one to articulate the narrative (i.e. to speak). This produces a reciprocal exchange of sensory signals that, formally, induces a generalised synchrony between internal (neuronal) brain states generating predictions in both agents. We develop the arguments behind this perspective, using an active (Bayesian) inference framework and offer some simulations (of birdsong) as proof of principle. --------------------------------------------------------------------------------","One of the most intriguing issues in (social) neuroscience is how people infer the mental states and intentions of others. In this paper, we take a formal approach to this issue and consider communication in terms of mutual prediction and active inference (De Bruin & Michael, 2014; Teufel, Fletcher, & Davis, 2010). The premise behind this approach is that we model the causes of our sensorium – and adjust those models to maximise Bayesian model evidence or, equivalently, minimise surprise (Brown & Brün, 2012; Kilner, Friston, & Frith, 2007). This perspective on action and perception has broad explanatory power in several areas of cognitive neuroscience (Friston, Mattout, & Kilner, 2011) – and enjoys support from several lines of neuroanatomical and neurophysiological evidence (Egner & Summerfield, 2013; Rao & Ballard, 1999; Srinivasan, Laughlin, & Dubs, 1982). Here, we apply this framework to communication and consider what would happen if two Bayesian brains tried to predict each other. We will see that Bayesian brains do not predict each other – they predict themselves; provided those predictions are enacted. The enactment of sensory (proprioceptive) predictions is a tenet of active inference – under which we develop this treatment. In brief, we consider the notion that a simple form of communication emerges (through generalised synchrony) if agents adopt the same generative model of communicative behaviour. So how do prediction and generative models speak to theory of mind? The premise here is that we need to infer – and therefore predict – the causes of sensations to perceive them. For example, to perceive a falling stone we have to appeal (implicitly) to a model of how objects move under gravitational forces. Similarly, the perception of biological motion rests on a model of how that motion is caused. This line of argument can be extended to the perception of intentions (of others or ourselves) necessary to explain sensory trajectories; particularly those produced by communicative behaviour. The essential role of inference and prediction has been considered from a number of compelling perspectives: see (Baker, Saxe, & Tenenbaum, 2009; Frank & Goodman, 2012; Goodman & Stuhlmuller, 2013; Kiley Hamlin, Ullman, Tenenbaum, Goodman, & Baker, 2013; Shafto, Goodman, & Griffiths, 2014). This Bayesian brain perspective emphasises the role of prediction in making inferences about the behaviour of others; particularly in linguistic communication. We pursue exactly the same theme; however, in the context of active inference. Active inference takes the Bayesian brain into an embodied setting and formulates action as the selective sampling of data to minimise uncertainty about their causes. This provides a natural framework within which to formulate communication, which is inherently embodied and enactivist in nature. In brief, we hope to show that active inference accounts for the circular (Bayesian) inference that is inherent in communication. The very notion of theory of mind speaks directly to inference, in the sense that theories make predictions that have to be tested against (sensory) data. In what follows, we focus on the implicit model generating predictions: imagine two brains, each mandated to model the (external) states of the world causing sensory input. Now imagine that sensations can only be caused by (the action of) one brain or the other. This means that the first brain has to model the second. However, the second brain is modelling the first, which means the first brain must have a model of the second brain, which includes a model of the first – and so on ad infinitum. At first glance, the infinite regress appears to preclude a veridical modelling of either brain’s external states (i.e., the other brain). However, this infinite regress dissolves if the two brains are formally similar and each brain models the sensations caused by itself and the other as being generated in the same way. In other words, if there is a shared narrative or dynamic that both brains subscribe to, they can predict each other exactly, at least for short periods of time. This is basic idea that we pursue in the context of active inference and predictive coding. In fact, we will see that this solution is a necessary and emergent phenomenon, when two or more (formally similar) active inference schemes are coupled to each other. Mathematically, the result of this coupling is called generalised synchronisation (aka synchronisation of chaos). Generalized synchrony refers to the synchronization of chaotic dynamics, usually in skew-product (i.e., master–slave) systems (Barreto, Josic, Morales, Sander, & So, 2003; Hunt, Ott, & Yorke, 1997). However, we will consider generalized synchrony in the context of reciprocally coupled dynamical (active inference) systems. This sort of generalized synchrony was famously observed by Huygens in his studies of pendulum clocks – that synchronized themselves through the imperceptible motion of beams from which they were suspended (Huygens, 1673). This nicely illustrates the action at a distance among coupled dynamical systems. Put simply, generalised synchronisation means that knowing the state of one system (e.g., neural activity in the brain) means that one can predict the states of the other (e.g., another’s brain). The sequence or trajectory of states may not necessarily look similar but there is a quintessential coupling in the sense that the dynamics of one system can be predicted from state of another. In this paper, we will illustrate the emergence of generalized synchronization when two predictive coding schemes are coupled to each other through action. In fact, we will illustrate a special case of generalized synchrony; namely, identical synchronization, in which there is a one-to-one map between the states of two systems. In this case, there is a high mutual predictability and the fluctuations in the states appear almost the same. In short, we offer generalized synchronization as a mathematical image of communication (of a simple sort) that enables two Bayesian brains to entrain each other and, effectively, share the same dynamical narrative. This paper comprises five sections. The first provides a brief review of active inference and predictive coding, with a focus on the permissive role of sensory attenuation when acting on the world. In predictive coding, sensory attenuation is a special case of optimising the precision or confidence in sensory (and extrasensory) information that is thought to be encoded by the gain of neuronal populations encoding prediction errors (Clark, 2013a; Feldman & Friston, 2010). This takes us into the realm of cortical gain control and neuromodulation – that may be closely tied to synchronous gain and the oscillatory dynamics associated with binding, attention and dynamic coordination (Fries, Womelsdorf, Oostenveld, & Desimone, 2008; Womelsdorf & Fries, 2006). In the second section, we introduce a particular model that is used to illustrate perception, action and the perception of action in the context of communication. This model has been used previously to illustrate several phenomena in perception; such as perceptual learning, repetition suppression, and the recognition of stimulus streams with deep hierarchical structure (Friston & Kiebel, 2009; Kiebel, Daunizeau, & Friston, 2008). In the third section, we provide a simple illustration of omission related responses – that are ubiquitous in neurophysiology and disclose the brain’s predictive proclivity. In the fourth section, we use this model to illustrate the permissive role of sensory attenuation in enabling action by simulating a bird that sings to itself. The purpose of this section is to show that song production requires the attenuation of sensory input during singing – that would otherwise confound perceptual inference due to sensorimotor delays. This leads to the interesting (and intuitive) notion that the sensory consequences of acting have to be attenuated in order to act. This means that one can either talk or listen but not do both at the same time. Armed with this insight, we then simulate – in the final section – two birds that are singing to themselves (and each other) and examine the conditions under which generalised synchrony emerges. These simulations are offered as proof of principle that communication (i.e., a shared dynamical narrative) emerges when two dynamical systems try to predict each other.","Recent advances in theoretical neuroscience have inspired a (Bayesian) paradigm shift in cognitive neuroscience. This shift is away from the brain as a passive filter of sensations towards a view of the brain as a statistical organ that generates hypotheses or fantasies (from Greek phantastikos, the ability to create mental images, from phantazesthai), which are tested against sensory evidence (Gregory, 1968). This perspective dates back to the notion of unconscious inference (Helmholtz, 1866/1962) and has been formalised in recent decades to cover deep or hierarchical Bayesian inference – about the causes of our sensations – and how these inferences induce beliefs, movement and behaviour (Clark, 2013b; Dayan, Hinton, & Neal, 1995; Friston, Kilner, & Harrison, 2006; Hohwy, 2013; Lee & Mumford, 2003). Predictive coding and the Bayesian brain ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Modern formulations of the Bayesian brain – such as predictive coding – are now among the most popular explanations for neuronal message passing (Clark, 2013b; Friston, 2008; Rao & Ballard, 1999; Srinivasan et al., 1982). Predictive coding is a biologically plausible process theory for which there is a considerable amount of anatomical and physiological evidence (Friston, 2008; Mumford, 1992). See (Bastos et al., 2012) for a review of canonical microcircuits and hierarchical predictive coding in perception and (Adams, Shipp, & Friston, 2012; Shipp, Adams, & Friston, 2013) for a related treatment of the motor system. In these schemes, neuronal representations in higher levels of cortical hierarchies generate predictions of representations in lower levels. These top–down predictions are compared with representations at the lower level to form a prediction error (usually associated with the activity of superficial pyramidal cells). The ensuing mismatch signal is passed back up the hierarchy, to update higher representations (associated with the activity of deep pyramidal cells). This recursive exchange of signals suppresses prediction error at each and every level to provide a hierarchical explanation for sensory inputs that enter at the lowest (sensory) level. In computational terms, neuronal activity encodes beliefs or probability distributions over states in the world that cause sensations (e.g., my visual sensations are caused by a face). The simplest encoding corresponds to representing the belief with the expected value or expectation of a (hidden) cause. These causes are referred to as hidden because they have to be inferred from their sensory consequences. In summary, predictive coding represents a biologically plausible scheme for updating beliefs about states of the world using sensory samples: see Fig. 1. In this setting, cortical hierarchies are a neuroanatomical embodiment of how sensory signals are generated; for example, a face generates luminance surfaces that generate textures and edges and so on, down to retinal input. This form of hierarchical inference explains a large number of anatomical and physiological facts as reviewed elsewhere (Adams et al., 2012; Bastos et al., 2012; Friston, 2008). In brief, it explains the hierarchical nature of cortical connections; the prevalence of backward connections and explains many of the functional and structural asymmetries in the extrinsic (between region) connections that link hierarchical levels (Zeki & Shipp, 1988). These asymmetries include the laminar specificity of forward and backward connections, the prevalence of nonlinear or modulatory backward connections (that embody interactions and nonlinearities inherent in the generation of sensory signals) and their spectral characteristics – with fast (e.g., gamma) activity predominating in forward connections and slower (e.g., beta) frequencies that accumulate evidence (prediction errors) ascending from lower levels. Precision engineered message passing ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ One can regard ascending prediction errors as broadcasting ‘newsworthy’ information that has yet to be explained by descending predictions. However, the brain has to select the channels it listens to – by adjusting the volume or gain of prediction errors that compete to update expectations in higher levels. Computationally, this gain corresponds to the precision or confidence associated with ascending prediction errors. However, to select prediction errors, the brain has to estimate and encode their precision (i.e., inverse variance). Having done this, prediction errors can then be weighted by their precision so that only precise information is accumulated and assimilated at high or deep hierarchical levels. As for all expectations, expected precision maximises Bayesian model evidence (see Appendix). The broadcasting of precision-weighted prediction errors rests on gain control at a synaptic level (Moran et al., 2013). This neuromodulatory gain control corresponds to a (Bayes-optimal) encoding of precision in terms of the excitability of neuronal populations reporting prediction errors (Feldman & Friston, 2010; Shipp et al., 2013). This may explain why superficial pyramidal cells have so many synaptic gain control mechanisms; such as NMDA receptors and classical neuromodulatory receptors like D1 dopamine receptors (Braver, Barch, & Cohen, 1999; Doya, 2008; Goldman-Rakic, Lidow, Smiley, & Williams, 1992; Lidow, Goldman-Rakic, Gallager, & Rakic, 1991). Furthermore, it places excitation-inhibition balance in a prime position to mediate precision engineered message passing within and among hierarchical levels (Humphries, Wood, & Gurney, 2009). The dynamic and context sensitive control of precision has been associated with attentional gain control in sensory processing (Feldman & Friston, 2010; Jiang, Summerfield, & Egner, 2013) and has been discussed in terms of affordance in active inference and action selection (Cisek, 2007; Frank, Scheres, & Sherman, 2007; Friston et al., 2012). Crucially, the delicate balance of precision at different hierarchical levels has a profound effect on veridical inference – and may also offer a formal understanding of false inference in psychopathology (Adams, Stephan, Brown, Frith, & Friston, 2013; Fletcher & Frith, 2009; Friston, 2013). In what follows, we will see it has a crucial role in sensory attenuation. Active inference ~~~~~~~~~~~~~~~~ So far, we have only considered the role of predictive coding in perception through minimising surprise or prediction errors. However, there is another way to minimise prediction errors; namely, by re-sampling sensory inputs so that they conform to predictions; in other words, changing sensory inputs by changing the world through action. This is known as active inference (Friston et al., 2011). In active inference, action is regarded as the fulfilment of descending proprioceptive predictions by classical reflex arcs. In more detail, the brain generates continuous proprioceptive predictions about the expected location of the limbs and eyes – that are hierarchically consistent with the inferred state of the world. In other words, we believe that we will execute a goal- directed movement and this belief is unpacked hierarchically to provide proprioceptive and exteroceptive predictions entailed by our generative or forward model. These predictions are then fulfilled automatically by minimizing proprioceptive prediction errors at the level of the spinal cord and cranial nerve nuclei: see (Adams et al., 2012) and Fig. 1. Mechanistically, descending proprioceptive predictions provide a target or set point for peripheral reflex arcs – that respond by minimising (proprioceptive) prediction errors. The argument here is that the same inferential mechanisms underlie apparently diverse functions (action, cognition and perception) but are essentially the same; for example, in the visual cortex for vision, the insula for interoception, the motor cortex for movement and proprioception. Crucially, because these modality-specific systems are organised hierarchically, they are all contextualised by the same conceptual (amodal) predictions. In other words, action and perception are facets of the same underlying imperative; namely to minimize hierarchical prediction errors through selective sampling of sensory inputs. However, there is a potential problem here: Action and sensory attenuation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ If proprioceptive prediction errors can be resolved by classical reflexes or changing (proprioceptive) expectations, how does the brain adjudicate between these two options? The answer may lie in the precision afforded to proprioceptive prediction errors and the consequences of movement sensed in other modalities. In order to engage classical reflexes, it is necessary to increase their gain through augmenting the precision of (efferent) proprioceptive prediction errors that drive neuromuscular junctions. However, to preclude (a veridical) inference that the movement has not yet occurred, it is necessary to attenuate the precision of (afferent) prediction errors that would otherwise update expectations or beliefs about the motor plant (Friston et al., 2011). Put simply, my prior belief that I am moving can be subverted by sensory evidence to the contrary; thereby precluding movement. In short, it is necessary to attenuate all the sensory consequences of moving – leading to an active inference formulation of sensory attenuation – the psychological phenomena that the magnitudes of self-made sensations are perceived as less intense (Chapman, Bushnell, Duncan, & Lund, 1987; Cullen, 2004). In Fig. 1, we have omitted the (afferent) proprioceptive prediction error from the hypoglossal nucleus: see (Shipp et al., 2013) for discussion of this omission and the agranular nature of motor cortex. This renders descending proprioceptive predictions motor commands, where the accompanying exteroceptive predictions become corollary discharge. The ensuing motor control is effectively open loop. However, the hierarchical generation of proprioceptive predictions is contextualised by sensory input in other modalities – that register the sensory consequences of movement. It is these sensory consequences that are transiently attenuated during movement. In summary, to act, one needs to temporarily suspend attention to the consequences of action, in order to articulate descending predictions (Brown, Adams, Parees, Edwards, & Friston, 2013). Later, we will see that sensory attenuation plays a key role in communication. Birdsong and attractors ~~~~~~~~~~~~~~~~~~~~~~~ This section introduces the simulations of birdsong that we will use to illustrate active inference and communication in subsequent sections. We are not interested in modelling birdsong per se – or the specifics of birdsong communication. Birdsong is used here as a minimal (metaphorical) example of biologically plausible communication with relatively rich dynamics: noting that there is an enormous literature on the neurobiology and physics of birdsong (Mindlin & Laje, 2005), some of which is particularly pertinent to active inference; e.g., (Hanuschkin, Ganguli, & Hahnloser, 2013). We have used birdsong in previous work to illustrate perceptual categorisation and other phenomena. Here, we focus on omission-related responses to illustrate the basic nature of predictive coding of hierarchically structured sensory dynamics. The basic idea here is that the environment unfolds as an ordered sequence of states, whose equations of motion induce attractor manifolds that contain sensory trajectories. If we consider the brain has a generative model of these trajectories, then we would expect to see attractors in neuronal dynamics that are trying to predict sensory input. This form of generative model has a number of plausible characteristics: Models based upon attractors can generate and therefore encode structured sequences of events, as states flow over different parts of the manifold. These sequences can be simple, such as the quasi-periodic attractors of central pattern generators or can exhibit complicated sequences of the sort associated with itinerant dynamics (Breakspear & Stam, 2005; Rabinovich, Huerta, & Laurent, 2008). Furthermore, hierarchically deployed attractors enable the brain to predict or represent different categories of sequences. This is because any low-level attractor embodies a family of trajectories. A natural example here would be language (Jackendoff, 2002). This means it is possible to generate and represent sequences of sequences and, by induction sequences of sequences of sequences etc. (Kiebel, von Kriegstein, Daunizeau, & Friston, 2009). In the example below, we will try to show how attractor dynamics furnish generative models of sensory input, which behave much like real brains, when measured electrophysiologically: see (Friston & Kiebel, 2009) for implementational details. A synthetic songbird ~~~~~~~~~~~~~~~~~~~~ The example used here deals with the generation and recognition of birdsongs. We imagine that birdsongs are produced by two time-varying control parameters that control the frequency and amplitude of vibrations of the syrinx of a songbird (see Fig. 2). There has been an extensive modelling effort using attractor models at the biomechanical level to understand the generation of birdsong (Mindlin & Laje, 2005). Here we use attractors at a higher level to provide time-varying control over the resulting sonograms. We drive the syrinx with two states of a Lorenz attractor, one controlling the frequency (between two to five kHz) and the other controlling the amplitude or volume. The parameters of the Lorenz attractor were chosen to generate a short sequence of chirps every second or so. To give the generative model a hierarchical structure, we placed a second Lorenz attractor, whose dynamics were an order of magnitude slower, over the first (seconds as opposed to 100 ms or so); such that the states of the slower attractor change the manifold of the fast attractor. This manifold could range from a fixed-point attractor, where the states collapse to zero; through to quasi-periodic and chaotic behaviour. Because higher states evolve more slowly, they switch the lower attractor on and off, generating songs, where each song comprises a series of distinct chirps (see Fig. 2). Omission and violation of predictions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To illustrate the predictive nature of predictive coding, song recognition was simulated by integrating the predictive coding scheme above (equations in Fig. 1). These simulations used a standard integration scheme (spm_ADEM.m) described in detail in the appendix. The simulations reported in this paper can be reproduced by downloading SPM (http://www.fil.ion.ucl.ac.uk/spm/) and typing DEM to access the graphical user interface for the DEM Toolbox (Birdsong duet). A sonogram was produced using the above composition of Lorentz attractors (Fig. 2) and played to a synthetic bird – who tried to infer the underlying hidden states of the first and second level attractors. These attractors are associated with the higher local centre and area X in Fig. 1. Crucially, we presented two songs to the bird, with and without the final chirps. The corresponding sonograms and percepts (predictions) are shown with their prediction errors in Fig. 3. The left panels show the stimulus and percept, while the right panels show the stimulus and responses to omission of the last chirps. These results illustrate two important phenomena. First, there is a vigorous expression of prediction error after the song terminates abruptly. This reflects the dynamical nature of the recognition process because, at this point, there is no sensory input to predict. In other words, the prediction error is generated entirely by the predictions afforded by the dynamic model of sensory input. It can be seen that this prediction error (with a percept but no stimulus) is larger than the prediction error associated with the third and fourth stimuli that are not perceived (stimulus but no percept). Second, there is a transient percept when the omitted chirp should have occurred. Its frequency is too low but its timing is preserved in relation to the expected stimulus train. This is an interesting stimulation from the point of view of ERP studies of omission-related responses; particularly given that non-invasive electromagnetic signals arise largely from superficial pyramidal cells – which are the cells thought to encode prediction error (Bastos et al., 2012). Empirical studies of this characteristic response to violations provide clear evidence for the predictive capacity of the brain (Bendixen, SanMiguel, & Schröger, 2012). Creating your own sensations – a synthetic soliloquy ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We now consider the production of birdsong using the same model used for perception above. In the preceding simulations, descending predictions of exteroceptive (auditory) sensations were used to construct prediction errors that enabled perceptual inference and recognition of deep (hierarchical) structure in the sensory stream. However, in Fig. 1, there are also descending proprioceptive predictions to the hypoglossal region that elicit action through classical reflex arcs. This means the synthetic bird could, in principle, sing to itself – predicting both the exteroceptive and proprioceptive consequences of its action. This was precluded in the above simulations by setting the precision or gain of efferent proprioceptive prediction errors (that drive motor reflexes) to a very low value (a log precision of minus eight), in contrast to auditory prediction errors (with a log precision of two). This means that the precision weighted prediction errors do not elicit any action or birdsong, enabling the bird to listen to its companion. So what would happen if we increased the precision of proprioceptive prediction errors? One might anticipate that descending (multimodal) predictions from the higher vocal centre would cause the bird to sing – and predict the consequences of its own action – thereby eliciting a soliloquy. In fact, when the proprioceptive precision is increased (to a log precision of eight) something rather peculiar happens: Fig. 4 shows that a rather bizarre sonogram is produced, with low amplitude, high-frequency components and a loss of the song’s characteristic structure. The explanation for this failure lies in the sensorimotor delays inherent in realising proprioceptive predictions. In other words, descending auditory predictions fail to account for the slight delay in self-made sensations. This results in perpetual (and precise) exteroceptive prediction errors that confound perceptual synthesis and associated action (singing). This resonates with the well-known disruptive effect of delayed auditory feedback on speech (Yates, 1963). Put simply, action and perception chase each other’s tails, never resolving the discrepancy between the actual and predicted consequences of action. Although it would be possible to include sensorimotor delays in the generative model – see (Perrinet & Friston, 2014) for an example oculomotor control – a simpler solution rests on sensory attenuation: Sensory attenuation and action ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ If auditory predictions are imprecise, by virtue of sensorimotor delays, then their precision should be attenuated. This is an example of sensory attenuation or attenuation of sensory precision (Brown et al., 2013). If we attenuate the auditory precision by a factor of exp(2) – by reducing the log precision from 2 to 0 – the bird is now able to produce a well-formed soliloquy that would be recognised by another bird: see Fig. 5. The attenuation of sensory precision corresponds effectively to attending away from the consequences of action. A nice example of this is our inability to perceive (attend to) optical flow produced by saccadic eye movements: when visual motion or flow is produced exogenously – say by gently palpating the eyeball – they can be perceived but not when produced by oculomotor action. This simple but remarkable fact was first noted by Bell in 1823 and subsequently Helmholtz in 1866 (Wade, 1978). Put simply, we cannot speak and listen at the same time (Numminen, Salmelin, & Hari, 1999). From the perspective of the oculomotor system, this suggests that saccadic eye movements are imperceptible palpations of the world that are only attended to once complete; c.f., active vision (Wurtz, McAlonan, Cavanaugh, & Berman, 2011). Heuristically, this suggests that active inference presents in one of two modes; either attending to sensations or acting during periods of sensory attenuation. It also suggests that behaviour such as speech (that rests upon preordained sequences) may only be articulated with open loop control – a loop that is opened by sensory attenuation. Summary ~~~~~~~ These (and many other) simulations suggest that action and perception depend upon a delicate balance between the precision of proprioceptive and exteroceptive prediction errors that orchestrate the perceptual synthesis of sensations (caused by others) or realising sensory predictions (caused by self). It further suggests that certain behaviours (such as communication) are mediated by (transient) open loop control, which depends on sensory attenuation. This seems a plausible perspective on speech, in which violations or failures are generally recognised post hoc (e.g., slip of the tongue phenomena). It is interesting to speculate on the psychopathology that might attend a failure of sensory attenuation. In the first simulation above, we illustrated the failure to realise intended utterances when the precision of both proprioceptive and exteroceptive prediction errors were high. This might provide an interesting metaphor for failures of articulation; e.g., stuttering (Kalinowski, Armson, Rolandmieszkowski, Stuart, & Gracco, 1993). The converse pathology would be when both proprioceptive and exteroceptive prediction errors had too little precision, leading to psychomotor poverty and bradykinesia (e.g., Parkinson’s disease). See (Adams et al., 2013) for a more general discussion of precision in the context of psychopathology. One might ask whether sensory attenuation means that one cannot hear oneself. In quantitative terms, simulations like those above suggest the attenuation of sensory position is only quantitative: in other words, if the precision of prediction errors at extrasensory levels is greater than sensory precision, then descending proprioceptive predictions can be enacted with impunity – despite irreducible sensory prediction errors. This means that one can register and subsequently attend to violations in the consequences of action. Furthermore, it is possible to simulate talking to oneself by noting that higher levels of the cortical hierarchy (e.g., area X) are receiving prediction errors from lower areas (e.g., the higher vocal centre). In this sense, higher areas listen to lower areas, which embody sensorimotor content that may, or may not be, articulated (depending upon whether proprioceptive gain or precision is high or low). In the next section, we exploit the notion of sensory attenuation and model two birds that listen and sing to each other.","Finally, we turn to the perceptual coupling or communication by simulating two birds that can hear themselves (and each other). Each bird listened for two seconds (with a low proprioceptive precision and a high exteroceptive precision) and then sang for two seconds (with high proprioceptive precision and attenuated auditory precision). Crucially, when one bird was singing the other was listening. We started the simulations with random initial conditions. This meant that if the birds cannot hear each other, the chaotic dynamics implicit in their generative models causes their expectations to follow independent trajectories, as shown in Fig. 6. However, if we move the birds within earshot – so that they can hear each other – they synchronise almost immediately. See Fig. 7. This is because the listening bird is quickly entrained by the singing bird to correctly infer the hidden (dynamical) states generating sensations. At the end of the first period of listening, the posterior expectations of both parents display identical synchrony, which enables the listening bird to take up the song, following on from where the other bird left off. This process has many of the hallmarks of interactive alignment in the context of joint action and dialogue (Garrod & Pickering, 2009). Note that the successive epochs of song are not identical. In other words, the birds are not simply repeating what they have heard – they are pursuing a narrative embodied by the dynamical attractors (central pattern generators) in their generative models that have been synchronised through sensory exchange. As noted above, this means that both birds can sing from the same hymn sheet, preserving sequential and hierarchical structure in their shared narrative. It is this phenomenon – due simply to generalised (in this case identical) synchronisation of inner states – we associate with communication. The reason that synchronisation is identical is that both birds share the same prior expectations. In a companion paper we will illustrate how they learn each other’s attractors to promote identical synchronisation. Summary ~~~~~~~ In summary, these illustrations show that generalised synchrony is an emergent property of coupling active inference systems that are trying to predict each other. In this context, it is interesting to consider what is being predicted. The sensations in Fig. 7 are continuous and (for both birds) are simply the consequences of some (hierarchically composed and dynamic) hidden states. But what do these states represent? One might argue that they correspond to a construct that drives the behaviour of one or other bird to produce the sensory consequences that are sampled. But which bird? The sensory consequences are generated, in this setting, by both birds. It therefore seems plausible to assign these hidden states to both birds and treat the agency as a contextual factor (that depends on sensory attenuation). In other words, from the point of view of one bird, the hidden states are amodal, generating proprioceptive and exteroceptive consequences that are inferred in exactly the same way over time; irrespective of whether sensory consequences are generated by itself or another. The agency or source of sensory consequences is determined not by the hidden states per se – but by fluctuations in sensory attenuation (and proprioceptive precision). In this sense, the expectations are without agency. This agent-less aspect may be a quintessential aspect of shared perspectives and communication.","The arguments in this paper here offer a somewhat unusual solution to the theory of mind problem. This solution replaces the problem of inferring another’s mental state with inferring what state of mind one would be into produce the same sensory consequences: c.f., ideomotor theory (Pfister, Melcher, Kiesel, Dechent, & Gruber, 2014). Conceptually, this is closely related to Bayesian accounts of the mirror neuron system – in which generative models are used to both produce action and infer the intentions of actions observed in others (Kilner et al., 2007). The basic idea is that internal or generative models used to infer one’s own behaviour can be deployed to infer the beliefs (e.g., intentions) of another – provided both parties have sufficiently similar generative models. The example of communication considered above goes slightly further than this – and suggests that prior beliefs about the causes of behaviour (and its consequences) are not necessarily tied to any particular agent: rather, they are used to recognise canonical behaviours that are intermittently generated by oneself and another. In other words, the only reason that the simulations worked was because both birds were equipped with the same generative model. This perspective renders representations of intentional set and narratives almost Jungian in nature – presupposing a collective narrative that is shared among communicating agents (including oneself). For example, when in conversation or singing a duet, our beliefs about the (proprioceptive and auditory) sensations we experience are based upon expectations about the song. These beliefs transcend agency in the sense that the song (e.g., hymn) does not belong to you or me – agency just contextualises its expression. In this paper, we have focused on establishing the basic phenomenology of generalised synchrony in the setting of active inference. Interpreting this inference in terms of theory of mind presupposes that agents share a similar generative model of communicative behaviour. In a companion paper (Friston & Frith, 2015), we demonstrate how communication facilitates long-term changes in generative models that are trying to predict each other. In other words, communication induces perceptual learning, ensuring that both agents come to share the same model. This is almost a self evident consequence of learning or acquiring a model: if the objective of learning is to minimise surprise or maximise the predictability of sensations, then this is assured if both agents converge on the same model to predict each other. Clearly, we are not saying that theory of mind involves some oceanic state in which ego boundaries are dissolved. This is because using generative models to predict one’s own behaviour and the behaviour of others requires a careful orchestration of sensory precision and proprioceptive gain that contextualises the inference. This means, the generative model must comprise beliefs about the deployment of precision; including when to listen and when to talk; c.f., turn taking (Wilson & Wilson, 2005). We have not dealt with this aspect of communication here. However, it is clearly an important issue that probably rests on prosocial learning. It is interesting to note that a failure of sensory attenuation – in particular the relative strength of sensory and prior (extrasensory) precision – has been proposed as the basis of autism – whose cardinal features include an impoverished theory of mind (Happe & Frith, 2006; Lawson, Rees, & Friston, 2014; Pellicano & Burr, 2012; Van de Cruys et al., 2014). Furthermore, we have only considered a very elemental form of communication (and implicit theory of mind). We have not considered the reflective processing that may accompany its deliberative (conscious) aspects. This is an intriguing and challenging area (particularly from the point of view of modelling) that calls on things like deceit: e.g., (Talwar & Lee, 2008). Finally, we have not addressed non-verbal communication in motor and autonomic domains. A particularly interesting issue here is interoceptive inference (Seth, 2013), and how visual cues about the emotional states of others can be interpreted in relation to generative models of our own interoceptive cues: see (Kanaya, Matsushima, & Yokosawa, 2012). So far, in our discussion of communication, we have emphasised the importance of turn taking – and how this can lead to alignment of posterior expectations and identical synchronisation of inner states. However, one of the most important functions of communication is to enable one person to change another. For example, we can change another’s behaviour by giving them instructions in an experiment, or we can change their minds by helping them to understand a new concept (like free energy). For this aspect of communication there is an obvious asymmetry in the interaction, since the aim is to transfer information from one mind to another. Nevertheless, the basic features of the process remain as described above – and the asymmetric exchange will call upon generalised synchronisation. In this context, the naive speaker may assign greater salience or precision to predictions errors to update his (imprecise) priors about the nature of the new concept, while the knowledgeable agent will try to change the state of his listener. For the knowledgeable speaker, prediction errors indicate that his listener has still not understood (e.g., what must he be thinking to say that). For the ignorant speaker, prediction errors indicate that his concept is still not quite right: see Fig. 9 and (Frith & Wentzer, 2013) for discussion of the implicit hermeneutics. But the end result of the interaction will be a generalised synchronisation between the speakers. The emergence of such synchronisation indicates that the concept has been successfully communicated – and both parties can accurately predict what the other will say. As a result Chris’s concept (of free energy) will have changed a lot, but Karl’s will have changed a little as a result of communicating with Chris. From a mathematical perspective, we have promoted generalised synchrony (or synchronisation of chaos) as a mathematical image of communication. Furthermore, we have suggested that this synchronisation is an inevitable and emergent property of coupling two systems that are trying to predict each other. This assertion rests on the back story to active inference; namely, the free energy principle (Friston et al., 2006). Variational free energy provides an upper bound on surprise (i.e., free energy is always greater than surprise). This means that minimising free energy through active inference implicitly minimises surprise (or prediction errors), which is the same as maximising Bayesian model evidence. It is fairly easy to show that any measure-preserving dynamical system – including ourselves – that possesses a Markov blanket (here sensations and action) will appear to minimise surprise (Friston, 2013). The long-term average of surprise is called entropy, which means minimising surprise minimises (information theoretic) entropy. This is important because it means action (and perception) will appear to minimise the entropy of sensory samples, which are caused by external states. In turn, this means internal states will inevitably exhibit a generalised synchronisation with a system’s internal states. All that we have done in this paper is to associate the external states with the internal states of another agent. So why is generalised synchrony inevitable? This follows from the measure-preserving nature of coupled dynamical systems, which implies that internal states, external states – and the Markov blanket that separates them – possesses something called a random dynamical attractor. This attracting set of states plays the role of a synchronisation manifold that gives rise to generalise synchrony. The synchronisation manifold is just a set of states to which states are attracted to and thereafter occupy. All other points in the joint state space of internal and external states are unstable and will eventually end up on the synchronisation manifold. The simplest example of a synchronisation manifold would be the X equals Y line on a graph plotting an external state against an internal state (see Fig. 8). This corresponds to identical synchronisation. In other words, external and internal states track each other or – in the current context – both agents become identically synchronised. Crucially, the attractor (which contains the synchronisation manifold) has a low measure or volume. This is usually characterised in terms of the (fractional) dimensionality of the attractor. For example, in Fig. 8 the synchronisation manifold collapses to one dimension with the emergence of generalised synchrony. A measure of the attractor’s volume is provided by its measure theoretic entropy (Sinai, 1959). Although formally distinct from information theoretic entropy, both reflect the volume of the attracting set of dynamical states occupied by the systems in question. This means that minimising free energy (information theoretic entropy) reduces the volume (measure theoretic entropy) of the random dynamical attractor (synchronisation manifold); thereby inducing generalise synchrony. In conclusion, if the universe comprised me and you – and we were measure-preserving – then your states and my states have to be restricted to an attracting set of states that is small relative to all possible states we could be in. This attracting set enforces a generalised synchrony in the sense that the state you are in imposes constraints on states I occupy. It is in this sense that generalised synchrony is a fundamental aspect of coupled dynamical systems that are measure (volume) preserving. Furthermore, if we are both trying to minimise the measure (volume) of our attracting set (by reducing surprise or entropy), then that synchronisation will be more manifest. The notion of generalised synchrony may lend a formal backdrop to observations of synchronisation of brain activity between agents during shared perspective taking (Moll & Meltzoff, 2011). For example, functional magnetic resonance imaging, suggests that internal action simulation synchronizes action–observation networks across individuals (Nummenmaa et al., 2014). The arguments above suggest that synchronisation is not only fundamental for a shared experience of the world – it is a fundamental property of the world that we constitute.","The authors declare no conflicts of interest."],["Several years ago, a friend was visiting some monks. As they were eating together, he asked “If you don't believe in killing animals, why do you eat meat?” The answer: “Oh, we don't kill animals, the butcher does that. But everybody hates that guy.” Such deflection of responsibility in circumstances where an individual or group benefits from an object or good, in this case meat, that is obtained through an action the individual or group prefers to avoid, in this case, killing of a animal, is not unusual. Many people today have similar attitudes towards products and services that they consume, which are produced through actions that harm the environment. For example, consumers demand energy derived from fossil fuels and other goods and services entailing emissions of carbon dioxide, even as they blame energy companies for their emissions, which contribute to climate change. Yet, in practice, consumers may have fewer choices and their impacts are individually small and dispersed. If a policy solution is to be implemented to reduce these emissions, where would it be more acceptable? At the source (the butcher shop) that is blamed for the problem, or at the point of consumption (the kitchen) where concerned individuals may be nudged to act on their values? To combat carbon emissions,1 where should a mandatory emissions tax, permit system, or offsetting requirement be applied? “Upstream” on fossil fuel production, or “downstream” on consumers? And how do consumers react to these different policies? This issue – how consumers respond to different carbon pricing policy descriptions – is directly relevant in the context of policies being developed to reverse climate change or mitigate its effects (Ivanovich, Ocko, Piris-Cabezas, & Petsonk, 2019; Lee, 2018). For example, the International Civil Aviation Organization (International Civil Aviation Organization, 2016) has developed the Carbon Offsetting and Reduction Scheme for International Aviation (CORSIA), which caps the net carbon dioxide emissions of flights between participating countries at the average of 2019–2020 levels, and allows airlines to compensate or “offset” their direct emissions above those levels by paying for emissions reductions achieved from activities outside of the aviation sector. CORSIA is voluntary for countries over its first phases (2021–2026). Many countries are still deciding whether or not to opt-in for these phases and thus impose carbon offsetting requirements on arriving and departing flights. Airlines and their regulators will also have to determine how best to communicate (explicitly or otherwise) the associated carbon costs to their customers. Many countries and airlines are currently looking to make carbon regulation more attractive to consumers, to garner political support for the regulation and to maintain consumer demand in competitive markets. In this paper, we explore the psychological impact of different carbon emissions control approaches on consumers, depending on whether that extra cost or “price” to consumers is communicated as applied “upstream” on fossil fuel producers and importers, or “downstream” on consumers, as well as whether the extra cost or price is communicated to be a “tax” or an “offset”. In the following two sections, we review the academic literatures on these two dimensions. Upstream vs downstream carbon pricing ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In theory, a carbon price may be imposed at different points in the production and usage of fossil fuels and have equivalent impact on carbon emissions; regulation can be applied “upstream” on the extraction and importation of fossil fuels, or more “downstream” on the usage of goods and services such as electricity and air transport (Matthews, 2010). According to conventional economic theory, in the absence of market power, transactions costs, or other distortions, regulation of environmental pollutants is most cost-effective when applied at the point (upstream or downstream) that ensures the greatest flexibility for reducing emissions (see Mansur, 2011). In the case of fossil fuels, which are the largest contributor to global greenhouse gas emissions, the carbon dioxide emissions at the point of combustion are proportional to the carbon content of the fuels, assuming there is no way to scrub out or remove the emissions before they are combusted (or to capture and store them after combustion). As a result, upstream or downstream regulations provide the same chances of reducing emissions, making the two approaches theoretically equivalent from a conventional economics point of view. Entities regulated upstream will tend to raise their prices and thus pass the carbon price through to consumers, transmitting the carbon price downstream. Similarly, if consumers facing a downstream carbon price respond by reducing their demand, producers will then see the same reduction in profits as if they had faced the carbon price upstream. Given this theoretical equivalence of points of regulation, an upstream point of monitoring and regulation for fossil fuels (for example, requiring permits for the mining and importation of fossil fuel based on its carbon content) has a number of practical advantages relative to more downstream monitoring and regulation (such as a tax levied at the point of sale to consumers). In particular, economists have argued for the value of upstream regulation based on its administrative simplicity (requiring fewer sources to be monitored) and ability to ensure a broader economic coverage of all sources consuming fossil fuels in the economy, improving cost-effectiveness and reducing potential for emissions leakage (e.g., Stavins, 2007).2 Despite this practical rationale for an upstream approach, policy makers in most existing and proposed emissions trading systems to date have chosen to regulate fossil fuel emissions at more downstream points, despite significantly higher administrative costs and a reduction of coverage for sources below certain thresholds and sectors such as transportation. These choices may have a behavioral justification (Matthews, 2010). One argument is that downstream regulation increases the salience of the price and has different consequences in terms of managerial attention (Hanemann, 2009). These will induce greater innovation and behavioral change to reduce emissions, reducing the costs of the policy, even as the overall cap or price is still maintained under either type of regulation. These behavioral arguments for downstream regulation are largely “anecdotal” in the context of carbon pricing policies (Aldy & Pizer, 2009), and systematic empirical research on this specific issue of upstream vs downstream carbon prices is lacking. We aim to address this in the current research. What might be predicted from previous research about the psychological effects of upstream vs downstream pricing on consumers? The implications are mixed. On one hand, studies in the tax arena have shown that consumers are in fact more responsive to taxes that are embedded in a product's price, versus placed on a bill as an-add on to the sticker price (Chetty, Looney, & Kroft, 2009). “Downstream” prices may seem more personally relevant, and therefore more influential on behavior. On the other hand, motivated reasoning (Kunda, 1990) and self- serving bias (Campbell & Sedikides, 1999) predict that consumers will tend to hold others responsible for problems when possible, and that this is especially likely under self- threat (Campbell & Sedikides, 1999). Climate change is a moral issue for many people, and as such, it is likely that when this moral dimension is made salient, people will seek to hold someone accountable. Furthermore, consumers’ self-concept may be threatened if they are engaging in a carbon polluting activity, and therefore motivate them to blame upstream parties for the emissions. Likewise, to the extent that consumers believe they have low self-efficacy to lower carbon emissions, they may make an external attribution of control (Ajzen, 2002) and hold others responsible for the emissions. If consumers do hold others responsible for carbon emissions, they may be more likely to support more “upstream” carbon regulations, as these are perceived to hold the responsible parties accountable for their actions. In summary, while standard economic theory suggests that consumers would respond equivalently to “upstream” and “downstream” prices, the psychological literature suggests that consumers may respond more strongly to one or the other, and this topic demands empirical study. Product attribute framing ~~~~~~~~~~~~~~~~~~~~~~~~~ Turning to the formulation of a carbon price as a “tax” or an “offset”: does it make a difference? Product attribute framing can have a substantial impact on consumers (Levin, Schneider, & Gaeth, 1998). In a classic study, consumers were willing to pay more for a burger labeled as 75% meat than one labeled 25% fat, even after tasting it (Levin & Gaeth, 1988). Similarly, costs labeled as “taxes” are more odious to people than other equivalent financial costs (Ericson & Kessler, 2016; Kessler & Norton, 2016), and this may be particularly true for political conservatives (Hardisty, Johnson, & Weber, 2010; Sussman & Olivola, 2011). However, these tax aversion studies have only been conducted with U.S. participants, and could well be smaller in other countries and cultures around the world where taxes are less stigmatized, such as in Northern Europe. Why does this carbon “tax” aversion happen? According to Query Theory (Johnson, Haubl, & Keinan, 2007), when people consider a novel choice and “construct” their preference, their first thought has undue influence because it biases subsequent information processing in a confirmatory manner. In the U.S., “tax” is a “dirty word” for consumers (especially for Republicans) that triggers an immediate negative evaluation, in turn leading the overall evaluation of the carbon tax to be negative (Hardisty et al., 2010). In contrast, other terms (such as a carbon “offset”) do not trigger the same immediate negative reaction and cascade of negative thoughts. Another psychological factor driving tax aversion could be goal framing (Lindenberg & Steg, 2007; Steg, Lindenberg, & Keizer, 2016); it may be that taxes trigger a personal “gain” frame (in the goal framing context, a “gain” frame means “to guard and improve one's resources”, and thus also applies to personal losses, Lindenberg & Steg, 2007), which focuses consumers on the financial cost they will pay (rather than the environmental benefit). If this is true, tax framing may influence choices (increasing the decision weight of personal costs) similarly across countries. Likewise, “downstream” frames may induce a “gain” goal frame or “hedonic” goal frame (more focused on personal resources and pain avoidance), whereas “upstream” frames (more distant from the self) may induce a “normative” goal frame, more focused on morally correct behavior. As such, upstream carbon pricing policies may induce more pro-environmental behavior than downstream policies. A third factor driving acceptance of carbon pricing could be the extent that it focuses consumers on the impact of the policy on climate change. Previous work has found that a climate change frame (vs a financial cost frame or an energy usage frame) increases environmental behavior intentions (Spence, Leygue, Bedwell, & O'Malley, 2014). Carbon “tax” frames may focus consumers on the cost, and therefore be less effective than carbon “offset” frames that draw attention to the environmental impact of the policy. In short, to the extent these policy descriptions are visible to them, U.S. consumers ought to support paying for an “upstream” carbon offset system more than an equivalent “downstream” tax. The present research ~~~~~~~~~~~~~~~~~~~~ We test U.S. consumer perceptions of upstream versus downstream carbon regulation approaches in the aviation industry, along with two different policies: tax and offset. In combination, the two point of regulation frames and two policy frames result in four unique carbon price frames. We compare these four frames with two control conditions: a “no-fee” control in which there is no carbon price, and a “no-information” control in which the amount of the carbon cost is embedded into the price but where this is not broken out or explained to consumers in any way. In Study 1, we compared these six conditions in an exploratory fashion to determine which combinations were most and least desirable to consumers. In supplemental Study S1 (reported in the Online Supplement), we tested a replication of Study 1 with minor variations in the study procedure. In Study 2, we sought to further replicate the key results, and to uncover the psychological processes driving consumer psychology in this area: specifically, we measured perceptions of environmental impact and accountability. We examined whether consumers believe that certain frames would have a bigger impact on climate change and would ensure that those responsible for carbon emissions would pay to reduce them, in turn predicting flight choices and support for policy.","Participants were 588 residents of the United States (60% men; M age = 34.82 years [SD = 11.29]) who were recruited through Amazon. com's Mechanical Turk (MTurk) website. While MTurk represents a broad, diverse sample of U.S. participants, it is generally younger, more educated, and lower income than the general U.S. population (Paolacci, Chandler, & Ipeirotis, 2010). A sample size target of roughly 600 was set in advance, aiming for 100 participants per experimental cell, to ensure adequate power for all pairwise comparisons. Sample sizes in each experimental cell were as follows: no fee control = 94, no info control = 96, upstream offset = 105, upstream tax = 88, downstream offset = 102, downstream tax = 103. Experimental manipulation: proposed carbon regulatory policies ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ All studies were reviewed and approved by a University ethics board. Participants were randomly assigned to one of six conditions: “upstream offset”, “upstream tax”, “downstream offset”, “downstream tax”, “no-information control”, or “no-fee control”. Participants in the four experimental conditions first read a brief description of the policy (between 35 and 38 words), detailed in Appendix A. Other than varying the stream (“upstream” or “downstream”) and frame (“tax” or “offset”) of the proposed carbon fee, the content of the regulatory policy in each of the four carbon-fee labelling conditions was identical. Participants in the control conditions did not read about any policy. Measure of pro-environmental flight preference ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Next, participants were asked to imagine that they were planning a vacation, and were presented with five pairs of similarly priced flights to small island vacation destinations. Flight B always included an additional $14.00 carbon fee which varied in label depending on condition, as seen in Appendix B. For example, “Flight A to the island of Nevis for $322.99” or “Flight B to the island of Bequia for $310.99 plus an additional $14.00 carbon offset [tax] on aviation fuel production and importation [on airplane travel and cargo]”. The five flight pairs differed only slightly in price difference (in each case, the price of Flight B differed from that of Flight A by no more than +/− $6). In the “no-information control” condition, the $14.00 carbon fee applied to Flight B was embedded into the price of Flight B with no additional information provided. In the “no-fee control” condition, the $14.00 carbon fee was not applied to Flight B. Our main dependent variable of interest was how likely participants would be to buy Flight B instead of Flight A on a 7-point scale (from “1- Definitely Not” to “7- Definitely”), averaged across the five flight pair items. 3 Finally, all participants provided demographic information.","Data is publicly available for download at https://tinyurl.com/y75k6hkm. To test whether experimental condition affected pro-environmental flight preferences (i.e., preference for flights which included a carbon fee), we conducted an ANOVA on continuous preference for Flight B (collapsed across all flight pair choices), with condition as the independent variable. Results revealed a significant main effect of experimental condition, F (5,582) = 6.07, η2 = 0.05, p < .001. Fig. 1 shows that the “upstream offset” condition was most effective at eliciting pro-environmental flight preferences. We followed up the ANOVA with a series of pairwise comparisons; as this was an exploratory study, we corrected for family-wise error by using Tukey's HSD. Preference for flights carrying a $14.00 carbon fee in the “upstream offset” condition was significantly greater when compared to the “no information” control, meanupoff-noinfo = 0.88, p < .001, CI95 [0.39,1.36], and when compared to the “downstream tax” meanupoff-downtax = 0.55, p = .01, CI95 [0.07,1.02]. Also, the “no fee” control was preferred to the “no information” control, meanupoff- downtax = 0.63, p < .01, CI95 [0.14,1.13]. No other pairwise comparisons were statistically significant. We also ran an additional replication Study S1 (reported in the Online Supplement), which found the same pattern of results: preference for the “upstream offset” was greater than preference for the “downstream offset”, “upstream tax”, and “downstream tax” conditions. As political affiliation of U.S. consumers has previously found to be a significant moderator of environmental policy support (with stronger framing effects found for Republicans than Democrats; Hardisty et al., 2010; Hart & Feldman, 2018), we also analyzed political affiliation and found that it did indeed moderate the results, as detailed in the Online Supplement, Appendix G. Specifically, there is a significant interaction between political party and carbon pricing condition, F (5,344) = 6.75, p < .001, ηp2 = 0.08, such that Democrats respond particularly well to the “upstream offset” condition, and Republicans respond particularly badly to the “Downstream Tax” and “Downstream Offset” conditions.","Individuals who had read a brief description of a proposed regulation for a carbon fee described as a “carbon offset on aviation fuel production and importation” consistently reported a greater preference to purchase flights carrying a $14.00 carbon fee (versus similarly priced flights with no carbon fee) than, a) individuals who read other brief descriptions describing the carbon fee as a “carbon tax on airplane travel and cargo”, and b) individuals who did not read any description of an additional carbon fee and for whom this $14.00 fee was embedded in the cost of the flight. In addition, individuals in the “upstream offset” condition equally preferred flights carrying a $14.00 carbon fee as compared to those in the “no-fee control” condition in which the $14.00 carbon fee was not even applied. In a separate study, reported in the web supplement Appendix E, we replicated these results with lengthier (~140 word) policy descriptions. Study 2 was conducted to explore some potential psychological mechanisms that might underlie the preference for an “upstream offset” policy description. Specifically, we aimed to test whether the noted effects were mediated by participants’ perceptions of an “upstream offset” having a greater environmental impact and/or better ensuring that those responsible for carbon emissions pollution are actually held accountable.","Participants were 1215 residents of the United States (53% women; M age = 36.99 years [SD = 11.88]) who were recruited through Amazon. com's Mechanical Turk website. Data was collected in two waves (n = 406 and n = 809), one month apart, combined with two other (unrelated) studies. A sample size of roughly 1215 was targeted, with the goal of at least 400 participants per cell, based on the modest sample sizes observed in Study 1 (e.g., d = 0.31) and the need for a larger sample size to test multiple, simultaneously estimated mediation models. Sample sizes in each between-subjects condition were as follows: upstream offset = 420, upstream tax = 378, downstream offset = 417. Experimental manipulation: proposed carbon regulatory policies ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Participants were randomly assigned to one of three conditions: “upstream offset”, “upstream tax”, or “downstream offset” and read a corresponding policy description, identical to those used in Study 1. Our aim was to focus on the “upstream offset” condition and learn why it was most preferred (in Study 1 and Study S1), by comparing it to the two most similar labelling conditions: “downstream offset” and “upstream tax”. (To increase statistical power, this study did not include the “downstream tax”, “no fee” or “no information” conditions of study 1.) Measures of pro-environmental flight preference and policy support ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In Study 2 participants’ pro-environmental flight preference was gauged using a single flight pair choice. How likely participants would be to buy “Flight B to the island of Bequia for $310.99 plus an additional $14.00 carbon [offset] (or tax) on [aviation fuel production and importation] (or on airplane travel and cargo)” instead of “Flight A to the island of Nevis for $322.99” on a 7-point scale (from “1- Definitely Not” to “7- Definitely”) served as our measure of pro-environmental flight choice. In addition, participants were asked, “How much would you support the implementation of this carbon regulatory program on a national level?” Responses on this 7-point scale (from “1- Strongly opposed” to “7- Strongly support”) served as our measure of policy support. Participants in the second wave of data collection (n = 809) also reported preference for Flight B on the additional flight pair choice items that were included in Study 1. For this subsample we were able to create a 5-flight composite collapsed across flight pair choices and run identical analyses. The results of these are very similar, and are reported in the online supplement, Appendix I. Measures of perceived environmental impact and carbon emissions accountability ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Next, participants were asked two additional questions: “Would you expect this carbon regulatory program to have a significant impact on climate change?” and “Does this carbon regulatory program ensure that those who are responsible for carbon emission pollution are the ones who pay to reduce it?” Responses on both of these 7-point scales (from “1- Definitely not” to “7- Definitely” and from “1- Not at all” to “7- Very much”, respectively) served as our potential mediators of interest. Participants in all three conditions also provided demographic information (e.g., age, sex, income).","Data is publicly available for download at https://tinyurl.com/y75k6hkm. To test whether experimental condition affected pro-environmental flight preference (i.e., preference for the flight which included a carbon fee), we conducted an ANOVA on continuous preference for Flight B, with condition as the independent variable. Results revealed a significant main effect of experimental condition, F (2,1212) = 4.25, η2 = 0.007, p = .011. Table 1 shows that the “upstream offset” condition appears to be most effective at eliciting pro- environmental flight preferences. The “upstream offset” condition inspired greater pro- environmental flight preference than the “upstream tax” and marginally greater pro- environmental flight preference than the “downstream offset” conditions; t (796) = 3.00, d = 0.21, p = .003, and t (835) = 1.90, d = 0.13, p = .058, respectively. To test whether experimental condition affected how much participants would support the policy on a national level, we conducted an ANOVA on policy support, with condition as the independent variable. Results revealed a significant main effect of experimental condition, F (2,1215) = 5.81, p = .003, ηp2 = 0.009. Table 1 shows that the “upstream offset” condition appears to be most effective at eliciting policy support. Policy support in the “upstream offset” condition was marginally greater when compared to the “upstream tax”, t (796) = 1.79, d = 0.13, p = .073, and significantly greater when compared to the “downstream offset” conditions, t (835) = 3.46, d = 0.24, p = .001. To test whether experimental condition affected how much participants believed that the carbon regulatory program would have a significant impact on climate change, we conducted an ANOVA on perceived environmental impact, with condition as the independent variable. Results revealed a significant main effect of the experimental condition, F (2,1212) = 4.95, ηp2 = 0.007, p = .008. Table 1 shows that the “upstream offset” condition appears to be most effective at eliciting high perceived environmental impact. Perceived environmental impact in the “upstream offset” condition was significantly greater when compared to the “upstream tax”, t (796) = 2.79, d = 0.20, p = .005, and the “downstream offset” conditions, t (835) = 2.70, d = 0.18, p = .007. To test whether experimental condition affected how much participants believed that the carbon regulatory program would ensure that those who are responsible for carbon emission pollution are the ones who pay to reduce it, we conducted an ANOVA on perceived carbon emissions accountability, with condition as the independent variable. Results revealed a significant main effect of the experimental condition, F (2,1212) = 8.83, ηp2 = 0.014, p < .001. Table 1 shows that the “upstream offset” condition appears to be most effective at eliciting perceptions of holding accountable those responsible for the carbon emissions. Perceived carbon emissions accountability in the “upstream offset” condition was not statistically different when compared to the “upstream tax” condition, t (796) = 1.27, d = 0.09, p = .21, but was significantly greater when compared to the “downstream offset” condition, t (835) = 4.14, d = 0.288, p < .001. Mediation analyses ~~~~~~~~~~~~~~~~~~ We conducted two sets of mediation analyses to examine the processes associated with flight preference and policy support. We now examine each of these in turn. Flight preference Both perceived impact and perceived accountability were correlated with flight preference, as seen in Table 2. Furthermore, in a GLM with frame (3-levels: upstream offset, downstream offset, and upstream tax), perceived impact, and perceived accountability predicting flight preference, both impact, F (1,1210) = 119.11, p < .001, ηp2 = 0.09, and accountability, F (1,1210) = 26.77.11, p < .001, ηp2 = 0.02, remained as significant predictors of preference, while frame was reduced to marginal significance, F (2,1210) = 2.31, p = .10, ηp2 = 0.004. We also analyzed political affiliation and found that it did indeed moderate the results, as described in the Online Supplement, Appendix G. Specifically, Republicans were more responsive to policy description differences than Democrats, consistent with previous research on U.S. participants (Hardisty et al., 2010; Hart & Feldman, 2018). Next, we conducted a simultaneous indirect effect analysis of the frame → impact → flight preference pathway and the frame → accountability → flight preference pathway. As seen in Fig. 2a, each leg of each mediation pathway is significant. We used a recommended bias-corrected bootstrapping procedure in R (Shrout & Bolger, 2002), with 5,000 bootstrap samples. To test the 3-level categorical framing variable, we used two dummy coded variables, and focus on the upstream offset condition (coded 1) versus the other two conditions (coded 0). This revealed significant mediation via both perceived impact, β = 0.13, CI95 [0.03,0.24], p < .01, and perceived accountability, β = 0.09, CI95 [0.05,0.16], p < .001. Policy support Turning to the processes associated with policy support, it is again apparent that both perceived impact and perceived accountability are correlated with policy support, as seen in Table 2. As seen in Fig. 2b, each leg of each mediation pathway is significant. Again, in a model with frame (3-levels: upstream offset, downstream offset, and upstream tax), perceived impact, and perceived accountability predicting policy support, both impact, F (1,1210) = 483.98, p < .001, ηp2 = 0.29, and accountability, F (1,1210) = 96.13, p < .001, ηp2 = 0.07, remained as significant predictors of support, while frame dropped to non- significance, F (2,1210) = 1.26, p = .29, ηp2 = 0.002. Next, we conducted a simultaneous indirect effect analysis of the frame → impact → policy support pathway and the frame → accountability → policy support pathway, as seen in Fig. 2b, using the same bootstrapping procedure with 5,000 bootstrap samples. This revealed significant mediation via both perceived impact, β = 0.19, CI95 [0.06,0.34], p < .01, and perceived accountability, β = 0.13, CI95 [0.07,0.21], p < .001. Political affiliation ~~~~~~~~~~~~~~~~~~~~~ As detailed in the Online Supplement, there was a significant interaction between political affiliation and carbon pricing condition, F (2,773) = 4.96, p < .01, ηp2 = 0.01, such that carbon pricing condition had a slightly stronger effect on the policy support of Republicans than Democrats. In particular, the “upstream tax” pricing condition was supported by Democrats much more than Republicans.","The results of Study 2 are consistent with those of Study 1, while adding important additional nuances which may hint at the psychological mechanisms underlying the noted success of the “upstream offset” policy description. When a proposed carbon fee was described as a “carbon offset on aviation fuel production and importation” (versus other descriptions), individuals were more likely to purchase a flight carrying the carbon fee instead of a similarly priced flight with no fee, and were more likely to support the proposed carbon regulation policy. These individuals were also more likely to perceive the policy as having a more significant impact on climate change, and believe that the policy would better ensure that those responsible for carbon emissions pollution were the ones held accountable for paying the fee. Mediation analyses revealed that participants’ increased pro-environmental flight preference and heightened policy support was driven by perceptions that the “upstream offset” policy would have a greater impact on climate change (versus “upstream tax” and versus “downstream offset” conditions) and, would better ensure that those who are responsible for carbon emissions pollution are actually the ones held accountable for covering the cost (versus the “downstream offset” condition only).","Across three studies, U.S. consumers were consistently more likely to choose to purchase a flight with a carbon price when the additional price was described as a “carbon offset for aviation fuel production and import” than when it was described using other frames, such as a “carbon tax for airplane travel”. Notably, this upstream offset frame was popular enough among consumers that it counteracted the expected preference to avoid the $14.00 additional cost of the fee. In addition to the two studies reported above, we ran another study (reported in the online supplement, Appendix E) which replicated the flight preference results. The description of the point of regulation (upstream vs downstream) does seem to have a notable impact on consumer preferences. A “downstream” carbon offset was not nearly as popular as an “upstream” carbon offset. Why? The upstream offset is perceived both to help the environment more and to hold accountable those responsible for the emissions. In other words, the upstream offset condition addresses both the perceived causes and consequences of the emissions. This is true even though the cost to consumers for each policy was held constant in our studies, regardless of whether the price was described as applying upstream or downstream. Our findings contribute to previous studies on tax aversion. In addition to replicating the general finding of tax aversion in U.S. consumers (Hardisty et al., 2010; Sussman & Olivola, 2011) and the effectiveness of environmental frames (Parag, Capstick, & Poortinga, 2011), we introduce and measure the psychological dimension of “upstream” vs “downstream” framing: not only are taxes disliked, but this is particularly true for “downstream” taxes on consumers, as opposed to “upstream” offsets for producers. We also contribute new explanations for tax aversion in environmental policy: that taxes are not perceived to be effective for making an environmental impact, and that taxes may not be seen as holding the right people accountable. In future research, it would be interesting to see if “upstream” vs “downstream” descriptions influence the order and balance of thoughts in the way predicted by Query Theory (Johnson et al., 2007): perhaps “downstream” costs induce a negative cognitive bias in the same way that taxes do, influencing subsequent information processing. An additional cost or penalty applied to “me” may be more likely to trigger an immediate negative evaluation. It would also be interesting to study how “upstream” vs “downstream” descriptions influence goal framing (Lindenberg & Steg, 2007; Steg et al., 2016), and whether perhaps the “upstream” and “offset” frames induce a “normative” goal of doing the right thing, whereas the “downstream” and “tax” goes may induce a “gain” goal of maximizing resources. The implications for policy are somewhat nuanced. Countries or states that wish to enact a carbon price may want to use the “upstream offset” frame (paired with a brief description), especially if competing with other countries without a carbon fee. Our results suggest that consumers may in fact prefer airline flights with an upstream carbon offset; this preference may be strong enough to counteract any additional cost to the country or state that implements it (though this should be tested in future research in field studies). The implementing country could potentially realize further benefits if the offset investment helps finance sustainable low-carbon development in that country. Furthermore, aviation consumers might be more accepting of “upstream offset” regulation than “downstream tax” regulation. In fact, our findings suggest that customers may be willing to purchase tickets that include appropriately described carbon offsets even if the cost is higher; these descriptions may still be effective even when they are as short as 28 words in length. These findings have policy implications for the new market-based measure adopted for the international civil aviation sector, which has an offsetting program as a central feature, with a relatively upstream point of obligation at the level of airlines. Our research suggests that airlines might wish to highlight their carbon offset purchases for their customers in order to increase customer receptivity. Our findings also suggest that participation in CORSIA could thus bring overall economic benefits rather than costs for participating airlines and that countries may thus benefit from opting-in to the initial voluntary phase. If the goal of a carbon price is to limit the aviation sectors’ climate impact, a carefully vetted upstream offset program with strong offset quality requirements may be an effective solution. Future research could further examine if customer preferences vary with the type of offset and how best to communicate about offsets to customers to elicit the most favorable response. The studies reported here have several important limitations. First, the scenarios and answers are hypothetical, and therefore participants may have tried to give answers that “look good”, but that do not represent their real views and behavior. For example, in Study 1, we found that the flight including a $14 upstream offset was equally preferred (or even slightly more preferred) as compared to a flight with no carbon fee. When spending actual money, consumers may weight that $14 cost more heavily, and prefer the “no-fee” flight. Therefore, these results need to be tested with real flight choices in the future. Second, we only sampled U.S. consumers, and the results may be different in other regions and cultures. For example, taxes may be less odious in other countries (such as Northern European countries), and environmental regulation may be more popular. This should be explored in future research. Third, the sample population we used, Amazon Mechanical Turk, is not representative of the U.S. population. Specifically, MTurkers tend to be younger, more educated, and lower income than the general U.S. population (Paolacci et al., 2010). As such, this group may be more price sensitive, and also more pro-environmental than the U.S. population, both of which could affect our results. Fourth, while the key findings of the present research were replicated across three studies (the two studies in the main manuscript plus one in the supplemental material), the effect sizes are small. As such, while framing of carbon prices may be a helpful tool, it should not be considered a panacea for consumer acceptance of carbon prices, and should be combined with other measures. Up/downstream framing is a useful new dimension in environmental consumer psychology. It can readily be applied in the airline industry, as well as in other domains and contexts more broadly, opening new avenues for research and practice. For example, when products that include a charitable donation (e.g. “$1 from every purchase is donated to Green Peace”), who gets the credit for the good deed, the consumer, the firm, or both? We suspect that “downstream” pricing may be more effective in these contexts, but this should be tested in future research."],["The Intermodal Preferential Looking paradigm provides a sensitive measure of a child's online word comprehension. To complement existing recommendations (Fernald, Zangl, Portillo, & Marchman, 2008), the present study evaluates the impact of experimental noise generated by two aspects of the visual stimuli on the robustness of familiar word recognition with and without mispronunciations: the presence of a central fixation point and the level of visual noise in the pictures (as measured by luminance saliency). Twenty-month-old infants were presented with a classic word recognition IPL procedure in 3 conditions: without a fixation stimulus (No Fixation - noisiest condition), with a fixation stimulus before trial onset (Fixation, intermediate), and with a fixation stimulus, a neutral background and equally salient images (Fixation Plus - least noisy). Data were systematically analyzed considering a range of data selection criteria and dependent variables (proportion of looking time towards the target, longest look, and time-course analysis). Critically, the expected pronunciation and naming interaction was only found in the Fixation Plus condition. We discuss the impact of data selection criteria and the dependent variable choice on the modulation of these effects across the different conditions. --------------------------------------------------------------------------------","Over the last four decades, a considerable amount of energy and creativity has been devoted to designing and testing numerous experimental methods to investigate early speech perception and language comprehension in young children. One of the most popular methods is the head-turn preference paradigm (Polka & Bohn, 1996; Werker, Polka, & Pegg, 1997; Werker & Tees, 1983, 1984) which is primarily used with infants aged from 5 to 16 months of age to investigate listening preferences and discrimination. The study of word recognition or word learning from the age of 12 months (Schafer & Plunkett, 1998) relies on two paradigms, the Switch task (Stager & Werker, 1997; Werker, Fennell, Corcoran, & Stager, 2002) and the Intermodal Preferential Looking paradigm (IPL, Bailey & Plunkett, 2002; Golinkoff, Hirsh-Pasek, Cauley, & Gordon, 1987; Swingley & Aslin, 2000), also called looking-while-listening procedure (Fernald, Zangl, Portillo, & Marchman, 2008). In principle the IPL procedure is more versatile than the Switch task, whereby a new label is presented alongside a new visual item until a looking time threshold is reached, followed by a trial where the label is maintained or replaced by another (this is the switch trial). Indeed, whereas the Switch task is designed for novel word learning situations, the IPL allows for the investigation of novel word learning (e.g., Gurteen, Horne, & Erjavec, 2011; Schafer & Plunkett, 1998; Swingley & Aslin, 2002, 2007) together with familiar word recognition (e.g., Fernald, Pinto, Swingley, Weinberg, & McRoberts, 1998; Fernald et al., 2008; Houston-Price, Mather, & Sakkalou, 2007; Mani, Coleman, & Plunkett, 2008; Mani & Plunkett, 2007, 2008, 2011a; Ramon-Casas, Swingley, Sebastián-Gallés, & Bosch, 2009; White & Morgan, 2008), mutual exclusivity (Houston-Price, Caloghiris, & Raviglione, 2010) and most recently has been adapted for priming tasks (e.g., Arias-Trejo & Plunkett, 2009, 2013; Mani, Durrant, & Floccia, 2012; Styles & Plunkett, 2009). The standard procedure consists in presenting pairs of images horizontally on a screen for several seconds and, mid-trial, playing a target word or a carrier sentence. A trial is thus divided in a pre-naming and a post-naming phase (for longer post-naming windows, see Zangl, Klarman, Thal, Fernald, & Bates, 2005). For the duration of the experiment, eye movements are recorded by cameras mounted above each image (or eye-tracking when available). These eye movements are then time-locked onto each trial and traditionally manually coded frame by frame. With the growing use of eye trackers, gaze coding is now often automatic. Fernald and colleagues (2008) provided a comprehensive review of the procedure and improvements added progressively since the first introduction of the paradigm, and listed factors that need to be controlled: both images have to be matched for size and visual salience; auditory stimuli have to be controlled for duration across items; target side has to be counterbalanced overall. One of their recommendations is that across all participants both objects in a given trial should be used as target and as distracter, since it is the best control to avoid any preference for one stimulus over another. Although desirable, such a control is not always possible given the restricted choice of stimulus items in young children and the need for a sufficient number of trials per participant. Some experiments have controlled for this possible preference effect (Mani & Plunkett, 2007; Swingley & Aslin, 2000, 2002), presenting the same visual stimuli at least twice, while others have not (Durrant, Delle Luche, Cattani, & Floccia, 2014; Floccia, Delle Luche, Durrant, Butler, & Goslin, 2012; Mani et al., 2008) but still found comparable results. One possible compromise, as suggested by Fernald et al. (2008), is to ensure an equal preference for both pictures during the pre-naming phase by monitoring looking times in silence during a pilot experiment. However, to our knowledge, such a pretest or control for the absence of a pre-naming image bias is not reliably reported in the literature (with the exception of the results section in Swingley, Pinto, & Fernald, 1999). Another way of controlling for pre-naming visual preferences is to take into account looking behaviour in the pre-naming phase in the statistical analyses, which is frequently reported (e.g., Mani & Plunkett, 2007; Meints, Plunkett, & Harris, 1999). Finally, Fernald et al. argue that, contrary to adult visual experiments, a central fixation point right before naming is not necessary as children would not follow such an implicit instruction (note that White & Morgan, 2008, used a centering stimulus before the pre- and the post-naming phases, while Gurteen et al., 2011, used a centering light before post-naming). Despite the excellent review by Fernald et al. (2008) the relatively recent addition of the IPL paradigm to the field of developmental psycholinguistics means that researchers often face choices regarding the procedure itself and the methods of analyses, all of which can have important consequences on the observation of an experimental effect. For example, as demonstrated by Arias-Trejo and Plunkett (2010), choosing a distracter image which is perceptually close to the target image (e.g., a balloon paired with an egg) can result in uninvited interference effects so that 18- to 24-month-olds fail to identify the target image. The consequences of other methodological choices (such as the use of a fixation point) on the robustness of the experimental effects are largely unknown. As we will show below, a review of the recent literature reveals a great deal of variation in many aspects of the procedure, as well as in the selection of the dependent variables used for the analysis of looking times. Appendix A provides details on methodological aspects such as the presence of a central fixation stimulus or the duration of the pre- and post- naming phases across a range of studies that have used the IPL methodology. The aim of the current study is to complement and extend Fernald et al.’s review by examining how the different methodological choices and the different methods of looking time analyses impact on the observation of significant results. In three experiments testing familiar word recognition with 20-month-olds, we manipulated the level of visual noise (with saliency as measured by luminance and the presence of a central fixation point), and provided a systematic and thorough analysis of looking times using those methods most representative of the current literature. The main objective of this paper is to provide researchers with some data-grounded recommendations about the best practices when using the IPL procedure. Central fixation point ~~~~~~~~~~~~~~~~~~~~~~ A review of the literature using the preferential looking paradigm shows that around half the experiments use a centering stimulus at the beginning of each trial, visual or auditory (Curtin, 2010; Dittmar, Abbot-Smith, Lieven, & Tomasello, 2008; Meints et al., 1999; Meints, Plunkett, Harris, & Dimmock, 2002; Meints, Plunkett, Harris, & Dimmock, 2004; Schafer & Plunkett, 1998, see Appendix A) while the other half do not report such a practice (Bailey & Plunkett, 2002; Ballem & Plunkett, 2005; Mani et al., 2008; Mani & Plunkett, 2007, 2008). This stands in contrast with ERP studies, and adult experiments generally, where a fixation stimulus is systematically presented to centre the participant's attention before trial onset (Kuipers & Thierry, 2011). When a central fixation point is used, trials are always triggered by the experimenter, but not systematically when such cue is not used, in which case trials are sometimes automatically interspaced (Ramon-Casas et al., 2009; Swingley, 2003, 2007). Although it is not possible, from these studies, to draw direct comparisons between results obtained with and without fixation points given the variety of investigated topics, one can estimate that centering attention, even furtively, should benefit the procedure and the quality of the data, especially since it ensures that the child is attentive to the screen immediately before trial onset. The presentation of an image in the centre of the screen after termination of a trial has multiple advantages: (i) since the trial is triggered only if the child is looking at the centre, it ensures the child is attentive and active; (ii) attention is attracted back to the middle of the screen, giving the same weight to the probability that the first look will be at the target or the distracter once the trial begins; (iii) looking at the centre right before trial onset should encourage the child to explore all new stimuli that appeared in her peripheral vision. This is of particular importance when considering that trials where the child does not look at both images in the pre-naming phase can be discarded in the statistical analyses (e.g., Mani & Plunkett, 2007); and (iv) by having something to look at for the whole duration of the experiment, the entire procedure becomes dynamic and eventful, maintaining the child's interest. One of the aims of this study will be to verify if looking behaviour is affected by the potential noise reduction provided by a fixation point in a classic IPL task. Quality of visual stimuli ~~~~~~~~~~~~~~~~~~~~~~~~~ Fernald et al. (2008) recommended controlling the visual stimuli for size, animacy and salience (in the sense of visually engaging images, especially by matching objects for animacy). Regarding saliency, experimenters decide on the pictures without any objective measures. Visual stimuli are often static colour photographs of objects on a white or grey background (e.g., Mani & Plunkett, 2007; Swingley & Aslin, 2000), sometimes a mix of realistic drawings and photographs (Swingley, Pinto, & Fernald, 1999), or quite exceptionally line drawings, coloured (White & Morgan, 2008; White, Morgan, & Wier, 2005) or plain (with made up animals, Mather & Plunkett, 2011). So as to enhance interest in the visual stimuli, pictures sometimes move in synchrony on a vertical axis (Swingley, 2003; Swingley & Aslin, 2002, 2007). In word learning studies, made up objects (obtained by editing colour photographs as in Schafer & Plunkett, 1998) are visually comparable to real objects. To our knowledge, only Gurteen et al. (2011) have presented real objects to the participating infants. Such variability in the selection of visual stimuli in the literature has been enabled by the possibility of retrieving images and photographs from the internet, departing from the perceptually controlled line drawings from Snodgrass and Vanderwart (1980) or their coloured version (Rossion & Pourtois, 2004) to achieve more naturalistic representations. One way to control for the quantity of information provided by photographs is to remove any background, even though the percept is less naturalistic looking. This practice is supported by Meints et al. (2004) who showed that background affected word recognition. Younger children (15 months) do not recognize a sheep when it is pictured with a naturalistic background (e.g., a sheep on grass), while they do when the sheep is presented without background or on an unusual – or less rich – background. Older children recognize the sheep regardless of background, although the distracting effect of the typical background was still observed to a certain extent. Perhaps estimating the basic visual salience of the stimuli would be another, quantifiable, step towards ensuring that target images are not more attractive than their corresponding distracters, to complement experimenter judgments. Note that this is a purely visual control of salience, away from the subjective salience discussed by Fernald et al. (2008) or the more cognitive saliency maps (for a review, see Althaus & Mareschal, 2012). This will be achieved in the current study by performing cross-correlations of the pairs of images presented (Chinga & Syverud, 2007). Images are transformed into a matrix containing the luminance of each pixel, then into a vector. The cross-correlation compares then the two vectorized images: a high correlation score will mean that the two images are similar in salience. By contrasting a more or less visually noisy set of pictures, we will examine its effect on looking behaviour. Methods of analysis ~~~~~~~~~~~~~~~~~~~ Regardless of the task the participant is engaged in, all experiments in the IPL literature divide an experimental trial into a pre- and a post-naming phase, based usually on the onset of the target word (with the exception of priming studies in which there is no pre-naming phase, e.g., Arias-Trejo and Plunkett, 2009). The selection of analysable data as well as the choice of dependent variable varies according to research groups and, on occasion, differs for a single experiment (see Appendix A). The most typical time window of interest starts 367 ms after word onset (and less frequently word offset): indeed it has been established that a minimum of 233 ms is necessary to obtain stimulus- related saccades, and that this latency depends on age, vocabulary size or task complexity (Fernald et al., 2008; Fernald, Perfors, & Marchman, 2006; Zangl & Fernald, 2007; Zangl et al., 2005). This time window usually ends at 2000 ms, as it is generally considered that later looking behaviour is no longer related to the processing of the auditory stimulus. When plots of the time course for proportions of looks to the target are included, usually at end of the result section (Arias-Trejo & Plunkett, 2010; Fernald et al., 2008; Swingley & Aslin, 2000), the time window can then be justified a posteriori, the visual inspection confirming that roughly 2000 ms after word onset, looking behaviour resumes to chance level (that is, equal looks to the target and the distracter). However, since latency is a function of at least age (Zangl & Fernald, 2007) and vocabulary size (Fernald et al., 2006), the whole looking behaviour can also be influenced by task difficulty. This is the case for example when mispronunciations of words are minor (Mani et al., 2008; Mani & Plunkett, 2011a; Swingley, 2007; Swingley & Aslin, 2002). A systematic time window, regardless of data distribution in the post-naming phase, may overlook meaningful looks if children are still looking at the target after 2000 ms. It seems recommendable (see Fernald et al., 2008) that the first, and not the last, step in data analysis should be the systematic plot of the unfolding looking behaviour, so as to ascertain that the window of analysis comprises all the word processing and task related looks. This is common practice in EEG or MEG experiments, since electrophysiological markers can vary in location and time period (e.g., Bastiaansen, van der Linden, ter Keurs, Dijkstra, & Hagoort, 2005; Hagoort & Brown, 2000). Perhaps a way to enhance the precision of the results would be to determine statistically the exact time window when the two conditions differ (target vs. distracter in simple naming tasks, correct vs. incorrect pronunciation in mispronunciation experiments). To our knowledge, with the IPL paradigm, only one instance of such an analysis has been published so far (von Holzen & Mani, 2012, see Maris & Oostenveld, 2007 for a detailed explanation); the authors showed that in addition to standard comparisons of looking times averaged across pre- and post-naming trials, it was possible to identify accurately a specific time window where performance between conditions differed (in their case, 1140–1580 ms after target word onset). There are two advantages for this method: first, representing time course plots gives a dynamic evaluation of visual/linguistic processing; secondly, this data driven method is objective and prevents a priori judgement of the data. In relation to the looking behaviour, two types of dependent variables are usually considered: proportion of target looking (taking into account, or not, the pre-naming phase, see Appendix A), and latency of shifts to the target. They are respectively assimilated to a correct response and a reaction time (see Fernald et al., 2008). Other measures have also been used, and usually provide comparable direction of results, such as total looking time (that is, the sum of looks to the target during the post-naming phase minus those to the distracter), or the longest look (longest single fixation to the target). Data filtering or pre-processing is where the greatest variation across experiments is observed (see Appendix A). In word recognition tasks, with or without a mispronunciation element, words are selected so that they are likely to be known a priori by all participants according to standardized norms (thus reducing considerably the number of potential stimuli, re. Ramon-Casas et al., 2009; Swingley et al., 1999), or by at least 50% of children of the corresponding age (e.g., Styles & Plunkett, 2008). Then, some authors further filter the data by analyzing only trials where parents report the words as known by the child (e.g., Mani & Plunkett, 2011b). Whether it is necessary or not to check for infants’ knowledge of distracter will be addressed here. On the one hand, if the child does not know the distracter, she might be looking more at the target once it has been named simply because the target object is the only object for which she has a name, and not because she recognizes the link between the label and this object; this would artificially inflate the target looking time. On the other hand, a child who does not know the distracter's name would look longer at its picture in mispronunciation trials, in the spirit of the study by White and Morgan (2008) whereby unknown objects were presented as distracters. This may be an advantage for strengthening mispronunciation effects. The impact of filtering out trials where the child does not know both the target word and the distracter will be evaluated in the current study. The criteria used to select the attended trials is also variable, with some considering only long enough fixations to the images (1500 ms in each phase, Bailey & Plunkett, 2002), or fixation to both images in the pre-naming phrase or at least, throughout the trial (Mani & Plunkett, 2007). Finally, data cleaning is achieved by excluding participants not contributing to all experimental conditions (Fernald et al., 2006; second analysis in Styles & Plunkett, 2008), or whose data points fall outside normality (Fernald et al., 2006; see also Mani & Plunkett, 2007). Although all these types of pre-processing or filtering allow for cleaner datasets, comparability across experiments would benefit from consistent practice, preferably on the measures that are the most conservative. Falling on an agreement on exclusion criteria or on the best-suited age-specific time window would be desirable. Indeed, some analyses only include trials where the child is looking at a picture (final analysis in Fernald et al., 1998), while others reject children not contributing to all conditions (e.g., Fernald et al., 2006; second analysis in Styles & Plunkett, 2008), set a minimum looking time (e.g., Bailey & Plunkett, 2002; Ballem & Plunkett, 2005) or only include trials were both pictures are fixated (Mani & Plunkett, 2007, 2008). The goal of the present research is to examine how the different methods of data analyses are resistant to methodological alterations, or noise, such as image quality and the presence of a central fixation stimulus. For this purpose, we ran three versions of a classic IPL procedure testing the detection of mispronunciation of familiar words, varying the pictures’ saliency and background, and the presence of a fixation point. For each experiment, we evaluated how the degree of visual noise, together with the different criteria for data selection, modified the different dependent variables. Here, the stimuli were comparable to those in Mani and Plunkett (2007), in which 15-to-24-month-olds were presented with two images on both sides of a screen and heard, mid-trial, correctly pronounced in a carrier sentence such as “Look, dog!” for half of the trials, or as a mispronounced version of the target word (“Look, bog!”) in the other half of trials. If children recognize lexical entries of familiar words only if they are pronounced correctly (as is expected from the age of 18 months, e.g., Mani & Plunkett, 2007), a naming effect should be observed in correctly pronounced trials, but not (or significantly less) in mispronunciation trials resulting in an interaction between naming and pronunciation – the key result in these studies. The participants in the current study are 20 month olds, and so results comparable to 18 month olds can be expected, if not stronger, because their lexical repertoire increases steadily and their phonological sensitivity seems stable around these ages (for a developmental Switch task, see Werker et al., 2002). As presented earlier, we hypothesize that adding a fixation stimulus between trials should enhance the quality of the data. Indeed, since experimental trials are then only triggered when the child is attentive to the screen, post-hoc measures of attentiveness (by checking the videos, re. Fernald et al., 2008; or excluding trials with looks <1500 ms, re. Bailey & Plunkett, 2002) are less critical yet desirable. The added advantage of having a centred fixation stimulus is that it should entice the participants into looking at the two images that appear in their peripheral vision at trial onset. The second manipulated methodological choice is the uniformity of picture background and the choice of visual stimuli. As tested by Meints et al. (2004), a typical background is distracting and reduces the naming effect in younger children, and to a certain extent in 18-month-olds. To capitalize on these findings and fully appreciate the extent of any distraction relating to a naturalistic background on the naming effect, we compared two sets of images, one where the content was naturalistic (mostly with a background), and one where there was no background, leaving the stimuli in isolation. In addition, we examined the effect of image saliency on looking behaviour, especially in the pre-naming phase, which, to our knowledge, has never been investigated experimentally in infants, despite recommendations to control for it (Fernald et al., 2008). To sum up, we manipulated the presence/absence of a central fixation point together with the uniformity of picture background and picture salience, to evaluate the effect of experimental noise in an auditory word recognition task. In the first condition (No Fixation), no central fixation stimulus was used and no attempt was made to control for the picture background colour or salience. This is the noisiest condition. In the second condition (Fixation, intermediate noise condition), a central fixation point was added and the same images were used as in the No Fixation condition. In the third condition (Fixation Plus), the central fixation point was augmented by a systematic absence of background colour for pictures and target- distracter pairs were selected so that their salience was highly correlated. This is the least noisy condition. We expect that adding the fixation stimulus should encourage more looks to both images in the pre-naming phase, leading to cleaner data in terms of trial rejection, and more balanced looks towards target and distracter (as seen for example by shorter “longest look” measures). The adjunction of better controlled image background and saliency should also contribute to enhance the identification of objects, reducing the “back and forth” between target and distracter, especially during the post-naming phase. This would translate into longer “longest look” measures but also in a larger naming effect in correct trials. To fully estimate the impact of experimental stimuli manipulations on data quality, and therefore get a clear picture of the robustness of the method, we will provide different types of analyses. First, we will evaluate results for the classic 367–2000 ms post-naming time window and will examine effects of conditions on the different dependent variables used in the literature. We will also look at the impact of the different criteria used for trial rejection on the results. Then we will present a newer type of analysis that takes the time course of looking behaviour into consideration and thus avoids averaging over the whole post-naming window (von Holzen & Mani, 2012). While time course is seemingly the crucial aspect of the IPL paradigm or other visual paradigms, it has been rarely exploited in IPL experiments yet should provide with informative and complementary results.","In this experiment the classic mispronunciation IPL paradigm is used: two objects are presented side by side on a screen and one is named halfway through the trial. Pronunciation is correct for half the trials (e.g., “Look! Bed!”), and incorrect for the other half (“Look! Bud!”, for bed). We developed three versions. In the No Fixation condition, no central fixation stimulus was used and images were simply controlled for suitability and size but not for background, colour or perceptual saliency (they mostly had a naturalistic background). In the second, Fixation condition, a fixation stimulus was added before the start of every trial, and in the last condition, Fixation Plus, this was augmented by controlling the background colour of the pictures and their relative saliency.","Participants from the three conditions were matched for gender, age and vocabulary scores as measured by the Oxford Communicative Development Inventory (Hamilton, Plunkett, & Schafer, 2000). Twenty 20-month-olds were successfully tested for the Fixation condition (mean age 20;07; range 19;19 to 21;13; SD 3 days; 8 females and 12 males). Average OCDI scores were 241 words (SD = 19.1) in comprehension. One additional participant was excluded for fussiness. The population in the Fixation condition served as the baseline for the selection of participants in the two other conditions as they were matched on their OCDI scores in comprehension. Thus, 20 children out of an initial group of 39 children constituted the No Fixation condition (M = 19;16; range 18;20 to 20;29; SD 4 days; 8 females and 12 males). Their comprehension score was 234 words on average (SD = 17.7). Participants for the Fixation Plus condition (M = 19;27; range 18;14 to 21;5; SD = 5; 8 females and 12 males) were selected from the participants of Durrant et al. (2014, 32 toddlers), with an average comprehension score of 241 words (SD = 15.6). Stimuli The stimuli were 32 monosyllabic consonant initial words taken from the OCDI (see Appendix B), half targets and half distracters. All are imageable nouns. Target words were judged as known by at least 40% of 20-month-olds (database from the Oxford Babylab), and they were paired with a distracter sharing the same onset consonant and phonemic structure (e.g., dog and duck). Out of the 16 targets, participants heard 8 labels that were correctly pronounced and 8 that were mispronounced. Following Mani and Plunkett (2007), mispronunciations were obtained by changing one phoneme on one or more dimension, either on the onset consonant or the vowel (4 trials each), leading to a pseudoword or a very rare word in infant-directed speech (e.g., bud). The speech stimuli were recorded in an enthusiastic child friendly manner by a female native speaker of British English. Recordings were conducted in a sound-attenuated booth, digitized at a rate of 44.1 kHz and a resolution of 16 bits. The recorded tokens were matched so that the duration, amplitude and F0 of the correctly pronounced labels and their mispronunciations did not differ significantly. The tokens were then spliced onto a carrier sentence “Look! Target word!”, with onset of the test word starting 2500 ms into the trial. Note that the auditory stimuli are identical across conditions. The visual stimuli were photographs of the targets and distracters retrieved from the web and judged by the experimenters as good exemplars of the chosen categories. For the Fixation and Fixation Plus conditions, a smiley face was presented between trials at the centre of the screen until fixated by the participant, and followed by the next trial (triggered by the experimenter). For the Fixation Plus condition, image quality was manipulated: as exemplified in Appendix C, we systematically removed background colour and matched pictures (pairs of target and distracter) for visual saliency. This was checked with Pearson's correlation conducted on the vectorized image pairs (Chinga & Syverud, 2007). The saliency of target-distracter pairs presented to the No fixation and Fixation conditions was not well correlated (0.13 on average, for the absolute value of the Rs), whereas a reasonable correlation for the Fixation Plus condition was found (0.35 on average, which is a good correlation magnitude, Hemphill, 2003). Images were projected onto a screen 1.20 m away from the child, each of the stimuli image measured 52 cm diagonally and were separated by 43 cm, so that both images were comprised within 48° of visual angle with 10° of gap between them. The smiley face was centred and measured 14 cm diagonally.","After written consent was obtained, both the participant and the parent were invited into the room set up for the IPL. An image was presented on the screen in the dimly lit room so as to entice the child into sitting in a high chair. A short animated cartoon was played on the screen to keep her entertained and looking at the screen while the experimenter adjusted the cameras on her face. The parent sat behind the child and was asked not to intervene in any way so as not to influence the child. The experimenter could hear the stimuli being presented (so that she could also hear possible parental intervention) but could not see the screen. Each of the 16 trials (plus two for training) were started manually by the experimenter when the child was looking at the screen (anywhere for the No Fixation condition) or at the centre of the screen, that is, at the smiley face (for the Fixation and Fixation Plus condition). Participants were presented with 8 correct labels and 8 mispronounced labels. Order of presentation, pronunciation type and side of the target were counterbalanced across participants. Scoring ~~~~~~~ The digital scoring system developed by Meints and Woodford (2008) was used to synchronize videos with trial onsets. Eye movements (left image, right image, middle or away) frame by frame (40 ms) were scored by skilled coders trained by the first author and naïve to the items being presented to the participants. Each eye fixation was coded, and an independent skilled coder scored 10% of the pool of data randomly for each group. Agreement between coders was high with an intraclass correlation coefficient of 0.936 (Shrout & Fleiss, 1979). Looking times on each image were then automatically extracted for the pre- and post-naming phases, providing the proportions of looks to the target and the distracter in the pre- and post-naming phases, as well as longest look measures and frame by frame eye position for the time-course analysis. Time course plot and windows of analysis ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Following recommendations by Fernald et al. (2008), proportions of looks to the target as a function of time and pronunciation type were plotted in Fig. 1 for each condition (No Fixation, Fixation and Fixation Plus). Visual inspection of the plots suggests that the usual 367–2000 ms window post-naming seemed adequate, although it looks like the naming effect extended after 2000 ms post-naming for the No Fixation and the Fixation conditions. The final time course analysis (as in von Holzen & Mani, 2012) will be the best test to determine when pronunciation effects occur. It can be seen from the plots that at the end of the pre-naming phase, looks to the target were, on average, below the expected 50% (corresponding to no preference for targets or distracters) for the No fixation and the Fixation conditions, while the proportion was more balanced for the Fixation Plus condition. While Fernald et al. (2006) recommend an average of 50% of looks to the target in the pre-naming phase to avoid any bias, it is often normalized by subtracting pre- naming measures from post-naming measures (see, among others, Mani & Plunkett, 2007) or computing a salience score like in Swingley and Aslin (2007). We will return to this observation in the discussion. Selection of trials ~~~~~~~~~~~~~~~~~~~ First of all, we only analyzed trials where the target was known to the participant. As children were matched for vocabulary knowledge across the three conditions, we expected the proportions retained for each condition to be comparable: 95.0% for the No Fixation condition (304/320 trials), 87.8% for the Fixation condition (281/320), and 93.4% for the Fixation Plus condition (299/320). However, a chi-square test run on the raw scores revealed a significant difference between conditions (χ2(2) = 12.55, p = .002). A general linear model of the data with vocabulary (target known vs. unknown) and condition (No Fixation, Fixation, Fixation Plus) as factors confirms that more words were known in the No Fixation condition compared to the Fixation condition (z = 3.001, p = .003), but not between the No Fixation and the Fixation Plus conditions (z < 1, n.s.). This was very likely due to a sampling effect, since the Fixation group is the only one where two participants did not know 6/16 words, while in the other groups the maximum words that were unknown did not exceed 4/16. The χ2 test excluding these two participants shows that then the three groups no longer differed (χ2(2) < 1, n.s.). The second step in data selection often involves an inspection of looking time distribution trial per trial. The procedure presented by Fernald et al. (2008) includes a pre-screening of the recorded videos. Another way of retaining attended trials is to select trials where the child is looking at both pictures, either necessarily in the pre-naming phase (strict criterion), or at some point throughout the trial (lax criterion). The strict criterion is likely to result in cleaner data; however it could be that the lax criterion is preferable: training trials should be sufficient for the child to understand that pictures appear in the two corners of the screen. Consequently, upon hearing “Look! Dog!”, shifts to the target in the post-naming phase when the child did not look at the target picture in the pre-naming phase can still be considered as a sign of word recognition (re. mutual exclusivity in monolinguals, Houston-Price et al., 2010). In what follows, we present data based on the lax criterion, but we also provide in Appendix C.1-3 the analyses based on the strict criterion. With the lax criterion (looks at the target and distracter at some point throughout the trial), we retained 98.7% of the known trials in the No Fixation condition (300/304), 95.4% in the Fixation condition (268/281) and 97.0% in the Fixation Plus condition (290/299). A Pearson's chi-square test with Yates correction shows that the rejection rate did not differ across conditions (χ2(2) = 4.52, p = .10). The strict criterion (looks to both target and distracter during the pre-naming phase) retained 88.8% of the trials in the No Fixation (270/304), 90.4% of the Fixation trials (254/281) and 88.0% of the Fixation Plus trials (263/299). Again, a chi-square test revealed no main effect of condition (χ2(2) < 1, n.s.). Note that all children were included in the analysis as they contributed to all conditions (re. Fernald et al., 2006). Looking behaviour in the pre-naming phase ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Preliminary analysis of the pre-naming phase is meant to control for a potential bias towards the image that will be named later in the trial. Although averages were close to 50% during the whole pre-naming phase, they were significantly below chance for No Fixation (M = 44.4%, SD = 6.0%, t(19) = −4.15, p = .001) and Fixation (M = 42.9%, SD = 5.9%, t(19) = −5.34, p < .0001), showing a bias towards the distracter image, which was similar in both conditions (t(38) < 1, n.s.). In the Fixation Plus condition however this bias was no longer significant (M = 49.1%, SD = 4.5%, t(19) < 1, n.s.), and differed significantly from that in the No Fixation condition (t(38) = −2.79, p = .008) and the Fixation condition (t(38) = −3.72, p = .001). Dependent variables ~~~~~~~~~~~~~~~~~~~ IPL studies in the literature have presented a whole range of different dependent variables (see Table 1), and included pre-naming as a factor (Mani & Plunkett, 2007; Styles & Plunkett, 2008) or as a salience score (Swingley & Aslin, 2007; White & Morgan, 2008). In the first case, the two resulting independent variables are naming (pre- vs. post-naming) and pronunciation (correct vs. incorrect), and in the second case only pronunciation. Results from the pronunciation effect with the salience score are then identical to the post-hoc analysis of the naming × pronunciation interaction. As such, we will not present the salience score analysis. It is worth noticing that sometimes the pre- naming phase is not explicitly analyzed (Fernald et al., 2008; Ramon-Casas et al., 2009; Schafer & Plunkett, 1998; Swingley, 2003). The most widely used dependent variables are either the longest look measure (LLK) or the proportion of total looks towards the target (PTL). The PTL is usually calculated by dividing the total looking time to the target (T) by the total amount of looks to the target and distracter, that is T/(T + D), in the pre- and post-naming phases. A significant increase in the PTL in the post- compared to the pre-naming phase is taken as evidence of a naming effect. A significant decrease or absence of difference will show an absence of naming effect, that is, no evidence of word recognition. The LLK, on the other hand, represents the longest single fixation on the target (and also distracter). Successful word recognition, or the naming effect, should lead to an increase of LLK in the post-naming compared to the pre-naming phase. Descriptive statistics are presented for the analysis where children looked at both images at some point during the whole trial (lax criterion), with proportion of looks to the target (PTL, Fig. 2) and longest look (LLK, Table 1) measures for the target and distracter, in the pre- and post-naming phases. Statistics are reported for the overall LLK and PTL as dependent variables, with pronunciation (correct, incorrect) and naming (pre-, post-naming) as within-participant factors, and condition (No Fixation, Fixation, and Fixation Plus) as between-participant factor. We also report effects and interactions of pronunciation and naming for each condition separately, as each condition could potentially constitute a stand-alone experiment. Since the p values appear to be comparable for both measures (with the exception of the triple interaction, significant with LLK, and driven by the interaction between naming and pronunciation for the Fixation Plus condition), in the first following section we only report values for the PTL measure. Detailed descriptive statistics for the other measure (LLK) can be found in Table 1, and the corresponding ANOVAs are described in the following sections. Comparisons of measures will be discussed in the next section. Comparison between dependent measures So far, we ran a first series of analyses with condition, pronunciation and naming as factors, followed by a second series of analyses broken down by condition, with pronunciation and naming as factors. For the two dependent variables, we observed an overall agreement in the statistical tests, in particular for the naming variable, where the largest effect was expected overall. The more sensitive result, namely the interaction between pronunciation and naming, is clearly absent in the No Fixation and Fixation conditions, and robust in the Fixation Plus condition. A rerun of all these analyses on trials retained with the strict attention criterion (that is, trials where children look at both pictures in the pre-naming phase) produced a comparable pattern of results (see Appendix D), except for the main effect of pronunciation that became weaker (from marginal to non-significant for the No Fixation condition, and from significant to marginal for the Fixation Plus condition). Time course analysis Time course plots allowed us initially to visualize the time window where the naming effect was most likely taking place, ensuring that using the classic 367–2000 ms window would not miss out interesting data points. The following tests will provide us with a more precise measure as to when the two pronunciations elicit quantitatively different looking time behaviour. In order to analyze the time course of the proportions of look, we followed the methods advocated by von Holzen and Mani (2012), a non-parametric random permutation analysis (Maris & Oostenveld, 2007) to test effect of pronunciation across time, so as to identify the time period when looking times were significantly different. With this method however, and contrary to the classic PTL analysis including naming as a factor, pre-naming behaviour is not taken into consideration in relation to post-naming. This is why we considered another analysis of the time course, focusing on post- naming data corrected by the pre-naming data (PTL during the post-naming phase – average PTL in the whole pre-naming phase). In short, we ran two analyses, one similar to von Holzen and Mani (2012), and one where the average PTL to the target in the pre-naming phase (between 367 and 2000 ms) is subtracted from each data point. The latter will reveal any potential pronunciation effect that would have been masked by a pre-naming bias. The procedure, best described in Maris and Oostenveld (2007), identifies the time period where looking behaviour differs between the correct and incorrect pronunciation, that is, the naming effect. In the first step of the procedure individual paired sample t-tests were performed at each time sample, and used to identify significant (p < .05) t-values. In step- two, clusters were identified by finding significant t-values that were contiguous across time. For each such cluster, a cluster-level t-value was calculated as the sum of all single sample t-values within the cluster. Analysis thereafter was based on these clusters and their associated cluster level t-value, rather than the individual (and highly non-independent) t-values. Since cluster level t-values could not be tested for significance against a standard t distribution, in step three of the procedure, the significance of each cluster was calculated by comparing its cluster-level t-value to a Monte Carlo distribution of cluster level t-values generated from the cluster with the largest cluster-level t-value. To do this each of the original paired sample t-tests that were used to generate this cluster were repeated, but with the data items of each pair randomly assigned between the two conditions. This was performed 1000 times to generate a Monte Carlo distribution of 1000 summed t-values corresponding to the null hypothesis. The summed t-values of these randomized tests provided a null distribution against which the actual cluster-level t statistic of each of the observed cluster could be compared. Thus, for each observed cluster, a Monte Carlo p-value was calculated as the proportion of the null distribution which had a cluster-level t statistic that exceeded the actual cluster-level t-statistic. The first series of analyses were conducted on the whole trials (pre- and post-naming phases) and revealed no significant difference between correct and incorrect trials for the No Fixation condition (identified cluster: 900–1060 ms after target word onset; cluster t statistics = 12.13, Monte Carlo p = .50) and the Fixation condition (here, the two pronunciation conditions do not differ significantly from each other in any time window, therefore no Monte Carlo estimate was calculated) (Fig. 1a and b). For the Fixation Plus condition however (Fig. 1c), the two types of pronunciation differed significantly between 900 and 1580 ms post stimulus onset (cluster t statistics = 48.73, Monte Carlo p = .04), with an increase in looks at the target when it was correctly named. The second series of analyses was conducted on the post-naming phase only, after subtracting the overall PTL to the target in the pre-naming phase. This is a similar approach to the inclusion of the salience score by Swingley and Aslin (2007), where they analyze PTL in the post-naming phase, minus PTL in the pre-naming phase; instead this time it is applied on each time frame. Again, no significant difference was obtained in the No fixation condition (all t-tests reveal ps > .05). In the Fixation condition a cluster between 2420 and 2460 ms post word onset was identified but was, however, non significant (cluster t statistics = 4.33, Monte Carlo p = .96). For the Fixation Plus condition, we found a significant difference window which was comparable to the first series of analyses, with more looks to the correctly pronounced target from 740 to 1940 ms after stimulus onset (cluster t statistic = 105.34, Monte Carlo p = .002). The same analyses conducted on the dataset selected with the strict criterion reveal rather comparable results, with perhaps more sensitivity. On the whole duration of the trial, the No Fixation condition reveals difference cluster between 980 and 1700 ms after the word onset, but Monte Carlo simulations failed to confirm that this difference is significant (cluster t statistics = 12.76, p = .53). For the Fixation condition no cluster was identified (all t-tests reveal ps > .05). For the Fixation Plus condition, we replicate the significant effect of pronunciation, from 980 to 1700 ms after stimulus onset (cluster t statistics = 54.55, p = .03).","The goal of the present study was to evaluate how infants’ looking behaviour in the Inter- modal Preferential Looking paradigm varies as a function of noise generated by two simple methodological modifications, and how different methods of analysis best account for the resulting behavioural changes. Three groups of 20-month-olds were tested for recognition of correctly and incorrectly pronounced familiar words, in conditions that varied in terms of level of visual noise. In the first, No Fixation condition, no central fixation stimulus was used and images were simply equated on size and judged as good exemplars, with no attempt to control for background and salience. In the second and third conditions (Fixation and Fixation Plus), a central fixation stimulus appeared between trials to attract the child's gaze to the centre before the onset of the next trial (and not at the onset of the post-naming phase as in Portillo et al., 2007, cited in Fernald et al., 2008). In addition, in the third, Fixation Plus condition, children were presented with images without any background and target-distracter pairs were matched for visual salience. To analyze the resulting data we examined the impact of different trial selection criteria, compared two dependent measures (LLK and PTL) and performed a time course analysis using a combination of Monte Carlo estimate and cluster analysis, to identify the time period where mispronunciation affected behaviour. Following numerous studies using a similar paradigm (e.g., Mani & Plunkett, 2007; Swingley, 2003; Swingley & Aslin, 2000; White & Morgan, 2008), the expected result at 20 months was a naming effect restricted to the correctly pronounced targets, that is, an increase in looks to the target in the post-naming phase as compared to the pre-naming phase, only when the target word is pronounced correctly. We also expected that each methodological modification (addition of a central fixation and increased control of images) would contribute to enhance the quality of the data. The central point of the study was to determine which method of analysis would prove the most robust across methodological variations, and which would be the most sensitive. Results overall revealed that all three groups showed a main effect of naming, that is, children fixated the target image longer after it was named. However, only the Fixation Plus group behaved as predicted by the literature, they showed a naming effect restricted to words correctly pronounced, just like in Mani and Plunkett (2007) or Swingley and Aslin (2000). Our main interpretation of these results is that the combination of the central fixation point and the selection of better-controlled images contributed to enhance the quality of data, and to reduce experimental noise, in the pre- naming phase, which in turn resulted in less variable post-naming data, as will be discussed below. We suspect that other parameters could act to reduce similarly the level of unwanted noise, such as the use of the same items to act as targets and distracters (as recommended by Fernald et al., 2008, see Mani & Plunkett, 2007; Swingley & Aslin, 2000) or the selection of the most frequent words in a child's vocabulary (e.g., Swingley & Aslin, 2000). By progressively reducing the noise in the visual stimuli in the Fixation Plus condition, exploration of the visual stimuli during the pre-naming phase was more balanced with a PTL to the target around the expected 50%, despite a slight bias towards the distracter across all conditions. This bias must be due to reduced familiarity with the distracter objects (across all groups, participants were reported knowing the names of the target in 283 trials, and the name of the distracter in 246 trials), which thus worked as initial attracters. However, the pre-naming imbalance was not strictly comparable in the No Fixation (Fig. 1a) and the Fixation (Fig. 1b) conditions: whereas it was observed from the very onset of the pre-naming phase in the No Fixation condition, in the Fixation condition children looked equally long at targets and distracters from the onset of the pre-naming phase, and it is only after about 300 ms that the preference for distracters emerged. Given that the only methodological difference between the No Fixation and the Fixation conditions was the adjunction of a central fixation stimulus at the onset of each trial, it is quite likely that this central fixation point contributed to reducing the imbalance between target and distracter looks during the pre-naming phase. However, controlling for the quality of images had a cumulative positive effect, as seen in the Fixation Plus condition. Not only were pre-naming looks between targets and distracters more balanced, but the expected naming effect was obtained earlier, and was more robust than in the Fixation condition. What changed between the two conditions was a disappearance of the background and a quantitatively controlled balance in visual salience of the target-distracter pairs. Whether background control contributed more than saliency control to the sharpening of infants’ behaviour remains undetermined in this study. At this point, these results allow us to add to the recommendations of Fernald et al. (2008) and from the literature using the IPL paradigm: a fixation stimulus helps by centring the child's attention before trial onset, and carefully selected images even out the probability of looking at both images in the pre-naming phase, enhancing the sensitivity of the method. Crucially, our central aim was to compare how these different methodological choices would impact on the robustness of data through the lens of different analytical choices such as the criteria for data selection and the dependent variables. The literature shows that the criteria used for data inclusion or rejection in the pre-processing phase varies substantially across experimenters. A most reasonable practice – which we did not question – is to include only trials where the target word is known by the participant, as attested by parental questionnaire. More questionable is the practice of rejecting trials during which the child has not looked at both the distracter and the target at some point: does it have to be at some point during the entire trial (lax criterion), or during the pre-naming phase only (strict criterion)? We have shown that the two criteria, which measure the level of attentiveness during each trial, do not lead to fundamentally different results. Unsurprisingly more trials were rejected due to the application of the strict criterion (2.9% for the lax criterion vs. 11.0% for the strict criterion), resulting in a loss of experimental power. However, a close inspection of results in the PTL section in the Results section and Appendix D.2 shows that the size of the main effect of condition is larger when the strict criterion is applied. In contrast, applying the lax criterion results in larger sizes for all other effects, including pronunciation and naming as well as the crucial naming × pronunciation interaction. This suggests that the overall behavioural adjustments due to methodological changes may be enhanced with the use of the strict criterion, but not the quality of the key effects (naming modulated by pronunciation). With the strict criterion, we ensure that children have seen the target and the distracter during the pre-naming phase. Upon hearing the label they would know that a mispronounced name does not correspond to any picture; they can then use a ‘better match’ strategy based on phonological overlap and look slightly longer at the target. This translates into a relative decrease in the size of the pronunciation effect as compared to the same data analyzed with the lax criterion. In contrast, with the lax criterion, we also include trials where the child has only checked the target during pre-naming,1 and therefore, can reasonably assume that the mispronounciation can refer to the unchecked item. This results in a slightly higher number of looks towards the distracter during the post-naming phase. To sum up, data may suggest that we do not measure the same behaviour or strategy if we apply the strict or the lax criterion: for the former, we may measure a better-fit strategy based on the degree of phonological overlap, whereas in the latter, children may produce a response based on a Mutual Exclusivity-type principle (Halberda, 2003). If this speculative assumption was corroborated by further research, this should be kept in mind when deciding for one criterion over the other, depending on the theoretical goals of the experiment. We have seen that it is common practice to exclude trials in which children do not know the name for the target object; is it justified to also exclude trials where the child does not know the distracter (as done by, among others, Swingley et al., 1999)? On the one hand, this could have some advantage: the children would be more likely to look away from an incorrectly named target image and attach the mispronounced label to the distracter (White & Morgan, 2008), strengthening the mispronunciation detection effect. On the other hand, unknown distracters could result in children showing a familiarity effect rather than a naming effect (looking at the named target simply because they have a name for it, not because it has been named with its specific label). Quite pragmatically, excluding trials in which the child does not know the distracter would possibly result in a loss of experimental power. This was indeed the case for the No Fixation and Fixation conditions, but not in the more robust condition where the critical interaction was replicated. Therefore it appears that the application of this criterion does not substantially modify the quality of the data, at least not in the current study. It should be kept in mind however that, similarly to what was discussed above for the use of the lax vs. strict criteria, knowing, or not knowing, the distracter label may modify the strategy that the child uses in the procedure. An unknown distracter promotes the use of the Mutual Exclusivity principle whereas a known distracter encourages the use of a best-match strategy based on the degree of phonological overlap. Regarding the choice of the dependent variable, the literature often presents side by side analyses based on the proportion of looks to the target (PTL) and on the longest look to the target (LLK), as they usually show similar results. The same conclusion can be applied here, although with a caveat. A close inspection of statistics in the result section shows that in most analyses, effect sizes are larger for the LLK measures than for the PTL ones. This could be due to the fact that in the vast majority of cases, the longest look is also the first look towards the target, and during that period which lasts about 700 ms (see Table 1), the child computes all the information that is needed to correctly identify the target. Possibly all further looks towards the target are either verification or random noise, which is incorporated in the PTL measure but not in the LLK variable, resulting in less variable data in the latter than in the former measure. Finally, we questioned the importance of adjusting the time window of analysis, and investigated the relevance of a time course analysis. Depending on the age and/or vocabulary size of the participants, it is common practice to adjust the onset of the post-naming phase to account for variation in gaze shift latency (Fernald et al., 2008). In addition, many factors can influence the processing time of the target and distracter pictures, starting with the nature and complexity of the auditory stimulus, the visual properties of the stimuli (e.g., Arias- Trejo & Plunkett, 2010), or the type of distracter (familiar vs. unfamiliar; White & Morgan, 2008). Therefore a time window fixed a priori may not be the most accurate. Of course, selecting for each experiment a time window based on the visual inspection of the data would be unacceptable as it would lead to a strong human bias. One way around this is to generalize the use of the time course analysis as reported here which allows the identification of time windows where the naming or pronunciation effects are indeed significant. It seems that the statistical analysis of the time course provided an accurate estimate of looking behaviour, since it allowed us to distinguish between a very short lived pronunciation effect (in the No Fixation condition) and an enduring one (with the Fixation Plus condition). This mirrored the outcome of the classic mean-based looking times analyses, namely a robust interaction pronunciation × naming in the Fixation Plus condition and none in the No Fixation condition. While very promising to estimate the speed of word recognition (like Durrant et al., 2014; Fernald et al., 2006), this approach needs to evolve to establish the minimal temporal window where pronunciation differences are meaningful (re. the very short lived pronunciation effect in the No Fixation condition). In summary, observations based on infants’ behaviour in a classic IPL task can vary quite substantially depending on the methodological parameters chosen during the pre- processing period or data analysis. Perhaps the vulnerability of the data are best illustrated in Fig. 2c which displays the results of the Fixation Plus condition. Correctly and incorrectly pronounced words produced different looking times for about 700 ms, as revealed by the time course analysis. This is a rather short window of interest as compared to the entire duration of the trials, for example as compared to head-turn procedures which typically generate differences of about 2 s of looking times between conditions (e.g., Mattys, Jusczyk, Luce, & Morgan, 1999; Jusczyk & Aslin, 1995). It can be argued that IPL is a more direct and precise measure of auditory processing than head turn paradigms as it does not rely on an experimenter's intervention during the session (whereas head turn set-ups usually do: Floccia et al., 2012; Nazzi, Jusczyk, & Johnson, 2000; Schmale & Seidl, 2009). Yet this augmented precision perhaps makes the IPL tool more prone to vary with methodological noise. To borrow an example from physical instruments, a digital thermometer might be more precise than a mercury one, yet thanks to its inertia the latter is more likely to give robust repeated measures than the former. It is our hope that this methodological study will contribute to sharpen the use of this invaluable paradigm in the quest of infants’ representation and processing of visual and auditory information."],["Researchers have questioned whether there is a relationship between personality and patterns of online self-presentation. This paper examined, more specifically, whether personality predicts profile choices as well as image choice behaviour on two different SNSs: Twitter and Facebook. We found that personality does, to some extent, predict choices regarding profile images; however, not always in the direction we predicted and results differed across sites. We found that participants who scored higher on conscientiousness and lower on extraversion were more likely to change their Facebook profile image. Participants who scored lower on extraversion were more likely to choose a Twitter profile image that included a photograph of themselves compared to participants who scored higher on extraversion. For participants whose Facebook profile image was a photograph of themselves, a greater proportion of participants selected a recent photograph from the past six months. However, this was not the case for Twitter. We conclude that personality can predict some image choices and behaviours that might be useful for future work on authentication and identification, although other predictor variables are potentially also important when considering the types of individual characteristics which might predict online behaviour on SNSs. --------------------------------------------------------------------------------","Social networking sites (SNSs) have become a popular medium for communication and networking for individuals of all ages (Nadkami & Hofmann, 2012; Subrahmanyam, Reich, Waechter, & Espinoza, 2008; Valkenburg & Peter, 2009). These sites can potentially provide valuable information about a person that could assist in authentication and identification. Moreover, they might provide the user with rich information (e.g., about potential dates or employees). In contrast, of course, users can perform different versions of the ‘self’ on these platforms (e.g., Turkle, 1995) or, as others would contend, express their ‘true selves’ in this space (e.g., Marriott & Buchanan, 2014; McKenna & Bargh, 2000). Although researchers are beginning to learn more about what online personal data presented on SNSs might tell us about a person, there is scope to learn much more about digital identities and what they reveal about the person behind the profile. This study focused, in particular, on whether personality predicts profile choices as well as image choice behaviour on two different SNSs: Twitter and Facebook. Drawing from Goffman's (1959/1997) work, it has been theorised that SNSs provide an ideal environment for impression management. SNS users can be very selective in the information they choose to present on these sites, including, for example, their interests, activities, opinions and emotions. It has been argued that some users consciously select information to present in order to convey a certain impression (Wu, Change, & Yuan, 2015). Equally, however, some details might leak additional personal information about a person, unintended and sometimes unbeknownst to the user; for example, ethnicity, education or class (Whitty & Young, 2017). Individuals who score higher on certain types of personality characteristics might choose to represent themselves in distinct ways (e.g., Back et al., 2010; Krämer & Winter, 2008). Extraverts, for example, are: more likely to select self-representative photographs (Wu et al., 2015), less likely to post conservative pictures of themselves, have a greater number of online friends, and are more likely to use the communication function on SNSs (Krämer & Winter, 2008; Wang, Jackson, Zhang, & Su, 2012). Neurotic individuals are more likely to use the status update feature and agreeable individuals are more likely to write comments on others' profiles (Wang et al., 2012). Women low in agreeableness are more likely to use the instant messaging features on social networking sites, while men low in openness play more games via social networking sites (Muscanell & Muscanell & Guadagno, 2012). Individuals who score high on narcissism are more likely to select pictures that are more physically attractive (Kapidzic, 2013; Wu et al., 2015). Images and photographs, in and of themselves, might be used intentionally to convey a particular identity, including personality features (e.g., Wu et al., 2015). Kapidzic and Herring (2015), for example, found that females are more likely than males to select a seductive photograph as their profile picture. Zheng, Yuan, Chang, and Wu (2016) found that women are more likely to emphasize emotional expression in their profile pictures compared with men, whilst men were more likely to emphasize having fun. Users, however, might also unintentionally leak aspects about a person; such as age, ethnicity, hobbies, relationship status (Lee-Won, Shim, Joo, & Park, 2014). For example and perhaps unsurprisingly, it has been found that individuals who upload dyadic profile pictures on Facebook report feeling more satisfied with their relationship and closer to their romantic partners compared with individuals who do not (Saslow, Muise, Impett, & Dublin, 2012). Tifferet and Vilnai-Yavetz (2014) have found that men were more likely to have profile pictures which accentuate status and risk taking, while women were more likely to have photographs which including familial relations and showed emotional expression. The research presented in this paper builds upon this previous literature by examining whether personality predicted users' profile image choices on SNSs. We focused on two different types of SNSs in our study, which serve different social purposes — Facebook and Twitter. Facebook is more privately oriented with a focus on maintaining connections with existing friendship groups and communicating with them via multiple methods (Raacke & Bonds-Raacke, 2008). Conversely, Twitter is more publically oriented and users' friends are more likely to be unknown compared with Facebook friends (Hughes, Rowe, Batey, & Lee, 2012; Marwick, 2011). We investigated whether personality predicts: how often users change their profile image; if individuals choose to use an avatar (an icon or figure that represents that person, but is not a photograph of that person) or a photograph of themselves; if users include a recent photograph of themselves; and if users select a profile image they believe represents their personality. Hypotheses ~~~~~~~~~~ We employed the Five Factor model (Costa & McCrae, 1992; Goldberg, 1990), which is commonly employed to measure personality. We hypothesised that those who score high on extraversion (H1) and conscientiousness (H2) will be more likely to change their images more frequently compared with those who score low on these two scales. Our first hypothesis is based on the assumption that because extraverts are outgoing and social they will be more likely to want to update how they present themselves. Our second hypothesis is based on the assumption that those who score high on conscientiousness will be more motivated to keep their profile up-to-date. It was further hypothesised that individuals who score low on openness to experience (H3) and high on extraversion (H4) will be more likely choose a photograph than an avatar. These hypotheses are based on the assumption that people who score higher on openness tend to be more creative and open to new ideas and therefore an avatar might be perceived as a more creative depiction of themselves, whilst, introverts might choose an avatar in preference to a photograph to divert attention away from themselves. The fifth hypothesis (H5) is that those who score high on conscientiousness will be more likely to include a recent photograph of themselves. Again, this is based on the assumption that conscientious people might be more likely to keep their profile up-to-date. The sixth and final hypothesis (H6) is that individuals who score low on openness to experience will choose an image that more closely represents their self-concept. This is based on the assumption that because individuals low on openness to experience are far less likely to be creative and experimental they will chose an image that more closely represents their self-concept.","One thousand, two hundred and twenty four individuals were recruited from a ‘Qualtrics’ online panel. Of these individuals, 357 individuals reached the end of the study (29% completion rate), in the main due to screener questions, which eliminated participants due to not having a Twitter and a Facebook account (43%). This meant the drop out rate, due to incompletion of the survey was only fairly small (28%). Another 150 individuals were then excluded for no longer using their Facebook or Twitter account (136), for withdrawing their consent for the study (8), for repetitive or inappropriate responding (4), or for incomplete responses (2). After exclusions, the final sample for this study consisted of 207 participants (115 male; 92 female). The mean age of the sample was 40.1 years (SD = 13.1) ranging from 19 to 77 years. For men, the mean age was 41.0 years (SD = 12.8) ranging from 19 to 68 years. For women, the mean age was 39.0 years (SD = 13.3) ranging from 20 to 77 years. All participants had both an active Facebook and Twitter account.","Data were collected using a questionnaire hosted on the Qualtrics online survey platform. Big 5 Personality traits were measured using a Five-Factor Personality Inventory validated for use online (Buchanan, Johnson, & Goldberg, 2005). This 41-item inventory gives measures of Openness to Experience, Conscientiousness, Extraversion, Agreeableness, and Neuroticism. The subscales have high reliability and have been successfully used in a range of Internet-mediated studies. In our study, Cronbach's α ranged from 0.69 to 0.84. Self-rating items Participants were asked several questions about their Facebook and Twitter activities, including: frequency of use (on a 7-point scale from ‘once a year or less’ to ‘more than once a day’); how regularly they changed their profile image (on a 7-point scale from ‘once a year or less’ to ‘more than once a day’); whether the profile image was an avatar or a photograph of themselves (with three categories to choose from: ‘yes, the image includes a recent photograph of me from the last six months’; ‘yes, the image contains a photograph of me, although is more than six months old’; ‘no, the image does not contain a photograph of me’); and how closely their profile image represented their personality with three categories to choose from (‘not at all representative’; ‘somewhat representative’; ‘very representative’) (see Table 1 for correlations between these items). We appreciate that the question that asks participants how closely their profile image represented their personality might not be interpreted as personality representativeness per se. However, by this question we were more interested in the participants' perceptions of self-concept. We opted to use this phrasing given our piloting of questions suggested that participants felt more comfortable with the term personality, as this term was more commonly used in ‘everyday speech’. Furthermore, we believed it was important to ask this question to all users (i.e., those who included an avatar or a photograph). This is based on previous research that has examined representation of self via the use of photographs (e.g., Fernandez, Stosic, & Terrier, 2017; Leikas, Verkasalo, & Lönnqvist, 2013) and avatars (e.g., Dunn & Guadagno, 2012). Procedure We commissioned Qualtrics to recruit participants from their online panel. Participants were required to reside in the UK and have both a Facebook and a Twitter account. Participants were presented with information about the study and asked to indicate informed consent before proceeding. On the subsequent pages they were asked to complete a number of demographic items; self-rating items (described above), and complete the Five- Factor Personality Inventory.","Our first two hypotheses were concerned with the frequency participants updated their profile imagine on Facebook and Twitter. To test our hypotheses a forced-entry multiple ordinal regression was calculated with the frequency of avatar change as the outcome variable and the Big Five personality traits as the predictor variables. Ordinal regression was considered appropriate given that the frequency of avatar change measure was ordered-categorical in nature. Separate regression models were calculated for Facebook and Twitter (frequency of participants' updates is shown in Table 2). The frequency with which participants used an online service was also controlled for in the regressions because participants who more frequently used a service would have more opportunity to change their profile images compared to those participants that used a service less frequently. In line with the second hypothesis, participants who scored higher on conscientiousness more frequently updated their Facebook profile compared with those who scored lower on conscientiousness. However, although we obtained a significant finding for our first hypothesis, it was in the opposite direction to what we predicted. Participants who scored lower on extraversion were more likely to change their profile image compared with those who scored high on extraversion. None of the Big Five personality measures predicted how often participants changed their Twitter profile image (see Table 3). The third and fourth hypotheses predicted that those low on openness to experience as well as extraversion will be more likely to select a photograph of themselves in preference to an avatar. A Chi-square test indicated that on Facebook significantly more individuals choose a photograph to represent themselves (n = 134; 64.7%) than an avatar; n = 73; 35.3%), χ2(1) = 17.98, p < 0.001. However, there was no statistically significant difference in the proportion of individuals who chose a Twitter profile image that included a photograph of themselves (n = 108; 52.2%) compared with an avatar (n = 99; 47.8%), χ2(1) = 0.39, p = 0.532. A forced-entry binary logistic regression was used to explore the relationship between personality and whether a profile image was a photograph of themselves or an avatar. Separate models were calculated for Facebook and Twitter (see Table 4). The logistic regression model for Facebook profile images was not statistically significant (χ2(5) = 5.63, p = 0.344) and explained a very small proportion of the variance (Nagelkerke R2 = 0.037). Consequently, the third and fourth hypotheses were rejected for the Facebook profile images. In contrast, a logistic regression model that predicted the content of Twitter profile images was statistically significant (χ2 (5) = 11.84, p = 0.037), provided a good fit to the data (Hosmer & Lemeshow χ2(8) = 4.59, p = 0.801) and explained a small proportion of the variance (Nagelkerke R2 = 0.074). Contrary to the direction predicted in our fourth hypothesis, participants who scored lower on extraversion were more likely to choose a Twitter profile image that included a photograph of themselves compared to participants who scored higher on extraversion (B = − 0.07, S.E. = 0.03, Wald = 7.38, p = 0.007; ExpB = 0.93) (H4). However, our third hypothesis was not supported. For participants whose Facebook profile image was a photograph of themselves, a greater proportion of participants selected a recent photograph from the past six months (n = 82; 61.2%), in preference to an older photograph that was taken more than six months ago (n = 52; 38.8%), χ2(1) = 6.72, p = 0.010. For participants whose Twitter profile image was a photograph, there was no statistically significant difference in the proportion of participants who used a photograph of themselves from the past six months (n = 51; 47.2%) and those who used an older photograph (n = 57; 52.8%), χ2(1) = 0.33, p = 0.564. The relationship between personality and whether the profile contained a recent photograph of the profile owner was tested using a forced entry binary logistic regression. Separate models were used for Facebook and Twitter profile images (see Table 5). Only participants who indicated that their profile image was a photograph of themselves were included in the model. The frequency with which participants used an online service was controlled for because participants who more frequently used a service have more opportunity to change their profile image compared to those who use a service less frequently. The logistic regression model for Facebook profile images was statistically significant (χ2(6) = 27.63, p < 0.001), provided a good fit to the data (Hosmer & Lemeshow χ2(8) = 10.46, p = 0.880), correctly predicted a good number of cases (71.6%), and explained a moderate proportion of the variance (Nagelkerke R2 = 0.234). The model for Twitter profile images was also statistically significant (χ2(6) = 23.92, p < 0.001), provided a good fit to the data (Hosmer & Lemeshow χ2(8) = 7.07, p = 0.529), correctly predicted a moderate to good number of cases (68.5%), and explained a moderate proportion of the variance (Nagelkerke R2 = 0.265). However, none of the Big Five personality traits predicted whether participants selected a recent or an older photograph of themselves for either their Facebook or their Twitter profile images (Table 5). Consequently, our fifth hypothesis was rejected for both the Facebook and Twitter profile images. The sixth and final hypothesis was that individuals who score low on openness to experience would choose an image that is closer to their self-concept (see Table 6). A forced-entry ordinal regression was used to explore the relationship between personality and the extent that participants selected a profile image representative of their self-concept. The ordinal regression model for Facebook profile images failed to reach statistical significance (χ2(5) = 11.36, p = 0.05, Nagelkerke R2 = 0.067; goodness of fit χ2(371) = 354.98, p = 0.716). The ordinal regression for Twitter profile images also failed to reach statistical significance (χ2(5) = 11.27, p = 0.05, Nagelkerke R2 = 0.071; goodness of fit χ2(371) = 343.86, p = 0.358). Consequently, our sixth hypothesis was rejected.","In this paper we were interested in whether personality predicts profile image choices and behaviours carried out two SNSs: Facebook and Twitter. In line with previous research, we found personality does, to some extent, predict users' choice of profile images as well as online behaviours (e.g., Back et al., 2010; Dunn & Guadagno, 2012; Hughes et al., 2012; Kapidzic, 2013; Kapidzic & Herring, 2015; Lee-Won et al., 2014; Leikas et al., 2013; Saslow et al., 2012; Wang et al., 2012; Wu et al., 2015; Zheng et al., 2016), but not always in the direction we hypothesised. Given that the behaviours and the image choices we examined are not immediately obvious indicators of personality, the work here suggests that users leak information about themselves online, without necessarily intending to elucidate particular personality characteristics. These findings could be applied in work on authentication, identification or pre-screening (e.g., by cybersecurity professionals, employers) or by organisations that wish to be more selective in the types of information they target individuals (e.g., organisations wishing to carry out targeted advertising on SNSs). On the other hand, these findings also suggest that users might want to be wary of how they behave and what information they present about themselves to avoid personality characteristics being leaked to others. Our second hypothesis, that people who are more conscientious are more likely to update their profile images compared with those who are less conscientious, was the only hypothesis that was supported by our data (although we note that significant findings were obtained for many of the other hypothesises, but in the opposite direction to what was predicted). This result might be particularly interesting to employers who are looking to employ conscientious workers. Given that employers are increasingly drawing from online sources to assist in employment decisions (Davison, Maraist, Hamilton, & Bing, 2012), this finding suggests there might be some utility in carrying out further research to examine other online indicators of conscientiousness. Our research adds to the research that has found that extraverts behave on SNSs in distinct ways (e.g., Krämer & Winter, 2008; Wang et al., 2012; Wu et al., 2015); however, our findings were not in the direction we predicted. In our study, introverts were more likely than extraverts: to update their profile images (Facebook only) and use a photograph in preference to an avatar to represent themselves (Twitter only). Perhaps our findings can be explained by previous work which suggests that introverts feel protected and safe using the Internet to socially interact with others, often preferring to communicate with others online (e.g., Amichai-Hamburger & Ben-Artzi, 2000). The introverts in our study might be focusing more of their energy on online communications and relationships compared to offline, thereby spending more time updating their profile images and feeling safer to display their ‘real’ image online. In addition to our significant results, we found that hypothesis three and five were not significant for either SNSs. Hypothesis three predicted that individuals who score low on openness to experience will be more likely to select a photograph than an avatar. Perhaps this null result might be explained by recent research on SNSs. Wu et al. (2015), for example, found that profile pictures were believed by their users to reflect their personality and that personality could be reflected in both photographs and avatars. According to their findings, users believed that a variety of impressions could be made using their profile pictures, such as: socialising, playing sport, romantic, family, etc. The distinction, therefore, between photograph and avatar, as we have made in this paper, could be too limiting and future researchers might wish to categorise images when examining the type of profile images users with different personality choose to use to represent themselves. Hypothesis five predicted that users who score high on conscientiousness will be more likely to include a recent photograph of themselves. The null result for this hypothesis might be explained in a similar way. In addition, given that those who score high on conscientiousness were more likely to update their profile image, as we predicted, the issue of using an updated photograph might not be as a concern. Moreover, these findings suggests that other predictor variables, in addition to personality, should be considered when examining how individuals characteristically select profile images and image choice behaviour on SNSs. Previous studies on predictability of behaviour on SNS, for example, have examined Narcissism, cultural differences and the true self (see Whitty & Young, 2017, for a more in-depth discussion). Our sixth and final hypothesis that individuals who score low on openness to experience will choose an image that more closely represents their self-concept, was not significant for Facebook or Twitter. In fact, none of the personality traits predicted whether a participant felt their chosen image more closely represented their self-concept. Perhaps other personality traits are more important to consider when examining the likelihood of chosen an image that represents one's self- concept (e.g., self-monitoring). As a further point, our findings were not always consistent across SNSs. This is an important finding given that many previous studies have focused on one SNS (typically Facebook), assuming findings can be generalised to other SNSs. This study found that the same participants behaved somewhat differently across sites, suggesting that the platform itself plays a fairly substantial role. For example, individuals were far less likely to change their Twitter profile than their Facebook profile image (suggesting these platforms serve a different function). While this finding is perhaps not surprising, they do highlight the importance that future research into detecting personality online ought to pay attention to platform itself and not just the behaviours per se. Although our study examined more than one SNS, this narrow focus might be perceived as a limitation and we would encourage future studies to focus on a wider selection of SNSs. Our findings, at least, suggest that this is a worthwhile exercise.","In conclusion, there is a growing interest in academia as well as in lay people's practices to predict psychological characteristics of a person by examining their online personal data. Our study suggests that our online behaviours can, perhaps unknowingly, leak aspects about our personality. Our study provides some evidence that personality predicts the types of profile photographs/images individuals select as well as the likelihood of updating profile images. Future research might consider image choices in more detail and across other types of sites."],["Objectives: The overall purpose of this study was to examine Canadian Provincial Sport Organization representatives' research priorities in order to provide directions for future research and knowledge translation initiatives in youth sport. Design: Qualitative description methodology. Method: Interviews were conducted with 60 representatives of Canadian PSOs from five provinces. Analysis followed the process of data condensation, data display, and drawing conclusions. Results: The most frequently reported research priorities were athlete development systems, participation and retention, parenting, benefits of sport, and coaching. Conclusions: Research that addresses the priorities of stakeholders may increase the adoption of research evidence in practice, program, and policy contexts. This study provides directions that may help inform future research agendas in youth sport. Conversely, the study also revealed that extensive previous research exists in several of the priority areas identified. These findings suggest that knowledge translation research is required in these areas. Knowledge tailoring may be a useful strategy for improving knowledge translation in these areas. Ultimately, it appears that efforts to increase connections between researchers and other stakeholders in youth sport are important. --------------------------------------------------------------------------------","Qualitative description methodology was used in this study, which is suited to obtaining “answers to questions of special relevance to practitioners and policy makers” (Sandelowski, 2000, p. 337). Consistent with a pragmatic philosophy, qualitative description is a tradition of research that can be used to describe participants' perceptions of phenomena and is suitable for problem identification (Neergaard, Olesen, Andersen, & Sondergaard, 2009). Qualitative description focuses on providing ‘straight’ descriptions of events or perceptions with low inference interpretation (Sandelowski, 2000; 2010). Given the purpose of the current study, qualitative description was an appropriate methodological selection because we wanted to depict participants' views about their priorities for future research. Participant sampling and recruitment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In order to purposefully sample participants (Milne & Oberle, 2005; Sandelowski, 2000; 2010), we approached PSOs in five Canadian provinces via an e-mail explaining the purpose of the study. We explained that we wanted to speak with a PSO representative who was primarily responsible for youth sport policies and programming, emphasizing that the focus of our project was to examine issues relating to research priorities and the use of research evidence. We estimated, a priori, that approximately 10 participants per province would enable us to attain adequate data saturation (with the exception of Ontario, which is the province with the largest population, and as such we sought 20 participants from two locations in Ontario). A total of 60 PSO representatives (39 male, 21 female, M age = 43.5 years, SD = 12.6 years) participated in this study. Twelve participants were from Alberta, nine from Manitoba, 20 from Ontario, 10 from Québec, and nine from Prince Edward Island. Twenty-one participants were executive directors, 12 were presidents/CEOs, and the remainder were technical directors or held similar roles (e.g., youth program manager, sport program manager). Six participants held a high school diploma, 44 held undergraduate degrees, and 10 held graduate degrees as their highest level of education. They represented a range of team and individual sports, including soccer, football, lacrosse, volleyball, ice hockey, curling, athletics, tennis, gymnastics, and golf. They had, on average, 8.8 years of experience working for the PSO (SD = 8.5 years, range 1–30 years). Interviews ~~~~~~~~~~ Consistent with qualitative description (e.g., Neergaard et al., 2009), individual semi- structured interviews with open-ended questions were conducted. An interview guide was developed following guidelines proposed by Rubin and Rubin (2012) and piloted via an interview with a former executive from a PSO, which led to very minor modifications. The interview guide commenced with a preamble explaining the study and some demographic questions. The main questions included: What types of research would be most helpful for you/your organization? What types of research would you like to see people at universities doing? What types of studies in youth sport would be particularly useful for you/your organization? What are the opportunities you see in using research? Participants were also asked some questions about the use of research evidence, data from which were not used in the current study. Finally, in a summary section, participants were asked: If you had total control over research, what project would you like to see us complete? Trained interviewers (n = 5) completed the interviews in each province. The interviewers attended a three-day qualitative methods workshop conference that included sessions on principles of qualitative interviewing and training specific to the current study. When the five interviewers returned to their home provinces to undertake data collection, they remained in regular contact with each other via e-mail to discuss and debrief after interviews. Interviews in Québec were conducted in French. Other interviews were conducted in English. Interviews lasted, on average, 45 min.","A professional transcription company transcribed the audio files of the interviews conducted in English. The researcher who conducted the interviews in Québec transcribed those interviews in French and translated them into English. Participants were assigned a number and any identifying information (e.g., names of sport, names of individuals) was removed from the transcript. The principal investigator led data analysis with the assistance of one of the interviewers. Data from one province (Alberta) were analyzed first, and the process was repeated for data collected from each of the other provinces. Analysis followed the process of data condensation, data display, and drawing conclusions (Miles & Huberman, 1994; Miles, Huberman, & Saldaña, 2014). The first step, data condensation, is the process of selecting and simplifying data contained in the transcripts. This involved the principal investigator reading and re-reading transcripts and coding meaningful segments of information (meaning units) and assigning labels (codes). A ‘long list’ of codes was initially produced and rules of inclusion (i.e., descriptions of the meaning of the code and the data coded therein) were written for each code. Then, pattern coding was completed, which is a process of grouping codes into a smaller set of themes. The long list of codes was condensed into themes by comparing and combining data that reflected similar ideas (e.g., codes reflecting different ideas around the notion of athlete development were grouped into the theme of ‘athlete development systems’). Data coded into each theme were checked and re-checked by the principal investigator and one of the interviewers. The themes depict common patterns shared across the ‘cases.’ Data displays are visually formatted representations that systematically portray information and, according to Miles et al. (2014), “are a major avenue to robust qualitative analysis” (p. 13). A content analysis summary table was created (Table 1) depicting specific examples of research topics (i.e., the sub-themes) that were combined to form each theme. A summed indices data matrix was also created to illustrate participants' responses in each theme (Table 2). This matrix also provided an indication of the extent to which participants’ perspectives were shared and facilitated descriptive conclusion drawing about the patterns in the data (Miles et al., 2014). The numerical reduction and display of data (i.e., Table 2) is an acceptable approach when using qualitative description, but numerical summaries should be considered as a supplement to the textual results of the content analysis (Neergaard et al., 2009; Sandelowski, 2010). Hence, the final step involved creating a written description of themes. Methodological rigor ~~~~~~~~~~~~~~~~~~~~ Consistent with Neergaard et al.’s (2009) suggestions, our approach to methodological rigor for this qualitative description study followed guidelines provided by Milne and Oberle (2005), which focus on authenticity and credibility, along with criticality and integrity. Authenticity and credibility require purposefully sampling an adequate number of participants and conducting an analysis that is driven by their responses rather than pre-determined ideas. It is about ensuring participants' ‘voices’ are reported with relatively low inference (an emic perspective), probing for clarification and depth during interviews, accurate transcription, inductive analysis, and demonstrating an understanding of context. The extensive interviewer training also added to authenticity and credibility. Criticality requires the critical analysis of research decisions, which were discussed during regular meetings. Integrity involves reflecting on the role and influence of the researchers, which was addressed via regular debriefing among the interviewers and the principal investigator maintaining a journal of the research decisions. For instance, one factor that involved extensive reflection was the principal investigator's previous work in youth sport and the fact that he has conducted several studies on some of the topics identified by participants. It was important his research not unduly influence the focus on providing the participants' accounts of their priorities for future research (i.e., the emic perspective).","Table 1 provides a summary of the analysis depicting specific examples of research topics (i.e., the sub-themes) that were condensed to form each theme. These sub-themes are identified in italics in the following sections, where we describe the main themes reported by the majority of participants and discuss the pertinent findings in relation to previous research. Table 2 depicts the prevalence of all the themes identified across the cases. Athlete development systems ~~~~~~~~~~~~~~~~~~~~~~~~~~~ A total of 49 of 60 participants (81.7%) reported preferences for more research examining issues related to the athlete development systems for their sport. Examples coded in this theme referred to a desire to understand more about the effectiveness of the long-term athlete development (LTAD) models. Some participants specifically questioned the research evidence-base for the Canadian LTAD model. For example, Alberta PSO#7 (AB_PSO7) argued that “we have our LTAD models but I think a lot of it isn't really based on research, it's based on experience, people that are [sport] experts, but based more on their experience. I think we need more research.” Similarly, Ontario PSO (site A) #5 (ON_A_PSO5) said: OK, I'm not a big fan of LTADs. I don't believe that there is any documented evidence for a lot of what we say in LTADs … Where [is there evidence] that says that you're actually gonna create an Olympian or longevity [of participation]? … I don't know if there are any [research] groups that their sole purpose is to look at LTAD and evaluate it. I don't know if Sport Canada has a team that works with LTAD and evaluates it … Other participants, while generally having a favorable view of LTAD, also highlighted the need to evaluate its effectiveness. As ON_A_PSO2 said, “I think it's maybe a little too early to know whether or not it's successful but I guess it's good to have a good model that everybody can integrate with.” Manitoba PSO#4 (MB_PSO4) highlighted a need for studies examining the long-term impact of LTAD on athletes and coaches, explaining: It's been 10 or 12 years since LTAD became a buzzword and it would be interesting to see if that program in and of itself is having any impact on athlete and coach development … I think looking at the impact of LTAD from a longitudinal perspective would be useful. And I think it would be interesting to see if the changes to programs have yielded the results that we anticipated or if it's just an approach that doesn't really result in changes in the long term. There is, of course, a large body of research on the psychological, social, and physiological aspects of talent development in sport (see Gledhill, Harwood, & Forsdyke, 2017; Rees et al., 2016 for reviews) and, to a lesser extent, talent development environments (see Henrikson, Stambulova, & Roessler, 2010; Martindale et al., 2005 for reviews). However, it is possible that much of this research may not directly address the needs of the participants in the current study because they specifically expressed a desire for research examining LTAD models (which are not based on this body of talent development research). Studies of LTAD are scarce and, of the few published studies, there has been a focus on coaches’ perceptions of LTAD (e.g., Black & Holt, 2009; Chevrier, Roy, Turcotte, Culver, & Cybulski, 2016) and their views of its adoption and implementation (Beaudoin, Callary, & Trudeau, 2015). Ford et al. (2011) suggested that LTAD models are fundamentally based on physiological principles. However, in reviewing some principles associated with LTAD, specifically in terms of physical literacy, aerobic performance, and anaerobic performance (speed, strength, power), they concluded “there is little evidence to support the LTAD claims” (p. 398). Ford and colleagues further suggested that coaches should be better educated about how to interpret the recommendations of the model. The participants in the current study appeared to appreciate some of the limitations of the LTAD model, and their comments reinforce the need for more empirical examination of the foundational principles contained with LTAD models. It may be that sport organizations themselves could conduct (or commission) LTAD research. In fact, we did find some examples of sport organizations conducting their own evaluations of LTAD (in the absence of partnerships with university researchers). ON_A_PSO9 explained: What's happening right now is our NSO is rolling out an LTAD program. A year ago, they were rolling out a pilot for the LTAD program and they invited clubs across the country to volunteer to be part of the pilot program. So in terms of evaluation and ongoing assessment, they definitely are trying to do it in a very deliberate manner and make sure they're collecting all that information. We suspect that perhaps it is neither the sole responsibility of a PSO or NSO, nor that of research teams, to evaluate LTAD models. Rather, partnerships between sport organizations and researchers would seem to be a fruitful avenue for examining the effectiveness of LTAD models. Participation and retention ~~~~~~~~~~~~~~~~~~~~~~~~~~~ A total of 42 participants (70%) reported preferences for research examining athlete participation and retention. There was a need for research to track athlete participation rates. A participant from Prince Edward Island, PSO#3 (PEI_PSO3) highlighted some of the challenges in this respect: I don't think we have a very good handle on participation rates in this province. We have self-reporting from our member sport organizations but you've got overlap with that because there's kids that play multiple sports that you're duplicating the count … From our perspective, I would love a robust way of truly tackling our participation rates. In a similar vein, participants highlighted a need for research tracking athlete retention rates. AB_PSO6 said: I don't think we've done a lot on retention on any sport. No one can really tell you what their retention rates are … why are they quitting? Stuff like that would be beneficial, if you knew your population. You [can] just do it by postal code, you know if you're hitting those guys or not … No one can really tell you what their retention rates are. Some sport participation tracking data are available. For example, one national survey in Canada showed that 51% of children aged 5–14 years (approximately 2 million children) regularly took part in sport in the previous 12 months in 2005, compared to 57% of children in 1992 (Clark, 2008). Another Canadian survey showed that an estimated 77% of 5–19 year olds participated in organized sport or physical activity programs, and participation rates were relatively stable compared to the previous eight years in which this survey had been conducted (CANPLAY, 2015). A recent Australian survey showed that the participation of Australian children aged 5–14 years remained relatively stable at 64% participation in 2000 to 62% participation 2012 (Vella et al., 2016). However, as Both, Rowlands, and Dollman (2015) observed, sport participation research is littered with methodological issues such as differences in questionnaire design, sample age, study period, and definition of participation that make it difficult to compare participation trends over time. Nonetheless, the key issue for participants in the current study was not sport participation trends at a national level, but rather participation trends in their own sport and/or province. Presumably this issue could be addressed by building tracking mechanisms into their registration systems (e.g., assigning participants an identification number which is used each time they register for a program). In fact, while it would present a logistical challenge, all children in a province could be assigned a ‘sport identification number’ that would produce data regarding their movement between different sports, which would eliminate “duplicating the count” as PEI_PSO3 put it. As the quote from AB_PSO6 (above) alluded to, participants also highlighted the need to understand more about participation and retention of individuals from specific demographic sections of the population. For instance, ON_B_PSO8 said: Yeah, I think dwindling participation. I would say to try to figure out why. I think most sports are experiencing a bit of atrophy. I would say you wanna figure out what motivates each age and stage of the general population to participate in sport so that I would have more appropriate information to know how to target certain demographics. So I would say that would be very useful, something that our general membership could learn from, and we could learn how to adapt and innovate as a sport to remain viable. Functionally, sport participation requires two intertwined types of resources: opportunities to engage in sport and motivation to engage in those opportunities (Balish, McLaren, Rainham, & Blanchard, 2014). However, individuals from various demographic groups face numerous barriers that restrict their opportunities to engage in sport. For instance, in countries that use a ‘pay-to-play’ privatized model of youth sport, financial costs are a barrier to participation among children from low-income families (e.g., Holt, Kingsley, Tink, & Scherer, 2011). Further reflecting the idea of targeting particular groups, MB_PSO8 commented on two primary (related) issues: First of all, I think because the majority of our high performing athletes in Manitoba are male. And I think the second reason is because I think girls are up to six times more likely to drop out of sport than are boys in Canada. I think that there's a lot of reasons for it, but definitely some psychological issues that are attached to that, or some social stigmas that are attached to being a female athlete. And so it would be nice to see what we could do to combat that. Girls are generally less likely than boys to participate in sport (Clark, 2008). Reasons associated with reduced participation and dropout among girls are multifaceted. For example, Eime et al. (2015) showed that intrapersonal barriers (lack of time, lack of energy, and perceived competence), interpersonal factors (family and friend/peer support), and environmental/organizational (access, opportunity and resources) were associated with reduced physical activity and sport participation among adolescent girls. Such evidence notwithstanding, researchers have recently acknowledged that evidence gaps in this area still include a lack of participation information (e.g., with regards to dropouts and commencers of sport participation), lack of surveillance data, and measurement differences (Vella et al., 2016). Our participants, as well as recent commentaries by researchers, suggest a need for on-going research to track sport participation and retention in specific sports and across population subgroups. Parenting ~~~~~~~~~ In total, 42 participants (70%) reported preferences for research examining various aspects of parental involvement in sport. Some participants expressed concerns about parents exerting too much pressure on their children to achieve and parents interfering with coaching. ON_B_PSO9 said, “You know every parent would like the best for themselves and they're too eager to interfere, and parents' interference [is] one of the biggest challenges.” Furthermore ON_B_PSO10 said: I don't know what the hell we do with parents. If we could find a way, something to get into the psyche of the parents, I mean we mandated Respect in Sport [an online parent and coach education module] because we just don't wanna have poor parent behavior … The parents are a huge problem … I think if we could resolve the parental problems, and that's the biggest problem with our club teams, that's the biggest complaint I get from them, is the parents are killing my coaches. They coach for a year, they quit. There is an extensive body of research examining parenting in youth sport (see Holt & Knight, 2014; Knight, Berrow, & Harwood, 2017 for recent reviews). Indeed, researchers have examined the effects of parental pressure on young athletes for many years (e.g., Brustad, 1988; Leff & Hoyle, 1995). It appears that much of the research on parental pressure had not reached the participants in the current study, which suggests a knowledge translation issue. Participants focused on ways to enhance parent education to address concerns about parents. ON_A_PSO4 highlighted that: The parents are key … That to me would be very useful information because then that could be something that we could then turn around and give to the clubs in order to inform the parents that some of the things they're doing are not beneficial. Québec PSO#4 (QC_PSO4) said that “Educating the parents is a big challenge … What is the right attitude to have with your child for them to be happy? For them not to feel overwhelmed? For them to have a good self-esteem?” Similarly, PEI_PSO5 said: On the parenting side, helping to educate parents on how they can help their children, not only after a game or before a game, that kind of stuff, but even that they can be reaching out to professionals to help make them better in their sport …. I would love more research or guidance on how to deal with parents. Participants also expressed concerns about parent-coach relationships. QC_PSO 7 said: Yes. That's it. Sometimes it can happen. It's very, very close. If I give a lesson, of course, if the parent is behind me, he/she will hear everything that is going on during the lesson. Sometimes it can be good, but sometimes they can misunderstand things. [ Research] might be something about knowing how to deal with parents like that. There are only a small number of studies examining parent education approaches (e.g., Dorsch, King, Dunn, Osai, & Tulane, 2017; Thrower, Harwood, & Spray, 2017) and parent-coach relationships (e.g., Barber, Sukhi, & White, 1999). Hence, while there may be a research to practice gap in terms of research on the effects of parental pressure in sport, our findings also highlight important future directions for research examining parent education and parent- coach relationships. Such research would likely be highly relevant to sport organizations. Benefits of the sport ~~~~~~~~~~~~~~~~~~~~~ A total of 39 participants (65%) reported preferences for research that could demonstrate the personal and social benefits of the sport. By and large, participants wanted to know more about the specific benefits of participating in their sport, often in comparison to other sports. For example, PEI_PSO6 said there should be more research on: The social benefits of [name of sport]. It is kind of a unique sport in the fact that it’s a team sport with an individual [performance focus] in a team context. It's not like [ice] hockey where in order to win, the team moves the puck and passes. Like in [name of sport] you have to do one skill by yourself. ON_B_PSO6 highlighted the need for more research into and promotion of “the benefits of team sports off the court or off the ice or whatever it happens to be.” Similarly, AB_PSO4 said: … it would be neat from our perspective to have someone dig a little deeper into the specifics of the game and what it offers and why to choose [our sport]. So what's different about [our sport]? And what are the benefits? … What are the benefits of the intangible part that will hopefully enable you to be successful in a career or business or understanding team dynamics? Although researchers have questioned the assumption that sport participation leads to positive developmental outcomes (e.g., Coalter, 2013), and some research has revealed sport participation has been associated with negative outcomes (e.g., increased substance use; Veliz, Schulenberg, Kloska, McCabe, & Zarrett, 2015), there is a growing body of evidence depicting the reported benefits of sport participation (see Eime, Young, Harvey, Charity, & Payne, 2013; Holt et al., 2017 for reviews). For example, studies have revealed that compared to non- participants, youth sport participants score better on markers of physical health outcomes (Pate, Trost, Levin, & Dowda, 2000), educational attainment (Fredricks & Eccles, 2006), and – when sport participation is combined with participation in other youth development activities – positive psychological functioning (e.g., Zarrett et al., 2009). Research in the area of positive youth development (PYD) suggests that developmental benefits are not gained merely by participating in sport; rather, they appear to be largely dependent on social interactions that occur within sport settings (e.g., empathetic relationships with adult coaches and leaders, positive interactions with peers, and the supportive involvement of parents; Holt et al., 2017). Comparing the participants' desire for more research on the benefits of their sport with existing literature in this area, a plausible implication is that research efforts could focus on examining the contextual conditions that exist within the programs offered by particular sports and the extent to which they may contribute to, or detract from, participants' acquisition of positive developmental outcomes. Put differently, it is not a question of whether participation in a particular sport leads to positive outcomes, but rather a question of what contextual conditions exist with these sports that contribute to or detract from participants’ acquisition of positive outcomes. There may be some unique traditions and characteristics in some sports (e.g., certain martial arts; Chinkov & Holt, 2016) that contribute to participants obtaining benefits, and evidence from a systematic review showed that team sport participation was associated with improved mental health outcomes compared to individual activities, which the authors attributed to the social nature of team sport participation (Eime et al., 2013). Evans et al. (2016) suggested that sport types (e.g., amount of interdependence among teammates, physical contact), sport settings (e.g., community leagues, levels of competitiveness) and individual patterns of sport involvement (e.g., amount of organized sport, specialization or sampling pathways) may shape the psychosocial experiences of young athletes in different ways. Furthermore, outcomes may vary depending on the type of participation in said sport (e.g., at lower versus higher levels of competition) or the dimensions of involvement in terms of breadth, intensity, duration and engagement at different ages (Bohnert, Fredricks, & Randall, 2010; Côté & Hancock, 2016). Whereas it may be an oversimplification to merely study the benefits of a particular sport, more sophisticated studies that involve examining the effects of context and type of participation appear to be important future research directions that would also be relevant to the participants in the current study. Coaching ~~~~~~~~ A total of 39 participants (65%) reported preferences for research on coaching. Some PSOs wanted to know more about the effects of particular coaching approaches. MB_PSO2 said: It would be fascinating to see the direct impact of coaching. I know we see it first-hand in results, you know Olympic medals. But we'd like to have it at the youth levels. What kind of impact coaching has at a youth level? … Is there a correlation between a coach being really ethical and is his team being the same? Is there a correlation between really practicing good skills at an early age and those athletes continuing on? This finding was somewhat surprising, given that PSOs are responsible for providing coach education training schemes (which, presumably, are informed by research to some degree). Furthermore, there is an extensive body of research on coach education and coaching practices spanning several decades. Over 25 years ago Smith and Smoll's research showed, for example, that children with low self-esteem responded more positively to coaches who were reinforcing and encouraging and provided technical instruction, while they responded more negatively to coaches who were low on these dimensions (Smith & Smoll, 1990). More recently, research suggests that transformational leadership behaviors among coaches can have positive effects on participant outcomes by helping them to think more positively about themselves and their tasks, enhancing the quality of their relationships by creating environments that are fair, respectful, and supportive (Turnnidge & Côté, 2017). Thus, it appears that research on coaching effectiveness has not adequately reached the sport organization representatives sampled in the current study. This finding suggests a knowledge translation issue. The need for research examining on-going coach education was also highlighted. For instance, ON_A_PSO8 told us “[athletes] say they need better coaching, then obviously what kind of research can we do that would make sure that they're getting better coaching?” Similarly, AB_PSO12 said: Coaching in general. Any research on that would definitely be beneficial, how our coaches can become better coaches, what to stay away from, what to kinda strive to go towards type of thing would be all good stuff …. We're like a sponge we want to soak as much as we can in some of these areas and apply it to [name of sport] if possible. Here, the issue seems to reflect the continuing professional development of coaches (in the participants' words, are coaches getting “better”?). Griffiths, Armour, and Cushion (2016) claimed that there is a lack of “conclusive evidence” about the effectiveness of different types of coach education on coach learning and changes to practice, meaning that little is known about “what works” (p. 1). More specifically, researchers have tended to focus on specific learning activities and individuals, rather than institutional or organizational factors that mediate learning. In their evaluation of a sport organization's youth coach education program in the UK, Griffiths et al. (2016) found that a clear understanding of the ‘transmission model’ for education was missing and the different ‘communities’ that could support learning were not connected within the organization. From the perspectives of participants in the current study, it would appear that important future research directions could involve examining sport-specific coach education and development programs and their impact on both coaches and young athletes. Other participants talked about coach retention. MB_PSO7 suggested it would be important to learn more about: What motivates people to become coaches because becoming a coach, there's not a lot of money in it but it takes a lot of time and energy, so it's challenging to keep volunteer coaches involved. So what motivates people to become coaches and continue in sport coaching volunteer role? The retention of volunteer coaches is a critical issue for youth sport delivery systems in many countries. In a focus group study, Rundle-Thiele and Auld (2009) found enjoyment, success, and support from parents, clubs, and the league were key factors that contributed to Australian youth coaches’ decisions to stay involved as a volunteer coach. Results from a survey conducted in the United States showed that 94% of youth sport coaches from a municipal parks and recreation soccer program were motivated to stay involved because they wanted to instill positive values in children through their coaching (Busser & Carruthers, 2010). However, there does not appear to be extensive evidence on factors that influence coach retention over time, and this could represent a valuable area for future research. Finally, several participants highlighted the need for more research with parent-coaches. ON_A_PSO4 shared the following specific example: We have one parent in particular who's now star of the club and his son is one of the best [athletes] in Canada, but the kid has no break because his dad is coaching him. And his dad was not an athlete. So you've got that kid who's under so much stress 24-h and I think that they're practicing 15–18 hours a week. That's a lot of time. That's a lot of their free time. The rest of the time they should be home being a kid. It is noteworthy that there is very little previous research on parents as coaches (e.g., Weiss & Fretwell, 2013), despite the fact that parents often appear to fulfill a coaching role on their children's youth sport teams (Holt & Knight, 2014). The participants' priority for future research examining parent-coaches may represent an important area for future research that would be beneficial to Canadian PSOs and other sport systems that rely on parent-coaches.","The overall purpose of this study was to examine Canadian PSO representatives’ research priorities in order to provide directions for future research and knowledge translation initiatives in youth sport. The most frequently identified areas were athlete development systems, participation and retention, parenting, benefits of sport, and coaching. There was an underpinning preference for sport-specific research in several of these areas. We then compared findings to the existing youth sport literature, revealing future research directions that may be valuable for stakeholders in youth sport. Conversely, the results also revealed some of the reported priorities have already received extensive attention in the youth sport literature, which is instructive information for future knowledge translation initiatives. A promising feature of this study is that there does appear to be some enthusiasm and appreciation for the adoption of research evidence among the 60 PSO representatives sampled. These participants were able to articulate a wide range of research topics that would assist their work in youth sport. Hence, it appears that collaborative approaches that include involving stakeholders in the generation of research agendas are fruitful endeavors for pursue in the future, which may reduce ‘waste’ (Chalmers et al., 2014) and increase the adoption of research evidence in practice, program, and policy contexts (Innvaer et al., 2002; Straus et al., 2013). Returning to the KTA model (Graham et al., 2006) that guides our larger project, the current study was located at the start of the action cycle and focused on identifying problems and searching for existing relevant research knowledge. Similarly, the small body of previous knowledge translation research in sport can be located in the ‘early stages’ of knowledge translation models, given that the focus has been on identifying barriers to the use of evidence (a type of problem identification itself). Though the evidence base is small, a ‘picture’ is beginning to emerge; people who work in the sport sector (e.g., coaches, sport organization representatives) face numerous barriers that restrict their ability to access and use research evidence (Holt et al., 2018; Pain & Harwood, 2004; Reade et al., 2008b, 2008a; Williams & Kendall, 2007). Researchers themselves may face barriers to engaging in knowledge translation research. Traditional academic graduate training models may not adequately prepare researchers for effective knowledge translation (Gould, 2016; Greenwood & Abbott, 2001). Fully embracing knowledge translation in youth sport may require major training initiatives. Engaging stakeholders in the development of research agendas will be an important starting point. Other important steps include considering what knowledge should be disseminated, to whom, by whom, how, and with what effects? (Lavis, Robertson, Woodside, McLeod, & Abelson, 2003). By integrating such dissemination topics into graduate training and emphasizing the need to collaborate with stakeholders in all stages of the research process (especially in the generation of ideas), important steps toward creating a generation of researchers dedicated to enhancing the use of research evidence in youth sport will be made. These suggestions for future research and training notwithstanding, we also wish to critically reflect on the very premise of using research evidence to inform decisions in youth sport. To draw an illuminating historical comparison, evidence-based medicine (EBM) became a paradigm for medical practice in the early 1990s (Montori & Guyatt, 2008). EBM emphasizes the examination of evidence from clinical research, requiring physicians to have skills in formulating questions, searching and retrieving the best available evidence, and critically appraising studies to establish the validity of results (Evidence-Based Medicine Working Group, 1992). Yet, whatever the evidence, value and preference judgments are implicit in every clinical decision, not only in relation to patients' preferences but also physicians’ evaluations of the relative benefits/harms or costs of treatment (Guyatt et al., 2008). Enduring concerns in the adoption of EBM in medical and allied health professions have included excessive reliance on easily obtained but potentially misleading evidence and the increase in commercial interests to produce and interpret evidence for physicians (Als-Nielsen, Chen, Gluud, & Kjaergard, 2003; Montori & Guyatt, 2008). Some of these challenges may also apply to sport. For example, commercial entities often make unsubstantiated claims about the performance benefits of their products or services, a recent example being the use of genetic testing for talent identification (Vlahovich, Fricker, Brown, & Hughes, 2016). More recently, the notion of practice-based evidence has been introduced in the health care field as it begins to advance beyond the EBM paradigm. As Green (2008) remarked, “if we want more evidence-based practice, we need more practice-based evidence” (p. 23). Practice-based evidence acknowledges that efficacy or effectiveness research is just one of several sources of information needed to improve health care, and that understandings of the challenges faced by people who both receive and deliver interventions are needed (Green, 2008). Ammerman, Woods Smith, and Calancie (2014) suggested community-based participatory research approaches, whereby research and interventions are informed by the views of different stakeholders, as a useful strategy for developing practice-based evidence. Furthermore, they argued that systems-based models that consider dynamic, complex, and multilevel factors and include the views of multiple stakeholders hold great promise for generating practice-based evidence. Working collaboratively to set research agendas with appropriate stakeholders in youth sport may provide a useful way forward for bridging research to practice gaps, and there are some examples from public health that could provide models for future endeavors in this regard (e.g., Lenaway et al., 2006). Another useful concept to consider is evidence-based policy. Essentially, evidence-based policy is an extension of evidence-based practice, suggesting that if practitioners are expected to base their work on evidence then the same expectations should be placed on policy-makers (Ham, Hunter, & Robinson, 1995). Evidence-based policy is sometimes viewed as a linear process, whereby a problem is defined and then research evidence is amassed to provide policy options (Black, 2001). However, this linear process, whereby evidence leads to the development of policy, may be problematic. Policy change is not simply made based on the availability of evidence; it may depend on an appraisal of what is scientifically plausible, what is politically acceptable or achievable, and what is practical at a given point in time (Nutbeam & Milat, 2017). Indeed, representatives of sport organizations have expressed a preference for using research that supports the programs they are already planning, rather than using evidence as a starting point to inform new programs (Holt et al., 2018). One concept underlying several of the themes (most notably athlete development systems, participation and retention, and benefits of the sport) was a preference for sport-specific research. In this respect, the health care literature can again provide some guidance for future knowledge translation research in youth sport. Within the KTA framework (Graham et al., 2006) knowledge tailoring is a critical issue for youth sport researchers to consider, because ‘generic’ knowledge that is not specifically contextualized is “seldom directly taken off the shelf and applied without some sort of vetting or tailoring to the context” (Graham et al., 2006, p. 20). That is, the implications of relevant research, even when it is accessible, must be adapted to local contexts (e.g., sport systems). Our results regarding the preference for sport-specific research highlight the need to tailor ‘generic’ knowledge to local contexts (cf. Straus et al., 2013). From here, from a knowledge translation perspective, future projects that involve adapting implications to local contexts will provide opportunities for primary studies that monitor knowledge use, evaluate the outcomes of using knowledge, and examine sustained ongoing knowledge use (Graham et al., 2006). A key strength of this study was the sample, specifically that we were able to recruit 60 people from five provinces. This facilitated a high degree of data saturation. The thorough approach to training the interviewers was another strength. Nonetheless, several limitations of this study must be considered. Qualitative description, while an appropriate methodological selection for this study, is perhaps one of the most simplistic forms of qualitative research (Neergaard et al., 2009). Another limitation is that sport organizations are merely one type of stakeholder in the world of sport. There is a need for further research with other stakeholders (e.g., coaches, athletes, parents). Additionally, in considering the implications of this study, the extent to which the findings would apply to youth sport in other jurisdictions is unknown. However, it is quite likely the findings go beyond the Canadian context. For example, the most frequently reported theme, LTAD, is not only a fundamental feature of the Canadian sporting landscape; iterations of LTAD are used in numerous countries (including, but not limited to, the United Kingdom, Ireland, South Africa, and New Zealand) and, of course, other countries have some form of athlete development models in place. Thus, some of the current findings may well apply beyond the Canadian context. In addition to highlighting a need for more LTAD research, other implications for generating future research arising from this study include finding ways to track athlete participation and retention, working with sport organizations to develop parent education initiatives, understanding the contextual conditions and participation factors that may contribute to or detract from the acquisition of positive developmental outcomes in particular sports, examining the effects of on-going coach development, and studies of coach retention and parents as coaches. By identifying these topics our study may help inform the development of future programs of study that would be highly relevant to sport organizations. Implications for knowledge translation include the need to develop programs of research dedicated to knowledge translation, which involve engaging stakeholders from the beginning rather than merely developing dissemination strategies. Studies that examine the complexities surrounding the implementation and adoption of research evidence, and the evaluation of knowledge translation initiatives, are important next steps. Finally, our findings suggest that knowledge tailoring strategies (Graham et al., 2006) will be important in the future in order to contextualize the implications of ‘generic’ knowledge for particular sport organizations. We hope that the views of the participants in this study may help inform future research agendas, increase the applicability of findings, and ultimately their translation into the policies, programs, and practices of youth sport."],["Two challenges that face popular self-monitoring theories (SMTs) of auditory verbal hallucination (AVH) are that they cannot account for the auditory phenomenology of AVHs and that they cannot account for their variety. In this paper I show that both challenges can be met by adopting a predictive processing framework (PPF), and by viewing AVHs as arising from abnormalities in predictive processing. I show how, within the PPF, both the auditory phenomenology of AVHs, and three subtypes of AVH, can be accounted for. --------------------------------------------------------------------------------","The positive symptoms of schizophrenia include delusions of control (“somebody else is controlling my actions”), thought insertion (“somebody is putting their thoughts into my head”) and auditory verbal hallucination (AVH) (hearing voices in the absence of a speaker). Perhaps the most popular theories for understanding these disparate symptoms are self-monitoring theories, which attempt to explain them as the product of one abnormality, namely, a problem with self-monitoring. According to these theories, our nervous systems distinguish self-generated from externally generated stimuli, through a process of self- monitoring. When this monitoring goes awry, self-generated stimuli are erroneously attributed to an external cause. The various positive symptoms all involve faulty monitoring and simply differ insofar as that which is failing to be properly monitored differs. In delusions of control it is bodily action, whereas in AVH and thought insertion it is widely thought to be inner speech (Feinberg, 1978; Frith, 1992; Jones & Fernyhough, 2007; Seal, Aleman, & McGuire, 2004). Although ingenious, the breadth of application of SMTs has recently been questioned by several theorists (Gallagher, 2004; Jones, 2010; Stephens & Graham, 2000; Wu, 2012). Such criticisms tend not to take issue with the application of SMTs to symptoms involving bodily action, like delusions of control and (their merely experiential analogue) illusions of passivity. Rather, they claim that SMTs struggle to account for AVH and thought insertion. In this paper, I will focus on AVH, and on two challenges in particular. They are: The Auditory Phenomenology Challenge – How do you explain the auditory phenomenology of AVH if it is misattributed inner speech? The Varieties of AVH Challenge – How do you account for the varieties of AVH if it is (always) misattributed inner speech? In this paper, I suggest that both challenges can be met if we adopt a recently popular general framework for thinking about what the brain does (e.g. Clark, 2013; Friston, 2005, 2010; Hohwy, 2013) which we could call the predictive processing framework (PPF). It is worth mentioning that the application of predictive processing to psychosis is not new. Indeed, Chris Frith, perhaps the best-known proponent of SMTs, has suggested something along these lines in Fletcher and Frith (2009). Since then, Adams, Shipp, and Friston (2013) have also suggested accounts of psychosis within the PPF. This work, however, does not focus on AVHs to the extent that I do, nor does it focus on the two challenges that I address here. I proceed as follows. I start by presenting SMTs and show why they have been found attractive and plausible. I present the two challenges facing the application of SMTs to AVHs. I then introduce, motivate and clarify the PPF. I then present evidence suggesting that predictive processing might be disrupted in psychosis. Finally, I end by applying the PPF to voice-hearing, and show how it can, first, address the auditory phenomenology challenge, and second, nicely account for the three subtypes of AVH I present.","In this section I characterise SMTs, and describe the evidence that has been used to support them. Introducing self-monitoring theories (SMTs) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Perhaps the first theorist to make use of self-monitoring was Helmholtz (1866). His concern, however, was not with psychopathology, but with the following problem presented by healthy visual cognition. When an image moves across the retina, how does our brain know whether it is the world moving across our eyes or our eyes moving across the world? Helmholtz suggested that our brain can tell the difference because when our eyes move there is a motor command. More specifically, information about the motor command, which Sperry (1950) later dubbed the “corollary discharge”, is used by the brain to predict the sensory consequences that would be produced by the eye movement. If the predicted and actual sensory consequences match then the brain infers that the change was self-generated and the conscious percept is adjusted accordingly. We can see exactly what happens when there is no such motor command, and hence no such adjustment, when we press on our eye with our finger. When we do this, the world itself seems to tilt and shake. It took more than a hundred years for Helmholtz’s ideas to be applied to psychosis (Feinberg, 1978). Although Feinberg’s initial paper was on thought (which he took to involve ‘motor mechanisms’) and thought insertion, the easiest symptoms for which to introduce the account are delusions of control, since it is clear that, if anything involves motor commands, bodily actions do.1 In delusions of control, a subject may perform actions that are in keeping with her plans and intentions (for example, she might brush her hair), but she claims that somebody else is controlling her. Frith and Done (1989) took this to be a problem with self-monitoring. In particular, there is a mismatch between the predicted and actual sensory consequences of the bodily movement and so (as with Helmholtz’s ocular example) the movement is attributed to an external source. This, in principle, could be a problem with the generation of the prediction itself (e.g. the corollary discharge) or with the mechanism that compares expected and actual sensory (including proprioceptive) input, what became known as “the comparator”. Later manifestations of the self-monitoring theory (Frith, Blakemore, & Wolpert, 2000) saw comparator-based self-monitoring as the human body’s way of meeting a computational challenge, in particular, involving skilled reaching and online correction (see Wolpert, 1997 for a review of these computational approaches to motor control). This connection, in the late nineties, between the computational neuroscience of healthy cognition and the neuropsychology of schizophrenia undoubtedly contributed to the credibility of SMTs. Support for SMTs ~~~~~~~~~~~~~~~~ Whereas in Helmholtz’s example, the recognition by the nervous system that a certain stimulus is self-produced causes a correction of the conscious percept, in more typical bodily motor control, it results in sensory attenuation. The evolutionary benefit of this is clear enough: your nervous system needs to pay attention to stimuli that come from the outside, not the endogenous stimuli that (in a well-functioning system) will be harmless and irrelevant. Various data suggest that something goes wrong with this monitoring and subsequent attenuation (Blakemore, Smith, Steel, Johnstone & Frith, 2000). The most striking such datum is the reported finding that subjects with diagnoses of schizophrenia can tickle themselves. The postulated explanation for this is that there is a mismatch between expected and actual sensory consequences and the sensory consequences are not attenuated: the tickling sensation is like being tickled by somebody else. Typical subjects can’t tickle themselves because their nervous systems accurately monitor, and successfully attenuate, the sensory consequences of the tickling movements (Blakemore, Wolpert & Frith, 1999). It is not only with bodily action that support has been shown for the claim that self-monitoring goes awry in schizophrenia. Studies showed that patients with diagnoses of schizophrenia are more likely to misattribute their own voices than healthy controls (Cahill, 1996; Johns et al., 2001). For example, Johns et al. (2001) got subjects to read words aloud and played them feedback of their voices with mild acoustic distortions. The subjects with diagnoses of schizophrenia were considerably more likely than healthy controls to claim that the voice they heard was someone else’s. Indeed, support that this was directly related to the production of hallucinations was supported by the result that, within the “schizophrenia” group, those who experienced hallucinations were more likely to misattribute their voice than those who did not hear voices. Applying SMTs to AVH ~~~~~~~~~~~~~~~~~~~~ Several theorists (Feinberg, 1978; Frith, 1992; Jones & Fernyhough, 2007; Seal et al., 2004) have attempted to explain AVHs in terms of inner speech misattribution, based on self-monitoring abnormalities. Although from a pre-theoretical standpoint it is not obvious that inner speech involves motoric elements, this has been empirically supported by several electromyographical (EMG) studies (which measured muscular activity during inner speech) some of which date as far back as the early 30s (e.g. Jacobsen, 1931). Later experiments made the connection between inner speech and AVH, showing that similar muscular activation is involved in healthy inner speech and AVH (Gould, 1948; McGuigan, 1966). The involvement of motoric elements in both inner speech and in AVH is further supported by findings from Gould (1950), who showed that when his subjects hallucinated, subvocalisations occurred which could be picked up with a throat microphone. That these subvocalisations were causally responsible for the inner speech allegedly implicated in AVHs, and not just echoing it (as has been hypothesised to happen in some cases of verbal comprehension (cf. e.g. Watkins, Strafella, & Paus, 2003)), was suggested by Bick and Kinsbourne (1987), who demonstrated that if people experiencing hallucinations opened their mouths wide, stopping vocalisations, then the majority of AVHs stopped. To recap, SMTs have been presented as a unifying model for understanding the positive symptoms of schizophrenia. The idea is that there is one deficit concerning the monitoring of self- produced stimuli, and the symptoms differ because there are different kinds of self- produced stimuli. In delusions of control it is physical action, whereas in AVH and thought insertion it is inner speech.","As mentioned, criticisms of SMTs tend not to take issue with the role of self-monitoring deficits in delusions of control, but rather concern the aforementioned extension of self- monitoring to symptoms that don’t involve bodily action, namely, AVH and thought insertion. In this paper, I specifically focus on two challenges facing the application of SMTs to AVHs. The Auditory Phenomenology Challenge ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ How are we to account for the distinctly auditory phenomenology of certain AVHs? As Cho and Wu (2013) put it, if a theory claims that AVHs are misattributed inner speech, it must explain how we get a transformation from the experience of the subject’s own inner voice […] often lacking acoustical properties such as pitch, timbre, and intensity into the experience of someone else’s voice with acoustical properties. (p.2) As Wu (2012) puts it in an earlier paper, “we must explain this ‘transformation’ from the normal to the pathological” (p.94). Although one can question the premise that inner speech lacks auditory phenomenology (McCarthy-Jones and Fernyhough, 2011; Moseley & Wilkinson, 2014), even granting that this is the case, the PPF makes this ‘transformation’ less perplexing.2 The Varieties of AVHs Challenge ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Jones (2010) emphasises the heterogeneity of AVHs: The term AVH encapsulates a diverse phenomenological experience, which may involve single and/or multiple voices, who may be known and/or unknown, speaking sequentially and/or simultaneously, in the first, second, and/or third person and which may givecommands, comments, insults, or encouragement. (2010, p.566) In particular, Jones presents two models for understanding AVHs and argues that both are promising for understanding different subtypes of AVH. The first model he examines is the “memory-based” model proposed by Waters, Badcock, Michie, and Maybery (2006), which he suggests fits well with the phenomenological features of AVHs that appear to be the upshot of a traumatic experience. In these AVHs you often find, for example, instructions to self-harm in a voice that is recognisably that of the abuser. Even more strikingly, the contents of these AVHs “can be related to what was said during or surrounding these events (e.g. if you tell anyone I’ll kill you)” (Jones, 2010, p.568). However, these form only a relatively small subset of AVHs (e.g. 4 out of 24 in Fowler et al., 2006). A larger subset of AVHs seem to serve the function of regulating or commenting on current events (e.g. Nayani and David (1996) reported that for 46% of their sample, the voice had replaced their “voice of conscience”). For these, Jones (2010) suggests that his second candidate model is a better option. This model is precisely the inner-speech-based SMTs that we introduced in Section 1. The moral of this is that it is very difficult to account for all this variety within a model that explains AVHs in terms of misattributed inner speech. I am in complete agreement with Jones that we need different models for different subtypes.3 I would simply add to Jones’s suggestion, however, that we distinguish frameworks, theories, and models. Roughly, the distinction is as follows. Frameworks are very broad; they are ways of approaching a particular domain of inquiry (e.g. the brain and cognition). It is within them, whether implicitly or explicitly, that theories are built.4 Theories are falsifiable claims (rather than ways of approaching something) at a high level of generality that explain a phenomenon or class of phenomena by elucidating the fundamental nature of the phenomenon, or by postulating certain rules or principles that govern or define the phenomenon (e.g. a self-monitoring theory of schizophrenia). Models are at a lower level of generality; they explain how a particular kind of phenomenon arises, sometimes in terms of other, hopefully better understood, phenomena (e.g. a misattributed inner speech model of AVH). Theories might suggest candidate models. So a self-monitoring theory of the primary symptoms of schizophrenia might suggest a misattributed inner speech model for AVHs occurring in the context of schizophrenia. Models should be applicable to individual cases. Thus, given a theory that posits one kind of deficit, different models could well be needed for different symptomatic manifestations of that underlying deficit. In a nutshell, I will suggest that the different subtypes need (as Jones, 2010 suggests) to be explained in terms of different aetiological models. However, these models are to be built within the same framework, the predictive processing framework that I will present now.","I’d like to put psychosis to one side, and focus, in this section, on presenting a framework for understanding the brain and cognition generally. We will later see how this might apply to AVH. I will start with an initial presentation of the PPF, and then present some conscious effects that are suggestive that the brain is engaged in predictive processing. An initial presentation: “inference” and efficiency ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ According to the PPF, the brain’s main task is to “infer” from incoming signals, what the causes of those signals are, in other words, settle on a hypothesis about what is “going on”.5 However, the incoming signal, namely, the proximal stimulus on sensory receptors, underdetermines distal causes. Since inputs are noisy and ambiguous, there is no one-to- one mapping: the same stimulation can be brought about by two very different distal causes (and different stimulation in different circumstances can be caused by the same distal cause). Given that more than one hypothesis is compatible with the incoming signal, how does the brain settle on one hypothesis rather than another? It needs to take two things into account: first, the fit of the input with the hypothesis, and, second, how statistically likely that hypothesis is (the “prior probability”), at least as far as the brain is concerned (this is subjective rather than objective probability). A hypothesis could fit the input extremely well, but its prior probability could be so low that it isn’t even considered. Conversely, a hypothesis could have such a high prior probability, that, even though it doesn’t fit the input well, it is settled upon. Although this is a picture of what the brain has to do in order to do a good job of “inferring” what is going on, it is abstracted from the temporal dynamics of the brain’s functioning. In reality, the inputs, hypotheses and prior probabilities (priors) are in constant hierarchical interaction with each other. This is where the notion of prediction comes in. What the selection of a hypothesis does is that it determines a set of predictions about subsequent inputs, namely, inputs that are compatible with the hypothesis. If the hypothesis does a good job of predicting inputs, it will be kept. If it does a bad job, it will be tweaked or abandoned altogether in favour of another hypothesis. In other words, one hypothesis is selected rather than another if it better minimises prediction error. Some theorists (e.g. Hohwy, 2013) sum up the PPF by saying, for example, that the brain is in the business of minimising prediction error.6 This prediction error minimisation is not only taken to account for perception and cognition, but for action as well (see e.g. Adams et al., 2013). Instead of there being motor commands, as on the standard picture, what you have are predictions, which are then fulfilled by the subsequent bodily movement, thereby also being a case of prediction error minimisation. This is often called “active inference” (Friston, 2009), which Pickering and Clark (2014) helpfully gloss as follows: “the combined mechanism by which perceptual and motor systems conspire to reduce prediction error using the twin strategies of altering predictions to fit the world and altering the world (including the body) to fit the predictions” (p.1). This picture has interesting consequences for how we are to view the role of input on sensory receptors and its impact on higher cortical regions. According to the PPF, the only information that gets passed on up the hierarchy is prediction error. Indeed, as some recent theorists nicely put it: an expected event does not need to be explicitly represented or communicated to higher cortical areas which have processed all of its relevant features prior to its occurrence. This stands in sharp contrast to the standard (and admittedly intuitive) view of perception (what might be called a “bottom-up feature-detection view”). On such a view, inputs come in, are processed, and passed on. On the PPF, only the prediction error gets passed on, and therefore the vast majority of what determines your perceptual experience is what your brain has already predicted, your brain’s best hypothesis. In short, on the PPF, the role of the incoming signal in determining conscious perception is much less great than on standard bottom-up views. Precision weighting, second-order prediction and “attention” ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Although incoming signals are ambiguous, in different contexts, the degree of ambiguity will differ. To maximise its predictions, the brain needs to accurately estimate how much ambiguity (uncertainty) there will be. In contexts where low ambiguity is expected, higher precision will be demanded, and vice versa. This is called “precision weighting” in the literature (Feldman & Friston, 2010; Friston, 2009; Hohwy, 2012), and is taken to be modulated by neurotransmitters such as dopamine (Corlett, Taylor, Wang, Fletcher, & Krystal, 2010).7 Recently, theorists (Feldman & Friston, 2010; Hohwy, 2012) have equated this precision weighting with what we often call “attention”. As is well known, attention can be brought about endogenously or exogenously, namely, either because the subject is in a state of wanting to attend to something (endogenous attention), or because something in the environment attracts attention (exogenous attention). Suppose I’m talking to someone in a crowded bar. By using information from seeing the person’s lips move, my brain has certain expectations about what sounds are likely to be produced. That part of my visual field, and the corresponding part of my “auditory field”, now has relatively high precision attributed to it, largely because I am interested in what’s going to be said (endogenous focal attention). Everything else around that, both visual and auditory, has relatively low precision. My brain has a very gist-like hypothesis about what is going on there, and, any prediction error regarding that has low weighting. Of course, a loud peripheral bang or camera flash can grab my attention (a case of exogenous focal attention). That would constitute a strong enough input to overcome whatever down- modulation my brain has placed on that peripheral prediction error. To sum up, then, we can think about the directing of attention as “turning up” the expected precision of an incoming signal, which amounts to turning up the “gain” on any prediction error. This can be viewed normatively. That is to say, not only can your brain get its predictions wrong, it can also get the expected precision of its predictions wrong, i.e. its second-order predictions. Your brain can wrongly think that it is in a more or less ambiguous environment than it actually is in. And this will mean too much or too little weighting on prediction error. Binocular rivalry This is perhaps the most common example used in support of the PPF. It speaks strongly against a bottom-up view in favour of a more predictive or “inferential” view. Under experimental conditions, each eye is presented with a different, but meaningful, stimulus. One standard example involves presenting one eye with a picture of a house, and the other with a picture of a face. Subjects do not report visually experiencing, as one might expect, a mixture of face and house. Rather, they experience a “bi-stable” switching, from face to house, and back, and so on (the switching is often reported as a gradual “breaking through” of the other image). As Hohwy, Roepstorff, and Friston (2008) point out, this can be nicely explained within the PPF, as a reasonable response to highly un-ecological circumstances. In a nutshell, you experience the bi-stable state because your brain is switching between hypotheses about what is out there. When it settles on one hypothesis, say, the face hypothesis, inputs from the house image fail to accord with this hypothesis and prediction error is sent up the hierarchy. When enough prediction error accumulates, the hypothesis switches (and with that, what the subject consciously experiences) to the house hypothesis, but then the input from the face image doesn’t accord, and so on, and so forth. The fact that this bi-stable switching only occurs with certain well-chosen stimuli, suggests the operation of what might be called a “hyperprior”: an expectation about the world that is stable, and often at a high degree of abstraction. The “hyperprior” in this case is that faces and houses, being at different scales, cannot occupy the same region of the visual field at the same time. As a result, the “face-house” hypothesis about what is out there never presents itself. Note how binocular rivalry puts pressure on a “bottom-up” view of perceptual experience. As Hohwy (2013) puts it: During rivalry, the physical stimulus in the world stays the same and yet perception alternates, so the stimulus itself cannot be what drives perception (p.20) One might think that binocular rivalry could result from the operation of a different principle, and one that is compatible with a largely “bottom-up” view, namely the principle: “If your eyes each have coherent but incompatible inputs, don’t receive input from both at once.” This nicely illustrates how predictive processing concretely differs from a bottom-up view such as this one. When the subject experiences a house, on the bottom-up eye- selection hypothesis, this means that the eye being presented a house is passing on a signal, whilst the eye being presented a face is not. On the predictive processing view, it is the opposite: the eye being presented the face is passing on a signal, namely a prediction error signal. One reason for rejecting the eye- selection account is that it doesn’t (unlike the build-up of prediction error) explain the stable switching back-and-forth. A second, more striking, reason is that this account was falsified by Diaz-Caneja (1928) who discovered that if each stimulus picture is cut in half and swapped over, so that each eye is being presented with a half-face half-house image (split down the middle) there is the same experience as before, namely, a rivalry between a complete percept of a face and a house. The Hollow Mask Illusion When you are presented with a rotating mask that is slowly turned to present you with the concave back of the mask, your brain “corrects” the concave stimulus into a convex stimulus. (You yourselves can experience this illusion on several sites on the internet.) You experience the concave back of the mask, as convex. Again, this is due to a very strong prior (perhaps worthy of being called a hyperprior), that over-rides the incoming signal. The prior in question is that the faces you will encounter will always be convex (a fair expectation!). This prior is so strong that even though the “concave face” hypothesis would better match the input, it is never selected. Given the highly specialised processing that faces receive in human cognition, this is perhaps hardly surprising. The McGurk Effect Moving away now from uni-modal priors, there is an effect that deals with highly flexible, concrete, cross-modal priors. This is the fact that what you see will affect your brain’s expectations about what it hears. If a subject is played a video of someone appearing to say “ga, ga”, but you play them a synchronous audio track of someone saying “ba, ba”, you aurally experience “da, da” (something “in between”). This is called the McGurk Effect. The visual input changes your brain’s expectation concerning the auditory input, and interprets it differently. We can see the same cross-modal predictive processing (vision having an effect on audition) at work in more everyday cases. If you look across a crowded bar at someone ordering a drink, by looking at her lips, you can actually hear her order. Your brain uses the visual input to inform its priors about what is being said. If you had tried to hear her order with your eyes closed, there is no way you would have been able to single out her voice from the noise of the crowded room. Your brain just wouldn’t have had the visual cues to inform its expectations.8 Furthermore, since what she says is what matters to you at that moment, you “turn up the gain” on prediction error that is relevant to hypotheses concerning what she says. Clearing up potential confusion: what the brain does and what the person does ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Although this predictive processing is extremely relevant for our conscious experience, we need to be very careful to distinguish what our brains do from what we do. One nice way of seeing this distinction is by distinguishing “surprisal” (Tribus, 1961) and “surprise”, where the former is what is surprising to your brain (as signalled by prediction error), and the latter is what is surprising to you.9 These two come apart. If you look out of your window and see an elephant on the lawn, you might be very surprised. However, the fact that you see straight away that it is an elephant shows that your brain has already minimised the prediction error and settled on the elephant-on-lawn hypothesis. Conversely, when you are attempting a binocular rivalry task, and you’ve done it before, the switching doesn’t surprise you, however, your brain is switching between hypotheses precisely because it is struggling to keep “surprisal” to a sufficiently low level. Since we are concerned with AVHs, a conscious phenomenon, it is vital to understand this relationship between what the brain does, and what the person does, and, analogously, predictive processing on the one hand, and conscious experience on the other. Roughly, your conscious percept is determined by the overall hypothesis that your brain has adopted in order to minimise prediction error.","Now that I have presented the PPF, we can see what relevance it has for psychosis in general, and for AVH in particular. Two sources of evidence that support the hypothesis that predictive processing abnormalities might be implicated in psychosis are, first, that the conscious effects that have been taken to be suggestive of predictive processing are altered in subjects with diagnoses of schizophrenia, and, second, that there are behavioural results from eye tracking that are nicely explicable in terms of predictive processing. Conscious effects are altered in patients with diagnoses of schizophrenia ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Suppose that psychosis can be understood, in part, at least, as a breakdown in normal predictive processing. And suppose that, as we have suggested, the conscious effects just mentioned are the product of normal predictive processing. One would expect patients with schizophrenia diagnoses to experience these effects differently from the normal population. This is exactly what the evidence suggests. Schneider, Leweke, Sternemann, Weber, and Emrich (1996), Schneider et al. (2002) and Emrich, Leweke, and Schneider (1997) demonstrate that subjects with schizophrenia diagnoses do not experience the Hollow Mask Illusion. They see the mask, correctly, as hollow. Their brain does not “correct” the input. Pearl et al. (2009) showed that patients with diagnoses of schizophrenia experience the McGurk Effect less. In other words, the visual stimulus corrects the auditory input less, or not at all. Pearl et al. (2009) put this in terms of a “decreased reliance” on visual cues. Results with similar implications are reported by Ross et al. (2007) with regards to visual enhancement of speech comprehension in noisy environmental conditions. In other words, subjects with diagnoses of schizophrenia showed an impairment in their ability to enhance their auditory perception with visual cues. Within a PPF this would be put in terms of either decreased efficacy of auditory priors informed by the visual information, or in terms of excessively weighted bottom-up prediction error (or both). As for binocular rivalry, Heslop (2012) showed that switching rates in binocular rivalry were, on average, significantly slower in subjects with schizophrenia (0.28 switches per second, versus 0.54 in the non-clinical population). Evidence from eye tracking ~~~~~~~~~~~~~~~~~~~~~~~~~~ Following Adams, Perrinet, and Friston (2012), I distinguish three eye-tracking tasks where the performance of patients with diagnoses of schizophrenia differs significantly from healthy controls. First, there are cases of impaired tracking during visual occlusion (Hong, Avila, & Thaker, 2005; Thaker et al., 1998). In other words, when a moving target is temporarily occluded from view, controls are much better than subjects with diagnoses of schizophrenia at keeping track of the target when it comes back into view. Second, patients with diagnoses of schizophrenia show impaired “repetition learning.” When a target trajectory is repeated, healthy controls achieve optimal performance whereas subjects with schizophrenia diagnoses do not (Avila, Hong, Moates, Turano, & Thaker, 2006). Thirdly, there are cases of what, in the literature, is called “paradoxical improvement”, namely, where patients with a supposed illness or deficit perform better at a particular task than healthy controls (the fact that patients with schizophrenia do not experience the Hollow Mask Illusion might be thought of as a “paradoxical improvement” since their percept is more veridical). The task in question involves keeping track of a target that rapidly and unexpectedly changes direction. Subjects with schizophrenia are better at keeping track of the target than healthy controls (but only very shortly after the change in direction). Working within the PPF, Adams et al. (2012) explain these three discrepancies in terms of differing reliance on predictions (priors) and prediction error respectively. The first two tasks are improved by reliance on prediction, whereas the third task, being deliberately unpredictable, is hindered by prediction, and would be improved by a higher weighting on the prediction error. Note, however, that, like the Hollow Mask Illusion, the “paradoxical” improvement is shown on a stimulus that is statistically un-ecological. Just as we encounter convex faces and never concave ones, the world is full of statistical regularities and it is more adaptive to be able to exploit those regularities efficiently than to slightly outperform others in the rare instances when stimuli are totally unpredictable.","I will start by showing how the PPF addresses the Auditory Phenomenology Challenge. In particular, I will show how it circumvents the issue of explaining the “transformation” from the healthy phenomenon (e.g. inner speech) to the pathological phenomenon (i.e. AVH). I will then present three subtypes of AVH and show how the PPF accommodates different aetiological models for each. Dissolving the Auditory Phenomenology Challenge ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Viewed from within the PPF, this challenge is based on a mistaken understanding of how the brain works, and the relationship that this has to conscious experience. The PPF changes how one thinks of perceptual experience, and, by extension, radically changes one’s explanatory focus in trying to account for hallucinations. On a standard framework, where front-line sensory stimuli get gradually processed and passed on up the hierarchy, hallucinations make one wonder, “Where does this erroneous sensory stimulus come from?” Indeed, we can see SMTs as making attempts to answer this within a standard framework. Their answer is: they come from the quasi-sensory stimulation of inner speech, which is then misattributed. However, when, instead, we adopt the PPF, incoming stimuli play a much smaller role in determining the conscious percept, even where veridical perception is concerned. Given that a conscious percept is constituted by the hypothesis that best minimises prediction error, we don’t ask, “Where does the input come from?”, since the input alone doesn’t (and can’t) determine the percept. Rather we ask, “Why does this hypothesis minimise prediction error?” This general approach makes hallucinations both less perplexing, and less different from veridical perception.10 The point is that on the PPF, your conscious experience at any given time is the hypothesis that your brain has selected in order to minimise prediction error. Applied to AVHs, your experience will have auditory phenomenology if your brain has had to adopt the hypothesis that you are hearing something in order to minimise prediction error. This isn’t really a “transformation” because, on the PPF, it makes no sense to talk of inner speech as a “raw material” that needs transforming. Either you have an experience that is inner speech, because your brain has adopted the hypothesis that that is what is going on, or you have a different experience, corresponding to a different hypothesis. There is no experience of inner speech first, which is somehow then transformed. The question about whether inner speech is relevant for AVHs is not one of transformation. It is about whether relevant elements involved in the production of the inner speech experience are also involved in the production of some AVHs.11 It seems fairly clear, given all of the evidence in its support (some of which we have already seen, some of which we will touch upon below), that the answer to that is: yes. However, it is worth noting that on the PPF, all action will be couched in terms of active inference, namely, in terms of (hopefully) self-fulfilling sensory predictions (cf. Adams et al., 2013). How we are to think of inner speech within this framework is a fascinating direction for future inquiry. Given the plausible view that inner speech is developmentally derived from overt, private, speech (Berk, 1992; Vygotsky, 1934; Winsler, 2004), this could involve activating a prediction (as for any overt action), but over time somehow managing to dispense with many of the features of the action in question. This could involve learning to lower the weighting on the prediction error, or to make the prediction itself less demanding, or a combination of the two. Either way, prediction error could be minimised without having to go through with all aspects of the action (or, rather, its developmental precursor, private speech), although some aspects (e.g. subvocalisations, muscular activations) will remain. How this could lead to AVH is something that we will come to shortly. The key idea, roughly, is that inner speech is stripped down outer speech, but that this is built back up in AVH. However, it is not simply built back to overt speech, but to something different that involves the hearing of speech, but not its production. Accounting for three subtypes of AVH ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Jones (2010) expressed a concern that one model is unlikely to explain the various subtypes of AVH. He presented two subtypes. I will add a third but then gesture towards how the PPF provides a framework against which all three can be explained through different aetiological models. Inner speech hallucinations As we mentioned, the most popular model for understanding AVHs involves misattributed inner speech, where the misattribution results from an abnormality in self-monitoring. Although I want to abandon the framework within which SMTs tend to function, following Jones (2010), I think it is important to emphasise that there is great deal of empirical support for the view that a subset of AVHs implicate mechanisms involved in inner speech. There is support for the idea that inner speech might be relevant, other than the experiments we noted earlier in support of SMTs. Ford and Mathalon (2004) used EEG findings to suggest that during the production of healthy inner speech, corollary signals from speech production areas seem to suppress activation in speech processing areas. As Bentall and Varese (2013) nicely put it, “it might be said that when we talk to ourselves, the frontal language production areas of the brain tell the posterior areas not to bother listening” (2013, p.71). This attenuation should remind us of the support from bodily effects of monitoring, namely, the attenuation of sensation in (attempted) self-tickling. Both the attenuation pertaining to self-tickling and to inner speech can be accounted for within the PPF. Within the PPF, the fact that we cannot tickle ourselves is explained rather differently from how it is explained in SMTs that rely on motor commands and self-monitoring mechanisms. Brown and et al. (2013) have done work within the PPF that accounts for this in terms of a more general dampening of weighting on prediction error. The reason for this dampening is that movement involves implementing a currently false hypothesis (“I am moving”) in favour of the currently true hypothesis (“I am not moving”). The dampening swings the balance in favour of the false hypothesis, which then becomes true when the action is initiated.12 In inner speech, there is the prediction that you will speak, but the precision weighting will be turned right down, circumventing the need to minimise prediction error by actually overtly speaking (although, as we saw, some vestiges of the overt speech will remain). This will correspond to the suppression of activation in “speech processing areas” which are reported in both inner speech and in hearing oneself speak aloud (Ford & Mathalon, 2004). On the PPF there would be less activation in these areas, even when hearing oneself speak aloud, since there would be less prediction error (better predictions) than hearing someone else speak. The fundamental principle of prediction-error minimisation is the same whether the activity to be predicted is self-generated or not. One important difference, however, is that, if the activity is self-generated, the expected precision will be higher. Given this, only a relatively slight error in prediction would result in highly weighted prediction error, which would subsequently result in a more erroneous hypothesis to explain it away. This could account for cases where there are delusions of control and inner speech-based hallucinations in the absence of more outwardly directed perceptual abnormalities. Memory-based hallucinations Waters et al. (2006) distinguish their view from “standard cognitive accounts”, by which they mean SMTs, and which they call “single deficit accounts”, by proposing that: at least two cognitive deficits must be present to explain auditory hallucinations: (1) a fundamental deficit in intentional inhibition which leads to auditory mental representations intruding into consciousness in a manner that is beyond the control of the sufferer; and (2) a deficit in binding contextual cues, resulting in an inability to form a complete representation of the origins of mental events. This talk of “auditory mental representations”, “inhibition” and “binding of contextual cues” is nicely illustrative of how removed this is from the sort of framework that I am suggesting in this paper. However, this is not to say that they are incompatible, and cannot both be helpful. Using a couple of tasks (the Hayling Sentence Completion Task (Burgess & Shallice, 1996) and the Inhibition of Currently Irrelevant Memories Task (Schnider & Ptak, 1999)) Badcock, Waters, Maybery, and Michie (2005) demonstrated that hallucinating patients with diagnoses of schizophrenia have lower inhibition (more intrusion) of irrelevant associations and memories. The second deficit is needed to explain why these memories, which are failing to be inhibited, aren’t being recognised as memories. In research on episodic memory, there is a distinction between content (what is remembered) and context (e.g. temporal and “source” features of what is remembered). Waters et al. (2006) suggest that, although the content is remembered, there is a deficit in “context memory”, and, more specifically, “an impairment in combining contextual cues together to form an integrated representation of an event in memory”. Thus, the memory will present something that happened in the past, usually something traumatic, but it will not be experienced as a memory. It will be taken as a voice from the present. Let us look at what might be going on here through the lens of the PPF. In a case of an episode of healthy episodic memory, whether deliberate or unbidden, what is involved is the activation of relevant imagery, which within the PPF (in a way somewhat similar to our earlier discussion of inner speech) is not some kind of inner sensation, but a prediction. However, unlike the predictions at play in perception and action, this prediction is not answerable to – malleable in the light of – inputs. How is this “decoupling” from the environment achieved? What needs to happen is that the prediction need not to be “taken too seriously”, and so this will mean turning the precision right down on prediction error (this is similar to what happens in inner speech, minus the articulatory component). That way the “prediction” can be maintained even though the system is fully aware that it is not really predicting anything, that there is nothing really corresponding to the localised hypothesis generating the prediction. The general, overarching, hypothesis, is still that the subject is where she is (e.g. at a desk), but she is simply remembering something (e.g. that time when her father shouted at her). Now suppose that this prediction is activated, but there is a problem with keeping the weighting on prediction error low. This generates erroneous amounts of prediction error for which the brain has to adopt a hypothesis: this is perception, not memory. In a sense, then, if predictive processing, and, in particular the upwards and downwards modulation of precision on prediction error, goes badly awry then we get a blurring between mere imagery (as involved in episodic memory, imagination and, with relevant qualification, inner speech) and perception. The clever trick of decoupling, which enables us to remember episodically, or to imagine vividly, is disrupted. However, if this is correct, then we don’t need to appeal to integration of contextual cues (except insofar as there is a de facto failure to realise that this is not happening now). The recollective episode is not recognised by the subject as a memory, not because of some cognitive failure to remember the temporal or “source” context in which this happened, but because it doesn’t feel phenomenologically like remembering something: it feels like hearing something. Indeed, failure to remember temporal or source context cannot suffice to explain AVHs since there are cases where the voices are recognised by the subject as being exact replays of the past (12% in McCarthy-Jones et al., 2012), but they are still not experienced as memories, but as perceptual experiences. Indeed Wu’s Auditory Phenomenology Challenge applies to memory-based context- monitoring views almost as much as to inner-speech-based self-monitoring views. I say “almost as much” since perhaps (and we know very little about whether this is actually the case) episodic auditory memories have more phenomenological features in common with auditory perceptual experiences than episodes of inner speech (which arguably have an active, articulatory component). But, in any case, they certainly do not share all of them. So the challenge goes: How does one account for the auditory phenomenology of AVH if they are simply mis-identified memories? How does one account for the “transformation” from auditory-memory phenomenology to auditory-experience phenomenology? On the PPF account, since the conscious experience is a product of prediction-error minimisation, this question doesn’t arise. The initial event, the explanatorily relevant genesis of the experience, may be (as with the inner speech hallucinations) the same initial event as for a recollection in episodic memory, however, the erroneous prediction error causes the brain to hypothesise that something very different is going on. Hypervigilance Hallucinations The two varieties I have just elaborated on were introduced earlier in the paper through a consideration of Jones’s Varieties Challenge. This third kind has not yet been presented in this paper. Perhaps the first step towards an appreciation of the existence of hypervigilance hallucinations was made in a paper by Delespaul, DeVries, and Van Os (2002), who conducted an investigation into the contextual influences on AVHs. They found that the context where voice hearing is most likely to occur is either in the presence lots of people or alone. Building on this, Dodgson and Gordon (2009) suggested, on grounds of clinical case-studies, a kind of AVH called a “hypervigilance hallucination.” Here is a description of hypervigilance hallucinations from their paper: Michael was experiencing a series of stressors including break up with a girlfriend and exclusion from seeing their child, heavy street drug use, including amphetamine, and a pending court case for arson. He began to experience auditory hallucinations, hearing people calling him a “nonce”. Michael began to believe that people could read his thoughts, and that people thought he was a paedophile. The existence of hypervigilance hallucinations as a separate subtype was subsequently supported by Garwood, Dodgson, Bruce, and McCarthy-Jones (2013) who, based on a cluster analysis, showed that AVHs tend to occur when: attention is directed inward in quiet contexts attention is directed outward in noisy contexts They took (i) to suggest that inner speech hallucinations were occurring, and (ii) to suggest that hypervigilance hallucinations were occurring. Attention seems like a key factor here. So let’s view this notion of hypervigilance from within the PPF and its gloss on attention as precision-weighting. Let us contrast vigilance, with an illustratively opposing state: total calmness and lack of threat. Compare two cases of walking down a familiar woodland path at dusk. In the one case, you are walking home after an agreeable dinner party at a friend’s house, and you are feeling happy and relaxed. Because it is getting dark, you will find that your brain is forming rather vague hypotheses about what is out there, but since you are relaxed, and you have no reason to think that anything might be a threat, this is fine: precision-weighting (“attention”) is kept low. You get home safely having represented your environment in a vaguer, more coarse-grained way than you would have done in broad daylight, but this served your navigational purposes. Now suppose, instead, that you are walking home after having watched a horror movie. You are no longer relaxed, but in an emotional state of vigilance. As a result, you will be more likely to interpret something as looking like a person lurking amongst the trees. If this startles you, your attention will focus on that part of your visual field, and the precision-weighting of prediction error in that area will be turned right up. You may, due to this upwards-modulation, see that it is, in fact, only a tree trunk. Although as you look more closely, you get more incoming visual information suggesting the tree-trunk hypothesis, interoceptive information, namely, the fear that you feel, which is residual from the horror movie you saw, counts in favour of the lurking-person hypothesis (Pezzulo, 2013). You (or, perhaps more accurately, your nervous system) roughly think: “Why am I scared? There must be something to be scared of”, and then you get even more scared. In sharp contrast, in the first case, when you are happy and relaxed, the lurking-person hypothesis doesn’t even present itself. This fits nicely with hypervigilance hallucinations. In a state of hypervigilance, ambiguous inputs will be given a threatening interpretation; hypothesis selection will be biased towards something threatening, because this will also serve the purpose of explaining the interoceptive state (Seth, 2013). For the subject, this will give rise to experiences which misrepresent reality: from the ticking clock or the muffled sound of neighbours talking will emerge the experience of a voice telling the subject exactly what the subject, in his state of hypervigilance, is afraid of hearing (e.g. “nonce!”). Unlike the two other subtypes, hypervigilance hallucination have their genesis in the external world. One might of course claim that, since hypervigilance hallucinations do, strictly speaking, have some kind of worldly stimulus, then they are not, strictly speaking, hallucinations; they are rather illusions. I see this as a mainly terminological point, however, so will not go into it.13 A more important point is that, because of this, SMTs cannot account for hypervigilance hallucinations whereas the PPF can. The point is that, according to SMTs, normal monitoring exploits mechanisms (e.g. motor commands) linked to the production of self-generated events, and it is this that goes wrong and is misattributed. Thus, although it has a story to tell about how bodily actions, inner speech, and perhaps even episodic memories, can be badly monitored, and hence misattributed, it must remain silent about hypervigilance hallucinations, which are precisely not a case of a self-generated stimulus being misattributed, but rather an external stimulus being misinterpreted. To put it another way, unlike with SMTs, in the PPF, all stimuli, not just the self-generated ones, need to be predicted. Multiple models within one framework: personal history and attentional focus ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ I have gestured towards how these three subtypes could be accounted for within the PPF, but why would a given subject, at a given time, experience one subtype rather than another? This is an unavoidably difficult question. Although I cannot answer it in full, I think that two things are of particular explanatory relevance within the PPF. The first is the personal history of the subject; the second is the subject’s attentional focus, which will probably often be the accompaniment of certain moods, emotions and contexts. Personal history, for example, whether a subject underwent some form of traumatic abuse as a child, is likely to have an impact on whether a subject hears voices, and, as we have suggested, these voices are perhaps likely to be memory-based. Similarly, if a subject (like Michael, in Dodgson and Gordon’s case-study) has a history of amphetamine use, coupled with other so-called “stressors”, this might well lead to heightened risk of AVH, but probably not memory-based ones, since there is no single traumatic event to form the basis of the relevant episodic memory. Of course there could be interaction between the different factors that contribute to the (perhaps unrealistically clear) delineations of the subtypes. Thus childhood trauma and substance abuse could contribute to someone being at heightened risk of AVH. These might be, for example, hypervigilance hallucinations whose content could in some way be tied to past trauma. What, then, would make an individual more likely to experience AVHs that have an internal (viz. inner speech or memory) or an external (viz. hypervigilance) genesis? In answer to this question it may be useful to appeal to attentional focus. We saw that, within the PPF, following Hohwy (2012), we can think of attention as turning up the precision-weighting, and hence potential prediction error, on the attended stimulus. Attention can be directed outward, at incoming stimuli, or it can be directed inward, at thoughts, feelings and memories. Thus, if attention is directed outward in a state of anxious hypervigilance, there is more likely to be excessively weighted prediction error from the outside. This would correspond to hypervigilance hallucinations. As a result, traffic noise, or a ticking clock, or the sound of a group of people talking, might be embellished into the experience of a voice saying something. Alternatively, if attention is directed inwards, and the subject is socially isolated, ruminating in a state of shame or guilt, then there is more likely to be excessively weighted prediction error from the inside. This would correspond to inner speech or memory-based hallucinations. This nicely fits the findings by Garwood et al. (2013), concerning the prevalence of hallucinations in quiet and noisy contexts. It also allows that one patient, in different contexts and different moods, may experience different AVH subtypes. Indeed all three of the subtypes presented here may occur in one subject at different times. Of course, the picture is more complicated, but the framework is precisely the sort that is capable of accommodating such complications on a case-by- case basis.","I have suggested that we view AVHs through the lens of a predictive processing framework as a way of addressing two challenges, the first of which, in particular, plagues orthodox SMTs, the second of which is problematic for any single model of AVHs. The first challenge, of how the auditory phenomenology of AVHs can be accounted for, or at least rendered less perplexing, is addressed by the fact that the subject’s conscious percept is not determined by incoming stimuli, but by the brain’s best hypothesis as selected on the basis of how well it minimises prediction error. The question is not: Where does this erroneous stimulus come from? Rather the question is: Why does the brain adopt such a strange hypothesis? The answer to this question is (at least partially): erroneous precision weighting on prediction error. The second challenge, concerning the varieties of AVH, is addressed by realising that a new framework can accommodate elements of pre- existing models, but views the contribution of, for example, episodic memory or inner speech, in a way that is somewhat different. They are not “raw materials” to be transformed. Rather, mechanisms that are implicated in generating healthy experiences of episodic memory and inner speech, can generate erroneous prediction error that will result in erroneous hypotheses being adopted to minimise it. A major advancement of the PPF over SMTs is the realisation that, since prediction lies at the heart of cognition (broadly construed to include emotion and action), both internally and externally produced stimuli need to be predicted. SMTs had to be silent regarding external stimuli, since only self- generated stimuli can be monitored. Adopting the PPF, like adopting any framework, is only a first step. Not only is there, first of all, work to be done on better understanding predictive processing; for example, its neurobiological underpinnings and how it might go wrong at that level (see Corlett et al., 2010 for advancements in this direction). There is also exciting work to be done within the PPF with specific case-studies. It is a flexible framework within which we can better understand unique individuals, with different histories, problems and conditions.14"],["Acquiescence has been found to distort the psychometric quality of questionnaire data. Previous research has identified various determinants of acquiescence at both the individual and the country level. We aimed to synthesize the scattered body of knowledge by concurrently testing a multilevel model encompassing a set of presumed predictors of acquiescence. Based on a representative sample comprising almost 40,000 respondents from 20 European countries, we analyzed the effects of the country-level indicators economic wealth, corruption level, and collectivism and the individual-level indicators age, gender, educational attainment, and conservatism. Results revealed that 15% of the variance in acquiescence was due to country-level variations in corruption levels and collectivism. Differences among individuals within countries could be partially explained by conservatism and educational attainment. --------------------------------------------------------------------------------","Acquiescence—that is, the tendency to respond to descriptions of conceptually distinct attributes or attitudes with agreement/affirmation (agreement acquiescence) or disagreement/opposition (counter-acquiescence) regardless of their content—has been widely recognized as a threat to the validity of questionnaire-based data (e.g., Rammstedt, Goldberg, & Borg, 2010; Soto, John, Gosling, & Potter, 2008). Specifically, acquiescence can affect mean levels in item responding, thereby yielding misleading mean differences. For example, Van Vlimmeren, Moors, and Gelissen (2015) showed that country-level differences in trust in NATO differed substantially before and after controlling for acquiescence. Such effects of acquiescence on mean-level differences can occur if acquiescence differentially affects item responding across countries. Moreover, acquiescence may blur the intended factorial structure of a questionnaire by biasing item variances and covariances (Rammstedt et al., 2010). Finally, it has been shown that acquiescence can substantially bias the associations between personality items and behavioral criteria, thereby attenuating predictive validity (Danner, Aichholzer, & Rammstedt, 2015). Given the threats that acquiescence poses to the validity of questionnaire-based data, the overall aim of the study reported here was to summarize and integrate the available body of knowledge with regard to central socio-demographic and social indicators into one single conceptual model encompassing the presumed determinants of acquiescence. In what follows, we begin by summarizing the reported evidence on individual-level determinants and then address country-level predictors. Individual-level predictors of acquiescent responding ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Numerous studies have revealed that individuals differ systematically in their tendency to acquiesce. However, the empirical evidence is not univocal. While some studies have suggested that age is positively related to acquiescence (e.g., Meisenberg & Williams, 2008; Weijters, Geuens, & Schillewaert, 2010), others have failed to find evidence in support of this notion (e.g., Eid & Rauber, 2000). Findings with regard to possible effects of gender on acquiescent responding are even more heterogeneous. Some studies have suggested that women show, on average, a higher tendency toward acquiescent responding than men (e.g., Weijters et al., 2010), whereas others have found no gender effect (e.g., Marin, Gamba, & Marin, 1992). However, a broad consensus exists that educational attainment is a source of systematic differences in the tendency to acquiesce. Results of several studies have indicated that acquiescence appears to be more frequent among persons with a lower level of educational attainment (e.g., Narayan & Krosnick, 1996; Rammstedt et al., 2010; Rammstedt & Kemper, 2011). It has been suggested that persons with relatively low education have less clear self-concepts, smaller vocabularies, and less developed verbal comprehension skills than more highly educated persons. This may make them relatively uncertain when it comes to responding to questionnaire items, and may thus leave more room for the influence of systematic response biases (e.g., Goldberg, 1963). For some countries (e.g., Germany), this inverse effect of education on acquiescence has been widely replicated. Moreover, there is evidence to suggest that this effect can be replicated in several other countries, albeit with some exceptions (Danner et al., 2015; Rammstedt, Kemper, & Borg, 2013). However, results do not indicate a simple generalizability of the inverse effect of education on acquiescence across all countries (Meisenberg & Williams, 2008; Rammstedt et al., 2013). Rather, countries appear to differ systematically in this regard. Moreover, Smith and Fischer (2008) were able to show that individual-level interdependence, used as a proxy for a collectivistic cultural orientation, was positively related to acquiescence. Taken together, the literature on individual-level determinants of acquiescence partially supports the role of age, gender, level of educational attainment, and degree of conservatism in acquiescent responding. Country-level predictors of acquiescent responding ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In addition to individual differences in acquiescent responding, recent research has identified cross-national differences in the tendency to acquiesce, as reflected by mean- level differences (e.g., Javeline, 1999; Johnson, Kulesa, Cho, & Shavitt, 2005). For example, Van Herk, Poortinga, and Verhallen (2004) investigated acquiescent response tendencies in six European countries. The results revealed that respondents in the Mediterranean countries scored higher on acquiescence than those in the Northwestern European countries. A worldwide investigation of acquiescence was conducted by Meisenberg and Williams (2008). Based on the World Value Survey conducted in 80 countries, they showed that response styles were most prevalent in less developed countries and that—at the country level—acquiescence could best be explained by the country's corruption level. The authors interpreted their findings by suggesting that people who live in corrupt societies tend to be subservient to powerful others—a tendency that carries over into their survey responses. A similar effect was reported by Smith (2004), suggesting that acquiescence is significantly less pronounced in certain European countries than in countries with lower levels of economic development such as Panama, Nigeria, or the Philippines. In addition, there is a broad consensus that response styles are systematically related to cultural variables (Hofstede, 2001; Schwartz, 1994) and that they tend to be more pronounced in traditional cultures (Javeline, 1999). Specifically, several studies have suggested that the prevalence of acquiescence differs across countries and depends on cultural orientations. For example, a study by Johnson et al. (2005) indicated that collectivistic cultures were especially prone to acquiescent responding. The authors hypothesized that members of collectivistic nations experienced greater cultural pressure to acquiesce (Smith & Fischer, 2008). Support for this association was also provided by Harzing (2006), who investigated 26 countries from all major cultural clusters in the world. However, Grimm and Church (1999) could not confirm the effect of collectivism on acquiescence response style. In sum, the results of cross- national comparative research suggest that there are systematic differences between countries with regard to the mean tendency to acquiesce and that these differences are a function of the country's social and economic situation and its cultural orientations—in particular, the degree to which collectivistic values are endorsed. Thus, we expect that individual differences at the country level can be explained by these variables. Assessing acquiescent responding ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Even though the nature of, and the reasons for, acquiescence are still unclear, different approaches are used to investigate a person's tendency toward acquiescence. Some studies—especially those that use only positively keyed items—use the percentage or ratio of items agreed with (e.g. Harzing, 2006). For this approach, too, different methods of including and weighting the responses are employed across studies. Instead of using only positively keyed items, recent studies (e.g. Johnson et al., 2005, Rammstedt & Kemper, 2011; Rammstedt & Farmer, 2013; Rammstedt et al., 2010, 2013; Soto et al., 2008) have used, whenever possible, pairs of positively and negatively coded items assessing the same construct (e.g., Prefer to be with others and Like to be all by oneself). Persons with a high tendency toward acquiescence should have comparatively higher mean scores across these item pairs than those with a lower tendency to acquiesce. Even though some studies report only a weak consistency of acquiescence across different scales in general (e.g. Ferrando, Condon, & Chico, 2004), other studies report latent correlations r > 0.71 between acquiescence indicators of such pairs of negatively and positively keyed items (Danner et al., 2015). The present study ~~~~~~~~~~~~~~~~~ As summarized above, past research has yielded evidence of individual-level determinants (age, gender, and educational attainment) and country-level predictors (economic development, degree of collectivism, corruption level) of the tendency to acquiesce. However, previous studies have yielded inconsistent findings with regard to these characteristics. These inconsistencies may be due to the fact that most of these studies used highly selective samples that were not representative of the respective populations. In addition, to date no study has concurrently and systematically investigated these country-level and individual-level characteristics, taking into account the multilevel interrelationships between them. The present study aimed to fill this gap by investigating potential determinants of acquiescence by simultaneously analyzing the different country- level and individual-level characteristics and by relying on data that were representative of the population in 20 European countries. Specifically, we investigated the impact on the tendency toward acquiescent responding of the country-level characteristics economic wealth (GDP), corruption level, and level of collectivism in combination with the individual characteristics age, gender, educational level, and degree of conservatism. Individual-level and country-level predictors can be combined in one multilevel model where respondents (i) are nested within countries (j), and differences in acquiescence at the respondent level are modeled as acquiescence ij = β0j + β1(age) + β2(gender) + β3(education) + β4(conservatism) + εij, and differences at the country level are modeled as β0j = γ00 + γ01(wealth) + γ02(corruption) + γ03(collectivism) + υ0j.","The present analyses are based on data of the European Social Survey (ESS; www.europeansocialsurvey.org/data). The ESS is a cross-national survey that investigates changes in social structure, conditions, and attitudes in Europe. A key aim of the ESS is to implement high quality standards in its methodology. These high quality standards are especially relevant for the translation and adaptation of the questionnaires to guarantee comparability across the different countries. The survey has been conducted every two years since 2001. To test our conceptual multilevel model, we selected Round 1 of the ESS (European Social Survey Round 1 Data, 2002) as a data source because it included several contrasting item pairs that had already been used in an earlier study as an indicator for acquiescence (Johnson, Mohler, Harkness, & Braun, 2010). Samples The 2002 round of the ESS collected data from 22 European countries. For the present analyses, only those countries for which all relevant indicators were available were included. Therefore, Italy and Luxembourg were excluded from our analyses because conservatism was not assessed in these countries. A list of the countries included in our analyses can be found in Table 1. In each country, a sample representative of the population aged 15 years and over was drawn. Design weights provided in the data set were applied to adjust for different selection probabilities. The number of interviews conducted ranged between 1360 in the Czech Republic and 2919 in Germany, with a total sample size of 39,600 respondents. ESS questionnaires are administered as face-to-face interviews. Participation is voluntary and not usually incentivized—although some countries offer small incentives in order to increase participation. (For a detailed description of the samples and the assessment design, see European Social Survey Round 1 Data, 2002). Demographics Respondents' age was assessed in ESS Round 1 via the year of birth. Gender was assessed in ESS Round 1 with a dichotomous variable (1 = men, 2 = women). Levels of educational attainment were assessed in all countries on the basis of the national education system. Later, the ESS researchers harmonized these national data into one variable with reference to the International Standard Classification of Education (ISCED) 1997 levels, which yielded five categories: (1) “Less than lower secondary education (ISCED 0–1)”; (2) “Lower secondary education completed (ISCED 2)”; (3) “Upper secondary education completed (ISCED 3)”; (4) “Post-secondary, non-tertiary education completed (ISCED 4)”; (5) “Tertiary education completed (ISCED 5)”. Conservatism As a supplement to the ESS, the Portrait Values Questionnaire (PVQ; Schwartz, Melech, Lehmann, Burgess, & Harris, 2001) was administered, allowing, among other things, the assessment of the higher order value dimension conservatism, which overlaps strongly with the cultural orientation collectivism as defined by Hofstede (see, e.g., Schwartz & Ros, 1996). The conservatism dimension includes the lower order scales conformity, tradition, and security. It includes items such as “She/he believes that people should do what they're told” (conformity). All items are to be answered on a six-point scale ranging from not like me at all to very much like me. Cronbach's alpha ranged between 0.67 (Hungary) and 0.78 (Austria). Economic wealth As an indicator for the economic wealth of each country, the logarithm of its gross domestic product (GDP, adjusted for purchasing power) averaged across the years 1990 to 2002 (as suggested by Meisenberg & Williams, 2008), was used. Information was retrieved from the World Development Indicators of the World Bank (data.worldbank.org/data-catalog/world-development-indicators). Corruption level A measure of the corruption level per country was obtained based on averaged scores of Transparency International's Corruption Perceptions Index for the years 1999–2005 (transparency.org). Collectivism Country-level scores on an individualism-collectivism scale for the 20 countries investigated in the present study were taken from Hofstede (2001). Acquiescence As an indicator of acquiescent responding, an acquiescence measure was constructed on the basis of eight pairs of survey responses from the ESS questionnaire. These item pairs, which constituted sets of statements that clearly represented opposing opinions, stem from different question blocks within the ESS questionnaire and are intended to measure different constructs ranging from socio-political evaluations to attitudes toward migrants. A full list of these item pairs can be found in Table S1. All items were to be answered on a five-point Likert scale ranging from agree strongly to disagree strongly. In line with earlier research (Aichholzer, 2014; Billiet & McClendon, 2000; Danner et al., 2015; Rammstedt et al., 2010; Rammstedt & Kemper, 2011; Rammstedt & Farmer, 2013), we scored the acquiescence scale by averaging all 16 items. Given that eight items were positively keyed and eight items were negatively keyed, the resulting mean score does not reflect any construct variance but rather only acquiescence (and random measurement error). We estimated the reliability of the acquiescence index by subtracting each item from its opposing item and subsequently calculated Cronbach's alpha. Cronbach's alpha was 0.47 on average and varied between 0.39 (Denmark) and 0.58 (Greece).1 Table 1 provides an overview of the distribution of the different indicators investigated in the 20 countries.","In a first step, we investigated whether the 20 countries differed in their tendency toward acquiescent responding. We specified a baseline multilevel model with respondents (i) being nested within countries (j), acquiescenceij = β0j + εij that allowed acquiescence differences between countries, β0j = γ00 + υ0j. The model parameters were estimated using the residual maximum likelihood estimator implemented in SAS 9.3. The model revealed an intra-class correlation of ICC = 0.15, which indicates that 15% of the total acquiescence variance can be explained by systematic country differences. Hence, the 20 European countries investigated differed systematically in their mean tendency toward acquiescent responding. On average, the highest acquiescence scores were found in Greece, followed by Portugal and Poland; the lowest mean acquiescence scores were found in Norway, the Netherlands, and Denmark (see Table 1). In a second step, we concurrently investigated different possible determinants of these between-country and within-country differences in acquiescent responding: We conducted a second multilevel analysis with age, gender, educational level, and the degree of conservatism as individual factors: acquiescenceij = β0j + β1(age) + β2(gender) + β3(education) + β4(conservatism) + εij. At the country level, we investigated economic wealth, corruption level, and degree of collectivism: β0j = γ00 + γ01(wealth) + γ02(corruption) + γ03(collectivism) + υ0j. The intercorrelations between the individual-level predictor variables and between the country-level predictors are shown in Table S2. They indicate that GDP and collectivism, in particular, are moderately interrelated (− 0.42). The results of the second multilevel model fitted to the data are shown in Table 2. With the exception of GDP, all country-level predictors contributed significantly to explaining the tendency to acquiesce. The highest—and following Cohen (1992) a large — effect on acquiescence was estimated for the country's corruption level (η = 0.473): The higher the level of corruption, the greater was the tendency toward acquiescence in that country. A country's level of collectivism also had a large effect on acquiescence (η = 0.187): The more collectivistic the country, the higher was the overall tendency to acquiesce. In addition to the country's cultural orientation, the individual's degree of conservatism was the strongest predictor of acquiescence at the individual level (η = 0.045) although this effect was small to moderate in size. Further variance in acquiescence could be explained by the individual's level of education (η = 0.029). Thus, at the individual level, respondents with a low degree of conservatism and with a lower level of educational attainment exhibited a higher tendency toward acquiescent responding. Most likely due to the large sample size, the effect of age and gender was also statistically significant. However, age contributed only marginally (η = 0.004) to explaining the variance in our model, and gender did not explain a substantial portion of the variance at all. Overall, at respondent level, education, age, gender, and degree of conservatism explained 10% of acquiescent responding, and at country level, wealth, degree of collectivism, and corruption level explained 74% of acquiescent responding. To ensure that our results were not due to capitalizing on chance and that they were not biased by the correlation between GDP, collectivism, and conservatism at the country level, in particular, we conducted a set of robustness checks. First, we cross-validated the results of the multi-level analyses by randomly splitting each country sample into two subsamples and then replicated the analyses for both splits. The pattern of results was very similar for both splits, which suggests no effect of capitalization on chance (see Table S3). In addition, we investigated the biasing effect of the intercorrelations of our predictors at the country level. We therefore excluded each of these three variables in turn from the analyses. The pattern of the results was highly comparable to that originally found (see Table S4). We also investigated whether the association between acquiescence and age, gender, education, or conservatism differed between countries and specified the associations as random effects which did not change the pattern of results (all ∆ β < 0.002). Finally, we also investigated whether our measure of acquiescence and our conservatism measure were invariant across countries. In particular, we investigated configural invariance (same factor structure), metric invariance (same factor loadings), and scalar invariance (same factor loadings and same intercepts) with multi-group structural equation models. The metric invariance models showed acceptable model fits (RMSEA < 0.07, SRMR ≤ 0.05), whereas the scalar invariance model revealed a poor model fit for both scales (RMSEA > 0.11, SRMR > 0.10). Likewise, the change in RMSEA (∆ RMSEA < 0.015) and SRMR (∆ SRMR < 0.030) also suggest accepting metric invariance for both scales (Chen, 2007).2","The aim of the present study was to develop and test a comprehensive model encompassing individual-level determinants and country-level predictors of the tendency toward acquiescent responding in surveys. Our results corroborate systematic differences in acquiescence between countries: 15% of the variance in acquiescence could be explained by country differences, while the remaining 85% was due to individual variations among the respondents within countries. Thus, the European countries included in our study differed substantially in their tendency toward acquiescence. With our set of predictors we were able to explain three-quarters of the country-level differences. The most central indicator explaining these differences in acquiescence between countries was the corruption level of the respective countries, followed by differences in their respective cultural orientations. The higher the corruption level and the level of collectivism of a country, the greater was the overall tendency to acquiesce. By contrast, the economic wealth of the country had no significant additional effect on acquiescence. These results support Meisenberg and Williams' (2008) assumption that people living in corrupt societies are more subservient to powerful others. However, it has to be kept in mind that the present study investigated only European and thus comparatively wealthy countries. It needs, therefore, to be investigated in future studies if economic wealth has an additional effect on acquiescence based on a more heterogeneous set of countries. Individual indicators such as educational level and degree of conservatism differed systematically across the countries and thus contributed to the explanation of the between-country differences in this response style. Moreover, besides these differences between countries, our results revealed systematic differences between subpopulations: Some subpopulations within countries were more prone to acquiescence than others. Ten percent of these individual differences could be explained by our set of predictors. In line with previous research (e.g., Narayan & Krosnick, 1996; Rammstedt et al., 2010), our results revealed that—across all countries—individuals with lower levels of educational attainment and less conservative individuals were more prone to acquiescent responding than higher educated and more conservative individuals. In contrast, the effect reported in earlier studies (see Weijters et al., 2010) that females have, on average, a higher tendency toward acquiescence could not be replicated. Rather, we found a statistically significant but not substantial effect of gender that indicated a slightly greater tendency toward acquiescence on the part of males compared to females. Earlier findings with regard to the effects of age on acquiescence were inconsistent. While some studies suggested that older individuals have a greater tendency toward acquiescence (e.g., Meisenberg & Williams, 2008; Weijters et al., 2010), others identified no substantial effect of age (e.g., Eid & Rauber, 2000). Our results support these albeit somewhat contradictory findings insofar as we found a small statistically, but not practically, significant effect of age in the suggested direction. In sum, our results support the notion that the corruption level and the cultural orientation of a country, in particular, explain cross-national differences in acquiescent responding. Taking these indicators into account, the economic wealth of a country was not found to contribute to the explanation of cross-national differences. This contrasts with findings of previous studies that did not control for other indicators. Overall, our study substantially contributes to an understanding of cross-national differences in acquiescent responding. Systematic cross- national response artifacts, as detected in our study, may be misinterpreted as substantive differences across countries or cultures. Therefore, before drawing any conclusions from cross-national surveys, due caution should be taken to reduce acquiescence ex ante (e.g., by using balanced scales; Billiet & McClendon, 2000) or to control for acquiescence ex post (e.g., by using ipsatized data; Brown & Maydeu-Olivares, 2011). Typical areas of application for such ex-ante or post-hoc measures aimed at reducing acquiescence are all questionnaire-based data collections performed in multinational, multicultural, and multiregional contexts. Examples include, but are not limited to, international employee surveys (e.g., Johnson et al., 2005), cross-national media usage studies (e.g., Kuru & Pasek, 2016), and cross-cultural health surveys (Shavitt et al., 2016). In all these areas, variations in acquiescence carry the danger of being misinterpreted as differences in organizational commitment, media exposure habits, or health-related compliance. Therefore, mindfully considering and correcting for acquiescence should yield a more accurate picture about ‘true’ differences, and help to prevent deriving inappropriate subsequent measures."],["Background Depression in the postpartum period involves feelings of sadness, anxiety and irritability, and attenuated feelings of pleasure and comfort with the infant. Even mild- to- moderate symptoms of depression seem to have an impact on caregivers affective availability and contingent responsiveness. The aim of the present study was to investigate non-depressed and sub-clinically depressed mothers interest and affective expression during contingent and non-contingent face-to-face interaction with their infant. Methods The study utilized a double video (DV) set-up. The mother and the infant were presented with live real-time video sequences, which allowed for mutually responsive interaction between the mother and the infant (Live contingent sequences), or replay sequences where the interaction was set out of phase (Replay non-contingent sequences). The DV set-up consisted of five sequences: Live1-Replay1-Live2-Replay2-Live3. Based on their scores on the Edinburgh Postnatal Depression Scale (EPDS), the mothers were divided into a non-depressed and a sub-clinically depressed group (EPDS score ≥ 6). Results A three-way split-plot ANOVA showed that the sub-clinically depressed mothers displayed the same amount of positive and negative facial affect independent of the quality of the interaction with the infants. The non-depressed mothers displayed more positive facial affect during the non-contingent than the contingent interaction sequences, while there was no such effect for negative facial affect. Conclusions The results indicate that sub-clinically level depressive symptoms influence the mothers’ affective facial expression during early face-to-face interaction with their infants. One of the clinical implications is to consider even sub-clinical depressive symptoms as a risk factor for mother-infant relationship disturbances. --------------------------------------------------------------------------------","The caregiver displays multimodal emotional expressions during early social interaction with the infant, such as affective facial expression, gaze, physical touch, and vocalization (Stern, 1974). The affective quality and contingency are important components during early face-to-face interaction (Bornstein & Manian, 2013). The infant is, through her/his intersubjective awareness (Trevarthen, 2011), receptive to the caregiver’s affective expressions, as well as attuning his/her internal state for emotional sharing with the caregiver during early interactions (Trevarthen, 1980, 2001; Trevarthen & Aitken, 2001). An increasing body of research have shown how the infant is intuitively social, and how this shapes the affective relations with shared meaning between the caregiver and the infant (for a review see Trevarthen and Delafield-Butt, 2017). Postpartum depression, both in mothers (Goodman et al., 2011) and fathers (Ramchandani, Stein, Evans, & O'Connor, 2005), is associated with adverse child outcomes. Based on a meta-analysis of 193 studies Goodman et al. (2011) found that maternal depression was significantly related to higher levels of internalizing- and externalizing behavioral problems and general psychopathology in children, and that these associations were strongest for younger children. The mechanisms that explain these associations are not clear, but regulation of emotions in the early transactions between the mother and the infant is suggested as possible mechanisms (Field, Healy, Goldstein, & Guthertz, 1990). Maternal depression is associated with impairment in the mother’s capacity to synchronize with the infant’s positive affective state, as well as increased negative affect and irritation or intrusiveness during face-to-face interaction (Cohn, Campbell, Matias, & Hopkins, 1990). Based on findings from their meta-analytic review, Lovejoy and colleagues suggested that the parenting difficulties experienced by depressed mothers may not be associated with depression per se, but rather to more general distress associated with having psychological problems (Lovejoy, Graczyk, O'Hare, & Neuman, 2000). Thus, even sub-clinical levels of depression may influence the early mother-infant interaction negatively. It has been suggested that a mild-to-moderate depression has a more selective impact on maternal behavior during early mother-infant interaction, in terms of the mothers contingent responsiveness and affective availability (Hoffman & Drotar, 1991). In an earlier experimental perturbation study with sequences of real-life contingent interactions and with sequences where the interactions were set out of phase, we found that infants of sub- clinically depressed mothers were less sensitive to the interruptions of the contingency than were infants of non-depressed mothers (Skotheim et al., 2013). A possible explanation of this finding is that the infants of the sub-clinically depressed mothers were more accustomed to non-contingent interaction. The purpose of the present study was to investigate if even sub-clinical levels of depression may impact maternal behavior and affective responsiveness during sequences of mutually face-to-face interaction with their infants and when the interaction are set out of phase, using the double video (DV) set-up. Field et al. (2005) found in a similar perturbation DV design that clinically depressed mothers displayed the same amount of positive affect both before and after the interaction with their infant was set out of phase, while the non-depressed mothers displayed reduced positive affect after the perturbation of the interaction. Nadel et al. (2005) found, also using the same design, that both the non-depressed and the clinically depressed mothers evidenced a similar decrease in positive affect after the perturbation of the interaction. However in both studies (Field et al., 2005; Nadel et al., 2005), the mothers were told in advance about of the perturbation, which may have influenced the interactions with their infants. In addition, the clinically depressed mothers in Field et al’s (2005) study displayed less positive affect than the non-depressed mothers already from the start of the experiment, which suggests that the results may be attributable to an effect of maternal depression and not the manipulation of contingency. In the present study the mothers focus of gaze and facial expression of affect were examined in groups of sub- clinically depressed and non-depressed mothers during face-to-face interaction with their infants. This was examined, drawing on the same sample of sub-clinically and non-depressed mothers as in Skotheim et al. (2013). In the present study, we examined the mothers’ responses to live (contingent) and non-contingent interaction with their infants. More specifically, we investigated if sub-clinically depressed mothers would change their focus of gaze and facial expression of affect in response to shifts between contingent interaction and non-contingent interaction in a DV-experiment.","Fifty-one mothers and their three months old infants took part in this study. The participants were drawn from a prospective longitudinal population-based study on nutrition, mental health and infant development (Markhus et al., 2013) where all the mothers (n = 105) were asked to participate in the current study. Informed consent was obtained from 51 mothers. Twelve mother-infant dyads had to be excluded from the analyses, either because of excessive infant crying (n = 5), or because the mother and the infant did not manage to establish mutual gaze contact at the start of the experiment (n = 7). Thus the final sample comprised 39 mothers (mean age = 32 years, range 21–41 years) and their three months old infants (mean age = 13 weeks, range 9–20 weeks; Boys = 16 (41%). None of the mothers fulfilled the DSM-IV criteria for major depression, based on the clinical interview with the Norwegian version of the MINI International Neuropsychiatric Interview (MINI; Leiknes, Malt, & Malt, 2007; Sheehan et al., 1998). The mothers were divided into two groups, based on their levels of self-reported depressive symptoms on a validated Norwegian version of the Edinburgh Postnatal Depression Scale (EPDS; (Berle, Aarre, Mykletun, Dahl, & Holsten, 2003), using an EPDS total cut-off score ≥6 (Nadel et al., 2005). Twenty-four mothers were categorized as non-depressed (mean EPDS score 3.26, range 0–5) and 15 mothers were categorized as sub-clinically depressed (mean score of 8.26, range 6–13). There were no significant differences between the two groups, neither for infant age or gender nor maternal marital status, parity, educational level or yearly income. All the mothers received written information about the study prior to giving their written consent on behalf of themselves and their infant. The study was approved by the Regional Committee for Medical and Health Research Ethics and the Norwegian Social Science Data Services. The procedures followed were in accordance with the Helsinki Declaration of 1975 (revised in 2008). The double video set-up The double video set-up and equipment were identical to the Skotheim et al. (2013) study. The mother and the infant were placed in two separate sound-proof rooms, where they were seated in front of a one-way mirror that enabled them to see a life size image of each other’s faces/upper bodies (see Fig. 1). Each one-way mirror was positioned in a 450 gradient in front of both the mother and the infant, with a LCD-monitor lying flat below. A SONY HDR-HC7E digital video camera was placed behind each of the two one-way mirrors, allowing the infant and the mother to have direct eye contact with each other while they were looking straight into the camera. The cameras in each room recorded the video stream with 25 frames-per-second. The video was transmitted over firewire to the two computers PC1 and PC2 respectively and further in an uncompressed format over a high speed wired network to a control-PC (PC3). During piloting the DV set-up, calculation of the length between each sequence showed a pause from 150 to 250 ms. The partners could not be aware of this pause when interacting together. The mothers and the infants vocalizations were captured by microphones integrated in the video cameras and was embedded in the video stream. The mothers used earplugs to avoid audio feedback. At PC3, the researcher (SS) could monitor the video captured by the two computers. This design enabled the researcher to monitor the mother- infant interaction and start and stop the experiment. All recorded videos were stored on PC1 and PC 2 during the experiment, and then copied to an external hard disk for further analyses. As in our previous studies (Braarud & Stormark, 2006, 2008; Skotheim et al., 2013; Stormark & Braarud, 2004), the DV experiment comprised five sequences in a fixed Live1-Replay1-Live2-Replay2-Live3 order. All five sequences consisted of a 60 s interval, in which video-recordings were made of the mother’s and the infant’s behavior. The three live sequences all consisted of uninterrupted face-to-face interaction. In the first replay sequence, the mother saw and heard the infant live while the infant was responding to a replay of the mother from the previous live sequence. In the second replay, the mother received a replay of the infant’s behavior from the first live sequence. These modifications enabled us to control for possible memory bias, since the replay sequences represented two different types of non-contingency. The mothers in the present study were naive with respect to the experimental manipulation, in contrast to Nadel et al. (2005) and Field et al. (2005).","Before entering the study, the mother’s level of depressive symptoms had been assessed with the Norwegian version of The Edinburgh Postnatal Depression Scale (EPDS; Berle et al., 2003; Cox et al., 1987), as part of their visit to the well-baby clinic. The EPDS is a 10-item self-report scale that assesses current (last week) postpartum depressive symptomatology. Each item is rated on a 4-point scale (0–3), yielding a total score ranging from 0 to 30, with higher scores indicating increased symptomatology of postpartum depression. A Norwegian (Leiknes et al., 2007) version of MINI (Sheehan et al., 1998) was administrated as part of the study to determine if any mothers fulfilled axis I diagnoses in the Diagnostic and Statistical Manual of Mental Disorders (DSM IV). Procedure When arriving at the laboratory, the mothers were first interviewed with the Norwegian version of MINI 5.0.0. After the interview, the mothers were explained that the purpose of the study was to investigate face-to-face interaction between mothers and infants, using a closed-circuit video and computer system. The mothers were told to engage with their infant as during normal interaction. When the mother and the infant had adjusted to the laboratory setting, they were seated in each of the rooms where they could see and hear each other through PC1 or PC2, respectively. When the experimenter (SS) judged that the mother and the infant had established a stable eye-to-eye contact, the experiment started. Coding and reliability ~~~~~~~~~~~~~~~~~~~~~~ The mothers’ gaze foci and facial affective expression were scored independently, using The Observer XT 9. The mutually exclusive categories were Focus of gaze (looking at partner, looking away from partner, indefinite focus; see Stormark and Braarud (2004)). For Facial expression, the mutual exclusive categories were positive affect (smiling; simultaneous raising of lip corners), negative affect (frowning; lowering of eyebrows, eye squeeze and deepening of the nasolabial furrow) and neutral facial expression (in accordance with Nadel et al. (2005)). Two independent coders were trained to score the mothers’ focus of gaze and facial expression of affect. They were blind to which sequence they were scoring, and blind to the clinical assessment of the individual mother. Calculation of inter-rater reliability based on eight mothers gave a Cohen’s kappa for gaze focus (k = 0.97) and facial affective expression (k = 0.67). Statistics ~~~~~~~~~~ The data on maternal gaze focus was subjected to a separate two-way split-plot ANOVA involving Group (non-depressed mothers vs. sub-clinically depressed mothers) x Sequence (Live1, Replay1, Live2, Replay2, Live3). The data on positive and negative facial affective expression was subjected to separate three-way split-plot ANOVAs involving Group (non-depressed mothers vs. sub-clinically depressed mothers) x Affect (positive vs negative affect) x Sequence (Live1, Replay1, Live2, Replay2, Live3). To investigate if the mothers behavior were related to shifts between contingent and non-contingent infant behavior, rather than general fatigue, the number of live sequences had to exceed the number of replay sequences. As a result, the research question could not be fully evaluated in an omnibus ANOVA. Therefore the dataset was examined in a priori comparisons (planned comparisons of contrast variables (Hays, 1988), so the live and replay sequences could be weighted equally. The statistical analyses were carried out with STATISTICA 64, StatSoft, Inc. Mothers’ gaze focus ~~~~~~~~~~~~~~~~~~~ There was a significant main effect of gaze focus (F(1.37) = 23618, p< 0.001), independent of Sequence and Group, which reflects that the mothers looked at their infants almost exclusively during the experiment (see Fig. 2). There were no main effect of Sequence and no interaction effect between Group and Sequence. However, planned comparisons revealed a non-significant (p = 0.06) linear trend reflecting an increased maternal gaze focus at the infants during the experiment in the sub-clinically depressed mothers. Mothers’ facial expression of positive and negative affect ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There was no main effect of Group, but there was a significant main effect of Affect, reflecting that the mothers, independent of group and sequence displayed significantly more positive than negative affect during the experiment (F(1.37) = 409.98, p < 0.0001), see Fig. 3 and Fig. 4. There was also a main effect of Sequence (F(4148) = 5.66, p < 0.001), and an interaction effect of Affect and Sequence (F(4148) = 6.55, p < 0.001), reflecting that there was overall more positive affect during the Replay than the Live sequences (F(4, 144) = 4.64, p < 0.05), while there was no such sequential effect of negative affect. Planned comparisons revealed further that there was a linear decrease in positive affect across sequences, in both the sub-clinically depressed (F(1, 37) = 9.75, p < 0.01), and the non-depressed mothers (F(1, 37) = 4.72, p < 0.05). There were no significant differences for negative affect, neither for groups nor across sequences. Further planned comparisons revealed that the non-depressed mothers displayed significantly more positive affect during the Replay sequences (Replay 1–2 averages) than during the Live sequences (Live1-3 averages; F(1, 37) = 5.28, p < 0.05), while there was no such difference in the sub-clinically depressed mothers. There were no differences in positive and negative affect between the two groups, neither for the Live nor the Replay sequences.","The main findings in this study were first that the sub-clinically depressed mothers did not discriminate with respect to the manipulation of the contingency, but expressed the same amount of facial affect during contingent and non-contingent interactions with their infants. The non-depressed mothers evidenced significantly higher amount of positive affect during the replay sequences compared to the live sequences. Both groups of mothers showed a high amount of gaze focus at their infant, and both groups of mothers showed a linear decrease in positive affect during the experiment. In line with earlier studies that focus on the interactive behavior of depressed mothers (Cohn et al., 1990; Field, 1984), we could expect that the sub-clinically depressed mothers, compared to the non- depressed mothers, would express less positive and more negative affect during the experiment. Instead, our results show that the sub-clinically depressed mothers expressed the same amount of positive affect as the non-depressed mothers, but without discriminating between the sequences that allowed for contingent interaction and the sequences that did not. In fact there were no overall differences in amount of positive affect between non-depressed and sub-clinically depressed mothers, which accords with Nadel et al.’s (2005), but not with Field et al.’s (2005) study. However, while the sub- clinically depressed mothers in the present study did not discriminate between contingent and non-contingent interactions with their infants, the non-depressed mothers were affected by the subtle experiential differences between live- and replay sequences. While earlier studies have shown a decrease in infant directed speech in the replay-sequences (Braarud & Stormark, 2008; Murray & Trevarthen, 1986), we found more positive affect in the replay than the live sequences during this double video experiment. Pechtel et al. (2013) suggested that sensitive and attuned mothers would be able to continue with maternal engagement and sensitive regulation of the infant, even in emotional arousing caregiving situations evoked by infant distress. The non-depressed mothers behavior responses could therefore be related to their infants affective display during the replay sequence, and not necessarily to the lack of contingent responses from the infants. Thus, when the non-depressed mothers experienced a reduction of smiling from their infants in Replay 1, the mothers could respond with sensitive regulation in an attempt to regulate the infants toward a more positive state (Feldman, 2003). However, since the replay- sequences impede on establishing mutual synchrony, one may speculate if the non-depressed mothers decrease in smiling in the following live-sequence could be understood as a response to the failure of regulating their infants. Replay2 would represent yet another and equal perturbation as the mothers received a replay of the infants behavior from Live1. The mothers saw and heard the infants behavior from what was an optimal social situation, but where the contingency was interrupted. An alternative and more general explanation could be that the mothers are used to a greater degree of variability in infant responsiveness, and need more time before they detect the perturbation of contingency in the replay sequences (Henning & Striano, 2011). Results from earlier double video studies (Field et al., 2005; Nadel et al., 2005; Stormark & Braarud, 2004) accords with the present finding that both groups of mothers in the present study expressed high amount of gaze at their infants independent of the quality of contingency. However, mothers infant-directed speech (content, stylistic features and the fundamental frequency (F0)) seems to be modulated according to changes in social contingency during mother- infant face-to-face interactions (Braarud & Stormark, 2008; Murray & Trevarthen, 1986). If one considers that the internal processes of reactivity and regulation in the caregiving system are related to the infant’s affective cues, one could speculate whether the mothers’ affective facial responses, gaze focus and infant-directed speech have different meaning during early social interactions. Van Egeren et al. (2001) found that maternal and infant vocalization was the primary response in their microanalyses of mother-infant responsiveness when the infants were 4 months old. The results accentuate that maternal sub-clinical depression influence sensitivity to social contingency and affective behavior during mother-infant face-to-face interaction, which is in line with Hoffman and Drotar (1991). The unexpected high amount of positive affect in the group of sub-clinically depressed mothers may be attributed to the experimental setting. Even if the mother and the infant are physically separated during the experiment, the DV set-up optimizes face- to-face interaction, without any disturbances found in natural circumstances. This interpretation is supported by the fact that the mothers had nearly 100% amount of gaze focus at their infants in all sequences. So, even if the mothers and the infants in the sub-clinically depressed group may have more dyadic difficulties in their daily interactions (Reck et al., 2004), the DV set-up could have been a positive interactional experience for the mothers. However, the lack of sensitivity to social contingency found in this study highlights the selective impact of maternal sub-clinical depressive symptoms on early face-to-face interaction. One limitation of the study concerns the DV design and ecological validity. Three months old infants and their caregivers spend most of their awakening hours in physically close contact during normal situations (Trevarthen, 1979). In this study, the infant and the mother sat in two separate rooms where they can interact via PCs, which might be experienced as artificial. The linear decrease in positive affect for both groups of mothers could be understood as a sign of minor, but increasing distress during the experimental procedure. Still, even in this artificial environment, the experiment also demonstrated three-month-olds interest in face-to-face communication, which also could explain why only five recordings had to be stopped because of excessive crying. We also acknowledge that there are factors that affect the mother-infant interaction and dyadic regulation that are not addressed in this study. Infant temperament is one such factor, even though we have earlier reported that there was no significant difference in temperament between the two groups of infants (Skotheim et al., 2013).","To sum up the sub-clinically depressed mothers displayed the same amount of positive affective facial expression independently of the quality of the interactions with their infants, while the non-depressed mothers displayed a decrease in positive affect after the sequences where the interaction with their infants were set out of phase. There was a linear decline in positive affect across the five sequences for both groups of mothers, which suggests that both groups experienced an increase in distress during the experiment. These results strengthen the findings in Skotheim et al.’s (2013) study, where the infants of the sub-clinically depressed mothers were less sensitive to the quality of the affective interaction with the mothers than the infants of the non-depressed mothers were. Both studies illustrate the importance of the mother’s well-being for early mother-infant interactions and optimal infant development."],["Newborns produce spontaneous movements during sleep that are functionally important for their future development. This nuance has been previously studied using animal models and more recently using movement data from sleep resting-state fMRI (rs-fMRI)scans. Age-related trajectory of statistical features of spontaneous movements of the head is under-examined. This study quantitatively mapped a developmental trajectory of spontaneous head movements during an rs-fMRI scan acquired during natural sleep in 91 datasets from healthy children from ∼birth to 3 years old, using the Open Science Infancy Research upcycling protocol. The youngest participants studied, 2–3 week-old neonates, showed increased noise-to-signal levels as well as lower symmetry features of their movements; noise-to-signal levels were attenuated and symmetry was increased in the older infants and toddlers (all Spearman's rank-order correlations, P < 0.05). Thus, statistical features of spontaneous head movements become more symmetrical and less noisy from birth to ∼3 years in children. Because spontaneous movements during sleep in early life may trigger new neuronal activity in the cortex, the key outstanding question for in vivo, non-invasive neuroimaging studies in young children is not “How can we correct head movement better?” but rather: How can we represent all important sources of neuronal activity that shape functional connections in the still-developing human central nervous system? --------------------------------------------------------------------------------","A complex process helps human newborns experience a structured world. This mechanism is supported in part by the neuronal substrates with early-onset myelination, innately-driven sensitivities as well as pre- and post-natal learning playing an important role in development. Currently, little is known about the character and the contribution to this process by the spontaneous, endogenous head movements produced during rest or sleep, even though newborns naturally spend the majority of their time sleeping (Roffwarg, Muzio, & Dement, 1966)—an important knowledge gap in our understanding about how the central nervous system (CNS) develops in humans during early critical time period (early postnatal life). Importantly, diverse types of motility occur throughout sleep, across the different state epochs according to Prechtl (1974). Movements, including startles and mouthing, occur during quiet sleep, QS (corresponding to the behavioral state 1) as well as active sleep, AS (state 2), and twitches occur particularly during AS sleep (see, e.g. p. 209 in Prechtl, 1974). Developmental or maturational character of spontaneous, endogenous movements during sleep has not been mapped in human infants. An insight from biology is that spontaneous, endogenous movements are a source of cortical stimulation in early life and trigger, or are associated with, neuronal activity (Khazipov et al., 2004). In anesthetized neonatal rats, the majority of the bursts in the primary somatosensory cortex are “associated with overt movement, including isolated muscle twitches, limb and whole- body jerks, crawling and sucking” (Khazipov et al., 2004). Recent work has focused on the role of spontaneous movements during sleep for the developing CNS. Subtle, abundant, and functionally highly important movements during sleep have been detected and described in sleeping infant rats (Blumberg et al., 2015; Tiriac, Del Rio-Bermudez, & Blumberg, 2014). It has been argued that twitches or jerky spontaneous movements in newborn rats, measured during REM sleep (Rapid Eye Movement; corresponds to AS stage), lack corollary discharge during sleep (compared to wakefulness), and as such, these movements would be expected to affect the sensory system, generating concurrent neuronal bursts, and aiding in building environment-appropriate sensorimotor maps for the developing infant (Blumberg et al., 2015; Tiriac & Blumberg, 2016; Tiriac et al., 2014). Neuronal spiking increases after twitches (Sokoloff, Uitermarkt, & Blumberg, 2015). It can be briefly noted that newborn (P0, Postnatal day 0) or neonatal rats’ maturational stage corresponds to the second half of gestation in humans, including the pre-term and near-term periods (Colonnese & Khazipov, 2012). Generally, P10 in rodents corresponds to the third trimester in humans and P13 corresponds to term, although variation in the maturational correspondence exists for different brain regions. For example, term in humans is P5–7 for the hippocampus and P12–14 for the visual cortex (Colonnese & Khazipov, 2012). Tiriac et al.’s studies included P8–P10 infant rats, a range which approximately corresponds to near-term or term in humans. Intriguingly, because new neuronal events may accompany spontaneous, endogenous movements (whether such movements occur during quiet or active sleep in human neonates), these neuronal events may assist with synapse formation and help build cortical connectivity maps in the developing human CNS, in particular, the sensorimotor maps. None of the non-invasive, in vivo functional neuroimaging approaches in use today with human neonates, infants, and toddlers consider the importance of endogenous, self-generated spontaneous movements while resting or sleeping for the developing mind and brain. Why is this significant? The relevant caveat is that head or body movements have a deleterious impact on the homogeneity of the magnetic field, and researchers use motion correction strategies to correct high-movement volumes acquired during MRI scans, albeit with downstream consequences for functional connectivity metrics and inference (Denisova, 2018). The problem is that such a motion correction strategy is incompatible with extensive knowledge on the role of sensory inputs from spontaneous movements that are triggers for neuronal bursts and synapse formation in early development (Khazipov et al., 2004), the drivers of hemodynamics and hence functional connectivity measures that developmental researchers seek to investigate. That is, functional connectivity studies, which acquire data during natural sleep or rest in human neonates, infants, and toddlers, set out to investigate correlations in subtle hemodynamic fluctuations between different regions, but these studies suppress data acquired while subjects are moving before applying statistical metrics. In summary, despite knowledge from animal work that neuronal events of substantial importance to the maturing central nervous system occur due to spontaneous movements, the potential functional role of these movements and of the new neuronal events they generate, at least very early in infancy, is not given due consideration in MRI research with human children. For instance, spontaneous movements in humans could contribute significantly toward nascent functional connectivity. Do such movements play a generative, rather than auxiliary or inconsequential, role in the developing human CNS? Considering this problem in the context of biology while using real rs-fMRI datasets permits us to meaningfully pose questions when studying how the mind and brain develop in humans. Biology also provides a pertinent framework to seriously evaluate and to revolutionize the methods used to answer questions about the neurobiological underpinnings of cognitive development in children. As a first step toward this goal, the aim of this study is investigate how one important variable for developmental studies, the age of the child, has an effect on statistical features of head movements while children between birth and 3 years of age underwent resting-state fMRI (rs-fMRI) scan during natural sleep. In humans, recent research has detected informative differences in infants’ spontaneous head movements during a resting-state functional Magnetic Imaging Resonance (rs-fMRI) scan acquired during natural sleep. Researchers studied infants from two cohorts: at high vs low familial risk for developing Autism Spectrum Disorder (ASD; high risk, HR is defined by virtue of having a biological sibling with a confirmed ASD diagnosis) (Denisova & Zhao, 2017). In particular, Denisova and Zhao (2017) found that relative to low risk (LR) infants, HR infants as young as 1–2 months old have more random and less symmetric features of head movements during an rs-fMRI scan obtained during natural sleep. Atypical head movements of the HR cohort at 1–2 months stayed abnormal at 9–10 months and these features were associated with flatter or delayed future early learning developmental trajectory (Denisova & Zhao, 2017). It remains unknown whether the statistical character of subtle movements of the head during sleep is expected to normalize with age during the first few years of life in normal, healthy children at a low genetic risk for atypical development, since the study by Denisova and Zhao (2017) considered neither participants at the intervening ages, nor neonates and older toddlers. One question to ask is: do spontaneous head movements during sleep in humans change as children age? Here it may be noted that the capacity to respond flexibly, whether or not the behavior is accompanied by higher noise (here, “higher noise” refers to higher variance relative to mean) levels, at least in early infancy, seems paramount for later normative development. Behavior at either end of the ‘normative’ spectrum, that is, either excesses (for example, high levels of motor noise: e.g., (Therrien, Wolpert, & Bastian, 2018)) or conversely, a lack thereof (for example, pertaining to culture-specific rearing practices, e.g., (Adolph, Hoch, & Cole, 2018)), may be detrimental for normative functioning. Further, investigators have argued for the importance of variability for normative human behaviors (Latash, 2016). In early life, this issue is complex and depends on whether movements are considered during wakefulness vs. sleep, in infants at low vs. high risk for autism, and on the age of the child. For example, the study by Denisova and Zhao (2017) found that infants’ head movements differ depending on the experimental condition during the neuroimaging scan, such as whether they were sleeping or awake, while native language speech was presented to them, and on the age of assessment (Denisova & Zhao, 2017). The goal of this study is to investigate the developmental pattern of spontaneous head movements during sleep in a normative sample of children between birth and 3 years of age, as children underwent an rs-fMRI scan during natural sleep. Quantitative characterization of an age-related trajectory of spontaneous head movements during sleep in a sample of healthy children can inform our understanding of sensorimotor development in humans and would provide context for studies of young children who are not developing normally. In addition, given that researchers often acquire neuroimaging data in young children during sleep, a characterization of a developmental trajectory can inform future brain imaging work in children. Thus, this work quantitatively mapped statistical features of head movements in neonates, infants, and toddlers in a new cohort of 86 children, 5 of whom provided longitudinal data, for a total of 91 datasets from National Database for Autism Research (NDAR). Of note, researchers can distinguish between genetic and environmental contributors to development in humans (McGue & Bouchard, 1998). The current work excludes datasets from subjects at known genetic risk for atypical development; datasets include those from healthy children, who were deemed to be free from severe medical concerns and who are free of genetic risk for atypical development. Assessments at all study-sites used a strictly standardized protocol required for MRI acquisitions, such that all children lay supine on the scanner bed, with head and body movements minimized using cushioning and blankets during the 5-minute session. The main questions for the current study are as follows: whether spontaneous head movements during sleep in young healthy children are (i) qualitatively and quantitatively different among individual children even when the head is relatively stable and well-cushioned during a sleep rs-fMRI scan, (ii) whether there is a cross-sectional maturational pattern from birth to 3 years old, such that neonates show the noisiest and least symmetrical movements relative to older children, and (iii) whether the maturational pattern is robust, that is, somewhat independent of the absolute magnitudes of head movements. Datasets ~~~~~~~~ Data used in the preparation of this study were obtained from the NIH-supported National Database for Autism Research (NDAR). NDAR is a collaborative informatics system created by the National Institutes of Health to provide a national resource to support and accelerate research in autism. Dataset identifier(s) (along with Submitters): NDARCOL0002329 (Scott Holland & Jennifer Vannest), NDARCOL0001890 (Claudia Buss & Pathik Wadhwa), and NDARCOL0002340 (Mary Phillips). As NDAR is part of the NDA (NIMH Data Archive), available child datasets include those from research studies that do not solely focus on children at risk for atypical development, such as autism. The inclusion criteria for selecting specific collections (#2329, 1890, and 2340) that contributed datasets in the current study consisted of (i) availability of at least 1 resting-state functional Magnetic Resonance Imaging (rs-fMRI) scan, with data in original format with no motion correction or other post-scan processing applied, (ii) a wide age of infants, ranging from birth up to approximately 3 years old (36 months) during the scan, and (iii) healthy infants, in particular those who were not known to be at high familial risk for autism or atypical development. These inclusion criteria yielded the three study-sites above. Study-sites were excluded if scans did not have comparable temporal resolution relative to the current datasets (2 s; please see below for details), and/or if participants were at high risk for atypical neurodevelopment. Specifically, researchers can distinguish between genetic and environmental risk factors as potentially affecting behavior (McGue & Bouchard, 1998). The current work excludes datasets known to belong to infants at a genetic risk for atypical development; no datasets known to be from infants at high familial risk for autism are part of the current sample. Additional collection-specific inclusion and exclusion criteria are listed below. All data are de-identified in compliance with U.S. Health Insurance Portability and Accountability Act (HIPAA) guidelines. Signed, written informed parental consent was obtained by original study investigators in accordance with U.S. 45 CFR 46 and Declaration of Helsinki for participation and study procedures, including neuroimaging protocols, were approved by the IRB of each institution. Analyses of these de-identified data were reviewed and approved by, and a waiver of informed consent was obtained from, the Institutional Review Board of Columbia University Medical Center. The current study includes datasets from the May 2017 NDAR data release. The total sample included 91 datasets (including N = 86 unique participants, with N = 5 who were scanned longitudinally for a total of 2 times). This sample includes neonates, infants, and toddlers. All children underwent an rs-fMRI scan during natural sleep (no sedation was used). For N = 86 unique participants, the Male to Female ratio (M/F) is 40M/46F. Overall for this dataset (N = 91), the age ranged between 2 and 3 weeks and 35 months. Note that NDAR provides ages in months; for #1890, ages were available in weeks for participants less than 1 month old. Birthweight data were available for 63 (including the five children who contributed longitudinal datasets) out of 86 children. The average birthweight (mean and standard deviation) was 3282.69 (442.78) grams. Of these children, only four had birthweights below 2500 g (these infants’ birthweights are: 2410 g, 2410 g, 2380 g, and 2350 g). Additional information on inclusion and exclusion criteria at the three study-sites ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Considering datasets by each of the three study-sites (collections), datasets from collection #2329 comprised N = 35 participants (14M/21F), infants and toddlers (1–36 months; N = 5 with longitudinal scans). For N = 5 participants who were scanned longitudinally, their ages at first relative to second scan were: 3 months/14 months (Male), 13 mo/27 mo (Male), 17 mo/31 mo (Female), 17 mo/31 mo (Female), 20 mo/32 mo (Female). Thus, together with the longitudinal participants, there were N = 40 datasets (16M/24F) from this site. All were required to meet inclusion criteria of low-risk, healthy pregnancy and to have no siblings who were diagnosed with ASD, following recruitment procedures of the Cincinnati MR Imaging of Neurodevelopment (C-MIND) study. All participants underwent a single resting-state scan. Datasets from collection #1890 dataset comprised N = 22 healthy, typically developing neonates and infants (10M/12F) meeting inclusion criteria of a low-risk pregnancy as per original study protocol during recruitment at University of California, Irvine. Out of these N = 22, N = 14 participants were 2–3 weeks old (7M/7F), N = 4 were 1 month old (2M/2F) and N = 4 were 11–13 months old (1M/3F). All participants underwent a single resting-state scan. Datasets from collection #2340 comprised N = 29 participants (2–5 mo-old infants, 16M/13F) who underwent 2 resting- state scans each during a single visit to Pitt. All infants met inclusion criteria for normal pregnancy, but infants were recruited into the Pitt study if they were at risk for atypical maternal care. However, this dataset was not excluded from the current study, because the goal of the current work was to characterize a representative and relatively large cross-sectional sample of young, healthy children. Both genes and environment can influence behavior (McGue & Bouchard, 1998). In the current work, the important exclusion criteria pertained to excluding known genetic risk factors that place children at risk for atypical development, such as familial risk due to having an older biological sibling diagnosed with ASD. To ensure that total scan time is comparable among all datasets from all sites, only the first resting-state scan from this dataset is analyzed in the current study. Thus, all available rs-fMRI datasets from all infants were analyzed, except for participants from #2340 who were scanned 2 times; only the 1st scan was used to ensure that the total scan time is comparable among all datasets. MRI acquisition parameters ~~~~~~~~~~~~~~~~~~~~~~~~~~ Resting-state MRI (rs-fMRI) Blood Oxygenation Level-Dependent (BOLD) data were acquired at 3 Tesla at all three sites, on an Achieva (Phillips Medical Systems, Best, The Netherlands) MR scanner (C-MIND #2329) or on Tim Trio Siemens (Siemens Healthcare, Erlangen, Germany) MR scanners (#1890 and #2340). At C-MIND (#2329), rs-fMRI EPI parameters were: [TR/TE: 2000/35 ms, flip angle = 90°, FOV = 240 × 240 mm, 80 × 80 matrix, 35 axial 4 mm slices], with scan duration of 5 min (150 volumes). At UC Irvine (#1890) rs- fMRI scan EPI parameters were: [TR/TE: 2000/30 ms, flip angle = 77°, FOV = 220 × 220 mm, 64 × 64 matrix, 32 axial 4 mm slices], with scan duration of 6.5 min (195 volumes). At Pitt (#2340), rs-fMRI EPI parameters were: [TR/TE: 2020/32 ms, flip angle = 80°, FOV = 256 × 256 mm, 64 × 64 matrix, 32 axial 4 mm slices], with scan duration of 5.05 min (150 volumes) of each run. Thus, all scan sessions lasted about 5 min. Data were acquired during natural sleep, with eyes closed. During assessment, all children (neonates, infants, toddlers) lay supine on the scanner bed, with the head cushioned to minimize movement. Datasets from #1890 were truncated (only the 1st 150 volumes were used) in order to ensure that the number of samples per se does not affect statistical measures. This procedure ensured that the number of volumes in this dataset is equivalent to the number of acquired volumes in the other two datasets. Thus, exactly N = 150 data samples (N = 149 speed data points; please see below for details) per each participant from each of the datasets were used in the current study. Estimating head movements during the scan Well-established techniques (Friston, 2008; Friston et al., 1995) to estimate head movements during resting-state scans were followed and are presented in detail elsewhere (Denisova et al., 2016). Image volumes were pre-processed using Statistical Parametric Mapping (SPM8), a free open-source software to process neuroimaging data (http://www.fil.ion.ucl.ac.uk/spm/software/spm8/) running MATLAB version 8.3 (R2014a) (The MathWorks, Inc., Natick, MA). Briefly, realignment of scanned volumes in SPM involves estimating the six parameters of an affine ‘rigid- body’ transformation (b-splines interpolation using least-squares approach) that minimizes the differences between each successive scan and a reference scan (Friston, 2008; Friston et al., 1995). The linear (affine) position transformations specified for some portion of the image hold for all other portions of the image (Friston et al., 1995) by virtue of the treatment of this problem as one involving rigid body transformations (Denisova & Zhao, 2017). For each participant, rs-fMRI volume image files in NIFTI format were pre-processed, producing an output with the six motion parameters (3 linear translations in x, y, z directions, and 3 rotations: pitch (about x-axis), roll (about y-axis), and yaw (about z-axis)) which is recorded as an rp_%s.txt file in SPM8; rotational coordinates were in radians and converted to degrees (*180/pi) (Denisova & Zhao, 2017).","Overall, a similar analytic strategy—estimating realignment parameters of head movements during an fMRI or rs-fMRI scan, calculating a summary metric, and applying distributional analyses on the movement time series—was used in several recent studies (e.g., Denisova & Zhao, 2017; Denisova et al., 2016), and is referred to as an ‘upcycling’ strategy in the context of Open Science Infancy Research protocol. Thus, Gamma fits were performed on individual participants’ speed data, and because Gamma has two parameters, each fit yielded 2 parameter estimates per dataset, the shape and the scale. In addition, mean, variance, skewness, and kurtosis were also calculated for each individual dataset. Quality checks performed for all individual fits included checking that MATLAB's maximum likelihood estimation (MLE)-based fitdist algorithm converged on a solution for every participant; successful convergence occurred for all fits reported. For non- parametric Spearman's rank-order rho analyses (two-tailed) of parameter estimates and age, I performed 5000 random bootstrap samples (with replacement) to obtain the uncertainty of the rho estimate (95% Confidence Intervals, CIs). Engle's ARCH test was used to assess for violations of homoscedacity. Linear fits for parameter estimates vs. age are plotted with a Working-Hotelling 95% confidence band covering the entire range of values. The functions and tools including the cftool in the Statistics and Machine Learning and Curve Fitting Toolboxes in MATLAB 8.3 (R2014a) (MathWorks, Natick, MA, USA) were used for all statistical analyses.","Young children—from neonates to 3 year olds—show diversity in their spontaneous head movements during a sleep resting-state fMRI. Fig. 1 presents individual parameter estimates for shape (a parameter, denoting whether the distribution is best characterized by Exponential or skewed-to-Gaussian) and scale (b parameter, denoting noise-to-signal levels) for each participant; the mean, variance, skewness and kurtosis are presented in Supplementary Figure 2). When inspecting individual parameter estimates on shape (a) and scale (b) parameters shown in Fig. 1, note that infants with higher levels of noise-to- signal levels (b parameter) also have more random, noisy signatures (a parameter, tending toward Exponential distributions) in their spontaneous head fluctuations, following a power–law relation f(x) = a*xb (consistent trend for normalized peaks data: Supplementary Figure 3). Supplementary Table 1 and Supplementary Table 2 present goodness-of-fits for Power and Exponential fits for raw time series and normalized peaks, respectively). Taken together, these results illustrate the presence of substantial variability in these cross- sectional data. Can maturational processes account for some of the diversity in the patterns we observed thus far? The prediction is that infants with the highest shape values and the lowest scale parameter estimate values (datasets in the right-most lower location in Fig. 1) are overall older relative to those showing reverse pattern (i.e., those datasets with the lowest shape and the highest scale values that are located toward the left-most upper location in Fig. 1). Age-related patterns in spontaneous head movements ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ For sample individual data from several participants, including a neonate, several infants, and older toddlers, please see Supplementary Figure 4 and also Fig. 2a (data shown are for male children; both the number of speed data samples (n = 149) per dataset as well as the sampling rate are equal across all participants). Note that the youngest child's (neonate; 2–3 weeks old) movements shown in Supplementary Figure 4 have more peak maxima and a more skewed distribution (Fig. 2a) relative to movements of a 28 mo-old toddler. Significant associations are detected between statistical features of spontaneous head movements and the age of participants. While no evidence of heteroscedacity was detected (P > 0.05), linear fits in all plots are shown with Working-Hotelling 95% Confidence bands. Spontaneous angular head movements in older children are significantly less skewed in older children, located toward more symmetric values (lower data points on the y-axis) (non-parametric Spearman's rank-order correlation: rs(89) = −0.4901 (−0.632 to 0.322, 95% Confidence Intervals, CIs), P = 8.1892e−07) (Fig. 2b, left panel) (Consistent pattern is found for skewness of linear spontaneous movements and age, rs(89) = −0.3805 (−0.556 to 0.176, 95% CIs), P = 1.9902e−04). Congruent with this pattern, spontaneous angular head movements in older children are quantified by a significantly higher a-shape parameter, located toward more Gaussian parameter values on the y-axis on the Gamma parameter plane (non-parametric Spearman's rank-order correlation: rs(89) = 0.4611 (0.261–0.635, 95% CIs), P = 4.2389e−06) (Fig. 2b, right panel). This association with age is consistent for parameter estimates obtained from normalized angular peaks data (P = 0.0086) (Supplementary Figure 5a; Supplementary Results). When considering parameter estimates of linear speed, a significant pattern of higher a-shape parameter with age is detected: rs(89) = 0.3376 (0.120–0.542, 95% CIs), P = 0.0011 (consistent for normalized peaks linear speed, P = 0.0424; Supplementary Figure 5a; Supplementary Results). Further, significantly lower b-scale (noise-to-signal level) parameter estimates (angular speed) are associated with increasing age: rs(89) = −0.2878 (−0.487 to 0.060, 95% CIs), P = 0.0057) (Supplementary Figure 5b; Supplementary Results) (consistent results for normalized angular peaks, P = 0.0080). For linear speed, a significant decrease in b-scale parameter with age is detected: rs(89) = −0.2169 (−0.425 to 0.005, 95% CIs), P = 0.0354 (Supplementary Figure 5b; Supplementary Results) (association is consistent for normalized linear peaks, P = 0.0464). Thus, older toddlers have a significant increase in the symmetry as well as a decrease in noise-to-signal levels in their spontaneous head movements during a sleep resting-state fMRI (Fig. 2; Supplementary Figure 5) (Note that this finding is particularly robust for parameters derived from angular speed, and is consistent for analyses using linear speed). This developmental pattern is uncoupled from peak amplitudes of movements during the scan per se. Specifically, the overall trend for increased symmetry and reduced noise levels with age remains when using normalized peaks data (right panels in Appendix A).","The current work quantitatively describes, for the first time, the developmental progress of a set of parameters of spontaneous speed of head movements during sleep rs-fMRI scans in 91 children's datasets with ages ranging from ∼birth to about 3 years old. The signatures of spontaneous head movements, even when the head is relatively well-stabilized during a neuroimaging scan, are diverse among a group of neonates, infants, and older toddlers, up to approximately 3 years old. The youngest participants studied, 2–3 week-old neonates, showed less symmetric and more noisy structure in their movement signatures, tending toward more Exponential distributions and away from more symmetrically Gaussian distributions shown by older toddlers. The current findings strongly suggest that the developmental pattern of head movements during sleep rs-fMRI scan in healthy infants is different from the one to be found in a sample of infants at a familial or genetic risk for atypical development. The role of spontaneous movements during sleep in early life in humans is not well-understood, and several implications and interpretations of the current findings are considered. There are two ways to think about spontaneous head movements during sleep in early life. First, the traditional view is that head movements are a nuisance for neuroimaging data acquisition and analyses. Continued research efforts are still needed to understand the impact of head movements, as well as of motion correction itself, on MRI analyses and inference, especially in studies with young children. Second, the new view is that head movements are actually an important source of neuronal activity in the developing cortex, that this activity builds synapses and functional connections between the different regions in the developing brain, and thus new, transformative research is urgently needed to incorporate all major drivers of nascent functional connections into new methods and studies that measure functional connectivity in early life. Implications for brain imaging work in young children ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In particular, variations in participant movement during MRI acquisition have unintended, deleterious consequences (Friston, Williams, Howard, Frackowiak, & Turner, 1996) on image quality and inference in MRI research (cf. (Denisova, 2018)) and in this case would confound higher-level age-related hypotheses. In this regard, these findings extend recent study reporting significant differences in head movement in infants during MRI as a function of familial autism risk (Denisova & Zhao, 2017). Because the child's age alone is a significant factor in the extent of movement shown by the participant, age-related trajectories of head movements during the scan need to be accounted for in neuroimaging analyses and inference with neonates, infants and toddlers. Age and head motion correction Based on the current findings, we can expect that children of different ages will have different patterns of movement spikes and periods of stillness, such that neonates will have the most variable patterns relative to older children. This effect could be due to differences in how different spontaneous movements, including startles and twitches, naturally occur during sleep at younger vs. older ages. Since removing volumes exhibiting motion, in order to correct for the deleterious impact of movement on acquired EPI data (i.e. via zeroing out or scrubbing; (Power, Barnes, Snyder, Schlaggar, & Petersen, 2012)) affects spectral features of the remaining time series, a less veridical, and more variable, input for functional connectivity analysis is available at the youngest ages. Put simply, due to motion correction strategies, the spectral features of the new (i.e. motion-corrected) EPI time series would vary systematically across age. For this reason, age-related rs-functional connectivity MRI (rs-fcMRI) analyses and inferences are problematic. Increasing scanning duration while obtaining EPI data for rs-fcMRI analyses, as well as multi-echo EPI acquisitions, are among realistic strategies available today for acquiring more good quality data in young children. Nevertheless, the outstanding question is not, “How can we correct movement better?”, but a more serious and compelling one: How can we measure and represent all important sources of neuronal activity that generate functional connections in the developing CNS? New questions for the study of functional connections in the developing human brain Even if it turns out that spontaneous head movements trigger or associate with neuronal activity in a different way than what has been shown by animal researchers, a significant and important issue remains for rs-fcMRI studies with humans in early life. Zeroing out and/or scrubbing EPI volumes with high movement means that researchers do not take into account co-occurring neuronal activity, which could contribute toward building synapses and functional connectivity in the still-developing CNS, during periods of spontaneous movements during sleep (A different problem occurs, as noted earlier, if high movement volumes are not removed, because EPI time series used in the analyses will exhibit age-driven differences in movement, which could lead to spurious age-related functional correlations and could be misinterpreted as brain maturational effects). Existing approaches and methods that aim to non-invasively measure functional connectivity fall short because an important contributor to neuronal activity, for example, for sensorimotor development, is missing from data analyses. The unique, transient biology of the developing CNS must be taken into account as we develop new non- invasive imaging approaches and methods when aiming to reveal nascent functioning of the developing human CNS. Interpretations of age-related movement patterns and targets for future research directions in early brain development ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The key finding of an association between age and improvement in spontaneous movement signatures should be considered with respect to the special needs of neonates for sleep. Because neonates sleep more than older toddlers and may also produce more frequent spontaneous movements than toddlers, this may explain, in part, the finding of more variable signatures in neonates’ spontaneous movements. As noted in the Introduction, spontaneous movements during sleep in children include twitches and startles and these movements occur throughout sleep, that is, during AS sleep as well as quiet sleep (QS) stages which correspond to behavioral state 2 and 1, respectively (Prechtl, 1974). One potential interpretation that is consistent with the current data is that as infants successfully build sensorimotor maps of their world (Blumberg et al., 2015; Tiriac et al., 2014), their requirements for such movements would decrease over time. Another way to understand these findings is to consider that spontaneous movements permit the neonate to gain a sense of physical space in the extrauterine environment. Fetuses have been argued to show proto-predictive capacity and ability to control their movements (Zoia et al., 2007), prenatal learning and perception occurs in-utero (DeCasper & Spence, 1986; Reid et al., 2017), and there is decrease in movements before birth womb (Myowa-Yamakoshi & Takeshita, 2006); there is also the availability for symmetric, non-chaotic movements—while swimming—in newborns (McGraw, 1939)). The change in the statistical character of movements with age observed here may reflect in part the gradual adaptation to the differing physical constraints of the extra-uterine world. It is warranted, in future work, to direct greater research focus to probe the maturation of brain structures that are involved in both cognition and movement, such as the cerebellum. Perhaps unexpectedly, the cerebellum has not received substantial attention in infancy research relative to the cerebral cortex and subcortical structures, even though this structure contains 80% of the neurons in the human brain (Herculano-Houzel, 2010) with evidence of microscopic myelination already present at birth (Kinney, Brody, Kloman, & Gilles, 1988). With regard to the specific features of the neuronal architecture available early in life, although axonal myelination (the sheathing or covering of neurons’ axons in the nervous system that supports neuronal conduction speeds) is initially weak at birth in humans, postmortem studies indicate that the onset, rate, as well as the direction of myelination varies (Paus et al., 2001) not only by a specific brain structure (e.g., the cerebellum is the first major structure to myelinate, before the cerebral cortex (Kinney et al., 1988)) but also locally within the cerebral cortex. Recent work detected structural abnormalities in the cerebellar cortex in boys with ASD diagnoses relative to typically developing children (Zhao, Walsh, Long, Gui, & Denisova, 2018). As such, it would be instructive to investigate the relation between the improvement of movement signatures with age relative to the anatomical maturation of the cerebellum in order to better understand the important anatomical brain-based and developmental behavior-based contributions to cognition. Finally, advancing knowledge about early sensorimotor development is very important for understanding normative brain development. In particular, normal development of sensorimotor systems and their connections are important for normative development of vocal learning and speech (e.g., (Bruderer, Danielson, Kandhadai, & Werker, 2015)). Atypical functioning of these systems and connections could herald neurodevelopmental disorders in children, including ASD. Further, different systems (e.g., auditory, visual, limbic) and structures (e.g., cerebellum) could interact in different ways, exactly during the critical time when synapses and connections are being built within and across these systems. Limitations ~~~~~~~~~~~ During MRI scans, young infants are swaddled and their heads are well-cushioned, in order to restrict their movements during signal acquisition. Considering that these constraints are in place during acquisition of sleep rs-fMRI scans, infants’ spontaneous head movements may be likely different from movements that would be produced by these infants during ‘normal’, non-constrained sleep. For example, one may expect that there would be increased movement, such as increases in terms of magnitude, during unconstrained sleep. In future work it would be helpful to measure statistical features during unconstrained sleep in natural environments using motion capture systems as well as using polysomnography (PSG) and electroencephalography (EEG) methods. Nevertheless, the current study established an age-driven pattern of reduced noise-to-signal and increased symmetry also using normalized data, meaning that this developmental pattern holds after accounting for magnitude of movements per se in individual infants.","This cross-sectional study revealed an age-related pattern in statistical features of head movements during a sleep rs-fMRI scan in a sample comprised of healthy children between 0 and 3 years of age, including neonates, infants, and toddlers. Specifically, a significant attenuation of noise-to-signal and an increase in symmetry in head movements during sleep rs-fMRI was detected in older infants and toddlers, relative to the youngest participants studied. The current quantitative characterization of a functionally important and heretofore under-examined aspect of normative development during sleep establishes that statistical features of spontaneous head movement signatures during rest or sleep are expected to become less noisy and more symmetrical in ∼3 year old children. For in vivo, non-invasive neuroimaging studies that aim to reveal how young brains work and learn, the key outstanding question is not “How can we correct head movement better?” but rather: How can we represent all important sources of neuronal activity that shape functional connections in the still-developing human CNS?","I have no competing interests, including any financial interests, activities, relationships, and affiliations that could have influenced decisions or work on this manuscript.","KD conceived and designed study, performed analyses, and wrote paper.","This work is supported by the Sackler Award in Developmental Psychobiology to KD. The sponsor had no role in study design; in the collection, analysis and interpretation of data; in the writing of the report; and in the decision to submit the article for publication."],["Disrupted sleep is a transdiagnostic factor characterising a multitude of psychiatric conditions. Although this is well-recognised, the cause of poor sleep across conditions is unclear. One possibility is that poor sleep is driven by traits which also co-occur with multiple conditions. Previous research suggests that alexithymia (an inability to identify and describe one's emotions) is a candidate trait, as it is linked to poor sleep quality and elevated levels of alexithymia are seen across multiple diagnostic groups. The association between alexithymia and poor sleep quality has been questioned however, with studies arguing that it is depression and anxiety, rather than alexithymia, which impact sleep quality. Problematically, such studies typically utilise measures of depression and anxiety which include items relating to sleep – meaning that apparent associations between depression and anxiety may be due to measurement issues, rather than to depression and anxiety per se. Study 1 confirmed the relationship between alexithymia and subjective sleep quality, whilst Study 2 utilised an independent sample to replicate the association between alexithymia and sleep quality, and to demonstrate that it is not a product of co-occurring depression or anxiety. Results therefore support the suggestion that alexithymia may explain disrupted sleep across multiple psychiatric conditions. --------------------------------------------------------------------------------","Poor sleep quality is reported across multiple psychiatric conditions (e.g., Freeman et al., 2017) including depression (see Tsuno, Besset, & Ritchie, 2005) and certain subtypes of anxiety disorders (see Papadimitriou & Linkowski, 2005). Given the importance of sleep for wellbeing (Freeman et al., 2017), identifying factors responsible for poor sleep across the disorders in which it is experienced is an important research aim. In recent years there has been a growing appreciation that transdiagnostic symptoms, including poor sleep, may be explained by traits that co-occur with multiple disorders, rather than by the disorders themselves. In particular, the contribution of alexithymia, a sub-clinical condition characterised by difficulties identifying and describing one's own emotions and an externally orientated thinking style, has begun to be appreciated (Sifneos, 1973). Alexithymic individuals often exhibit poor emotional functioning in a number of domains outside their defining inability to identify and describe their own emotions, including emotion recognition (Brewer, Cook, Cardi, Treasure, & Bird, 2015; Cook, Brewer, Shah, & Bird, 2013; Grynberg et al., 2012; Heaton et al., 2012) and empathy (Bird et al., 2010). Perhaps as a consequence, individuals with alexithymia also experience problems with interpersonal relationships and exhibit increased rates of mental and physical ill-health (Taylor, Bagby, & Parker, 1999). Importantly, increased rates of alexithymia are observed across many psychiatric conditions (e.g., Brewer, Cook, & Bird, 2016; Murphy, Brewer, Catmur, & Bird, 2017), and alexithymia has been demonstrated to be responsible for a range of symptoms across these conditions (e.g., Brewer et al., 2015; Cook et al., 2013). With respect to sleep, previous studies have identified an association between alexithymia and poor sleep quality using both subjective (e.g., Bauermann, Parker, & Taylor, 2008) and objective measures of sleep quality (e.g., Bazydlo, Lumley, & Roehrs, 2001). For example, using subjective measures alexithymia has been associated with sleep-related problems in community samples (e.g., Bauermann et al., 2008; Hyyppä, Lindholm, Kronholm, & Lehtinen, 1990) as well as in clinical groups such as men with depression (e.g., Honkalampi, Saarinen, Hintikka, Virtanen, & Viinamäki, 1999) and individuals with depression (e.g., Aydin, Ozdemir, & Selvi, 2012); and rates of alexithymia are reported to be higher in those with insomnia (e.g., Engin, Keskin, Dulgerler, & Bilge, 2010). Whilst few studies have employed objective measures, there are also reports that alexithymia is associated with increased light sleep (Bazydlo et al., 2001; but see De Gennaro et al., 2002). Furthermore, recent evidence links alexithymia to atypical interoception (perception of one's internal bodily state; Brewer et al., 2016; Herbert, Herbert, & Pollatos, 2011; Shah, Hall, Catmur, & Bird, 2016; Murphy, Catmur, & Bird, 2018), and poor interoception has also been linked to poor self-reported sleep quality in clinical conditions (Ewing et al., 2017). These studies are therefore consistent with the proposal that heightened levels of alexithymia may confer risk of disrupted sleep. Although a body of evidence links alexithymia to poor sleep, it should be acknowledged that individuals with higher levels of alexithymia are more likely to report increased symptoms of depression and anxiety (e.g., Hendryx, Haviland, & Shaw, 1991; Honkalampi, Hintikka, Tanskanen, Lehtonen, & Viinamäki, 2000), which have both been associated with poor sleep (see Tsuno et al., 2005; Papadimitriou, & Linkowski, 2005). As a consequence, it is possible that it is depression or anxiety, and not alexithymia, that is responsible for poor sleep quality. Indeed, previous studies examining the specific contribution of alexithymia, depression and anxiety to poor sleep quality have produced mixed results, with some studies suggesting that the relationship between alexithymia and sleep disturbance is driven by anxiety (Lundh & Broman, 2006; with anxiety quantified using the Karolinska Scales of Personality; KSP; Gustavsson, Weinryb, Göransson, Pedersen, & Åsberg, 1997) or depression (De Gennaro, Martina, Curcio, & Ferrara, 2004; with depression assessed using the Center for Epidemiological Studies Depression scale; CES-D; Radloff, 1977), and some not (Kronholm, Partonen, Salminen, Mattila, & Joukamaa, 2008; with depression assessed using the Beck Depression Inventory; BDI; Beck, Ward, Mendelson, Mock, & Erbaugh, 1961). Problematically, however, many measures of depression and anxiety include items relating to sleep, and all of the previous studies investigating the relationship between alexithymia, anxiety, depression and sleep have utilised measures that include items assessing sleep. For example, the most commonly used measure of depressive traits, the BDI (Beck et al., 1961) includes items relating to both sleep (e.g., “I don't sleep as well as I used to”) and tiredness (e.g., “I get tired more easily than I used to”). The same is true for other routinely-used measures such as the CES-D (Radloff, 1977) which also includes items relating to sleep (e.g., “My sleep was restless”); and measures of anxiety that include items relating to tiredness (e.g., “Quite often, especially when I am tired, I get the feeling that either I or the world around me is changing - a feeling of unreality”; KSP; Gustavsson et al., 1997). The association between depression and anxiety and sleep, and whether anxiety and depression are responsible for the association between alexithymia and sleep quality, may therefore depend on the degree to which items relating to sleep contribute to assessment of depression and anxiety on a particular measure. Specifically, studies utilising depression or anxiety measures that have a greater focus on sleep quality than other measures may inflate relationships between anxiety, depression and sleep, which in turn obscure, or reduce, associations between alexithymia and sleep quality (De Gennaro et al., 2004; Kronholm et al., 2008; Lundh & Broman, 2006). The degree to which the association between alexithymia and sleep quality is suppressed will, in turn, depend on the proportion of individuals in a particular sample with co-occurring clinical symptoms and sleep problems. These factors make it very difficult to determine whether alexithymia contributes to poor sleep quality independent of its relationship with depression and anxiety when measures of depression and anxiety are utilised that include items relating to sleep. Consequently, the aim of the present set of studies was to 1) confirm the relationship between alexithymia and poor sleep quality and 2) examine whether these associations are driven by anxiety or depression using measures which do not include items that assess sleep quality.","Participants were recruited via pre-existing databases of individuals who had indicated an interest in taking part in psychological research and via social media advertisements. Participants were informed that the study aimed to investigate links between emotional and bodily awareness, and physical health. 86 participants took part in Study 1. Of these, 70 participants fully completed the questionnaires, had English as their first language and identified their gender as male or female (Mage = 42.93, SDage = 21.80, range 18–91, 26 Males) and were included in analyses. 13.95% were removed as they had English as their second language, 3.49% were removed for a failure to complete all measures and 1.16% were removed as they identified as a non-binary gender. Ethical approval was granted by the local ethics committee, all participants gave informed consent, and were fully debriefed upon completion.","Participants completed the Toronto Alexithymia Scale (TAS-20; Bagby, Parker, & Taylor, 1994) and the Pittsburgh Sleep Quality Index (PSQI; Buysse, Reynolds, Monk, Berman, & Kupfer, 1989) in a randomised order online via Qualtrics (Provo, UT). High scores on these measures indicate elevated alexithymic traits and poor sleep quality, respectively. The TAS-20 is comprised of three subscales, difficulties describing feelings, difficulties identifying feelings and externally orientated thinking.","Where directional predictions are made, one-tailed tests are utilised. Alexithymia scores ranged from 21 to 79 (M = 46.90, SD = 13.55). As predicted, total alexithymia scores were associated with reduced sleep quality (r(68) = 0.462, p < .001; one-tailed). Analysis of the TAS-20 sub-factors indicated a significant association between reduced sleep quality and each sub-factor (all p < .006; one-tailed). As the relationship between alexithymia and sleep quality may vary across genders (e.g., Honkalampi et al., 1999; Kronholm et al., 2008), and levels of alexithymia have been reported to vary across the lifespan (e.g., Mattila, Salminen, Nummi, & Joukamaa, 2006), additional analyses were conducted controlling for these variables. Regression analyses predicting sleep quality from age (years), gender (0 = female, 1 = male) and alexithymia total scores suggested that both alexithymia (standardised β = 0.534, t = 4.903, p < .001; one-tailed) and age (standardised β = 0.265, t = 2.408, p < .010; one-tailed) were predictive of poor sleep quality and the overall model was significant (F(3,66) = 8.420, p < .001). To uncover whether the relationship between alexithymia and sleep varied as a function of gender, the regression analysis was re-run including the interaction term between gender (−1 = female, and 1 = male) and mean-centred alexithymia scores. The interaction term did not predict sleep quality (standardised β = 0.154, t = 1.302, p > .05; two-tailed). An adequate range of alexithymia (Mage = 50.19, SDage = 12.16, Range 24–79), depression (Mage = 10.77, SDage = 10.0, Range 0–38) and anxiety scores (Mage = 8.27, SDage = 7.62, Range 0–30) were obtained, with scores on the anxiety and depression DASS subscales ranging between normal and extremely severe (Lovibond & Lovibond, 1995) and scores on the TAS-20 ranging from very low to high alexithymic traits (Bagby et al., 1994). Likewise, a large range of PSQI scores were obtained (Mage = 6.49, SDage = 3.60, Range 0–18). As expected, significant correlations were observed between these variables (all p < .001; one-tailed; Table 1.). As in Study 1, all three subscales of the TAS-20 were correlated with poor sleep quality (p < .05; one-tailed). Analysis of the data was carried out using hierarchical regression. Participant age (years), gender (0 = female, 1 = male) and anxiety scores were entered into the first step, depression into the second step, and alexithymia total scores into the third step of a regression model predicting sleep quality. Visual examination of the residuals confirmed a normal distribution and no multicollinearity (all VIF < 1.88). At step one only anxiety scores predicted sleep quality (standardised β = 0.408, t = 3.513, p < .001; one-tailed) and the overall model was significant F(3,69) = 5.041, p = .003. At step two only depression scores predicted sleep quality (standardised β = 0.428, t = 3.328, p < .001; one-tailed), and the overall model was significant (F(4,68) = 7.103, p < .001). Anxiety no longer predicted sleep quality (standardised β = 0.131, t = 0.959, p > .15; one-tailed). When alexithymia was added (step three), both alexithymia (standardised β = 0.261, t = 2.133, p < .02; one-tailed) and depression predicted sleep quality (standardised β = 0.338, t = 2.561, p = .007; one-tailed). The inclusion of alexithymia significantly increased the variance accounted for by the model (by 4.5%; F(1, 67) = 4.551, p = .037), and the overall model was significant (F(5,67) = 6.889, p < .001). Although residuals appeared normally distributed, the Kolmogorov–Smirnov test indicated that the residuals were not normally distributed (D = 0.183, p < .001). Accordingly, the analysis was re-run using robust regression (entry method), which is less sensitive to departures from normality (Field & Wilcox, 2017). This analysis confirmed the same pattern of results, whereby depression (standardised β = 0.107, t = 2.353, p = .011; one-tailed) and alexithymia (standardised β = 0.096, t = 2.753, p = .004; one-tailed) were the only predictors of poor sleep quality. The original regression analysis was repeated including the interaction between gender (−1 = female, 1 = male) and mean centred alexithymia scores in step four. This had little effect on the results obtained, and the interaction term was not a significant predictor of sleep quality (standardised β = 0.065, t = 0.554, p > .250).","Having confirmed an association between alexithymia and poor sleep quality, the aim of Study 2 was to determine whether this association was driven by depression or anxiety using measures which do not include items relating to sleep quality.","A power analysis using the effect size observed in Study 1 for the correlation between alexithymia and sleep quality indicated an N of 27 would provide 80% power to detect an effect at α = 0.05 (one-tailed). Participants were recruited as per Study 1 and the same ethical guidelines were adhered to. 108 participants completed the online questionnaires anonymously via Qualtrics (Provo, UT). Of these, 73 fully completed the surveys, had English as their first language and identified as male or female (Mage = 38.84, SDage = 15.82, Range 18–79, 26 Males). 26.9% of the sample were removed as they had English as their second language, 2.8% of the sample were removed for non-completion and 2.8% of the sample were removed as they identified as a non-binary gender.","Participants completed the TAS-20 and PSQI as well as the Depression, Anxiety and Stress Scale (DASS-21; Lovibond & Lovibond, 1995) online. The DASS-21 provides separate subscales for depression and anxiety, and no items relating to sleep are present in either subscale. As is typical, scores from the DASS-21 were doubled to equate to the DASS-42 to allow ease of comparison with studies using this measure (Lovibond & Lovibond, 1995), with higher scores indicative of greater depression or anxiety.","The present set of studies sought to examine the association between alexithymia and sleep quality. Consistent with previous evidence, Study 1 confirmed that alexithymia is associated with poor sleep quality and, in this sample, the impact of alexithymia on sleep quality did not vary across genders. In Study 2, the association between alexithymia and poor sleep quality was replicated, and analyses confirmed that the association was not a product of co-occurring depression or anxiety. As expected, anxiety, depression and alexithymia were all associated with reduced sleep quality (e.g., Tsuno et al., 2005; Papadimitriou, & Linkowski, 2005; Bauermann et al., 2008). However, no relationship between anxiety and sleep quality was observed once depression was controlled for, suggesting that the relationship between anxiety and sleep quality may be driven in part by the high comorbidity between depression and anxiety (see Papadimitriou, & Linkowski, 2005). It remains a possibility, however, that a unique effect of anxiety on sleep quality may be observed for distinct subtypes of anxiety disorders. In contrast to previous evidence suggesting associations between alexithymia and sleep quality are driven by heightened rates of anxiety and depression (De Gennaro et al., 2004; Lundh & Broman, 2006) in the alexithymic population, this study found an independent effect of alexithymia on sleep quality. Such findings are consistent with data from Kronholm et al. (2008), who found an independent effect of alexithymia on sleep not explained by depression despite using a depression measure that included items relating to sleep (the BDI). Whilst these findings conflict with some previous reports (e.g., De Gennaro et al., 2004; Lundh & Broman, 2006), as previously noted it is possible that the measures of depression and anxiety used in previous studies that specifically ask about sleep quality or refer to tiredness (e.g., the BDI, CES-D, and KSP) masked the relationship between alexithymia and sleep in samples with particular patterns of covariance between clinical symptoms and sleep quality. These data therefore underscore the importance of ensuring independence of measurement when selecting measures of depression and anxiety, and highlight the usefulness of the DASS-21 (which does not include items that assess sleep-related factors) when assessing correlations between depression, anxiety, and sleep quality. The finding that heightened alexithymia is associated with poor sleep suggests that alexithymia may contribute towards the often-observed sleep disturbances across a range of disorders (e.g., Freeman et al., 2017) that are also characterised by heightened rates of alexithymia (e.g., Murphy et al., 2017). Whilst the mechanism by which alexithymia confers risk of disrupted sleep remains unclear, suggestions include increased nocturnal arousal as a result of poor verbalisation of emotions (Hyyppä et al., 1990), and increased light sleep (e.g., Bazydlo et al., 2001; see Bauermann et al., 2008). More recent associations between alexithymia and interoception (e.g., Brewer et al., 2016; Herbert et al., 2011; Murphy et al., 2018; Shah et al., 2016), and evidence that low interoceptive accuracy is associated with poor sleep quality in certain disorders (Ewing et al., 2017), also raise the possibility that poor interoception may be the mechanism by which alexithymia confers risk of disrupted sleep. However, the direction of causality remains unclear (see also Ewing et al., 2017); it is also possible that poor interoception results in heighted alexithymia which in turn results in poor sleep quality and a greater risk of psychiatric disorders, or that poor sleep may impact upon interoception which results in increased alexithymia and risk of poor mental health. As is clear from this speculation, future studies using longitudinal designs should examine the relationship between alexithymia, interoception, and sleep, using both subjective and objective measures of interoception and sleep quality. It is important to acknowledge certain limitations of these studies. First, although several potential confounds were controlled for, body composition (e.g., obesity), often associated with both sleep (Beccuti & Pannain, 2011) and alexithymia (Pinna et al., 2011), was not controlled for. However, given that associations between alexithymia and sleep quality have been reported when body composition is controlled for (Kronholm et al., 2008), it is unlikely this would have changed the pattern of results. Second, although sleep quality is often quantified using self-report measures (e.g., the PSQI), discrepancies between objective and subjective measures of sleep quality have been reported, with some suggestion that PSQI scores may be influenced by general negativity (Grandner, Kripke, Yoon, & Youngstedt, 2006). However, as previous reports suggest that alexithymia is associated with objective measures of poor sleep even when depression (CES-D) is controlled for (Bazydlo et al., 2001), it is likely that alexithymia also makes a contribution to objective measures of sleep disturbance. Future research, therefore, should investigate the association between alexithymia, depression, anxiety, and sleep quality, using complementary measures of both subjective and objective sleep quality. Third, in the present study, depression and anxiety were assessed over the previous week whereas sleep quality was rated over the previous month. Whilst this method has been utilised previously (e.g., De Gennaro et al., 2004), it is possible that these differing time frames may influence the relationship between these measures. It is therefore important that future studies equate the measurement period for each factor assessed. Finally, in comparison to previous work, the sample size recruited for the present study was small. However, it is important to note that an adequate range of scores was present for all key variables, the association between alexithymia and sleep quality was replicated across independent samples, and sample sizes were above those determined using formal power analyses. In summary, these data demonstrate that alexithymia is an independent predictor of poor self-reported sleep quality, raising the possibility that alexithymia may contribute towards sleep problems observed across a number of psychiatric disorders."],["Correctly estimating the confidence we should have in our decisions has traditionally been viewed as a perceptual judgement based solely on the strength or quality of sensory information. However, accumulating evidence has demonstrated that the motor system contributes to judgements of perceptual confidence. Here, we manipulated the speed at which participants’ moved using a behavioural priming task and showed that increasing movement speed above participants’ baseline measures disrupts their ability to form accurate confidence judgements about their performance. Specifically, after being primed to move faster than they would naturally, participants reported higher confidence in their incorrect decisions than when they moved at their natural pace. We refer to this finding as the adamantly wrong effect. The results are consistent with the hypothesis that veridical feedback from the effector used to indicate a decision is employed to form accurate metacognitive judgements of performance. --------------------------------------------------------------------------------","Humans are unique amongst animals in being able to provide explicit reports on the reliability of, or confidence in their decisions. Previous studies have demonstrated that our confidence in our decisions or opinions plays a key role in group interactions (Bahrami et al., 2010; Koriat, 2012). Whenever people express an opinion, they are likely to also communicate their confidence in that opinion, be this explicitly through what they say or implicitly in their movements and facial expressions (Aitchison, Bang, Bahrami, & Latham, 2015). Accurate understanding of confidence has obvious implications for high-risk decision making domains such as financial investment (e.g. Broihanne, Merli, & Roger, 2014), medical diagnosis (e.g. Berner & Graber, 2008), jury verdicts (e.g. Tenney, MacCoun, Spellman, & Hastie, 2007), and politics (Johnson, 2004). Theoretical models of perception have proposed that confidence is related to the quality or strength of sensory processing (Barthelmé & Mamassian, 2010; Kepecs, Uchida, Zariwala, & Mainen, 2008; Kiani & Shadlen, 2009; Vickers, 1979; Zylberberg, Barttfeld, & Sigman, 2012; see Yeung & Summerfield, 2012, for a review) and speak to a domain-specific formation of confidence judgements. However, there is increasing evidence that perceptual-decision signals are also seen in neural circuits specialised for motor actions (Cisek & Kalaska, 2005; Freedman & Assad, 2011; Hernández, Zainos, & Romo, 2002; Romo, Hernández, & Zainos, 2004; Shadlen & Newsome, 2001), suggesting a contribution of the motor system to estimates of confidence, and supporting the idea of metacognition as a domain-general process. Indeed, it has recently been shown that disruption of the motor system, specifically the dorsal premotor cortex, reduces metacognitive ability when performing a perceptual discrimination task (Fleming et al., 2014). In addition, Allen et al. (2016) report the results of an interoceptive priming manipulation where autonomic arousal modulates subjective confidence on a motion-discrimination task. Thus, it is has been suggested that movement parameters proprioceptive and interoceptive states may also serve as a useful cue for the inference of confidence in our own decision-making. Indeed, previous research has shown that the speed at which a participant makes a forced choice decision is correlated with their confidence, with faster reaction times associated with more confident decisions (Fleming, Weil, Nagy, Dolan, & Rees, 2010). Moreover, subjects are able to infer the subjective confidence of another person simply by the observation of their actions (Patel, Fleming, & Kilner, 2012), with faster movements rated as more confident and vice versa. This is reliant on the motor system as subjects with movement disorders have difficulty inferring the confidence of others moving at speeds very different from their own (Macerollo, Bose, Ricciardi, Edwards, & Kilner, 2015), and disrupting activity in the motor system reduces healthy individuals sensitivity to infer confidence from the kinematics of others (Palmer, Bunday, Davare, & Kilner, 2016). These findings suggest that an individual may in part infer their confidence in their decisions from their own movement parameters. Here, we tested this hypothesis using a behavioural priming task to alter movement speed, while participants performed a perceptual contrast discrimination task. We recorded the speed at which the participant made their decisions. After each trial of the perceptual decision task, we asked the participant to rate their confidence in their performance, and calculated their metacognitive ability, as a measure of the relationship between their confidence and accuracy.","Forty-eight healthy participants with normal or corrected-to-normal vision were recruited (31 female, 17 male), with a mean age of 27 (range 18–53, median 24). Forty-four reported being right-handed, four reported being left-handed. The experiment was fully explained to participants, apart from the aim of the project and the true aim of the priming task, which were not disclosed until debriefing to prevent bias. The experiment was approved by the University College London Ethics Review Board. Informed written consent was obtained from all participants and procedures were conducted in accordance with the Declaration of Helsinki. Equipment ~~~~~~~~~ Participants were seated at a table, 60 cm in front of a Dell laptop computer, and responded using the standard QWERTY keyboard, a marble and three touch-sensitive containers (Fig. 1). Stimulus display and response collection were controlled by MATLAB 7.8.0 (Mathworks Inc., MA, USA) using the Cogent 2000 toolbox (http://vislab.ucl.ac.uk/cogent.php). Stimuli and procedure The experiment was carried out at the Institute of Neurology, University College London. All participants were tested individually, in the presence of the experimenter. Participants completed two blocks of a metacognition task described below (50 trials per block), followed by a movement speed prime (50 trials), a third block of the metacognition task, a second movement speed prime, and finally a fourth block of the metacognition task (Fig. 2a). The first block of the metacognition task was used as a practice session and was not included in the analyses. Metacognition task The metacognition task was a perceptual contrast discrimination task used in previous studies (Fleming et al., 2010; Patel et al., 2012). The stimuli were comprised of two images shown in quick succession on the laptop computer screen. Each image comprised a circular clock-face with six Gabor gratings (circular patches of light and dark bars) arranged around a central fixation point (Fig. 2b). The background was uniform grey, with a luminance of 3.66 cd/m2. In one of the two images, all the Gabor gratings were set to the same contrast, that is, a ‘baseline’ Gabor grating. In the other image, one of the Gabor gratings was set to a higher contrast than the other five baseline gratings, causing it to appear as a ‘pop-out’. Pop-out gratings were drawn from a stimulus set that varied in contrast between 23 and 80% in increments of 3%. The pop-out Gabor grating would appear at random in either the first or second image and at a random spatial position orientated around a central fixation point. After viewing both images, participants were prompted to make a decision as to which image (first or second) contained the pop-out grating by the appearance of the message “1 or 2” on the screen. Participants made their response by moving a marble from the ‘homepad’, positioned centrally on the table, to one of two containers, marked “1” and “2”. The container marked “1” was always located on the right of the homepad, and “2” on the left. After indicating their decision, participants returned the marble to the homepad. A red square framed their selection on the screen and participants were then prompted to indicate their confidence in their decision by typing a number between 1 and 99. A lower value indicated low confidence and a higher value indicated high confidence. No feedback was given to participants about the actual accuracy of their decision. Reaction times were recorded from participants while they indicated their decision on the contrast discrimination task, allowing us to operationalise movement speed as the time between the marble being removed from the homepad and placed in a response container. The contrast of the pop-out Gabor grating was varied throughout the experiment using a two-up, one-down staircase procedure to keep accuracy consistent (Fleming et al., 2014; Patel et al., 2012). The staircase operated such that after two consecutive correct decisions the contrast was decreased by one increment, whereas after one incorrect decision the contrast was increased by one increment. A correct response would thus be made on approximately 70% of trials, ensuring that the analysis of movement speed and confidence was not affected by performance. In addition to movement speed and confidence rating, we also recorded accuracy (correct or incorrect response) and signal strength (pop-out Gabor grating contrast strength) on each trial. Movement prime During this task, participants were prompted to move the marble from the homepad to either of the two containers marked “1” and “2”, and then return it to the homepad. The desired location was instructed on a trial by trial basis, by the appearance of a 1 or 2 on the screen. On completing the movement, participants were given visual feedback on the screen about their movement speed (“just right”/“too slow”/“too fast”) (Fig. 2c). Movement speed was calculated from the same reaction time parameters as those used in the metacognition task. The fast prime was designed to increase movement speed above natural average parameters, while the slow prime was designed to decrease movement speed below natural parameters. In the fast prime task, the goal was for participants to remove the marble from the homepad container and move it to the indicated container as fast as possible. A message of “just right” appeared on the screen if the time it took to complete the task was less than 1500 ms. If the time between the marble being removed from the homepad container and being placed in another container was greater than 1500 ms, a message of “two slow” appeared. In the slow prime task, participants were required to complete the movement of the marble from the homepad container to the indicated container in over 2300 ms. If they accomplished this, a message of “just right” appeared on the screen. If their movement speed was less than 2300 ms, a message of “too fast” was presented. To prevent participants simply delaying the onset of the movement, a message of “start time too slow” was presented if the participant did not remove the marble from the homepad container in under 1600 ms. The order in which participants completed the movement prime task (fast or slow first) was counterbalanced. Analyses ~~~~~~~~ As the first block of the metacognition task was used as a practice session, the analyses presented here were conducted on blocks two through four. All trials with a recorded movement time of zero were removed from analyses. Greenhouse-Geisser was used to correct for violated sphericity and a more stringent alpha criterion of p = .01 was used to correct for multiple comparisons, where required. Analysis 1: Relationship between movement speed and confidence We conducted a bivariate correlation analysis of confidence and movement speed at baseline (second block of the metacognition task). Movement speeds and confidence ratings were mean-corrected and then ranked by movement speed from fastest to slowest. These values were then divided into ten equal sized bins in order of increasing magnitude. Each bin was then averaged to produce ten mean values for confidence and ten mean values for movement speed. A significant negative relationship would replicate previous findings of an inverse interdependence (Baranski & Petrusic, 1998; Fleming et al., 2010; Patel et al., 2012), whereby higher confidence is associated with faster movement. Analysis 2: Effect of movement prime on movement speed An analysis of variance (ANOVA) was used to determine if the movement primes had worked. That is, was there was a significant effect of movement prime on movement time, in the expected direction. In other words, had the fast prime made participants move faster and had the slow prime made participants move slower? Analysis 3: Effect of movement prime on confidence An analysis of variance (ANOVA) was used to determine if the movement primes had produced a significant effect on confidence ratings. Analysis 4: Effect of movement prime on metacognitive accuracy Thus, the area below the major diagonal in an individual’s ROC curve is a measure of their ability to link confidence to perceptual performance (AROC), in other words, their metacognitive accuracy. Analysis of variance (ANOVA) was used to determine if the movement primes had produced a significant effect on metacognitive accuracy. Secondly, analysis of variance (ANOVA) and nonparametric permutation were used to determine if there was an interaction effect between correct and incorrect responses and movement primes on confidence. Analysis 1: Relationship between movement speed and confidence ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A bivariate correlation analysis revealed a significant negative correlation between confidence and movement speed [r2 = −0.870, p = .001, df = 47], before the application of any movement prime tasks (Fig. 3). This replicates previous findings of a negative relationship between confidence and movement speed (Baranski & Petrusic, 1998; Fleming et al., 2010; Patel et al., 2012). Analysis 2: Effect of movement prime on movement speed ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A one-way repeated-measures ANOVA with three levels (baseline, fast, slow) revealed a significant effect of condition on movement speed [F(1.62,76.11) = 8.720, p = .001, ηp2 = 0.156]. Pairwise comparisons revealed significantly faster movement speeds after the fast prime (mean = 706.394, SD = 145.070) compared to baseline measures (mean = 745.460, SD = 165.241) [t(47) = 3.161, p = .002, d = 0.469] and slower movement speeds after the slow prime (mean = 776.829, SD = 203.351) compared to after the fast prime [t(47) = 3.948, p < .001, d = 0.638] (Fig. 4a). The difference in movement speed after the slow prime relative to baseline measures was not significant [t(47) = 1.597, p = .117, d = −0.239]. The difference between movement speeds post fast prime and post slow prime was confirmed to be significantly different with a nonparametric sign test (p < .001). One large outlying data-point (>500 ms) can be seen in the Slow condition (Fig. 4a). Removing this participant’s data from the analysis revealed the same result, with a main effect of condition [F(2,92) = 13.542, p < .001, ηp2 = 0.227], driven by significant differences between baseline and fast conditions [p = .007] and fast and slow conditions [p < .001]. Analysis 3: Effect of movement prime on confidence ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A one-way repeated-measures ANOVA with three levels (baseline, fast, slow) revealed no significant main effect of condition on confidence [F(1.79,84.32) = 0.070, p = .916, ηp2 = 0.001] (Fig. 4b). One large outlying data-point (>15) can be seen in the Slow condition (Fig. 4b). Removing this participant’s data from the analysis revealed the same result, with no significant main effect of condition [F(1,46) = 0.057, p = .812, ηp2 = 0.001]. Analysis 4: Effect of movement prime on metacognitive accuracy ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A one-way repeated-measures ANOVA with three levels (baseline, fast, slow) revealed no significant main effect of condition on [F(1.56,73.50) = 2.389, p = .111, ηp2 = 0.048] on AROC. Specifically looking at the difference in metacognitive accuracy between the two primed conditions revealed significantly higher AROC after the slow prime (mean = 0.756, SD = 0.080) than after the fast prime (mean = 0.725, SD = 0.089) [t(47) = 2.659, p = .011, d = 0.385) (Fig. 4c). Neither of these values was significantly different to baseline (mean = 0.747, SD = 0.105) (Baseline vs Slow [t(47) = 0.534, p = .596, d = −0.077], Baseline vs Fast [t(47) = 1.611, p = .114, d = 0.239]). One large outlying data-point (>0.2) can be seen in the Slow condition (Fig. 4c). Removing this participant’s data from the analysis revealed the same result, with no main effect of condition [F(1.637,75.298) = 2.296, p = .117, ηp2 = 0.048, Greenhouse-Geisser corrected]. As AROC has previously been found to be influenced by both Type-I accuracy/d′ and criterion (Maniscalco & Lau, 2012; Fleming & Lau, 2014), Spearman’s correlation analyses were conducted to confirm that changes observed in AROC were not the result of conditional changes in these measures. There was no significant correlation between AROC and d′ [rs = 0.175, p = .234] or AROC and criterion [rs = −0.077, p = .602] after the fast prime or between AROC and d′ [rs = −0.123, p = .407] or between AROC and criterion [rs = 0.189. p = .199] after the slow prime (see Supplementary Fig. 1). (See also Supplementary Fig. 2 for subject wise Type-II ROC curves for each condition). As we had found a significant difference in metacognitive accuracy between the primed conditions in the absence of significant changes in confidence, we investigated whether changes in confidence were dependent on participant’s accuracy. To this end, a two-way repeated-measures ANOVA was conducted with a factor of response (2 levels, correct or incorrect) and a factor of condition (3 levels: baseline, fast prime, slow prime) was performed. This revealed a significant main effect of response, with significantly higher confidence on correct trials than incorrect trials [F(1,47) = 169.559, p < .001, ηp2 = 0.783]. There was no significant main effect of condition [F(1,47) = 0.823, p = .369, ηp2 = 0.005]. There was a significant interaction between the two [F(1.99, 93.66) = 5.487, p = .006, ηp2 = 0.188]. To investigate the location of the interaction effect, we removed the main effect of response by mean- correcting the confidence data. Confidence ratings for correct and incorrect responses were mean-corrected separately. As a result, the main effect of response was no longer significant [F(1,47) = 0.823, p = .269, ηp2 = 0.021] but the results of the main effect of condition and interaction effect remained identical [F(1,47) = 0.823, p = .369; F(1.99,93.66) = 5.487, p = .006, ηp2 = 0.188, respectively]. Pairwise comparisons revealed significantly higher confidence for incorrect responses (mean = 1.519, SD = 5.068) than correct responses (mean = −1.523, SD = 4.455) after a fast prime [p = .002, d = 0.469] (Fig. 5). All other comparisons were insignificant after correcting for multiple comparisons. Furthermore, we tested whether this interaction effect was also significant using non-parametric analysis. To this end, we ran a permutation test on the interaction. We simulated 10,000 permutations to determine the distribution of values for the interaction of interest. In each simulation, the average confidence ratings for each of the four conditions were randomly assigned to the four conditions: Fast Prime correct, Slow Prime incorrect, Fast Prime Incorrect, Slow Prime correct. This was performed for each participant independently. The average interaction value was then calculated across participants as above. Once this had been repeated 10,000 times, the distribution of the interaction terms could be plotted, and the point of the true interaction value on this distribution was calculated. This test also revealed a significant interaction effect (less than 0.1% of the data had permutations at a higher value than the true value, i.e. p < .001). In other words, this interaction effect occurred by chance in less than 0.1% of circumstances, indicating this to be a robust effect.","Here we tested the hypothesis that altering participants’ movement speed will disrupt their ability to form accurate confidence judgements of their performance. The results supported this hypothesis by indicating that after being primed to move faster than they would naturally, participants report higher confidence in their incorrect decisions than their correct ones. Overall, metacognitive accuracy was significantly higher when participants were slowed than when they were quickened. These findings suggest that the metacogntive ability to match confidence to accuracy relies on a veridical signal from the motor system. These data are consistent with previous findings on the role of the motor system in judgements of perceptual confidence. We replicated the finding that the speed at which an individual makes a forced choice decision is correlated with their confidence, with faster reaction times associated with more confident decisions (Fleming et al., 2010). Consistent with the finding that disrupting activity in the motor system reduces an individual’s ability to form accurate judgements of confidence (Fleming et al., 2014) and to infer confidence from the kinematics of others (Palmer et al., 2016), here we find that altering an individual’s kinematics behaviourally reduces their ability to infer their own confidence. The absence of a significant direct effect of movement time on confidence in the presence of a relationship between confidence, accuracy and movement, suggests that the relationship between metacognition and movement is not a simple linear one. In this study, moving faster does not increase confidence and moving slower does not decrease confidence directly. However, moving faster when you are incorrect raises your confidence above and beyond what it is following a correct response. It is important to note that feedback on performance was not provided to participants in the research reported here. Thus, the only information participants had to use in their judgement about their confidence was the decision making process itself, including the movement made to indicate their response. It is possible that if participants possess an internal model of their decision making, including the movement made to indicate that decision, modifying movement parameters may disrupt the learnt relationship between how movement time relates to confidence. In effect, participants are no longer able to successfully ‘read out’ their confidence from their movement time. As a result, their metacognitive ability can no longer be predicted. The results presented here point towards a role of kinematics in higher-order cognitive functions, such as the metacognitive monitoring of performance. Previous research had shown that cognition has a top-down influence on an individual’s motor behaviour, i.e. that we move faster when we’re feeling more confident (Audley, 1960; Baranski & Petrusic, 1998; Henmon, 1911; Johnson, 1939; Patel et al., 2012; Volkmann, 1934), but the results of the current research also suggests that we may need an accurate and veridical representation of our motor behaviour in order to form accurate and veridical cognitive judgements. This supports a more ‘closed-loop’ view of action and cognition, whereby there are bottom-up influences on cognition from the body, in addition to the more closely studied top-down cogntive influences on behaviour. As such, the present findings hold implications for traditional views of cognition, as operating on a distinct plane from the motor system. While the results of the current research do suggest a trial by trial correspondence between confidence and accuracy that is influenced by ascending information from the effector about the movement speed used to indicate the decision, the present design does not allow us to comment on whether continued processing of this information takes place. According to decision-locus models of metacognition confidence is solely based on evidence available at the point of the judgement (Kiani & Shadlen, 2009; Vickers, 1979). Alternatively, according to postdecisional-locus models, metacognitive evidence continues to be accumulated after the decision has been made (Pleskac & Busemeyer, 2010). Evidence from transcranial magnetic stimulation (TMS) indicated that stimulation applied after the decision affected confidence judgements and does support the latter hypothesis (Fleming et al., 2014). The time course of this effect, however, is not known. Further research might investigate for what length of time after movement modulation metacognitive disruption is seen. In this study we demonstrated a causal role of the motor system in metacognition, specifically, that how an individual moves when they make a decision impacts their ability to relate their feelings of confidence to their performance. This finding suggests that humans obtain information about their confidence levels by monitoring their movement kinematics and that these inferences can be manipulated by changing how subjects move. It would be interesting for future research to investigate the role that organic changes in movement speed, such as that seen in Parkinson’s disease, play in metacognitive ability and estimates of subjective states, such as confidence. This research suggests that highlights the need for researchers should to consider the inclusion of kinematic information in subsequent models of metacognition. One such model is presented by Fleming and Daw (2017), which accounts for action in self-evaluation judgements proposes that individuals form confidence estimates by applying an analogous computation to their own actions.","This work was supported by a PhD studentship awarded to E.R.P. from the Economic and Social Research Council (ESRC) and the Medical Research Council (MRC). A.K. was supported by a ‘European Research Council Starting Investigator Award’ [ERC-2012-STG GA313755]."],["Difficulties with handwriting are reported as one of the main reasons for the referral of children with Developmental Coordination Disorder (DCD) to healthcare professionals. In a recent study we found that children with DCD produced less text than their typically developing (TD) peers and paused for 60% of a free-writing task. However, little is known about the nature of the pausing; whether they are long pauses possibly due to higher level processes of text generation or fatigue, or shorter pauses related to the movements between letters. This gap in the knowledge-base creates barriers to understanding the handwriting difficulties in children with DCD. The aim of this study was to characterise the pauses observed in the handwriting of English children with and without DCD. Twenty-eight 8-14 year-old children with a diagnosis of DCD participated in the study, with 28 TD age and gender matched controls. Participants completed the 10. min free-writing task from the Detailed Assessment of Speed of Handwriting (DASH) on a digitising writing tablet. The total overall percentage of pausing during the task was categorised into four pause time-frames, each derived from the literature on writing (250 ms to 2 s; 2-4 s; 4-10 s and >10 s). In addition, the location of the pauses was coded (within word/between word) to examine where the breakdown in the writing process occurred. The results indicated that the main group difference was driven by more pauses above 10 s in the DCD group. In addition, the DCD group paused more within words compared to TD peers, indicating a lack of automaticity in their handwriting. These findings may support the provision of additional time for children with DCD in written examinations. More importantly, they emphasise the need for intervention in children with DCD to promote the acquisition of efficient handwriting skill. --------------------------------------------------------------------------------","Developmental Coordination Disorder (DCD) is the term used to refer to children who present with motor coordination difficulties, unexplained by a general medical condition, intellectual disability, sensory or neurological impairment (American Psychiatric Association [APA], 2013). Handwriting difficulties are mentioned in the formal diagnostic criteria for DCD (APA, 2013), are frequently mentioned in parent and teacher reports and are the most common reason for referral to occupational therapy services for this population (Asher, 2006). While a knowledge base surrounding their handwriting deficits has emerged in recent years (Jolly & Gentaz, 2013; Prunty, Barnett, Wilmut, & Plumb, 2013; Rosenblum & Livneh-Zirinski, 2008), the underlying mechanisms of the handwriting deficits are still unclear. In particular, there is a limited understanding of the excessive pausing during handwriting reported in this population (Rosenblum & Livneh-Zirinski, 2008). This creates barriers for healthcare professionals in terms of prescribing intervention, as it is unclear what the pausing behaviour in the handwriting process actually represents. In a recent study we found that children with DCD produced less text than their typically developing (TD) peers in four handwriting tasks from the Detailed Assessment of Speed of Handwriting (DASH; Barnett, Henderson, Scheib. & Schulz, 2007; Prunty et al., 2013). However, closer inspection of the handwriting process through the use of a digital writing tablet revealed that the slowness in production was not due to slower movement of the pen, but due to a higher percentage of the task spent pausing (with the pen either in the air, or resting on the page) (Prunty et al., 2013). This ‘pausing phenomenon’ in the handwriting of children with DCD was initially revealed in a study in Israel by Rosenblum and Livneh-Zirinski (2008), where children with DCD were found to spend considerably more time than controls with the pen in the air. However, beyond the findings of these studies, little is known about the behaviour of pausing and the possible cognitive or physical explanations for why they occur in the handwriting of children with DCD. It is imperative to understand this pausing behaviour, as it has been shown to significantly impact on the production of text in children with DCD (Prunty et al., 2013). The concept of ‘pausing’ during writing is complex in nature and cannot be considered without recognising handwriting as a component of the writing process. Handwriting is ‘language by hand’ (Berninger & Graham, 1998; Berninger, Abbott, Abbott, Graham, & Richards, 2002) and there are many cognitive processes which occur before, during and after the pen is placed on the page (Kandel, Soler, Valdois, & Gros, 2006; Van Galen, 1991). Van Galen's (1991) psychomotor model of handwriting which informs the current investigation illustrates this, by describing the process from the transformation of language into the sequencing of handwriting movements. At the highest level of the model is the activation of the intention to write followed by semantic retrieval, syntactical construction and spelling. The first step in the motor process is to select the appropriate allograph, which according to Van Galen (1991) is the activation of the motor programme (retrieval of an allograph action pattern from long-term motor memory). Following the activation of the motor programme the module of size control and speed is activated (Van Galen, 1991). The muscle synergies from both the agonist and the antagonistic muscles are then recruited during the muscular adjustment module, which results in the real time movement of the pen (Van Galen, 1991). Van Galen's model of handwriting, although the most complete in the literature, is not without limitations. According to Kandel and Spinelli (2010) it fails to account for parallel processing at different levels of the model, which results in an increased duration of the handwriting movements. In addition, research by Kandel, Soler, Valdopis and Gros (2006) in French writers has demonstrated that letters are not programmed individually, but rather in chunks, and the temporal profile is determined by the number of syllables in the word. This emphasises the relationship between the production of handwriting and the cognitive and linguistic aspects of writing and must be accounted for when examining pausing behaviour in more detail. The definition of a pause in writing is inconsistent in the literature. Rosenblum and Livneh-Zirinski (2008) defined a pause as a pen lift from the writing tablet. However, it was not clear as to how long the pen needed to be raised from the surface in order to be classified as a pause. In the writing literature Alamargot, Chesnet, Dansac, and Ros (2006) and Alamargot, Plane, Lambert, and Chesnet (2010) defined a pause as a period of 15 ms or more with the pen not in contact with the tablet. The rationale for such a short pause threshold was to include all writing events that occurred; including raising the pen to think of an idea or briefly to dot an ‘i’. Other authors have omitted to define a pause and made reference only to the fact that pauses occurred during the writing task (Accardo, Genna, & Borean, 2013). In addition to the debate on how pauses are classified, it is also unclear even in the literature on writing what exactly a pause represents. Thus, the pause thresholds used in the current study are grounded in evidence where possible, while some aspects of the analysis are exploratory in nature. In relation to pausing in children with DCD, Rosenblum and Livneh-Zirinski (2008) proposed anomalies in the lower level processes of handwriting, such as between stroke muscular adjustments, as reasons for excessive pausing. Alternative explanations for the excessive pausing include physiological factors such as fatigue (Rosenblum & Livneh- Zirinski, 2008). However, none of these theories have been tested and it remains unclear whether children with DCD pause excessively for short periods of time (i.e. <1 s), or whether they pause for longer periods possibly as an indication of fatigue or higher-level writing processes such as planning (i.e. >4 s). In a study by Alamargot et al. (2010) on French writers it was found that longer pauses possibly reflected processes such as planning. In their study, five participants with varying levels of writing expertise were asked to compose a text by extending a narrative provided to them. There were no time constraints and they were asked to write as much as they felt was necessary to finish the story. The participants included three school students, one in grade 7 (12 years old), one in grade 9 (14 years old) and the third in grade 12 (17 years old). The remaining two participants included a university graduate student (22 years old) and an established expert author. The writing task was completed on a writing tablet and eye-tracking was used to infer processes that occurred during pauses in the writing. Alamargot et al. (2010) analysed four types of pausing activity for each writer which they labelled as ‘quartiles’. The quartiles consisted of pauses between 78 and 189 ms, 129 and 416 ms, 194 and 624 ms, 695 and 23,248 ms. The participant in grade 7 had the most pauses in quartile 4 (longer pauses) compared to the other four writers. In fact, the 7th grade writer had pauses as long as 13–18 s at times. According to Alamargot et al. (2010) the longer pauses were due to a strategy known as step-by-step production of text, where the child switches between planning and formulation of the text to cope with the cognitive demands of handwriting. Indeed as the level of writing expertise increased, the number of longer pauses decreased substantially. Using eye-tracking Alamargot et al. (2010) were able to investigate the longer pauses based on gaze fixations. These were classified based on whether the participant was looking back at text, looking at the handwriting area or looking away from the task. They found that the least experienced writers were inclined to look away from the task, which according to Alamargot et al. (2010) was an indication of planning. Despite the detailed analysis of pausing in typically developing children, the nature of pausing has not been investigated in the handwriting of children with DCD and knowing whether the pauses are driven by many pauses of small duration or a few of longer duration is needed to inform approaches to intervention. Another issue alongside the duration of pauses is their location within text. Previous research by Kandel et al. (2006) on children found that words tend to be programmed prior to execution, followed by online processing during the execution phase. According to Alamargot et al. (2010) if the cognitive demands of handwriting exceed working memory capacity, the word cannot be processed online, therefore a pause occurs within the word. The pause would allow the information to be processed before completing the next word segment (Alamargot et al., 2010). Indeed issues with word level pausing have been found in other developmental disorders, where children with dyslexia paused for greater periods within words and particularly around misspellings (Sumner, Connelly, & Barnett, 2012; Sumner, 2013). According to Sumner (2013), this was indicative of difficulties at the word level due to the constraints that spelling difficulties had on word processing. Although spelling difficulties are distinctly different to motor difficulties, both spelling and handwriting form the basis of transcription skills in models of writing and handwriting (Berninger & Amtmann, 2003; Van Galen, 1991). According to Kandel et al. (2006) both spelling and motor planning for handwriting are processed prior to writing the word and online thereafter. It is therefore plausible that the motor difficulties associated with DCD would manifest in a within word pause, similar to that of dyslexia, as the child would be unable to plan words online given the strain of the handwriting on cognitive resources. Prunty et al. (2013) hypothesised that a lack of automaticity in the handwriting of children with DCD contributed to the lack of production of text. Indeed examining whether children with DCD pause within words would be a method of examining automaticity of handwriting and their ability to manage cognitive processes in parallel. Exploring this area in more detail would provide a better description of the handwriting process in children with DCD and in doing so inform an evidence base to support the provision of intervention in this group. Therefore the aim of this study was to investigate the pausing behaviour demonstrated by the same group of children with DCD in Prunty et al. (2013) compared to their TD peers. To do so, their pausing behaviour on a 10 min free-writing task was analysed and categorised using four time-frames chosen from the available literature on handwriting in DCD and writing. Since children with DCD have motor difficulties, it was hypothesised that the DCD group would pause for a greater percentage of time within smaller time-frames due to possible difficulties manipulating the pen to form the letters. In addition, it was anticipated that the DCD group would also pause more within-words as a result of the lack of automaticity in their handwriting. Additional measures of pause frequency and mean pause duration were also examined separately to compare findings from previous studies on the handwriting process in children (Alamargot et al., 2010; Rosenblum & Livneh-Zirinski, 2008).","Twenty eight children with DCD (27 boys, 1 girl) and 28 age (within 4 months) and gender matched typically developing (TD) controls were included in the study. DCD group All children met the DSM-5 diagnostic criteria for DCD (APA, 2013) and a diagnostic assessment was in line with recent European guidelines (Blank, Smits- Engelsman, Polatajko, & Wilson, 2012). The children had significant motor difficulties, with performance below the 10th percentile (24 below the 5th, 4 below the 10th) on the Movement Assessment Battery for Children 2nd edition Test (MABC-2; Henderson, Sugden, & Barnett, 2007) which examines three components of motor competency; manual dexterity, ball skills and balance. These motor difficulties had a significant impact on their activities of daily living, as reported by their parents and evident on the MABC-2 Checklist (Henderson et al., 2007). A developmental, educational and medical history was taken from the parents, which confirmed that there was no history of neurological or intellectual impairment and no medical condition that might explain the motor deficit. The British Picture Vocabulary Scale 2nd edition (BPVS-2, Dunn, Dunn, Whetton, & Burley, 1997) was used to give a measure of receptive vocabulary, which correlates highly with verbal IQ (Glenn & Cunningham, 2005). This was in at least the average range for all children, confirming the absence of a general intellectual impairment. The Strengths and Difficulties Questionnaire (SDQ; Goodman, 1997) was also used to note other behavioural difficulties reported by the parent, which commonly occur with DCD such as attention deficit hyperactivity disorder (ADHD) (Miller, Missiuna, Macnab, Malloy-Miller, & Polatajko, 2001). No child had a diagnosis of ADHD, but hyperactive behaviour was noted on the SDQ for seven children. The children were also assessed on the reading and spelling components of the British Ability Scales 2nd Edition (BAS-II; Elliott, 1996). These revealed that eight children with DCD had literacy difficulties (1 in reading, 7 in spelling), as defined by a standard score of less than 85 on the BAS-II components, although none had a formal diagnosis of dyslexia or other language impairment. Typically developing (TD) control group The control group was recruited through local primary and secondary schools in Oxfordshire, England. Teachers were asked to use their professional judgement to identify children without any motor, intellectual or reading/spelling difficulties. To ensure the children identified were free of these difficulties, they were individually tested on the MABC-2 Test (Henderson et al., 2007), BPVS-2 (Dunn et al., 1997) and the reading and spelling components of the BAS-II (Elliott, 1996). Children were included in the control group if they scored at least at the level expected for their age on all measures. Children from both groups with a diagnosis of dyslexia, and/or those who had English as a second language were excluded from the study. Children in both groups who had a reported physical, sensory or neurological impairment were also excluded. This was to ensure that handwriting difficulties could not be attributed to other disorders. See Table 1 for performance profiles of both groups. The study was approved by the University Research Ethics Committee at Oxford Brookes University. Parents were required to sign a consent form and children were asked to either assent (below 11 years), or counter sign the parent consent form (over 11 years). The handwriting component of this study took place over one 60-min session. The detailed assessment of speed of handwriting (DASH; Barnett et al., 2007) The free-writing task from the DASH (Barnett et al., 2007) was chosen for an extended analysis of pausing data presented in Prunty et al. (2013). This is an ecologically valid task in terms of its similarity and relevance to handwriting demands in the classroom. Free-writing involves the integration of all aspects of the writing process including idea generation, production of language, spelling and handwriting. Examining the pauses in the context of a free-writing task provides an overall view of the child's ability to cope with the demands of the writing system at work. It also provides an opportunity to examine whether the writing process is forced to succumb to high cognitive loading by imposing pauses where text would have otherwise been processed online. The DASH free-writing task requires the child to write about the topic of ‘my life’ for 10 min using their everyday handwriting. They are presented with a spider diagram prior to the beginning of the task which provides topics that they could write about. They are then given 1 min to plan their writing and are asked to write continuously once the task starts (Barnett et al., 2007). The handwriting product scores for the DCD and TD groups on this measure have been previously reported by Prunty et al. (2013). The DASH has UK norms for children aged 9–16 years. The internal reliability of the total score is between α = .83–.89 and the inter-rater reliability for all four tasks is .99, as reported in the test manual.","When completing the DASH free-writing task the participants wrote with an inking pen on paper placed on a Wacom Intuos 4 digitising writing tablet (325.1 mm × 203.2 mm) to record the movement of the pen during handwriting. The writing tablet transmits information about the spatial and temporal data of the pen as it moves across the surface. The data was sampled at 100 Hz via a Celeron Dual Core CPUT3500@2.10GHz laptop computer. Eye & Pen version1 (EP1) software (Alamargot et al., 2006) was used to analyse the pauses. The pauses were examined in three separate analyses in order to address independent questions. The rationale for the three analyses is provided below and included the following: The categorisation of pauses into time-frames and taken as a percentage of the overall pause time. An analysis of the location of pauses to ascertain where in the writing process the pauses occurred. An examination of the frequency and duration of pauses. 1. Categorising the pauses The first analyses included the categorisation of pauses into time-frames to examine the pauses which attributed to the overall percentage of pause time. The analyses included an examination of the following. Pausing at the letter level The first analysis in categorising the pauses examined short pauses of between 30 and 250 milliseconds (ms). This time-frame was chosen from the literature, as it is thought to represent the graphomotor component of handwriting. Alamargot et al.’s (2010) study mentioned above established a link between short pauses and graphomotor execution, particularly in pauses between 78 and 189 ms. However, it is not clear from the literature exactly what is meant by ‘graphomotor’ activity. It could include for example, the transition between individual letters, or a split second pause between letter strokes. Nevertheless, short pauses are thought to represent the pauses that occur specifically at the letter level. Since the study by Alamargot et al. (2010) included a small sample size of five participants ranging from a novice, grade 6 writer to an expert published author, the time-frame for letter level pauses was adjusted to 30–250 ms for the current analysis to capture variation in the writing process of children. The second time-frame used to examine between letter pauses was 250 ms to 2 seconds (s). This was chosen based on previous research by Rosenblum and Livneh-Zirinski (2008), where children with DCD were found to pause for longer between letter strokes. Rosenblum and Livneh- Zirinski (2008) reported in-air time (pause time) ranging from .37 s to 1.27 s on the alphabet task, suggesting that this was the pause time which occurred between letters. Given the difference in the production of Hebrew and Latin based letters a 250 ms to 2 s time-frame was used to capture any variation due to kinematic differences in the production of Latin based letters. The pauses between 250 and 2 s were analysed and calculated as a percentage of the overall pause time. Pausing at the word-level To examine word level pauses, the time frame for analysis was between 2 and 4 seconds (s). This was chosen from the literature on writing, where a 2 s pause in typically developing (Alamargot et al., 2010; Alves et al., 2007; Wengelin, 2007) and atypically developing (Sumner, 2013) writers is considered to represent a pause from formulating the text in order to access a higher-level writing process such as planning It was important to capture pauses at or above 2 s to examine pauses at the word level. However, it was also important to restrict the pause time- frame to below 4 s, so lengthy pauses possibly due to fatigue could be measured separately. Two time-frames were analysed for the longest pauses in writing. A pause that was between 4 and 10 s was considered to represent a higher level writing process (generating ideas) or resting due to fatigue. A pause above 10 s was considered to be a significant halt in the writing activity, possibly due to fatigue or a lack of writing ideas. Alamargot et al. (2010) found that the younger, 6th grade writer paused at times for over 10 s which they hypothesised as representative of planning during writing. 2. Location of pauses The second set of analyses examined the location of pauses in order to evaluate the breakdown in the writing process in greater detail. The analyses included an examination of the following. All pauses above 2 s were analysed to examine pauses at the word level. This is the threshold previously used to explore the issue of word level pausing in relation to spelling difficulties in children with dyslexia (Sumner, 2013). For the calculation of between-word pauses, the total number of opportunities to pause between words was always one less than the total words produced. For example, if a child produced 60 words, there were 59 opportunities to pause between words. Since the children with DCD produced fewer words (M = 119.60 SD = 59.59) than the TD group words (M = 156.53 SD = 43.65) (see Prunty et al., 2013), the between-word pauses and within-word pauses were analysed as a percentage of each participants’ word count. When coding a between word pause, if a child paused twice between two words, only one of them was coded as a between-word pause ‘1′, the other was labelled as miscellaneous ‘−1’. This was to differentiate between pauses due to programming the following word versus those used to revise or edit previously written text. Pauses that occurred at or above 2 s were examined and coded as within word pauses where appropriate. Not every pause was coded. For example, if a child paused more than once in a word then only one pause was coded, the rest were coded as miscellaneous, ‘−1’. Similar to between word pausing, this was to differentiate between pauses due to programming the remainder of the word versus those used to edit previously written text. Equally if a child finished writing a word and paused to go back and dot an ‘i’ this was also coded as ‘−1’ as it was not considered to be a between word or within word pause (i.e. it was not considered to be planning the execution of the next word). Also, if a child paused within a word to go back and edit a previous word and then paused within the previous word during the edit, only one pause was coded, the rest were coded as ‘−1’. The pauses within words were also coded based on the accuracy of spelling and legibility. Pauses within correctly spelled words were coded as ‘2’, pauses within misspelled words were coded as ‘3’ and pauses within illegible words (identified by DASH scoring instructions) were coded as ‘4’. The amount of time spent over 2 s within these locations were calculated as a percentage of overall pause time (Fig. 1). Longer 10 s pauses The longer pauses were divided into two categories to distinguish between those related to fatigue and those related to a lack of ideas for writing. To do so, the pauses over 10 s that occurred within a sentence or writing topic/idea were identified and were coded as ‘1’, suggesting that the child had already generated ideas to write about. In this instance, having to stop within a sentence may indicate fatigue. The code ‘2’ was assigned to a pause over 10 s that occurred between the end of a writing topic and the beginning of a new one. A pause before a new topic of writing would perhaps suggest that the pre-writing pause was due to planning. Fig. 2 illustrates an example of code ‘1’ and ‘2’. 3. Frequency and duration of pauses The third and final set of analyses included an examination of behaviours reported in the literature on writing and to report their presentation in children with DCD. The analyses included an examination of the following. The frequency of pausing was also considered here as Alamargot et al. (2010) found no difference in the frequency of pausing between a novice writer and an expert author. However, it is not known whether children with DCD demonstrate similar frequencies of pausing compared to TD children. It remains unknown whether their pausing behaviour is driven by the length of their pauses rather than the frequency. Therefore, the frequency of pauses was calculated for three time-frames, those over 250 ms (letter level based on Rosenblum & Livneh-Zirinski, 2008 and above), between 4 and 10 s (longer pauses possibly due to planning) and those over 10 s (longer pauses possibly due to planning or fatigue). The mean pause duration over 2 s was examined in order to ascertain whether children with DCD paused for longer on average during a word level pause compared to their TD peers. This analysis was based on Rosenblum and Livneh-Zirinski (2008) where children with DCD demonstrated a longer mean pause duration than TD peers. However, it is not known whether this is the case within English children with DCD. Statistical analysis For comparisons between the DCD group and TD group, tests of normality were conducted initially and descriptive statistics for the dependent variables examined. T-tests were used to examine the differences in the mean values between the groups for all normally distributed measures. Those measures which did not meet the normal distribution assumptions were compared using the nonparametric Mann–Whitney-U test. Since age was often a significant co-variate for the handwriting product measures in our previous work (Prunty et al., 2013) and many variables in the pausing analysis violated normal distribution, Spearmans bivariate correlations were used to examine the relationship between age and the pausing measures. Both groups were analysed together and separately with a significance level set at p < .05. Regression analyses were undertaken for children with DCD to ascertain what factors best predict pausing behaviour. In order to ensure that any difficulties with reading and/or spelling did not influence the results, a sub analysis of the DCD group was undertaken prior to the main analyses. Results for the 20 children from the DCD group with at least average reading and spelling scores were compared to the eight children with poor reading and/or spelling (standard scores below 85). There was no significant difference between these two groups in any of the pausing analyses; therefore the two groups were combined to form one DCD group for all subsequent analyses. A similar sub-analysis was completed for those children who demonstrated slightly raised profiles on the SDQ for attention. However there was no significant effect of group for any of the pause analysis therefore the children were combined into one DCD group (n = 28).","1. Categorising the pauses Table 2 illustrates the total overall pause time for each group during the 10-min free-writing task along with the breakdown of each pause time-frame for both groups. The total pause time and breakdown of pauses within specific time-frames are reported in minutes. There was a significant group difference for the overall total pause time (t(54) = 2.34, p < .023), as previously reported by Prunty et al. (2013). Pausing at the letter level There was no effect of group for the amount of time (U = 387.0, Z = −.082, p = .935) or percentage of pause time spent pausing within the 30–250 ms range (t(54) = .359, p = .721). For the 250 ms to 2 s time-frame there was no effect of group for the amount of time spent pausing within this time-frame (t(54) = −.887, p = .380). However, there was a significant effect of group for the percentage of overall pause time, with a greater percentage of pausing for the TD group (t(54) = −2.21, p = .032). Pausing at the word-level There was no significant group difference for the time spent pausing at the word level (U = 380.0, Z = −.197, p = .844) or the percentage of overall pause time attributed to pauses between 2 and 4 s (U = 325.0, Z = −1.09, p = .272). There was no significant group difference for the amount of time spent pausing between 4 and 10 s, (t(54) = .420, p = 676) or for the percentage of pausing within this time frame (t(54) = −.160, p = 874). The DCD group paused for significantly longer over 10 s (U = 323.5, Z = −2.19, p = .029) and a greater percentage of their pausing occurred above 10 s compared to the TD group (U = 265.5, Z = −2.14, p = .032). Fig. 3 highlights all pauses above 250 ms in the writing of a 13 year old male participant with DCD. Fig. 4 highlights the same pauses but in a typically developing male participant. The figures illustrate that both participants have a high frequency of pauses. The difference is in the size of the circles, as larger circles indicate longer pauses. Fig. 3 illustrates a higher percentage of longer pauses. 2. Location of pauses There was no significant group difference in the time spent pausing between words (U = 303.5, Z = −1.45, p = .147) or the percentage of between word pauses (U = 326.0, Z = −1.08, p = .279). The DCD group paused within 22% of the words produced during the free-writing task, compared to 16% for the TD group. However, this difference was not statistically significant (U = 299.0, Z = −1.52, p = .127). In terms of the duration of time spent pausing, there was a significant group difference in within word pausing when all three categories were combined (within correctly spelled words, misspelled words and illegible words) (t(54) = 2.28, p = .026). Individually, there was no significant effect of group for the duration of time spent pausing within correctly spelled words (U = 363.5, Z = −.468, p = .640) or misspelled words (U = 322.0, Z = −1.55, p = .121), but there was a significant effect of group for within illegible word pauses (U = 270.0, Z = -2.08, p = .037). Table 3 illustrates the duration of pauses within words and also shows the percentage of overall pause time spent within words and the locations of word level pausing. Longer 10 s pauses Sixty-eight percent of children with DCD had pauses of over 10 s compared to 50% of the TD group. There was a significant effect of group for the location of pauses, as the DCD group paused more frequently within an idea compared to the TD group (U = 230.5, Z = −2.885, p = .004). In terms of pausing to think of new ideas, there was no significant effect of group for 10 s pauses before a new topic of writing (U = 342.0, Z = −.948, p = .343). A Wilcoxon signed-rank test revealed that the DCD group produced more of their 10-s pauses within an idea (Z = 2.78, p = .006), whereas there was no distinction between pause locations within the TD group (Z = −.162, p = .871). 3. Frequency and duration of pauses There was no effect of group for frequency of pausing above 250 ms (U = 363.5, Z = −.467, p = .640) or above 4 s (U = 265.0, Z = −1.12, p = .260) but there was a significant group difference for frequency of pausing over 10 s (U = 258.0, Z = −2.27, p = .023). For mean pause duration, there was a significant group difference, as children with DCD paused for longer with a mean pause duration of 5.33 s (SD = 1.90) (Mdn = 5) compared to 4.15 s (SD = 1.16) (Mdn = 4) (U = 265.0, Z = −2.081, p = .037). Correlations between age and pausing measures ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 4 shows the Spearmans bivariate correlations that examined the relationship between pausing and age (years and months). As can be seen, for children with DCD and their TD peers, pause time over 4 s, mean pause duration and frequency of pausing over 10 s were all significantly negatively related to age. In addition, for children with DCD a significant negative correlation was found between age and the time spent pausing within misspelled words. A significant positive relationship was found between age and time pausing between 250 ms and 2 s for the TD group only indicating that as TD children increase in age, the amount of time pausing within 250 ms to 2 s increases. The measure of frequency of pauses over 250 ms revealed no significant correlations for either of the two groups. Regressions ~~~~~~~~~~~ Regression analyses were undertaken for children with DCD to ascertain what factors best predict the frequency of pausing over 10 s. Spelling ability was included due to the close relationship between spelling and handwriting in models of writing (transcription skills). Although the participants were closely matched for age, age was entered into this regression due to the wide age range of participants and the correlation shown between age and frequency of 10 s pauses. In addition, manual dexterity as measured by the manual dexterity component of the M-ABC-2 test was included to see if this explained any of the variance. The regression model was a predictor of frequency of pauses above 10 s in the DCD group, R2 = .32, adjusted R2 = .27, F(2, 25) = 5.98, p = .008. It was found that age significantly predicted frequency of 10 s pauses (β = −.372, p = .034), as did manual dexterity (β = −.387, p = .028) indicating that better manual dexterity was associated with a decrease in the frequency of pauses. Coefficients can be found in Table 5.","Until now, the most detailed studies on handwriting in children with DCD were those of Rosenblum and Livneh-Zirinski (2008) and Prunty et al. (2013), where children with DCD were found to spend considerable extra time pausing, compared to their typically developing peers. However, information was lacking on the exact nature of the pausing in children with DCD. The aim of this study was to pinpoint and characterise the pausing behaviour by locating the exact time-frames and locations of pauses in the handwriting of children with DCD. This type of analysis was the first of its kind in the field of DCD and was developed in an attempt to form a greater understanding of the pausing behaviour and its impact on the writing process. While debates continue to surround the categorisation and interpretation of pauses, where possible the different pause categories in this study were selected in line with the available literature on writing. The first two research questions investigated pausing at the grapho-motor level of the handwriting process to examine whether children with DCD spent a greater amount of time than their TD peers pausing at the letter level. The justification for this analysis lay in previous research in the field of DCD where Rosenblum and Livneh-Zirinski (2008) proposed three possible reasons for pausing, including difficulties with the perceptual aspect of the movement, difficulties with motor memory for letter formation and/or difficulties in visualising the letters prior to forming them. While there may well be difficulties in these areas in children with DCD, they did not appear to contribute to the excessive pausing in this study, as there were no group differences in the amount of time spent pausing between 30 and 250 ms and they were not shown to pause more than the TD group in the time frame of 250 ms to 2 s. This finding was also supported in our previous study (of which this is a more detailed analysis) through the lack of group differences that emerged on the alphabet writing task (involving writing the alphabet from memory as quickly as possible for 60 s) (Prunty et al., 2013). In-fact the alphabet task in the previous study was the only handwriting task that did not reveal group differences in pausing. However, it is important to note that there may be separate issues that occur at the letter level, as the DCD group spent more time pausing within illegible words. This calls into question the quality of the movement and the possibility that although the children with DCD were able to transition between letters as quickly as their TD peers, they produced poorer quality letters. Legibility is a crucial element of handwriting performance that needs to be examined in more detail empirically in the future. Meanwhile, it may be necessary for clinicians to examine both speed and legibility in children with DCD given their behaviour of pausing within illegible words. In terms of the behaviour of pausing, the group difference did not appear to be represented at the letter level. The second time-frame that was analysed focused on word level pauses and whether the excessive pausing in children with DCD lay within a 2–4 s time-frame. The 2-s threshold has been previously used in the writing literature, where a boundary of 2 s is thought to represent higher units of processing rather than letter production (Alamargot et al., 2010; Alves et al., 2007; Wengelin, 2007). Word level pausing was also of interest here, as previous research by Kandel et al. (2006) has shown that spelling and motor programming of words occur prior to the execution phase and online thereafter. This suggests that children with DCD might be forced to pause within a word to cope with the excessive cognitive load imposed by handwriting. However, the findings revealed no group differences in the amount of time or percentage of time pausing within a 2–4 s time-frame. With regards to the location of word level pausing, there was also no group difference for between-word pausing, which may suggest that the DCD group did not take longer than their TD peers to programme words prior to executing them. However, a lack of fluency in the writing process did emerge in the DCD group through within word pausing, as they spent a greater amount of time pausing within words compared to their TD peers. Therefore, although there were no differences prior to executing the words, there were halts within words indicating difficulties with processing information on-line (Kandel et al., 2006). Figs. 3 and 4 provide a visual representation of this behaviour, as the sample of writing from the child with DCD clearly exhibits pauses within-words compared to the behaviour of the TD child in Fig. 4. This finding is in contrast to studies on other developmental disorders, where it was found that children with dyslexia paused more both between and within-words compared to TD peers (Sumner, 2013; Wengelin, 2007). In dyslexia, within-word pausing occurred in correctly spelled and misspelled words (Sumner, Connelly, & Barnett, 2014). However, this was not found to be the case here, as the DCD group did not pause more within correctly or incorrectly spelled words, but did pause more within illegible words. Although difficulties with spelling are distinctly different from difficulties with motor skill, it may be that within word pausing occurs in dyslexia and DCD as a means of coping with difficulties in transcription. In the case of DCD, this study found that the within word pausing occurred within illegible words, a concept which should be considered in more detail in the future. The analysis of longer pauses between 4 and 10 s was addressed in order to distinguish between shorter, letter-level pauses and those of longer duration attributed to higher-level writing processes. In Alamargot et al.’s (2010) study it was found that the grade 7 (12 year old) writer had the longest pause times of the five writers and paused for as long as 13–18 s at times. According to Alamargot et al. (2010) the longer pauses were due to a strategy known as step-by-step production of text, where the child switches between planning and formulation of the text to cope with the cognitive demands of handwriting. Indeed as the level of writing expertise increased (22+ years), the number of longer pauses decreased substantially. In the current study, this was also found to be the case, as a relationship was found between age and duration of pauses above 10 s. However using eye-tracking Alamargot et al. (2010) were able to investigate the longer pauses based on gaze fixations. These were classified based on whether the participant was looking back at text, looking at the handwriting area or looking away from the task. They found that the least experienced writers were inclined to look away from the task, which according to Alamargot et al. (2010) was an indication of planning. In the current study, there was a significant group difference in longer pauses above 10 s, as children with DCD not only had more pauses above 10 s, but also paused for longer than the TD group when doing so. This could suggest that the difference in the groups lay in the fact that the handwriting skill in the children with DCD was not automatic enough to concurrently process higher-level writing components. Instead, the DCD group may have been forced to take longer pauses to plan the next phase of text. However, Chang and Yu (2010) suggested that a decrease in strength and endurance was a possible contributor to poorer handwriting control in children with DCD. This is possibly supported by our finding that the children with DCD paused more above 10 s within ideas, particularly a sentence, rather than before starting a new topic. If they were planning the content in a similar way to that of the 7th grader in Alamargot et al.’s (2010) study, they perhaps would have been more inclined to pause before a new topic rather than within a sentence or current idea. Further research needs to be done to investigate this in more detail in an effort to rule out physiological factors such as fatigue. However, practical implications follow from this particular finding as it provides some empirical evidence to support additional time for students with DCD to complete examinations. This type of support is termed ‘access arrangements’ in the UK, where schools are allowed to offer the child extra time to complete exams, if they satisfy specific criteria, usually in relation to handwriting (Department for Education (DfE), 2013). Another interesting finding was that the TD group spent 46% of their pause time within a time-frame of 250 ms to 2 s compared to 37% in the DCD group. In the TD group a positive relationship between age and pausing within this time-frame was found, possibly indicating that as TD children become more experienced at writing, they are able to manage most of the lexical and spelling processes within this time. However, this relationship was not found in the DCD group. This suggests that the DCD group may need longer to process this information given that 14% of their pause time was above 10 s compared to 5% in the TD group. This difference in the distribution of pauses suggests that while the TD group had the ability to process the lexical and spelling components within a range of 250 ms to 2 s, the DCD group were unable to do so and were forced to take longer pauses as a result. From a practical perspective this highlights the need for handwriting intervention in children with DCD, where they are given an opportunity to acquire competency in the skill of handwriting so they can process information on-line, in a similar way to that of their TD peers. Another area addressed in this study was the overall frequency of pausing. According to Alamargot et al. (2010) the frequency of pauses was not an indication of a less experienced writer. In fact, in their study the expert author exhibited the greatest number of pauses, but the key factor lay in their duration. The excessively frequent pausing in the expert author did not have a costly effect on the writing, as they were short pauses allowing for on-line processing to occur quickly. In the current study there were no group differences in the frequency of pauses above 250 ms indicating that the costly effect of the pausing was not attributed to pausing frequency, but due to their duration. The only exception for this was in pauses over 10 s, where the DCD group did pause more frequently in this time-frame. Regression analyses indicated that the frequency of pauses over 10 s was predicted by both the age of the child and the manual dexterity score. Given the role of manual dexterity in the frequency of pauses over 10 s, the role of fatigue should be investigated in future research. The final area for consideration was the overall mean-pause-duration. Rosenblum and Livneh-Zirinski (2008) found that children with DCD had a longer mean-pause-duration than TD peers. However, without knowing how a pause was defined it is difficult to interpret their findings. In a more specific instance, Alamargot et al. (2010) found that less experienced writers exhibited a longer mean-pause-duration than the more experienced authors. In the current study, the DCD group had a longer mean-pause-duration than their TD peers. However, it seems that this may have been driven by the pauses over 10 s. Similarly to Alamargot et al. (2010) a negative relationship with age was found for mean- pause-duration for both groups, indicating that as children get older and possibly more experienced at handwriting, the duration of the pauses decrease. The possible limitations of the current analysis lay in the lack of clarity in the writing literature to justify using particular time-frames to examine pauses. Although this study went some way in categorising the pauses into timeframes, it is still unclear given the novel nature of pausing analysis what the pauses actually represent. Further research needs to be conducted in the field of writing to support this type of analysis. Within the available research on pauses in writing, it is generally accepted that longer pauses are capturing higher-level writing processes (Alamargot et al., 2010; Olive et al., 2009), while shorter pauses reflect transcription (Alamargot et al., 2010). In the present study a range of time-frames were used to capture the pausing patterns of children with DCD in an attempt to characterise them in a way that has not been done before. Further research is needed on handwriting in DCD to provide insight into the longer 10 s pauses in children with DCD, which seems to be the more influential time-frame emerging from this study."],["This study tested whether overimitation is subject to an audience effect, and whether it is modulated by object novelty. A sample of 86 4- to 11-year-old children watched a demonstrator open novel and familiar boxes using sequences of necessary and unnecessary actions. The experimenter then observed the children, turned away, or left the room while the children opened the box. Children copied unnecessary actions more when the experimenter watched or when she left, but they copied less when she turned away. This parallels infant studies suggesting that turning away is interpreted as a signal of disengagement. Children displayed increased overimitation and reduced efficiency discrimination when opening novel boxes compared with familiar boxes. These data provide important evidence that object novelty is a critical component of overimitation. --------------------------------------------------------------------------------","Children are predisposed to copy the actions of others with high fidelity even when they are visibly unnecessary (Lyons, Damrosch, Lin, Macris, & Keil, 2011; Lyons, Young, & Keil, 2007). Strikingly, this “overimitation” is pervasive, occurring when children are directly instructed to complete only necessary actions (Lyons et al., 2007) and when unnecessary actions are performed on simple familiar objects (Marsh, Ropar, & Hamilton, 2014). Despite a decade of research, there is little consensus on why children engage in overimitation. Social signaling theory suggests that overimitation is akin to mimicry and serves as a signal to others, conveying likeness or willingness to interact. Consistently, children overimitate more in scenarios that have increased social relevance to them (Marsh et al., 2014; Nielsen, 2006; Nielsen & Blank, 2011; Nielsen, Simcock, & Jenkins, 2008; Over & Carpenter, 2009). However, evidence suggests that overimitation also occurs in the absence of social drivers (Whiten, Allan, Devlin, Kseib, & Raw, 2016) and regardless of whether irrelevant actions are demonstrated communicatively or noncommunicatively (Hoehl, Zettersten, Schleihauf, Grätz, & Pauen, 2014). Failures in causal encoding may also play a role in overimitation (Lyons et al., 2007, 2011), and overimitation increases with task opacity (Burdett, McGuigan, Harrison, & Whiten, 2018). However, causal misunderstanding is unlikely to be the sole determinant given that overimitation increases with age and into adulthood when causal reasoning is fully matured (Marsh et al., 2014; McGuigan, Makinson, & Whiten, 2011; Whiten et al., 2016). Alternatively, overimitation could reflect a bias to generate and defer to normative rules when observing intentional actions (Kenward, 2012; Kenward, Karlsson, & Persson, 2011; Keupp, Behne, & Rakoczy, 2013; Schmidt, Rakoczy, & Tomasello, 2011). However, it remains unclear whether children defer to norms because they are driven to signal their similarity to others or because they find it intrinsically rewarding to do so regardless of whether a social signal is sent. Evolutionary accounts propose that overimitation is adaptive; by copying when uncertain and refining one’s behavioral repertoire later, imitation serves dual functions of learning about the causal properties of objects in addition to learning social conventions (Burdett et al., 2018; Wood et al., 2016). Each of these theories is supported by a strong set of studies but is also refuted by others, leading to an empirical impasse. A potential explanation for this lack of consensus is that two key features of experimental paradigms used to study overimitation have varied between studies: the audience during the response phase and the complexity of objects used in overimitation tasks. This study sought to systematically manipulate these factors within a single study to examine their impact on overimitation. People change their behavior under conditions in which they feel like they are being observed, and this audience effect has been linked to a change in self-focus or reputation management (Bond & Titus, 1983). Audience effects have been studied in many domains, but recently there has been an increased focus on audience effects as a marker of reputation management (Izuma, Matsumoto, Camerer, & Adolphs, 2011) or social signaling (Hamilton & Lind, 2016). If overimitation is a signaling phenomenon, then it should be modulated by the presence of an audience because it is not worth sending a signal if there is no audience available to perceive it. However, if overimitation reflects a learning process (either causal or normative rules), then the demonstration phase of the study when children gain new information about the task is critical. If children extract a causal or normative rule from the demonstration, then they will faithfully replicate the demonstration regardless of their audience. There are methodological differences in previous overimitation studies regarding whether the participants are observed during the response phase. In some studies children were directly observed by the demonstrator (Nielsen, 2006; Nielsen & Blank, 2011) or by a separate experimenter (Burdett et al., 2018; Marsh et al., 2014; Wood et al., 2016), but in other studies the demonstrator turned his or her back on the children during their response (Keupp et al., 2013) or left the children entirely alone (Hoehl et al., 2014; Kenward et al., 2011; Lyons et al., 2007, 2011; Schleihauf, Graetz, Pauen, & Hoehl, 2018). To date, no single study has directly compared these conditions. In this study, we directly compared the rates of overimitation when children are alone, when they are in the presence of a demonstrator who turns his or her back on the participants, or when their actions are directly observed. A second research question relevant to overimitation is the extent to which the type of object used in a given study influences our estimates of overimitation. Tasks used to elicit overimitation vary with regard to the type of objects and tools that are used, ranging from simple familiar objects to complicated puzzle boxes (see Marsh et al. (2014) and Taniguchi & Sanefuji (2017) for discussions), although the puzzle box designed by Horner and Whiten (2005) has dominated the field. Traditionally, overimitation was demonstrated by comparing rates of imitation on transparent and opaque puzzle boxes under the assumption that the causal properties of a puzzle box are apparent if it is transparent. Indeed, research suggests that imitation is more prevalent when interacting with an opaque puzzle box compared with an otherwise identical but transparent box (Burdett et al., 2018; Horner & Whiten, 2005), although manipulating the opacity of the reward container had no effect (Schleihauf et al., 2018). These studies used novel puzzle boxes, but we posit that encountering any novel object is likely to cause some uncertainty about the way in which it is operated regardless of its physical transparency. This uncertainty may lead to increased overimitation (Rendell et al., 2011; Wood et al., 2016). Here we examined the effects of object novelty on overimitation and uncertainty by directly comparing overimitation on matched novel and familiar boxes while also examining children’s understanding of the efficiency of the actions the children witness. If a “copy when uncertain” bias is present, then we predict reduced efficiency discrimination for novel objects and a corresponding increase in overimitation.","A sample of 86 4- to 11-year-old children were randomly assigned to one of three experimental conditions (see Table 1). The sample was recruited and tested at the University of Nottingham Summer Scientists event, which attracts middle-class families from a mid-sized city in England. Stimuli Two sets of six puzzle boxes were used: a novel set and a familiar set. Each box was a simple transparent container with a removable lid; there were no hidden mechanisms or latches. In the familiar set, the boxes were not modified further. In the novel set, identical boxes were slightly modified to create a simple box that participants had not encountered before. Buttons, switches, or additional decorations were added to each box (see Fig. S1 in online supplementary material). Importantly, none of these decorations changed the function of the boxes, but we anticipated that these decorations would affect children’s certainty about how the objects should be operated. A small toy was put inside each box for children to retrieve. Design ~~~~~~ This study adopted a 3 × 2 mixed design, with children randomly assigned to one of three between-participants audience conditions: Act Alone, Disengagement, or Audience. Object novelty was manipulated within participant. To rule out poor memory as an explanation for why young children overimitate less, a memory control task was included. Four of the six boxes were selected to be the overimitation trials (two novel and two familiar), and two boxes were selected to be memory trials (one novel and one familiar). Boxes were counterbalanced for novelty and task between participants.","Testing took place in a partitioned section of a room introduced to children as a “den.” Poster boards and colored fabric were arranged such that the den was not visible to anyone waiting outside. A hidden camera was positioned behind a hole in the fabric wall to record the sessions without children’s awareness (see Fig. S2). Children were tested alone, with parents waiting in the room outside. They sat at a small table, opposite the experimenter, and completed a warm-up task (see supplementary material) before completing three experimental tasks in a fixed order. Overimitation task The experimenter demonstrated a sequence of three actions (two necessary and one unnecessary) to open the box and retrieve the object (see supplementary material for verbatim instructions and details of the actions). The experimenter then reset the box behind a screen and handed it to the child with this instruction: “When I say ‘GO,’ I would like you to get the [duck] out of the box as quickly as you can.” In the Act Alone condition, the experimenter left the den, shut the door behind her, and called “GO” to the child to signal the start of his or her turn. In the Disengagement condition, the experimenter turned around in her seat, called “GO,” and sat still facing away from the child until the box had been opened. In the Audience condition, the experimenter called “GO” and continued to sit and watch the child while he or she retrieved the object. This sequence of events was repeated for four overimitation trials. Memory task Children completed a warm-up copying task before watching the experimenter open two more boxes (see supplementary material). The memory trials were completed under the same audience conditions as the overimitation trials, so the child was instructed, “When I say ‘GO,’ can you get the [duck] out of the box. Remember to copy me exactly.” Efficiency discrimination task Children rated the efficiency of one necessary action and one unnecessary action from each trial on a scale from 1 (very sensible) to 5 (very silly) as described in Marsh et al. (2014) (see also supplementary material). Efficiency discrimination scores were calculated by subtracting the unnecessary action rating from the necessary action rating on each trial. This score could range from −4 (poor efficiency discrimination) to +4 (good efficiency discrimination), with a zero (0) score indicating no discrimination. Data coding and analysis ~~~~~~~~~~~~~~~~~~~~~~~~ All responses were coded from video. Overimitation on each trial was coded as 1 if the child made a definite and purposeful attempt to replicate the unnecessary action described in Table S1 in the supplementary material or was coded as 0 otherwise. The same criterion was applied to the memory trials, and a total memory score was calculated for each child. Preliminary analyses showed that overimitation did not vary as a function of trial order, F(3, 343) = 1.85, p = .138, or puzzle box, F(5, 343) = 0.421, p = .834), so these variables are not considered further. Mixed-effects models were run using the lme4 package (Bates, Mächler, Bolker, & Walker, 2015) and the MuMIn package (Barton, 2013) in R (Version 3.4.2). Separate models were used to predict propensity to overimitate and efficiency discrimination scores. For overimitation, a full model was constructed that included predictors of interest (audience, novelty, and efficiency discrimination) and control predictors (age, gender, and memory) as fixed effects plus an audience by novelty interaction. Random intercepts for child ID were included to account for the nested structure of the data. The full model was compared with a null model that included only control predictors and random effects using a likelihood ratio test. If the full model outperformed the null model (i.e., a significant difference in model fit), then a reduced model (full model minus the interaction term) was compared with the full model to ascertain whether the interaction term significantly contributed. Efficiency discrimination scores were analyzed using the same protocol. The full model included predictors of interest (age and novelty), control predictors (gender and memory), and random effects (child ID). The full model was compared with a null model (control predictors + random effects) using a likelihood ratio test.","Children in each of the three experimental conditions were matched for age (see Table 1). The rate of overimitation was high, with 80.2% of children overimitating on at least one trial. Memory scores were also high, with 73.3% of children performing at ceiling. Only 3.5% of children scored 0 (see Table 1). The reduced model was the best fit to the overimitation data, explaining 12.0% of the variance by fixed effects and 69.3% of the variance by random effects (see Tables S3 and S4 for model summaries and comparisons). Audience was a significant predictor of overimitation, χ2(2) = 7.85, p = .020. Children in the Disengagement condition (M = 1.83, SD = 1.64) overimitated less than those in the Audience condition (M = 2.77, SD = 1.43, odds ratio = 7.82, 95% confidence interval [CI] = [1.63–50.66]) and Act Alone condition (M = 2.77, SD = 1.42, odds ratio = 6.34, 95% CI = [1.23–42.68]). There was no difference between the Audience and Act Alone conditions (odds ratio = 1.23, 95% CI = [0.22–7.51]). Novelty also significantly predicted overimitation, χ2(1) = 6.04, p = .014, such that unnecessary actions on familiar objects (M = 1.13, SD = 0.88) were imitated less frequently than unnecessary actions on novel objects (M = 1.31, SD = 0.88, odds ratio = 0.46, 95% CI = [0.24–0.86]) (see Fig. 1). Age, gender, efficiency discrimination, and memory did not predict overimitation (see Table S3). There was no interaction between audience and novelty, χ2(2) = 0.20, p = .905, indicating that novelty had the same effect on overimitation behavior regardless of audience. The full model best predicted the efficiency discrimination data, explaining 10.6% of the variance with fixed effects and 54.2% with random effects (see Tables S4 and S5 for model comparisons and model summaries). Age predicted efficiency discrimination, χ2(1) = 13.33, p < .01, such that the older children were better at discriminating necessary and unnecessary actions compared with the younger children (odds ratio = 1.68, 95% CI = [1.28–2.20]). Novelty also predicted efficiency discrimination, χ2(1) = 4.30, p = .038. Children were worse at discriminating the efficiency of necessary and unnecessary actions when objects were novel (M = 2.20, SD = 1.85) compared with when they were familiar (M = 2.44, SD = 1.88, odds ratio = 1.26, 95% CI = [1.01–1.57]). There was no effect of gender or memory on efficiency discrimination scores (see Table S5).","The effect of audience on overimitation was assessed for novel and familiar objects across a broad developmental spectrum. We demonstrate a clear effect of audience, but not an increase with increasing level of observation (i.e., Audience > Disengagement > Act Alone). Instead, we report similar levels of overimitation when children were observed and when they were alone, but reduced imitation when the demonstrator turned away. This intriguing set of results can be explained in two ways. First, we could argue that an audience effect was found in the Audience > Disengagement comparison but that some distinctive feature of the Act Alone condition disrupted this pattern. High levels of overimitation in the Act Alone condition could be explained by an “omniscient adult phenomenon” or “Monika effect,” whereby children falsely attribute knowledge to unseen adults (Wimmer, Hogrefe, & Perner, 1988). When children were left alone, perhaps they were uncertain about whether they were being observed and, by default, acted as though they were. This mirrors other findings in developmental psychology (Meristo & Surian, 2013; Rubio-Fernández & Geurts, 2013). Children expect a third party to act consistently with the children’s knowledge even if there is no evidence that the third party shares this knowledge. However, if children are provided with direct evidence that a third party does not share the same knowledge, they will predict behavior based on the agent’s knowledge. For example, Rubio-Fernández and Geurts (2013) demonstrated that 3-year-olds passed a standard false-belief task when the protagonist turned her back (giving direct evidence that the protagonist could not see the location change) but failed when the protagonist left the scene entirely (see also Meristo & Surian, 2013). Perhaps this bias extended to the children in our study. When children had direct evidence that they were not being watched (Disengagement condition), they reduced their overimitation. However, when there was no such evidence (Act Alone condition), they assumed that they were observable and acted similarly to those children who were directly observed (Audience condition). These findings hint toward an interesting distinction in our processing of others’ minds when those individuals are physically present, but not watching, and when they are completely absent. An alternative interpretation is that there was a reduction in overimitation in the Disengagement condition because the experimenter’s actions caused the children to feel less rapport or motivation to engage. By turning her back on the children as they acted without excusing herself, the experimenter gave a strong signal of disinterest that could be interpreted as ostracism (Wirth, Sacco, Hugenberg, & Williams, 2010). As a result, the children may experience reduced rapport with the experimenter and, thus, a reduced drive to overimitate (Nielsen, 2006). This is contrary to several other findings indicating that children actually increase their imitative fidelity following exposure to third-party experience of ostracism (Over & Carpenter, 2009; Watson-Jones, Legare, Whitehouse, & Clegg, 2014) and first-person experience of ostracism (Watson-Jones, Whitehouse, & Legare, 2016). However, previous work primed ostracism indirectly via computer animations, which do not directly depict the demonstrator, prior to the imitation task. It is possible that disengagement from a live model during the interaction has the opposite effect on imitation due to either the proximity or time course of ostracism. Further work is needed to disentangle these effects. A neater direct test of the ostracism account could be to compare rates of imitation when the experimenter excuses herself and turns around to complete a task with those when the experimenter simply disengages without excuse (as in this study). Both interpretations are consistent with the signaling theory of overimitation, in which children are motivated to send a signal to people who are watching and with whom they have a rapport. The data may also be consistent with a normative account in which children opt to disregard the newly learned norm following ostracism, although further research examining the flexibility of norm adherence is required to support this argument. However, it is not clear how causal encoding can account for the differences among the three social conditions in this study. Another striking finding was that children were more likely to overimitate and were less able to discriminate the efficiency of actions when interacting with a novel box compared with a familiar box. Thus, it seems that altering the perceived novelty of the boxes reduced children’s certainty about the causal properties of the objects, leading children to overimitate more even though only very minor decorative changes distinguished familiar and novel boxes. This is consistent with emerging work that illustrates increased overimitation when task complexity increases (Taniguchi & Sanefuji, 2017). These results reflect an element of causal understanding in any overimitation task, and using novel objects can contaminate social effects. Alternatively, it is possible that the children in this study interpreted the novel boxes in this task as more playful, which led them to copy more and rate the unnecessary actions as less “silly.” However, given that the object novelty manipulation was presented within participant and in a randomized order, it is unlikely that the children interpreted the exact same instructions differently on each trial. Regardless, previous studies have varied in the use of transparent and opaque puzzle boxes, with some including redundant mechanisms, hidden catches, and superfluous decorations. This lack of consistency increases response variability and could account for the disparity in results from different labs. We stress that future work examining the social effects of overimitation should carefully evaluate the findings with regard to the type of objects that have been used to elicit overimitation. To conclude, this study provides evidence that observation of participants during their response and stimulus familiarity can affect overimitation. These factors may account for discrepancies among previous studies. The finding that children reduce overimitation when the experimenter disengages is consistent with social signaling and social rapport accounts of overimitation, and we look forward to further studies that distinguish these motivating factors."],["Can personality traits predict willingness to fight or even die for one's heritage culture group? This study examined insecure attachment dimensions - avoidance and anxiety - as predictors of perceived rejection from heritage culture members and, in turn, greater endorsement of extreme pro-group actions. Expressing extreme commitment for the heritage culture may represent an attempt by insecure individuals to reduce their perceived marginalisation and reaffirm their heritage culture membership and identity. Participants completed measures of attachment dimensions, intragroup marginalisation, and endorsement of extreme pro-group actions. Individuals who were high in anxiety or avoidance reported heightened intragroup marginalisation from family and friends. In turn, friend intragroup marginalisation was associated with increased endorsement of pro-group actions. Our findings provide insight as to why insecurely attached bicultural individuals may be drawn to endorse extreme pro-group actions. --------------------------------------------------------------------------------","At approximately 8:50 am on July 7th, 2005, a series of explosions in London killed 52 people and left more than 800 injured. For the first time in modern terrorism, the threat was not wholly external – all four men responsible were British citizens who had been integrated into the mainstream culture. On the surface, they did not appear excluded from the mainstream culture, as one might expect from their radicalisation (Silber & Bhatt, 2007). While there has been much speculation about their attitudes towards their British mainstream culture, comparatively little attention has been given to their interactions and identification with their heritage cultures (BBC News, 2005). By the same token, the heritage culture experiences of the estimated 3000–4000 bicultural individuals with an EU nationality who have travelled to Syria to fight with ISIS as part of their radicalisation have also received little attention (Traynor, 2014). These examples suggest that radicalised individuals who are citizens of Western countries may be motivated in part by the struggle to find an identity (Silber & Bhatt, 2007). This struggle may stem from the extent to which one feels accepted by their mainstream and heritage cultures. Mainstream culture is defined as the dominant culture where one currently lives (Berry, 2001). Heritage culture is defined as the culture of one's birth or upbringing, or the culture that had a significant impact on previous generations of one's family. Conceptualisations of marginalisation remain focused on exclusion by members of the mainstream culture and an enforcement of culture loss (e.g., Kosic, Mannetti, & Sam, 2005); fewer studies have examined the influence of perceived exclusion by heritage culture friends and family, in spite of its importance for maintaining one's identity. Our study aimed to address this research gap by investigating the association of perceived marginalisation from heritage culture members with extreme pro-group actions, defined as willingness to commit, fight, or even die in aid of one's heritage culture. Thus, this research may shed light on some of the reasons why Westernised bicultural individuals might be drawn to joining extremist groups as a compensatory reaction in response to perceived rejection from their heritage culture. What role does perceived rejection play in the construction of our identity? Humans share a fundamental need to form meaningful interpersonal attachments (Baumeister & Leary, 1995). Rejection from other heritage culture members – defined as intragroup marginalisation – can occur when individuals develop ties to two or more cultures, and as a result, no longer conform to the expectations of the heritage culture identity (Castillo, Conoley, Brossart, & Quiros, 2007). These detrimental impacts and experiences of intragroup marginalisation may be shaped by personality; those who are insecurely attached and chronically perceive rejection report increased intragroup marginalisation (Ferenczi & Marshall, 2014). A compensatory response to intragroup marginalisation may be to reaffirm one's heritage culture identity through endorsing pro-group actions. In an effort to gain acceptance and avoid rejection, do insecurely attached individuals endorse pro-group actions that are extreme? Attachment ~~~~~~~~~~ Attachment is conceptualised as an internal working model of self and others that informs their interactions over the course of their life (Bowlby, 1969). Views of self and other are essential for constructing bicultural identity and perceived rejection from in-group members. Secure attachment is typified by an internalised positive model of the self; that is, one feels worthy of love, and also a positive model of ‘other’, as significant others are thought of as being available and trustworthy (Bartholomew & Horowitz, 1991). Secure attachment is conceptualised as low anxiety and avoidance (Mikulincer, Shaver, Sapir- Lavid, & Avihou-Kanza, 2009). Anxious individuals have a negative model of self, and tend to be preoccupied with winning affection from others (Mikulincer, 1998). They endorse positive models of other, which results in the individual feeling unworthy of love (Mikulincer, 1995). For those high in anxiety, the attachment system is hyper-activated in response to perceived rejection threats (Campbell & Marshall, 2011). Anxious individuals are sensitive to rejection, recalling emotionally painful memories with ease whilst unable to repress the resulting negative effects (Mikulincer & Orbach, 1995). Their difficulty in recovering from past experiences of rejection (Marshall, Bejanyan, & Ferenczi, 2013) may generalise to intragroup marginalisation. Conversely, individuals who are avoidant perceive others as untrustworthy and unreliable, and hold positive views of their self, resulting in exaggerated self-reliance (Li & Chan, 2012). Avoidant individuals engage in deactivating strategies in response to threat, such as moving away from attachment figures and suppressing emotions to pre-empt the frustration and pain arising from rejection (Shaver & Mikulincer, 2002). Although they may appear to have high self-esteem (Mikulincer, 1998), it may be little more stable than a house of cards. Highly distressing events can unearth anxiety (Mikulincer & Orbach, 1995) and result in difficulties coping with rejection (Birnbaum, Orr, Mikulincer, & Florian, 1997). Although they may be adept at suppressing negative impacts of mild threats, they nonetheless experience heightened psychological distress in response to stress (Stanton & Campbell, 2013). Despite their defences, they report a need to belong to close others (Carvallo & Gabriel, 2006). We argue that perceiving rejection from close others qualify as severe threats. Thus, we expected that individuals high in avoidance would also perceive greater intragroup marginalisation from close others such as family and friends. Intragroup marginalisation ~~~~~~~~~~~~~~~~~~~~~~~~~~ Social rejection can be conceptualised as a social death (Williams & Nida, 2011). The negative effects of rejection remain even if it is merely perceived (Smith & Williams, 2004). Individuals who perceive rejection in the form of intragroup marginalisation may face accusations of betraying their heritage culture, such as by assimilating into the mainstream culture (Castillo, Zahn, & Cano, 2012). They may perceive that family and heritage cultural friends view them as threatening the distinctiveness of the cultural group through deviating from the prescribed social identity (Castillo et al., 2007). Thus, no longer meeting the expectations of the heritage culture, individuals may feel rejected, regardless of their own wishes to maintain their heritage culture identity (Ferenczi & Marshall, 2014). We hypothesised that those individuals high in anxiety or avoidance, who have a heightened sensitivity to rejection (Shaver & Mikulincer, 2002), would report greater intragroup marginalisation (Ferenczi & Marshall, 2014). In turn, what compensatory actions would they endorse in striving for acceptance and positive perceptions of their self from their heritage culture in-group? Extreme pro-group actions ~~~~~~~~~~~~~~~~~~~~~~~~~ Individuals strive for others to perceive them as they do themselves (Swann, 1983). In fact, we engage in a continuous construction of ourselves that draws the feedback from close others (Swann & Brooks, 2012). Reflected appraisals – perceptions of how others perceive oneself – play an important role in constructing identity, in particular for individuals who may experience ambiguity resulting from a dual identity (Khanna, 2004). If there is an indirect threat to one's opportunity to self-verify, then they may engage in compensatory self-verification to re-establish coherence (Swann & Brooks, 2012). Self- verification can occur at the level of the collective self – the evaluation of the self in relation to one's in-group (Chen, Chen, & Shaw, 2004). In the context of the current study, if a British Asian woman perceives herself as Punjabi, yet finds her family criticising her Punjabi language skills, then she might come to question her knowledge of herself. To avoid threats to the very foundation of her identity, what can she do? We hypothesised that those who experience intragroup marginalisation will self-verify by endorsing extreme pro-group actions, in the hope of reaffirming their heritage culture identity. Thus, by supporting attitudes which are extreme, individuals can demonstrate their loyalty and commitment. Our study is, to our knowledge, the first to link attachment, intragroup marginalisation, and extreme pro-group actions. We hypothesised that insecure attachment would be linked with intragroup marginalisation, and, in turn, with greater endorsement of extreme pro-group actions.","208 participants (Mage = 30.29, SD: 11.74; female: 105, male: 100; missing: 2, transgender: 1) completed the measures. Inclusion criteria for the study required each participant to have a different heritage and mainstream culture (i.e., they were a first- or later-generation migrant). 49% of participants reported that they were first-generation migrants (Myears residing in mainstream culture = 11.27, SD: 8.47); 51% were born and raised in a mainstream culture that was different to their heritage culture (second- or later-generation migrants). Participants reported the following heritage cultures: European (23%), Latin American (18%), East Asian (17%), South Asian (9%), Southeast Asian (9%), Middle Eastern/North African (8%), African (7%), Jewish (3%), Native American/First Nations (3%), Caribbean (1%), Mixed (1%), and North American (1%). The majority of participants reported living in a North American mainstream culture (83%); they also reported living in Europe (15%), East Asia (1%), and the Middle East (1%). The majority of participants reported being in a relationship (65%). Participants were recruited online via Amazon MTurk (paid $0.30), or through the Social Psychology Network (no reward). As an attention-check measure, we asked participants to report the date, and compared this with their timestamp. All materials were in English. Berkeley personality profile Neuroticism refers to emotional instability and correlates with increased insecure attachment (Shaver & Brennan, 1992). We included seven items from the neuroticism subscale (Harary & Donahue, 1994; α = .82; e.g., “I worry a lot”; 1 = Disagree strongly, 5 = Agree strongly) to establish that the association of insecure attachment with intragroup marginalisation could not be attributed to neuroticism. Experiences in close-relationships – short form (ECR-S) The ECR-S (Wei, Russell, Mallinckrodt, & Vogel, 2007) is a twelve-item scale (1 = Strongly disagree, 5 = Strongly agree) that measures general attachment style to a partner. It is composed of two subscales: anxious (α = .69; “My desire to be very close sometimes scares people away”) and avoidant attachment (α = .80; “I am nervous when partners get too close to me”). Intragroup marginalisation inventory (IMI) We included two subscales from the IMI (Castillo et al., 2007): one focused on rejection from family members (11 items, α = .84; “Family members laugh at me when I try to speak my heritage/ethnic culture group's language”), and one focused on rejection from heritage culture friends (16 items, α = .91; “Friends of my heritage culture group tell me that I am a ‘sell-out’”; 1 = Never/Does not apply, 7 = Extremely often). Perceived rejection can be in the form of criticism, mocking, or perceived differences between the self and family/friends on dimensions that are important to heritage identity. Extreme pro-group actions Swann, Gómez, Seyle, Morales, and Huici (2009) developed this scale to investigate participants' willingness to endorse extreme pro-group actions, using it as a proxy for extreme pro-group behaviour (Swann, Jetten, Gómez, Whitehouse, & Bastian, 2012). By measuring the level to which individuals endorse such actions, we could highlight one of the ways in which individuals may choose to self-verify in the face of rejection. The subscale measuring willingness to fight for the group was composed of five items (e.g., “I would fight someone insulting or making fun of my heritage country as a whole”). A two-item subscale centred on participants' willingness to die for their heritage culture (e.g., “I would sacrifice my life if it gave the heritage culture group status or monetary reward”). We created three items to extend the range of pro-group actions to include increased commitment to the heritage culture (“The needs of my heritage culture group come before my own”, “I would be willing to attend a protest if my heritage culture group were threatened”, and “I would be willing to donate money to organisations promoting interests of my heritage culture group if need be”). We collapsed the ten items into one overall measure of extreme pro-group actions (α = .91; 1 = Totally disagree, 7 = Totally agree). Descriptive statistics Descriptive statistics and Pearson's correlations are reported in Table 1. We ran preliminary hierarchical regressions controlling for gender (− 1 = male, 1 = female), age, cultural background (− 1 = 2nd/+ generation migrants, 1 = 1st generation migrants), and neuroticism. Results revealed that neuroticism was positively correlated with marginalisation from family, β = .28, p < .005. Over and above the control variables, avoidant and anxious attachment were linked with increased marginalisation from friends (β = .27, p < .001 and β = .19, p < .05, respectively) and family (β = .23, p < .005, and β = .15, p = .05, respectively). Anxious and avoidant attachment were correlated with greater endorsement of extreme pro-group actions, (β = .19, p < .05 and β = .19, p = .05, respectively). Friend intragroup marginalisation positively correlated with extreme pro- group actions, β = .35, p < .005. We then used structural equation modelling to investigate the pathways between anxious and avoidant attachment, family and friend intragroup marginalisation, and endorsement of extreme pro-group actions.1 Hu and Bentler (1999) recommended that for an acceptable model fit, the following requirements must be met: the chi-square statistic should be non-significant, the comparative fit index (CFI) should be .95 or greater, and the root mean square error of approximation (RMSEA) should be .06 or less. We used item parcelling as it requires estimation of fewer parameters and thus results in a more parsimonious and stable model (Little, Cunningham, Shahar, & Widaman, 2002). Items were assigned to one of two parcels at random for each latent variable (Little et al., 2002); this method provides a superior model fit in comparison to others (Landis, Beal, & Tesluk, 2000). Measurement model ~~~~~~~~~~~~~~~~~ We conducted a Full Information Maximum Likelihood (FIML) procedure in AMOS to include partially missing data. The measurement model provided a good fit to the data [χ2(25) = 33.41, p > .05, CFI = .99, RMSEA = .04 (CI = .00, .07)]. All of the indicators loaded significantly onto their respective latent variables (all βs ≥ .58, p < .001). Structural model ~~~~~~~~~~~~~~~~ A fully saturated structural model was tested initially; we included structural covariances between anxious and avoidant attachment, and family and friend intragroup marginalisation. Two structural coefficient pathways were non-significant: the pathway between avoidant attachment and extreme pro-group actions, and the pathway from family intragroup marginalisation to extreme pro-group actions. To create a more parsimonious model, we removed these pathways in order of lowest standardised regression weights. Chi- square difference tests did not yield significant differences in model fit, respectively: [χ2D(1) = .09, p > .05], and [χ2D(2) = .65, p > .05]. The final model (see Fig. 1) provided a good fit [χ2(27) = 34.06, p > .05, CFI = .99, RMSEA = .04 (CI = .00, .07)].2 Tests of indirect effects ~~~~~~~~~~~~~~~~~~~~~~~~~ Bootstrap procedures in AMOS3 were used to test the indirect effects of anxious and avoidant attachment on endorsement of extreme pro-group actions via friend intragroup marginalisation. Inspection of 95% bias-corrected confidence intervals (CI) from 1000 bootstrap samples found support for the indirect effects of anxious [β = .09, p < .005 (CI: .03, .15)] and avoidant [β = .14, p < .005 (CI: .18, .25)] attachment on endorsement of extreme pro-group actions via increased friend intragroup marginalisation. Attachment and intragroup marginalisation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our findings contribute to understanding the wide-reaching impact of attachment dimensions on chronic tendencies of perceiving rejection from others. Anxious individuals may perceive greater intragroup marginalisation because they are sensitive to rejection (Campbell & Marshall, 2011) and perceive themselves as having poor capabilities (Mikulincer & Florian, 1995) in their heritage culture. For example, they may feel that they are not proficient in their heritage culture language. The parallel findings for individuals high in avoidance provide further support for the fragility of their positive self-view (Mikulincer, 1995). Although avoidant individuals claim to be highly self- reliant, they still experience the need to belong (Carvallo & Gabriel, 2006). They too may perceive pressure and rejection from those in their closest social circle on the basis of their heritage culture identity. The current findings provide a link between the rejection-sensitivity of insecurely attached individuals to endorsement of extreme pro- group actions, partially through experiences of friend intragroup marginalisation. Intragroup marginalisation and extreme pro-group actions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ One reason why only rejection from friends was linked to greater endorsement of pro-group actions may be because of the more fragile nature of friendships compared to family bonds (Allan, 2008). Whilst individuals may feel rejected from their family on the basis of their heritage cultural identity, they may nonetheless perceive acceptance on binding cultural conceptions of ‘blood’. As friendships are largely voluntary and based on equality (Hays, 1988), intragroup marginalisation may be more detrimental for their maintenance. Close friendships are often considered a ‘chosen’ family, fulfilling an individual's needs for belonging and closeness (Wrzus, Wagner, & Neyer, 2012). Thus, individuals may endorse extreme pro-group actions to reaffirm their heritage cultural identity, and in doing so, realign balance in the friendship after perceiving rejection. Through measuring endorsement of, as opposed to actual pro-group actions, we could investigate the responses of individuals who have ostensibly not engaged in extreme actions. Indeed, many of the EU citizens who joined ISIS are young, often well-educated, bicultural individuals with no prior history of criminal offences. Research needs to shift from conceptualising terrorism as a result of personality disorders or irrationality removed from the general populace (Crenshaw, 2000; Kruglanski & Fishman, 2006). Our findings indicate that individuals sampled from a broad demographic may indeed endorse actions that appear extreme. Individuals are more likely to identify with radical groups when they perceive uncertainty in the form of threat to their values and behavioural practises (Hogg, Meehan, & Farquharson, 2010). We posit that individuals who report intragroup marginalisation experience uncertainty in terms of their heritage culture identity. Thus, they may shift to endorse more radical ideas as a method of alleviating uncertainty and re-establishing their membership within the heritage culture group. Limitations and further research ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Future research should seek to further validate these relationships in several ways. First, additional research should focus on priming rejection from one's heritage culture. Second, research should investigate whether the process outlined here is unique to bicultural individuals, or whether monocultural individuals may also perceive generalised rejection from the mainstream culture and, in turn, endorse extreme pro-group actions. However, this process would be different to intragroup marginalisation as monocultural individuals would not experience a tension between two cultural identities. Third, other antecedents of intragroup marginalisation should be measured, such as general negativity towards others. Fourth, future research should investigate the direct association of intragroup marginalisation with attitudes towards militant organisations, such as Al- Qaeda, as well as with overt pro-group behaviours. Finally, although we measured endorsement of extreme behaviour, in part because of the ethical difficulties of measuring actual extreme behaviour, attitudes can predict intentions and behaviour according to the theory of reasoned action (Ajzen & Fishbein, 1980).","Malik (2015) reports that the Westernised individuals who join ISIS are as removed from Muslim communities as from the mainstream cultures in which they live. Our findings provide support that measures which encourage biculturalism are crucial on both the parts of mainstream cultures and of heritage cultural communities to prevent individuals who are struggling to find an identity in perceiving rejection from their heritage culture. Clinical and community interventions which focus on individuals' heritage identities alongside their integration to the mainstream society could ameliorate experiences of rejection which may lead to detrimental attitudes and, ultimately, to tragedies similar to the London bombings."],["Visual information has been observed to be crucial for audience members during musical performances. The present study used an eye tracker to investigate audience members’ gazes while appreciating an audiovisual musical ensemble performance, based on evidence of the dominance of musical part in auditory attention when listening to multipart music that contains different melody lines and the joint-attention theory of gaze. We presented singing performances, by a female duo. The main findings were as follows: (1) the melody part (soprano) attracted more visual attention than the accompaniment part (alto) throughout the piece, (2) joint attention emerged when the singers shifted their gazes toward their co-performer, suggesting that inter-performer gazing interactions that play a spotlight role mediated performer-audience visual interaction, and (3) musical part (melody or accompaniment) strongly influenced the total duration of gazes among audiences, while the spotlight effect of gaze was limited to just after the singers’ gaze shifts. --------------------------------------------------------------------------------","What do audiences look at when they watch and listen to an ensemble music performance? Many studies have explored what factors attract visual attention, when they attract the attention, and how (reviewed in Carrasco, 2011). In particular, gaze has been observed as an indicator of visual attention. However, the gaze of the audience of a musical performance remains unclear even though visual information has a great impact on an audience appreciating an audiovisually presented musical performance (e.g., Platz & Kopiez, 2012). Investigations of gaze in various situations thus contribute to a holistic understanding of human behavior. Taking into account auditory attention, while appreciating audio-visually presented music, a complicated audience gaze is likely to emerge. This assumption is supported by findings that visual attention interacts with auditory attention. For example, in brain activity, the circuit related to auditory attention is linked with the circuit related to vision (Winkowski & Knudsen, 2006). In Driver and Spence’s review of crossmodal attention, auditory stimuli affect visual attention, while visual stimuli do not influence auditory attention (Driver & Spence, 1998). These findings suggest that visual attention is associated with the cognition of musical sound. In the present study, we particularly focused on the gaze of the audience while the audience is appreciating an ensemble performance in order to explore the psychological process of music appreciation as audiovisual experience. An ensemble performance involves multiple musical parts and performers that can potentially divide or integrate the audience’s attentions, as well as inter-performer interaction cues (reviewed in Keller, 2014) that could influence performer-audience visual interactions. Despite such intriguing aspects, fundamental perspectives on audience gaze are insufficient. What attracts the gaze of an audience during the appreciation of a musical performance? Do audiences evenly look at the whole of an ensemble, or do they look at specific performers? If so, why? To find out, we explored gazing behavior during ensemble music appreciation. First, in accordance with the dominance of a musical part in auditory attention when listening to multipart music that contains different melody lines, we hypothesize that auditory attention, depending on the musical part (melody or accompaniment, namely soprano or alto in the present study), influences visual attention. Soprano is the highest female voice (Jander, 1980), while alto is lower than soprano. Ensemble performance often involves multipart that could potentially yield a bias of auditory attention. For example, prior studies have shown a higher-pitched effect, in which higher voices attracted listeners’ attention in multipart music (e.g. Fujioka, Trainor, Ross, Kakigi, & Pantev, 2005). Nevertheless, in audiovisual music appreciation, no study has examined whether such auditory attention corresponds to visual attention. Second, we hypothesized that inter- performer gazing affects an audience’s visual attention. In accordance with the joint- attention theory, we highlighted, in particular, one potential function of the gaze as spotlight, as suggested by Kawase’s study (Kawase, 2009a), in which a singer in a pop band directed her gaze toward a co-performer who played a solo part. Since previous studies have mostly focused on performer-audience visual interactions in solo performances, this hypothesis regarding ensemble performances differs from available findings. In light of the reciprocal interaction model including performer-audience communication flows during musical performances (Hargreaves, MacDonald, & Miell, 2005; Kawase et al., 2007), gaze interactions between performers (Davidson, 2005; Kawase, 2009a; Kawase, 2014a; Kawase, 2014b; Moran, 2010) can influence audiences’ visual attention. This hypothesis is also supported by the joint attention theory that revealed that a person’s visual cue can attract other person’s visual attention (e.g. Baron-Cohen, 1995; Frischen, Bayliss, & Tipper, 2007). If this theory is applicable to musical performances, gazing interactions between performers could influence the visual attention of audiences. Data of audiences’ gazes during musical performances could shed light on additional aspects of joint attention. Third, in the context of audiovisual interactions during multimodal music appreciation (e.g. Platz & Kopiez, 2012), it would be beneficial to investigate how audio and visual attention are associated with each other. We hypothesized that certain audiovisual interactions alter the audience’s gaze. Therefore, we investigated how the audience’s visual attention reacted when auditory attention (e.g. dominance of a musical part) and visual attention (joint attention between the performer and the audience) simultaneously emerge. Given that ensemble musical performance potentially causes attention allocation in both auditory (multipart) and visual (multi-party) aspects, this attempt could lead to an understanding of attention in simulated multimodal situations in the real world. To investigate these issues, we used eye tracking to analyze audience members’ gaze when watching an actual performance by a singing duo. Given that multipart music is often played in an ensemble, and that gaze between performers naturally emerges in music performance, our attempt would be in agreement with ecological validity. Factors that support our hypotheses are reviewed below. Auditory attention in multipart music ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Prior studies have explored how listeners recognize each melody in multipart music. In terms of auditory attention, higher-pitched parts have been observed to attract more listeners’ attention. Gregory’s (1990) study of melody recognition in multipart music found that participants could recognize each melody in multipart music with significant accuracy and that factors such as the relationships and distances between keys and pitches affected recognition; specifically, higher melodies were better recognized than lower ones when there were significant gaps in pitch between melodies. With respect to the relationship of melody and accompaniment part, Uhlig, Fairhurst, and Keller (2013) examined the influence of structural and temporal aspects on auditory attention in multipart music and showed that the melody part was perceived as leading compared to accompaniment part. Furthermore, cognition research has elucidated high-voice superiority in polyphonic music, meaning that listeners pay more attention to higher pitches (Fujioka et al., 2005; Trainor, Marie, Bruce, & Bidelman, 2014), even among very young (three- month-old) children (Marie & Trainor, 2014). By demonstrating the model, Trainor et al. (2014) suggested that such high voice superiority was derived from characteristics of the auditory nervous system. Other studies have explored how ensemble performers integrate multipart melodies by measuring behavioral (Bigand, McAdams, & Forêt, 2000) and brain activity (Ragert, Fairhurst, & Keller, 2014). Although these findings yield important perspectives on perceptions of multipart music, it remains unclear how musical parts (melody and accompaniment) affect the visual attention of audiences during ensemble performances. Thus, the present study examined the influence of musical parts on the gaze of audience members. Joint attention as one of the functions of gazing in everyday communication ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Joint attention represents the sharing of attention among other person(s) in communication when the other person(s) are referring to something. This function’s tendancy to alter attentions has been investigated in daily communication (e.g., Frischen et al., 2007), child development (e.g., Tomasello, 1995), and human-robot communication (Staudte & Crocker, 2011). Before discussing the validity of focusing on joint attention between performers and audience members, joint attention in daily life should be reviewed. Joint attention, one of the fundamental functions of gaze, has been described as a phenomenon in which a person’s gaze direction shifts another person’s attention (e.g., Baron-Cohen, 1995; Frischen et al., 2007). Langton, Watt, and Bruce’s meta-analysis of gaze functions (2000) suggested that joint attention is a phenomenon in which eye information provides substantial cues for another’s gaze direction, and that gaze direction causes reflexive shifts in an observer’s visual attention based on neural mechanisms. Other studies have also reported the “following gaze.” For example, adults attend to objects that others attend to and switch others’ attention through gazing (Shepherd, 2010). Such phenomenon of joint attention can be examined by measuring gaze. Previous studies explored gaze through observation (e.g., Tomasello & Farrar, 1986), eye tracking via a head mount eye tracker, (e.g., Hayhoe & Ballard, 2005), or measurement of gaze on a monitor (e.g., Schilbach et al., 2010). In the present study, we observed gaze on a monitor through an eye tracker. This joint attention theory could predict gaze as a spotlight between performers in ensemble music performance. Kawase (2009a) measured the gazes of a popular music band’s members during a concert and found that the performer’s gaze differed according to instrument and musical structure. Kawase (2009a) suggested that the gaze of the vocalist—a frontwoman performing a central role in the band—functioned to spotlight soloists by attracting joint attention from audience members. In the present study, we examined this assumable function of gaze. The influence of audio-visual interactions on audiences during music performance ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Visual information such as gaze (Antonietti, Cocomazzi, & Iannello, 2009; Kawase, 2009a; Kurosawa & Davidson, 2005), body movement (Broughton & Stevens, 2009), facial expression (Thompson, Russo, & Livingstone, 2010), clothing (Griffiths, 2008), and physical appearance (Wapnick, Mazza, & Darrow, 1998) strongly influence the way audiences decode music performances (Platz & Kopiez, 2012). Visual information has also been observed to influence audience perceptions of expression (Davidson, 1993), emotion (Dahl & Friberg, 2007), tone duration (Schutz & Lipscomb, 2007), and performance proficiency (Tsay, 2013) during musical performances. Audiences also attempt to ensure good visibility of the performers when selecting seats in a concert hall (Kawase, 2013). Meanwhile, the influence of auditory information in audiovisually presented music has been reported for audience judgments of tension during a clarinet performance (Vines, Krumhansl, Wanderley, & Levitin, 2006), emotions evoked by drums and saxophone (Petrini, McAleer, & Pollick, 2010), portrayed age and gender of characters in Gidayubushi, traditional Japanese performing arts (Kawase, 2009b), and expression in opera (Silveira & Diaz, 2014). Given most of these findings of the dominance of audio or visual cues during the appreciation of audiovisually presented music focused on a solo music performance, it would be beneficial to investigate ensemble performance that potentially causes complications in the attention of audience due to its multifaceted auditory and visual factors. Aims and research questions ~~~~~~~~~~~~~~~~~~~~~~~~~~~ The present study aimed to quantitatively investigate audience members’ gaze data during an audiovisually presented performance by focusing on the auditory (melody or accompaniment) and visual aspects (joint attention) of singers. The main research questions were as follows: (1) Do singers’ musical parts (melody or accompaniment, namely soprano or alto in the present study) affect audience members’ visual attention? (2) Does gazing between performers affect audience’s visual attention? (3) How does the audience’s visual attention react when auditory attention (e.g. dominance of the part) and visual attention (joint attention) simultaneously emerge? Research design ~~~~~~~~~~~~~~~ We showed participants a videotaped performance and analyzed their gaze points using an eye tracker.","Thirty people participated in the experiment. We analyzed data for 20 (Mage = 20.8; 12 female and 8 male) due to low data clarity for the other 10 (e.g., the duration of measurement was extremely short due to obstacles such as eyeglass frames). All participants were non-music majors.","We selected a performance by a singing duo as audiovisual stimulus. We made this selection for the following reasons: (1) a duo is the simplest form of multipart (i.e., melody and accompaniment) ensemble, (2) the absence of instruments makes it easy to control gazing in the experiment, and (3) vocal performance is not accompanied by additional body movements—such as hand movements for instrument manipulation—that could potentially attract visual attention. Six audiovisual stimuli of the female duo singing (two singing parts: soprano and alto, in three gaze conditions: unreciprocated gaze, one from each singer, and a third with absence of eye contact) that were generated in the recording experiment were used. To generate audiovisual stimuli, we recorded a two proficient female singing duo that utilized both melody and accompaniment (Mage = 28.5). The singers sang the Spanish ballad “Juanita” (Norton, 1951), which was separated into melody (soprano) and accompaniment (alto). The singers used “na” vocalizations without instrumental accompaniment. The average total duration of each stimulus was about 26 s, which incorporated the first half for 13 s and the latter half for 13 s. To examine gaze shifts according to musical parts, in the middle of the piece where the first half (approximately 13 s from the onset of the performance) has passed and the latter half (the rest of approximately 13 s of the performance) began, the melody and accompaniment parts switched. Thus, the singer assigned the melody part switched to accompaniment after the first half and vice versa. To avoid attracting visual attention, we prepared the experimental condition as follows: First, we put a blackout curtain behind the singers to conceal any other objects. Second, the singers were instructed to deliver a natural performance but maintain a stationary position to avoid any significant movements that might attract visual attention. Third, both singers wore white shirts and adopted modest hairstyles (i.e., they did not adopt the showy style one might see at a live concert). The stimuli were created by combining videotaped images (GZ-MG40, Victor) with sounds recorded through two microphones (SM57, Shure) with a multitrack recorder (SX-1, TEAC). One microphone was placed in front of each singer, and the recorded sounds were presented as stimuli (Fig. 1a). In presenting the stimuli to the participants, the voices were panned left and right: the right speaker presented the recording of the singer on the right, and the left speaker presented the singer on the left. The images were videotaped by a single camera placed in front of the singers. To examine whether a bias of visual attention exists toward one side of the singer, on another day, we conducted a supplemental experiment in which other participants rated the stimuli used in the main experiment in the visual-only mode (Appendix A). The result showed no differences between the two singers in terms of attracting visual attention from the participants in 9 of 12 stimuli presented in the visual-only mode. The singers’ gaze conditions were manipulated as follows: four gaze conditions (unreciprocated gaze from each singer, absence of eye contact, and existence of eye contact) × two musical conditions (melody in the first half of the performance or vice versa). The results for the eye-contact condition were eliminated since that condition was not established for the purposes of the present study, although they were included in the appreciation experiment. Under the unreciprocated gaze condition, at the moment the parts were switched, the singer looked once toward the co-performer for approximately 2 s. The reasons for using such a brief gaze were that (1) a long duration would have required the performer to sing with her head turned toward her co-performer, which discords with ecological validity, and (2) an observer’s attention shifts with shifts in the face of the person being looked at, even if the duration is less than 1 s (Langton, Watt, & Bruce, 2000). We regarded the singer’s face direction as similar to gaze direction because (1) in terms of ecological validity, singers are likely to directly face co-performers, (2) it could be difficult for participants to perceive gaze shifts through singers’ gazes alone since the video stimuli focused not on the singers’ faces but on most of their bodies (Fig. 1b), (3) head direction plays a crucial role in the directional attention of others (Langton et al., 2000), and (4) head or nose direction has also been observed to affect joint attention, which is the topic of the present study (Langton, Honeyman, & Tessler, 2004). Accordingly, we checked the videotape frame by frame and determined the duration of the gaze toward the co-performer from its onset until the performer returned to a front- facing orientation. To examine whether participants clearly discerned the melody and accompaniment parts, after the experiment, four audio-only stimuli were presented. Twenty- eight participants were asked how the melody and accompaniment parts switched between the right and left speakers. Twenty-six participants correctly responded in all four trials, recognizing the soprano part as the melody. One participant correctly responded to 3 of the 4 stimuli, and another responded that he did not understand the meaning of “melody part.” Two of the 30 participants did not participate in this listening test because the duration of the setting of the experiment became so long that it would exceed the duration of the experiment. Overall, the stimuli used in the experiment were perceived as sufficiently separated. Because participants recognized soprano as the melody part, hereinafter, the soprano part represents the melody part. Although Fig. 3 portrayed the appearance of the singers through line drawing, the participants saw recorded actual singing of a real female duo in front of a blackout curtain. Procedure The experiment was conducted in a quiet room. Each participant was rewarded with a book token for 500 Japanese yen. After providing informed consent, the experiments began. Participants sat approximately 65 cm away from a 15.6-in. monitor (Fig. 2). As a dummy task, after watching the stimuli, they rated their preference on an 11-point scale, from very undesirable (0) to very preferable (10). The stimuli were presented randomly to each participant using a display and two loudspeakers (SB-CH150, Panasonic), and each participant’s gaze points were recorded using a noninvasive eye tracker (GP3 Eye Tracker, Gazepoint). Each participant’s gaze-point coordinates on the monitor were analyzed. After two practice trials, we conducted the experiment. The singing was presented through two speakers on either side of the monitor. The right speaker presented sounds recorded by the microphone pointed at the right-side singer, while the left speaker presented those recorded by the microphone pointed at the left-side singer. Accordingly, participants predominantly heard the right-side singer from the right and the left-side singer from the left. After this experiment, another experiment was conducted. The whole experiment took about 40 min, including a short break. Data analysis We used the eye tracker to measure participants’ gaze points. As shown in Fig. 3, participants’ gazes concentrated on the faces, on the bodies, or around the bodies of the singers. When participants looked at 40% of either side of the display (i.e., left or right side of the display), we regarded them as watching near around the singer on that side. There were three reasons for this criterion. First, since the present study aimed to examine approximately which singer participants looked at, namely, the direction of gaze was important, a detailed analysis of precise gaze points on areas such as singers’ eyes, heads, or bodies in the stimuli was not required. Second, accuracy of gaze point was likely to slightly differed among participants. Third, given that the background was blackout curtain, there was nothing to attract visual attention. Taking account of these reasons, this measurement was adopted. We eliminated data for 20% of the gaze points around the center of the display because those points represented participants shifting between left and right. We also eliminated data for gazing beyond the monitor. The sampling rate for the gaze data was 60 Hz.","Fig. 4 shows the results for the directions of participants’ gazes when watching stimuli in which both singers sang facing forward throughout the piece. When the right-side singer sang the melody in the first half of the piece, two-way ANOVA on the duration of gaze (switch of musical part × first or second half) revealed interaction, a trend that approached significance (F(1, 19) = 3.730, p = 0.068, ηp2 = 0.164). The main effect was not significant. When the left-side singer sang the melody in the first half, two-way ANOVA on the duration of gaze (switch of musical part × first or second half) did not show interaction. The main effect was significant (F(1, 19) = 6.026, p = 0.024, ηp2 = 0.241). Participants gazed significantly longer at the right-side singer than the left-side one, regardless of assignment of musical part (melody or accompaniment). Fig. 5 showed the directions of participants’ gazes at the point of the largest pitch interval and the highest pitch of the score. Fig. 5a and b show participants’ gaze durations at two points, during which the pitch interval between the melody and the accompaniment was the largest (hereinafter, PI1 and PI2). The t-test results showed that participants significantly looked at the left-side singer longer at PI1 when that singer sang the melody in the last half (t(19) = 2.346, p = 0.030, Cohen’s d = 0.525, n.s. at PI2), and participants significantly looked at the right-side singer at PI2 when that singer sang the melody in the latter part of the piece (t(19) = 4.117, p = 0.001, Cohen’s d = 0.921) and at PI1, a trend that approached significance (t(19) = 1.998, p = 0.060, Cohen’s d = 0.447). Fig. 5c and 4d shows participants’ average gaze duration when the highest tone (F5, i.e. 698 Hz) was sung (hereinafter, HP1 and HP2). The t-test results showed that gaze duration toward the left side at HP2 was longer than that of the right side when the left-side singer sang the melody (t(19) = 2.283, p = 0.034, Cohen’s d = 0.511; t(19) = 0.768, n.s. at HP1), and gaze duration toward the right-side singer at HP2 was longer than that of the left side when the right-side singer sang the melody (a trend that approached significance: t(19) = 1.825, p = 0.084, Cohen’s d = 0.408; t(19) = 1.704, n.s. at HP1). To examine the influence of singers’ gaze shifts on participants’ gazes, we conducted three-way ANOVA (switch of musical part × gaze direction × first or second half) by subtracting the duration participants gazed toward the left from the duration they gazed toward right. Fig. 6 shows the participants’ gazes under each stimulus condition. Three-way interaction was not significant. Simple interaction (switch of musical part × first or second half) was significant (F(1, 19) = 13.474, p = 0.002, ηp2 = 0.415). In the case where the left-side singer looked toward the right, two-way ANOVA (switch of musical part × first or second half) showed that interaction was significant (F(1, 19) = 9.907, p = 0.005, ηp2 = 0.343). The main effect of the first/s half was significant (F(1, 19) = 5.639, p = 0.028, ηp2 = 0.229). The main effect of the switch of musical part was marginally significant (F(1, 19) = 3.837, p = 0.065, ηp2 = 0.168). In the case where the right-side singer looked toward the left, two-way ANOVA (switch of musical part × first or second half) showed that interaction was significant (F(1, 19) = 9.525, p = 0.006, ηp2 = 0.334). The main effect of the switch of musical part was marginally significant (F(1, 19) = 3.325, p = 0.084, ηp2 = 0.149). Thus, in terms of the whole duration of gazes throughout the piece, participants’ gazes toward the singer assigned the melody part increased when the shift in the melody part occurred. Fig. 7 shows the average duration of participants’ gazes when each singer shifted her gaze toward her co-performer. For example, as the white circle in Fig. 7 represents, while watching the stimulus in which the left side singer looked at the right singer, and then the melody part switched from the left singer to the right singer, participants looked at the left singer. Subsequently, 3–5 s after this gaze shift, participants’ gaze did not concentrate on either side of the singer, and 6 s after this gaze shift, participants looked at the right singer for a longer period. To investigate participants’ visual attention while the singers looked at each respective co-performer, we conducted a two-way ANOVA on the duration of gaze (switch of musical part × gaze direction) by subtracting the duration participants gazed toward the left from the duration they gazed toward the right. The result showed that interaction was not significant, while the main effect of the gaze-shift factor was significant (F(1, 19) = 18.544, p < 0.01, ηp2 = 0.494). Thus, participants looked longer at the singer who shifted gaze direction toward her co-performer than the singer who did not, regardless of the assignment of the melody part. To analyze participants’ time-lapse gazing behavior, we conducted a three-way ANOVA (time (from “during gazing” to 2 s) × gaze direction × switch of musical part) by subtracting the duration participants gazed toward the left side from the duration they gazed toward the right. Three-way interaction was not significant. Simple interaction (time × gaze direction) was significant (F(2, 38) = 12.153, p < 0.001, ηp2 = 0.390). In the case where the right-side singer sang the melody in the first half, the two-way ANOVA (time × gaze direction) showed that the interaction effect was significant (F(2, 38) = 9.138, p = 0.001, ηp2 = 0.325). The main effect was not significant. In the case where the left-side singer sang the melody in the first half, two-way ANOVA (time × gaze direction) showed that the interaction effect was significant (F(2, 38) = 7.311, p = 0.002, ηp2 = 0.278). The main effect was also significant (F(1, 19) = 5.372, p = 0.032, ηp2 = 0.220). Thus, the singers’ gaze shifts influenced the time- series alternation of the participants’ gazes. The overall tendency in participant gaze shifts was as follows: First, participants looked at the singer who shifted gaze. Subsequently, participants shifted their gaze toward the singer who was looked at by her co-performer. Finally, participants looked at the singer assigned the melody part.","The present study investigated the gazing of audience members when watching an audiovisual presentation of a duo-singing performance by focusing on the dominance of the musical part in auditory attention and visual joint attention. The main results were as follows: (1) regardless of whether it was the left or right singer, the one assigned the melody part (soprano part) attracted more visual attention throughout the piece, with the exception of two stimuli, (2) joint attention emerged when each singer shifted her gaze toward her co- performer, suggesting that inter-performer gazing interaction, when playing a spotlight role, mediated performer-audience visual interaction, and (3) the musical part strongly influenced the total duration of participants’ gazes, while the spotlight effect of gaze was limited to just after the singers directed their gazes toward the co-performer. First, the present study obtained quantitative evidence that musical parts—namely, melody (soprano) or accompaniment (alto)—significantly affect where audiences direct their gaze during an ensemble performance. Overall, the singer assigned the melody part attracted longer durations of gazing. Even when the singer’s gaze shifted toward her co-performer, participants were still totally inclined to look at the singer assigned the melody. While prior studies have yielded findings regarding auditory factors in multipart music (Bigand et al., 2000; Ragert et al., 2014; Uhlig et al., 2013), the results of the present study provide a new perspective suggesting that the visual attention of audience members depends on musical parts. At the same time, the present result aligns with Gregory’s (1990) finding that melody parts attract attention from listeners. Why did audiences allocate more visual attention to the singer attracting auditory attention? As our hypothesis, the effect of high-pitched voices could account for this visual attention to melody (soprano) parts. When the singer assigned the melody sang the highest pitch in the piece, which was also the highest pitch in relation to the pitch of the accompaniment (alto) part, participants looked toward the performer singing the higher part. Similarly, at the points in which pitch interval between melody and accompaniment part was largest, the singer assigned the melody attracted the gaze of the audiences. These results correspond to other findings showing that music listeners pay more auditory attention to higher pitches (Fujioka et al., 2005; Gregory, 1990; Marie & Trainor, 2014; Trainor et al., 2014). If such high voice superiority was derived from characteristics of auditory system (Trainor et al., 2014), it can be inferred that the processing of auditory information influenced the visual attention toward the soprano part in the present results. Still, careful consideration is necessary to argue the dominance of the melody part in visual attention, because the melody part that is generally assigned as a higher pitch part is not always higher than the accompaniment part in all ensemble performances. Owing to the fact that participants in the present study recognized the soprano part as the melody part through the listening experiment, the dominance of the melody part in visual attention in the present results was proved. However, given that the aim of the present study was not designed to examine whether the effect of high-pitched voices is coincides with the dominance of melody part, it remains unclear whether an accompaniment part involving higher pitch would also attract visual attention. If an implication that the melody part in western music dominates the accompanying harmony in auditory aspects (Uhlig et al., 2013) can be applicable to visual aspects, a melody part incorporating a lower pitch could attract more visual attention than the accompaniment part did. Thus, further research would be useful in examining whether the melody part always attracts visual attention from audiences, especially in cases where the pitch is lower than that of the accompaniment. Second, joint attention emerged when the singers shifted gaze directions. Thus, gaze shift between the singers played a role as a temporary spotlight. Participants first looked at the singer who looked toward her co-performer and then (approximately 2 s after the gaze shift) looked at the singer who was being looked at by her co-performer. After that, (approximately 6–7 s after the gaze shift), participants looked at the singer assigned the melody part. This phenomenon can be explained as follows: Participants first gazed at the singer who looked at her co-performer because the singer, who had been facing frontward, suddenly took a new action. The participants’ second gaze shift toward the singer being looked at by her co-performer can be explained in terms of joint attention, in which gazing attracts the visual attention of observers (Baron-Cohen, 1995; Frischen et al., 2007; Langton et al., 2000). The present results thus provided evidence of joint attention during musical performance, while previous findings focused on daily communication (e.g. Carrasco, 2011), child development (e.g. Tomasello, 1995), or human-robot communication (e.g. Staudte & Crocker, 2011). The present result also confirms Kawase’s (2009a) implication that a performer’s gaze attracts visual attention from audience members, like a spotlight. Over all, the present results provided an implication for the cognitive function of paying attention to performers during an actual musical performance. For both performers and audiences, gaze can be a useful interaction cue. Performers can employ gaze as an effective spotlight that leads an audience to direct attention toward certain performer(s) or objects. Audiences also can use the spotlight of a performers’ gaze as a cue for what they should direct attention to or for notification of performance process, e.g. the boundary of solo part, especially when upcoming events are unpredictable in performances such as improvisation. Yet, it is necessary to consider the situational differences between our results and ensemble concerts in the real world. In the present study, the duration of joint attention was short (approximately 2 s), which could have been caused by the singer’s immediate gaze shift from her co-performer back to the front. In Kawase’s (2009a) field study, the singer directed her gaze toward the guitarist playing a solo for a longer duration (approximately 20 s), suggesting that a longer gaze duration between performers affects audience fixation on the performer being gazed at by co- performers. Thus, further research is needed to explore the relationships between the duration of inter-performer gazing and the duration of audiences’ joint attention. Third, in terms of audiovisual interaction, the present results point to the strong influence of auditory information on the total duration of participant gazing, regardless of shifts in the singer’s gaze direction. The present result is supported by the supplementary experiment in Appendix A, which found no differences between the two singers in terms of attracting visual attention from the participants in 9 of 12 stimuli presented in the visual-only mode. In summary, the process of audio-visual interactions that influenced an audience’s gaze in the present study is as follows: Musical parts (melody or accompaniment) predominantly affected the gaze of audiences throughout the performances. The singer assigned the melody part attracted longer durations of gazing. Meanwhile, gaze shift between performers induced the joint attention of audiences. However, this spotlight effect of gaze was limited to just after the singers directed their gaze toward the co- performer, subsequently, the effect of musical parts returned. A combination of two different pitch parts, the existence of two singers, and gazing between the singers in the present study provided additional aspect of crossmodal attention, because we may hardly ever see these factors simultaneously occur in daily communication, e.g. conversation. The present results that auditory attention toward the dominant musical part corresponded to visual attention toward the singer who was assigned as that part support previous findings that auditory attention affect visual attention (Driver & Spence, 1998), or that circuit that is related to auditory attention is linked with the circuit related to vision in the brain (e.g., Winkowski & Knudsen, 2006). Given that music appreciation often occurs in daily life, the present results provided evidence of more natural attentional behavior in crossmodal situations, in accordance with ecological validity. In terms of audience attention during audiovisually presented music, the present results yielded another perspective by focusing on an ensemble music performance that potentially causes attention allocation in both auditory and visual aspects, while previous studies have mostly focused on performer-audience visual interactions in solo performances. During a singing duo performance containing not only two musical parts and two performers, but also inter- performer interactions through multiple cues (e.g. Keller, 2014), for example gaze, (Davidson, 2005; Kawase, 2009a; Kawase, 2014a; Kawase 2014b; Moran, 2010) which often occur, it makes audiences’ attention more complicated. The present result revealed aspects of this complicated attention of audiences and confirmed an assumable relationship between inter-performer interactions and performer-audience interactions (Kawase et al., 2007). Meanwhile, the present result for joint visual attention supports other findings regarding the powerful impact of visual information on audience members who judge audiovisual music performances (Platz & Kopiez, 2012). However, since the visual cues were controlled to focus on gaze effect, the influence of visual information on participants’ visual attention was limited. The singers in the present study were instructed to restrict salient visual expressions such as body movement (Broughton & Stevens, 2009), facial expressions (Thompson et al., 2010), clothing (Griffiths, 2008), and overall physical appearance (Wapnick et al., 1998). In addition, they were not instructed to manipulate performance aspects such as manner of expression (Davidson, 1993), emotion (Dahl & Friberg, 2007), tone duration (Schutz & Lipscomb, 2007), and performance proficiency (Tsay, 2013). Such control of visual information helps to highlight the influence of auditory information on the audience’s gaze, while audiences at concerts actively try to receive visual information (Kawase, 2013). Conversely, the present result accords, to some degree, with other studies showing the strong influence of auditory information upon audience members when judging music performance elements in audiovisual presentation (Kawase, 2009b; Petrini et al., 2010; Silveira & Diaz, 2014; Vines et al., 2006). Further study would be useful to explore whether, and if so, how, visual information, apart from gaze such as body movement, facial expressions, or physical appearance, that was employed in previous studies also attracts audiences’ gaze. Future studies should investigate whether these results can be applied to larger ensembles, such as orchestras, that might involve more complicated interactions between performers. The absence of inter-performer gazing due to instrument manipulation or performance etiquette should be examined as well. In such cases, visual attention of audience members could differ from the present results. A field study would also be fruitful for exploring audience gazing in real concerts since performer-audience interactions in ensemble concerts involve multifaceted nonverbal cues other than gazing (Kawase et al., 2007; Kurosawa & Davidson, 2005) that might attract visual attention. Such an attempt could contribute to the elucidation of a holistic perspective on musical communication.","We would like to thank Dr. Hiroshi Kinoshita, Dr. Kenji Katahira, Kei Eguchi and all the performers. Furthermore, we are grateful for the valuable comments from the anonymous reviewers. This research was funded by JSPS KAKENHI Grant Number 26870735 and Yamaha Music Foundation."],["One theory of visual awareness proposes that electrophysiological activity related to awareness occurs in primary visual areas approximately 200 ms after stimulus onset (visual awareness negativity: VAN) and in fronto-parietal areas about 300 ms after stimulus onset (late positivity: LP). Although similar processes might be involved in auditory awareness, only sparse evidence exists for this idea. In the present study, we recorded electrophysiological activity while subjects listened to tones that were presented at their own awareness threshold. The difference in electrophysiological activity elicited by tones that subjects reported being aware of versus unaware of showed an early negativity about 200 ms and a late positivity about 300 ms after stimulus onset. These results closely match those found in vision and provide convincing evidence for an early negativity (auditory awareness negativity: AAN), as well as an LP. These findings suggest that theories of visual awareness are also applicable to auditory awareness. --------------------------------------------------------------------------------","How does the brain enable us to experience seeing a picture or hearing a tone? This question has been studied extensively in vision. A common strategy has been to use threshold tasks: If a single, weakly visible image is shown repeatedly at the awareness threshold, subjects typically report that they are aware of the image on half of the trials, even though the image remains constant. The differences in associated neural activity between images that subjects report that they are aware of versus images that they are unaware of represent the neural correlates of visual awareness (Aru, Bachmann, Singer, & Melloni, 2012; Crick & Koch, 1998). Electrophysiological correlates of visual awareness ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To understand visual awareness, the location and timing of its neural mechanisms are presumed to be critical. Some theories suggest that awareness is mediated by early recurrent processing in the visual cortex (Lamme, 2010), whereas other theories suggest that it is mediated by late activation in the fronto-parietal network (Dehaene & Changeux, 2011; Lau & Rosenthal, 2011). Because of their excellent temporal resolution, electrophysiological measures derived from electroencephalography (EEG) have been immensely useful in addressing this issue (Eklund & Wiens, 2018; Koivisto & Grassini, 2016; Koivisto, Salminen-Vaparanta, Grassini, & Revonsuo, 2016). In several electrophysiological studies, a threshold task was used, and awareness was measured with subjective ratings. Critically, these studies did not use objective measures of performance (such as those used in a detection task) because these may not capture awareness equally well, as subjects may detect a stimulus without being aware of it. For example, subjects with blindsight can perform visual detection tasks well while claiming not to have any visual awareness (Stoerig, 2006; Weiskrantz, 1996). In these electrophysiological studies, awareness ratings were used to separate images into those that subjects reported that they were aware of and those that subjects reported that they were unaware of. An event-related potential (ERP; the average EEG response across images) was computed for each awareness rating, and a difference ERP was computed between them. Results of these studies suggest two ERP correlates of visual awareness: visual awareness negativity (VAN) and late positivity (LP; Koivisto & Revonsuo, 2010). VAN is a negative wave at occipital electrodes that occurs about 200 ms after stimulus onset. LP is a positive wave at parietal electrodes that occurs at least 300 ms after stimulus onset (Koivisto & Revonsuo, 2010). VAN has been suggested as the earliest electrophysiological correlate of visual awareness (Koivisto & Grassini, 2016), and this claim has been corroborated with magnetoencephalography (MEG; Andersen, Pedersen, Sandberg, & Overgaard, 2016). However, it is unclear whether LP is correlated with awareness (Dehaene & Changeux, 2011; Salti, Bar-Haim, & Lamy, 2012), post-perceptual processes (Andersen et al., 2016; Koivisto et al., 2016), or both. In terms of mechanisms, the current understanding is that during visual processing, activity from lower visual areas is fed forward to higher areas and then fed back to form recurrent loops. The early recurrent loops occur in lower areas (Lamme & Roelfsema, 2000) and are captured by VAN (Koivisto & Revonsuo, 2010). As later recurrent loops involve higher areas, including the fronto-parietal network, global recurrent processing ensues. This process may be captured by LP (Koivisto & Grassini, 2016), and it enables subjects to report their awareness (Lamme, 2006). Neural correlates of auditory detection ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ According to recurrent processing theory (Lamme, 2010), the processes that result in auditory awareness might be similar to those that result in visual awareness. Previous electrophysiological studies with threshold sounds have mainly used detection tasks with and without confidence ratings (Hillyard, Squires, Bauer, & Lindsay, 1971; Parasuraman & Beatty, 1980; Squires, Hillyard, & Lindsay, 1973). These studies provide some evidence for neural correlates of auditory detection in an early interval (i.e., N1) and a late interval (i.e., P3). In the first study, subjects (N = 3) were asked to detect threshold tones in white noise (Hillyard et al., 1971). Correctly detected tones (hits) were associated with an N1 (an early negativity at vertex, at about 100 ms) for one subject and a P3 (a late positivity at vertex, at about 300 ms) for all three subjects. In contrast, undetected tones (misses) were not associated with any N1 or P3. In a follow-up study, the task was similar, but subjects also rated their confidence regarding each reported detection (Squires et al., 1973). When subjects correctly detected tones and rated their confidence in the detection as high, these tones were associated with an N1 and a P3. Undetected tones were not associated with an N1 or P3. Similar effects were observed when subjects were asked to detect and identify tones with different frequencies and also provide a confidence rating (Parasuraman & Beatty, 1980). More recently, the neural correlates of informational masking were studied with MEG (Gutschalk, Micheyl, & Oxenham, 2008). The task was to detect target tones within series of tones that were masked by background tones (multitone masking). Compared with undetected tones, detected tones resulted in a negativity in the MEG between 50 and 250 ms after tone onset. This negativity was localized in the same area as the N1. These findings have been corroborated in subsequent studies with MEG (Dykstra & Gutschalk, 2015; Giani, Belardinelli, Ortiz, Kleiner, & Noppeney, 2015) and electrocorticography (Dykstra, Halgren, Gutschalk, Eskandar, & Cash, 2016). Similarly, studies that recorded EEG in a change deafness task found an enhanced N1 and P3 to detected changes versus undetected changes (Gregg & Snyder, 2012; Puschmann et al., 2013). Taken together, previous studies support the idea that detection is correlated with both early and late activity. However, these studies used methods that are not optimal for measuring awareness. Although awareness is largely captured in a detection task, awareness may be misclassified if subjects have only two ratings to choose from (Hillyard et al., 1971). For example, subjects may categorize tones as not heard even though they heard the tone faintly. Although tasks with confidence ratings allow for graduated ratings (Parasuraman & Beatty, 1980; Squires et al., 1973), they remain an indirect measure of awareness. Thus, they do not reflect subjects’ experiences as well as asking subjects directly about their experience does. Therefore, awareness should be measured directly with a rating scale that has several alternatives to allow subjects to rate their level of awareness (Sandberg, Timmermans, Overgaard, & Cleeremans, 2010). Neural correlates of auditory awareness ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The goal of the present study was to establish ERP correlates of auditory awareness. The study was modelled after previous studies (Hillyard et al., 1971; Parasuraman & Beatty, 1980; Squires et al., 1973) with the addition of using explicit awareness ratings, similar to previous studies in vision (Eklund & Wiens, 2018; Koivisto & Grassini, 2016; Koivisto & Revonsuo, 2010). Tones were presented at each subject’s awareness threshold, and subjects reported their auditory awareness of the tones by using an awareness rating scale with three levels while EEG was recorded. In vision, this approach has found VAN. We expected a similar correlate in hearing, the auditory awareness negativity (AAN).","The method and analyses were preregistered in detail before any data were collected. Deviations from the preregistration are noted below. All data and scripts are available elsewhere (Wiens & Eklund, 2019). We collected data from two groups of subjects (group A, doi: 10.17605/OSF.IO/HN8KJ; group B, doi: 10.17605/OSF.IO/SHVZD). For group A, we preregistered an interval for the AAN on the basis of the responses to louder tones (control trials). However, as explained in Section 3.2.1, the data suggested that this preregistered interval preceded the AAN. Therefore, we preregistered a later interval and used a new sample (group B) to test for AAN in this interval.","We preregistered to recruit at least 20 subjects in each group. If the Bayes Factor (BF) exceeded 3 or was below ⅓ for our hypotheses, recruitment would end. Otherwise, subject recruitment would continue for another week. This process was repeated until the BF reached the criterion, a maximum of 60 subjects were tested, or a predetermined end date was reached. Group A consisted of 24 healthy subjects (age: M = 25.63 years, SD = 4.43), of whom 11 were male and 22 were right-handed. Group B consisted of 25 healthy subjects (age: M = 26.52 years, SD = 6.18), of whom 10 were male and 24 were right-handed. Subjects had normal or corrected-to-normal vision and were recruited from local universities and through online billboards. They were compensated with either movie vouchers or course credits. Before starting the experiment, subjects provided written consent in accordance with the Declaration of Helsinki. The research was conducted in accordance with the principles of the regional ethics board. Although we preregistered several exclusion criteria (e.g., noisy EEG data), only three subjects from group B were excluded because they did not show a negativity in the N1 interval in the ERP extracted from aware control trials, as defined below. Critically, the decision to exclude these subjects was made before the ERPs from the critical trials were analyzed. The final samples consisted of 24 subjects in group A (11 male; age: M = 25.63, SD = 4.43) and 22 subjects in group B (8 male; age: M = 26.32, SD = 6.38).","Stimuli were 100-ms tones (f = 1000 Hz, 5 ms fade-in and fade-out). Tones were presented binaurally through in-ear tubephones (ER2; Etymotic Research Inc., IL; www.etymotic.com). Instructions were displayed on a BenQ XL2430T, 24-inch gaming monitor (at 144 Hz, 1920 × 1080 resolution). PsychoPy v 1.85.3 (Peirce, 2007) was used to generate tones and to collect behavioral data. A Cedrus StimTracker (Cedrus Corporation, San Pedro, CA) was used to generate triggers to the actual tones. These triggers were used to define tone onsets and served to compensate for any timing errors between the actual presentation of the tones and their event markers from the presentation computer. Procedure Subjects performed a tone-detection task while seated in front of a computer screen with their chin in a chinrest. Fig. 1 shows the time course of a trial. On each trial, a black fixation cross (0.5 visual degrees) was shown for 1000 ms. Critical trials contained a tone at the individual subject’s auditory awareness threshold (see below), and control trials contained a tone at 10 dB above the calibrated threshold level. On critical and control trials, a tone was played binaurally 500 ms after trial onset. On catch trials, no tone was played. The fixation-cross remained visible for the duration of the trial. Afterwards, subjects rated their subjective awareness of the tone by using one of three buttons (“1,” “2,” and “3,” corresponding to “I did not hear any stimulus,” “I heard the stimulus weakly,” and “I heard the stimulus clearly,” respectively). Subjects were instructed to focus on rating their awareness accurately rather than responding quickly. The task comprised 800 trials (640 critical, 80 control, and 80 catch). The trials were divided into eight blocks of 100 trials each (80 critical, 10 control, and 10 catch), with a short break between each block. For each subject in group A, the order of critical, control, and catch trials were randomized within each block. For each subject in group B, the order of trials was randomized within each set of 10 trials (with 8 critical, 1 control, and 1 catch trial per set). Before the experiment began, subjects performed a short practice task. This task was identical to the main task but with clearly audible stimuli. After the practice task was completed, an interleaved staircase was used to calibrate the tone to a level at which the subject reported being aware (weakly or clearly) on approximately 50% of the trials (i.e., individual auditory awareness threshold). The staircase procedure consisted of three interleaved staircases with 40 trials each (36 tone present trials and 4 catch trials). One staircase started at the threshold estimate obtained from pilot subjects (4 dB), another started 20 dB above this estimate, and another started 20 dB below. The staircase procedure was as follows: If the subject reported awareness when a tone was presented, the level decreased. If the subject reported no awareness when a tone was presented, the level increased. For each staircase, reversal steps were 8, 8, 4, 4, 2, and 2 for the first six reversals, and 1 dB for subsequent reversals. After the calibration, a validation block was run with 100 trials (80 critical, 10 control, and 10 catch trials). The level of the critical tone in the validation block was determined from both the convergence of the three staircases (from visual inspection) and the mean of the final six reversals for each staircase. If 45–55% of the critical trials in the validation block were rated as aware, the experiment began. If a subject reported awareness on less than 45% of the critical trials, another validation block was run at a higher sound level. Similarly, if a subject reported awareness on >55% of the critical trials, another validation block was run at a lower sound level. This validation was repeated until the 45–55% aware criterion was met. If this criterion was not met within five validation blocks, subjects were tested at the level that was closest to their awareness threshold. However, all subjects met the criterion within five validation blocks. EEG recording ~~~~~~~~~~~~~ EEG data were recorded from six electrodes at standard 10/20 positions (Fpz, Fz, Cz, Pz, P9, and P10) and two additional electrodes (one on the tip of the nose, and one on the cheek) with an Active Two BioSemi system (BioSemi, Amsterdam, Netherlands). Fpz, Fz, Cz, Pz, P9, and P10 were recorded with pin electrodes in a 64-electrode EEG cap; the tip of the nose and the cheek were recorded with flat electrodes attached with adhesive disks. Because the left and right mastoids (M1 and M2) were not available in the EEG cap, we used the nearby positions P9 and P10 for convenience. Two additional, system-specific channels were recorded with pin electrodes in the EEG cap: The common mode sense (CMS; between PO3 and POz) served as the internal reference electrode, and the driven right leg (DRL; between POz and PO4) was used as the ground electrode. Data were sampled at 1024 Hz and were filtered with a hardware low-pass filter at 104 Hz. Data analysis The data were processed and analyzed with Matlab (The MathWorks, Inc.) and R (R Core Team, 2016). Physiological data were processed offline with the toolbox FieldTrip (version 20181003) in Matlab (Oostenveld, Fries, Maris, & Schoffelen, 2011). The behavioral analyses included all trials, whereas in the EEG data analyses, some trials were excluded (see below). In the EEG data analyses, tone onset was indexed by the Cedrus StimTracker, which eliminated any timing errors in tone onset. Offline, continuous EEG data were high- pass filtered with a 0.1 Hz Butterworth fourth degree two-pass filter. Fpz, Fz, Cz, Pz, and the mastoids were re-referenced to the tip of the nose, and Fpz was re-referenced to the cheek electrode (for a combined measure of vertical and horizontal electrooculography). Epochs were extracted from 100 ms before tone onset to 600 ms after tone onset. Each epoch was baseline corrected to the mean of the 100-ms interval before tone onset (−100 to 0 ms). For each subject, maximum amplitude ranges were extracted for individual epochs, and the distribution of these amplitude ranges was inspected. Individual trials that were apparent outliers (such as eyeblinks) were excluded. The exclusion thresholds were set for each individual because subjects showed substantial variability in these amplitude ranges. Critically, inspection of trials was blinded to trial type (critical, catch, and control) and awareness ratings to avoid bias (Keil et al., 2014). ERP analysis ~~~~~~~~~~~~ Two event-related potentials (ERPs) were derived from critical trials: Aware trials were tones rated as “I heard the stimulus clearly” or “I heard the stimulus weakly,” and unaware trials were tones rated as “I did not hear any stimulus.” Note that for aware trials, clearly and weakly heard tones were combined because subjects rarely rated their awareness as clear (<1%, see Section 3.1), similar to studies in vision (Eklund & Wiens, 2018; Koivisto & Grassini, 2016). Difference waves were calculated by subtracting the ERP to unaware trials from the ERP to aware trials. We predicted that this difference wave would be negative in the N1 interval (AAN) and positive in the P3 interval (LP). Because in vision VAN overlaps in latency with the visual N1 interval (Eklund & Wiens, 2018; Koivisto & Revonsuo, 2010), we initially expected to observe AAN in the auditory N1 interval. Because we expected N1 at centrally located electrodes (Parasuraman & Beatty, 1980), we combined Fz and Cz electrodes (after 30-Hz low-pass filtering). For group A, the relevant interval for AAN was preregistered as the peak (±50 ms) of the N1 in the grand mean ERP of the control trials. This definition of the N1 interval was independent from the critical trials and thus did not bias the main results. However, as discussed in Section 3.2.1, this interval preceded an apparent negative difference wave between aware and unaware trials. Therefore, we preregistered this later interval for group B. For both groups, mean AAN amplitudes were computed for the N1 interval across Fz and Cz electrodes, and mean LP amplitudes were computed for the P3 interval between 350 and 550 ms for the Pz electrode. We selected the Pz electrode based on previous findings in vision (Eklund & Wiens, 2018). We conducted Bayesian hypothesis testing to determine the degree of evidence for or against the alternative hypothesis (Dienes, 2008). The Bayes Factor (BF10) expresses the likelihood of the data given the alternative hypothesis relative to the likelihood of the data given the null hypothesis, whereas the BF01 shows the reverse (Dienes, 2008, 2016; Wagenmakers, Marsman et al., 2017; Wiens & Nilsson, 2017). Although the BF is a continuous measure of evidence, we adopted a common interpretation scheme (Wagenmakers, Love et al., 2017). The BF was calculated with Aladins Bayes Factor in R (Wiens, 2017). These scripts compute and plot the BF for mean differences in raw units if the alternative hypothesis is modelled as a normal, t, or uniform distribution, and the likelihood is modelled as a normal or t distribution (Dienes & McLatchie, 2018). For group A, the alternative hypotheses for AAN and LP were modelled as uniform distributions with the limits of −2 µV to +2 µV. For group B, the alternative hypothesis for AAN was modelled as a t distribution derived from group A. Specifically, we preregistered the following t distribution: M = −0.58 (SEM = 0.22, df = 22, 2-tailed). We note that the preregistration states a positive mean value (i.e., 0.58), but the reason is that the R scripts (Wiens, 2017) require the theoretical effect to be positive. Further, we note that when we checked our analyses after the preregistration, we realized that one subject was missing from this previous analysis. With this subject included, the correct t distribution should have been as follows: M = −0.67 (SEM = 0.25, df = 23, 2-tailed). Although we report the results for the preregistered analysis, results were similar with this corrected alternative hypothesis. These and other additional analyses are available elsewhere (Wiens & Eklund, 2019). Although we meant to use the specific t distribution (i.e., M = −0.58) only for the alternative hypothesis with regard to AAN, the wording of the preregistration suggests that we intended to use it also for the LP. However, this would be incorrect, because LP amplitudes are positive and much larger than AAN amplitudes. For simplicity, the Bayesian analyses reported below use the same alternative hypothesis for LP as in group A (i.e., −2 µV to +2 µV). Note that results are comparable if the alternative hypothesis for LP is modelled as a t distribution derived from the LP results from group A (Wiens & Eklund, 2019). Behavior ~~~~~~~~ Table 1 shows the descriptive statistics for the behavioral data. Subjects performed the task as intended: At the individual awareness threshold, close to fifty percent of the critical tones were rated as aware (of these, 99.7% were rated as weakly heard in both groups). Most control tones were rated as aware, and most catch trials were rated as unaware. The average level of the critical tone was 4.7 dB (SD = 4.1) for group A and 6.7 dB (SD = 7.9) for group B. Auditory awareness negativity Table 2 shows the descriptive and inferential statistics for the mean amplitudes. Fig. 2 shows the mean ERPs across subjects in group A. In the left panel, a clear negativity (N1) to control trials (black line) can be seen with its peak at 150 ms after stimulus onset. Our initial, preregistered hypothesis for AAN was that a negative difference wave of aware minus unaware critical trials would occur in the same interval as the N1 to control trials, as in vision (Eklund & Wiens, 2018; Koivisto & Revonsuo, 2010). However, for this interval (94–194 ms), the Bayesian one-sample t test provided anecdotal evidence against a negativity (BF10 = 0.37, which corresponds to BF01 = 2.70). Nonetheless, the left panel in Fig. 2 suggests that there is an apparent negative difference wave (green line), but its peak occurred later, about 190 ms after stimulus onset. On the basis of these findings, we preregistered this interval (140–240 ms) and tested it on a new sample. Fig. 3 shows the mean ERPs across subjects in group B. In the left panel, a clear negativity for the difference wave (green line) is visible with its peak around 200 ms after stimulus onset. This peak closely matched that of group A, as confirmed by the Bayesian one-sample t test that supported a negativity in the revised interval (BF10 = 5.90). These findings provide moderate evidence for AAN. Late positivity Table 2 shows the descriptive and inferential statistics for the mean amplitudes. The right panels in Figs. 2 and 3 show the mean ERPs for the two groups. A positive difference wave of aware minus unaware critical trials (green line) was apparent for both groups after 300 ms. For the preregistered interval (from 350 to 550 ms), evidence for LP was very strong to extreme (group A: BF10 > 45,000, and group B: BF10 = 81.10).","When tones were presented at the individual awareness threshold, subjects reported being aware of about 50% of these tones. From the electrophysiological recordings, a difference wave was computed for these aware minus unaware tones. This wave showed a central negativity in the early interval (140–240 ms after tone onset) and a central positivity in the late interval (350–550 ms after tone onset). Bayesian hypothesis testing provided moderate evidence for the early negativity (AAN) and very strong evidence for the late positivity (LP). Auditory awareness negativity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our initial, preregistered hypothesis for AAN was that a negative difference wave of aware minus unaware tones would occur in the same interval as the N1 to control tones, as in vision (Eklund & Wiens, 2018; Koivisto & Revonsuo, 2010). For this interval, the Bayesian analysis provided anecdotal evidence against a negativity, opposite to our predictions. However, visual inspection of the data suggested an apparent negative difference wave at a later latency (see green line in Fig. 2A). In hindsight, this finding of a delayed negativity may not be surprising: Because the latency of the N1 peak is delayed for quieter tones (Picton, Woods, Braun, & Healey, 1977) and the tones at the awareness threshold were quieter than the control tones, it is reasonable that the N1 was delayed to tones at the awareness threshold. Because this N1 delay made theoretical sense, we used the obtained interval for the N1 to preregister a new interval (140–240 ms). In the new sample, the Bayesian analysis provided moderate evidence for a negativity and thus AAN. Specifically, the BF = 5.90 implies that the presence of AAN is almost six times more likely than the absence of AAN. Because few electrodes were used in the present study, no source localization is possible with the present data. However, previous studies that recorded MEG during a multitone-masking task suggest that the main sources may be in auditory cortex (Gutschalk et al., 2008). Late positivity ~~~~~~~~~~~~~~~ For both data collections, our preregistered hypothesis for LP was that a positive difference wave for aware minus unaware trials would occur between 350 and 550 ms after stimulus onset. Indeed, Bayesian analyses provided very strong to extreme evidence for LP. In vision research, the P3 was larger to aware than to unaware trials; thus, there have been consistent reports of an LP as a positive difference wave (Dehaene & Changeux, 2011; Koivisto & Revonsuo, 2010). In hearing research, similar results were obtained when subjects provided binary detection responses or rated their confidence (Dykstra, Cariani, & Gutschalk, 2017; Hillyard et al., 1971; Parasuraman & Beatty, 1980; Squires et al., 1973). In the present study, subjects were required to rate their awareness explicitly. Therefore, our findings of an LP extend previous findings and demonstrate that the late neural correlates of hearing are similar to those in vision. Implications for theories of awareness ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the philosophical debate about what awareness or consciousness is, two concepts are central: phenomenal consciousness, which refers to “what it is like” to have an experience, and access consciousness, which refers to the state of introspection or reporting about an experience (Block, 1995). It remains unclear if and how the early process (VAN in vision, AAN in hearing) and the later process (LP in both modalities) map onto these concepts. According to recurrent processing theory (Lamme, 2006), local recurrent processing is necessary and sufficient for phenomenal consciousness. Therefore, if VAN and AAN are indirect measures of local recurrent processing, they index phenomenal consciousness (Lamme, 2018). Furthermore, global recurrent processing enables reporting and introspecting on an experience (Koch, Tononi, Massimini, & Boly, 2016; Lamme, 2006) and thus indexes access consciousness. This is a post-perceptual process according to recurrent processing theory (Koivisto & Grassini, 2016; Lamme, 2010). Therefore, if LP is an indirect measure of global recurrent processing, it indexes access consciousness (Koivisto & Grassini, 2016). According to global workspace theory, local recurrent processing is a preconscious process (Dehaene, Changeux, Naccache, Sackur, & Sergent, 2006; Naccache, 2018). Thus, VAN and AAN may at best index only preconscious processing. In contrast, if LP is an indirect measure of global recurrent processing, it indexes access consciousness, which subsumes phenomenal consciousness according to global workspace theory (Cohen & Dennett, 2011; Naccache, 2018). Although non-response tasks may be a promising approach to separating neural correlates of awareness from those of post- perceptual processes (Tsuchiya, Wilke, Frässle, & Lamme, 2015), the present findings do not resolve this discussion. Nonetheless, they provide convincing evidence that the ERP correlates of auditory awareness (AAN) are analogous to those in visual awareness (VAN).","In terms of neural correlates of auditory awareness, very little empirical research has been done previously (for reviews, see Dykstra et al., 2017; Snyder, Yerkes, & Pitts, 2015). The present results expand on previous studies that used a detection task (Hillyard et al., 1971) or detection and confidence ratings (Parasuraman & Beatty, 1980; Squires et al., 1973) to investigate the neural correlates of detection in hearing. To capture awareness in hearing, our study used awareness ratings, because they have been shown to capture subjective experiences accurately in vision (Sandberg et al., 2010). The present findings show that AAN is an early neural correlate of awareness in hearing, similar to VAN in vision."],["Higher levels of well-being are associated with longer life expectancies and better physical health. Previous studies suggest that processes involving the self and autobiographical memory are related to well-being, yet these relationships are poorly understood. The present study tested 32 older and 32 younger adults using scales measuring well-being and the affective valence of two types of autobiographical memory: episodic autobiographical memories and semantic self-images. Results showed that valence of semantic self-images, but not episodic autobiographical memories, was highly correlated with well-being, particularly in older adults. In contrast, well-being in older adults was unrelated to performance across a range of standardised memory tasks. These results highlight the role of semantic self-images in well-being, and have implications for the development of therapeutic interventions for well-being in aging. --------------------------------------------------------------------------------","The way we remember our past is thought to influence both our sense of self and our general well-being. However, the mechanisms underpinning the links between memory, self and well-being are little understood. This study aimed to investigate the relationships between well-being and two types of autobiographical memory: semantic self-images (e.g., autobiographical knowledge about the self, comprising traits, family roles and group- membership) and episodic autobiographical memories (detailed memories of specific, personally-experienced events). Furthermore, in light of established age-related changes in memory function (e.g., Levine, Svoboda, Hay, & Winocur, 2002; Piolino, Desgranges, Benali, & Eustache, 2002), we tested younger and older adult groups in order to examine whether these relationships differ between age groups. A better understanding of the role that different types of memory play in well-being has important implications. People with higher levels of well-being benefit from more than simply positive mood and good social relationships (Diener, 2013), they also have increased life expectancies and better physical health (Diener & Chan, 2011; Kok et al., 2013). Crucially, this study took the novel approach of comparing the roles of semantic and episodic autobiographical memory, addressing the recent call for more detailed investigation of the function of semantic components of autobiographical memory (e.g., Haslam, Jetten, Haslam, Pugliese, & Tonks, 2011; Prebble, Addis, & Tippett, 2013; Thomsen, 2009). Traditionally, the main focus of autobiographical memory research has been on episodic, rather than semantic, memories (see Thomsen, 2009). Episodic autobiographical memories relate to a specific moment in time and typically include sensory-perceptual details (e.g., visual imagery and sensory information) and are characterised by a sense of ‘mental time travel’ (Tulving, 1983, 2002). In contrast, semantic autobiographical memories involve simply knowing about an event or fact, without a sense of mentally travelling back and re-living a specific event (Tulving, 1983, 2002). These semantic autobiographical memories include knowing the locations of countries visited on holiday and names of family members. They also include the sets of traits, roles and beliefs that form semantic self-images, such as knowing one is a retired accountant and a mother of two children. Semantic self-images are thus a particularly self-relevant subsection of semantic autobiographical memory. Research suggests that as we get older, sensory-specific episodic memories become less accessible, whilst semantic autobiographical memories are retained (e.g., Levine et al., 2002; Piolino et al., 2002). Specific episodic autobiographical memories arguably play a range of important roles, including promoting social intimacy (Pillemer, 1998), acting as landmark events marking transitions in the life story (Shum, 1998; Thomsen, Pillemer, & Ivcevic, 2011), and engaging people with their long-term goals (Singer & Salovey, 1993). However, recent work (e.g., Haslam et al., 2011; Thomsen, 2009) has emphasised the important role of semantic autobiographical memory. Key to the present study, Piolino et al. (2006) proposed that the retention of semantic autobiographical memories may play a central role in enabling older adults to maintain a sense of diachronic unity. Haslam et al. (2011) suggest semantic self-knowledge is a bi-directional mediator between episodic autobiographical memory and identity. They proposed that episodic autobiographical memories provide the basis for semantic autobiographical memories, and it is these semantic “facts” that support the self. Prebble et al. (2013) developed this idea further, proposing that semantic autobiographical memories may help form a scaffold for both episodic recollection and imagination of future events (e.g., Irish, Addis, Hodges, & Piguet, 2012). Prebble et al. suggest that self-continuity in older adults is maintained by semantic autobiographical memory. Semantic autobiographical memories may provide a more efficient means of promoting narrative continuity, possibly by virtue of the fact that they can be used to synthesise and organise large amounts of information into a coherent life story (Prebble et al., 2013). In support of this organisational account, Thomsen (2009; Thomsen et al., 2011) proposed that people use semanticised life story chapters (e.g., time at university X) to organise autobiographical retrieval and shape a narrative life story. Semantic facts structure the way we remember the past and also the way children imagine the future (Bohn & Berntsen, 2011). In short, a growing number of researchers have emphasised the role that semantic autobiographical memory plays in organising memory and promoting a coherent sense of self. Semantic autobiographical memory is the most resilient form of autobiographical memory, preferentially preserved in healthy aging (Levine et al., 2002; Piolino et al., 2002) retrograde amnesia (Klein & Lax, 2010; Rathbone, Moulin, & Conway, 2009), depression (Dalgleish et al., 2007), autism (Crane & Goddard, 2008) and Alzhiemer’s disease (Martinelli, Anssens, Sperduti, & Piolino, 2013). In addition to highlighting the dissociation between episodic and semantic memory, these studies also raise the possibility that semantic self-images might be a useful starting point for rehabilitation in a range of clinical groups. The present study was particularly focused on the emotional valence of semantic self-images and episodic autobiographical memories. This is because there are established age-related changes in the emotional ratings of autobiographical memory (e.g. the positivity effect; Kennedy, Mather, & Carstensen, 2004), but little is known about how these might relate to well-being. Well- being can be conceptualised at both the eudaimonic and hedonic level. Eudaimonic well- being is associated with viewing one’s life with meaning, purpose and a sense of growth (Bauer, McAdams, & Pals, 2008) and using and developing the best aspects of oneself (Huta & Ryan, 2010). Eudaimonic well-being is often conceptualised as “psychological well-being” (e.g. Ryff & Keyes, 1995), although other conceptualisations exist (for a review, see Huta, 2013). Hedonic well-being focuses on both the experience of pleasure and a more cognitive evaluation of life satisfaction (Diener, 2013). Recent empirical work suggests that hedonic and eudaimonic well-being may overlap, with greatest overall well-being perhaps stemming from the pursuit of both hedonic and eudaimonic aims (Huta & Ryan, 2010). We suggest that semantic self-images play a crucial role in supporting the self. The present study aimed to explore the role of semantic self-images in well-being in aging, and centred on three main research questions, which will be explored in detail below: (1) How does the emotional valence of episodic autobiographical memories and semantic self- images differ between younger and older adults; (2) is well-being correlated more strongly with the valence of semantic self-images or episodic autobiographical memories and does this differ between younger and older adults; (3) in older adults, is the valence of semantic self-images more strongly correlated with well-being than objective measures of memory performance. The first research question focused on the phenomenological ratings of episodic autobiographical memories and semantic self-images in the two age groups. Older adults tend to rate their episodic autobiographical memories more positively than younger adults (the “positivity effect”; Kennedy et al., 2004; Schlagman, Schulz, & Kvavilashvili, 2006; Schryer, Ross, St-Jaques, Levine, & Fernandes, 2012). Compared to younger adults, older adults tend to reappraise negative memories in a more positive light (Comblain, D’Argembeau, & Van der Linden, 2005) and report a larger proportion of positive than negative autobiographical memories (e.g., Mather & Carstensen, 2005). The first prediction, in line with the positivity effect (e.g., Kennedy et al., 2004), was that older adults would rate both their semantic self-images and episodic autobiographical memories more positively than younger adults. The second research question centred on the relationships between well-being and the emotional valence of episodic autobiographical memories, and semantic self-images. In particular, it aimed to elucidate the role that semantic self-images play in well-being. Swann and Buhrmester (2012) suggest that stable self-representations are vital for social interactions, making goals, and enabling a sense of belonging. The present study sought to ascertain whether the valence of semantic self- images was more closely related to well-being than the valence of episodic autobiographical memories. Well-being may be associated with the way that autobiographical memories are framed (Philippe, Koestner, Beaulieu-Pelletier, Lecours, & Lekes, 2012). For example the presence of ‘growth’ memories (e.g., goal-related memories that lead to a richer understanding of oneself and one’s life) has been associated with increased ratings of well-being (Bauer et al., 2008; Bauer, McAdams, & Sakaeda, 2005; Philippe et al., 2012). Whilst there is a general consensus that judgements of mood and well-being are affected by the emotions associated with autobiographical memories (e.g., Hart, 2013), it is unclear which level of autobiographical memory (e.g., semantic or episodic) is involved. This question has particular relevance to aging, as research suggests that older adults may use semantic self-images to buffer against the negative effects of aging (Heidrich & Ryff, 1993). Older adults with negative self-perceptions of aging tend to show a decline in physical (Sargent-Cox, Anstey, & Luszcz, 2012) and psychological well-being (Mock & Eibach, 2011). Indeed, the self-system has been conceptualised as a “coping resource in the aging process” (Diehl, Hastings, & Stanton, 2001, p. 644). The present study thus sought to elucidate the relationship between the emotional valence of semantic self-images and well-being. Based on recent work suggesting that semantic autobiographical memories support identity and organise autobiographical retrieval (Prebble et al., 2013; Thomsen, 2009), the second prediction was that in both younger and older adults, the emotional valence of semantic self-images would correlate more closely with measures of well-being than the emotional valence of episodic autobiographical memories. The third research question focused only on the older adult group. We aimed to establish whether memory performance on a range of standardised cognitive tasks was related to well-being. Previous work has suggested that cognitive decline in aging is associated with decreased well-being (Llewellyn, Lang, Langa, & Huppert, 2008; Wilson et al., 2013). However, well- being in older adults tends to be rated relatively highly, in spite of declines in physical and cognitive health (Scheibe & Carstensen, 2010). Research testing older adults with dementia indicated that identity (measured as strength of personal identity) mediated the effects of memory loss on well-being, suggesting that it is not memory loss per se that leads to lower well-being (Jetten, Haslam, Pugliese, Tonks, & Haslam, 2010). The design of the present study enabled a clear comparison between the roles of both episodic autobiographical memories and semantic self-images, and the role of standardised memory performance, in well-being. This question has implications for our aging population. In essence, do the changes in memory associated with aging necessarily relate to how positive people feel about themselves and their lives? It was predicted that well-being would correlate more closely with how positively participants viewed themselves (e.g. emotional valence of semantic self-images) than how well they performed in standardised memory tasks.","Participants consisted of 32 younger adults (26 female; mean age = 20.25; SD = 2.99; range = 18–28) and 32 older adults (19 female; mean age = 70.22; SD = 3.02; range = 65–75,). The younger adults were psychology students at a UK university who received credits for participating. The older adults were recruited from a database of local adults interested in participating in research projects. All participants were community-dwelling and provided details of mental and physical health and any previous history of substance abuse. Five of the older and five younger adults reported a history of anxiety and/or depression. Cognitive function was measured using the Trail Making Tests A and B (completion time in seconds; Reitan, 1958), vocabulary sub-test of WAIS-R (Wechsler, 1981), and category (animal) and letter (F) fluency (Tombaugh, Kozak, & Rees, 1999). For details of participant characteristics see Table 1.1 In line with previous research, compared to younger adults, the older adults performed better on the vocabulary task (Verhaeghen, 2003) but worse on the trail making tasks (Kennedy, 1981) and category fluency task (Tomer & Levin, 1993).","All participants completed the IAM task (Rathbone, Moulin, & Conway, 2008), three measures of cognitive function (described above), and five measures of well-being in one lab-based session that lasted between 1.5 and 2 hours. The older adult group also attended a second 1-hour session in which they completed an additional set of memory and cognitive function tests. This second session took place within two weeks of session 1. IAM task Participants generated up to 10 ‘I am’ statements and then selected their two most important statements as cues for up to five episodic autobiographical memories per statement (generating up to 10 memories in total).2 All episodic autobiographical memories were dated for age at event, and rated on an 11-point scale for vividness (0 = not at all vivid, 10 = very vivid), emotional valence (−5 = very negative, +5 = very positive), personal significance (0 = not at all personally significant, 10 = very personally significant), rehearsal (0 = never think about it, 10 = think about it all the time) and on a dichotomous scale for imagery perspective (observer or field). Memories were rated in two ways for episodic specificity. First, participants were provided with standardised Remember/Know/Guess instructions (Gardiner, 1988) and asked to rate whether each memory was something they remembered (R), knew (K) or guessed (G). All responses rated as R were probed for episodic details, and were rated as Justified Remember (JR) if such details were provided (following Piolino et al., 2003). Memories were also rated for specificity by the researcher using a 0 to 4 scale (Baddeley & Wilson, 1986), in which 4 = a specific event with details and situated in time and space, 3 = a specific event without any detail but situated in time and space, 2 = a repeated or extended event situated in time and space, 1 = a repeated or extended event not situated in time and space, and 0 = no memory given/only general information about the topic. All ‘I am’ statements (e.g. semantic self-images) were rated by the participant for emotional valence (−5 = very negative, +5 = very positive), age of identity formation, and importance (0 = not at all important, 10 = very important). Thus, the emotional valence scores for both semantic self-images and episodic autobiographical memories were based on comparable participant-reported rating scales. Well-being scales A range of scales were employed to assess eudaimonic, hedonic, and clinically- relevant measures of well-being. Eudaimonic well-being was measured using the 42-item Psychological Well-being Scale (PWB; Ryff & Keyes, 1995; total scale score). Hedonic well-being was measured with scales that tapped into emotional experience (e.g., The Positive and Negative Affect Schedule, PANAS trait version [participants rated to what extent they “generally feel this way…on average”] Watson, Clark, & Tellegen, 1988), evaluation of how positive or negative one’s life has been (Satisfaction with Life Scale, SWLS; Diener, Emmons, Larsen, & Griffin, 1985), and optimism (Life Orientation Test-Revised Optimism Scale; LOT-R; Scheier, Carver, & Bridges, 1994). Finally, the Hospital Anxiety and Depression Scale (HADS; Zigmond & Snaith, 1983) provided a more clinically-relevant measure of well-being. Additional measures for older adult group The Autobiographical Memory Interview (AMI; Kopelman, Wilson, & Baddeley, 1989) assessed episodic and semantic autobiographical memory, the Rey-Osterrieth Complex Figure Test (Osterrieth, 1944; Rey, 1941) assessed visuospatial memory, the Logical Memory Test (LMT; subtest of Wechsler Memory Scale, Wechsler, 1997) assessed recall and recognition for narrative memory, the Graded Naming Test (GNT; McKenna & Warrington, 1983) measured semantic memory for object names, and the Everyday Memory Questionnaire (EMQ; Sunderland, Harris, & Gleave, 1984) was a self-report measure of participants’ day-to-day memory problems. General cognitive function was assessed using the Montreal Cognitive Assessment (MOCA; Nasreddine et al., 2005). Statistical procedures Comparisons of episodic autobiographical memory and semantic self-image ratings between younger and older adults were based on ANOVAs with ‘age group’ as a between subjects factor. In order to examine the relationships between well-being scales, and episodic autobiographical memory and semantic self-images ratings within each age group, Pearson product-moment correlation coefficients were calculated.","The first aim of this study was to examine whether the phenomenological ratings for semantic self-images and episodic autobiographical memories differed between younger and older adults. Mean ratings and ANOVA results for older and younger adults are shown in Table 2. For completeness, we also include comparison of well-being scores between groups. Results suggest that age has no effect on the emotional valence, importance, or frequency of semantic self-images generated. In contrast, older adults’ episodic autobiographical memories were significantly more positive, and rehearsed less frequently, than those of younger adults. As one would expect, older adults’ episodic memories were dated at a higher mean age at event, and were rated as containing less episodic detail than younger adults’ memories. There was also an age effect for proportion of memories viewed from a field perspective; older adults were more likely to remember events from a field perspective than younger adults, but both age groups used a field perspective for the vast majority of their memories. There were no significant differences in well-being measures between age groups. To examine whether gender had any effects on these data, it was included as covariate in the ANOVAs shown in Table 2. Gender had no significant effects on any of the key variables in this study (well-being scales and valence of episodic autobiographical memories and semantic self-images). The second aim was to explore the relationships between the emotional valence of semantic self-images and episodic autobiographical memories and participants’ well-being scores. Semantic self-image valence scores were calculated using participants’ mean emotional valence ratings for all self- images generated in the IAM Task. Episodic autobiographical memory valence scores were calculated using participants’ mean emotional valence ratings for all episodic memories generated in the IAM Task. To better understand how these relationships might differ between age groups, correlational analyses were run within groups for younger and older adults (see Table 3). As one would expect, both age groups showed multiple significant correlations between well-being measures, but of particular interest were the correlations between well-being and semantic self-image valence, as well as well-being and episodic autobiographical memory valence. All correlations were in the expected direction, with higher (i.e. more positive) emotional valence ratings for semantic self-images associated with increased well-being (e.g., higher scores on the PWB Scale, PANAS positive, SWLS, and LOT-R, and lower scores on the PANAS negative and HADS). There were no significant correlations between episodic autobiographical memory valence and well-being in the older adult group, whilst episodic autobiographical memory valence correlated with only the PANAS positive measure in the younger adult group. In contrast, the valence of younger adults’ semantic self-images correlated with three well-being measures and in the older adult group significant correlations were shown between semantic self-image valence and all measures of well-being. This critical finding suggests that, in both younger and older adults, semantic self-image valence is more closely correlated with well-being than episodic autobiographical memory valence. To examine this further, the correlation co- efficients in Table 3 were entered into Hotelling William method t tests. This tested whether there was a significant difference between the correlations for well-being and valence of episodic autobiographical memories compared to between well-being and valence of semantic self-images. Results (shown in summary in Table 4) indicated that, within the older adult group, five out of the six measures of well-being were more significantly correlated with valence of semantic self-images than valence of episodic autobiographical memories. In contrast, within the younger adult group there were no significant differences between correlations for well-being and episodic autobiographical memory valence compared to semantic self-image valence. The final aim was to examine whether increased well-being scores were associated with better cognitive performance in the older adult group. For this purpose, correlational analyses were run on scores obtained on a battery of cognitive tasks, including objective memory measures (e.g., AMI, LMT, GNT) and the subjective self-report EMQ, alongside well-being scores (see Table 5). The EMQ was significantly negatively correlated with the PWB Scale and PANAS positive, and significantly positively correlated with the PANAS negative and HADS. High scores on the EMQ indicate self-reported experience of a high number of everyday memory problems, thus these correlations are in the expected direction. None of the objective measures of memory were correlated with any of the measures of well-being, suggesting that (at least within this non-clinical group) there was no relationship between memory performance and well- being.","The findings of this study suggest that well-being is more closely linked to how positively people view themselves at a semantic level, rather than at an event-specific level – an effect that was particularly pronounced in older adults. The first aim of this study was to examine how the emotional valence of semantic self-images and associated episodic memories might differ between younger and older adults. There were few differences between age groups on the measures of episodic memory, and no differences between age groups in semantic self-image or well-being scores. As predicted, older adults rated their episodic autobiographical memories more positively than younger adults. In contrast, counter to predictions, semantic self-images were rated equally positively by adults in both age groups. This suggests that the positivity effect in ageing established for episodic autobiographical memories (e.g. Kennedy et al., 2004) may not extend to semantic self-images, replicating recent work by Chessell, Rathbone, Souchay, Charlesworth, and Moulin (2014). The second aim was to examine the correlations between well-being and the emotional valence of both semantic self-images and episodic autobiographical memories. As predicted, the emotional valence of semantic self-images was significantly related to participants’ well-being. This effect was shown in both younger and older adults, but was particularly pronounced in the older adult group, for whom semantic self-image valence correlated with all six measures of well-being. In contrast, the emotional valence ratings for episodic autobiographical memories were not correlated with any measures of well-being in older adults, and were only correlated with one measure of well-being in the younger adult group. When correlations between well-being and episodic autobiographical memory valence were compared to those between well-being and semantic self-image valence, we found – counter to our predictions – pronounced aging effects. In younger adults, there was no difference between the correlations for well- being and episodic autobiographical memory valence compared to well-being and semantic self-image valence. However, in older adults, well-being was significantly more related to semantic self-image valence than episodic autobiographical memory valence in five of the six measures of well-being. This may reflect enhanced emotional-regulation processes associated with aging (Carstensen, 2006) – a point we discuss later in this section. This key finding suggests that, in older adults, well-being is more closely linked to the emotional valence of semantic self-images, rather than episodic autobiographical memory. The final aim was to examine the relationship between well-being and memory performance in older adults. In line with predictions, objective memory performance across a range of standardised tasks bore little relation to measures of well-being. In contrast, self- reported memory problems showed multiple correlations with well-being. These results suggest that poor memory performance per se does not necessarily equate with diminished well-being. The current findings add to a growing body of research highlighting the important role of semantic autobiographical memory (e.g., Bohn & Berntsen, 2011; Klein & Lax, 2010; Prebble et al., 2013; Thomsen, 2009). Previous theorists have emphasised the role that this form of autobiographical memory might play in structuring retrieval of past and construction of future events (Bohn & Berntsen, 2011; Irish et al., 2012; Prebble et al., 2013), structuring a life story (Thomsen, 2009) and maintaining identity in the face of episodic deficits (Klein & Lax, 2010; Rathbone et al., 2009). Here we suggest that this factual knowledge about the self might play an additional role, through its close relationship to our judgements of well-being. Further experimental research is needed to unpick the causal mechanisms between semantic self-image valence and well-being. However, the results of this study suggest that there are very close ties between perceptions of well-being and the semantic facts that comprise self-knowledge. This finding relates to current theories on the central role that imagery plays in emotion (e.g., Holmes & Mathews, 2010). Indeed, recent studies suggest that positive imagery acts to scaffold positive affect (Blackwell et al., 2013; Torkan et al., 2014). Within the autobiographical memory literature, imagery is more typically a hallmark of episodic, rather than semantic, autobiographical memory (e.g., Tulving, 1983, 2002). However, there is a wide literature on the clinical relevance of many different forms of imagery (Holmes & Mathews, 2010), including the construction of negative semantic images of the self (e.g. Conway, Meares, & Standart, 2004; Hirsch, Clark, Mathews, & Williams, 2003; Ng & Abbott, 2014). More research is needed to understand the nature of mental imagery associated with semantic self-images. Semantic self-images, by definition, reflect a particularly self-relevant form of imagery. Although they represent an amalgam of semantic information (e.g. being confident, sociable, keen on attending football matches), rather than discrete episodic events, the emotional valence of these semantic self-images appears to be closely linked to well-being. Future work based on cognitive bias modification paradigms may shed more light on the potential therapeutic benefits of modifying semantic self-images, in light of the success of using this approach with mental images that are episodic in nature (e.g., Torkan et al., 2014). The present findings support the results of previous research. Firstly, across both age groups, participants’ emotional ratings for both episodic autobiographical memories and semantic self-images were predominantly positive, replicating previous work (e.g., Chessell et al., 2014; Thomsen, 2009) and supporting the theory that most people tend to view themselves and their memories in a positive light (Walker, Skowronski, & Thompson, 2003). Socioemotional selectivity theory (SST; Carstensen, 2006) may explain why, for older adults in particular, more positive views of the self were associated with increased well-being. According to SST, the age-related decrease in subjective sense of future time (e.g. how much time one has left to live) leads to increased emotional regulation. Furthermore, the fact that well-being showed strong associations with the emotional valence of semantic self-images is of particular relevance to older adults, as well as various clinical groups who demonstrate episodic deficits (e.g., Dalgleish et al., 2007; Klein & Lax, 2010; Martinelli et al., 2013). The older adults in this sample represented a healthy community-based sample, none of whom demonstrated severe memory deficits. It is thus possible that more advanced memory loss would have a greater impact on well-being (e.g., Llewellyn et al., 2008; Wilson et al., 2013). The question of how memory, well-being, and identity are inter-related in dementia is complex, and recent papers on the topic have highlighted the importance of studying identity using a range of approaches (e.g., Caddell & Clare, 2013). The results of the present study also extend previous findings. Although we found few relationships between ratings of episodic autobiographical memories and well-being, other researchers describe significant relationships between well-being and episodic memory. For example, Philippe et al. (2012) found that young adults who retrieved memories that were rated more highly for need satisfaction tended to demonstrate increased levels of well-being. However, Philippe et al. analysed a specific component of episodic memory, related to achieving growth- related goals, whereas in the current study (in order to allow comparison of self-images and episodic memories) we used an emotional valence scale to analyse the relationship between memory and well-being. It is thus possible that other measures of the content and affective properties of episodic memories, such as redemption and contamination themes (McAdams, Reynolds, Lewis, Patten, & Bowman, 2001) or intrinsic versus integrated themes (Bauer et al., 2005), would yield significant relationships with well-being. Crucially, few studies that measure the relationship between episodic memory and mood or well-being (e.g., Gillihan, Kessler, & Farah, 2007; Philippe et al., 2012) also include measures of semantic autobiographical memory, such as self-images. Thus it is possible that when memories have been shown to have an impact on well-being, they may have done so via their relationship with semantic facts about the self. This mediating role of semantic autobiographical memory has been proposed in models of memory and identity (e.g., Haslam et al., 2011). It is possible, therefore, that semantic self-images play a similar mediating role in promoting well-being, and this again would be an interesting avenue for future research. The present research was subject to several limitations and suggests multiple avenues for future work. First, the study employed a correlational design and as such it cannot speak to the causal direction between the constructs of semantic self-image valence and well-being. Whilst it is possible that viewing the self in a positive light promotes well-being, it is equally possible that higher levels of well-being lead people to take a more positive view of themselves. Future studies employing experimental designs will be needed to unpick the direction of effects, however it seems likely that a bi- directional relationship exists between semantic self-image and well-being (as would be predicted by models such as the self-memory system; Conway, 2005). Indeed, the self-system is widely regarded to be dynamic and likely to change depending on mood and context (e.g., Markus & Kunda, 1986; Showers, Abramson, & Hogan, 1998). What the present research clearly indicates is that the emotional appraisal of semantic self-images is closely related to a range of well-being measures. Regardless of bi-directional effects, if focusing on positive self-images could be shown to raise levels of well-being this would have important therapeutic implications. A second point relates the lack of counter-balancing between the set of well-being measures and the memory tasks. To avoid priming particular aspects of identity, the semantic self-image generation task was completed first. Thus, future studies will be required to rule out potential order effects.","In their recent review, Prebble et al. (2013) suggested that research needs to explore the semanticised summaries of life narratives, as well as specific event memories, in order to develop a more complete picture of the roles of episodic and semantic autobiographical memory. The present paper addressed this aim and found that the valence of semantic self- images may be more fundamental to conceptions of well-being than that of episodic autobiographical memories. In essence, we found clear relationships between how positively people perceive themselves and their overall well-being. This relationship was particularly pronounced in older adults, for whom a wide range of well-being measures correlated with ratings of semantic self-image valence. In addition, older adults’ performance on standardised measures of cognition and memory bore no relation to ratings of well-being. These findings have implications for promoting well-being in aging. It seems that well-being does not depend on what you remember, or even how good your memory is – what is key is how you conceptualise your sense of self in the present moment."],["Viewing the brain as an organ of approximate Bayesian inference can help us understand how it represents the self. We suggest that inferred representations of the self have a normative function: to predict and optimise the likely outcomes of social interactions. Technically, we cast this predict-and-optimise as maximising the chance of favourable outcomes through active inference. Here the utility of outcomes can be conceptualised as prior beliefs about final states. Actions based on interpersonal representations can therefore be understood as minimising surprise - under the prior belief that one will end up in states with high utility. Interpersonal representations thus serve to render interactions more predictable, while the affective valence of interpersonal inference renders self-perception evaluative. Distortions of self-representation contribute to major psychiatric disorders such as depression, personality disorder and paranoia. The approach we review may therefore operationalise the study of interpersonal representations in pathological states. © 2014 The Authors. --------------------------------------------------------------------------------","The sense of self may be experienced at many levels – from the elementary, pre-verbal ‘minimal self’ that accompanies all conscious perception, to the purposeful, historically constructed ‘narrative’ self who takes action under the conscious guidance of goal and context. Research based on the idea of the brain as a probabilistic inference device has seen great advances in recent years (Chater & Oaksford, 2008; Friston & Stephan, 2007), allowing important aspects of the minimal and narrative self during perception and action to be considered in the light of how probabilistic prediction interacts with sensory evidence. The computations that brains perform to predict and hypothesis-test underlie what it is like to be an I who expects the consequences of acting and perceiving – now and through time (Hohwy, 2007). In this article, we extend this work to self-perception in the interpersonal domain, while acknowledging that the sense of self is important, even in the absence of interactions with others. Simple observation suggests that the interpersonal self is as complicated in its detailed mechanics as it is blatant about its presence. While the minimal self at the core of near-instantaneous perception is difficult to put into words, there is, in the first instance, nothing difficult about putting the interpersonal self into words: ‘am a kind person’, ‘am not as good as her’. The English language describes this powerful self-perception with expressions such as ‘he is a terribly self-conscious’. We claim that the interpersonal self is actively inferred during social exchanges and that many of its properties correspond to the means and ends of a machinery of probabilistic inference. Inferring self-representations may thus help achieve desirable ends (social outcomes). The evidence that we marshal to develop this argument comes from a wide variety of sources, including computational neuroscience, brain imaging, psychiatry, social and clinical psychology. We develop our claim as follows. First, we provide evidence that some high-level, affectively coloured (pain) perception is well described in terms of basic Bayesian reasoning. Second, we describe an extended framework of approximate Bayesian reasoning, namely active inference, which encompasses agency and decision-making. Third, we review the kinds of psychological construct upon which active inference may operate – and show that inferring about these constructs subserves important goals. Fourth, we suggest a model of interpersonal exchange that could form the basis for empirical study. Finally, we examine the psychiatric relevance of making affectively charged inferences, especially about the self. We conclude by discussing the limitations of our approach and the implications for future research. Using Bayesian inference to make sense of experience ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The Bayesian approach considers probabilities to be degrees of belief, so that Bayesian inference has the following form. If I make an observation o, what should become of my belief P (S = s) that some relevant aspect of the world is in state s? For example,1 if o = ‘Emil gave me a present’, what should become of my belief ‘I am a bad person’? If the new observation is surprising – with respect to the existing belief framework – the framework is poor at predicting the observation. It therefore needs to be updated if it is to describe the world more adequately. This updating of beliefs is the essence of Bayesian inference, which adjusts the agent’s model of the world so as to render new observations (data) less unpredictable. Although a full description of this well-established formal approach is outside the scope of the present article, the interested reader is referred to (Chater & Oaksford, 2008; Friston & Stephan, 2007; Friston et al., 2013; King-Casas et al., 2008). The claim we make in this paper is that this inferential framework applies to all beliefs – including beliefs about the self. In a Bayesian framework what the brain minimises as it makes inferences, including inferences about the self, is unpredictability and not, for example, proximal discomfort. We will consider an example of this below, in the case of perception of pain. We reformulate the principle of psychological economy as follows: the primary gain of a representation is its power to predict outcomes that matter under some prior beliefs. Maximising predictability is equivalent to minimising surprise. Clearly, surprising outcomes rest upon prior beliefs. In our case, these beliefs will be about the self (and others). Crucially, surprise can be quantified as the negative log (Bayesian) evidence for a model. This means that minimising surprise maximises the evidence for a model or representation of interpersonal exchange. We now turn to a simple but informative application of the Bayesian framework, the understanding of placebo responding. Placebo responding crucially depends on an interaction between prior beliefs about analgesia and sensory evidence (Morton, El-Deredy, Watson, & Jones, 2010). This case study will help structure further discussion in two ways: on the one hand, its limitations will motivate the need for goal-directed, active inference; but on the other, placebo- responding provides important lessons for inference about self-representations. The Bayesian model of pain perception ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The Bayesian model of pain perception2 (El-Deredy, Trujillo- Barreto, Watson, & Jones, 2010; Watson, El-Deredy, Bentley, Vogt, & Jones, 2006) provides proof-of-principle that humans perform high-level Bayesian inference to form affectively charged percepts. These researchers modelled pain perception in two groups of healthy people, ‘placebo responders’ and ‘placebo non-responders’. Pain ratings were collected through three phases. In both groups: Painful (‘active’) stimuli were initially administered in the absence of any treatment, establishing an expectation of stimulus – pain perception. Placebo analgesia was then administered while stimulation was covertly switched to sham stimuli. These were visually identical but non-painful, giving the impression that the ‘treatment worked’. Stimulation was finally covertly switched back to active stimuli. The above Bayesian updating model of pain rating allows one to address the following question: given beliefs subjects already entertained (priors) about how painful percepts are generated, how are current percepts integrated into new beliefs (posteriors)? Crucially, ‘new beliefs’ include the precision or the weight that should be attached to past reports. The role of precision is crucial, because agents come to rely more on cues that have the greatest precision. Notice that the previous response enters as a new observation in this Bayesian updating scheme – in other words, the model is observing and trying to explain its own responses. The authors found that the full model explained the pain-perception data of placebo non-responders, but a reduced model – which neglects sensory information (ot) and makes predictions based only on past reports – best accounted for the data from placebo- responders.3 In other words, one’s inference about the causes of nociceptive input can be based purely upon previous reports or, equivalently, behaviour that belies one’s own inferences. Bayesian updating thus provides a good account of behaviours involving high- level beliefs about analgesic properties. This example also provides proof-of-principle that apparently irrational, affectively charged human beliefs can be described quantitatively in terms of the balance of prior beliefs relative to sensory evidence. Interpersonal interactions also involve apparently irrational, affectively charged beliefs and the related negative affect is neurally and subjectively related to pain (Eisenberger, 2012). Neither physical pain nor social ‘pain’ stand in a one-to-one relation with damage: people have widely varying sensitivities to the same physical manipulation (Eisenberger, Jarcho, Lieberman, & Naliboff, 2006). The experience of physical pain is not unrelated to self-awareness (‘Pinch me, am I dreaming?’) but interpersonal ‘pain’ is quite directly related to self- and other- representation. When I am deceived, for example, beliefs like ‘I am a fool’ and ‘he is devious’ immediately gain weight. Much as inferences about pain relate to the risk posed by stimuli (and, more subtly, my own physical vulnerability), so during interpersonal inference the self and others are vividly experienced as vulnerable (or not), noxious (or safe), etc. We will consider examples of inferring aspects of the self to be noxious, when we consider psychiatric conditions later on. This brings us to a key limitation of this Bayesian model of pain perception, which is its silence as to the functional role of inference about pain (or, in our case, about the self). In fact the optimal readiness with which pain is to be inferred depends on context: organisms can tune their own pain perception according to both their prior beliefs and the specific biological goals they believe are attainable in that context (Boureau, 2005). This would be an irrational anomaly if pain were a raw datum, but not if pain was a motivational force: if there is little of biological use that can be done, it is better to reduce pain sensitivity. We therefore need a Bayesian framework that explicitly represents the person’s agency and goals, namely active inference (Friston et al., 2013). In order to study self-awareness as inference, we need to quantify the relevant aspects of self- representation. We will consider this in more detail below; for now, suffice it to say then we can cast traits like ‘fairness’ or ‘jealousy’ as social preferences: how much another’s pain or gain can act as a motivating force (Camerer, 2003). We propose that such sensitivities (preferences) may stand to social outcomes as pain sensitivity stands to somatic ones. That is, their usefulness may lie in helping the agent reach their goals. Using active inference to reach desirable goals ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ During perceptual inference, we infer those descriptions of the world that are most consistent both with our sensory data and with our prior beliefs about what the world should be like. During active inference, however, we can entertain beliefs about alternative scenarios that lie ahead – scenarios that depend on our own fictive choices in the future. Policies are then chosen in a Bayes optimal fashion that depends upon these prior beliefs. Our ‘prior beliefs’ are not just about what we believe the world to be like, but which alternative outcomes we realistically expect (hope) that we can reach: our goals. From now on we will refer to prior beliefs about goals simply as ‘goals’. Conversely, once a subject has experienced part of the scenario, but has not as yet reached the final outcome, they can be said to hold empirical priors about their policy (i.e. the choices that are shaped by their goals). These empirical priors pertain to (a) goal attainability from the current state and (b) beliefs about future choices (Friston & Stephan, 2007; Friston et al., 2013). Empirical priors are a necessary aspect of inference in hierarchical models. Technically speaking, they are prior beliefs that are conditioned on (i.e. depend on) other unknown variables. In our case, beliefs about the policy depend upon beliefs about hidden states of the world, where beliefs about future states depend upon the goal. Whether I have to get off the bus may not just depend on whether I see a park from the bus window, it also depends on which bus I took! In this sense, beliefs about the policy entail beliefs about hidden states in a hierarchical sense and are therefore empirical priors. They are essentially prior beliefs about prior beliefs. To make the distinction between goals and empirical priors over policies clearer, consider going to the pub for a drink. My favourite drink is actually fairtrade hot chocolate; I also like Guinness, but less so. If they were in front of me I would choose fairtrade chocolate ¾ of the time, Guinness ¼ of the time, so these are my goals: [0.75, 0.25]. This illustrates how – in our framework – all utilities are necessarily relative: the utility of a particular outcome is always defined in relation to allowable alternatives. The computational purpose of defining goal priors can be seen as separating out the various task-dependent probabilities used in active inference from probabilities that do not depend on where we are in the task or on the agent’s behaviour during the task. In other words, we separate state-dependent beliefs about what we will do next from the goals that define beliefs about the final outcome. In contrast to goals, empirical priors entail beliefs about the dynamics of the task or, more simply, the consequences of a particular action in terms of the transition from one state to another: where will I end up if I take this sequence of actions? And how does the disparity between that outcome and my goals influence my choice of a subsequent policy? If I take a bus towards Barnet (action), am I likely to get to a nice pub (state) and look for hot chocolate there (policy)? In short, empirical priors entail the agent’s knowledge about the dynamics of her/himself in the world. They are ‘priors’ in the sense that the actual policy that we will choose (shall I look for hot chocolate?) will depend on observations about where we are and where we can go. Crucially, in the active inference framework, these choices will also depend on the confidence we have in reaching our goals from the current state. This is because action depends upon beliefs over policies and beliefs always have a precision. Our inference framework is thus able to decide about actions motivated by goals that embody the raw power of gain and loss, contentment and pain. The key device to achieve this, introduced in this review, is to convert the utilitarian formulation of classical economic theory, in which choices are assumed to maximise expected utility, into a pure (Bayesian) inference problem. One can do this by representing the utility of outcomes as prior beliefs about final states. This means that instead of making choices to maximise expected utility, one simply minimises surprise – under the prior belief that one will end up in desirable states. Practically, replacing utilities with prior beliefs means that one can appeal to well-established inference schemes such as (variational) free-energy minimisation in order to prescribe normative behaviour. Casting utility functions as prior beliefs means that one can understand utility – which depends on interpersonal factors – in terms of beliefs about oneself and others. As an example, assigning high utility to gains acquired by a partner who resembles me translates into holding a higher prior belief that the partner will be of a similar type to me and that they will gain from the exchange. So far we have only used common-sense examples; in order to apply the formalism of active inference to self- and other- representation, we must determine – at least roughly – the categories of belief about self and others that people actually use. The nature of beliefs about the self (and others) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ During self- and other-perception, empirical evidence suggests that people form beliefs about both the states and traits, of both self and others (Bem, 1972; Bentall, 2003; Fowler et al., 2006). Traits are characteristic patterns of emotion, behaviour or thinking that subjects engage across many situations. People naturally infer such traits in each other, in order to anticipate intentional mental states. Psychological research suggests that people independently score ‘positive attributes’ and ‘negative attributes’ of both ‘self’ and ‘other’. A modern clinical research instrument, used to assess such representations, is shown in Table 1. The items relate to (i) (social) outcomes; e.g. ‘respected’, ‘a failure’; (ii) Capabilities; e.g. ‘weak’, ‘talented’ and (iii) Preferences for acting according to social values; e.g. ‘good’, ‘fair’, ‘trustworthy’. Note that the instrument is asymmetric: ‘other’ attributes relate more to social-value preferences, such as fairness, harshness, etc.; whereas ‘self’ attributes relate more to social outcomes and abilities. This division is likely to be an artefact of the clinical focus of this scale; namely, depression and paranoia. If people are asked to describe ‘what kind of person’ they desire to be and ‘what kind of person they try to avoid being’, they give a mixture of success-, motivation- and ability-related traits for themselves too (Francis, Boldero, & Sambell, 2006). We will call a ‘type’ the traits of a person that are relevant to the current context (e.g. an interpersonal task). Building on the work on inferring an other’s type (Ray, King-Casas, Montague, & Dayan, 2008), we hypothesize that people harvest observations to update their beliefs about their own type for good reason: Self (and other) representations are tools facilitating the efficient computation of interpersonal behaviour. An important psychological theory that makes contact with these issues is the ‘sociometer theory of self-esteem’ (Leary, Tambor, Terdal, & Downs, 1995). According to the sociometer account, self-esteem has an important functional role, which is to indicate one’s likely evaluation by the social milieu. The crucial corollary is that actions that improve self-esteem are rewarding because – if all goes well – they subserve socially sanctioned goals. We generalise the ‘sociometer theory of self-esteem’ is to a ‘sociometer theory of self-representation’. Type-based interpersonal representations, of which (trait) self-esteem is only a subset, serve to optimise context-dependent social computations.4 Aspects of self-representation important for prediction may be: (i) how successful I have been in a given context so far; (ii) an appraisal of my capabilities; and (iii) what my (interpersonal) preferences are. Constructs such as ‘a bad person’ or ‘fair’, can inform optimal decision-making in a formal and fundamental fashion, despite the emotive and informal nature of these concepts. Our contention that (interpersonal) preferences should be the object of inference parallels Hohwy’s hypothesis that desires are inferred in the process of applying generative models of ourselves (Hohwy, 2007). The coherence of these models would be important for the construction of the ‘narrative self’. In an interpersonal context, traits such as ‘talented’, ‘harsh’ etc. can feed directly into the outcomes that a person wants to reach (or indeed to avoid). In terms of active inference, agents can inform goals by the desirable self-representations that they are likely to reach as outcomes. An example is: “If the outcome of this exchange is that I cooperate and my partner runs off with the money, this would be evidence that I am a fool. This is highly undesirable – I’ll attach very low probabilities to outcomes implying that I’m a fool”. We hypothesise that different people are equipped with different prior beliefs, acquired during their upbringing and built on a base of genetic preparedness. This allows modelling of individual variation, including variation in psychopathology. It provides a simple and graceful way to account for different types (phenotypes) in a quantitative and formal sense. Clearly, people may have a vast lexicon of potential traits to consider. We expect the ensuing paradigm could the used to inform those traits that are inferred, especially in clinical applications. Self-representation and desirable goals of social interaction ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Type-based representations provide a simple and formal belief space that people may use in categorising themselves and others – a belief space that can be tested empirically. These representations can be linked to the returns (e.g. utilities or prior beliefs) that people expect themselves and others to derive from social actions. An agent’s task is then to infer, or estimate, ‘types’ linked to utilities. If, for example, I infer that my partner is a ‘fair person’, this summarises an expectation: that they will avoid actions leading to inequitable returns for all involved – finding such actions aversive. If I estimate myself to be ‘highly competitive’, I expect to derive high utility from actions that lead to me doing better than others, even if my material returns are somewhat lower than under some other outcome (e.g., where I come second). But why bother using person- representations as heuristics? Why not just calculate out how everyone should behave based on their self-interest? The tractability of optimal Bayesian inference is a key issue here; real agents must perform the requisite computations, which can be very hard. Type- based interpersonal representations can be likened to heuristics that facilitate approximate Bayesian inference. Here type-based representations can be a shortcut to tracing out a complex tree of possible future states and outcomes. In the ‘stag hunt’ game, for example, agents with limited depth-of-thought but equipped with ‘prosocial’ preferences make decisions equivalent to having greater depth-of-thought but no social biases (Yoshida, Dolan, & Friston, 2008). In such cases, explicit optimal solutions are often prohibitively complicated to compute. When framed in terms of heuristics or prior beliefs, approximate Bayesian inference provides a tractable and – in some instances – neuronally plausible scheme for optimisation. This approach marks a significant departure from the traditional (normative) behavioural-economic approach, where an agent’s beliefs about their own type do not depend on their actions and vice versa. It also differs from traditional psychological approaches, where self-representations – including theories based on self-esteem or on unconscious representations – are seen as the object of a more- or-less self-contained homeostasis, rather than an explicit heuristic of how to behave optimally in social exchanges. Self- and other-representation in a model Trust Task ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Game-theoretical paradigms can induce people to infer traits in others and express their own traits in beliefs and behaviour. The ‘Trust task’ is a prototypical paradigm (King- Casas et al., 2008; Ray et al., 2008). In each of several rounds one of the participants, in the role of the Investor, is given a certain ‘wage’. They can invest any fraction fI of that ‘wage’ with the Trustee. The Trustee makes a profit, which is usually taken to be 200%. They then keep as much of this as they want and return to the Investor any fraction fT of the ‘wage’. How much will the Investor entrust to the Trustee? Repeated rounds allow each player to try to deduce the ‘type’ of their partners. Are they disposed to cooperate (fT > fI), or maybe ‘grab the money and run’ (fT ≪ fI)? Behavioural studies demonstrate healthy people often ‘repair cooperation’ that falters, but patients with borderline personality drift towards the Nash equilibrium, suppressing overall incomes (King-Casas et al., 2008). The Nash equilibrium has a paranoid flavour: in this game, it says that in the last round the Trustee will ‘take the money and run’, and therefore should not be trusted. The same logic holds for the penultimate round, and so on through a process of iterative backward inference to the very first one. Therefore the Investor should never trust any money to the Trustee. The Trust Task has been extensively studied, and is attended by a large amount of psychological and neurobiological data (Xiang, Ray, Lohrenz, Dayan, & Montague, 2012). However, existing models do not make adequate contact with clinical psychology. Here, we consider a framework for modelling this task that is equipped with minimal clinically relevant self- and other-representations. This can be used to analyse the behaviour of people who might hold unstable or distorted interpersonal representations. Our framework extends classical models (Camerer, 2003; Ray et al., 2008) that lack flexible, computationally functional self-representations. Let us consider agents who think about alternative future scenarios to a limited depth-of-thought (‘one, two steps …’) and then consider how they and their partners would ‘come across’ in the long term (‘… infinity’). This gives interpersonal representations a functional role. Unlike classic neuroeconomic models – and closer to clinical psychology – at the end of each round subjects have to update their own self-image based upon their actions and those of their partner. We therefore have a Bayesian Attribution – Self-Representation (BASR) model, formalising the attribution theory suggested by Bentall (2003). In the BASR model, working out each possible scenario two moves into the future enables the ‘type’ of each partner at that future point to be estimated (Fig. 1). There are at least two interesting ways in which the estimated types of self and other may be useful. First, people may have intrinsic preferences as to the type of person they want to be, and the type of person they want their partner to be. The framework of active inference allows us to think about ‘wanting the partner to be’ as ‘believing that the partner will indeed turn out to be’. The preferences inform goals over the sort of person they are. We could, furthermore, hypothesize that such goal preferences are context specific: contexts may, for example, be labelled as competitive or cooperative, and stronger priors attached to corresponding person types. Second, types can be exploited in explicit approximations of the consequences of behaviour. An expectation that long-term behaviour will be consistent, on average, with the inferred player ‘types’ would lead to an estimate of long term returns. This would then be absorbed into appropriate goal priors. Overall, the BASR model of a simplified Trust Task would run as follows. First, agents know what actions are available. Let’s allow them just two options in each round of 30 rounds, a high (cooperative) or low (uncooperative) contribution fhigh or flow. Second, agents hold beliefs about how ‘cooperative’ each player is, i.e. the player’s types. We can think of ‘cooperative’ being a positive, high-esteem trait in this context. Third, agents estimate the likely evolution of the game a small number of moves into the future – e.g. ‘I play X, she plays Y, then I play Z’. This evolution may be more or less likely according to the agents’ empirical priors. It may also be more or less compatible with their goal priors, as we saw. Agents also have a sense of how precise their prediction of reaching their goals may be; this ‘precision’ is important for the mechanics of active inference but outside the scope of this review. The interested reader is referred to our technical exposition (Friston et al., 2013). Finally, the player whose turn it is to play infers an updated view of the interaction. This maximises the consistency between their prior beliefs about the world, their observations, their goals, their confidence (precision) in reaching those goals – and of course their policy. Their action then flows from their preferred policy, and it’s the other player’s turn. The interpersonal self in psychiatric disorders ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Severe psychiatric disorders often illustrate terrible distortions and deficits in self- perception. Let us consider two examples: “I cannot live and I cannot die, because I have failed so much, I shall bring my husband and children to hell ... I shall go to the convict prison and my two girls as well, if they do not make away with themselves because they were born in my body.” This extract is from a letter written by a severely depressed patient of Emil Kraepelin – the patient who failed to update her self-image in response to being given a nice present. It expresses a nihilism that remains familiar to contemporary clinicians, who are all too aware that a distorted representation of the self can be a debilitating symptom in severe depression. In the second example, self-perception is deficient: “Isn’t America supposed to be the land of the free? How come, if I’m free, I can’t deprive a stupid f∗∗∗ing dumbs∗∗t of his possessions if he leaves them on the front seat of this f∗∗∗ing van out in plain sight...? Natural selection. F∗∗∗er should be shot.” This is a quote from Mr. E. Harris, a mass murderer who attracted a diagnosis of psychopathic disorder (Cullen, 2009). It illustrates how he both talked and acted as if his representations of other people carried no value. This was most striking when the prospect of their suffering seemed to hold no aversive value for him. Moreover, he was devoid of any concern that his beliefs and behaviour would devalue him as a person, de facto rendering him despicable in any community. Contemporary psychiatric practise acknowledges the importance of self-representation. Descriptive diagnostic criteria, such as the DSM-5, define a disturbance of the representation of self and the evaluation of others as core features of Personality Disorders (Skodol et al., 2011). Moreover, several empirically validated therapies assume that self-representations are important links in a causal chain that leads to psychiatric disorders. For example, cognitive-behavioural therapy postulates that ‘core beliefs’ related to the self, such as “I am unworthy”, underpin depression (Waller, Shah, Ohanian, & Elliott, 2001). Many mentalization-based therapies propose that infants internalise representations that caregivers make available to them – a process which, if it goes awry, may lead to conditions such as borderline personality disorder (Allen, Fonagy, & Bateman, 2008). The less evidence-laden approach adopted by psychoanalysis posits fragmented representations of others (part-objects) as underlying severe mental disorders. In Borderline Personality disorder, patients often fail to predict the damage that their actions cause, in terms of the way they are perceived by other people and the ruinous consequences this has for the person themselves (Allen et al., 2008). When self and others are perceived as noxious entities ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Many psychiatric conditions might be described as ‘nocebo states’. In nocebo states physiologically inert stimuli have aversive effects, e.g. an inert tablet causing nausea. This is a mirror image of the placebo effect. In psychiatry, features of the self or the environment that most outsiders would regard as innocuous are often perceived as toxic. A common example of perceiving a facet of the self as toxic is as follows. A patient with borderline personality quarrels with a relative about an everyday matter. This activates loathing of the self to such a degree that the patient hides away and carves insults on her skin with a razor. However most common, distressing ‘nocebo’ perceptions in psychiatry are unwarranted (to the outside observer) versions of everyday concerns – such as that others are devious and the self is worthless, as exemplified by the study by quoted in Table 1 (Fowler et al., 2006). In many clinical accounts it is not only self- representations that are dysfunctional: efforts to bolster specific aspects of self- representation are also seen as maladaptive. However, the precise role of self- representation is highly controversial, partly due to the fact that it is difficult to quantify and access in a strictly empirical manner. This highlights the fact that despite a focus on self-representation over recent years, its normative function is poorly understood. There is obvious value for interpersonal exchange in asking “what sort of person is the Other?” However, the value of healthy inference about “what sort of person do these actions make Me” has not being addressed so thoroughly. The subtle clinical concept of ‘mind-blindness’ is relevant here. Clinicians mean by this that borderline patients, especially under stress, show an inability to reflect about other people’s minds or indeed their own. Increasing this ability through gradual learning is the basis of Mentalization Based Therapy (Allen et al., 2008). As we saw, placebo responders are characterised by the absence of a term reporting sensory information about the outside world (no term in ys,t−1 in Eq. (1)). There are indications that people with borderline personality disorder may have an analogous but more subtle deficit in interpersonal interaction. When interpersonal Trust falters they fail to ‘signal’ to their opponent, by risking part of their income, that they are trustworthy. This leads to an unravelling of trust. In addition, healthy people playing borderline partners display a reduced level of theory-of-mind, as would be expected if the borderline partner were deficient in higher- order theory-of-mind terms (Xiang et al., 2012). The ‘missing term’ in these patients’ model of the world would be exactly “what sort of person would these actions make Me?” We thus have a computational neuroscience counterpart of the clinical concept of ‘mind- blindness’. Estimation of social threat and opportunity using active inference requires not only prior beliefs about self and others, but also their approximate Bayesian updating – including the evaluation of alternative future scenarios. Several cognitive deficits and biases, such as mildly reduced IQ and working memory or a ‘jumping to conclusions’ cognitive style, are associated with paranoid psychosis (Bentall et al., 2009). How such non-specific cognitive deficits may contribute to the aetiology of abnormal beliefs is the subject of much debate. Affective biases are likely to involve prior beliefs about the nature and likelihood of noxious states in self and others. Could nonspecific cognitive deficits shift interactions from cooperative to suspicious in the presence of a specific distributions of prior beliefs? The cognitive-deficit part of this would be analogous to the way that play can shift from cooperative to non-cooperative under the influence of reduced depth-of-thought (Yoshida et al., 2008). Paranoia may result from prior beliefs about self and others that are activated by specific, aversive contexts – and maintained by a reduced cognitive ability to consider more complicated but benign scenarios. An important factor in generative models of paranoia may be the belief that persecutors do not attach aversive motivational value to the persecuted person’s distress. Such a model is actually true if one’s partner has psychopathy. The hallmark of this is that the partner is unable to attach aversive motivational value to the suffering of others (like Mr. Harris above) despite being able to infer its presence. The embarrassment of Bayesian riches ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Active inference assumes that subjects optimise a probabilistic representation of the environment. This assumption is not one of perfect or rational behaviour but of a type of bounded rationality based upon a potential universe of prior beliefs or heuristics. Previous research has also used this philosophy, of perfect optimisation within a suboptimal cognitive model (Moutoussis, Bentall, El-Deredy, & Dayan, 2011). When it comes to psychopathology, however, we encounter a difficult problem: Is ‘deviant’ behaviour the result of inefficient model optimisation or due to an ‘inefficient’ model of the world? In other words, are prior beliefs (the model) abnormal, or is psychopathology the result of ‘broken’ active inference? This is a particularly important and challenging question as on the one hand, any behaviour can be explained as optimal Bayesian inference given some set of priors and utility function; while some behaviours are undoubtedly the result of what we may call ‘broken’ active inference. Examples of the latter are likely to include drug- induced psychosis or dementia. How can one’s type be unknown? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ If, as we maintain, a person’s type enters into their decisions, isn’t that person’s type tautologically the one which determined those decisions? In what sense can such a type be unknown to the subject, and have to be inferred? At least three alternatives should or could be tested. First, there may be separate decision-making and self-evaluating parts of the self, something akin to an actor-critic architecture in reinforcement learning (Sutton & Barto, 1998). Parenthetically, no self-deception needs to be invoked here. Beliefs posterior to observing the world, including the self, need not accord with the maximum likelihood of observations: other constrains enter inference too, as we have seen in some detail. Second – in a related vein – there may be no well-defined ‘type’. The history of psychology and psychiatry is littered with abandoned models of personality traits and types, while stereotypes about groups have little predictive value. Like these models, people’s self-representations (e.g. ‘I am a music-lover’) may simply be instances of the fundamental attribution error. In this case, the approach suggested above might help delineate stable patterns of interactions and attributions, including self-fulfilling prophesies, which describe the situation at hand but have no deeper explanatory power. Third – and most interesting – types may have to be inferred and applied in relation to social norms. How much work a person has to do for others to be considered ‘kind’, or what offers a ‘fair’ person might make in an ultimatum game? (Xiang, Lohrenz, & Montague, 2013). If ‘kind’ or ‘fair’ people are treated on average better than ‘unkind’ or ‘unfair’ ones, then inferring one’s own type becomes even more important, and more challenging, for each social partner.","In conclusion, we have reviewed two key hypotheses. First, that people make inferences about themselves – and others – to minimise interpersonal surprise enabling them to make decisions that are most consistent with their model of the interpersonal world. Second, we suggest that the framework of active inference can capture the iterative dynamic of ‘where do I desire to get’, ‘what do I decide’ and ‘what do my actions make me’. The programme of research that this analysis speaks to is to characterise different prior beliefs (about contingencies, utilities, prosocial utilities and types) in relation to choice behaviour, under the assumption that subjects engage in approximate Bayesian inference. Bayesian subjects are assumed to update of models of their world, while Bayesian researchers infer models of individual variability of choice behaviour (we could call the latter meta- Bayesian fitting). Both can be implemented through free energy minimisation. The updating of self-representations by subjects could then be mapped onto neurobiological substrates, e.g. via functional neuroimaging. This approach could help advance our understanding of distorted self- and other-representations in many conditions – including depression, personality disorder and paranoia. Practically, we envisage that the first step will entail characterising the updating of beliefs over other people’s type, which in turn results from exposure to others’ actions, under the simplification that only the immediate returns matter. One could then consider prior beliefs about outcomes in the long-term future, and their influence on self-representation. Finally, one might simulate and study the behaviour of interacting self-representing agents that can even exert choice over the types of ‘games’ they engage in. We hope to pursue these studies and the ideas reviewed above to help contextualise and motivate a formal focus on how we represent ourselves and how these representations determine our behaviour."],["We investigated early behavioural markers of autism spectrum disorder (ASD) using the Autism Observational Scale for Infants (AOSI) in a prospective familial high-risk (HR) sample of infant siblings (N= 54) and low-risk (LR) controls (N= 50). The AOSI was completed at 7 and 14 month infant visits and children were seen again at age 24 and 36 months. Diagnostic outcome of ASD (HR-ASD) versus no ASD (HR-No ASD) was determined for the HR sample at the latter timepoint. The HR group scored higher than the LR group at 7 months and marginally but non-significantly higher than the LR group at 14 months, although these differences did not remain when verbal and nonverbal developmental level were covaried. The HR-ASD outcome group had higher AOSI scores than the LR group at 14 months but not 7 months, even when developmental level was taken into account. The HR-No ASD outcome group had scores intermediate between the HR-ASD and LR groups. At both timepoints a few individual items were higher in the HR-ASD and HR-No ASD outcome groups compared to the LR group and these included both social (e.g. orienting to name) and non-social (e.g. visual tracking) behaviours. AOSI scores at 14 months but not at 7 months were moderately correlated with later scores on the autism diagnostic observation schedule (ADOS) suggesting continuity of autistic-like behavioural atypicality but only from the second and not first year of life. The scores of HR siblings who did not go on to have ASD were intermediate between the HR-ASD outcome and LR groups, consistent with the notion of a broader autism phenotype. --------------------------------------------------------------------------------","Younger siblings of children with an autism spectrum disorder (ASD) represent a high-risk group for ASD, with recent estimates of the recurrence rate in siblings as high as 18.7% (Ozonoff et al., 2011). This allows prospective study of development from the first few months of life in infants who will later go on to receive a diagnosis of ASD. Within such prospective infant sibling designs, identifying the earliest differences or markers in those who go on to develop autism is a research priority (Zwaigenbaum, Bryson, & Garon, 2013). Current aetiological models of autism propose that typical developmental trajectories are derailed by complex interactions between underlying genetic and neurological vulnerabilities, environment and behaviour, with cascading developmental effects (Gliga, Jones, Bedford, Charman, & Johnson, 2014). However, the details of these developmental processes are poorly understood. It is hoped that understanding the ordering and interactive influences of the earliest biological and behavioural perturbations will elucidate developmental mechanisms that lead to the pattern of symptoms and impairments that characterise the clinical phenotype, as well as protective mechanisms that differentiate those at familial risk who go on to have non-ASD outcomes. This may in turn point to targets for treatment as well as improving identification of those infants at highest risk for the disorder in infancy, allowing very early intervention to be put in place (Green et al., 2013, in press). Early behavioural signs of autism in high-risk studies ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Recent reviews of high-risk (HR) sibling studies find convergent evidence for the emergence of overt behavioural markers between 12 and 18 months of age that distinguish, at a group level, those infants who go on to receive an ASD diagnosis from other HR infants and low-risk (LR) groups (Jones, Gliga, Bedford, Charman, & Johnson, 2014; Zwaigenbaum et al., 2013). At this age, clinically relevant behavioural differences span both social communication (e.g. gaze following (Bedford et al., 2012); social referencing (Cornew, Dobkins, Akshoomoff, McCleery, & Carver, 2012)) and stereotyped/repetitive behaviour (e.g. repetitive play (Christensen et al., 2010); repetitive movement (Loh et al., 2007) domains, as well as motor (Flanagan et al., 2012) and attentional atypicalities (Elsabbagh et al., 2013), appearing to represent early manifestations of later ASD symptoms. However, there is considerable heterogeneity at the individual level in the pattern of symptom emergence. Prior to 12 months of age, however, few overt behavioural markers for autism have been identified (Jones et al., 2014). In a recent report, Jones and Klin (2013) found that in a small sample of HR infants who went on to an ASD diagnosis, fixation on the eyes declined between 2 and 6 months of age. Other behavioural signs in the first year of life have included reduced gaze to people (Chawarska, Macari, & Shic, 2013) and vocal atypicalities (Paul, Fuerst, Ramsay, Chawarska, & Klin, 2010). Experimental studies have detected atypical neural response to social stimuli such as dynamic eye gaze from as young as 6 months of age (Elsabbagh et al., 2012). The Autism Observation Scale for Infants (AOSI) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Whilst the research reviewed above and elsewhere (Gliga et al., 2014; Jones et al., 2014; Zwaigenbaum et al., 2013) has mostly used experimental tasks/paradigms and observational and parent-report methods, there is also a clinical need for an instrument that allows systematic observation of early-emerging atypicalities in infants at-risk for ASD. The Autism Observational Scale for Infants (AOSI; Bryson, Zwaigenbaum, McDermott, Rombough, & Brian, 2007) is a semi-structured, experimenter-led behavioural assessment designed to measure early behavioural markers of ASD in infants aged between 6 and 18 months. These include atypicalities or delays in social communication behaviours (e.g. anticipatory social response, social babbling, orientation to name, eye contact) and non-social behaviours (e.g. disengagement of visual attention, motor control and behaviour, atypical sensory behaviours) as well as aspects of temperament (e.g. reactivity, ease of transitions between activities). Preliminary findings from the instrument's authors’ HR sibling cohort suggested that AOSI scores by 12 months but not at 6 months were promising as a predictor of later ASD outcomes based on 24 month ADOS classification (Zwaigenbaum et al., 2005). The same group have reported that whilst many individual AOSI items across domains (e.g. orients to name, eye contact, reactivity) at both age 6 and 18 months differentiated the group of HR siblings who go on to have an ASD classification at 36 months and LR controls, only atypical motor behaviour differentiated HR siblings who go on to have ASD from HR siblings who do not at both timepoints (Brian et al., 2008, 2013). Early behavioural atypicalities measured by the AOSI at 12 months have also been shown to characterise nearly one fifth of HR siblings who do not go on to have ASD (Georgiades et al., 2013), consistent with the notion of sub-clinical manifestations of ASD being present at an enhanced rate in family members of individuals with ASD, referred to as the broader autism phenotype (BAP; Bolton et al., 1994). The current study ~~~~~~~~~~~~~~~~~ The present study sought to replicate in an independent sample whether predictive associations exist between AOSI scores in early (at around 7 months) and later infancy (around 14 months) and ASD outcome at 36 months. In the present study we analysed AOSI data from both 7 and 14 month timepoints in a cohort of HR siblings and LR controls subsequently followed up at 24 and 36 months to answer the following questions: Do scores on the AOSI differ between HR siblings and LR controls at 7 and 14 months? Do scores on the AOSI differ between those HR siblings who go on to have a diagnosis of ASD from those HR siblings who do not? Within the HR group are there associations between AOSI scores at the 7 and 14 month timepoint and later scores on the Autism Diagnostic Observational Schedule (ADOS-G; Lord et al., 1999) at 24 and 36 months?","Ethical approval for the BASIS study was obtained from NHS NRES London REC (08/H0718/76). One or both parents gave informed, written consent for their child to participate.","One hundred and four children (54 HR, 50 LR) were recruited as part of the British Autism Study of Infant Siblings (BASIS; www.basisnetwork.org). They were seen on four visits when aged 6–10 months (mean = 7.35, SD = 1.21; hereafter 7 m), 11–18 months (mean = 13.79, SD = 1.46; hereafter 14 m), and then around their 2nd birthday (mean = 23.9 months, SD = .95; hereafter 24 m), and third birthday (mean = 37.93 months, SD = 3.02; hereafter 36 m). Each HR infant had an older sibling (in 4 cases, a half-sibling) with a community clinical ASD diagnosis (hereafter, proband), confirmed on the basis of information in the Development and Wellbeing Assessment (DAWBA; Goodman, Ford and Richards 2000) and the Social Communication Questionnaire (SCQ; Rutter, Bailey and Lord 2003) by expert clinicians on our team (TC, PB). Most probands met ASD criteria on both measures (n = 44). While a small number scored below threshold on the SCQ (n = 4), no exclusions were made due to meeting the DAWBA threshold and expert opinion. For two probands, data were only available on one measure, and for four probands, neither measure was available (aside from parent-confirmed local clinical diagnosis). Parent-reported family medical histories were examined for significant conditions in the proband or extended family members (e.g., Fragile X syndrome, tuberous sclerosis) with no such exclusions deemed necessary. LR controls were full-term infants (gestational ages 37 to 42 weeks; 3 born 32 to 36 weeks) recruited from a volunteer database at the Birkbeck Centre for Brain and Cognitive Development. Medical history review confirmed lack of ASD within first-degree relatives. All LR infants had at least one older sibling (in three cases, only half-siblings). The SCQ was used to confirm absence of ASD in these older siblings, with no child scoring above instrument cut-off (≥15; n = 1 missing data). The Autism Observational Scale for Infants (AOSI) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The Autism Observation Scale for Infants (AOSI; Bryson et al., 2007; revised version used in this study; Brian et al., 2008) is an experimenter-led, semi-structured observational assessment, developed to study the nature and emergence of ASD-related behavioural markers in infancy (6–18 months). A standard set of objects and toys are used across five activities – each with a specified series of presses for a particular behaviour – and two periods of free-play. Responses to presses and observations made throughout the assessment are used to code nineteen items (see Brian et al., 2008; Bryson et al., 2007; for full description of presses, items and coding; items listed in the Appendix). Each item is coded on a scale from 0 to 2 or 0 to 3. A rating of 0 denotes typical behaviour and higher scores denote increasing atypicality. In the current study the 19 item version of the AOSI reported by Brian et al. (2008) was used. The AOSI yields a Total Score (sum of all codes; max score 44). The AOSI is administered by a trained examiner who sits at a table opposite the infant who is held on the parent's lap. AOSIs were administered by research-reliable research staff and the majority of administrations were double-coded by the examiner and an observer. Agreement between the two coders was excellent at both 7 months (n = 92, intraclass correlation coefficient = .97) and 14 months (n = 85, intraclass correlation coefficient = .95). When codes differed between researchers, they discussed and agreed on a consensus code, where no observer codes were available the examiner's code was used. Developmental assessments and outcome groups ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ All participants were assessed at all visits on the Mullen Scales of Early Learning (MSEL; Mullen, 1995)—a measure of developmental abilities yielding an Early Learning Composite (ELC) standardised score (mean = 100; SD = 15). In order to explore the association between verbal and nonverbal developmental abilities and AOSI scores we calculated mean T-scores from the two Verbal (Receptive Language, Expressive Language) and two Nonverbal (Fine Motor, Visual Reception) Mullen subscales. Of the 54 HR infants recruited, 53 were retained to the 36 m visit when comprehensive diagnostic assessment was undertaken. At 36 m parents of HR siblings completed the Autism Diagnostic Interview—Revised (ADI-R; Lord, Rutter, & Le Couteur, 1994) and the SCQ (Rutter et al., 2003), and both HR and LR toddlers were assessed with the Autism Diagnostic Observation Schedule (ADOS-G: Lord et al., 1999; 24 m module 1N = 50, module 2N = 2; 36 m module 2N = 50 toddlers, module 1N = 3) and the revised Social Affect and Repetitive and Restrictive Behaviours subtotal and calibrated severity scores computed (Gotham et al., 2007). Assessors were not blind to risk-group status. Assessments were conducted by or under the close supervision of clinical researchers (i.e., psychologists, speech therapists) with demonstrated research-level reliability. Different teams of researchers saw participants at the first two visits and the second two visits. Those assessing developmental outcomes were blind to infants’ performance on the AOSI. In determining diagnostic outcome status, four clinical researchers (KH, SC, GP, TC) reviewed information across the 24 m (including an ADOS-G assessment for the HR siblings) and 36 m visit (including both ADOS-G and ADI-R administration for the HR siblings). Seventeen toddlers (11 boys, 6 girls) met ICD-10 (World Health Organisation, 1993) criteria for an ASD (combining ICD-10 childhood autism and pervasive developmental disorder (hereafter, HR-ASD subgroup)). The remaining 36 toddlers (10 boys, 26 girls) did not meet diagnostic criteria for ASD (HR-No ASD). For the LR control group, in the absence of a full developmental history (no ADI-R was administered) no formal clinical diagnoses were assigned but none had a community clinical ASD diagnosis at 36 months. It is worth noting that the recurrence rate reported in the current study (32.1%) is higher than that reported in the large consortium paper published by Ozonoff and colleagues (18.7%; Ozonoff et al., 2011). This is likely to reflect the modest size at-risk sample in the current study (N = 53). Whilst recurrence rates approaching 30% have been found in other moderate size samples (Landa & Garrett-Mayer, 2006; Paul et al., 2010) these rates are sample specific and will likely not be generalizable as findings from larger samples where autism recurrence rates converge between 10% and 20% (Constantino et al., 2010; Sandin et al., 2014). Similar procedures combining all information from standard diagnostic measures and clinical observation and arriving at a ‘clinical best estimate’ ICD-10 diagnosis was used in the present study in line with other familial at-risk studies and was conducted by an experienced group of clinical researchers. HR siblings and LR controls did not significantly differ from each other in age at any visit (t-tests, all ps > .34), nor did the HR outcome groups or LR controls differ from one another in age at any visit (all ps > .20; see Table 1 for descriptive statistics on all measures). Whilst the ELC scores of the HR siblings were in the average range, at each visit their scores were lower than those for the LR controls (t-tests, all ps < .01). The HR-ASD group had lower ELC scores than the LR group at all four visits (all ps < .05) and lower ELC scores compared to the HR-No ASD group at the 14 m and 36 m visits (both p < .05) but not the 7 m and 24 m visits. Analysis ~~~~~~~~ Due to the skewed distribution of the AOSI Total Score a square root transformation was applied and the transformed data met assumptions of normality, with the exception of the LR group at 14 m of age (p < .05). HR versus LR scores were compared using ANOVA and HR- ASD, HR-No ASD and LR scores were compared using a one-way ANOVA and post-hoc Tukey HSD tests. Cohen's d effect sizes are reported (Cohen, 1988). Following this in order to control for verbal and nonverbal developmental level the Mullen mean Verbal and Nonverbal T-scores were covaried and post-hoc least significant difference (LSD) tests conducted. For individual AOSI items, HR versus LR scores were compared using Mann–Whitney tests and HR-ASD, HR-No ASD and LR scores were compared using Kruskal–Wallace tests and significant differences followed-up using post-hoc Mann–Whitney tests. Given the larger number of items but also allowing for the exploratory nature of the analysis, a moderately conservative significance level of p < .01 was used. Correlations between AOSI scores and the total scores on the ADOS (Social Affect and Restricted and Repetitive Behaviour subtotals combined) at 24 months and 36 months in the HR group only were examined using Pearson's product moment correlations. HR versus LR group differences ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As shown in Table 2, at 7 m the HR group had a higher AOSI score than the LR group F(1, 102) = 4.10, p < .045 (d = .39). At 14 m the HR group had a higher AOSI score than the LR group but the difference was not statistically significant F(1, 99) = 3.36, p = .07 (d = .36). However, when Mullen Verbal and Nonverbal T-score at each age point was covaried, the difference between the HR and LR groups was no longer significant (7 m: F(3, 99) = .66, p = .37; 14 m: F(3, 96) = 2.24, p = .14). At 7 m but not 14 m the covariate effect for Mullen Verbal T-score was significant (F(3, 99) = 4.48, p < .05). HR outcome group differences ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Comparing AOSI scores for the HR-ASD and HR-No ASD outcome groups and the LR group, at 7 m the one-way ANOVA just missed significance F(2, 98) = 2.94, p = .058 (see Table 2 and Fig. 1). At 14 m, the three outcome group comparison was significant F(2, 96) = 4.43, p = .014. Post-hoc Tukey tests showed that the HR-ASD group scored higher than the LR group (p = .01; d = .81). The HR-ASD group also had a marginally but not significantly higher AOSI score than the HR-No ASD group (p = .07; d = .65). These analyses were repeated covarying for Mullen Verbal and Nonverbal T-score at each age. The HR outcome groups and LR group did not differ from each other at 7 m of age, F(4, 95) = .61, p = .47 but did differ from each other at 14 m of age, F(4, 93) = 3.85, p = .03. Post-hoc LSD tests showed that at 14 m the HR-ASD group scored higher than both the LR group (p < .01) and the HR-No ASD group (p = .03) and that the HR-No ASD group and LR group did not differ from each other (p = .53). At 7 m but not 14 m the covariate effect for Mullen Verbal T-score was significant (F(4, 95) = 4.31, p < .05). Individual AOSI items ~~~~~~~~~~~~~~~~~~~~~ In terms of individual items, at 7 m the HR and LR groups did not differ on any items but the HR group scored higher than the LR controls on one item at 14 m (orientation to name; U = 931.00, p < .01). In terms of HR-ASD vs. HR-No ASD outcome groups vs. LR group differences, at 7 m there were significant differences for the following items: visual tracking (χ2 = 9.99, p < .01; HR-No ASD > LR: U = 634.50, p < .01) and Social Referencing (χ2 = 10.78, p < .01; HR-ASD > LR: U = 206.50, p < .01). At 14 m there were significant differences for the following items: orientation to name (χ2 = 11.50, p < .01; HR-No ASD > LR: U = 596.00, p < .01; HR-ASD > LR: U = 282.50, p < .01), Engagement of Attention (χ2 = 9.75, p < .01; no significant post-hoc tests) and Social Referencing (χ2 = 9.75, p < .01; no significant post-hoc tests). Associations between AOSI scores at 7m and 14m and ADOS scores at 24m and 36m ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To examine associations over time between behavioural atypicality as measured by the AOSI in infancy and the ADOS in toddlerhood correlations were examined in the HR group. Since the AOSI measures attentional disengagement and atypical motor behaviours, as well early social communication behaviours, AOSI scores were compared to the ADOS ‘total score’. AOSI score at 7 m was not associated with ADOS score at either 24 m (r = .23, p = .10) or 36 m (r = .16, p = .27). However, AOSI score at 14 m was significantly associated both with ADOS score at 24 m (r = .30, p = .03) and at 36 m (r = .38, p = .005).","Consistent with previous reports (Brian et al., 2008, 2013; Georgiades et al., 2013) behavioural atypicalities as measured by the AOSI differentiated the HR and LR groups in the first year (7 months) and (marginally) early in the second year (14 months) of life. At 14 months (but not 7 months) these behavioural markers also discriminated between those HR infants who went on to an ASD diagnosis and LR controls and marginally between HR who did and did not go on to an ASD diagnosis. The HR versus LR group comparisons were attenuated when verbal and nonverbal developmental level was controlled but the HR-ASD outcome group differences remained significant. The HR-No ASD outcome group scored intermediate between the HR-ASD outcome group and the LR control group. Several aspects of this pattern of differences are worthy of comment. First, the HR versus LR group comparisons are modest only in effect size (.39 at 7 months and .36 at 14 months). However, in those HR infants who went on to meet diagnostic criteria for ASD at 36 months by the 14 month timepoint the effect sizes were large (.81 compared to the LR controls; .65 compared to HR infants who did not go on to meet criteria for ASD at 36 months. The current sample is modest in size but, notwithstanding this, significant sub-group differences were still found because of the relatively large behavioural differences on this observational measure of early autistic atypicality. Although the AOSI combines developmental abilities (e.g. imitation) and the presence of atypical or unusual behaviours (e.g. atypical motor and sensory behaviours) the effects of developmental level as measured by the Mullen verbal and nonverbal subscales did not predominate, with only verbal ability at 7 month being significantly associated with the AOSI total score. Finally, the HR-No ASD group performed intermediate between the HR-ASD group and LR group, both on overall total scores (Table 2) and on individual items (see below). The boxplots in Fig. 1 show that the interquartile range of this group span across those of both the HR-ASD and LR groups, suggesting that characteristics that might be considered as aspect of the early ‘broader autism phenotype’ (BAP) are seen in some but not all of the infants at familial high-risk who do not go onto an ASD presentation at 36 months of age (see also Georgiades et al., 2013). Individual items ~~~~~~~~~~~~~~~~ Although the initial report on the AOSI found that scores were predictive of a diagnosis at 12 months but not 6 months (Zwaigenbaum et al., 2005), subsequent studies have found that some behaviours at 6 months differentiate HR siblings and LR controls (e.g. eye contact, social referencing) and also those HR siblings who go on to have an ASD diagnosis from both LR controls (e.g. reactivity, disengagement of attention) and HR siblings who do not go on to have an ASD (e.g. atypical motor control) (Brian et al., 2013). In terms of individual AOSI items we found no HR vs. LR group differences at 7 m and HR siblings only scored higher than LR controls on one item at 14 m (orientation to name). We found that social referencing at 7 m and orientation to name also at 14 m differentiated the HR siblings who went on to have an ASD from LR controls. Furthermore, in line with the BAP concept (Georgiades et al., 2013; Ozonoff et al., 2014), HR siblings who did not go on to have an ASD also showed higher levels of atypical behaviour at 7 m (visual tracking) and 14 m (orientation to name), compared to LR controls. However, these item-level findings should be considered exploratory given that a strict correction for multiple testing was not undertaken and require replication in other samples. In contrast to previous reports (Brian et al., 2008, 2013; Georgiades et al., 2013), we conservatively adjusted for the differences in verbal and non-verbal developmental ability between the groups. It is increasingly apparent, consistent with the phenotype of ASD, that both developmental and language delays are part of the BAP at a group level (Messinger et al., 2013; Ozonoff et al., 2014) and, whilst covarying for these differences one might be taking out some of the variance of interest, early atypical behaviours still discriminated between the infants who went on to have ASD and those who did not. Associations between the AOSI and later ADOS scores ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Within the HR sibling group we examined the association between early behavioural atypicalities as measured by the AOSI and later early symptoms of autism as measured by the ADOS. AOSI scores at 7 months were not associated with later ADOS scores but AOSI scores at 14 months were moderately associated with ADOS scores at 24 and 36 months. Although the two instruments are not identical there is a considerable overlap in the concepts, behaviours and scoring systems, and this suggests a moderate degree of continuity of autistic-like behavioural atypicality from the beginning of the second year of life into the toddler years. However, as has been found with many experimental measures (Jones et al., 2014), with a few exceptions, this continuity is not apparent from as early as 6 to 8 months of age. As such, it appears that the AOSI is successfully capturing very early emerging autistic behaviours. Larger samples will be required in order to test both how predictive such early behavioural markers are at an individual, as opposed to a group, level and in order to trace the longitudinal trajectory of the emergence of such behaviours over this early time course (Landa et al., 2012, 2013). This work is much needed as increasingly in some communities concerns about possible autism are raised about some children in the second year of life, in particular younger siblings of a child with an ASD given the now well-established recurrence rate of between 10% and 20% (Constantino et al., 2010; Ozonoff et al., 2011; Sandin et al., 2014). Limitations ~~~~~~~~~~~ We consider these findings preliminary due to the modest sample size and they will require confirmation in larger and other independent samples. However, it is the first independent report on the AOSI and replicates some of the findings from the instrument's originators (Brian et al., 2008, 2013; Zwaigenbaum et al., 2005). We also note several limitations in the design, including non-blind assessment (to risk status) at both the infancy and toddler assessments, although the team conducting the toddler visits were blind to infant AOSI scores.","The current findings confirm the emerging picture that early behavioural atypicalities in emergent ASD include both social and non-social behaviours (see Jones et al., 2014; for a review). Some of these atypicalities are found only in HR siblings who go on to have ASD but others are also found in HR siblings who do not, supporting the notion of an early broader autism phenotype (Georgiades et al., 2013; Ozonoff et al., 2014). Understanding the interplay between different neurodevelopmental domains across the first years of life and the influences on these will be important both to understand the developmental mechanisms that lead to the ASD behavioural phenotype and to inform approaches to developing early interventions (Green et al., 2013, in press; Wallace & Rogers, 2010)."],["This paper discusses the ecological case for epistemic innocence: does biased cognition have evolutionary benefits, and if so, does that exculpate human reasoners from irrationality? Proponents of 'ecological rationality' have challenged the bleak view of human reasoning emerging from research on biases and fallacies. If we approach the human mind as an adaptive toolbox, tailored to the structure of the environment, many alleged biases and fallacies turn out to be artefacts of narrow norms and artificial set-ups. However, we argue that putative demonstrations of ecological rationality involve subtle locus shifts in attributions of rationality, conflating the adaptive rationale of heuristics with our own epistemic credentials. By contrast, other cases also involve an ecological reframing of human reason, but do not involve such problematic locus shifts. We discuss the difference between these cases, bringing clarity to the rationality debate. --------------------------------------------------------------------------------","Like any other biological organ, the human brain is a product of evolution by natural selection: The brain secretes thought as the liver secretes bile, wrote the 18th century French physiologist Pierre Cabanis in his Des Rapports du Physique et du Morale de l’Homme. But does the brain’s biological provenance suggest it must be successful at producing true beliefs and making rational judgments? Several generations of psychologists have suggested otherwise, documenting the myriad flaws and foibles of human reasoning (Gilovich, Griffin, & Kahneman, 2002; Kahneman, 2011; Kahneman, Slovic, & Tversky, 1982). Especially in popular summaries of this research, humans come off rather badly – prone to all sorts of biases and fallacies, woefully inadequate at dealing with probability and uncertainty, and inclined to persist in making errors of social judgment, even after these have been clearly spelled out (Ariely, 2009; Piattelli-Palmarini, 1996; Shermer, 2011; Singer & Benassi, 1981; Sutherland, 2007). The psychologist John Kihlstrom (2004) has called this the ‘People are Stupid’ school of psychology (PASSP). Recently, however, a new wave of research, inspired by evolutionary ideas, has posed a challenge to this bleak picture of human reason. This school of thought heralds an ecological conception of rationality, aligning itself with research in evolutionary psychology. If we approach the human mind as a collection of cognitive heuristics, tailored to the structure of the ecological environment, many alleged biases and fallacies arguably emerge as artefacts of narrow norms and artificial set-ups. After introducing this ‘ecological rationality’ research program, we identify a ‘locus shift’ in some attributions of rationality. In these cases, individual reasoners are being praised as rational simply because the heuristics employed in their reasoning show adaptive design (i.e. are “rational” from an evolutionary perspective). In other words, the adaptive rationale of cognition (evolutionary locus) is conflated with our own epistemic credentials (personal locus). We then evaluate whether the normative categories of (ir)rationality can still be applied at the evolutionary level of analysis, and conclude that this facet of the ‘ecological defence’ is confusing: (a) human reasoners do not deserve the “credit” for the adaptive designs they have been equipped with, and (b) evolution cannot exculpate human irrationality. Finally, we discuss cases in which the program of ecological rationality has succeeded in rehabilitating human reason, showing that earlier charges of irrationality had been premature. These cases also involve a form of ecological reframing, but they do not involve adaptive locus shifts. The aim of this paper is to provide some much needed clarification to the debate about rationality. In particular, by clarifying which ‘defences’ of putative irrationality are legitimate and which illegitimate, we hope to provide a robust conceptual framework within which to situate and evaluate future empirical evidence. The debate about rationality is not merely of academic interest: irrationality is a topic of critical personal and social import. Biased beliefs about the self and the future may promote individually harmful behaviours like smoking, unsafe sex and overspending, as well as potentially precipitating global catastrophes such as sectarian violence, world wars, exploding financial bubbles and environmental disasters (Johnson & Fowler, 2011; Sharot, 2011). Given these wide-ranging outcomes, getting clear about the nature and extent of human irrationality is a critical philosophical and psychological project. Recasting rationality in the environment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the heuristics and biases program instigated by Daniel Kahneman and Amos Tversky in the 80s (Kahneman, 2011; Kahneman et al., 1982), heuristics are viewed as mental short-cuts leading to imperfect and often outright irrational inferences. Though Kahneman and Tversky pointed out that even fallible heuristics often lead to successful inferences, over time this sense of balance was lost, and a negative slant began to dominate the research program they founded (Krueger & Funder, 2004).1 The program of ecological rationality, by contrast (Fawcett et al., 2014; Gigerenzer, Hertwig, & Pachur, 2011; Gigerenzer & Todd, 1999; Hertwig & Hoffrage, 2013; Todd & Gigerenzer, 2012), aims to effect a gestalt switch, urging us to rethink the norms of what counts as rational. The canons of ‘classical’ rationality, to which Kahneman & Tversky’s subjects were held accountable, consist of a general, formal, content-free framework for valid reasoning (e.g. modus ponens). Advocates of ‘ecological rationality’, by contrast, uphold a radically different view of heuristic reasoning. Heuristics, they argue, provide us with an ‘adaptive toolbox’ (Gigerenzer, 2008; Gigerenzer & Todd, 1999), each suited to a particular set of challenges endemic to a particular environment. In contrast with unbounded models of classical rationality, which typically assume unlimited resources both with regard to information gathering and computational processing, ecological rationality is ‘fast and frugal’. According to Gigerenzer and Todd (1999, p. 13), a heuristic is ecologically rational ‘to the degree that it is adapted to the structure of an environment’. Heuristics are quick and computationally cheap, requiring few and simple computational steps, and operating on a limited input domain. For example, when confronted with two objects, only one of which is recognized, people infer that the familiar one will be more important or have a greater value. Although this so-called ‘recognition heuristic’ – on which more later – leads us astray in artificial set-ups, it appears to be a surprisingly accurate guide to real world problems (Goldstein & Gigerenzer, 2002). Proponents of ecological rationality claim that many apparent demonstrations of irrationality in the heuristics and biases program are artefacts of inappropriate standards and narrow norms, such as coherence criteria and the axioms of probability theory. Indeed, many experiments are expressly designed to mislead or distract participants, placing them in situations where their usually successful strategies and heuristics lead them astray. If one canvasses the same phenomena in a broader framework, however, taking into account the structure of real environments, computational limitations and various trade-offs, the human mind emerges as more rational than many psychologists suppose. Two strands ~~~~~~~~~~~ We detect two major strands in the re-appreciation of (apparent) human irrationality as ecologically rational. On the one hand, as the ecological rationalist points out, traditional cognitive psychologists have too often delighted in tripping up their subjects with artificial set-ups that truncate the complexity of real life and human intelligence (first strand). On the other hand, ecological rationalists have argued that heuristics are not designed to perform well in artificial (non-ecologically valid) environments. If putatively ‘irrational’ subjects were tested in an ecologically relevant context, the kind to which their evolved heuristics are attuned, they would behave rationally (second strand). In both strands, ecological rationality focuses our attention on the match between our reasoning and the real world. The sensible argument behind this is that we should evaluate heuristics in their ‘proper’ context. You should not use a screwdriver as a crowbar and be surprised that it breaks.2 Whereas the first strand focuses on artificial set-ups and uncharitable experimenters, the second strand invokes evolutionary rationales. In the latter case, advocates of ecological rationality point out that heuristics are crafted by evolution to deal with particular adaptive problems in an ancestral environment. Therefore, they have a “proper domain” of application (Sperber, 1996) – what Millikan (1984) termed “Normal conditions” – outside of which we should not expect them to perform well. Blaming them for “malfunction” in unnatural contexts, without proper appreciation of their design features, is like blaming the screwdriver. Nevertheless, as we will argue, while the former approach is (often) warranted, the latter involves problematic locus shifts. Evolutionary psychology ~~~~~~~~~~~~~~~~~~~~~~~ This second evolutionary strand runs through much of the literature on ecological rationality, (Gigerenzer & Selten, 2002). And indeed, the program connects with work in the field of evolutionary psychology on the adaptive value of cognitive bias (Aktipis & Kurzban, 2004; Cosmides & Tooby, 1994; Haselton & Nettle, 2006; Haselton, Nettle, & Andrews, 2005; Mercier & Sperber, 2011; Nesse, 2001), the main thrust of which is also to show the ecological validity of human reason. As Haselton & Nettle write: Some cognitive illusions disappear or greatly attenuate when the task is presented in an ecologically valid format (Cosmides & Tooby, 1996; Gigerenzer & Hoffrage, 1995). Ecological validity, a long-standing but undertheorized term in psychology, may in effect be equated to the task format approximating some task that humans have performed recurrently over evolutionary time. Apparent biases and illusions are recast in an evolutionary light, as serving some adaptive rationale. From an adaptationist point of view, human reason is not the botched and bias-riddled device many psychologists take it to be, but emerges as an extremely effective, highly specialized set of adaptive tools for dealing with recurring problems in the Environment of Evolutionary Adaptedness (EEA). In an overview of the research program of ecological rationality, Gerd Gigerenzer wrote that ‘what appears to be a fallacy can often also be seen as adaptive behaviour, if one is willing to rethink the norms’ (Gigerenzer, 2008, p. 13). In a similar vein, Haselton and Nettle (2006, p. 59) have argued that ‘bias in cognition is no longer a shortcoming in rational behaviour, but an adaptation of behaviour to a complex, uncertain world’. Elsewhere, Haselton and colleagues wrote ‘an adaptationist perspective suggests that the mind is remarkably well designed for important problems of survival and reproduction, and not fundamentally irrational.’ These assessments seem to provide solace for the pessimistic view of human reason voiced by many psychologists. But can evolution really get us off the hook? Minimal conditions for rationality ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In order to assess what exactly ecological/adaptive rationality amounts to, we need a serviceable conception of rationality. While attempting to formulate a full blown definition seems doomed to fail, given the loose and varied use of the term, some constraints on the meaning of ‘rationality’ will guide us through this treacherous terrain. In this paper, we will be mainly concerned with epistemic rationality (belief formation that is accurate and truth-tracking) and not so much with instrumental rationality (actions that maximize the probability of success, according to some standard). According to De Sousa (2009, p. 289), rationality presupposes intentionality. An inanimate object or mere mechanism cannot be rational. It would be absurd to claim that a rock falling to the ground is rationally obeying the laws of gravity, much as it would be strange to call a calculator rational because it yields correct answers. Likewise, irrationality is a concept that can be meaningfully applied only to intentional creatures – entities capable of full-blown rationality (de Sousa, 2007). Irrationality is a term of epistemic disapprobation, used when someone fails to live up to some expected normative standard. As ought implies can, it makes no sense to blame a mindless process: It makes sense to criticize a person, but not a cell, for having made a mistake in computation, or with having failed to foresee what should have been foreseen, or with having acted on reasons that fell short of the best set of reasons. In an extended, metaphorical sense, however, standards of rationality can be fruitfully applied to the domain of functionality, including biological adaptations and artefacts, even when no intentionality is involved. Although natural selection is a mindless process, it can produce an appearance of intentionality by crafting design solutions to recurring adaptive problems (Williams, 1966/1996). Evolution by natural selection gives rise to what Dennett coined ‘free-floating rationales’ (Dennett, 1983, p. 351). In this extended sense of rationality, the relevant measure of success is equated with reproductive success or inclusive fitness (Over, 2000), and standard frameworks of rationality such as game theory (Maynard Smith, 1974) or expected utility theory (see Section 3.2) can be applied (Kacelnik, 2006). Because of the absence of intentionality, however, not all dimensions of full-blown rationality can be transferred to evolution. Whereas individual reasoners are capable of foresight, evolution cannot look into the future. For example, selection can easily get stuck on local peaks in an adaptive landscape, unable to reach nearby higher summits. If evolution had the capacity of foresight, it would backtrack and make a detour, as a human reasoner would. Evolution, therefore, provides a mere ‘simulacrum of rationality’ (de Sousa, 2009, p. 298). Locus shifts ~~~~~~~~~~~~ Given that rationality (in the literal sense) can only be attributed to person-level decision making (see above), we question whether the ecological reframing succeeds in dispelling the charge of irrationality. In order to get clear about the notion of ‘locus shifts’, we first discuss an example from evolutionary psychology in which a similar calculus of rationality may be implemented either at the individual level or on an evolutionary scale: error management. This will allow us to pull apart the two loci of rationality, while holding the measure of success (e.g. protection from bodily harm) constant. Expected utility theory is one of the hallmarks of classical rationality: it describes how to maximize utility under uncertainty, as a function of the probability of different outcomes and their associated costs and benefits. Normally, the model is used to describe individual choices, for example betting preferences. In recent years, however, the logic of expected utility theory has been extended to evolutionary adaptations, under the guise of error management theory (Haselton & Buss, 2000; Haselton & Nettle, 2006; Nesse, 2001).3 In a range of different situations, organisms need to process ambiguous stimuli that may or may not signal adaptively relevant situations, such as the presence of a predator or pathogen. If the costs associated with the respective errors, measured in terms of (inclusive) reproductive fitness, are asymmetric, then expected utility theory can be applied to predict optimal behaviour. If the costs associated with one type of error are larger than those associated with the opposite kind, then a strategy biased in favour of making the less costly error may pay off in the long run, even if that increases the absolute numbers of errors (Arkes, 1991). As Haselton et al. put it: ‘it is better to make more errors overall as long as they are of the relatively cheap kind’ (Haselton et al., 2005, p. 731). Error management is displayed in a range of biological adaptations that are clearly beyond our voluntary control, such as inflammation, disgust, fever and coughing. In all of these cases, the locus of rationality is clearly evolution by natural selection, not the individual subject. Things get more interesting when considering applications of error management theory to human cognition. In such cases, the locus of rationality may reside either at the individual level or at the evolutionary level. In most discussions of error management, human behavioural biases are interpreted in terms of biased belief, with evolution as the rational bookkeeper. For example, people buy into various superstitions because they evolved pattern detection modules that err on the side of caution. In order not to miss out on any important causal patterns in the world, evolution has ‘decided’ to run the risk of occasional false positives, which are mostly relatively harmless. In a commentary on error management theory, however, McKay and Efferson (2010) argued that human behavioural bias does not automatically entail biased belief. In principle, the same behavioural bias may be arrived at through accurate belief formation combined with judicious action policies, based on expected utility theory. If I react to a rustle in the leaves as if there is a predator lurking in the bushes, I do not need a strong belief that a predator is there. Likewise, when approaching a girl that I like, I do not need to be certain that she fancies me. I may decide to act upon an admittedly remote possibility, in light of the high opportunity costs associated with false negatives (ending up as lunch, or missing a romantic – read reproductive – opportunity). As McKay & Efferson noted, ‘there are infinitely many ways to accomplish the required behavioural change … inferences about cognition can be radically underdetermined when one observes an interesting behavioural bias.’ (see also Marshall, Trimmer, Houston, & McNamara, 2013; McKay & Efferson, 2010, p. 312). The logic underlying the behavioural bias is the same, but our normative appraisal of human reasoners depends on where the locus of rationality resides. If the behavioural bias is the outcome of a prudential action policy, without the involvement of biased belief, we should not accuse reasoners of irrationality. For example, people fasten their seatbelts before starting the car not because they strongly believe they will crash into other cars, but just because they would rather be safe than sorry: the cost of a false negative (being flung through the windscreen) is orders of magnitude larger than that of a false positive (the energy expended in fastening one’s seatbelt and the mild discomfort of being strapped in it), large enough to offset the small probability of a serious car accident. This is a case of full-blown (intentional) rationality, situated at the level of individual decision making. An alien scientist interpreting this prudential human behaviour as a symptom of irrational paranoia would be very uncharitable indeed. On the other hand, if a behavioural bias stems from some attendant belief distortion, the charge of irrationality may be apposite. For example, if an unattractive man approaches every woman at the bar, believing that he is irresistible to the opposite sex, we may be justified in calling him irrational. If biased behaviour translates to bias in belief formation, however, with evolution keeping the books of fitness costs and benefits, we think our assessment should be different (for a discussion, see Galperin & Haselton, 2012). From the adaptationist point of view, the ecological ‘problem’ of exploiting causal regularities in an organism’s environment can be solved by different means: a fixed and genetically encoded action sequence, the capacity to form conditional reflexes and make associations, a belief formation system with a pre- programmed belief bias, or the genuine capacity for causal thought and expected utility reasoning (Dennett, 1996; Lorenz, 1941/2009). Only in the latter case is rationality displayed at the locus of the individual decision-maker. For example, many people engage in superstitious rituals because they strongly believe in the causal connection. If error management theorists are right, there is indeed method in the madness of such superstition, since the fitness cost of overlooking causal links outweighs the fitness cost of detecting non-existent causal relationships. In such cases, however, the locus of ‘rationality’ pertains to the adaptive design of our reasoning faculties. Although it is tempting to collapse both levels of rationality, as though the adaptive rationale of our superstitious behaviour redounds to the human reasoner, the ‘method’ of natural selection does not obviate – indeed it requires – the ‘madness’ of the individual. Accordingly, we defend the following claims: We should not credit human reasoners simply because the adaptive design of their reasoning faculties conforms to some standard principles of rationality.4 We do not congratulate people for their sophisticated immune systems or ingenious thermoregulation, nor do we credit cicadas for having discovered prime numbers,5 or spiders for the beautiful geometry of their webs. There is neither cognitive transparency – the organism has no clue why it behaves as it does – nor intentional action. According to Dennett, If we discover that an animal is too simple-minded to harbour an adaptive rationale, we do not discard the rationale but are simply forced to ‘pass the rationale from the individual to the evolving genotype’ (Dennett, 1983, p. 351). Conversely, if we have good grounds for calling someone’s beliefs irrational (see further), then merely pointing out the larger evolutionary rationale of the faculties generating those beliefs will not get him off the hook. Having biologically adaptive cognition is not an alibi where the charge of irrationality is concerned. If superstitious belief is adaptive, because it motivates ‘better safe than sorry’ behaviours, that does not make superstitious people rational. For example, while people who believe that black cats are harbingers of bad luck may be heeding to the ‘wisdom’ of time-honoured cognitive adaptations, that does not make their belief rational or truth-tracking. Statistical rationales In the classical heuristics-and-biases framework, heuristics are perceived as imperfect shortcuts in an uncertain world, strategies that execute a trade-off between cost and accuracy. Gigerenzer and his colleagues, however, have documented several cases where this trade-off can be escaped, and ‘less information and computation lead to more accurate judgments’ (Gigerenzer & Sturm, 2012, p. 261). The showpiece of these ‘less is more’ effects is the recognition heuristic, where people take advantage of their own ignorance to make intelligent inferences. In judging which out of two cities has a larger population, for example, people use name recognition as a cue for population level. When the task consists of German cities, American subjects perform better than their German colleagues, because the latter recognize too many cities to exploit their own ignorance. A measure of ignorance turns out to be profitable in this case (Goldstein & Gigerenzer, 2002), while knowing too much is a burden. Apparently, everyday life presents us with many situations in which the recognition heuristic proves successful. It provides us with a powerful and accurate tool for making inferences about the environment. Fast and frugal heuristics escape a trade-off between accuracy and speed because ‘they make a trade-off on another dimension: that of generality versus specificity’ (Gigerenzer & Todd, 1999, p. 18). In later work, Gigerenzer and Brighton (2009) have explored another dilemma (between bias and variance) for explaining when and why such ‘less-is-more’ effects occur. The upshot of this statistical analysis is that the simplicity of fast and frugal heuristics (few computational steps and few parameters) protects them against overfitting (i.e. mistaking noise for underlying patterns) and hence ensures robustness in the face of environmental change. The success of the recognition heuristic derives from a probabilistic rationale. In order for the trick to work, there needs to be a correlation between the probability that an item is recognized and the value of interest. This ‘recognition validity’ needs to outweigh the ‘knowledge validity’, which is the probability of giving the correct answer when both items are recognized (Goldstein & Gigerenzer, 1999, 2002). In other words, your ignorance must be more valuable than your knowledgeability. In addition, success depends on the degree of uncertainty, the number of alternatives, and the size of the learning sample. Goldstein & Gigerenzer conclude: [M]any scholars, psychologists included, have mistrusted the power of these heuristic principles, and saw in them simple-mindedness and irrationality. This is not our view. The recognition heuristic is not only a reasonable cognitive adaptation because there are situations of limited knowledge in which there is little else one can do. It is also adaptive because there are situations … in which missing information results in more accurate inferences than a considerable amount of knowledge can achieve. Now what should we make of this contrast between simple-mindedness and ‘reasonable cognitive adaptation’? Can we attribute rationality to heuristic-wielding human thinkers when their heuristics yield accurate inferences? Cognitive transparency The answer to this question depends on the criteria of cognitive transparency and intentionality. As pointed out, attribution of (literal) rationality requires intentional deliberation on the part of the actor, not just the execution of hard- wired action patterns or mindless heuristics. (cf. 3.1 Minimal conditions for rationality). To avoid regress of justification, of course, we should not require that the agent has contemplated every step in the chain of belief-formation. Perceptual beliefs may arise automatically, without deliberation and thus without rationality, but more complex beliefs can be formed consciously on the basis of basic perceptual beliefs. On the latter level, attributions of rationality and irrationality can be meaningfully made.6 Without delving into empirical details, we can be confident that many people are ignorant of the statistical principles underlying the success of the recognition heuristic. This is hardly surprising if the heuristic is a ‘cognitive adaptation’, as Goldstein and Gigerenzer (2002) suggest. Many cognitive modules in our brain perform their proper function without us having the slightest awareness of their operation, let alone their design features and the reasons for their success. If the recognition heuristic is an evolved adaptation, activated under appropriate circumstances by a mechanism to which human cognizers have no access, it would be strange to credit people for its success. The locus of rationality, in this case, is at the level of adaptation. Evolution, rather than human cognizers, has exploited recurrent statistical correlations in the environment. However, recent research shows that some people who use the heuristic articulate a justification that captures the ecological validity (Gigerenzer, 2007, pp. 125–126), and they will suspend the use of the heuristic if they have good counter-indications (Pachur, Todd, Gigerenzer, Schooler, & Goldstein, 2011). Even if such reasoners have no full understanding of the statistical principles involved, it seems that the criterion of transparency is satisfied, and it would be uncharitable to deny them rationality at the personal level. In general, however, the evidence points in the other direction Research suggests that people hardly ever make conscious decisions about which heuristic to use but that they quickly and unconsciously tend to adapt heuristics to changing environments, provided there is feedback. In another striking example from the catalogue of fast and frugal heuristics, both humans and dogs use the following rule of thumb to catch a ball in flight: keep your gaze fixed at the ball, and adjust your running speed such that the angle of the ball in your visual field remains constant (Gigerenzer & Todd, 1999). If baseball players are asked to predict where a ball will land, but are not allowed to run towards it, they perform very poorly (Babler & Dannemiller, 1993). However, when asked how they proceed to catch a ball on the fly, people are typically oblivious of using the heuristic: ‘most fielders are blithely unaware of the gaze heuristic, despite its simplicity’ (Gigerenzer, 2007, p. 11). This, it should be emphasised, does not mean that heuristic-wielding humans are irrational, but that the normative categories of (ir)rationality simply do not apply insofar as humans are working on automatic pilot. Interestingly, on some occasions, Gigerenzer & Todd explicitly credit evolution for the design of heuristics: ‘evolution would seize upon informative environmental dependencies such as this one and exploit them with specific heuristics if they would give a decision-making organism an adaptive edge’ (Gigerenzer & Todd, 1999, p. 19). Such statements, however, are in tension with the main thrust of their research programme, which is to dispel accusations of ‘simple-mindedness’ and irrationality by pointing to the larger ecological picture. As we saw in the example of error management, adaptive design is perfectly compatible with dumbness. In this regard, the novel insights gleaned from Gigerenzer et al.’s research program conflict with their aspiration to rehabilitate human reason. Indeed, if the rationale of the recognition heuristic or the gaze heuristic were transparent to those profiting from it, the results of Gigerenzer et al. would not have been so informative. The effectiveness of these heuristics is surprising even to the researchers themselves. For example, another deceptively simple heuristic, ‘take the best’, which takes the first discriminative cue between two items and ignores the other cues, has outperformed weighted statistical decision procedures such as multiple regression. This is how Gigerenzer & Goldstein relate their amazement at their own result: ‘When we first obtained these results, we could not believe them.’ (Gigerenzer & Todd, 1999, p. 89). Because people use their heuristics without cognitive transparency, they are also often insensitive to local conditions where their heuristics lead them astray.7 For example, many advertisers exploit the recognition heuristic by investing in brand recognition rather than product quality, thereby undermining the ecological validity of the correlation between renown and product quality and fooling consumers. In the ‘overnight fame’ experiments by Jacoby, Kelley, Brown, and Jasechko (1989), people who had been presented with unknown names the day before were tricked into believing that those names were famous. This should not come as a surprise. Evolution exploits recurring statistical properties of the environment of evolutionary adaptiveness (EEA), but it is not responsive to any environmental contingency. For evolution it is ‘rational’ to implement the recognition heuristic in an organism as long as overshooting or misfiring has been sufficiently rare or inconsequential in the EEA. Error management on the evolutionary scale abstracts away from local contingencies. Why adaptive ‘rationality’ does not align with personal rationality ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the section on error management, we noted that adaptively ‘rational’ cognition may translate into irrational belief, such as superstition. There are two additional reasons why evolutionary rationality and personal rationality may diverge. The first is the so- called problem of adaptive mismatch. If heuristics are evolved adaptations for managing cost asymmetries in the EEA, there is no guarantee that the same balance sheet still applies to modern environments. Probabilities may have changed over time, as may have the respective costs and benefits. This mismatch hypothesis is one of the main tenets of evolutionary psychology (Pinker, 1997; Tooby & Cosmides, 1994).8 A heuristic may succeed in the EEA but fail in modern environments, or only match some limited aspects of modern environments (e.g. to predict which city will have the largest population). A second reason concerns the issue of who – or rather what – is the ultimate beneficiary of natural selection (cui bono? see Dennett, 1995). As Dawkins (1976) has pointed out, following up on pioneering work by Williams and Hamilton (Hamilton, 1963; Williams, 1966/1996), genes can be seen as the ultimate beneficiaries of evolutionary adaptation. Individual organisms are best seen as ‘vehicles’ for genetic propagation (Dawkins, 1982). Take, for instance, the suicidal sting of a bee. Such behaviour is ‘rational’ from the point of the evolutionary processes that gave rise to such behaviour, because the bee shares 75% of its genes with its sisters. Sacrificing its life to enhance the survival chances of the hive therefore makes sense from a gene-centric perspective. But it certainly does not benefit the individual bee. Evolutionary rationality – which is situated at the level of the replicators – therefore does not always coincide with rationality at the level of the vehicle (Skyrms, 2000; Stanovich & West, 2000; Stanovich & West, 2003).","Let us now have a closer look at two famous examples from the irrationality literature, where advocates of ecological rationality have performed a locus shift: the gambler’s fallacy and the hot hand fallacy. The former refers to the belief of roulette players that numbers which have not turned up yet are more likely to win in the next round, because they are ‘overdue’. In reality, of course, each turn of the roulette wheel is independent of the next. The wheel has no memory, so the probability distribution remains the same at every trial. More generally, the gambler’s fallacy describes the belief that statistical deviations in a given direction will soon be counterbalanced by deviations in the opposite direction. The hot hand fallacy can be seen as the mirror image of the gambler’s fallacy. To commit the hot hand fallacy is to believe that, if some deviation from chance has been observed, further deviations in the same direction are more likely on subsequent trials. The name derives from basketball, where there is a widespread belief among players and fans that players score in streaks (Gilovich, 1983; Gilovich, Vallone, & Tversky, 1985). If one player has already scored well, he has a ‘hot hand’ and is more likely to make further points during the rest of the game. If he makes a bad start, however, there is a jinx on him and he will be less likely to perform well during the rest of the match. Statistical analyses of basketball scoring data have revealed no such clustering (but see Raab, Gula, & Gigerenzer, 2012, for evidence of streakiness in volleyball), but belief in the hot hand persists.9 The fact that each of these forms of inference has been bestowed its proper label in the literature, despite their being one another’s mirror image, is a source of amusement to Gigerenzer and his colleagues. For them, it exemplifies the bad habit of psychologists to label any observed deviation from accepted rationality models as a ‘fallacy’ or ‘bias’, without any understanding of the underlying cognitive mechanisms (Gigerenzer & Brighton, 2009). Are these unflattering labels justified? Steven Pinker has argued that the gambler’s fallacy is not really a fallacy, because in most natural environments, a succession of events is not statistically independent: the probability of an event changes in function of earlier occurrences. Many events work like that. They have a characteristic life history, a changing probability of occurring over time which statisticians call a hazard function. An astute observer should commit the gambler’s fallacy and try to predict the next occurrence of an event from its history so far … Roulette wheels, with their smooth and radially symmetric design, are expressly designed to foil this heuristic, according to Pinker: ‘in any world without casinos, the gambler’s fallacy is rarely a fallacy’ (Pinker, 1997). Indeed, our ability to spot statistical dependencies between events in real-life may often prove very useful. That morale even holds for chance games outside the idealized world of casinos. If I play a dice game at a fair and the dice land on 12 a couple of times in a row, it is not unreasonable to predict that they will do so the next time too – because they are probably rigged. In general, as Bennis et al. have written, ‘casino games [are] exquisitely designed to exploit otherwise adaptive heuristics to the casino’s advantage’ (Bennis, Katsikopoulos, Goldstein, Dieckmann, & Berg, 2009, p. 421). Taleb (2008) has referred to the inappropriate application of pure and simplified models of probability to real life as the ‘ludic fallacy’. Intuitions that lead us astray in the rarefied and idealized world of casinos, may prove very useful in real-life situations. If I am presented with some strange device that churns out black and white balls, and I do not have access to its inner working, I may be forgiven for thinking that there is a statistical dependence between different trials. I may even be perfectly justified in doing so, extrapolating from earlier experience. An experienced gambler in a casino, by contrast, possesses all the requisite evidence he needs to conclude that this is not a machine to which his pattern detection heuristics will apply. If even a crash course in statistics and a careful inspection of the roulette wheel fail to cure him of his habit of thought, then we may, pace Pinker, call his reasoning fallacious. The natural way to construe Pinker’s argument is as a point about adaptation. Given that pure randomness is rare in the EEA (as it most probably still is today in natural environments), it made sense for evolution to endow us with a knack for spotting patterns, with the minor side-effect of making us vulnerable to (fair) dice games and roulette wheels. After enduring five straight days of rainy weather, as Pinker notices, it would have made sense for the optimistic early hominid to expect a sunny spell. As weather fluctuation often follows statistical patterns, the sun can really be ‘overdue’. For evolution, it makes sense to ‘gamble’ on the kind of environments with a specific hazard function, and to disregard the abstractions of casino wheels, which did not feature in ancestral environments. In a similar vein, the hot hand fallacy may arise from a heuristic adapted to environments with positive temporal autocorrelation or ‘clumping’ (Fawcett et al., 2014; Wilke & Barrett, 2009). But what does that say about us? Are we forgiven for using the heuristic anyway, even in the face of countervailing information? Personal-level rationality involves the ability to use contingent information about the problem task in order to understand which strategies are conducive to success. A rational person should reflect upon the deliverances of her intuitions and rules of thumb, and disregard them when she has good reasons to believe that they lead her astray (Over, 2000). Irrationality often stems from what Fiedler and Wänke (2004, 2013) have called ‘meta-cognitive myopia’, or the inability to critically evaluate intuitive heuristics and use them in appropriate settings. In the case of casinos, there is indeed an alternative course of action available. Many experienced gamblers, though still feeling the intuitive pull of the pattern-seeking heuristic, suppress or resist it in practice, because they realize that the wheel has no memory and every turn is independent of all the others. Surely (Dennett, 2013) such players are more ‘rational’ than the ones who succumb to the gambler’s fallacy, in all of the senses we explored: they understand the justification of their beliefs, they intentionally aim for a specific goal, and they will be more successful in attaining it. Indeed, one of the hallmarks of (human) rationality is our ability to reflectively evaluate and cross-validate the output of our cognitive modules, and to overrule their intuitive hunches when the context calls for it. In this regard, whereas Pinker emphasises that the gambler’s fallacy is not a fallacy in a world without casinos, we emphasise that it is nevertheless a fallacy in a world with casinos. As de Sousa points out, the rational ‘strategies’ attributed to natural selection are conceived in the most general terms, on the basis of phylogenetic ‘experiences’ and without precise reference to the circumstances under which they may be implemented. By contrast, individual rational agents face particular problems in possibly unique circumstances (de Sousa, 2007, p. 137). Rationality, in this perspective, consists of using our reasoning abilities appropriately to deal with the situation at hand, not blindly following heuristics of which – with hindsight – we can appreciate the adaptive rationale. Conjunction fallacy? ~~~~~~~~~~~~~~~~~~~~ Does the program of ecological rationality always involve locus shifts, and their attendant problems of transparency and intentionality? No. The second strand of the program of “ecological rationality” (Section 2.2), does not depend on adaptive considerations, and hence does not involve locus shifts. The ‘ecological’ dimension of rationality, in this strand, concerns the real-life contexts of human reasoning, which are typically truncated in artificial lab experiments. To illustrate this difference, we will briefly consider two examples of alleged human bias that have been reinterpreted under the banner of “ecological rationality”, but which do not involve adaptive locus shifts and which satisfy the minimal requirements of personal-level rationality. Consider the following classic test: Linda is 31 years old, single, outspoken, and very bright. She majored in philosophy. As a student, she was deeply concerned with issues of discrimination and social justice, and also participated in anti-nuclear demonstrations. Which is more probable? Linda is a bank teller. Linda is a bank teller and is active in the feminist movement. In their original research on what became known as the ‘Linda problem’, Kahneman & Tversky held human reasoners accountable to the conjunction rule: B cannot be more probable than A, because it is an accepted rule of probability theory that the conjunction of two events can never be more probable than either of its members. Still, the majority of subjects answered B. Gigerenzer, however, pointed out that terms such as ‘and’ and ‘probable’ are polysemous. ‘Probable’ can also mean plausible, sensible, or supported by evidence (Gigerenzer, 1996; Hertwig & Gigerenzer, 1999). Kahneman & Tversky expect subjects to interpret ‘probable’ in the sense of mathematical probability. As a number of researchers have pointed out, however (Adler, 1984; Dulany & Hilton, 1991; Hilton, 1995), this construal violates pragmatic rules of conversational inference, in particular the maxim of relevance (Grice, 1989). According to this maxim, it makes little sense to adopt the notion of mathematical probability in answering the Linda problem, given all the additional (and hence presumably relevant) information presented. If subjects interpret the task as Kahneman & Tversky want them to, the little vignette about Linda is rendered completely irrelevant to the task, and the ‘correct’ answer depends solely on the logical operator ‘AND’ and the mathematical meaning of ‘probable’. Why would the researcher present such an elaborate description of Linda’s character and background, if it were completely irrelevant? In addition, subjects can reasonably interpret the first statement as implying that Linda is not active in the feminist movement. These ambiguities can be ruled out by presenting the task in terms of natural frequencies, or by including another question that makes the description of Linda relevant, so that the (overall) task does not violate the maxim of relevance. After such modifications, as it turns out, subjects reason in accordance with the conjunction rule (Hertwig & Gigerenzer, 1999).10 Given that simple application of the conjunction rule glosses over the conversational subtleties and ambiguities of the situation, a number of critics have argued that Kahneman & Tversky’s interpretation is uncharitable (e.g. Margolis, 1987), and even that they, and not their subjects, have failed to grasp the logic of the Linda problem (Hertwig & Gigerenzer, 1999): ‘human intelligence reaches far beyond narrow logical norms. In fact, the conjunction problems become trivial and devoid of intellectual challenge when people finally realize that they are intended as a content-free logical exercise’ (Gigerenzer, 2008, p. 73). Base rate fallacy? ~~~~~~~~~~~~~~~~~~ Another case of alleged irrationality that (at least partly) disappears when cast in an appropriate real world environment, instead of an artificial and misleading set-up, is the infamous base rate fallacy (Casscells, Schoenberger, & Graboys, 1978; Kahneman et al., 1982, p. 154). For example, when estimating the probability that a patient has contracted a particular disease, physicians are (sometimes) found to ignore the base rate of the disease in question. Information about the reliability of the test appears to be the sole determinant of their judgment, while the statistical prevalence of the disease in the population group is ignored. However, as Gigerenzer noted, real life is more complicated ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Clinicians, however, know that patients are usually not randomly selected—except in large survey studies—but rather ‘select’ themselves by exhibiting symptoms of the disease. In the absence of random sampling, it is unclear what to do with the base rates specified. If patients are not randomly sampled, Gigerenzer notes, physicians cannot be accused of committing the base rate fallacy. Indeed, given that real world environments, as opposed to convenient mathematical idealizations, do not typically proffer random samples, the judgment of the clinicians is perfectly defensible. In a series of experiments conducted by Gigerenzer and et al. (1988), people were found to be capable of factoring in base rates, provided they were ‘convinced’ that the sample was randomly drawn from the population. Inviting subjects to blindly draw patients’ names from an urn proved effective in overruling their prior beliefs, whereas merely pointing out random sampling in the problem description did not. Even though non-random sampling appears to be an intuitive default assumption, these experiments show that people can and will – given sufficient commitment – account for base rates when necessary. Indeed, as Gigerenzer points out, with Tooby and Cosmides (1994), the neglect of base rate information is by no means a systematic defect of human reasoning, as the label ‘base rate fallacy’ suggests. If Bayesian problems are couched in terms of frequencies, as opposed to abstract mathematical probability, it turns out that the ‘fallacy’ largely disappears (but see Kahneman & Tversky, 1996; Sloman, Over, Slovak, & Stibel, 2003). In other words, it is not that most people lack the requisite skills to integrate base rate information (Zhu & Gigerenzer, 2006). Provided that there is sufficiently strong indication of random sampling, and that base rate information is presented in a concrete, accessible way, people’s reasoning approximates Bayesian theory (Cosmides & Tooby, 1996; Gigerenzer & Hoffrage, 1995). Once again, Gigerenzer and colleagues argue that the ‘fallacy’ is elicited by the artificial or experimental set-ups, and evaporates when transferred to real life. It is important to note that whether and to what extent the conjunction and base rate fallacies disappear when problems are presented in ecologically valid formats is still hotly debated. This, however, is a matter of experimental psychology. The point we are making is merely that the approach taken in these cases does not involve a problematic locus shift. Analysis ~~~~~~~~ Wherein lies the difference between the latter cases and the problematic ones we discussed earlier, given that both involve a form of ‘ecological’ reframing? In Gigerenzer’s deflation of the ‘conjunction fallacy’ or the ‘base rate fallacy’, the gestalt switch also involves adopting a broader view, casting human inference-making in a rich and real world environment. In these examples, however, the force of the ecological gestalt switch derives not from any shift to the adaptive rationale of their heuristics, but from the fact that human reasoners had not been given sufficient information to disambiguate the problem at hand and to home in on the interpretation intended by the experiments. In particular, subjects did not know – or were not properly committed to the belief – that they were dealing with an artificial situation in which the normal richness and ecological complexity of problem solving should be ignored. How were Kahneman & Tversky’s subjects supposed to know that their experimenters were interested in a ‘silly’ question that violates an accepted rule of sensible conversation? Why should they have interpreted the problem in the sense of mathematical probability, contra their intuition that this made the whole story about Linda irrelevant? Similarly, why should the physicians have realized they were expected to take a step back from their regular practice, considering an unrealistic case in which a subject is randomly drawn from the population at large. Remember that, when Gigerenzer (1991) committed subjects to the mathematical interpretation, their reasoning conformed to the conjunction rule. The force of the argument from ecological rationality here is not that ‘we can design experiments in which cognitive illusions disappear’ (Kahneman, 2003, p. 711), as Kahneman construes Gigerenzer’s point, but that there was no illusion in the first place.","In order to understand the strengths and frailties of human reason within a naturalized framework, an evolutionary perspective is invaluable. We applaud the program of ecological rationality for taking the evolutionary roots of (human) cognition seriously. If one tries to dispel apparent instances of human irrationality, however, one should not take recourse to ultimate, evolutionary alibis. The foibles of reason cannot be exculpated simply because they display some evolutionary rationale, whether arising from error management, adaptive bias, or some mismatch between evolved heuristics and modern environments. It is no criticism of a screwdriver (or its human designer) to note that it makes a poor crowbar – but this does not exculpate the DIY enthusiast who uses a screwdriver as a lever. Likewise, we can admire the adaptive or functional design of a heuristic while still impugning the rationality of individuals who blindly misapply it – and should know better. Conversely, evidence of adaptive cognitive design is not evidence of human rationality. We cannot fully congratulate ourselves on our ‘rational’ behaviour unless we have some cognitive access to the goal we want to achieve, our strategy for attaining it, and our reasons for thinking this strategy is or might be successful. Insofar as humans are working on automatic pilot, profiting from (or being misled by) the wisdom of their evolved heuristics, the normative categories of (ir)rationality do not apply. Genuine rationality presupposes an intentional striving for success. The program of ecological rationality has pursued two distinct projects, although the distinction between these has been underappreciated. In its best moments, it has convincingly demonstrated that the rigid application of content-blind logical norms as a benchmark of rationality produces an uncharitable view of human cognition, a view insensitive to the subtlety and richness of human reasoning. On the other hand, advocates of ecological rationality, together with evolutionary psychologists, have also pointed out the adaptive rationale of some of our cognitive biases. That research is certainly fascinating, but we should remain wary of subtle locus shifts between different levels of rationality. In some cases (e.g. superstition, gambler’s fallacy), having an adapted mind is compatible with misbelief and irrationality. Even sheer stupidity."],["Using carers to help assess, monitor, or promote health in people with intellectual disabilities (ID) may be one way of improving health outcomes in a population that experiences significant health inequalities. This paper provides a review of carer-led health interventions in various populations and healthcare settings, in order to investigate potential roles for carers in ID health care. We used rapid review methodology, using the Scopus database, citation tracking and input from ID healthcare professionals to identify relevant research. 24 studies were included in the final review. For people with ID, the only existing interventions found were carer-completed health diaries which, while being well received, failed to improve health outcomes. Studies in non-ID populations show that carers can successfully deliver screening procedures, health promotion interventions and interventions to improve coping skills, pain management and cognitive functioning. While such examples provide a useful starting point for the development of future carer-led health interventions for people with ID, the paucity of research in this area means that the most appropriate means of engaging carers in a way that will reliably impact on health outcomes in this population remains, as yet, unknown. © 2014 The Authors. --------------------------------------------------------------------------------","Physical and mental health inequalities have been well documented for people with intellectual disabilities (ID) (Cooper, Smiley, Morrison, Williamson, & Allan, 2007; Emerson, 2011; Emerson, Baines, Allerton, & Welch, 2010; Janicki et al., 2002; Kerr et al., 2003; Servais, 2006; Straetmans, van Schrojenstein Lantman-de Valk, Schellevis, & Dinant, 2007; Underwood et al., 2012). Although there have been substantial increases in life expectancy for people born with ID over the past 60 years (Bittles et al., 2002; Hollins, Attard, von Fraunhofer, McGuigan, & Sedgwick, 1998; Puri, Lekh, Langa, Zaman, & Singh, 1995; Yang, Rasmussen, & Friedman, 2002), their median age at death remains far below that of those without ID, with the disparity increasing with the severity of the ID (Bittles et al., 2002; Glover & Ayub, 2010; Thomas & Barnes, 2010). The most recent data published on life expectancy comes from a confidential inquiry into the premature deaths of people with ID in the UK (Heslop et al., 2013). The review covered all deaths (N = 247) between 1st June 2010 and 31st May 2012, of people with ID aged 4 years or older, who were registered with a GP in one of 5 areas in South West England. The median age at death for men with ID was 13 years below that of men in the general population – for women the difference increased to 20 years. Just under half (48%) of these deaths were deemed to have been ‘avoidable’, meaning they could have been avoided through good-quality healthcare (‘amenable’ deaths) or through public health interventions (‘preventable’ deaths) (Heslop et al., 2013). Health checks ~~~~~~~~~~~~~ In light of the increasing awareness of health inequalities, and of the barriers to accessing good quality healthcare often experienced by people with ID (Alborz, McNally, & Glendinning, 2005; Backer, Chapman, & Mitchell, 2009; Krahn, Hammond, & Turner, 2006; Redley, Banks, Foody, & Holland, 2012), several countries have in recent years introduced primary care health checks for this population (Barr, Gilgunn, Kane, & Moore, 1999; Lennox et al., 2007; NHS, 2008; Webb & Rogers, 1999). Research has found that health check programmes for adults with ID identify unmet health needs (Baxter et al., 2006; Lennox et al., 2007), however their longer-term impact on health outcomes remains to be established. Carer-led interventions ~~~~~~~~~~~~~~~~~~~~~~~ For those adults who have difficulty recognising and gaining treatment for their health needs, carers, whether paid or voluntary staff, family, or spouse carers, may be in a position to monitor illness symptoms, promote healthy lifestyles, or advocate between the adults they care for and their healthcare providers (Langan, Whitfield, & Russell, 1994). By holding such a key role in the daily lives of people with ID, carers could potentially provide a useful resource in terms of more formally assessing and monitoring health needs and promoting positive health outcomes in the person they care for. Certain existing ID health checks include carers in the health check process, to provide support, advocacy and information about the patients’ past and present health status (Lennox et al., 2007; Turk et al., 2010), however the impact of such interventions on health outcomes is unclear, and adherence tends to be poor. The primary aim of this study was to conduct a rapid, systematic literature review of existing health interventions led by carers. We conducted a broad search, across age-groups which was not limited to the ID population, in order to find examples of, and outcomes from, carer-led health interventions in a variety of settings. The results of this review may inform future research to develop and evaluate carer-led interventions to improve the primary care and health outcomes of individuals with ID. Search strategy ~~~~~~~~~~~~~~~ We used rapid review methodology (Ganann, Ciliska, & Thomas, 2010; Watt et al., 2008). Literature was identified via four main sources: the Scopus database, citation tracking, hand-searching reference lists and expert input. Several people working and publishing in the field of health care provision and ID were asked to identify any evidence-based carer- led interventions that they were aware of. Scopus was chosen as the search database, as it is currently the largest abstract and citation database of peer-reviewed literature, and has 100% Medline coverage (“SciVerse Scopus Facts & Figures,” 2010). Initial searches were conducted in February 2013. Citation tracking included following ‘cited by’ links in Scopus searches, and setting up the searches themselves such that any articles meeting the search criteria that were published after our initial search, would automatically be added to the list. 31 May 2013 was deemed the end date for adding new literature. We included the following search terms: TITLE-ABS-KEY (“carer-led” OR “caregiver-led” OR “carer- assisted” OR “caregiver-assisted” OR “carer-directed” OR “caregiver-directed” OR “parent- led” OR “parent-assisted” OR “parent-directed” OR “spouse-led” OR “spouse-assisted” OR “spouse-directed”) AND health AND (“intervention” OR “check” OR “monitor*” OR “program” OR “review”). Criteria for selection ~~~~~~~~~~~~~~~~~~~~~~ Inclusion criteria were: original research; published in English-language, peer-reviewed, academic journals; concerning carer-led interventions for improving, monitoring, screening or promoting physical or mental health, in healthy or clinical populations. Adult and child participants older than 2 years were included. ‘Carers’ included paid or voluntary care staff, family or spouses/partners. Review papers were excluded from the final list, but they were used for hand-searching original research articles. Additional exclusion criteria were single case studies, method papers not reporting outcomes from the intervention, infant participants (less than two years old), interventions not led by a carer, carer-led interventions not targeting physical or mental health, grey literature and papers published in languages other than English. The titles and abstracts of all articles were initially assessed for inclusion by the lead researcher (RH). Potentially qualifying papers were read in full to extract details about the study design and method, participants, carer group, intervention, controls used and main findings. Final decisions on inclusion of the studies were agreed by the review team (RH, AS and MB) and all authors contributed to interpretation and integration of the findings. The reviewers were not blind to the authors, institutions or journal of publication when assessing the eligibility of the papers. Quality assessment ~~~~~~~~~~~~~~~~~~ The level of evidence provided by each study was agreed by two researchers (RH and AS), using the levels of evidence hierarchy published by the Centre for Evidence-Based Medicine (Howick et al., 2011). The scale runs from 1A to 5, where 1A is the highest level of evidence possible (i.e. a systematic review of randomised controlled trials (RCTs)) and 5 the lowest (expert opinion without explicit critical appraisal). These ratings were used to help describe the evidence found, with greater attention given to those studies that provided a higher level of evidence. The interventions themselves, and the methodology and outcomes used to assess them, varied widely. It was not possible to quantitatively assess the efficacy of carer-led interventions, thus results will be presented in the form of a narrative review, with studies combined by common themes. Overview of the included studies ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 1 shows the results of the search process and reasons for excluding studies. 24 papers were included in the final review. An overview of these studies, including the carer group, patient group, setting, study method, main findings and evidence rating is provided in Table A1. Summary of results The interventions included in the final list were divided into the following themes (number of papers; total number of participants included): carer-led pre- health check questionnaires and patient-held records (3; N = 503), carer-led interventions for health promotion (12; N = 3498), carer-led symptom monitoring and management (4; N = 445), carer-led interventions for mental health (4; N = 267), screening delivered by carers (1; N = 108). Carer-led pre-health check questionnaires and patient-held records ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The Comprehensive Health Assessment Programme (CHAP) and Advocacy Skills Kit Diary (ASK Diary) were developed in Queensland, Australia, by Lennox and Colleagues (Lennox, Rey- Conde, & Faint, 2008; Lennox et al., 2010). The CHAP includes an extensive pre-check questionnaire that is completed by a carer, followed by a GP-led health check. The ASK diary combines a physical health record and information for GPs working with people with ID, with advocacy skills training provided for the person with ID if able to self- advocate, or those advocating on behalf of someone in their care. The intention is to improve communication between health professionals, people with ID and their carers and the diary has been designed for continual use in all healthcare environments. An initial feasibility study with 30 adolescents with ID found that the CHAP was effective at identifying unmet health needs and increasing health-related activity. However the ASK diary, although popular with the participants, their parents and their teachers, did not lead to improved communication levels: despite initially stating their intention to ‘speak up’ in the health check, not one of the adolescents in the study reported asking a question, or requesting further information, during medical visits (Lennox et al., 2008). A later four-arm cluster RCT compared the effects of the CHAP and ASK Diary with care as usual in adults with ID (Lennox et al., 2010). This trial provides strong evidence for the efficacy of the CHAP in identifying previously unmet health needs and increasing health- related activity. However, neither this RCT nor the earlier pilot study considered the impact of the initial carer-led questionnaire in isolation from the GP-led section of the CHAP. As such, conclusions about the effect of including carers in the health check process are difficult to draw. The RCT found that the ASK diary had no quantifiable impact on measured health outcomes, which included vaccination rates and clinical activity such as vision checks or weight monitoring. However, exit interviews with 94 carers revealed that 19% of them had never used the diary: any effects could therefore have been under- estimated. Poor adherence was also a key concern for Turk and colleagues, who developed and trialled a personal health profile (PHP) for adults with ID in London, UK (Turk et al., 2010). The PHP aimed to improve health knowledge in the individual with ID, their carers and health professionals, and could be updated by anyone involved in the individual's care. Despite more than 90% of adults with ID and their carers expressing satisfaction in having the PHP and 87% saying it helped them know more about their health or the health of the person they cared for, in exit interviews only 66% of adults with ID and 55% of carers reported actually using the profile. Furthermore, when adults with ID or their carers were asked whether they had a particular condition, there was no difference in valid responses between the PHP group and controls during the follow-up assessments. This indicated that the PHP did not lead to significant increases in personal or general health knowledge, in either adults with ID or their carers. The profile did not lead to significant increases in attendance at GP or other healthcare services. Taken together, these three studies provide little support for the use of carer-completed health records as a means of improving communication with health professionals or short-term health outcomes for people with ID. Parent-led interventions for managing childhood overweight and obesity In a pilot RCT involving 50 families, Moens and Braet (2012) found that a combination of healthy lifestyle, behaviour change and parenting training for parents of overweight children led to significant reductions in their child's BMI. The BMI in a wait-list control group (WLC) did not change significantly. Comparisons with a reference group, composed of children whose parents had chosen not to take part in the main study, found that members of the reference group were no different to the intervention group in terms of the child's gender, age or BMI at baseline, but showed significant increases in BMI 12 months later. While showing that developing a variety of skills in parents can lead to positive health outcomes for their children, this research did not consider how the different components of the intervention might have impacted on BMI reduction. Two earlier RCTs were more comprehensive: Golley, Magarey, Baur, Steinbeck, and Daniels (2007) randomised the parents of 111 overweight pre-pubertal children to parenting skills plus lifestyle education (PS + LE), parenting skills (PS) alone, or WLC. Participants in the PS + LE and PS groups both showed significant reductions in BMI Z-score and waist circumference over the 12-month study period, whereas those in the WLC showed a smaller reduction in BMI, and no significant change in waist circumference. Magarey et al. (2011) composed their groups in the converse fashion: parenting skills training was combined with healthy lifestyle training (N = 85) and compared to lifestyle training alone (N = 84). Participants achieved an average 10% weight loss in both groups, indicating no additional benefit from combining the healthy lifestyle training with parenting skills. A further RCT from Golley, Magarey, and Daniels (2011) on overweight pre-pubertal children, used the same three conditions (PS, PS + LE and 12-month WLC) with the child's food intake and activity levels reported by parent completed questionnaires as the outcome measures. Mean daily consumption of ‘extra’ food (defined as energy-dense, nutrient-poor foods) significantly reduced in both intervention groups, with reductions being maintained for up to 6 months. Food intake was unchanged in the WLC group and there were no significant differences between the groups in terms of activity levels, at any stage of the study. Resnick et al. (2009) randomly assigned 46 parents to a health education intervention, delivered through the post, or via contact with a health worker. Intervention materials included a cookbook, physical activity book, pedometer and information about making healthy eating and activity choices. Although parent-report measures showed no significant increase in parental knowledge after the intervention, their children's BMI percentile reduced by an average of 5 points over the 41-week period. No significant difference in BMI percentiles was seen between the groups, indicating that the materials had a similar effect whether read alone by the parents or delivered by a health worker. A non-randomised contemporaneous control group was found to have no change in BMI during the same period. Small and colleagues conducted a pilot study of an educational intervention for 45 mothers of 4–6 year old children (Small et al., 2012) and found that their educational intervention, consisting of nutritional information and practical serving tips, also failed to increase parental nutrition knowledge, but did lead to reductions in the average amount of calories served to, and consumed by, the children. These combined results show that targeting parents can lead to improvements in health promotion and healthier lifestyles in their children. Parenting or healthy lifestyle skills training may lead to reductions in children's food consumption and BMI and can be achieved by providing parents with appropriate educational materials, even when these materials do not appear to be improving the parents’ knowledge. However, in the current studies such improvements have been small and the mechanisms behind such health improvements, in the absence of a clear impact on parental knowledge, remain unexplored. Parent-led interventions to improve feeding problems A pilot study by Dovey and Martin (2012) targeted problem feeding in 17 children whose problems resulted from sensory defensiveness. Following a 6-month, parent- led, contingent reward desensitisation intervention, which included children receiving reward tokens after trying problem foods, the children's population- corrected height and weight had increased significantly, although BMI changes did not reach significance. A previous RCT of 156 parents of 2–6 year old healthy children by Wardle et al. (2003) had found that children whose parents were taught to lead a 14-day, graduated food exposure programme showed greater increases in liking, ranking and consumption of target vegetables than children whose parents only received nutritional advice or WLC. In the exposure group, these increases were all significant. The control group showed small, but significant increases in liking and ranking of the target food, but decreased their consumption, whereas those whose parents had only received nutritional advice showed no significant increases in any of the measures. These results suggest that parent-led graduated exposure techniques can be successfully used to increase their children's consumption of healthy foods. Parent-led drug and alcohol education interventions Beatty, Cross, and Shaw (2008) assessed their parent-focused, tobacco and alcohol education intervention using a school group randomisation procedure. 1201 parents of 10–11 year old children were randomised to the intervention that consisted of five sets of learn-at-home drug education materials and activities, or to the control group. Post-intervention questionnaires revealed that intervention group parents were more likely to have recently spoken to their child about the hazards of tobacco and alcohol, covering more of the essential topics outlined in the educational materials and to have reported higher levels of engagement when talking with their child, than control parents. However, no behavioural outcome measures were included, thus the relationship between improved communication and the children's alcohol or tobacco use remains unclear. Jackson and Dickinson (2011) randomised 1183 non-smoking parents to either an anti-smoking parenting programme or a control group who received basic factsheets. Their original unpublished results found no significant group differences in smoking-specific outcomes in the children. Additional analyses reported in the current paper, however, revealed a dose–response pattern: greater parental engagement with the materials led to their children recalling more of the 14 behaviours targeted by the programme, and experiencing fewer pro-smoking risk factors (such as perceiving easy access to cigarettes, or feeling that there would be no negative consequences if their parents found them smoking) when interviewed both six months and three years following the initial intervention. This increased awareness of anti-smoking advice and pro-smoking risk factors were not, however, sufficient to change smoking behaviour: the children's likelihood to start smoking three years later showed no association with the amount their parents engaged with the study materials. Parent-led educational interventions appear to be a potentially effective way of improving parent–child communication and increasing children's knowledge of substance-use associated risk, but in the studies described, such improvements did not lead to positive changes in behaviours. Spouse-assisted lifestyle change interventions Voils et al. (2013) conducted an RCT comparing a telephone-delivered spouse- assisted lifestyle change intervention, with treatment as usual, for 255 adult patients with high cholesterol. Their intervention included information about hypercholesterolemia and joint goal setting to improve the healthiness of their lifestyle. Cholesterol levels did not change post-intervention, however patients in the intervention group had significantly lower self-reported caloric intake than those in the control group. The intervention group also reported 10% longer, and 20% more frequent, physical activity sessions than the controls. Carer-led symptom monitoring and management ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Porter and colleagues’ RCT (Porter et al., 2011) compared caregiver-assisted Coping Skills Training (CST) to an educational control, for patients with early stage lung cancer and their carers. Caregiver-assisted CST involved 14 telephone sessions, over 8-months, which taught caregivers how to help patients acquire and maintain coping skills over the illness trajectory. The control condition provided patients and caregivers with information about lung cancer and treatment options. Significant improvements in ratings of worst pain, physical and functional well-being, lung cancer symptoms, depression and self-efficacy were seen in the patients in both groups, with no significant differences between them. Thus in this study, caregiver-assisted CST was no more beneficial than providing patients and their carers with helpful information. Keefe et al. (1996, 1999) conducted a three-arm RCT with patients with osteoarthritis of the knee, comparing spouse-assisted CST with conventional CST with no spousal involvement and also with a control group receiving a spousal support arthritis education package). At the end of the 10-week study period, patients in the spouse-assisted CST condition had significantly lower levels of pain, psychological disability, and pain behaviour, and higher scores on measures of coping attempts, marital adjustment, and self-efficacy, than controls. There were no significant post-intervention differences between those receiving spouse-assisted CST and conventional CST (Keefe et al., 1996) and after 12 months, patients in both CST groups had had lower levels of physical disability and higher levels of self-efficacy than the controls (Keefe et al., 1999). Abbasi et al. (2012) compared a spouse-assisted multi-disciplinary pain management programme (SA-MPMP), consisting of seven weekly two-hour group sessions covering dyadic pain coping and couple skills, with a conventional patient-oriented pain management programme (MPMP) and standard medical care (SMC), for 36 patients with chronic lower back pain. All groups showed significant improvements in disability scores post- intervention, however at the 12 month follow-up, while both intervention groups’ scores remained lower than baseline levels, those who received SMC had higher levels of disability than at the start of the study. Taken together, these studies suggest that including carers in health monitoring and symptom management in long-term physical health conditions offers little benefit over and above that obtained through existing interventions or educational packages. Carer-led interventions for mental health ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Quayhagen and Quayhagen (2001) reported on two RCTs assessing 12 and 8 week versions of their caregiver-led cognitive stimulation intervention (CSI) for people with dementia. This involved hour-long sessions five times a week, focussing on memory, problem solving and communication skills. Controls included placebo (watching TV with a carer) and WLC. The first study (N = 56) showed improvements in immediate memory following the 12-week intervention, whereas the second (N = 30) found improved problem-solving after 8 weeks of CSI. In both studies, CSI groups also showed improved verbal fluency, whereas performance across all measures declined in both control groups. An RCT of 156 dementia patients currently receiving the dementia medication Donepezil compared the effects of a carer-led reality orientation programme, with treatment as usual (Onder et al., 2005). The programme involved three 30-minute sessions per week, during which the caregiver directed attention to the current date, time and location. This was followed by discussing topics such as historical events, and exercises designed to target attention, memory and visuospatial skills. Following the 25-week intervention period, significant group differences were seen in the patients’ cognitive functioning. Mini mental state examination (MMSE) scores and scores in the cognition subscale of the Alzheimer's disease Assessment Scale (ADAS-Cog) showed slight improvements in patients receiving the programme, compared to substantial decline in those receiving care as usual. Although such improvements were small, the authors argue that a difference in just one point in the MMSE is enough to substantially change the cost of caring for that patient. There were no significant differences in caregiver burden, anxiety, or depression after the study (Onder et al., 2005) and a recent feasibility study by Milders, Bell, Lorimer, MacEwan, and McBain (2013) reported that delivering a cognitive stimulation intervention was beneficial not only to the dementia patient, but to carers as well. Finally, a small single group pilot study assessed a parent-led cognitive behavioural therapy (CBT) programme for 26 young children with anxiety disorder (van der Sluis, van der Bruggen, Brechman-Toussaint, Thissen, & Bögels, 2012). This small single group pilot study found that parent-led CBT decreased both parent- and teacher-reported measures of child anxiety and behavioural inhibition, and increased mothers’ use of positive parenting methods. This group of studies suggests that carers are capable of delivering psychological interventions that lead to significant mental health improvements in those that they care for. Screening delivered by carers ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A two-part study by Paysse, Camejo, Hussein, and Coats (2004) assessed parents’ reliability when delivering visual acuity testing to their children. In the initial stage, parents used an electronic visual acuity tester (EVA) to assess the visual acuity of their child (N = 64), with reliability assessed by comparison with an ophthalmic technician's results. In the second stage, 44 children were randomly assigned to either group ‘A’ who had their parents assess their visual acuity prior to the technician confirming the result, or group ‘B’ in which visual acuity was assessed using the full protocol by the technician. The researchers found that parents could not only reliably assess their child's visual acuity using the EVA, but that by doing so, the amount of time required with a qualified health professional was significantly reduced. This study provides a positive example of carers reliably performing an essential part of an eye exam, which could set a precedent for including carers in other health check situations. Summary of findings ~~~~~~~~~~~~~~~~~~~ This review summarises the findings from a broad range of carer-led health interventions. Our initial aim was to identify successful intervention models that involved carers in screening, monitoring, or promoting health in adults with ID. However, with such a paucity of research in this field, more general examples were also sought. In non-ID populations, we found examples of carer-led interventions for health promotion, symptom monitoring and management, mental health and health screening procedures. Such diverse interventions and findings make generalisation difficult, however these studies do allow an insight into some aspects of carer-led interventions that do, and do not, seem to have a positive impact on health outcomes for the person being cared for. Such details may prove helpful for developing future carer-led interventions for people with ID and will be explored in more depth. Of the three papers involving participants with ID, all assessed health outcomes following the use of a health diary or profile completed by carers. Despite such records being apparently well received by people with ID, their carers and health professionals, they were not used by a significant minority of those they were designed for and none of the studies found quantifiable improvements in health outcomes following their use (Lennox et al., 2008, 2010; Turk et al., 2010). With such poor adherence to these interventions, it is unclear whether this type of intervention is ineffective, or whether the reported lack of improvement in communication and health related outcomes reflects the fact that too few in the intervention groups engaged with the materials. Authors for all three papers also noted that health outcomes from such interventions may take much longer to come to light than the timescale of these studies allowed (Lennox et al., 2008, 2010; Turk et al., 2010). In terms of health promotion, a series of studies found that developing parenting and/or healthy lifestyle skills could lead to reductions in both BMI and the consumption of unhealthy foods in children (Golley et al., 2007, 2011; Magarey et al., 2011; Resnick et al., 2009). In two studies where parental knowledge of nutrition and health were assessed before and after the intervention, there appeared to be no significant improvement in what they knew (Resnick et al., 2009; Small et al., 2012). Nevertheless, in both studies, positive health outcomes were recorded in their children. One could therefore conclude that some other aspect of the interventions was leading to behavioural change. These findings link well with those of Davison, Jurkowski, Li, Kranz, and Lawson (2013). Their paper reported on the process of including parents as equal collaborators in the development and evaluation of a family-led childhood obesity intervention and they explored factors impacting on the families’ health from their perspective. In analysing their intervention, nutritional knowledge was not felt to be an important factor; rather the parents felt that they needed improved skills in social networking, advocacy, communication skills and conflict resolution. It was a lack of these skills rather than a lack of knowledge that they felt prevented them from providing a healthy lifestyle for their children. When considering the mechanisms that could lead to improvements in patient outcomes following carer interventions, it may therefore be that a focus on skills, rather than knowledge acquisition, is key. Further evidence that improving knowledge is ineffective at modifying behaviour comes from a study assessing a parent-led tobacco education programme (Jackson & Dickinson, 2011). Variation in engagement with the study materials predicted children's knowledge, such that the children whose parents engaged most with the materials, had the highest knowledge of smoking risks and the lowest experience of particular risk factors. However, increased engagement with the programme had no impact on the child's likelihood to have started smoking three years later. In terms of symptom monitoring and management, none of the studies included found that carer-led interventions were any more effective than existing interventions (Keefe et al., 1996, 1999; Porter et al., 2011; Voils et al., 2013). However, one argument for including carers in the healthcare of those they support is that it may help them to use health resources more appropriately. Although these studies suggest that using carers offers no additional benefit to the interventions that already exist, the results indicated that they were no worse. As such, training carers to deliver interventions may be a cost-effective way of ensuring that patients (such as those with severe ID) get the care they need. We also found one promising example of health screening. Parents of children awaiting an optician's appointment were found to be capable of accurately and reliably assessing their child's visual acuity (Paysse et al., 2004). This is a positive example of carer input successfully reducing the need for clinician time, without negatively impacting on the reliability of the procedure. When considering the inclusion of carers in the health check process for people with ID, examples such as these may work as a precedent, highlighting the ability of carers to perform basic health screens, potentially increasing the effectiveness of interventions and health services. For patients with dementia, we found that carers were not only capable of delivering interventions that led to positive improvements in the patient's cognitive functioning (Onder et al., 2005; Quayhagen & Quayhagen, 2001), but that delivering such interventions did not increase caregiver burden or anxiety (Onder et al., 2005) and in many cases, was felt to be beneficial to the carer as well (Milders et al., 2013). Limitations of existing research ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Several studies reported here were small scale, with short follow-up periods and varying levels of quality in their study control. Future studies need to be fully powered to test efficacy. As already noted, poor adherence to, or engagement with, interventions was a relatively commonly reported issue in these studies (Jackson & Dickinson, 2011; Lennox et al., 2010; Turk et al., 2010). Such variability in engagement with an intervention will inevitably impact on the results. Developing interventions in collaboration with carers may be one way of ensuring that the interventions are acceptable to those delivering them, and in turn, improve engagement (Davison et al., 2013). A further limitation highlighted by Turk et al. (2010) was the high turnover of carers experienced by adults with ID in their study. The authors argued that this in itself highlighted a need for a patient-held health record, to ensure that necessary information was readily available, regardless of how well care staff knew the individual in question (Turk et al., 2010). In many of the examples we found from outside the field of ID, the carer in question was either a parent or spouse. While this may too be the case for some people with ID, others will rely on paid or voluntary care staff. The studies presented here have provided good support for being able to train carers to deliver health interventions, but in contexts where care staff may change frequently the acceptability of the intervention from carers’ and patients’ points of view, and time and financial costs involved in training new staff when required, would need to be incorporated into the research programme. Participant attrition was a significant problem for many of the studies. The centre for evidence-based medicine argues that retention of less than 80% in a study reduces the quality of the study (Howick et al., 2011). Where it was investigated, for example by Jackson and Dickinson (2011) on their anti-smoking programme, differential attrition occurred, such that those who dropped out were more likely to be from the BME community, and to have high school education or lower, than those who chose to continue. Identifying factors involved in participants discontinuing with a study is important as such factors would need to be addressed and improved in order to make the intervention generalisable if successful. A final limitation is that none of the studies included a cost analysis of the interventions presented. Using carers to deliver health interventions could save clinician time, and therefore money (Paysse et al., 2004). However, the cost effectiveness of such interventions will depend on the cost of training carers, and ensuring that they maintain reliability and accuracy in the intervention. In a context of frequently changing care staff, such costs could accumulate more quickly than in contexts where carers are more consistent. Strengths and limitations of this review ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As far as the authors are aware, this is the first review to consider carer-led health interventions as a means of improving the health of the people they care for, including those with intellectual disabilities. Our intention was to focus on adults with ID, however a lack of existing research led to the inclusion of carer-led health interventions from a variety of patient populations, with different medical conditions, of differing age groups and in different settings. Such variation in the samples, the interventions themselves, and the outcome measures used to assess their efficacy, meant that a quantitative review of carer-led interventions was not possible. As such, over-arching conclusions concerning the efficacy of carer-led health interventions as a means of improving health are difficult to make. Our research was time-limited, thus rapid review methodology was used. To reduce search time, we used just one database (Scopus), which may have resulted in missing some papers. This review also excluded studies that were not in English, and did not include interventions published in the grey literature. It is likely that other carer-led interventions exist, and are in use, but have yet to be evaluated. Nevertheless, several successful models have been presented, and barriers to research, such as high carer turnover and a lack of engagement with the interventions have been identified. This review may therefore help future development of carer-led health interventions in general, not just in the ID population. Directions for future research ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ With such a broad range of studies presented here, future research needs to aim to define the elements of carer-led interventions that are most likely to make a positive impact on a person's health. We need to establish how best to involve carers to ensure that high quality interventions are delivered, and that the interventions are acceptable to both those who are delivering, and those who are receiving them. Identifying factors that would improve participant adherence and engagement with any given intervention is essential for improving the quality of future interventions. Including carers and the people they care for in the planning and development stages may be one way of producing interventions that are better adhered to. Most crucially, it is essential that future interventions have a positive impact on health outcomes in the target population. Future studies must include sufficiently long follow up periods, to assess longer-term impacts, and health outcomes should be presented in ways that allow the reader to assess their impact in terms of noticeable changes for the patients, in addition to statistically significant results.","Carers have successfully delivered healthcare interventions, including screening, monitoring and health promotion, in a range of health settings. Nevertheless, for people with ID, there remains a lack of research in this area. With such varied interventions and outcomes found in the studies described in this review, it is hard to draw clear conclusions as how best to use carers to impact positively on the health of the person they care for in the ID community. In addition, engagement with the intervention appeared to be a common issue. Involving carers and the people they care for in the research process may lead to better adherence to, and engagement with, carer-led interventions and should be an integral part of future research.","The authors declare that there are no conflicts of interest.","This article presents independent research funded by the National Institute for Health Research (NIHR) under its Programme Development Grants programme (Reference Number RP- DG-0611-10003). The views expressed are those of the authors and not necessarily those of the NHS, the NIHR or the Department of Health."],["We explored the development of sensitivity to causal relations in children's inductive reasoning. Children (5-, 8-, and 12-year-olds) and adults were given trials in which they decided whether a property known to be possessed by members of one category was also possessed by members of (a) a taxonomically related category or (b) a causally related category. The direction of the causal link was either predictive (prey. →. predator) or diagnostic (predator. →. prey), and the property that participants reasoned about established either a taxonomic or causal context. There was a causal asymmetry effect across all age groups, with more causal choices when the causal link was predictive than when it was diagnostic. Furthermore, context-sensitive causal reasoning showed a curvilinear development, with causal choices being most frequent for 8-year-olds regardless of context. Causal inductions decreased thereafter because 12-year-olds and adults made more taxonomic choices when reasoning in the taxonomic context. These findings suggest that simple causal relations may often be the default knowledge structure in young children's inductive reasoning, that sensitivity to causal direction is present early on, and that children over-generalize their causal knowledge when reasoning. © 2013 Elsevier Inc. --------------------------------------------------------------------------------","Children make category-based inductions when they infer properties and features in novel categories based on what they know to be true about familiar related categories (for reviews, see chapters in Feeney & Heit, 2007, and Hayes, Heit, & Swendson, 2010). Many different types of relations between categories can support such inferences. For example, the fact that tigers have a property or that antelopes have a property may be equally good evidence that lions have the property. The first inference might be strong because lions and tigers are taxonomically related, whereas the second may be strong because lions eat antelopes and this food chain relation provides a plausible causal mechanism for property transmission. This example is consistent with claims based on structured Bayesian approaches to inductive reasoning (see Kemp & Tenenbaum, 2009) that our knowledge about the relations that hold between categories of objects can be structured in a variety of ways. One of our aims in this study was to examine whether causal or taxonomic relations are more privileged in young children’s category-based inductive reasoning. It was unclear which knowledge structure might serve as the default because some researchers suggest that taxonomic reasoning is a default strategy (e.g., Kemp & Tenenbaum, 2009; Shafto & Coley, 2003), whereas others emphasize the primacy of causal knowledge (e.g., Rehder, 2006; Rehder, 2009). Because they are inductive, category-based inferences are probabilistic, but they effectively reduce uncertainty about the world. Understanding the constraints placed on inductive inferences by the underlying structure of different knowledge sources is crucial if we want to understand the processes that allow inductive inferences to be flexible yet effective. Several recent studies (Kemp & Tenenbaum, 2009; Shafto, Kemp, Bonawitz, Coley, & Tenenbaum, 2008) show that adults’ inferences are especially sensitive to knowledge about how causal relations are structured. However, little is known about whether children and adults use causal knowledge in similar ways to support their inductive inferences. Our second aim of this study was to examine whether, like adults (see Rehder, 2009; Shafto, Coley, & Baldwin, 2007; Shafto, Coley, & Vitkin, 2007; Shafto et al., 2008), children are sensitive to the direction of the causal relation that holds between categories. Thus, in addition to examining when children’s inductive inference becomes sensitive to causal relations, we examined how sophisticated children are in their use of such knowledge for reasoning. Causal knowledge in inductive reasoning ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The effects of causal knowledge on reasoning are not very well captured by older models of category-based induction that emphasize featural similarity (Sloman, 1993) and/or class membership (Osherson, Smith, Wilkie, Lopez, & Shafir, 1990). Such similarity-based models are powerful at accounting for patterns of inductive reasoning about taxonomic properties (i.e., properties such as genes whose distribution in the population may depend on taxonomic relations) and about blank properties (i.e., properties that participants possess no knowledge about). However, they fail to capture induction across a broader variety of properties and in expert populations (see Medin, Coley, Storms, & Hayes, 2003; Rehder & Hastie, 2001; Shafto & Coley, 2003), especially when there is a causal explanation for the occurrence of shared properties (Rehder, 2006). Causal knowledge plays a vital role in cognition from infancy onward (Sobel & Kirkham, 2007). The ability to understand causal structures provides children with tools that help them to successfully predict future events and understand the outcome of active intervention, allowing them to gain increasing control over their environment (Gopnik et al., 2004). By 4 years of age, children are capable of understanding simple causal mechanisms across the domains of biology (Wellman, Hickling, & Schult, 1997) and psychology (Flavell, Green, & Flavell, 1995) as well as causal explanations in social and physical domains (Hickling & Wellman, 2001). Similarly, children use causal knowledge to classify objects (Ahn, Gelman, Amsterlaw, Hohenstein, & Kalish, 2000) and natural kinds (Meunier & Cordier, 2009). The fact that children make use of causal information across diverse domains and tasks underscores its potential importance in children’s category-based reasoning. Indeed, evidence suggests that children can use causal knowledge when making inductive inferences. For example, Hayes and Thompson (2007) taught children (5- and 8-year-olds) and adults about features of two artificial base creatures, followed by a target that was more similar to one base but shared a causal antecedent with the other base. Results indicated that when the causal link was explicit, all age groups preferred to make causal rather than similarity-based inductions. That is, they preferred to project a property to the target from the causally related base creature than from the more similar base creature. When the causal relation was implicit, 5-year-olds did not yet show a preference for choosing the causally related items, unlike the older children and adults who made predominantly causal choices. Similarly, Opfer and Bulloch (2007) demonstrated that 5-year-olds were capable of ignoring perceptual similarity in favor of relational similarity when the latter had a causal antecedent but not when it was non-causal. Both of these studies suggest that children’s category-based inductive reasoning may be affected by knowledge about causal relations, although they do not allow us to conclude that causal knowledge structures are the default. Moreover, such studies do not address the extent to which children’s inductive reasoning is sensitive to the underlying structure of causal knowledge (e.g. Pearl, 2000; Sloman, 2005). Consequently, we cannot know whether children’s use of causal knowledge in induction is mediated by the same underlying processes as in adults. Structure of causal knowledge: Causal asymmetry effects ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ An important feature of causal knowledge is that it is directional, with causes always preceding or at least co-occurring with their effects (Waldmann, 2000). There is evidence that adults take such structural features into account because their causal reasoning tends to display distinctive asymmetry effects (Fenker, Waldmann, & Holyoak, 2005; Fernbach, Darlow, & Sloman, 2011; Sloman, 2005). Reasoning in line with how we experience cause–effect relations (predictive causal reasoning) appears to be less effortful than reasoning backward (diagnostic causal reasoning), suggesting that computational complexity is a key determinant of the causal asymmetry effect (Kahneman & Tversky, 1973). For example, when people are asked to verify whether two words are causally related, they respond faster when the words are presented in a predictive order compared with a diagnostic order (Fenker et al., 2005). Causal direction also affects inductive inferences. Thus, people find category-based inductive arguments more convincing when reasoning in a predictive direction from cause to effect than when reasoning diagnostically from effect to cause (Medin et al., 2003; Shafto et al., 2008). To illustrate, people are more confident in the conclusion that bees have an unknown property given that it is present in the premise category flowers than when the roles of the two categories are reversed. People’s judgments appear to accord well with the causal asymmetry effects that are predicted by formal models of causal-based property generalization (Rehder, 2009; Shafto et al., 2008), although there are arguments that, relative to a normative standard, people have too much confidence in arguments with a predictive structure (Fernbach et al., 2011). Despite the convincing evidence that the causal asymmetry effect is a robust phenomenon in adults’ category-based reasoning, and evidence that children’s predictive reasoning about novel mechanical systems is better than their diagnostic reasoning (see Bindra, Clarke, & Shultz, 1980; Hong, Chijun, Xuemei, Shan, & Chongde, 2005), it is unknown whether children display similar asymmetry effects when they use causal knowledge to support their inductive inferences. Our study was designed to answer this question. Context-sensitive reasoning ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Although it is clear that causal knowledge is important in induction, adults are adept at tailoring their inferences to the particular reasoning context. Consider the example with which we opened this article. If we ask a participant to decide whether lions have a certain gene on the basis that tigers have that gene, then the taxonomic relation that holds between lions and tigers is relevant. On the other hand, if we ask whether lions have a disease on the basis that antelopes have that disease, then the causal mechanism for disease transmission from antelope to lions via predation is relevant. There is much evidence that adults, and sometimes children, can show inductive selectivity by selecting the appropriate knowledge with which to evaluate category-based inductive arguments. For example, Heit and Rubinstein (1994) showed that people make stronger inferences about shared properties when the nature of the property is in accord with the nature of the relationship between two categories. According to Shafto, Coley, and Baldwin (2007); Shafto, Coley, and Vitkin (2007), changes in context affect the acute availability of different knowledge structures. Thus, when reasoning about properties such as diseases, people are more inclined to use knowledge about causal or ecological relations, whereas taxonomic knowledge tends to be invoked when reasoning about anatomical properties. There is some evidence that children are also sensitive to the nature of the property that they are reasoning about (for a full review, see Hayes, 2006). Nguyen and Murphy (2003), for example, showed that by 7 years of age, children reason taxonomically about biochemical features but use thematic relations to guide inferences about situational properties. However, in their study younger children’s use of different relations was heavily influenced by task format, making it difficult to evaluate the degree to which these children were using different knowledge structures in a systematic manner. Coley (2012) demonstrated that there are environmental influences on the age by which children show inductive selectivity; by 10 years of age, children raised in urban environments selectively project disease properties between categories that are ecologically related and so suggest a causal mechanism for disease transmission, whereas rural children show the same pattern as early as 6 years of age. In an open-ended category generation task, Vitkin, Coley, and Hu (2005) also showed that children’s inferences are constrained by the nature of the property that they are asked to reason about. Thus, when told that the two categories share “stuff inside,” children generated categories based on similarity, whereas they gave interaction-based responses when reasoning about shared diseases. In this study we manipulated the nature of the property that participants were asked to reason about. Although we expected adults to show property effects, our aim in this study was not to assess the age by which children demonstrate inductive selectivity because any answer to this question will depend on the paradigm employed. Indeed, inductive selectivity effects have been demonstrated in young children (Heyman & Gelman, 2000a; Heyman & Gelman, 2000b) and even infants (see Mandler & McDonough, 1998; Rakison & Hahn, 2004). Instead, we wanted to examine, in the domain of folk biology, how children reason before selectivity emerges and whether a particular knowledge structure is more important. On the basis of the literature, it is very difficult to answer questions about default knowledge structures. There appears to be reason for assuming that taxonomic knowledge structures are the default; for example, using a speeded response paradigm, Shafto and colleagues (2007) showed that access to knowledge about causal mechanisms for disease transmission due to ecological relations between categories is restricted when adults are asked to respond quickly, whereas taxonomic knowledge seems unaffected by response time manipulations. This finding, that taxonomic knowledge is primary, is consistent with suggestions that taxonomic knowledge structures are the default (Kemp & Tenenbaum, 2009). However, there is also reason to believe that taxonomic knowledge structures are not the default; Rehder (2006), Rehder (2009) emphasized the primacy of causal knowledge over taxonomic knowledge, and there is evidence that young children prefer to use thematic knowledge rather than taxonomic knowledge when reasoning about categories (see Greenfield & Scott, 1986; Nguyen & Murphy, 2003; Smiley & Brown, 1979). Because the developmental literature is divided on the question of how young children reason inductively, there are also important theoretical questions about the development of inductive reasoning to which questions of default knowledge are relevant. Sloutsky (2010) theorized that inductive selectivity depends on a maturationally late learning system. Sloutsky suggested that, prior to the emergence of inductive selectivity, children rely on knowledge acquired through a simple learning system that exploits co-occurrences and similarity. He cited evidence that before 7 years of age children rely largely on such similarity between base and target category to guide their inferences (Fisher & Sloutsky, 2005; Rakison & Lupyan, 2008; Sloutsky & Fisher, 2004). This would suggest that prior to the emergence of inductive selectivity, children ought to rely heavily on superficial similarity or category membership rather than on causal knowledge. However, there is evidence that even children as young as 5 years can ignore featural similarity when explicit causal knowledge is available (Hayes & Thompson, 2007; Opfer & Bulloch, 2007). Here we examined whether children as young as 5 years prefer to reason on the basis of taxonomic or causal relations. This group was younger than the youngest participants that Coley (2012) found to display inductive selectivity. The contradictory findings in the literature made it hard to predict whether these children will prefer arguments with taxonomic or causal relations and whether they will be sensitive to causal structure. If young children’s inductive inferences are dominated by taxonomic relations, then we would except to see few effects of causal knowledge. If, on the other hand, causal knowledge dominates early in reasoning development, then the question arises as to whether young children will exhibit causal asymmetry effects.","In total, 44 Year 1 primary school children (Mage = 5.4 years), 40 Year 4 primary school children (Mage = 8.4 years), 54 Year 8 secondary school children (Mage = 12.6 years), and a control group of 26 adults from Durham University (Mage = 24.6 years) in the United Kingdom took part in the experiment. Year 1 and Year 4 students were recruited from a suburban state primary school in the North East of England. Year 8 students were recruited from a suburban state secondary school in the same geographical area. There were approximately equal numbers of male and female participants in the various age groups. Design ~~~~~~ The experiment had a 4 (Age Group: 5-year-olds, 8-year-olds, 12-year-olds, or adults) × 2 (Property: disease or cells) × 2 (Direction: predictive or diagnostic) × 2 (List: A or B) mixed design, with repeated measures on the direction variable only.","Participants reasoned about 12 trials, each consisting of a triad of categories: a base and two targets. On each trial, participants were presented with a base category, were told that it possessed a feature, and were asked which of the target categories they thought was most likely to share that feature. One of the targets was causally related to the base in the form of a food chain relation. This causally related target category always belonged to a different superordinate category from the base. The other target category was taxonomically related to the base category and was from the same superordinate category as the base. In the predictive condition, the causal direction was ecologically consistent. For example, the base category might be a banana and the causally related target might be a monkey. In the diagnostic version of this example, the directionality of the causal link was reversed so that the monkey served as the base and the banana served as the target. To ensure that children understood the causal mechanisms by which two categories might be connected, the nature of the causal relationship was limited to transparent food chains. An example with pictures, all of which were selected from Rossion and Pourtois (2004) database, is shown in Fig. 1. The direction of the causal link was counterbalanced across items. For List A, the causal direction was predictive for Items 1 to 6 and was diagnostic for Items 7 to 12; for List B, the opposite was the case. Thus, although direction was a within-participants factor, participants never saw any of the causal pairs twice. The task was self-paced and run on a laptop using purpose-built software. So that the causally related pairs could be presented in either a predictive direction (e.g., carrot → horse) or a diagnostic direction (e.g., horse → carrot), different categories served as the base (e.g., carrot or horse). Because the taxonomically related target category belonged to the same superordinate category as the base category, this meant that the taxonomically related target category was different on diagnostic and predictive trials. Thus, it was crucial to equate the similarity between the base instance and the taxonomic alternative targets across predictive conditions (e.g., carrot → onion) and diagnostic conditions (e.g., horse → sheep). A group of 18 Durham University students rated the similarity on a scale from 1 (not at all similar) to 9 (highly similar) between pairs of categories presented verbally in a pretest. Only triads in which the similarity ratings were approximately equal across both conditions were selected (all paired-sample t tests had ps > .05), resulting in 12 triads. Mean similarity ratings by triad are presented in Appendix Table A1). Although we did not collect similarity ratings from children, there is evidence that children and adults perceive similarity relations in similar ways. For example, using materials pretested for similarity with adults by Osherson and colleagues (1990), López, Gelman, Gutheil, and Smith (1992) showed effects of similarity on reasoning in 5-year-olds that were identical to the similarity effects originally found in adults by Osherson and colleagues. The strength of the causal relation between the base and target categories was examined in a separate pretest on 19 Durham University students who rated the extent to which properties might be transmitted between the base and both target categories in each triad on a scale from 1 (not at all) to 9 (a great deal). In all cases, ratings were stronger for causally related categories than for taxonomically related categories. These differences were statistically significant (p < .05) in 9 of 12 cases and were marginally significant in the remaining 3 cases (see Appendix Table A2 for ratings). Procedure Parents received written information about the study with an opt-out form in case they did not want their children to take part. Children gave their assent at the beginning of the session. Participants were randomly allocated to either the cell or disease condition. The task was explained verbally, and written instructions were presented on-screen. Participants completed 2 practice trials before commencing the main task. They saw a base picture at the top of the screen and were told that the animal, plant, or object had an unfamiliar disease or special cells (e.g., “This horse has a disease called talio/has talio cells inside”). A different artificial cell or disease name was used for each trial. The base picture was followed by two arrows pointing to two target pictures. Participants decided which of the two pictures was more likely to share the disease or special cells with the base category. The location of the causally and taxonomically related targets on the screen was randomized. Participants pressed 1 if they thought that the left-hand picture was more likely to share the property and pressed 9 if they chose the right-hand picture.","For each individual, the proportion of causally related targets that were chosen was calculated for the six predictive and six diagnostic triad items. These were analyzed using a 2 (Direction) × 4 (Age Group) × 2 (Property) × 2 (List) mixed-design analysis of variance (ANOVA) with direction as the only within-participants variable. For the item analysis, proportions of causal choices were averaged across items rather than participants. All effects involving the counterbalancing list variable were nonsignificant (all ps > .10), so there is no further reference to this variable. As predicted, there was a main effect of direction, Fs(1, 146) = 20.8, p < .0005, effect size f = .38, Fi(1, 11) = 21.4, p = .001, and a significant main effect of age group, Fs(1, 146) = 6.5, p < .0005, effect size f = .37, Fi(1.9, 21.3) = 22.56 (Greenhouse–Geisser adjustment for non- sphericity), p < .0005. Participants made significantly more causal inferences when the link between the causally related categories was predictive (M = .60, SE = .026) than when it was diagnostic (M = .47, SE = .028). Despite the significant main effects, the interaction between age group and direction did not approach significance, Fs(2, 146) = 0.21, p = .89, effect size f = .006, Fi(3, 33) = 0.55, p = .66. Rather, as may be seen in Table 1, children were as strongly influenced by the structure of causal knowledge as were adults. Post hoc comparisons using Bonferroni adjustments on the means involved in the significant effect of age group showed that 8-year-olds made more causal choices than both 12-year-olds (p = .05) and adults (p < .0005). There was no significant difference in the proportion of causal choices made by 12-year-olds and adults (p = .17), and 5-year-olds made significantly more causal choices than adults (p = .014). Both 5-year-olds, t(43) = 2.69, p = . 001, and 8-year-olds, t(39) = 3.91, p < .0005, made more causal choices than predicted by chance alone. The trend line in Fig. 2 suggests that causal induction shows a curvilinear development in the shape of an inverted U, which was confirmed with a polynomial contrast. This involves starting with the linear contrast and then sequentially carrying out higher power contrasts until the highest-power contrast is nonsignificant. There was a significant linear trend (p < .0005) as well as a significant quadratic term (p = .014). The cubic contrast was nonsignificant (p = .18). The quadratic effect across age group was also highly significant across items, Fi(1, 11) = 15.7, p < .0005. This pattern confirms that the use of causal knowledge in inductive reasoning follows a curvilinear developmental trajectory, peaking at 8 years and decreasing thereafter. The results of the ANOVA contained a significant main effect of property, Fs(1, 146) = 8.6, p = .004, effect size f = .24, Fi(1, 11) = 60.9, p < .0005. Causal choices were significantly more frequent when participants reasoned about diseases (M = .60, SE = .033) than when they reasoned about cells (M = .47, SE = .033). However, this was qualified by a significant interaction between property and age group, Fs(2, 106) = 6.2, p = .001, effect size f = .36, Fi(3, 33) = 62.7, p < .0005. As Fig. 2 shows, both 4- and 8-year-olds made a similar number of causal inductions in both the cell and disease conditions (ps > .25, effect size d < 0.2), suggesting that they are not exhibiting strong inductive selectivity effects. In contrast, both adults (p = .031, effect size d = 0.8) and 12-year-olds (p = .0005, effect size d = 1.2) showed strong inductive selectivity effects, making significantly more causal inductions in the disease condition than in the cell condition. Furthermore, the proportion of causal choices was similar across all age groups when reasoning about diseases (all pairwise comparison ps > .20). In contrast, when reasoning about cells, 4- and 8-year-olds made significantly more causal choices than both 12-year- olds and adults (ps < .022). The latter two age groups were not significantly different from each other (p = .64). Thus, it appears that the increased selectivity is driven by a developmental decrease in causal choices when reasoning about cells rather than a change in the application of causal knowledge when reasoning about diseases. There was no significant interaction between direction and property, Fs(1, 146) = 2.26, p = .13, effect size f = .12, Fi(1, 11) = 2.94, p = .11, and no three-way interaction among age group, property, and direction, Fs(3, 146) = 0.93, p = .43, effect size f = .13, Fi(3, 33) = 0.89, p = .46. In this equation, P(w1 & w2) represents the probability that a text contains both words w1 and w2, and P(w1) represents the probability that it contains w1 on its own. To calculate the conditional probability, one can simply count the number of times w1 and w2 co-occur and divide this by the number of times w1 occurs by chance in the same text sample. This is repeated for w2. We took the mean of the two conditional probabilities for each pair and calculated z scores. Using an independent-samples t test, we compared the mean co-occurrence index for the causally related category pairs (mean z score = −0.3, SD = 0.4) with the mean co-occurrence index for the taxonomically related category pairs (mean z score = 0.15, SD = 1.2). This showed that there was no significant difference between the two co-occurrence indexes, t = 1.68, df = 31.3 (adjusted for unequal variances), p = .10, effect size d = 0.5. However, as indicated by the medium effect size, if anything the taxonomic category pairs co-occurred slightly more frequently, strengthening our claim that children and adults were drawing on structured causal knowledge rather than simple associative knowledge. Another possible alternative explanation of our results is that the youngest participants preferred causally related targets because they did not possess the knowledge about taxonomic relations required to recognize the strength of arguments based on such relations. To rule out this possibility, we checked taxonomic and causal knowledge levels in a separate sample of the same age as our youngest age group. A group of 20 5-year-olds were shown the color pictures of the category pairs used in the main experiment and were asked about biological group membership (“Do this [name of first category] and this [name of second category] belong to the same group?”) and causal relatedness (“Does this [name of first category] eat this [name of second category]?”). Paired-samples t tests across the 20 children and by items showed that there was no difference between the causal relatedness endorsement proportion (M = .78, SD = .13) and the biological group membership endorsement proportion (M = .79, SD = .17), ts(19) = −0.68, p = .51, effect size d = 0.15, ti(11) = −0.037, p = .97. Thus, it is unlikely that the 5- and 8-year-olds in the main experiment preferred causal targets because, relative to their causal knowledge, they lacked taxonomic knowledge.","We had two aims in carrying out this study. First, we wanted to examine whether causal or taxonomic knowledge is the default prior to the age at which children develop inductive selectivity. Second, we wanted to see whether children are sensitive to aspects of causal structure when evaluating category-based inductive arguments. In showing that 5- and 8-year-old participants preferred arguments that were strong because of a causal mechanism for transmission via a predation relation even when the property reasoned about was a cell, our findings suggest that causal knowledge structures are the default in children’s reasoning prior to the emergence of inductive selectivity. In addition, 5-year-olds are just as sensitive to causal direction as are adults. In fact, all age groups in our experiment endorsed significantly more predictive arguments than diagnostic arguments. Interestingly, causal reasoning seems to show a curvilinear development, peaking at 8 years of age and decreasing thereafter due to the emergence of inductive selectivity. Although we did not anticipate this last finding, as we show below, it has parallels in the literature on children’s understanding of folk biology. The pattern of results we observed has very interesting implications for theories about how inductive reasoning develops. Sloutsky (2010) suggested that “adult-like” inductive reasoning depends on a maturationally late learning system. This suggestion can explain why we did not observe inductive selectivity until the age of 12 years, driven by an increase in taxonomic inferences when reasoning about cells. Under Sloutsky’s account, reasoning about cells might require more abstract conceptual knowledge about mechanisms that are not perceptually observable such as genetics and inheritance. However, this account also seems to suggest that prior to the emergence of inductive selectivity children ought to rely heavily on superficial similarity or category membership rather than on causal knowledge. Our data suggest otherwise, with 5- and 8-year-olds making the most causal inductions regardless of context. Alongside recent evidence from Hayes and Lim (2013), who showed that inductive selectivity even in relatively transparent contexts depends on conscious awareness of the relevance of contextual clues, the current findings cast doubt on the claim that early category-based induction is driven exclusively by simple learning mechanisms based on co-occurrence and perceptual similarity. At the very least, it requires a substantial downward revision of the age at which conceptual knowledge such as simple causal structures can support category-based inductive reasoning. In addition to their tendency to prefer causal relations rather than taxonomic relations as the basis for inference, 5- and 8-year-old participants, like older participants, made more causal inferences when the causal relation between the categories was predictive than when it was diagnostic. The influence of different kinds of knowledge seems to be fundamentally related to the way in which this knowledge is organized in long-term memory (Fenker et al., 2005). Although causal knowledge itself does not have a homogeneous structure and can vary in complexity from direct cause–effect links to more complex common cause relations (Sloman, 2005), one might expect the abstract organization of knowledge about food chains involving simple causal transmission to be similar in adults and children. For example, children as young as 3 years show an understanding of the importance of causal order in which effects follow or co-occur with their causes (Bullock & Gelman, 1979; Bullock, Gelman, & Baillargeon, 1982). Reasoning about relations that are in line with this ordering of events might be cognitively simpler than needing to reason about possible causes given an outcome (Kahneman & Tversky, 1973). Shafto and colleagues (2008) suggested that causal asymmetry effects are driven by a multilevel understanding of causal knowledge—concrete knowledge about the existence of causal relations and more conceptual knowledge about how and when causal relations are most likely to warrant an inference from one agent to another. The simple causal structure used in the current task might explain why some of our findings are different from previous results. For example, whereas the current study showed no age-related changes in the frequency of causal reasoning about diseases, work on ecological reasoning by Coley, Vitkin, Seaton, and Yopchick (2005) demonstrates a clear developmental trend in the use of non-taxonomic knowledge. They argued that this is driven by an experiential increase in ecological knowledge. It is conceivable that the structure of the causal relations underlying the ecologically related categories in Coley and colleagues’ study (e.g., between tiger and parrot) was more complex, involving indirect causal pathways and common causes rather than simple and direct cause–effect relations. Thus, increased complexity may render non-taxonomic knowledge less available to reasoning processes in younger children. There may be a number of reasons why the younger children in our experiment relied on causal knowledge rather than taxonomic knowledge when reasoning. The properties used required knowledge about two quite distinctive biological domains: (a) food chain relations and disease contagion and (b) taxonomic relations and genetics. One likely possibility is that knowledge of the genetic basis of shared properties, inheritance, and biological processes is less elaborate in 5- and 8-year-olds compared with 12-year-olds and adults (Au & Romo, 1999; Hatano & Inagaki, 1994; Hatano & Inagaki, 1997). Indeed, educational research suggests that children have conceptual gaps and erroneous concepts in their explanations for genetics (Lewis & Kattmann, 2004; Smith & Williams, 2007) that become more elaborate and accurate only with explicit instruction (Venville & Donovan, 2007). On the other hand, the mechanism by which members of a common food chain may transmit diseases is obvious and observable and constitutes a central part of the science curriculum for 8-year-olds. In line with other work showing that children like to have an explanation for phenomena (Callanan & Oakes, 1992), the children in our study may have preferred to base their reasoning on relations for which they have a mechanistic explanation and a more coherent theory (Gutheil, Vera, & Keil, 1998). Children as young as 3 years have a basic appreciation that invisible agents can cause illness (Kalish, 1996) and understand simple mechanisms by which contamination may come about (Siegal & Share, 1990; Springer & Belk, 1994), suggesting that children do have some naive but systematic theory about biological disease transmission. One striking aspect of our results was that younger participants over-generalized their knowledge of disease transmission mechanisms. A relevant example of knowledge over-generalization comes from Keil and colleagues (1999), who showed that increases in knowledge about the mechanisms of disease transmission simultaneously led to more accurate inductions for physical illness contagion but less accurate reasoning about the causes of mental illness. Interestingly, the 8-year-old participants in Keil and colleagues’ study made the most inaccurate transmission choices. This over-generalization parallels our current finding in which simple causal transmission is also seen as a basis for sharing physiological properties such as cells. Thus, inaccurate beliefs or inappropriate causal explanations seem to be supplanted only when children have more appropriate coherent explanatory systems at their disposal. Because they show that causal knowledge dominates taxonomic knowledge in young reasoners, our results are contrary to the claim made by proponents of the structured Bayesian approach that taxonomic knowledge structures are the default (see Kemp & Tenenbaum, 2009; Tenenbaum, Kemp, & Shafto, 2007). Nonetheless, theory-driven Bayesian approaches have the potential to explain developmental changes in patterns of inductive reasoning. The domain-specific theories captured in Bayesian knowledge structures are minimalist and open to revision, both on encountering direct evidence and through verbal instruction. These characteristics of the Bayesian account, in conjunction with the observation that children are likely to have less well- developed domain-specific theories, may allow the Bayesian approach to account for the divergence between 8-year-olds’ and adults’ patterns of inductive selectivity. However, it aims to explain inference at a computational level rather than in terms of process; the Bayesian approach does not address the issues of availability and processing effort (Shafto, Coley, & Baldwin, 2007; Shafto, Coley, & Vitkin, 2007) that are likely to be important when explaining developmental changes in reasoning ability. Finally, our research made use of non-blank properties such as cells and diseases, with inductive selectivity driven by the decrease in the use of causal relations when reasoning about cells. A future interesting question will be to explore the developmental use of causal knowledge for inductive problems that use blank properties. It seems likely that changes in the application of causal knowledge would be driven by cultural factors and expertise with causal choices likely to decrease in Western and/or urban samples, whereas the use of causal knowledge is likely to persist in non-Western and more rural samples (e.g., Coley, 2012; Lopez, Atran, Coley, Medin, & Smith, 1997). To summarize, young children strongly prefer inductive arguments where there is a plausible causal mechanism for property transmission, even where older children and adults appear to judge that mechanism irrelevant to the property to be generalized. Recruitment of this salient source of knowledge may be supplanted only by more contextually appropriate knowledge when formal education offers children an alternative feasible mechanism by which two categories come to share properties. In addition, even 5-year-olds are influenced by the directional nature of causal relations. All of this suggests that causal relations are at least as important to young children’s reasoning as they are to adults’ reasoning and that, in some respects, the manner in which causal reasoning influences reasoning is similar across development."],["Attempting to understand how humor styles relate to psychological adjustment by correlating these two constructs fails to address the emerging understanding that individuals use combinations of humor styles, and that different combinations may be differentially associated with psychosocial adjustment. Indeed humor types have been identified in adult samples (Galloway, 2010; Leist & Müller, 2013). The main aim of the study was to explore whether similar humor types are evident at a younger age and whether these types can be distinguished in terms of children's psychological and social well-being. Participants were 1234 adolescents (52% female) aged 11-13 years, drawn from six secondary schools in England. Self-reports of humor styles and psychosocial adjustment were collected at two time points, 6 months apart. A cluster analysis was performed using the child humor styles scores at Time 1. Four humor types were identified: 'Interpersonal Humorists' (high on aggressive and affiliative humor, low on self-defeating and self-enhancing humor), 'Self-Defeaters' (high self-defeating humor, low on the other three), 'Humor Endorsers' (high on all four humor styles), and 'Adaptive Humorists' (high on self-enhancing and affiliative humor, but low on aggressive and self-defeating humor). 'Self-Defeaters' scored highest in terms of maladjustment across all of the outcomes measured. Our analyses support the presence of distinctive humor types in childhood and indicate that these are related to psychosocial adjustment. --------------------------------------------------------------------------------","Over the past two decades, there has been a steady accumulation of research on the topic of humor. Far less research has focused on the social/emotional functions of humor in children or the way that these functions develop through childhood/adolescence (Martin, 2007). It is recognised that among children and adults, there are four main types of humor style, and these reflect the use of humor in everyday life (Fox, Dean, & Lyford, 2013; Martin, Puhlik-Doris, Larsen, Gray, & Weir, 2003). Self-enhancing humor is the ability to maintain a humorous perspective in the face of stress and adversity; it is closely aligned to coping humor (e.g. ‘My humorous outlook on life keeps me from getting too upset or depressed about things’). Aggressive humor also enhances the self, at least in the short- term, but is done at the expense of others (e.g. ‘If someone makes a mistake I often tease them about it’). Affiliative humor enhances one's relationships with others and reduces interpersonal tensions (e.g. ‘I enjoy making people laugh’). Finally, self-defeating humor, largely untapped by previous humor scales, is used to enhances one's relationships with others, but at the expense of the self (e.g. ‘I often try to make people like or accept me more by saying something funny about my own weaknesses, blunders and faults’). The ability to distinguish between different components of humor has brought with it a clearer picture of the relationships between humor and adjustment. This is evidenced by the stronger correlations between humor and psycho-social adjustment which are reported when using the adult Humor Styles Questionnaire (HSQ) as compared to previous research using unidimensional measures (Martin et al., 2003). Among adults, affiliative and self- enhancing humor are negatively correlated with anxiety, depression, and suicidal ideation, and positively correlated with self-esteem and life satisfaction. In contrast, self- defeating humor is associated with high levels of anxiety, depression, and suicidal ideation, and lower self-esteem and lower life satisfaction (Dyck & Holtzman, 2013; Kuiper, Grimshaw, Leite, & Kirsh, 2004; Martin et al., 2003; Tucker et al., 2013). Aggressive and self-defeating humor styles are both associated with hostility and aggression (Martin et al., 2003). In addition, aggressive humor is not associated with psychological adjustment but is strongly negatively correlated with social adjustment measures (Yip & Martin, 2006). Whether such associations can be generalised to adolescent populations is not yet clear, and data relating to this would clarify the ways in which humor as a coping strategy develops across the lifespan. Using a series of studies, the HSQ was adapted for use with children and young people aged 11–16 years (Fox et al., 2013). An adaptation of the HSQ was needed to enable the examination of all four styles of humor styles in children and adolescents. The HSQ could not be used in its adult form as prior administration with adolescents aged 12–15 (Erickson & Feldstein, 2007) demonstrated unsatisfactory internal reliability for the two maladaptive sub-scales (α = .65 and .58). When used with 11–16 year olds, there were acceptable levels of reliability for all four sub-scales (all α > .70), and both principal components analysis and confirmatory factor analysis identified a clear four-factor structure. In addition, it was found that affiliative humor was positively correlated with global self-worth and self-perceived social competence, and negatively correlated with anxiety and depressive symptoms. Conversely, self-defeating humor was negatively correlated with global self-worth and self-perceived social competence, and positively correlated with anxiety and depressive symptoms. In a more recent longitudinal study using the child HSQ, self-defeating humor was associated with an increase in depressive symptoms and loneliness and a decrease in self-esteem. In addition, depressive symptoms predicted an increase in the use of self- defeating humor over time, thus suggesting a bi-directional relationship. Self-esteem was associated with an increase in the use of affiliative humor over the school year but not vice-versa (Fox, Hunter, & Jones, unpublished manuscript). This was the first study to examine the longitudinal relationships between the four humor styles and aspects of psychosocial adjustment. While such studies can be more difficult to implement, it is important that they are utilised more frequently so that we can better understand the development of humor and its potential negative consequences. Indeed, it could be that particular styles of humor (e.g. adaptive) lead to improved psychological adjustment or that better mental health and well-being facilitates the greater use of adaptive styles of humor. More recently researchers have begun to consider the combination of humor styles characteristic of any given person (Galloway, 2010; Leist & Müller, 2013). These analyses, it is believed, can extend knowledge about the four humor styles beyond what we already know from examining them individually. By identifying different humor types or profiles, new associations might emerge and advance understanding of the four humor styles and the ways in which they are related to psychological well-being. Indeed, as stated by Martin et al. (2003), the four humor styles are closely interrelated, but not equally adaptive for well-being. Using a sample of 318 Australian adults, Galloway (2010) identified four clusters: 1) Those who scored above average on all four humor styles; 2) Those who scored below average on all four; 3) Those who scored above average on the adaptive humor styles and below average on the maladaptive humor styles, and 4) Those who scored above average on the maladaptive humor styles and below average on the adaptive humor styles. Differences between the groups were identified, for example, those in Cluster 1 scored above average for extroversion and openness and below average for conscientiousness and agreeableness. Cluster 3 scored highest for self-esteem and agreeableness, and those in Cluster 4 were below average for self-esteem, extroversion, openness, and agreeableness. However, the study was limited by the small sample size and the cross-sectional nature of the design. A similar study by Leist and Müller (2013) with a German sample identified three, not four clusters: 1) Those who scored above average on all four; 2) Those who scored below average on all four, and 3) Those who scored above average on the adaptive humor styles and below average on the maladaptive humor styles. Those in the third cluster reported higher self-esteem in comparison to the other two clusters. In addition, those in Cluster 2 reported poorer life satisfaction in comparison to the other two clusters. The same limitations that apply to Galloway's study (2010) apply here, namely, the small sample size (N = 348) and the lack of longitudinal data. Moreover the sample was two- thirds female, which is potentially problematic given that gender differences in the four humor styles have been consistently identified (Fox et al., 2013; Führ, 2002; Martin et al., 2003; Saraglou & Scariot, 2002). Thus, the main aim of the current study was to identify humor types in children and improve on some of the limitations identified in previous studies, by using a large sample size and a longitudinal design (drawing on the same data set as used by Fox et al., unpublished manuscript). Differences between the humor types were examined based on measures of psychosocial adjustment: depressive symptoms, loneliness and self-esteem. As well as looking at cross-sectional associations we used residual scores to compare the clusters in relation to changes in adjustment across time. We also looked at the distribution of gender across the clusters.","We recruited 1234 pupils aged 11–13 years (school years 7 and 8; 680 children aged 11–12 years, and 554 children aged 12–13 years), from six state secondary schools in the Midlands, UK. In terms of gender, 599 participants were male and 620 female (with missing data for 15 participants). The mean age of the sample at Time 1 was 11.68 years (SD = 0.64). The ethnic composition of each school (M = 93% white) was a reflection of the region in which the research was located; the sampling strategy took into account both rural/urban and SES profile to achieve a range of schools representative of the area from which they were recruited. Parents or carers of all children in the relevant year group at each school were invited to allow their child to participate, using the opt-out method of consent. Pupils who did not participate in the first session of data collection at Time 1 were not permitted to take part in the second session of data collection at Time 1. Across the time points of the study, the participation rate ranged from 70% to 85% of eligible young people registered in the schools. Participant recruitment and data collection were conducted during school hours. Participants assented to take part in the study during class time. Classes varied in size from 10 to 31 with a modal class size of 24 pupils. Participants who were not taking part completed an alternative activity.","Students completed an answer booklet at each session in which they recorded their name, age, school class, gender and ethnicity, prior to completion of the measures pertinent to that session. Humor styles Participants completed the self-report child Humor Styles Questionnaire (child HSQ; Fox et al., 2013), which is an adapted version of the adult HSQ (Martin et al., 2003). Using a 4-point response scale (1 = strongly disagree to 4 = strongly agree), participants rated their agreement with the 24 statements. There are six items per sub-scale with four sub-scales in total: Self-Defeating (e.g. ‘I often put myself down when I am making jokes or trying to be funny’), Aggressive (e.g. ‘When I tell jokes I'm not worried if it will upset other people’), Affiliative (e.g. ‘I don't have to try very hard to make people laugh — I seem to be a naturally funny person’) and Self-Enhancing (e.g. ‘I find that laughing and joking are good ways to cope with problems’). When used with 11–16 year olds, Fox et al. (2013) found acceptable levels of internal reliability for all four sub-scales (all α > .70), and confirmatory factor analysis identified a very clear four- factor structure. The child HSQ also has acceptable levels of test re-test reliability (rs range from .65 to .75 across one week). For the present study, reliability coefficients were all above .70, apart from aggressive humor at Time 1 (Time 1: αaggressive = .66; αself-defeating = .73; αself-enhancing = .75, αaffilitative = .85; Time 2: αaggressive = .71; αself-defeating = .81; αself- enhancing = .82, αaffilitative = .88). In addition, the four-factor structure was confirmed. Mean scores were calculated for each sub-scale, with higher scores reflecting greater use of that form of humor. Depressive symptoms The 10-item, self-report Children's Depression Inventory — Short Form (Kovacs & Beck, 1977) for ages 7–17 years was administered. For each symptom, participants are required to indicate which of three items best describes them over the preceding two weeks, and responses were scored from 0 (no symptom), 1 (mild symptom), or 2 (moderate/severe symptom). An example item is: “I am sad once in a while”, “I am sad many times”, and “I am sad all the time.” Sum scores were then calculated, with higher scores reflecting the presence of greater symptomatology. This measure showed acceptable internal consistency with the current sample (α = .86, and α = .88, at T1 and T2 respectively). Loneliness This was assessed using the four-item, self-report Loneliness and Social Satisfaction scale (Asher, Hymel, & Renshaw, 1984; Rotenberg, Boulton, & Fox, 2005). A 5-point Likert scale ranging from 1 = not at all true to 5 = really true was used and a mean score was calculated such that higher scores reflected greater loneliness. An example item is ‘I am lonely.’ In the current study, this measure showed acceptable internal consistency (α = .86, and α = .88, at T1 and T2 respectively). Self-esteem Rosenberg's (1965) 10-item, self-report self-esteem measure for adolescents and adults was used with participants judging each item on a 4-point scale from 1 = strongly disagree to 4 = strongly agree. An example item is “I am able to do things as well as most people”. With the current sample, reliability coefficients were α = .87 and α = .89 at T1 and T2 respectively. Sum scores were calculated with higher scores reflecting higher self-esteem. Procedure Prior to data collection, the study was approved by the relevant University Ethics Committee. Data collection took place in the Fall (Time 1) and Summer (Time 2) terms of the school year, in school classrooms with a class teacher present. Data collection took approximately half an hour. A range of other variables were measured but are not the central focus of this paper. Sessions began with the researchers introducing themselves and explaining the measures that would be collected that day, and explaining the confidential nature of the questionnaires. Pupils were asked to complete the questionnaire booklets in silence; they were asked to keep their answers private and not look at what other children were doing. Following data collection pupils were thanked and fully debriefed as to the aims and purpose of the study. Descriptive statistics and intercorrelations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Descriptive statistics for all the Time 1 and Time 2 variables can be seen in Table 1. Table 2 shows the intercorrelations for all the T1 and T2 variables. As might be expected there are significant correlations between the four humor styles and depression, loneliness and self-esteem, at Time 1 and Time 2 and from Time 1 to Time 2. For example, the adaptive styles of humor (affiliative, self-enhancing) are positively correlated with self-esteem and negatively correlated with depression and loneliness. In contrast, self- defeating humor is negatively correlated with self-esteem and positively correlated with depression and loneliness. Cluster analysis ~~~~~~~~~~~~~~~~ We decided to use a series of K-mean cluster analyses as used by Galloway (2010) with an adult sample. When using cluster analysis, the investigator typically has to decide how many clusters to extract from the data. A degree of trial and error is often used to identify the most optimal solution. Normally this is done taking into account the interpretability and parsimony of the different solutions being considered. First, the child HSQ scores were transformed to Z scores to help with the interpretation of the results. A series of K-mean cluster analyses were conducted with the four-cluster solution being the one considered the most appropriate. The Z-scores for the four humor styles within the four clusters can be seen in Table 3 and in Fig. 1. The first cluster, which we have called the ‘Interpersonal Humorists’ includes those who scored above average on aggressive and affiliative humor, but below average on the other two humor styles. Cluster 2 represents the ‘Self-Defeaters’ who scored high on this style of humor, but low on the other three. Thirdly, the ‘Humor Endorsers’ scored above average on all four humor styles, and finally, the ‘Adaptive Humorists’ scored high on the two adaptive styles of humor, but low on aggressive and self-defeating humor. These clusters were replicated when performing a cluster analysis on the Time 2 data (see Table 3). Group differences ~~~~~~~~~~~~~~~~~ A Chi-square analysis to test for gender differences across the four clusters was significant (χ2 = 35.30, df = 3, p < .001). The percentage of males and females in Clusters 1 and 2 were roughly equivalent. However, there were more males then females in Cluster 3 (‘Humor Endorsers’) and more females than males in Cluster 4 (‘Adaptive Humorists’) (Table 4). A series of one-way ANOVAs tested for differences between the groups in terms of the measures of psychosocial adjustment at Time 1 and using standardised residual change scores from Time 1 to Time 2 (by regressing the T2 scores on respective T1 scores). Table 5 shows that the main differences were between Cluster 3 (the ‘Self-Defeaters’) and the other three clusters. Those in Cluster 3 scored highest for depressive symptoms, loneliness, self-esteem and change in loneliness, compared to the other three clusters. In addition, those in Cluster 3 showed a greater change in depressive symptoms, in comparison to the ‘Interpersonal Humorists’ (Cluster 1) and the ‘Adaptive Humorists’ (Cluster 4). For the self-esteem change scores there were no differences between the four clusters. When looking at the cross-sectional data only, some other group differences emerged. The ‘Adaptive Humorists’ (Cluster 1) scored highest for self-esteem in comparison to the ‘Self-Defeaters’ (as above) but also in comparison to the ‘Interpersonal Humorists’ (Cluster 1) and the ‘Humor Endorsers’ (Cluster 3). There was also a significant difference in depressive symptoms between the ‘Adaptive Humorists’ and the ‘Humor Endorsers’. For loneliness, the ‘Humor Endorsers’ (Cluster 3) reported significantly less loneliness then the Self-Defeaters (Cluster 2), but significantly more loneliness than the ‘Adaptive Humorists’ (Cluster 4) and the ‘Interpersonal Humorists’ (Cluster 1) (Table 5).","This is the first study to identify humor types using a large sample of adolescents and a longitudinal design. It builds on the research by Galloway (2010) and Leist and Müller (2013) who were the first to explore combinations of humor styles with adult samples. In the same way, the current study extends knowledge beyond what we already know about humor styles from examination of simple correlations between the four humor styles and measures of psychosocial adjustment. Indeed, our findings suggest that simple correlations can be misleading. Of note is that self-defeating humor, when used in combination with other humor styles, is not necessarily maladaptive. Of all the clusters identified, those in Cluster 2 (the ‘Self-Defeaters’) scored highest in terms of psychosocial maladjustment. They scored highest for depressive symptoms, loneliness, self-esteem and change in loneliness, compared to the other three clusters. For change in depressive symptoms, the ‘Self-Defeaters’ scored the highest in comparison to the ‘Interpersonal Humorists’ (Cluster 1) and the ‘Adaptive Humorists’ (Cluster 4). However, there was no difference between the ‘Self-Defeaters’ and the ‘Humor Endorsers’ (Cluster 3), who appeared to show no change. In contrast, those in Cluster 4 (the ‘Adaptive Humorists’) scored highest for self-esteem in comparison to the three other clusters. They were also less lonely than the ‘Self-Defeaters’ and the ‘Humor Endorsers’; these two groups were also more lonely than the ‘Interpersonal Humorists’. This suggests that the negative effects of using self- defeating humor can be offset to some extent when it is used in combination with other humor styles, but not always. Future research could examine a wider range of adjustment variables over time to explore the similarities and differences between these two groups. Crucially though, the ‘Humor Endorsers’ were not found to be more maladjusted than the ‘Adaptive Humorists’ and the ‘Interpersonal Humorists’ when examining change in adjustment across time. However, future studies could extend this research to see whether any differences do emerge over a longer period of time. In the same way it would be interesting to examine the ‘Interpersonal Humorists’ over a longer time period. Our findings seem to suggest that aggressive humor is not typically used in isolation, and when it is used with affiliative humor, it could be viewed as adaptive, at least for the protagonist. The ‘Interpersonal Humorists’ were less lonely in comparison to the ‘Self- Defeaters’ and ‘Humor Endorsers’, although this effect disappeared when we looked at the data across time. Dyck and Holtzman (2013) noted that of all the four humor styles, aggressive humor has the weakest and most inconsistent findings, with many studies that have used the adult HSQ failing to find significant associations between aggressive humor and measures of psychological adjustment. As argued by Martin (2007), over the long-term, aggressive humor may be detrimental to the self because it tends to alienate others. Studies over a longer time-frame may be able to identify significant associations between aggressive humor types and psychosocial maladjustment. There are some similarities between these findings and those of Galloway (2010) and Leist and Müller (2013). Similar to the present study, both studies identified a group scoring high on all four humor styles and a group who scored above average on the adaptive humor styles, but below average on the maladaptive humor styles. However, these previous studies also identified a group who scored low on all four. In addition, Galloway (2010) identified a group who scored above average on the maladaptive humor styles and below average on the adaptive humor styles. In contrast, for the present study, aggressive humor and affiliative humor were combined, leaving a group of children who scored high on self-defeating humor and low on the other three humor styles. This may be due to the difference in age between the samples, the different cultures and/or it may cast some doubt on the findings of these previous studies which drew on relatively small samples of participants (300–400). There is a need to replicate these humor types using similar samples but in different contexts and countries. The adult HSQ has received cross-cultural validation among European (Saraglou & Scariot, 2002), Chinese (Chen & Martin, 2007) and Arabic samples (Kazarian & Martin, 2004). It demonstrates good psychometric properties and the findings provide support for the four- factor structure. In addition, there is cross-cultural stability in terms of the associations between the four humor styles and measures of psychosocial adjustment; gender differences are also consistent. However, there are differences when looking at the interrelations between the sub-scales. Unlike with North American samples (Martin et al., 2003), there is typically no correlation between affiliative and aggressive humor in other cultures (Chen & Martin, 2007; Kazarian & Martin, 2004; Saraglou & Scariot, 2002). Given that Cluster 1 with the current sample includes those who scored high on affiliative and aggressive humor, it is important that these types are investigated further in other cultures. One of the main limitations of the study is the use of self-reports to measure humor styles and children's psychosocial adjustment, which raises the possibility that many of the associations were confounded by shared method variance. Further research should consider gathering data on children's adjustment from different sources, such as teachers and parents. There is also a need to examine associations between humor types and a wider range of adjustment variables. However, it is not clear that shared method variance necessarily leads to inflated estimates (Conway & Lance, 2010). In addition, the use of humor types to some extent does guard against this claim though because we are not examining simple correlations between positive/negative humor styles and positive/negative aspects of psychosocial adjustment; instead we are looking at differences between humor types, which involve combinations of adaptive and maladaptive humor styles. A potential useful way forward would be to test some of these hypotheses experimentally, similar to the studies by Kuiper, Kirsh, and Leite (2010) and Ziegler-Hill, Besser, and Jett (2013). For example, Ziegler-Hill et al. (2013) provided evidence that adult targets displaying the more benign styles of humor are perceived more positively by others. Similarly, Kuiper et al. (2010) found that both adolescents and adults are less willing to continue an interaction with someone displaying maladaptive humor (i.e. aggressive or self-defeating humor). For example, it would be useful to know how others perceive those who use all four styles of humor, in comparison to those who use only self-defeating humor. Results from an experimental approach such as this could advance our understanding, possibly identifying perceived motives behind the different uses of humor. In conclusion, the current study has extended knowledge beyond what we already know about the simple correlations between the four humor styles and aspects of psychosocial adjustment in children. Furthermore, it is the first study to examine differences in humor types across time. The findings suggest that the negative effects of self-defeating humor could be offset if it is used in combination with other humor styles. However, further research is needed to examine children's humor types in other cultures as well as differences between humor types over a longer period of time."],["Inferences are crucial to successful discourse comprehension. We assessed the contributions of vocabulary and working memory to inference making in children aged 5 and 6. years (n = 44), 7 and 8. years (n = 43), and 9 and 10. years (n = 43). Children listened to short narratives and answered questions to assess local and global coherence inferences after each one. Analysis of variance (ANOVA) confirmed developmental improvements on both types of inference. Although standardized measures of both vocabulary and working memory were correlated with inference making, multiple regression analyses determined that vocabulary was the key predictor. For local coherence inferences, only vocabulary predicted unique variance for the 6- and 8-year-olds; in contrast, none of the variables predicted performance for the 10-year-olds. For global coherence inferences, vocabulary was the only unique predictor for each age group. Mediation analysis confirmed that although working memory was associated with the ability to generate local and global coherence inferences in 6- to 10-year-olds, the effect was mediated by vocabulary. We conclude that vocabulary knowledge supports inference making in two ways: through knowledge of word meanings required to generate inferences and through its contribution to memory processes. --------------------------------------------------------------------------------","Skilled comprehenders make sense of written and spoken language by constructing a coherent memory-based representation of the state of affairs described by the text, commonly referred to as a situation model (Kintsch, 1988). The text does not always explicitly state all of the information needed for coherence. Therefore, readers and listeners regularly make inferences to integrate information within the text and to fill in details that are only implicit. These inferences are incorporated into the situation model that the comprehender constructs and result in a more accurate and complete understanding of the text (Graesser, Singer, & Trabasso, 1994). Although children make inferences from an early age, they do not typically make as many as do older children and adults (e.g., Ackerman, 1986; Casteel, 1993). Our focus in this study was to understand better why this is the case by examining the role played by two critical factors related to young children’s inference generation: vocabulary and working memory. Furthermore, we explored whether these factors make different contributions to distinct types of inference. A greater understanding of the factors that support early inference making will inform models of comprehension development and also the literacy curriculum and targeted intervention programs for children with weak inference-making and comprehension skills. When considering different types of inference, one of the key distinctions that can be made is between inferences that establish local coherence and those that establish global coherence (Graesser et al., 1994). Inferences that are necessary for local coherence typically involve the integration of separate propositions within the text and are usually cued by a pronoun, synonym, or category exemplar; for example, “He finished the orange juice quickly. The drink was very refreshing.” Local coherence inferences often require a mapping between related words, for example, between synonyms or category exemplar pairings as in the example above. In contrast, inferences necessary for global coherence can involve inferring goals that motivate particular actions or establishing an overall theme of a text, and this often relies on the ability to connect ideas that are not explicitly signaled by a single word and that can be distributed throughout the text. For example, readers can infer the likely setting of a story through links between semantically related concepts such as “building sandcastles,” “paddling in the water,” and the presence of a “pier” (which together indicate that the setting is the seaside). In this way, inferences that are necessary for global coherence typically draw more heavily on information that is external to the text than do local coherence inferences (Cain & Oakhill, 1999, in press). We note that it is likely that local and global coherence inferences are not truly categorical distinctions and that, rather, different inferences draw on information in the text and background knowledge to differing degrees (Florit, Roch, & Levorato, 2011). Furthermore, some authors distinguish local and global coherence inferences in terms of distance between elements of the text (e.g., McKoon & Ratcliff, 1992), but we did not manipulate this feature of our texts. Critically, both local and global coherence inferences are necessary for a full understanding of the text, in line with Cain and Oakhill’s (1999) study of children and Long and Chong’s (2001) study of adults. Adults routinely make inferences to establish local and global coherence (e.g., Albrecht & O’Brien, 1993; Bloom, Fletcher, van den Broek, Reitz, & Shapiro, 1990; Nicol & Swinney, 1989; Sanders & Noordman, 2000). With regard to local coherence inferences, children are capable of making them from an early age, but developmental improvements are clear. Ackerman (1986) found that 6-year-olds were as sensitive as 9-year-olds to the need to establish local coherence in short texts but made fewer local coherence inferences than did 9-year-olds when required to support comprehension. Barnes, Dennis, and Haefele- Kalvaitis (1996) demonstrated substantial developmental improvements in the ability to make local coherence inferences between 6 and 11 years of age, with smaller gains thereafter up to 15 years (the oldest age group in their study). Likewise, research examining children’s ability to generate inferences to establish global coherence demonstrates early sensitivity as well as developmental gains, with 4-year-olds making fewer inferences about narrative themes and characters’ goals than 6-year-olds (Lynch et al., 2008). These studies provide clear evidence that, despite early sensitivity to the need to establish both local and global coherence when processing text, significant gains in inference-making ability occur during early to middle childhood. When we examine the wider literature on inference making in adults, as well as in children, we find evidence that vocabulary and working memory influence an individual’s inference-making ability. We next turn our attention to a consideration of these skills and how they might be related to developmental improvements in inference making. Obviously, text comprehension could not occur without knowledge of individual word meanings, and for that reason vocabulary is routinely shown to be related to general measures of reading (Oakhill & Cain, 2012) and listening comprehension (Florit, Roch, Altoe, & Levorato, 2009). Vocabulary is also significantly correlated with measures of inference making in both children (Lynch et al., 2008; Oakhill & Cain, 2012) and adults (Dixon, LeFevre, & Twilley, 1988). For local coherence inferences, knowledge of word meanings may be particularly important when a synonym, paraphrase, or category member refers back to an earlier mentioned object (Perfetti, Yang, & Schmalhofer, 2008). In addition to knowledge of specific word meanings, the production of some inferences also draws heavily on background knowledge and the interrelations between words (Cain & Oakhill, 2014; Casteel, 1993). In the case of global coherence inferences, one cannot infer that a furry animal that barks and likes going for walks is a dog unless one possesses the requisite knowledge about dogs and their characteristics. Thus, not just knowledge of individual word meanings but also rich semantic networks with robust connections between the meanings of words associated by topic may be important for easy and accurate inference making. Therefore, there are clear mechanisms through which vocabulary may support both local and global coherence inference making. Working memory refers to the memory systems used for the simultaneous storage and processing of information (Baddeley & Hitch, 1974). Measures of verbal working memory (sometimes referred to as verbal complex memory span) are more strongly related to reading comprehension in young children than are measures that simply tap storage of verbal information (or short-term span) (Leather & Henry, 1994). Furthermore, working memory explains unique variance in general measures of listening and reading comprehension after controlling for vocabulary in children between 4 and 10 years of age (Cain, Oakhill, & Bryant, 2004; Florit et al., 2009; Seigneuric & Ehrlich, 2005). Working memory may be particularly important for inference generation because a reader or listener needs to maintain activation of previously processed information while relating this to the piece of text currently being processed. Younger children might routinely fail to do so if their working memory capacity limits the amount of information that they can store when processing text. This may account for the observation that younger children tend to process text in a piecemeal manner, not always making connections between ideas, particularly if they are not presented in succession (Schmidt & Paris, 1983). However, studies that have included assessments of working memory and either inference making or more general reading comprehension within a developmental framework report contradictory results. Chrysochoou and Bablekou (2010) found a reduction in the influence of working memory on inference making between 5 and 9 years of age, whereas other work suggests that the relation between working memory and reading comprehension either stays constant between 7 and 10 years of age (Seigneuric & Ehrlich, 2005) or increases during that period (Cain et al., 2004). One factor influencing the strength of any relation between inference and working memory may be the nature of the materials; the influence of working memory on inference making for 9- and 10-year-olds is strongest when the pieces of information supporting the inference need to be integrated over large units of text (Cain, Oakhill, & Lemmon, 2004). In the current study, we were interested in the contributions made by both vocabulary and working memory to local and global coherence inferences in different age groups. Recent work supports this focus by showing that from as young as 4 to 6 years, higher level comprehension skills such as text integration (particularly relevant for local coherence inferences) and knowledge accessibility (important for both inference types but particularly for global coherence inferences) emerge as separate skills (Hannon & Frias, 2012). To date, there is only one published study that has explored the relative influences of both working memory and vocabulary to different types of inference making (Chrysochoou, Bablekou, & Tsigilis, 2011). The two inference types studied by Chrysochoou and colleagues (2011) were required to establish coherence in the text. They broadly map onto the local and global coherence distinction that we focused on here. We note that the authors referred to the latter as elaborative inferences, but on examination these appear to be necessary to ensure a full and coherent understanding of the text as intended by the authors of the original article (Cain & Oakhill, 1999). Chrysochoou and colleagues (2011) found small, but significant, associations between complex working memory span and inferences that were required for local coherence in their sample of 9-year-olds. However, they found a much stronger association between working memory and inferences required to establish global coherence. The relations between their measures of short-term span (e.g., measures of the phonological loop) and both types of inference were weak. Critically, Chrysochoou and colleagues (2011) demonstrated that vocabulary fully mediated the relations between working memory and the inferences required for local coherence but only partially mediated the relations between working memory and the inferences required to establish global coherence. One possible reason for this distinction is the way in which vocabulary can support memory. The theory is that children and adults with richer vocabulary knowledge have more accurate and available representations of words in their long-term memory than those with poorer vocabulary knowledge, and this better knowledge supports accurate maintenance of information in verbal working memory (Nation, Adams, Bowyer-Crane, & Snowling, 1999; Walker & Hulme, 1999). Thus, vocabulary may be important for inference making in two ways. First, vocabulary knowledge is important because inferences involve word knowledge; local coherence inferences involve mapping between synonyms and category exemplars, and global coherence inferences tap knowledge about the interrelations between word meanings. Second, vocabulary knowledge may support inference making because it can provide a boost to accurately maintain the contents of working memory, necessary to aid the integration of information from different parts of the text. In relation to the findings of Chrysochoou and colleagues (2011), the full mediation of working memory by vocabulary for local coherence inferences may reflect the lower memory demands associated with this type of inference, which involve integration between successive sentences, and the partial mediation for global coherence inferences may reflect the higher memory demands of global coherence inferences, which involve integration of ideas throughout the text and also with background knowledge external to the text. Another possibility for this distinction is that Chrysochoou and colleagues (2011) used a single measure of receptive vocabulary, which is regarded by some as a measure of the number of words known and which does not tap deeper knowledge of the interrelations between words (Tannenbaum, Torgesen, & Wagner, 2006). We assessed vocabulary in two ways to achieve a more complete assessment of this complex construct; we used the same single word comprehension measure as Chrysochoou and colleagues and also a measure that assessed knowledge of word networks (semantic fluency). We believed that this broader conceptualization of vocabulary knowledge would provide a robust test of its relation to inference making. The current study sought to build on previous research on children’s inference making by looking specifically at the distinction between inferences required to establish local and global coherence (a distinction that has not been examined in previous work with younger age groups of children) with the aim to determine how vocabulary and working memory influence this ability in 6- to 10-year-olds. Our review of the literature indicates that both vocabulary and working memory will be associated with inference making in this age range. We predicted that vocabulary would be more strongly associated with global coherence inferences than with local coherence inferences because the former rely on richer and better connected semantic networks, in line with Cain and Oakhill (2014). We also predicted that memory would be more strongly associated with global coherence inferences than with local coherence inferences because global coherence inferences involve maintaining activation of a larger amount of text. Regardless of their age, all children in our sample were presented with the same texts, which were read aloud to them by the assessor. It is widely believed that the same skills underpin comprehension of written and spoken text (Hoover & Gough, 1990). This is supported by research demonstrating the role of inference for each (Barnes et al., 1996) and strong relations between performance on reading and listening comprehension tasks when word reading ability is suitably controlled (Cain, Oakhill, & Bryant, 2000; Kendeou, van den Broek, White, & Lynch, 2009; Stothard & Hulme, 1992). Because we used the same texts, we expected to find developmental improvements in inference making and the strongest relation between memory capacity and inference making in the youngest age group, where the processing demands would be greatest. In general, we expected stronger associations between inference making and measures of complex memory span (tasks that tap both storage and processing) than between inference making and short-term memory span measures (which tap only storage), similar to the results found for 8- and 9-year-olds by Chrysochoou and colleagues (2011). There are contradictory findings for the roles of vocabulary and memory; studies of 8- to 11-year-olds find a unique role for working memory in the prediction of concurrent reading comprehension that is independent of vocabulary (Cain et al., 2004; Seigneuric & Ehrlich, 2005), whereas other work suggests that working memory has a unique role in the prediction of reading comprehension, over and above vocabulary knowledge, for 5- and 7-year-olds but not for 9-year-olds (Chrysochoou & Bablekou, 2010; Florit et al., 2009). A key aim of this study was to determine whether vocabulary and working memory predicted unique variance in inference making or whether any relation between working memory and inference making was mediated by vocabulary as found by Chrysochoou and colleagues (2011; see also Dixon et al., 1988, for work with adults). This finding is in contrast to studies of discourse comprehension that identify a unique contribution for working memory on reading and listening comprehension independent of vocabulary (Cain et al., 2004; Florit et al., 2009; Seigneuric & Ehrlich, 2005). Full mediation would indicate that the relation between working memory and inference making is due to the support that vocabulary provides for maintaining activation of relevant information in memory. Thus, our analyses were constructed to determine whether this relation exists across all ages for both types of inference.","The participants were 130 children from schools in the northwest of England. Of this total sample, 44 children were enrolled in Year 1 classrooms (5–6 years of age, M = 6;2 [years;months], SD = 4 months, range = 5;7–6;8; 26 boys and 18 girls), 43 were from Year 3 classrooms (7–8 years of age, M = 8;3, SD = 3 months, range = 7;9–8;8; 21 boys and 22 girls), and 43 were from Year 5 classrooms (9–10 years of age, M = 10;2, SD = 4 months, range = 9;8–10;8; 24 boys and 19 girls). Consent was obtained from headteachers and parents, and assent was received from children, prior to each assessment session. The schools served socially mixed catchment areas. All children spoke British English as their first language, and children with a statement of special educational needs did not take part in the study.","All children completed assessments of inference making, vocabulary, and memory. They were assessed individually over four sessions, with each session lasting no longer than 15 min. Local and global coherence inferences Each child listened to four short stories modified from materials developed for another project (Language and Reading Research Consortium, in press) that were read aloud by the experimenter. The stories were 148 to 161 words in length. Each story had an episodic structure that began with a setting (to introduce characters, time, and/or place) and had categories of story units, including events, goals, attempts, outcomes, and/or reactions (Stein & Glenn, 1982). The Coh-Metrix Text Easability Assessor (Graesser, McNamara, & Kulikowich, 2011) metrics showed that the texts had high narrativity (the extent to which the text is story-like, M = 63%), had a high number of concrete and imageable content words (M = 94%), and were syntactically simple (M = 94%) but had low referential cohesion (M = 35%). The latter was to be expected because the texts were constructed to require inferential processing for adequate comprehension. Children were instructed to listen carefully to the stories so that they could answer the memory questions after each one. At the end of each story, children were asked four questions that tapped the ability to make local coherence inferences and four questions that tapped global coherence inferences. For each story, the four local coherence inferences required the listener to integrate information from two sentences in the text and the four global coherence inferences required the listener to understand motivations and infer themes, settings, or central character identity. Examples are provided in Table 1. A second rater categorized each of the 32 questions. There was 94% agreement for the local versus global coherence distinction. The two questions categorized differently were resolved through discussion. Cronbach’s alpha on this task for this sample of children was adequate (α = .77). The order of the questions followed the order of information presentation in the story.1 If a child gave an incomplete or vague answer, the experimenter prompted the child. This was either a repetition of the question or encouragement to be more specific (see Table 1 for examples). Each inference question was scored as correct, correct after a prompt, or incorrect. Correct responses were awarded 2 points for a full response and 1 point for a partial response. When a child did not provide an acceptable response to a global coherence question, the child’s background knowledge for that information was checked with a follow-up question (see Table 1). Although both local and global inferences draw on relevant background knowledge, the global coherence inferences examined in the current study were more knowledge dependent in that they drew on a range of facts to work out a setting or a character’s identity. Therefore, we made additional checks that children possessed the necessary background knowledge for this inference type so that we could determine whether knowledge differences were the source of any developmental differences found. There were very few instances where children did not have the required background knowledge; only one child in each year group answered one background knowledge question incorrectly (<0.3%). An adjusted total score was calculated taking into account only those items for which the child possessed the relevant background knowledge. This was a proportionate score based on the child’s actual score and the maximum possible score (for only those items where the background knowledge was known). To enable ease of comparison, the raw correct scores for the local coherence questions are reported as proportions correct. Vocabulary Each child completed two measures of vocabulary, one that tapped breadth of vocabulary knowledge and one that tapped depth of vocabulary knowledge, in order to obtain a more complete assessment of this construct. The British Picture Vocabulary Scale–Third Edition (BPVS; Dunn et al., 2009) provided a measure of breadth of vocabulary knowledge. Each child was asked to select one of four pictures that best showed the meaning of a word spoken aloud by the experimenter. The test was administered and scored according to the guidelines in the manual. Internal consistency was not reported in the manual. Therefore, Cronbach’s alpha was calculated for this sample of children (α = .97). The Word Associations subtest from the Clinical Evaluation of Language Fundamentals–Fourth Edition (CELF-IV; Semel, Wiig, & Secord, 2006) tapped semantic fluency and depth of vocabulary knowledge. In each trial, the child was asked to provide as many words as possible in 1 min from a specified category (items of clothing that people wear, animals, foods that people eat, or jobs that people do). The first category was a practice trial. The test was administered and scored according to the manual. We could not calculate internal consistency for this task, but the test–retest reliability value reported in the manual for this age range is good (.96–1.00). Working memory measures Each child completed four assessments of verbal working memory. There were two measures of the ability to store information (simple span) and two measures of the ability to store and process information (complex span). Internal consistency was not reported in the manual for the working memory measures, so it was calculated for the current sample. Simple span measures Each child completed the Word List Recall and Digit Recall subtests from the Working Memory Test Battery for Children (WMTB-C; Pickering & Gathercole, 2001). In these tasks, children were asked to recall either lists of words or strings of digits spoken by the experimenter. The tasks were administered and raw scores were calculated according to the test manual. There were six trials at each level of difficulty. Once a child had correctly completed four of the trials on a level, the assessor moved on to the next level, as directed by the manual. The child received credit for any trials not completed on the previous level. The assessor stopped testing once three trials on a level were answered incorrectly. Raw scores (total number of trials correct, including those given credit for as a result of moving on to the next level) and standardized scores are presented in Results. Cronbach’s alpha based on the current sample of children was good (α = .85) for the word span and excellent (α = .90) for the digit span. Complex working memory span Each child completed the Listening Recall and Counting Recall subtests from the WMTB-C (Pickering & Gathercole, 2001). In the Listening Recall task, children listened to short sentences, made a true/false judgment about each one, and then recalled the final word in each one. This test began with one sentence and increased in difficulty by the addition of sentences in each set (two sentences, three sentences, etc.). In the Counting Recall task, children counted the number of dots on a page and recalled the total. The test increased in difficulty in that children were required to count more than one pattern of dots over successive pages and then to recall the totals in the correct order. Both tests were administered according to the manual and followed the same progression and discontinuation rules as described above. Cronbach’s alpha based on the current sample was good (α = .88) for both the listening span and counting span (α = .90).","The results are reported in two sections. The first section concerns developmental comparisons of performance on all tasks. The second section concerns analyses that explore the interrelations between vocabulary, working memory, and inference performance. Before analysis, the “Q–Q plots” for all measures were examined separately for each year group. Q–Q plots provide a means of assessing deviation from a normal distribution (Cohen, Cohen, West, & Aiken, 2003). The plots indicated that the data were normally distributed. The data distributions were also checked for outliers. All analyses were run on the original dataset and also on a dataset that was adjusted to remove outliers following the recommendations of Tabachnick and Fidell (2007), whereby outlier data points are changed to the next highest/lowest (non-outlier) number. Less than 1% of datapoints were replaced in this way. The patterns of analyses were the same for both analyses, and so the analyses reported here were those conducted on the original data. Developmental comparisons on all measures ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ First we report the performance on the inference task, followed by performance on the vocabulary and memory measures. Performance on local and global inference questions The proportion of total correct items (adjusted to take general knowledge into account, as described above) for local and global coherence inferences were the dependent variables in the analyses reported in this section. Before analysis, the data were arcsine transformed, as is recommended to stabilize variance and normalize proportional data (Sheskin, 2003). The pattern for the analyses of the raw scores was the same as that of the arcsine-transformed scores, so the analyses reported here are those conducted on the raw proportionate scores. Two analyses are reported below. In the first analysis, the dependent variable is performance when the question was first asked. In the second analysis, the dependent variable is performance that includes correct responses provided after a prompt. The mean proportions of correct responses on the local coherence and global coherence inference questions are shown in Table 2. Performance on local and global inference questions: First responses A mixed factor ANOVA with age (6, 8, or 10 years) as a between-participants factor and inference type (local or global) as a within-participants factor was conducted. There was a significant main effect of age, F(2, 127) = 14.78, p < .001, ηp2 = .19. Post hoc Tukey tests (p < .05) indicated that the 6-year- olds obtained significantly lower scores than the 8- and 10-year-olds and that the 8-year-olds obtained significantly lower scores than the 10-year-olds (Ms = .58, .65, and .74 in order of increasing ages). There was also a significant main effect of inference type, F(1, 127) = 106.89, p < .001, ηp2 = .46, because the children performed best on the global inference questions (Mlocal = .58, Mglobal = .73). The interaction was not significant, F(2, 127) < 1.0, ns. Performance on local and global inference questions: After prompts As is evident from Table 2, performance improved with the use of prompts. The difference between unprompted and prompted scores (collapsed across age group) was significant for both types of inference: local, t(129) = 15.72; global, t(129) = 12.08; ps < .001. A mixed factor ANOVA with the same design as before was conducted on the scores that included prompted responses. The pattern of results was the same as for the analysis of unprompted scores, with main effects of age, F(1, 127) = 14.21, p < .001, ηp2 = .18, and inference type, F(1, 127) = 127.82, p < .001, ηp2 = .50, and no interaction between the factors, F(2, 127) < 1.0, ns. Performance on vocabulary and working memory measures The mean raw scores (total number of items correct for the vocabulary measures and total number of trials correct for the word, digit, listening, and counting working memory measures) for each age group are reported in Table 3. The mean standardized scores were available for all measures with the exception of the Word Associations task, for which standardized scores are not reported in the manual. The standardized scores were all within an age-appropriate range. Criterion reference scores were available for each age group on the Word Associations task, and all children met the minimum criterion score for their age. A small number of children (n = 10) had missing data due to absence at a testing session. A one-way ANOVA on the raw scores for each measure was conducted with age group as a between-participants factor. As predicted, there were age differences for all measures shown by significant F values (see Table 3; all ps < .001). Post hoc Tukey tests (p < .05) revealed that for most measures there were significant differences in performance between each successive age group in the following order: 6 < 8 < 10 years. The two exceptions were the Word Associations and Counting Recall tasks. Here, both the 8- and 10-year-olds obtained significantly higher scores than the 6-year-olds, but these two older groups did not differ from each other. The 6-year-olds performed poorly on both complex span measures of working memory, with 27.3% of these children not passing the first level of the Listening Recall task and a further 61.4% not progressing beyond the one-span level. For the Counting Span task, 22.7% did not progress beyond the one-span level. This is in line with Gathercole, Pickering, Ambridge, and Wearing’s (2004) observation that the task demands for these complex span measures are too high for very young children and that these complex span measures might not be sensitive to individual differences. On that basis, the complex span tasks were not included in further analyses for this age group. Correlations between inference, vocabulary, and working memory ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ First, zero-order correlations were conducted for each age group separately to examine the interrelations between variables. Composite scores were produced to ensure a broad and comprehensive indicator of the vocabulary and memory constructs: vocabulary (BPVS and Word Associations), simple memory span (Word List Recall and Digit Recall), and complex memory span (Listening Recall and Counting Recall). The working memory composites support the distinction between tasks tapping short-term storage in the phonological loop and those tapping the central executive (Pickering & Gathercole, 2001). Reducing the number of measures to form composites also met the requirement of 10 data points per predictor for the multiple regression analyses reported next (as recommended by Tabachnick & Fidell, 2007). The scores on the first response to the inference questions were used. All scores were converted to z scores (for each age group separately) to ensure that measures were on a comparable scale (see Table 4). A small number of children (n = 10) had missing data on one of the two subtests used to produce the composites. In these cases, the mean performance on the missing subtest for the relevant age group was used so that a composite could still be calculated for each of these children. This amounted to 18 data points in total (2.36% of all data points).2 Because of the small sample size, we discuss correlations both in terms of statistical significance and in relation to effect sizes where .10 represents a small effect, .30 a moderate effect, and .50 a large effect (Field, 2005). In all age groups, the two inference measures were significantly correlated and the effects were moderate to large. Of note, vocabulary was more strongly related to inference making than was memory. The vocabulary composite was significantly correlated with local coherence inference making for the 6- and 8-year-olds and with global coherence inference making for all age groups. All effect sizes were large. For the youngest age group, simple span was significantly correlated with local coherence inference making, with a moderate effect size, but a significant relation was not found for the older age groups. The simple and complex working memory measures (where used) were significantly correlated with global coherence inference making for all age groups, where the effects were moderate (with the exception of simple span for the 8-year-olds, where the effect was small and did not reach significance). Do working memory and vocabulary each explain independent variance in inference making in different age groups? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A series of multiple regression analyses were conducted to determine the relative contributions of memory and vocabulary to the two types of inference. Fixed-order hierarchical multiple regression analyses were conducted for each inference type and each age group separately. Vocabulary, simple span, and complex span (8- and 10-year-olds only) were entered as predictors in the analyses. Simple span was added at the first step, followed by complex span (8- and 10-year-olds only) and vocabulary. Previous research has found that vocabulary mediates the relationship between memory and inference generation (Chrysochoou et al., 2011); therefore, the influence of the memory measures before the inclusion of vocabulary could be determined using this order of variables. In addition, complex span tasks have been found to be more highly related to reading comprehension than simple span tasks (Daneman & Merikle, 1996). Therefore, this order of memory variables also ensured that any influence of simple span could be identified. Table 5 shows the results for the multiple regression analyses for the local coherence inferences. For local coherence inferences, simple span and vocabulary explained variance in performance for the 6-year-olds, although (as indicated by the beta values) only vocabulary was a significant predictor when both variables were included in the model (see Table 5). For the 8-year- olds, only vocabulary significantly accounted for significant variance in local coherence inference making. In contrast, for the 10-year-olds, none of the measures significantly explained variance. For global coherence inferences, simple span and vocabulary explained additional variance in performance for the 6-year-olds, although (similar to the local coherence analyses) only vocabulary predicted unique variance when both variables were taken into account (see Table 6). For the 8-year-olds, complex span and vocabulary significantly accounted for additional variance in global coherence performance, although (again) only vocabulary was a significant predictor when all variables were taken into account. For the 10-year-olds, simple span and vocabulary explained additional variance in global coherence performance, but only vocabulary significantly predicted unique variance. Does vocabulary mediate the influence of working memory on inference making in different age groups? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As outlined in the Introduction, the literature indicates that vocabulary may mediate the influence of memory on inference making. The data in Tables 5 and 6 demonstrate that both vocabulary and short-term memory made significant contributions to local inference making in the 6-year-olds and that vocabulary and memory (short-term span for 6- and 10-year-olds and complex span for 8-year-olds) predicted global inference making in all age groups. The analyses reported next assess whether or not vocabulary mediated the influence of memory on inference making. The total effect of working memory on inference making is presented in Fig. 1, the mediation path assessed is presented in Fig. 2, and the labels used to denote each path are also referred to in Tables 7 and 8. A set of mediation analyses was conducted based on the causal steps strategy (see Baron & Kenny, 1986) and Preacher and Hayes’s (2008) bootstrap (1000) resample procedure using an SPSS macro for this procedure, which gives point estimates (PEs) and bias-corrected confidence intervals (BC CIs) for the indirect effect. Where zero is not included in the 95% confidence interval for the mediated effect, the effect is statistically different from zero and mediation is confirmed. For local coherence inferences, simple span predicted significant variance in local coherence inference making for the 6-year-olds (see Table 5). The mediation analyses demonstrated that vocabulary mediated its effect; the effect of working memory on inference making (c′ path) was not significant after taking vocabulary into account. Critically, mediation was confirmed because zero was not included in the 95% confidence interval for the mediated effect (see Table 7). Mediation analyses were not conducted for the 8- and 10-year-olds because the working memory measures did not predict unique variance in local coherence scores (see Table 5). For global coherence inferences, vocabulary mediated the relationship between working memory and inference for all age groups. As shown in Table 6, although simple span predicted inference for the 6- and 10-year-olds, the effects of working memory on inference were mediated by vocabulary at both ages because the path coefficient for the regression of this working memory measure on inference was no longer significant when vocabulary was entered into the equation (c′ path). Mediation was confirmed because zero was not included in the 95% confidence interval for the mediated effect (see Table 8). For the 8-year-olds, complex span was related to global coherence inference making (see Table 6). However, as for the analysis for the two other age groups above, vocabulary mediated this effect because the path coefficient for the regression of complex span on inference was no longer significant when vocabulary was entered into the equation (c′ path). The mediated effect of complex span via vocabulary was significant because zero was not included in the 95% confidence interval (see Table 8).","The main aims of this study were to explore developmental differences in local and global coherence inference making and to determine the extent to which vocabulary and working memory influenced performance. In line with expectations and previous research, the 6-year-olds were able to generate these inferences from short narrative text, but significant improvements were found with age. Critically, our work indicates that vocabulary and working memory differed in the extent to which they predicted performance on both types of inference between age groups. Specifically, there was evidence that vocabulary knowledge was critical both independently and through its role in supporting the memory processes required for global inference making in particular. First, let us consider performance on the local coherence inferences. In line with previous research, performance improved with age (Ackerman, 1986). To integrate information between sentences in order to establish local coherence, children needed to understand and make use of pronouns and also synonyms. The previous literature on pronoun and synonym comprehension has focused on individual differences in relation to measures of reading comprehension skill in general. This body of work finds that children who differ in reading comprehension skill differ in their understanding and use of pronouns and synonyms to integrate different propositions in a text (Ehrlich & Remond, 1997; Megherbi & Ehrlich, 2005; Oakhill, 1983; Oakhill & Yuill, 1986; see also Cain & Nash, 2011, for developmental improvements in knowledge of other cohesive devices). We propose that poorer use of these signaling cues may have limited local coherence inference making for the children in our study and could, in part, explain the developmental differences. The source of the poorer use of these signals appears to be different for the 6-, 8-, and 10-year-olds. Vocabulary was a particularly important predictor of local coherence inference making for the two youngest age groups, a finding supported by previous research on inference making (Chrysochoou & Bablekou, 2010; Chrysochoou et al., 2011) and reading comprehension in general (Seigneuric & Ehrlich, 2005). The current study is the first to show that vocabulary knowledge is specifically related to 6- to 10-year-olds’ ability to make inferences to establish local coherence. We cannot determine whether children lacked the category-relevant knowledge, failed to activate this, and/or failed to use this information during text processing from our data. However, we believe that a lack of relevant knowledge per se is an unlikely source of these developmental differences because few children failed the background knowledge check questions for the global inference measures. Thus, we propose that speed of knowledge retrieval and integration during text presentation are more likely sources of failure to establish local coherence (as proposed by Perfetti et al., 2008, for adults). We are currently exploring online text processing in these age groups to examine these possibilities. Although the zero-order correlations and regression analyses indicated that working memory played a role in the determination of the youngest children’s local coherence inference performance, working memory was not a unique predictor after taking vocabulary knowledge into account. The youngest children’s poorer memory skills may have limited their ability to retain the information from two successive sentences in order to integrate their meanings when constructing their situation model of the text. However, the mediation analysis confirmed that for the 6-year-olds the influence of memory was fully mediated by vocabulary, similar to the findings in other work with older children (Chrysochoou et al., 2011). Children with richer vocabulary knowledge will be better able to retain information in verbal working memory (Cain, 2006; Nation et al., 1999), which supports our finding of a mediation effect. A complex interaction between vocabulary and verbal working memory may underpin developmental differences in this type of inference making, and further work is needed to test the proposal that speed of knowledge retrieval and/or maintenance of accurate representations of meaning were the source of local coherence inference difficulties. In contrast to the pattern of prediction for the youngest age groups, neither vocabulary nor working memory explained performance on the local coherence inferences for the oldest age group. The older children had superior memory and vocabulary skills, which may have been sufficient for them to establish local coherence. However, the 10-year-olds were not at ceiling on the task, so this explanation is unsatisfactory. One possibility is that they had sufficient vocabulary and working memory skills to perform the task, but additional gains on these questions could come from strategic processing. For example, older children may have greater awareness of which information in a text is critical for the generation of local and global inferences. We are currently investigating developmental differences in children’s use of text information in the generation of both inference types to explore this issue further. Work with poor comprehenders (Yuill & Oakhill, 1988) also shows that children can be taught how to make inferences and that this boosts both inference making skill and reading comprehension. We now turn to a discussion of the findings on global coherence inferences. As predicted, and in line with other work, there were developmental improvements on the global coherence inferences (Chrysochoou & Bablekou, 2010; Schmidt & Paris, 1983). It has been suggested that younger children tend to view statements in a text as independent pieces of information and do not always link successive ideas together, particularly if they are not presented in succession (Schmidt & Paris, 1983). The local coherence inferences were signaled, indicating when successive pieces of text should be linked. In contrast, the global coherence inferences were not signaled and typically required the listener to link information across a number of different nonconsecutive sentences. Thus, one reason for failing to generate these inferences is a piecemeal processing style. Another reason is the processing demands of this type of inference. The global coherence inferences required the integration of information presented in different nonconsecutive sentences. The association between working memory and performance on this type of inference probably arose because of this demand on memory. Age differences in global coherence inference ability persisted even when differences in the necessary background knowledge were controlled (indeed, there were only three instances where a child failed to correctly answer a global coherence inference and lacked the necessary knowledge). Thus, this study adds to a growing body of work demonstrating that the ability to make coherence inferences is not solely dependent on having the relevant background knowledge (Barnes et al., 1996; Cain & Oakhill, 1999; Cain, Oakhill, Barnes, & Bryant, 2001). The findings support the idea that children need to know when and how to draw on background knowledge during text comprehension and that this ability might improve with age (Cain & Oakhill, 1999). Our work advances our understanding of the development of inference making by demonstrating that at all three ages, vocabulary was the most important factor explaining performance on the global coherence measures and vocabulary fully mediated the relationship between global coherence inference making and working memory. This finding is supported by previous work assessing this inference type with older children (Chrysochoou et al., 2011). Thus, although children demonstrated relevant background knowledge when assessed, they did not always use this to generate or encode the inference. Furthermore, the information to be linked was typically presented in words that were semantically associated; for example, “pet,” “furry,” “playful,” and “kennel” were all used to indicate that the story was about a dog. It may be that, for some children, the associations between these links were not sufficiently strong to activate the concept to which they were all related. The developmental improvements found in vocabulary knowledge and the richness of semantic networks, in which the associations between co-occurring concepts are encoded (Metzger et al., 2008), may assist in inference making that relies on background knowledge. Future work, using online measures to assess whether children use successive pieces of information, may help to further our understanding of the production of this inference type. In this study, children made a greater number of global coherence inferences than local coherence inferences, in contrast to other work that has compared the generation of these inference types during reading (Cain & Oakhill, 1999). Although the questions tap different elements of the story and, therefore, are not directly comparable, this finding requires consideration. One explanation could be that the global coherence inferences (particularly those that were thematic in nature, e.g., the setting of a story or a character’s identity) were more foregrounded because the text made a number of references to these elements. Support for this explanation comes from work on the centrality effect, which finds that information that is more central to the overall story meaning is more likely to be remembered than information that is peripheral to overall story meaning (e.g., see Albrecht & O’Brien, 1993; Miller & Keenan, 2009). Future work comparing these inference types should consider this difference. There are several educational implications that stem from this work. First, performance for all age groups on both types of inference improved after prompting underspecified responses. This finding identifies prompting as a useful means to demonstrate a child’s potential (or competence) and also as a tool to indicate the standard of coherence required of comprehension (van den Broek, Risden, & Husebye-Hartman, 1995). Second, the finding that all age groups made both global and local coherence inferences indicates that children of this age engage in the processing required to construct a representation of the entire text rather than engaging in piecemeal line-by- line processing. Third, the importance of vocabulary knowledge to inference making and the way in which vocabulary mediated the relations between working memory and inference making supports the critical role of background knowledge in text comprehension (Eason, Goldberg, Young, Geist, & Cutting, 2012). It also lends weight to other research proposing that children need to be taught how to use both the text and background knowledge as sources of information to guide inference generation and text comprehension (Brandão & Oakhill, 2005; Raphael & Wonnacott, 1985). There are several limitations to this work that need to be addressed in future work in order to understand fully the roles of vocabulary and working memory in children’s inference making. First, although our sample size had adequate statistical power to detect the contributions of key skills, replication with a larger sample size is required to confirm the developmental differences. Second, we did not include independent assessments of strategic knowledge, so we do not know to what extent different age groups were aware of different sources of knowledge and strategies that may influence language comprehension (Brandão & Oakhill, 2005). Third, although we drew on a large literature indicating that the same processes underpin the comprehension of written and spoken text, we note the evidence for a greater influence of attention on listening comprehension (Cain & Bignell, 2014). Future work should consider the extent to which this might specifically affect inference making. Finally, all of our measures were offline. Language comprehension is a dynamic process. Thus, it is essential to investigate how both vocabulary and memory (both short-term and working memory) interact during online language processing to identify when and where they influence the inference-making process. We are pursuing this line of inquiry in ongoing work. In summary, few studies have experimentally contrasted children’s ability to generate different types of inference, and none has developmentally explored the unique roles of vocabulary and working memory in local and global inference making. This study has highlighted different roles for vocabulary and working memory in relation to these two types of inference at different ages. This not only adds to our understanding of inference development but also indicates factors that may limit language comprehension in the classroom."],["De Villiers (Lingua, 2007, Vol. 117, pp. 1858–1878) and others have claimed that children come to understand false belief as they acquire linguistic constructions for representing a proposition and the speaker's epistemic attitude toward that proposition. In the current study, English-speaking children of 3 and 4 years of age (N = 64) were asked to interpret propositional attitude constructions with a first- or third-person subject of the propositional attitude (e.g., “I think the sticker is in the red box” or “The cow thinks the sticker is in the red box”, respectively). They were also assessed for an understanding of their own and others’ false beliefs. We found that 4-year-olds showed a better understanding of both third-person propositional attitude constructions and false belief than their younger peers. No significant developmental differences were found for first-person propositional attitude constructions. The older children also showed a better understanding of their own false beliefs than of others’ false beliefs. In addition, regression analyses suggest that the older children's comprehension of their own false beliefs was mainly related to their understanding of third-person propositional attitude constructions. These results indicate that we need to take a closer look at the propositional attitude constructions that are supposed to support children's false-belief reasoning. Children may come to understand their own and others’ beliefs in different ways, and this may affect both their use and understanding of propositional attitude constructions and their performance in various types of false-belief tasks. --------------------------------------------------------------------------------","A large number of studies have shown that language plays a facilitative role in children’s development of false-belief understanding (for overviews, see Astington & Baird, 2005; Milligan, Astington, & Dack, 2007). However, it is still unclear which aspects of language are responsible. Some researchers have looked to discourse because children must constantly confront mismatches between what they and their interlocutor know or do not know for pragmatically appropriate communication (e.g., Harris, 1996, 1999; Peterson & Siegal, 2000; Tomasello & Rakoczy, 2003). Others have looked to the more representational aspects of language, in particular (a) children’s mastery of mental state terms, such as think and know (e.g., Bartsch & Wellman, 1995; Olson, 1988; Ruffman, Slade, & Crowe, 2002), and (b) children’s mastery of propositional attitude constructions in which these mental state terms prototypically occur, such as I think that he will be late again (e.g., de Villiers & de Villiers, 2000; Hale & Tager-Flusberg, 2003; Lohmann & Tomasello, 2003; Low, 2010). In the remainder of this article, we use linguistic terminology and refer to these propositional attitude constructions as complement-clause constructions or just complements. These complement-clause constructions contain a main clause expressing the attitude toward a proposition (e.g., I think) and a subordinate clause expressing that proposition (e.g., he will be late again). In the current study, we took a closer look at children’s understanding of these complement-clause constructions and the parallel development of their false-belief understanding. Around the age of 4 years, children typically start to pass explicit tests of false belief (Wellman, Cross, & Watson, 2001). To investigate whether this developmental achievement is equally supported by different kinds of complement-clause constructions, we compared 3- and 4-year-olds’ comprehension of complement-clause constructions with first- and third-person subjects in the main clause (henceforth, first-person complements, such as “I think the sticker is in the red box”, and third-person complements, such as “The cow thinks the sticker is in the red box”). We further investigated how the same children perform in tasks testing their understanding of their own and others’ false beliefs. Our main hypothesis was that third-person complements are more tightly related to false belief because they are more likely to encode genuine references to mental states. Linguistic research suggests that whereas first-person complements can refer to mental states, they are also regularly used as epistemic parentheticals, which are produced to alert the listener to the relative (un)certainty of a proposition. Phrases like I think can be translated as maybe (e.g., Thompson & Mulac, 1991; Verhagen, 2005). In other words, first-person complements are ambiguous because they can either refer to mental states or function as (un)certainty markers (see also Manson, 2002). Moreover, even when used as epistemic parentheticals, first-person complements are ambiguous on another level. That is, a phrase like I think in I think it’s in the red box can express either certainty or uncertainty, and it has been suggested that children up to the age of 5 years treat I think as if it means I know (e.g., Bassano, 1985; Miscione, Marvin, O’Brien, & Greenberg, 1978; Naigles, 2000). Therefore, first-person complements are highly ambiguous, and this ambiguity may affect the way in which children interpret them and the way in which they are related to their false-belief understanding. Production of complement-clause constructions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The idea that first-person complements do not necessarily refer to mental states is supported by studies of spontaneous speech demonstrating that children produce first- person complements considerably before they typically show an explicit understanding of false belief. When Diessel and Tomasello (2001) looked at young English-speaking children’s production of complement clauses, they found that at around the age of 3 years children used many mental verbs only in the first person, as in I think it’s in there (see also Bloom, Rispoli, Gartner, & Hafitz, 1989). Following functional linguists (e.g., Thompson & Mulac, 1991; Verhagen, 2005), they argued that in this case the I think phrase is not being used to refer to a mental state or mental activity but rather is being used as a kind of epistemic marker to alert the listener to the speaker’s relative uncertainty (for similar findings in German children’s spontaneous speech, see Brandt, Lieven, & Tomasello, 2010). Comprehension of complement-clause constructions and false belief ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To investigate the developmental gap between children’s production of first-person complement-clause constructions at around the age of 3 years (e.g., Diessel & Tomasello, 2001) and their successful performance in explicit false-belief tests at around the age of 4 years (Wellman et al., 2001), Moore, Bryant, and Furrow (1989) and Moore, Pure, and Furrow (1990) tested whether and when young English-speaking children correctly interpreted first-person complements and whether they understood them before they passed explicit false-belief tests. Moore and colleagues (1989) developed a hidden-object task where two puppets indicated which of two boxes contained some candy. The puppets produced first-person complements only. For example, one puppet said I think it’s in the red box, and then the other puppet said I know it’s in the blue box. In this case, children who understood these verbs and sentence types were expected to pick the blue box. Moore and colleagues (1989) contrasted three mental state verbs: guess, think, and know. Overall, 3-year-olds performed at chance level, whereas 4-year-olds tended to perform at above chance level but still performed worse than 6- and 8-year-olds. Looking at the think–know contrast in particular, the 3-year-olds, on average, went for the correct box in 2 of 4 trials (M = 2.07), whereas the 4-year-olds, on average, went for the correct box in 3 of 4 trials (M = 3.14). Moore and colleagues (1990) showed that, in addition, 4-year-olds’ performance in this hidden-object task was strongly correlated with their performance in various explicit tests of false belief. These findings suggest that, even though children produce first-person complements with mental verbs before they show an explicit understanding of false belief, their comprehension of first-person complements develops at the same time as they start to show an explicit understanding of false belief. A similar developmental pattern and dissociation between production and comprehension has been observed in children’s use and understanding of epistemic modals. Like first-person complements, modal verbs, such as must and will, are ambiguous. For example, must can express necessity (e.g., you must go to bed now) or epistemic modality (e.g., saying that must be the postman on hearing the doorbell). Studies looking at children’s production and comprehension of modal verbs suggest that children first use them to express notions like necessity (e.g., Wells, 1979). At around the age of 3 years, they start using the same verbs to express epistemic modality in apparently appropriate contexts. However, when children are tested on their comprehension of the epistemic meaning of modal verbs, they do not seem to understand them in any systematic way until the age of 4 or 5 years, and Papafragou (1998) suggested that children’s full understanding of the epistemic functions of modal verbs depends on their theory-of-mind understanding (see also Moore, Pure, & Furrow, 1990). For mental verbs and complement clauses, it has been suggested that children’s comprehension and correct use of the more complex (i.e., mental state) functions do not just depend on, but also support, their theory-of-mind development (e.g., Astington & Baird, 2005; de Villiers, 2007; Milligan et al., 2007). In the current study, we explored the possibility that children’s theory-of-mind development is mainly associated with children’s understanding of third-person complements. This assumption is suggested by a number of studies that found that caregivers’ talk about their own mental states (e.g., I think this is a golf ball) shows weaker links to children’s false-belief understanding than their talk about others’ mental states, including those of the children (e.g., you think this is a golf ball) (Adrián, Clemente, & Villanueva, 2007; Booth, Hall, Robison, & Kim, 1997; Howard Gola, 2012; Taumoepeau & Ruffman, 2006). Howard, Mayeux, and Naigles (2008) also found that mothers’ use of first-person complements does not support children’s ability to systematically distinguish between the epistemic functions of I think, expressing relative uncertainty, and I know, expressing certainty. This is probably due to the fact that, when used with first-person subjects, mental verbs (such as think and believe) can express either certainty or uncertainty. For example, when a child hears an utterance like I think it’s bedtime now from their mother, the mother is pretty certain that it indeed is bedtime (cf. Howard et al., 2008; Naigles, 2000). In this case, the meaning of I think cannot easily be distinguished from the meaning of I know (see also Bassano, 1985; Miscione et al., 1978). Howard and colleagues (2008) found that in the linguistic input of 3- and 4-year-old English-speaking children, more than half of the phrases containing the verb think were used to express certainty rather than uncertainty, that think was most often used with first-person subjects, and that these phrases did not directly support children’s understanding of false belief and the epistemic functions of mental verbs. To summarize, previous research suggests that first-person complements do not necessarily refer to mental states, and thus their use in children’s own language and in their input does not directly support children’s false-belief understanding. Neither does the everyday use of first-person complements seem to support children’s understanding of the semantics of mental verbs and the complement-clause constructions in which these verbs are used (Howard et al., 2008). What has not been investigated yet is (a) whether young children find it easier to distinguish the semantics and epistemic functions of mental verbs when they are used in third-person complements (e.g., she thinks the sticker is in the red box vs. she knows the sticker is in the blue box) as opposed to first-person complements and (b) how young children’s understanding of third-person complements is related to their false-belief understanding. First-person versus third-person complements and false-belief understanding ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the current study, we modified the hidden-object task (cf. Moore, Bryant, & Furrow, 1989) to directly compare 3- and 4-year-olds’ comprehension of first- and third-person complements. For example, in the first-person version, one puppet said I think the sticker is in the red box, and the other puppet said I know the sticker is in the blue box. In the third-person version, the experimenter spoke for the puppets: The cow thinks the sticker is in the red box. The pig knows the sticker is in the blue box. In addition, we used a more balanced set of false-belief tasks that allowed us to directly compare children’s understanding of their own and others’ false beliefs and how this relates to their understanding of first- and third-person complements. Based on the assumption that phrases like I think in first-person complements can express either certainty or uncertainty (e.g., Howard et al., 2008), we expected children to perform worse on the first-person I think–I know contrast than on the third-person the cow thinks–the pig knows contrast. Based on the assumption that first-person complements can either refer to mental states or function as (un)certainty markers (e.g., Diessel & Tomasello, 2001), we also expected stronger developmental associations between third-person complements and false-belief understanding than between first-person complements and false belief.","The children were recruited through a child participant database and tested in a quiet room at the university of a medium-sized English city. In total, 32 young 3-year-olds (M = 3 years 4 months, range = 3;1 [years;months] to 3;5, 17 girls) and 32 young 4-year-olds (M = 4 years 4 months, range = 4;1 to 4;4, 15 girls) participated in the study. One additional 3-year-old was tested but needed to be excluded from the analyses because she failed the pretest (described below). All children were English-speaking monolinguals. None of the participants had any known language impairment. Design and materials ~~~~~~~~~~~~~~~~~~~~ All children started with the hidden-object task (cf. Moore et al., 1989) and did four false-belief tests afterward. As described in more detail below, we used the classic unexpected content and change of-location tests (Perner, Leekham, & Wimmer, 1987; Wimmer & Perner, 1983) as well as a new version of the change-of-location test, which tests children’s understanding of their own false belief (Buttelmann, 2016). In the hidden- object task, we had two conditions that were tested between participants. In each age group, 16 children (8 girls) were tested in each condition. In the first-person condition, for each trial children heard two contrastive statements from two hand puppets (a cow and a pig). The complement clauses were used with one of two mental verbs and a first-person singular subject in the main clause (e.g., pig: I think the sticker is in the blue box; cow: I know the sticker is in the red box). As in one of Moore and colleagues’ (1989) conditions, the two mental verbs that were contrasted were think, marking relative uncertainty, and know, marking certainty. In the first-person condition, the test sentences were prerecorded and played from little speakers hidden under the hand puppets. In the third-person condition, the experimenter spoke for the puppets and the children heard, for example, The pig thinks the sticker is in the blue box; The cow knows the sticker is in the red box. Each child received eight trials. Across trials, we counterbalanced the order of the statements (whether the first statement contained think or know in the main clause), the assignment of the statements to the hand puppets (whether the pig or cow knew or thought), and the assignment of the statements to the boxes (whether it was known or thought that the sticker was hidden in the red or blue box).","All children were tested by a female experimenter who sat opposite them at a small table. For each trial, the child saw two opaque boxes on the table—always a red one and a blue one. The red box was always placed to the left of the child. The experiment started with a pretest. The experimenter told the child that she, the pig, and the cow had hidden many stickers in the boxes and that the pig and cow would help the child to find these stickers. She also told the child that the pig and cow might not remember all the hiding places. For each of four trials a new pair of boxes was placed on the table in front of the child. After the experimenter put the two boxes on the table, she asked the puppets, Can you help X [child’s name] find the sticker? Which box is the sticker in? In the pretest, the child heard two non-contrastive statements about the location of the sticker from the two hand puppets. One statement was affirmative and the other was negated (e.g., pig: The sticker is in the red box; cow: The sticker is not in the blue box). Whether the cow or the pig used the affirmative or negated statement, and which statement came first, was counterbalanced. In the pretest, children were allowed to choose and open one of the two boxes right away. Children who picked the right box in at least three of four trials continued with the experiment. Children who scored lower than three of four trials received two additional trials. If they then picked the right box in four of six trials, they also continued with the experiment. As indicated above, we needed to exclude one 3-year-old because she did not reach this criterion. In each experimental trial, children needed to choose one box but were not allowed to look into any of the boxes before they were finished with all eight trials. In the first-person condition, the hand puppets uttered the prerecorded test sentences. In the third-person condition, the hand puppets first whispered into the experimenter’s ear, and the experimenter then uttered the test sentences. The whispering did not contain any actual words, and the experimenter produced the third-person complements right after the whispering. No additional instructions were given to the child. Because the task was quite demanding, after four experimental trials there was a break during which the experimenter played a puzzle game with the child. They played this game for approximately 5 min and then continued with the second set of experimental trials. Before they did the false-belief tests, children were allowed to look into the boxes they chose and collect their stickers. Note that, finally, all boxes contained stickers so that children were not differently rewarded before they began the subsequent false-belief tests. For testing children’s understanding of false belief, we presented participants with four different tasks. In the classic change-of-location test (Wimmer & Perner, 1983), children needed to answer a test question about another person’s false belief. The experimenter told the story and acted it out with two little dolls and props: Sally puts a ball into her basket, Sally leaves the room, Ann transfers the ball from Sally’s basket into her box, and Sally returns. The experimenter then asked three questions: (a) the test question (Where will Sally look for her ball?), (b) the reality control question (Where is the ball really?), and (c) the memory control question (Where did Sally put her ball in the first place?). The change-of-location own-belief test is based on the classic change-of-location paradigm but tests children’s understanding of their own false belief (Buttelmann, 2016). For this, things were arranged such that children searched in the incorrect location for a small toy. The experimenter placed two boxes (a green one and a pink one) on the table and told the child that she was going to hide a small toy ball in one of them. She then put an occluder on the table to block the child’s view and put the toy ball into one of the boxes (the pink one in Fig. 1A and B). At the same time, she slightly manipulated the position and the cover of the other box (the green one in Fig. 1). So, after the occluder was removed, it looked like the experimenter had manipulated one box, but had not touched the other one (see Fig. 1C and Buttelmann, 2016, for details regarding the procedure of this test). She then asked the child (a) the manipulation control question (Where is the ball?). Except for one 3-year old, all children pointed to the box that looked like it had been manipulated and thus held a false belief concerning the location of the ball. The experimenter then showed the child that this box was actually empty and that the ball was hidden in the other box. She then put the ball back in the same box (i.e., the one that did not look like it had been manipulated) and asked (b) the test question (Where did you first think the ball was?). As in the classic change-of-location test, the experimenter also asked (c) the reality control question (Where is the ball now?). The unexpected content (or “Smarties”) test included questions about both children’s own false belief and another person’s false belief (Perner et al., 1987). We followed the original procedure of the task and used a Smarties tube that was filled with crayons. The experimenter first asked the child what he or she thought was is in there. After the child said Smarties, chocolate, sweets, or something similar, the experimenter showed the child the actual content of the box (crayons). She then put the crayons back into the tube, closed it, and asked the child three questions: (a) the memory control question (Can you remember what’s inside here?), (b) the own-belief test question (What did you first think was inside here?), and (c) the other-belief test question (What will pig think is inside this box?). The order of the own-belief test question (b) and the other-belief test question (c) was counterbalanced across children. Thus, we were able to test children’s understanding of their own and others’ false beliefs in scenarios involving unexpected contents and the change of an object’s location. The unexpected content and change-of-location tests were presented as blocks, and the order was counterbalanced. Within these blocks, we also counterbalanced the order of own and others’ false-belief questions. Test sessions lasted 20 to 25 min. Scoring ~~~~~~~ Children’s choices in the hidden-object task were scored as correct when participants chose the box marked by I know in the first-person condition and the box marked by the pig knows or the cow knows in the third-person condition. For each false-belief task, children got a score of 1 (pass) if they correctly answered both the test question and the corresponding control question(s). Overall, eight 3-year-olds failed to correctly answer the reality control question (about the actual location of the object) in the change-of- location task (five in the change-of-location own-belief test and three in the change-of- location other-belief test). In addition, six 3-year-olds did not give a correct answer to the memory control question (about the actual content of the Smarties tube) in the unexpected content task. Among the 4-year-olds, only two gave incorrect answers to the reality control questions (about the actual location of the object) in the change-of- location own-belief test. Trials in which children did not answer the control question(s) correctly were dropped from the analyses. That is, for that specific false-belief test, children did not get a score of either 0 (fail) or 1 (pass). For the unexpected content task, this means that they did not get a score for either the own- or other-false-belief test question. However, no child failed all control questions. So, each child got a score of 0 (fail) or 1 (pass) for at least one false-belief measure and was considered in the subsequent analyses.","First, we analyzed children’s performance in the hidden-object task. A 2 (Age: 3- vs. 4-year-olds) by 2 (Condition: first- vs. third-person complements) analysis of variance (ANOVA) suggested that there was a significant interaction between age and condition, F(1, 60) = 10.56, p = .002. Therefore, we ran separate analyses for the two age groups. The 3-year-olds performed at chance in both the first- and third-person conditions (first- person: M = 52.3% of trials, SD = 13.1; Wilcoxon test: Z = .577, N = 16, p = .564; third- person: M = 50.0% of trials, SD = 12.9; Wilcoxon test: Z = .054, N = 16, p = .957) (see Fig. 2). Similarly, there was no significant difference between the younger children’s performances in the first- and third-person conditions (Mann–Whitney U-test: U = 117.0, Z = .443, Nfirst-person = 16, Nthird-person = 16, p = .657). The 4-year-olds, in contrast, performed above chance in both conditions (first-person: M = 58.6% of trials, SD = 16.3; Wilcoxon test: Z = 2.08, N = 16, p = .038; third-person: M = 79.7% of trials, SD = 15.1; Wilcoxon test: Z = 3.434, N = 16, p = .001). Unlike the 3-year-olds, the older children also performed significantly better in the third-person condition than in the first-person condition (Mann–Whitney U-test: U = 46.0, Z = 3.157, Nfirst-person = 16, Nthird-person = 16, p = .002) (see Fig. 2). When we directly compared the two age groups, we found that the 4-year-olds performed better than the 3-year-olds in the third-person condition (Mann–Whitney U-test: U = 238, Z = 4.23, N3-year-olds = 16, N4-year-olds = 16, p < .001). For the first-person condition, however, we did not find any significant age differences (Mann–Whitney U-test: U = 148, Z = .793, N3-year-olds = 16, N4-year-olds = 16, p = .468). Table 1 shows the number and percentage of children who passed each false-belief test in each age group. Remember that we had to exclude a number of 3-year-olds and two 4-year- olds from some trials because they did not answer the corresponding control question(s) correctly. In addition, one 3-year-old did not have a false belief in the own-false-belief version of the change-of-location test. Therefore, depending on the task, the total number of 3-year-olds included in the analyses ranged from 26 to 29 and the number of 4-year-olds included in the analyses ranged from 30 to 32. The 4-year-olds performed better in the own-false-belief tests than in the other-false-belief tests. No significant differences were found between the own-false-belief versions of the unexpected content test and the change-of-location test (McNemar test: N = 50, p = .804). Similarly, there was no significant difference between the other-false-belief versions of the two tasks (McNemar test: N = 56, p = .839). Therefore, both tasks were combined for percentages of trials passed in the own- and other-false-belief tests. The 4-year-olds performed above chance in the own-false-belief tests (M = 84.4% of trials, SD = 29.6; Wilcoxon test: Z = 4.315, N = 32, p < .001) and at chance in the other-false-belief tests (M = 56.3% of trials, SD = 39.7; Wilcoxon test: Z = .894, N = 32, p = .371). The 3-year-olds performed below chance in both the own- and other-false-belief tests (own: M = 31.3% of trials, SD = 37.6; Wilcoxon test: Z = 2.558, N = 32, p = .011; other: M = 25.8% of trials, SD = 31.3; Wilcoxon test: Z = 3.441, N = 31, p = .001). In accordance with these results, we found that the 4-year-olds performed better than the younger group in both own-false-belief tests (Mann–Whitney U-test: U = 857, Z = 4.98, N3-year-olds = 31, N4-year-olds = 32, p < .001) and other-false-belief tests (Mann–Whitney U-test: U = 704, Z = 3.06, N3-year-olds = 31, N4-year-olds = 32, p = .002). The pairwise comparisons presented so far suggest that between the ages of 3 and 4 years, children develop a better understanding of third-person complements and of their own false belief. Their understanding of others’ false beliefs also develops, but even the 4-year-olds still performed at chance in tasks testing the understanding of others’ false beliefs. To further investigate whether there is a relationship between children’s understanding of first- and third-person complements and their understanding of false belief, we ran two regression models in R (R Core Team, 2014). In the first model, we entered children’s performance in the hidden-object task (percentage trials correct), condition (first- vs. third-person complements), and age (3- vs. 4-year-olds) to predict their performance in own-false-belief tasks. Condition and age were coded as categorical variables (third-person and 4-year-olds, respectively). Performance in the hidden-object task was coded as a continuous numerical variable. We found main effects for age, performance in the hidden-object task, and condition as well as a three-way interaction among all the factors (see Table 2). Together with the pairwise comparisons presented above, these main effects and complex interaction suggest that it is the older children’s growing understanding of third-person complements that is related to their improved understanding of their own false belief. When we entered the same variables to predict children’s performance in other-false-belief tasks, we found a main effect for their performance in the hidden-object task and a two-way interaction between age and performance in the hidden-object task. Together with the pairwise comparisons presented above, this suggests that only the 4-year-olds’ understanding of complement clauses was positively related to their understanding of others’ false beliefs (see Table 3).","In the current study, we found developmental differences between 3- and 4-year-old English-speaking children’s understanding of third-person complements and between 3- and 4-year-olds’ understanding of false belief. In particular, the older children performed above chance level and also significantly better than the younger age group in the third- person condition of the hidden-object task (with third-person complements) and in the own- false-belief tests. Unlike the younger age group, the 4-year-olds also performed just above chance in the first-person condition of the hidden-object task (with first-person complements), but their performance in this condition did not significantly differ from that of the 3-year-olds. In addition, the 4-year-olds performed at chance level in the other-false-belief tasks, whereas the 3-year-olds were below chance and significantly worse than the older age group. Together with these pairwise comparisons, the main effects and complex interactions in the regression analyses suggest that the older children’s understanding of own false belief was positively related to their understanding of third- person complements. Their developing understanding of others’ false beliefs seems to be related to their understanding of both first- and third-person complements. These findings support our main hypothesis that the developmental link between third-person complements and false-belief understanding should be stronger than the relation between first-person complements and false belief because third-person complements are more likely to encode genuine reference to mental states and mental processes (e.g., Diessel & Tomasello, 2001; Howard et al., 2008; Manson, 2002). In addition, the older children found it easier to distinguish the semantics of the mental verbs think and know in a third-person context than in a first-person context, which supports the assumption that the semantics of mental verbs is less ambiguous when used with third-person subjects (cf. Howard et al., 2008). We found no clear developmental links between first-person complements and false-belief understanding. In the hidden-object task, the 4-year-olds did not perform significantly better with first-person complements than their younger peers. However, we did find developmental differences for children’s understanding of false belief in the sense that the 4-year-olds were significantly better than the 3-year-olds despite the older group being only at chance on others’ false beliefs. It could also be possible that the understanding of first-person complements was linked to children’s understanding of their own false beliefs. But the regression analysis and pairwise comparisons suggest that only the third-person complements, not the first-person complements, were positively linked to children’s understanding of their own false beliefs. These findings provide further support for the assumption that it is third-person complements that are intrinsically related to children’s development of false-belief reasoning. We suggest that first-person complements show weaker links to children’s false-belief development because they do not necessarily refer to mental states and are ambiguous even when they are used as epistemic markers (cf. Diessel & Tomasello, 2001; Howard et al., 2008). However, previous studies did find correlations between children’s understanding of first-person complements and false belief as well as developmental differences between 3- and 4-year-olds’ understanding of first-person complements with mental verbs (e.g., Howard et al., 2008; Moore et al., 1989, 1990). The discrepancies between the current study and previous studies might be due to methodological differences. For example, Moore and colleagues’ (1989) finding that there were developmental differences for children’s understanding of first-person complements might be due to the fact that the age ranges applied in that study were much wider than those applied in the current study (i.e., the 4-year-olds were between the ages of 4 and 5). Therefore, the effect might have been driven by the older participants. Another possible explanation is that in Moore and colleagues’ original study the experimenter produced all test sentences, whereas we used prerecorded test sentences in the first-person condition. However, when we did a similar study with German-speaking children, the test sentences were produced live in both conditions of the hidden-object task as in Moore and colleagues’ study. The pattern of results was similar to that in the current study; we found developmental differences only for third-person complements and false-belief understanding (Brandt & Buttelmann, 2015). As mentioned before, previous studies also found correlations between children’s understanding of first-person complements and their general understanding of false belief (Howard et al., 2008; Moore et al., 1990). However, unlike the current study, previous investigations have not systematically distinguished between children’s understanding of their own and others’ false beliefs. For example, Howard and colleagues (2008) also used the unexpected content test but gave children a combined score for their answers to the questions about their own and someone else’s false beliefs. When we systematically distinguished between children’s understanding of their own and others’ false beliefs, we found a positive relation between the older children’s developing understanding of others’ false beliefs and their comprehension of both first- and third-person complements (see regression analysis in Table 3). However, the older children’s understanding of their own false beliefs was more advanced than their understanding of others’ false beliefs and was positively related only to their understanding of third-person complements (see regression analysis in Table 2). The relationship between first-person complements and false-belief understanding is probably due to the fact that although phrases like I think and I know are often used just like adverbials expressing (un)certainty (e.g., Diessel & Tomasello, 2001; Thompson & Mulac, 1991; Verhagen, 2005), this function is not completely independent of the more complex meanings of these mental state verbs. Most of the time, we do not use these phrases to explicitly refer to mental states. Still, even using them to express different degrees of certainty requires some concept of mind. Indeed, recent proposals suggest that even if first-person complements are not used to refer to mental states directly, their mastery may bootstrap young children into understanding true reference to mental states. This might be because children notice that the verbs they use and comprehend as a signal of, for example, (un)certainty are being used in a slightly different way and might be used and comprehended as more or less explicit reference to mental states (Gordon, 1995; Tomasello & Rakoczy, 2003). Children’s comprehension of mental verbs in first-person contexts is also likely to be informed by their understanding of mental verbs in other contexts, such as third-person complements, where these verbs are more likely to refer to mental states. However, as has been shown for a variety of linguistic terms and constructions (for an overview, see Tomasello, 2003), developing a more abstract, context- independent representation of mental verbs takes time. And, as our current and previous findings suggest, the acquisition of an abstract representation of mental verbs also interacts with children’s false-belief development. A similar developmental story has been put forward for the acquisition of modal verbs and other forms of epistemic markers and evidentials, such as sentence-final particles, where children use apparently semantically and/or syntactically complex terms and structures appropriately before they understand the full range of concepts behind these terms and structures (e.g., Aksu-Koç, Ögel-Balaban, & Alp, 2009; Matsui, Yamamoto, & McCagg, 2006; Papafragou, 1998; Papafragou, Li, Choi, & Han, 2007). For example, children start using the modal auxiliary will at around the age of 2.5 years (Wells, 1979). However, early in development, this modal auxiliary is most likely to be used to communicate intention. Only at around the age of 5 years do children use will to express how certain they are about something (e.g., saying that will be the postman on hearing the doorbell) (Wells, 1979). This latter use is referred to as epistemic modality, and Papafragou (1998) argued that children’s comprehension and correct use of modal verbs with an epistemic function depends on their theory-of-mind development. To grasp the epistemic function of modal verbs, children need to have developed a “representational model of mind” (Forguson & Gopnik, 1988, as cited in Papafragou 1998, p. 383). However, Papafragou also suggested that this epistemic function is related to other, more basic functions of modal verbs and that it is, indeed, not always easy to distinguish between epistemic and other kinds of modals when looking at spontaneous speech. It seems possible that once children have acquired a theory of mind, they extend the more basic root functions of modal verbs expressing intention, ability, obligation, and so on to the more complex epistemic functions of modal expression. For mental verbs and complement clauses, it has been suggested that children’s comprehension and correct use of the more complex functions not only depend on but also support their theory-of-mind development (e.g., Astington & Baird, 2005; de Villiers, 2007; Milligan et al., 2007). When children start using mental verbs and complement clauses, they tend to use first-person complements with restricted phrases and fixed discourse functions (Köymen, Lieven, & Brandt, 2016). Most important, it has been claimed that children’s first uses of mental verbs and complements do not refer to mental states (e.g., Bartsch & Wellman, 1995; Diessel & Tomasello, 2001; Shatz, Wellman, & Silber, 1983). Nevertheless, children start to talk about their own mental states at around the age of 3 years (Bartsch and Wellman, 1995). Results from the current study and previous research suggest that this talk about own mental states might be related to children’s developing understanding of (others’) false beliefs. However, third-person (and possibly also second-person) complements show a stronger developmental link with children’s understanding of false belief. As has been suggested by de Villiers (2007), complement clauses serve as representational tools for children’s (and adults’) false-belief understanding because “the complement is embedded under the verb and takes the particular perspective or point of view of the subject, not the speaker, licensing also the subject’s terms of reference even when these are not the speaker’s” (p. 1868). In other words, third-person complement-clause constructions (e.g., she thinks he’ll be late) allow us to distinguish between our own and someone else’s perspective (e.g., she). First-person complements, on the other hand, express only one perspective because the subject (e.g., I in I think he’ll be late) refers to the speaker herself. The current data do not allow us to make substantial claims about the causal relationship between developments in language and theory of mind, but previous studies suggest that language supports explicit false-belief understanding rather than the other way around (see meta-analysis by Milligan et al., 2007). Results from the current study allow more detailed hypotheses, which need to be tested in follow-up training and longitudinal studies. Our findings suggest that even though first-person complements also play a role in children’s developing understanding of false belief, it is the understanding of third-person complements that shows a parallel development with that of false belief. In particular, children’s understanding of their own false belief develops together with their understanding of third-person complements. A more general and explicit understanding of both own and others’ false beliefs might develop out of children’s understanding of own beliefs, very likely supported by their understanding of first- and third-person complements. In both linguistic and sociocognitive development, children develop more abstract representations of mental verbs and belief as they discover commonalities across different discourse contexts."],["Children between 5 and 8. years of age freely intervened on a three-variable causal system, with their task being to discover whether it was a common cause structure or one of two causal chains. From 6 or 7. years of age, children were able to use information from their interventions to correctly disambiguate the structure of a causal chain. We used a Bayesian model to examine children's interventions on the system; this showed that with development children became more efficient in producing the interventions needed to disambiguate the causal structure and that the quality of interventions, as measured by their informativeness, improved developmentally. The latter measure was a significant predictor of children's correct inferences about the causal structure. A second experiment showed that levels of performance were not reduced in a task where children did not select and carry out interventions themselves, indicating no advantage for self-directed learning. However, children's performance was not related to intervention quality in these circumstances, suggesting that children learn in a different way when they carry out interventions themselves. --------------------------------------------------------------------------------","Most developmental studies of causal learning (e.g., Bullock, Gellman, & Baillargeon, 1982; Gopnik, Sobel, Schulz, & Glymour, 2001; Shultz, 1982) have required children to judge whether an event is causally efficacious. However, in learning about the world, what is at issue is often not just whether a specific variable has a particular causal power but also the structure of the causal relations between a set of variables. You might observe that Events A, B, and C tend to co-occur (e.g., that when you feel stressed you are likely to drink more heavily and also that your blood pressure is raised). This co- occurrence is consistent with a variety of different causal structures; for example, the structure may be a common cause in which A independently causes both B and C, B ← A → C (stress causes heavier drinking and also independently causes raised blood pressure), or a causal chain in which A causes B which causes C, A → B → C (stress causes heavier drinking which raises blood pressure). How do we distinguish between these possibilities? Intervening on a causal system potentially provides very important information (Hagmayer, Sloman, Lagnado, & Waldmann, 2007; Sloman & Lagnado, 2005; Steyvers, Tenenbaum, Wagenmakers, & Blum, 2003). Observing what happens if we intervened on B only would allow us to distinguish between the two suggested causal structures; if, assuming no background causes, we make B occur on its own (engage in heavy drinking when not stressed), and C does not occur (blood pressure is not elevated), we can rule out the A → B → C causal chain. A variety of studies have required adults to infer the structure of the relations between sets of variables (e.g., Fernbach & Sloman, 2009; Kushnir, Gopnik, Lucas, & Schulz, 2010; Lagnado & Sloman, 2004, 2006; Sobel & Kushnir, 2006; Steyvers et al., 2003). In several of these studies, participants learned the causal structure by deciding what interventions to make on elements in the system, carrying out these interventions and observing their effects (Bramley, Lagnado, & Speekenbrink, 2015; Lagnado & Sloman, 2004, 2006; Sobel & Kushnir, 2006; Steyvers et al., 2003). Such interventions are assumed to reveal conditional dependencies or independencies between variables, and adults’ success on these tasks has been interpreted as being consistent with the causal Bayes net approach to causal learning (e.g., Glymour, 2001; Gopnik et al., 2004) that captures causal learning in terms of the construction of causal models based on conditional probability information. A major advantage of this approach is that it specifies how and why interventions on a system yield richer information about the causal (in)dependencies between variables than that which is available through observation of patterns of covariation (Hagmayer et al., 2007; Steyvers et al., 2003; Waldmann & Hagmayer, 2005). The causal Bayes net approach has been extensively adopted by developmental psychologists interested in explaining children’s learning about causation (Gopnik, 2012; Gopnik & Wellman, 2012; Gopnik et al., 2004). The majority of studies in this tradition have involved children learning whether an object possesses a particular causal power, usually on the basis of observing the experimenter’s actions (e.g., whether an object makes a box light up and play a tune; Gopnik & Sobel, 2000; Kushnir & Gopnik, 2007; Sobel, Tenenbaum, & Gopnik, 2004). Relatively few studies have used tasks in which children themselves decide which interventions to carry out in order to discover the causal structure of a system (e.g., whether it has a causal chain or common cause structure). Such studies are particularly important because they can be used to assess young children’s effectiveness in generating and testing hypotheses about the causal relations between sets of variables. Moreover, a key advantage of the causal Bayes net approach over most other accounts of causal learning is that it can capture this more complex type of learning, distinguishing between different causal paths as well as identifying variables’ ultimate effects. One study that did examine children’s ability to learn causal structure by means of making interventions on a system is that of Schulz, Gopnik, and Glymour (2007, Experiment 3), in which 4- and 5-year-olds intervened on a causal system involving a box with two gears. Children needed to decide whether each gear moved by itself or whether one of the gears caused the other one to move. Children could remove each gear in turn from the box to examine whether the other gear worked on its own when the box itself was switched on. They gave their answers about the relations between the gears by selecting from a set of anthropomorphized pictures of the two gears that depicted different possible relations between them. Performance on this task was mixed. Children did not all reliably generate the right interventions to distinguish between the different possible relations that might hold between the gears. Even among those who did make appropriate interventions, children were not successful at identifying instances in which one of the other gears caused the other one to move. Schulz and colleagues’ (2007) study provides limited support for the claim that children will generate informative interventions and use this information to distinguish between different causal structures. Not only was performance relatively weak, but children were only required to make judgments about the dependencies between pairs of variables rather than to distinguish among three-variable causal structures. Children gave their responses by pointing to pictures of the two gears that showed whether they turned themselves or one turned the other. There was a third variable—a switch—that was important in the system, but children did not need to represent how its relation to the gears varied for the different causal systems, and it did not feature in the pictures depicting causal relations between the gears. Thus, this study does not allow us to draw firm conclusions about whether children can use interventions to distinguish between, for example, common cause and causal chain structures. However, the findings of some other studies suggest that we should expect even very young children to be good at crafting appropriate interventions and using them to learn about causal systems (Bonawitz, van Schijndel, Friel, & Schulz, 2012; Cook, Goodman, & Schulz, 2011; Schulz & Bonawitz, 2007). Indeed, Schulz (2012) argued that the ability to select appropriate interventions and use the evidence generated from such interventions may be developmentally basic. In various studies, she and her colleagues showed that young children will appropriately explore a causal scenario when given ambiguous information (Cook et al., 2011; Schulz & Bonawitz, 2007). In these scenarios, children’s behavior did suggest that children were trying to figure out whether an object possessed a certain causal property. However, children did not need to make interventions to disambiguate the structure of the relations between different variables and then use this information to decide, for example, whether a system was a common cause or causal chain. A further study by Sobel and Sommerville (2010) tried to address this specific issue. Children viewed a box with four colored lights—A, B, C, and D—and were told that some of the lights could make other lights turn on. The box was configured so that the relations between the lights took the form of either a common cause, B ← A → C, or a causal chain structure, A → B → C. Children could interact freely with the box by switching on lights and observing their effects. They were then asked a series of questions about the relation between pairs of lights. Sobel and Sommerville found that children performed above chance on these questions, which could be interpreted as indicating that they were able to use the information generated from their interventions to decide on the structure of the causal relations. There are, however, two difficulties with this interpretation. First, before children answered the questions, the experimenter pressed each of the buttons in turn and narrated what it did; arguably, this provided children with the answers to the test questions (indeed, children performed above chance, although less accurately, when given just this narration). Second, it was not clear that to answer correctly children needed to have an integrated representation of how the three variables in the system were related to each other rather than just knowledge of pairwise relations. Indeed, Sobel and Sommerville did not include in their analyses the answers that children gave to the question of whether A makes C go in the case of the causal chain, arguing that answers to this question are hard to interpret. However, by questioning children only about the other pairwise relations between A and B and between B and C, it is impossible to know whether children actually understood the nature of the overall causal structure. The general point here is that we can distinguish between learning structurally local pairwise links and integrating such links to form a representation of causal structure. This distinction is important because learning localized pairwise relations is likely to be easier than learning global structure (Fernbach & Sloman, 2009). Moreover, learning local pairs only is often liable to lead to the wrong global model; for example, when one connection “explains away” the dependence between two others, a pairwise learning strategy would still attribute a connection between these two variables, whereas a global strategy would not. A number of other studies of children’s causal learning can also be interpreted as studying children’s learning of pairwise relations rather than global causal structure (e.g., Schulz, Goodman, Tenenbaum, & Jenkins, 2008; Sobel & Sommerville, 2009), meaning that we still have limited evidence about children’s ability to learn causal structure. Uncertainty as to whether children are adept at appropriately generating interventions and using them to learn causal structure comes from two sources. First, research on children’s scientific learning has for many years suggested that younger children may have great difficulty in generating appropriately informative interventions and learning the nature of relations between variables from the evidence generated by these interventions (e.g., Klahr & Dunbar, 1988; Klahr, Fay, & Dunbar, 1993; Kuhn, 1989; Schauble, 1996; Zimmerman, 2000, 2007). On the face of it, this body of findings seems at odds with recent findings from the causal Bayes net tradition. One possible explanation of the differing findings lies in the role of the knowledge base in scientific learning studies. For example, preexisting, and sometimes erroneous, beliefs can hamper children’s ability to generate appropriate interventions and interpret statistical data (e.g., Amsel & Brock, 1996; Kuhn et al., 1988). Indeed, Schulz and colleagues suggested that this type of factor, along with task complexity, may mask children’s basic learning skills (Bonawitz, van Schijndel, Friel, & Schulz, 2012; Cook et al., 2011), which may be better demonstrated in the tasks used in the causal Bayes net tradition where domain-specific knowledge is of limited importance and statistical evidence is simple. However, the findings of a recent study by McCormack, Frosch, Patrick, and Lagnado (2015) provide a second reason for being unsure about children’s ability to learn from interventions on a causal system. This study, like most of those in the causal Bayes net tradition, involved children learning about a novel mechanical system. The only relevant data for causal learning were supplied by two types of domain-general cues: statistical information provided through interventions on the system and the temporal patterns of event occurrence. Children needed to learn the causal structure of the system—a box with three separate shapes (A, B, and C) on its surface that rotated. Across two experiments, children watched while the experimenter intervened on components of the system. In one experiment, the experimenter carried out interventions in which she disabled one of the shapes by preventing it from moving before moving each of the other shapes in turn. Children did not find it straightforward to use the patterns of evidence provided by these interventions to discriminate between causal structures even when the system operated deterministically. Although 6- and 7-year-olds were able to use the evidence from the more complex interventions to accurately infer when the system was one of the causal chains, children younger than this could not do so, and even 7- and 8-year- olds were unable to use information from these interventions to accurately judge when the system was a common cause. McCormack and colleagues argued that children’s difficulties may stem from integrating pieces of evidence provided across a number of separate observations of the causal system. At first sight, McCormack and colleagues’ (2015) findings seem to be more consistent with the conclusions stemming from research on children’s scientific learning that has emphasized its limitations. It might be argued, however, that this study did not provide children with an optimal opportunity to demonstrate their abilities. Children watched while the experimenter made a series of interventions rather than making the interventions themselves. Sobel and Kushnir (2006) argued that participants find it easier to learn causal structure when they decide what interventions to conduct, largely because this provides an opportunity for them to engage in more active hypothesis testing (but see Lagnado & Sloman, 2004). In particular, they suggested that when participants craft interventions themselves, they obtain evidence in a structured way that makes it more apparent whether it supports specific hypotheses. Moreover, children might be particularly likely to benefit from being allowed to explore how a system operates in that hands-on interventions may ensure they stay engaged with the task. In this study, we used a task very similar to that of McCormack and colleagues (2015) in which children needed to decide whether a three-element causal system was a common cause, B ← A → C, an A → B → C causal chain, or an A → C → B causal chain. Children intervened on the system themselves in order to learn its structure. Shapes on top of a box rotated when children moved them by hand, or shapes could be moved by rotating another shape that was causally connected to them. For example, for the A → B → C causal chain, spinning A initiated the rotation of both B and C, and spinning B rotated C; all of the shapes always moved simultaneously in the tasks to minimize temporal cues. Children needed to select and carry out a series of interventions; these could be simple interventions in which they made one of the three shapes spin, or they could be more complex interventions in which children prevented one of the three shapes from moving by disabling it and then spun one of the other two shapes. Note that we were not attempting to faithfully recreate a free-play situation because it was important for our analyses that we were able to exhaustively categorize children’s actions on the system. Although they were completely free to choose their interventions, the only actions children could carry out were interventions on the system. Furthermore, it was made clear to children that their job was to learn the causal structure of the system and that they could not make an unlimited number of interventions. This allowed us to look in our modeling work at the efficiency with which children produced informative interventions. We examined two aspects of performance: the nature of the interventions that children selected and children’s causal structure choices. Not all interventions provided useful information to discriminate among the three possible causal structures, which allowed us to examine whether the tendency to choose informative interventions changes with age. We also examined whether there was any relation between the quality of children’s interventions and the likelihood that children chose the correct causal structure at test. It is possible to try to examine these issues without formal modeling (see Sobel & Kushnir, 2006) by, for example, simply distinguishing between two broad classes of informative and non-informative interventions. However, we chose to model children’s learning in a Bayesian framework. Doing so has two key advantages. First, it allows us to properly assess whether there are developmental changes in the extent to which children resemble idealized Bayesian learners. This is important because, within the currently dominant causal Bayes net tradition, young children’s learning is often characterized as approximating to such an ideal, particularly with regard to causal learning from statistical information (e.g., Gopnik, 2012; Gopnik & Wellman, 2012). Formal modeling allows us to assess the extent to which this characterization is appropriate by assessing children’s performance against the standards set by the Bayesian tradition itself. Second, although in this study we can (and do) classify interventions broadly as informative or non-informative, the learning task itself is sequential. This means that how informative an intervention is depends on what children have already observed and what they can remember about such observations. However, figuring out the informativeness of each intervention that a participant makes on a trial- by-trial basis would be a formidable task without a formal model. Indeed, without such a model, it is hard to see how one would operationalize the notion of informativeness under such circumstances. Our Bayesian model allowed us to capture the sequential nature of the learning task by assuming that the most informative interventions were those that maximally reduced uncertainty about which was the correct hypothesis at any particular point in the learning sequence given some level of forgetting.","Children were from three different school years: 21 5- and 6-year-olds (M = 72 months, range = 64–80), 31 6- and 7-year-olds (M = 86 months, range = 80–93) and 25 7- and 8-year-olds (M = 98 months, range = 93–103). Children were tested individually in their schools.","The study used a wooden box, 41 cm (long) × 32 cm (wide) × 20 cm (high), which had an on/off switch at the front. There were three different colored lids for the box. Two of these had three colored/patterned shapes (e.g., circle, rectangle, star) inserted on their surface that rotated independently on the horizontal plane; a separate lid was used in pretraining and had only two shapes (see Fig. 1). The colors and shapes of the components were varied across participants and causal structures. On each of the two lids used in testing, the three shapes formed an equilateral triangle of 24-cm sides. Each shape had a small hole that aligned with a hole in the lid of the box. There was a miniature red-and-white “Stop” sign affixed to a metal rod that could be inserted through the hole on any shape into the corresponding hole in the box, preventing it from moving. Each of the shapes could be rotated by hand; the rotation of the other shapes was controlled by a laptop hidden inside the apparatus that participants were unaware of. A set of photographs was used during the learning phase that participants used to indicate which intervention they were going to make; these photographs depicted each shape on the box, and in addition there were photographs of each of the shapes alongside the stop sign. Photographs of the whole box with its shapes depicting three possible causal structures were used at test for children to indicate their judgment of the causal structure: one common cause and two causal chains (i.e., depicting B ← A → C, A → B → C, or A → C → B). The photographs for use at test were overlaid with pictures of hands to indicate causal links (following Frosch, McCormack, Lagnado, & Burns, 2012). Procedure Children completed two test trials: one common cause and one causal chain (order counterbalanced). There was a pretraining phase to ensure that children knew what their task was and how to give their answer at test. The pretraining procedure used a lid on the box that had only two colored shapes inserted on its surface; its purpose was to demonstrate that some shapes caused others to move but that the stop sign could be used to prevent a shape from moving. Children were initially asked to name the colors of the shapes to ensure that they would know which shapes the experimenter was referring to, and the experimenter drew children’s attention to the on/off switch at the front set at the “off” position. She then switched the box on and manually rotated one of the two shapes (X). This had no effect on the other shape (Y), which remained stationary, and the experimenter pointed this out to children. She then rotated the other shape (Y), which resulted in the first shape (X) simultaneously rotating. She explained to children, “Some shapes are made to move by others.” The experimenter then switched the box off and introduced children to the stop sign, which she inserted into X to stop it from moving, saying, “See this stop sign, it can be used to stop a shape from moving. See the [color of X] one cannot move now.” She then switched the box on again and rotated Y, which this time had no impact on the movement of X because it was prevented from moving by the stop sign. Following this, the lid was removed from the box and replaced by a different colored lid with three different shapes for the first test trial. Children were asked to name the colors of the three shapes and were told that their job was to figure out how the box worked. They were introduced to the three test pictures depicting the three different causal structures, with the experimenter saying, “In a moment I will ask you to figure out how the box works, but first I want to show you some pictures of the box which show different ways in which the box may be working. Only one of them is right, and you’ve got to work out which is the right one. It won’t change halfway through, and it is definitely only one of the pictures. You’ll have to use your detective skills to work out which picture shows what the box does.” The experimenter described each of the three pictures (e.g., “In this picture, the red one makes the blue one go, and the blue one makes the white one go, and the hands show that”). Following these three descriptions, children were then asked a set of three comprehension questions. For each causal chain picture, the experimenter asked, “Can you show me the picture where the [color of A] one makes the [color of B/C] one go and the [color of B/C] one makes the [color of B/C] one go?” For the common cause picture, the experimenter asked, “Can you show me the picture where the [color of A] one makes both the [color of B] one and the [color of C] one go?” The majority of children answered these questions correctly the first time, but if they did not answer all three questions correctly, the experimenter repeated the initial descriptions and asked the comprehension questions again. This procedure was repeated again if necessary. Following this pretraining, the experimenter said, “I am going to switch the box on now, and I want you to figure out how the box works.” Children were told that they could do one of two things (order counterbalanced): either “You can move a shape to see if it makes other shapes move” or “You can stop a shape from moving by putting the stop sign in and then see what happens when you move another shape.” It was explained to children that before they carried out each intervention, they needed to point to a card indicating what they intended to do. The experimenter said, “Before you try anything on the box, I want you to point to one of these cards. This card means you want to spin the [color] one, and you point to this card if you want to stop the [color] one. See, we also have the cards for spinning and stopping the [color and color] ones. So, each time you want to do something, you point to one of these cards first.” Children were told that they had 12 “goes” to start with and that each time they moved a shape counted as 1 “go.” It was made clear that using the stop sign did not count as a go by itself; children needed to then in addition move one of the other shapes. The procedure with cards was used to ensure that children interacted with the box in a controlled way and to make clear that they could not make an unlimited number of interventions. It also ensured that all children made a fixed minimum number of interactions before attempting to answer the test question. Children were told that they did not need to keep track of the number of goes that they had with the box because the experimenter would count this for them. Before children began, the experimenter said, “Remember, you’ve got to figure out which picture shows how this box really works.” She then demonstrated what happened when the A shape was moved, which was that the other two shapes also moved simultaneously, and pointed out that they did not know yet “which ones make other ones go.” Participants were subsequently allowed to make interventions on the box by first selecting the appropriate card and then making the intervention. So, for example, if they wanted to see what happened when C was moved if B was disabled, they needed to point to the card depicting B with the stop sign in it and then to the card depicting C. They then carried out their intervention. After participants completed 12 interventions, the experimenter said, “You have had your 12 goes now. Do you want to choose which picture you think shows what the box did, or do you want to have another 6 goes?” The majority of participants opted to choose after 12 interventions. Children completed a short filler task (a paper-and-pencil maze) in between the common cause/causal chain trials. It was made clear that the second box might work in the same way as the first box or it might work in a different way. The second box always had a lid of a different color and different shapes.","In both trials, 69 of the 77 participants stopped after 12 interventions. The remaining 8 opted for an additional 6 interventions in one or the other trial. Of these, 4 participants opted for the additional 6 interventions on both trial types. Initial data analyses examined participants’ responses for each of the two trial types. Fig. 2 shows the percentage of participants who chose each response type for each trial type. The majority of participants in each group, except for the youngest group, chose the correct answer for the causal chain trial. The majority of participants in all groups chose the common cause response for the common cause trial. Binomial analysis showed that each group of participants chose the correct response more often than chance (all ps < .01) except for the 5- and 6-year-olds who did not select the correct causal chain more often than chance. This group tended to select the common cause response for both structures. Performance on the causal chain structure was associated with age, χ2(2) = 6.91, p < .05, with the number of correct responses improving with age. Performance on the common cause structure was marginally significantly associated with age, χ2(2) = 5.66, p = .056, although in this case the 6- and 7-year-olds gave more correct responses than each of the other groups. Analysis of interventions Subsequent analyses examined the nature of participants’ interventions on the system. We initially discriminated between whether an intervention was informative or not given the three possible causal structures. There were three interventions that were never informative: A+, B+C–, and B–C+, adopting the notation of “+” to mean that a particular shape was moved by the participant and “–” to mean that a shape was disabled. Potentially informative interventions were A+B–, A+C–, B+, C+, A–B+, and A–C+. We also classified interventions as simple or complex; A+, B+, and C+ were classified as simple, and those involving initially disabling one of the components before moving another component were classified as complex. Table 1 shows the percentage of times that participants in each age group chose each of these interventions. The most popular intervention tended to be A+, which, although it was uninformative, did make all of the three shapes spin. Propensity to select a complex intervention increased significantly with age, F(2, 74) = 7.22, p < .002, η2 = .16, with 7- and 8-year-olds being the most likely to pick the complex interventions (65% of the time vs. 46% for 5- and 6-year-olds and 45% for 6- and 7-year-olds). We examined whether participants chose informative interventions more often than chance by conducting a one-sample t-test with a test value of .67 given that two thirds of the nine possible interventions were informative. Only the 7- and 8-year-olds were significantly more likely than chance to select informative interventions, t(24) = 2.83, p < .01, both ps > .10 for the younger groups. A logistic regression showed that the proportion of informative interventions significantly predicted the probability of a participant getting the chain trial correct (z = 2.73, p < .01) (see Table 2), but this was not the case for the common cause trial (z = −0.62, p > .50). One potential explanation for the latter finding is that the children were overall more likely to select the common cause, doing so 56% of the time. Thus, some of the correct responses on the common cause test trial are likely to have been made by the weaker participants purely by virtue of their favoring the common cause structure. Modeling interventions So far, we have looked at proportion of informative intervention choices without considering the sequential nature of the task or whether and how efficiently children produced a set of informative interventions sufficient to discriminate between causal structures. A child who did not produce such a set but repeatedly produced a single informative intervention would score 100% on this measure. Moreover, how useful an intervention is depends on what learners already know (in this case what they have already learned from their previous interventions). For example, A+B– and C+ are both informative interventions in this task provided that you do not know anything yet. But suppose that you have already performed A+C– and observed that this made B spin. This evidence effectively rules out the ACB chain, leaving only the ABC chain and the common cause as possibilities. Now, on subsequent trials, performing C+, or repeating A+C–, will not tell you anything new because both of these interventions simply distinguish the ACB chain from the other two. To capture how efficiently children’s intervention choices allow them to home in on the true structure, we can analyze the interventions sequentially by looking at how effectively these interventions reduce uncertainty, assuming that initially children are perfectly able to remember past outcomes and integrate new information. For all age groups, for both structures, children’s interventions were significantly much more efficient than the chance level (mean efficiencies for chain and common cause, respectively: 5- and 6-year-olds = .71 and .66, 6- and 7-year-olds = .79 and .63, and 7- and 8-year-olds = .82 and .80; all ts > 6, ps < .0001). Children’s efficiency for the common cause changed significantly with age, F(2, 74) = 4.47, p ⩽ .02, η2 = .11, but there was no effect of age on efficiency on the causal chain trials, F(2, 74) = 1.14, p = .32, η2 = .03. Unlike proportion informative interventions, efficiency did not predict accuracy on the chain (see Table 2) and was in fact negatively related to accuracy on the common cause (z = –2.62, p < .01). Whereas proportion informative interventions did not take into account the sequential nature of the task, arguably interventional efficiency has the opposite shortcoming. By assuming, implausibly, that children have a perfect memory for the outcomes of their previous interventions and perfect ability to make inferences from this information, it ignores what they do on subsequent interventions once they have, in principle, enough information to potentially identify the correct structure. An inspection of the modeled data found that children obtained sufficient information for certainty after an average of only 2.75 interventions; this means that our measure of efficiency ignores a large proportion of the data and makes no allowances for noise, forgetting, or uncertainty in learning. A more balanced way to assess the quality of participants’ interventions is achieved by adding some noise, encapsulating the idea that learning is likely to be somewhat leaky or error prone. The exact level of “forgetting” in the model turned out not to be particularly important. We found qualitatively the same results setting it to 10%, 25%, 50%, 75%, and 90%, although the results were clearer for the lower levels of forgetting. Here we report results assuming 25% forgetting after each test. We established chance levels of intervention quality, again through simulation over 1000 trials. The 5- and 6-year-olds’ intervention quality was not significantly above chance on either the chain or common cause; the 6- and 7-year-olds were above chance on the chain, t(30) = 2.88, p < .01, and marginal on the common cause, t(30) = 1.73, p = .09; and the 7- and 8-year-olds’ intervention quality was above chance for both trials (chain: t(24) = 2.46, p < .02; common cause: t(24) = 4.87, p < .001). Averaged over trial types, we found that intervention quality improved with age, F(2, 74) = 4.03, p < .03, η2 = .10, with 7- and 8-year-olds significantly more efficient that 5- and 6-year-olds (p < .01), but no significant difference between 5- and 6-year- olds and 6- and 7-year-olds. Breaking this into responses for the two structures, regardless of forgetting rate, intervention quality was a significant predictor for correct identification of the chain structure (z = 2.61, p < .01), but not for the common cause structure (see Table 2).","Our findings provide important information about developmental changes in children’s ability to learn causal structure through intervention. Children’s ability to learn a causal chain structure improved developmentally, with the youngest children not managing to learn this structure at above-chance levels. However, we need to consider why the 5- and 6-year-olds correctly identified the common cause structure as accurately as the 7- and 8-year-olds. Our view is that the good performance on this second trial type is due to a tendency even among the youngest children to assume that, when events happen simultaneously, the underlying structure is a common cause. Previous studies have found that both children and adults make use of this simple temporal heuristic when they observe a three-variable system with this sort of temporal schedule (Burns & McCormack, 2009; Fernbach & Sloman, 2009; Lagnado & Sloman, 2006; McCormack et al., 2015). Indeed, McCormack et al. (2015) demonstrated that children will use this type of temporal heuristic even when faced with contradictory statistical information provided either through observing the operation of a probabilistic causal system or through observing the effects of interventions on a deterministic system. Thus, the good performance of the younger children on the common cause structure is likely to reflect use of this heuristic rather than use of statistical information derived from interventions on the system. This would also straightforwardly explain the lack of a relation between intervention quality and performance on the common cause structure. The analyses of children’s intervention choices provide insights into why performance improved developmentally on the causal chain trial. Interventions could be initially classified as informative or non-informative given the three possible causal structures. Over all trials, unlike the oldest group, younger groups of children did not choose informative interventions more often than chance. It proved to be fruitful, however, to further examine intervention choices and how these related to performance in more detail by means of our modeling work. The initial analysis of how efficient participants were at producing a set of interventions that could in principle discriminate between the different hypotheses showed that all groups of children produced such a set more quickly than would be expected if they were simply choosing between interventions at random. This means that even the youngest children had the evidence available to them to make the appropriate causal inferences. However, intervention efficiency was not a predictor of performance. This demonstrates that good performance does not hinge on simply initially choosing interventions that are as a matter of fact disambiguating. Children may forget or fail to make use of what they have observed, and the subsequent interventions they make may also influence their judgments. Our modelling work suggested that this was indeed the case because our measure of the quality of children’s interventions that took into account the complete sequence of interventions predicted performance on the causal chain structure, under the assumption that there was some degree of forgetting. Moreover, unlike efficiency, intervention quality improved with age, with older children being more likely to consistently choose interventions that would help to disambiguate the causal structures given what they had already observed. These results indicate that with development children become more discerning in their choice of interventions, and this has an impact on their causal structure learning. How do our findings fit with what is already known about developmental changes in children’s use of interventions to learn about causal systems? In designing our study, we sought to ensure that domain-specific knowledge was not relevant for task performance. However, this did not rule out children exploiting a type of preexisting, albeit domain-general, heuristic about the nature of the causal system, namely that when multiple events occur immediately following an intervention, the underlying structure is likely to be a common cause (McCormack et al., 2015). When the evidence generated from interventions was consistent with this assumption (i.e., in the common cause trial), even young children performed well. However, younger children had difficulty in discarding this assumption on the basis of the contradictory evidence provided by their interventions. This is consistent with evidence from the scientific learning literature indicating that children have difficulty in discarding a preexisting hypothesis and may routinely ignore statistical evidence that fails to support such a hypothesis (Amsel & Brock, 1996; Kuhn et al., 1988). Furthermore, an inspection of developmental changes in the pattern of children’s intervention choices (Table 1) yields some further interesting additional parallels with findings from the scientific learning literature. Our task is very different from those used in research on children’s scientific learning; it is simpler, and the children we tested are younger than those typically used in such studies (but see Koerber, Sodian, Thoermer, & Nett, 2005; Piekny, Grube, & Maehler, 2014). Nevertheless, some of our findings confirm broad developmental patterns that are well established in that research. Younger children tended to prefer making the A+ intervention and did so repeatedly. This intervention is the most causally effective (it makes all of the events happen) but does not discriminate among the three available hypotheses. However, it reinforces any existing hypothesis that the causal structure is a common cause by providing the temporal pattern of all events happening simultaneously. Young children’s preference for this intervention has parallels with demonstrations in the scientific learning studies showing that children attend most to the variable already believed to be causal, focus more on producing an effect than on generating disambiguating evidence, and produce evidence that is consistent with their existing hypothesis rather than seeking to disconfirm it (e.g., Klahr & Dunbar, 1988; Klahr et al., 1993; Kuhn, 1989; Schauble, 1990, 1996). Although younger children’s patterns of interventions led to poorer performance, it is interesting to note that recent formal analyses have demonstrated that whether their type of approach should be viewed as inefficient depends on the learning context. First, the tendency to intervene on variables already believed to be causal in order to confirm an existing hypothesis is not necessarily always the wrong strategy. This type of strategy has been shown to be rational under the assumption that causal connections in the world are sparse (Navarro & Perfors, 2011), meaning that competing causal hypotheses do not generally share the same effect variables. In such circumstances, “positive tests,” operationalized as intervening on the variable thought to be the root cause (Coenen, Rehder, & Gureckis, 2015), are highly diagnostic. Hence, younger children’s pattern of interventions could be interpreted as due to a tendency to act in a way that has proved to be an effective general-purpose method for learning causal relationships in the past but is not appropriate given the specific learning context in which they find themselves. Second, we also found that younger children were less likely than older children to produce the more complex interventions that involved disabling one of the components in the system. This type of intervention can be particularly informative because it can be used to exclude a variable as being necessary for production of an effect. However, separate Bayesian modeling work with adults has demonstrated that producing simple rather than complex interventions is not always an inefficient strategy. Bramley, Lagnado, et al. (2015) showed that simple interventions tend to be more informative than complex interventions with respect to a broader hypothesis space (e.g., all possible three- variable causal models), with more complex interventions becoming more useful once the space of possibilities narrows to favor only a few overlapping causal hypotheses. In our task, children needed to discriminate among just three competing hypotheses, so it is one in which complex interventions are likely to be useful. In summary, the observed developmental changes can be interpreted as supporting the idea that whereas younger children used simple strategies that may have proved to be effective in other contexts, older children were more able to adjust their learning strategy in a way that was appropriate for the task—that is, to use a control of variables strategy (Chen & Klahr, 1999; Dean & Kuhn, 2007) in which confounding variables are experimentally controlled.","Experiment 1 showed that by 6 or 7 years of age children can generate informative interventions and use these to derive the structure of a three-variable causal system and that the quality of their interventions predicts their performance. What we have not yet shown, however, is that children benefited from generating interventions themselves. Following Sobel and Kushnir’s (2006) study with adults and Sobel and Sommerville’s (2010) related study with children, we might predict that being able to self-generate interventions facilitates active hypothesis testing. In our second experiment, we tested additional groups of children using a “yoking” procedure similar to Sobel and Kushnir’s procedure. Children did not select interventions themselves; rather, each child was individually matched with one of the children from Experiment 1 and saw the outcomes of the set of interventions made by that child. There were two conditions; children in the yoked–self condition carried out a series of interventions on the system but did not select which interventions to make, and children in yoked–observe condition simply watched interventions being carried out by the experimenter. If children benefit from selecting interventions themselves, we would expect performance to be worse in the yoked conditions compared with the self-generated interventions in Experiment 1. If there are benefits from simply carrying out interventions oneself, even if one has not selected them, we would expect performance in the yoked–self condition to be better than performance in the yoked–observe condition. In addition to comparing performance across conditions, we examined whether our measures of intervention quality, taken from the interventions made by the yoking participants from Experiment 1, were predictive of performance even in the yoked conditions. It is important to note that all participants in Experiment 2 received sufficient information to disambiguate the causal structures. Because all children in Experiment 1 obtained sufficient information to make correct inferences, the relation between intervention quality and performance found in that experiment can be explained in two potentially independent ways. It may be that this relation was obtained because higher quality interventions resulted in a larger quantity of useful information that was obtained in a sequence that facilitated learning. Alternatively, it may be that this relation was obtained because children who were engaged in productive hypothesis testing selected useful interventions that allowed them to assess these hypotheses. If the former is the case, we might expect to see a relation between intervention quality and performance in the yoked conditions of Experiment 2 because children in those conditions received identical sequences of information to the children who selected the interventions themselves. However, if this relation hinges on the role of active hypothesis testing, it might not be obtained in circumstances where children do not select interventions themselves.","Three groups of children took part in this part of the study: 28 5- and 6-year- olds (M = 75 months, range = 70–81), 62 6- and 7-year-olds (M = 87 months, range = 81–93), and 50 7- and 8-year-olds (M = 99 months, range = 91–105), who were recruited from the same year groups and from schools in the same area as the child participants in Experiment 1. Half of the children were assigned to the yoked–self condition and half to the yoked–observe condition.","These were identical to those used in Experiment 1. Procedure The procedure for each of the conditions was identical to that of Experiment 1 up until the learning stage of the task. In the yoked–self condition, the experimenter explained that she was going to ask children to do things to the shapes in the box in order to figure out how the box worked, using cards to give instructions. She explained that when she pointed to a picture of a specific shape, children needed to move that shape to see what happened to the other shapes, and when she pointed to a picture of a shape with a stop sign in it, children needed to initially put the stop sign in that shape and then see what happened when they moved whichever shape was depicted in the next picture that she pointed to. Children were told that the experimenter would give them 12 instructions of this sort to start with. During the learning phase, the experimenter then instructed children to carry out the interventions in the same order as their yoked participants from Experiment 1. For those children who were yoked to children from Experiment 1 who completed 12 interventions (the majority of children), after those interventions were completed the experimenter said, “You have had your 12 goes now. Which picture do you think shows how the box works?” For the remaining children, after 12 interventions the experimenter said, “We have 6 more things to try before you give your guess about how the box works,” and then showed the remaining 6 interventions before asking the test question. The procedure for the yoked–observe condition was very similar except that the experimenter explained that she would point to the pictures of the shapes herself before moving them. During the learning phase, the experimenter then pointed to the appropriate pictures before each intervention and carried out all interventions herself with children observing.","Fig. 4A shows the distribution of responses for each trial type for the yoked–self condition, and Fig. 4B shows the distribution for the yoked–observe condition. As can be seen from the figure, correct responses were relatively similar to those in Experiment 1, although the 7- and 8-year-olds seemed to perform less well on the causal chain. We compared performance in these conditions with performance from the yoked participants from Experiment 1; Table 3 shows the percentages of correct responses for the causal chain and common cause trials as a function of condition. There was no significant association between condition and numbers of correct responses for the causal chain trial, χ2(2) = 1.48, p = .48, or for the common cause trial, χ2(2) = 3.89, p = .14; the association between condition and numbers of correct responses for each trial type was also not significant if each age group was examined separately or if each condition was compared separately with the other two conditions. We then examined whether the factors that were predictive of performance in Experiment 1 predicted performance in the yoked conditions. Neither the proportion of informative interventions presented to participants, the efficiency of the sequences of interventions for identifying the true structure, nor their overall quality (assuming 25% forgetting) predicted performance for either the yoked–self or yoked–observe groups on either trial type (all ps > .05). The relation between these measures of interventions and performance on the causal chain identified in the 77 participants in Experiment 1 were still significant for the 70 whose data were yoked to the yoked–self and yoked–observe groups (informative interventions: z = –2.69, p < .005; intervention quality: z = –2.50, p < .01), indicating that the fact that the measures do not predict performance in the yoked groups is not a problem of slightly reduced statistical power.","Experiment 1 examined children’s causal structure learning under circumstances in which they selected and carried out interventions on a simple three-variable causal system. To the best of our knowledge, the analyses reported here of its data constitute the first attempt to model the quality of children’s interventions when learning causal structure within a Bayesian framework. Our findings regarding children’s interventional learning varied depending on whether children were learning a causal chain or a common cause structure. With regard to the former, there were clear developmental improvements not only in terms of accuracy of structure learning but also in terms of the quality of the interventions that children produced, as assessed in our Bayesian modeling. The key advantage of the modeling is that it provided us with a quantitative measure of the quality of children’s interventions, allowing us to formally examine the extent to which children’s interactions with the system were optimal. This Bayesian measure of intervention quality predicted performance. Put simply, the findings suggest that with development children increasingly resemble idealized Bayesian learners, although we note that the best predictor of performance from our modeling results was a measure of interventional quality that assumed some degree of noise in the Bayesian learning process. The same pattern of findings was not obtained for the common cause structure, and the most plausible interpretation of this is that younger children’s inferences in this task were based on a simple temporal heuristic (“assume a common cause if effects happen simultaneously”) rather than on use of statistical information provided from interventions. Use of such temporal heuristics is widespread in both children’s and adults’ causal structure learning (Burns & McCormack, 2009; Fernbach & Sloman, 2009; Lagnado & Sloman, 2004, 2006; White, 2006), with McCormack et al. (2015) demonstrating that younger children’s causal structure inferences are highly influenced by the temporal pattern of events. Their findings are consistent with those from the current study insofar as those authors also found no developmental improvements in the likelihood that children would give a common cause judgment under circumstances in which all events happened simultaneously. Children’s tendency to recruit temporal heuristics is likely to be due to the heuristics’ low demands on information processing in comparison with using statistical information (Fernbach & Sloman, 2009). For example, in the current study, use of such a heuristic would have been based on the observation of a single intervention—the temporal pattern of events following A+. This intervention was the most common one made by the younger groups; we interpreted this as suggesting that these children focus on producing an effect rather than systematically testing the competing hypotheses and in doing so are provided with evidence (i.e., the temporal pattern of events) that they take to be consistent with their existing hypothesis. Younger children were also less likely to disable components in the system, suggesting that they were less likely to try to exclude any variables. Although younger children’s interventions in the system had these characteristics, all children produced a set of interventions that could in principle have allowed them to correctly judge the causal structure. However, the Bayesian analysis proved to be useful in establishing that simply initially producing interventions that could potentially disambiguate the causal structure was not predictive of good performance. Rather, children’s performance was related to how informative their interventions were as they moved through the task sequentially, with the Bayesian modeling making it possible to operationalize the informativeness of sequential intervention choices. In general, as we argued above, our developmental findings are broadly consistent with some well-established findings in the literature on scientific learning. One of the aims of the scientific learning literature is to explore the wide variety of cognitive and metacognitive processes that are important at each stage of the reasoning process, from formulating initial questions, generating and evaluating evidence, to constructing and revising theories. Our aim in this study was not to match such studies in the breadth of cognitive processes that they explore or in the depth of analysis that they typically provide regarding, for example, the range of strategies that participants recruit. For a start, although we have described children as deciding between hypotheses, we do not believe that it is helpful to consider children to be engaged in theory construction in our task. Nevertheless, our findings are valuable in terms of what they add to our knowledge specifically about basic aspects of causal structure learning. As already emphasized, in many instances scientists must uncover not just which variables are causal but also the overall structure of a causal system. Most studies of children’s causal and scientific learning have focused on their ability to learn whether specific variables are causally efficacious. Although in scientific learning studies such variables sometimes have additive or interactive effects (e.g., Kuhn & Pease, 2008; Kuhn, Pease, & Wirkala, 2009), typically participants do not need to distinguish between, for example, common cause and causal chain structures. It is an important strength of the causal Bayes net approach that, unlike more traditional models of causal learning, it can model this sort of learning of structure. Our modeling work demonstrates the utility of examining such learning within a broadly Bayesian framework, albeit to allow us to conclude that younger children might not be the idealized Bayesian learners they are sometimes assumed to be within the causal Bayes net tradition. Active versus passive learning ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our use of a yoking procedure in Experiment 2 provided insights into whether active participation assists children’s causal learning. We found no clear learning benefits of either selecting or carrying out interventions oneself, a finding that contrasts with that of Sobel and Kushnir’s (2009) adult study but is consistent with the findings of Lagnado and Sloman (2004). We note that there is a long-standing debate in the literature on children’s scientific learning over whether “hands-on” learning is more beneficial than more passive learning (e.g., Bruner, 1961; Kirschner, Sweller, & Clark, 2006; Mayer, 2004). The debate regarding active versus passive learning is multifaceted, encompassing both cognitive and motivational issues, but in the context of children’s scientific learning the central question has been whether children show benefits in terms of the amount, depth, and persistence of knowledge gained when allowed to engage in self-guided discovery learning. So, for example, Klahr and Nigam (2004) compared how well children learned a control-of-variables strategy in two conditions: either direct instruction or a condition in which they designed and carried out experiments themselves. On the basis of their findings, these authors argued that direct instruction is more beneficial than discovery learning. Unsurprisingly, such a conclusion has attracted much attention because it is believed to have important consequences for how children should be taught science (Hmelo-Silver, Duncan, & Chinn, 2007; Kuhn, 2007). In interpreting the contribution of our findings to this debate, it is important to remember that our task is much simpler than those typically used in studies of the benefits of active learning. Moreover, children did not need to learn a theory or generalize from their learning, and the task was not one in which we could contrast hands-on learning with, for example, teacher instruction. Nevertheless, our task provided a context in which we could examine whether there were measurable benefits to actively selecting one’s own interventions in circumstances where children need to decide between different hypotheses. In the self-select condition, unlike in the other two conditions, children freely decided what information to sample and when to sample that information; this freedom with regard to information sampling can be seen as the most important difference between active and passive learning conditions (Gureckis & Markant, 2012). We found no evidence that such freedom yielded benefits. This suggests that, at least on the face of it, in learning causal structure, children do not benefit from being able to actively generate interventions in order to test specific hypotheses that they may be entertaining, as opposed to either just observing such interventions or carrying out interventions that were decided by another person. However, the results also indicate a more nuanced conclusion because we examined not just whether levels of performance differed depending on learning condition but also whether the quality of interventions predicted performance even in yoked conditions. Interestingly, we found that this predictive relation did not hold in our yoked conditions. This finding is reminiscent of the results of correlational analysis conducted by Sobel and Kushnir (2006) on their adult data. They calculated the proportion of interventions that participants self- generated (or were exposed to in yoked conditions) that would provide information that was critical for learning a particular causal structure (i.e., a simple measure of intervention quality). They found that the correlation between the proportion of critical interventions and accuracy in causal structure judgments was significant only in conditions where participants generated the interventions themselves, not in yoked conditions. Their interpretation of this finding is that participants who generate critical interventions are better placed to interpret the outcome that results from an intervention and use it to distinguish between hypothesized structures. Although we did not find that self-generating interventions boosted learning, our findings are at least compatible with Sobel and Kushnir’s (2006) suggestion that learning may proceed in a different way under such conditions. When selecting interventions themselves, children who were testing hypotheses in an effective way were able to select informative interventions specifically to test particular hypotheses (consistent with Sobel & Kushnir’s characterization of learning under this condition), and this may have underpinned the relation between quality of interventions and performance. Gureckis and Markant (2012) argued that when participants are selecting for themselves, they can preferentially select information that reduces their current uncertainty. Indeed, the measure of quality of interventions produced in our Bayesian modeling can be seen as a formal measure of the extent to which participants are adept at selecting interventions to reduce uncertainty, and it is this that predicted performance in the self-select condition. However, this mode of learning was not available to children in the yoked conditions, which may be why there was no relation between quality of interventions and performance in these conditions. Note, however, that the unavailability of this mode of learning in the yoked conditions did not significantly impair performance. It is possible that the benefits of active hypothesis testing were outweighed in our study by the information processing demands of selecting and implementing interventions, demands that were reduced in the yoked conditions. Such a possibility is consistent with suggestions that there may be disadvantages associated with self-directed learning (for discussions, see Gureckis & Markant, 2012; Kirschner et al., 2006; Markant & Gureckis, 2014). Moreover, as Gureckis and Markant (2012) pointed out, active learning might not necessarily be beneficial under circumstances where learners are biased in their information seeking and focus, for example, on confirming preexisting erroneous hypotheses. As we have discussed, there is evidence that the younger children adopted such an approach to their intervention choices. There may be other circumstances, however, where children’s causal learning would benefit from the opportunity to engage in active hypothesis testing, but the broader literature on self-directed learning has yet to clearly identify the situations where any advantages may outweigh potential disadvantages (Gureckis & Markant, 2012).","The findings of this study point to two clear directions for future work in this area. First, the fact that children become increasingly Bayes-efficient information seekers in their causal learning raises the question of what cognitive changes underpin this developmental shift. There has been considerable discussion over the role of Bayesian modeling not only in a developmental context but also in providing an explanatory framework within cognitive psychology more generally (e.g., Bowers & Davis, 2012; Jones & Love, 2011; Schlesinger & McMurray, 2012). Here, we have used a Bayesian approach simply to provide a more formal analysis of how the quality of children’s interventions improves developmentally. Questions regarding the developmental changes that account for these improvements cannot be directly addressed by the modeling reported here, which does not model cognitive processes (although we note that directly inspecting the patterns of children’s interventions suggested that some developmental changes have parallels to those shown in the scientific learning literature). It is possible, however, that in the future models based on approximate Bayesian inference that attempt to be more psychologically plausible (Bramley, Dayan, et al., 2015; Kemp, Tenenbaum, Niyogi, & Griffiths, 2010; Sanborn, Griffiths, & Navarro, 2010; Shi, Feldman, & Griffiths, 2008) may play a role in addressing this question. Importantly, the developmental improvements found in our study highlight the need for Bayesian models that not only capture idealized learning but also can accommodate, and potentially explain, developmental changes in the quality of children’s causal learning. Explaining developmental changes will require additional research that builds on the current findings but also tries to examine in more detail the role of particular types of cognitive processes such as children’s strategies. The second direction for future research stems from our finding that although active learning, in the form of self-selecting interventions, did not seem to benefit children’s causal structure learning (beyond simply observing someone else’s interventions), our modeling raised the possibility that learning may proceed in a qualitatively different way when children do not have the opportunity to choose interventions themselves. This finding demonstrated the potential benefits of examining not just whether causal learning is affected by whether or not individuals can engage in more active learning but also how it may be differentially related to the sort of information generated in active learning. Further examination of potential qualitative differences looks like an important goal for future empirical and modeling work on causal structure learning."],["Sex differences on the WISC-R in Chinese children were examined in a sample of 788 aged 12. years. Boys obtained a higher mean full scale IQ than girls of 3.75 IQ points, a higher performance IQ of 4.20 IQ points, and a higher verbal IQ of 2.40 IQ points. Boys obtained significantly higher means on the information, picture arrangement, picture completion, block design, and object assembly subtests, while girls obtained a significantly higher mean on coding. The results were in general similar to the sex differences in the United States standardisation sample of the WISC-R. Boys showed greater variability than girls. --------------------------------------------------------------------------------","The question of sex differences in intelligence has been debated from the early years of the twentieth century. The almost unanimous consensus has been that there is no sex difference in “general intelligence” defined as the average of the major cognitive abilities and measured by tests like the Wechslers and the Binets. There are, however, sex differences in a number of specific abilities. The conclusion that there is no sex difference in “general intelligence” was reached in the second decade of the twentieth century by Terman (1916, pp. 69–70) on the basis of his American standardisation sample of the Stanford–Binet test. In recent decades this conclusion was endorsed by many leading authorities. Thus “it is now demonstrated by countless and large samples that on the two main general cognitive abilities – fluid and crystallized intelligence – men and women, boys and girls, show no significant differences” (Cattell, 1971, p. 131); “gender differences in general intelligence are small and virtually non-existent” (Brody, 1992, p. 323); “there is no sex difference in general intelligence worth speaking of” (Mackintosh, 1996, p. 567); and “sex differences have not been found in general intelligence” (Halpern, 2000, p. 218). The only challenge to this consensus has come from Lynn (1994, 1998, 1999), who has argued that males have larger average brain size than females, that brain size is positively correlated with intelligence at a magnitude of approximately .40 (Vernon, Wickett, Bazana, & Stelmack, 2000), and hence that there is a theoretical expectation that males should have higher average intelligence than females. To examine this theoretical expectation, Lynn (1994) proposed that the Wechsler intelligence tests could be taken as among the best measures of general intelligence on the grounds that they provide measures of the major cognitive abilities of verbal, numerical, perceptual, reasoning, spatial, immediate memory, perceptual speed and general knowledge. He then examined the sex difference in eight standardization samples of the Wechsler intelligence tests for children aged 6–16 and showed that boys obtained a higher mean Full Scale IQ by an advantage of 2.25 IQ points. He also showed that in six standardization samples of adults, men obtained a higher mean Full Scale IQ by an average of 3.08 IQ points. Despite these results, it has continued to be asserted that “females and males score identically on IQ tests” (Halpern, 2012, p. 233) and that “there is no evidence, overall, of sex differences in levels of intelligence” (Sternberg, 2014, p.178). However, Ellis et al. (2008) recently argued in their book that studies have shown that, although small, there are significant sex differences in intelligence over the years throughout the world. It has also been consistently asserted for approximately a century that while males and females have the same average intelligence, males have greater variability of intelligence than females. An early first statement of this proposition was made by Ellis (1904, p.425): “It is undoubtedly true that the greater variational tendency in the male is a psychic as well as a physical fact”. This sex difference in variability was reaffirmed by Thorndike (1910) and by many subsequent authorities including Penrose (1963, p. 186) “the consistent story has been that men and women have nearly identical IQs but that men have a broader distribution…the larger variation among men means that there are more men than women at either extreme of the IQ distribution”. Others who have asserted this conclusion include Herrnstein and Murray (1994, p. 275), Lehrke (1997, p. 140), Jensen (1998, p. 537), and Ceci and Williams (2007, p. 223): “all sides in the gender wars agree that there is greater variability in male distributions of many abilities.” This conclusion has recently been affirmed once again by Deary, Penke, and Johnson (2010): “Males have a slight but consistently wider distribution than females at both ends of the range.” There is a large research literature on sex differences in various cognitive abilities. Kimura (1999, 2002) lists five abilities on which males obtain higher average means than females: spatial orientation, visualization, line orientation, mathematical reasoning, and throwing accuracy; and five abilities on which females obtain higher average means than males: object location memory, perceptual speed, verbal memory, numerical calculation, and manual dexterity. In this paper we examine sex differences in intelligence in China with a view to determining how far these are consistent with those found in studies in the United States and other western countries on which most of the conclusions have been based.","A Chinese sample of 788 children in the sixth grade aged 11–13 years with a mean age of 12.5 was tested with the Chinese version of the Wechsler Intelligence Scale for Children- Revised (WISC-R) in 2011–13. The sample was obtained from the Jintan Child Cohort Study. The Wechsler (1974) original sample in this study comprised 1656 Chinese children (55.5% boys, 44.5% girls) consisting of 24.3% of all children in this age range in the Jintan region, which is located in Jiangsu, China. This sample includes children from city, town, and village populations; in addition, the demographics of Jiangsu are similar to those found on the national level, making this sample likely to be fairly representative in terms of sex ratio, urban versus rural population ratio, ethnic majority, and others. The Jintan Child Cohort Study is an on-going prospective longitudinal study with the aim of exploring early health risk factors in the development of child cognition and behavior. Details of this study have been described in a previous publication (Liu, McCauley, Zhao, Zhang, & Pinto-Martin, 2010). The cohort took their first IQ test (the WPPSI) at age 6 years and the sex differences have been given by Liu and Lynn (2011, 2013) and Liu, Yang, Li, Chen, and Lynn (2012). The Chinese version of the Wechsler Intelligence Scale for Children-Revised (WISC-R) with which the children were tested was standardized in China in 1985 and has shown good reliability in Chinese children (Yue & Gao, 1987). The WISC-R consists of six verbal subtests, namely Information, Comprehension, Arithmetic, Vocabulary Similarities and Digit Span, that are summed to give the Verbal IQ, and of six non-verbal subtests, namely Picture Arrangement, Picture Completion, Object Assembly, Block Design, Coding and Mazes tests, that are summed to give the Performance IQ. The Verbal IQ and Performance IQs are combined to give a Full-Scale IQ. The WISC-R IQs of the 2011–13 Chinese sample were collected between spring 2011 and summer 2013 when the participants were in sixth grade or had just graduated from sixth grade. Participants were invited to the laboratory, where research assistants, who participated in an intensive training course, administered the Chinese WISC-R. Ten of the subtests were used, Digit Span and Mazes being omitted. The research assistants were supervised by a Ph.D. trained clinical psychologist who specializes on cognitive brain assessment at Nanjing Brain Hospital. The same training procedure as described in detail in Liu and Lynn (2013) was followed. The IQ test was administered over the course of one hour in a quiet room in Jintan Hospital. Each test was scored by two individuals to minimize scorer bias. This procedure for data collection was approved by the research ethics committee of Jintan Hospital and the University of Pennsylvania. Written consent was obtained from parents and written assent from children was collected prior to initiation of the study.","Table 1 gives the mean scaled scores and standard deviations for boys and girls on the subtests, and the verbal, performance and full scale IQs on the Chinese WISC-R of the 2011–2013 Jintan sample. Also given are the differences between the means of the boys and girls expressed as ds (the difference between the means divided by the pooled standard deviation, with minus signs showing that girls obtained higher means than boys), the t values using independent sample t-tests for the statistical significance of the differences between the means of the boys and girls, and the variance ratios (VR) as a measure of the sex differences in variability calculated as the standard deviation of the males divided by the standard deviation of the females. Thus, a VR greater than 1.0 indicates that males had greater variance than females. Table 2 gives sex differences on the WISC-R in China and in the standardization sample (N = 2200) in the USA given by Jensen and Reynolds (1983).","The results provide six points of interest. First, it is shown in Table 1 that in the present Chinese sample boys obtained a significantly higher Full Scale IQ than girls by 0.25d, the equivalent of 3.75 IQ points. This figure is higher than the average boys’ advantage of 2.25 IQ points on the Wechsler Full Scale IQ in eight standardization samples of the Wechsler tests for children noted in the introduction. In the present Chinese sample boys obtained a significantly higher Performance IQ than girls by 0.28d, the equivalent of 4.2 IQ points, and a significantly higher Verbal IQ than girls by 0.16d, the equivalent of 2.40 IQ points. These results provide additional evidence that modest but significant sex differences exist in intelligence, thus refuting continued assertions that no differences exist (e.g., Halpern, 2012, p. 233; Sternberg, 2014, p. 178) Second, there are six statistically significant sex differences on the subtests of the WISC-R in the present Chinese sample shown in Table 1. Boys obtained significantly higher means than girls on Information, Picture Arrangement, Picture Completion, Block Design and Object Assembly, and girls obtained a significantly higher mean than boys on Coding. Third, on several of the subtests, the sex differences in the present Chinese sample were consistent with those in the American standardization sample shown in Table 2. The advantage of boys in the present Chinese sample on Information is virtually identical to that in the United States with statistically significant ds of .44 and .37, respectively. These results confirm those of several studies of the Wechsler information tests among adults and of other studies finding that among adults men have significantly higher means than women on information and general knowledge (Lynn & Irwing, 2002; Lynn, Irwing, & Cammock, 2002). The advantage of boys in the present Chinese sample on Picture Arrangement is consistent with that in the American standardization sample with statistically significant ds of .19 and .11, respectively. The advantage of boys on Object Assembly in the present Chinese sample is also consistent with that in the United States with statistically significant ds of .38 and .18, respectively. Boys obtained higher scores on Picture Completion in the present Chinese sample (.19) and in the U.S. (.15) and on Block Design with ds of .19 and .15, respectively. The higher means obtained by boys in both China and the United States on Picture Arrangement, Object Assembly, Picture Completion and Block Design are explicable because these are all measures of visual–spatial abilities on which males typically obtain higher means than females (Linn & Peterson, 1985; Voyer, Voyer, & Bryden, 1995). The statistically significant advantage of girls on Coding in the present Chinese sample is consistent with the higher mean obtained by girls in the United States with ds of .41 and .53, respectively. Fourth, on Comprehension and Similarities the sex differences in the present Chinese sample, where both boys and girls score similarly at .00 and −.01, is consistent with those in the American standardization sample with ds of .01 and .07. Fifth, there is some inconsistency in the sex difference in Vocabulary, where there was no significant difference between boys and girls in the Chinese sample (d = −.03) but boys obtained a significantly mean in the American sample (d = .14). Sixth, the frequent assertion that males have greater variability of intelligence than females is generally confirmed in the present Chinese sample. Boys had greater variability than girls on the Verbal, Performance and Full Scale IQs and in six of the ten subtests. However, girls had greater variability than boys in Comprehension, Vocabulary and Block Design, and there was no difference in the variability of boys and girls on Similarities. Future studies might consider controlling for sociodemographic variables to further validate this finding."],["Convergent research points to the importance of studying the ontogenesis of sustained attention during the early years of life, but little research hitherto has compared and contrasted different techniques available for measuring sustained attention. Here, we compare methods that have been used to assess one parameter of sustained attention, namely infants' peak look duration to novel stimuli. Our focus was to assess whether individual differences in peak look duration are stable across different measurement techniques. In a single cohort of 42 typically developing 11-month-old infants we assessed peak look duration using six different measurement paradigms (four screen-based, two naturalistic). Zero-order correlations suggested that individual differences in peak look duration were stable across all four screen-based paradigms, but no correlations were found between peak look durations observed on the screen-based and the naturalistic paradigms. A factor analysis conducted on the dependent variable of peak look duration identified two factors. All four screen-based tasks loaded onto the first factor, but the two naturalistic tasks did not relate, and mapped onto a different factor. Our results question how individual differences observed on screen-based tasks manifest in more ecologically valid contexts. © 2014 The Authors. --------------------------------------------------------------------------------","Research is increasingly suggesting that early-developing, domain-general aspects of attentional control may mediate subsequent skill acquisition in a variety of areas (e.g. Heckman, 2006; Karmiloff-Smith, 1998; Wass, Scerif, & Johnson, 2012). For example, aspects of domain-general attentional control have been shown to predict, on starting school, children's’ subsequent learning on literacy and numeracy tasks (e.g. Welsh, Nix, Blair, Bierman, & Nelson, 2010). And research into the development of attentional control within clinical disorders suggests that early disruption to attentional control may play a key role in impairing early learning in social settings, for example during word learning, leading to subsequent catastrophic developmental cascades (e.g. Karmiloff-Smith, 1998). This suggests the importance of researching the ontogenesis of attentional control during the first few years of life. Cohen suggested that infant attention involves at least two different mechanisms: an attention-getting process which determines whether an individual will orient towards a stimulus presented in his periphery, and an attention-holding process which determines how long his attention will be maintained once he fixates (Cohen, 1972). This second phase, the attention-holding process, is commonly described as ‘sustained attention’ (Richards, 2011). However, although individual differences in attention are frequently reported in applied and developmental psychology, the terms used are rarely precisely defined and are conventionally assessed using a variety of methods. Historically, the most widely used technique for measuring infants’ looking behaviour involves presenting static stimuli using a slide projector or computer screen across a number of discrete but contiguous trials; the infant's viewing behaviour is coded either live by an experimenter viewing the infant on a video feed, or post hoc (Colombo & Mitchell, 2009). Two variables are typically derived: peak look duration, the duration of the longest unbroken look to the screen, and habituation rate, i.e. the rate of change of looks over time. Colombo and Mitchell argued in favour of peak look duration as the better metric of individual and developmental differences in visual attention during infancy because it is more reliable, and shows more robust relationships with long-term cognitive outcomes (Colombo & Mitchell, 1990). Previous research has demonstrated that peak look duration to novel, static, screen-based stimuli show a U-shaped trajectory over the first year of life (Colombo & Mitchell, 2009; Colombo & Cheatham, 2006; Courage, Reynolds, & Richards, 2006). Research has also robustly demonstrated that peak look duration to novel stimuli during the first year of life relates negatively with long-term cognitive outcomes: shorter look duration during the first year is associated with better performance on later IQ and language measures (Colombo, 1993; McCall & Carriger, 1993; Tami-LeMonda & Bornstein, 1989) and recognition memory (Rose, Feldman, & Jankowski, 2003a, 2003b). Shorter looking is also associated with higher pre-existing knowledge bases and general arousal levels (de Barbaro, Chiba, & Deak, 2011; Dixon & Smith, 2008). An alternative technique for assessing looking durations during infancy involves presenting dynamic stimuli on a computer screen (Courage et al., 2006; Shaddy & Colombo, 2004; see Richards, 2010 for a review). This work has generally used either TV clips (e.g. Richards & Anderson, 2004) or specially filmed naturalistic or semi-naturalistic dynamic scenes (Wass, Porayska-Pomsta, & Johnson, 2011). These techniques have been used to investigate how autonomic indices change in different attention states (Richards, 2011; Richards & Cronise, 2000), how looking behaviour towards the screen changes over time (Anderson, Choi, & Lorch, 1987; Richards & Anderson, 2004), and how these changes are different in children with Attention Deficit Hyperactivity Disorder (ADHD) (Lorch et al., 2004). To our knowledge, no research has investigated whether individual differences in look duration are consistent across static vs. dynamic looking time paradigms. A third paradigm that has been used to assess looking durations involves presenting a number of unfamiliar objects consecutively or concurrently in a table-top setting, and performing video coding post hoc to analyse looking behaviour. For example, Kannass and Oakes (2008) videoed 9-month-old and 31-month-old infants playing with toys, in both single-object (objects presented consecutively) and four-object (objects presented concurrently) conditions; they also measured 31-month language performance in the same children (see also Sarid & Breznitz, 1997). They found that shorter look durations in the single-object task correlated with larger vocabularies at 31 months (Kannass & Oakes, 2008). For the multiple object condition, however, they found the opposite relationship: longer durations at 9 months correlated with larger vocabularies at 31 months (see also Choudhury & Gorman, 2000). Despite the strong face similarities between these paradigms, no previous research has assessed whether individual differences using one type of looking time paradigm are consistent across different assessment techniques. A number of studies have addressed this indirectly, but none directly. Kagan and Lewis (1965) examined the relationship between looking behaviour towards static stimuli at 6 and 13 months and the amount of free-play locomotor activity at 13 months, and found that infants with long fixation times at 6 and 13 months were more sedentary during free play. Coldren found that infants’ attention to stimuli in laboratory tasks correlated with the attention to their caregiver in face-to- face interactions at 3- and 4-month-olds but not at 6 months (Coldren, unpublished data, described in Colombo & Mitchell, 1990). Pêcheux and Lécuyer (1983) found with 4-month-olds that fixation time towards static stimuli was positively correlated with their visual exploration of a toy (Fig. 1). This gap in the literature is important for a number of reasons. As we note in Part 2, there are a number of marked differences between these different looking time paradigms, such as: the size of the target towards which attention is being directed, the presence or absence of movement in the target or periphery of the visual field of the child, and the relative luminance of the target relative to other elements within the infants’ field of view (Fig. 2). In the absence of data showing cross- paradigm consistency, we cannot be sure how individual differences in attention as assessed using screen-based tasks might relate to individual differences in attention in naturalistic settings. Are the dissimilarities between screen-based and naturalistic attention tasks documented in Fig. 2 incidental to the individual differences that are assessed on these tasks? Or are they central to them? Within the habituation literature, shorter looking to static stimuli during the first year is frequently described as an index of ‘faster processing speed’; this is frequently posited as an explanation for the negative correlations noted between look duration during the first year and long-term outcomes (Colombo & Cheatham, 2006). One question that follows from this is: does ‘faster processing’ as assessed using screen-based attention tasks also manifest as different (‘better’, or ‘more efficient’) orienting in naturalistic contexts? Or is shorter looking to screen-based stimuli associated with better long-term outcomes because both measures tap some underlying, ‘pure’ aspect of cognition that is entirely independent of naturalistic orienting? The present study is intended as a small step towards addresing these questions. The present study ~~~~~~~~~~~~~~~~~ As described above, there exists to our knowledge no previous research that has addressed whether individual differences in peak look duration are consistent across different assessment techniques. The present study was conducted in order to address this question. We presented four screen-based assessments, namely: (i) looking behaviour towards ‘interesting’ (complex) static stimuli, (ii) ‘boring’ (non-complex) static stimuli, (iii) mixed static and dynamic stimuli and (iv) to videos under conditions of distraction (during the recording of EEG data). We also presented two semi-naturalistic looking assessments involving the presentation of novel objects in a table-top setting, in (i) a single-object condition (novel objects presented one by one) and (ii) a four-object condition (four novel objects presented concurrently) (following (Kannass & Oakes, 2008). The six measures were presented in different testing rooms and inter-leaved in order, to a single cohort of typically developing 11-month-old infants. 11-months was chosen as the age for the present study because this has been characterised as an age that shows the first emergence of endogenous attentional control (Colombo & Cheatham, 2006; Courage et al., 2006). Across all six paradigms, the single dependent variable we assessed was peak look duration. This was selected because it has previously been argued to be the most stable assessment of looking behaviour during infancy–in comparison for example to habituation rate (the rate of change of looks over time), which is less reliable, and shows less robust relationships with long-term cognitive outcomes (Colombo & Mitchell, 1990). As far as possible, peak look duration was assessed identically across the six paradigms we administered. From reviewing the literature we were able to find no discussions suggesting that different factors might influence peak look duration differentially between screen-based and semi-naturalistic settings. Therefore we predicted that individual differences in peak look duration would be consistent across all the paradigms administered.","42 typically developing 11-month-old infants participated in the study. Mean age at testing was 337 days (range 312–259, standard deviation 9). Gender ratios were 26 male/16 female. Of note, other aspects of these data have already been published elsewhere (Wass et al., 2011; Wass & Smith, 2014). The current data contain, however, completely novel analyses which do not overlap with previous publications.","The six peak look assessments were administered in three sections. Section A consisted of the ‘static non-complex’ assessment, the ‘static complex’ assessment and the ‘mixed dynamic/static’ assessment. Section B consisted of the ‘structured free play’ assessment. Section C consisted of the ‘videos during EEG’ assessment. Presentation order. All three sections were administered during a single visit, which generally lasted c. 90 min with breaks. Section A was presented in two halves (‘A1’ and ‘A2’). The order in which the sections were administered was: Section A1, then Section B, then Section A2, then Section C. The naturalistic ‘structured free play’ assessment (Section B) was therefore presented between the other screen-based tasks. This design was chosen in order to preclude the possibility of order effects being responsible for the results observed. Testing rooms. Sections A–C were each presented in different rooms. Of note, therefore, the screen-based tasks included in Section C were presented in a different room to the screen-based tasks in Section A. In the detailed descriptions of the methods that follows, materials are described section by section, together with the data processing techniques that were used. Section A – ‘static non-complex’/‘static complex’/‘mixed dynamic/static’ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Materials. For the three peak look assessments contained in section 1 infants were seated on their caregiver's lap while the viewing material was presented on a Tobii 1750 eyetracker subtending 24° of visual angle. The three assessments were presented interleaved with each other. Static non-complex images. Two different still images were presented at different stages of the testing protocol. The two ‘non-complex’ images were both monochromatic objects presented against a white background (see Fig. 1 for example). Trials were presented concurrent with child-friendly music, such as songs from Sesame Street. Four different songs were used that were paired randomly with the different images. All infants heard the same four songs over the course of all experiments. Trials were presented using a gaze-contingent infant-controlled habituation protocol procedure: images were presented and remained on-screen for as long as the infant looked to the screen. Following cessation of a look, the image was re-presented until two successive looks had taken place that were less than 50% of the longest look so far. In order to confirm eyetracker contact, a small (c. 0.4°) re-fixation target was briefly presented every 15 s; subsequent analyses (described in the Supplementary Materials) suggested that this did not influence the timing of peak look duration measure. Peak look was calculated independently for each image and then averaged. Static complex images. The two ‘complex’ images were polychromatic scenes (see Fig. 1 for example). The testing procedures used were identical to those used for the static non-complex assessment. For practical reasons, individual trials were capped at 120 s; 11 out of the 152 individual trials included reached this cap (see Fig. S1). To confirm our classification of images into ‘complex’ and ‘non-complex’ feature congestion was calculated for each image using Matlab scripts from Rosenholtz, Li, and Nakano (2007). Feature congestion quantifies local variability across different first-order features such as colour, orientation and luminance; see SM for a more detailed description. For the two ‘non-complex’ images, average feature congestion across the whole frame was found to be 1.7 and 1.6; for the two ‘complex’ images, average feature congestion was 7.6 and 5.1 (see Fig. S2). This confirmed our classification of the stimuli into ‘complex’ and ‘non-complex’. Mixed static/dynamic images. 3 blocks of mixed static and dynamic images were presented at different stages of the testing protocol. Each block lasted 65 s. Each block consisted of a mixture of: head shots of actors (single and in groups) reciting nursery rhymes, still images of actors’ faces, and shots of toys and birds accompanied by background music (see Fig. 1 for example). The individual stimuli within each 65-s block each lasted 4–12 s. As with the static images, a small re-fixation target was briefly presented c. every 15 s in order to confirm eyetracker contact (see analyses in SM). Data processing. Infants’ looking behaviour was coded from a camera on top of the monitor. Gaze was coded in 1-s bins, as either looking at the screen or not. Total percentage looking time and the length of each unbroken look to the target were calculated. Instances in which the participant looked away from and then back to the screen within 1 s were treated as constituting one continuous look rather than two discrete looks. Coder 1 coded 76%, coder 2 48%; 25% of the videos were double coded. Cohen's kappa was calculated to assess inter-rater reliability and was found to be 0.71. Section B – ‘structured free play’ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Materials. The two peak look assessments contained in Section 2 were conducted in a puppet theatre with attractive surrounds, and a stage behind which experimenter and the camera were visible. Infants sat on their caregiver's lap, close enough to the stage so that they could reach to and touch the objects on it. Between each trial, the curtains of the puppet theatre were closed and new objects were placed on the stage; reopening them marked the start of the next trial. The two assessments were presented consecutively: Free play – 1-object condition. In the single-object condition, five objects (an plastic figure/a basting pipette/a glitter lamp/a lion mask/a rabbit mask) were presented in randomised order consecutively for 30 s each. Fig. 1 shows an example of the objects used; Fig. S3 shows images of all the objects used. The objects used varied in size from 5–20 cm. Free play – 4-object condition. The four-object condition was presented immediately after the one-object condition. Four objects (a rubber duck, a plastic train, a plastic teddy bear, a tiger finger puppet) were presented concurrently in a line across the stage, in a randomised order, for 90 s. The objects used varied in size from 5 to 10 cm. Data from 10 participants was unusable for the four-object condition due to changes made to the experimental protocol during testing. Data processing. Infants’ looking behaviour was recorded from a camera positioned behind the stage. The coding protocol used was based on that used by Kannass and Oakes (2008). Infants’ looking behaviour was coded for whether the infant was looking at the object or not. Sections where the object was not on the stage (because the infant had knocked or thrown it off) were excluded. All coding was conducted in 1-s bins. Data were triple coded. Coder 1 coded 70%, coder 2 50% and coder 3 24%; 40% were double coded by coders 1 and 2 and 24% by coders 1 and 3. Cohen's kappa was calculated to assess inter-rater agreement. This was found to be 0.72 between coders 1 and 2 and 0.78 between coders 1 and 3. Section C – ‘videos during EEG’ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Materials. The peak look assessment contained in Section 3 was presented with infants sitting on their caregiver's lap while viewing materials were presented on a cathode ray TV subtending 30° of visual angle. Simultaneously with the administration of this task, infants were having EEG data recorded using a 128-channel EGI hydrocel net (Wass, 2011). Only one assessment was presented in this section: Videos during EEG. Three videos were presented sequentially in rotation during EEG recording. These videos were: (i) a series of actresses reciting nursery rhymes to camera; (ii) videos of toys spinning; (iii) a short TV clip. Videos lasted 32–44 s each. Each video was presented twice. Data processing. Infants’ looking behaviour was recorded from a camera positioned below the monitor. Videos were coded according to whether the infant was looking to or away from the screen, using an identical coding scheme to that used in sections and 2. Coder 1 coded 76%, coder 2 48%; 24% were double coded. Cohen's kappa was calculated to assess inter- rater reliability and was found to be 0.88.","The results section is in two parts. Firstly, descriptive statistics of the looking time data obtained from the different paradigms are presented. Secondly, analyses are presented that examine the inter-relationships in looking time across the different assessments administered. Specifically we wished to evaluate the hypothesis that individual differences in looking time would be consistent across the six assessments. Part 1 – Descriptive statistics of looking time data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Descriptive statistics for the entire data set are shown in Table 1. For each assessment, mean peak look oberved across all infants has been reported, together with the Standard Error of the Mean (S.E.M.) and range. Additionally, for comparison, identical data have been reported for mean look duration (i.e. the average of all looks recorded towards the stimuli). For each assessment, the number of participants who provided usable data is shown in the final column. With the exception of the free-play 4 object task, for which (as described above) changes were made to the experimental protocol during testing, drop- out rates are acceptable (maximum 4/42). These were due to fussiness and non-compliance during testing. Fig. 3a–f shows histograms of all the individual looks collected on the different tasks. Fig. 3g shows plotted lognormal fittings. Lognormal distributions were calculated as these are generally reported to be the best fit on infant looking time data (Pempek et al., 2010; Richards & Anderson, 2004). Marked differences in the patterns of look durations observed on different tasks can be seen: both peak and mean look duration were higher for all of the screen-based tasks than for the structured free play tasks. Within the screen-based tasks, markedly longer peak looking times were observed in the static complex and mixed dynamic-static categories than in the other categories. The between-participant distributions of peak look durations were found to be positively skewed, in common with all looking time assessments (see e.g. Richards & Anderson, 2004); therefore all subsequent analyses have been calculated based on log-transformed data (following e.g. Frick, Colombo, & Saxon, 1999). Part 2 – Analyses to examine the inter-relationships in looking time across the different assessments administered ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We wished to evaluate the hypothesis that individual differences in looking time would be consistent across the six assessments we administered. In order to examine this, two analyses were conducted. First, zero-order correlations were calculated. Second, a factor analysis was performed. First, histograms and scatterplots were calculated to assess whether per-participant peak look values derived from the log-transformed data were parametrically distributed, and whether any bivariate relationships observed were robust. Fig. 4 shows these scatterplots. All parameters were found to be normally distributed. Zero-order correlations. Fig. 4 shows the zero-order bivariate correlations that were observed between the variables entered into the factor analysis. The four screen-based tasks (static complex, static non-complex, mixed dynamic-static, videos during EEG) all show significant correlations (r = .33 to .56, all ps < 05) with the exception of the static complex to dynamic during EEG (r = .24, p < .10). Inspection of the scatterplot (Fig. 4) suggests that this relationship is weakened by an outlier. In comparison the two FP tasks do not correlate with any of the screen-based tasks (negative in 5 of the 9 comparisons conducted, and r < = .12 in the remaining 4). The scatterplots in Fig. 4 suggest that this not attributable to the presence of outliers. One explanation that was considered for the low zero-order correlations observed with the free play data was that these data were inherently more ‘noisy’ than the screen-based looking time data. In order to evaluate this possibility, data were examined from an overlapping dataset that has been published previously (Wass, 2011; Wass et al., 2011). In this paper, an identical task to that presented here was presented twice at fifteen days’ interval to a smaller cohort (N = 21) of infants. Analyses assessed the number of total attentional reorientations and attentional shifts from object to person. Test-retest reliability between the two testing sessions was r = .53, p < .01 for total attentional reorientations, and r = .52, p < .05 for attentional shifts from object to person. This suggests that these measures are relatively stable as indices of individual differences. Factor analysis. Factor analyses were conducted to examine the factorial structure underlying our data in more detail. Our analytical approach was based on that used in previously published research (Rose, Feldman, & Jankowski, 2004, 2005). First, the sample size was examined. The ratio of participants to variables for the factor analysis was found to be 6.3, which is above the prescribed ratio of 5 suggested by Hair, Tatham, Anderson, and Black (1998). To maximise the sample size for factor analysis, missing values were imputed based on the mean of the subscale to which that item belonged (following Blair & Razza, 2007). The factor analysis yielded a two-factor solution (Eigenvalues >1.0) representing 59% of the total variance (see Table 2). These two factors were submitted to a principal axis rotation (oblimin) and the scree plot was inspected, supporting a two-factor solution. Thresholds were set at 0.70 for principal loading and 0.50 for secondary loading (Guadagnoli & Velicer, 1988). The first factor, which had an Eigenvalue of 2.29 and accounted for 38% of the variance, was defined by three of the screen-based tasks, with the fourth screen-based task allotted a secondary loading. The second factor, with an Eigenvalue of 1.23, was loaded onto by the two FP variables (4-object as primary loading and 1-object as secondary loading), and (negatively) by the static complex variable (primary loading).","Our analyses were designed to assess whether individual differences in peak look duration are consistent across different looking time measurement paradigms. To 42 typically developing 11-month-old infants we administered six assessments of peak look duration, including four screen-based assessments and two free-play based assessments. We predicted that results obtained would be consistent across all paradigms. The results were not as predicted. The factor analysis suggested a two-factor solution. The first factor was defined by the four screen-based tasks (‘static non-complex’, ‘mixed dynamic-static’ and ‘videos during EEG’, with ‘static complex images’ as a secondary loading). The second factor was defined by the two free play tasks (one as a secondary loading) and also (with a negative loading) by the static complex screen task. The four screen-based tasks were administered across different testing rooms, and interspersed with the free play tasks, which precludes the possibility of room or order effects being responsible for our results. Looking time to static screen stimuli and to dynamic screen stimuli showed strong correlations. Strikingly, we also found that looking behaviour towards a TV screen during recording of EEG data, which has the additional variance of tightness of fit of the EEG net, reactivity to testing and so on, mapped onto the same factor as the other three screen-based tasks, that were administered using a different screen in a different room. In contrast the two FP assessments mapped onto a separate factor, and showed non- significant (max r = .12) zero-order correlations with each of the screen-based tasks. The zero-order correlations observed between the FP and screen-based tasks were negative in 5 out of 8 comparisons. In the factor analysis, the only screen-based task (static complex) that loaded on to the second factor loaded on negatively (higher looking time to static complex images associated with lower looking time during structured free play). Further analyses were conducted to assess the possibility that these findings might be attributable to other factors such as increased measurement error during the administration of the free play tasks, with negative results. There are a number of limitations to this study. The sample size was relatively small (N = 42), and a number of techniques used by other researchers to complement looking time measures (such as heart rate measurement and focused attention coding) were not applied. Furthermore, measurements were only taken with one age group (11-month-olds), whereas the limited data available suggests that different results may have been observed if the experiment were repeated with younger infants (Coldren, unpublished data; discussed in Colombo & Mitchell, 1990). Nevertheless, our results suggest that, in 11-month-old infants, individual differences in peak look duration are constant across different screen-based tasks but not between screen-based and semi-naturalistic tasks. We were able to find no discussion in the literature suggesting that different factors might influence peak look duration between screen-based and semi-naturalistic settings. What kinds of differences might these be? The following discussion is structured around a number of factors commonly thought to influence peak look duration. The first factor commonly associated with peak look duration is processing speed. Sokolov argued that the initial presentation of a novel stimulus produces a conflict between a “neural model” of the current environment and the sensory processes occurring in the brain; prolonged exposure to that stimulus allows the viewer to form an internal representation of it, which is why looking durations decline over time (Sokolov, 1963). ‘Faster processors’ are thought to require less time to form an internal representation; this is frequently linked to the finding that shorter peak look duration to static stimuli during the first year correlates negatively with long-term cognitive outcomes (e.g. Rose et al., 2002; Rose, Feldman, Jankowski, & Van Rossem, 2008). Advocates of the importance of processing speed in influencing peak look duration might predict that individual differences would not be stable between looking time to static screen and dynamic screen stimuli, since one involves static visual information and the other constantly changing information. They might also suggest that individual differences might be stable between static screen and our free play task, since both require forming internal representations of static targets (on-screen pictures and ‘real-world’ objects). In fact we found the opposite pattern: individual differences in peak look duration were consistent across the static screen and dynamic screen stimuli but not with the free play task. A second factor related to peak look duration is ease of disengaging of visual attention. Frick and colleagues measured the relationship between experimentally assessed attentional disengagement latencies and spontaneous looking behaviour to static screen stimuli in typically developing 3- and 4-m-os (Frick et al., 1999). They found that long- looking infants showed greater variability in their response latencies. This suggests that, at least in younger infants than the 11-month-olds studied here, attentional disengagement may play a role in mediating spontaneous looking behaviour. This is one area where differences can be noted between our screen-based and semi-naturalistic paradigms (see Fig. 2). Screen-based paradigms tend to be designed with the screen occupying a relatively large proportion of the infant's visual field (typically c.25° of visual angle, as here), whereas in FP paradigms the target is generally much smaller (c. 5° in our case). In screen-based tasks the target is generally much more luminant than the surrounds (which are typically dark); in our free play paradigms, in contrast, this was not the case (see Fig. 2). Lastly, in screen-based tasks there are sharp luminance contrasts between the edge of the screen and the surrounds; again, these were not present in the FP task (see Fig. 2). These differences may be important because previous research has noted that viewers tend to dwell on areas of high luminance contrasts such as object boundaries. Although this effect has been reported at all ages from 6-week-old infants (Bronson, 1994) through to adults (Henderson & Smith, 2009) its effect has been reported to decline with increasing age (Frank, Vul, & Johnson, 2009; Karatekin, 2007). The high-contrast and prominent luminance contrasts present in our screen-based but not in our naturalistic tasks may influence behaviour in the current study, and perhaps more for some infants than others. A third factor that may relate to peak look duration is autonomic arousal. This can be assessed in both phasic (i.e. event-related) and tonic contexts (de Barbaro et al., 2011; Richards, 2011). Richards and colleagues explored changes in heart rate variability and peak look during object examination; they found a decrease in variability during attention, which was interpreted as consistent with a model of phasic parasympathetic vagal influence on the heart during sustained attention phases (Richards & Casey, 1991; Richards, 2011). Of note, our screen-based tasks (particularly the static screen stimuli) contained abrupt changes in luminance coincident with the onset of each trial: the screen transitioned from dark to bright in an otherwise darkened room and an auditory stimulus was presented; such changes were completely absent in the naturalistic task. It is possible that these abrupt changes in luminance are associated with phasic changes in sympathetic/parasympathetic nervous system balances, and that some infants are more susceptible to these changes than others (Alkon et al., 2006). This factor would influence looking behaviour in the screen-based but not the naturalistic looking time tasks. Aston- Jones and colleagues suggested a role for brainstem, Norepinephrine modulated arousal systems in shifting between attention states; they distinguish between a ‘scanning’ mode, in which look durations are short and the focus of visual attentiveness is wide, and a ‘focused’ mode, in which look durations are longer and the spatial distribution of attention is narrower (Aston-Jones, Rajkowski, & Cohen, 1999; Aston-Jones, Iba, Clayton, Rajkowski, & Cohen, 2007; cf. Pannasch et al., 2008). In our semi-naturalistic tasks, other targets (objects and people) are present within the peripheral visual field of the infant, whereas attempts were made to ‘black out’ all peripheral objects for the screen- based tasks (as is typical in other labs) (see Fig. 2). The shifting between attention states that Aston-Jones and colleagues describe may therefore be a factor in our semi- naturalistic tasks but not in our screen-based tasks. A fourth factor that may relate to peak look duration is executive control. Aspects of executive control have been reliably associated with sustained attention in older children (e.g. Reck & Hund, 2011). (Note however that sustained attention in these studies with children is assessed not using looking time measures but with tasks such as the Continuous Performance Task, whose relationship with peak look duration has, to our knowledge, not been studied.) Colombo and Cheatham point out that positive correlations are observed between long-term cognitive outcomes and peak look to static stimuli after the first year of life, whereas negative correlations are observed between the same two variables during the first year. They suggest that this may be attributable to the emergence of effortful control as a factor mediating behaviour at about the 12-month boundary (Colombo & Cheatham, 2006; see also Courage et al., 2006). Note, however, that Kochanska and Aksan (2006) found that focused attention (not peak look duration) during the second half of the first year correlated positively with effortful control at 22 months. One difference between our screen-based and our semi-naturalistic tasks may be relevant here. This is that, for the semi- naturalistic tasks, a number of other informative gaze targets (such as the experimenter) are within the field of view of the child – whereas the screen-based tasks were conducted in a darkened room. There are a variety of reasons why looks away from the object may have an adaptive value in our free play paradigm but not in our screen-based paradigm (e.g. Rueda, Posner, & Rothbart, 2005; Sheese, Rothbart, Posner, White, & Fraundorf, 2008). It may be therefore that executive control relates more strongly to peak look duration in the free play than in the screen-based tasks, although future work is required to investigate this in more detail (cf. Rothbart, Ellis, Rueda, & Posner, 2003; Sheese et al., 2008).","To a cohort of typically developing 11-month-old infants we presented several assessments of peak look duration, including some that assessed looking behaviour on screen-based tasks and others that assessed behaviour on semi-naturalistic tasks. We found that the four screen-based tasks (looking to static non-complex, to static complex, to mixed dynamic/static and to dynamic stimuli during EEG recording) all mapped onto a single factor, whereas the two free play assessments mapped onto a separate factor. In our discussion we noted a number of ways in which factors such as susceptibility to high luminance contrasts and abrupt stimulus onset–offset changes may be key factors mediating individual differences on screen-based tasks, but relatively unimportant in more naturalistic contexts. Future research should exploit recent technological advances such as head-mounted eyetrackers (e.g. Aslin, 2009) to increase our understanding of how individual differences in naturalistic attention relate to individual differences in infant attention as assessed using screen-based paradigms."],["Background: Although cord blood (CB) stem cell research is being conducted for treatment of cerebral palsy (CP), little is known about children with CP and stored CB. Aims: To compare demographic and clinical characteristics of children with CP and stored CB to children with CP identified in a population-based study. Methods and Procedures: The Longitudinal Umbilical Stem cell monitoring and Treatment REsearch (LUSTRE®) Registry recruited children from the largest US private CB bank. Demographics, co-morbidities, and gross motor function (GMFCS level and walking ability) were collected and, where possible, compared with the CDC's Autism and Developmental Disabilities Monitoring (ADDM) Network. Outcomes and Results: 114 LUSTRE participants were compared to 451 ADDM participants. LUSTRE participants were more likely to be white, but sex distribution was similar. Co-morbidities (autism and epilepsy) and functional mobility were also similar. Conclusions and Implications: The results of this analysis suggest that while children diagnosed with CP and with access to stored CB differ from a broader population sample in terms of demographics, they have similar clinical severity and comorbidity profiles. As such, LUSTRE may serve as a valuable source of data for the characterization of individuals with CP, including individuals who have or will receive CB infusions. --------------------------------------------------------------------------------","To date, little research has been done on families that choose to privately bank their children’s cord blood (CB) and whose children are subsequently diagnosed with cerebral palsy (CP). This analysis was performed to help understand the extent to which the results of research within the private CB-storing population might be relevant to broader populations of children with CP. The study compared children with CP who were enrolled in the Longitudinal Umbilical Stem cell monitoring and Treatment REsearch (LUSTRE®) registry at the largest private, US cord blood bank to children with CP identified in the CDC’s Autism and Developmental Disabilities Monitoring (ADDM) Network. While the demographics of children with CP differed, these samples were similar when compared across several measures of functional mobility and for the prevalence of co-occurring medical conditions. Based on these results, LUSTRE may be useful for the characterization of children with CP, including individuals who have or will receive CB infusions.","Umbilical cord blood (CB) contains a mixed cell population and is a rich source of hematopoietic stem cells (Gluckman et al., 1989). CB derived cellular therapy has been used for decades to treat pediatric and adult patients with hematologic malignancies and disorders (Butler & Menitove, 2011; Gluckman et al., 1989; Prasad & Kurtzberg, 2009). Based on evidence that CB cells may have the capacity to facilitate repair of damaged tissues outside of the blood and immune system, it is also in early phases of research as a treatment for a number of additional conditions, including cerebral palsy (CP) (Chez et al., 2018; Dawson et al., 2017; Kurtzberg, Durham, Laskowitz, Balber, & Bennett, 2016; Sun et al., 2015; Sun et al., 2016; Sun et al., 2017). The collection and private storage of umbilical cord blood, recently estimated at approximately 4 million units worldwide, has become an increasingly popular choice for parents of newborn children (Ballen, Verter, & Kurtzberg, 2015). However, little is known about the medical conditions that affect families that choose to privately store their newborns’ CB. The Longitudinal Umbilical cord blood Stem cell monitoring and Treatment REsearch (LUSTRE®) Registry identifies and follows families that have both stored their children’s CB in the largest US CB bank and have children with conditions that are currently treated with, or under research for treatment with, CB. Over time, LUSTRE is designed to help better understand the clinical characteristics and disease severity of these children, describe the treatments they receive, compare the long-term clinical and quality of life outcomes associated with these treatments, and characterize the types of participants who are receiving stem cell therapy. CP, one of several neurological conditions currently being studied for potential treatment with CB, is a LUSTRE target condition. Little research exists regarding the characterization of children diagnosed with CP who also have access to privately stored CB; nor is there published research on how this population compares to the broader CP population. This study was undertaken to describe the LUSTRE-CP cohort and to assess, in terms of demographics and disease severity, how children in LUSTRE with CP compared to a nationally representative sample of children with CP. We used data from the Autism and Developmental Disabilities Monitoring (ADDM) Network for this comparison group. ADDM is funded by the U.S. Center for Disease Control and Prevention (CDC) to estimate the number of children with autism spectrum disorder and other developmental disabilities living across the United States. The ADDM Network has routinely provided valuable insight into the epidemiology of various developmental disabilities, including CP (Christensen et al., 2014; Kirby et al., 2011; Yeargin-Allsopp et al., 2008). In order to determine how children with CP and stored CB compare to children with CP in the general public, we compared children in LUSTRE to children captured in the most recent assessment of CP conducted by ADDM in 2008. Study design ~~~~~~~~~~~~ LUSTRE is an observational disease registry open to all families who have stored CB in the largest, U.S.-based, private cord blood bank, CBR Systems, Inc. LUSTRE assembles data in three distinct phases, each collecting more detailed information from a more targeted group of families (Fig. 1). The Surveillance Phase (Phase I) collects data on the prevalence of medical conditions among all families storing CB in the bank. It uses a cross-sectional survey to identify whether families have any of 32 conditions currently treated with CB or under investigation for treatment with CB. The Monitoring Phase (Phase II) collects detailed data annually on children with selected target conditions identified in Phase I, including disease severity, medical and treatment history, comorbidities, demographics, and quality of life. This phase of the registry consists of an observational, prospective, longitudinal cohort study that collects data on children with target medical conditions and with access to their own (autologous) stored CB. Finally, the Infusion Phase (Phase III) collects additional clinical details on Phase II patients who receive a CB infusion. LUSTRE participants are followed over time, regardless of whether their families ever elect to use their stored CB.","The initial participants in Phase I of LUSTRE have been described previously (Mazonson et al., 2017). Briefly, in the original cohort, over 94,000 families provided information about whether their child with stored CB, or that child’s first-degree relatives, had been diagnosed with any of 32 conditions of interest. By offering the survey to more families as they store CB in the bank, new observations are added to the Phase I dataset, which now contains information on more than 120,000 families. Once families identified an eligible family member with one or more LUSTRE conditions, they were invited via email to join the Monitoring Phase of the registry. The Monitoring and Infusion Phases began enrollment in January of 2014 for children whose own CB had been stored and who had a CP diagnosis. LUSTRE’s study protocol was approved by Ethical and Independent Review Services (E&I #13-130). Informed consent was obtained before administering the Phase II study questionnaire. Selection of the Autism and Developmental Disabilities Monitoring Network comparison cohort ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The ADDM Network offers the largest cross-sectional assessment of children living with CP in the US (Christensen et al., 2014). These data are frequently referenced for CP prevalence rates in the US (Christensen et al., 2014). In 2008, 451 children with CP were identified via data abstracted from organizations that diagnosed, treated, and provided services to developmentally disabled children served by one of four ADDM surveillance sites in Alabama, Georgia, Missouri, and Wisconsin. Children were eligible if they were 8 years old at the time of case ascertainment. The ADDM Network surveyed 8-year-old children because it focused heavily on ASD and other developmental disabilities and previous work showed that, by this age, most children with ASD had been identified for health care services. Although all ADDM Network study subjects were 8 years of age, GMFCS level and walking ability category were abstracted based on available descriptions of study subjects at or after 4 years of age, when their motor functioning levels were believed to have stabilized. To maximize available data points but minimize the chance of misclassifying motor functioning, we, similarly, compared the GMFCS-derived study measures from LUSTRE children of at least 4 years of age with data published by the ADDM Network. Wider and narrower LUSTRE age groupings were also compared to the ADDM cohort to see if our findings were affected. Assessment of functional mobility and other clinical characteristics ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To optimize the collection of parent-reported data when designing the Phase II questionnaire, LUSTRE used existing, validated instruments and methodologies where possible. To assess functional mobility, primary caregivers of study subjects were asked to select which of five age-appropriate descriptions of functional mobility best described their child, using the Parent Report Gross Motor Functioning Classification System (GMFCS) (Palisano, Rosenbaum, Bartlett, & Livingston, 2008). Similarly, the ADDM Network abstracted GMFCS level based on available descriptions of study subjects. GFMCS level determinations were then used to classify study subjects into three aggregate categories describing walking ability: walks independently, walks with handheld mobility device, or limited or no walking ability. Specifically, in ADDM, GMFCS levels I and II were classified as “walks independently;” level III as “walks with handheld mobility device;” and levels IV and V as “limited or no walking ability.” To maintain concordance with the ADDM Network, LUSTRE reported data on both GMFCS levels as well as the aggregate walking ability categories used in ADDM. In addition to assessing their child’s functional mobility, primary caregivers with children enrolled in LUSTRE provided information regarding demographics and the presence of co-occurring medical conditions. For this study, we focused on co-occurring autism spectrum disorders (ASD) and epilepsy/seizures to make clinical comparisons since these were the only comorbid conditions publicly available from ADDM Network data. Statistical methods ~~~~~~~~~~~~~~~~~~~ Frequency distribution of characteristics in the LUSTRE data set, such as age at enrollment, parental education, and family income level were calculated for all participants with a CP diagnosis for whom CB was stored at birth. Additionally, Pearson χ2 tests were used to examine differences in characteristics of children enrolled in the LUSTRE and ADDM Network cohorts, including sex, ethnic group, co-occurring ASD, co- occurring epilepsy, walking ability, GMFCS level, and ambulatory status. Participants < 4 years of age (28.9%) were excluded from the LUSTRE cohort when comparing clinical characteristics and comorbidities, in order to focus on children with more stable condition diagnoses and gross motor function. A p-value of <0.05 was considered significant. All data analyses were performed using STATA statistical software program (version 13.0, STATA Corp LP, College Station, TX). Identification of children with CP in bank ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As of November 2016, 121,411 families had completed a LUSTRE surveillance questionnaire. 429 families (0.35%) indicated a possible CP diagnosis in a child with access to their own stored CB. 221 families chose to enroll in Phase II of LUSTRE, 158 (71.5%) were identified via the surveillance questionnaire and 63 (28.5%) were previously identified by the CB bank. Of the 221 enrolled families, 114 of these families reported a formal CP diagnosis. Data on the 114 subjects with CP are presented here. Demographic and SES overview of LUSTRE participants with CP ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As displayed in Table 1, slightly over half of LUSTRE subjects with CP were male (56.1%). The average age of LUSTRE children was 5.7 (SD 3.5), with a range at time of enrollment of 1–17 years. In addition, the majority of LUSTRE children were white (79.8%). The majority of fathers (64.0%) and mothers (60.5%) of children with CP had at least a four-year college degree (Table 1). 55.4% of LUSTRE households providing income information reported household incomes greater than or equal to $100,000. Of note, 35% of the LUSTRE cohort did not provide financial information. Demographic comparison of LUSTRE and ADDM network subjects ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 2 presents data on the 114 LUSTRE subjects compared to the 451 ADDM Network subjects. Both cohorts were majority male. (56.1% vs 60.8%, p = 0.369). However, the ethnic profile of the LUSTRE population showed a significantly higher percentage of white subjects compared to the ADDM Network cohort (78.8% vs. 51.6%, p < 0.001). Parental education and income levels were not available from the ADDM Network cohort for comparison with LUSTRE. Disease severity comparison of LUSTRE and ADDM network subjects ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The frequency of co-occurring ASD among LUSTRE participants was 10.3% (95% CI 3.4–17.1%) and among ADDM participants was 6.9% (95% CI 4.9–9.6%). The frequency of co-occurring epilepsy/seizures was 34.2% (95% CI 23.5–44.9%) in LUSTRE versus 41.0% (95% 36.4–45.7% CI) in ADDM (Table 2). The prevalence of ASD and epilepsy was not significantly different between ADDM and LUSTRE (p = 0.291 and 0.252, respectively). In addition, LUSTRE and ADDM Network subjects did not differ significantly on the GMFCS, measured across all 5 levels (p = 0.354), collapsed into walking ability categorization (p = 0.610) or summarized as ambulatory (GMFCS levels I-III) and non-ambulatory (GMFCS levels IV-V) (p = 0.473). No clinical differences were found when comparing LUSTRE participants across subgroups of ages (ages 4–8 or 6–10) to the ADDM participants (Table 3). Demographic comparison of respondent and non-respondent families ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The participation rate among families responding to Phase I of LUSTRE was 33.3%. After adjusting using inverse proportional weighting of characteristics available for both responders and non-responders, no material differences in condition prevalence rates were observed (Mazonson et al., 2017; Seaman & White, 2013). Among families invited to participate in Phase II, 41.9% enrolled. No significant differences in the available characteristics of participating and non-participating families were observed (e.g., mean age of primary parent, mean age of children with stored cord blood, mean CB units stored) (Table 4). Missing data ~~~~~~~~~~~~ Missing data were rare among LUSTRE participants. Two participants did not provide GMFCS level data. In contrast, due to ADDM Network’s reliance on medical records, clinician reviewers were unable to assign GMFCS levels or walking ability categories for 142 out of 451 children. However, efforts were made to maximize the number of children with a gross motor function assessment. For 28 of these 142 children, enough information was available for study clinicians to assign a walking ability category without a formal GMFCS level, reducing the number of observations missing a gross motor function assessment to 114 (25.3%).","LUSTRE is a unique registry of children with privately stored CB who have conditions treatable with, or undergoing research for treatment with, CB. When compared, subjects with CP in LUSTRE had similar clinical profiles to those of the ADDM Network’s CP sample. Specifically, the two CP cohorts displayed similar motor functioning levels and prevalence of comorbidities. Both LUSTRE and the ADDM Network reported that just over half of children walked unaided while just over a quarter were unable to walk. When comparing GMFCS levels, over two-thirds were ambulatory (GMFCS I-III) in both LUSTRE and the ADDM Network. In addition, the prevalence of co-occurring ASD and epilepsy/seizures were similar between the two samples. Both cohorts reported approximately 10% of children with CP had been diagnosed with co-occurring ASD, while over a third of children with CP had been diagnosed with epilepsy/seizures (Christensen et al., 2014). The LUSTRE cohort was less ethnically diverse than ADDM Network’s population-based sample. No significant difference in sex distribution was observed. While we were not able to compare parental education levels or household income directly, if the ADDM Network sample represented the average US population, LUSTRE families were likely to be significantly more educated and affluent than the ADDM Network sample. Specifically, 55.4% of LUSTRE CP households providing income information reported greater than or equal to $100,000 in family income, whereas, 26.4% of US households had equivalent income according to 2015 census data (United States Census Bureau, 2015a). Furthermore, among all US adults, only 32.3% of males and 32.7% of females have at least a college bachelor’s degree, whereas, among the parents of the LUSTRE cohort, 64.0% of fathers and 60.5% mothers reported this level of education (United States Census Bureau, 2015b). Limitations ~~~~~~~~~~~ There are potential limitations to this analysis related to both the ADDM Network and the LUSTRE data. With regard to the ADDM Network, limitations have been reported elsewhere (Christensen et al., 2014). Most notably, 25% of children with CP did not have sufficient information on gross motor functioning to classify them according to the GMFCS or walking ability categorization. To address this concern, ADDM Network investigators used a multiple imputation model to infer children’s walking ability based on other data available in their medical records. The distribution of walking ability among children with imputed data for missing values was similar to what was found for children with observed data. Thus, it is likely that data was missing at random, unrelated to severity level (Christensen et al., 2014; Maenner et al., 2012). With regard to LUSTRE, as previously published, the participation rate among families responding to Phase I of LUSTRE was 33.3% and the participation rate among families invited to take part in Phase II was 41.9%, raising the possibility of selection bias among the respondents. However, further analyses did not show evidence of selection bias in either group of participants. Nonetheless, some degree of selection bias could have occurred. Sampled populations also differed between LUSTRE and ADDM. While ADDM was limited to four sites, primarily in the south (Alabama, Georgia, Missouri, and Wisconsin), LUSTRE is open to families across the US. In addition, families participating in LUSTRE must pay to store cord blood at a private cord blood bank. This potential financial barrier to private cord blood banking likely results in a less socio-economically diverse LUSTRE cohort. A previous study of a Canadian-based CP cohort showed that lower socioeconomic status may be associated with lower GMFCS level (Oskoui, Messerlian, Blair, Gamache, & Shevell, 2016). This study also suggested that socioeconomic status is positively correlated with full-term pregnancies, which might suggest different etiologies of CP in full term pregnancies. Interestingly, despite differences in socioeconomic factors, we didn’t detect differences in functional mobility between the LUSTRE and ADDM cohorts. An additional, potential limitation of LUSTRE is its use of parent-reported data, which could have introduced some misclassification of conditions and comorbidities. However, many widely cited, peer- reviewed epidemiological studies utilize parental report in place of clinical report to provide descriptions of neurodevelopmental and other clinical conditions (Blanchard, Gurka, & Blackman, 2006; Gurney, McPheeters, & Davis, 2006; Kogan et al., 2009; Russ, Larson, & Halfon, 2012). In addition, evidence is growing that parents can accurately report diagnoses on web-based surveys (Daniels et al., 2012). Another potential limitation in both studies was in GMFCS level assignment. Traditionally clinicians determine these levels, but LUSTRE used parent report while the ADDM Network used chart review by clinicians who did not meet or examine the child. The validity of the parent-reported method has been documented as very high (ICC> = 0.9) in similar studies, and the author of the GMFCS supported the use of the one-page version for caregiver reporting in this study (Dietrich, Abercrombie, Bartlett, & Fanning, 2005; McDowell, Kerr, & Parkes, 2007; Morris, Galuppi, & Rosenbaum, 2004; Morris, Kurinczuk, Fitzpatrick, & Rosenbaum, 2006). Similarly, the ADDM Network has validated its “arm’s length” GMFCS level assignment method (Christensen et al., 2014). In further support of these findings of similar clinical severity levels within the two cohorts, as discussed above, the relationships held when the comparisons were made across different LUSTRE age subgroups.","Based on the findings from this study, it appears that children with CP and access to their own stored CB are less socio-economically diverse than the broader CP population. This is likely a result of the fact that families that choose to privately store CB must have the necessary financial resources to pay for private storage. However, because clinical severity and co-morbidity prevalence are similar, these groups may be more comparable than they first appear. LUSTRE may serve as a valuable source of data for the study and characterization of individuals with CP, including individuals who have or will receive CB infusions.","HH and HB are both employees of CBR Systems, Inc. PM, MK, KC, AM and CS received financial support from CBR Systems, Inc. to carry out this study. AS received compensation from CBR Systems, Inc. for consultation services.","This work was supported by CBR Systems, Inc."],["Investigating the limits of unconscious processing is essential to understand the function of consciousness. Here, we explored whether holistic face processing, a mechanism believed to be important for face processing in general, can be accomplished unconsciously. Using a novel \"eyes-face\" stimulus we tested whether discrimination of pairs of eyes was influenced by the surrounding face context. While the eyes were fully visible, the faces that provided context could be rendered invisible through continuous flash suppression. Two experiments with three different sets of face stimuli and a subliminal learning procedure converged to show that invisible faces did not influence perception of visible eyes. In contrast, surrounding faces, when they were clearly visible, strongly influenced perception of the eyes. Thus, we conclude that conscious awareness might be a prerequisite for holistic face processing. © 2014 The Authors. --------------------------------------------------------------------------------","Most people are good at recognizing faces, which is a capacity that is usually taken for granted. Yet, given that all faces are essentially very similar (e.g., all faces have a nose, eyes, and mouth; relative position of features is largely the same), the cognitive task of face recognition is far from straightforward. A key component of efficient facial processing is considered to be holistic processing (Farah, Wilson, Drain, & Tanaka, 1998; Richler, Cheung, & Gauthier, 2011) – the ability to perceive a face as a whole and not as a set of independent features. Probably the most spectacular demonstration of this phenomenon is the composite face effect (Young, Hellawell, & Hay, 1987) – that is, when a facial image is composed of the bottom and top halves of two different faces, where recognition of one half of a face (e.g., the top part) is modulated by that of the other half (e.g., the bottom part). Holistic face processing in general, and the composite face effect in particular, has been extensively explored in healthy populations using behavioral measures (for review: Rossion, 2013), functional MRI (e.g., Andrews, Davies- Thompson, Kingstone, & Young, 2010; Axelrod, 2010; Axelrod & Yovel, 2010, 2011; Schiltz, Dricot, Goebel, & Rossion, 2010; Schiltz & Rossion, 2006) and event-related potentials (e.g., Jacques & Rossion, 2009; Wiese, Kachel, & Schweinberger, 2013) as well as in participants with impaired face recognition (prosopagnosia) (Avidan, Tanzer, & Behrmann, 2011; Busigny & Rossion, 2011). However, whether conscious awareness is required for holistic face processing is not known. Understanding the role of consciousness and conscious awareness is one of the fundamental challenges of the cognitive sciences (Baars, 1993; Dennett, 1993; Koch, 2004). Empirically, the level of conscious awareness is usually evaluated by an introspective, “subjective report” (whether a participant was aware of a stimulus) and “objective measure” (forced-choice discrimination, even for subjectively unaware stimulus) (Merikle & Daneman, 1998). While qualitative differences between conscious and unconscious perception have been debated for years (e.g., Cheesman & Merikle, 1986; Peremen & Lamy, 2014; Vorberg, Mattler, Heinecke, Schmidt, & Schwarzbach, 2003), numerous behavioral (e.g., Marcel, 1983; Mudrik, Breska, Lamy, & Deouell, 2011; Sklar et al., 2012) and neuroimaging (e.g., Axelrod, Bar, Rees, & Yovel, 2014; Dehaene et al., 2001; Fahrenfort et al., 2012; Sterzer, Haynes, & Rees, 2008) studies demonstrate that information can be processed unconsciously. Unconscious face processing is one of the widely explored types of unconscious processing. A large body of evidence suggests that emotional aspects of face processing (for reviews: Pessoa & Adolphs, 2010; Tamietto & de Gelder, 2010), gaze (Chen & Yeh, 2012; Stein, Peelen, & Sterzer, 2012; Stein, Senju, Peelen, & Sterzer, 2011) and face familiarity (de Gardelle, Charles, & Kouider, 2011; Henson, Mouchlianitis, Matthews, & Kouider, 2008; Kouider, Eger, Dolan, & Henson, 2009) can be processed without conscious awareness; however, several studies have shown that facial identity (Moradi, Koch, & Shimojo, 2005; Stein & Sterzer, 2011; Stone & Valentine, 2005) and face gender/race (Amihai, Deouell, & Bentin, 2011) cannot be processed unconsciously. In the present study, we addressed the question of unconscious face processing from another angle, while asking whether holistic face processing can take place unconsciously. Based on the previous negative results of unconscious face identity/gender processing, and given that face recognition and holistic processing might share common underlying mechanisms (Richler, Cheung et al., 2011; Wang, Li, Fang, Tian, & Liu, 2012), one possibility is that holistic face processing cannot be accomplished without conscious awareness. Alternatively, it is also possible that holistic face processing is a more basic type of processing than face recognition, which implies that holistic processing might still occur unconsciously. In addition, given that holistic processing has been suggested to be an automatic process (Richler, Wong, & Gauthier, 2011), it therefore possibly could be executed unconsciously (Hasher & Zacks, 1979). In the present study, we devised a novel “eyes-face” composite stimulus that was composed of a pair of eyes plus the remaining part of the face [c.f., top and bottom image face parts (Young et al., 1987)]. Participants had to discriminate between pairs of eyes in two consecutively presented composite images while the rest of the face was either the same or different (Fig. 1A and B). Critically, in the subliminal version of the paradigm, while the eyes were always visible, the rest of the face was rendered invisible using Continuous Flash Suppression (CFS; Fig. 1B) (Harris, Schwarzkopf, Song, Bahrami, & Rees, 2011; Tsuchiya & Koch, 2005). We asked whether invisible faces influenced discrimination of the visible eyes.","In this experiment, we used three different sets of composite faces. The first set was comprised of five male composite faces (examples of faces: Fig. 1B left side). To increase the perceptual differences between images, a second set included three male and three female composite faces (examples of faces: Fig. 1C, top). The third set of images was comprised of three female faces, either with or without eyebrows (examples of faces: Fig. 1C, bottom). The motivation to include this third composite image set was that eyebrows are the closest facial feature to the eyes and, consequently, have a higher chance of being attended to when the task is eyes discrimination.","were presented with two consecutive composite faces that contained visible eyes and faces. The face part of each image was rendered invisible by CFS (Fig. 1B). The pairs of eyes in the two images were always the same, and the invisible faces were either the same or different. The task was to report whether successive pairs of visible eyes were the same or different. Participants were told that this eyes discrimination task was very difficult, and were encouraged to look out for the smallest differences between the two stimuli. Notably, because the stimuli did not appear in the exact same screen position (a small amount of spatial jitter was added; see Methods) and the eyes were surrounded by a constantly changing CFS mask, it was not evident that sequential eye stimuli were actually identical. Effect size was defined in percent units as the percent of trials answered “eyes same” when the invisible faces of the two images were the same minus the percent of trials answered “eyes same” when the invisible faces of the two images were different. An effect size larger than zero was taken as evidence of a subliminal influence of the invisible faces on judgments of the visible eyes. Participants Fifteen healthy volunteers (age 20–27 years, 10 females) participated in this experiment: all participants participated in the experiment with image set 1 (male faces), 14 of the same set of participants participated in the experiment with image set 2 (male and female faces), and 12 of the same set of participants took part in the experiment with image set 3 (faces with and without eyebrows). Two participants were excluded from the analysis of all three experiments because they reported that they could see the masked face. The experiment was approved by Tel- Aviv University ethics committee, and all participants gave informed consent to participate in the experiment.","For stimuli presentation, a CRT 17-in. color monitor was used. Screen resolution was 1024 × 768, and refresh rate was 85 Hz.","were presented using MATLAB 7.6 with Psychtoolbox (Brainard, 1997). Participants sat in a comfortable chair at a distance from the monitor of 30 cm. During the experiment, the room lights were turned off. Stimuli All image manipulations were performed in Adobe Photoshop CS2. Face stimuli (neutral face expressions) were taken from the Karolinska Directed Emotional Faces (Lundqvist & Litton, 1998). Image set 1 consisted of five male face identities (examples of faces: Fig. 1B, left side), and image set 2 consisted of three male and three female face identities (examples of faces: Fig. 1C, top). Image set 3 consisted of three female face identities – three original images and three images where the eyebrows were removed using Adobe Photoshop program (examples of faces: Fig. 1C, bottom). Five pairs of eyes (rectangle with the eyes, Fig. 1B) were cropped from different face images (not used in image sets 1–3) and integrated into each face image of the experimental image sets 1–3. The same pairs of eyes were used for all image sets. The size of the rectangle containing the eyes was identical for all pairs of eyes (see below). The resultant image sets included 25 stimuli each (five face identities with five pairs of eyes) for image set 1, and 30 stimuli each (six face identities with five pairs of the eyes) for image sets 2 and 3. None of the composite faces had original (“native”) eyes. The images were in color. Face images were positioned in the center of the rectangular background image frame, which was a light grey color (RGB: 97, 97, 97) (see Fig. 1B). Dimensions of the stimuli were as follows (in degrees of visual angle): background square frame (vertical: 27, horizontal: 27); face image (vertical: 21, horizontal: 16.5); eyes rectangle (vertical: 3.3, horizontal: 14.5). Invisibility manipulation To render stimuli invisible, we used Continuous Flash Suppression (CFS) (Tsuchiya & Koch, 2005). The paradigm was designed in such a way that the eyes rectangle was fully visible with both eyes, whereas the remaining part of the face was invisible (Harris et al., 2011). During the experiment, participants wore cardboard anaglyph red/cyan glasses. The visible part of the stimulus was projected using all three RGB colors. The invisible part of the stimulus was projected using the red color channel (visible using red filter), while a Mondrian mask was projected using green/blue channels. The Mondrian mask appeared over the whole background rectangle frame with the exception of the visible stimulus part (see Fig. 1B). The position of the visible eyes rectangle was fixed relative to the Mondrian image. In addition, on the Mondrian masks, we drew an elliptic contour line at the location of the invisible face (see Fig. 1B, right side). The motivation for adding this contour ellipse was to urge participants to make eye judgments as though the eyes were part of a face. The elliptic contour line was created by increasing the brightness of the corresponding mask pixels by 20%. The ellipse contour was created at the same position for all mask images and was at a fixed position relative to the eyes rectangle. Mondrian masks were prepared by randomly scrambling a kaleidoscope image. Our preliminary pilots showed that use of this pattern achieved higher invisibility effects than the geometrical shapes usually used (e.g., Tsuchiya & Koch, 2005). Mondrian masks were replaced continuously at a frequency of 10 Hz (every 100 ms). During the invisible part of Experiment 1, the face composite stimuli were always projected to the non-dominant eye of the participant while the CFS mask was presented to the dominant eye. Eye dominance was tested by asking participants to view a distant object through a hole made by the fingers of their two hands (“Miles test”) (Mendola & Conner, 2007; Miles, 1930). In the visible sessions, the participants wore specially prepared glasses with two red lenses (stimuli visible through both eyes and mask invisible through both eyes). Using this approach, we preserved the same quality of stimulation for visible and invisible sessions. Experimental design For each one of three image sets, there were three experimental sessions: eyes discrimination session with invisible face images, awareness invisibility test of discriminating invisible face images, and eyes discrimination session with visible face images. To minimize the number of switches between tasks, participants first performed three sessions of eyes discrimination with invisible face images (all three image sets), then three sessions of awareness invisibility (all three image sets; explained below), and finally three sessions of eyes discrimination with visible face images (all three image sets). The order of the image sets within these three sessions was counterbalanced across participants. From the side of stimuli presentation (computer code), all three experimental sessions were exactly the same for each image set. A trial consisted of two consecutively presented stimuli (each stimulus duration = 0.3 s, interstimulus interval [ISI] = 0.1 s) (Fig. 1A). There was a random position jitter between two stimuli (2% of stimulus size). The eyes in two images of a trial were always the same, whereas the faces were either the same (50% of the trials) or different (50% of the trials). For image set 1 (males only set), the different faces were of different male face identities. For image set 2 (males and females set) the different faces were each composed of different identities from opposite genders (either male–female or female–male, counterbalanced). For image set 3 (with and without eyebrows), the different faces were the same identity presented twice with and without eyebrows (either no eyebrows – eyebrows or eyebrows – no eyebrows). In the “invisible” sessions, only the eyes were visible; in the “visible” sessions (at the end of the experiment), both the eyes and faces were visible. Participants were asked to make their response after the second stimulus disappeared. In the eyes discrimination experiment, participants had to press ‘1’ if the two pairs of eyes were the same and ‘2’ if the two pairs were different. The fact that the pairs of eyes were always the same was not known by participants. Participants were told before the experiment that the task would be very difficult and that any slight difference that they perceived between the two eye pairs should be taken as an indication of the presence of a difference between the eyes. Our preliminary pilot tests showed that presenting the two stimuli with a slight spatial jitter and a constant change in CFS mask around the eyes created an impression that the eyes were not actually identical. Indeed, at the informal debriefing after the experiment, participants indicated that they had seen differences between the pairs of eyes. Awareness test sessions were very similar to the eyes discrimination sessions; the only difference was that participants had to discriminate between invisible faces (‘1’: same faces, ‘2’: different faces); since the participants admitted seeing nothing, they were encouraged to guess. Participants were asked not to use the visible eyes in the discrimination; instead, they were encouraged to try to discriminate based on the non-eye part of the image (inside the ellipse). Each session consisted of 50 trials (25 trials with the same faces and 25 with different faces). At the beginning of the experiment, before the first session with invisible face images, participants underwent a short training session (10 trials) of eyes discrimination with invisible faces. As face stimuli, we used two male identities that were not used afterwards in the experimental sessions.","Data were analyzed using MATLAB and SPSS 17 software. The effect size for the eyes discrimination task was calculated as the percent of trials answered “same” when the face identities of the two images were actually the same minus the percent of trials answered “same” when the face identities were different. An effect size larger than zero provided evidence that a change in face influenced perception of the eyes. Significance at the group level was established using non-parametric Wilcoxon Signed Rank test (signrank MATLAB function). In the visual awareness tests, the individual d-prime values [signal detection theory (Macmillan, 2002)] of discrimination invisible stimuli (correct/incorrect) were assessed using non- parametric Wilcoxon Signed Rank test vs. zero. For all the parametric tests (ANOVAs and correlations), the data was first tested for normality using Lilliefors test (lillietest MATLAB function). Reaction times of participants were not analyzed since the participants were asked to respond only when the second stimulus of the trial disappeared, and they were asked to maximize accuracy and not to minimize response time. To calculate the Bayes factor, we used the online calculator of Zoltan Dienes (http://www.lifesci.sussex.ac.uk/home/Zoltan_Dienes/inference/Bayes.htm). Results: Experiment 1 ~~~~~~~~~~~~~~~~~~~~~ Judgment of visible eyes was not influenced by invisible faces in any of the image sets (Fig. 2, left side) [males only set 1: effect size: −3.4%, MSE: 3.7%, p = .26, Wilcoxon sign-rank = 29.5, one Sample Wilcoxon Signed Rank test vs. 0; males and females set 2: effect size: 4.1%, MSE: 3.7%, p = .29, Wilcoxon sign-rank = 27; with and without eyebrows set 3: effect size: −6.4%, MSE: 4.8%, p = .19, Wilcoxon sign-rank = 11.5]. For raw response rates of “same eyes” answers, see Table 1, first row. Invisibility of faces was tested in separate sessions, with stimulus configurations identical to the main experiment, but the participants were required to discriminate between the invisible faces. Discrimination of faces did not differ significantly from chance for all three image sets [males only set 1: d-prime: 0.082, MSE: 0.11, p = .47, Wilcoxon sign-rank = 25; males and females set 2: d-prime: 0.15, MSE: 0.2, p = .50, Wilcoxon sign-rank = 25.5; males and females set 3: d-prime: 0.028, MSE: 0.13, p = .95, Wilcoxon sign-rank = 22] (raw response rates, Table 1, second row). Contrarily to invisible faces, judgment of visible eyes was strongly influenced by visible faces in all image sets (Fig. 2, right side) [males only set 1: effect size: 51.7%, MSE: 7%, p = .001, Wilcoxon sign-rank = 0; males and females set 2: effect size: 54.2%, MSE: 7.2%, p = .002, Wilcoxon sign-rank = 0; with and without eyebrows set 3: effect size: 42.8%, MSE: 7.6%, p = .005, Wilcoxon sign-rank = 0] (raw response rates, Table 1, third row). The difference between invisible and visible faces was confirmed by the two-way repeated measured ANOVA analysis with image set and face visibility as factors: there was a highly significant main effect of face visibility: F(1, 9) = 118.425, p < .001 but non-significant main effect of image set [F(2, 18) = 3, p = .075] and non-significant interaction between image set and face visibility [F(2, 18) < 1]. Experiment 1 found no evidence that invisible faces influence the perception of visible eyes. Notably, the “null effect” for invisible faces may stem either from a lack of sensitivity or, alternatively, from the phenomenological absence of unconscious processing (Dienes, 2011, in press). Since the orthodox (Neyman and Pearson) statistical approach is unable to distinguish between these two possibilities, we adopted a Bayesian approach (Dienes, 2011). That is, the Neyman and Pearson approach, by definition, might only find support for the alternative hypothesis (rejecting H0 and accepting H1), but cannot provide support for the H0 hypothesis (when H1 is not accepted). A Bayesian approach (Bayes factor), in contrast, can provide support for either H0 or H1 (Dienes, in press). As the data input parameters, we used effect size and mean square error (Jiang et al., 2012). To model the expectation parameters, following the recommendation of Dienes (Dienes, 2011), we consulted previous studies that also explored unconscious face processing using the CFS paradigm (Amihai et al., 2011; Moradi et al., 2005; Yang, Hong, & Blake, 2010). In particular, the effect size during unconscious face processing in these studies was at least two times smaller compared to that during conscious processing (Yang et al., 2010), whereas in some cases it was five to ten times smaller (Amihai et al., 2011; Moradi et al., 2005). Therefore, given that the effect size in the conscious condition of our study was on average 50%, the unconscious mean prediction was defined as 7.5% and the standard deviation was set to 5% (two-tailed normal distribution). The results of the analysis revealed the following: for the males-only set 1, the Bayes factor was 0.07; for the males-and-females set 2, the Bayes factor was 0.95; and for the eyebrows set 3, the Bayes factor was 0.23. Thus, the Bayes factor values in sets 1 and 3 provide strong support for the phenomenological absence of unconscious processing (Bayes factor < 0.33); the results of set 2 should be interpreted as a “lack of sensitivity” (0.33 < Bayes factor < 3) (Dienes, 2011).","In the current experiment, we asked whether there was a way to improve unconscious holistic processing by means of subliminal learning with feedback (e.g., Atas, Faivre, Timmermans, Cleeremans, & Kouider, 2014; Di Luca, Ernst, & Backus, 2010; Nishina, Seitz, Kawato, & Watanabe, 2007; Rosenthal & Humphreys, 2010; Watanabe, Náñez, & Sasaki, 2001).","underwent half an hour of subliminal learning with feedback, which was based on same/different invisible faces (see Methods). We hypothesized that by means of learning, we could induce the invisible face to modulate perception (discrimination) of the visible eyes. The flow of this experiment is shown in Fig. 3. This experiment used the image set with five male faces (two versions: intact and shifted eyes; described below). To verify that, as a result of repetitive exposure during learning, the invisible part of the stimulus (face) did not become visible (e.g., Atas, Vermeiren, & Cleeremans, 2013; Schwiedrzik, Singer, & Melloni, 2009, 2011), participants underwent an invisible face awareness test before and after learning. In addition, unrelated to subliminal learning, participants were tested with a set of stimuli in which the original five male stimuli were manipulated in such a way that the rectangle of the eyes in all images was shifted (Fig. 4A). Similar manipulations are effective in disrupting holistic face processing (e.g., Axelrod & Yovel, 2010; Maurer, Grand, & Mondloch, 2002; McKone, Kanwisher, & Duchaine, 2007). Our plan was to compare the magnitude of any influence of invisible faces between an intact and shifted eyes stimuli set. Critically, in order to conclude that any effect, if found, was related to holistic processing, the magnitude of this effect must be higher for the intact eyes than for the shifted eyes stimuli set. Participants Eighteen healthy volunteers (ages 19–30, 13 females) participated in this experiment. Two participants were excluded from the analysis because they reported consciously seeing the masked face. The experiment was approved by Tel-Aviv University’s ethics committee, and all participants gave informed consent to participate. Apparatus The same as in Experiment 1. Stimuli The image set with five male identities from Experiment 1 was used in this experiment. In addition, based on this image set, we created a second image set where the eyes rectangle was shifted to the left 6.5 degrees of a visual angle (see Fig. 4A). The empty eye position was filled with a uniform average color taken from the surrounding facial features (RGB: 217, 112, 67). The horizontal size of the background rectangle frames for both intact and shifted eye sets were increased to 35° of a visual angle. Invisibility manipulation The same as in Experiment 1. Experimental design The flow of the experiment is presented in Fig. 3. For each of the two image sets (intact and shifted eyes), there were five experimental sessions: two sessions with invisible face images, where eyes had to be discriminated (one before and one after learning), two awareness invisibility tests of discriminating invisible face images (one before and one after learning), and an eyes discrimination session with visible face images. The awareness invisibility tests always followed the eyes discrimination sessions with invisible face images. The eyes discrimination session with visible face images was always the last session of the experiment. The order of intact and shifted eyes sessions was counterbalanced between participants. All sessions of this experiment (including the learning sessions; see below) consisted of 30 trials (15 trials with same faces and 15 with different faces). The learning procedure included 14 sessions. These learning sessions had the same design as the testing sessions (eyes discrimination test) with the exception of a correct/incorrect indication after each trial (feedback) and overall score at the end of the session. While participants’ task was to discriminate between pairs of eyes, the correct/incorrect answer indication was based on same/different invisible faces. In particular, the answer was defined as correct if the participant answered “same” when the two invisible identities were the same or answered “different” when the two invisible identities were different. At the end of each learning session, participants received their session scores (percentage of correct answers). At the informal debriefing after the experiment, participants were asked based which parameters had served as the basis for their development of the ability to discriminate between the pairs of eyes. All participants indicated that their decision had been based on either eye shape, distance between the eyes, or eye color.","The analysis procedure of non-learning sessions was the same as in Experiment 1. The significance of a learning effect was evaluated using linear regression analyses, which were applied for group (averaged) and individual data. The procedure for the group-level linear regression analysis was as follows: (1) effect size values for each session were averaged across participants resulting in averaged effect size values (Fig. 4C); (2) these effect size values were submitted to linear regression; and (3) a regression line slope coefficient significantly different from zero was used to support the learning effect (Weisberg, 2005). The procedure for the individual-level linear regression analysis was as follows: (1) for each participant, the effect size values were submitted to linear regression; (2) the individual linear regression slope coefficients were obtained and submitted to Wilcoxon Signed Rank test (slope coefficients vs. zero test). Results: Experiment 2 ~~~~~~~~~~~~~~~~~~~~~ In line with the results of Experiment 1, before learning, invisible faces did not influence the judgment of visible eyes for both intact and shifted eye sets (Fig. 4B, leftmost bars) (intact eyes: effect size: −2.1%, MSE: 3.9%, p = .69, Wilcoxon sign-rank = 60.5; shifted eyes: effect size: −5%, MSE: 4.3%, p = .25, Wilcoxon sign-rank = 29) (raw response rates, Table 2, first row). The Bayes factor, modeled with the same parameters as in Experiment 1, was 0.23 for the intact eyes set and 0.21 for the shifted eye set. Thus, for both image sets, the results of the Bayesian analysis (Bayes factor) suggested the absence of unconscious processing. Participants then completed a subliminal learning task with feedback using the intact eyes image set. Average learning results across participants are shown in Fig. 4C; as seen, performance gradually improved across sessions (slope coefficient = 1.12, significantly different from zero: t(13) = 5.56, p < .001). To examine the effect of learning at an individual level, a linear regression model was estimated for each participant; the resultant individual slope coefficients were significantly above zero [Wilcoxon Signed Rank test: p = .0086, Wilcoxon sign-rank = 5.5]. After learning, participants were tested again (sessions without feedback), and we observed a significant effect of the invisible face on visible eyes judgments for both intact and shifted eyes sets (Fig. 4B, middle bars) [intact eyes: effect size: 10.4%, MSE: 4.3%, p = .029, Wilcoxon sign-rank = 26; shifted eyes: effect size: 11.2%, MSE: 3.4%, p = .013, Wilcoxon sign-rank = 7.5] (raw response rates, Table 2, second row). Repeated- measures ANOVA, with eye position (intact/shifted eyes) and time of test (before/after learning) as factors, revealed a highly significant main effect of time of test [F(1, 15) = 11.826, p = .004]; however, no significant effect of eye position [F(1, 15) < 1] and no significant interaction [F(1, 15) < 1] were found, which confirms the similar effect size for intact and shifted eyes. In addition, using the Bayesian approach, we tested whether the absence of difference between unconscious processing for intact and shifted eyes should be interpreted as an insufficient experimental sensitivity or as positive evidence in support of the absence of unconscious holistic processing. The data mean value was the difference in effect size between intact and shifted eyes (equal to 0.83%), and the mean standard error of difference was 0.7%. The modeling parameters were estimated based on the reduction of the effect for unconscious processing compared to conscious processing (as in Experiment 1) and the effect size for intact vs. shifted visible faces (around 50%). Thus, the model’s mean was set to 5% and the standard deviation was set to 5% (two-tailed normal distribution). The Bayes factor we found was 0.16, unequivocally suggesting the absence of unconscious holistic processing. Interestingly, as can be seen in Table 2 (second vs. first raw), the learning effect was associated with an increase in the proportion of “same” responses for the “same” invisible faces (but no major change for “different” invisible faces). To test this statistically, for raw responses (“same” responses), we ran three-way repeated-measures ANOVA with eyes position (intact/shifted eyes), time of test (before/after learning) and invisible stimulus type (same/different faces) as factors. We found significant two-way interaction between time of test and invisible stimulus type F(1, 15) = 11.826, p = .004]. The follow-up repeated-measures ANOVA only for “same” invisible faces with eyes position (intact/shifted eyes) and time of test (before/after learning) as factors revealed a significant main effect of time of test [F(1, 15) = 6.025, p = .027], no significant effect of eye position [F(1, 15) = 2.596, p = .128] and no significant interaction [F(1, 15) < 1]. Similar repeated-measures ANOVA only for “different” invisible faces revealed no significant effects (insignificant main effect of time of test [F(1, 15) < 1], eye position [F(1, 15) = 2.163, p = .162] and interaction between them [F(1, 15) < 1]). Thus, we conclude that the learning was associated with increasing the number of “same” responses for “same” invisible faces, but no change for “different” invisible faces. To ensure that faces were genuinely invisible, we ran awareness tests before and after learning (Fig. 3). These awareness tests confirmed that the faces were invisible before [intact eyes: d-prime: 0, MSE: 0.1, p = .84, Wilcoxon sign-rank = 36.5; shifted eyes: d-prime: 0.023, MSE: 0.1, p = .97, Wilcoxon sign-rank = 45] and after [intact eyes: d-prime: 0.069, MSE: 0.096, p = .43, Wilcoxon sign-rank = 46; shifted eyes: d-prime: 0.097, MSE: 0.12, p = .39, Wilcoxon sign-rank = 45] training. For both types of eyes, no differences existed in visibility level before or after learning [intact eyes: p = .63, Wilcoxon sign-rank = 51.5; shifted eyes: p = .68, Wilcoxon sign- rank = 60] and no significant correlation existed across participants between the level of face discrimination (awareness test) and level of eyes influence effect (eye judgment main task) [intact eyes: r(15) = 0.15, p = .56; shifted eyes: r(15) = 0.28, p = .28]. Taken together, these results suggest that because of learning, participants’ judgment of eyes was influenced by invisible faces. Critically, because the results for intact and shifted eyes were similar, we conclude that the effect is not related to unconscious holistic face processing [the classical composite face effect (Young et al., 1987)]. Finally, we tested the “eyes-face” stimulus but now with visible faces (Fig. 4B, rightmost bars). The findings revealed that, while in the intact eyes condition, faces influenced the perception of the eyes significantly [effect size: 48.3%, MSE: 6.6%, p < .001, Wilcoxon sign-rank = 1.5], the effect was completely abolished for the shifted eyes [effect size: 2.9%, MSE: 3.5%, p = .44, Wilcoxon sign-rank = 34.5]. This finding confirms that the shifted eyes manipulation disrupted holistic face processing. To examine the difference between visible and invisible perception (after learning), we ran a two-way repeated- measures ANOVA with visibility level (visible/invisible) and eye position (intact/shifted) as factors. The results showed a significant main effect of visibility level [F(1, 15) = 11.7, p = .004], a significant main effect of eye position [F(1, 15) = 26.5, p < .001], and a significant interaction [F(1, 15) = 22.9, p < .001]. To explore different patterns of visible and invisible processing further, we ran a post hoc non-parametric Wilcoxon Signed Rank test, which revealed a larger effect size for visible compared to invisible for intact eyes [p < .0012, Wilcoxon sign-rank = 5.5] and a trend for a larger effect size for invisible compared to visible for shifted eyes [p < .077, Wilcoxon sign-rank = 24.5].","The objective of the current study was to test whether holistic face processing can occur outside of conscious awareness. We used a novel “eyes-face” stimulus, where discrimination between sets of eyes was affected by the presentation of a visible congruent or incongruent face. Using different sets of composite face stimuli, we showed that visible but not invisible faces influenced perception of visible eyes. Moreover, even after subliminal learning, when invisible faces biased judgments of visible eyes, this effect was not found to be related to holistic face processing. Thus, we conclude, conscious awareness may be necessary for holistic face processing. The question of whether face recognition and face identity-related processing can be accomplished unconsciously has been at the focus of research in recent years. In particular, several studies have shown that neither facial identity (Moradi et al., 2005; Stein & Sterzer, 2011; Stone & Valentine, 2005) nor face gender/race (Amihai et al., 2011) is processed unconsciously. In the current study, we hypothesized that, if holistic processing is a more basic component of face recognition, then it might be possible to find evidence for unconscious holistic face processing. To increase the chances of identifying this effect, several steps were taken. First, the “eyes-face” stimulus used was optimized to generate a strong effect for visible faces. That is, studies with a classical composite face illusion stimulus (top and bottom face halves) often use response time to index holistic processing because response accuracy measures are often not sensitive enough (e.g., de Heering & Rossion, 2008; Wang, Li et al., 2012). Here, given the strong response accuracy effect for visible faces (Fig. 4B), relatively large amount of room is left for a potential effect with invisible faces. Second, in Experiment 1, we used different sets of faces while aiming to maximize perceptual differences between the faces (e.g., male and female faces). Third, by using the image set in Experiment 1 with and without eyebrows, we ensured that the facial features (eyebrows) which were essential for inducing the effect were as close as possible to the visible eyes rectangle. Finally, we employed a subliminal learning procedure, which was successful in the sense that, as a result of learning, invisible faces influence responses regarding visible eyes (discrimination of visible eyes). However, the fact that similar effects were found for intact and shifted eyes ruled out the interpretation that the effect was related to holistic face processing. Thus, given that despite all aforementioned steps, no unconscious holistic face processing effect could be found, we suggest that this process might not be processed unconsciously, at least for the condition of dichoptic stimulation employed here. The absence of holistic face processing outside conscious awareness is also interesting to consider in light of the recent proposal that the holistic phenomenon, as it is measured in a composite task (e.g., Young et al., 1987), is a result of automatic processing with a failure to allocate covert attention (Richler, Cheung et al., 2011). Accordingly, given that some define automatic processes as unconscious (e.g., Hasher & Zacks, 1979), one could have expected that face holistic processing might be executed unconsciously. Several lines of evidence can reconcile our result with these expectations. First, the link between automatic and unconscious processing is frequently made for the highly learned processes, like car driving (e.g., Charlton & Starkey, 2011), which might involve awareness mechanisms different from the sensory (un)awareness for masked stimuli used in our study. Second, the relationship between attention and consciousness is highly debated (for reivews: Koch & Tsuchiya, 2007; Marchetti, 2012; Van Boxtel, Tsuchiya, & Koch, 2010) and it has been shown, for example, that two processes can be dissociated (e.g., Naccache, Blandin, & Dehaene, 2002). Yet, it is not clear whether a failure to allocate covert attention, which might be a result of automatic processing (Richler, Cheung et al., 2011), can occur for invisible, unconscious stimuli. Finally, a similar hypothesis had been proposed with regard to the Stroop task (Stroop, 1935), which is also automatic (MacLeod, 1991) and therefore can be processed unconsciously (Marcel, 1983). Yet, the empirical support for this hypothesis is rather controversial: while some studies do find an unconscious Stroop effect (e.g., Marcel, 1983), others claim that such an effect can be explained by conscious awareness (Tzelgov, Porat, & Henik, 1997) or partial awareness (Kouider & Dupoux, 2004). Thus, in light of the evidence provided, the absence of unconscious holistic face processing might not be that surprising. The results reported here are interesting to consider in the context of unconscious processing of visual context in general. The invisible surroundings can influence the orientation of a centrally presented visible grating (Clifford & Harris, 2005), and the invisible surrounding luminance can modulate the perceived brightness of the centrally presented visible circle (Harris et al., 2011). In addition, two studies have explored the perception of illusory contours surrounded by an invisible context, and reported mixed results. One study found, using breaking continuous flash suppression (Jiang, Costello, & He, 2007) that an invisible Kanizsa triangle emerged into awareness faster than did a control stimulus (Wang, Weng, & He, 2012). The second asked participants to indicate the direction of the Kanizsa triangle induced by the invisible surroundings and reported only chance performance (Harris et al., 2011). Notably, the invisible context and the stimuli explored in the studies described above were relatively simple stimuli, which are known to be processed at a relatively low level within the visual hierarchy (e.g., V1 for line orientation, V4 for color processing). In contrast, faces are processed within high-level regions of the occipito-temporal cortex (Haxby, Hoffman, & Gobbini, 2000); therefore, it is plausible that low-level, but not high-level, context can be processed unconsciously. Finally, Mudrik and colleagues recently demonstrated an interesting example of invisible context processing (Mudrik et al., 2011), where the authors showed that invisible scenes containing incongruent invisible objects emerged into awareness faster than did congruent scenes (see also: Mudrik & Koch, 2013). Yet, as the invisible context in this study was semantic and not purely visual, there is no straightforward way to relate their findings to ours. What might be the possible underlying neural mechanisms of the effect we observed? The neural network of face processing is well characterized (Haxby et al., 2000), while the commonest and the most reproducible region across participants is the Fusiform Face Area (FFA) (Kanwisher, McDermott, & Chun, 1997). The FFA exhibits various properties pertinent to face processing, such as partial view-invariance (e.g., Axelrod & Yovel, 2012; Grill-Spector et al., 1999; Kietzmann, Swisher, König, & Tong, 2012) and discrimination between face identities (e.g., Gilaie-Dotan & Malach, 2007; Nestor, Plaut, & Behrmann, 2011; Rotshtein, Henson, Treves, Driver, & Dolan, 2005). More critical for the current discussion, however, is that the FFA has been suggested to be responsible for holistic face processing (Andrews et al., 2010; Axelrod & Yovel, 2010, 2011; Schiltz & Rossion, 2006; Zhang, Li, Song, & Liu, 2012). If this is indeed the case, and assuming that unconscious information can reach the FFA (e.g., Fahrenfort et al., 2012; Moutoussis & Zeki, 2002; Sterzer et al., 2008), then our result suggests that the FFA operations responsible for holistic processing are associated with awareness; however, one should also consider that the network of face-processing regions is distributed (Gobbini & Haxby, 2007) and spans not only the occipito-temporal cortex (Op de Beeck, Haushofer, & Kanwisher, 2008) but also the frontal lobes (Axelrod & Yovel, 2013; Ishai, Schmidt, & Boesiger, 2005). As such, holistic processing might also require interactive processing between several different brain regions – the type of distributed, long-range processing that was proposed to be associated with awareness (e.g., Dehaene & Changeux, 2011; Dehaene, Kerszberg, & Changeux, 1998). Overall, future research will be needed to gain a deeper understanding of the underlying neural correlates. A special note should be made regarding the subliminal learning procedure employed here. As discussed, the fact that, during and after learning, invisible faces influenced both intact and shifted eyes suggests that the effect was not related to holistic processing. We also found that the learning was associated with a more frequent “same” answer for the “same” invisible faces, but no change for “different” invisible faces. Yet, as the goal of the current study was not to explore the mechanisms of subliminal learning but rather to use this type of learning as a tool, it is not possible to clearly determine exactly what type of information was learnt and how the learning occurred. It is plausible that, in the course of learning, participants unconsciously learnt to associate between the “change” or “no change” of invisible images (faces) and the required response for visible eyes. In other words, the same/different faces were treated as just same/different images. In addition, it is also possible that motor response mapping associative learning occurred (Damian, 2001) while participants learnt to press ‘1’ when two invisible images were the same and ‘2’ when they were different. Notably, the learning procedure did not influence face awareness, which was equally invisible before and after learning. Finally, it should be noted that according to an alternative view of holistic processing, the effect that is measured by a composite task might not be of a perceptual but rather a decisional nature (e.g., Richler, Gauthier, Wenger, & Palmeri, 2008; Richler, Palmeri, & Gauthier, 2012; Rossion, 2013). In our study, to increase the sensitivity of the design, we included only the pairs of the same eyes. That is, by requiring participants to discriminate between the same eyes, we imposed the participants to set the same-different decision boundary at a low level. Indeed, this design was very sensitive, as we found that context had large effects on eye perception for visible faces. Yet, this design had its downside as well; since no different pairs of eyes were included, our ability to estimate potential response bias was limited (e.g., Richler et al., 2012; Rossion, 2013). Critically, the main goal of the present study was to explore whether invisible faces influence the discrimination of visible eyes, regardless of the nature of holistic processing. To conclude, in the current study we explored the question of whether holistic processing can take place outside of conscious awareness. Using three different sets of face images and applying a procedure of subliminal learning, we demonstrated that conscious visual awareness might be a prerequisite for holistic processing."],["The typical empirical approach to studying consciousness holds that we can only observe the neural correlates of experiences, not the experiences themselves. In this paper we argue, in contrast, that experiences are concrete physical phenomena that can causally interact with other phenomena, including observers. Hence, experiences can be observed and scientifically modelled. We propose that the epistemic gap between an experience and a scientific model of its neural mechanisms stems from the fact that the model is merely a theoretical construct based on observations, and distinct from the concrete phenomenon it models, namely the experience itself. In this sense, there is a gap between any natural phenomenon and its scientific model. On this approach, a neuroscientific theory of the constitutive mechanisms of an experience is literally a model of the subjective experience itself. We argue that this metatheoretical framework provides a solid basis for the empirical study of consciousness. --------------------------------------------------------------------------------","The current scientific view about the relationship between consciousness and neural activity is ambivalent. On the one hand, the enterprise to solve the mystery of consciousness using neuroscience is built on the physicalistic idea that consciousness is identical with neural activity: We know that the brain is necessary for consciousness, and neuroscience has begun to unravel the neural mechanisms that enable subjective, conscious experience (Dehaene & Changeux, 2011; Koch, Massimini, Boly, & Tononi, 2016; Laureys, 2005). On the other hand, scientists and philosophers often assume that science can merely observe correlations between neural activity and consciousness, not consciousness per se. After all, how could neural activity be identical with consciousness given that the two appear so different? For this reason, neuroscientists search for the neural correlate of consciousness (NCC), which has been defined as the minimal set of neural processes that are together sufficient for a specific conscious experience (Crick & Koch, 1990). Here, we attempt to shed light on this discrepancy. The present article is an attempt to formulate a plausible and solid metatheoretical foundation for the neuroscientific study of consciousness. This is, in our view, an imperative objective for consciousness science. While consciousness science has established itself as a specific field of empirical science, it seems to be still searching for its identity: Arguably, it is a field that does not yet have a clearly defined explanandum. This is likely to foster false disagreements between theories of consciousness, and contribute to the perception that consciousness science is a less rigorous field than other fields of science (Michel et al., 2018). In this paper, we defend a form of physicalistic monism, which holds that all concrete phenomena are physical—this is typically the starting point in modern science. Complex phenomena, such as those studied by chemistry, biology, or neuroscience, belong to the same ontological category as fundamental physical processes (those studied by fields such as particle physics), and are based on them. We could say that “physical” denotes the set that contains all concrete phenomena, and terms like “biological” or “neural” denote subsets of that all-encompassing set. By the term “physical” we mean the ontological category of phenomena studied by physics—thus, on our reading “physical” is a placeholder for whatever is the fundamental nature of phenomena studied by physics. As we will argue below, we take subjective experiences or consciousness to be concrete natural phenomena (Revonsuo, 2006; Searle, 2002; Strawson, 2008b). If we endorse the physicalistic premise that all concrete natural phenomena are physical, it follows that our experiences are physical phenomena. This might appear to be a contradictory claim, given that the term “physical” is often defined in opposition to “mental”. For instance, according to the Oxford Dictionary, “physical” is something “relating to the body as opposed to the mind”, or as “relating to things perceived through the senses as opposed to the mind” (emphasis added) (Physical, n.d.). Here we do not adhere to this standard usage of “physical”, because it a priori commits us to dualism. On our reading, terms like “experiental” or “conscious” denote certain natural phenomena which belong to the class of physical phenomena. The challenge for physicalists who consider experiences to be identical with some neural processes is to explain why the two appear so different, or why there is an epistemic gap between the two. For instance, a person with total congenital achromatopsia arguably cannot know what it is like to see the vivid colors of a sunset on a clear evening, even if she perfectly knew all the details of how the human brain works. It appears that science cannot tell us anything about the qualitative or subjective aspects of experiences—the colors of the sunset, the bitterness of coffee, or the pain of having a manuscript rejected—and thus cannot completely capture their nature (e.g. Chalmers, 1996; Jackson, 1986; Nagel, 1979). The difficulty of conceiving subjective experiences—or “qualia”, to use the philosophical term referring to the “qualitative feelings” associated with experiences—as an object of scientific study has led some empiricists to explicitly avoid referring to them (Dehaene, 2014), or to claim that such aspects do not exist (Dennet, 1991; Dennett, 2018). Many attempts in philosophy have been made to explain where the epistemic gap stems from, and the so-called phenomenal concepts strategy is probably the most popular alternative (for a good overview, see Balog, 2012). It aims to reduce the gap into differences between how phenomenal and non-phenomenal concepts refer. Non- phenomenal concepts have a mode of presentation that is distinct from their referent (e.g., we recognize water based on its perceptual features, but not everything that appears to be water is water), but phenomenal concepts are suggested to be more intimately tied with their referent. For instance, it is claimed that phenomenal concepts are recognitional concepts whose mode of presentation is identical with their referent (e.g. Loar, 1990); that phenomenal concepts are similar to indexicals like “here” and “now” (e.g. Ismael, 1999); or that they are quotational concepts, so that e.g. “pain” refers to “that state:___”, where the blanks are filled in with pain itself (e.g. Papineau, 2002). Our main focus in this paper is not on differences between how phenomenal and scientific concepts refer, but instead on the question of how, and to what extent, science can model consciousness. In this paper we take it as our premise that experiences, in all their qualitative richness, are concrete physical phenomena (Strawson, 2008a). To emphasize that empirical explanations of consciousness aim to characterize the causal physical processes that are identical with consciousness—not processes that merely correlate with it—we use the concept “constitutive mechanisms of consciousness” (CMC; Revonsuo, 2006). As discussed later in detail, by CMC we mean all those physical processes that together make up consciousness. Similar to biologists who try to describe the hierarchical physical processes that constitute, for example, a living cell or an amoeba, consciousness science aims to describe the hierarchical constitution of the processes that constitute consciousness. This approach is superficially similar to the aim of describing the NCC. The NCC and CMC approach share the idea that consciousness can be described as a hierarchical process: lower level phenomena (molecules, neurons, their activity) combine to produce a system-level phenomenon (a network of neurons and brain areas), which corresponds to consciousness. The crucial difference between these two views is that the CMC approach implies that the physical phenomenon the scientific theories describe is consciousness. The NCC approach, in contrast, suggests that the theories describe a physical phenomenon that merely correlates with consciousness. This brings us back to the aim of the present paper: What does it mean to claim that certain neural activation patterns are identical with experiences, given that the two appear so different? As a summary, we aim to characterize how subjective experiences are related to empirical observations and models about their constitutive mechanisms, and why there appears to be an epistemic gap between the two. We propose that the epistemic gap is distinctness between the scientific CMC-model and the concrete experience. A scientific model is always distinct from the concrete phenomenon it models, and in this sense there is an epistemic gap between any phenomenon and its model, not just experiences. As suggested by Hawking and Mlodinow (2010), through science we can never know the nature of the world in itself, we can only model it based on observations. This is because scientific models (like our everyday models of the world) are phenomena in the scientists’ brains: they are mental models that aim to explain and predict scientific observations of external phenomena. Thus, scientific models of subjective experiences are themselves conscious phenomena in the minds of scientists. Our framework implies that our subjective experiences are different from other natural phenomena solely because they are the only phenomena in the universe that constitute our subjective realm. If we assume that all the properties of experiences (including what they feel like) are physical and causally efficacious, then they can causally interact with measuring devices. Thus, it is possible to observe and model consciousness itself, not just its correlates. We call our approach Naturalistic Monism, because it shares some key components with Russellian Monism (RM), but is completely naturalistic: it does not postulate the existence of properties beyond the scope of natural science. This article is structured as follows. In the second part, we explain how our approach is motivated in particular by Strawson's (2008a, 2008b) version RM, but also how our view significantly differs from it. In the third part of the article, we describe in what sense scientific theories are distinct from the phenomena they model, and what implications this has for consciousness research. In the fourth section, we discuss the implications of Naturalistic Monism from the perspective of science of consciousness, and conclude that neuroscience can study consciousness, not just its correlates. Finally, in the fifth section we briefly discuss the philosophical implications of our approach.","Among philosophers, RM has recently gained popularity as an explanation of how experiences are related to their neural mechanisms (e.g. Alter & Nagasawa, 2012; Chalmers, 2016; Goff, 2017; Montero, 2015; Schneider, 2017). It promises to account for why science appears to have trouble in explaining the subjective aspects of experiences, while maintaining that experiences are nevertheless identical with their neural mechanisms. Naturalistic Monism shares two premises with RM: First, it holds that experiences are concrete natural phenomena; Second, it accepts that science has in some sense limited access to their nature. The crucial difference between Naturalistic Monism and RM concerns how the second premise should be interpreted. Strawson (2008a, 2008b) calls his version of RM “Real Materialism”, because it is both materialistic (or physicalistic; Strawson uses the terms interchangeably) and realistic about experiences. He writes: “Realistic physicalists, then, grant that experiential phenomena are real concrete phenomena—for nothing in life is more certain—and that experiential phenomena are therefore physical phenomena. It can sound odd at first to use ‘physical’ to characterize mental phenomena like experiential phenomena, and many philosophers who call themselves materialists or physicalists continue to use the terms of ordinary everyday language, that treat the mental and the physical as opposed categories. It is, however, precisely physicalists (real physicalists) who cannot talk this way, for it is, on their own view, exactly like talking about cows and animals as if they were opposed categories. Why? Because every concrete phenomenon is physical, according to them. So all mental (experiential) phenomena are physical phenomena, according to them; just as all cows are animals.” (Strawson, 2008c) Naturalistic Monism endorses this premise, as it is formulated in the above quote. The existence of consciousness is the most certain thing in the world: we can doubt the existence of the whole external world, but not our own experiences. If we are physicalists and take it that experiences feel like something, then we must admit that there exist at least some physical phenomena in the universe that feel like something. However, this is where our agreement with Strawson and other Russellians largely ends. For Strawson continues: “I am happy to say, along with many other physicalists, that experience is ‘really just neurons firing’ […] But when I say these words I mean something completely different from what many physicalists have apparently meant by them. I certainly don’t mean that all characteristics of what is going on, in the case of experience, can be described by physics and neurophysiology or any non-revolutionary extensions of them. That idea is crazy. It amounts to radical ‘eliminativism’ with respect to experience, and it is not a form of real physicalism at all.” (ibid) Strawson’s idea is probably the following: if it was possible to characterize the nature of experiences completely in terms of science, no room would be left for phenomenal qualities or what-it-is-likeness, because science appears to say nothing about such properties. Strawson appears to claim that if experiences could be perfectly characterized scientifically, we would be philosophical zombies (Chalmers, 1996). We strongly disagree with this claim: as we will argue, it conflates experiences with their scientific models. Strawson’s view is based on a metaphysical distinction between extrinsic (roughly: relational, structural, and dispositional) and intrinsic (non-relational, non-structural, and non-dispositional) properties (for a useful discussion of the distinction, see Seager, 2006). On this view, science is only concerned with extrinsic properties—how objects behave, which causal dispositions they have, which relations they stand in, what is their structure, and so on. For instance, in Newtonian physics, force is defined in relation to mass and acceleration. For Russellians, this leaves open the question: what is the non-relational, non- dispositional, and non-structural nature of a phenomenon? Russell (1927) himself argued that this intrinsic nature is ontologically “neutral” (that is, neither mental nor non- mental; see Stubenberg, 2016), but Strawson takes it that the intrinsic nature of all physical phenomena is (proto)mental (see also Goff, 2017). To use the current terminology, Strawson would argue that what an experience feels like is part of the intrinsic (science- transcendent) nature of the experience, but CMC characterizes the extrinsic (scientifically observable) nature of the same phenomenon. The intrinsic-extrinsic distinction could be criticized in many ways (see e.g. Hiddleston, 2019; Howell, 2015; Kind, 2015), but here we simply note that the postulation of intrinsic properties is metaphysically promiscuous and unscientific. If we suppose that there is an unobservable and causally impotent intrinsic property corresponding to each observable, causally efficacious property, we double the number of physical properties in the universe solely based on armchair reflection. Moreover, as argued by Ellis (2001), among others, there may be no need to postulate any categorical properties to “ground” causal dispositions or relational properties, because all natural phenomena can be considered as essentially causal-dispositional, or processes. For instance, scientifically it makes no sense to speak of the nature of quarks independently of how they interact with other subatomic particles. According to the standard version of quantum chromodynamics, it is impossible for quarks to exist in isolation, so it is nomologically impossible for them to possess a non-relational nature. Because RM implies that experiences are beyond the scope of science, it renders the scientific study of consciousness impossible in principle: science can only study the extrinsic correlates of consciousness, not consciousness per se. This resembles the dualistic notion that “conscious experience involves properties of an individual that are not entailed by the physical properties of that individual, although they may depend lawfully on those properties” (p. 110; Chalmers, 1996). This type of reasoning in part underlies the use of the term “correlate” to describe the neural mechanisms that underlie consciousness. If we could account for the epistemic gap without assuming an ontological gap between different types of properties, we should do so.1 This is the aim of Naturalistic Monism. It holds that experiences are just like any other physical phenomena in that they can be observed and scientifically modelled. The epistemic gap between experiences and their scientific models does not reflect the existence of any separate types of properties, but instead simply the distinctness between a phenomenon and its scientific model.","Naturalistic Monism shares with Russellian Monism (RM) the Kantian assumption that science has in some sense limited access to the nature of the phenomena it studies. However, as explained above, whereas RM holds that there exist two distinct types of properties (extrinsic and intrinsic), Naturalistic Monism takes it that there exists only one class of physical phenomena. The limits of science are not due to the existence of ontologically distinct types of properties, but instead to the fact that we can only know the world through observations and scientific models. Schneider (2017), a proponent of RM, motivates the existence of science-transcendent intrinsic properties by quoting the theoretical physicist Stephen Hawking, who famously asked: “What is it that breathes fire into the equations and makes a universe for them to describe? The usual approach of science of constructing a mathematical model cannot answer the questions of why there should be a universe for the model to describe.” (Hawking, 1988, 174). Strawson (2008b), in turn, quotes the astrophysicist Arthur S. Eddington, who noted that “[…] science has nothing as to the intrinsic nature of the atom. The physical atom is, like everything else in physics, a schedule of pointer readings. The schedule is, we agree, attached to some unknown background.” (Eddington, 1929, 259) But do Hawking and Eddington intend that there would exist some properties of physical phenomena completely beyond the scope of science, as Russellians would have it? Or did they simply mean that we cannot know the nature of physical phenomena without relying on observation (“pointer readings”) or without utilizing scientific models and equations? The former claim is metaphysical, the latter is purely epistemological. At least Hawking intended to make a purely epistemological claim, which he and Mlodinow formulated more explicitly when they introduced the notion of Model- Dependent Realism (MDR) (Hawking & Mlodinow, 2010). According to MDR, we can never know the nature of the world as it is in itself, but instead we can only know it through how it affects our senses. Based on observations, we can formulate models of the world in our minds, but we can never step outside the models and compare them to model-independent reality. We can never be certain if our models are true in a strict sense of the word; we can only evaluate whether they are elegant, can predict observations, or can be used to manipulate phenomena. Hawking and Mlodinow consider scientific models on par with our non- scientific, everyday models of the world: both are in the mind/brain of the organism and based on constant causal interaction between the organism and its environment. Consciousness itself can be considered as an internal model of the world, which affords us to predict and explain what happens in our surroundings, increasing our chances of survival (Friston, 2010; Hobson & Friston, 2014; Revonsuo, 2006). In this sense, there is only a matter of degree between a person forming an inner representation of the world based on their everyday interaction with it, aided by nothing but naked senses and their body, and a researcher building a scientific model based on more sophisticated, instrumentally aided, and theoretically-driven observations and manipulations of the phenomena. In both cases, the model is a natural phenomenon in the brain of the modeler, and can itself be an object of scientific inquiry (more about this in Section 5.1.). It is a philosophical question whether a model can ever be considered as “true”, but from a purely naturalistic perspective we can judge to what extent an organism’s internal model of the world affords it to interact with its environment efficiently. A mouse has a valid internal model of its surroundings if the model affords the mouse to seek shelter from the approaching cat in the nearest burrow. Likewise, our scientific atomic model can be considered as valid if it affords us to build functioning nuclear reactors. Thus, internal models—irrespectively of whether they are of a scientific or everyday variety—can be minimally considered as valid or even “true” in a pragmatic sense. Whether they are true in some stricter sense is a philosophical problem that is beyond the scope of the present paper. According to Naturalistic Monism, the epistemic gap between a subjective experience and its scientific description is distinctness between two phenomena: the experience in the mind of a participant and the scientific model in the mind of the researcher. Naturalistic Monism implies that an experience is the concrete phenomenon that a scientific model of its constitutive mechanism describes; it is what underlies scientific observations of its constitutive mechanisms. Scientists can model experiences, but the participants’ experiences are always distinct from their scientific models, which are in the minds of researchers. Whereas the participant has “immediate” or first-person access to their experiences, the researcher can know them “mediately” through observations and models. The researcher does not have direct access to the participant’s experiences; they can only access the participant’s consciousness indirectly by causally interacting with it (and directly experiencing a model of it in their own consciousness). In this respect, consciousness is an exceptional object of research: it is the only phenomenon that we can at least partly know or experience directly, not only based on observations and models. Thus, in the case of experiences, we can directly compare the model and the concrete phenomenon that is modeled—something we can never do in the case of, say, atoms. Science can model all the aspects of experiences—including what they feel like—but only model. For instance, suppose that subjective pain could be modelled as T-type interaction between neural modules M1 and M2 (or T-interaction, for short). “T-interaction” is just pain’s scientific model, formed in the mind of a scientist based on observations, such as which neural events are correlated with reports of pain, avoidance behavior, certain facial expressions, and so on. Even a scientific realist, who assumes that scientific theories can truthfully capture the nature of natural phenomena, cannot hold that pain is literally nothing but “T-interaction”, because “T-interaction” is a theoretical model of pain, and distinct from pain itself. Instead, the realist should be interpreted as claiming that pain can be truthfully modelled as “T-interaction”. This is compatible with the fact that when the concrete process scientifically modelled as “T-interaction” happens in a subject, the occurrence of the process feels like something for her. Crucially, the subjective feel of pain is nothing distinct from the process described by the model; it is the happening of the concrete process itself. Accordingly, a scientific realist can accept that pain has a qualitative feel; what their realism implies is only that the feel of pain can be truthfully modelled as “T-interaction”. In this sense, we can eliminate the theoretical notion of “qualia”, conceived of as a non-dispositional, non-relational, and non- structural property, without claiming that experiences would not feel like anything (in the everyday sense of “feeling like something”) (cf. Dennett, 2018). A neuroscientific CMCE-model of an experience E is never identical with E; it only refers to E. Whereas science can only model experiences through observation, we know our own experiences immediately. My pain is not something outside of me that I observe, it is part of my subjective realm. It constitutes the physical process that is my consciousness, which could be observed and modelled by a scientist as CMCpain. To paraphrase Edelman and Tononi (2000), “Unlike any other entity, […] with consciousness we are what we describe scientifically” (p. 14, italics in the original). The reason why we easily think that there is an epistemic gap only for experiences is because our experiences constitute us, and thereby we have non-empirical access to their nature—access that is not mediated through observations and models. Naturalistic Monism implies that there is an epistemic gap between any concrete phenomenon and its scientific model, due to the distinctness between the model and the phenomenon. As discussed later in the article (Section 5.2.), how wide we take the gap to be depends on which philosophy of science we endorse. Naturalistic Monism implies that what an experience feels like is part of its nature in itself, independently of observations, scientific models, or other theoretical characterizations. In fact, describing an experience as “feeling like something” is already a conceptualization, and it could be debated whether it adequately or truthfully captures the nature of the experience. Descriptions of consciousness typically depend on the person’s favorite philosophical theory, religion, culture, or level of education. Consciousness could be said to be “neural activation”, “qualitative”, “phenomenal”, “feeling like something”, “non-relational”, “non-structural”, consisting of “qualias”, “soul”, “Brahman”, and so on. All such definitions of experiences could be questioned, but that does not amount to questioning the nature of experiences themselves. If you take René who believes in an immaterial soul, David who believes in qualia, and Daniel who denies the existence of qualia, and stick them with a needle, the same kind of natural phenomenon is instantiated in all of them. The concrete phenomenon that occurs in them is of the same type, even if they define it in very different ways—nature does not care about our definitions.2 The qualitative feel of experiences is often considered as problematic or even impossible to model scientifically. Naturalistic Monism implies that it is possible to model what experiences feel like—but only model. The discrepancy between what, say, pain feels like for a subject S and how a scientific model characterizes it is simply distinctness between pain and its scientific model. Pain is a natural process happening in S, whereas its scientific model is a highly complex conceptual representation in the minds of scientists. What pain feels like for S is not something over and above the process described by the model; instead, it is the happening of the described process in S. When pain happens in me, I can minimally say it feels like “this”, referring to the happening of my pain experience (Ismael, 1999; Papineau, 2002). Consciousness science aims to model such processes, and can in this sense model what experiences feel like. In the form of an argument: What pain feels like for S = the happening of pain in S. The happening of pain in S is a concrete physical process. Science can model all concrete physical processes. Thus, science can model what pain feels like. It is common to suppose that knowing a scientific model of pain does not give knowledge of what pain feels like. In contrast, according to Naturalistic Monism, knowing the pain-model does give knowledge of what pain feels like, but only knowledge in terms of the scientific model, which is distinct from concrete pain. What concrete pain feels like is simply the happening of pain in us, a process that science aims to model with the pain-model. We can say that pain is nothing but “T-interaction” (in a de re sense, meaning that “T-interaction” refers to pain), but this does not amount to reducing concrete pain into the theoretical T-interaction-model. Instead, pain, including what we call its “what-it-is-likeness”, is the concrete phenomenon that science models as T-interaction. The take-home message is that we need not postulate the existence of any properties (“qualia”) for the what-it-is-likeness of pain over and above the properties described by the scientific pain-model. Once we have a scientific model of pain that is perfectly isomorphic with pain as experienced, we have a scientific model of pain’s phenomenology. For instance, the experienced intensity or sharpness of pain could be modelled as parameters in the pain-model that correspond to such qualities—any change in the experienced sharpness of pain corresponds to a change in the corresponding parameter (see Section 4.1.). However, this does not mean that pain itself would be just the theoretical parameters; it is what the parameters model.3","We have argued, first, that consciousness should be viewed as a concrete physical process in the world in the same sense as lightning, photosynthesis, and life—experiences are of the same ontological type as all physical phenomena, there is only difference in complexity. Second, we have stressed that scientific models are themselves mental phenomena and always distinct from the phenomena they model. Because experiences are concrete phenomena, they can causally interact with other physical phenomena, including observers and their measuring devices. Science can model experiences as CMCs. For a neuroscientist, the experience of a subject is the “external” phenomenon that affects the scientist’s measuring devices, and which they aim to describe with the CMC-model. From the subject’s perspective, their experience is an “internal” phenomenon, part of their subjective realm. We will next illustrate Naturalistic Monism with the help of a neuroscientific example. Consider a visual experience of a cat (Fig. 1). On our approach, a subjective cat experience is the physical process in the world that underlies neuroscientific observations that we interpret in terms of the CMCcat-model (a model of the constitutive mechanisms of a cat experience). Thus, the terms “cat experience” and “CMCcat” refer to the same concrete phenomenon.4 Whereas the term “experience of cat” is typically used to refer to an experience when it is part of a subject’s consciousness, the neuroscientific term “CMCcat” is used in scientific contexts, where it refers to the experience as the external phenomenon that causes scientific observations. Experiences are complex physical processes that are composed of lower-level phenomena, and the composition can be described scientifically. Science can describe how experiences are composed of lower level processes in the same way as it explains how hydrogen and oxygen form water. This framework is summarized in Fig. 1. The cat-experience is depicted in Fig. 1 as a complex physical process that is composed of lower-level constituent processes. This can be empirically modelled by defining how the CMCcat-process is related to its constituent processes. It is often argued that consciousness cannot be fully explained by analyzing the interactions between its lower-level constituents (this is the so-called “combination problem” that RM faces; see Chalmers, 2017). This argument conflates phenomena with their scientific models: one can never “reduce” any phenomenon to its description or descriptions of its lower level processes (reduce Xn to Yn or Yn-1 in Fig. 1); one can only reduce descriptions to other descriptions (reduce Yn to Yn-1). To give an example, one can explain how the functioning of an amoeba is based on the interactions between the molecules that constitute it, and in this sense reduce the scientific model of the living amoeba to models of lower-level phenomena, such as interactions between proteins, enzymes, and so on. This does not mean that the actual, concrete amoeba is “reduced” to anything else; the concrete amoeba is the complex system composed of the interaction between molecules, enzymes, and so on. In the same way, one can model how an experience is based on lower level processes (how Yn is based on Yn-1), but this does not amount to “reducing” the concrete experience (Xn) into a description. When we explicate how a cat-experience is composed of lower level processes (the lower axis in Fig. 1), we do it in terms of describing how the CMCcat is composed of lower-level processes, such as the activation of individual neurons (the upper axis in Fig. 1). Crucially, when we scientifically reduce an experience into lower-level processes, we do so only descriptively. “Reducing” the concrete experience—literally tearing it to parts—would amount to destroying it. The crucial point here is that if we accept that scientific models of experiences are literally models of experiences, then it is possible to scientifically model how subjective experiences are based on lower-level processes. The black boxes in Fig. 1 might appear to motivate some kind of mysticism about the nature of (non-mental) physical phenomena, similar to Strawson’s “intrinsic nature” that cannot be empirically captured. However, this mysticism is inherent to all science: we can never know the nature of the world as it is in itself, we can only model it based on how it affects our senses. In Fig. 1, the lower level boxes are mostly black to emphasize that the phenomena exist independently and outside of how we scientifically model them, but this does not imply that they are completely epistemologically opaque to us: we can know their nature, but only through observations and models. Moreover, it is important to notice that whether a box is black or not depends on perspective: my consciousness is not a black box for me, but it is for you, because your knowledge of my consciousness is based solely on observations and models (your model of my mind). Towards an empirical model of consciousness ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Because Naturalistic Monism implies that an experience is a concrete physical phenomenon, and that all the properties of experiences are physical properties, it must be possible to observe and scientifically model all the characteristics of an experience. Hence, based on the current framework, empirical science can study consciousness, not just its correlates. This is an important upshot, as it collapses the problem of consciousness into a standard problem of science. Specifically, we suggest that scientific models of consciousness should describe the following5 (Fig. 2): Constituents. Because consciousness is a complex physical process, science should describe its constituents. This explanation describes what consciousness is “made of”. Isomorphy. A complete model of consciousness should accurately describe all the details about the contents of consciousness, including its phenomenology. Etiology. Understanding the etiological basis of consciousness, from immediately preceding causes to its phylogenetic basis, is crucial for understanding phenomenology. Causal power. Being a concrete physical phenomenon, consciousness can causally interact with other physical processes inside the brain. A key part of explaining consciousness is explaining what it can do. Lower-level constituents According to Naturalistic Monism, a successful empirical theory of consciousness must specify the constituent processes that are necessary for consciousness, but which are not themselves conscious until they interact in some specific way to form the CMC. The constituents are for CMC what hydrogen and oxygen are for water: hydrogen and oxygen do not cause water, they constitute it. The constituents of consciousness are often called the “mechanism of consciousness”: Which neural processes determine whether an individual is conscious or unconscious? Which neural processes determine whether a visual stimulus will be consciously perceived? Examples of theoretical attempts to answer such question include the global workspace model (GWT; Baars, 1988; Dehaene & Changeux, 2011), integrated information theory of consciousness (IIT; Oizumi, Albantakis, & Tononi, 2014), and higher-order thought theory of consciousness (Lau & Rosenthal, 2011). These theories aim to describe how consciousness “emerges” from neural activity. According to Naturalistic Monism, these theories aim to describe the CMC. Crucially, because the CMC is a model of consciousness, explaining the constituents of CMC amounts to explaining the constituents of consciousness itself—not just something that correlates with consciousness. Consider the GWT, for example. According to the theory, consciousness is a process where a large number of individual modules share their information with each other to enable flexible behavior and maintenance of information (Baars, 1988). According to the neural model of the theory, this corresponds to synchronous interactions between specific neural populations in frontal, parietal and sensory cortical areas (Dehaene & Changeux, 2011). These individual processes are not conscious alone, but become conscious when they interact in the way that the theory describes. These constituent processes are further based on lower-level processes: oscillatory activity requires a certain balance of excitatory and inhibitory neurons, neural activity requires a resting membrane potential, neurons are made up of specific molecules, and so on. The constituents can, in principle, be traced to fundamental physical entities; in this sense, the CMC (the experience) is continuous with other physical phenomena. GWT does not make claims about how the subjective aspects of consciousness (phenomenology) are enabled by activation of the global workspace, but according to Naturalistic Monism, the concrete phenomenon whose functional aspects GWT models is consciousness and has a certain phenomenology for the subject in whom it happens. Accordingly, the functional properties modelled by GWT are not independent of phenomenal consciousness; they are the functional aspects of phenomenal processes (causal power of consciousness is discussed further in Section 4.1.4.). Isomorphism Because the phenomenology of an experience is a concrete physical process, any change in phenomenology can in principle be observed and modelled by a specific change in the model. Thus, a complete CMCE- model of an experience E must be perfectly isomorphic with the experience: there must be a one-to-one correspondence between the empirically observed physical phenomenon (CMCE) and the subjective experience E (Fingelkurts & Fingelkurts, 2004; Revonsuo, 2000). This means that an empirical theory of consciousness, if sufficiently detailed, can provide a complete scientific description of an experience and its phenomenology. To do this, the theory must be able to define which empirically observed physical processes correspond to different contents of experiences, and how these contents combine to form complex experiences. For instance, from the perspective of conscious vision, the requirement of isomorphy means that the model must predict the phenomenology of vision from the egocentric perspective, not from the retinotopically-centered frame of reference (Land, 2012). According to the integrated information theory of consciousness (Oizumi et al., 2014) the structure of recurrent networks (“qualia space”, to use IIT’s terminology) defines phenomenology. According to Naturalistic Monism, such models are descriptions of the concrete physical phenomenon that is consciousness. The reader may, at this point, insist that isomorphic descriptions of CMCs would not explain the “qualitative feelings” of experiences, or answer why certain neuronal activity feel like something, and hence the description would not be an explanation of phenomenology. Such counter-arguments are ill-posed, because they conflate concrete phenomena with their scientific models. For instance, what blue feels like for S is the happening of CMCblue in S, a concrete process that can be perfectly modelled scientifically—but just modelled. Grasping the CMCblue-model affords knowledge of how an experience of blue can be scientifically described, whereas what blue feels like is just the happening of the concrete phenomenon, which is distinct from the scientific model. No description alone can ever cause the happening of the concrete phenomenon it describes, but this does not imply that it could not perfectly describe the phenomenon and its happening in a subject. Science can model the constituents, etiology, etc. of blue experiences, and there is no further question of why blue (or the corresponding neural process) feels like blue. Feeling like blue is part of the nature of blue experiences, determined by their constituent processes, just like having a certain molar mass is part of the nature of helium. Science can model what blue feels like and how that feel is determined by lower-level processes, just like it can model the molar mass of helium and its constituent processes. Etiology In trying to explain what must happen in the brain so that, for instance, visual information crosses the threshold to consciousness, neuroscience examines the causal chain of events that lead to it (Dehaene & Changeux, 2011; Railo, Koivisto, & Revonsuo, 2011). In other words, they describe the etiology of conscious vision. The aim is to disentangle the CMC from other processes that correlate with conscious perception, but do not constitute it (e.g. Aru, Bachmann, Singer, & Melloni, 2012). For example, neural processes that precede the presentation of a stimulus help to predict whether or not the stimulus is consciously perceived – in other words, prestimulus activity correlates with the subsequent conscious perception (e.g. Iemi, Chaumon, Crouzet, & Busch, 2016). However, these prestimulus correlates of conscious vision are obviously “just correlates”, and the neural processes that constitute the conscious perception of the stimulus happen later. One of the challenges of empirical science is to characterize this etiological chain of events that leads to conscious vision, that is, the formation of the CMC that corresponds to experience of that stimulus. In addition to explaining the immediately preceding causes that lead to consciousness, science can describe longer time scales such as individual development, or evolution (see Revonsuo, 2006). A theory of consciousness should explain how consciousness has evolved from earlier products of evolution. For example, the phenomenology of visual perception could be accounted in empirical terms by characterizing how frequently different stimulus combinations have been encountered during the evolution (Purves, Wojtach, & Lotto, 2011). Similarly, pain can be considered as a complex biological adaptation that has helped organisms to avoid potentially damaging stimuli, and whose bio-chemical basis could possibly be traced back to single celled organisms (Bray, 2009). According to Naturalistic Monism, this does not mean that only the functional aspects of pain (or visual perception) would be products of evolution, but also that the different types of subjective unpleasant sensations we call “pain” (or visual phenomenology) are products of evolution. This view is built on the assumption that consciousness is not an epiphenomenon (cf. Jackson, 1982), but instead that it has power to influence other physical processes. Causal power The idea that experiences are something different from other physical phenomena is deep-rooted in philosophy and science. For instance, Edelman (2003), who builds an elegant theory of consciousness on a bio-physical foundation, argues that conscious states are “informational even if not causal” (p. 5523). This way, consciousness easily appears epiphenomenal, which is strongly at odds with our subjective experience as well as the empirical scientific assumption that human consciousness is a product of evolution, and plays a causal role in behavior. Naturalistic Monism ties together the phenomenal and functional aspects of consciousness. Because consciousness is a physical process, it can causally interact with other phenomena in the brain and drive the organism’s behavior. We propose that Naturalistic Monism may help to dismantle some stubborn conflicts about whether consciousness should primarily be described in functional or phenomenal terms (Block, 1995; Carter et al., 2018; Dehaene & Changeux, 2011). While it is an empirical question to find out whether specific cognitive functions (e.g. global access or error detection; Dehaene, Lau, & Kouider, 2017) are necessary for consciousness, our framework implies that experiences never take place in a causal void. An experience E (CMCE) always causally interacts with other processes in the brain, both conscious and non-conscious: it can cause, for instance, thoughts, emotions, bodily states, or actions, and interact with the outside world. One of the goals of empirical science of consciousness is to explain what gives conscious processes their power to potentially elicit fast and accurate actions, and flexible voluntary behavior6. Theories such as the Global Workspace theory are explicitly built around this foundation (Dehaene & Changeux, 2011). According to Naturalistic Monism, such “functional explanations” can provide key insights into the phenomenal qualities of experiences as well (cf. the example about pain, in previous subsection).","In this section, our aim is to briefly address the main philosophical implications of Naturalistic Monism. The discussion is not intended as exhaustive, but rather as pointing towards future research. Relativity of “subjective” and “objective” ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Naturalistic Monism implies that terms like “subjective”/”objective”, or “first- person”/”third-person” are relative. When a subject S1 perceives an apple, the apple is objective and third-person observable. The experience of the apple in S1 is subjective or first-person observable for S1, but because it is a natural phenomenon, it is third-person observable for a neuroscientist (S2). The neuroscientist’s observations and theoretical models of the neural mechanisms of S1’s apple-experience, in turn, are phenomena in the scientist’s mind/brain and (at least partly7) subjective or first-person observable for the neuroscientist—the scientist is aware of grasping the model. The scientist’s model is in principle objectively or third-person observable for a “metaneuroscientist” (S3), and so on ad infinitum. In short: any subject’s experience is another subject’s external phenomenon. Husserl’s notions of phenomenological and natural attitude can be used to illustrate this relativity. In everyday life, we typically endorse the natural attitude, focusing on external objects instead of our experiences of them. When I perceive an apple, my consciousness is directed at the apple and I am not necessarily aware of the details of my experience per se, such as what the greenness of the apple feels like. In Hurssel’s terminology this is called the natural attitude. In contrast, when I endorse the phenomenological attitude, I focus on my apple-experience as an experience. A scientist, in turn, is concerned with the external phenomena they study and thus automatically endorse the natural instead of the phenomenological attitude—they rarely focus on the nature of their scientific models as conscious phenomena in their minds. However, even the conceptual aspects of a scientific model—e.g. that the helium atom has two protons—are conscious phenomena in the mind of the scientist when they grasp the model. As such, the scientist can endorse the phenomenological attitude towards them and consider them as experiences. This illustrates that we can never escape our own point of view: even a scientific model is only a representation in the scientist’s mind. However, even though a scientist S1 cannot themselves “step outside” their model, an external observer S2 can: in principle, such a “meta-observer” could investigate the relationship between an object X and the scientist S1’s mental model of it from an external perspective, treating both the object and the model as external phenomena (“noumena” in Husserl’s terms). But the meta- observer S2 would then be confined to their own perspective (and so on for any further meta-observer). Philosophies of science and consciousness ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ If we consider consciousness as a concrete physical phenomenon, then the question about the relationship between an experience and its mature scientific model comes at least partly down to which philosophy of science we endorse (in addition to the details of the model itself). A scientific model is always distinct from the concrete phenomenon it models, and there is always a gap between the two in this sense. How wide we take the gap to be depends on whether we endorse realism or antirealism about science. Roughly, according to realism, mature scientific theories can be true, and truthfully describe the nature of the world (Psillos, 1999). Thus, the epistemic gap between a phenomenon and its mature scientific model is non-existent or narrow.8 According to antirealism, on the other hand, we always have to remain agnostic about the truth of scientific models. For instance, according to constructive empiricism (van Fraassen, 1980), scientific theories can be considered as empirically adequate if they are isomorphic with observations, but we can never say whether they are literally true. We could say that, on this view, there is always an indefinitely wide epistemic gap between a phenomenon and its model, or that it makes no sense to speak of such a gap. Scientific realism implies that if we said that our experiences involve purely qualitative aspects that cannot be modelled structural- relationally, we would be mistaken. According to scientific realism, all natural phenomena can be truthfully modelled structural-relationally. On this view, what we call a “qualitative feeling” can be truthfully modelled as a structural-relational process. Importantly, this approach does not deny that experiences feel like something; it only holds that the feel can be modeled purely structural-relationally. If we characterized the feel as something purely qualitative, we would be mistaken—not about the feel itself, but instead about how to describe it. Importantly, terms like “qualitative” or “feeling” are only concepts which refer to certain concrete processes that happen in us. The processes are the way they are, no matter how we refer to them. Scientific realists would hold that experiences can be truthfully modelled completely in structural-relational terms, but this does not, and cannot, amount to eliminating any aspects of the experiences themselves. It may be difficult to see how, say, the redness of red could be a structural-relational process, given that introspection about redness does not reveal that it has such nature. However, from the fact that a model M applies to an experience E, it does not follow that a subject experiencing E would introspectively see that M applies to it (cf. Dennet, 1991). If we endorse antirealism, such as van Fraassen’s empiricism, we can only say whether a neuroscientific theory of consciousness is compatible with observations, not whether it “captures the nature” of what it models. It could be said that, for the antirealist, there is always an indefinitely wide epistemic gap between any concrete phenomenon and its scientific model, and that experiences are not different in this respect. It is noteworthy that Naturalistic Monism is not committed to either realism or antirealism; this is an independent question. Naturalistic Monism minimally implies that there is always distinctness between a phenomenon and its model, but whether we take there to be an epistemic gap may depend on which philosophy of science we endorse. Detailed discussion of how consciousness is related to different philosophies of science is beyond the scope of this paper. The main point is that if we consider experiences as being of the same (ontological) type as any other physical phenomenon, then the question about the relationship between an experience and its scientific model boils down to the general question about the relationship between any phenomenon and its model. Panpsychism ~~~~~~~~~~~ According to Naturalistic Monism, experiences are concrete physical phenomena that are based on lower-level physical processes, and thereby ontologically continuous with them. Does it follow that all physical phenomena are of an “experiential nature”, because they are of the same ontological type as experiences (cf. the black boxes in Fig. 1)? In our view, this is not a necessary conclusion. An experience E can be scientifically modelled as CMCE, a system-level phenomenon in the brain that is based on the activity of individual neurons. In this sense, consciousness is continuous with other physical phenomena, all the way down to subatomic particles. However, insofar as individual neurons (not to speak of quarks) alone do not realize the CMCE-process, they are not conscious, at least in the same sense as we are. On this line of argumentation, it is possible for consciousness to emerge from something that is not conscious, and it is the task of empirical theories of consciousness to describe exactly how this happens (but see Strawson, 2008b). On the other hand, we have no non-empirical or non-model-mediated access to the nature of other physical phenomena besides our experiences. We can extrapolate from our own case and infer that at least cats and cows are likely to have similar experiences as us, given that they show similar behaviors and have similar neuroanatomy. It is more difficult to say anything about the possible inner lives of insects, amoebas, or quarks. This difficulty, however, stems from the limits of science, not from consciousness being a special type of phenomenon. Zombies, Mary, and the “hard problem” ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Philosophical zombies (Chalmers, 1996) are creatures that appear just like normal human beings in every functional and observable respect, but for whom “the lights are out”: it does not feel like anything to be a zombie. When we imagine zombie-pain, we imagine something that would cause pain-observations and fit our scientific model of pain, but which would not be real pain. If such fake-pains are conceivable, then arguably we could conceive of any fake natural phenomena. For instance, we can imagine that we live in a massive computer simulation, where our observations of electrons are not caused by real electrons, but instead by computer algorithms. If such imaginings are possible at all, what they teach us is that our scientific models and observations are always distinct from the concrete phenomena that underlie them. It does not tell us anything about the nature of experiences or electrons, except for that they are concrete phenomena and independent of their scientific models and observations. The same reasoning can be applied to the Mary case, a neuroscientist who has never experienced colors (Jackson, 1986). Mary knows everything about the visual color system in the brain, but she does not know what it is like to see red. The conclusion that the thought experiment is supposed to demonstrate is that there are some facts about experiences that cannot be described by scientific models. But the conclusion is false: Mary does know what it is like to see red, but only in terms of a scientific model. What it is like for S to see red is the happening of a red experience in S. Because science can model such processes, it can thereby model what it is like to see red. When Mary sees red for the first time, the concrete phenomenon whose model she already knew happens in her. What the Mary case teaches us is that a concrete phenomenon can be quite different from what a scientific model says it is. But this does not imply that there would be some properties (“qualia”) that cannot be scientifically modelled, it only shows that a concrete phenomenon is distinct from its model.","We have argued that once we clearly distinguish between concrete phenomena and their scientific models, the epistemic gap – why our experiences appear so different from neural activity – dissolves into distinctness between the model and the phenomenon. From the perspective of Naturalistic Monism, science can observe and model consciousness, not just its correlates. We hope that our proposal helps to establish a solid ground for the empirical science of consciousness, as it collapses the problem of consciousness into a standard problem of science. Scientific explanation never affords us knowledge of the nature of phenomena in themselves. However, if we accept that our subjective experiences are concrete physical phenomena, then science can provide a complete and exhaustive model of experiences, including what they “feel like”."],["Most humans share to some degree. Yet, from middle childhood, sharing behavior varies substantially across societies. Here, for the first time, we explored the effect of self-construal manipulation on sharing decisions in 7- and 8-year-old children from two distinct societies: urban India and urban United Kingdom. Children participated in one of three conditions that focused attention on independence, interdependence, or a control. Sharing was then assessed across three resource allocation games. A focus on independence resulted in reduced generosity in both societies. However, an intriguing societal difference emerged following a focus on interdependence, where only Indian children from traditional extended families displayed greater generosity in one of the resource allocation games. Thus, a focus on independence can move children from diverse societies toward selfishness with relative ease, but a focus on interdependence is very limited in its effectiveness to promote generosity. --------------------------------------------------------------------------------","Sharing of resources is vital and common in all human societies (Henrich et al., 2010). Yet, young children across societies tend to maximize their own outcomes and only from middle childhood begin to conform to the sharing norms of their respective societies (e.g., Cowell et al., 2017; House et al., 2013; Rochat et al., 2009). Children’s sharing behavior is influenced by social information such as explicit normative instructions (House, 2018; McAuliffe, Raihani, & Dunham, 2017) and demonstrations by adult models (Blake, Corbit, Callaghan, & Warneken, 2016; Over & Carpenter, 2013). Whereas children from different societies behave more selfishly if selfish behavior is modeled by an adult, Western children appear to be less flexible in adopting more generous behavior (Blake et al., 2016; Weltzien, Marsh, & Hood, 2018). For example, children in both urban United States and rural India reduced their giving when stingy behavior was modeled, but only Indian children increased their giving when generous behavior was demonstrated (Blake et al., 2016). Children’s abilities to copy others and follow instructions have been widely documented (Blake et al., 2016; House, 2018; McAuliffe et al., 2017; Over & Carpenter, 2013). What remains puzzling is why children from Western societies respond less to social influences aimed to increase generosity compared with children from other populations. An emphasis on child autonomy and independence in Western middle-class families (Blake et al., 2016; Grossmann & Na, 2014; Keller, Borke, Chaudhary, Lamm, & Kleis, 2010) may make children more reluctant to go against their self-interest. In contrast, children from societies that emphasize child obedience and interdependence (Clegg & Legare, 2016; Keller et al., 2010) may respond more readily to social influences even when they contravene children’s self-interest. Recent studies applying priming of self-construals in children offer tentative support for these claims. Experimentally manipulated self-focus reduced British children’s willingness to relinquish personal possessions (Hood, Weltzien, Marsh, & Kanngiesser, 2016) and their sharing and helping behavior (Weltzien et al., 2018), whereas focusing on others or close friends was ineffectual. These findings may reflect a relative strength of independent versus interdependent self-construals in British children. Yet, cross-cultural evidence from populations with predominantly interdependent self-construals is lacking. The United Kingdom and India are two countries that have traditionally differed along the dimension of individualism versus collectivism (Santos, Varnum, & Grossmann, 2017; Verma, 1999). The consensus across the cross-cultural literature is that Western middle-class families socialize their children toward autonomy, individuality, and self-sufficiency (Greenfield, Keller, Fuligni, & Maynard, 2003; Kärtner, Crafa, Chaudhary, & Keller, 2016; Keller et al., 2010). Indian middle-class families, on the other hand, emphasize interpersonal responsibilities, interdependence, and shared experiences to a larger degree (Kärtner et al., 2016; Keller et al., 2010; Miller, Bersoff, & Harwood, 1990). These differences also manifest in the extent to which the self is defined in relation to others. Arguably, in individualistic societies, people have more independent self-construals, defined largely in terms of internal attributes such as attitudes and abilities (Grossmann & Na, 2014; Markus & Kitayama, 1991). Conversely, in collectivistic societies, people have more interdependent self-construals, mainly defined in terms of group membership and close relationships (Markus & Kitayama, 1991). Although independent and interdependent self-construals coexist in all individuals, their prominence and accessibility vary due to different cultural conventions and socialization practices (Singelis, 1994). Across societies, middle childhood has been identified as an important phase in prosocial development where children begin to conform to societal norms (Cowell et al., 2017; House et al., 2013; Rochat et al., 2009). Therefore, the current study investigated the influence of independence and interdependence priming on the sharing behavior of 7- and 8-year-old children from urban locations in the United Kingdom and India. In the current study, children took part in one of three self-construal manipulations in the form of semistructured interviews: independence focus, interdependence focus, or a control condition. Sharing was subsequently assessed in a task that is commonly used to explore sharing behavior in the current age group (e.g., Fehr, Bernhard, & Rockenbach, 2008; House et al., 2013; Sheskin, Bloom, & Wynn, 2014). Each child completed four repetitions of three different sharing games. In each game, the child needed to choose between two mutually exclusive options that distributed stickers between the child and an anonymous recipient. We used a forced- choice sharing procedure with preset ratios to reduce the problem of individual biases in decision making (Green & Swets, 1966). Moreover, retaining the anonymity of the recipient avoided possible confounds due to reputation management, reciprocity, and/or differing relationships between the participant and the recipient. In the “other-advantage” game, the choice was between an equal split and an unequal split favoring the recipient: low stake (1:1 vs. 1:2) or high stake (2:2 vs. 2:4). In the “self-advantage” game, the choice was between an equal split and an unequal split favoring the participant: low stake (1:1 vs. 2:0) or high stake (2:2 vs. 4:0). The other-advantage and self-advantage games have been used to measure disadvantageous inequity aversion and advantageous inequity aversion, respectively (Blake et al., 2015; Fehr et al., 2008; Sheskin et al., 2014). In addition, we introduced a new game category, termed “that’s life,” where no fair option was available and self-advantage was pitted directly against other-advantage: low stake (1:0 vs. 1:2) or high stake (2:0 vs. 2:4). Specifically, a child could either maximize or minimize the other child’s payoff without incurring a cost to himself or herself. This dilemma has not previously been studied but presents an ecologically valid predicament given that inequality in real life is often difficult to avoid. We predicted that both British and Indian children would be less generous in the independence condition (as compared with the control condition) in all three games (e.g., Blake et al., 2016; Hood et al., 2016; Weltzien et al., 2018). In other words, we expected an increase in 1:1 (2:2) choices in the other-advantage game, an increase in 2:0 (4:0) choices in the self- advantage game, and an increase in 1:0 (2:0) choices in the that’s life game. Moreover, based on the differences in self-construal orientation that exist between these two societies, we predicted that Indian children, but not British children, would be more generous in the interdependence condition (as compared with the control condition) in all three games (Blake et al., 2016). That is, we predicted an increase in 1:2 (2:4) choices in the other-advantage game, an increase in 1:1 (2:2) choices in the self-advantage game, and an increase in 1:2 (2:4) choices in the that’s life game.","The Indian sample consisted of 90 7- and 8-year-old children enrolled in an English- speaking middle- to upper-class private school in Pune, India (Mage = 97.29 months, SD = 3.86, range = 91–105; 51 boys). The British sample consisted of 90 7- and 8-year-old children enrolled in middle- to upper-class schools in Bristol, United Kingdom (Mage = 92.68 months, SD = 6.46, range = 94–107; 40 boys). The sample size was specified a priori in accordance with previous research that used a similar design and procedure (Weltzien et al., 2018). An additional 12 Indian children and 15 British children were tested but excluded from the analyses because they failed to pass one or more of the control trials. All Indian participants had a high level of English language proficiency. Testing in both societies, therefore, was conducted in English by the first and second authors. All children were tested individually in suitable locations at their respective schools. Informed consent was obtained in written form from the parents of all children who participated in this study. All children provided verbal agreement that they wished to partake in the research. Interview procedure Each child took part in one of three interview conditions: an independence interview, an interdependence interview, or a control interview. Questions asked during independence and interdependence interviews were constructed and adapted for children based on the list of independent and interdependent self-construal primes developed by Kühnen and Hannover (2000). During the independence interview, the experimenter asked a series of questions about the child himself or herself, aimed at focusing attention on the child’s uniqueness and individuality (e.g., “What makes you special?” and “How are you different from other people?”). The experimenter also used second-person singular pronouns (e.g., you, your, yours, [child’s name]) whenever apt in order to further steer the child’s attention toward his or her self. During the interdependence interview, the experimenter asked questions about the child’s relationships with family and friends to focus attention on relatedness and closeness to others (e.g., “Is there anyone in your life that you feel close to?” and “Why is it important to have a family?”). In addition, the experimenter used second-person plural pronouns (e.g., you, your [together]) whenever apt in order to further steer the child’s attention toward relationships with and dependence on others. In the neutral control interview, the experimenter asked questions about animals and was careful to avoid the use of personal pronouns (see Appendix A for full interview scripts). Distribution game The experimental design was adapted from previous studies (Fehr et al., 2008; House et al., 2013; Sheskin et al., 2014). Across 12 trials, children decided between two mutually exclusive options for distributing tokens to the self and to “another child” (recipient). First, two opaque boxes were presented: one for the participant and one for “another child from a different school” (recipient). Next, the experimenter explained that tokens would be used in the task and that at the end of the task the tokens would be exchanged for stickers. The child was then introduced to two identical boards. On each board, there were two circles with arrows. The boards were placed so that one arrow pointed toward the child’s box, illustrating that the tokens in that circle would go to the child, whereas the other arrow pointed to the recipient’s box, illustrating that the tokens in that circle would go to the recipient (see Fig. B1, top and bottom, in Appendix B). Training Each child was given two training trials. The first training trial presented the child with a 1:1 versus 2:2 decision. Thus, one board delivered only one coin to the participant and one coin to the recipient, whereas the second board delivered two coins to each of them. This trial functioned as a control and assessed how often the child would choose the highest payoff when both options provided equal outcomes for both the participant and the recipient. The second training trial presented the child with a 1:0 versus 2:0 decision and assessed how often the child would choose the highest self-benefit. The two training trials were designed to introduce children to two important features of the game, namely that payoffs are influenced by the child’s decision and the recipient does not always need to obtain payoffs. Additional control questions ensured that all children had fully understood the training trials (see Appendix C for full testing script). Test trials The next 12 test trials were split into three games with two identical low-stake trials and two identical high-stake trials for each game: Other-advantage game (1:1 vs. 1:2 or 2:2 vs. 2:4): This game explored children’s propensity to choose a fair distribution over a distribution that provides a benefit to the recipient at no cost to the self. Self-advantage game (1:1 vs. 2:0 or 2:2 vs. 4:0): This game explored children’s propensity to choose a fair distribution at a cost to the self. That’s life game (1:0 vs. 1:2 or 2:0 vs. 2:4): This game explored whether children would opt to maximize or minimize the recipient’s payoff when there was no difference to the child’s own payoff. The 12 test trials were divided into two blocks that contained two repetitions of each game. Within each block, three low- stake trials were followed by three high-stake trials presented in a random order. After the first block, the control trial (1:1 vs. 2:2) was repeated to ensure that the child was still paying attention to the task (see Table D1 in Appendix D). The entire procedure lasted approximately 20 min per child (interview: approximately 4 min; distribution game: approximately 15 min). Data coding and statistical analyses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Time spent on the interviews did not differ significantly across societies or interview conditions (see Appendix E and Table E2 for analysis and summary statistics). Therefore, this factor was not considered further. All sessions were videotaped, and children’s behavior was coded live and from videotape. Sharing data Statistical analyses of sharing behavior were run in R Version 3.0.2 (R Development Core Team, 2013), using the lme4 package (Bates, Maechler, Bolker, & Walker, 2013) and the lsmeans package (Lenth, 2016). Because there were multiple trials per participant, observations could not be considered independent of each other. To account for this, the data were analyzed using generalized linear mixed models (multilevel logistic regressions), permitting the inclusion of random effects to model the nested structure of the data (Baayen, 2008; Bates, Maechler, Bolker, & Walker, 2013). Data from the three games were analyzed separately. In all models, sharing decision was entered as a binary response variable (other- advantage trials: fairness = 1, other-advantage = 0; self-advantage trials: fairness = 1, self-advantage = 0; that’s life trials: self-advantage = 1, other- advantage = 0). All full models included the predictor variables of interest (society and condition), their two-way interaction, and the control variables gender (male or female), stake (high or low), block (1 or 2), and age in months (Z transformed) as fixed effects. Reduced models were identical to the full models but without the interaction term. Null models included only control variables and random effects. Initially, likelihood ratio tests were used to explore to what extent including random effects to model between participant variation would significantly improve each of the model’s fit to the data. We compared three types of models: no random effects, only random intercepts, and random intercepts and random slopes for block and stake. For all three games, the inclusion of random intercepts significantly improved the model fit, whereas the inclusion of random slopes did not. Consequently, all analyses reported contain random intercepts but no random slopes. The following analytical procedure was carried out for each game. First, the fit of the full model was compared with the null model with a likelihood ratio test to determine whether the inclusion of the complete set of predictor variables (main effects and interaction) significantly improved model fit. Second, the fit of the full model was compared with that of the reduced model. Subsequent analyses were conducted on either the full or reduced model, conditional on whether inclusion of the two-way interaction significantly improved model fit (for all model comparisons, see Appendix F). To assess the significance for the remaining predictors in the model with the best fit to the data (either full or reduced model), likelihood ratio tests were used. Interpretation of predictors was done by examining model coefficients and, in the case of categorical predictors with more than two levels, by conducting pairwise comparisons using the lsmeans package. An exploratory analysis of the family structure data was also included for a subset of the data (self-advantage game, Indian sample; see Results for rationale). The procedure for this analysis was identical to the procedure described above. However, instead of including society as a predictor variable, this analysis included family structure (joint or nuclear). Interview data To explore children’s verbal responses to independence and interdependence questions, interviews were transcribed verbatim. From the British population, 4 independence interviews could not be transcribed due to faulty video recordings. From the Indian population, 12 independence interviews and 9 interdependence interviews could not be transcribed due to poor audio quality. All remaining scripts were coded by the first author using two coding categories: self- references and social relationship references. The self-reference category included the words I, me, mine, myself, and my own. In addition, the word my was included when related to personal possessions (i.e., my car, my house) but not when related to personal relationships (i.e., my mum, my friend). The social relationship reference category included the words he, him, himself, she, her, herself, his, hers, we, us, our, ourselves, they, them, their, each other, dad, mum, sister, and brother as well as name(s) of friend(s). See Table G3 in Appendix G for mean self-references and social relationship references by condition and society. For reliability, a second coder, blind to the purpose and predictions of the study, coded 25% of the transcripts (balanced by gender, condition, and society). Agreement between the first and second coders was excellent (κ = .96). The responses from the control interview were not included in this analysis because the questions were specifically designed to steer children away from speaking about themselves or their social relationships. Indeed, the authors intended to exclude data from children who made such references in the control interview, although this was not necessary because no references to self or social relationships were made in this condition. To account for individual variability in the overall number of self-references and social relationship references used by children, a proportional self-reference score was calculated for each child. This was done by calculating the proportion of self-references from the total number of self-references and social relationship references. Self-reference scores were analyzed in SPSS using an analysis of covariance (ANCOVA) with condition (independence or interdependence) and society entered as between- participant factors and age in months entered as a covariate. Other-advantage game (1:1 vs. 1:2 or 2:2 vs. 2:4) Children predominantly selected the fair option (72.9%) over the generous option (27.1%). Model comparisons showed that the reduced model had the best fit to the data (see Table H4 in Appendix H for model outputs). This model revealed that sharing decisions were significantly predicted by society, χ2(1) = 29.64, p < .001 (see Fig. 1). Irrespective of condition, British children were more likely to choose the fair (1:1 or 2:2) option (predicted probability = .90) compared with Indian children (predicted probability = .64), odds ratio = 4.85, 95% confidence interval (CI) = 2.86–9.03. In addition, sharing decisions were significantly predicted by condition, χ2(2) = 6.17, p = .046, although pairwise comparisons revealed no significant differences between the conditions after Tukey corrections for multiple comparisons (ps > .058). Sharing decisions were also significantly predicted by block, χ2(1) = 8.85, p = .003. Specifically, children were more likely to choose the generous (1:2 or 2:4) option in the first block (predicted probability = .26) compared with the last block (predicted probability = .16), odds ratio = 1.78, 95% CI = 1.23–2.59. This pattern is consistent with previous findings showing that repeated trials reduce generosity over time, although the reasons for this remain unknown (Kogut, 2012). Self-advantage game (1:1 vs. 2:0 or 2:2 vs. 4:0) Children predominantly selected the selfish option (75%) over the fair option (25%) in this game. The full model had the best fit to the data (see Table H5 in Appendix H for model outputs). First, the results revealed a significant interaction between society and condition, χ2(2) = 6.90, p = .031 (see Fig. 2). Pairwise comparisons (Tukey corrected) showed that British children were significantly more likely to choose the selfish (2:0 or 4:0) option in the independence condition (predicted probability = .99) than in the control condition (predicted probability = .82), odds ratio = 16.63, Z = 2.93, p = .038. Conversely, Indian children were significantly more likely to choose the fair (1:1 or 2:2) option in the interdependence condition (predicted probability = .39) than in the control condition (predicted probability = .03), odds ratio = 0.05, Z = −3.31, p = .012. In other words, British children were nudged toward selfishness after talking about their independence, whereas Indian children were nudged toward costly sharing after talking about their connectedness with others. Sharing decisions in the self-advantage game were also significantly predicted by age, χ2(1) = 5.88, p = .015, suggesting that older children were more likely to choose the fair (1:1 or 2:2) option compared with younger children, odds ratio = 1.99, 95% CI = 1.13–4.14. Moreover, sharing decisions were predicted by block, χ2(1) = 7.47, p = .006. Children chose the fair (1:1 or 2:2) option more often in the first block (predicted probability = .10) compared with the last block (predicted probability = .05), indicating that children in both societies gradually increased their selfishness over the course of the experiment, odds ratio = 0.50, 95% CI = 0.28–0.82. Finally, sharing decisions were significantly predicted by stake, χ2(1) = 15.56, p < .001. Specifically, children chose the fair (1:1 or 2:2) option more often in low-stake trials (predicted probability = .12) than in high-stake trials (predicted probability = .05), odds ratio = 2.72, 95% CI = 1.67–5.16, suggesting that across societies children found it harder to resist a self-benefit when the payoffs were higher. That’s life game (1:0 vs. 1:2 or 2:0 vs. 2:4) Children predominantly chose the selfish option (61%) over the generous option (39%) in this game. The reduced model had the best model fit (see Table H6 in Appendix H for model outputs). In this game, there was no effect of society, χ2(1) = 2.55, p = .11. However, sharing decisions were significantly predicted by condition, χ2(2) = 23.11, p < .001 (see Fig. 3). Pairwise comparisons (Tukey corrected) showed that children were more likely to choose the selfish (1:0 or 2:0) option in the independence condition (predicted probability = .83) than in the control condition (predicted probability = .59), odds ratio = 0.30, Z = −3.57, p = .001. Thus, activating children’s independent self-construals reduced generosity in both British and Indian children. There was no significant difference in sharing behavior between children in the control condition and children in the interdependence condition (predicted probability = .50). Sharing decisions were again predicted by block, χ2(2) = 21.49, p < .001. An examination of the model coefficients revealed that children were more likely to choose the generous (1:2 or 2:4) option in the first block (predicted probability = .44) compared with the last block (predicted probability = .25), odds ratio = 2.36, 95% CI = 1.62–3.60. Finally, sharing decisions were predicted by stake, χ2(1) = 10.28, p = .001, such that children chose the selfish (1:0 or 2:0) option more often in high-stake trials (predicted probability = .72) compared with low-stake trials (predicted probability = .59), odds ratio = 0.55, 95% CI = 0.39–1.25. Exploratory analyses of self-advantage game, Indian sample (1:1 vs. 2:0 or 2:2 vs. 4:0) Why did we find an effect of interdependence focus only in Indian children in the self-advantage game? One possibility that became apparent from the demographic data was that, unlike the British sample, many of the Indian children lived in joint (extended) families. Traditional Indian families typically harbor three or more generations, including members of the extended family. In such families, “collective responsibility” is highly valued, with the needs of the family superseding the needs of the individual (Chadda & Deb, 2013). Living in a large family unit also fosters the pooling of resources, with family members sharing everything from food to property (Chettiar, 2015). Self-construals begin to form during early childhood (Singelis, 1994). Children from joint families, as opposed to Western-style nuclear families, may have more salient interdependent self- construals and, thus, may be more susceptible to the interdependence manipulation. Of the current Indian participants, 54 children lived in joint families, whereas 36 children lived in nuclear families. This allowed us to carry out an exploratory analysis to investigate whether family structure might influence Indian children’s sharing decisions in the three conditions in this game. The full model had the best fit to the data (see Table H7 in Appendix H for model outputs). There was a significant interaction between family structure and condition, χ2(2) = 9.69, p = .008 (see Fig. 4). Pairwise comparisons (Tukey corrected) revealed that in the interdependence condition Indian children from joint families were significantly more likely to choose the fair (1:1 or 2:2) option in the interdependence condition (predicted probability = .73) than in the control condition (predicted probability = .07), odds ratio = 37.55, Z = 4.07, p < .001. Conversely, Indian children living in nuclear families were significantly less likely to choose the fair (1:1 or 2:2) option (predicted probability = .13) than children living in joint families (predicted probability = .73) in the interdependence condition, odds ratio = 0.06, Z = −3.05, p = .027.","Here, for the first time, we demonstrated that children from two societies with different self-construal orientations varied in their sharing behavior following a focus on independence or interdependence. In both societies, activating children’s independent self-construals reduced their generosity in two games (other-advantage and that’s life), where they could provide a benefit to recipients at no cost. This is consistent with previous findings that children from diverse societies will readily adjust their behavior toward selfishness following stingy modeling by adults (Blake et al., 2016). Moreover, the current results extend previous evidence of reductions in trading, sharing, and helping behavior following self-priming in British children (Hood et al., 2016; Weltzien et al., 2018) to a different population. Interdependence priming had an effect on Indian children in the self-advantage game (1:1 vs. 2:0 or 2:2 vs. 4:0) exclusively. In comparison with the other two games, sharing in the self-advantage game came at a cost to participants. Past work has shown that societal differences in sharing behavior are pronounced in costly sharing contexts (House et al., 2013). Here, we expanded on this finding by demonstrating societal differences in susceptibility to interdependence primes within these contexts. Enhanced reactivity of Indian children to interdependence primes is likely attributable to a relatively greater societal emphasis on interpersonal responsibilities (Kärtner et al., 2016; Keller et al., 2010; Miller et al., 1990). Moreover, our exploratory analysis revealed that the effect of interdependence primes in the Indian sample was qualified by family structure. Specifically, Indian children living in traditional extended families were significantly more likely to choose the fair option over the selfish option than children living in Western-style nuclear families. This suggests that growing up in an extended family may lead children to be more susceptible to interdependence primes. Over the past few decades, the sociocultural milieu of India has been going through a rapid change, with a gradual disintegration of the traditional extended family system and a corresponding increase in nuclear families, particularly in urban areas (Singh, 2014). This transformation has led to changes in family functions and values, with a growing emphasis on privacy and independence (Singh, 2014). As a consequence, parents in nuclear families may be more likely to instill individualistic values in their children. This could explain why no effects of interdependence primes were found for Indian children from nuclear families, who responded similarly to the British children. Our findings have implications for cross-cultural researchers because they suggest that it is critical to consider the appropriate level of analysis. Here, societal differences were apparent, but a closer look at family structure within our Indian sample provided a more nuanced picture. There are many other factors that could covary with family structure in India. These include but are not limited to parental income, level of education, and degree of foreign travel. Future work could be more informative not only by considering societal differences but also by paying attention to family-level factors. This could include a focus on measuring independent and interdependent beliefs at the family level. The results also revealed some societal differences in sharing behavior irrespective of priming condition. Specifically, in the other-advantage game (1:1 vs. 1:2 or 2:2 vs. 2:4), Indian children were more likely than British children to benefit the recipient even though a fair option was available. This finding contradicts previous research showing that children from the United States and India display similar levels of generosity (Blake et al., 2016) and are equally averse to receiving less than others (Blake et al., 2015). Instead, the current results support evidence that children from more collectivistic societies tend to share resources more generously than children from more individualistic societies (e.g., Scharpf, Paulus, & Wörle, 2016; Stewart & McBride-Chang, 2000). Previous studies using the self-advantage game have found an aversion to receiving more than others (advantageous inequity aversion) from 7 or 8 years of age in children from Switzerland, the United States, Canada, and Uganda (Blake et al., 2015; Fehr et al., 2008) but not in children from rural India (Blake et al., 2015). We found, however, that neither British nor Indian children showed such an aversion (control condition of self-advantage game). These divergent findings may be due to methodological differences between studies. It is well documented that the presence of a recipient or sharing with a known other influences sharing behavior (e.g., Leimgruber, Shaw, Santos, & Olson, 2012; Martin & Olson, 2015; Schäfer, Haun, & Tomasello, 2015). Whereas previous studies used either peer-to-peer encounters or a photograph of in-group recipients (e.g., Blake et al., 2015; Fehr et al., 2008), children in the current study shared with an unidentified child “from a different school.” Virtually all cross-cultural sharing studies, including the current study, have used tasks where sharing takes place in view of the experimenter (e.g., Blake et al., 2015; House et al., 2013; Rochat et al., 2009; Schäfer et al., 2015). There is evidence that both adults and children employ reputation-enhancing strategies and behave more generously when others witness their sharing (Alpizar, Carlsson, & Johansson-Stenman, 2008; Haley & Fessler, 2005; Leimgruber et al., 2012). In addition, reputational concerns may differ between societies (Callaghan & Corbit, 2018; Gächter & Herrmann, 2009). In the current study, the experimenter’s presence may have influenced both sharing decisions in general and the effects of the focus manipulations. For example, interdependence primes may have triggered a stronger awareness of being observed in the Indian participants than in the British participants. Future studies could investigate this possibility by using the current paradigm and comparing sharing in a fully anonymous sharing task that also controls for the presence of the experimenter. In conclusion, the current findings reveal that subtle self-construal manipulations have striking effects on children’s sharing in different societies. A focus on independence easily moved children in the United Kingdom and India toward selfishness, but a focus on interdependence was very limited in its effectiveness to promote generosity. Specifically, a focus on interdependence increased generosity only in Indian children living in traditional joint families during a costly sharing game. Our findings indicate the importance of considering other levels of analysis beyond broad societal comparisons."],["The aims of this study were to examine and compare the development of parenting cognitions and principles in mothers following preterm and term deliveries. Parenting cognitions about child development, including thinking that is restricted to single causes and single outcomes (categorical thinking) and thinking that takes into account multiple perspectives (perspectivist thinking), have been shown to relate to child outcomes. Parenting principles about using routines (structure) or infant cues (attunement) to guide daily caregiving have been shown to relate to caregiving practices. We investigated the continuity and stability of parenting cognitions and principles in the days following birth to 5 months postpartum for mothers of infants born term and preterm. All parenting cognitions were stable across time. Categorical thinking increased at a group level across time in mothers of preterm, but not term, infants. Perspectivist thinking increased at a group level for first-time mothers (regardless of birth status) and tended to be lower in mothers of preterm infants. Structure at birth did not predict later structure (and so was unstable) in mothers of preterm, but not term, infants and neither group changed in mean level across time. Attunement was consistent across time in both groups of mothers. These results indicate that prematurity has multiple, diverse effects on parenting beliefs, which may in turn influence maternal behavior and child outcomes. -------------------------------------------------------------------------------- CONSISTENCY OF MATERNAL COGNITIONS AND PRINCIPLES ACROSS THE FIRST FIVE MONTHS FOLLOWING PRETERM AND TERM DELIVERIES -------------------------------------------------------------------------------- A large literature documents the importance of parenting beliefs about infants and caregiving (Bornstein, 2015; Bornstein et al., 2007; Dichtelmiller et al., 1992; Goodnow, 1988; Miller, 1988; Miller-Loncar, Landry, Smith, & Swank, 2000; Moorman & Pomerantz, 2008; Pomerantz & Dong, 2006). However, few studies have examined the parenting beliefs of parents of preterm infants, despite preterm deliveries occurring in around 12–13% of live births in the United States and around 5–9% in Europe and other developed countries (Goldenberg, Culhane, Iams, & Romero, 2008). This gap in the literature is significant because infants born prematurely may be at risk due to early non-optimal caregiving environments (e.g., Clark, Woodward, Horwood, & Moor, 2008; Feldman & Eidelman, 2006; Forcada-Guex, Pierrehumbert, Borghini, Moessinger, & Muller-Nix, 2006). The current study compared two types of maternal beliefs–cognitions about child development and principles of caregiving–in comparable mothers of term and preterm infants and did so across two time points during infancy. Therefore, this study was able to examine both the nature, as well as the effect, of prematurity on the developmental trajectories of parenting cognitions and principles.","One aim of developmental research is to understand how constructs develop across time (Bates & Novosad, 2006; Wohlwill, 1970). In this study, we focused on two approaches to measure the development of parenting cognitions and principles: continuity and stability (Bornstein, 2002). Continuity is defined as consistency in group mean level performance across time. A continuous construct is one in which group means do not differ from one time point to a later time point, whereas changes in mean group performance across time would demonstrate that a construct is discontinuous. Individual differences have a complementary focus on variation around the mean. Stability in individual variation is defined as consistency in the relative rank or standing of individuals within a group across time. A stable construct is one on which some individuals rank at relatively high levels at one point in time and again display at relatively high levels at a later point in time, whereas other individuals display lower levels at both times. An unstable construct is one in which individuals do not maintain their rank order across time. Studying the continuity and stability of variables provides both a descriptive and explanatory account of development. Continuity and stability not only tell us about individual differences but also about the developmental origins, nature, and future of constructs (Bornstein & Putnick, 2012). Stability and continuity are mainstay concepts in developmental science and represent statistically and theoretically independent spheres of development (Bornstein & Bornstein, 2008). Developmental scientists are not only interested in how constructs manifest themselves but also in group and individual development across time and, therefore, in continuity and stability. Parenting cognitions and principles ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We chose to study the continuity and stability of both maternal cognitions and principles because parenting is multidimensional, modular, and specific (Bornstein, 2002, 2006). That is, different parental behaviors and beliefs serve different functions, have different developmental trajectories, and have different effects on children, with different domains not necessarily related. We selected the cognitions and principles described below based on their documented relations with parenting behavior and practices as well as child outcomes that may be particularly important for preterm infants (see Section 1.3). First, we studied the complexity of mothers’ thinking about development. Specifically, we focused on two levels of reasoning – categorical and perspectivist thinking – as well as an additional summary variable that reflects the balance between these two levels – complexity of thought (Sameroff & Feil, 1985). Perspectivist thinking reflects flexible reasoning that involves multiple perspectives and takes into account reciprocal influences and transactional perspectives on development, and is therefore more complex. Categorical thinking reflects reasoning that attributes behavior to a single cause and views the child as an extension of parents without individual needs and is therefore less complex. These two levels of complexity of thought describe the broader context of how mothers conceive of children and the parenting role (Miller-Loncar et al., 2000). A parent's ability to think complexly has been shown to relate to both parental behavior and child outcomes. Parents who think more complexly about child development show more sensitive and responsive parenting behaviors (Landry, Garner, Swank, & Baldwin, 1996; Miller-Loncar et al., 2000; Pratt, Hunsberger, Pancer, Roth, & Santolipo, 1993). The preterm and term infants of these parents, in turn, show higher levels of social responsiveness during childhood (Miller-Loncar et al., 2000). However, parents who rely on lower levels of thinking, who have fewer conceptual resources and perspectives to draw on, tend to show a more rigid and authoritarian behavioral style (Deković & Gerris, 1992). Second, we measured parenting principles of caregiving, which are specific personal codes that guide caregiving during infancy (Winstanley & Gattis, 2013) rather than the general level parents can think about development (as described above). Caregiving principles reflect how parents make decisions about infant care. Specifically, we focused on two principles: structure and attunement. Structure reflects mothers’ support of schedules and routines to guide their infants’ day-to-day lives. Attunement reflects mothers’ attention to and reliance on their infants’ cues to guide daily caregiving. These two caregiving principles are independent, and therefore some parents support attunement and oppose structure (or vice versa), whereas others support or oppose both principles. Structure and attunement are related to parenting practices. For example, attunement is positively related to bed- sharing, breastfeeding, and holding in parents of infants under 18 months (Winstanley & Gattis, 2013). Prematurity ~~~~~~~~~~~ There are several reasons to hypothesize that complexity of thought and caregiving principles could be important to understanding the social environment of infants following prematurity. First, the behaviors that have been documented in interactions between mothers and their preterm infants are the same as those that are related to lower perspectivist and higher categorical thinking. That is, mothers of preterm infants have been described as more intrusive and controlling and less sensitive and responsive (Feldman & Eidelman, 2006; Feldman & Eidelman, 2007; Forcada-Guex et al., 2006). The increased intrusiveness and reduced sensitivity seen in mother-preterm infant interactions could therefore reflect differences in maternal cognitions about development (such as, higher levels of categorical thinking). One study did show that mothers of 4-year-olds born preterm tended to score higher on categorical thinking (Pearl & Donahue, 1995). However, we do not yet know how early these differences appear and how categorical and perspectivist thinking develop. An additional reason to study complexity of thought and caregiving principles following prematurity is that parenting beliefs could be particularly meaningful for preterm infants’ development. For example, using preterm infants’ cues and states (sleepiness, arousal, hunger) to guide caregiving is crucial to ensure that such care is developmentally appropriate, as evidenced by the focus of many NICU-based interventions on new parents learning to use infants’ cues about hunger, distress, and sleepiness (Browne & Talmi, 2005; Graven & Browne, 2008; Kaaresen, Ronning, Ulvund, & Dahl, 2006; Landry, Smith, Swank, & Guttentag, 2008). NICU-based interventions have been developed on findings that parents who attend to the behavioral cues of their preterm infants to provide supportive early interactions have infants, and later children, with more positive outcomes (Bozzette, 2007; Landry, Smith, Miller-Loncar, & Swank, 1997). Support of the use of infants’ cues in this way is a central focus in the caregiving principle of attunement. In addition, the parenting practices of breastfeeding and holding (in particular, skin-to-skin touch) are advocated when caring for preterm infants (e.g., Flacking, Ewald, & Wallin, 2011; Tessier et al., 1998), and breastfeeding and holding practices are related to stronger support of attunement (Winstanley & Gattis, 2013). Therefore, parents’ support of attunement (with or without structure) could be important for positive infant outcomes following prematurity. Parenting cognitions and principles have generally been found to be stable in samples of parents of term infants of middle-SES European American background (Cote & Bornstein, 2003; Holden & Miller, 1999; Rubin & Mills, 1992). However, less is known about the stability and continuity of maternal complexity of thought and caregiving principles following preterm deliveries. Measuring complexity of thought and caregiving principles at one time point provides a static picture of how mothers approach and think about caregiving and their child. Looking at just one time point is inadequate because complexity of thought and caregiving principles are likely to change with time as early caregiving moves from the hospital to the home, and as children develop. Very early caregiving of preterm infants often occurs in the hospital, and during this time some parents report feelings that the medical staff are more capable of caring for their preterm infant than they are (Cleveland, 2008; Goldberg & DiVitto, 1983; Howson, Kinney, & Lawn, 2012). In addition, because preterm birth is often unexpected, parents are suddenly forced, ill-prepared, into parenthood, and so they may not have had ample opportunity to attend antenatal classes, read books about parenting and child development, or develop principles about how to care for their infants (Goldberg & DiVitto, 2002). It is therefore important to examine how complexity of thought and caregiving principles change or remain consistent from the days following a premature delivery into later infancy once caregiving routines have been established. As such, prematurity offers the interesting opportunity to understand how parenting complexity of thought and caregiving principles develop under different circumstances. Methodological issues ~~~~~~~~~~~~~~~~~~~~~ This study aimed to chart maternal beliefs following preterm birth, and so child chronological age (calculated from date of birth) was used to schedule visits. This decision was based on our plan to equate amounts of extrauterine experience across dyads at both time points. By contrast, using corrected age (calculated from the due date) to ensure equivalent biological maturity, mothers of preterm and term infants would necessarily differ in the quantity of postnatal and dyadic experience. Corrected age is problematic when studying social development and, in particular, the effects of preterm birth on early parent-infant interactions (Brachfeld, Goldberg, & Sloman, 1980; Wilcox, Weinberg, & Basso, 2011). When studying parenting following premature delivery, it is also imperative to distinguish between prematurity itself and the other factors associated with prematurity (Anderson & Doyle, 2008). For example, preterm birth occurs more often among mothers of low socioeconomic status, who are under 15 years of age, or who have had many pregnancies close together in time (Behrman & Butler, 2006). These demographic factors, without the consideration of premature deliveries, may be related to maternal cognitions and principles. Therefore, differences in beliefs between mothers of preterm and term infants may be ascribable to reasons other than the birth status of their infant. For example, modest relations between complexity of thought and maternal education have been reported (Gutierrez & Sameroff, 1990; Pratt et al., 1993), and lower educational attainment is a risk factor for premature delivery (Behrman & Butler, 2006). In addition, biological risk and neonatal experience of infants (for example, 5-minute Apgar score or days on ventilation) can affect outcomes for children and their social interactions (Aylward, 2002; Hintz et al., 2005; Landry et al., 1997; Vohr et al., 2000). Confounding variables must therefore be carefully explored (Aylward, 2002). To isolate the impact of prematurity on maternal complexity of thought and caregiving principles, we examined multiple demographic and medical covariates for the developmental trajectories of complexity of thought and caregiving principles following preterm and term deliveries. This study ~~~~~~~~~~ This study examined how maternal complexity of thought and caregiving principles develop from the days following a premature or term delivery to a time when the parenting role is more established, 5 months later. Therefore, we assessed the continuity and stability of complexity of thought and caregiving principles in mothers of preterm vs. term infants. Our first aim was to examine whether the continuity and stability of maternal complexity of thought and caregiving principles differed by birth status when measured at birth and again 5 months later. We chose to schedule the follow-up data collection at 5 months to ensure parents had become established in the parenting role and infants were settled (St James-Roberts et al., 2006). We expected to see differences in the developmental trajectories of preterm and term mothers’ complexity of thought and caregiving principles from birth to 5 months. Such results would demonstrate that the dynamics of complexity of thought and caregiving principles look different for mothers of preterm and term infants. Our second aim was to examine potential predictors of any observed change in complexity of thought and caregiving principles. To do this, we determined whether medical and demographic factors accounted for any changes observed. Few studies have systematically examined basic developmental properties of maternal complexity of thought and caregiving principles following premature deliveries. These assessments will increase our understanding of the development of maternal complexity of thought and caregiving principles in general and specifically in response to premature deliveries. Furthermore, this study examined whether prematurity had uniform or differentiated effects on complexity of thought and caregiving principles.","A total sample of 105 mothers completed questionnaires within the first month from delivery and 5 months later as part of a study about mothers’ and their preterm (n = 41) or term (n = 64) infants’ development. Parents of infants born between 30 and 42 weeks gestational age were recruited. Mothers were divided into two groups by the birth status of their infant based on infant's gestational age; infants below 37 completed weeks of gestation were in the preterm sample (up to and including 36 weeks and 6 days; Howson et al., 2012) and those 37 weeks and above were in the term sample. The majority of participants were recruited during the hospitalization period following delivery through the Department of Child Health at University Hospital Wales (UHW, n = 90), with the remaining 15 parents recruited through the Cardiff city registry office and other community links, such as the National Childbirth Trust (recruited either soon after delivery or prenatally but completed the questionnaires soon after delivery). An additional 3 dyads were excluded due to a history of maternal depression or placement on the child-in-need register by local authorities to monitor the child due to concerns about the social environment of the child—for example, exposure to domestic abuse. Families were not approached if their infants had serious medical conditions beyond prematurity alone or congenital abnormalities that could affect growth and development, including requiring surgical intervention during hospitalization. Additionally, multiple births and parents under 16 years old were excluded. Suitable infants were identified as fitting inclusion and exclusion criteria through discussions with medical staff (midwife, nurse, or doctor) responsible for their care. At birth, 148 mothers completed questionnaires. Attrition between birth and 5 months was 29% resulting in the final sample of 105 dyads at 5 months. The majority of parents who did not participate were not contactable at 5 months (53%); others had difficult life circumstances (27%), their infant had developed health problems (10%), or no longer wished to participate (10%). Table 1 compares health and demographic information about participating mothers and their preterm or term infants. The preterm and term samples did not differ on any of the demographic variables (infant birth order, maternal age, ethnicity, marital status, maternal education, and family income) or on 5-min Apgar scores. The samples naturally differed on medical status—preterm infants were born at younger gestational age and lower birthweight and spent more days on ventilation and in the hospital after birth. These variables were all checked as potential covariates and predictors of change in complexity of thought and caregiving principles.","All study procedures were reviewed by the Cardiff University School of Psychology's research ethics committee, the National Health Service's Research & Development, and Local Research Ethics Committee. Mothers consented to participate soon after delivering their baby at which time they completed the Baby Care Questionnaire (BCQ; Winstanley & Gattis, 2013) and Concepts of Development Questionnaire (CODQ; Sameroff & Feil, 1985) for the first time. Mothers completed these two questionnaires again 5 months later. Health and demographic information was collected from a combination of a self-report measure completed by the mother and inspection of medical records by research assistants. Five- month study visits were scheduled based on postnatal chronological age with a window of ±15 days. Gifts were given to infants and parents for participation (worth approximately USD $10). The Baby Care Questionnaire (BCQ) The BCQ (Winstanley & Gattis, 2013) asks parents to rate 30 statements about caregiving on a 4-point Likert scale ranging from strongly disagree (1) to strongly agree (4). The BCQ contains 2 subscales—structure and attunement. Subscale scores were calculated by averaging across relevant items (17 items for structure and 13 items for attunement). Structure represents parent support of regularity and routines in their infant's daily life. For example, It is important to introduce a sleeping schedule as early as possible. Structure showed adequate internal consistency at birth (preterm: α = 0.81; term: α = 0.87) and 5 months (preterm: α = 0.79; term: α = 0.84). Attunement represents parent trust and attention to their infant's cues and support of close physical contact. For example, Responding quickly to a crying baby leads to less crying in the long run. Three items for the attunement subscale had poor distributions and did not relate with overall attunement or other items making up the attunement subscale and so were not used in calculating the average attunement score. The resulting 10-item attunement scale showed adequate internal consistency at birth (preterm: α = 0.60; term: α = 0.75) and 5 months (preterm: α = 0.68; term: α = 0.75). Concepts of Development Questionnaire (CODQ) The CODQ (Sameroff & Feil, 1985) asks parents to rate 20 statements about child development on a 4-point Likert scale ranging from strongly disagree (0) to strongly agree (3). The CODQ measures parent cognitions about child development, in particular parent ability to think complexly about children. The CODQ contains two subscales – categorical and perspectivist – and a summary scale of complexity. For each subscale an average score is calculated. At the categorical level, parent cognitions are restricted to single determinants and single outcomes. For example, Parents must keep to their standards and rules no matter what their child is like. Three items for the categorical subscale had poor distributions and did not relate with the subscale overall or with the other items making up the subscale and so were not used to calculate the categorical subscale score. The remaining 7-item categorical subscale showed adequate internal consistency at birth (preterm: α = 0.68; term: α = 0.77) and 5 months (preterm: α = 0.65; term: α = 0.53). At the perspectivist level, child development is viewed from multiple perspectives, allowing parents to understand that multiple factors can interact and change over time to result in different outcomes. For example, Parents change in response to their children. Cognizing at the perspectivist level allows parents to view and evaluate a large range of developmental possibilities. Three items for the perspectivist subscale had poor distributions and did not relate with the subscale overall or with the other items making up the subscale, so were not used to calculate the perspectivist subscale score. The remaining 7-item perspectivist subscale showed adequate internal consistency at birth (preterm: α = 0.60; term: α = 0.65) and 5 months (preterm: α = 0.60; term: α = 0.67). Previous studies have found similar internal consistency coefficients when using the CODQ (e.g., Benasich & Brooks Gunn, 1996; Landry et al., 1996; Lee, 2005; Manlove, Vazquez, & Vernon-Feagans, 2008). Complexity is calculated as (perspectivist − categorical + 3.0)/2. Therefore, complexity represents the balance between categorical and perspectivist thinking and also ranges from 0 to 3. To ensure the BCQ and CODQ used equivalent scales, after calculating complexity, all subscales of the CODQ were transformed to range from 1 to 4 (by adding 1 to all scores). Therefore, a complexity score of 4 means that parents strongly agree with items related to the perspectivist subscale and strongly disagree with items related to the categorical subscale. Conversely, a complexity score of 1 means that parents strongly disagree with items related to the perspectivist subscale and strongly agree with items related to the categorical subscale. Correlations among maternal structure, attunement, categorical, perspectivist, and complexity for birth and 5-month variables by birth status indicated that structure was not related to any of the cognitions, but attunement was positively related to perspectivist scores and to a smaller extent complexity scores. However, the correlation coefficients did not indicate singularity for any of these measures (shared variance = 0–28%) and so were treated independently. Descriptive and explanatory variables ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Mothers completed a demographic questionnaire after delivery that collected information about demographic variables, previous pregnancies, and the current or most recent pregnancy and delivery. In addition, researchers accessed the medical records of the infant after discharge and created records of the infants’ health during their hospitalization. Analysis plan ~~~~~~~~~~~~~ Prior to data analysis, distributions of categorical thinking, perspectivist thinking, complexity of thought, structure, and attunement at both time points were examined for normality, homogeneity of variance, and influential outliers. All variables met assumptions for parametric tests. The first aim was to assess the continuity and stability of complexity of thought and caregiving principles in mothers of preterm and term infants. Therefore, the effects of child age (birth vs. 5 months) and birth status (preterm vs. term) were tested using repeated-measures Analysis of Variance (RM-ANOVA). Child age (birth vs. 5 months) was treated as a within-subjects variable and birth status (preterm vs. term) was treated as a between-subjects variable. In addition, stability estimates were reported across ages for maternal complexity of thought and caregiving principles by birth status using correlation. Table 2 presents means and standard deviations for maternal complexity of thought and caregiving principles by infant age and birth status. Table 2 also presents correlations between subscales at birth and 5 months by infant birth status. Follow-up analyses were run to control for mothers’ previous experience of having a child (birth order: firstborn vs. laterborn) and having previously had a preterm delivery (no vs. yes). The uncontrolled analyses are reported below as all results controlling for experience of mothers with previous children were the same as uncontrolled analyses except where noted. The next aim was to identify what variables may account for any developmental changes observed in complexity of thought and caregiving principles. Multiple regression analyses were used to examine predictors of differential stability across groups, and RM-ANCOVAs were used to examine predictors of discontinuity. The predictors of change in mean level or discontinuity were examined using RM-ANCOVAs. First, the main effects of infant age (birth vs. 5 months) and birth status, and the effect of the infant age × birth status interaction on cognitions or principles at 5 months were examined. Then, covariates were included to examine their effect on the significance of the main effect of age and the infant age × birth status interaction. Reducing either the main effect of age or the infant age × birth status interaction to nonsignificance would indicate that the covariates, at least in part, accounted for the discontinuity of maternal cognitions or principles. For the predictors of differential stability of cognitions or principles, we used a similar process with multiple regressions. In the first step, cognitions or principles at birth (centered), birth status (preterm vs. term), and the interaction between cognitions or principles at birth and birth status were included as predictors, and cognitions or principles at 5 months was the outcome variable. Covariates were then included in the second step to examine whether their inclusion attenuated the cognitions or principles at birth or interaction term to nonsignificance. Potential health and demographic covariates were selected by examining zero-order correlations with 5-month scores for cognitions or principles of interest. Complexity of thought ~~~~~~~~~~~~~~~~~~~~~ For categorical thinking, there was a main effect of age, F(1, 103) = 12.28, p < 0.001, ηp2 = 0.107, no main effect of birth status, F(1, 103) = 0.00, p = 0.960, ηp2 = 0.002, and a significant interaction between infant age and birth status, F(1, 103) = 4.34, p = 0.040, ηp2 = 0.040. Simple effects analyses examined the interaction between infant age and birth status on categorical thinking. A main effect of age was found for mothers of preterm, F(1, 40) = 15.36, p < 0.001, ηp2 = 0.277, but not term infants, F(1, 63) = 1.17, p = 0.283, ηp2 = 0.018. Categorical thinking was continuous in mothers of term, but not preterm, infants. Mothers of preterm and term infants did not differ in their categorical thinking at birth, F(1, 103) = 1.03, p = 0.312, ηp2 = 0.010, or 5 months later, F(1, 103) = 1.61, p = 0.208, ηp2 = 0.015. Mothers of preterm infants increased their categorical thinking with age, whereas mothers of term infants did not change with age. We examined possible explanations for the discontinuity of categorical scores in mothers of preterm infants. Apgar scores at 5 min, number of siblings, and whether the mother was living with a partner (no vs. yes) were included as covariates because these variables were related to categorical thinking at 5 months. Apgar scores at 5 min were not recorded in medical records for 11 infants (2 preterm and 9 term), and therefore these dyads were excluded, resulting in a final sample of 94 (39 preterm and 55 term). Table 3 presents the F, p, and ηp2 values for step 1—infant age, birth status, and infant age × birth status interaction – and step 2—adding covariates – in predicting categorical thinking. In step 1, there was a significant main effect of age, F(1, 92) = 11.83, p < 0.001, ηp2 = 0.114, but no main effect of birth status, F(1, 92) = 0.32, p = 0.573, ηp2 = 0.003, and a trend level interaction between infant age and birth status, F(1, 92) = 3.77, p = 0.055, ηp2 = 0.039, on categorical scores. Adding covariates attenuated the main effect of age to nonsignificance, F(1, 89) = 0.15, p = 0.705, ηp2 = 0.002 (a 0.112 reduction in the ηp2), but only very slightly reduced the interaction between age and birth status, F(1, 89) = 3.46, p = 0.066, ηp2 = 0.037 (a 0.002 reduction in the ηp2). None of the covariates independently predicted discontinuity in categorical thinking. The RM-ANCOVAs for mothers of preterm and term infants individually showed that the main effect of age was nonsignificant for both groups (preterm: F(1, 35) = 0.97, p = 0.332, ηp2 = 0.027; term: F(1, 51) = 0.41, p = 0.523, ηp2 = 0.008). The only difference in the RM-ANCOVAs was that Apgar scores at 5 min independently predicted categorical scores at 5 months for mothers of preterm infants, F(1, 35) = 4.43, p = 0.042, ηp2 = 0.112 (higher Apgar scores related to higher categorical thinking), but not term infants, F(1, 51) = 0.01, p = 0.932, ηp2 = 0.000. Together, Apgar at 5 min, number of siblings, and mother living with a partner accounted for discontinuity in categorical thinking for both groups, but only Apgar scores at 5 min independently predicted categorical scores at 5 months in mothers of preterm infants. For perspectivist thinking, there was a main effect of age, F(1, 103) = 12.53, p < 0.001, ηp2 = 0.108, a main effect of birth status at a trend level, F(1, 103) = 3.70, p = 0.057, ηp2 = 0.035, but no significant interaction between infant age and birth status, F(1, 103) = 0.48, p = 0.491, ηp2 = 0.005. Mean differences demonstrated that perspectivist thinking was lower, on average, at birth (mean difference = −0.12, SE = 0.03, p < 0.001) and in mothers of preterm infants (mean difference = −0.11, SE = 0.06, p = 0.057). When controlling for previous children and previous preterm birth, the main effect of age remained, F(1, 99) = 17.39, p < 0.001, ηp2 = 0.149, as did the main effect of birth status, F(1, 99) = 4.00, p = 0.048, ηp2 = 0.039. However, there was also a significant interaction between age and birth order, F(1, 99) = 5.11, p = 0.026, ηp2 = 0.049. Simple effects analyses indicated that the main effect of age was only significant for mothers of firstborns, F(1, 61) = 20.39, p < 0.001, ηp2 = 0.251, and not mothers of laterborns, F(1, 36) = 0.02, p = 0.894, ηp2 = 0.000. Mothers of firstborns increased in their perspectivist thinking from birth to 5 months (birth: M = 2.78, SD = 0.35; 5 months: M = 2.96, SD = 0.35) but perspectivist thinking did not change for mothers of laterborns (birth: M = 2.92, SD = 0.35; 5 months: M = 2.94, SD = 0.35) for both groups of mothers. In addition, mothers of preterm infants were lower on perspectivist thinking than mothers of term infants. We examined potential explanations for the discontinuity of perspectivist scores in the mothers of firstborn preterm and term infants. Mother's highest educational qualification, attending antenatal classes (no vs. yes), and breastfeeding at birth (no vs. yes) were included as covariates because these variables were related to perspectivist thinking at 5 months. Only first-time mothers were included, and one dyad was excluded due to missing data for breastfeeding at birth, resulting in a final sample of 63 (22 preterm and 41 term). Table 3 presents the F, p, and ηp2 values for step 1—infant age, birth status, and infant age × birth status interaction – and step 2—adding covariates – in predicting perspectivist thinking. In step 1, there were main effects of age, F(1, 61) = 16.20, p < 0.001, ηp2 = 0.210, but no significant main effect of birth status, F(1, 61) = 2.19, p = 0.144, ηp2 = 0.035, or interaction between infant age and birth status, F(1, 61) = 0.84, p = 0.364, ηp2 = 0.014, on perspectivist scores. Adding covariates attenuated the main effects of age, F(1, 58) = 2.83, p = 0.098, ηp2 = 0.046 (a 0.164 reduction in the ηp2) to nonsignificance. None of the covariates independently predicted perspectivist thinking. Therefore, highest maternal qualification, attending antenatal classes, and breastfeeding together accounted for the discontinuity of perspectivist thinking. Finally, there were no main effects of age, F(1, 103) = 0.01, p = 0.937, ηp2 = 0.000, or birth status, F(1, 103) = 1.82, p = 0.181, ηp2 = 0.017, or interaction between infant age and birth status, F(1, 103) = 0.93, p = 0.336, ηp2 = 0.009, for complexity of thought. Complexity of thought did not differ by infant age or birth status and was therefore continuous for mothers of preterm and term infants. Stability was found for all three CODQ variables in mothers of term and preterm infants (Table 2). Z-tests indicated that the stability coefficients did not differ between mothers of preterm and term infants for categorical scores (z = 1.02, p = 0.308, two-tailed test), perspectivist scores (z = −0.33, p = 0.741, two-tailed test), or complexity scores (z = 0.00, p = 1.00, two-tailed test). Caregiving principles ~~~~~~~~~~~~~~~~~~~~~ For structure, there was no main effect of age, F(1, 103) = 0.12, p = 0.729, ηp2 = 0.001, or birth status, F(1, 103) = 3.28, p = 0.073, ηp2 = 0.031, or interaction between infant age and birth status, F(1, 103) = 0.18, p = 0.671, ηp2 = 0.002. When controlling for previous children and previous preterm birth, there was a significant main effect of birth status, F(1, 99) = 3.96, p = 0.050, ηp2 = 0.038, but still no main effect of age, F(1, 99) = 0.83, p = 0.364, ηp2 = 0.008, or interaction between infant age and birth status, F(1, 99) = 0.07, p = 0.789, ηp2 = 0.001. The main effect of birth status reflected that mothers of preterm infants (M = 2.77, SD = 0.30) supported structure more, on average, than mothers of term infants (M = 2.65, SD = 0.30). For attunement, there was no main effect of age, F(1, 103) = 0.89, p = 0.349, ηp2 = 0.009, or birth status, F(1, 103) = 2.06, p = 0.154, ηp2 = 0.020, or interaction between infant age and birth status, F(1, 103) = 0.04, p = 0.838, ηp2 = 0.000. Therefore, structure and attunement did not differ across age or birth status and were continuous for mothers of preterm and term infants. However, when controlling for previous experience of caregiving, mothers of preterm infants were higher on structure than mothers of term infants. Stability was found for structure for mothers of term infants but not mothers of preterm infants (Table 2). Z-tests indicated that the stability of structure was significantly lower for mothers of preterm infants than term infants (z = −2.50, p = 0.012, two-tailed test). Stability was also found for attunement for mothers of preterm and term infants (Table 2). The stability of attunement did not differ between mothers of preterm and term infants (z = −0.45, p = 0.653, two-tailed test). We examined potential explanations for the differential stability of structure in mothers of preterm and term infants. Family income was included as a covariate because this variable was related to structure at 5 months. Three dyads were missing data on family income and therefore were excluded from these analyses resulting in a final sample of 102 (40 preterm and 62 term). Table 4 presents the regression coefficients for step 1—structure at birth (centered), birth status, and structure at birth × birth status interaction – and step 2—adding family income – in predicting structure at 5 months. Step 1 accounted for 34% of the variance in structure scores at 5 months, R2 = 0.34, F(3, 98) = 16.72, p < 0.001. Neither structure at birth nor birth status predicted structure at 5 months, but the interaction between structure and birth status predicted structure at 5 months. This result reflects the earlier finding that structure was only stable in mothers of term infants and not in mothers of preterm infants. Step 2 did not account for significantly more variance than step 1, ΔR2 = 0.00, F(1, 97) = 0.53, p = 0.467. Adding family income as a covariate attenuated the interaction term to nonsignificance (model 1: β = −0.21, p = 0.037; model 2: β = −0.19, p = 0.065) but only just (a 0.02 absolute reduction in the β). Family income did not independently predict instability in structure. Therefore, family income accounts for some of the differential stability in structure between mothers of preterm and term infants but only to a limited degree.","The first aim of the study was to examine developmental trajectories of maternal complexity of thought and caregiving principles following preterm and term deliveries in the first 5 months of their infant's life. We found prematurity had a differentiated effect on complexity of thought and caregiving principles, with different levels of complexity of thought and caregiving principles affected differently. The second aim was to examine demographic and medical variables that might explain changes in complexity of thought and caregiving principles. Below we first discuss the findings related to complexity of thought and then caregiving principles. Complexity of thought ~~~~~~~~~~~~~~~~~~~~~ Categorical thinking was continuous for mothers of term, but not preterm, infants. Mothers of preterm infants showed increasing levels of categorical thinking from the delivery of their infant to 5 months later (and therefore an increase in less complex thinking). In combination, Apgar scores at 5 min, number of siblings, and whether mothers were living with a partner accounted for the discontinuity of categorical thinking in mothers of preterm infants. The only independent predictor of categorical thinking was 5-min Apgar scores for mothers of preterm, but not term, infants. For mothers of preterm infants, higher 5-min Apgar scores were related to higher categorical thinking scores at 5 months but not at birth. However, it must be noted that 5-min Apgar scores of the preterm infants reflected the low-risk nature of the sample, as scores ranged from 7 to 10 (but around 86% scored 9 or 10). Therefore, this finding needs to be further examined in a higher risk sample with a wider range of Apgar scores. For both groups of mothers of firstborn infants, perspectivist thinking increased over time and overall levels were lower in all mothers of preterm infants regardless of parity. The discontinuity of perspectivist thinking for first-time mothers was accounted for by a combination of highest maternal qualification, having attended antenatal classes, and breastfeeding at birth (however, none of these variables independently predicted perspectivist thinking). Finally, complexity of thought was continuous for both birth status groups. Stability was found at equal levels for all three complexity of thought variables across birth status groups. Despite equal levels of complexity of thought, prematurity had an impact on categorical and perspectivist thinking. Categorical thinking increased in mothers of preterm infants from birth to 5 months but remained continuous in mothers of term infants. In addition, mothers of preterm infants were less able (at trend levels) to demonstrate perspectivist thinking. However, complexity of thought did not differ by the birth status or age of the infant. Despite being able to think as complexly about child development as mothers of term infants, mothers of preterm infants appear to increasingly rely on lower levels of thinking. This result could have important implications because lower levels of thinking have been shown to relate to more rigid and authoritarian parenting (Deković & Gerris, 1992), and more complex levels of thinking have been shown to relate to warm and sensitive parenting (Landry et al., 1996; Miller-Loncar et al., 2000; Pratt et al., 1993). Furthermore, this result could inform the design of NICU interventions. That is, early interventions could focus on helping mothers to take multiple, flexible perspectives when thinking about children and development (and so become less reliant on lower levels of thinking). Further work needs to examine whether the similar levels and trajectories of complexity of thought for mothers of preterm infants can prove protective despite differences in categorical and perspectivist thinking (compared to mothers of term infants). Therefore, relations between complexity of thought (both overall level and trajectories) with later parenting behaviors and child outcomes need to be studied. Caregiving principles ~~~~~~~~~~~~~~~~~~~~~ Structure and attunement were both stable and continuous in mothers of term infants. In contrast, structure was continuous but unstable, and attunement was continuous and stable, for mothers of preterm infants. Therefore, only the caregiving principle of structure appeared to be affected by premature birth. In mothers of preterm infants, structure at birth did not predict structure 5 months later. Therefore, mothers of preterm infants appeared to change their support of structure over the first 5 months but not in a uniform way (i.e., no mean increase or decrease). The differential stability of structure by birth status was reduced to nonsignificance by controlling for family income. However, the reduction of the interaction term of structure at birth × birth status was small, and family income did not add to the variance accounted for in the regression model. Therefore, this result should be considered with caution. Perhaps characteristics of the infant (for example, sleep state stability or clarity of cues) would provide a better explanation of the differential stability of structure than demographic or medical factors. For example, it could be that the lower stability of structure following preterm delivery reflects that these mothers are more willing to be flexible in structure and base that principle on their infant's willingness or ability to fit into a schedule. Alternatively, as preterm birth is often unexpected and accompanied by periods of hospitalization, mothers’ support of structure at birth may not be a representation of their true principles but instead reflect the structure of the hospital and medical staff. As such, mothers of preterm infants may only develop their own principles about structure on leaving the hospital and becoming fully responsible for their infants’ daily care. Both of these hypotheses fit with findings that, when controlling for previous experience with children, structure appeared to be higher in mothers of preterm, as compared to term, infants. Future work should examine such hypotheses. Attunement, in comparison, did not appear to be affected by premature deliveries as attunement showed stability and continuity for mothers of preterm and term infants. Many interventions implemented following premature deliveries or during NICU stays focus on teaching parents to respond to the cues and unique characteristics of their infant (Browne & Talmi, 2005; Kaaresen et al., 2006; Landry et al., 2008). In addition, more optimal outcomes have been found for children who had parents who followed their interests and cues as infants (Landry et al., 1997). Understanding a parent's support of attunement at the start of one of these interventions or following a premature delivery may prove useful in planning how to best approach helping parents. For example, parents who already support the principle of attunement may only require guidance in identifying their infants’ unique cues, whereas parents who show little or no support of attunement may first need the significance of using their infants’ cues to guide caregiving explained. Therefore, this finding of similar levels and trajectories of attunement demonstrates that mothers of preterm infants, at least in principle, do not differ from mothers of term infants in the value they place in using and trusting infant cues and signals to guide caregiving. Therefore, attunement may reflect an enduring caregiving principle of parents that is less affected by external forces, whereas structure is more responsive to environmental factors. More work is needed to examine the predictors of these principles. Limitations ~~~~~~~~~~~ There were relatively high levels of attrition between data collection at delivery and at 5 months. Loss of participants was primarily due to difficulties contacting families when attempting to schedule the 5-month visit. In addition, this sample relied on families consenting to take part in the study and as such was not representative of all preterm infants born in UHW. The subscales of the CODQ showed relatively low internal consistency; however, these consistency estimates are similar to those reported previously with samples of parents of term, as well as preterm, infants (e.g., Benasich & Brooks Gunn, 1996; Landry et al., 1996; Lee, 2005; Manlove et al., 2008). Future work might include a larger, more heterogeneous sample. The current sample was relatively homogeneous, including primarily white mothers of higher income and education. Furthermore, we excluded three families due to a history of depression or concerns about the social environment of the child. We were unable to include these families in the sample because the small number meant that we were not able to control for these social risk factors. Future studies, with larger, more diverse samples, should include such variables in analyses. This study is, however, an important first step to examining the effects of infant prematurity on the development of maternal cognitions and principles.","This study demonstrates differences in the trajectories of some of the studied cognitions and caregiving principles following preterm, as compared to term, deliveries. The studied maternal cognitions and caregiving principles were not affected equally by prematurity. Mothers increasingly relied on categorical thinking in establishing the caregiving role for their preterm infants and were less stable in reports of structuring their caregiving. However, both structure and categorical thinking showed similar levels for mothers of preterm and term infants at both time points. This study highlights the need to examine developmental trajectories of cognitions and principles. In addition to differences in the trajectories of complexity of thought and caregiving principles, mothers of preterm infants tended to report lower perspectivist thinking than mothers of term infants following delivery and after 5 months of caring for their infant. These results provide insight into the development of complexity of thought and caregiving principles following premature deliveries and tell us about the broader context in which cognitions about child development and caregiving principles develop in parents. For example, first-time mothers (at a group level) increased in their perspectivist thinking as their infants grew older, while mothers of laterborns showed consistency across time. In addition, given that different cognitions about child development and caregiving principles showed different trajectories these results further highlight that parenting is multidimensional (Bornstein & Cote, 2004; Bornstein & Tamis-LeMonda, 1990; Bornstein, Tamis-LeMonda, Hahn, & Haynes, 2008). Further work needs to examine the change and consistency of these cognitions about child development and caregiving principles later in infancy, as well as their impact on parenting behaviors and child outcomes.","Alice Winstanley, Department of Psychology, University of Cambridge; Rebecca G. Sperotto and Merideth Gattis, School of Psychology, Cardiff University; Diane L. Putnick and Marc H. Bornstein, Eunice Kennedy Shriver National Institute of Child Health and Human Development; and Shobha Cherian, Department of Child Health, University Hospital of Wales."],["This paper examines how office-based lighting and computer use behaviours relate to similar behaviours performed by the same individuals in a household setting. It contributes to the understanding of energy use behaviour in both household and organisational settings, and investigates the potential for the 'spillover' of behaviour from one context to another. A questionnaire survey was administered to office-based employees of two adjacent local government organisations ('City Council' and 'County Council') in the East Midlands region of the UK. The analysis demonstrates that the organisational or home setting is an important defining feature of the energy use behaviour. It also reveals that, while there were weak relationships across settings between behaviours sharing other taxonomic categories, such as equipment used and trigger for the behaviour, there was no evidence to support the existence of spillover effects across settings. © 2014 The Authors. --------------------------------------------------------------------------------","In recent years, concern about environmental impacts and the cost, availability and security of energy supplies has led to heightened interest in ways to reduce energy use within buildings. For psychologists, work in this area has frequently focused on understanding the determinants of energy use behaviours, or on testing the effectiveness of intervention strategies aimed at changing behaviours (Abrahamse, Steg, Vlek, & Rothengatter, 2005). Much of the research into the determinants of energy use behaviours has focused on household settings (Abrahamse, Steg, Vlek, & Rothengatter, 2007; Owens & Driffill, 2008; Steg, Dreijerink, & Abrahamse, 2005). However, non-domestic buildings account for around one quarter of total UK energy use (Brown, Wright, Shukla, & Stuart, 2010), with local government buildings alone estimated to consume 26 billion kWh of energy annually (Carbon Trust, 2007). Interest is now growing in understanding energy use behaviours in non-domestic, organisational settings such as offices and other workplaces (Lo, Peters, & Kok, 2012; Matthies, Kastner, Klesse, & Wagner, 2011; Murtagh et al., 2013, in press; Scherbaum, Popovich, & Finlinson, 2008). At the same time, many behaviour change interventions include, explicitly or otherwise, the notion of ‘spillover’ – that encouraging people to take up one pro-environmental behaviour may lead them to take up further pro-environmental behaviours (Thøgersen & Ölander, 2003). By exploring how office- based lighting and computer use behaviours relate to similar behaviours performed by the same individuals in a household setting, this paper contributes to the understanding of energy use behaviour in both household and organisational settings, and investigates the potential for ‘spillover’ of behaviour from one context to another. Energy saving behaviours such as turning off equipment when it is no longer in use are not necessarily motivated by pro-environmental intentions; they may be the result of, for example, habit or routine, organisational practice, a personal dislike of waste, or a fear of electrical faults. Literature exploring these behaviours from an environmental standpoint can nevertheless provide insights. Stern (2000) identifies and describes four classes of pro- environmental behaviour: environmental activism such as involvement in environmental organisations; non-activist public behaviour such as support for or acceptance of public policies; private-sphere environmentalism including the purchase, use and disposal of household products; and other environmentally-significant behaviour including behaviour within organisations. This classification distinguishes behaviours performed in household settings from those performed in organisational settings. In particular, it identifies that individuals may affect the environment by influencing organisations to which they belong, or by how they carry out their role within an organisation. Much of the literature examining individual environmentally-significant behaviour focuses on behaviours that could be classed as private-sphere environmentalism: waste and recycling (e.g. Barr, 2007; Tudor, Barr, & Gilg, 2007), energy demand (e.g. Abrahamse et al., 2005) and travel mode choice (e.g. Anable & Gatersleben, 2005; Bamberg & Schmidt, 2003). For much of this research, the context of the behaviour is a household setting, where individual control over the performance of behaviours is likely to be relatively high. While even in households individuals do not have complete autonomy (their behaviour may be influenced or constrained by the people they live with, or by the finances, time or facilities available to them) it is still likely that an individual will have greater control over these behaviours in their own home than in an organisational setting such as an office. In offices, behaviours are shaped by the physical context of the office (the presence of controls over building systems or equipment), but also by the social context (the needs, expectations or norms of the people they share the office with) and by the organisational context (the policies and expectations of the organisation that employs them). However, many pro-environmental behaviours within organisational settings could fit into more than just Stern's (2000) fourth category of ‘other behaviour including within organisations’. Non-activist public behaviour within an organisation could include support for a company's environmental policies, while private-sphere environmentalism choices could affect an employee's actions within the workplace. For such behaviours to be classified separately to similar behaviours performed in a household setting, the setting that the behaviour occurs within would need to be a defining feature of that behaviour. A number of researchers have considered how environmentally-significant behaviours in one setting relate to similar behaviours in different settings. Barr, Shaw, Coles, and Prillwitz (2010) found that people tend to behave in a less pro-environmental manner when on holiday than when at home, often finding it difficult to transfer commitment to environmental action into other, more problematic contexts. The problematic aspects of other contexts are likely to vary according to the nature of the context in question. This is an area that has not yet been fully explored by researchers. However, it has been identified that the influencing factors most relevant to a particular behaviour are specific to each context (Stern, 2000). For example, Siero, Bakker, Dekker, and Van Den Burg (1996) argue that it is not possible to generalise from household energy saving behaviour to workplace energy saving behaviour because expenditure is experienced more directly by the household, while employees only benefit indirectly from financial benefits of energy saving at work. However, this suggests that cost is an overriding factor in the decision-making process, while other research has identified a wide range of factors that may influence environmentally-significant behaviour, including situational characteristics, prior awareness and experience of the behaviour, habits and routines, environmental beliefs and values, social and personal norms, and perceptions of behavioural control and self- efficacy (Bamberg & Möser, 2007; Barr, 2007; Clayton & Brook, 2005). Who pays for the energy used, then, is only one difference between the home and workplace settings, and not necessarily the decisive difference. Where connections have been found between behaviours performed in household and organisational settings, prior experience of the behaviour has been shown to be important. Studies of waste and recycling behaviour found that office workers who actively recycled at home were more likely to recycle paper (Lee, De Young, & Marans, 1995) and textiles (Daneshvary, Daneshvary, & Schwer, 1998) at work than colleagues who did little home recycling, while a sample of hospital workers reported recycling similar items in the workplace to those they recycled at home (Tudor et al., 2007). Tudor et al. (2007) suggest that similarities between specific recycling items may act as a cue to prompt the behaviour in each location. Barr (2007) suggests that the link identified by Daneshvary et al. (1998) between behavioural experiences in one setting and action in another implies a ‘behavioural snowball effect’, with participation in one behaviour leading to uptake of others. This has also been identified as a ‘spillover’ effect in the context of behaviour change interventions (Thøgersen & Ölander, 2003). Much of the evidence suggesting the existence of a spillover effect is correlational (Barr, Gilg, & Ford, 2005; Poortinga, Whitmarsh, & Suffolk, 2013; Thøgersen & Noblet, 2012; Whitmarsh & O'Neill, 2010), with evidence that correlations between behaviours increase with the similarity (Bratt, 1999) and the perceived similarity (Thøgersen, 2004) of the behaviours. Thøgersen and Noblet (2012) argue that behaviours in the same taxonomic categories (time and place of behaviour, skills employed etc.) tend to be more strongly correlated than behaviours within different taxonomic categories. For similar behaviours in household and organisational settings, however, it is not clear whether prior experience of the behaviour in one setting will encourage the performance of the behaviour in the other setting, leading to spillover effects, or whether differences between the household and organisational contexts will lead to differences in the performance of the behaviour. This question is important because the concept of spillover is influential in the design of many public behaviour change campaigns, which encourage people to take small steps to mitigate environmental impacts in the hope that small actions will lead to more and larger pro-environmental actions (Thøgersen & Crompton, 2009). If such an effect does exist and can be encouraged across contexts, this could add to the potential influence of behaviour change campaigns, with workplace-based campaigns able to influence home behaviours and vice versa. However, Nye and Hargreaves (2010) argue that different mechanisms drive behaviour change in workplace and household settings, with normative influences particularly influential in the workplace. Furthermore, the notion of spillover is problematic. Thøgersen and Noblet (2012) criticise behaviour change programmes and policies that attempt to trigger spillover, arguing that there is little evidence that ‘wedge’ or ‘catalyst’ behaviours lead to large behavioural changes, beyond a weak ‘foot in the door’ effect. This effect suggests that performing pro-environmental behaviours can ‘prepare the ground’ for acceptance of more far-reaching pro-environmental changes, but that this is likely to only work when the original behaviours are considered pro- environmental, rather than common, socially mandated or providing individual benefits (Poortinga et al., 2013; Thøgersen & Noblet, 2012). This is problematic in organisational settings such as offices, where other considerations such as carrying out tasks related to the job role, meeting the expectations of the employing organisation or interacting with colleagues in a shared environment may lead to multiple or competing motivations. This paper, then, addresses two questions: Is there a fundamental difference between energy use behaviours performed in the organisational setting of an office and energy use behaviours performed in a household setting? Does the performance of an energy use behaviour in the organisational setting of an office spill over to influence the performance of related behaviours in a household setting? These questions are addressed by examining responses to a questionnaire survey on lighting and computer use in office and household settings. This allows the connections between the performance of similar behaviours by the same individuals across organisational and household settings to be examined. The study ~~~~~~~~~ This paper discusses responses to a questionnaire survey administered to office-based employees of two adjacent local government organisations (‘City Council’ and ‘County Council’) in the East Midlands region of the UK. The independence of the study from the Councils and the confidentiality of responses given were emphasised in the invitation to take part and in the questionnaire's introductory text. Results of the study were shared with the Councils, but only at an aggregate level once the analysis was complete. Responses were examined in two stages. Stage one investigated whether there was a fundamental difference between energy use behaviours performed in an office setting and in a household setting. Stage two investigated whether there was evidence for the spillover of energy use behaviours from the office to the household setting. Respondents from the City Council (n = 337) were based in a single modern open-plan office building with predominantly centrally-controlled or automated lighting (hereafter ‘City Central Building’). Respondents from the County Council (n = 296) were based in 32 separate office buildings, but with 226 respondents concentrated in four main buildings. The remaining respondents were mostly based in small offices within specialist buildings such as libraries and children's centres. Most of the County Council respondents (259 of 296) reported that they had some individual-level control over lighting within their office building. In stage one, all respondents were included. In stage two, the 337 responses from the City Central Building (19% of the 1785 occupants) were compared with the largest group of responses from a single building within the County Council sample (n = 144, 32% of the 450 occupants). This building (hereafter ‘County Individual Building’) was an older office building where occupants had a higher level of individual control over lighting than in the City Central Building. While occupants of the City Central Building were only able to turn lights off in meeting rooms, occupants of the County Individual Building were able to turn lights off in meeting rooms, toilets, and open plan offices (using light switch cords hanging from ceilings above the desks). Including the full County Council sample for stage one rather than just those included in stage two gave a larger sample (296 rather than 144), which was better for conducting Principal Components Analysis (Field, 2009). Limiting the County Council sample in stage two to those from the County Individual building meant that levels of individual control over energy use were consistent across the whole of each sample. Survey design ~~~~~~~~~~~~~ The questionnaire was administered as a web-based survey, with respondents invited to take part via all-staff emails and through advertising on each organisation's intranet. The wording of these invitations was provided by the researchers and emphasised their independence from the Councils, although the all-staff emails themselves were sent by an employee from each Council's Sustainability team. Sections in the questionnaire covered socio-demographics, self-reported lighting and computer use in the office and in the home setting, and a selection of attitude-behaviour items. This included items measuring constructs taken from the Theory of Planned Behaviour (Ajzen, 1991). In this paper, these were used to identify how variables measuring different aspects of behaviours in the office and home settings grouped together; an investigation of the relationships between constructs proposed by the Theory of Planned Behaviour is not presented here (Littleford, 2013). Office behaviours examined relate to lighting and computer use. These behaviours had quite a small impact in terms of the amount of energy the appliances consumed, or the potential energy savings that could be made through changes in individual behaviour. However, individual office occupants had a greater level of control over the use of computers and lighting than they did over other building systems that consumed more energy, particularly heating and cooling. In the City Central building, temperature was controlled by a Building Management System, with no opportunity for individual control. In the County Individual building, the heating system could be controlled on each floor but worked poorly, with large fluctuations in internal temperature in different parts of the building making this a contentious issue. These circumstances meant that no comparisons could be drawn between office and home heating behaviours for participants from the City Central building as they could not control heating in the office, while participants in the County Individual building's heating behaviours were dominated by experiences of discomfort, making it unlikely that meaningful comparisons could be drawn with their heating behaviours at home. Furthermore, the agreement reached with the two Councils participating in the study was to focus on behaviours within the office buildings, so other energy intensive behaviours such as driving were outside the scope of the study. The study aimed to develop insights into behaviour in the office and home settings, and behaviours that were commonly performed in both settings, such as using lighting and computer equipment, allowed the study to identify processes and relationships across settings, even if the behaviours in question were not the most energy intensive behaviours performed in each setting. Respondents from the County Council were asked to report their performance of three lighting behaviours: turning office lights off when they were not needed, turning meeting room lights off when they leave the room empty, and turning toilet lights off when leaving them unoccupied. In the City Central Building, office and toilet lights were controlled centrally by the Building Management System, but occupants were able to turn off lights in meeting rooms, so were only asked about these. All respondents had the same level of individual control over their computer use, and were asked about three computer-related behaviours: turning off the computer when they finished for the day, turning off the computer monitor when they finished for the day, and turning off the computer monitor when away from their desk for more than ten minutes. Respondents were asked how often they performed each behaviour, with five response categories: Never, Rarely, Half the time, Frequently, and Always. Home behaviours examined in the questionnaire were chosen for their similarities to the office behaviours. All were asked how often they performed two lighting-related and two-computer related behaviours, and again were given five response categories. The two lighting behaviours were turning off lights in a room when they weren't needed, and turning off lights in a room when leaving the room empty; these matched the wording of the office lights and meeting room lights questions respectively. The two computer-related behaviours were turning off the home computer when finished using it, and turning off the computer monitor when away for more than ten minutes. The first question referred to ‘when finished using it’ rather than ‘when finished for the day’ used in the office setting, reflecting that in the office context the occupant left the vicinity of the computer at the end of the working day (by going home), but in the home context was more likely to remain in the same environment. The question about turning off the computer monitor when away for more than ten minutes matched the wording in the office setting. An additional item in the Home setting stated ‘I turn the main TV off fully instead of leaving on standby’; while this was similar to the computer-related behaviours (turning off the equipment when it was finished with), there was no direct comparison in the Office setting. Respondents were also asked to state their level of agreement with a range of attitudinal statements, on a five-point scale (‘Strongly disagree’ to ‘Strongly agree’). These included items measuring constructs within the Theory of Planned Behaviour (Ajzen, 1991), and additional items relevant to the household and organisational settings. The constructs from the Theory of Planned Behaviour were Attitude towards the behaviour (ATT), Subjective Norm (SN) and Perceived Behavioural Control (PBC). These were measured at a behaviour and setting-specific level, with six items measuring ATT for each behaviour in the Office setting, six in the Home setting, six measuring SN in the Office setting and four in the Home setting, and two measuring PBC in each of the Office and Home settings. Further attitude statements related to the respondent's sense of their own responsibility for saving energy at work and at home, their sense of moral obligation to save energy in each setting, and whether they saw reducing the Council's energy use as ‘good’ or ‘important’. Additional items addressed the respondent's perceptions of the organisation's expectations of its employees, the organisation's commitment to energy conservation, and the importance placed on energy conservation by senior management.","In stage one, Principal Components Analysis (PCA) was conducted on reported behaviours in both settings, and on responses to items measuring constructs within the Theory of Planned Behaviour. Direct Oblimin rotation, an oblique rotation method, was used, as it could not be assumed that the variables were fully independent. The analysis was conducted separately on the City Council (n = 337) and County Council (n = 296) samples; as the non- normal distribution of the results for many variables limited the generalisability of the findings from a single sample, results from two separate samples helped to confirm the findings (Field, 2009). Stage two used the whole of the City Council sample (n = 337) based in the City Central Building, and a sub-set of the County Council sample (n = 144) based in the County Individual Building. A non-parametric correlation technique, Spearman's rho, was utilised to identify significant associations between office and home based attitudes and behaviours within each sample. Subsequently, hierarchical multiple regression was used to identify significant differences between the two samples in both the office and home settings. This allowed differences in the gender make-up of each sample to be controlled for, enabling the analysis to better identify effects resulting from differences in the office buildings the respondents were based in. If such differences existed in the office setting but not in the home setting, this could indicate that behaviour was not spilling over between the two settings. Characteristics of the samples ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 1 presents demographic details about the samples (including for the sub-sample of the County Individual Building, taken from the County Council sample). For most demographic items, the proportions of respondents in each category were similar across all of the samples. Most respondents were full time employees who were not in a managerial role and were owner-occupiers of their home. The largest difference between samples was in the gender split, with females making up 57.9% of the City Council/City Central Building sample, but only 39.6% of the County Individual Building sample, possibly reflecting a difference in the kinds of departments based in each building. While the City Central Building housed the majority of the City Council's office-based employees, the County Individual Building housed a sub-section of the County Council's office-based employees and included some technical services (highways, transport) which have been noted nationally to be dominated by men, with women making up only 5% of senior local government roles in highway services and 9% in transport across the UK (LocalGov.co.uk, 2008). Stage one: the distinction between behaviour in the office and home settings ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 2 presents summary statistics for each of the self-reported behaviours included in the questionnaire for respondents from both Councils. High levels of performance were reported for all behaviours apart from turning off computer monitors when away in both the office and home settings. The highest reported performances of behaviours came in the office setting. There was a near-universal performance of turning off a computer at the end of the day, with 97.3% (City Council) and 96.8% (County Council) of respondents reporting that they ‘Frequently’ or ‘Always’ performed this behaviour. These high levels of reported enactment are a useful finding but resulted in too little variance in results to be used in stage two of the analysis. Reported performance of turning meeting room lights off was slightly lower, with 93.8% (City Council) and 90.6% (County Council) selecting ‘Frequently’ or ‘Always’. Turning office lights off when they were not needed, which only the County Council respondents were able to perform, was reported much less frequently (74.1% ‘Frequently’ or ‘Always’). Respondents from the County Council consistently reported a more frequent performance (albeit to a small extent) of each office-based computer-related behaviour, while respondents from the City Council reported slightly higher frequencies of performance of turning off meeting room lights. However, the overall patterns of computer-related behaviours in the office setting were similar for both Councils, with high reported frequencies of turning off computers and monitors at the end of the day, but low reported frequencies of turning off monitors when away for more than ten minutes. Turning off monitors when away for more than ten minutes was performed less frequently in the office than at home, with 14.1% (City Council) and 16.8% (County Council) reporting that they ‘Frequently’ or ‘Always’ turned off their monitor when away more than ten minutes in the office, but 32.8% (City Council) and 36.0% (County Council) reporting the same in the home setting. However, turning off the computer when finished using it at home was performed less often than the equivalent, near-universal office behaviour of turning off the computer at the end of the day, with 83.9% (City Council) and 79.5% (County Council) reporting that they ‘Frequently’ or ‘Always’ turned off the home computer. This suggests that there is a difference in reported performance between similar behaviours in the office and home settings. To explore this further, Principal Components Analysis (PCA) was conducted to identify how the reported behaviours grouped together. Table 3 presents the results of the first PCA, conducted on the reported performance of nine behaviours (four in the office and five at home) in the County Council sample. All factor loadings above .3 (or below −.3) are presented, with loadings used in factor identification in bold. Data was excluded pairwise to minimise losses due to missing responses, providing a sample of between 216 and 285 for each behaviour. The PCA was conducted using Direct Oblimin rotation, and identified three components with eigenvalues above 1, explaining 29.3%, 14.8% and 12.6% of the variance respectively. The items clustering on the same components suggest that component 1 represents Home behaviours, component 2 represents Computer Monitor behaviours, and component 3 represents Office Lighting behaviours. The Home behaviours component had a reasonably strong reliability, α = .644, but the components for Monitor behaviours (α = .598) and, particularly, Office lighting behaviours (α = .421) were weaker. These components are made up of a small number of items, which has been noted to weaken the results of Cronbach's α tests for the internal reliability of a scale (Field, 2009). PCA was also conducted on items measuring behaviour in the City Council sample, using the same parameters. Initial testing revealed low levels of correlation and communality for three variables in this sample: turning off meeting room lights, turning off the home computer, and turning the main TV of fully instead of leaving on standby. These were excluded from the analysis, leaving only four items to test. Despite this, the City Council sample provided some support for the factor structure identified in the County Council sample, identifying components representing Home behaviours (Home lights when not needed and Home lights when leave room empty, eigenvalue 1.938, 48.5% of variance, α = .843) and Monitor behaviours (Office monitor when away from desk and Home monitor when away from desk, eigenvalue 1.204, 30.1% variance, α = .575). The results of these PCA support the presence of a stable factor structure and identify that behaviours group together at a specific level, based on the type of equipment (lighting, computer monitor). For the lighting behaviours, these also grouped according to the setting that the behaviour occurs within, providing some evidence that energy demand behaviours performed in an office setting are different to energy demand behaviours performed in a home setting. However, it was not possible to see whether this was also true for the computer monitors as there were too few items to form separate components in each setting. For this reason, further PCA (again using Direct Oblimin rotation) were conducted using responses to items measuring constructs within the Theory of Planned Behaviour. Using these statements gave a larger number of measured items for each behaviour. Statements measured ATT Attitude towards the behaviour, SN Subjective Norm and PBC Perceived Behavioural Control, using multiple statements tailored to each behaviour in each setting. Each sample was asked about four behaviours. As the City Council sample had no individual control over office lighting, they were asked about turning off meeting room lights when leaving the room empty, while the County Council sample was asked about turning off office lights when they weren't needed. The other behaviours were the same for both samples: turning off computer monitors when away for more than ten minutes in both the office and home settings, and turning off home lights when they were not needed. Table 4 presents the results of the PCA conducted on the City Council sample. The analysis identified 11 components with eigenvalues over 1, explaining between 2.7% and 21.5% of the variance. Most components clustered around items relating to the same equipment, Theory of Planned Behaviour construct and setting for the behaviour. Where the components were not consistent with this (components 7, 8, 10 and 11), they reflected a weakness in the measurement of the Perceived Behavioural Control construct, which was only measured with two items for each behaviour, with one of those items being reverse-worded. The greater number of items relating to each specific behaviour in this analysis reveals a distinction between monitor behaviours in the office and home setting that could not be identified in the PCA of behaviours. For all of the behaviours presented here, then, setting is a defining feature on which the components cluster. The analysis also reveals similarities between behaviours performed in the same setting. Component 4 combines items measuring the Subjective Norm (SN) for both switching off home monitors and switching off home lighting, although the factor loadings form distinct groupings by equipment within that component. This suggests that the relationship between the Subjective Norm and the reported performance of the behaviours is similar for both pieces of equipment in the home setting. However, the influence of SN in the office setting appears to be different, with SN statements for computer monitor and lighting behaviours in the office setting clustering separately from each other and from their equivalents in the home setting. A Principal Components Analysis was also conducted using responses from the County Council sample, and the results supported those found for the City Council sample. In the County Council sample, 9 components were identified, accounting for between 3.0% and 17.4% of the variance, with similar patterns of clustering according to construct, equipment and setting for each specific behaviour. The results of Principal Components Analysis for both samples support the contention that behaviours in the office and home setting are different, even when the other taxonomic categories relating to those behaviours (e.g. equipment, action, trigger for the action) are very similar. Stage two: connections and spillover between office and home settings ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Stage one of this analysis identified that the setting in which an energy use behaviour occurs is an important factor in defining that behaviour. Stage two examines relationships between energy use behaviours in the office and home setting, and explores whether there is evidence of the spillover of behaviour between the two settings. The analysis focuses on two specific buildings, the City Central Building with high levels of centralised control over lighting through a Building Management System, and the County Individual Building with a higher level of individual occupant control over lighting, including light switch cords hanging from ceilings above individual desks. Correlations were calculated using the non-parametric Spearman's rho. Table 5 presents the results of the correlations for the samples from both buildings. A large number of significant correlations (r) were identified, reflecting the high levels of reported performance for most behaviours. There was a greater number of significant correlations (p < .05) between behaviours in the same setting (17) than in different settings (12); of the most highly significant (p < .01), 13 were in the same setting and five in different settings. However, the setting was not the only distinguishing feature. There were also a greater number of significant correlations between behaviours in the office and home settings for the County Individual Building (9) than for the City Central Building (3). Excluding the office lighting behaviour that only the County Individual occupants could perform reduced this to a difference of seven to three. The relationships between the office-based behaviour of turning off meeting room lights and the four home behaviours were the clearest difference between the two samples, with all four correlations for the City Central Building non-significant and for the County Individual Building highly significant. The meeting room lighting behaviour in County Individual Building correlated with all of the home-based behaviours, while the meeting room lighting behaviour in City Central Building did not correlate significantly with their performance of the home behaviours. Effect sizes were calculated using r2, indicating the proportion of variance in the ranked data explained by the correlation. The effect sizes for the correlations for County Individual Building were small, explaining between 3% (Home lights when not needed) and 9% (Home computer when finished) of shared variance in the ranks. Nevertheless, this highlights the lower level of individual control over lighting in the City Central Building than in the County Individual Building. However, the correlations for turning office lights off when not needed, which only respondents in County Individual Building could perform, were only significant for two of the four home-based behaviours (Home lights when not needed, and Home computer when finished). The correlation with the other home lighting behaviour, Home lights when empty, was not significant. This suggests that the correlation is not only related to the type of equipment (lights, computer monitors), but also to the triggers for behaviour (when not needed, when leaving a room empty). Given this, it is of no surprise that the strongest correlation between behaviours in different settings in both samples was between the office-based and home-based versions of turning off the computer monitor when away for more than ten minutes, explaining 12.2% of the variance for City Central Building and 9.6% of the variance for County Individual Building. These effect sizes are quite small, but nevertheless significant. In both locations, the behaviours share the type of equipment and the triggers for the behaviour. The second biggest effect size between locations for County Individual Building (explaining 9% of the variance) was between turning meeting room lights off when leaving the room empty, and turning the home computer off when finished using it. These are different types of equipment, but could arguably share a trigger of their use having finished. However, this relationship is non-significant for the City Central Building, suggesting that other differences between the samples are influencing reported behaviour. To investigate some possible differences between the samples from each building, responses to a number of statements measuring attitudes and organisational factors were examined using hierarchical multiple regression. As the gender make-up of each sample was markedly different, this method of analysis allowed the effects of gender to be controlled for when identifying relationships with the building that respondents were based in. Table 6 presents the results of the regression analysis, revealing that, once gender was controlled for, significant differences between the two samples were found for two of the three office-based behaviours and for three of the four office-based attitude statements. Once gender was controlled for, the building that the respondent was based in did not have a significant relationship with any of the three organisational variables measured. Significant differences were seen for nine of the ten variables examined, although all were small effects; the largest, for the attitude item ‘Reducing the Council's energy use is a good thing’, accounted for just 5% of the variance. Significant relationships between reported behaviours and the building that the respondents were based in were found for two of the three office-based behaviours. Respondents in the County Individual Building were significantly more likely to report turning off their monitor at the end of the day than respondents in the City Central Building (explaining 4% of the variance). Conversely, respondents in the City Central Building were significantly more likely to report turning off meeting room lights, although this only accounted for 1% of the variance. Significant relationships were also found between the building the respondents were based in and their responses to three attitude statements. Respondents in the City Central Building were more likely to agree that reducing the Council's energy use was ‘a good thing’ (explaining 5% of variance) and ‘important’ (2%), while respondents in the County Individual Building were more likely to agree that they ‘should do what they can’ to help the Council save energy (2%). The remaining attitude statement, ‘I should do what I can to help the Council save energy’, was significantly related to both gender and the building the respondent was based in; the building alone explained 2% of variance, rising to 3% once gender was accounted for, revealing that women and respondents in the County Individual Building were more likely to agree. These results distinguish between an assigned responsibility to act (‘not my responsibility’), a moral sense of obligation to act (‘should do what I can’), and an assessment of the value of acting (‘important’ and ‘a good thing’). Differences in behaviours across the two buildings were accompanied by differences in the respondents' sense of moral obligation to act (with respondents in County Individual Building feeling this more strongly) and their assessment of the value of acting (with respondents in City Central Building feeling this more strongly). Gender alone accounted for differences between the samples found in three variables, with women being significantly more likely to agree that Council employees were expected to try to conserve energy (explaining 1% of variance), that the organisation was committed to saving energy (2%), and that senior management saw this as a priority (3%). As the two samples originate from different organisations, these three statements were designed to examine whether organisational differences accounted for differences between the samples, measuring respondents' perceptions of organisational commitment to energy saving. The first, ‘People who work for the Council are expected to try to conserve energy’, measured perceptions of the expectations placed on respondents by the organisation. The second, ‘The Council is committed to saving energy’, measured perceptions of the organisation's commitment to saving energy. The third, ‘Senior management see conserving energy as an important priority’, measured perceptions of the importance of energy saving to the organisation's leadership. Responses to these statements were not as positive as for attitudes and behaviours. For the first two statements, most respondents selected either ‘3 = Neither agree nor disagree’ (City Central Building, 36.2% and 26.1%; County Individual Building, 35.3% and 39.1%) or ‘4 = Tend to agree’ (City Central Building, 39.9% and 44.2%; County Individual Building, 35.3% and 39.1%). For the third statement, responses were even more ambivalent, with 46.3% (City Central Building) and 42.0% (County Individual Building) selecting ‘3 = Neither agree nor disagree’. While the results did identify small differences relating to gender, the results did not distinguish between the two buildings, suggesting that the respondents' perceptions of their organisation's commitment to energy saving did not explain differences between the two samples' office-based attitudes and behaviours. To test whether differences between responses from each building carried over into the home setting, respondents were also asked to report on home-based behaviours and respond to home-based attitude statements. Hierarchical multiple regression was again conducted on the results to test whether the building the respondent was based in or their gender was related to their responses (Table 7). Only one variable was found to be significantly related to either the building or gender. Women and those based in the County Individual Building were significantly more likely to agree with the attitude statement ‘I should do what I can to save energy at home’, explaining 3% of variance. However, gender had a greater effect than the building, which alone accounted for just 1% of variance. The results reveal that, while there were significant differences between the two samples for attitudes and behaviours in the office, there were no significant differences for behaviours at home, and only one small difference on attitudes at home. If there had been significant differences in behaviours reported in the home setting consistent with those seen in the office setting, this would have provided some evidence of factors influencing behaviour across different settings. Instead, no evidence has been found to support the existence of spillover effects between behaviours reported in the office and home settings, beyond weak correlations between behaviours, and only small evidence of consistency in attitudes across settings.","This paper addressed two questions that explore the relationships between energy use behaviours performed in office and home settings. Firstly, the paper examined whether there was a fundamental difference between energy use behaviours performed in the organisational setting of an office and in the home. Using Principal Components Analysis (PCA), it was found that factor loadings grouped energy use behaviours according to similarities in the equipment involved, the trigger for the behaviour, and the setting that the behaviour took place within. Furthermore, even where the energy use behaviours in each setting involved the same action and the same trigger for performing the action, the factor loadings still clearly distinguished between settings. This suggests that the organisational or home setting is an important defining feature of the energy use behaviour. The second stage of the paper examined relationships between the energy use behaviours performed in the office and the home setting, and whether this provided any evidence for the spillover of energy use behaviours between settings. A greater number of correlations between behaviours were found within the same setting than was found across different settings. Correlations were found between behaviours that shared the same type of equipment (lighting, computer monitors) and triggers for the behaviour (leaving a room, finishing using the equipment) as well as the setting for the behaviour. The strongest correlation across settings was found for turning off a computer monitor when away more than ten minutes, with the behaviours sharing the same type of equipment and trigger for the behaviour in each setting. The correlations identify that relationships between behaviours are strongest when they share defining features; the PCA analyses identify that setting is a particularly important defining feature. This suggests that spillover effects across settings would be most likely to occur where other taxonomic categories (equipment, trigger for the behaviour) were similar. However, this study did not identify evidence of such an effect. Indeed, differences between the behaviours reported by respondents across the two buildings provided a further opportunity to identify connections between behaviours in different settings, but again did not find evidence of spillover. Respondents from the City Central Building were significantly more likely to report turning off meeting room lights than respondents from the County Individual Building, while those from the County Individual Building were significantly more likely to report turning off their computer monitor at the end of the day. However, no significant differences were found between the two samples for reported behaviours in the home setting; the causes of differences in reported behaviour in the office setting did not carry over into the home setting. The correlations revealed further differences between the behaviours reported by respondents from each office building. There was a greater number of correlations between all behaviours reported by respondents from the County Individual Building than by respondents from the City Central Building. Across settings, this difference was particularly marked, with the office behaviour of turning off meeting room lights correlating significantly with all four home behaviours for County Individual respondents, but with no home behaviours for City Central respondents. The main difference identified between these buildings was the level of individual control that occupants had over lighting, with those in the County Individual Building having a higher level of control. The greater number of correlations between office and home behaviours for respondents from the County Individual Building suggests that people behave more consistently across settings when they have greater control over their own behaviour. However, in settings such as offices, individual control is more than just physical control; it is also normative, reflecting the influence of an environment shared with colleagues and shaped by the expectations of the employing organisation. Connections between behaviours across settings, then, depend on the features of the behaviours in question (equipment, trigger for the behaviour) and on the nature of the context, both physical and normative. With setting an important defining feature of behaviour, and with the different constraints created by different types of setting shaping the behaviours reported, any spillover effects between behaviours in different settings seem likely to be weak at best. This has implications for the design of behaviour change interventions, suggesting that they will be most effective when they recognise the specific features of the target behaviours – the equipment involved, the triggers for the performance of the behaviour, and most importantly, the nature of the setting that the behaviour occurs within. Interventions within organisational settings such as offices, this suggests, cannot be expected to result in behaviour change within households (and vice versa), unless other defining features of the target behaviours (equipment, trigger for the behaviour) are very similar. One reservation about these findings does need to be noted, however. The level of performance of each behaviour reported by respondents was very high, resulting in data that was skewed and low levels of variance in the data. This limits the level of variance in the home setting that could be explained using the data from the office setting, and could potentially have masked further effects between settings. The high levels of correlation between behaviours in different settings could be a sign of the existence of such effects. Further research in this area could address some of the limitations of this study, by identifying behaviours with greater variance that may reveal effects not seen in the data presented here. Further research could also compare behaviours in different office buildings within the same organisations rather than between two similar organisations to eliminate any effects originating from differences in organisational culture that have not been identified here. Research that compared the results of behaviour change interventions in different settings would also be able to identify further evidence for the existence, or otherwise, of a spillover effect between contexts. There is also potential for research to usefully focus on how similar behaviours need to be in each setting for influences on the performance of one behaviour to also influence the performance of a behaviour in a different context. The research presented in this paper has highlighted the importance of identifying the defining features that make behaviours similar. Further work is needed to understand how such defining features interact with context and action to create triggers for the performance of behaviours, in order to understand the influence that the type of setting has on the performance of individual energy demand behaviours."],["Theory and correlational research suggest that connecting with nature may facilitate prosocial and environmentally sustainable behaviors. In three studies we test causal direction with experimental manipulations of nature exposure and laboratory analogs of cooperative and sustainable behavior. Participants who watched a nature video harvested more cooperatively and sustainably in a fishing-themed commons dilemma, compared to participants who watched an architectural video (Study 1 and 2) or geometric shapes with an audio podcast about writing (Study 2). The effects were not due to mood, and this was corroborated in Study 3 where pleasantness and nature content were manipulated independently in a 2×2 design. Participants exposed to nature videos responded more cooperatively on a measure of social value orientation and indicated greater willingness to engage in environmentally sustainable behaviors. Collectively, results suggest that exposure to nature may increase cooperation, and, when considering environmental problems as social dilemmas, sustainable intentions and behavior. --------------------------------------------------------------------------------","We clearly face significant environmental challenges (e.g., climate change, pollution, accelerating extinctions). Although the causes and solutions are obviously multifaceted and complex, many have suggested that modern lifestyles contribute to environmental destruction—not only via excessive consumption, but also by disconnecting people from nature. This scholarship often draws on Wilson's (1984) biophilia hypothesis, which posits that humans have an innate need to associate with other living things due to our evolutionary history. We evolved in natural environments and, thus, they still support optimal human functioning (Kellert, 1997). We do not need to accept the specific innate need posited by biophilia to see a gap between humans' evolutionary environments and the current living conditions of people in modern societies. This gap may be a source of suboptimal well-being. Consistent with this idea, living near greenspace predicts higher happiness (White, Alcock, Wheeler, & Depledge, 2013) and longevity (Mitchell & Popham, 2008), and spending time in nature seems to provide a variety of cognitive, mood, and physiological benefits (reviewed by Hartig, Mitchell, de Vries, & Frumkin, 2014 and Selhub & Logan, 2012). Despite nature's apparent benefits, most people spend the majority of their time indoors away from nature (MacKerron & Mourato, 2013). This physical disconnection may also foster a problematic psychological disconnection. That is, when humans do not feel like they are part of larger ecosystems, they may be less inclined to protect the natural environment (Schultz, 2000). Supporting this idea, individual differences in subjective connectedness with nature consistently predict pro-environmental attitudes and behaviors, as well as happiness (Capaldi, Dopko, & Zelenski, 2014; Mayer & Frantz, 2004; Nisbet, Zelenski, & Murphy, 2009; Tam, 2013). Ironically, our threatened natural environments may be critical to fostering the deep concern that would protect them. Although suggestive, past research linking nature with sustainable behavior is mostly correlational, qualitative, or relies on subjective self-reports. In this research we take an experimental approach by manipulating exposure to nature and observing effects on a laboratory analog of sustainable behavior: a fishing-themed commons dilemma (Gifford & Gifford, 2000). Dawes (1980) described environmental problems as social dilemmas with two key features: individuals benefit by behaving selfishly (e.g., over-harvesting resources, polluting) regardless of others' choices, and where all would benefit if everyone cooperated instead of pursuing immediate or narrow self interest (see also Parks, Joireman, & Van Lange, 2013). Said another way, broad participation and cooperation are essential to resolving many environmental problems. We hypothesize that participants exposed to nature will make more cooperative, and thus sustainable, choices. We view cooperative behavior as that which contributes to collective benefits (not necessarily without simultaneous personal benefit), and, in this context, sustaining resources. This prediction is similar to ideas prevalent in environmental psychology—that time in nature and strong subjective connections with nature promote sustainable attitudes (Gifford, 2014). Nonetheless, it departs from most research in the area by suggesting that these effects can be observed over the course of a few minutes in the laboratory. The processes involved in a lifetime of accumulated nature experience may well differ, but we nonetheless draw on the personality-level correlations as part of the rationale for our prediction. Fleeson (2001) has suggested that associations at the trait level often apply at the state level too (e.g., trait extraversion predicts high positive affect and most people experience positive emotions when they behave in extraverted ways). Regarding nature and sustainability, part of the link has been established. Brief exposures to natural settings increase momentary feelings of nature relatedness (Mayer, Frantz, Bruehlman-Senecal, & Dolliver, 2008; Nisbet & Zelenski, 2011; Schultz & Tabanico, 2007). Because trait nature relatedness is strongly associated with sustainable attitudes (Tam, 2013), state nature relatedness, caused by nature exposure, may be too. Research on the short-term consequences of nature exposure also suggests some reasons that nature could promote sustainability, particularly when we think of sustainable behaviors that are also cooperative behaviors. For example, nature exposure is often associated with good moods (Mayer et al., 2008; Nisbet & Zelenski, 2011). Intuitively, and generally consistent with the ‘broaden and build’ view of positive emotions (Fredrickson, 2001), good moods may facilitate cooperative or prosocial behavior, actions that would also be sustainable in resource dilemmas. Research on mood and cooperation, however, suggests that the link may be complex and depend on context (Hertel, Neuhof, Theuer, & Kerr, 2000). Considering another route, Kaplan and Berman (2010) reviewed nature's effects on attention restoration, crime reduction, subjective energy, frustration tolerance, etc., and suggested that they share the common core of improved self-control. Nature may facilitate cooperation in commons dilemmas by improving self-control, thus curtailing temptations to cheat or overharvest. Perhaps even more relevant, Weinstein, Przybylski, and Ryan (2009) manipulated nature exposure with photographs (nature vs. built environments) or plants (present or absent) and found that nature increased participants' intrinsic aspirations and generosity, and decreased extrinsic, materialistic aspirations. That is, nature caused people to report valuing others and prosocial behavior more, and wealth and fame less. This extended to actual behavior in the ‘trust game’; participants exposed to nature gave more actual money to another person that they could have kept for themselves without negative consequence. These effects were mediated by feelings of (state) nature relatedness and autonomy, and were strongest among participants who felt most immersed in the nature. Similar effects may not require deep immersion, however. Mazar and Zhong (2011) found that participants merely exposed to green products in a consumer study gave away more money than participants who viewed more conventional products. Such effects contrast with findings that money primes make people more self-sufficient and less prosocial (Vohs, Mead, & Goode, 2006); nature may function oppositely (Nisbet & Zelenski, 2009). Although suggestive, none of this research has examined sustainability attitudes or behaviors. Commons dilemmas provide a link between nature effects and sustainability because they channel cooperation, trust, and prosocial motivations into sustainable behaviors. To be clear, cooperative behavior is not always sustainable. Humans often cooperate in ways that ultimately threaten natural environments; most current environmental crises result from economic activity that requires some cooperation among individuals and groups. Moreover, not every sustainable behavior requires cooperative intentions. The environmental benefits may be diffuse (e.g., a reduction in greenhouse gasses benefits all), but the intentions may be completely local and selfish (e.g., thinking, ‘a tree would look nice in my backyard’). Said another way, altruism is not required for cooperation or sustainable behaviors. Our primary focus is the confluence of cooperation and sustainability. Environmental problems are classic examples of commons dilemmas, and, thus, research on commons dilemmas has much potential to inform environmentally sustainable behavior and decision making. We have focused on an environmentally themed commons dilemma because it allows us to bridge different literatures in suggesting nature exposure as a potential aid to cooperative or sustainable behavior. We extend the theory and mostly correlational research that suggests a strong link between connecting with nature and sustainability by adding experimental manipulations that speak to causal direction more directly. We extend experimental studies' suggestive hints about nature's effects on mood, self-control, prosocial motivation, and trust by testing them in contexts more relevant to sustainability. In sum, there are theoretical and empirical reasons to suspect that exposure to natural (vs. built) environments may promote cooperative, sustainable behavior. To test these ideas, we conducted three studies. In the first, we randomly assigned participants to view videos of almost exclusively natural or built environments.","were later asked to ‘play a fishing game’, an iterative, fishing-themed commons dilemma where they were paid for each fish harvested. We also included measures of mood, state nature relatedness, and state trust (as possible mediators), and trait measures of nature relatedness and trust as exploratory predictors or moderators. Study 2 reports a close replication. Study 3 provides a conceptual replication and extension; it begins to disentangle cooperation from sustainability by measuring these outcomes independently. We report how we determined our sample size, all data exclusions, all manipulations, and all measures in all studies (Simmons, Nelson, & Simonsohn, 2012). Participants Undergraduate students (n = 111) were recruited for a study titled ‘Personality and Media’ via our department online subject pool system. Our goal was n = 120 for an exploratory study, and we collected data to the end of a semester. The sample was 70.3% female with a mean age of 20.81 (SD = 3.10). Participants received course credit as compensation. They were also paid based on fishing performance, but learned this only after arriving for the study.","Videos. To manipulate exposure to natural vs. built environments, participants viewed one of two 12-min videos that included educational narration and background music. The nature video excerpted BBC's Planet Earth series, beginning in tundra forest with images of trees and animals. It then proceeded to areas around the world and showcased the plants and animals native to those areas, ending in a jungle. We chose this particular excerpt because there are no mentions of marine life or fish, as well as to avoid explicit appeals for conservation. Planet Earth is easily described as a ‘nature documentary.’ It represents nature as environments relatively untouched by humans (lack of buildings), with abundant and beautiful fauna and flora (see Vining, Merrick, & Price, 2008). The built video excerpted Landmark Media's Walks with an Architect series, and featured in-depth looks at buildings, their history, and locations in New York City. The buildings in New York City arguably include some of the world's finest architecture, and we chose this video to contrast with the nature in Planet Earth while still conveying impressive content. The fact that nearly every part of the video contains human- built spaces makes it antithetical to common conceptions of nature (Vining et al., 2008). Thus, these excerpts were relatively ‘pure’ representations of nature and non-nature (cf., a cultivated garden). No pilot data were collected on the effects of these videos, yet some of their impacts (e.g., on mood and subjective connection with nature) are described by this study. Participants viewed videos on desktop computers with 17-inch screens and received sound via headphones. Commons dilemma. To assess cooperation, participants engaged in a fishing-themed commons dilemma, specifically FISH 3.1 (Gifford & Gifford, 2000; see http://web.uvic.ca/∼rgifford/fish/). In this microworld simulation, participants make choices about how many fish to harvest across multiple ‘seasons’. In our application, participants harvested from an ocean shared by three other fishers who were, unbeknownst to participants, actually simulated. Between each season, fish regenerated at a rate of 1.5, and fishing continued until fish were gone or 15 rounds had passed, but participants were not informed of this limit. The ocean began with 50 fish, and participants were paid $.10 per fish harvested. A fee of $.05 was charged to go to, and return from, sea, thus making it necessary to catch at least two fish to profit in any one season (though participants could stay ‘on shore’ for free). Information about the number of fish harvested, fish remaining in the ocean, profits, and other fishers' catches were all displayed continuously on the screen. Simulated fishers were programmed to behave relatively cooperatively (an average of .5 on the 0 to 1 ‘greed’ setting). FISH yields measures of fish harvested, total seasons, profits, as well as calculated indexes of efficiency and restraint (see Gifford & Gifford, 2000). The indexes were calculated in each season, and then averaged across seasons so each participant received a single score. Restraint tracks the raw number of fish harvested while accounting for group size (but not regeneration rate); higher numbers (between 0 and 1) are necessary for a sustainable resource. Efficiency tracks the number of fish taken relative to the current size of the fish population and the regeneration rate. Scores above 1 indicate ‘unnecessary’ efficiency, i.e., taking less than would regenerate, whereas scores below 1 indicate that more fish were taken than could be regenerated in the next season, thus shrinking the population. State Scales. To assess state nature relatedness, participants completed the single-item Inclusion of Nature in Self measure (INS; Schultz, 2002). Participants were presented with seven pairs of circles labeled self and nature that differ in the degree of physical overlap. Participants choose the pair that represented, “… your relationship with the natural environment at this point in time. How interconnected are you with nature right now?” As distractors, participants also rated pairs of circles labeled “self” and “people, family, friends, community, an urban center, and to all humanity.” A mood questionnaire included the Positive and Negative Affect Scales (PANAS; Watson, Clark, & Tellegen, 1988). Participants rated adjectives on a Likert scale of 1 (very slightly or not at all) to 5 (extremely) to describe their feelings “in the moment”. Because the 10-item positive affect (α = .85) and negative affect (α = .81) scales assess only high arousal affects (e.g., enthusiastic, proud, interested, and afraid, nervous, distressed, respectively), we added adjectives that were lower in arousal, yet still pleasant, and intuitively associated with nature experiences: fascination, peaceful, content, in awe, curious, and relaxed. Although this scale is admittedly ad hoc, it may capture aspects of the ‘soft fascination’ described by Kaplan (1995). Nature is also a prototypical trigger of awe (Keltner & Haidt, 2003). This pleasant affect scale had good internal consistency (α = .79). Following a practice round of the FISH simulation, participants also completed a 3-item ad hoc measure of trust in other fishers (α = .79). They rated items like, “I expect that my group members will be trustworthy” on a 5-point scale of agreement. Trait Scales. Participants completed the 5-item Faith in People Scale (trust; Rosenberg, 1957), and the 6-item Short Nature Relatedness Scale (Nisbet & Zelenski, 2013). These were embedded in the 44-item Big Five Inventory (John & Srivastava, 1999), a broad measure of personality traits, to avoid suggesting our interest in specific individual differences. Procedure Participants arrived at the lab and were ushered to a small testing room. One or two participants were tested simultaneously, but the lab layout suggested the possibility of more. After informed consent and a brief description of the study (framed as being about personality and perceptions of media), participants completed the trait measures. They were then randomly assigned to watch the Planet Earth or Walks with an Architect video. Following the video, they completed a brief questionnaire about their liking of the video (cover story), the INS, and mood questionnaire. Next, participants received detailed written and verbal instructions about FISH, including a three-season practice session. They then completed the state trust measure and began the actual FISH simulation of up to 15 seasons. Following FISH, participants completed a brief questionnaire about their impressions of the ‘fishing game’ (cover story), a demographics questionnaire, and a questionnaire that probed for suspicion.1 Finally, participants were debriefed and paid according to their FISH performance.","Our primary hypothesis was that exposure to natural (vs. built) environments, operationalized as Planet Earth (vs. architecture) videos, would produce higher rates of cooperative and sustainable behavior. We tested this hypothesis across various FISH indicators. Many had substantial skew, kurtosis, and outliers, so we present comparisons in two ways: as t-tests with 15 outliers2 excluded from all tests, and nonparametric Mann–Whitney's U with all participants (see Table 1). Results generally supported hypotheses. Participants who watched the Planet Earth video harvested significantly fewer fish per season and had commons pools that lasted more seasons than those who watched the architecture video. These differences are mirrored in the indices of restraint and efficiency with the Planet Earth condition showing more of both. Participants who viewed the architecture video made significantly more money, suggesting that this scenario favored a short-term, unsustainable strategy (i.e., large harvests across a few rounds). By season 15, 49.09% of the architecture condition's oceans went extinct, compared to 28.57% in the Planet Earth condition, χ2 (1, N = 111) = 4.92, p = .03. We anticipated that the Planet Earth video might produce better moods and more feelings of nature relatedness and trust compared to the architecture video. These suspicions were only partially confirmed (see Table 2). Participants who saw the Planet Earth video reported significantly more pleasant affects and less negative affect, but groups did not differ significantly on (high arousal) positive affect, state nature relatedness, or trust. Given that the Planet Earth video produced somewhat more pleasant states, we conducted exploratory bootstrapping mediation analyses (Preacher & Hayes, 2008) with all FISH indicators as dependent variables. In no case was negative affect or pleasant affect a significant mediator. (Said another way, controlling for mood made no difference.) In other exploratory analyses, we tested whether relevant personality variables (nature relatedness and trust) predicted FISH behavior or moderated the effect of our experimental manipulation, but results were almost uniformly not statistically significant.","The results of Study 1 provide preliminary evidence that exposure to nature can promote cooperative or sustainable decisions. Participants who viewed the Earth video harvested more sustainably (i.e., fewer fish per season and extinctions) than participants who watched the architecture video. Although the mechanisms of this effect are not clear, data suggest that mood, trust, and subjective feelings of nature relatedness do not account for differences.","In Study 2 we attempted a close replication with minor alterations. First, we adjusted the FISH parameters slightly to see whether results would extend to a context that did not favor a short-term strategy (in terms of profits). That is, we reduced the cost of going out to sea so each fish was more profitable, especially in small catches. We also increased the maximum number of seasons, giving fishers more time to benefit from a sustainable, long-term strategy—small profits in each season add up when there are more seasons. In addition, we added another more neutral control condition to confirm that the action of the effect was not entirely due to the architecture video. Finally, we omitted some of the exploratory measures from Study 1.","With procedures identical to Study 1, 121 students (71% female) were recruited and randomly assigned to either the nature, built, or neutral conditions. Our goal was 40 subjects per condition, allowing good power to detect the effect sizes observed in Study 1. All received course credit (and money) as compensation.","The study was identical to Study 1 with the following exceptions: A new control condition consisted of watching the iTunes visualizer (full screen) while listening to the Grammar Girl podcast, The Rules of Story. In FISH, the cost of going out to, and returning from, sea was reduced to $.02, and the maximum number of seasons was increased to 25. Trait nature relatedness and both trust measures were omitted. A Big Five personality measure remained (recall the cover story about personality and media), but was changed to the 40-item ‘mini markers’ (Saucier, 1994) as this was helpful to an unrelated project.","As in Study 1, our primary hypothesis was that the Planet Earth video would produce more cooperative and sustainable FISH decisions compared to the architecture video and grammar podcast. Skew, kurtosis, and outliers were again concerns with FISH variables, so we conducted parametric analyses with three outliers excluded and nonparametric analyses with no exclusions. (See Supplementary Materials for more Study 2 analysis details.) When comparing omnibus tests across the three conditions, we found differences that were often marginally significant (e.g., ANOVA ps from .06 to .14 and Kruskal-Wallace ps from .03 to .08), with the exception of profits where there were somewhat smaller differences (ps = .27 and .15). Across indicators, the architecture and grammar conditions were most similar (comparisons produced ps all > .26), and both tended to differ from the Earth condition. Table 3 provides means, SD, and an indication of where differences between two conditions are statistically significant. Unless one is rigid about the p < .05 criterion, results replicate Study 1's finding that Planet Earth produces more sustainable fishing behavior, though effect sizes are somewhat smaller. Also, as anticipated, adjustments to FISH parameters resulted in better outcomes for a sustainable strategy; the Planet Earth condition now harvested more fish than architecture or grammar conditions. Mood differences were also similar to Study 1 with nature somewhat more pleasant; the new grammar control produced moods similar to the architecture condition. Exploratory mediation analyses again failed to provide any evidence that mood was responsible for the effects of videos on FISH outcomes.","To build on two studies with very similar methods and findings, Study 3 addressed some issues of generalizability (e.g., going beyond the particular Planet Earth clip) and took a stronger approach to ruling out mood as a possible (yet increasingly unlikely) explanation for nature's effect on cooperation. To accomplish this, we created a new video manipulation that independently varied pleasantness and nature content with a 2 × 2 design. In addition, we replaced FISH, as a measure of cooperation, with social value orientation, a conceptually similar and empirically related measure (Balliet, Parks, & Joireman, 2009) that includes no environmental connotations, followed by some more explicit questions about sustainable behaviors.","Undergraduate students were recruited via our department online portal until 250 had completed the study. To ensure valid responses, analyses excluded participants who finished in less than 10 min (median time was 21 min) and who did not comply with two requests to leave items blank; thus, n = 228.","Videos. Drawing on videos from YouTube, we created 2 min clips designed to independently manipulate pleasantness and natural (vs. built) context. We used criteria similar to Study 1 to determine nature and built contexts. To enhance the valence manipulation, we replaced original sound with upbeat, pleasant music or minimalist, foreboding music. Visual content included: Las Vegas strip (pleasant, built) which included images and video from this street at night with neon signs and edited in a fast, ‘upbeat’ way; old-growth forest (pleasant, nature) which showed a time lapse clip of undergrowth sprouting and aerial shots of large, mature trees; abandoned, decrepit house (unpleasant built) which slowly toured a clearly abandoned and distressed building with minimal furnishings; a flood (unpleasant nature 1) which showed expansive and fast moving water in a clearly flooded landscape that included some occasional damaged houses; a wolf pack (unpleasant nature 2) which showed wolves antagonizing a bison and bear, and hunting and killing an elk. The ‘extra’ unpleasant nature video was included as a comparison because the flood video included brief built elements, i.e., houses being washed away. They are sometimes combined in results for efficiency, but yield similar results individually. Questionnaires. Cooperation was assessed with the social value orientation slider measure (SVO; Murphy, Ackermann, & Handgraaf, 2011). Across six items, participants allocated points (imagined as money) to themselves and a hypothetical other via a forced choice of nine alternatives that vary benefits to self vs. other. Although typically conceptualized as an individual difference measure, instructions do not imply anything trait-like, and similar measures are sensitive to context (Bekkers, 2004). High scores indicate more pro-social choices. SVO scores were missing for two otherwise complete cases. Willingness to behave sustainably was assessed with a 30-item questionnaire developed by Ferguson, Branscombe, and Reynolds (2011). Participants indicated their willingness to engage in a variety of sustainable behaviors, such as, “Reduce the amount of warm and hot water used” on a 7-point Likert scale ranging from 1 (extremely unwilling) to 7 (extremely willing). Items cover transportation, energy and water use, social advocacy, tax support, and regulation support. Scores reflect mean ratings, α = .94. Similar to Study 1, mood was assessed with the 20-item PANAS (PA and NA αs = .91) interspersed with 6 vitality items; state nature relatedness was assessed with the INS (plus family and society as foils). The full 21-item trait nature relatedness scale (α = .90) was embedded in a 100-item IPIP (ipip.ori.org) Big Five personality questionnaire, and, as a validity check, “Please leave this item blank” was inserted twice. Procedure Participants were directed to a Qualtrics webpage that administered the study. Following consent, they completed the personality measures and were then randomly assigned to view one of five videos. Following the video, they completed measures of mood, INS, cooperation, and willingness.","Manipulation checks suggested that videos altered pleasantness and nature exposure relatively independently. In 2 × 2 (valence by environment) ANOVAs, we found significant effects of valence on positive affect, F(1, 224) = 12.61, p < .001, ηp2 = .053, and negative affect, F(1, 224) = 33.70, p < .001, ηp2 = .13 (see Table 4 for means). Corresponding effects of environment were null for positive affect, F(1, 224) = .01, p = .92, ηp2 < .01, and marginally significant for negative affect, with built videos producing slightly higher ratings, F(1, 224) = 3.49, p = .06, ηp2 = .015. We also observed a marginally significant effect of environment on INS, F(1, 224) = 3.46, p = .06, ηp2 = .015, with nature videos producing higher levels of subjective nature relatedness. Thus, manipulations functioned largely as expected. Our primary hypothesis was that nature videos (forest, wolves, and flood) would produce more cooperative choices than built videos (Las Vegas and old house). A 2 × 2 ANOVA with SVO as the dependent variable revealed that environment had a significant effect, F(1, 222) = 4.43, p = .04, ηp2 = .020, with nature videos producing more pro-social responses (and no effect of valence or interaction).3 We also tested whether this extended to willingness to engage in environmentally sustainable behaviors, and found a similar effect of environment, F(1, 224) = 4.51, p = .04, ηp2 = .020. However this was qualified by an interaction with valence, F(1, 224) = 3.96, p = .05, ηp2 = .017.4 Essentially, the built, pleasant, Las Vegas video produced particularly low levels of willingness compared to the other groups, which were similar (see Table 4). Exploratory bootstrapping analyses suggested that state nature relatedness (INS) mediated the effect of videos (nature vs. built) on SVO (95% CI: .002, 1.45) and willingness (.01, .25; see Fig. 1), though ‘significance’ depended somewhat on using this approach (cf. Baron & Kenny) and combining the negative nature videos (see Supplement). Thus, state nature relatedness may account for nature's effects on cooperation and sustainability, but the evidence is somewhat inconsistent. There was also a significant correlation between SVO and willingness in this study (r = .28). Thus, we tested the possibility that SVO mediated the effect of videos on willingness. Bootstrapping indicated a possible mediation path when negative nature conditions were combined (95% CI: .002, .16; see Fig. 1), though ‘significance’ again depended somewhat on this particular approach (see Supplement). The films' effect on sustainability may be due to shifts in cooperation, but evidence is again somewhat inconsistent. Finally, we also explored the role of trait nature relatedness and found that it correlated significantly with SVO (r = .28), INS (r = .61), willingness (r = .64), and positive affect (r = .23), but typically did not interact with manipulations in predicting these things. In sum, this study provides a conceptual replication supporting the idea that exposure to nature (in this case forest, flood, and wolf videos) promotes cooperative decisions, even absent an environmental context. These effects did not depend on nature's pleasantness, and also extended to explicit statements about environmental behavior under some circumstances.","After finding that short walks in nature produced both pleasant moods and feelings of nature relatedness, Nisbet and Zelenski (2011) suggested nature as a ‘happy path to sustainability’. This research takes the critical next step by explicitly testing the link between nature exposure and sustainable behaviors, rather than inferring this from the trait–level association between nature relatedness and sustainable behavior. Across three studies, we found consistent evidence for the idea that exposure to nature (videos) can produce cooperative behavior, which was also sustainable behavior in the context of commons dilemmas. Viewing environmental problems as social dilemmas underscores the link between cooperation and sustainability. Environmental issues are classic examples of social dilemmas, and cooperation is essential in solving them. In our lab analog, exposure to nature increased sustainable fishing and helped determine whether or not fish stocks collapsed. These effects appear to be due to nature per se as both built and neutral control comparisons produced similar results. Moreover, although pleasant moods are typically associated with nature, they did not explain its effect on cooperation. Results held using statistical mood controls, and when we directly manipulated pleasant and unpleasant representations of nature. The mediation results for state nature relatedness were inconsistent across studies (significant only in Study 3), and not robust enough to provide strong evidence for this as the only path. Although it remains plausible that time in nature fosters connectedness and sustainable attitudes over the long term (Schultz, 2000), different processes may explain nature's effect on cooperative or sustainable behavior in the moment. That is, repeated experiences in nature, especially pleasant ones, may foster a more stable sense of nature relatedness, and then a desire to protect nature, habits of spending more time in nature, associating with individuals and groups that value nature and sustainable practices, etc. (e.g., Bragg, 1996; Kals, Schumacher, & Montada, 1999; Mayer & Frantz, 2004; Nisbet et al., 2009; Orr, 1993). Such connections likely develop over time. A single exposure to nature will probably not permanently change a person's attitudes or behavior, and it is entirely possible that momentary feelings of connectedness with nature do not cause sustainable choices in the same way that a more stable sense of a nature related self does. Previous correlational research on these topics likely speaks more to the stable contents of personality, whereas the studies we report here deal more with shifts in processing. Momentary nature exposure may produce a set of changes in emotion and cognition that are temporary and relatively distinct from personality-level processes. Our studies found little support for improved mood or trust as reasons that nature influences momentary cooperation or sustainability. Other research has shown that nature exposure can help restore attention or self-control resources (see Kaplan & Berman, 2010). Perhaps related to this, two recent articles have found that nature reduces temporal discounting (Berry, Sweeney, Morath, Odum, & Jordan, 2014; van der Wal, Schade, Krabbendam, & van Vugt, 2013). That is, nature appeared to shift people's preferences from immediate gratification to larger but more distant payoffs. Nature's ability to improve self-control in this way seems very consistent with the more sustainable strategies we observed in the FISH studies (though our studies did not guarantee a higher payoff with a long-term strategy). Additional research is needed, however, to determine which of these, or other, changes are more fundamental to nature exposure. Said another way, the improvements in sustainable behavior that we observed in these studies may be due to (mediated by) more basic shifts in delay of gratification. On the other hand, the measures in Study 3 suggest that there is more than temporal discounting at stake. The SVO measure is about immediate allocations and seems to measure prosociality more than self-control (see also Guéguen & Stefan, 2014). In addition, Study 3 included self-reported willingness to perform sustainable behaviors. Nature seemed to increase these, but the effect was qualified by an unexpected interaction. In addition, the effect of nature on willingness appeared to be mediated by SVO (prosociality or cooperation). Thus, we suggest a cautious interpretation of possible momentary changes in explicit environmental attitudes. Said another way, it is possible that nature's ability to promote sustainable behavior ‘in the moment’ applies primarily to cooperative contexts or commons dilemmas (or possibly even an explicit understanding of the choice as a social dilemma). Nature may shift cooperation more than sustainability per se when the two are dissociated. As with all meditation claims, determining the exact mechanism(s) of nature's effect on sustainability and cooperation will take considerable (more) research (see Bullock, Green, & Ha, 2010). This research bridged gaps between experimental studies of nature exposure that have not considered sustainable behavior, and correlational studies that suggest a link between nature relatedness and sustainable attitudes. Like most experimental studies, our methods trade apparent external validity for methodological control. Although we argue that they maintain psychological realism, testing generalizability is clearly an issue for future research. Our participants were all Canadian students exposed to representations of (not actual) nature, ‘fishing’ in a simulation for relatively low stakes and in relatively cooperative contexts, or completing questionnaires in Study 3. In addition, studying the effects and active ingredients of ‘nature’ is tricky given its nearly infinite exemplars and the absence of clear or direct controls. We began to deal with this issue of stimulus sampling (Wells & Windschitl, 1999) by comparing four nature videos with four comparison conditions across our studies (i.e., Planet Earth, forest, wolves, and flood vs. architecture, grammar podcast, Las Vegas, and decrepit house). One could develop hypotheses about why, for example, Las Vegas or wolves might have particular effects beyond the natural-built distinction (indeed, we believe they must in some ways), but given a 4 vs. 4 comparison, the more parsimonious explanation seems to be an effect of nature—this is the consistent theme across all stimuli. Still, despite some breadth across these studies, this research begs for future conceptual replications, falsification attempts, and search for boundary conditions. Collectively our results suggest that nature may increase cooperation and sustainable choices, yet additional exemplars should be tested in future work before considering the matter settled. As one specific and potentially important example, our nature videos may have primed the idea of conservation, thus creating demand. Although possible, some design choices argue against this particular problem: 1) we chose nature videos that excluded marine life and pro-environmental messages that might make the link obvious; 2) we paid our research participants for each fish harvested, giving them an incentive inconsistent with pleasing the experimenter; 3) we used a between-subjects design so that the video content could not be compared across conditions or become an obvious cue; 4) we crafted a plausible cover story about the study's purpose (i.e., about personality and reactions to media) that included distracting measures to support that story. In addition, Study 1 included a funneled debriefing for suspicion, and even with a liberal criterion, less than 10% of participants could identify the study's purpose, and removing these participants' data had little effect on the results. Finally, there is nothing ‘environmental’ about the SVO measure in Study 3 where nature again produced cooperative responses. Thus, narrow priming or demand (i.e., nature videos to conservation) seems like an incomplete explanation for our results. It seems more plausible that nature could prime cooperation broadly, similar to, but opposite, the self-sufficiency primed by money (Vohs et al., 2006). With these caveats in mind, we turn to more speculative implications. As a methodological note for researchers, nature images or videos are often used as ‘neutral’ stimuli in control conditions. Our results suggest that nature can produce psychologically meaningful effects, and, thus, its ‘neutrality’ may be unwisely assumed in some research contexts. For example, a recent registered trial of ‘brain training’ failed to find much benefit for the cognitive exercises, but participants in a control condition, which consisted of watching short nature videos, reported significant improvements in psychological well-being and decreased stress over time (Borness, Proudfoot, Crawford, & Valenzuela, 2013). Also, three of nine studies in the money priming article just mentioned included control conditions that seem like nature (e.g., images of fish, a flower poster, and a seascape poster; Vohs et al., 2006). Appropriate comparison conditions depend on context, but nature may be appropriate less often than many assume. Outside the lab, conservation activists have long used nature imagery in persuasive appeals, but recent messaging around climate change often prefers economic or security arguments. Given the effects of priming money vs. nature, nature imagery may produce more persuasive appeals or better reminders to behave sustainably–environmental problems are social dilemmas, and cooperation is key to sustainable solutions. Future empirical work could compare the relative efficacy of such appeals more directly. Our research also contributes to a growing body of work that suggests nature's benefits extend beyond individual well-being, for example, to prosocial aspirations (Weinstein et al., 2009) and behavior (Guéguen & Stefan, 2014) and reduced aggression and crime (Kuo & Sullivan, 2001). Such findings, combined with nature's salubrious effect on socioeconomic health disparities (Mitchell & Popham, 2008), suggest that societies might consider investing more in nature. Similar to arguments for public education, providing nature access to all citizens could possibly provide a net social or financial benefit."],["Objectives: This experiment investigated the extent to which independent action observation, independent motor imagery and combined action observation and motor imagery of a sport-related motor skill elicited activity within the motor system. Design and method: Eighteen, right-handed, male participants engaged in four conditions following a repeated measures design. The experimental conditions involved action observation, motor imagery, or combined action observation and motor imagery of a basketball free throw, whilst the control condition involved observation of a static image of a basketball player holding a basketball. In all conditions, single pulse transcranial magnetic stimulation was delivered to the forearm representation of the left motor cortex. The amplitude of the resulting motor evoked potentials were recorded from the flexor carpi ulnaris and extensor carpi ulnaris muscles of the right forearm and used as a marker of corticospinal excitability. Results: Corticospinal excitability was facilitated significantly by combined action observation and motor imagery of the basketball free throw, in comparison to both the action observation and control conditions. In contrast, the independent use of either action observation or motor imagery did not facilitate corticospinal excitability compared to the control condition. Conclusions: The findings have implications for the design and delivery of action observation and motor imagery interventions in sport. As corticospinal excitability was facilitated by the use of combined action observation and motor imagery, researchers should seek to establish the efficacy of implementing combined action observation and motor imagery interventions for improving motor skill performance and learning in applied sporting settings. --------------------------------------------------------------------------------","An a priori power analysis was conducted using G*Power software to determine the number of participants required for this experiment. The power analysis was based on the data reported by Wright et al. (2014), who compared differences in corticospinal excitability between AO, MI and AOMI of an index finger abduction-adduction movement against a static hand control condition, and obtained effect sizes ranging from d = 0.68–1.78. Based on the lowest effect size (d = 0.68) and α set at 0.05, the power analysis indicated that in order to find significant differences between the experimental conditions and the control condition, at least 15 participants would be required to achieve a power of 80%. In order to allow for possible participant dropout or other loss of data, 18 participants were recruited to participate in this experiment. All 18 participants were male and aged between 19 and 32 years (mean age 22.61 ± 3.45 years). They were all right-handed, as assessed by the Edinburgh Handedness Inventory (Oldfield, 1971), and all had at least moderately good imagery ability, as assessed by the Vividness of Movement Imagery Questionnaire-2 (Roberts, Callow, Hardy, Markland, & Bringer, 2008; See Table 1). All participants were novice basketball players, in that they had some experience of playing the sport in practical physical education lessons at school, but had never played competitively. Furthermore, none of the participants were susceptible to possible adverse side-effects of transcranial magnetic stimulation, as assessed by the TMS Adult Safety Screen (Keel, Smith, & Wassermann, 2001). All participants provided written informed consent to take part in the experiment, which had been granted ethical approval by the University Ethics Committee at the host institution. Electromyography and transcranial magnetic stimulation procedure ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ EMG was recorded throughout the experiment using a Delsys Bagnoli 2-channel EMG system (Delsys Inc, Boston, MA). Prior to electrode attachment, participants were asked to repeatedly flex and extend their right wrist, mimicking the action of shooting a basketball free throw, whilst the experimenter felt the participants' right forearm to identify the flexor carpi ulnaris (FCU) and extensor carpi ulnaris (ECU) muscles. Once identified, the sites were cleaned using alcohol wipes and surface EMG electrodes were attached over the belly of both muscles, and a reference electrode was placed on the olecranon of the ulna bone. Recordings were taken from the FCU and ECU muscles as they are both active when flicking the wrist to release the ball from the hand during the execution of a basketball free throw. The EMG signal was recorded using Spike 2 (version 6.18) software, with a sampling rate of 2 kHz, bandwidth of 20–450 kHz, 92 dB common mode rejection ratio and >1015 Ω input impedance, received by a Micro 1401–3 analogue-to- digital converter (Cambridge Electronic Design, Cambridge, UK). Single-pulse TMS was delivered to the left primary motor cortex using a figure-of-eight shaped coil (two 70 mm diameter loops), orientated at a 45° angle to the central line between the nasion and inion landmarks of the cranium (Brasil-Neto et al., 1992), and connected to a Magstim 2002 magnetic stimulator (Magstim, Whitland, Dyfed, UK). The optimal scalp position (OSP) was identified as the scalp location that produced MEPs of largest amplitude in both muscles, using a stimulation intensity of 60% maximum stimulator output (e.g., Clark et al., 2004; Williams et al., 2012; Wright et al., 2014). Once identified, the OSP was marked on a tightly fitting polyester cap worn by the participants by drawing around the coil with a marker pen. The coil was held fixed against the OSP using a mechanical arm, and accuracy of coil placement was ensured throughout the experiment by checking the coil position frequently in relation to the marking and adjusting the positioning if necessary. Resting motor threshold (RMT) was then determined for each participant by gradually reducing or increasing the stimulation intensity, until the minimum intensity capable of producing MEPs with peak-to-peak amplitudes in excess of 50 μV in five out of 10 trials was identified (Rossini et al., 1994, 2015). This stimulation intensity, plus 1% of the maximum stimulator output was identified as the RMT (see Rossini et al., 2015 for guidelines on TMS procedures). Based on the recommendations of Loporto, Holmes, Wright, and McAllister (2013) the stimulation intensity for the experiment was set at 110% of each participant's RMT to reduce the likelihood of direct wave stimulation. Modal values for the location of the OSP, and mean values for the RMT and experimental stimulation intensity can be found in Table 1. Experimental procedure ~~~~~~~~~~~~~~~~~~~~~~ Participants were seated at a desk in front of a 32-inch Samsung flat-screen TV, positioned at eye-level at a distance of 90 cm. Their head was placed comfortably in a custom-built head-and-chin rest, and their hands and forearms rested in a pronated position on the desktop. The lighting in the room was dimmed and blackout curtains were drawn along either side of the desk to eliminate any potentially distracting visual stimuli in the surrounding area. Whilst seated in this position, participants took part in four experimental conditions within a single testing session on the same day. As shown in Fig. 1, the four conditions were termed: static observation (control), action observation (AO), motor imagery (MI) and combined action observation and motor imagery (AOMI). Each condition required participants to complete a block of 30 repetitions of a 10-s duration video. One stimulation from the TMS device was delivered per video, resulting in a total of 30 stimulations per condition. Thirty stimulations per condition were administered as this is recommended as a sufficient number of stimulations to provide a reliable measure of corticospinal excitability (Cuypers, Thijs, & Meesen, 2014; Goldsworthy, Hordacre, & Ridding, 2016). Each stimulation was delivered 4640 ms after the onset of each video, as this corresponded with the point at which the model flicked their wrist to release the ball in the action observation video. There was a 3 s rest period between each video, resulting in a 13 s inter-stimulus interval between trials. This duration inter-stimulus interval is consistent with TMS safety guidelines and is sufficient time for the effects of previous stimulations to have subsided (Chen et al., 1997). The total duration of each experimental condition was 6 min 30 s. Following each condition participants were given a rest period of approximately 3 min before starting the next condition. This duration rest period between conditions was appropriate as MEP amplitudes return the baseline levels after 1 min (Baldi, Perretti, Sannino, Marcantonio, & Santoro, 2002), and it is also consistent with previous TMS experiments exploring AO or MI processes (e.g., Loporto, McAllister, Edwards, Wright, & Holmes, 2012; Wright et al., 2014). The entire experiment, including participant familiarization, completion of consent forms and questionnaires, EMG preparation, OSP and RMT procedures, experimental conditions and participant debriefing, lasted approximately 90 min. Static observation (control) condition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the static observation condition, participants were shown a silent video, filmed from a third-person visual perspective (see Fig. 1), depicting a male basketball player standing still on a basketball free throw line and holding a basketball. Participants were instructed to observe the videos and were reminded of this instruction verbally every 10 trials. Action observation (AO) condition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the AO condition, participants were shown a video of the same male basketball player shooting a successful basketball free throw. In this video, filmed from the same third- person visual perspective as the static observation condition, the model bounced the basketball twice before shooting a right-handed free throw that went straight through the hoop without hitting the rim or backboard. The sounds of the ball being bounced twice during the model's pre-performance routine, the ‘swish’ of the ball going through the net, and the ball bouncing after the shot landed were all audible in the video. Participants were instructed to observe the videos and were reminded of this instruction verbally every 10 trials. Motor imagery (MI) condition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the MI condition, participants were shown a video of a black screen, but heard the same audio recording as in the AO condition. Participants were instructed to actively imagine themselves shooting a successful basketball free throw in time with the audio recording. No specific instructions were provided regarding which perspective participants should image from, but they were instructed to focus specifically on imagining the feelings and sensations associated with flicking the wrist as they released the ball. Kinesthetic imagery instructions were emphasized explicitly as these have been shown to facilitate corticospinal excitability to a greater extent than visual imagery (Stinear, Byblow, Steyvers, Levin, & Swinnen, 2006). Participants were reminded of this instruction verbally every 10 trials. In addition, prior to beginning this condition, participants were asked to keep their eyes open during their imagery to maintain consistency across conditions. They were also reminded that they could refer to their mimicking of the wrist flick action during the EMG preparation procedures to recall the kinesthetic sensations associated with executing the movement. Combined action observation and motor imagery (AOMI) condition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the AOMI condition, participants were presented with the same visual and auditory stimuli they had seen in the AO condition, but were instructed to actively imagine themselves shooting a successful basketball free throw in time with the video. As in the MI condition, participants were instructed to focus specifically on imagining the feelings and sensations associated with flicking the wrist as they released the ball. Participants were reminded of this instruction verbally every 10 trials. Order of conditions ~~~~~~~~~~~~~~~~~~~ Experimental conditions were always presented in the following fixed order: static observation, AO, MI, AOMI. The decision to utilize a fixed order of conditions was taken for several reasons. The static observation condition was always presented first to acquire a baseline MEP amplitude value before any action observation or motor imagery had taken place. This was important to reduce the chance of participants engaging in spontaneous or deliberate motor imagery during control trials, by virtue of having already being exposed to imagery instructions or the action observation stimuli. Similarly, the AO condition was presented next, before any imagery instructions were provided, in an attempt to reduce the likelihood of knowledge of prior imagery instructions eliciting forms of imagery in this condition. The MI condition was then presented as the third condition, after AO, as it was deemed necessary to have first exposed participants to basketball free throw stimuli to allow them to image the action, due to their novice status. The AOMI condition was then presented last as it was important for participants to have experienced both AO and MI conditions independently to allow them to be combined effectively. This approach to ordering conditions is consistent with previous TMS research using a similar experimental design (Wright et al., 2014, 2016).","An increase in EMG activity at the time of stimulation can result in an increase in the amplitude of the subsequent MEP (Devanne, Lavoie, & Capaday, 1997; Hess, Mills, & Murray, 1987). As such, the amplitude of each participant's EMG activity in the 200 ms prior to each stimulation was measured in both muscles. Any trials in which this value was greater than 2.5 SD above the mean of that participant's baseline EMG for that muscle were removed from the analysis (Loporto et al., 2013; Wright et al., 2014). This resulted in a mean of 2.35 (±0.86) trials being removed per participant from each muscle in each condition, and so no participants were removed from the experiment due to excessive loss of data. The peak-to-peak amplitude of MEPs in the remaining trials was then measured. Due to large intra- and inter-participant variability in MEP amplitude, these data were normalized using the z-score transformation commonly used in TMS action observation and imagery research (e.g., Aglioti, Cesari, Romani, & Urgesi, 2008; Fadiga et al., 1995; Wright et al., 2014). The normalized MEP amplitude data were then analyzed with a 2 (muscle) x 4 (condition) repeated measures analysis of variance (ANOVA), using the IBM SPSS Statistics 21 software package. Whilst no significant differences were predicted between muscles due to both muscles having similar involvement in the execution of a basketball free-throw, muscle was included as a factor in the ANOVA as it was prudent to first examine whether differences between muscles existed before exploring differences between conditions. Where Mauchly's test indicated that the assumption of sphericity had been violated, the degrees of freedom were corrected using the Greenhouse-Geisser method. The alpha level for statistical significance was set at α = .05 and effect sizes are reported as Cohen's d. Post-hoc pairwise comparisons with the Bonferroni adjustment were used to explore significant effects.","Table 2 shows the raw MEP amplitude data obtained from each muscle in each condition. Due to the large variability both within and between participants in the raw MEP amplitudes, this data was normalized using the z-score transformation. Fig. 2 shows the normalized z-score MEP amplitude data. In this figure, a value of zero indicates the mean MEP amplitude across all conditions, with the positive and negative values indicating by how many standard deviations a particular condition was above or below the mean of all conditions, respectively. The 2 (muscle) x 4 (condition) repeated measures ANOVA performed on the z-score MEP amplitude data showed no significant main effect for muscle, F(1, 17) = 0.59, p = .45. There was, however, a significant main effect for condition, F(3, 51) = 6.21, p = .001. Post-hoc pairwise comparisons with the Bonferroni adjustment showed that MEP amplitude in the AOMI condition was significantly larger than in both the static observation (p = .03, d = 0.75) and action observation (p = .05, d = 0.71) conditions (see Fig. 2). No other pairwise comparisons were statistically significant. The muscle × condition interaction was not significant, F(1.8, 30.61) = 0.07, p = .91.","The aim of this experiment was to establish the effects of different action observation and motor imagery conditions on corticospinal excitability for a sport-related motor skill. Specifically, the amplitude of MEPs obtained during AOMI, independent AO and independent MI of a basketball free throw task were compared against a control condition. MEP amplitudes were significantly larger during AOMI, compared to both the control condition and the independent AO condition. There was no difference in MEP amplitude between either the independent AO or independent MI conditions and the control condition. As the amplitude of the MEP provides a marker of corticospinal excitability (Naish et al., 2014), these results indicate that in the current experiment corticospinal excitability was only facilitated by AOMI, but not by independent AO or MI. This finding of increased activity in the motor system during AOMI supports previous research showing increased neurophysiological activity in various motor regions of the brain during AOMI conditions using TMS (e.g., Mouthon et al., 2015; Ohno et al., 2011; Sakamoto et al., 2009; Wright et al., 2014; Wright et al., 2016), EEG (e.g., Berends et al., 2013; Eaves, Behmer, et al., 2016) and fMRI (e.g., Nedelko et al., 2012; Taube et al., 2015; Villiger et al., 2013). The findings of the current experiment add to this body of literature by demonstrating this effect in a sport-related motor skill, as opposed to simple hand movements or activities of daily living. The facilitation of corticospinal excitability during AOMI is likely to reflect increased activity in the premotor cortex in this condition. Meta- analyses of neuroimaging data have shown that the primary motor cortex, to which TMS was delivered in this experiment, is not reliably activated during MI or AO (Caspers, Zilles, Laird, & Eickhoff, 2010; Hardwick et al., 2017; Hétu et al., 2013). The primary motor cortex, however, is linked to the premotor cortex by strong cortico-cortical connections (Fadiga, Craighero, & Olivier, 2005). Hardwick et al. (2017) recently demonstrated that the dorsal (PMd) and ventral (PMv) premotor cortices are activated consistently by both AO and MI. Although both simulation states can evoke activity in the premotor regions, multi- voxel pattern analysis has shown that AO and MI produce activity in topographically distinct regions of the premotor cortex. For example, Filimon, Rieth, Sereno, and Cottrell (2015) reported that anterior regions of PMd and posterior regions of PMv are more active during MI, whilst lateral and posterior regions of PMd and anterior regions of PMv are more active during AO. It is therefore possible that instructing participants to engage simultaneously in AOMI would produce increased and more widespread activity throughout the premotor cortex than independent AO or MI. This would manifest in an elevated MEP response via cortico-cortical connections linking the premotor and motor cortices (Fadiga et al., 2005) and would explain why the greatest facilitation in corticospinal excitability in this experiment was found during the AOMI condition. Based on this finding, it is conceivable that greater improvements in the performance and learning of motor skills may be obtained through AOMI interventions, compared to the more established use of independent AO or MI (Holmes & Wright, 2017). Specifically, the increased activity obtained during AOMI may promote functional connectivity and plasticity within the brain, facilitating a more efficient motor execution as learning progresses (O'Shea & Moran, 2017; Ruffino, Papaxanthis, & Lebon, 2017). Although longitudinal research incorporating both neurophysiological and performance measures is required to verify this claim, some preliminary evidence indicates that AOMI interventions can modulate behavioral outcomes (see Eaves, Riach, et al., 2016). For example, Romano-Smith, Wood, Wright, and Wakefield (2018) reported that AOMI interventions can improve aiming performance in a dart throwing task. In addition, AOMI has been shown to influence automatic imitation effects (Bek, Poliakoff, Marshall, Trueman, & Gowen, 2016; Eaves, Behmer, et al., 2016; Eaves, Haythornwaite, & Vogt, 2014) and improve balance (Taube, Lorch, Zeiter, & Keller, 2014), grip strength (Sun, Wei, Luo, Gan, & Hu, 2016) and hamstring strength (Scott, Taylor, Chesterton, Vogt, & Eaves, 2017). Whilst further research is required to examine the effect of AOMI on the performance and learning of motor skills, there are possible explanations for why AOMI interventions may provide an effective tool for sport psychologists and athletes. One possibility is that AOMI interventions may contribute to improvements in motor performance and learning by developing athletes' mental representation of a skill. Mental representations are cognitive representations for motor actions comprising a compilation of body postures and associated sensory consequences, known as basic action concepts, that are related functionally and biomechanically to the successful execution of a motor skill (Frank, Land, & Schack, 2013; Schack, 2012). These mental representations are encoded in long-term memory and guide motor skill execution (Land, Volchenkov, Bläsing, & Schack, 2013; Schack & Mechsner, 2006). According to Schack and Mechsner (2006), expert performers have mental representations that are highly organized and closely related to the functional demands of the skill, whereas the mental representations of novices are comparatively less organized and less closely related to the functional demands of the skill. Frank et al. (2013) demonstrated that mental representations of novices became functionally more organized as performance improved following physical practice. Recent research indicates that the structure of novices' mental representations can also be developed through both AO (Frank, Kim, & Schack, 2018; Kim, Frank, & Schack, 2017) and MI (Frank, Land, Popp, & Schack, 2014; Kim et al., 2017) interventions. Although both AO and MI contribute to the development of mental representations of action, it is possible that they do so through different mechanisms (Kim et al., 2017). AO provides a visual representation of an action, typically without the deliberate generation of associated kinesthetic sensations. As such, AO may enhance the structure of mental representations primarily through developing the sequencing and timing of different basic action concepts. In contrast, MI involves the generation of visual and kinesthetic aspects of a movement and so may enhance an individual's mental representation primarily by developing the sensory consequences associated with different basic action concepts. By combining the two techniques, AOMI interventions may develop the mental representation of a skill by enhancing both the sequencing between basic action concepts and the associated sensory consequences, and this in turn may lead to improvements in motor skill performance and learning. In this experiment, corticospinal excitability was not facilitated by either independent AO or independent MI. This finding was somewhat unexpected as it is well-established in the TMS literature that both AO (e.g., Naish et al., 2014) and MI (e.g., Grosprêtre et al., 2016) usually facilitate corticospinal excitability, relative to control conditions. This effect has been demonstrated in sport-related tasks for both AO (e.g., Aglioti et al., 2008; Wrightson, Twomey, & Smeeton, 2016) and MI (e.g., Fourkas, Bonavolontà, Avenanti, & Aglioti, 2008; Wang et al., 2014). Although this finding conflicts partially with our hypothesis and with previous TMS research on this topic, it could be explained by the choice of stimuli used for the control condition in this experiment. There are inconsistencies in the choice of control conditions used across experiments exploring AO, MI or AOMI with TMS (Loporto, McAllister, Williams, Hardwick, & Holmes, 2011). Rest (e.g., Wang et al., 2014), observation of blank screens (e.g., Wrightson et al., 2016), fixation crosses (e.g., Sakamoto et al., 2009) or static images of the body or a body part (e.g., Aglioti et al., 2008; Wright et al., 2014) are all common choices of control stimuli. Loporto et al. (2011) suggest that the use of a static image of a body or body part as the control condition is the most appropriate as it ensures that any facilitation of corticospinal excitability during observation or imagery conditions is related to the observation or imagery of biological movement. In contrast, when using a fixation cross or blank screen it is not possible to determine whether a facilitation effect is due to the observation or imagery of biological movement per se, or rather just the presence of some form of visual stimuli on screen or the involvement of some form of cognitive activity (Loporto et al., 2011). The use of a static image of the body was, therefore, chosen deliberately for this experiment to provide a more stringent control condition against which the effects of the three different interventions could be compared. The fact that only AOMI produced a facilitation of corticospinal excitability relative to this stricter control condition provides justification for the use of AOMI, rather than independent AO or independent MI interventions. Although this experiment is the first to demonstrate the effects of AOMI of a sport-related motor skill on corticospinal excitability, it is important to acknowledge several possible limitations associated with the experiment. First, the four conditions were presented in a fixed order, rather than being randomized or counterbalanced throughout the experiment. As participants always completed the AOMI condition last, it is possible that the enhanced MEP amplitude in this condition was due to either increased familiarity with the stimuli, or a carry-over effect whereby MEP amplitude was enhanced during the final condition due to residual corticospinal activity from the previous conditions. Although these explanations are plausible, Loporto et al. (2012) showed in two experiments that MEP amplitude did not change over the course of observing five blocks of the same action observation stimuli. In addition, the current experiment utilized a 3 min rest period between conditions as Baldi et al. (2002) showed that MEP amplitudes return to baseline levels after only 1 min. Taken together, it is therefore unlikely that the increased effect reported in the AOMI condition is due to familiarity with the stimuli or carry-over effects. Instead, the finding for the AOMI condition is likely to reflect increased activity in the premotor cortex resulting from combining the two simulation states. Second, imagery perspective may have differed between the MI and AOMI conditions. Participants were told to image the feelings and sensations associated with executing the free throw, but were not instructed to use a specific imagery perspective in either condition. The third-person perspective of the video in the AOMI condition may have encouraged imagery from this perspective, whereas first- or third-person perspectives may have been used in the MI condition, depending on an individual participant's perspective preference. Imagery from a third-person perspective may produce MEPs of larger amplitude (Fourkas, Avenanti, & Aglioti, 2006), although it may also be more difficult to generate kinesthetic imagery from this perspective (Callow & Hardy, 2004). Given this conflict, future research should provide a stricter control of imagery perspective. It may also be worthwhile to investigate the effects of manipulating different AO and MI perspective combinations within AOMI interventions on various neurophysiological and behavioral measures. A final issue to be acknowledged is that the present experiment used novice participants rather than experienced basketball players. Neurophysiological activity during AO and MI differs between experts and novices. Specifically, expert performers in a variety of skills typically exhibit increased neurophysiological activity during AO and MI compared to novices (Aglioti et al., 2008; Calvo-Merino, Glaser, Grezes, Passingham, & Haggard, 2005; Fourkas et al., 2008; Mizuguchi & Kanosue, 2017). As such, the direction of the effects reported here would likely replicate in an expert sample, although the magnitude of the effects may be enhanced. This would be a worthwhile area for future research investigating the neurophysiological effects of AOMI interventions. In conclusion, the main finding of this experiment is that AOMI of a basketball free throw facilitated corticospinal excitability relative to the control condition, but independent AO or MI had no such effect. This finding has important implications for the design and delivery of sport psychology interventions aimed at improving sport performance and enhancing motor skill learning. Independent AO (Ste-Marie et al., 2012) and independent MI (Cumming & Williams, 2012) are well-established techniques that are used widely for improving motor skill performance and learning. The mechanism by which these methods are effective is through producing activity in brain regions that are involved in motor execution (Jeannerod, 2001). The findings of the current experiment indicate that greater activity in the motor system occurs when AO and MI are combined into a single intervention strategy. As such, implementing AOMI interventions may offer a more effective method for improving motor skill performance and learning than the independent use of either technique. There is, however, currently a lack of research examining the effects of AOMI interventions on the performance and learning of motor skills. Future research should therefore first attempt to identify the efficacy of AOMI interventions for improving movement outcome and technique across a range of skill types, for both novice and expert performers. Research could then explore optimal methods for delivering AOMI interventions by, for example, establishing the efficacy of different visual perspectives for AOMI or the effects of introducing MI alongside AO gradually in a layered manner."],["Lisa Bortolotti is interested in the potential benefits of irrational beliefs, and she focuses on pathological beliefs. In particular, she discusses those delusions that have been construed as playing a defensive function, such as Reverse Othello syndrome, erotomania, and anosognosia. Such delusions are wildly implausible, but at the time at which they are endorsed, they may carry both psychological and epistemic benefits. They act as a defence protecting agents from low self-esteem and the potentially disruptive consequences of overwhelming negative emotions. In virtue of such benefits, they also allow agents to avoid depression and continue interacting with the surrounding physical and social environment in a way that may be conducive to feedback from social exchanges and to the acquisition of useful information. To characterise cognitions that are typically false and irrational but may also carry benefits of this sort, Bortolotti introduces the notion of epistemic innocence which captures the status of cognitions that have some significant epistemic benefit and whose benefit could not be attained by other means. To show that at least in some circumstances there may be no alternatives to a delusional belief, Bortolotti argues that in anosognosia people do not have direct evidence of their impairments and are unable to integrate indirect evidence about their impairments in their concept of themselves. This means that one plausible alternative to the delusional belief that they are not impaired, that is, the belief that they are impaired, may not be available to them. Bortolotti concludes that, in the case of motivated delusions, psychological benefits can turn into epistemic ones. In their paper, Maarten Boudry, Michael Vlerick, and Ryan McKay revisit the contribution of the friends of ecological rationality to the rationality debate in cognitive science. The rationality debate concerns the implications of people failing simple inductive and deductive reasoning tasks in experimental settings. The friends of ecological rationality correctly point out that failure in solving the reasoning tasks may be explained in some circumstances by the fact that the tasks are presented in a misleading way. But they also defend a more general and stronger claim, the claim that heuristics regarded as epistemically flawed or biased can be shown to be ecologically rational. This is the claim Boudry and colleagues find problematic. Some of the heuristics responsible for reasoning mistakes can be adaptive, but they cannot be redeemed as rational. Boudry and colleagues illustrate the difference between adaptiveness and rationality with the example of superstitious beliefs and fast-and-frugal heuristics. Superstitious beliefs may be adaptive (as in some environments genetic fitness may be enhanced in organisms that avoid risks) but this does not make them rational (as they are badly supported by the evidence and fail to track the truth). A fast-and-frugal heuristic such as the recognition heuristic is useful when people are asked to make a choice in a situation of ignorance. They select better-known versus less well-known items, and this can lead them to making the right choice in some environments. But in advertising the recognition heuristic is exploited: the presence of a known brand determines a consumer’s choice, and other relevant factors such as product quality are not taken into account. Heuristics are effective in some domains, and misfire in others. Their local adaptiveness is definitely a benefit, but it is not a good indication of their epistemic rationality.","Jordi Fernandez argues that memories have (at least) two types of functions and two types of benefits. They preserve information about the past (narrative function) and they are reconstructions of events that engage the same capacities involved in imagination and are aimed to build a narrative (reconstructive function). They can provide good evidence about the past, thereby allowing the subject to represent the past accurately (which is epistemically beneficial); and they contribute to the formation of beliefs about the past that have an instrumental value for the agent, thereby allowing the agent to satisfy some of her goals (which is adaptively beneficial). Depending on the agent’s goals, the epistemic benefits and the adaptiveness of memories can be related. In the paper, Fernandez asks whether two forms of memory distortions—observer memories and fabricated memories—can be adaptive. In the context of trauma, observer memories enable agents to obtain some affective relief in the short term, but may hinder their capacity to develop a coherent and healthy self-concept in the long term. In the context of false memories of abuse, memories can respond to a need for explanation thereby relieving internal tensions, but can cause emotional damage and compromise personal relationships. The interesting result is that, if we take memory to have only a narrative function, then observer memories and fabricated memories are not distorted. They have been produced to further the goals of the agent, not to represent the past correctly. But if we take memory to have only a preservative function, then observer memories and fabricated memories have no benefits, because the only benefits that count are epistemic ones. Fernandez argues that both conclusions are unattractive and that the case of beneficial memory distortions suggests that we should take an inclusive approach to the functions of memory. Martin Conway and Catherine Loveday reach a similar conclusion to Fernandez, that false memories can have significant benefits for an agent, but start from a more radical position in that they downplay the preservative function of memory, based on empirical investigations of how memory works. Memories and imagined events are constructed in a similar way, inferentially, via the so-called “remembering–imagining system”, and the accuracy of autobiographical memories is understood in terms of the relationship between correspondence (how the memory captures an experienced event) and coherence (how the memory coheres with other beliefs about the self). Whereas a memory can succeed in its coherence, it can never fully succeed in its correspondence as it will always be partial and to some extent distorted. No memories represent events “literally” and maybe they are not supposed to do so. The main function of memory is to “generate personal meanings”, that is, to provide an understanding of the world that allows agents to successfully adapt to it.","Aikaterina Fotopoulou argues that, just as the past self is known via inference from autobiographical memory, so the present self is known via perceptual inference. Due to its indirect nature, representations of the past and present self are imperfect, both in the context of normal perceptual inference and of pathological conditions such as anosognosia for hemiplegia. Consistent with the prediction–error model, in Fotopoulou’s account the brain predicts the hidden causes of sensory inputs and revises such predictions in order to minimise errors. There some delusional beliefs emerging in anosognosia, such as the illusion of movement and the adherence to the denial of paralysis even after the paralysis has been acknowledged. How can these beliefs be explained? A temptation is to rely on a multiplicity of distinct factors, where the hypothesis is that both perception and reasoning are damaged. But Fotopoulou argues instead that we should just focus on perceptual inference. In anosognosia prediction errors are absent or unreliable and this results in patients making inferences from out-of-date models of their motor abilities. More specifically, in Fotopoulou’s account, anosognosia is due to an inability to update bodily awareness in the light of new information about the affected body parts and to integrate first- and third-person perspectives on the body. Anosognosia for hemiplegia is just an exaggeration of the imperfection of bodily awareness. Kengo Miyazono also considers the costs of delusional beliefs, and attempts to account for their pathological nature. What is the difference between everyday irrational beliefs, such as the unjustified belief in the infidelity of one’s spouse or common instances of self- deception, and delusional beliefs that are symptoms of schizophrenia and delusional disorders? Miyazono critically assesses several answers provided in the literature and dismisses them: delusions are not necessarily more bizarre, more irrational, or less understandable than non-delusional beliefs. Moreover, it is not always the case that a person with delusions lacks responsibility for the actions guided by her delusional beliefs. Miyazono’s positive account is that delusions are pathological because they involve a harmful biological malfunction. Delusions are harmful because they disrupt good functioning and often negatively affect the quality of life of people who report them. Delusions are malfunctioning beliefs because the processes by which they are formed are abnormal in some important respect (that is, in some respect other than statistical normality), where the respects in which the process malfunctions vary according to one’s preferred aetiological account of delusions. In the person who forms delusions, either experience is abnormal or, additionally, there are deficits concerning attention and reasoning.","Jules Holroyd investigates responsibility for implicit biases and the actions resulting from them. Holroyd considers three epistemic conditions for responsibility for implicit bias, endorsing the third, which is that one should have observational awareness of the effects of implicitly biased behaviour. Observational awareness is being aware that one’s behaviour has some morally undesirable property, for example, the property of being discriminatory. According to Holroyd, we should not be asking whether agents do, as a matter of fact, have observational awareness of their biased behaviour, but rather, whether they ought to have this awareness, and whether they are culpable for not having it. Drawing on empirical work, Holroyd argues that agents can have observational awareness of their discriminatory behaviours that are based on implicit attitudes. She also argues that agents ought to have this awareness, and resists the claim that agents are not responsible for behaviours manifesting biases because biased actions are guided by implicit cognitions. Finally, Holroyd considers the role of other imperfect cognitions in relation to implicit biases, specifically, failures of attentiveness, and self-deception. She suggests that an investigation into whether agents are responsible for actions guided by implicit biases may in part depend on the relationship between implicit biases and other imperfect cognitions. Ema Sullivan-Bissett is interested in the epistemic status of confabulatory explanations of decisions or actions guided by implicit bias. She is keen to resist the trade-off view of imperfect cognitions; that a cognition enjoys pragmatic benefits at the expense of epistemic ones. To this end, she too appeals to the notion of epistemic innocence. She focuses on two imagined cases of decisions or actions guided by implicit biases. Via an analysis of these cases, she argues that at least sometimes, confabulatory explanations of decisions or actions guided by implicit bias are epistemically innocent. First, they may be epistemically beneficial. They fill an explanatory gap, potentially leading to the acquisition and retention of true beliefs and knowledge, and they also help to maintain consistency between an agent’s beliefs. Second, alternative (more epistemically worthy) explanations that could confer these benefits are unavailable. Sullivan-Bissett concludes that when we are in the business of epistemic evaluation, we should consider both the epistemic benefits of an imperfect cognition, and the context in which it occurs.","We believe that the eight papers in this issue initiate a much needed interdisciplinary dialogue on imperfect cognitions, and make substantial progress in answering key research questions about the types of costs and benefits that such cognitions in the clinical and non-clinical population may have. Hopefully the ideas presented here will also stimulate further empirical research and conceptual investigation into different forms of imperfect cognitions, and help sketch a more psychologically realistic account of human agency and cognition."],["The music genre of jazz is commonly associated with creativity. However, this association has hardly been formally tested. Therefore, this study aimed at examining whether jazz musicians actually differ in creativity and personality from musicians of other music genres. We compared students of classical music, jazz music, and folk music with respect to their musical activities, psychometric creativity and different aspects of personality. In line with expectations, jazz musicians are more frequently engaged in extracurricular musical activities, and also complete a higher number of creative musical achievements. Additionally, jazz musicians show higher ideational creativity as measured by divergent thinking tasks, and tend to be more open to new experiences than classical musicians. This study provides first empirical evidence that jazz musicians show particularly high creativity with respect to domain-specific musical accomplishments but also in terms of domain-general indicators of divergent thinking ability that may be relevant for musical improvisation. The findings are further discussed with respect to differences in formal and informal learning approaches between music genres. © 2014 The Authors. --------------------------------------------------------------------------------","Within the field of music, jazz is commonly considered as a particularly creative discipline (e.g., Barrett, 1998). This appraisal is related to the fact that jazz music involves a high degree of improvisational playing. Jazz improvisation can range from the simple embellishment of the melody of the theme to e.g. the continuous extemporization of entirely new melodies that fit to the sequence of chords (Johnson-Laird, 2002; Pressing, 1988). Jazz musicians who are highly skilled in improvising hence may possess traits that are different from those of musicians in other disciplines such as classical music. So far, only little is known about the individual differences between musicians devoted to different music genres. Therefore, this study compared jazz musicians with musicians of classical and folk music with respect to their musical activities, creativity and personality. Only few studies have investigated specific differences in attitudes, and learning approaches of musicians specialized in different music genres (e.g., Bézenak & Swindells, 2009; Creech et al., 2008; Papageorgi, Creech, & Welch, 2013; Welch et al., 2008). Classical musicians are reported to acquire musical skills mainly in formal educational settings involving one-to-one instruction and by practicing alone, whereas non-classical musicians devote more time to extra-curricular activities such as playing music for fun with others or having professional conversations (Bézenak & Swindells, 2009; Welch et al., 2008). Additionally, classical musicians attach greater importance on technical proficiency involving sight-reading, notation, and quality of tone, whilst non- classical musicians appear to attach greater importance to skills such as memorization or improvisation (Bézenak & Swindells, 2009; Creech et al., 2008). Bézenak and Swindells (2009) found that jazz musicians show higher intrinsic motivation and experience more pleasure in musical activities than classical musicians. In contrast, classical musicians report higher levels of performance anxiety than other non-classical musicians (Papageorgi et al., 2013). These findings already suggest important differences in the general approach towards learning and playing music between different genres such as jazz and classical music. Research also addressed the question what factors lead to expert performance in music and more specifically in improvisational skills. It is now widely accepted that the cumulative amount of deliberate practice but also the quality of practice is highly predictive of mastery in the domain of music (Ericsson, Krampe, & Tesch-Römer, 1993; Williamon & Valentine, 2000). Additionally, there is evidence that individual differences in domain-general cognitive abilities also contribute to expert performance (Hambrick et al., in press). Beaty, Smeekens, Silvia, and Kane (in press) report a study where ten jazz students were video-taped during improvisation performances on a piece unknown to them, which then was rated for creativity by three professors of jazz studies. They found that creativity of improvisation was independently predicted by practice hours and divergent thinking ability (i.e., a common indicator of creative potential) of the jazz students. The findings suggests that divergent thinking, commonly defined as the ability to fluently generate original and appropriate ideas, may represent a relevant ability supporting improvisational creativity. This notion is in line with formal models of jazz improvisation stating that improvisation requires the continuous generation and evaluation of musical ideas (Pressing, 1988). Similarly, divergent thinking is considered as a central factor underlying creative thinking in music according to Webster’s model (2002), together with certain differences in personality and motivation. As a consequence, jazz musicians who are highly skilled in improvisation may differ in their creativity and personality from musicians of other genres. The aim of this study is to formally test this hypothesis by comparing Jazz musicians with musicians specialized in classical and folk music.","A total of 120 students enrolled in the study of instrumental pedagogy at the University of Music and Arts in Graz participated in this study. They majored in various different musical instruments (e.g., piano, violin, voice), but were enrolled in one of three tracks related to a specific genre of music: classical music, jazz music, or folk music. The study curriculum is largely the same for all three genres, but classical and folk musicians have more courses on analyzing theoretical aspects of music as compared to Jazz musicians, who attend more courses focused on improvisational skills, ensemble playing and developing practical musical skills. The curriculum of folk musicians specifically requires the playing of at least two folk instruments and offers supplementary classes on folk dance or yodeling. We excluded seven participants who were enrolled in more than one music program and hence could not be attributed unambiguously to one music genre. Moreover, we included only students who indicated to have good to excellent language skills, leading to the exclusion of another 14 participants. The remaining sample consisted of 99 students, including 52 students of classical music, 25 students of jazz music, 21 students of folk music. On average, students had an age of 24.8 years (SD = 5.6), and have been studying music for 2.6 years (SD = 1.8). The sex distribution was fairly balanced with 47% females. The music groups did not differ in their age (F[2,95] = 2.25, p = .11), nor sex ratio (χ2[2] = .08, p = .96), but jazz students on average reported a longer duration of study (F[2,84] = 9.69, p = .001, partial-η2 = .19; classical music: 2.1 years; jazz music: 3.9 years; folk music: 2.4 years). For analyses involving speeded creativity tests we only included participants with German as mother tongue, resulting in 70 students (30 classical music, 22 jazz music, 18 folk music). Study and practice activities We assessed relevant socio-demographic information including age, sex, nationality, selected study programs, and students gave a self-assessment of language skills (“excellent”, “good”, “fair”, or “bad”). They were asked how many hours they typically practiced their instruments at every single day of the week. This data was used to compute a reliable estimate of the practice hours per week. Finally, participants indicated how many concerts they play per semester, and how many competitions they had participated, how often they had won competitions, and how many productions they had published so far. Creativity assessment Creative cognitive potential in the verbal domain was assessed with four divergent thinking tasks taken from a well-known German creativity test (Verbaler- Kreativitätstest; VKT; Schoppe, 1975). The tasks included two alternate uses tasks asking participants to generate different creative uses for a “tin can” and a “simple string”, and two instances tasks which asked to generate many things that could be used “for faster locomotion” or that are “bendable”. In all tasks, participants were instructed to find as many and as creative ideas as possible within the given time (120s, or 90s for the alternate uses and the instances task, respectively). The performance in the divergent thinking tasks was scored for ideational fluency (i.e., number of ideas), and ideational creativity. For the scoring of ideational creativity we created lists of pooled, alphabetically sorted, non-redundant responses for each task. Four experienced raters rated each idea for creativity on a four-point scale (“0, uncreative”, “1, somewhat creative”, “2, fairly creative”, and “3, very creative”). We then computed a top-3 creativity score by averaging the creativity ratings of the three top-most creative ideas within each task (Benedek, Mühlmann, Jauk, & Neubauer, 2013). This scoring method was found to yield valid scores that show to discriminant validity with regard to fluency measures (Benedek, Franz, Heene, & Neubauer, 2012; Benedek et al., 2013; Silvia et al., 2008). We averaged scores of the two alternate uses tasks and the two instances tasks to obtain one fluency score and one creativity score per task type. Additionally, creative potential in the figural domain was assessed with a picture completion task taken from the imagination subscales of the Berliner-Intelligenz-Test (Jäger, Süß, & Beauducel, 1997). Participants were shown a series of abstract lines which had to be completed in an original way to form meaningful objects. This task was scored for ideational fluency following the instructions of the test manual. Besides creative potential, we also assessed real-life creative activities and achievements of the students using the inventory of creative activities and achievements (ICAA; described in Jauk, Benedek, & Neubauer, 2014). This inventory assesses creative activities and achievements in eight domains, including literature, music, arts and crafts, creative cooking, sports, visual arts, performing arts, and science and engineering. In the activities scale, participants report on a 5-point scale how often they carried out certain activities within the last 10 years. In the achievements scale, participants marked achievements they had already attained in each domain ranging from “I have never been engaged in this domain” (0 points) to “I have already sold some of my original work in this domain” (10 points), and values of all achievements are summed. Activities and achievements scores can be analyzed separately for each domain or as a composite score, after summing across domains. Personality assessment Personality was assessed with respect to the Big Five using the NEO-FFI (Borkenau & Ostendorf, 1993). We assessed schizotypy using the German 17-item version of the Schizotypical Personality Questionnaire (SPQ; Klein, Andresen, & Jahn, 1997). Participants also completed the Error Orientation Questionnaire (EOQ; Rybowiak, Garst, Frese, & Batinic, 1999). This questionnaire contains 37 items asking about individual attitudes towards errors at work. In this study work was defined as practicing and performing activities as a musician. The EOQ consists of eight scales, including error competence, learning from errors, error risk taking, error strain, error anticipation, covering up errors, error communication and thinking about errors. Further questionnaires include the German version of the Frost Multidimensional Perfectionism Scale (FMPS; Stöber, 1998), the German Achievement Motivation Scale (AMS; Lang & Fries, 2006), rumination scale of the Perfectionism Inventory (PI; Hill et al., 2004), the student version of the SELLMO (a German questionnaire on learning and achievement motivation; Spinath, Stiensmeier- Pelster, Schöne, & Dickhäuser, 2012), and a set of self-devised questions on error behavior.","Participants were tested in groups of 10–25 people in lecture rooms. First, they provided general information on their person, studies and practicing habits. They then worked on the divergent thinking tasks, the NEO-FFI, the SPQ, the EOQ, completed a self-devised questionnaire on error behavior, the ICAA, the FMPS, the AMS, the PI rumination scale and the SELLMO. The total session took about 90 min.","Potential group differences between musical genres (classical music, jazz, or folk music) were analyzed by means of ANOVAs with the between-subject factor music genre. In case of significant group effects, LSD posttests were employed to further examine differences between group means. Genre-related differences in general musical activities ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Students of folk music devote on average 12 h per week on practice, which is significantly less than the typical practice periods of students of classical music (p = .01) or jazz music (p = .08), who spend about 18 h a week (F[2,92] = 3.39, p = .04, partial-η2 = .07; see Table 1). Jazz musicians played significantly more concerts per semester than classical musicians (p = .001) and folk musicians (p = .001; F[2,92] = 10.78, p = .001, partial-η2 = .19). On the other hand, jazz musicians participated in a lower number of music competitions (F[2,92] = 4.16, p = .02, partial-η2 = .08) than classical (p = .01) and folk musicians (p = .02), and also won a lower number of competitions (F[2,92] = 4.60, p = .01, partial-η2 = .09). Finally, folk musicians published a significantly higher number of works than classical musicians (p = .03; F[2,92] = 3.08, p = .05, partial-η2 = .06). Genre-related differences in creativity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Students of classical, jazz and folk music were compared in their levels of creative potential, creative musical activities and creative musical achievements. As can be seen in Table 2, the three groups differed significantly in ideational creativity as measured by the alternate uses task (F[2,67] = 3.61, p = .03, partial-η2 = .10) and the instances task (F[2,67] = 3.96, p = .02, partial-η2 = .11). In both tasks, Jazz musicians showed higher ideational creativity than folk musicians (p = .01, and p = .01) and classical musicians (p = .09, and p = .02). No additional group differences were observed with respect to ideational fluency in the divergent thinking tasks (alternate uses task: F[2,67] = 0.34, p = .71; instances task: F[2,67] = 1.97, p = .15; picture completion task: F[2,67] = 0.17, p = .85). We then analyzed potential group differences in creative activities and achievements with a focus on the musical domain. Jazz musicians reported to have engaged in a significantly higher number of creative musical activities over the last years (F[2,94] = 20.20, p < .001, partial-η2 = .30) than classical musicians (p < .001) and folk musicians (p < .001). Moreover, jazz musicians also showed higher creative achievements in the musical domain (F[2,94] = 15.03, p < .001, partial-η2 = .24) than classical musicians (p < .001) and folk musicians (p = .003). Notably, these group differences remained highly significant even after statistically controlling for differences in age and duration of study. As a side analysis, we also looked for group differences in other domains measured by the ICAA. The only significant finding was that folk musicians showed higher creative achievements in the domain of arts and crafts (F[2,94] = 8.92, p < .001, partial-η2 = .16) than classical musicians (p = .02) and jazz musicians (p = .02). Genre-related differences in personality ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Analyses of group differences in personality structure revealed a significant effect for extraversion (F[2,96] = 3.88, p = .02, partial-η2 = .08) and an effect by trend for openness (F[2,96] = 2.42, p = .06, partial-η2 = .05), but not effects for neuroticism, agreeableness, or conscientiousness (see Table 3). Specifically, folk musicians were found to be more extraverted than classical musicians (p = .007) and jazz musicians (p = .05). Classical musicians tend to be less open to new experiences than jazz musicians (p = .06) and folk musicians (p = .11). The music genre groups did, however, not differ in schizotypy (F[2,96] = 0.34, p = .72). No further significant group differences were observed in the motivational measures of this study, including all sub-facets of error orientation and perfectionism, and indicators of achievement motivation, learning motivation and rumination.","The analysis of general musical activities revealed that the participants, although still studying at the Arts College, already were very accomplished musicians, completing a high weekly pensum of practice, playing a considerable number of concerts per year, regularly participating in music competitions, and having a substantial number of music productions published. Moreover, the music students already have attained very high levels of creative achievement in the domain of music, with 8% having accomplished every single musical achievement listed in the employed inventory. Interestingly, students of different music genres differed substantially in the relative amount of engagement in these activities. Classical musicians practice a lot and participate in a high number of competitions, but do not publish as many music productions as non-classical musicians. In contrast, jazz musicians perform a larger number of concerts per year but do not participate as often in music competitions. This confirms recent research showing that classical musicians are more focused on achievements related to solo professional work, whereas jazz musicians are more engaged in informal ways of practice by playing lots of concerts (Bézenak & Swindells, 2009; Creech et al., 2008; Welch et al., 2008). Finally, folk musicians showed a lower amount of weekly practice but have many performances in concerts, competitions and music productions. Taken together, these findings support the view that music learning is not only the product of formal educations systems but also includes different informal ways of practice (Green, 2002), and that approaches to music learning may differ between music genres (e.g., Welch et al., 2008). Besides differences in musical activities, music genre was also related to individual differences in psychometric creativity. As expected, Jazz musicians showed higher divergent thinking ability (i.e., creative cognitive potential) in terms of ideational creativity than classical and folk musicians. The ability to fluently generate original ideas can be considered highly compatible with the improvisational skills that are required and trained in jazz music, and they may be of relatively lower significance in classical music or folk music (Pressing, 1988; Webster, 2002). Interestingly, Fink and Woschnjak (2011) reported a similar finding from the domain of dance. They found that modern/contemporary dancers, who are often required to improvise on stage, showed higher creative potential than ballet dancers who are normally obliged to adhere to well-structured choreographies. Of course, one can only speculate about the causality in the relationship between divergent thinking and improvisation abilities: Is high divergent thinking ability a precondition for becoming a good jazz musician, or does continuous improvisation training implicitly increase divergent thinking ability? There is evidence in support of both perspectives. On the one hand, divergent thinking was shown to predict improvisational creativity beyond the mere amount of practice (Beaty et al., in press). On the other hand, extensive engagement in divergent thinking can increase divergent thinking performance (Benedek, Fink, & Neubauer, 2006) and even have effects on relevant brain activation patterns (Fink, Grabner, Benedek, & Neubauer, 2006). Thus, both perspectives may apply to some extent. It would, however, need longitudinal studies to properly disentangle the relative effects in the causal relationship of creative potential and improvisation training. As another result, jazz musicians reported to engage in a much larger number of creative activities in the domain of music than classical or folk musicians. In this context, it is important to point out that the employed measure of creative musical activities specifically reflects activities that involve creating something new but it does not consider playing music as a creative activity per se. Sample items of this scale include “I reinterpreted a piece of music in a creative way”, “I made up a melody”, “I made up a rhythm”, or “I artificially created sounds”. These creative activities are very common during jazz improvisation, whereas classical music usually involves a flawless reproduction with focus on technical excellence (Bézenak & Swindells, 2009; Creech et al., 2008). To be sure, playing e.g. a Bach sonata also involves individual expression by giving it a personal note in terms of temper and atmosphere, but this individuality usually does not go as far as changing the rhythm or melody of the piece in a substantial way. A similar argument may also explain why jazz musicians showed much higher creative achievements in the domain of music. Again, the scale explicitly focuses on achievements related to original pieces of work (e.g., musical compositions or rearrangements). We also observed differences in personality between musicians of different genres of music. First of all, folk musicians are more extraverted than classical and jazz musicians. This finding may be related to the fact that folk music is commonly played at sociable events involving regular interactions with the audience. Therefore, the genre of folk music may more likely attract extraverted musicians that enjoy social interactions as an integral part of their performance. As an interesting additional finding, folk musicians were more achieved in the domain of arts and crafts. This may refer to stronger bonds of folk musicians to traditions and related skills in arts and crafts. We also observed a weak group effect for openness suggesting that jazz and folk musicians are more open to new experiences than classical musicians. Openness to new experiences reflects a preference for variety and the readiness to leave beaten paths. It is consistently related to creativity in the literature (e.g., Feist, 1998; Jauk et al., 2014) and may also promote the readiness to seek variation in musical play as required during improvisation. Finally, we did not observe group differences in error orientation or motivational variables between music genres in this study. This is an interesting finding as one might have expected that jazz musicians e.g. are more comfortable with risk taking during their improvisational play. It should be noted, however, that the EOQ does not differentiate between errors occurring during practice or performance (Kruse-Weber & Parncutt, 2013). It hence is possible that the questions were rather attributed to the process of learning rather than to stage performances and thus errors were conceived as equally important.","This study revealed evidence that jazz musicians show higher divergent thinking ability, and a higher number of creative activities and achievements in the musical domain as compared to musicians from other genres such as classical music or folk music. These findings support the view that the music genre of jazz is highly associated with creativity, both in terms of musical activities and psychometric aspects of musicians. The observed differences may be related to differences in the formal and informal ways of practice and learning, with Jazz musicians attaching more importance on informal practice and playing for fun and lower value on technical perfection and competitions. Finally, the findings add to the evidence that individual differences in domain-general abilities (i.e., creative potential) may be relevant for the realization of domain-specific creative activities and achievements (Jauk et al., 2014; Kaufman & Beghetto, 2009)."],["Anecdotal reports link alcohol intoxication to creativity, while cognitive research highlights the crucial role of cognitive control for creative thought. This study examined the effects of mild alcohol intoxication on creative cognition in a placebo-controlled design. Participants completed executive and creative cognition tasks before and after consuming either alcoholic beer (BAC of 0.03) or non-alcoholic beer (placebo). Alcohol impaired executive control, but improved performance in the Remote Associates Test, and did not affect divergent thinking ability. The findings indicate that certain aspects of creative cognition benefit from mild attenuations of cognitive control, and contribute to the growing evidence that higher cognitive control is not always associated with better cognitive performance. --------------------------------------------------------------------------------","Can alcohol consumption support creative thought by inducing disinhibition, or will it just impair cognitive control and similarly affect creative cognition? The idea about a positive relationship between alcohol and creativity has been popularized by reports associating eminent creativity with excessive alcohol consumption (Knafo, 2008). But empirical evidence is sparse, and the association between alcohol and creativity seems at odds with the relevance of cognitive control for creative thought (e.g., Benedek, Jauk, Sommer, Arendasy, & Neubauer, 2014). Therefore, this study examined the effect of alcohol on executive control and on standard measures of creative cognition. Creative cognition is assumed to rely on both controlled, goal-directed and spontaneous, undirected cognitive processes (Beaty, Silvia, Nusbaum, Jauk, & Benedek, 2014; Benedek & Jauk, in press; Sowden, Pringle, & Gabora, 2015). Pertinent research mostly focused on divergent thinking (viz. creative idea generation) and creative problem solving (i.e., problems that can be solved either analytically or insightfully, which typically implies a restructuring of the problem representation). The relevance of cognitive control for divergent thinking is evidenced by consistent correlations with intelligence (Kim, 2005; Silvia, 2015), particularly with fluid intelligence (Jauk, Benedek, Dunst, & Neubauer, 2013; Nusbaum & Silvia, 2011) and broad retrieval ability (Avitia & Kaufman, 2014; Benedek, Franz, Heene, & Neubauer, 2012; Silvia, Beaty, & Nusbaum, 2013). At the level of executive abilities, divergent thinking has been associated with working memory capacity and cognitive inhibition (Benedek et al., 2012, 2014; De Dreu, Nijstad, Bass, Wolsink, & Roskes, 2012; Zabelina, Robinson, Council, & Bresin, 2012). Divergent thinking requires overcoming prepotent, uncreative response tendencies and involves cognitive strategies (Gilhooly, Fioratou, Anthony, & Wynn, 2007), which was shown to be facilitated by intelligence (Beaty & Silvia, 2012; Nusbaum & Silvia, 2011; Nusbaum, Silvia, & Beaty, 2014). While much of the empirical evidence on creative cognition and cognitive control is based on divergent thinking, similar evidence also exists for creative problem solving. Creative problem solving tasks like Duncker’s candle problem or the Remote Associates Test can be achieved in a strategic way (Fleck & Weisberg, 2004; Smith, Huber, & Vul, 2013), and higher performance again has been related to intelligence and executive control (Gilhooly & Fioratou, 2009; Lee, Huggins, & Therriault, 2014). Creativity has also been associated with disinhibition and spontaneous insight (Eysenck, 1995; Kounios & Beeman, 2014). Empirical evidence for the relevance of spontaneous, undirected cognitive processes in creative thought mostly comes from research on incubation processes. Creative problem solving sometimes leads to an impasse of thought, also known as mental fixation, were goal-directed solving attempts are no longer fruitful. Incubation research has demonstrated that breaks from deliberate problem solving can benefit creativity by refreshing inadequate mindsets while leaving room for unconscious work (Hélie & Sun, 2010; Sio & Ormerod, 2009). Similarly, while expertise typically supports problem solving by guiding search through problem space, it can also be detrimental when misdirecting search efforts to salient but inadequate concepts (Wiley, 1998). Together, these findings suggest that cognitive control generally supports creative cognition by facilitating the effective implementation of goal-directed processes, but focused attention may sometimes be ineffective and potentially even harm creative problem solving (Wiley & Jarosz, 2012). The role of cognitive control in creative cognition has also been addressed by experimental studies examining the effects of low to moderate doses of alcohol (i.e., usually inducing a blood alcohol concentration <0.08) on different measures of creative ability. One study reported that the fluency of idea generation was reduced in both an alcohol group and a placebo group compared to the control group (Gustafson, 1991). Another investigation found that intoxicated writers and non-writers showed reduced idea flexibility but an increased number of non-obvious, original ideas (Norlander & Gustafson, 1998). Yet another study observed no notable effects of alcohol on divergent thinking performance, but participants evaluated their performance as more creative when they thought that they had received alcohol (Lang, Verret, & Watt, 1984). A more recent study demonstrated that moderate alcohol intoxication impaired working memory performance, but the intoxicated group showed higher performance in the Remote Associates Test (RAT) compared to a control group not receiving any drinks (Jarosz, Colflesh, & Wiley, 2012). Together, the available research provides partial support for a positive effect of alcohol on creative cognition, but evidence is still sparse and inconsistent. Part of the inconsistency might be attributed to missing placebo control groups, and the focus on single measures of creative potential. People tend to overestimate their creative performance and even become more creative when they think they have consumed alcohol, which points to the importance to include placebo control groups in order to dissociate pharmacological effects from expectation effects (Lang et al., 1984; Lapp, Collins, & Izzo, 1994). Moreover, findings may be specific to certain aspects of creative cognition such as insight problem solving and divergent thinking, and even the scoring of divergent thinking tasks can be an issue when focusing on summative uniqueness, which is known to be severely confounded with response fluency (Silvia et al., 2008). The present study thus tested the effects of mild alcohol intoxication in a placebo-controlled design using alcoholic and non-alcoholic beer. We examined the effects of alcohol on objective and subjective levels of intoxication, executive control, and two standard tasks of creative potential: creative problem solving in the Remote Associates Task and divergent thinking ability, scored for rated creativity, fluency, flexibility and novelty. Reduced cognitive control via mild alcohol intoxication could be expected to attenuate fixation effects (Smith & Blankenship, 1991) and thus support cognitive flexibility in the Remote Associates Test (Jarosz et al., 2012). The available research allows no clear prediction regarding the effect of alcohol on divergent thinking ability (Gustafson, 1991; Norlander & Gustafson, 1998). On the one hand, divergent thinking is known to involve executive processes similar to intelligence tasks (Benedek et al., 2014; Silvia, 2015). On the other hand, lower cognitive control might increase disinhibition and unusualness of thought (e.g., Babor, Berglas, Mendelson, Ellingboe, & Miller, 1982; Eysenck, 1995; Field, Wiers, Christiansen, Fillmore, & Verster, 2010), and thereby support the exploration of new and unusual parts of the idea space in the given problem (as measured by fluency, flexibility and novelty). Together, these mechanisms might facilitate the generation of unusual and potentially even creative ideas.","132 people participated in an online screening, which aimed to identify those eligible for participation in the main study. The online screening asked about age, potential pregnancy, heart or liver diseases, psychiatric disorders, and included the Alcohol use disorders identification test (AUDIT; Saunders, Aasland, Babor, De la Fuente, & Grant, 1993). 89 people from the screening met all criteria for participation, as they were at least 18 years old (the local age limit for legal consumptions of any type of alcoholic drinks), reported no pregnancy or relevant disease, and were casual drinkers of alcohol but with an AUDIT level < 8 reflecting no risk for alcohol-related problems. A total of 70 young adults (54 % female), aged between 19 and 32 years (M = 23.3; SD = 2.8), finally participated and completed all measures. Experimental design and procedure ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This study investigated the effect of alcohol on cognition in a randomized placebo- controlled pretest-posttest design. We used beer as experimental intervention, because it is available in alcoholic and non-alcoholic form and represents a highly popular drink among male and female University students. Participants of the alcohol group received Gösser Zwickl® (5.2% alcohol by volume) and participants of the non-alcoholic group received Gösser Naturgold® (<0.5% alcohol by volume). These two beers were selected because they are very similar in taste and visual appearance (i.e., golden, naturally cloudy). The amount of beer was individually adjusted for weight, age and gender (Watson, Watson, & Batt, 1980; Widmark, 1932) targeting at a blood alcohol concentration (BAC) of 0.03 (in case of alcoholic beer). For a male with 22 years, 75 kg and 182 cm this results in about 500 ml of beer, whereas a female with 22 years, 65 kg and 165 cm received about 350 ml beer. Drinks were cooled and prepared in a separate room and served in neutral drinking glasses. Participants were asked not to consume alcohol or other drugs 24 h before the experiment and not to eat and drink caffeinated drink 2 h before the experiment. They were tested in groups of two people, who were randomly assigned to the experimental groups (1 alcohol and 1 placebo). The participants were blind to the experimental condition, and the employed group setting was intended to avoid any experimenter effects. After a general instruction, participants signed informed consent. A first measurement of alcohol concentration with the alcohol tester ACE Neo (ACE Handels- & Entwicklungs GmbH; Freilassing, Germany) confirmed that all participants were sober at the start of the experiment. In the pretest, participants completed an executive function test (2-back task), two measures of creative potential (Remote Associates Test and divergent thinking test), and a test of creativity evaluation skills. Then participants were given their drinks and they watched a documentary about South Africa for half an hour to allow the BAC level to reach its maximum. In the posttest, participants worked on different versions of the same tests as in the pretest. The BAC was measured right after the executive function test, with the result being concealed from the participants. Additionally, all participants were asked indicate their subjective level of alcohol intoxication on a four-point rating scale (0 = not at all, 1 = a little, 2 = quite a bit, 3 = very much). After the experiment, the participants were informed about the existence of the placebo-control group and what group they had been assigned to. The total experiment took about two hours. The procedure had been approved by the ethics committee of the local university. Self-reported alcohol use The alcohol use disorders identification test (AUDIT; Saunders et al., 1993) is a 10-item screening of individual drinking behavior. It asks for the frequency of alcohol consumption, and for indicators of alcohol dependence and harmful alcohol use. Total scores of 8 or more are recommended as indicators of potentially hazardous and harmful alcohol use (Babor, Higgins-Biddle, Saunders, & Monteiro, 2001). Executive control Executive control was measured with a verbal 2-back task (Baddeley, 2003). The computer-based task presented a sequence of single white letters (K, M, R, or T) on black background with an inter-trial-interval of 1.5 s (inter-stimulus-interval = 0.2 s). Participants had to decide whether the current letter is identical with the one presented two stimuli ago by pressing a button for each target. Participants completed 20 practice trials, followed by the actual test with 100 trials (25% targets). The final score reflected the total number of correct responses to targets and non-targets. Creativity measures Creative thinking was assessed with the Remote Associates Test (RAT; Mednick, 1962) and divergent thinking tasks (Guilford, 1967), two common measures of creative potential (Kaufman, Plucker, & Baer, 2008). The RAT presents three unrelated words (e.g., cottage, blue, cake) and ask for a solution word that provides an unexpected connection between them (cheese: cottage cheese, blue cheese, cheesecake). Pretest and posttest used different sets of 10 items with increasing difficulty, but matched for item difficulty across tests (Landmann et al., 2014). The items were presented on a computer and participants entered the solution via a keyboard (timeout: 30 s). Divergent thinking (DT) was assessed with a computer-based version of the alternate uses task, which asks to find creative uses for common objects within 2.5 min (pretest: umbrella, shoe; posttest: car tire, fork). Task performance was scored for rated creativity as well as for fluency, flexibility, and novelty. All responses were evaluated for creativity (a holistic rating reflecting the novelty and usefulness of ideas) by six independent judges on a 4-point Likert-like scale ranging from (0 = uncreative, to 3 very creative). Inter-rater-reliability ranged from 0.75 to 0.81. DT creativity was defined as the average creativity evaluation of the three most creative ideas per task according to judge’s ratings. This top-3 creativity score was shown to avoid confounds with the fluency of responses (Benedek, Mühlmann, Jauk, & Neubauer, 2013). DT fluency reflected the number of ideas per task. DT flexibility was defined as the number of categories into which ideas fall. To this end, two judges classified all responses for the 28 categories provided by the manual of the Torrance Test of Creative Thinking (Torrance, 1974). The inter-rater-reliability of the flexibility assessments ranged from 0.86 to 0.95. Finally, DT novelty was defined as the average statistical infrequency across ideas. We computed the novelty of each idea as the inverse of its frequency within our sample (e.g., ideas generated by five people get a score of 0.2, whereas unique ideas get a score of 1). For all DT scorings, pretest and posttest scores represented the average performance in the two respective DT tasks. Creativity evaluation skills were assessed with a short-version of the creativity evaluation test (CET; Benedek et al., 2016). The CET presents different ideas that have to be judged as either common, inappropriate, or creative. Short versions for pretest and posttest were defined by selecting 8-item task blocks from the original CET. CET performance was scored by the informedness of judgements, a standard index from signal detection theory that equally accounts for sensitivity and specificity of judgements (Powers, 2011). Manipulation check and control analyses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The alcohol and the placebo groups each consisted of 35 participants, with a similar ratio of females (49% and 60%, respectively; χ[1] = 0.92, p = 0.47). Participants who had consumed alcoholic beer had an average BAC of 0.026 (SD = 0.08), which is close to the targeted BAC level of 0.03. As expected, participants who had consumed non-alcoholic beer had a BAC of 0 (SD = 0.00; t[66] = −19.5, p < 0.001). Importantly, however, the two experimental groups did not differ in the level of subjectively perceived intoxication (U = 711.5, p = 0.13); drinkers of alcoholic and non-alcoholic beer all predominantly indicated to feel a little bit intoxicated (both groups: mode = 1, median = 1). We observed no significant differences between gender groups: In the alcohol group, the BAC was largely similar between females (0.028; SD = 0.08) and males (0.025, SD = 0.08; t[33] = 1.64; p = 0.11). Moreover, gender groups did not differ in the subjective level of intoxication, neither in total (U = 565, p = 0.51), nor within experimental groups (alcohol group: U = 121.5, p = 0.30; placebo group: U = 149.5, p = 0.93; all modes and medians = 1). Effects of alcohol intoxication on cognition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We analyzed effects of time (from pretest to posttest, i.e., T1 to T2) and experimental group (alcohol versus placebo) on cognitive performance with two-way mixed ANOVAs. In the 2-back task, participants showed practice effects (time: F[1, 68] = 20.81, p < 0.001, partial-ɳ2=0.23), but this effect tended to be moderated by the experimental group (interaction: time ∗ group: F[1, 68] = 3.95, p = 0.05, partial-ɳ2 = 0.06): only the placebo group significantly increased their performance from T1 to T2 (t[34] = −4.40, p < 0.001), whereas the intoxicated group showed only a trend towards improvements (t[34] = −1.93, p = 0.06; see Fig. 1); yet, the groups did not differ significantly at T1 (t[68] = −1.37, p = 0.18) or T2 (t[68] = −0.43, p = 0.57). In the Remote Associates test (RAT), participants also showed increased performance at T2 (time: F[1, 68] = 24.47, p < 0.001, partial-ɳ2 = 0.27), and this effect was again moderated by experimental group (time ∗ group: F[1, 68] = 4.76, p = 0.03, partial-ɳ2 = 0.07): groups did not differ in performance at T1 (t[68] = 0.14, p = 0.89), but following the intervention, intoxicated participants showed higher RAT solution rates than those from the placebo control group (t[61.82] = −2.50, p = 0.02, d = 0.59; see Fig. 1). Regarding divergent thinking (DT) performance, we observed no time effects and, importantly, no significant moderation by experimental group for all DT scores, including DT creativity (time: F[1, 68] = 2.13, p = 0.15; time ∗ group: F[1, 68] = 1.73, p = 0.19; see Fig. 1), DT fluency (time: F[1, 68] = 3.77, p = 0.06; time ∗ group: F[1, 68] = 1.50, p = 0.22), DT flexibility (time: F[1, 68] = 1.01, p = 0.32; time ∗ group: F[1, 68] = 0.23, p = 0.63), or DT novelty (time: F[1, 68] = 0.89, p = 0.35; time ∗ group: F[1, 68] = 0.55, p = 0.46). Finally, creativity evaluation performance showed a significant time effect (time: F[1, 68] = 5.78, p = 0.02, partial-η2 = 0.08; T1: M = 0.53, SD = 0.49; T2: M = 0.69, SD = 0.29), but no significant interaction with experimental group (time ∗ group : F[1, 68] = 2.17, p = 0.24).","This study investigated the effect of low alcohol intoxication on creative cognition in a placebo-controlled design. Consumption of a low dose of alcohol tended to impair executive control, but facilitated creative problem solving (viz. insight problem solving), and did not affect divergent thinking ability. These findings replicate and extend recent research in this field. For example, Jarosz et al. (2012) also reported reduced working memory capacity and increased performance in the Remote Associates Test after moderate alcohol consumption. As a difference, Jarosz and colleagues did not include a placebo control group, because this may not be credible for moderate to high doses of alcohol (Martin & Sayette, 1993). However, assuming to have consumed alcohol can lead to higher evaluations of one’s creative output, and potentially even to more original performance (Lapp et al., 1994). Therefore, this study used a lower dosage (target BAC of 0.03, versus 0.07 in Jarosz et al., 2012), which effectively induced similar levels of subjectively experienced intoxication in alcohol and placebo groups. We found that intoxicated participants actually improved their performance in the RAT beyond the level of the placebo group, which rules out the possibility that performance gains were caused by expectancy effects alone. This study thus replicated effects of alcohol on creative problem solving for a lower BAC level, while more clearly attributing them to the (additional) pharmacological effect of alcohol. Alcohol intoxication neither improved nor harmed divergent thinking (DT) performance. This is consistent with some previous findings (Lang et al., 1984), while other studies found reduced idea fluency, flexibility, and increased originality at higher levels of intoxication (Gustafson, 1991; Norlander & Gustafson, 1998). Given the important role of executive control for divergent thinking ability (e.g., Benedek et al., 2014), it seems possible that negative effects on controlled processes are compensated by positive effects on spontaneous processes at low levels of alcohol, but one might expect reduced DT creativity for higher alcohol levels. We also observed no effect of alcohol intoxication on creativity evaluation skills, which is in line with the idea that intoxicated people do not differ in their creativity evaluation from people who assume to have consumed alcohol (Lang et al., 1984). Finally, it is interesting to mention that the employed N-back task proved to be a sensitive indicator of mild intoxication effects on cognitive control. Previous studies have often used complex span tasks (Colflesh & Wiley, 2013; Jarosz et al., 2012), but findings do not always align between these tasks (Redick & Lindsey, 2013). Therefore, it is useful to note N-back and complex span tasks both appear to reliably capture effects of mild to moderate alcohol intoxication on cognitive control. Considering the findings together, a moderate attenuation of cognitive control seems to benefit creative problem solving but not divergent thinking. Creative cognition is generally assumed to rely on the interaction of controlled and spontaneous processes, but the optimal balance may differ between tasks (Benedek & Jauk, in press). While most cognitive activities usually benefit from high cognitive control, some may actually suffer from too much focus (Chrysikou, Weber, & Thompson-Schill, 2014; Radel, Davranche, Fournier, & Dietrich, 2015; Wiley & Jarosz, 2012). Our findings suggest that high cognitive control is relatively more important to divergent thinking than to creative problem solving. This interpretation is consistent with the observation that creative problem solving tasks are often solved by spontaneous insight and accompanied by “Aha”-experiences (Kounios & Beeman, 2014). While attentional control typically supports the cued search of memory (Unsworth & Engle, 2007), looking for remote associations may benefit from by a slightly reduced attentional focus and more flexible integration of semantic concepts (Rowe, Hirsh, & Anderson, 2007; Wiley, 1998). Alcohol may particularly play a role in mitigating fixation effects. In creative problem solving, problems can often only be solved after a restructuring of the problem representation. When initial solution attempts get on the wrong track, this can cause blocks to immediate problem solving, which is known as mental fixation (Smith & Blankenship, 1991). These fixations typically fade with time, which is considered a central mechanism behind incubation effects (Storm & Koppel, 2012; Vul & Pashler, 2007). In a similar way, alcohol may reduce fixation effects by loosening the focus of attention and hence impeding the building and maintenance of dominant but inappropriate mental representations. Thereby, alcohol may facilitate a broader associative search and the effective solving of creative tasks that are prone to fixation effects.","Our study corroborates the notion that small attenuations of cognitive control may facilitate certain aspects of creative cognition while not affecting others. It contributes to our understanding of the interplay between controlled and spontaneous processes in creative thought and of their relative importance in different types of creative cognition. The findings, however, should not be overgeneralized by assuming that creativity is generally supported by alcohol. Beneficial effects are likely restricted to very modest amounts of alcohol, whereas excessive alcohol consumption typically impairs creative productivity (Kerr, Sheffer, Chambers, & Hallowell, 1991). Moreover, positive effects appear limited to specific phases in the creative process that are not fully tractable by goal-directed thought, while other phases such evaluation and implementation of ideas usually suffer from reduced cognitive control (Norlander, 1999). Our findings thus add to the increasing evidence suggesting that higher cognitive control is not always equivalent to better cognitive performance (Amer, Campbell, & Hasher, 2016; Beilock, Carr, MacMahon, & Starkes, 2002; Chrysikou, et al., 2014; Wiley & Jarosz, 2012)."],["Our daily lives involve high levels of repetition of activities within similar contexts. We buy the same foods from the same grocery store, cook with the same spices, and typically sit at the same place at the dinner table. However, when questioned about these routine activities, most of us barely remember the details of our actions. Habits are automatically triggered behaviours in which we engage without conscious awareness or deliberate control. Although habits help us to operate efficiently, breaking them requires great effort. We have developed a 27-item questionnaire to measure individual differences in habitual responding in everyday life. The Creature of Habit Scale (COHS) incorporates two aspects of the general concept of habits, namely routine behaviour and automatic responses. Both aspects of habitual behaviour were weakly correlated with underlying anxiety levels, but showed a more substantial difference in relation to goal-oriented motivation. We also observed that experiences of adversity during childhood increased self-reported automaticity, and this effect was further amplified in participants who also reported exposure to stimulant drugs. The COHS is a valid and reliable self-report measure of habits, which may prove useful in a number of contexts where discerning individuals' propensity for habit is beneficial. --------------------------------------------------------------------------------","Global challenges such as poverty, obesity and climate change require large parts of the general population to change the way we behave in order to make steps towards addressing these problems. Thus far, educational approaches and attempts to appeal to individuals' insight into the pressing need for change have largely failed (Webb & Sheeran, 2006). One reason for the lack of success may be that the targeted behaviours are largely habitual in nature, occurring outside conscious awareness. A better understanding of the mechanisms underlying habitual responses and individual variations in forming and breaking habits is needed in order to develop more effective strategies to address these global challenges (Marteau, Hollands, & Fletcher, 2012). Habits constitute response patterns that a person repeatedly exhibits in a specific situation (Lally & Gardner, 2013; Wood & Runger, 2016). These responses are learned and become automatically activated when the individual enters the associated environment. Examples could be making breakfast on coming into the kitchen in the morning, or putting the mobile phone onto charge when coming home from work. Such automatic responses are generally triggered by environmental cues, allowing us to perform routine actions highly efficiently whilst focussing our attention on other things. Meanwhile, the original motivation for these habitual actions becomes increasingly irrelevant and, once initiated automatically without intention, habits continue without conscious control. As habits are highly stable, they are difficult to change or break altogether. However, within a different environment (e.g. in a friend's kitchen), the same actions involved in making breakfast may suddenly run less smoothly, requiring conscious attention, and we may likewise run the risk of forgetting to charge the phone at the end of a day off work. Substantial experimental evidence has shown that habits develop though instrumental learning (Thorndike, 1898). The repetition of reinforced actions, if performed within the same environment, results in contextual stimulus-response associations in memory that trigger the behaviour automatically within that environment (Dickinson, 1985). These stimulus-response associations seem to overshadow the purpose that initially motivated the behaviour, rendering the behaviour insensitive to changes in the value or the contingency of the consequences. When habits are formed, control over the behaviour gradually shifts away from being guided by our intentions to being automatically triggered by cues in the environment. Consequently, once formed, habits are no longer motivated by a goal, and are thus difficult to break with goal-oriented intentions or knowledge of the consequences of habitual actions (Wood & Neal, 2007). There is significant variation in the degree to which different individuals show a propensity for developing habits. While some people delight in novelty and change in their lives, others go so far as to even describe themselves as ‘creatures of habit’, an expression that reflects their appreciation of routine and regularity in their lives. What underlies these individual differences in habit formation is still largely elusive, but may provide important insight into differences in the strategies needed to change habits in different people. Nevertheless, a number of factors have already been identified that can influence the switch of initially goal-directed actions into habitual responses. These include prolonged practice (Boakes, 1993; Dickinson, Balleine, Watt, Gonzalez, & Boakes, 1995; Neal, Wood, & Quinn, 2006), experiences of acute or chronic stress (Dias-Ferreira et al., 2009; Schwabe & Wolf, 2011), or exposure to stimulant drugs (Nelson & Killcross, 2006; Corbit, Chieng, & Balleine, 2014). By contrast, strong executive functions seem to promote goal-directed behaviours (Otto, Raio, Chiang, Phelps, & Daw, 2013), and possibly facilitate the regain of control over behaviours that have become habitual. A core question meriting consideration revolves around the extent to which ‘creature of habit’ traits might represent a vulnerability marker for the development of clinical conditions in which habitual behaviours have spiralled out of control, such as drug addiction, gambling, obsessive-compulsive disorder, or eating disorders. Indeed, a number of mental health problems involve rigid and inflexible routines, and actions performed in response to particular triggers regardless of negative consequences (American Psychiatric Association, 2013). Clarification around a role for proneness to habits in these conditions may shed light on more successful treatments than those currently available. Habitual behaviour can be assessed by experimental paradigms that manipulate the value or contingencies of the outcome to identify behaviour patterns that persist irrespective of such manipulations (Ersche et al., 2016; de Wit, Niry, Wariyar, Aitken, & Dickinson, 2007; Mckim, Bauer, & Boettiger, 2016; Gillan et al., 2013). Evidently, self-report measurements of behaviours that largely occur without awareness is not without criticism (Sniehotta & Presseau, 2012). The self-reported habit index (SRHI) is one of the few questionnaires that evaluate individuals' perception of a particular behaviour with respect to frequency, automaticity, efficiency, and self-reference using a 12-item rating scale (Verplanken & Orbell, 2003). The focus of the SRHI lies on a specific recurring behaviour that has been identified by the researcher, not by the scale. This presents a major drawback of the SRHI, as it excludes individuals who, due to a different lifestyle, do not engage in the behaviour in question. To the best of our knowledge, there are currently no tools available to assess more generally how individuals differ in their engagement in habits in daily life. The aim of the present study, therefore, was to develop a scale that reflects variations in individuals' tendencies towards responding in a habitual manner in everyday life. Variations in proneness to habit may be driven by a need for structure and predictability, which may reassure anxious individuals who worry about uncertainty and the possibility of things going wrong in novel situations (Evans et al., 1997; Connors, Bisogni, Sobal, & Devine, 2001). We therefore hypothesized that increased habitual tendencies are associated with higher levels of anxiety and obsessive-compulsive traits. Conversely, sensation-seeking traits and goal-striving personalities are likely to run counter to regularity and repetition (Dunn, 2000). We therefore predicted that low levels of habitual behaviours in daily life are associated with high levels of sensation-seeking and goal pursuit. An ancillary aim of the study was to examine whether exposure to stress or stimulant drugs, which have been shown to promote habitual responding in experimental settings, also affect participants' self-reported habitual tendencies. Scale development ~~~~~~~~~~~~~~~~~ For the first step in developing a scale measuring characteristic behaviours and attitudes for ‘creature of habit’ traits, we generated a pool of 59 items based on a thorough review of the literature, and interviews and discussions with experimental and health psychologists. On compiling the questionnaire, we noticed that half of the generated items related to tendencies describing regular behaviours (e.g. I park my car always in the same place), mental attitudes surrounding the minimisation of effort (e.g. I quite happily work within my comfort zone), or the establishment of safety/predictability (e.g. I rely on what is tried and tested), as well as emotional reactions when faced with irregularity (e.g. I hate it when the grocery store re-arranges the aisles). The other half of the items were behaviours occurring in the context of eating, such as describing behaviour motivated by preferences (e.g. I have a preferred sandwich), automatic responses (e.g. I always follow a certain order when preparing a meal), and behaviours characterised by a lack of planning (e.g. I tend to cook more than I eat). Participants were required to indicate for each statement their level of agreement on a 5-point Likert scale, ranging from strongly disagree (1) to strongly agree (5). We extensively piloted the questionnaire within the local community and conducted face-to-face interviews about the meaning of the items. Items that were consistently misunderstood were either reworded or removed. Further piloting showed that administering the entire 59-item questionnaire presented a challenge to participants, so that we subsequently divided it into two parts. Although the categorisation of general habits and food-related habitual responses was initially unintended, it provided a rationale for splitting the COHS into two parts with similar numbers of items (see Appendix A). Study sample ~~~~~~~~~~~~ We used Amazon's Mechanical Turk (MTurk), a crowdsourcing internet marketplace, to collect data from 406 individuals in the online community. Forty-four participants (11%) were excluded due to either incompletion, invalid responses or duplication of data, leaving a total sample of 362 participants (47% male), whose identity remained anonymous to the research team. Participants had to be at least 18 years of age [mean age 39.7 years ± 11.5 standard deviation (SD)] and based in the United States of America. All participants received $2.00 for completion of the study, which included the two parts of the COHS with the items in each part being presented in random order, and a selection of validated questionnaires to assess personality traits of anxiety, compulsivity, sensation-seeking, and goal-pursuit. We also collected background information, including ethnicity, native language, education level, and employment status. Moreover, we asked participants to indicate whether they have ever had any experience with stimulant drugs (either for recreational purposes or as medication) and to complete the Childhood Trauma Questionnaire (CTQ, Bernstein et al., 2003). The characteristics of the full study sample and the subgroups are shown in Table 1. As recommended by Meade and Craig (Meade & Craig, 2012), we also included two attention check items to safeguard against careless participants. The study was approved by the Psychology Research Ethics Committee (Pre.2015.124; PI:KDE). Personality measures ~~~~~~~~~~~~~~~~~~~~ Anxiety personality traits: The trait version of the Spielberger State-Trait Anxiety Inventory (STAI, Spielberger, Gorsuch, Lushene, & Jacobs, 1983) assesses variations in trait anxiety of a long-standing nature. It consists of 20 questions surrounding worry, tension, apprehension, and nervousness that are rated on a 4–point scale ranging from almost never (1) to almost always (4). Obsessive-compulsive Personality Traits: The Obsessive-Compulsive Inventory–Revised (OCI-R, Foa et al., 2002) is an 18-item questionnaire to assess obsessive-compulsive symptoms in both clinical and non-clinical samples. Participants rate the degree to which they have been bothered or distressed by obsessive-compulsive symptoms in the past month on ranging from not at all (0) to extremely (4). Sensation-Seeking Personality Traits: The Sensation-Seeking Scale Form-V (SSS-V, Zuckerman, 1996) is a widely-used psychological instrument for measuring individuals' need for novel and complex sensations along with the willingness to take risks for the sake of having such experiences. It is composed of a series of 40 pairs of dichotomous statements from which one must be selected. Goal-Oriented Personality Traits: The Habitual Self-Control Questionnaire (HSCQ, Schroder, Ollis, & Davies, 2013) is a 14-item instrument measuring variations in people's commitment to completing tasks and their self-perceptions around their drive for goal-pursuit. Participants express their level of agreement to each item on a 5-point Likert scale, ranging from disagree strongly (1) to agree strongly (5). Psychometric analysis of the creature of habit scale (COHS) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We first applied graph theory networks (Harary, 1969) to the correlation matrix to visualise the structure of the item pool. We then used the Mokken Homogeneity Model (MHM, Mokken, 1971), a non-parametric method based on Item Response Theory, to identify redundant items, meaningful subscales and summary scores (Stochl, Jones, & Croudace, 2012). The goal of this method is to cluster the items into scales that meet the three key assumptions of MHM: unidimensionality, monotonicity, and local independence. The summary scores of the items within a subscale should then allow to order people with respect to severity of the measured trait. Unidimensional subscales from the original item pool (meaning that all items measure the same latent trait) were identified using Loevinger's scalability coefficients (Loevinger, 1947) to assess scalability of a single item in relation to the other items of the scale as expressed by Hi, and the scalability of the total scale, as expressed by H. To meet the unidimensionality assumption of MHM, none of the Hi's should drop below 0.30 (Hemker, Sijtsma, & Molenaar, 1995). Monotonicity was assessed using methods available in R package mokken (van der Ark, 2007, 2012). Local independence assumption states that an individual's responses to items are independent of each other and is usually only examined by critical review of item wording. Items violating any assumption of MHM were discarded from the final set. Structure and construct validity of the final set of items in the COHS was subsequently assessed by confirmatory factor analysis. Reliability of subscales was assessed by McDonald's omega (McDonald, 1999). For reasons of convention, we also computed Cronbach's alpha (Cronbach, 1951) and estimated the graded response model (Samejima, 1969) in order to examine measurement error in detail. Finally, we evaluated convergent and discriminant validity using Pearson's correlations with personality traits that have been linked to variations in habitual tendencies. The aforementioned analyses were conducted using R (R Core Team, 2016) and the packages qgraph (Epskamp, Cramer, Waldorp, Schmittmann, & Borsboom, 2012), mokken (van der Ark, 2007, 2012) and ltm (Rizopoulos, 2006). Effects of childhood adversity and stimulant drug exposure ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Depending on whether participants had been exposed to stimulant drugs, they were allocated into two categorical subgroups (no stimulants, stimulant exposure), since no further information with respect to amount and duration of use was available. We also divided the sample based on the CTQ scores into participants with no or mild type of adversity experiences and those with moderate to severe adversity experiences, following the method described by Bradley et al. (2008). Possible adverse experiences included verbal assault, humiliation, intimidation, domestic violence, or sexual abuse. We used t-tests and chi- square to explore differences in demographics between the subgroups. We also assessed moderating and mediating effects of stimulant drug use between childhood adversity and the two COHS scales (automaticity, routine). Gender, ethnicity and education were included as covariates to control for subgroup differences in these variables. Data structure and item correlations of the COHS ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Item correlations and structure are shown in Fig. 1 and suggest that items can be roughly divided into two clusters. Mokken scaling indicates that 39 of the 59 items grouped into seven subscales and 20 items failed to either cluster with any subscale or create a subscale on their own. The largest subscale consisted of 16 items, describing behaviours and attitudes that favour order, familiarity and regularity, which we consequently termed routine. The second largest subscale included 11 items, creating a relatively separate cluster of items describing automatic behaviour patterns, which we consequently termed automaticity. Although the items within the remaining scales 3–7 were strongly correlated, each scale only consisted of just two to three items, which was considered insufficient for a meaningful measurement (Thissen & Wainer, 2001). These items were therefore discarded from the final set, in addition to the 20 items that did not fit into any subscale. Psychometric properties of the COHS ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 2 provides additional results from Mokken Scale Analysis on the assessment of homogeneity of subscales. For both scales, homogeneities (H coefficients) were calculated at 0.35 (routine) and 0.41 (automaticity), suggesting medium strength (Sijtsma & Molenaar, 2002). We did not detect any significant violations of monotonicity in either scale, suggesting compliance with the Mokken MHM for all items. Consequently, we ascertained that individuals should be assessed on the basis of the score of each scale separately, rather than by an overall score of both scales (Molenaar, 1982). Both dimensionality of the final item set and construct validity of the items were evaluated using confirmatory factor analysis. Considering the results of the Mokken Scale Analysis and network approach, we hypothesized a 2-factor (correlated) structure for the final item set. The path diagrams of the model are shown in Fig. 2. Model fit was good (Comparative fit index = 0.966, Tucker-Lewis index = 0.963, Root Mean Square Error of Approximation = 0.057 (95% CI 0.051–0.062)). All standardized factor loadings were significant and, with the exception of item COHS_16, higher than 0.5, suggesting good construct validity of items. Factor analysis also confirmed that the two scales, routine and automaticity, were only weakly correlated domains (r = 0.14, p < 0.001), which further justified separate scoring of the two scales. Reliability, measurement error, and validity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ McDonald‘s omega and Cornbach's alpha for the COHS routine (ω = 0.92, α = 0.89) and automaticity (ω = 0.91, α = 0.86) suggest satisfactory reliability for both scales. Detailed examination of measurement error can be found in Fig. 3, showing that both scales have reasonably small measurement error across a wide range of score distributions. As shown in Table 3, the COHS routine scale was significantly inversely correlated with sensation-seeking (SSS-V) and weakly correlated with trait levels of anxiety (STAI-T) and compulsivity (OCI-R). COHS automaticity, on the other hand, showed a significant inverse relationship with goal-pursuit (HSCQ), in addition to weak correlations with trait-anxiety and compulsivity. We further correlated both subscales with age to examine whether older individuals show a preference for routines (Bergua et al., 2006), but the results were non-significant (both p > 0.1). Effects of childhood adversity and stimulant drug exposure ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Participants who reported previous exposure to stimulant drugs were predominantly male (χ2 = 5.1,p = 0.024) and of white ethnic origin (χ2 = 11.2,p = 0.047), but were not different from stimulant-naïve participants in terms of education level or employment status. The ANCOVA model, which included gender and ethnicity as covariates, did not reveal any significant subgroup differences in terms of routine behaviour (F1,358 = 0.1,p = 0.921) or automaticity (F1,358 = 0.8, p = 0.365). For childhood adversity, significantly more women than men reported traumatic experiences during childhood (χ2 = 9.7,p = 0.002). Participants without adverse childhood experiences were more likely to attend university compared with participants who had such experiences (χ2 = 9.0,p = 0.011), but they did not differ with respect to ethnicity or employment status. When the two subgroups for childhood adversity were compared, with gender, ethnicity and education included as covariates, participants with childhood adversity showed significantly higher levels of automaticity compared with their counterparts without such traumatic experiences (β = 0.126,p = 0.019) (Fig. 4). The two subgroups did not, however, differ in terms of routine behaviours (β = − 0.018,p = 0.746). The comparison of participants with and without stimulant drug exposure, again accounting for covariates, revealed no group differences either for automaticity (β = − 0.001,p = 0.982) or routine (β = − 0.049,p = 0.362). However, we did identify a significant interaction effect between stimulant drug exposure and childhood adversity on automaticity (β = 0.172, p = 0.035) and an interaction effect at trend level for routine (β = 0.145,p = 0.082) suggesting that stimulant drug exposure further exacerbated the increased levels of automaticity observed in those individuals with childhood adversity. We did not find any mediation effect of stimulant drug exposure on the relationship between childhood adversity and either automaticity (total standardized indirect effect = − 0.004,p = 0.671) or routine (total standardized indirect effect = − 0.007, p = 0.426).","We present a novel questionnaire to assess individual variations in habitual tendencies in everyday life. The COHS is theoretically sound and has shown good psychometric properties. As identified by mokken scaling, the COHS differentiates two distinct features of habits: routine behaviours and automatic responses. Questionnaire items that were associated with volition, such as personal preferences or lack of planning, did not survive the analysis, further supporting the notion that habits are not mediated by a goal (Wood & Neal, 2007). The two scales for routines and automaticity both constitute forms of habits: they share with habits their implicit nature, in the sense that within daily life, they may both be initiated without conscious awareness. Habits are behaviours learned by repetition, not by insight (Boakes, 1993) - another critical feature shared by both routines and automaticity. The regular nature of routines involves repetition, and repeated practice is likewise a prerequisite for the transformation of purposeful behaviour into an automatism (Orbell & Verplanken, 2010). Yet, routine and automaticity also distinctly differ from one another, specifically in terms of their function and control over behaviour, as reflected by the weak correlation between the two COHS scales. This differentiation is further supported by the influence that adverse life events and stimulant drug exposure appear to have on the degree to which individuals express automaticity, whilst not affecting routine behaviours. How routines and automaticity relate to habits ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Routines are generally described as familiar action patterns that involve regularity, which are likely to be performed on a daily basis (Gallimore & Lopez, 2002). Routines have a relatively fixed temporal pattern of sequenced actions that are executed voluntarily in order to make daily life more orderly and efficient (Clark, 2000). This functional purpose and their relative independence from the immediate environment are critical features of routines, which may explain why routines are maintained as long as they deliver the desired outcome. Once this is no longer the case, e.g. due to changing circumstances, a conscious decision is typically made to either drop or amend the routine (Clark, 2000; Denham, 2002). Likewise, meaningful routines may also be deliberately combined and turned into rituals in order to increase cohesion between members of a group (Doherty, 1997; Denham, 2003). Although routines fall within the realm of habits, their inherent functionality (either implicit or explicit) and their strong link with a time frame, may explain why routines and habits are not synonymous. Habits thus only include those routine behaviours that are performed automatically without serving a specific purpose and are no longer restricted to a fixed temporal pattern. One possibility is that the degree to which people are inclined to apply routines in their daily lives might be one factor underlying their proneness to form habits. Personality traits associated with seeking or avoiding situations of novelty and excitement may further explain the spread of individual variations in routine practices (Dunn, 2000). Clearly, individuals with high sensation- seeking traits would be less inclined to engage in rituals, which would make them feel bored and stuck in a rut, whilst those more prone to anxiety in novel situations may relish the predictability and comfort of routine. Just as routine behavioural patterns can reduce effort and cognitive load, so can automatic behaviours, thereby enhancing functional efficiency of daily activities. However, in contrast to routines, automaticity neither has to be sequential in nature nor does it involve any kind of deliberation, cognitive direction or dependency on utility of the outcome (Saling & Phillips, 2007). In fact, automatic actions are initiated by environmental cues without a deliberate intention, and they may even continue without the involvement of conscious control (Bargh, 1994). Habits, by contrast, extend beyond simple automatic reactions, involving complex patterns of behaviours that are performed repeatedly and relatively automatically with very little variation (Clark, 2000). Environmental cues not only trigger the behavioural response but also the mental representation of the habit (Aarts & Dijksterhuis, 2000; Wood & Runger, 2016). This may also explain the inverse relationship between automaticity and goal-striving in the COHS. As such, goal-striving individuals manage their behaviour with consideration for the consequences of their actions in mind, and are less inclined to allow environmental stimuli to take over control. Given that automatic actions are dependent on the context that triggers them, automatic action patterns necessarily vary enormously with respect to individuals' lifestyles. This poses a particular challenge for self-report measures, which require questionnaire items that capture the specificity of the environment while applying to a large number of people. As the consumption of food is largely automatic (Cohen & Farley, 2008), it is not surprising that many automaticity items are related to eating. Yet, this should not detract from the fact that other non- food related behaviours are equally likely to be automatic but are just more difficult to assess by questionnaire. It is of further note that eating habits also involve routines, as traditionally exemplified by the eating pattern of three meals a day (Jastran, Bisogni, Sobal, Blake, & Devine, 2009). The involvement of the routine component of habit in food- related behaviour is also captured by a number of items in the routine subscale of the COHS. The benefits and risks of habits ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A major hallmark of habits is that they are not mediated by a goal (Dickinson, 1985; Wood & Neal, 2007). Although goal pursuit might have motivated the behaviour initially, once the habit has been acquired, the goal is no longer required to motivate or guide the actions, thereby making habits less effortful and cognitively demanding than goal- orientated behaviours. Habits thereby allow individuals who are affected by cognitive decline or impaired motivation, such as older adults or drug-addicted individuals, to behave highly efficiently despite their deficits (Bergua et al., 2006; de Wit, van de Vijver, & Ridderinkhof, 2014; Ersche et al., 2016), albeit at the expense of flexibility. Furthermore, as habits are decoupled from goal-orientated motivation, habitual behaviour is less susceptible to motivational urges, offering therapeutic opportunities for individuals who particularly struggle with resisting temptations and cravings (Lin, Wood, & Monterosso, 2016). This critical component of habits, namely their insensitivity to reinforcement contingencies or outcome, is, however, not captured by the COHS, but may be explored by studies using the COHS together with experimental paradigms. Importantly, although habits continue even though these actions are no longer needed, they do not necessarily continue in an uncontrolled manner; for example, they will not be carried out in exactly the same way when there is a change in the environment or temporal configuration of the situation (Lombo & Gimenez-Amaya, 2014). However, if habits lose the specific link to the context and co-occur with compromised inhibitory control, an increased risk arises of habitual patterns being repeatedly practised over and over again, becoming more and more deeply entrenched, which is the case in clinical conditions involving maladaptive habits such as obsessive-compulsive disorder and addiction (White, 1997; Mishkin, Malamut, & Bachevalier, 1984; Graybiel & Rauch, 2000). Accordingly, both COHS factors were positively correlated with OCI-R scores, a clinical measure of compulsive symptoms (Foa et al., 2002). Although the present sample was not drawn from a clinical population, the positive relationship suggests that both aspects of habits captured in the COHS lend themselves to susceptibility for development of compulsivity. The underlying causes for the development of compulsivity are still elusive, but investigation into them is of critical importance given their apparent relevance to a number of disorders. Neural substrates of the creature of habit components ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The differentiation of routine behaviour and automaticity, as assessed by the COHS, might also be reflected in distinct neural substrates underpinning different components of habits. Hitherto, the basal ganglia have been considered to play a key role in the formation of habits (Ashby, Turner, & Horvitz, 2010; Yin & Knowlton, 2006; Graybiel, 2008). During the initial phases when a behaviour is learned, the associative part of the striatum (caudate nucleus and anterior putamen) and the adjunct limbic structures and medial prefrontal cortex are critically involved (Doyon et al., 2009; Ashby et al., 2010). In as much as these brain systems are critical for both learning and memorising the instrumental contingencies as well as the differential effects of action-outcomes (Liljeholm, Tricomi, Doherty, & Balleine, 2011; Tanaka, Balleine, & O'Doherty, 2008), goals remain relevant during this phase and might still be involved in the development of routines. However, with prolonged practice, sensorimotor regions of the striatum (putamen) that are primarily connected to sensory and motor cortices take over control (Tricomi, Balleine, & O'Doherty, 2009), resulting in behaviour patterns becoming more automatic and eventually taken over by the cerebellum, possibly contributing to automaticity (Lang & Bastian, 2002; Doyon et al., 2009). Whether the processes for developing routines and automatic response are reflected in variations in connectivity in cortico-striatal and cortico-cerebellal pathways, respectively, are hypotheses to be tested using the COHS in combination with neuroimaging technology. Creature of habit: Moderator effects ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Participants who reported adverse childhood experiences scored significantly higher on the COHS automaticity scale than those who did not report such experiences. This effect was further exacerbated in individuals who also reported using stimulant drugs at least once in their lives. These results are in keeping with prior preclinical work suggesting that the development of automatic responses could be accelerated by exposure to stress and psychostimulant drugs. Our findings are consistent with preclinical work showing that both stress and stimulant drugs induce sensitisation of dopaminergic systems, thereby promoting a more rapid transition to stimulus-response action patterns (Dias-Ferreira et al., 2009; Nelson & Killcross, 2006; Nordquist et al., 2007). From a psychological perspective, a predominance of automatic responses under stress is clearly beneficial. As long as the context is stable, the individual may rely on cerebellar circuits, ensuring that complex actions are performed efficiently even under stressful conditions, (Doyon et al., 2009). In light of the role automaticity plays during stress, it is noteworthy that most of the COHS automaticity items are food-related, which might point to the stress-induced neuroendocrine functions that have also been linked with automatic eating patterns (Kandiah, Yake, & Willett, 2008; Tryon, DeCant, & Laugero, 2013). This further supports our finding of higher automaticity scores in people who report having experienced stressful childhoods.","The 27-item COHS is a measure of habitual tendencies in daily life that demonstrates good validity and reliability (Appendix B). The questionnaire has been designed to measure creature of habit traits, but the findings also point to sub-dimensions of habits related to automatic responses and routine behaviours. This differentiation sheds light on habitual behaviours and may inform interventions to break habits in populations with abnormal maladaptive habits. For example, it is conceivable that individuals with highly expressed routine behaviours would particularly benefit from training alternative routines to replace maladaptive habits. Conversely, for individuals with a particular affinity for automaticity, interventions would be more effective if they were to focus on breaking the maladaptive stimulus-response relationships by linking the triggering stimulus to a more desirable response and practising this new stimulus-response relationship. A shortcoming of the current questionnaire is that it does not assess the critical feature of habits concerning insensitivity to reinforcer devaluation – a feature that would be difficult to assess by self-report. Thus, further benefit could be derived from administering the COHS in combination with other diagnostic tools, specifically instrumental learning tasks. Moreover, cross-validation of the COHS in an independent sample is also warranted. Generally, the potential usefulness of this questionnaire may be better appraised when the measure has been used in further empirical studies. The use of neuroimaging methods may be of particular benefit in elucidating the different neural pathways involved. Specifically in populations with dysfunctional habit formation, the COHS may help to clarify which aspects of the construct are abnormal, and may then help to determine the most appropriate therapeutic strategy.","All authors declare that they have no competing or potential conflicts of interest in relation to this work."],["Objectives The decisions made by officials have a direct bearing on the outcomes of competitive sport contests. In an exploratory study, we examine the interrelationships between the decisions made by elite netball umpires, the potential contextual and environmental influences (e.g., crowd size), and the umpires’ dispositional tendencies – specifically, their propensity to deliberate and ruminate on their decisions. Design/Method Filmed footage from 60 England Netball Superleague matches was coded using performance analysis software. We measured the number of decisions made overall, and for home and away teams; league position; competition round; match quarter; and crowd size. Additionally, 10 umpires who officiated in the matches completed the Decision-Specific Reinvestment Scale (DSRS). Results Regression analyses predicted that as home teams’ league position improved the number of decisions against away teams increased. A model comprising competition round and average league position of both teams predicted the number of decisions made in matches, but neither variable emerged as a significant predictor. The umpire analyses revealed that greater crowd size was associated with an increase in decisions against away teams. The Decision Rumination factor was strongly negatively related to the number of decisions in Quarters 1 and 3, this relationship was driven by fewer decisions against home teams by umpires who exhibited higher Rumination subscale scores. Conclusions These findings strengthen our understanding of contextual, environmental, and dispositional influences on umpires’ decision-making behaviour. The tendency to ruminate upon decisions may explain the changes in decision behaviour in relation to the home team advantage effect. --------------------------------------------------------------------------------","In competitive sports, officials are required to make rapid and complex decisions, often in a highly pressured environment (Helsen & Bultynck, 2004). Moreover, their decisions often directly affect the outcome of competitions (Plessner & MacMahon, 2013). For example, during the final minutes of the 2015 Rugby World Cup quarter-final between Scotland and Australia, referee, Craig Joubert, decided to award a controversial penalty to Australia for a deliberate knock-on, resulting in a 35–34 victory for Australia, which enabled them to progress to the semi-final of the competition. Such decisions invariably attract negative evaluations by aggrieved players, coaches, spectators and the media, so the importance of consistent and impartial officiating is unquestionable (Stulp, Buunk, Verhulst, & Pollet, 2012). Decision-making can be influenced by a variety of factors (MacMahon et al., 2015), such as home advantage and crowd noise (e.g., crowd noise contribution to the home advantage effect, Nevill, Hemingway, Greaves, Dallaway, & Devonport, 2016; Unkelbach & Memmert, 2010), competition level (Souchon, Cabagno, Traclet, Trouilloud, & Maio, 2009; Souchon et al., 2016), reputation (e.g., expectation bias in gymnastics, Plessner, 1999) and time (e.g., decision accuracy and frequency thoughout games, Emmonds et al., 2015; Mallo, Frutos, Juárez, & Navarro, 2012). In the current paper, we employ an exploratory approach to examine the decisions made by netball umpires and the influences of contextual and environmental factors on the number of decisions made. Moreover, we investigate umpires' self-reported tendency to reinvest in, and ruminate upon, their decisions. Many researchers have focused upon the home advantage in sports – a phenomenon whereby there is an apparent advantage conferred to the home team. Four major determinants have been suggested to cause the home advantage effect namely, familiarity, territoriality, travel fatigue, and crowd noise (Pollard, 2008). It has been suggested that home advantage fluctuates throughout the game. For example, in basketball, Jones (2007) demonstrated that the home advantage (difference in points scored by the home and away teams) was greatest in the first quarter. In volleyball, home teams had a greater advantage at the beginning (1st set) and towards the end of the game (4th and 5th sets); this effect has been attributed to familiarity with the venues and crowd effects (Marcelino, Mesquita, Palao, & Sampaio, 2009). In relation to the referee's influence on the home advantage, Boyko, Boyko, and Boyko (2007) examined data from 5244 English Premier League soccer matches involving 50 referees. They found that referees differed in their susceptibility to the home advantage effect; hypothesising this was due to variations in the referees' ability to deal with social pressure. However, Johnston (2008) replicated Boyko et al.’s (2007) approach and found no evidence of such individual differences when removing referees who only officiated a few matches. To investigate this discrepancy further, Page and Page (2010) analysed footage from 37,830 national and international soccer matches across 58 competitions, between 1994 and 2007. Their analyses showed that not only did the size of the home advantage differ significantly between referees, but also, in line with Boyko et al. (2007), their decisions were moderated by crowd size – lending support to the notion that referees cope differently with the social pressure exerted by home crowds. Using a video-based protocol, Nevill, Balmer, and Williams (2002) manipulated crowd noise presence (“loud” or none) and found that soccer referees made more decisions in favour of the home team, and in line with the original match referee. Unkelbach and Memmert (2010) identified the inherent limitation of testing crowd noise (“natural conditions”) versus no crowd noise (“unnatural conditions”). The authors highlighted that Nevill et al.'s (2002) findings merely indicate that home crowd noise biases decisions compared to no crowd noise, rather than crowd noise influencing referee decisions in favour of the home team. Subsequently, Unkelbach and Memmert (2010) tested the hypothesis that louder crowd noise would lead to more yellow cards awarded compared to low crowd noise. Twenty referees viewed 56 foul scenes, in which 50% led to the award of a yellow card and 50% did not. The high-volume crowd noise led to substantially more yellow cards than low-volume crowd noise. Further evidence in soccer indicates that home teams were awarded more penalties (e.g., Nevill, Newell, & Gale, 1996; Scoppa, 2008; Sutter & Kocher, 2004), and fewer yellow and red cards (Buraimo, Forrest, & Simmons, 2010) with the size of the attending crowd moderating these effects (Boyko et al., 2007). The mediating effect of competition level has received scant attention, whilst stage of competition (e.g., Round 1, playoffs, finals, etc.) has yet to be investigated. Souchon et al. (2009) proposed that the level of competition is a stereotyping heuristic used by referees to form their decisions, interpreting fouls differently according to their preconceptions regarding the standard of play. Souchon et al. (2009) investigated this notion in handball (e.g., lower versus higher standard), predicting the level of competition effects would be greater for more difficult, ambiguous handball transgressions (“pushing offences”, opposed to clearer “holding back” offences) and anticipating that referees would be more lenient in higher-standard competition. They reported that referees intervened less frequently at higher levels of competition and allowed play to continue without intervention more frequently following more ambiguous transgressions (pushing offences compared to holding offences). Similarly, Souchon et al. (2016) observed that referees intervened less often when higher-level players transgressed. The authors suggested that a reduction in decisions made may be the culmination of a number of factors: referees trying to maintain the flow of a match; referees making fewer calls to maintain the game's value as a spectacle (e.g., Mascarenhas, O'Hare, & Plessner, 2006); that a greater number of fouls may be more ambiguous in high-level competition, due to the high speed of play; that greater levels of player aggressiveness may make it more difficult to identify transgressions; or that referees may assume that certain players can continue their actions despite the seriousness of the foul committed (e.g., gender stereotype and males superior physical ability, Souchon et al., 2010). In this study, we aim to examine potential changes in the number of decisions made across progressive competition rounds (perceived match importance arguably increases as the rounds progress). Few researchers have focused on the effect of the competing teams' abilities on sports officials' judgements. However, Plessner (1999) examined the idea of an expectation bias in team gymnastics, where gymnasts normally perform in a ranked order, worst to best. Plessner predicted that when the same routines, placed in either first or fifth position, will score higher when the judges view them in the latter position. Forty-eight gymnastic judges, with prior expectations of coaches' rank order of the gymnasts, judged videotapes of a men's team competition. Their results supported the notion of an ability expectation bias, whereby, for difficult tasks (e.g., pommel horse, vault, and horizontal bar) the judges awarded greater scores when the target routines were presented fifth than if they were presented first. Findlay and Ste-Marie (2004) explored athlete reputation bias in figure skating judgments. Twelve judges evaluated performance of 14 skaters, half of whom were known to the judges. The performance of skaters with a pre-existing positive reputation were scored more highly than those of the unknown skaters. It is possible that similar unconscious biases relating to perceived athlete ability may also exist in team sports; hence, we also took the competing teams' pre-eminence (i.e., their league position) into account in this study. To date, a limited body of research has investigated the effect of the match period on sports officials' decision-making. Mallo et al. (2012) assessed the soccer referees' decision quality and quantity in relation to match periods. Mallo et al. reported that a greater number of incidents occurred in the last 15- minute period of matches – but the lowest referee decision accuracy (77%) was also observed during this period. They suggested that physical and mental fatigue occurs during the final stages of a match leading to impaired decision-making. Similarly, Emmonds et al. (2015) found a drop in penalty judgement accuracy in rugby league referees in the last 10 min of matches. Conversely, Mascarenhas, Button, O'Hare, and Dicks (2009) reported that soccer referees were less accurate in the opening 15 min of each half than they were at any other period. They attributed poorer decision-making to warm up decrements, whereby their physical warm-up was not accompanied by any mental warm up techniques. Finally, Elsworthy, Burke, and Dascombe (2014) investigated decision-making demands of Australian Football referees, and reported that the number of free kicks awarded and free kick accuracy did not differ across each quarter of the match. Accordingly, in the present study, we analysed differences in the number of decisions made by netball umpires across each of the four match quarters. Published reports using qualitative methods have identified several sources of pressure and anxiety for sports officials (such as game importance, Hill, Matthews, & Senior, 2016; time, Morris & O'Connor, 2016; social pressure, Schnyder & Hossner, 2016). Morris and O'Connor (2016) found that National Rugby League (NRL) referees identified the time during a match as an influence on their game management strategies and decision-making ability. For example, one referee stated “certain decisions can have a greater impact at different stages in a game which can increase media scrutiny” (Morris & O'Connor, 2016, p. 854). Schnyder and Hossner (2016) interviewed high-level soccer referees regarding decision-making and the difficulties they face. Several of the referees identified social pressures, including pressure from the media, teams, football associations and even themselves. Hill et al. (2016) interviewed seven expert rugby referees and noted that avoidance coping behaviours were regularly employed to deal with multiple stressors that influence their performance including: unfamiliarity (e.g., new situations); performance errors (e.g., mistakes that ‘harm’ players, coaches and own career prospects); interpersonal conflict (e.g., manging player hostility); game importance (e.g., when the match outcome held significant consequence for players such as a final, or for themselves such as games close to renewal of contracts) and self-presentational concerns (e.g., fear of negative evaluation by selectors, avoiding criticism that could damage their confidence and reputation). The avoidance behaviours manifested themselves as denial after performance errors, rushing or withdrawal during the game, and a lack of preparation leading into games. Similarly, overt and maladaptive changes in behaviour under anxiogenic conditions have been observed in soccer (Jordet & Hartman, 2008) in climbing (Nieuwenhuys, Pijpers, Oudejans, & Bakker, 2008), dart throwing (Nibbeling, Oudejans, & Daanen, 2012), golf (Hill, Hanton, Matthews, & Fleming, 2010), and police arrest procedures (Renden et al., 2014). Decision avoidance has been described as “a tendency to avoid making a choice, by postponing it or by seeking an easy way out that involves no action or no change” (Anderson, 2003, p. 139). Selection difficulty has been identified as a major contributor to decision avoidance including factors such as: reasoning; preference uncertainty; attractiveness of options; attentional focus; time limitation; negative emotion (associated with blame and regret); and conflict type (Anderson, 2003). Researchers have shown that decision averseness occurs when situations have inequitable outcomes for others – particularly when the decision maker is held accountable (Beattie, Baron, Hershey, & Spranca, 1994); and the likelihood of negative outcomes also increases negative emotions associated with such decisions (Luce, Bettman, & Payne, 1997). In this study, we explored the notion that withdrawal of decisions (fewer decisions made) may be an example of decision avoidance behaviour. Several theories have been proposed to explain performance decrements under pressure. A prominent example is Reinvestment Theory (Masters, 1992). Reinvestment is defined as the “propensity for manipulation of conscious, explicit rule based knowledge, by working memory, to control the mechanics of one's movements during motor output” (Masters & Maxwell, 2004, p.208). Consequently, the use of explicit knowledge to consciously control normally automatic movements typically results in performance decrements or outright failure. Researchers have demonstrated that, when performing well-learnt motor skills or complex cognitive tasks, individuals who have a strong tendency to reinvest (as measured by the Reinvestment Scale, Masters, Polman, & Hammond, 1993) (as measured by the Reinvestment Scale) are more susceptible to poor performance under pressure (Jackson, Kinrade, Hicks, & Wills, 2013; Kinrade, Jackson, & Ashford, 2010). To address potentially differential effects of reinvestment on motor skill execution and decision-making, Kinrade, Jackson, Ashford, and Bishop (2010) modified the original scale to create a decision-specific version focusing on individuals' propensity to deliberate, and ruminate, on their decisions – the Decision- Specific Reinvestment Scale (DSRS). Kinrade et al. (2010) proposed two explanations for the breakdown of decision-making under pressure. First, that conscious processing of explicit information results in poor decision-making, by interfering with normal automatic processes (Decision Reinvestment; e.g., “I'm aware of the way my mind works when I make a decision”). Secondly, ruminative thoughts (e.g., over past poor decisions) lead to poor decision-making by drawing processing resources away from the task at hand (Decision Rumination; e.g., “I remember poor decisions I make for a long time afterwards”). Kinrade et al. (2010) described rumination as a thought process that typically involves repetitive negative thoughts about past events or current mood states. Higher decision reinvesters and ruminators tend to exhibit poorer working memory task performance (Laborde, Furley, & Schempp, 2015), and poorer decision-making performance in complex tasks (Kinrade, Jackson, & Ashford, 2015). Kinrade et al. (2015) suggested that ruminative thoughts may occupy working memory capacity at a time when executive functions are already in great demand to complete the primary task. Poolton, Siu, and Masters (2011) used the DSRS to examine soccer referees' susceptibility to the home advantage effect. Twenty-eight experienced referees were asked to make decisions when viewing game footage of two opposing players competing for the ball, by stating which player committed the foul. Referees that emerged as ‘high decision ruminators’ disproportionately made decisions in favour of the home team. We aim to explore this link further in the present study, in the context of netball officiating. In order to more fully understand contextual and dispositional influences on the decision-making of netball umpires, we used performance analysis to examine decisions made by umpires during matches in the England Netball Superleague – the highest echelon of competitive netball in the UK. We explored not only environmental and contextual influences such as crowd size, but also the umpires' self-reported tendency to reinvest in, and ruminate upon, their decisions. The number of decisions made provided an overt manifestation of the observed umpires' behaviour, a technique previously used to categorise observational data into approach- and avoidance-type behaviours (Jordet & Hartman, 2008). In accordance with previous research (Anderson, 2003; Hill et al., 2016; Jordet & Hartman, 2008; Nevill et al., 2002; Poolton et al., 2011; Souchon et al., 2016), we tentatively hypothesised that umpires' decision frequency would be mediated by environmental/contextual influences such as home team status, crowd size, match prominence, league position, and time during the match. More explicitly, we predicted that, home teams in the presence of larger crowds, greater match significance, more prominent teams, and early match quarters would each be associated with lower decision frequencies (i.e., avoidance behaviour). We also predicted that a tendency to reinvest and ruminate would be associated with inhibited decision-making.","Altogether, 15 umpires officiated in the Superleague during the 2014 season, umpiring approximately eight matches each (M = 8.067, SD = 3.77). From this original sample 10 umpires (M age = 39.6 yrs, SD = 9.38 yrs) with a mean total years' experience of 14.5 years (M = 14.5 yrs, SD = 7.66 yrs), qualified at international (International Umpire Award) or national level (A-award), completed the DSRS. On average, they officiated almost nine matches each throughout the season (M = 8.80, SD = 2.859). Data acquisition Video footage from sixty Netball Superleague 2014 season matches was obtained. Crowd size (number of people present in the crowd) data were collected from the individual teams for their home fixtures and from England Netball for all ‘neutral’ venues (i.e., those for which there was no home team). League table data for each round were obtained from England Netball. Approval was obtained from the lead institution's local ethics committee. Variables All coded variables were derived from discussions with a panel of experts (an England Netball Officiating Manager, a retired international umpire and assessor, a current national level umpire and tutor) and in accordance with variables previously shown to be pertinent with regard to sports officials' decision-making (e.g., match importance, Hill et al., 2016; Decision Rumination and the home advantage effect, Poolton et al., 2011). The primary dependent variable was the number of observable decisions made (NoD), split into three subcategories: overall; those against the home team; and those against the away team. Other coded variables included: infringement type (contact, obstruction, offside, breaking, out of court, and other infringement); and sanctions imposed (penalty pass, advantage, throw in, advantage goal, other sanction.). Additionally, we recorded six variables that were hypothesised to have a potential influence on umpires' decision-making: crowd size; competition round number (e.g., 1 = 1st round); league positions (of home teams, of away teams, and average; 1 = top of the league); and match quarter (e.g., Q1 = 1st quarter). Decision specific Reinvestment Scale Altogether, 10 umpires completed the Decision-Specific Reinvestment Scale (DSRS, Kinrade et al., 2010), a 13-item scale, comprising two subscales (Decision Reinvestment and Decision Rumination). Participants responded to each of the 13 items using a 5-point Likert scale anchored by 0 (“extremely uncharacteristic”) and 4 (“extremely characteristic”). The Decision Reinvestment subscale comprises 6 items, assessing the individual's propensity to consciously monitor their decision-making processes, with scores ranging from 0 to 24. The Decision Rumination subscale comprises 7 items, assessing tendency to negatively evaluate previous poor decisions, with scores ranging from 0 to 28. Kinrade et al. (2010) reported an internal consistency of 0.89 for the Decision Reinvestment subscale items and 0.91 for the Decision Rumination subscale items.","The matches were analysed using digital performance analysis software (Sportscode Elite Version 9, Sportstec, Australia). A self-devised code window was designed to collect the number of observable decisions, based on arm signals and vocalisations made by the umpires during the matches. Observable decisions were infringements that were registered and acted upon by the official by either a whistle blow or signalling advantage (this did not include time calls e.g., injury, blood). Also, umpires can decide not to interfere with play (Helsen & Bultynck, 2004) and these non-observable decisions were not recorded. Situations in which decisions were unclear were coded separately (accounting for 1.4% of total decisions made). Two researchers independently coded all the footage; intraclass correlation coefficients were used to test for inter and intra-observer reliability (ICC >0.90 for all). Data analyses ~~~~~~~~~~~~~ Preliminary screening of all data, using univariate z-scores (>± 3.29) and multivariate Mahalanobis distance values revealed one outlier from both the match and umpire data set which were removed. The data were normally distributed. A repeated-measures ANOVA was completed to compare differences in the NoD made across quarters. The relationships between contextual/environmental influences, dispositional tendencies, and decision-making were examined using two different analyses: one in which matches were treated as cases (n = 59), and another in which umpires were cases (n = 15 [all umpires] or n = 10 [DSRS completer's only, accounting for 72% of all matches, n = 42]). Pearson's product moment correlation coefficient was calculated for all bivariate combinations of the following variables in the match analyses: NoD; per match and per quarter; overall, in favour of home teams and in favour of away teams; crowd size; competitive round number; and home, and away team league positions, and their average. For the umpire analyses, bivariate correlations included total years of experience, Reinvestment, Rumination and number of games umpired. For the match-level analysis, all variables that were significantly related to NoD were entered as predictors into two stepwise multiple regression analyses and one linear regression, in which backward elimination was used in order to find a model that best explained the data. NoD, NoD Away, and NoD Home were the criterion measures for each of the three models. Alpha was set at 0.05 for all statistical tests. Due to the exploratory nature of the study, and accordingly tentative but directional nature of the hypotheses, we made no correction for multiple comparisons.","The descriptive statistics are presented in Table 1. On average, umpires made 120 observable decisions per game (M = 120.41, SE = 4.07). A repeated-measures ANOVA indicated that more decisions were made in the first quarter (M = 33.02, SE = 1.14) than in the third (M = 29.63, SE = 1.16) and fourth (M = 27.72, SE = 1.61) quarters, (F (3, 39) = 4.811, p = 0.006, ηp2 = 0.270). The most common infringement type was contact (M = 45.69, SE = 1.04), and the most frequently awarded sanction was a penalty (M = 48.77, SE = 1.37). Descriptive statistics revealed that DSRS scores ranged from 15 to 35 (DSRS Global M = 25.50, SD = 6.67), and Reinvestment subscale score from 7 to 16 (Reinvestment M = 12.8, SD = 2.82), and Rumination subscale score from 4 to 20 (Rumination M = 12.7, SD = 5.42). Total NoD All match-level bivariate correlations are presented in Table 2. NoD decreased as the average league position of the two teams increased (r = -0.269, p = 0.040); that is, the higher the positions of the two teams, the greater the NoD. Similarly, the higher the home team league position (NB: top position in the league = 1), the greater the NoD (r = -0.259, p = 0.047). As the teams progressed through the competition rounds, NoD increased (r = 266, p = 0.042). A backward stepwise regression was completed to identify the best predictors for NoD (variables entered: average league position, round, and home league position). The model that best predicted NoD included round and average team position (F (2, 58) = 3.919, p = 0.026, R2Adjusted = 0.091), although, when considered individually, neither predictor contributed significantly; they only approached significance (round p = 0.078, average team position p = 0.074) (see Table 3). NoD home NoD Home increased with the away team's league position (r = -0.340, p = 0.008). A linear regression indicated that away league position was a significant predictor of NoD (Home) (F (1, 54) = 6.255, p = 0.016, R2Adjusted = 0.089) (see Table 3). NoD away NoD Away increased as home teams' positions improved (r = -0.424, p = 0.001). As away teams progressed through rounds (r = 0.344, p = 0.008) or played in front of larger crowds (r = 0.312, p = 0.023) the NoD against them increased. A multiple regression was run to identify the best predictors for NoD Away (variables entered crowd size, round, and home league position) using the backward method. After the exclusion of crowd size and round, home team league position was shown to best predict NoD Away (F (1, 48) = 7.940, p = 0.007, R2Adjusted = 0.126). (See Table 3). Total NoD The total number of match decisions was not significantly correlated with any of the influences. As the average league position improved the number of decisions were greater (r = -0.573, p = 0.032). NoD home NoD Home increased as the competition progressed (i.e. later rounds, r = -0.618, p = 0.018) and the away team's league position became more prominent (r = -0.603, p = 0.022). NoD away As crowd size increased so did the NoD Away (r = 0.560, p = 0.037) (see Table 4). DSRS The correlations completed with the DSRS subscales include only the data from the ten umpires who completed the scale. The Rumination subscale score was significantly negatively associated with NoD Q1 (r = -0.795, p = 0.006), NoD Q3 (r = -0.709, p = 022), NoD Home Q1 (r = -0.717, p = 0.020) and NoD Home Q3 decisions (r = -0.660, p = 0.038); that is, higher Rumination subscale scores were associated with fewer decisions. Reinvestment subscale scores were not significantly correlated with any NoD variables.","In an exploratory study, we examined the influence of contextual and dispositional differences on decision-making of umpires in actual match settings. We hypothesised, based on existing literature, that environmental and contextual influences (i.e., larger crowds, more prominent teams, greater match significance, and early quarters) would be associated with lower decision frequencies. Furthermore, we predicted that inhibited decision-making would be associated with a dispositional tendency to reinvest and ruminate. In line with our hypotheses, match prominence and league position were associated with a reduction in the number of decisions. The Decision Rumination factor was linked with inhibited decision making; but contrary to our hypothesis, the Reinvestment factor was unrelated. In contrast to our hypotheses, increasing crowd size was associated with a greater number of decisions, particularly against away teams; and the number of decisions diminished throughout a match. Our data indicated that more decisions were made in Q1 (33 decisions) than in Q3 (29 decisions) and Q4 (27 decisions), incongruent to our hypothesis and the findings by Mallo et al. (2012) and Elsworthy et al. (2014). These differences could be related to physical fitness and fatigue of umpires; for example, Paget (2015) found that the distance covered by netball umpires was significantly reduced in the fourth quarter. It is possible that, if umpires are physically fatigued and not covering the same distances as they did in the early stages of a match, the fewer decisions later in the game could be those missed or avoided as a result of incorrect positioning. Multiple researchers have highlighted the link between position (distance and angle) of soccer referees and decision performance (e.g., Gilis, Helsen, Catteeuw, & Wagemans, 2008; Mallo et al., 2012; Oudejans et al., 2000, 2005). For example, Mallo et al. (2012) demonstrated referees had a lower number of incorrect decisions when the referees were positioned in the central area of the field. Research in medical and military settings has shown that fatigue and physical exertion have a detrimental effect on decision-making (e.g., Kovacs & Croskerry, 1999; Larsen, 2001). However, in sport contexts, decision-making performance was shown to be unaffected by physical exertion in Australian football umpires (Elsworthy, Burke, Scott, Stevens, & Dascombe, 2014; Paradis, Larkin, & O'Connor, 2015), fatigue in English Premier League assistant referees (Catteeuw, Gilis, Wagemans, & Helsen, 2010) or physical performance of New Zealand Football Championship referees (Mascarenhas et al., 2009). Thus, it is possible the change in the number of decisions is in response to the reducing work rate of the players or level of performance. For example, Weston and colleagues (Weston, Bird, Helsen, Nevill, & Castagna, 2006; Weston et al., 2012) found that soccer referees and players high intensity running distance, ball travel, and total distance covered were correlated. However, further research is required to understand the link between player and referee physical performances and their impact on referee decision-making. As suggested by Poolton et al. (2011), higher Rumination subscale scores, and not Reinvestment scores, were strongly associated (r > -0.7) with fewer decisions in Q1 and Q3. Notably, higher ruminators made fewer decisions against home teams during those quarters. Burke, Joyner, Pim, and Czech (2000) demonstrated that basketball officials' cognitive anxiety was higher pre-game, and at half time when compared to post-game. It is possible that prior to the start of the game, where officials arrive at the venue early and watch the teams' warm-up pre-game, and during the half-time break, there is greater potential for officials to engage in ruminative thoughts than during the smaller breaks taken between Quarters 1 and 2, and 3 and 4. To our knowledge, no researchers have investigated the timing of sports officials' decision ruminations. However, Roy et al. (2016) explored the timing of rumination by asking hockey players to rate on a 5-point scale whether they would continue to think about the play when it was over and their role in the play (past play), and how the team and individual would perform in the rest of the match (future play). Their results indicated that participants were unlikely to think about previous play after it was over, or about how the game would unfold; however, they were more likely to think about past play than future play. The authors suggested that the low rumination observed in successful field hockey players could reflect that people low in rumination do best in tasks requiring quick shifts of attention (such as dynamic team sports). Alternatively, a possible explanation might be that umpires engage in avoidance behaviours to reduce the chance of scrutiny of their decisions (Anderson, 2003). Contrary to our hypothesis, but consistent with Poolton et al. (2011), Reinvestment subscales scores were not related to the number of decisions. A home advantage effect was observed; the descriptive statistics indicated that more decisions were awarded against away teams, supporting findings in soccer, that home teams were awarded more penalties (Nevill et al., 1996) and that more yellow cards were awarded to away teams (Goumas, 2014). Factors purported to contribute to the home advantage include travel (i.e. greater time and distances for the away team), referee bias, familiarity and crowd size (Pollard, 2008). Furthermore, the correlations suggested that for matches in later rounds, where there is often greater importance due to more matches influencing final placings, play-offs and finals, fewer decisions were awarded against home teams. One explanation could be that officials exhibit avoidance-type behaviours to cope with the increases in anxiety resulting from increased perceived importance. Hill et al. (2016) found that rugby referees highlighted the importance of the game as one of the stressors affecting their performance, and that some referees use avoidance coping methods (e.g., Jordet & Hartman, 2008) to manage this stressor. It is possible that umpire experience could have confounded these figures, however a correlation between round and the umpires years of experience, where you might expect the most experienced umpires to officiate the latter rounds, was non-significant (r = 0.126, p = 728). Our results are consistent with previous research (e.g., Boyko et al., 2007; Page & Page, 2010) where increases in crowd size were associated with an increase in the number of decisions against away teams. One possible explanation is that when faced with a difficult decision, officials draw on other salient cues (e.g., crowd noise), particularly when placed under time constraints (Balmer et al., 2007). In order to reduce the complexity of a decision (Souchon et al., 2010) umpires' may use simple heuristics (Raab, 2012). For example, if two opposing players contested a ball and the umpire was unsure of the penalty decision, they may place equal weight on the auditory crowd cues as they do their visual information. Crowd noise typically favours the home team, resulting in more decisions against away teams (Nevill & Holder, 1999). This finding is reflected in our data, with larger crowd sizes associated with more decisions against away teams. Alternatively, researchers have reported that crowd noise induces a reluctance to penalise the home team (Nevill et al., 2002) (i.e., an absence of crowd noise indicates to the referee that no serious offence has been committed). The number of years' experience was not associated with the number of decisions made. This may be due to the number of years' experience umpiring at Superleague level (which was not recorded) or that there was little to no difference in qualification (Hancock & Ste-Marie, 2013). Other researchers have found the referee's experience to influence decision -making. Nevill et al. (2002) found as referees experience increased, that more fouls were awarded against home players, until a peak of 16 years, where upon a decline was then observed. However, the number of games umpired was positively associated with Reinvestment subscale scores. Potentially, those umpires who deliberate more on their decisions are deemed more effective and are therefore requested to umpire more often. League position predicted fewer decisions against home teams when playing lower positioned away teams, and for away teams playing lower positioned home teams. This finding may be similar to the reputation bias of judges found by Findlay and Ste-Marie (2004) and Plessner (1999) whereby teams with a better performance reputation may be sanctioned less. Alternatively, it is possible that the results of this study could be explained by the differences in players (e.g., lower ability teams or less competitive matches), or players' susceptibility to pressure, and not that of the officials. Previously, researchers have reported that yellow cards against away players in soccer could be a consequence of a poorer psychological state when compared with playing at home (Bray, Jones, & Owen, 2002; Terry, Walrond, & Carron, 1998). There were several limitations that need to be acknowledged. First, we had incomplete data for crowd size, resulting in six matches being excluded from the crowd size analyses. Similarly, not all umpires who officiated the season completed the DSRS and were therefore excluded from the correlational analyses. However, those who did complete the DSRS officiated 72% of the matches analysed. Second, the accuracy of decisions was not recorded, preventing insight into the performance change of umpires exposed to different contextual and environmental conditions or comparisons between those with greater or lesser disposition to ruminate. However, it was not practically possible to obtain objective assessments of every decision made by the officials across the season. We also acknowledge that rumination is often seen as a negative process (referring to passive self-critical worrisome or anxious thinking, Trapnell & Campbell, 1999; Treynor, Gonzalez, & Nolen-Hoeksema, 2003), whereas self-reflection (considered to be a motivated process aimed at understanding in the self and overcoming problems and difficulties, Trapnell & Campbell, 1999; Treynor et al., 2003) on performance is an important post-game learning tool used by sports officials (MacMahon et al., 2015). Although the DSRS items refer to negative ruminative thoughts, our study design did not allow us to collect data on the types or timings of rumination/reflection. Further investigation is required to examine the relationship between rumination and performance in sports officials, with reference to the types (rumination versus reflection) and timings (before, during, and after performance) of ruminations officials' make through self-report or stimulated recall. Third, we cannot isolate the influence of each potential bias using the current study design. The number of decisions umpires make may be a result of a combined effect of crowd sizes, league position, round, and time. For example, you might expect later rounds to have greater crowd sizes, which could have confounded our data. However, a correlation between round and crowd size, was not significant (r = 0.136 p = 0.326). It would be beneficial to investigate these effects in isolation in a controlled environment in order to draw clearer conclusions regarding the potential influence of these factors. Furthermore, we cannot be certain that the players' performance was not affected by the same contextual, environmental or dispositional influences, leading the umpires to adjust their decision-making accordingly. Finally, we used observational data and descriptive and correlational analyses. An advantage of the use of observational data is the high external validity, making the results easily interpretable and applicable in the real world. While our approach is novel and the study presents the first empirically based analysis of netball officiating behaviour we cannot infer causality from the findings. In future, controlled experiments are required to establish any causal links that may be implied in our data. For example, future research should examine the specific crowd factors that lead to changes in decision-making behaviour such as examining the impact of volume on decision-making, where crowd size has been linked to crowd noise (Hayne, Taylor, Rumble, & Mee, 2011); or investigating the semantics of crowd members (e.g., relevant or irrelevant to the decision, Bishop, Moore, Horne, & Teszka, 2014). In summary, we explored putative contextual/environmental and dispositional influences on netball umpires' decision-making. We observed a home advantage effect, whereby more decisions were awarded against away teams when crowd sizes were greater. We found a reduction in the number of observable decisions made, against teams with higher status, in more important matches, as the time played in a match decreased and as a function of increasing levels of Decision Rumination. Our study presents the first empirically-driven task analysis of the demands of refereeing in netball and highlights a number of key areas for which follow-up research comprising experimental designs and manipulations may be employed."],["Delusional beliefs are typically pathological. Being pathological is clearly distinguished from being false or being irrational. Anna might falsely believe that his husband is having an affair but it might just be a simple mistake. Again, Sam might irrationally believe, without good evidence, that he is smarter than his colleagues, but it might just be a healthy self-deceptive belief. On the other hand, when a patient with brain damage caused by a car accident believes that his father was replaced by an imposter or another patient with schizophrenia believes that \"The Organization\" painted the shops on a street in red and green to convey a message, these beliefs are not merely false or irrational. They are pathological. What makes delusions pathological? This paper explores the negative features because of which delusional beliefs are pathological. First, I critically examine the proposals according to which delusional beliefs are pathological because of (1) their strangeness, (2) their extreme irrationality, (3) their resistance to folk psychological explanations or (4) impaired responsibility-grounding capacities of people with them. I present some counterexamples as well as theoretical problems for these proposals. Then, I argue, following Wakefield's harmful dysfunction analysis of disorder, that delusional beliefs are pathological because they involve some sorts of harmful malfunctions. In other words, they have a significant negative impact on wellbeing (=harmful) and, in addition, some psychological mechanisms, directly or indirectly related to them, fail to perform the jobs for which they were selected in the past (=malfunctioning). An objection to the proposal is that delusional beliefs might not involve any malfunctions. For example, they might be playing psychological defence functions properly. Another objection is that a harmful malfunction is not sufficient for something to be pathological. For example, false beliefs might involve some malfunctions according to teleosemantics, a popular naturalist account of mental content, but harmful false beliefs do not have to be pathological. I examine those objections in detail and show that they should be rejected after all. --------------------------------------------------------------------------------","Delusional beliefs are typically pathological.1 Being pathological is clearly distinguished from being false or being irrational. Anna might falsely believe that her husband is having an affair but it might just be a simple mistake. Again, Sam might irrationally believe, without good evidence, that he is smarter than his colleagues, but it might just be a healthy self-deception. On the other hand, when DS, a patient with brain damage caused by a car accident, believes that his father was replaced by an imposter (Hirstein & Ramachandran, 1997) or Peter with schizophrenia, believes that “The Organization” painted the shops on a street in red and green to convey a message (Chadwick, 2001), these beliefs are not merely false or irrational. They are pathological. What makes delusional beliefs pathological? This paper explores the negative features because of which delusional beliefs are pathological. In Section 2, I critically examine the proposals according to which delusions are pathological because of (1) their strangeness, (2) their irrationality, (3) their resistance to folk psychological explanations or (4) impaired responsibility-grounding capacities of people with them. I present some counterexamples as well as theoretical problems for these proposals. In Section 3, I argue, following Wakefield’s harmful dysfunction analysis of disorder, that delusional beliefs are pathological because they involve some sorts of harmful malfunctions. In other words, they have a significant negative impact on wellbeing (=harmful) and, in addition, some psychological mechanisms, directly or indirectly related to them, fail to perform the jobs for which they were selected in the past (=malfunctioning). An objection to the proposal is that delusional beliefs might not involve any malfunctions. For example, they might be playing psychological defence functions properly. Another objection is that a harmful malfunction is not sufficient for something to be pathological. For example, false beliefs might involve some malfunctions according to teleosemantics, a popular naturalist account of mental content, but harmful false beliefs do not have to be pathological. I examine those objections in detail in Section 4 and show that they should be rejected after all. The central question of this paper is about what makes delusional beliefs pathological. Before starting, I have several remarks on the idea that delusional beliefs are pathological. First, when I use the term “pathological” in talking about mental states, I refer to the property of the mental states in virtue of which they constitute, together with other symptoms, mental disorders. Delusional beliefs, together with other positive and negative symptoms, constitute schizophrenia, for example. Second, the idea that a belief is pathological is different from the idea that it is delusional. Unfortunately, there is no uncontroversial definition of delusionality. According to DSM-5, a delusion is “a false belief based on incorrect inference about external reality that is held despite what almost everyone else believes and despite what constitutes incontrovertible and obvious proof or evidence to the contrary” (American Psychiatric Association, 2013, p. 819). This definition is, however, very controversial. Delusions might be accidentally true. Some delusions are not about external reality but rather about internal mental states. Some delusions might not be based on inference of any sort, and so on. In this paper, I simply stipulate that a belief is delusional if sufficient psychiatrists regard it as delusional. Third, I assume that delusionality and pathology can come apart, at least, in principle. First, some delusional beliefs might not be pathological. For example, it is often argued that healthy individuals can have delusional beliefs or delusion-like ideas.2 A person without any psychiatric diagnosis might have a paranoia belief that his colleagues are trying prevent him from being promoted. The belief is delusional (i.e. regarded as delusional by psychiatrists) but not pathological (i.e. does not constitute a mental disorder). Again, some pathological beliefs might not be delusional. For example, some instances of confabulations or obsessive thoughts might involve non-delusional pathological beliefs. It is conceivable that a person has obsessive thoughts about being contaminated by gems, but the thoughts are not delusional because he perfectly recognizes their implausibility. The thoughts in such a case are pathological (i.e. constitute a mental disorder, such as OCD) but not delusional (i.e. not regarded as delusional by psychiatrists).3","(1) Strangeness: Anna’s belief that his husband is having an affair is false but it is not very strange. Many married women can have the same belief at some point. DS’s belief that his father was replaced by an imposter, on the other hand, is not only false but also strange.4 The same thing is true about Peter’s belief that The Organization painted the shops to convey a message. This observation motivates the first proposal, according to which delusions are pathological because their strange content. In other words, the pathology of delusion comes from the abnormality of the content. A problem of this proposal is that it is not obvious that all delusions are significantly stranger than healthy beliefs. Peter’s belief is certainly strange. But, there are some healthy beliefs that are as strange as his. For example, Murphy (2013) discusses a community in Sudan where it is believed that ebony trees provide important social information. The belief about ebony trees is culturally normal and hence not pathological. Nonetheless, it seems to be as strange as Peter’s delusional belief. One might think, however, that this problem can be solved by introducing a culture-relative notion of strangeness. The idea, for example, is that the belief about ebony trees is not strange relative to the culture in the community. Peter’s belief, on the other hand, is strange relative to the culture in the western, modern community to which he belongs. But, this response does not solve all the problems, because pathological delusions and healthy beliefs with similar content can exist in the same cultural contexts. For example, it can be difficult to distinguish Anna’s belief from the delusional jealousy in Othello syndrome by content alone. Presumably, the main difference between healthy beliefs about the partner’s infidelity and delusional jealousy is not about the content but rather about the sensitivity to the cues. As Easton, Schipper, and Shackelford noted, delusional jealousy “can be thought of as hypersensitive jealousy, as these individuals experience jealous reactions at a much lower threshold than normal individuals” (Easton, Schipper, & Shackelford, 2007, p. 399). Another problem is that being strange is not sufficient for a belief to be pathological. Philosophers seriously believe very strange things. But, typically, these philosophical beliefs are not the expressions of a mental disorder, but rather of remarkable insights and argumentative skills. For example, there are some philosophers who seriously believe that every single object in the universe is conscious (panpsychism), that, for any objects, however arbitrary they are chosen, there is a further object that is composed by them (unrestricted composition), that there are facts about the boundaries of a vague predicate which we can never discover (epistemicism about vague predicates), and so on. (2) Irrationality: Sam’s belief that he is smarter than his colleagues is irrational but, presumably, DS’s belief that his father was replaced by an imposter and Peter’s belief that The Organization painted the shops to convey a message are more irrational. Maybe, they are too irrational. According to the second proposal, delusion is pathological because of their extreme irrationality.5 However, it is not obvious that delusional beliefs are extremely irrational. According to empiricist accounts of delusion formation, which is very influential recently, delusions are formed in response to abnormal experience. Given the fact that the abnormal experience can be understood as a kind of evidence for delusional beliefs, it is not obvious at all that, according to empiricism, delusions are extremely irrational. Indeed, a number of empiricist researchers support the view that a delusion is a reasonable response to abnormal experience. Maher famously argued that delusions “are derived by cognitive activity that is essentially indistinguishable from that employed by non-patients, by scientists, and by people generally” (Maher, 1974, p. 103). Coltheart, Menzies, and Sutton (2010) support this claim and argue that it is perfectly Bayesian rational for a Capgras patient to adopt the imposter hypothesis rather than the competing, realistic hypotheses given the abnormal data. (They do not use the term “experience” because they think that the “data” are not consciously accessible.) The imposter hypothesis actually gets a higher posterior probability than the competing hypotheses. Similarly, Corlett and colleagues argue that it is hard for a Capgras patient to avoid the imposter hypothesis given the abnormal experience they have; “the phenomenology of the percepts are such that bizarre beliefs are inevitable; surprising experiences demand surprising explanations” (Corlett, Taylor, Wang, Fletcher, & Krystal, 2010, p. 360). Still, the view that a delusion is a rational response to abnormal experience is controversial. Stone and Young (1997), for instance, argue that delusional beliefs are produced not only by abnormal experience but also by the irrational reasoning with the bias toward observational adequacy; people with delusion irrationally put more emphasis on incorporating new observations into belief system (observational adequacy) than keeping existing beliefs as long as possible (doxastic conservatism). McKay (2012) follows this suggestion and argues, in response to Coltheart et al. (2010) that delusional beliefs are produced through the Bayesian-irrational reasoning process with the bias of discounting prior probabilities of the hypotheses; people with delusion irrationally put more emphasis on likelihoods (which summarize how nicely hypotheses explain the observation) than prior probabilities (which summarize how probable the hypotheses are prior to the observation). But, even if one of those views is correct, it is still not obvious that delusional beliefs are extremely irrational. Similar irrational biases might be found in healthy beliefs as well. For example, the famous study by Kahneman and Tversky on the base-rate neglect (1973) shows that normal people have the strong tendency to neglect the base-rate information that is relevant for given hypotheses. Here, the base-rate information gives the prior probabilities of the hypotheses at issue. Thus, the tendency to neglect the base-rate information can be understood as the tendency to neglect prior probabilities. One of the basic principles of statistical prediction is that prior probability, which summarizes what we knew about the problem before receiving independent specific evidence, remains relevant even after such evidence is obtained. Bayes’ rule translates this qualitative principle into a multiplicative relation between prior odds and the likelihood ratio. Our subjects, however, fail to integrate prior probability with specific evidence. [ …] The failure to appreciate the relevance of prior probability in the presence of specific evidence is perhaps one of the most significant departures of intuition from the normative theory of prediction (Kahneman & Tversky, 1973, p. 243). Experimental studies revealed some “biases” in judgment processes of people with delusion. However, the studies do not necessarily support the idea that delusions are very irrational. The term “bias” can be used in, at least, two different ways. It might refer to the deviation from the norm of (logico- mathematical) rationality. When we say that the tendency to neglect the base-rate information is “biased”, the term is used in this sense. It deviates from the Bayesian norm of rationality. Alternatively, the term might refer to the deviation from the performance of normal people. In the context of the delusion research, the term “bias” is often used in the second way. So, it is perfectly possible that the “biased” performance of people with delusion does not actually deviate from the norm of rationality. This possibility is nicely illustrated by the well-known study of the “jumping-to-conclusion bias” (Huq, Garety, & Hemsley, 1988). It was found in the study that people with delusion have the tendency to “jump to conclusion”; they require less evidence before coming to conclusions than people in control groups (healthy people and non-delusional people with schizophrenia). At the same time, though, it was also found that the performance of people with delusion is more rational from a mathematical point of view than that of people in control groups. People with delusion reach the conclusion when the probability of the hypothesis is reasonably high, while people in control groups do not reach the conclusion until the probability is unnecessarily high. So, Huq et al. wrote; “[i]t may be argued that the deluded sample reached a decision at an objectively ‘rational’ point. It may further be argued that the two control groups were somewhat over cautious” (Huq et al., 1988, p. 809). (3) Understandability: Sam’s belief that he is smarter than his colleagues is irrational, but it is “understandable” in the sense that we can give a simple folk psychological account of it. Sam comes to believe it because he wants it to be the case that he is smarter than his colleagues. In other words, his belief is driven by the desire to be smarter than the colleagues. On the other hand, DS’s belief that his father is replaced by an imposter and Peter’s belief that The Organization painted the shops to convey a message are not “understandable” in this way. There are no easy folk psychological explanations of those delusional beliefs. According to the third proposal, delusions are pathological because of the “ununderstandability” or the resistance to folk psychological explanations. But, this view is not fully satisfactory. First, it is not clear that the all delusions resist folk psychological explanations. Sam’s belief is “understandable” because we can identify the motivational factors that play crucial roles in the formation of the belief. But, then, when motivational factors play crucial roles for some delusions, at least those delusions are “understandable” by identifying those motivational factors.6 Butler (2000) reported the case of B.X. who had the delusion about the fidelity of his former romantic partner (reverse Othello syndrome). B.X. was a gifted musician who had been left quadriplegic following a car accident. One year after his injury, he developed delusional beliefs about the continuing fidelity of his former romantic partner, N., who had in fact severed all contact with him soon after the accident. A pretty straightforward folk psychological explanation of this case would is that B.X. formed his delusional beliefs because he desperately wanted it to be the case that N. still loved him. In other words, his delusional beliefs were driven by the desire that N. still loved him. Second, resisting folk psychological explanation does not seem to be sufficient for beliefs, or mental states in general, to be pathological. The so-called “twisted self-deception” is a good example. Twisted self-deceptive beliefs are, roughly speaking, the unwelcome irrational beliefs. For example, if it turns out that Anna’s belief about the husband’s affair is not supported by the available evidence at all, it is a twisted self-deceptive belief. Presumably, there is no straightforward folk psychological explanation of twisted self-deceptive beliefs. Mele (1999) provides an influential account according to which twisted self-deceptive beliefs are produced by the people’s tendency to avoid costly errors. If, on one hand, Anna falsely believes that her husband is having an affair, then the falsity is not very costly. For example, it just annoys the husband. If, on the other hand, she falsely believes that the husband is not having an affair, then the falsity is very costly. The relationship is seriously threatened in that case. This account is not purely folk psychological. The idea that people choose their beliefs so that they can avoid costly errors does not seem to be a part of folk psychology. Indeed, the idea comes from the scientific, not folk, psychological theory by Friedrich (1993). Another example comes from Gendler (2008). Many people experience an extreme fear when they walk on the horseshoe-shaped transparent walkway on the 4000 feet above the floor of the Grand Canyon. There is nothing abnormal or pathological in this experience. As Gendler noted, “the basic phenomenon—that stepping onto a high transparent safe surface can induce feelings of vertigo—is both familiar and unmysterious” (Gendler, 2008, p. 635). The people who walk on the walkway seem to believe that the walkway is safe. But, then, why do they feel the extreme fear? The fear seems to be ungrounded if they seriously believe that the walkway is safe. Are they, then, somewhat sceptical about the safety? But, in that case, we cannot explain the fact they step onto the walkway in the first place. Presumably, there is no easy folk psychological explanation of the fear. Gendler argues that they feel the extreme fear because they alieve that the walkway is not safe, although they believe that it is safe. Alief is, roughly speaking, “a mental state with associatively linked content that is representational, affective and behavioral, and that is activated—consciously or nonconsciously—by features of the subject’s internal or ambient environment” (Gendler, 2008, p. 645). Although alief sounds like another folk psychological state, Gendler’s account of the case is not purely folk psychological. After all, alief is not a part of the conceptual repertoire of folk psychology. Presumably, it is best understood as an extended folk psychological account.7 (4) Responsibility: Suppose that Anna, on the basis of her belief, acts violently to the woman who is mistakenly regarded as the affair partner. The falsity of the belief does not change the fact that Anna is responsible for what she does. Anna is clearly responsible for her violence. On the other hand, if DS, because of his delusional belief, had acted violently to his father, he would not have been fully responsible for the violence. People with Capgras delusion sometimes act violently on the basis of their delusional beliefs. In a tragic case in Louisiana in 2011, for instance, a father beheaded his disabled son because he believed that the son was had been replaced by a CPR dummy. Following the testimony by forensic psychiatrists, he was ruled not guilty by reason of insanity. The forth proposal is that delusions are pathological because of responsibility-grounding capacities, such as decision-making capacity or autonomous agency, are significantly impaired in people with delusion. It is, however, not obvious that responsibility-grounding capacities are always impaired in people with delusion. Certainly, it is very likely that belief-forming or belief-checking capacities are impaired somehow in them. But, the impairment in belief-forming/checking capacities might be dissociated from the impairment in responsibility-grounding capacities. A possible response to this challenge is that belief-forming/checking capacities are relevant to the attribution of responsibility. In other words, responsibility-grounding capacities include belief-forming/checking capacities. For example, according to M’Naghten rules, a person is not responsible for what he does when he is ignorant of “the nature and quality of the act he was doing.” Presumably, people with delusion are ignorant of the nature and quality of their acts because of impaired belief-forming/checking capacities and, hence, they are not responsible for the acts. However, this response is not very convincing. Certainly, there is a sense in which a violent Capgras patient is ignorant of the nature and quality of his act; he thinks that he is attacking the imposter, which is false. But, this is also true about Anna; she thinks that she is attacking the affair partner, which is false. So, if “being ignorant” means “having the false belief” or “not having the true belief”, then Anna is as ignorant as violent Capgras patients. The phrase might be interpreted in different ways, but it is not clear that there is an interpretation according to which violent Capgras patients are significantly more ignorant than Anna. Another response is that delusion is just a tip of iceberg. People with delusion often have other kinds of abnormalities at the same time and these abnormalities impair responsibility-grounding capacities. For instance, delusions in the context of schizophrenia are accompanied by other positive and negative symptoms that directly or indirectly impair responsibility-grounding capacities. However, we cannot assume a priori that responsibility-grounding capacities are compromised in people with schizophrenia in general. There might be serious individual differences about the quality of responsibility-grounding capacities of people with schizophrenia given the fact that schizophrenia is an extremely heterogeneous condition. Maybe, the fact that one gets the diagnosis of schizophrenia in itself does not tell us much about his responsibility- grounding capacities. As Bortolotti, Broome, and Mameli (2013) pointed out, “[t]he assumption that people who have psychotic symptoms or have received a diagnosis of schizophrenia lack responsibility or have reduced responsibility for action is especially problematic, as the behavior of two people with psychosis or schizophrenia can differ almost entirely. Some people with schizophrenia are able to function well, cognitively and socially, and to control their delusions to some extent.”","I believe that the main problem of the previous proposals is that they are detached from the considerations about what, in general, makes a condition pathological. In order to explain why a condition X is pathological, first, we need to have a general account of the features that make conditions pathological and, then, show that X has the features as well. But, the previous proposals skip the first step and simply point out some remarkable negative features of delusion. This invites all sorts of counterexamples and difficulties. Wakefield (1992a, 1992b) presented a general account of disorder, which is very influential and, in my view, more plausible than its rivals. It is often called “Harmful Dysfunction Analysis of Disorder.” I will call it “HDA.” According to HDA, disorders are “harmful malfunctions” or “harmful dysfunctions.” Being “harmful” means having a negative impact on wellbeing. The harmfulness condition is important because “disorder is in certain respects a practical concept that is supposed to pick out only conditions that are undesirable and grounds for social concern” (Wakefield 1992b, p. 237). A “malfunction” is, roughly speaking, the failure to perform an etiological function,8 where the etiological function of something is the performance for which it was selected in the past.9 For instance, a heart malfunctions when it fails to pump blood, a kidney malfunctions when it fails to filter metabolic wastes from blood, a corpus callosum malfunctions when it fails to facilitate interhemispheric communications, and so on. My proposal, which relies on HDA, is that delusions are pathological because (1) they are harmful and (2) they involve some etiological malfunctions directly or indirectly. I will call (1), and (2) “the harmfulness thesis” and “the malfunction thesis” respectively. There are many objections to HDA. It should be noted, however, that most (if not all) objections are aiming at refuting the necessity claim of HDA, namely, the claim that a harmful etiological malfunction is necessary for a disorder. This means that those objections are not very serious for my purpose. What is crucial for my account is that a harmful etiological malfunction is sufficient for a disorder. We can say that delusions are pathological because they involve harmful etiological malfunction as long as a harmful etiological malfunction is sufficient for a disorder. For example, Tengland (2001) argues that viral infection is a counterexample to HDA. Viral infection is a disorder but its symptoms (e.g. fever, cough, sneezing) are very often biological defences, not malfunctions. Again, Murphy and Woolfolk (2000) argue that appendicitis is another counterexample. Appendicitis is clearly a disorder, but vestigial organs such as appendix cannot fail to perform its etiological functions because they do not them in the first place.10 Wakefield argues, in response to Tengland, that viral infection involves etiological malfunctions at the level of cell (Wakefield, 2011) and, in response to Murphy and Woolfolk, that appendicitis involves etiological malfunctions at the level of tissue (Wakefield, 2000). In my opinion, Wakefield’s responses are convincing. But, even if they are not, those counterexamples are not very serious for my account, because they are the counterexamples to the necessity of a harmful etiological malfunction for a disorder. The harmfulness thesis ~~~~~~~~~~~~~~~~~~~~~~ Delusions seem to be harmful in all sorts of ways. Delusions do not only cause psychological stress and anxiety but also harmful consequences in the life of the people with them, such as losing a job or failing to maintain relationship with other people. McKay and colleagues propose the following definition of delusions, which explicitly mentions the harmfulness of delusions; “A person is deluded when they have come to hold a particular belief with a degree of firmness that is both utterly unwarranted by the evidence at hand, and that jeopardises their day-to-day functioning [emphasis added]” (McKay, Langdon, & Coltheart, 2005, p. 315). Delusions can be especially harmful when people act on them. There are many reported cases where people perform harmful actions on the basis of their delusions. The case in Louisiana that I already mentioned, for example, a father killed his disabled son on the basis of his Capgras delusion, which is not only harmful to the son (being killed), but also to the father (killing his own son without knowing).11,12 Some clarifications are in order. First, the “harm” at issue does not have to be the harm to people with delusion themselves. It might be the harm to people around them, such as family members, friends, colleagues, or neighbours. It is conceivable that some people with delusions are happy because of them at some point. For example, the grandiose delusion about special abilities given by God might make some people happy at some point. Still, the delusion might be harmful overall because of the serious troubles it causes for other people. Second, the harmfulness thesis is not about hedonistic pains or pleasures. The life on the experience machine (the machine that produce the perfect illusion of leading whatever kind of life one desires, while one floats about in a tank) is extremely pleasurable from the hedonistic point of view. But, typically, we do not find the life very attractive. Nozick argues that the life is not attractive because “we want to be a certain way, to be a certain sort of person. Someone floating in a tank is an indeterminate blob. There is no answer to the question of what a person is like who has long been in the tank. Is he courageous, kind, intelligent, witty loving? It’s not merely that it’s difficult to tell; there’s no way he is” (Nozick, 1974, p. 43). Similarly, it is conceivable that the life with the grandiose delusion about special abilities is really pleasurable from the hedonistic point of view. But, again, it does not mean that the life is really attractive. It is not attractive because, as a matter of fact, the person with the delusion fails to be the person he really wants to be. He wants to have some special abilities that distinguish him from others. However, he does not have such abilities. His real life might not be less miserable the one in the tank. Third, there is an implicit ceteris paribus clause in the harmfulness thesis. Delusions are ceteris paribus harmful. For example, we can imagine the case where Peter avoided a fatal plane crash because he had changed the reservation due to his delusional belief that The Organization wants him to do so. In this case, his delusional belief is beneficial rather than harmful. Actually, the same thing is true about disorders in general. For example, we can easily imagine the case where Peter avoided the plane crash because he had changed his reservation due to his flu. But, of course, the case does not show that flu is not harmful. The claim that flu is harmful has an implicit ceteris paribus clause and these tricky cases do not contradict the ceteris paribus claim. The malfunctioning thesis ~~~~~~~~~~~~~~~~~~~~~~~~~ According to the malfunctioning thesis, delusions involve some etiological malfunctions. What kinds of malfunctions are they exactly? A fully satisfactory answer to this question is not available yet because the process of delusion formation and maintenance has not been fully understood. Still, the current understanding of the process is informative enough to help us to identify good candidates.13 First, according to the empiricist theories, delusions are formed in response to abnormal experience. In this view, there might be some etiological malfunctions in the mechanisms that are responsible for the abnormal experience. For example, Ellis and Young (1990) argue that Capgras delusions are formed in response to the abnormal experience of seeing familiar faces which is caused by the disconnection between autonomic nervous system and face recognition system.14 This hypothesis is supported by the finding that people with Capgras delusion do not show the asymmetrical autonomic responses between familiar faces and unfamiliar faces (Ellis, Young, Quayle, & De Pauw, 1997). The experience of Capgras patient is abnormal in that it lacks the affective component that is usually a part of the experience. A Capgras delusion is formed as an explanation of this abnormal perceptual-affective experience. Second, there might be some etiological malfunctions in attention mechanisms. According to prediction-error theories (Corlett et al., 2010; Fletcher & Frith, 2009), people with delusion have problems in the allocation of attention caused by aberrant prediction-error signals. Prediction-error signals are the indicators of the mismatch between expectations and actual inputs. They guide the process of attention allocation in such a way that attention is paid to the things or events that defy expectations. Due to aberrant prediction-error signals, people with delusion pay attention to the things or events that are not actually important. A delusion arises as the explanation of the apparent significance of these things and events. The study by Corlett et al. (2007) supports this hypothesis. In the study, two groups of participants (people with and without delusion) were tested with a task involving learning the association between certain foods and allergic reactions, while the activity of the right prefrontal cortex (which was identified as a reliable marker of prediction-error processing in previous studies) was monitored with fMRI. In the delusional group but not in the control group, the magnitude of the activity of the right prefrontal cortex was not significantly different between the cases where expectations about allergic reaction were confirmed and the cases where they were violated. In other words, in the delusional group, the right prefrontal cortex fails to distinguish prediction-error from prediction-confirmation.15 Third, the two-factor theories (Coltheart, 2007; Davies, Coltheart, Langdon, & Breen, 2001) posit the so-called “second factor”, namely, the factor that explains the adoption and/or maintenance of delusional hypotheses. The second factor is often associated with the damage to right frontal lobe. If the second factor really exists, then there can be some etiological malfunctions that are responsible for it. The main reason for positing the second factor is that it explains the difference between people with delusions and people without delusions who are experientially equivalent. For example, patients with the damage to the ventromedial prefrontal cortex do not adopt imposter hypotheses but are often regarded as experientially equivalent to people with a Capgras delusion. Unfortunately, there is no agreement on the nature of the second factor. Here is a recent proposal. Coltheart et al. (2010) argue that the second factor is the bias of neglecting contradictory evidence after adopting delusional beliefs.16 After a Capgras patient adopts the imposter hypothesis, he will be exposed to contradictory evidence such as the behavior of the “imposter”, the testimony of trustworthy friends, and so on. The contradictory evidence however is neglected because of the bias and, consequently, the delusional hypothesis is maintained.17,18 Delusions might not be malfunctional ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There are at least two kinds of possible objections to the claim that delusions are pathological because of X. First there might be some objections according to which it is not the case that all delusions are X.19 For example, an objection to the view that delusions are pathological because of their strangeness is that it is not the case that all delusions are strange. Delusional jealousy is not. Second, there might be some objections according to which X is not sufficient for a belief to be pathological. For example, an objection to the view that delusions are pathological because of the resistance to folk psychological explanations is that resisting folk psychological explanations is not sufficient for a belief to be pathological. Twisted self-deceptive beliefs resist folk psychological explanations, but they are not pathological. In this section, I examine these two kinds of objections to my own proposal. According to the first group of objections (4.1), it is not the case that all delusions are malfunctional. Some of them are perfectly functional. According to the second (4.2), a harmful malfunction is not sufficient for a belief to be pathological. There are some beliefs that are harmful and malfunctional but are not pathological. Psychological defence ~~~~~~~~~~~~~~~~~~~~~ One might think that some delusions are not malfunctional but rather are successfully performing a function, namely, the psychological defence function. In the case of B.X. that I mentioned earlier, for instance, his delusion about the fidelity of N. plays the psychological defence function. It defends B.X. from the stark reality that his body is paralyzed and N. does not love him anymore. Indeed, the idea that some delusions have defensive roles is popular recently (Bentall & Kaney, 1996; McKay, Langdon, & Coltheart, 2007). For instance, it has been suggested that persecutory delusions are produced by the so-called “externalizing attribution bias,” which is the bias of attributing negative events to other agents rather than themselves in order to defend self-esteem. Psychological defence objection, however, is not very persuasive. I do not rule out the idea that some delusions play psychological defence roles. The problem of this objection is rather that even if they play such roles, it does not imply that they are successfully playing biological, etiological functions. Certainly, defending self-esteem is a good thing for us. It brings psychological comfort. But it is not obvious that it is not only psychologically good, but also biologically good. In other words, it is not obvious that defending self-esteem does not just bring psychological comfort, but also brings reproductive success. Stich famously argues that, “natural selection does not care about truth; it cares only about reproductive success” (Stich, 1990, p. 62). Similarly, we can also say that natural selection does not care about psychological comfort; it cares only about reproductive success. Etiological functions and psychological comfort do come apart in some cases. For example, the negative emotions, such as fear or anxiety, are psychologically negative, but they play important biological functions, such as the function of avoiding dangers or threats. Presumably, we can even say that they play those functions exactly because they are psychologically negative, in the same way that pain plays the function of defending body from damage exactly because it is psychologically negative. (Pain would fail to defend body if it were psychological positive.) Furthermore, there are some conditions that are psychologically positive, but etiologically malfunctional. Nesse (1998) argues that insufficient anxiety, which is psychologically positive, is etiologically malfunctional. Anxiety has important etiological functions, and insufficient anxiety is the failure of performing these functions. People with insufficient anxiety never visit psychiatric clinics. They do not think they have to do so. Nonetheless, their anxiety mechanisms are etiologically malfunctioning and, presumably, they should be medically treated in some cases. Doxastic shear pin ~~~~~~~~~~~~~~~~~~ McKay and Dennett (2009) consider an interesting hypothesis according to which delusions are “doxastic shear pins.” A shear pins is a metal pin installed in complex mechanistic systems, and it is designed to break in certain circumstances in order to protect other, more expensive parts of the systems. It is conceivable that some delusions play similar roles. A possible hypothesis is that there is a mechanism whose function is to prevent motivational factors from influencing belief forming processes. But, in the situation where one faces extreme psychological stress, the mechanisms is designed to break and let motivational factors influence belief forming processes in order to protect more important cognitive mechanisms. For example, in the case of B.X., the mechanism is broken, in accordance with the design, in the face of the extreme psychological stress and, consequently, his desire for the continuing fidelity of N. has a significant impact on belief forming processes, which leads to his delusion. Mishara and Corlett (2009) propose another version of shear pin hypothesis on the basis of the prediction-error theory. According to the theory, a delusion is formed in response to prediction-error signalling abnormalities. Due to aberrant prediction-error signals, trivial things or events become abnormally salient and attention-grabbing. This is the stage of the so-called “delusional mood.” A delusion in the end arises as the explanation of the abnormal salience attached to the things and events. A delusion, according to Mishara and Corlett, can be understood as a kind of doxastic shear-pin; The delusions appear as an Aha-Erlebnis, or “revelation”, concerning what had been perplexing during delusional mood. [ …] The delusions are not primarily a defensive reaction to protect the self, but involve a “reorganization” of the patient’s experience to maintain behavioral interaction with the environment despite the underlying disruption to perceptual binding processes. At the Aha-moment, the “shear-pin” breaks, or as Conrad puts it, the patient is unable to shift “reference-frame” to consider the experience from another perspective. The delusion disables flexible, controlled conscious processing from continuing to monitor the mounting distress of the wanton prediction error during delusional mood and thus deters cascading toxicity. At the same time, automatic habitual responses are preserved, possibly even enhanced (Mishara & Corlett, 2009, p. 531). Now, I do not have a priori reasons to rule out these hypotheses. What is crucial for me is that they are perfectly compatible with my proposal. They are compatible because there might still be some etiological malfunctions that are directly or indirectly related to delusional beliefs in those hypotheses. This is very likely in the hypothesis by Mishara and Corlett. In the hypothesis, it is assumed that prediction-error signalling is abnormal and it causes abnormalities in attention allocation processes. The role of delusions is to help people to maintain behavioral interactions with the environment despite these abnormalities. They do not eliminate the abnormalities. In the hypothesis by McKay and Dennett, the mechanism that normally constrains the influence of motivational factors on belief formation fails to perform one of its functions. It is broken. Certainly, it successfully plays another function, namely, the function of defending more important mechanisms. Presumably, the mechanism is best understood as having two incompatible functions. On one hand, it has the function of constraining the influence of motivational factors on belief formation processes. On the other hand, it has the function of defending more important cognitive mechanisms. Those functions are incompatible because the mechanism successfully performs the latter function only by failing to perform the former and vice versa. In the case where the “shear pin” breaks, the mechanism is functional in the sense that it successfully performs the latter function, but it is malfunctional in the sense that it fails to perform the former. The error management theory ~~~~~~~~~~~~~~~~~~~~~~~~~~~ The error management theory (Haselton & Buss, 2000) is the view recurrent asymmetries in the costs of false alarms shaped varieties of cognitive and behavioral biases over evolutionary history. For example, it is well-established that, compared with women, men have the stronger tendency to overperceive sexual interest. According to the error management theory, this tendency is explained by the recurrent asymmetry in the costs of errors. Smoke detectors are designed to be activated more often than they really need to be. This is because false positives are not very costly (i.e. some unnecessary evacuations), while false negatives are extremely costly (i.e. the building will be burnt down). Similarly, natural selection designed the men’s sexual perception system so that it is activated more often than it really needs to be. This is because “false alarms typically result in trivial expenditures of wastes courtship effort for men: Although rejected men may experience social embarrassment, women generally do not respond antagonistically to men’s overperception of sexual interest. The costs of missed mating opportunities, on the other hand, were substantial for men over the course of human evolution, because men’s reproductive success can be directly affected by the access to fertile mates” (Perilloux, Easton, & Buss, 2012, p. 146). The error management theory might cause some problems for my proposal. Maybe, some delusions are produced by the biases that evolved due to recurrent cost asymmetries. Those biases are not malfunctional. Rather, they come from the very design of relevant mechanisms. Delusional jealousy is a good candidate. False positives about the infidelity of partners do not seem to be very costly (i.e. partners will be annoyed), while false negatives are very costly (i.e. the partners might leave the subjects). So, the error management theory predicts that that jealousy is, by design, activated more often than it is really needs to be. In other words, people are designed to be oversensitive to the infidelity of partners. It is conceivable that this bias does not only explain normal jealousy but also delusional jealousy. Consistent with this hypothesis, delusional jealousy has some important characteristics in common with normal jealousy. For instance, men with delusional jealousy are especially upset about the partner’s sexual infidelity, whereas women with delusion of jealousy are especially upset about the partner’s emotional infidelity, which is consistent with the pattern that is seen in normal jealousy (Easton et al., 2007). If it is true that delusional jealousy is the product of the error management theoretic bias, then there is nothing malfunctional about it. As Easton and colleagues suggested, “morbid jealousy does not meet the dysfunction criterion and therefore should not be considered a mental disorder” (Easton, Schipper, & Shackelford, 2006, p. 412). Error management theory in itself is a very plausible view. Still, I do not think that it causes serious troubles for my account. First, delusional jealousy might be pathological in the same sense that fever as the symptom of viral infection is pathological. Fever as the symptom of viral infection itself is not malfunctional. It is rather a designed defensive response. When we regard fever as pathological, we do so in virtue of the fact that it is a symptom of viral infection which involves etiological malfunctions. In other words, fever indirectly involves etiological malfunctions even though it is perfectly functional in itself. The same thing might be true about delusional jealousy. Indeed, delusional jealousy often occurs as a symptom of the conditions that are expected to involve some etiological malfunctions. It occurs, for example, in the contexts of schizophrenia, bipolar disorder, Parkinson’s disease, brain injuries, and so on. When we regard delusional jealousy as pathological in those cases, we do so presumably in virtue of the fact that it is a symptom of the conditions that involve etiological malfunctions. Delusional jealousy, thus, indirectly involves etiological malfunctions even if it is perfectly functional in itself. Second, even if normal people have the error management theoretic bias of being oversensitive to the infidelity of partners, the bias might not be sufficient to explain delusional jealousy. A possibility is that the bias is pathologically exaggerated in people with delusional jealousy. As McKay and Dennett suggested, “the most that can presently be claimed is that delusions may be produced by extreme versions of systems that have evolved in accordance with error management principles, that is, evolved so as to exploit recurrent cost asymmetries. As extreme versions, however, there is every chance that such systems manage errors in maladaptive fashion” (McKay & Dennett, 2009, p. 502). Indeed, if delusional jealousy is just an expression of an error management theoretic bias, then we cannot explain the fact that it is often seen in the contexts of schizophrenia or bran injuries. Presumably, delusional jealousy is more similar to rheumatoid arthritis, which is the product of pathologically exaggerated immune responses, than to the fever in viral infection. Another possibility is that the designed bias is only a factor of delusional jealousy. There is the second factor, which might be the bias of neglecting contradictory evidence after adopting delusional hypotheses (Coltheart et al., 2010). Maybe, the bias explains the fact that delusional jealousy is maintained, after adopted, in the absence of supportive evidence. Indeed, the two-factor theorists tend to think that the second factor is shared in all types of delusions. If so, again, it is likely that there are some malfunctions that underlie the second factor.20 Harmful malfunction is not sufficient ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ According to the second objection, a harmful etiological malfunction is not sufficient for a belief to be pathological. For example, it is sometimes said that the fundamental idea of teleosemantics is that misrepresentations, such as non-veridical perceptions or false beliefs, involve the failures of etiological functions. The basic idea behind teleological theories of content is that this normative notion – and its distinction between proper functioning and malfunctioning – might somehow underwrite the normative notion of content – and its distinction between representation and misrepresentation. (Neander, 1995, p. 112). Much of the original appeal of teleosemantics was its ability to employ teleo- functional notions of purpose in order to deal with apparently normative aspects of semantic phenomena. In particular, the biological notion of failure to perform a proper function was used to attack the problem of misrepresentation, which had caused a lot of trouble for information-based theories (Godfrey-Smith, 2006, p. 62). If it turns out that all misrepresentations involve etiological malfunctions, then, according to my proposal, all harmful misrepresentations are pathological. All harmful non-veridical perceptions and false beliefs are pathological. But, this creates too many mental disorders! Obviously, it is not the case that all harmful misrepresentations are pathological. One might falsely believe that he is not as smart as his colleagues, and the false beliefs might have a negative impact on wellbeing (e.g. the loss of self-esteem, psychological stress amnesia, etc.). This can happen to a perfectly healthy person (although it could lead to pathological conditions such as depression). This seems to show that my proposal is wrong. More precisely, it shows that a harmful etiological malfunction is not sufficient for a mental state to be pathological. Healthy harmful misrepresentations involve harmful etiological malfunctions, according to teleosemantics, but they are not pathological. There are some possible responses to this objection. First, I might simply reject teleosemantics. However, I do not find this option very attractive because the incompatibility with teleosemantics, which is a popular account of mental representations among naturalist philosophers of mind, is a disadvantage of my proposal. Second, I might argue that healthy harmful misrepresentations are not harmful enough to be pathological. For example, the healthy false belief that I am not as smart as my colleagues is not a counterexample my claim because it is not harmful enough. This option is committed to the view that the difference between healthy false beliefs and delusions is about the harmfulness condition of HDA. Both involve some kinds of etiological malfunctions. Thus, they are equivalent in terms of the malfunction condition. On the other hand, they are different in the degree of harmfulness. Delusions are harmful enough to be pathological. Healthy false beliefs are not. This response might or might not work. But I do not agree with the idea that delusions and healthy false beliefs are different only in terms of the harmfulness condition. I do believe that they are also different in the malfunction condition. Teleosemantics, properly understood, does not imply that all misrepresentations involve etiological malfunctions. It is a misunderstanding of teleosemantics that all misrepresentations involve etiological malfunctions according to the theory. Thus, even if we accept teleosemantics, we do not have to accept the view that all false beliefs involve etiological malfunctions. There are different versions of teleosemantics and they need different discussions. In the following, I will talk about two notable examples; Millikan (1984), Millikan (1989) consumer-based teleosemantics and Neander (1995, 2013) informational teleosemantics. Let us begin with the following famous example by Dretske. Some marine bacteria have internal magnets, magnetosomes, that function like compass needles, aligning themselves (and, as a result, the bacterium) parallel to the Earth’s magnetic field. Since the magnetic lines incline downward (toward geomagnetic north) in the northern hemisphere, bacteria in the northern hemisphere, oriented by their internal magnetosomes, propel themselves toward geomagnetic north. Since these organisms are capable of living only in the absence of oxygen, and since movement toward geomagnetic north will take northern bacteria away from the oxygen-rich and therefore toxic surface water and toward the comparatively oxygen-free sediment at the bottom, it is not unreasonable to speculate, as Blakemore and Frankel do, that the function of this primitive sensory system is to indicate the whereabouts of benign (i.e. anaerobic) environments (Dretske, 1991, p. 63). What does the state of magnetosome represent? Does it represent magnetic north or the oxygen-free sediment? Millikan thinks that it represents the oxygen-free sediment, not magnetic north. In Millikan’s view, what the state of magnetosome represents is determined by what the state needs to correspond to in order for the consumer of the state to perform its function successfully in its normal way. For the successful performance of the consumer (i.e. motor mechanism), the state needs to correspond to the oxygen-free sediment. After all, what is crucial for the successful functioning of the motor mechanism is to lead the bacteria to the oxygen-free sediment. Then, Millikan’s account actually allows for misrepresentations without etiological malfunctions. Suppose that I use a bar magnet to lead a bacteria upward and, consequently, the bacteria dies because of the exposure to oxygen-rich surface water. In this case, the state of the magnetosome misrepresents without etiological malfunctions. The state misrepresents because, on the one hand, it represents the oxygen-free sediment and, on the other hand, it is tokened when the oxygen-rich surface water is there instead. It does not involve any etiological malfunctions because nothing is wrong about the bacteria itself. Rather, it is just unlucky. Still, Millikan insists that the magnetosome in the case fails to perform its function in a certain sense. […] Dretske is right that the magnetosome that directs that bacterium in the wrong direction because someone holds a bar magnet overhead is not broken or malfunctioning. In that sense, it is functioning perfectly properly. But it doesn’t mean that it is succeeding in performing all of its functions, any more than a perfectly functional coffeemaker is performing its function when no one has put any coffee in it. Very often things fail to perform their functions, not because they are damaged, but because the conditions they are in are not their normal operating conditions (Millikan, 2004, p. 83). The magnetosome fails to perform its functions in the same way that a coffee maker fails to perform its function (of making coffee) when nobody puts coffee beans in it. Millikan carefully distinguishes the cases where something malfunctions or, in other words, it fails to perform its function due to the intrinsic damage from the cases where something fails to perform its function doe to the environmental misfortune. For the sake of avoiding confusion, I will use the term “misfunction” for the second type of failures. The coffee maker does not malfunction but misfunctions when nobody puts coffee beans in it. The magnetosome does not malfunction but misfunctions when it is fooled by my bar magnet. What is crucial here is that it might be the case that all misrepresentations involve some etiological misfunctions in Millikan’s version of teleosemantics (Millikan, 1997), but it is not the case that all misrepresentations involve etiological malfunctions. Unlike Millikan, Neander seems be committed to the idea that all misrepresentations involve etiological malfunctions. She discusses the example of a frog (Rana pipiens) that catches and eats flies. The frog, however, responds not just to fries, but to other small, dark, moving things that are not flies, such as BBs. Let us call the frog’s representation of its target “R.” Neander thinks that R represents small, dark, moving things rather than flies. This means that R does not misrepresent as long as it is caused by small, dark, moving things. It does not misrepresent, for instance, when it is caused by a BB. When does R misrepresent, then? It misrepresents, according to Neander, when it is caused by something which is not a small, dark, moving thing. It misrepresents, for example, when it is caused by a snail. Neander discusses a challenge according to which this view does not allow for the possibility of misrepresentation at all. After all, it is very unlikely that R is caused by a snail. In response, she argues that R will never be caused by a snail as long as perceptual systems of the frog are healthy, but “[a] sick frog might R-token at a snail if it was dysfunctional in the right way. Damaging the frog’s neurology, interfering in its embryological development, tinkering with its genes, giving it a virus, all of these could introduce malfunction and error” (Neander, 1995, p. 109). These are the cases where R misrepresents, according to Neander. Then, it looks as though all misrepresentations involve some malfunctions in this account after all. This commitment, however, is problematic. Obviously, one can have a false belief without having any neurological or genetic abnormalities. Neander recognizes this problem. Her answer to it is that the claim that misrepresentations always involve malfunctions is true only for primitive representations in the early stages of visual processing. Consider the case where we see a skinny cow in the dim distance and mistakenly represent it as a horse (Fodor’s example). Here, we may suppose, we misrepresent without malfunctioning, and clearly the content of our perceptual representation goes beyond the physical parameters of the environmental features measured. But this sophisticated representation occurs after much visual processing has already taken place, at least, this is so on computational theories of vision. In such theories, early visual processing does not represent the cow as a horse (or as a cow) but as something which looks a certain way – as having a certain outline texture, color and so on. That is, according to conventional computational theories of perception, initially there is a representation of the physical parameters of the environment as measured by the visual system. It is much plausible that there is no misrepresentation without malfunction at this level (Neander, 1995, p. 132). The claim that misrepresentations always involve etiological malfunctions is true only in the early stages of visual processing. It is not true for sophisticated representations such as beliefs or the perceptual representations in the later stages of visual processing. In sum, both versions of teleosemantics, Millikan’s and Neander’s, are actually free from the view that all misrepresentations involve etiological malfunctions. Misrepresentations might involve etiological misfunctions in Millikan’s account. But, they do not always involve etiological malfunctions. In Neander’s theory, misrepresentations involve etiological malfunctions in the early stages of visual processing. But, it does not generalize to other kinds of misrepresentations.","I have argued that delusions are pathological because they are harmful and malfunctional. They have significant negative impacts on wellbeing. And, some psychological mechanisms, or the connections between them, that are directly or indirectly related to delusions fail to perform their etiological functions due to intrinsic problems. In Section 2, I discussed some possible explanations of the pathology of delusional beliefs. The explanations are not fully satisfactory primarily because they are detached from the considerations about what, in general, makes a condition pathological. My explanation, on the other hand, is an application of a general account of disorders by Wakefield, which successfully explains various kinds of physical and mental disorders. Two types of objections were critically examined in Section 4; (1) it is not the case that all pathological delusions are malfunctional and (2) involving harmful malfunctions are not sufficient for a belief to be pathological. The first type of objections come from the ideas that delusions are playing psychological defence functions, that they are doxastic shear pins, and that they are produced by error-management theoretic biases. In response, I argued that those ideas are perfectly compatible with the claim that delusions involve some etiological malfunctions. The second type of objection comes from the worry that all misrepresentations involve some etiological malfunctions if teleosemantics is correct. In response, I showed that the objection is based upon a misunderstanding about teleosemantics. If we are careful enough about the distinction between malfunction and misfunction, it is not very difficult to see that notable teleosemantics theories are free from such a commitment about misrepresentations."],["We investigated the development of theory of mind use through eye-tracking in children (9–13 years old, n = 14), adolescents (14–17.9 years old, n = 28), and adults (19–29 years old, n = 23). Participants performed a computerized task in which a director instructed them to move objects placed on a set of shelves. Some of the objects were blocked off from the director's point of view; therefore, participants needed to take into consideration the director's ignorance of these objects when following the director's instructions. In a control condition, participants performed the same task in the absence of the director and were told that the instructions would refer only to items in slots without a back panel, controlling for general cognitive demands of the task. Participants also performed two inhibitory control tasks. We replicated previous findings, namely that in the director-present condition, but not in the control condition, children and adolescents made more errors than adults, suggesting that theory of mind use improves between adolescence and adulthood. Inhibitory control partly accounted for errors on the director task, indicating that it is a factor of developmental change in perspective taking. Eye-tracking data revealed early eye gaze differences between trials where the director's perspective was taken into account and those where it was not. Once differences in accuracy rates were considered, all age groups engaged in the same kind of online processing during perspective taking but differed in how often they engaged in perspective taking. When perspective is correctly taken, all age groups’ gaze data point to an early influence of perspective information. --------------------------------------------------------------------------------","Over the last couple of decades, cognitive neuroscience research has shown that brain areas involved in theory of mind (the “social brain”)—our ability to attribute the beliefs, thoughts, desires, intentions, and feelings of others—undergo changes not only during childhood but also during adolescence (Burnett, Sebastian, Cohen Kadosh, & Blakemore, 2011). A substantial number of studies have now provided evidence for structural and functional changes in the social brain during childhood and adolescence. In addition, there is a body of evidence that theory of mind (ToM) is applied more robustly or accurately with age through middle childhood (Devine & Hughes, 2013; Epley, Morewedge, & Keysar, 2004; Lecce, Bianco, Devine, Hughes, & Banerjee, 2014; Surtees & Apperly, 2012; Wang, Ali, Frisson, & Apperly, 2016) and adolescence (Dumontheil, Apperly, & Blakemore, 2010; Vetter, Altgassen, Phillips, Mahy, & Kliegel, 2013). Previous studies have shown that children’s performance on paradigms such as false-belief tasks reaches ceiling at around the age of 5 years (Surian, Caldi, & Sperber, 2007; Wellman, Cross, & Watson, 2001). The question arises as to what factors affect older children and adolescents’ successful application of such abilities. The Director paradigm has been used to investigate the ability to take the perspective of another individual into account in a communicative context (Apperly, Back, Samson, & France, 2008; Brown-Schmidt & Hanna, 2011; Fett et al., 2014; Keysar, Barr, Balin, & Brauner, 2000; Keysar, Lin, & Barr, 2003). In these studies, the participant interacts with another agent (a “director”) to act on a set of objects (Director paradigm; Fig. 1). Crucially, some of the objects are blocked off from the director’s point of view and are visible only to the participant. Thus, when the director talks about an object (e.g., “the large ball”; Fig. 2), the participant should ignore any object that is not visible to the director and instead select a referent from what is in the “common ground,” that is, what is visible to both the participant and the director. This paradigm requires the participant to infer the speaker’s referential intention (a mental state) based on beliefs that differ from his or her own due to the speaker’s ignorance of the presence of an object that would be a potential referent for the instruction given. For example, in the setup shown in Fig. 2, the participant, but not the director, sees a third ball that best fits the description “the large ball” (the basketball) and needs to discount it as the intended referent because the director does not know about that ball. The Director paradigm is useful for the study of the development of the social brain because it can be used to measure the application of aspects of ToM without asking participants to make an explicit judgment about their own or someone else’s perspective or thoughts. Given that even adults manifest less than ceiling performance on the Director paradigm (Brown-Schmidt & Hanna, 2011; Keysar et al., 2000), it is well suited for exploring how ToM abilities develop across childhood and adolescence. Early visual world eye-tracking studies using the Director paradigm demonstrated that adult participants are not able to ignore objects to which they have privileged visual access and that they are less accurate in choosing the object that is visually available to both participants and the director compared with control conditions (Epley et al., 2004; Keysar et al., 2000, 2003). Keysar and colleagues explained these results by suggesting that participants have an initial bias to take an egocentric perspective and that their initial interpretation of the instruction is later adjusted according to the speaker’s knowledge state by a second process that corrects any manifest errors. In contrast, studies using very similar paradigms have found evidence that participants are able to integrate information about the speaker’s beliefs immediately (Hanna, Tanenhaus, & Trueswell, 2003; Heller, Grodner, & Tanenhaus, 2008). This is comparable to findings that individuals may rapidly and automatically compute what other people see (Samson, Apperly, Braithwaite, Andrews, & Bodley Scott, 2010). It seems clear that adults are able to integrate information to some extent about what their interlocutor knows from the earliest stages of referential language processing even though they are liable to attend to objects that are not known to the speaker (Brown-Schmidt & Hanna, 2011). Among adults, performance on perspective-taking tasks has been shown to be affected by factors such as the nature of the verbal stimulus (Barr, 2008), the extent to which the interaction offers information about the interlocutor’s mental state (Brown-Schmidt, 2009a), cultural background (Wu & Keysar, 2007), inhibitory control (Brown-Schmidt, 2009b; Nilsen & Graham, 2009), and mood (Converse, Lin, Keysar, & Epley, 2008). Dumontheil and colleagues (2010) tested 7- to 27-year-old female participants’ performance using a computerized version of Keysar and colleagues’ (2000) Director task.","were presented with a 4 × 4 grid that contained various objects in different slots. A director (an avatar) gave participants instructions about which objects to move and where to move them. As in the original design, some of the slots on the shelves were occluded so that the director could not see what was present on those shelves. The study also included a No-Director condition where participants performed the same task in the absence of a director. In this condition, they were told that the instructions would not refer to items in slots with a gray background. According to Dumontheil and colleagues (2010), the difference between the two conditions is that whereas only executive function processes are required for participants to perform well on the task in the No-Director condition, both ToM processes and executive function processes are required in the Director condition because participants need to take someone else’s perspective into account in order to perform well. Dumontheil and colleagues (2010) found that error rates were higher in the late adolescent group (14- to 17-year-olds) than in the adult group in the Director condition but that they did not differ between the late adolescent and adult groups in the No-Director condition, which relies on executive functions only. This suggests that late adolescents differ from adults in their application of ToM inference when controlling for certain executive function demands. In contrast to the error rates, Dumontheil and colleagues (2010) reported that reaction times on the critical trials in the Director condition remained the same across all age groups. Furthermore, they found that reaction times were longer in the No-Director condition than in the Director condition. According to the authors, this difference suggests that participants approach the Director task in a way that is more efficient than just applying an arbitrary rule. They proposed that the difference in accuracy between the late adolescent and adult groups might stem not from how efficiently or rapidly they can make referential inferences but rather from their propensity to integrate information about the speaker’s beliefs in making those inferences. These suggestions point to a response to the question raised above as to what factors are responsible for the better application of ToM abilities during this stage of development. The current study used visual world eye- tracking measures on a variant of the Director task employed by Dumontheil and colleagues (2010). It is widely accepted that eye fixations can indicate how we process information during spoken language comprehension (Altmann & Kamide, 1999; Tanenhaus, Spivey-Knowlton, Eberhard, & Sedivy, 1995). By measuring when participants fixate their gaze on an object, we can identify which object they are considering as a possible referent at a given point in time. Although eye-tracking has previously been used in variants of the Director task, samples of these studies have been limited to either adults (Keysar et al., 2003; Lin, Keysar, & Epley, 2010; Wu & Keysar, 2007) or younger children (5- and 6-year-olds in Nadig & Sedivy, 2002; 4- to 12-year-olds in Epley et al., 2004). The current study afforded an opportunity to examine and compare how children, adolescents, and adults apply ToM in real time through monitoring their eye gaze while they were performing the task. Specifically, the aim was to explore the time course of the decision process that yields both correct and incorrect responses in participants. Furthermore, the current study also aimed to understand whether ongoing development of executive function processes, specifically inhibition, plays a role in perspective taking. Previous studies have found a correlation between inhibition and the application of ToM inference in children (3- and 4-year-olds in Carlson, Moses, & Claxton, 2004; 4- to 10-year-olds in Hansen Lagattuta, Sayfan, & Harvey, 2014; 4- to 9-year-olds in Lagattuta, Sayfan, & Blattman, 2010; 2- to 5-year-olds in Nilsen & Graham, 2009) and in adults (Brown-Schmidt, 2009b). A meta-analysis of 100 studies with 3- to 6-year-olds by Devine and Hughes (2014) highlighted the association between executive function and ToM. Inhibitory control, like other executive skills, is still maturing during adolescence (Leon-Carrion, García-Orza, & Pérez-Santamaría, 2004; Luna, Padmanabhan, & Hearn, 2011). A study by Vetter and colleagues (2013) found that inhibitory control, as measured with an anti-saccade task, predicted age-related variance in affective ToM during adolescence. No other study, however, has looked at the role of inhibitory control in adolescents’ ability to inhibit one’s perspective and consider someone else’s perspective. Therefore, we measured participants’ inhibitory control through two variants of the Go–NoGo task in order to investigate whether individual differences in inhibitory control can account for their performance on the Director task. Based on the evidence of ongoing functional development of neural processes related to ToM and the results of Dumontheil and colleagues (2010), we expected that age-related differences in participants’ accuracy would be greater in the Director condition than in the No-Director condition. We also expected to see age-related differences in inhibitory control, which may partially account for participants’ performance on the Director task. We measured the extent to which participants considered the target object relative to the distractor by computing the target advantage score, which is the average probability of looks at the target object minus the average probability of looks to the distractor (Brown-Schmidt, 2009b; Kronmüller & Barr, 2015). This allowed us to examine and compare how each age group used ToM to identify the referent. According to the perspective adjustment model proposed by Keysar and colleagues, participants have an initial egocentric bias, which is then adjusted, during a second process, to the speaker’s perspective. Therefore, this theory predicts that participants will fixate on the distractor first and, if the second process of adjusting is successful (i.e., on correct trials), then participants will fixate their gaze on the target object and give a correct response. However, if the second process fails and perspective taking does not occur, then we would expect participants’ gaze to remain on the distractor, not to fixate on the target, and participants to eventually give an incorrect response. As such, it is predicted that although both adolescents and adults would fixate on the distractor object initially, demonstrating an initial egocentric bias, adults would then be more likely to switch to fixating their gaze on the target object and do so more quickly than adolescents. In other words, both adolescents and adults would have a smaller target advantage initially, but adults would have a bigger target advantage score than adolescents in earlier time regions. Other models of perspective taking found in Hanna and colleagues (2003) and Brown-Schmidt, Gunlogson, and Tanenhaus (2008) suggest that perspective information is available from the outset but that the extent to which it is integrated online depends on how sensitive participants are to such information (Brown- Schmidt, 2012) or on participants’ ability to integrate that information with linguistic and visual contextual information (Hanna et al., 2003). According to these proposals, an incorrect response may reflect a simple failure to integrate perspective from the start of a trial. If this is the case, then patterns of eye gaze may differ from the onset of the critical phase of the instruction between correct and incorrect trials. Participants ~~~~~~~~~~~~ A total of 65 participants took part in this study, constituting three age groups: child (n = 14, 9–13 years old, M = 11.2 years, SD = 1.2), adolescent (n = 28, 14–17.9 years old, M = 16.2 years, SD = 1.1), and adult (n = 23, 19–29 years old, M = 23.0, SD = 2.8). The age groups were divided to match the age groups of Dumontheil and colleagues (2010). Adult participants were recruited from the University College London (UCL) Psychology participant pool, whereas children and adolescents were recruited from London schools. All participants were native English speakers. Data from three adults were excluded due to technical errors. We also excluded data from two adolescents who did not understand the instructions of the No-Director condition. Verbal ability was measured in all participants using the vocabulary subtest of the Wechsler Abbreviated Scale of Intelligence (WASI; Wechsler, 1999). Average verbal IQ scores were 124 (SD = 8.0) for adults, 118 (SD = 6.9) for adolescents, and 123 (SD = 10.2) for children. A one-way analysis of variance (ANOVA) showed a significant difference in verbal IQ scores between groups, F(2, 57) = 5.97, p = .03. Bonferroni-corrected post hoc comparisons showed that adolescents performed worse than adults (p = .04), but verbal IQ did not significantly differ between children and adolescents (p = .22) or between children and adults (p = 1.00). Parents/guardians of all child and adolescent participants as well as the adult participants were given information sheets prior to the study. Informed consent was obtained from the parents/guardians for all child/adolescent participants and from all adult participants. This study was approved by the UCL research ethics committee. Director task The current experiment had two within-participant factors: condition (Director or No Director) and trial type (Control or Experimental). It had one between- participants factor: age group (child, adolescent, or adult). The task design and stimuli were based on a previous study by Dumontheil and colleagues (2010; see also Apperly et al., 2010). Participants were presented with a visual scene of a 4 × 4 set of shelves containing eight different objects and were asked to move one of the objects in each trial. In the Director condition, a “director” was shown standing behind the shelves, viewing the shelves from behind (Fig. 1). Participants were asked to listen to the director’s instructions to move one of the objects (e.g., “Move the large ball up”) and respond. Participants were told that they should take the director’s viewpoint into account when following the director’s instructions. They were told that objects in slots with a gray background were visible only to them, whereas the other objects could be seen from either side of the shelves. In Experimental trials, the instruction referred to one object (“target”) given the director’s point of view but would refer to another object (“distractor”) if one assumed participants’ perspective (Fig. 2A). As such, participants needed to take the director’s perspective into account in order to respond correctly in an Experimental trial. In Control trials, the distractor object was replaced by an irrelevant object and the instruction referred to an object that was visible to both participants and the director (Fig. 2B). In Filler trials, instructions referred to single objects that were visible to both participants and the director (e.g., the turtle in Fig. 2). In the No- Director condition, the only difference was that the director was removed from the stimuli (Fig. 2C and D). Participants were told that the auditory instructions would refer only to items in clear slots and not items that were in slots with a gray background. A total of 48 pairs of shelf configurations were created as Experimental and Control trials. All shelf configurations depicted eight objects and included either three (Experimental trials) or two (Control trials) exemplars of the same object that differed in position (top/bottom) or size (large/small). In Experimental trials, the exemplars were distributed such that the distractor object (the top-most, bottom-most, smallest, or largest object) was in a slot with a gray background, whereas the target object (the second top-most, bottom-most, smallest, or largest object) and the third object (“Object 2”) were in a clear slot (Fig. 2A and C). In the Control trials, the distractor was replaced by an irrelevant object (Fig. 2B and D). The rest of the objects were unique objects distributed among three slots with a gray background and two clear slots. Another 48 pairs of shelf configurations were created for the Filler trials, in which the instructions referred to single objects that were in a clear slot. Taken together, the stimuli were divided such that half were presented in the Director condition and the other half were presented in the No-Director condition (counterbalanced across participants), and each participant saw 12 Experimental trials, 12 Control trials, and 24 Filler trials in the Director and No-Director conditions. The order of stimulus presentation was counterbalanced between participants. The materials and design of the current study differed from those Dumontheil and colleagues’ (2010) study in three ways. First, we included a greater number of Experimental and Control trials to increase statistical power and allow eye-tracking data analysis. Second, whereas each shelf configuration was used with three successive instructions in Dumontheil and colleagues’ study, a different shelf configuration was used in each trial in the current study to minimize participants’ learning strategies. Third, whereas participants needed to pretend to drag the object in the previous study, they were able to click and drag any object on the grid in the current study. Inhibitory control tasks To measure participants’ inhibitory control, we used two different Go–NoGo tasks, both of which had one within-participant factor trial type (Go or NoGo). The Simple Go–NoGo task was based on the standard Go–NoGo paradigm (Simmonds, Pekar, & Mostofsky, 2008). A colored square was presented on the left or right side of the screen in each trial. Participants needed to indicate which side of the screen it appeared on if the square was green (a Go trial) and to inhibit their response if the square was red (a NoGo trial). The Complex Go–NoGo task was identical except that it used yellow and blue squares and included a 1-back working memory (WM) requirement (see Simmonds et al., 2008), such that participants needed to indicate on which side of the screen the square was shown (Go trials) except when a blue square was preceded by a yellow square (NoGo trials).","Participants were tested individually in one session lasting approximately 45 min. They completed the various tasks in the following order: (1) the Director task (Director condition and then No-Director condition), (2) the inhibitory control tasks (Simple Go–NoGo task and then Complex Go–NoGo task), and (3) the vocabulary subtest of the WASI. Eye movements during the Director task were recorded using a Tobii TX300 eye-tracker at a sampling rate of 60 Hz. Stimuli for the Director task were presented using E-Prime 2, and the Go–NoGo tasks were programmed in Cogent (http://www.vislab.ucl.ac.uk/Cogent/index.html) running in Matlab 7.0 (MathWorks). Standardized instructions were read to participants prior to the Director task. For the Director condition, participants were told that the director had a different view of the shelves (Fig. 1) and that the director’s point of view must be considered when following the director’s instructions to move objects. We asked participants to give an example of an object that both they and the director could see as well as an object that the director could not see to ensure that participants understood the task. For the No-Director condition, participants were told that the director was no longer present and that instructions would refer only to items in clear slots, such that they should ignore items in slots with a gray background when performing the task. The Director condition was always presented prior to the No-Director condition in order to prevent participants from applying the rule provided in the No-Director condition (Dumontheil et al., 2010). Participants were presented with two Filler practice trials before each condition. In each trial, the visual stimulus and auditory instructions were presented over a period of 2.2 s. The visual stimulus remained on the screen for another 3.8 s, during which participants responded by clicking an object and dragging it to a different position. A response was considered correct if the target object had been selected. Response times (RTs) were measured from the onset of the display. In the Go–NoGo tasks, each square was presented for 400 ms, with a 600- to 800-ms jittered fixation cross intertrial interval (ITI). Participants responded by pressing the left or right key using their right index or middle finger, respectively (adapted from Watanabe et al., 2002). Participants performed two practice blocks prior to each task; the first practice block presented 10 Go trials to establish a habitual response, and the second one presented 6 Go trials and 4 NoGo trials. Practice was repeated if participants made two or more errors. Each participant performed 80 test trials on each task (25% NoGo). Behavioral data analyses All analyses were processed in SPSS and R. Statistical significance was set at p < .05. Bonferroni-corrected post hoc t-tests were performed to explore significant main effects and interactions further. Director task Mixed repeated measures ANOVAs with two within-participant factors (condition and trial type) and one between-participant factor (age group) were performed on participants’ mean accuracy (percentage errors) and reaction times (RTs). Data for Filler trials were not analyzed; participants made fewer than 2% errors on these trials on average. Because verbal IQ differed among groups, we conducted an additional mixed repeated measures ANOVA that included verbal IQ as a covariate. Inhibitory control tasks For both the Simple and Complex Go–NoGo tasks, mean accuracy (percentage errors) was calculated for each participant in each trial type (Go or NoGo) and median RT was calculated for correct Go trials. A 2 × 3 mixed ANOVA with a within-participant factor trial type (Go or NoGo) and a between-participant factor age group (child, adolescent, or adult) was performed for each task. One-way ANOVAs were performed to examine the effect of age group on participants’ RTs in Go trials in each task. Association between Director and inhibitory control tasks Regression analyses were performed to investigate whether age-related changes in Director task performance were associated with performance on the inhibitory control tasks. The difference in percentage errors between Director and No-Director Experimental trials, which is the critical measure of interest, was entered as the dependent variable. In a first step, age was entered as a continuous variable. In a second step, the four percentage error measures of the Simple and Complex Go–NoGo trials were entered in a stepwise regression. Finally, in a third step, IQ was entered to assess whether it accounted for variance associated with age in this model. Eye-tracking data analyses Eye movements were analyzed by computing a target advantage score, which is the average probability of looks at the target object minus the average probability of looks at the distractor (Kronmüller & Barr, 2015). A look to the target was defined as a look to the object or the slot in which the object was located. Likewise, a look to the distractor (or the irrelevant object in Control trials) was defined as a look to the object or the slot in which the object was located. Note that because the distractor was an irrelevant object in the Control trials, target advantage scores in the Control trials represented the difference between looks to the target and looks to the irrelevant object. Target advantage scores were calculated over 50-ms time bins for the figures where 0 ms was the noun onset. For statistical analyses, the target advantage scores were calculated for five different time windows (regions) during a trial. The five regions were time- locked to the onset of words in the auditory stimuli. Region 1 was the verb (“move”), Region 2 was the article (“the”), Region 3 was the modifier (“large”), Region 4 was the critical noun region (“ball”), and Region 5 was the directional preposition (“up”). Analyses of data in these regions were offset by 200 ms to account for the time required for planning and launching an eye movement (Hallett, 1978). Data in each region were analyzed for the Director and No-Director tasks separately. Because we did not have predictions regarding participants’ eye gaze prior to the modifier (e.g., “large”), we focused our analyses on data in Regions 3 to 5 only. First, 2 (Condition) × 2 (Trial Type) × 3 (Age Group) mixed ANOVAs were performed on correct trials for each region. Data from 2 participants in the child group were excluded from this analysis because they did not have enough correct trials. Second, 2 (Accuracy) × 3 (Age Group) mixed ANOVAs were performed separately for correct and incorrect Director Experimental trials, the key trials of interest, for each region. Because children and adolescents did not have many correct trials (in some cases only one or two), and some adults did not have any incorrect trials, some participants (5 children, 5 adolescents, and 9 adults) needed to be omitted from this analysis due to the lack of eye gaze data. Given the clear differences in accuracy between groups, we believe that it was not suitable to combine correct and incorrect trials in the analysis. However, interested readers can find the results of such analyses in the online supplementary material. Accuracy Fig. 3 shows participants’ percentage error as a function of age group, and Table 1 shows the results of the statistical analysis. There were significant main effects of condition (ηp2 = .591), trial type (ηp2 = .675), and age group (ηp2 = .220), which were qualified by significant Age Group × Trial Type and Age Group × Condition two-way interactions and a significant three-way interaction (p = .013, ηp2 = .142). We explored this interaction further by analyzing the Experimental and Control trials separately. No significant effects were found in the Control trials (all ps > .08). In contrast, in Experimental trials there was a main effect of condition, F(1, 57) = 83.57, p < .001, ηp2 = .595, a main effect of age group, F(2, 57) = 8.10, p = .001, ηp2 = .221, and a significant interaction between condition and age group, F(2, 57) = 3.67, p = .032, ηp2 = .114. Follow-up analysis examining Experimental trials in the Director and No-Director tasks separately revealed a significant effect of age group in Director Experimental trials, F(2, 57) = 8.10, p = .001, ηp2 = .221, but not in No-Director Experimental trials, F(2, 57) = 1.85, p = .167. Follow-up t-tests for the Director Experimental trials showed that both the child and adolescent groups made significantly more errors than the adult group (child: t(32) = 3.73, p = .001, d = 1.321; adolescent: t(44) = 2.94, p = .005, d = 0.887), but there was no difference between the child and adolescent groups, t(38) = 1.36, p = .183. Because the age groups significantly differed in terms of their verbal IQ, we repeated the 2 × 2 × 3 mixed repeated measures ANOVA with standardized IQ (z-score) as a covariate. The same main effects and interactions were found. To summarize, children and adolescents made more errors than adults only in the Director Experimental trials, which required taking the perspective of the director into account, in contrast to the rule-based control condition (No-Director Experimental trials). This replicated the behavioral results reported in Dumontheil and colleagues (2010). Reaction time Analysis of RT showed that children responded more slowly than the adolescents and adults in all trial types except for the Director Experimental trials, which showed no difference among age groups (see supplementary material for analyses). These results are also in line with those reported in Dumontheil and colleagues (2010). Simple Go–NoGo The repeated measures ANOVA revealed significant main effects of trial type, F(1, 57) = 50.34, p < .001, ηp2 = .469, and age group, F(2, 57) = 5.29, p = .008, ηp2 = .156 (Fig. 4). More errors were made in NoGo trials than in Go trials and the child group made more errors than the adolescents (p = .027) and adults (ps < .03), who did not differ (p = 1.00). The interaction between age group and trial type was not significant (p > .20). Complex Go–NoGo We found significant main effects of trial type, F(1, 57) = 108.87, p < .001, ηp2 = .656, and age group, F(2, 57) = 11.44, p < .001, ηp2 = .286, in the Complex Go–NoGo task and no interaction between these factors (p > .10) (Fig. 4). Here, the child and adolescent groups did not differ (p > .10), but they both made more errors than the adult group (ps < .01). Simple and Complex Go RT Analyses of median RT in correct Go trials revealed that, in both the Simple and Complex Go–NoGo tasks, the child group was significantly slower than the adolescent and adult groups, who did not differ from each other (see supplementary material for RT analyses). Association between inhibitory control tasks and Director task Regression analyses showed that age significantly accounted for 14.4% of variance in the difference in percentage error between Director and No- Director Experimental trials (Table 2). Inhibitory control, as measured by the Simple NoGo percentage error, accounted for an additional 9.2% of variance, with more errors on Simple NoGo trials predicting more errors in Director versus No-Director Experimental trials. The other three measures of Go–NoGo percentage error were not significant predictors (all βs < .124 and ps > .370). The effect of age was still significant in the second model, suggesting partially independent effects of age and inhibitory control. Finally, entering verbal IQ as an additional regressor did not improve the model, and both the independent effects of age and Simple NoGo percentage error remained significant (Table 2). Eye-tracking results ~~~~~~~~~~~~~~~~~~~~ Table 3 shows mean target advantage scores across conditions, and Table 4 shows the results of the statistical analyses. Fig. 5A–D show plots of average target advantage over time. No significant effects were found in Region 3 (“large”) (all ps > .20). Main effects of condition and trial type were marginal in Region 4 (“ball”) (ps < .10) and were significant in Region 5 (“up”) (condition: ηp2 = .042; trial type: ηp2 = .234). The Condition × Trial Type interaction was significant in both Region 4 (ηp2 = .031) and Region 5 (ηp2 = .035). Follow-up analyses revealed no significant effects in Control trials in either region (ps > .30), whereas the target advantage in Experimental trials was significantly smaller in the No-Director task than in the Director task in both Region 4, F(1, 57) = 8.22, p < .01, ηp2 = .055, and Region 5, F(1, 57) = 10.84, p = .002, ηp2 = .090. No significant effects involving age group were observed in any of the regions (all ps > .15). To summarize, in correct trials, participants showed a smaller target advantage in No-Director Experimental trials than in Director Experimental trials on hearing the noun. Crucially, their eye movement patterns did not vary across age groups in any of the regions. We conducted additional analyses to examine the differences in eye gaze between correct and incorrect Director Experimental trials, the key trials of interest in the Director paradigm (Tables 5 and 6). We included an additional region (pre-response region) to check whether participants looked at the object they chose. This region was defined as the 200 ms before response time in a given trial. Graphs of target advantage again indicate a similar pattern across age groups (Figs. 5B and 6). Mixed 2 (Accuracy) × 3 (Age Group) ANOVAs revealed a main effect of accuracy across Regions 3 to 5 and the pre- response region (Region 3: ηp2 = .240; Region 4: ηp2 = .271; Region 5: ηp2 = .224; pre- response: ηp2 = .118). Participants had a bigger target advantage in incorrect trials than in correct trials in Regions 3 to 5; conversely, they had a bigger target advantage in correct trials than in incorrect trials in the pre-response region. No significant effects involving age group were observed (all ps > .30). To summarize, analyses of eye movement data in the Director Experimental trials revealed clear differences between correct and incorrect trials, but there was no significant effects involving age group in any of the regions. In other words, there were no significant age-related differences in participants’ eye gaze patterns in both correct and incorrect Director Experimental trials. Furthermore, the effect of accuracy observed in Regions 3 to 5 was reversed in the pre-response region, such that participants initially showed a greater target advantage in the incorrect trials than in the correct trials, and it was only right before they responded that they showed a greater target advantage in correct trials than in incorrect trials.","In this study, we collected behavioral and eye-tracking data to investigate online use of ToM during perspective taking in children, adolescents, and adults. Experimental trials of the Director task required participants to take into account the director’s perspective to determine the intended referent in the instructions (e.g., “large ball”). The No-Director condition required participants to follow an explicit avoidance rule to determine the correct referent. Whereas both the Director and No-Director conditions involved executive function, in particular inhibition, only the Director condition required an inference about the speaker’s intentions given the speaker’s ignorance of certain objects (Dumontheil et al., 2010). The current study had three main objectives: (a) to replicate the behavioral findings of Dumontheil and colleagues (2010), (b) to assess the role of inhibitory control in this task across age groups, and (c) to compare the time course of information processing among children, adolescents, and adults through eye-tracking in order to determine in what ways the deployment of ToM differs across age groups. Behavioral data ~~~~~~~~~~~~~~~ We found that in the Director Experimental condition, the child and adolescent groups performed worse than the adult group, but participants’ accuracy did not differ across groups in the No-Director Experimental condition. These results are in line with Dumontheil and colleagues (2010) findings and provide further evidence for age-related differences between adolescents’ and adults’ tendency to take someone else’s perspective into account. Conversely, response times decreased with age in all conditions except in Director Experimental trials, which did not vary with age, and are again similar to those observed by Dumontheil and colleagues. Inhibitory control We observed age-related differences in participants’ performance on both inhibitory control tasks. In the Simple Go–NoGo task, the child group made significantly more errors and was significantly slower than the adolescent and adult groups. These results are in line with previous studies suggesting that with a low-level cognitive load inhibitory control performance reaches adult level performance by approximately 14 years of age (Lamm, Zelazo, & Lewis, 2006; Leon- Carrion et al., 2004; Luna, Garver, Urban, Lazar, & Sweeney, 2004). In the Complex Go–NoGo task, the child and adolescent groups performed worse than the adult group. Because the Complex Go–NoGo task included a 1-back WM requirement, these results are also in line with studies investigating WM that found age-related differences between adolescents and young adults (Conklin, Luciana, Hooper, & Yarger, 2007; Luna et al., 2004). Association between inhibitory control tasks and Director task What was novel in this study was the investigation of the relationship between the inhibitory control tasks and the Director task. The results show that inhibitory control as measured by Simple NoGo accuracy accounted for some of the variance in accuracy difference between Director and No-Director Experimental trials. Critically, inhibitory control accounted for only some of the variance due to age given that age remained a significant predictor, suggesting that some additional factors are behind age-related changes. These results are in line with a previous study by Vetter and colleagues (2013), who found that 15% of the variance of affective ToM performance was uniquely explained by age, indicating independent effects of age and inhibition. Interestingly, no relationship was found between participants’ performance on the Complex Go–NoGo task and the Director task. This is surprising because performance on the Complex Go–NoGo task showed age-related differences between the adolescent and adult groups. This could be because the Complex Go–NoGo task placed greater WM demands than the Director task. A study that explored the relationship between WM and perspective taking (Lin et al., 2010) found that participants with lower WM capacity performed more poorly on the Director task than participants with greater WM capacity and that participants’ performance on the Director task was worse during high-WM load trials than during low-WM load trials. However, Lin and colleagues’ (2010) study did not include a No-Director condition with matched general executive function demands, whereas the current study used the difference in accuracy between Director Experimental and No-Director Experimental trials as the measure of interest. Future research might shed light on these differences in results by using separate measures of WM and inhibition (Vetter et al., 2013). A possible limitation for the interpretation of the behavioral results is the difference in verbal IQ among age groups. However, the significant interactions with age observed in the Director task were still present when verbal IQ was added as a covariate. Moreover, adding verbal IQ as a predictor in the multiple regression did not affect the results given that both the independent effects of age and Simple NoGo accuracy remained significant. Eye-tracking data ~~~~~~~~~~~~~~~~~ Through eye-tracking, we were able to investigate additional underlying factors in Director task performance. Examining correct and incorrect trials separately indicated that adults, adolescents, and children did not differ in their online processing of the task (Fig. 5). Analyses of Director Experimental trials data showed opposite effects of accuracy in the earlier regions (Regions 3–5) and the pre-response region, such that participants initially showed a greater target advantage in the incorrect trials than in the correct trials. They showed a greater target advantage in correct trails than in incorrect trials only right before they responded. These results seem inconsistent with the perspective adjustment model (Keysar et al., 2000, 2003), according to which one would expect that on both the correct and incorrect trials participants first attend to the distractor, the best-fit referent from an egocentric perspective. On correct trials participants would adjust to considering the director’s perspective and attend to the target, whereas on incorrect trials the second adjustment process would fail because it is costly and participants would remain focused on the distractor. However, as is evident from the incorrect trials for all age groups, participants appear to consider the target early on in the trial and then, before responding, their eye gaze shifts to the distractor. Based on the eye-tracking results, it seems that at the point where the director indicates which object should be moved (Region 4), participants on correct trials are already on a path to correctly considering the director’s perspective. However, participants’ strategy seems to be to first look at the objects that participants are not going to choose (or the objects they should not pick) as a process of elimination before focusing on the object that they will ultimately choose. This pattern has not been reported before. Other studies, such as Hanna and colleagues (2003), Heller and colleagues (2008), and Barr (2008), showed that bias in participants’ eye gaze builds steadily toward the target after initial interference from the distractor. What sets our study apart from these eye-tracking studies is the fact that the distractor is the best fit for the description, making the task of ignoring the privileged object particularly challenging. In addition, the description itself contains a relational modifier (e.g., “big,” “bottom”) that implicitly refers to a contrast set. Given these two factors, it should not be surprising that participants adopt a strategy of checking all objects of the same type (e.g., all balls on the display) to ensure that they choose the correct one. Our eye gaze data of Object 2 (e.g., the second commonly viewable ball) suggests this also. It shows that prior to focusing on the object they chose, participants paid similar attention to each of the other two objects they eliminated (see Figs. S3–S5 in supplementary material). The only other studies that used a setup similar to ours are those reported in Keysar and colleagues (2000) and Wu and Keysar (2007). Neither of those articles reported eye gaze data in full, but their results are consistent with ours in that they found more looks to the distractor overall and first looks to the distractor were earlier than first looks to the target. Adults, adolescents, and children in the current study seemed to follow a process of elimination strategy in control trials as well, where only the two commonly viewable objects denoted by the noun (e.g., “ball”) were present and only one fit the full description (“large ball”). We propose that participants focus first on the second object to exclude it before looking at the target (see Figs. S3–S6 in supplementary material). Therefore, there may be a consistent strategy across all conditions, and all age groups, of looking at the objects that are not going to be chosen prior to focusing on the object to be chosen. If we accept that participants adopt a process of elimination strategy, then our results suggest that the eye gaze pattern on correct trials is influenced by the actual beliefs of the director early on, as we see the pattern emerge in Region 4 when participants process the modifying adjective (“large”). The idea that eye gaze data reflect an influence of the speaker’s perspective at an early stage is in line with a number of previous perspective-taking studies mentioned above (Brown-Schmidt et al., 2008; Hanna et al., 2003; Heller et al., 2008). The results are also consistent with Nadig and Sedivy’s (2002) observation that 5- and 6-year-old children showed early sensitivity to a speaker’s perspective in a much simplified director task. These articles proposed an alternative to the perspective adjustment model, claiming that mental state information, like any other relevant information, is potentially available to be integrated into referential decision processes from the outset. On this constraint-based view, the extent to which mental perspective information is used depends on the extent to which other constraints are conflicting and how salient or available the perspective information is (Brown-Schmidt & Hanna, 2011; Samson et al., 2010). Hanna and colleagues (2003) argued that the Director paradigm used in Keysar and colleagues (2000) and in this article makes conflicting cues related to the linguistic form particularly strong because the occluded object (the distractor) is in fact the best fit for the description. Other studies such as Brown-Schmidt (2009b) have suggested that varying the strength (or quality) of cues to the speaker’s mental perspective can affect online referential processes. If varying cues to mental perspective can have an impact on ToM integration, it seems plausible from this constraint-based perspective that individuals may differ in the extent to which their referential processes exploit a given cue. Thus, an explanation of the age-related differences that are not accounted for by inhibitory control may lie in differences in participants’ sensitivity of online processes to mental perspective information. This is a hypothesis that requires further exploration. If we assume that accuracy differences in our Director task are a product of varying abilities to integrate perspective information in incremental referential processes, rather than at a later corrective stage, then we can make sense of the reaction time results reported here and in Dumontheil and colleagues (2010). These results showed shorter RTs in the Director task compared with the No- Director task and also no RT difference in the Director condition between age groups. The observation made about these results is that the Director task engages a more efficient or rapid process than simple explicit rule following and does so to the same extent across age groups. Taken together, the current results provide evidence for age-related differences between adolescents and adults in their online use of ToM. Contrary to perspective adjustment accounts of the Director task, our results suggest that all age groups appear to engage in the same kind of online processes during perspective taking but differ in how often mental state information informs incremental decision processes. Taking our results in the wider context of research into online use of ToM, we see one possible source of the difference between age groups as being their sensitivity to available cues to mental state information."],["Understanding the cultural commonalities and specificities of facial expressions of emotion remains a central goal of Psychology. However, recent progress has been stayed by dichotomous debates (e.g. nature versus nurture) that have created silos of empirical and theoretical knowledge. Now, an emerging interdisciplinary scientific culture is broadening the focus of research to provide a more unified and refined account of facial expressions within and across cultures. Specifically, data-driven approaches allow a wider, more objective exploration of face movement patterns that provide detailed information ontologies of their cultural commonalities and specificities. Similarly, a wider exploration of the social messages perceived from face movements diversifies knowledge of their functional roles (e.g. the ‘fear’ face used as a threat display). Together, these new approaches promise to diversify, deepen, and refine knowledge of facial expressions, and deliver the next major milestones for a functional theory of human social communication that is transferable to social robotics. --------------------------------------------------------------------------------","Are facial expressions of emotion universal across cultures or are they culture specific? That is, can Chileans understand the emotions of the Chinese from reading their facial expressions, and vice versa? Such questions (and more) have been at the center of one of the longest standing debates in Psychology — whether facial expressions of emotion are hard-wired and universal, or learned and thus subject to cultural variability. By virtue of the dichotomous nature of the debate — that is, nature versus nurture, essentialism versus constructivism — the direction and focus of the field has followed a cyclic, back- and-forth seesaw pattern for over a century (e.g. see [1,2]). Several major milestones have marked this era: Darwin's revolutionary theory of the biological and evolutionary origins of facial expressions that supported views of universality [3]; later counteractions by rising cultural relativism (e.g. [4]); Ekman's pioneering work showing the pan-cultural recognition of six face movement patterns as basic emotions (e.g. [5]) that cemented the recent dominant view that facial expressions of emotion are universal. Indeed, most introductory Psychology textbooks — a litmus test for the main thinking in the field — tend to report that six specific face movement patterns universally convey six basic emotions across all cultures, with cultural variance often consigned to a footnote (if at all). Consequently, research in the past 50 years or so has focused almost exclusively on these six facial expressions with little exploration of the cultural diversity in face movement patterns and the social messages they convey. Yet, in the last decade or so the emergence of an interdisciplinary scientific culture using new, imported methods and concepts is now pushing research boundaries toward a broader, deeper, and more refined understanding of facial expression communication. Consequently, several significant new advances have questioned the true universality of facial expressions of emotion, instead revealing a more complex account that combines traditionally distinct views (e.g. nature versus nurture). Such an approach sharply contrasts with the cyclic, seesawing patterns of past research, and mark the beginning of a new research culture that has the potential to deliver significant new milestones that cut across fundamental (e.g. Anthropology and Psychology) and applied (e.g. Computing Science and Social Robotics) disciplines of social communication. In this review, we will highlight two recent pieces of research that have used creative, out-of-the-box thinking to advance knowledge of facial expressions of emotion across cultures, and generate new questions that will guide future research directions. To appreciate the relevance and scope of these new approaches, it is first useful to outline the classic methods used to understand facial expressions across cultures. CLASSIC APPROACHES TO UNDERSTANDING FACIAL EXPRESSIONS OF EMOTION ACROSS CULTURES -------------------------------------------------------------------------------- Since the inception of the universality debate, a central goal has been to identify which face movement patterns are common across cultures and which are culture-specific. However, doing so is genuinely challenging because the human face can generate an incredible diversity of facial expressions. To illustrate, consider that the face can produce over 40 individual movements, measured as Action Units (AUs) [6], such as Upper Lid Raiser (AU5), Nose Wrinkler (AU9), and Lip Stretcher (AU20), each of which can be combined in different numbers to create a vast array of complex patterns. Each AU can also be activated with a specific movement pattern across time based on, for example, different acceleration, peak latency, and amplitude, which further magnifies the number of movement combinations the face can generate. Indeed, due to these complex variations Ekman noted that ‘it is exceedingly difficult to observe the common facial expressions of emotion across cultures’ ([7], p. 234). One of the most popular approaches to understanding facial expressions across cultures has involved selecting images of facial movement patterns thought to convey specific basic emotions based on theory and naturalistic observation, and testing their recognition across cultures (e.g. [8–10]). Most notably, Ekman and colleagues used this approach to show that six specific face movement patterns thought to represent basic emotions of happy, surprise, fear, disgust, anger, and sad elicited above chance recognition accuracy across several distinct cultures (e.g. [5]). Consequently, these six face movement patterns, each represented as a specific combination of AUs — for example, ‘happy’ involves Cheek Raiser (AU6) and Lip Corner Puller (AU12), whereas ‘sad’ involves Inner Brow Raiser (AU1), Brow Lowerer (AU4) and Lip Corner Depressor (AU15) — became widely considered as the gold standard in universal displays of emotions thought to be basic. However, the classic approach of using top-down, theory-driven methods to select and test specific face movement patterns (i.e. the AU patterns proposed by Ekman and colleagues) and the social messages they convey (i.e. six emotion categories) has substantially restricted knowledge of how the face communicates emotion messages. Specifically, such methods are typically grounded in the experimenter's culture and can thus reflect culture-specific intuitions and observations more than human behavior more broadly (i.e. a bias of cultrocentrism) — for example, see [11•,12–14]. Perhaps unsurprisingly then, numerous cross-cultural studies have shown that these ‘universal’ face movement patterns are in fact not universally recognized across cultures, at least in terms of equal performance levels (see [15,16] for recent reviews. See also Gendron in this special issue). Instead, these face movement patterns are best recognized by Westerners and elicit significantly lower performance in other cultures particularly for ‘fear,’ ‘disgust’ and ‘anger.’ Thus, while this approach has delivered recognizable representations of Western facial movement patterns of emotion, equivalents in other cultures remain largely unknown. Knowledge has been further restricted by limiting the exploration of the social messages that face movement patterns can convey. For example, classic approaches have focused primarily on only six emotion categories, which, in addition to representing a small proportion of the nuanced emotion messages required for the complex social exchanges of daily life, could instead reflect the main emotion concepts of Western culture (e.g. see [17•,18]). Furthermore, classic approaches have focused mostly on the inner emotional states of the transmitter — for example, a lowered brow with tightened lips and eyes indicates that ‘he is angry’ — rather than their predicted behaviors toward others — for example, ‘he will attack me’ — which overlooks key aspects of human social communication and interaction ([19]; see also [20]). Finally, face movements are complex dynamic information patterns (see [21] for a review) where the temporal order and activation of different AUs provide important diagnostic information for emotion categorization (e.g. [22] see also [23]). Classic approaches have mostly used static displays such as images of posed face movements, or created the illusion of movement by progressively morphing between two different static images (e.g. happy and sad). Yet, neither method can capture nor explore how the dynamic parameters of face movements — for example, AU amplitude, acceleration, or peak latency — influence the interpretation of face movement patterns. Classic approaches have undoubtedly advanced understanding of how face movements can convey different emotions, but knowledge remains limited to only a small and (Western) specific set of facial patterns and social messages. Consequently, substantial knowledge gaps remain both in the characterization of face movement patterns (in terms of AU composition and their respective timings) and the messages they convey within and across cultures. Rather, revealing the true diversity of dynamic face patterns along with their cultural commonalities and specificities first requires a broader understanding of the face movements used in different cultures and the messages they convey (see also [24] for further discussions). We will now outline two key studies that have made significant advances toward this goal.","In recent work, Jack and colleagues [25••] diversified and deepened knowledge of how face movements convey emotions across cultures using a novel data-driven approach to objectively and mathematically model dynamic face movement patterns. Figure 1(a) illustrates this approach. On each experimental trial, a dynamic face movement generator [26] creates a random facial animation by randomly selecting a subset of individual face movements (i.e. AUs; see colored labels on left) and applying a random dynamics to each AU (see color-coded curves). The cultural observer categorizes the facial animation by emotion (e.g. disgust) and rates its intensity (e.g. strong) when the face movement pattern correlates with their prior knowledge of that face movement pattern and its associated message (e.g. ‘strong disgust’). If the pattern does not correspond to one of the response options (here, the six classic emotions) the observer selects ‘other.’ After many such trials, measures of statistical association (e.g. regression, correlation, mutual information) are used to build a relationship between the dynamic patterns presented on each trial and the observer's responses. The analysis thus produces, for each observer independently, a mathematical model of the dynamic face movement patterns that convey these specific emotions to individuals in a given culture. These mathematical models can then be submitted to rigorous analyses to extract patterns that are common across cultures and those that are culture-specific. Such an approach provides several advantages, particularly in relation to the debate about the universality of facial expressions of emotion. First, data-driven methods typically make few a priori assumptions about which stimulus patterns will convey which messages to whom, thereby allowing a much broader and agnostic exploration of face movement patterns as carriers of relevant information. This approach also makes intuitive sense for the purposes of objective study, particularly of groups for which there may be little existing knowledge (e.g. Sentinelese society). Second, building detailed, quantitatively characterized facial movement patterns (i.e. an information ontology) enables precise and objective analyses and comparisons to show how face movement patterns are similar or different across cultures. Third, such methods are generic and can be used to sample any objectively measureable information space (e.g. face morphology and complexion, body movements [27], vocalizations [28,29]) to test against almost any perceptual category (e.g. attractive, trustworthy [30,31], interested, confused [32], delighted, embarrassed [25••]). Such methods therefore have significant potential to advance understanding of how the human face conveys different messages because they impose fewer (subjective) restrictions on empirical investigation (see also [33•] for further discussion). Jack and colleagues [25••] used this approach to explore cultural commonalities and specificities in facial expressions of emotion by modeling the dynamic face movement patterns associated with over 60 different emotions across two cultures — Western and East Asian. Using a multivariate data reduction technique applied to the resulting culturally valid face movement models, they revealed four latent and culturally common Action Unit (AU) patterns each associated with a specific combination of valence, arousal, and dominance. Figure 1(b) summarizes the results. Color-coded face maps show the four latent face movement patterns with red indicating stronger AU presence and blue indicating weaker AU presence (see also AU labels above each face map). Emotion words below each face show a sub-sample of the face movement models that the latent pattern contributes most to (see [25••] for full list of emotion words). Plots below each face show the distribution of average ratings of valence, arousal, and dominance for each emotion word associated with each latent movement pattern. Extracting these latent patterns from the set of 60+ culturally valid face movement models also revealed the specific face movements that accentuate each latent pattern to create complex facial expressions of emotion in each culture (see also [34] for discussion on cultural accents). Together, these data question the widely held view that six facial movement patterns universally convey the six emotions of happy, surprise, fear, disgust, anger, and sad, and instead suggest that four latent patterns are common across cultures. Furthermore, the combination of culturally common face movement patterns and culture- specific accents also suggests a symbiosis (not opposition) of biology and culture, thereby generating new predictions about the bio-cultural phylogeny and ontogeny of facial expressions. The projection of latent face movement patterns onto broad dimensions (e.g. valence, arousal) with specific accents that map more closely to specific categories (e.g. rage, disgust) also suggests a specific synergy between the dimensional and categorical perception of face movements [35,36].","In addition to characterizing the specific face movement patterns that are used for social interaction in different cultures, a central and related goal is to understand their communicative aspects. That is, what messages do face movements convey to others? While psychologists have typically focused on messages that reflect the inner states of the transmitter (e.g. ‘he feels angry’), behavioral ecologists have tended to consider face movement patterns as tools to influence the receiver's behavior (e.g. ‘I should submit’) [37]. Since mouting evidence now questions the traditional psychological view that specific face movement patterns are pan-cultural transmitters of ‘basic’ emotions (e.g. [38–40]), new opportunities now emerge to explore the broader range of messages that face movements convey within and across cultures. In a recent cross-cultural study [41••] Crivelli and colleagues stepped beyond the traditional set of six emotion categories to explore the social motives that could be attributed to face movement patterns. Across two complementary experiments, Trobriand Islanders of Papua New Guinea matched the classic face movement patterns of emotion with classic emotion labels (i.e. ‘happy,’ ‘surprise,’ ‘fear,’ ‘disgust,’ ‘anger,’ and ‘sad’) and with different social motives such as ‘social invitation,’ ‘protection,’ ‘threat,’ ‘submission’ and ‘rejection.’ Contrary to the view that these face movement patterns primarily convey emotions, Trobriand Islanders matched them with emotions and social motives. In further contrast to widely held views of universality, Trobriand Islanders consistently associated the classic ‘fear’ face movement pattern — that is, knitted brows, wide-open eyes, laterally stretched mouth — with ‘anger’ and ‘threat.’ Examination of the Trobriand Islanders’ material culture [11•] and observation of their traditional rituals and social interactions [42] further corroborated these findings by showing that classic ‘fear’ face movement patterns are consistently used as threat displays in their own culture as well as others (e.g. Maori, ! Kung Bushmen, Himba, Eipo). Together, these results show that face movements convey multi-component messages including behavioral intentions rather than a fixed set of emotion categories [43].","Here, we have highlighted two recent studies that have moved beyond the boundaries of traditional approaches to make significant new discoveries on how face movement patterns convey social messages across cultures. In doing so, each study demonstrates the power and potential of interdisciplinary approaches to access the corners of knowledge that have so far been overlooked or have remained inaccessible. In particular, mature data-driven methods imported from visual psychophysics combined with state-of-the-art dynamic 3D computer graphics can now characterize face movement patterns with unprecedented detail to deliver precise information ontologies and reveal how face movement patterns differ (or are similar) across cultures. Similarly, integrating perspectives from separately evolving fields (e.g. social face perception of emotions, personality, conversational messages, e.g. [44]) or across dichotomous debates (e.g. nature versus nurture) boosts progress in understanding the functional (e.g. see [45–47]) and perceptual ontologies of face movement patterns (e.g. personality traits [30,48], intelligence [49]. See also Niedenthal in this special issue). Applications of advanced technologies, interdisciplinarity, and creative thinking now mark the emergence of a new scientific culture that holds great potential to make significant new milestones, and to raise the profile and impact of Psychology to realize its potential in other fields (e.g. computer vision, social robotics; see [50])."],["Stimulus contrast and duration effects on visual temporal integration and order judgment were examined in a unified paradigm. Stimulus onset asynchrony was governed by the duration of the first stimulus in Experiment 1, and by the interstimulus interval in Experiment 2. In Experiment 1, integration and order uncertainty increased when a low contrast stimulus followed a high contrast stimulus, but only when the second stimulus was 20 or 30 ms. At 10 ms duration of the second stimulus, integration and uncertainty decreased. Temporal order judgments at all durations of the second stimulus were better for a low contrast stimulus following a high contrast one. By contrast, in Experiment 2, a low contrast stimulus following a high contrast stimulus consistently produced higher integration rates, order uncertainty, and lower order accuracy. Contrast and duration thus interacted, breaking correspondence between integration and order perception. The results are interpreted in a tentative conceptual framework. --------------------------------------------------------------------------------","Human perceptual awareness has some paradoxical properties. We are able to detect flashes of light that last only 1 ms, but we cannot reliably estimate just how brief that is (Efron, 1967). It has also long been known that despite our apparent sensitivity, rapid sequences of brief visual stimuli can outpace the visual system relatively easily. This difficulty does not seem to rest with any particular stimulus being too brief to process perceptually, but rather with the speed at which one stimulus is followed by the next. It has long been known that in extremis, at high succession speeds, stimuli are simply perceived as simultaneous (Exner, 1875). Before that unified state is reached, two presumably related phenomena occur: Confusion arises about which stimulus came first, and also, the identities of individual stimuli may get blended to the extent that they are perceived as parts of a single composite stimulus, which comprises all the features of its multiple constituents. Evidence for the first phenomenon comes from temporal order judgment (TOJ) tasks, in which observers are presented with two almost simultaneous stimuli, and are asked to decide which of the pair came first: When the stimulus onset asynchrony (SOA) between them falls from approximately 100 to 20 ms, observers drop from near-perfect order judgments to effectively guessing (e.g., Jaśkowski & Verleger, 2000). The accuracy of order judgments is thought to depend on a central, cognitive function, rather than modality-specific factors, since TOJ tasks involving visual, auditory, and tactile stimuli, all produce similar estimates of the critical interval (Hirsh & Sherrick, 1961; Sternberg & Knoll, 1973). The second phenomenon, the temporal integration of successive stimuli, is typically found in tasks that test the observer’s ability to respond to a feature that is only apparent from the combination of two individually shown stimulus displays. One exemplary procedure is the missing element task (MET; Akyürek, Schubö, & Hommel, 2010; Hogben & Di Lollo, 1974), which presents two brief, successive displays of simple stimuli such as dots or small squares within a regularly spaced grid of 5 by 5 positions, such that the first display contains 12 of these stimuli, and the second display another 12. One position in the grid is thereby left empty, for the observer to find. Trying to mentally compare the two partial displays from memory is not a feasible strategy in this task, but the missing element is easily found if the observer is able to perceptually integrate the two displays. Such temporal integration becomes increasingly difficult as the total display duration increases, particularly beyond 100 ms. Like temporal order judgments, it appears that integration is similar across modalities, suggesting it has a central source also (Saija, Andringa, Başkent, & Akyürek, 2014). Intuitively, it seems likely that temporal integration and order judgment are closely related. If two successive stimuli are temporally integrated, and perceived as a unitary event, surely a temporal order can no longer be assigned between them. Vice versa, if the stimuli appear to occur so close together in time that their order can only be guessed at, that would suggest a degree of simultaneity that would be associated with integration. The timing of the stimuli obviously strongly governs both the ability to assign order and the tendency to integrate. However, although a strong correlation between integration and order perception may indeed exist (for a demonstration in a rapid serial visual presentation [RSVP] task, see Akyürek et al., 2012), it may not always be perfect. In the MET, observers often report a sense of having seen multiple stimuli (i.e., they detected a temporal gap), which implies a minimal awareness of some order, even if the stimuli still appeared to ‘fall in line’, and integration succeeded. When Kinnucan and Friden (1981) measured MET performance (% correct localization) with stimuli of varying brightness, and subsequently asked observers to rate the degree to which the successive MET displays appeared as one, the outcomes differed as a function of their brightness manipulation. The authors suggested that actual temporal integration and the subjective appearance of unity may rely on different mechanisms, and the latter on discontinuity detection in particular—a function that is presumably central also to temporal order judgments. Further hints for a possible dissociation between integration and TOJ may be found in studies of stimulus intensity effects in integration tasks. Many studies report inverse intensity effects, that is, integration is found to be enhanced by less intense stimuli (e.g., Bowling & Lovegrove, 1981; Di Lollo & Bischof, 1995). A similar phenomenon of “inverse effectiveness” has also been found in multisensory integration tasks, where less salient stimuli are more easily integrated between modalities (Meredith & Stein, 1983). In TOJ tasks it has also been found that higher stimulus intensity facilitates temporal separation (Jaśkowski & Verleger, 2000), which reflects the same dynamics. Yet, there have been reports of an opposite relationship as well, in which higher stimulus intensity impedes separation (e.g., Ueno, 1983; Wilson, 1983). A complicating factor in the interpretation of many of these collective studies is the possible role of retinal afterimages, elicited by using bright stimuli on a dark background. Nonetheless, one account for the discrepant effectiveness that has been offered is that task characteristics vary, namely whether observers are (implicitly) asked to judge stimulus offset or total duration (Nisly & Wasserman, 1989; Wasserman & Nisly-Nagele, 2001, although see also Di Lollo & Bischof, 1995). It is conceivable that such characteristics may similarly underlie possible differences between integration (in the MET) and TOJ tasks. In line with this notion, Jaśkowski (1996) has argued that because stimulus intensity effects found on visible persistence are not necessarily mirrored in TOJ performance, TOJ may not strongly rely on perceived duration. When the strength of the first and second display is independently manipulated, such that a brighter stimulus follows a dimmer stimulus, or vice versa, the outcomes are even less uniform. In a TOJ task, Bachmann, Põder, and Luiga (2004) found that when the relative contrast of a pair of stimuli was manipulated, observers tended to report the stimulus with the lowest contrast as the first. Performance was thus best when the first stimulus was dim, which does not point to inverse effectiveness. Inverse effectiveness would predict higher perceptual latency and more integration with the second stimulus, and thus lower TOJ performance. In line with these findings, however, are reports by Kinnucan and Friden (1981; Experiment 1) and Johnson, Nozama, and Bourassa (1998), who found increased integration when the first stimulus was stronger than the second in a MET paradigm, again suggesting direct rather than inverse effectiveness. Johnson and colleagues nonetheless also showed that inverse effectiveness was again obtained when the two displays were matched in luminance. In a similar vein, Long and O’Saben (1989) also observed inconsistent intensity effects on integration. Finally, it is conceivable that temporal integration and temporal order judgments are differentially affected by the deployment of exogenous (stimulus-driven) and endogenous (volitional) modes of attention.1 Exogenous attention is engaged by stimulus- related manipulations, such as intensity, while endogenous attention responds to (learned) contingencies, such as predictable stimulus timing. Lawrence and Klein (2013) recently demonstrated that exogenous and endogenous factors can have different, dissociable effects on performance (reaction time and accuracy) in temporal attention tasks. Exogenous and endogenous factors are typically not explicitly controlled for in temporal integration and order tasks, but it is conceivable that they are differentially involved in these two tasks, which might lead to different response profiles. Summarizing, even though conceptually temporal integration and order judgment would appear to be two sides of the same coin in perceptual awareness, the collective body of studies on these phenomena shows relatively little consistency. The relationship between temporal integration and order judgment thereby remains underspecified. A closer examination of the correspondence between these measures of rapid visual perception seems called for, and to do so was the aim of the present study. To thus investigate whether temporal integration and order judgments are similarly affected by relative stimulus strength, the present study measured both integration performance and the accuracy, as well as the associated uncertainty of order judgments by means of a single, uniform task, in which stimulus contrast and duration were varied systematically. The inclusion of a measure of uncertainty in the TOJ task was motivated by a previous study by Ulrich (1987), who demonstrated that perceptual moment and triggered moment models do not account well for TOJ performance in a classic ternary task, in which the third response option is that of indicating simultaneity. Since the idea of a perceptual moment, whether it is externally triggered or not, is conceptually close to an interval during which temporal integration takes place, this may be taken as evidence for a dissociation between integration and order judgments. However, indicating simultaneity corresponds to a rather specific percept, while it is conceivable that within a particular range of SOA close to actual simultaneity, the perception of simultaneity is not elicited (e.g., because flicker is detected), but order still remains ambiguous. In other words, in this range there may be an interval during which a gap is detected, that is, some implicit order is recognized, but it may yet be impossible to determine what the order actually was. Temporal integration might still occur in this SOA range. To examine whether such impressions are indeed experienced, and whether these might correlate with integration frequency, observers were presently given the option to indicate uncertainty with regard to order. Finally, in the present study stimulus strength was not manipulated directly as a function of brightness or luminance, but by means of relative contrast, which entailed that when the first stimulus was high contrast, the second stimulus was low contrast, and vice versa. Individual stimulus contrast was furthermore defined such that high contrast corresponded to lower stimulus luminance (and vice versa), which made the stimulus contrast more strongly with the white background, thereby removing possible low-level confounds related to stimulus intensity, such as retinal afterimages, which can vary for both visual latency and persistence measures (Allik & Kreegipuu, 1998; Coltheart, 1980).","Experiment 1 examined the effect of stimulus contrast on temporal integration performance and the accuracy of temporal order judgments under task conditions commonly used in MET paradigms. To this end, an MET based on an existing paradigm (Akyürek et al., 2010), was designed to include two different stimulus onset asynchronies (SOAs) that were determined by the duration of the first stimulus display (cf. Di Lollo, 1977, 1980). These were furthermore crossed with two different contrast conditions. Either the first or the second stimulus display had higher contrast, while the other display had lower contrast. Finally, the duration of the second stimulus display was varied as well. Next to accuracy measures in both tasks, the frequency of uncertain responses in the ternary temporal order judgment task was also measured.","Twenty-one Psychology students (19 female) at the University of Groningen participated in the experiment. They could earn a small monetary compensation in exchange for good task performance (detailed below), and were informed about this opportunity beforehand. The study was conducted in accordance with the Declaration of Helsinki, and approved by the departmental ethical committee prior to its execution. All participants reported normal or corrected-to-normal visual acuity and gave written informed consent. The data of two female participants were excluded, because their overall task performance did not meet a minimum level of 10% correct in either task. In the final sample, mean age was 20 years (range 18–25 years).","Participants were seated individually in sound-dampened testing cabins, at a viewing distance of approximately 60 cm (not fixed) to the screen. Cabin lighting was dimmed. Stimuli were shown on a 22″ CRT screen, refreshing at 100 Hz, using a display resolution of 800 by 600 pixels and a color depth of 16 bit. The screen was driven by a standard Windows XP personal computer. The experiment was programmed in E-Prime 2.0 Professional, and logged responses that were entered by means of a standard PS/2 keyboard and mouse. As shown in Fig. 1, a white background (157 cd/m2) was maintained throughout the experiment. Stimuli consisted of colored squares of 10 by 10 pixels, centered in an invisible 20 by 20 field. These fields were arranged in a grid of 5 by 5 positions (25 total), which itself was centered on the screen. On each trial, one of these positions remained empty. The others were filled such that on each of the two successive stimulus displays, 12 squares appeared in the one color, with the squares in the other display having the second color. The order of these colors was randomly drawn, but equally distributed. The same logic was applied to the contrast of the squares in either display, which was either low or high. Thus, a high contrast stimulus in the one color would be followed by a low contrast stimulus in the other color, or vice versa. The four possible colors were high contrast red (RGB 213, 0, 0; 23 cd/m2), low contrast red (RGB 255, 159, 159; 71 cd/m2), high contrast blue (RGB 0, 0, 213; 12 cd/m2), and low contrast blue (RGB 170, 170, 255; 64 cd/m2). On the response screen in the temporal order judgment task, a brief text in black 18 point bold Courier New font prompted participants to enter which color had come first, using the z or c key for either color, or the x key if they were unsure. On the response screen in the temporal integration task, a full grid of black outlined squares appeared. Participants could use the mouse to click where no colored square had previously appeared. Procedure There were 1152 experimental trials in the experiment, divided across the two tasks in two successive blocks of 576 trials each. The order of blocks was counterbalanced between participants, and each block was preceded by 24 practice trials that were discarded prior to analysis. Each task block was further subdivided in 6 trial blocks, after which a summary of performance was given, and at which point participant could take a short break. Participants could initiate a new block by pressing the right mouse button, but within each trial block the trials continued without interruption. Good performance was rewarded such that a correct answer in the integration task yielded 10 points, while an incorrect or missing answer subtracted 5. In the temporal order judgment task a correct answer yielded 5 points, an incorrect answer cost 10 points, and participants had a third option: By indicating that they were unsure of the temporal order, a loss of points could be avoided (but nothing could be gained either). After the experiment, the count was settled such that €1 was paid per 500 points earned. Each trial started with a blank screen that lasted 600–1200 ms (600 + a ∗ 30 ms, where a was randomly varied between 0 and 20). The first stimulus display (S1) followed, with a duration of 30 or 70 ms, depending on the experimental condition. After a brief interstimulus interval of 10 ms, the second stimulus display (S2) followed in turn, lasting 10, 20, or 30 ms, again dependent on the condition. The response screen then appeared after a blank interval of 600 ms, and terminated upon user input, or when 1200 ms had passed by. The design featured three variables that were analyzed by means of repeated measures analysis of variance (ANOVA). The first variable was Contrast, with two levels (S1 high & S2 low or S1 low & S2 high). The second two-level variable was SOA (40 or 80 ms; S1 duration + 10 ms fixed ISI duration). The third variable was S2 duration, which had three levels (10, 20, or 30 ms). The full design thus comprised 12 cells (2 × 2 × 3). Analyses were performed separately for integration and order judgment tasks, and frequencies computed relative to all trials in the respective tasks. When a significant test of sphericity occurred, degrees of freedom were adjusted with the Greenhouse-Geisser epsilon correction.","In line with expectations, temporal integration frequency as well order judgments showed some clear common effects. Among these were the straightforward changes brought about by SOA: Shorter SOA increased integration and uncertainty with regard to temporal order, and decreased the accuracy of order judgments. S2 duration had comparable effect on temporal integration and temporal order uncertainty, although the effect on the former was mostly limited to the condition in which S2 was 30 ms. Furthermore, there was no overall S2 duration effect on temporal order judgment accuracy. Stimulus contrast, however, produced some unexpected outcomes, even though the observed trends seemed straightforward initially. Overall, there was a trend indicating that a high contrast S1 followed by a low contrast S2 increased integration, and there was reliable evidence for a decrease in the number of uncertain responses. Yet, contrary to what might be expected, high contrast at S1 was also associated with a trend towards increased temporal order accuracy. Stimulus contrast furthermore proved to depend on the duration of the shorter stimulus (i.e., S2) quite markedly. At 10 ms S2 duration, contrary to the overall trend in the present data described above, integration frequency and TOJ uncertainty were lower when S1 was high contrast, compared to when it was low contrast, while temporal order judgment accuracy was higher. However, at 20 and 30 ms S2 duration, both integration and temporal order accuracy were higher when S1 was high contrast; contrary to the reciprocal relationship between these measures that was observed at 10 ms S2 duration. Furthermore, temporal order uncertainty no longer clearly differed. Taken together, the results of Experiment 1 indicated that temporal integration and temporal order judgments can vary in different ways, depending on the interplay between stimulus contrast and duration. These findings support the notion that even when the stimulus material is identical, the integration and order judgment task may not tap fully identical cognitive processes, at least for the presentation conditions currently tested. It is similarly conceivable that different modes of attention (endogenous or exogenous) were being engaged. However, the discrepancy between tasks seemed to rest primarily with the accuracy of order judgments. The frequency of uncertain responses in the TOJ task showed a (mirrored) pattern that was quite similar to that of correct integrations, showing a correlation between the means of r = 0.875. This suggests that uncertainty with regard to order may thus rely more on sensory signals that also enable temporal integration, while the judgment that is otherwise made may include other, more top-down driven aspects.","Although having a variable and relatively long duration of the first stimulus display is common in temporal integration tasks, it is conceivable that the manipulations in Experiment 1 were specifically driven by the inequality in stimulus strength between the first and second stimulus display. The relatively long duration of the first stimulus presumably made it much more prominent than the second stimulus, if only because it is known that for near-threshold stimuli, perceived brightness and duration are related (Bloch, 1885). To investigate whether the relative strength of the first stimulus in Experiment 1 might have played a role in the outcomes, in Experiment 2 the duration of that stimulus was reduced so that it always matched the second stimulus, thereby equalizing their comparative strength.","Twenty-five new participants (21 female) took part in the experiment. Following the same criterion as in Experiment 1, the data of 6 participants (5 female) were excluded. Mean age was 20.3 years (range 18–28 years) in the final sample.","The experimental setup was identical to that of Experiment 1. Procedure and design The procedure and design were similarly unchanged, with the exception of S1 duration. Rather than 30 or 70 ms, S1 duration was either 10, 20, or 30 ms, equal to the duration of S2 on the same trial. SOA remained the variable of interest, and was preserved at either 40 or 80 ms, being equal to S1 duration + ISI. Thus, when S1 was 10 ms, the ISI was either 30 or 70 ms, when S1 was 20 ms, the ISI was either 20 or 60 ms, and when S1 was 30 ms, the ISI was either 10 or 50 ms.","The results of Experiment 2 were straightforward: Integration was facilitated by a high contrast S1 and a low contrast S2, while temporal order judgments were less accurate in this condition. Uncertainty with regard to order again followed the opposite pattern; with more uncertainty resulting from the high contrast S2 condition. There was evidence from all three response measures that contrast effects were more pronounced when SOA was short. Since the temporal intrusion is obviously higher at short SOA, these effects confirm that contrast does not generically affect the perception of stimuli, but specifically affects the perceptual process of temporal integration and/or separation. There were furthermore several marginal trends suggesting that longer S2 duration might also enhance contrast effects, to a more limited degree. These trends might have been observed since S2 was not followed by a mask, providing more opportunity for S2 (and its contrast) to leave an impression.","Temporal integration rate and order uncertainty varied similarly as a function of stimulus contrast and duration in both experiments, suggesting that these perceptual states may be similarly driven by sensory information. Temporal integration rate and order judgment accuracy also showed similar patterns in several of the presently tested conditions, yet there were also notable exceptions, in which wholly opposite effects were observed. In the following, a tentative account for the findings will be presented. To start with the most straightforward outcome: The results of Experiment 2 consistently showed that a high contrast S1 followed by a low contrast S2 resulted in more temporal confusion. Integration as well as uncertain temporal order responses were higher, while order accuracy was lower, in comparison with the reversed contrast. These results follow the same trend that has been observed for stimulus intensity in integration tasks (e.g., Eriksen & Collins, 1968; Kinnucan & Friden, 1981), and fit with previous observations that a low contrast stimulus tends to be perceived as having occurred earlier in time (Bachmann et al., 2004). This outcome seems compatible with the idea that low stimulus contrast evokes less brain activity than high contrast (e.g., Sclar & Freeman, 1982), based on which a conceptual model of temporal perception can be specified. In Panel A of Fig. 4, using as few additional assumptions as possible, resultant neural activity distributions are visualized as a function of time for each of the contrast conditions. Neural activity is plotted in arbitrary units, as the model is principally neutral with regard to the nature of such activity (e.g., firing rate or phase locking). Activity in the model is thought to have an exponential property, such that activity accelerates towards peak activity levels, but also drops back to baseline (the horizontal axis) more readily. This assumption of non- linear scaling is motivated by the idea that it would help efficient representation (Baddeley et al., 1997), but is not essential for the model to function. Although it is not essential for the model either, it is assumed that a certain degree of activity is needed for the brain to become perceptually aware of a stimulus; this is visualized by means of a threshold level (t). The idea of a (dynamic) threshold is shared with the influential Global Neuronal Workspace model developed by Dehaene and Changeux (2011) and Dehaene, Sergent, and Changeux (2003). In this model, once activity passed the threshold, the global workspace is “ignited” and self-amplifying recurrent activity occurs, constituting conscious awareness. The current model is nevertheless principally agnostic with regard to the question of whether awareness should involve wide-spread (non-specific) recurrent activity across brain regions, or whether more local activity would suffice (cf. Bachmann, 2007; Dehaene, Changeux, Naccache, Sackur, & Sergent, 2006; Lamme, 2006). It is furthermore reasonable to assume some time is needed before the threshold is reached, which is typically estimated to be around 100 ms (e.g., Wu, Busch, Fabre-Thorpe, & VanRullen, 2009). Critically, the model couples a degree of persistence with the contrast- induced difference in the magnitude of activity. The idea that the neural signal lingers after stimulus offset, and that it has perceptual consequences, is supported by classic behavioral experiments and remains uncontested (e.g., Hogben & Di Lollo, 1974; Sperling, 1960). Persistence is modeled by having neural activity subside more slowly than it rises. This produces the dynamics visualized in Panel A of Fig. 4 (by the solid lines; the dotted lines reflect an additional assumption further detailed below). When a high contrast stimulus is followed by a low contrast stimulus (solid lines, top plot), it differs in three ways from when a low contrast S1 is followed by a high contrast S2 (solid lines, bottom plot): The activity peaks are closer together, the interval in-between during which neither stimulus is above threshold is shorter, and there is more overlap between the activity distributions of the stimulus pair. These differences all point towards the same perceptual outcome, namely that the perceived temporal separation between the stimuli is lower when a high contrast S1 is followed by a low contrast S2, as expressed in increased integration rates, increased TOJ uncertainty, and reduced TOJ accuracy. It may be noted here that previously advanced formal models of persistence would presumably generate similar predictions (Dixon & Di Lollo, 1994; Loftus & Irwin, 1998). The present model not only shows which assumptions are necessary to produce the observed behavior, but also highlights at least one other candidate assumption that would actually be counterproductive, namely that brain activity associated with low contrast not only produces lower peak amplitude, but also takes longer to build than for high contrast stimuli. Panel B of Fig. 4 shows the resultant distributions. Including this assumption would appear to be justified on the basis of previous findings. For instance, Alpern (1954; see also Roufs, 1963) demonstrated that observers had more difficulty judging temporal order with lower stimulus intensity, and Kelly (1961) showed that flicker sensitivity increased when luminance was higher, suggesting that the visual system ‘speeds up’. A similar result was obtained for stimulus contrast in both flicker and motion tasks (Stromeyer & Martini, 2003). However, when this assumption is added to the model, it only counteracts the observed behavior. As shown in Panel B of Fig. 4 (solid lines), when a high contrast S1 is followed by a low contrast S2, S2 would have reached threshold relatively late, while the high contrast S1 would not have been delayed (top plot). In this case, maximal temporal separation should have been observed, resulting in more accurate and less uncertain TOJ, and less integration, and the reverse when a low contrast S1 is followed by a high contrast S2 (bottom plot). Neither of these predictions were confirmed by the present data, suggesting that task performance did not depend on such latency differences. Thus, these differences are best omitted from the model if it is to account for temporal integration and order judgment as a function of relative stimulus contrast. Relationship to brain electrophysiology ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There is some prior evidence from event-related potential (ERP) studies of temporal integration that supports the model dynamics advocated here. Akyürek et al. (2010; Akyürek & Meijerink, 2012) measured component amplitude during both early (N1, N2) and late (P3) phases of the ERP during temporal integration in a typical MET, similar to the current paradigm. In their task, SOA was determined mainly by S1 duration, as in Experiment 1 of the present paper. The authors found that increased component amplitude across N1, N2, and P3 was associated with successful integration of the two successive displays. By contrast, there was no indication that the latency of any of these components changed systematically. If component amplitude can be related to representation strength (assumption 1), this pattern of results would match the predictions of the model depicted in Fig. 4, provided that the ERP was driven mainly by the (onset of the) first stimulus (assumption 2). Some evidence for both assumptions may be found in the Appendix Figure published by Akyürek and Meijerink (2012), the relevant panel of which is reproduced here as Fig. 5. Here, the ERPs of different S1 duration conditions of their MET were overlaid (40, 70, and 100 ms, with 10 ms ISI). The figure firstly shows that increased component amplitude was elicited by the longer-lasting (stronger) S1s, at least as far as the earlier components are concerned. Secondly, it is apparent from this figure also that there was no consistent time-shift in the ERP, despite the varying SOA between the contrasted conditions. These findings held not only when the participants were doing an actual integration task, but also for another experimental condition in which no temporal integration, but only singleton detection was required, with identical stimulus timing (not shown). At a conceptual level, the ERP observed in this MET thus aligns well with the neural activation dynamics predicted by the current model. Similar direct ERP evidence for TOJ tasks under conditions comparable to those in the MET paradigm is not yet available. Such evidence would provide a further validation check of the assumptions underlying the present model. Eventually, a direct test of the effects of stimulus duration and relative contrast on the ERP in these tasks will be essential to further support theorizing on the underlying neural dynamics. Further research is thus clearly needed here. Some relevant results are nevertheless available from a previous ERP study of a different TOJ task in which attention was also modulated (McDonald, Teder-Sälejärvi, Di Russo, & Hillyard, 2005). An amplitude change, spanning the P1, N1, and P2 components, was also implicated in this study. Stimulus and task differences may affect the components that are modulated, but the activity dynamics were in line with the current model, even though other authors have additionally observed component latency shifts in a bimodal TOJ task (Vibell, Klinge, Zampini, Spence, & Nobre, 2007). Differences between tasks ~~~~~~~~~~~~~~~~~~~~~~~~~ Can the model depicted in Panel A of Fig. 4 be further modified to account for the deviant results obtained in Experiment 1? In this experiment, SOA was directly determined by S1 duration, as is common in temporal integration tasks (e.g., Akyürek et al., 2010; Di Lollo, 1977, 1980). Even though the present study used negative contrast (i.e., relatively dark stimuli on a white background), it seems likely that the impression made by the stimuli increases in strength as duration increases (cf. Bloch’s Law for perceived brightness; Bloch, 1885). Previous research furthermore suggests that the relatively long duration of S1 may have made its offset in particular more noticeable (Wilson, 1983). Two notable observations resulted in the present study. First, at 10 ms S2 duration, a pattern of performance was observed that was opposite to that found in all conditions of Experiment 2. A low contrast S1 followed by a high contrast S2 facilitated integration, decreased the accuracy of order judgments, and increased TOJ uncertainty, compared to the reverse contrast condition. Second, the accuracy of order judgments at 20 and 30 ms S2 duration remained higher for a high contrast S1 followed by a low contrast S2 than for the reverse contrast condition, even though integration performance reverted to match Experiment 2 at those S2 durations. In the following, possible explanations for these discrepant results will be offered, interpreting the results within the framework of the model. Primarily, to account for the deviance in temporal integration frequency and order judgments in the 10 ms S2 condition of Experiment 1, an additional assumption is required to alter the model dynamics. This assumption is that stimulus contrast and duration interact in some cases, such that a particularly weak stimulus not only attains lower peak activity, but also takes longer to get there. This may (only) happen in the most extreme condition currently tested in Experiment 1, when a brief (10 ms), low contrast S2 is preceded by a strong, and relatively long-lasting S1. The hypothesized role for S1 here is justified because the combination of its contrast and duration is likely to intensify forward masking of S2 (cf. Kirschfeld & Kammer, 2000). It may also cause S1 to be perceived as if it occurred later, due to its delayed offset (Jaśkowski, 1991), increasing the ambiguity of the stimulus sequence. In line with the above, it must also be assumed that when S1 has low contrast, and S2 has high contrast, the hypothesized slowing of S2-related activity does not, or only to a much lesser extent, occur. The resultant distributions are visualized with the dotted lines in Panel A of Fig. 4. Because the low contrast S2 (top plot) suffers more from short duration than the high contrast S2 (bottom plot), it gets delayed to the point where the temporal separation perceived between the stimuli is greater in the former case than in the latter. This would cause the observed reversal of performance seen at 10 ms S2 duration in Experiment 1. Note that applying the same transformations on the distributions in Panel B, which assume task performance should reflect a ‘main effect’ of contrast on activity rise time, does not as easily produce a reversal in outcomes, again suggesting that this assumption is less likely to be correct. Here, even though the slowing of a low contrast S2 might be less (because of the already- present slowing for low contrast stimuli), this would only reinforce the temporal distinctiveness advantage of the condition in which a high contrast S1 precedes it (top plot). The comparatively small delay that might occur for a high contrast S2 would be hard-pressed to overcome the temporal proximity caused by the slow rise of the low contrast S1 that precedes it (bottom plot). With regard to the observed facilitation of temporal order judgment accuracy in all S2 duration conditions when it followed a high contrast S1 in Experiment 1, the most parsimonious account appears to assume that when confronted with a relatively long S1 the observers relied on a different signal to judge order. Since the activity associated with S1 is expected to be high overall, the total area of its activity distribution above the detection threshold may also become quite large. This may demarcate the stimulus from the weaker S2 to such an extent that it is consequently seen (correctly) as having come first. Such an effect is reminiscent of prior entry (Spence & Parise, 2010; Titchener, 1908), an attention-related latency illusion in which stronger stimuli are seen as having occurred earlier in time (though see Jaśkowski, 1996, for an opposing view on prior entry in TOJ). It seems conceivable that while order judgments may normally rely on perceived temporal separation, it may be overruled by the strong discrepancy in stimulus (offset) clarity, as elicited by the relatively long duration of S1 in Experiment 1. The decision level that is needed to judge order, beyond a state of mere uncertainty, may provide the opportunity to take such additional evidence into account. This idea fits with previous evidence against an account of TOJ performance on the basis of perceptual latency (Ulrich, 1987), and a similar proposal that was fielded by McDonald et al. (2005), who studied attentional biases in temporal order judgments. The authors observed attentional amplitude enhancement over visual cortex, in the absence of latency change, suggesting that (attentional modulation of) TOJ in their task relied on signal strength. A similar reliance may have occurred in the present study. It has indeed been suggested previously that attentional involvement in temporal order judgment may primarily affect decision mechanisms (Aghdaee, Battelli, & Assad, 2014). Thus, if the relatively long duration of S1 caused differential involvement of exogenous and endogenous modes of attention (cf. Lawrence & Klein, 2013), that might be expected to result in a change at the decision level, thereby causing TOJ accuracy to deviate from the other measures. In view of previous evidence implicating parietal cortex in perceiving stimulus on-/offset and temporal order (Battelli, Cavanagh, Martini, & Barton, 2003; Davis, Christie, & Rorden, 2009), it might be a fruitful avenue for further neurophysiological research to assess the relative involvement of this area as a function of stimulus contrast and duration in both order and integration tasks, which might provide converging evidence for differential functional involvement. Taken together, not one single mechanism based on either perceived stimulus strength or latency seems able to account for all aspects of the present results. Behavior in trials in which a TOJ decision was made in particular seems to involve processing beyond what underlies both temporal integration and the experience of order uncertainty. A similar conclusion was reached by Di Lollo, von Mühlenen, Enns, and Bridgeman (2004), based on an examination of metacontrast masking as a function of target-mask (∼S1-S2) SOA and mask duration. These authors obtained evidence that different mechanisms might underlie the effects on task performance caused by these two factors. Although the tasks in the present study were different, and although the duration of S1, rather than S2, seemed the critical variable in these tasks, the outcomes do seem to converge on the idea that not all aspects of the perception of brief, successive stimuli follow uniform rules. As previously alluded to, the current results suggest that this non-uniformity might be attributed to the need for more (conscious) evidence weighing to reach a perceptual decision (beyond being uncertain about order) in the TOJ task than in the integration task, which might imply reliance on recurrent pathways in the brain, and possibly increased top-down control in the former task. This could be accounted for in models of perceptual awareness through interactions between stimulus-specific and top-down activations, as previously proposed by Bachmann (1997), which might specifically trigger deviations in some cases, such as when a relatively long S1 is presented.","Although temporal integration and temporal order judgments were presently found to change consistently with relative stimulus contrast in most conditions of the present study, this was clearly less the case when the duration of the first stimulus was relatively long. Thus it must be concluded that when it comes to perceiving simultaneity and episodic unity, relative stimulus contrast interacts with stimulus duration, such that at a given contrast level facilitation as well impedance can result. The present results also underscore that in the perception of rapid successive stimuli, seeing integrated percepts and temporal order do not always perfectly align, resulting in a ‘two-sided’ perceptual experience: It seems that sometimes perceived stimulus strength can be used, possibly post hoc, to infer order between a pair of stimuli that make an integrated impression."],["This article summarises part of the findings of our research-creation Dizziness––A Resource (2014-17) funded by the Austrian Science Fund and hosted at the Academy of Fine Arts Vienna. We will introduce the two underlying key concepts, related practice examples, and tell the story of experimental filmmaker Oskar Fischinger's wax-slicing machine. Furthermore, we will give insight into the philosophical background of the methodology that we have developed throughout the research process. This article underlines the creative potential of dizziness by introducing new ways it may be conceptualised, thinking with and through the frames of emotions and space. Dizziness is understood as a phenomenon of embodied knowledge (Manning, 2009; Varto, 2013), blurring categorisation between the perception and conception of dizziness. Taumel, German for ‘dizziness’, implies a broader semantic field including medical indications of vertigo and further notions of physical and emotional disequilibrium, exhilaration, confusion, uncertainty, disorientation, and turmoil. Taumel therefore includes positive, negative, and ambiguous connotations. Mirroring the research trajectory, this jointly written text oscillates between cross- disciplinary conceptualisation and practice, approaching dizziness in a metaphorical sense as object and method. Research-creation as the combination of arts-based research and research-based art (Loveless, 2015; Manning, 2008) allows for manifold approaches, heterogeneous formats, diverse outcomes, and contradicting methods as modes of ‘curiosity, sustained questioning, and analysis’ (Green, 2012: 272). As in dizziness, we believe that the strength of research-creation resides in its ambiguous, wide-stretched and diversity- affirming nature. The first key concept we introduce in this paper relates to dizziness as a possible resource for creativity. The second explores ‘compossibility’ as an actual and theoretical space for the experience of, and reflection on, dizziness. Both concepts are mutually dependent if they are to bring about their creative potential. Dizziness––A Resource started from the assumption that feelings of dizziness – of being lost or disoriented (Solnit, 2005; Ladewig, 2016) – are not only a part of the artistic and philosophical, but also of any creative process (Anderwald et al., 2013; Anderwald and Grond, 2015; Feist, 1998; Feyertag, 2015; Jullien, 2012; 2015; Montuori, 1994). Consequently, we inquired when, where, and how dizziness arises and how the experience of, and reflection on, dizziness and its conceptualisation can lead to a better understanding of inherent creative processes. Relating to the first key concept, our proposition is that the notion of dizziness could provide a critique of simplistic views on creative processes that describe them as logical circuits, e.g. ‘break in, break down, break through’ (Anderwald and Grond, 2015; Deleuze, 2003; Katzmair, 2015; Marks, 1998). Moreover, dizziness as a ‘concept in motion’ needs a mode of thinking infused by movement, not relying on fixed points but on moving relations and shifting anchor points (Strong, 2004). Our critique underlines the importance of the creation of ‘compossible spaces’ by the interaction and collaboration of persons who feel affected by dizziness in different or even conflicting ways, experiencing it as fear and/or pleasure. Within these ‘compossible spaces’ dizziness as movement becomes a resource, shaking and swiping away long- established oppositions and making room for the unfolding of seemingly contradictory feelings, processes, theories, matters, and disciplines.","Dizziness arises locally and combines various elements: theory and emotion, momentum and disorientation, time and space. It can clear, cause a great stir, or move heaven and earth – it destabilises. According to Plato, dizziness constitutes all philosophical thought by destabilising the basis of knowledge to a state of uncertainty: as an ontological state it can prompt transformation and innovation (Plato, Timaeus: 49e; Echterhölter et al., 2010). The dizzy individual experiences an emotional rollercoaster ride involving feelings of exhilaration, anxiety, and disorientation. Exposure to dizziness increasingly reduces predictability and our ability to exercise control. From a medical perspective, dizziness is considered a symptom, not a sign. Similar to vertigo, it can only be described by the experiencing subject and cannot be measured objectively. As a medical symptom, dizziness is ambiguous and can lead to a multitude of diagnoses. The vestibular system is the sensory apparatus that signals the coordinates of our spatial position to the brain, affording our sense of balance and spatial orientation. Together with the cochlea, it constitutes the labyrinth of the inner ear. As our movements consist of rotations and translations, the vestibular system comprises two components: the semi-circular canal system that indicates rotational changes in velocity and the otoliths that indicate linear changes in velocity. The vestibular system sends signals primarily to the neural structures that control our eye movements and to the muscles that keep us balanced. Psychobiologist Matti Mintz's research suggests a connection between our ability to maintain emotional and corporeal equilibrium and flexibility. His research particularly indicates comorbidity between anxiety disorders and a poor sense of orientation and balance (Mintz, 2016; Erez et al., 2002). Working on his film Failed States (2008) filmmaker Henry Hills literally used dizziness as his resource. Spinning around and becoming dizzy with his camera in hand enabled him to overcome a severe crisis in his work. Not only had the vertiginous perception of the world matched his uncertainties about his work, but his becoming dizzy also stimulated new sensations and brought back a childhood memory – spinning and falling into the grass while watching the world turning around him. Speaking of filmmaking, dizziness in its corporeal sense can be transmitted by unstable or rotating camera movements of which abundant examples exist in art films such as Stan Brakhage's Scenes Before Under Childhood (1967-70), Michael Snow's La Région Centrale (1971), Steve McQueen's Static (2009), or Catherine Yass' Lighthouse (2011). Moreover, dizziness in its metaphorical sense has been employed in cinema from its very beginning, as in the Lumière's gravity-defying first special effect in film (La démolition d'un mur, Auguste and Louis Lumière, 1896) or Georges Méliès’ early film Un homme de tête (1898) and was later elaborated in surrealist films such as Teinosuke Kinugasa's A Page of Madness (1926) or Hans Richter's collaborative film Dreams That Money Can Buy (1947). Furthermore, dizziness can be produced by abundant visual input such as the flickering of light, as used in Tony Conrad's film Flicker (1965) or Joachim Koester's This Frontier is an Endless Wall of Points (after the mescaline drawings of Henri Michaux) (2007). But dizziness can also be produced by a deprivation of visual stimuli, as seen in the ‘prisoner's cinema’ phenomenon reflected in, for instance, Melvin Moti's eponymous video work (2008). When a person – a prisoner for example – is subjected to prolonged visual deprivation, hallucinations in the form of colours and shapes might occur (Sacks, 2012). Therefore, dizziness indicates a situation in which the possibilities of reality can no longer be grasped in a habitual manner because of a lack or overload of stimuli, knowledge, or input. Whether frightening or enjoyable, by falling into dizziness we enter a stage of uncertainty, disorientation, and heightened vulnerability where we are unsure of our abilities, perceptions, and processing – uncertain of ourselves (Butler et al., 2016). This manifests through feelings of excitement caused by a distorted perception of time and space, loss of proportion, and an increasing feeling of lack of control and/or temporary loss of memory and self (Katzmair, 2015; Montuori, 1994). Dizziness can affect us as an individual, group, or society (Koller, 2014a,b) and the ensuing insecurity affects interaction with our environment (Lorey, 2015). To different degrees, these conditions of loss and feelings of insecurity are present in all dizziness processes, from crisis to flow experiences (Csikszentmihalyi, 1996), aporia to ecstasy, immersion in a film or book to philosophical pondering, or the creation of an artwork (Montuori, 1994). Moreover, the emotional spectrum of dizziness must be considered in order to comprehend its potential as a resource. The experience of dizziness contains ambiguous and even contradictory feelings. This inherent unpredictability makes clear why dizziness cannot be seen as a means of ‘self-design’ (Groys, 2008). In its reflection, dizziness exposes related emotions as movements propelling the individual into a certain direction or perspective. For the aforementioned filmmaker Henry Hills, memories of being dizzy generated a positive reminiscence of childhood, which helped him come to terms with a creative crisis. However, not all recollections of dizziness necessarily need to be positive to have a constructive effect on navigation through dizziness. Therefore, the combination of the physical, emotional, and metaphorical experience of dizziness with the more reflexive and theoretical framework of compossibility proved essential for this research if dizziness was to be seen as a resource for creativity. Dizziness represents the limit state of the challenged subject experiencing the vacillation between loss of control (staggering) and gain of control (equilibrium) (Echterhölter et al., 2010). In its metaphorical sense, dizziness starts with teetering and staggering at the limitations of knowledge, for instance when faced with a central problem or crisis (Alon, 2014) or aimed towards the creation of a new work of art (Anderwald et al., 2013). The compossibility of precipitancy and precision is what German philosopher Marcus Steinweg suggests for a situation involving dizziness in artistic or philosophical practice (2013). He further indicates a connection between the processes of thinking and art creation, both grounded in the groundless and the abyssal, starting from inherently aporetic moments (Kofman, 1988) and aiming at the impossible, in contrast to the self-reduction to the possible exemplified by politics (Steinweg, 2013). As Steinweg quotes Heiner Müller: ‘Something new can only develop when you are doing something you cannot do […] Art is what you want to do, not what you can do’ (2013: 48). Describing this movement as headless or blind, Steinweg uses the practice of writing as an example in which the author develops a distance from the universe of facts without ever fully detaching from it. This striving to ‘develop something new’ by ‘doing what you cannot do’ is further reflected in the following story about seminal filmmaker Oskar Fischinger that influenced our methodological approach.","Between 1918 and 1921, filmmaker Oskar Fischinger – literature aficionado and at the time an apprentice – prepared a lecture for his literary club in Frankfurt. He set out to analyse and compare the structure of two theatre plays: Fritz von Unruh's Platz and William Shakespeare's Twelfth Night. To supplement his talk, he drew graphic charts that illustrated the plays' dramatic developments as lines that collide, swirl, and break as the action unfolds. Trying to express his findings, he remembered: ‘In preparing this speech I began to analyse the works in a graphic way. [ …] On large sheets of drawing paper, along a horizontal line, I put down all the feelings and happenings, scene after scene, in graphic lines and curves […] that showed the dramatic development of the whole work and the emotional moods very clearly’ (Fischinger, 2006: 110). However, this graphic exposition of the dramas' content seemed to baffle his audience. Fischinger understood that he needed to add the element of temporal movement in order to express his thoughts more clearly. After being introduced to Walter Ruttmann's work by a fellow member of his literature club, he felt encouraged to explore the possibilities of moving images. Following his precursory experiments with coloured liquids and wax, Fischinger invented a compelling apparatus: the wax-slicing machine. First, coloured wax threads were cast into a block, which the machine pushed towards a revolving, fan-shaped blade. The blade would cut thin slices from the block as the film camera shot single frames through an apparatus in the blade, to which the camera shutter was synchronised (see Figs. 1 and 2). Fischinger explains: On a 2-dimensional plane, plastic forms were build up [sic] and formed in a block of colour wax—all kinds of forms and shapes and colour were imbedded in such a block of wax forms like pyramids or kegels [cones] or fantastic shaped forms like spirals etc. […] [A]fter such a wax block [was] finished squarely, [it] was put in a machine, which cut fine thin slices off the surface of the block. After each slice was cut off, a motion picture camera placed before the machine photographed the surface of the wax plane. Then the next slice was cut off and again one frame was photographed with the camera … Imagine the beauty of a polished cut through a wonderful stone […] somehow the camera records the cutting of the full stone from the beginning to the end. The camera would, so to speak, wander through and through the stone, the wonderful pattern would grow. (undated typescript) The methodology and research process of Dizziness––A Resource is inspired by the wax-slicing machine's animation process. Each cut in the wax block generates a shot from Fischinger's camera and thus the cut becomes a film frame. Setting these frames in motion animates the wax block's inherent dynamics. With every slice, Fischinger's abstract film expands, involving the viewer in its dramatic evolution. In contrast, the metaphorical wax block of our research-creation – a ‘block of sensations’ – grows by gradually incorporating different conceptual, emotional, and disciplinary approaches towards dizziness. According to Deleuze and Guattari: What is preserved – the thing or the work of art – is a bloc of sensations, that is to say, a compound of percepts and affects. Percepts are no longer perceptions; they are independent of a state of those who experience them. Affects are no longer feelings or affections; they go beyond the strength of those who undergo them. Sensations, percepts, and affects are beings whose validity lies in themselves and exceeds any lived. (1994: 163-64). Moreover, these percepts and affects are the slices or ‘snapshots’ of our research-creation created through giving the research process a momentarily distinct form through writing, making art, or staging cross-disciplinary events (Coleman and Ringrose, 2013). Like single film frames, these artworks and texts are preserved and recorded on our project blog. Comparable to animating the snapshots on film, the blog's visitor then animates the accumulated ‘snapshots’ of the research process and creates his or her own ‘animation’ by choosing what to see, hear, or read (http://on-dizziness.com). Moving through the blog, the viewer is able to experience artworks, animate knowledge, and gain a new perspective on dizziness. Fischinger's animation process inspired this methodology: describing and applying movement in order to connect and convey meaning or knowledge that cannot be exposed otherwise.","‘ … because in the end, dizziness, which I call ambiguity, is compossibility.’ (François Jullien, Interview, Paris, May 26, 2015) Falling into dizziness is a gradual process and enabling the experience of, and reflection on, dizziness requires a specific spatial and temporal setting. French philosopher François Jullien compares this process to passing through the ‘sas’ (French for ‘lock’, ‘sluice’, or ‘compression chamber’), a space of exchange and transformation, passage through which results in a change in motion and velocity (2015). The sas is the metaphorical space-time of the in-between (things, views, feelings, situations, definitions, theories, matters) where compossibility sets in (Bachelard, 1969; Jullien, 2012; Game and Metcalfe, 2011; Simmel, 1994). According to Jullien, the term ‘compossibility’ means the possible and inclusive togetherness of contradictory elements (2015). Out of this confrontational togetherness, an in-between space can emerge, creating the possibility of dissolving and/or re-arranging what has been so far. Evidence for such a compossible space is already found in Plato's notion of the chôra: a space, the formless form, literally the maternal womb or matrix (Burchill, 2011; Pechriggl, 2006). Chôra is neither being nor non-being but an interval in which ‘forms’ were originally held (Plato, Timaeus: 52d-53a). Another historical reference to compossible space is found in Leibniz's concept of possible worlds, constituted of compossible substances. These substances are only able to create a possible world when they do not contradict or exclude each other (Messina and Rutherford, 2009). However, by Jullien's contemporary definition, compossibility involves creating a space of ambiguity where established oppositions possibly collide, dissolve, and mix anew (2015). Compossible space may be used to describe a situation and condition that an individual, group, or society can transgress and designates the possible and inter-relational existence of several mutually exclusive and contradictory worlds or states, as in simultaneous experiences of fear and pleasure evoked by dizziness. Therefore, compossibility can be understood as a paradoxical space where established opposites and ideas are still recognisable but tend to dissolve. The different motions assigned to dizziness – falling into, passing through, and coming out – disturb and unbalance the constituting components of the compossible space, allowing for their re-combination. Fischinger's animation of the wax block may represent this passing through the compossible space. As motion, dizziness gradually becomes independent of the initial experience, still preserving it, but necessarily transforming the experience and ensuing emotions by reflecting on them, transforming thrill into fear into pleasure for instance. To reflect always means to create ruptures and pauses in the continuum of time and space; at the same time, reflection also requires cohesion in order to take the experience a step further towards understanding (Arendt, 2006). This reflexive process builds upon plateaus, it ‘slices the wax block’ of experience. Visualised in Fischinger's work, the slicing of the wax block and ensuing snapshots engender cohesion through their animation on screen and thus establish the compossible space for the passage through dizziness. In this sense, dizziness does not solely pertain to a theoretical concept, but also a physical sensation and an ambiguous emotional experience, shaking convictions and habits. At the same time ontological concept and symptom, method and object, sensation and metaphor, dizziness is apt to animate its contradicting constituting components, and this requires a space to unfold. The inclusion of feelings into the philosophical concept of dizziness – feelings preserved as affects like despair, fear, enjoyment or exaltation – contributes to the transformation of a static philosophy of being into a dynamic philosophy of becoming made of percepts, affects and concepts (Deleuze and Guattari, 1994). With the following example from our research-creation, we introduce artworks created through the experience of, and reflection on, dizziness. The perception of spinning, transformed into a photo-sculpture model, provided first insight into the interdependence of dizziness and compossible space. Concerning the sensation of spinning, the simple percept of dizziness proved insufficient, demanding a supplementary transformation of the subjective experience into an affect, as we will explain in the following chapter.","To elaborate on compossibility and dizziness, we became particularly interested in constituting the ephemeral moment of a spinning person's simultaneous standing and falling. It appears to be the turning point, the superposition of keeping upright and falling, the precise temporal moment of the compossibility of motion and standstill (Feyertag, 2015). Moreover, while falling we might believe we are suddenly falling from safety to uncertainty. But the basis on which we fall plays a significant part in our staggering and falling. The supposition that one is stable before falling is misleading; the staggering begins while we still feel safe and in control. To elucidate on these thoughts in practice, every morning for the first few weeks of the research-creation project we spun in turn until falling down in a dizzy state and noted our observations in a research diary. Soon it became clear that one can see the other person stumbling or falling, but no outward signs indicate the quality of the other's experience of dizziness or when or if the other will fall. The individual's experience of staggering and falling is peculiar. From an evolutionary point of view, staggering triggers reflexes that are very old, located in the lower region of the spinal cord rather than the brain. When staggering, we instantly relax the unsteady leg and tense the other in an effort to regain balance (Bähr and Frotscher, 2009). As adults unaccustomed to falling, we were hardly able to predict whether we will regain balance or fall. At times, we experienced our staggering and falling down as if it was happening in slow motion – we even felt detached from ourselves. The dizziness we encountered in our daily spinning was sometimes so strong that it impeded us to go on with our daily chores. Our notes describe that the dizziness felt slightly different every day, at times rather brief, at times a lingering experience. We felt particularly overshadowed by the ensuing dizziness when we were full of energy and eagerness to start the day. In the course of this self-experiment, the anticipation of feeling unwell after spinning resulted in a very unpleasant start of the day. Due to this, and the fact that we usually needed quite some time to recover after spinning, we stopped these proceedings after a few weeks. In lieu we started riding a merry-go-round whenever in a difficult spot in our research-creation, as we found that the uncomfortable memories of being dizzy increasingly impacted the actual experience. Clearly, our emotional conditions influenced our experience of dizziness, pointing to the situational and conditional character of experiencing dizziness. Conversely, whether pleasant or unpleasant, dizziness as an out-of-the-ordinary experience extracts the individual from the day-to-day and thus may present a freeing experience. Furthermore, within the experience of dizziness the compossibility of simultaneously standing (certainty) and falling (uncertainty) can be stated but not really observed except with still photography. Intrigued by the sensations of dizziness and the observation of our dizzy bodies, we began work on data for 3D models to compare the expressions of the dizzy body to the sensation of dizziness. Using photogrammetry, we tried to translate subjective sensations into percepts and affects with the help of ‘iconic’, a leading 3D studio in Istanbul. Surrounded by the studio's cage of cameras, we spun one after the other to the point of falling. After quite a few failed efforts, ‘iconic’ was finally able to record snapshots of us in the moment of simultaneously falling and standing, the data from which was used to print two photo-sculpture models (see Figs. 3, 4, & 5). These were created through the combination of a full-body colour photograph superposed on a 3D rendering of the body in motion. Anderwald's figurine depicts her staggering, precisely at the moment of falling. The statue cannot stand, velocity solely stabilised the moving body into an upright position in the captured moment, exposing one's illusion of control of the body's movements while already falling. Both statues visualise the unpredictability of the body's expression when the subject is dizzy. Indeed, Grond was falling backwards the moment his photo was taken. However, his figurine's poise seems balanced and stable because his movements before and after the snapshot cannot be anticipated from his composure. Both 3D snapshots designate a compossible point, a threshold where standing and falling are momentarily coinciding (Feyertag, 2015; Jullien, 2016). Nevertheless, the statues cannot transmit the feelings the individuals experienced. As a symptom, dizziness needs individual expression and cannot be sufficiently explained or understood by objective measurements or visualisation from/of the outside.","Dizziness––A Resource aims at a more holistic understanding of the potential of dizziness by describing it through cross-disciplinary practice, analysis, and reflection. In our conversations with artists, scientists, and experts in fields related to dizziness we realised that describing dizziness is dependent on the use of metaphorical language. We intensified the collaboration with philosopher Feyertag, with who we started co-authoring texts including an artistic text lending voice to a personification of dizziness. In this chapter we will provide an overview of the processes leading us to the creation of a sound installation that treats the exhibition Dizziness. Navigating the Unknown (Kunsthaus Graz and Ujazdowski Castle CCA Warsaw) as a film animation, including the works of fellow artists as snapshots that form a compossible space animated by the voice of personified dizziness. After capturing and analysing the dizzy body's expression, we shifted our attention to the feelings involved in experiencing dizziness using the generated 3D data as an artistic and reflective tool, considering this moment from different perspectives. Clearly, dizziness could not be sufficiently described from the outside alone, even if the data is cast into a sculpture and forms what Deleuze and Guattari term a ‘percept’. Studying dizziness as a symptom and process requires the consideration of individual expression and time. Therefore, we reanimated the captured data from the ‘iconic’ sessions as a series of still images and simultaneously searched for a narrative to translate the feelings of dizziness into another artistic language. In cooperation with graphic designer Christian Hoffelner, we produced The Bend (2015), a booklet that combines the still images from these sessions with a poetic text of the same name that emulates a spinning motion – walking in uncertainty – as it is written in a loop. This writing process transformed the percept of dizziness into an affect by adding a poetic narrative. (see. Fig. 6). Translating the research involving the 3D figurines further into creative writing, we started imagining dizziness as an archaic and hermaphroditic being, soliloquising its distress and desire. Now generated by the polyphony of the project's artistic and philosophical voices, our writing gave an account of the feelings that affected us while exposed to dizziness at various stages of the research process. Integrating narrative and personal forms of dizziness, we continued by producing the film Dizziness is My Name, using footage of the animated statues in addition to the co-authored monologue of ‘Dizziness’ in which the persona states: ‘Dizziness is my name and I am a pendulum without rope or gravity. My gravity is movement.’ This artistic work eventually led to the eponymous sound installation that guides viewers through the compossible space of the exhibition Dizziness. Navigating the Unknown (Kunsthaus Graz and Ujazdowski Castle CCA Warsaw). The narrating voice of ‘Dizziness’ entices the visitor to move erratically through the exhibition space, setting in motion the animation of dizziness exposed in the selected artworks. Our cooperation with Katrin Bucher Trantow, chief curator of Kunsthaus Graz, was not as much a search for a common vocabulary as it was a search for common ground in the work process. Our exchange over the research trajectory was continuous and lively, involving other artists, scientists, and curators. We discovered that we had to adapt our improvisational attitude, while Bucher Trantow adjusted her strategic approach. The resulting process of co-curating the aforementioned exhibition mirrored the research- creation's findings in historical and contemporary artworks. Bucher Trantow brought a focus on emotion to the research process, emphasising the importance of viewers' feelings towards individual artworks and their emotional journey through the exhibition space. Correspondingly, we divided the exhibition into three intersectional fields: falling into dizziness, navigating through dizziness, and coming out of dizziness. These specific foci are inspired by a conversation with artist Joachim Koester, who insisted on the importance of coming into and getting out of dizziness as the defining moments for future feelings towards this experience of dizziness (unpublished interview by Anderwald and Grond, 09/14/2014). DIZZINESS IN ARTISTIC WORK PROCESSES – COOPERATION WITH CREATIVITY RESEARCH AND KUNSTHAUS GRAZ -------------------------------------------------------------------------------- In cooperation with Mathias Benedek and Emanuel Jauk, creativity researchers at the University of Graz, we addressed dizziness' potential and what it might entail for creativity research. The first challenge was, as it is often in cross-disciplinary research, finding a common language. Merleau-Ponty's metaphor describing the search for creative expression as ‘a step taken in the fog’ (Merleau-Ponty, 1964: 3) was the starting point of discussion with Benedek and Jauk when explaining our understanding of dizziness as a possible resource. In our experience as artists, we have regarded dizziness as a resource for creating new artistic work (Anderwald et al., 2013), whereas the creativity researchers considered states of dizziness as unproductive. Therefore, we questioned a relevant group of artists on their experience. But how could their experience of dizziness be measured for creativity research? We decided to develop a survey on dizziness in creative work processes connected to an art competition in order to gather empirical insight and quantitative and qualitative data on artists' work processes in a valid setting as well as encourage new ideas and artworks on the topic. Together with Bucher Trantow, we designed ‘Living in a Dizzying World’, the competition that was tied to the survey for inclusion in the final exhibition of this research-creation at Kunsthaus Graz. Standardised surveys, personality tests (Rammstedt and John, 2005; Tubes and Christal, 1992), and divergent thinking tests (Nusbaum et al., 2014) were examined and, by updating some of the questions, a questionnaire applicable to artists in their studio environment was created. Another challenge was to establish a language that did not put the artists enrolled in the survey under additional stress, but at the same time was still valid for a creativity research sampling investigation. Despite our efforts, six of the forty-four artists dropped out of the competition and survey because they felt pressure responding to daily questions about their work process. Other artists found it advantageous, as this participant explains: Somehow I was comforted by the idea that other artists were reflecting on their process at the same time I was. It made me realize that the challenge of art making is not unique—it is difficult for all of us and can lead to different emotional states, etc. I thought that the survey was really beneficial for me in terms of paying attention to my process in an objective way. (Benedek et al., 2017: no pagination) Most artists appeared to have mixed feelings about their respective creative process, partially because the competing participants only had two weeks to create and submit their artwork to be judged by an international jury (Katrin Bucher Trantow, Sergio Edelzstein, Anna Jermolaewa). The daily questionnaire included seventeen recurring questions and an open section for comments and could be filled out via an app or online form. At the end, the artists were asked to upload their finished work to a university server. In order to re-start the applicants' work process at the beginning of the competition, a quote from David Bowie's song ‘Changes’, ‘turn and face the strange’ (1971), was sent to the participants as an additional inspiration for the work. Indeed, finding a timespan that was reasonable for the artists' work processes, but also accord with experience-sampling methods for the investigation of extended creative work was yet another methodological challenge. For two years of close cooperation, we gained insight into the emotions and experiences of dizziness in artists’ work processes. Our first conclusion is that how artists deal with states of dizziness is related to their personality structure, momentary condition, and past experience. Some artists stated that they never experience anything like dizziness in their work process, while others related the opposite. Between-person analysis revealed that artists with lower levels of agreeableness, higher levels of openness and a history of high artistic achievements generated artworks the jury found to be of superior quality (Benedek et al., 2017). The artworks in the competition examined dizziness mainly from an aleatoric or destructive perspective. The winning artwork Fractal Crisis (2016) by Viktor Landström and Sebastian Wahlforss follows a woman having a nervous breakdown, a situation of internal and external crisis (Bucher Trantow et al., 2017). Furthermore, the personality trait of openness allows for a greater ability to perceive two contradicting images simultaneously as separate and combined images (Antinori et al., 2017). This research-creation contextualises said ability with John Keats' notion of ‘negative capability’ coined in 1918 (Keats, 2014). It designates the capability of staying open and flexible when confronted with unknown, uncategorised or unpredictable knowledge or situations, – or in context with our research – the ability to enter and endure compossibility (Jullien, 2015). Within the compossible space the ability to use the movement of dizziness as creative potential becomes decisive. Philosophically speaking, this means opening up to the indefinite and shaping this opening that can be equated with defining the indefinite (Steinweg, 2013). In this sense the creative work process is considered oblivious to the aspects of impossibility. Creative processes therefore lean towards the aporetic and experimental, towards situations of unexpectedness and experiences of groundlessness and despair (Kofman, 1988; Alon, 2014). When we arrive at a dead end, with no possibility to carry on or through, are we forced to recombine all impossibilities – all which seems contradictory, exclusive, senseless, and disparate – in order to transform them into a compossible space and make for a yet unknown possibility.","In this article we defined the term ‘dizziness’ applied in Dizziness––A Resource and reflected on the concepts of dizziness and compossibility in theory and practice. The cross-disciplinary research-creation process of Dizziness––A Resource has led to the development of a methodology that draws from Fischinger's invention of the wax-slicing machine for film animation, an approach mirrored in the conception of the exhibition Dizziness. Navigating the Unknown. The experience of dizziness contains ambiguous and even contradicting feelings. In its reflection, dizziness exposes related emotions as movements that propel the individual in a certain direction or perspective. Moreover, understanding the emotional and spatial movements that constitute dizziness, requires a mode of thinking based on movement. If we are able to conceptualise dizziness as constant movement through spatial, emotional, and social surroundings, we may gain new perspectives on the affects and effects of thought. Therefore, we base dizziness, as well as this adapted process of research-creation, on ‘the idea of “movement”, movement conceived both as object and method, as syntagma and paradigm, as a characteristic of works of art and a stake in a field of knowledge claiming to have something to say’ (Didi-Huberman in Michaud, 2007) and relate it to Fischinger's animation process. We conclude that movement is at the core of our experience of, and reflection on, dizziness. Furthermore, only through movement that describes dizziness can the pure modal logic of compossibility be set in motion. As addressed artistically and philosophically over the research process, thinking-in-motion holds the potential to overcome the traditional oppositions of motion and standstill, certainty and uncertainty, knowing and not-knowing, because there is space and movement ‘in between’ professed opposites, which can become productive in moving towards new knowledge and meaning (Arendt, 2006). The concept of dizziness introduces an element of movement and openness to the ‘in between’ that could enable us to think the compossibility of opposites and acknowledge this grey zone or blandness for its creative and innovative potential (Bey, 1985; Deleuze and Guattari, 1994; Jullien, 2007; Manning, 2008). Therefore the development of ‘negative capability’ and enabling social and spatial surroundings are germane for transforming the ambiguity of dizziness into a resource. Hence, the experience of dizziness is never purely enjoyable, as anyone who has set foot on a rollercoaster can concur, but it can provide new sensations, stimulus, and input. This out-of-the-ordinary experience brings about feelings of excitement and exhilaration that can oscillate between elation and exasperation, and thus its concomitant unpredictability has to be dealt with on an individual and interpersonal scale. Encouraged by our cross-disciplinary cooperation, we plan to continue this research-creation to provide a more holistic understanding of how dizziness affects togetherness. The research-creation Dizziness––A Resource was supported by the Austrian Science Fund (AR-224), FWF-PEEK, AR 224. We are very grateful to Mathias Bendek and Emanuel Jauk, Institute of Psychology, University of Graz, and Katrin Bucher Trantow and Kunsthaus Graz (Universalmuseum Joanneum) for their valuable contributions to this research-creation. We are also thankful for the important comments of the peer reviewers that helped sharpen the focus of this contribution to ESS."],["This study investigated the mediation role played by children's executive function in the relationship between exposure to mild maternal depressive symptoms and problem behaviors. At ages 2, 3, and 6. years, 143 children completed executive function tasks and a verbal ability test. Mothers completed the Beck Depression Inventory at each time-point, and teachers completed the Strengths and Difficulties Questionnaire at child age 6. Longitudinal autoregressive mediation models showed a mediation effect that was significant and quite specific; executive function (and not verbal ability) at age 3 mediated the path between mothers' depressive symptoms (but not general social disadvantage) at the first time-point and children's externalizing and internalizing problems at age 6. Improving children's executive functioning might protect them against the adverse effects of exposure to maternal depressive symptoms. --------------------------------------------------------------------------------","Among parents of young children, symptoms of depression are common and often chronic (Field, 2011), such that McLennan, Kotelchuck, and Cho (2001) found that nearly a quarter (24%) of 17-month-olds were exposed to maternal depression, with a third of these children still exposed to depression a year later. This early exposure to maternal depressive symptoms predicts a plethora of negative child outcomes. Compared with children of nondepressed mothers, children of depressed mothers show elevated rates of both externalizing problems, such as hyperactivity (Ashman, Dawson, & Panagiotides, 2008), conduct disorder (Leschied, Chiodo, Whitehead, & Hurley, 2005), and violence (Hay, Pawlby, Waters, Perra, & Sharp, 2010), and internalizing problems, such as depression (Hammen & Brennan, 2003; Murray et al., 2011), anxiety (Gartstein et al., 2010), and social phobia (Biederman et al., 2001). Studies of the mechanisms underpinning these associations have, to date, focused on aspects of maternal functioning such as maternal regulatory processes (Dix & Meunier, 2009). Child functioning has received much less attention, which is surprising given that exposure to maternal depressive symptoms is related to cognitive abilities that are relevant for behavioral adjustment such as executive functions (Hughes, Roman, Hart, & Ensor, 2013) and language development (e.g., Quevedo et al., 2012). Research into the cognitive and neural mechanisms that may underlie childhood antisocial behaviors has highlighted the higher order processes associated with the prefrontal cortex that underpin flexible goal-directed action, collectively known as executive function (EF) (Hughes, 2011). The protracted development of the prefrontal cortex has led theorists to posit that EF might be particularly susceptible to environmental factors (e.g., Mezzacappa, 2004; Noble, Norman, & Farah, 2005). Studies of risk factors indicate that exposure to extreme adversity (i.e., maltreatment or neglect) has profound consequences for the functioning of the prefrontal cortex (for a review, see Belsky & De Haan, 2011). Until recently, however, few studies considered less extreme adversity such as exposure to mild maternal depressive symptoms (Odgers & Jaffee, 2013). The current study addressed this gap by focusing on children’s exposure to maternal depressive symptoms in a normative sample and by examining the relationship between exposure to maternal depressive symptoms and child EF over the course of early childhood (ages 2–6 years). The development of the prefrontal cortex is marked by growth spurts, with the first 3 years of life representing a time when the majority of myelination occurs and is paralleled by peaks in synaptic formation and dendritic growth (e.g., Spencer-Smith & Anderson, 2009). This heightened brain development translates into important EF developments through both refinements of acquired skills (e.g., Alloway, Gathercole, Willis, & Adams, 2004) and initial attempts to integrate and coordinate multiple functions (Garon, Bryson, & Smith, 2008). The emergence of toddlerhood as a critical period is further supported by research highlighting that individual differences in EF appear remarkably stable over time (Carlson, Mandell, & Williams, 2004; Fuhs & Day, 2011; Hughes & Ensor, 2007; Hughes, Ensor, Wilson, & Graham, 2010), with the implication that effects observed in preschoolers and school-aged children might simply reflect carry-on effects of early delays. Indeed, the complexity of EF development, whereby at each point emergent skills are reliant on the mastery of simpler abilities, lends support to the idea that EF development follows a cascading pathway model (e.g., Cummings, Davies, & Campbell, 2000). As such, this study focused on whether exposure to maternal depressive symptoms might translate into behavior problems specifically through delays in EF skills acquisition at a very early stage (i.e., age 3 years). The current article represents a secondary data analysis. To examine whether poor early EF is a mechanism through which exposure to maternal depressive symptoms translates into problem behaviors, the current study builds on several studies involving overlapping samples. Regarding the first path in this proposed mediation model, these earlier studies showed that children of mothers who had fewer depressive symptoms at child age 2 years or who displayed steeper recoveries from depressive symptoms over 4 years typically showed better EF at age 6 years (Hughes et al., 2013). This was true even when individual differences in children’s working memory at age 2 and maternal education and positive control at child ages 2 and 6 were accounted for (Hughes et al., 2013). In contrast, no relationship between maternal depression and child EF was found by two separate studies of older children in which maternal depression scores were dichotomized (Klimes-Dougan, Ronsaville, Wiggs, & Martinez, 2006; Micco et al., 2009). As discussed by Hughes and colleagues (2013), the most likely explanations for these discrepant findings relate to differences during the developmental period under focus (adolescence vs. early or middle childhood) and the sensitivity of the measures used. Regarding the latter, the previous studies that reported a relationship between maternal depression and child EF used self- reported continuous measures of depression that differentiated between mothers’ depression severity at all levels, whereas the studies that found no relationship used rigorous clinical diagnoses such that mothers with subclinical depression were placed in the “control” group together with nondepressed mothers. Regarding the second path of the mediation model, longitudinal analyses of data from the current sample have documented predictive relations between (a) poor EF at age 3 years and multi-informant ratings of problem behaviors at age 4 years (Hughes & Ensor, 2008) and (b) poor EF at age 4 years and low gains in EF from ages 4 to 6 years and teacher-rated problem behaviors at age 6 years (Hughes & Ensor, 2011). Importantly, even though these earlier studies demonstrated each of the two paths of the mediation model, they did not examine whether EF is a mechanism through which exposure to maternal depression transforms into problem behaviors. Indeed, significant relationships for different segments of a theoretical model of mediation do not conclusively establish a mediated effect (Kenny, 2012). In this case, elevated maternal depressive symptoms may independently predict both poor EF and child problem behaviors (i.e., multifinality of causes; Cicchetti & Toth, 1998). The only study to have investigated a mediation effect was an analysis of data from an enlarged sample of 235 4-year-olds comprising the study children and their best friends; although this analysis did show a mediation effect of child EF in the relationship between exposure to maternal depressive symptoms and child behavior difficulties (sample overlap: 66%; Hughes & Ensor, 2009a), the cross-sectional nature of the data limits the strength of conclusions that can be reached. To address this gap, the current study aimed to examine whether children’s EF at age 3 years mediates the relationship between mothers’ depressive symptoms at child age 2 years and children’s externalizing and internalizing problems at age 6 years.","At study entry, the sample comprised 143 families (out of a total of 192 eligible families), with a child aged between 24 and 36 months at the first visit and English as a home language, recruited in Cambridgeshire, United Kingdom. Starting with age 4 years, the sample was enlarged to include the study children’s best friends, yielding a new sample of 235 children. This implied that analyses focused on toddlerhood involved the core sample, whereas analyses focused only on preschool and later development involved the enlarged sample, with a sample overlap of approximately 66%. The current study focused on toddlerhood; therefore, data analysis was applied to members of the core sample for whom data were available at ages 2 and 3 years. Face-to-face recruitment was carried out at support groups for young mothers and at every mother–toddler group in wards within the highest quartile nationally of deprivation (Noble et al., 2008), thereby resulting in a socially diverse sample. Informed consent and ethical approval were obtained for all assessments. As a token of thanks for their participation, families received £20 (i.e., 20 British pounds) for each visit to the home. A copy of the video footage recorded in the lab and taxi/travel costs were also provided. More than 95% of the children were White/Caucasian and 56 (39%) were girls. At the study entry, 25% of women were single parents and 53% of mothers had no education qualifications or had education qualifications only up to GCSE (general certificate of secondary education) level (usually obtained at age 16 years). Regarding family size, 12% of children had no siblings, 53% had one sibling, 22% had two siblings, and the remaining 13% had three or more siblings. This study focused on data collected when children were 2 years old (Time 1; SDage 2 = 4 months), 3 years old (Time 2; SDage 3 = 4 months), and 6 years old (Time 3; SDage 6 = 4 months). Deprivation Following the example of Moffitt et al. (2002), deprivation at study entry was evaluated using eight markers: maximum education level per family was GCSE (usually obtained at age 16 years), head of household occupation in elementary or machine operation occupation, annual household income under £10,000, receipt of public benefits, council house as the home, home reported as crowded by the mother, single-parent family, and no access to a car. Of the 127 (89%) families without missing data, 31% of families showed no deprivation (zero markers), 32% of families showed moderate deprivation (one or two markers), and 37% of families showed high deprivation (three or more markers). Mothers’ depressive symptoms Maternal depressive symptoms were self-reported using the Beck Depression Inventory (BDI; Beck, Ward, Mendelson, Mock, & Erbaugh, 1961) at all time-points. The BDI includes 21 questions about affective, somatic, and cognitive symptoms. Item response categories range from 0 (minimal) to 3 (extreme). Item intraclass correlation was high at all time-points, with a Cronbach’s alpha ⩾ .82. Two- parameter logistic item response theory models applied at each time-point revealed that the BDI was especially successful in identifying mothers who experienced mild to moderate depressive symptoms. Some of the items that showed lowest power to discriminate between different levels of depressive symptoms were those about “weight loss,” “loss of interest in sex,” and “somatic preoccupation.” Some of the items with consistently high discrimination across assessments were those about “concentration difficulty,” “past failure,” or “loss of pleasure.” The most “difficult” items (i.e., items endorsed only by very depressed mothers) were “punishment feelings,” “suicidal ideation,” “somatic preoccupation,” and “past failure,” whereas the “easiest” items (i.e., items also endorsed by nondepressed mothers) were “tiredness or fatigue,” “irritability,” and “changes in sleeping patterns.” Children’s problem behaviors At child age 6 years, 74 teachers completed the Strengths and Difficulties Questionnaire (SDQ; Goodman, 1997) for 120 of the children. The SDQ response categories range between 0 (not true) and 2 (certainly true), and the 25 items are divided into 5-item subscales: a positive subscale (prosocial behavior) and four problem subscales (conduct problems, hyperactivity, emotional problems, and peer problems). This study focused on the problem subscales. Item intraclass correlation was high for the total problem behavior score (Cronbach’s alpha = .82) and for the items comprising the externalizing (Cronbach’s alpha = .84) and internalizing (Cronbach’s alpha = .74) problems subscales. Verbal ability At ages 2 and 3 years, children’s verbal ability was assessed using the Naming and Comprehension subtests of the British Abilities Scales (BAS; Elliott, Murray, & Pearson, 1983), which tap expressive and receptive language skills. Scores ranged between 0 and 20 for the Naming subtest and between 0 and 27 for the Comprehension subtest. At age 6 years, verbal ability was measured using the Revised British Picture Vocabulary Scale (BPVS; Dunn, 1997), which taps receptive language. Total scores were computed as a sum. Executive function At ages 2 and 3 years, children completed four tasks: Beads, Trucks, Baby Stroop, and Spin the Pots. At age 6 years, children completed three tasks: Beads, Day–Night, and Tower of London. The tasks were delivered by researchers (PhD students and research assistants) trained by the task developer. The Beads task, a part of the Stanford–Binet Intelligence Scales (Thorndike, Hagen, & Sattler, 1986), taps working memory. Four warm-up trials are followed by 10 trials where children are shown 1 or 2 beads (for 2 or 5 s, respectively) and must identify them in an image of 12 beads sorted by color and shape. In a further 16 trials, children are presented with a photographic representation of a bead pattern for 5 s, which they must reproduce with real beads arranged on a stick. Scores represent the number of correct trials and range between 0 and 26 points. The Trucks task taps rule learning and switching (Hughes & Ensor, 2005), each tested with an 8-trial phase. Children must guess which of two pictures of trucks will lead to a reward. The first truck chosen by children gives the rule in the first phase; the opposite truck gives the rule in the second phase. Scores represent the number of correct trials and range between 0 and 16 points. The Spin the Pots task (Hughes & Ensor, 2005) assesses working memory. Children are shown eight distinct “pots” (e.g., jewelry boxes, candy tins, wooden boxes) placed on a Lazy Susan tray and are invited to help the researcher place attractive stickers in six of the eight pots. Then, the tray is covered with a cloth and spun, after which the cloth is removed and children must choose a pot. The test is discontinued when children find all of the stickers or after 16 attempts. The score is 16 minus the number of errors made. The Baby Stroop task taps inhibitory control (Hughes & Ensor, 2005). Children are presented with a normal-sized cup and spoon and a baby-sized cup and spoon. In the control phase, children must name the large cup/spoon “mommy” and the small cup/spoon “baby.” In the second phase, children must use the labels incongruously. The Day–Night task is a Stroop task designed for older children (Gerstadt, Hong, & Diamond, 1994). This task is identical to the Baby Stroop task except for the props; the cup and spoon are replaced by two abstract patterns representing “Day” and “Night.” The 12 trials are presented in a pseudo-random order, with scores ranging from 0 to 12. The Tower of London task (Shallice, 1982), taps planning abilities. The props include a wooden board with three pegs of unequal size and three large spongy balls. The large peg can carry three balls, the middle peg can carry two balls, and the small peg can carry only one ball. Children must reproduce arrangements presented in an image by moving only one ball at a time and using the minimum number of moves needed. Warm-up trials with one- move problems are followed by two-, three-, and four-move problems (three problems each). Children achieve 2 points for success using the minimum number of moves, 1 point for success with the use of extra moves, and 0 points for failure to complete the problem or when more than 2n + 1 extra moves are necessary. Total scores range between 0 and 18 points.","Analyses were conducted using Mplus Version 7.11 (Muthén & Muthén, 1998–2010). Power analyses were performed using the macro developed by Preacher and Coffman (2006). Little’s (1998) MCAR tests applied to missing data patterns revealed that data were missing completely at random for all variables. Specifically, data were missing completely at random in relation to maternal depressive symptoms, such that depressed mothers were not more likely to drop out of the study than nondepressed mothers. Maternal depression scores were based on longitudinal factor analysis; as such, scores were estimated in instances with missing data. Therefore, factor scores of maternal depressive symptoms were available for all mothers and at all time-points. The same was true for child EF. Data were also missing completely at random in relation to child problem behaviors, and further missing value analysis revealed that missing information about problem behaviors was not more likely to come from boys than girls or from children of mothers with low education than other children. Missing data were avoided with regard to problem behaviors for those cases with data on some of the items through the use of factor analysis in the creation of final scores. To account for data missingness and skewness, we used robust estimators: WLSMV to obtain factor scores of depressive symptoms and child problem behaviors (here indicators were categorical) and MLR to specify factors of EF and mediation models (here all measures were continuous). Model fit was evaluated using the comparative fit index (CFI), the Tucker–Lewis index (TLI), and the root mean square error of approximation (RMSEA). Adequate fit was achieved for CFI and TLI values ⩾.90 and for RMSEA values ⩽.08 (Hu & Bentler, 1999). Good fit was achieved for CFI and TLI values ⩾.95 and for RMSEA values ⩽.06 (Bentler, 1990). The chi-square is also reported for all models but was not used in the evaluation of model fit due to the tendency of the chi-square to over-reject true models for large samples and/or models with many degrees of freedom (Bentler, 1990). Fully standardized coefficients are presented.","Across time-points, 65% to 73% of mothers exhibited no or minimal depression (i.e., scores < 10), 19% to 24% exhibited mild to moderate depression (i.e., scores of 10–18), and 7% to 12% exhibited moderate to severe levels of depressive symptoms (i.e., scores ⩾ 19). Mean levels of child problem behaviors (M = 8.07) fell within the normal range (i.e., scores of 0–11), but elevated problem behaviors were also noted (Max = 25), and on average scores were slightly higher (on all dimensions except peer problems) than a nationally representative sample of 5- to 10-year-old children in Great Britain (youthinmind, 2012). Specifically, borderline or abnormal levels of problem behaviors were exhibited by 13.8% of the children in relation to conduct problems, 34.7% of the children in relation to hyperactivity, 11% of the children in relation to emotional problems, 14.5% of the children in relation to peer problems, and 26.1% of the children in relation to total behavior difficulties. Table 1 summarizes children’s mean levels of EF and verbal abilities. As expected, children improved in their mean levels of verbal abilities and EF across time. Data reduction ~~~~~~~~~~~~~~ We created factor scores from the questionnaire-based variables to index mothers’ depressive symptoms at each time-point and child problem behaviors at the final time- point. We used the procedure described by Ensor, Roman, Hart, and Hughes (2012) to create factor scores of mothers’ depressive symptoms by dichotomizing BDI items due to low selection rates of the higher categories and applying single-factor confirmatory factor analysis with scalar invariance (i.e., equal structure, loadings, and thresholds) to factors obtained at each time-point. The model achieved good power (α = 1.00) and adequate fit, χ2(1967) = 2184.025, p < .01, RMSEA = 0.03, 90% confidence interval (CI) [0.02, 0.04], CFI = .92, TLI = 0.92. We used the procedure described by Goodman, Lamping, and Ploubidis (2010) to create factor scores of child problem behaviors by applying confirmatory factor analysis with four first-order factors (conduct problems, hyperactivity, emotional problems, and peer problems) and two second-order factors (externalizing problems [conduct problems and hyperactivity] and internalizing problems [emotional problems and peer problems]). The model achieved good power (α = .91) and adequate fit, χ2(166) = 291.74, p < .01, RMSEA = 0.08, 90% CI [0.06, 0.09], CFI = .91 TLI = 0.90. All item loadings onto first-order factors (β ⩾ .57, p < .01) and all first-order factor loadings onto second-order factors were significant (β ⩾ .45, p < .01). In addition, we used the findings of exploratory factor analyses previously reported by Hughes and Ensor (2008) at ages 3 and 4 years to specify a single factor of EF at each time-point. At ages 2 and 3 years, we specified a factor based on children’s scores on the Beads, Trucks, Baby Stroop, and Pots tasks. At age 6 years, we specified a factor based on children’s scores on the Beads, Stroop, and Tower of London tasks. To examine stability in EF over time, we specified regression paths between EF factors at consecutive time-points. To account for the known association between EF and verbal ability, we included indicators of verbal ability as time-varying covariates of EF and specified regression paths between verbal ability indicators at consecutive time-points. The model achieved adequate power (α = .73) and adequate fit, χ2(73) = 106.54, p < .01, RMSEA = 0.06, 90% CI [0.03, 0.08], CFI = .93, TLI = 0.91. The range of factor loadings obtained here (βrange = .24–.65, p < .01) was similar to that reported by studies where EF was constructed using slightly different tasks (e.g., .17–.63; Espy, Kaufmann, Glisky, & McDiarmid, 2001). Fully standardized parameter estimates indicated that inter-individual differences in EF (β2–3 = .91 and β3–6 = .80, p < .01) and verbal ability (β2–3 = .79 and β3–6 = .60, p < .01) were stable over time. In addition, EF and verbal ability were significantly positively related at each time-point (rage 2 = .84, p < .01, rage 3 = .56, p < .01, and rage 6 = .78, p < .01). Mediation analyses ~~~~~~~~~~~~~~~~~~ To test mediation effects, we specified the autoregressive longitudinal mediation model presented in Fig. 1. The model showed good fit to the data, χ2(160) = 203.215, p = .01, RMSEA = 0.04, 90% CI [0.02, 0.06], CFI = .95, TLI = 0.94. The model achieved good power (α = .95). The power of the model was enhanced by the large number of degrees of freedom of the model (df = 160 in the mediation model), which offset any power issues related to a small sample size (143 participants). Relative to girls, boys had poorer EF at age 2 years (β = −.21, p = .01) and higher externalizing problems at age 6 years (β = .13, p < .05). Net of these effects, higher maternal depressive symptoms at child age 2 predicted poorer child EF at age 3 (β = −.20, p < .01) even when accounting for stability in EF from age 2 to age 3 (β = .90, p < .01) and the significant association between EF and verbal ability at age 3 (β = .71, p < .05). In turn, poorer EF at age 3 predicted higher externalizing problems at age 6 (β = −.32, p < .01) and higher internalizing problems at age 6 (β = −.32, p < .01). This was true even given the significant concurrent relationship between child EF and externalizing problems (β = −.48, p < .01) and between child EF and internalizing problems (β = −.43, p < .05). Regarding externalizing problems, the indirect effect was significant (βind = .07, p < .05, 95% CI [0.001, 0.130]). Regarding internalizing problems, the indirect effect was marginally significant (βind = .06, p = .059, 95% CI [−0.002, 0.134]). To further probe the robustness of our findings, we tested two alternative models. The first model (Fig. 2) showed that verbal ability did not act as an alternative mediator, although higher maternal depressive symptoms at child age 2 years predicted marginally significantly lower child verbal ability at age 3 years, which in turn predicted higher externalizing and internalizing problems at age 6 years. The second model (Fig. 3) showed that the observed effects of depressive symptoms on behavior adjustment via poor EF did not simply reflect effects of deprivation more generally because, although exposure to higher levels of deprivation at age 2 directly predicted higher externalizing problems at age 6 and poorer EF at age 3, and poorer EF at age 3 predicted higher externalizing and internalizing problems at age 6, the indirect effects were not significant.","This study investigated whether the path from maternal depressive symptoms to young children’s problem behaviors operates, at least in part, through impairments in children’s EF. Building on a previous model, which tested the long-term associations between children’s exposure to maternal depressive symptoms during toddlerhood and EF at age 6 years (Hughes et al., 2013), the current findings support Cummings et al. (2000) cascading pathway model of development in that they demonstrate more immediate associations between exposure to maternal depressive symptoms and poor EF, which are then carried over time. Regarding the second path, our findings were consistent with those from a recent meta- analytic review (Schoemaker, Mulder, Deković, & Matthys, 2013), which showed a weak to moderate association between preschool children’s externalizing problems and their overall EF, inhibitory control, working memory, and attention shifting (Schoemaker et al., 2013). Our findings are also consistent with previous reports of associations between internalizing problems and overall EF (Riggs, Blair, & Greenberg, 2003), working memory (Brocki & Bohlin, 2004), and inhibitory control (e.g., Rhoades, Greenberg, & Domitrovich, 2009). Our main study finding was that individual differences in EF at age 3 years significantly mediated the relationship between mothers’ depressive symptoms at child age 2 years and children’s externalizing problems at age 6 years. To our knowledge, this is the first study to test such mediation effects within a rigorous autoregressive longitudinal design, whereby measures (all except child problem behaviors) were assessed repeatedly at each of the three time-points. The prospective longitudinal design and the selected time intervals add to the significance of findings; ages 2 and 3 are key periods when children improve on their EF skills and acquire more advanced types of EF, whereas age 6 follows the transition to school. Importantly, the current findings indicate that the mediation role of EF is unlikely merely to reflect associations with other cognitive abilities that underlie behavior adjustment such as verbal fluency. Verbal ability is highly related to EF (e.g., Hughes & Ensor, 2007), and until recently tasks measuring various EF components tended to be verbal in nature (e.g., the backward word span task, tapping working memory) (Carlson, Moses, & Breton, 2002). It is noteworthy, therefore, that verbal ability was not a mediator, probably because heightened maternal symptoms only predicted marginally significantly lower verbal ability. Although conclusions are limited by the inclusion of a single indicator of verbal fluency, this finding suggests that different cognitive domains might not be equally important in the relation between exposure to maternal depressive symptoms and problem behaviors and calls for future studies of child mediators. Along the same lines, it is also important to note that child EF did not mediate the relationship between general deprivation and problem behaviors. Depressive symptoms occur more often in the context of deprivation (Stansfeld, Clark, Rodgers, Caldwell, & Power, 2011), and previous studies have dealt with this convoluted relationship by showing that maternal depression mediates the relationship between deprivation and child problem behaviors (Rijlaarsdam et al., 2013) and that reduced economic resources mediate the relationship between depressive symptoms and child behavior (Turney, 2012). Alternatively, measures of deprivation and depressive symptoms have been combined into a single measure of family risk (e.g., Halligan et al., 2013). Not surprisingly, we found that greater deprivation at age 2 years predicted poorer EF at age 3 years and heightened externalizing problems at age 6 years. However, the indirect effect was not significant; this null finding is important because it suggests that the mediatory role of poor EF is specific to the effects of exposure to maternal depressive symptoms. Establishing the specificity of risk factors implicated in an EF-mediated pathway to problem behaviors is theoretically significant. In particular, a significant mediation effect of EF in the relationship between exposure to deprivation and problem behaviors would have been consistent with the suggestion that those variations in problem behaviors that are due to impairments in EF are mostly created by inadequate levels of overall stimulation and environmental inconsistency, supporting theoretical conceptualizations of EF as a regulatory process dependent on “optimal” stimulation. In contrast, our findings of a significant mediation effect of EF in the relationship between exposure to maternal depressive symptoms and problem behaviors is consistent with the suggestion that variations in problem behaviors due to impairments in EF might be caused by poor-quality interpersonal relationships, supporting models of EF development through modeling of adult behavior and observational learning (Hughes & Ensor, 2009b). Maternal depression has been associated with impairments in mothers’ own EF skills (e.g., Barrett & Fleming, 2011; Johnston, Mash, Miller, & Ninowski, 2012). Children are keen observers of adults’ everyday behaviors (Dunn, 1993) and so may internalize problem-solving strategies that do not rely on the use of high EF skills. For example, depressed mothers are less likely to use planning in order to optimize repetitive tasks (Hughes & Ensor, 2009b), which might translate into fewer opportunities for their children to observe and internalize such strategies. In addition, depression is known to deplete emotional resources by activating an oversensitized distress response system and reducing the threshold for what is considered aversive (for a review, see Dix & Meunier, 2009). Both deficits in maternal EF (Psychogiou & Parry, 2014) and deficits in emotion processing and emotion regulation (Dix & Meunier, 2009) have been hypothesized as mechanisms through which depression translates into low maternal cognitive flexibility. A reduced ability to respond contingently would adversely affect parental scaffolding of children’s goal-directed activities, which has been shown to predict EF development (Bernier, Carlson, & Whipple, 2010; Hughes & Ensor, 2009b; Schroeder & Kelley, 2010). Scaffolding also involves showing children how to solve (a part of) the next step, which children immediately imitate, internalize, and then apply to further steps in the process until the next stage of difficulty is achieved and the scaffolding process is repeated. In other words, to a large extent, scaffolding can be seen as a form of concentrated and focused observational learning. Future studies that take a close look at how scaffolding unfolds in interactions between children and mothers in the context of depression could greatly enhance our understanding of the processes that relate exposure to maternal depressive symptoms and child EF development. Other factors that might be expected to distort the mediation effect reported here include parenting and maternal functioning. These have been omitted from the current study due to a focus on child mediators. However, the inclusion of measures of parenting and the parent–child relationship is unlikely to have altered the results because previous analyses of data from this sample showed that including maternal positive control at ages 2 and 6 years did not weaken the negative effects of mothers’ depressive symptoms at child age 2 on child EF at age 6 (Hughes et al., 2013); likewise, there was no interaction between observed mother–child mutuality (at both ages 2 and 6) and mothers’ depressive symptoms at age 2 as predictors of child problem behaviors at age 6 (Ensor et al., 2012). Another important environmental factor not included here is the father, an omission largely caused by the absence of fathers in more deprived families (25% of the mothers were single parents at study entry). A key strength of this study is the use of independent assessments for each variable. Mothers’ depressive symptoms were self-reported, children’s EF abilities were tested through age-appropriate experimental tasks, and children’s externalizing and internalizing problems were reported by teachers. On the flipside, a limitation of this study is the use of a single measurement method for each variable. For example, mothers’ self-reported depressive symptoms were not corroborated by clinical interviews. Children’s EF abilities were measured by multiple measures, but each subcomponent (i.e., working memory, inhibitory control, or attention shifting) was measured by a single task. This said, all main constructs were factor analyzed and used either as latent variables or as resulting factor scores in all analyses, thereby reducing measurement error. The small sample size and the presence of relatively few mothers with depressive symptoms indicate that results are preliminary in nature and warrant further validation with larger samples. The current study has several implications for interventions aimed at reducing children’s externalizing problems. Combined “two-generation” interventions aimed at improving both mothers’ depressive symptoms and children’s externalizing problems are a potential avenue given that improvements in maternal depressive symptoms have been associated with more successful treatment of child externalizing problems (Van Loon, Granic, & Engels, 2011). On the downside, such interventions might be more susceptible to failure if the component aimed at reducing maternal depressive symptoms is not successful, as has been indicated by research showing that heightened maternal depressive symptoms interfere with the positive effects of interventions aimed at child adjustment (Beauchaine, Webster-Stratton, & Reid, 2005; Van Loon et al., 2011; but see Rishel et al., 2006, for null findings). From this point of view, the current study’s implication that interventions that improve children’s EF could help to reduce externalizing problems in children exposed to maternal depression (as indicated by the significant longitudinal mediation effect) is particularly noteworthy. Previous research has shown that children’s EF is improved by a variety of interventions, some of which can easily be implemented in schools (for a review, see Diamond & Lee, 2011). In addition, it may be possible to promote children’s EF through interventions aimed at improving the home environment such as by reducing chaos (e.g., Evans, Gonnella, Marcynyszyn, Gentile, & Salpekar, 2005) and encouraging parents to limit children’s exposure to fast-paced cartoons (Lillard & Peterson, 2011). In sum, despite the omissions identified above, and within the limitations imposed by the study design, by demonstrating that individual differences in child EF at age 3 years mediate the relationship between exposure to maternal depressive symptoms at age 2 years and teacher ratings of externalizing and internalizing problems at age 6 years, this study adds to our understanding of the self-regulatory mechanisms through which exposure to maternal depressive symptoms might translate into child problem behaviors. Moreover, our findings provide a potential avenue in the quest for factors that may buffer young children from the adverse effects of exposure to mothers’ depressive symptoms."],["In a range of contexts, individuals arrive at collective decisions by sharing confidence in their judgements. This tendency to evaluate the reliability of information by the confidence with which it is expressed has been termed the 'confidence heuristic'. We tested two ways of implementing the confidence heuristic in the context of a collective perceptual decision-making task: either directly, by opting for the judgement made with higher confidence, or indirectly, by opting for the faster judgement, exploiting an inverse correlation between confidence and reaction time. We found that the success of these heuristics depends on how similar individuals are in terms of the reliability of their judgements and, more importantly, that for dissimilar individuals such heuristics are dramatically inferior to interaction. Interaction allows individuals to alleviate, but not fully resolve, differences in the reliability of their judgements. We discuss the implications of these findings for models of confidence and collective decision-making. © 2014 The Authors. --------------------------------------------------------------------------------","There is a growing interest in the mechanisms underlying the “two-heads-better-than-one” (2HBT1) effect, which refers to the ability of dyads to make more accurate decisions than either of their members (e.g., Hill, 1982). One study (Bahrami et al., 2010), using a perceptual task in which two observers had to detect a visual target, showed that two heads become better than one by sharing their ‘confidence’ (i.e., an internal estimate of the probability of being correct), thus allowing them to identify who is more likely to be correct in a given situation. Sharing of confidence as a strategy for combining individual opinions into a group decision has also been established in non-perceptual domains (e.g., Sniezek & Henry, 1989). This tendency to evaluate the reliability of information by the confidence with which it is expressed has been termed the ‘confidence heuristic’ (e.g., Thomas & McFadyen, 1995). A recent study has shown that a simple algorithm based on the confidence heuristic – always opt for the opinion made with higher confidence – can yield a 2HBT1 effect in the absence of any interaction between individuals (Koriat, 2012). Intrigued by this finding, we tested whether this algorithm could in practice replace interaction in collective decision-making. Importantly, such a formula for collective choice – if effective – would not be susceptible to the egocentric biases that may impair interaction (e.g., Gilovich, Savitsky, & Medvec, 1998), and could readily be used by decision makers, such as jurors, medical doctors or financial investors, who have to combine different opinions in limited time. Indeed, the implementation of heuristics inspired by individual decision-making has proved very useful within professional contexts (e.g., Gigerenzer, 2008). Circumventing interaction ~~~~~~~~~~~~~~~~~~~~~~~~~ Building on Bahrami et al.’s (2010) study, Koriat (2012) asked isolated observers to estimate the degree of confidence in their perceptual decisions. Participants, all of whom had received the same sequence of stimuli, were afterwards paired into virtual dyads so that they matched each other in terms of their ‘reliability’ (i.e., the reliability of their individual decisions about the visual target). To remove individual biases in confidence, their confidence estimates were normalised, so that they shared the same mean and standard deviation, before being submitted to the Maximum Confidence Slating (MCS) algorithm, which selected the decision of the more confident member of the virtual dyad on every trial. While circumventing interaction, the MCS algorithm yielded a robust 2HBT1 effect. Interestingly, isolated observers’ confidence estimates are negatively correlated with their reaction times when responses are given in the absence of speed pressure (e.g., Patel, Fleming, & Kilner, 2012; Pleskac & Busemeyer, 2010; Vickers & Packer, 1982), raising the possibility that a Minimum Reaction Time Slating (MRTS) algorithm may be sufficient to yield a 2HBT1 effect. In this study, we tested the efficacy of the MCS and MRTS algorithms without matching dyad members in terms of their reliability, and compared the responses advised by the algorithms with those reached by the dyad members through interaction (henceforth ‘dummy’ versus ‘empirical’ dyads/decisions). In particular, we addressed three questions. First, does the success of the MCS and MRTS algorithms depend on the similarity of dyad members’ reliabilities? Bahrami et al. (2010) found that the success of interactively sharing confidence was a linear function of the similarity of dyad members’ reliabilities. For similar dyad members, two heads were better than one. However, for dissimilar dyad members, two heads were worse than the better one. Interestingly, Bahrami et al. (2010) found that these discrepant patterns of collective performance could be explained by a computational model in which confidence was defined as a function of the reliability of the underlying perceptual decision. We predicted that the efficacy of the MCS and the MRTS algorithms would also depend on the similarity of dyad members’ reliabilities. Second, do the algorithms fare just as well as interacting dyad members? People vary in their ability to estimate the reliability of their own decisions (e.g., Fleming, Weil, Nagy, Dolan, & Rees, 2010; Song et al., 2011); this ability is typically referred to as ‘metacognitive’ ability and, in social contexts, determines the credibility of people’s confidence estimates. While the algorithms are prone to error when people misestimate the reliability of their own decisions, interacting individuals may take such misestimates into account (e.g., Tenney, MacCoun, Spellman, & Hastie, 2007). We predicted that interacting dyad members would take into account the credibility of each other’s confidence estimates when making their joint decisions, and that interaction would be relatively more beneficial than the algorithms for dissimilar dyad members; they have more to lose from following the more confident but less competent of the two. Third, what is the effect of normalising confidence estimates before selecting the decision made with higher confidence? Koriat (2012) reported that the MCS algorithm performed equally well when using raw and normalised confidence estimates as its input. However, this analysis was limited to (virtual) dyad members of nearly equal reliability. Even though people vary in the ability to evaluate the reliability of their own decisions, confidence estimates are rarely uninformative about underlying performance (e.g., Lau & Maniscalco, 2010). As a consequence, normalising confidence estimates may remove statistical moments that reflect actual differences in underlying performance (e.g. differences in average confidence due to differences in average performance). We therefore predicted that submitting normalised confidence estimates to the MCS algorithm would be relatively more costly for dissimilar dyad members.","To test our predictions, we analysed data from an experiment (Bahrami et al., 2012a) in which dyad members estimated their confidence in individual decisions on every trial, but were also required to make a joint decision whenever their individual decisions conflicted. We used data from two experimental conditions: a ‘non-verbal’ condition in which dyad members made their joint decisions only having access to each other’s confidence estimates, and a ‘verbal’ condition in which dyad members also had the opportunity to verbally negotiate their joint decisions (the NV condition and the NV&V condition in Bahrami et al., 2012a, 2012b). In total, fifty-eight participants (29 dyads) took part in the non-verbal (14 dyads) and the verbal (15 dyads) conditions. All participants were healthy adult males (mean age = 23.5 years, SD = 2.8 years) with normal or corrected-to-normal vision. The members of each dyad knew each other before taking part in the experiment. We describe the key experimental details below (see Bahrami et al., 2012a, 2012b, for response, display and stimulus parameters). Experimental details ~~~~~~~~~~~~~~~~~~~~ Dyad members sat at right angles to each other in a dark room, each with their own screen and response device. For each trial, dyad members viewed two brief intervals, on which six contrast gratings were presented simultaneously around a central fixation point. In either the first or the second interval, one of the six contrast gratings had a slightly higher level of contrast. After the two viewing intervals, a horizontal line with a fixed midpoint appeared on each dyad member’s screen. The left side of the midpoint represented the first interval, the right side represented the second. An additional vertical confidence marker was displayed on top of the midpoint. Each dyad member made his private decision about which interval he thought contained the oddball target by moving the confidence marker to the left (first interval) or to the right (second interval) of the centre. The confidence marker could be moved along the line by up to five steps on either side, each step indicating higher confidence. There was no response time limit. Next, the dyad member’s private responses (decision and confidence) were shared. In the case of agreement (i.e., if dyad members privately selected the same interval), they received feedback and continued to the next trial. But in the case of disagreement, one of the two dyad members was randomly prompted to make a joint decision on behalf of the dyad. In the non-verbal condition, the nominated dyad member made the joint decision only having access to the declared responses. In the verbal condition, the nominated dyad member also had the opportunity to verbally negotiate the joint decision with his partner. Dyad members were free to ignore each other’s confidence estimates at this stage of the experiment. After one practice block of 16 trials, two experimental sessions were conducted. Each session consisted of 8 blocks of 16 trials (128 trials in each session and 256 trials in total). Within each session, one dyad member responded with the keyboard (colour-coded as blue) and the other dyad member responded with the mouse (colour-coded as yellow). The dyad member using the keyboard controlled the confidence marker by pressing the ‘n’ (move left), the ‘m’ (move right), and the ‘b’ (submit response) buttons. The dyad member using the mouse controlled the confidence marker by pressing the ‘left’ (move left), the ‘right’ (move right), and the ‘scroll’ (submit response) buttons. Dyad members switched places (and hereby response device) at the end of the first session. See Fig. 1 for a schematic of an experimental trial. Computing MCS responses ~~~~~~~~~~~~~~~~~~~~~~~ We used Goodman & Kruskal’s gamma to test whether high (or low) confidence was associated with correct (or incorrect) decisions. As expected, the gamma coefficients (mean = .16, SD = .09) were significantly positive across dyad members, t(57) = 14.04, p < .001. For each dyad, we then derived the joint decisions advised by the MCS algorithm. In line with Koriat (2012), we first normalised the dyad members’ confidence estimates and then selected the decision of the more confident dyad member on each trial; the normalised confidence estimates, cnorm, were related to the raw confidence estimates, c, via cnorm = (c − μc_dyad)/σc_dyad, where μc and σc were the mean and the standard deviation of the raw confidence estimates pooled from both dyad members. Computing MRTS responses ~~~~~~~~~~~~~~~~~~~~~~~~ Due to a programming error, the experimental code continued to sample the mouse response time until both mouse and keyboard responses had been registered. This error meant that mouse response times were either identical to keyboard response times (i.e., those trials in which the dyad member using the mouse actually made a faster response than the dyad member using the keyboard) or slower than keyboard response times (i.e., those trials in which the dyad member using the mouse made a slower response than the dyad member using the keyboard). Only including those trials in which dyad members responded with the keyboard (i.e., the first 128 trials for dyad member A and the last 128 trials for dyad member B – see Section 2.2), we regressed reaction times (milliseconds) against confidence to test whether faster decisions were associated with higher confidence. As expected, the unstandardised regression coefficients (mean b = −156.29) were significantly negative across dyad members, t(57) = −5.13, p < .001. Including all trials, for each dyad, we derived the joint decisions advised by the MRTS algorithm. As described above, we could infer which dyad member made the faster decision on each trial. The programming error meant that we could not test the effect of normalising reaction times. Computing sensitivity ~~~~~~~~~~~~~~~~~~~~~ Psychometric functions were created for each dyad member and for the dyad (empirical or dummy) by plotting the proportion of trials in which the target was reported to be in the second interval against the contrast difference between the two intervals (i.e., the contrast level in the second interval minus the contrast level in the first interval at the target location). The steepness of the slope provides an estimate of sensitivity. More sensitive observers were, by definition, more ‘reliable’ in their estimates of contrast. See Fig. 2 for an example psychometric function. We computed the similarity of dyad members’ reliabilities as the ratio of the sensitivity of the worse dyad member to that of the better dyad member (Smin/Smax), with values near zero corresponding to dyad members with very different reliabilities and values near one corresponding to dyad members of nearly equal reliability. We computed the collective outcome of the dyad (empirical or dummy) as the ratio of the sensitivity of the dyad (empirical or dummy) to that of the more sensitive dyad member (Semp/Smax, SMCS/Smax, and SMRTS/Smax), with values below 1 indicating a collective loss and values above 1 indicating a collective benefit. Normalised versus raw confidence estimates ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To test the effect of normalising confidence estimates before selecting the decision made with higher confidence, we also submitted raw confidence estimates to the MCS algorithm. The trials in which dyad members reported the same level of confidence but selected different intervals were resolved by randomly selecting the decision of one of the two dyad members; we note that there were no such confidence ties when submitting normalised confidence estimates to the MCS algorithm. To ensure that the random selection did not favour one of the dyad members by chance, for each dyad, we generated one hundred dummy dyads using raw confidence estimates and used their mean sensitivity (SrawMCS) to test the effect of normalising confidence estimates (SMCS/SrawMCS). See table 1 for a summary of the strategies for collective choice.","As none of our interest measures showed significant differences between the non-verbal and the verbal conditions, we only report results based on data collapsed across both conditions. In addition, as none of these measures changed over time (i.e., from the first to the second session of the experiment), we only reported data collapsed across both experimental sessions. The efficacy of the algorithms ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The similarity of dyad members’ sensitivities (Smin/Smax) significantly predicted the collective outcome of the MCS algorithm (SMCS/Smax), b = .68, t(27) = 5.17, p < .001, and explained around 50% of the variance, R2 = 49.8, F1,28 = 26.74, p < .001 (see Fig. 4A). For similar dyad members, the MCS algorithm yielded a collective benefit (SMCS/Smax > 1 when Smin/Smax > 0.6). However, for dissimilar dyad members, the MCS algorithm yielded a collective loss (SMCS/Smax < 1 when Smin/Smax < 0.6). The similarity of dyad members’ sensitivities also significantly predicted the collective outcome of the MRTS algorithm (SMRTS/Smax), b = .59, t(27) = 4.69, p < .001, and explained around 45% of the variance, R2 = 44.9, F1,28 = 22.02, p < .001 (see Fig. 4B). However, the MRTS only yielded a collective benefit for five dyads. The MCS algorithm was superior to the MRTS algorithm, both when normalised confidence estimates, t(28) = 4.99, p < .001, and raw confidence estimates, t(28) = 6.26, p < .001, were used as its input. The relative benefit of interaction over the algorithms ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To test whether interacting dyad members used the credibility of each other’s confidence estimates to guide their joint decisions, we regressed the fraction of disagreement trials in which the dyad eventually followed the decision of dyad member A instead of that of dyad member B (the choice ratio) against the ratio of the AROC for dyad member A relative to that of dyad member B (the AROC ratio). The AROC ratio significantly predicted the choice ratio, b = .70, t(27) = 3.58, p < .001, and explained around 30% of the variance, R2 = .32, F1,28 = 12.78, p < .001 (see Fig. 5). The similarity of dyad members’ sensitivities significantly predicted the performance of the empirical dyads relative to that of the MCS algorithm (Semp/SMCS), b = −.43, t(27) = −4.21, p < .001, and explained around 40% of the variance, R2 = .40, F1,28 = 17.73, p < .001 (see Fig. 6A). For similar dyad members, there was no relative benefit for interaction over the MCS algorithm (Semp/SMCS ≈ 1 when Smin/Smax > 0.8). However, for dissimilar dyad members, there was a relative benefit for interaction over the MCS algorithm (Semp/SMCS > 1 when Smin/Smax < 0.8). The similarity of dyad members’ sensitivities also significantly predicted the performance of the empirical dyads relative to that of the MRTS algorithm (Semp/SMCS), b = −.651, t(27) = −2.20 p = .037, and explained around 15% of the variance, R2 = .15, F1,28 = 4.83, p = .037 (see Fig. 6B). However, the MRTS algorithm only (marginally) outperformed two of the empirical dyads. The effect of normalising confidence estimates ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The similarity of dyad members’ sensitivities significantly predicted the effect of normalising confidence estimates (SMCS/SrawMCS), b = .17, t(27) = 2.70, p = .012, and explained around 20% of the variance, R2 = .21, F1,28 = 7.31, p = .012 (see Fig. 7). For similar dyad members, there was a relative benefit for normalising confidence estimates (SMCS/SrawMCS > 1 when Smin/Smax > 0.6). However, for dissimilar dyad members, there was a relative cost to normalising confidence estimates (SMCS/SrawMCS < 1 when Smin/Smax < 0.6). We note that this effect was relatively subtle (cf. change of scale on y-axis in Fig. 7), and that the relative benefit of interaction over the MCS algorithm was not affected using raw, instead of normalised confidence estimates, as input to the MCS algorithm.","Sharing of confidence as a strategy for combining individual opinions into group decisions has been established in a wide range of contexts. This tendency to evaluate the reliability of information by the confidence with which it is expressed has been termed the ‘confidence heuristic’. In this study, we tested two simple ways of implementing the confidence heuristic in the context of a collective perceptual decision-making task: the MCS algorithm, which opts for the decision made with higher confidence, and the MRTS algorithm, which opts for the faster decision, exploiting a negative correlation between confidence and reaction time. Our findings have important implications for the use of heuristics for collective choice and for models of confidence and collective decision- making. The efficacy of the MCS and the MRTS algorithms ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As expected, the MCS and the MRTS algorithms only yielded 2HBT1 effects for dyad members of nearly equal reliability. However, despite a negative correlation between confidence and reaction time, the MCS algorithm markedly outperformed the MRTS algorithm. The superiority of the MCS algorithm to the MRTS algorithm could be due to differences in the ‘pre-processing’ of their input. While we could submit both normalised and raw confidence estimates to the MCS algorithm, we could only submit raw reaction times to the MRTS algorithm because of a programming error (see Section 3.2). Since individuals vary with respect to their average response speed, pooling raw reaction times from two individuals might corrupt the link between confidence and reaction time. However, the MCS algorithm outperformed the MRTS algorithm even when raw confidence estimates were used as its input, suggesting that reaction time may be a very noisy substitute for confidence. The sign of the correlation between confidence and reaction time has been found to depend on response demands, with a negative correlation in the absence of speed pressure and a positive correlation under speed pressure (see Pleskac & Busemeyer, 2010, for a computational account of this phenomenon). While no speed pressure was enforced in the current task, dyad members may have paced their responses on a subset of trials, thus corrupting the negative correlation between confidence and reaction time. We note that, outside the context of the MCS and the MRTS algorithms, the normalisation of confidence estimates has a more straightforward interpretation than the normalisation of reaction times. The normalisation of confidence estimates is intended to remove biases in the use of a scale – here, how an internal variable is mapped onto a confidence scale – and could potentially capture important aspects of collective decision-making (see Section 5.3). It is less clear how the normalisation of reaction times would translate into other contexts. Taken together, our findings show that heuristics for collective choice are susceptible to individual differences in reliability, and suggests that reaction time cannot be substituted for confidence without incurring a considerable collective accuracy cost. In this light, we will limit the remainder of the Discussion to the MCS algorithm. The relative benefit of interaction over the MCS algorithm ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ If the assumptions of the WCS model were satisfied, the responses advised by the MCS algorithm should be just as accurate as those reached by dyad members through interaction. However, the ability to estimate the reliability of one’s own decisions (cf. the level of noise in one’s perceptual system) shows substantial individual differences (e.g., Fleming et al., 2010; Song et al., 2011); this ability is typically referred to as metacognitive ability and quantified as metacognitive accuracy. While the MCS algorithm is prone to error when people misestimate the reliability of their own decisions, interacting individuals may take such misestimates into account. For example, one study has shown that mock jurors find witnesses who are confident about erroneous testimony less credible than witnesses who are not confident about it (Tenney et al., 2007). We predicted that interacting dyad members would take into account the credibility of each other’s confidence estimates when making their joint decisions, and that interaction would be relatively more beneficial than the algorithms for dissimilar dyad members, because they have more to lose from following the more confident but less competent of the two. As for the first prediction, the fraction of disagreement trials in which the dyad followed dyad member A instead of dyad member B depended on their relative metacognitive accuracy (here measured as AROC – see Section 3.4), indicating that dyad members used the credibility of each other’s confidence estimates to guide their joint decisions (see Fig. 5). As for the second prediction, interaction was more robust than the MCS algorithm to differences in reliability. For similar dyad members, the decisions reached through interaction were no more accurate than those advised by the MCS algorithm. However, for dissimilar dyad members, the decisions reached through interaction were considerably more accurate than those advised by the MCS algorithm; this was true irrespective of whether normalised or raw confidence estimates were submitted to the MCS algorithm. While models of collective decision-making have identified the ‘arbitration’ of confidence estimates as key to collective performance (e.g., Bahrami et al., 2010; Koriat, 2012), our findings suggests that the ‘weighting’ of confidence estimates is equally important for collective performance. Without taking the credibility of confidence estimates into account, the MCS algorithm cannot fully replace interaction in collective decision-making. Our study highlights the social heterogeneity of credibility as an interesting avenue for computational research: how do we estimate the credibility of each other’s opinions, and how good are we at doing so? We note that the relative benefit of interaction over the MCS algorithm need not necessarily result from dyad members discounting the opinion of the dyad member with lower metacognitive accuracy. More specifically, while dyad members may assign less weight to the opinion of the more confident but less competent member (“bad but doesn’t know it”), they may also assign more weight to the opinion of the less confident but more competent dyad member (“good but does know it”). We believe that computational models of social learning (e.g., Behrens, Hunt, Woolrich, & Rushworth, 2008) are needed to tease apart such decision strategies. Role of feedback All participants received feedback about the accuracy of each decision, and could thus directly evaluate the credibility of each other’s confidence estimates. Indeed, previous research has shown that diagnostic feedback helps groups of individuals to identify their more accurate members (Henry, Strickland, Yorges, & Ladd, 1996). The relative benefit of interaction over the MCS algorithm might therefore not persist in the absence of diagnostic feedback. However, using the same visual perceptual task, Bahrami et al. (2012b) found that diagnostic feedback was not necessary for the accumulation of a 2HBT1 effect – diagnostic feedback only appeared to accelerate the process – indicating that individuals may rely on other signals when they learn the credibility of each other’s confidence estimates. The identification and incorporation of these signals will pose a major challenge for dynamic models of collective decision-making. Role of familiarity Here, for each dyad, one participant was recruited, and then asked to bring along a friend to the study. Dyad members might thus have used their interpersonal history to establish the credibility of each other’s confidence estimates. However, using a similar task, but in the domain of approximate numeration, Bahrami, Didino, Frith, Butterworth, and Rees (2013) found that familiarity had little impact on collective performance, suggesting the dyad members evaluated the credibility of each other’s confidence estimates in the context of their current task performance. While there is no evidence for a main effect of familiarity on collective performance, it may be the case that familiarity matters more for dissimilar than similar dyad members. Non-perceptual domains Research has shown that individuals are ‘overconfident’ about the accuracy of their knowledge-based judgements but ‘underconfident’ about the accuracy of their perceptual judgements (see Harvey, 1997, for a review). These discrepant patterns of confidence have led to the hypothesis that different types of information determine confidence in knowledge and perception. For example, Juslin and Olsson (1997) propose a model of confidence in which perceptual judgements are dominated by ‘Thurstonian’ uncertainty, internal noise such as stochastic variance in the sensory systems, whereas knowledge-based judgments are dominated by ‘Brunswikian’ uncertainty, external noise such as less-than-perfect correlations between features in the environment. This dissociation raises issues as to whether confidence can be used as a proxy for the reliability of decisions in non- perceptual domains. However, direct comparison of knowledge-based and perceptual judgements has found evidence for a common basis of confidence (e.g., Baranski & Petrusic, 1994; Pallier et al., 2002), suggesting that the efficacy of the MCS algorithm will generalise to non-perceptual domains (but see Koriat, 2012, for exceptional environments).","This work was supported by the Calleva Research Centre for Evolution and Human Sciences (DB, JYFL), the European Research Council Starting Grant NeuroCoDec 309865 (BB), the Danish Council for Independent Research – Humanities (KT, RF), the EU-ESF program Digging the Roots of Understanding DRUST(KT, RF), the European UnionMindBridge Project (DB, KO, AR, CDF, BB), the Gatsby Charitable Foundation (PEL), and the Wellcome Trust (GR)."],["Noise has repeatedly been shown to be one of the most recurrent reasons for complaints in open-plan office environments. The aim of the present study was to investigate if enhanced or worsened sound absorption in open-plan offices is reflected in the employees' ratings of disturbances, cognitive stress, and professional efficacy. Employees working on two different floors of an office building were followed as three manipulations were made in room acoustics on each of the two floors by means of less or more absorbing tiles & wall absorbents. For one of the floors, the manipulations were from better to worse to better acoustical conditions, while for the other the manipulations were worse to better to worse. The acoustical effects of these manipulations were assessed according to the new ISO-standard (ISO-3382-3, 2012) for open-plan rooms acoustics. In addition, the employees responded to questionnaires after each change. Our analyses showed that within each floor enhanced acoustical conditions were associated with lower perceived disturbances and cognitive stress. There were no effects on professional efficiency. The results furthermore suggest that even a small deterioration in acoustical room properties measured according to the new ISO-standard for open-plan office acoustics has a negative impact on self-rated health and disturbances. This study supports previous studies demonstrating the importance of acoustics in work environments and shows that the measures suggested in the new ISO-standard can be used to adequately differentiate between better and worse room acoustics in open plan offices. --------------------------------------------------------------------------------","In relation to other ambient factors, the impact of unwanted sound or noise is probably the most studied when it comes to office environments (Boyce, 1974; De Croon, Sluiter, Kuijer, & Frings-Dresen, 2005; Leather, Beale, & Sullivan, 2003; Leder, Newsham, Veitch, Mancini, & Charles, 2015; Navai & Veitch, 2003; Nemecek & Grandjean, 1973; Pejtersen, Allermann, Kristensen, & Poulsen, 2006; Sundstrom, Burt, & Kamp, 1980; Sundstrom, Town, Rice, Osborn, & Brill, 1994; Veitch, Charles, Farley, & Newsham, 2007; Veitch, Farley, & Newsham, 2002; Warnock, 2004). Noise has been suggested to cause interruption, irritation and lowered performance among employees (Roelofsen, 2008), and is one of the most common reasons for complaints in open-plan office environments (Kaarlela-Tuomaala, Helenius, Keskinen, & Hongisto, 2009). However, this study addresses something that is less known about noise, namely, how better or worse acoustical conditions in open-plan offices affect employees' perception of disturbances, cognitive stress, and professional efficacy. Why noise is a common reason for complaints can be explained by the changing state hypothesis (Jones, Madden, & Miles, 1992), which suggests that sounds varying over time cause more disruptions. A sound that is constant in intensity or timbre should therefore cause fewer disturbances than sounds that constantly change their characteristics. A more uniform sound source can be created by filtering out high frequency sound, so called low-pass filtering (Jones, Alford, Macken, Banbury, & Tremblay, 2000) or by introducing new sources of sound, which either can be competing voices (babble-effect) or speech neutral masking noises, e.g. from ventilation (Loewen & Suedfeld, 1992). Increasing the number of sounds beyond a critical level causes the overall degree of variability in sound to drop, hence the overall result is a more even sound level where peaks and troughs from individual sound sources are cancelled out (Perham, Banbury, & Jones, 2007). The degree of variability might also be expected to drop when reverberation time increases. For example, Beaman and Holt (2007) found that a reverberation time, i.e. the time it takes for sound to attenuate, of 5 s led to the same low amount of error in conducting an immediate recall test as in the quiet control condition. However, in an office environment the reverberation time seldom approaches 5 s but varies in lower ranges (between 0.4 and 1 s). Perham et al. (2007) investigated if more realistic differences in reverberation time can affect performance on a cognitive test measuring serial recall. They compared one quiet condition with two different noisy conditions. The two noisy conditions were comprised of noise from various sources in an office recorded in a room with a reverberation time of either 0.7 or 0.9 s. The respondents conducted the test while listening to the noises through headphones. Although they found an effect on performance between the quiet condition, where no noise was played, and the two noisy conditions, performance on the test did not differ between the two noisy conditions. Further analyses revealed that speech intelligibility did not differ between the two noisy conditions, and the authors concluded that “at least for typical office reverberation times, lower reverberation times do not increase intelligibility” (Perham et al., 2007, p. 843). It has also been found that different noise types, for example speech, music, and office noise in general, in comparison with quiet conditions, negatively impact different cognitive outcomes, such as memory performance, reading comprehension, and proofreading (see Hongisto, 2005 for an overview). Noise has also been extensively studied in field studies. Ringing telephones, air conditioning, and office machinery have all been suggested to cause disturbances in office environments. Human speech (Boyce, 1974; Pierrette, Parizet, Chevret, & Chatillon, 2014; Sundstrom et al., 1994) and its intelligibility is another common distracting factor. It is measured by the Speech Transmission Index (STI), which ranges from 0, meaning that the speech is not understandable, to 1, meaning that the speech is fully comprehendible. When STI exceeds 0.2 it begins to cause a decrease in performance with the highest decrement occurring around 0.6 (Hongisto, 2005). Furthermore, field studies also show that distractions and noise are present also in cell offices (Seddigh et al., 2015), even if open-plan office environments usually are associated with more noise and distractions (Kaarlela-Tuomaala et al., 2009; Seddigh, Berntson, Bodin Danielson, & Westerlund, 2014). Consequently, it would be more relevant to investigate the impact of different sound intensities or certain aspects of noise rather than comparing its presence with absence. In addition, another study by Pierrette et al. (2014) could not find any association between the A-weighted sound pressure level dBA (LeqA) and the perception of noise in the office as high or annoying. The authors emphasised the relevance of measuring behavioural outcomes to appraise the appropriateness of the noise in open-plan office environment instead of relying overly much on objective acoustical measures. This conclusion corresponds well with the definition of noise not as the particular type or magnitude of the sound, but rather as the perception of the sound by the listener, i.e. to what extent the sound is experienced as noise (Roelofsen, 2008). Additionally, knowledge workers – that is workers who create, develop, manipulate, disseminate or use knowledge to provide an outcome – depend to high degree upon processing information (Bosch-Sijtsema, Ruohomäki, & Vartiainen, 2010; Janz, Colquitt, & NOE, 1997). According to the Load theory of selective attention and cognitive control (Lavie, Hirst, de Fockert, & Viding, 2004) unwanted stimuli such as noise need to be first processed and then actively inhibited in order to not distract the person who is exposed to noise. Therefore, for knowledge workers noise competes for the same cognitive capacities that process task related information (see also Diamond, 2013; Seddigh et al., 2014). Hence, lower in comparison to higher noise levels in office environments should lead to fewer problems for knowledge workers. Furthermore, another relevant theory concerning supportive design (Ulrich, 1991) suggest that while certain physical characteristics may not affect employees negatively per se, they may intensify the negative impact of some other factor in the environment (Evans, 2001). Leather et al. (2003) found such effect and reported that high noise together with high job strain, in contrast to low job strain, was associated with lower job satisfaction, lower organisational commitment and increased rate of symptoms of infectious diseases. Low noise regardless of the level of job strain did not have a large effect on these measures. A comparable suggestion to the interaction of noise level and job strain can be made for the joint effect of open-plan office environments and noise levels. That is, even if the open-plan office environments per se do not affect employee health and performance, bad acoustical conditions in these environments might. It is important to investigate the total acoustical condition in the office rather than focussing on any single aspect that may affect the acoustical condition. Namely even if wall panels can affect the acoustical condition in an office environment, in research settings it is important to focus on the actual acoustical condition in the office instead of the presence or absence of panels per se. In fact in a recent study Leder et al. (2015) found that larger workstations in open-plan offices were associated with greater satisfaction with privacy, however the degree of enclosure of the workstation by partial-height partitions was not associated with the same outcome measure. Furthermore, in order to more thoroughly understand the impact of noise on office workers health and performance, different types of measures should be used. Except behavioural outcomes, we believe that a more comprehensive mapping of the objective sound environment, rather than too much reliance on a single sound measure, could give a more extensive understanding of how objectively measured sound is associated with the perception of noise. This idea is in fact raised in the International Standard of room acoustic parameters (ISO-3382-3, 2012), which suggests that rather than relying too much on single measures, such as reverberation time, a combination of measures including STI and background noise levels should be focused on in order to receive a more complete evaluation. Hence, the purpose of the present study is to test the effect of different acoustical environments on employee ratings on indicators of disturbances, health, and performance. This is done by a crossover design that compares two different types of sound absorbents installed in contrasting sequences on two similar floors within the same office building. In order to obtain a comprehensive understanding of the room acoustics, we collected objective acoustical data in accordance with the international standard regarding room acoustics parameters (ISO-3382-3, 2012). We also collected behavioural measures, in order to understand how the acoustical environment impacts on the employees. Aims and hypotheses ~~~~~~~~~~~~~~~~~~~ In this study the aim was to investigate if enhanced or worsened room acoustic characteristics in open-plan office environments are reflected in changes in the employees' own perception of disturbances, health and/or performance. The manipulation consisted of different acoustic elements in the office building, where one condition enhanced the acoustic environment (better condition) and one worsened the acoustic environment (worse condition) as compared with a baseline condition. Disturbances are examined from three different perspectives. Firstly, a broad measure of disturbances was used in order to understand the impact of our manipulation of the environment on disturbances in general. The second and third measure of disturbances focused on disturbances from sounds from nearby and distant sources, respectively. Apart from the self-rated measures, we also used objective acoustical measures. The respondents' perception of the environment is followed over three time-points (T1, T2, and T3). Our overall hypothesis was that the acoustical conditions would have an impact on the respondents' experiences regarding the outcome variables, that is within each floor: the better condition is associated with lower disturbances in general, the better condition is associated with lower nearby disturbances, the better condition is associated with lower distant disturbances, the better condition is associated with lower cognitive stress, the better condition is associated with higher professional efficiency. Participating organization and employees ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Two months before the study started the organisation had moved from an old office building with mostly private office rooms to a new renovated building with 6 floors, which was where this study was conducted. Two out of the six floors were used for the study (floors 4 and 5) as they had identical layouts, were similarly furnished, and the employees on these floors had similar work assignments. Each floor was highly open, with limited or no partitions, carpeted and with ceilings furbished with highly sound absorbent tiles. Each employee had his/her own designated desk. The sample consisted of 151 employees in a municipality office outside of Stockholm, Sweden. The improvements in acoustics were partly made through installation of wall absorbents. However, the wall absorbers had not been aired before they were set up and during the initial days three employees felt irritation in the form of smell and headaches. These three employees were excluded from the analyses. A number of randomly selected employees were asked if they had noticed any smell or symptoms, but no one else had. During data collection a fourth employee received a screen that would protect against glare. Because her acoustic environment may have changed due to the screen she was also excluded from the analyses. Further, two employees on the managerial level had been informed about the design of the study and were also excluded. After exclusion of these individuals the sample size consisted of 145 persons. 77% (n = 117) of the total sample completed the baseline survey in its entirety (T0), 70% (n = 106) the first survey (T1), 62% (n = 94) the second (T2), and 64% (n = 97) the third (T3). In total around 40 individuals had a full set of data for T1, T2 and T3. Study design and procedure ~~~~~~~~~~~~~~~~~~~~~~~~~~ This study employed a crossover design in an office environment to investigate if enhanced and worsened acoustical environment impact employees' perception of disturbances, self- rated health and performance. Before data was collected the employees were invited to a meeting where they were informed about the purpose and procedure of the study. They were told that four electronic surveys would be sent out and that during the total time of the study, we might change the acoustics of the office several times. They were also told that at the end of the study they would be given full information about the results of the study and what changes we had made. The baseline survey was collected just before any manipulations were made to the office environment. Each manipulation resulted in one of two conditions: In the better condition, sound absorbing wall panels were set up and the pre-existing, highly sound absorbent ceiling tiles were kept. In the so called worse condition, there were no sound absorbing wall panels and highly sound reflective ceiling tiles were installed, replacing 55% of the original highly absorbent tiles. Both types of tiles had similar colour and form and could not easily be distinguished from each other (See Zalyaletdinov, 2014 for the full acoustical report). During the weekend after the baseline survey (T0) had been collected, changes were made on floor 4 to create the better condition, and on floor 5 to create the worse condition. Two weeks after the first manipulations had been made the first survey was sent out. The surveys were always sent out on Mondays. During the weekend after, floor 4 was changed to create the worse condition and vice versa. After two weeks of exposure to the new conditions, the second survey was sent out. The three weeks following after the second survey contained many national holidays. In order to ensure that most employees had been exposed to the sound environment for two full weeks, the third survey were sent out six weeks after the second survey had been completed (see Fig. 1). Survey measures ~~~~~~~~~~~~~~~ All respondent data was collected by means of an electronic survey. Disruption in general was measured by four items. The questions were “To what extent have you in the past seven days been disturbed by ventilation sounds”; “… by sounds from computers”; “… by ringing phones”; and “… by colleagues' phone calls”. All questions concerning disruptions were measured by using a five-point rating scale (1 = “to a small extent”, 5 = “to great extent”). Cronbach's α for internal reliability from the first survey was 0.71, indicating satisfactory consistency. Nearby disturbances were measured by the question “To what extent have you in the past seven days been disturbed by speech and laughter from colleagues sitting near you (within a radius of 10 m)”. Distant disturbances were measured by the question “To what extent have you in the past seven days been disturbed by speech and laughter from colleagues who sit further away (beyond a radius of 10 m)”. Cognitive stress was measured by the cognitive stress scale (4 items) from the Swedish version of the Copenhagen Psychosocial Questionnaire (COPSOQ) (Kristensen, Hannerz, Høgh, & Borg, 2005). Sample question: How much of the time during the past week have you found it difficult to think clearly? All items were scored on a 5-point rating scale (1 = never, 5 = always). Cronbach's α for internal reliability from the first survey was 0.88, indicating satisfactory consistency. The professional efficacy subscale (6 items) of the Swedish version of the Maslach Burnout Inventory – General Survey (MBI-GS) was used to assess self-rated performance (Schutte, Toppinen, Kalimo, & Schaufeli, 2000). All items were scored on a 7-point rating scale (ranging from 1 = never, 7 = daily). Cronbach's α for internal reliability from the first survey was 0.85, indicating satisfactory consistency. See Table 1 for a correlation matrix between the dependent variables at T0. The covariates included in the model were age (continuous: ranging from 21 to 69), gender (0 = male, 1 = female), and educational level (dichotomized: 0 = low for those without an academic degree, 1 = high for those with an academic degree; see Tables 2a and 2b). Acoustic measurements ~~~~~~~~~~~~~~~~~~~~~ We included several acoustical measures in accordance with ISO 3382-3 guidelines (ISO-3382-3, 2012). These are D2,s, Lp,A,S,4 m, and radius of comfort (rc). D2,s is a rate of spatial decay of A-weighted sound pressure level of speech per distance doubling. D2,s is therefore a measure of how fast the decibel level has been attenuated at a certain point from the sound source. Lp,A,S,4 m is a nominal A-weighted sound pressure level of normal speech at a distance of 4.0 m from the sound source. In other words Lp,A,S,4 m shows how much normal speech sound has been attenuated at a distance of 4 m from the sound source. Radius of comfort (rc) is the distance from the sound source where the sound pressure level of speech meets 48 dBA, which is the targeted value of Lp,A,S,4 m according to (ISO-3382-3, 2012). The radius of comfort formula was suggested at EuroNoise 2012 with background from the field study report made by Nordic Innovation (Hellström & Nilsson, 2010). The formula for calculating rc is rc = 4 × 100.3(Lp,A,S,4m−Lc)/D2,s. These measurements were carried out for each condition in furnished rooms without staff along four measurement paths, two paths on each floor (please see Appendix for a more detailed description of the objective measurements). In addition, dBA levels were recorded from four points by two microphones on each floor. Point 1 and 2 on floor 5 and point 3 and 4 on floor 4. These microphones registered the equivalent dBA for every 30 min interval from 06.30 until 18.00 h. Our intention was to register dBA levels for the total period, however, technical difficulties prohibited us from collecting data at point 1 during the first period and at point 2 during the third period. The length of the data collection for the dBA levels were 11 days for the first period, 4 complete days for the second period, and 15 days for the third period. Weekends were not included in the analysis. At each point for each period, an equivalent dBA was calculated for every past half an hour starting from the first registration at 7 AM to the last at 6 PM. All objective acoustical data were gathered in order to confirm that the manipulations we had made to the physical environment had led to two distinguishable acoustical conditions on each floor. The acoustic conditions for each path were the same for T1 and T3, hence the objective measures concerning D2,s, Lp,A,S,4 m, and rc measured at T1 were assumed being the same at T3. Data analyses ~~~~~~~~~~~~~ 39 employees had at one or several time points worked in the open-plan office environment less than 15 h the latest 7 weekdays before answering the survey. The answers of these employees at the specific time point(s) were removed. That is, if an employee had worked less than 15 h the latest 7 days before answering the survey at T1 and T3 but more than 15 h the latest 7 days before answering the survey at T2, the responses at T1 and T3 were removed while the responses at T2 were kept. First, five 3 × 2 repeated ANCOVA analyses were carried out for each of the five outcome variables for T1, T2, and T3 in order to test if the different order of the better versus worse conditions generated a different development of the outcome measures over time. By investigating if the quadratic function of time and floor was significant, the repeated ANCOVAs test if the repeated manipulations to the different floors affected the outcome measures in the supposed direction. That is, the exposure for each floor either went from better to worse to better (floor 4), or from worse to better to worse (floor 5) which was hypothesised to yield approximately symmetrically different U-shape curves of the outcome variables for the two floors that significantly differed in their direction. Employees on floor 4 should rate Disruption in general, Nearby disturbances, Distant disturbances, and Cognitive stress as low (in the better condition) – high (worse condition) – low (better condition) creating a ∩-shaped pattern, while employees on floor 5 should rate the same outcomes as high (worse condition) – low (better condition) – high (worse condition) creating a U-shaped pattern. For professional efficiency employees on floor 4 should rate as high (better condition) – low (worse condition) – high (better condition) creating a U-shaped pattern, while employees on floor 5 should rate professional efficiency as low (worse condition) – high (better condition) – low (worse condition) creating a ∩-shaped pattern. A significant quadratic function of time and floor would mean that the better and worse conditions affected the employees according to intentions, which will allows us to conduct further analyses to test if the manipulations between the better and the worse conditions differed meaningfully within each floor. Second, on floor 4 and 5 separate repeated ANCOVAs were carried out. For floor 4 these tested if the first better condition (T1) significantly differed from the worse condition (T2) (the first contrast analysis) and if the second better condition (T3) significantly differed from the worse condition (T2) (the second contrast analysis). For floor 5 these tested if the better condition (T2) differed significantly from the first worse condition (T1) (the third contrast analysis), and if the better condition (T2) differed significantly from the second worse condition (T3) (the fourth contrast analysis). These additional analyses were carried out only for the outcomes that were significant in the first set of analyses investigating the quadratic interaction effect of time and floor on the outcomes. The analyses were conducted in SPSS version 21 by means of the General Linear Model. Sex, age, and educational level were included as covariates. The repeated ANCOVA analyses rely on non-missing-data for each respondent for T1–T3. Hence, for each ANCOVA analyses missing answers at T1, T2 and/or T3 for each outcome variable lead to case-wise deletion.","The difference between the better and the worse acoustical condition for the active parts of the working days and for each floor are shown in Fig. 2, which illustrates that in general throughout the days during data collection, both floors had a lower dBA level during the better condition in comparison to the worse. Floor 5 had a larger variation than floor 4. The figure also shows a trend that the dBA levels seem to increase from morning to the late afternoon. For the other objective measures please see Table 3. According to expectations, and as shown in Table 3, the condition with both absorbing tiles and wall absorbents, absorbed noise better than the condition with reflective tiles and no wall absorbents according to the latest ISO standard. Disruption in general ~~~~~~~~~~~~~~~~~~~~~ According to Wilks' criterion there were no significant main effects of time or floor. The interaction effects between time and the covariates were not significant. The time and floor interaction was significant for the hypothesised quadratic function (F[1, 38] = 7.29, p = 0.01, partial η 2 = 0.16). The manipulations on each floor yielded symmetrically different U-shaped curves for disruption in general which suggested lower disturbances in the better conditions in comparison to the worse. Contrast analyses comparing the conditions within each floor where carried out to test the first hypothesis. On floor 4 the change from the better (T1) to the worse (T2) condition was significant while the change from the worse (T2) to the better (T3) condition was not. On floor 5 the change between the worse (T1) to the better (T2) condition was not significant but the change between the better condition (T2) to the worse (T3) was significant (all p < 0.05; please see Fig. 3a). To conclude, the first hypothesis was supported in that the better acoustical condition is related to less reported disturbances in general. Nearby disturbances ~~~~~~~~~~~~~~~~~~~ With the use of Wilks' criterion there was no significant main effect of time or floor. The interaction effects between time and the covariates were not significant. The time and floor interaction was significant between time and floor for the hypothesised quadratic function (F[1, 40] = 16.69, p < 0.001, partial η 2 = 0.29). The manipulations on each floor yielded symmetrically different U-shaped curves for nearby disturbances, which suggested lower disturbances in the better conditions in comparison to the worse. Contrast analyses comparing the conditions within each floor showed that on floor 4 the change from the better (T1) to the worse (T2) condition was significant while the change from the worse (T2) to the better (T3) condition was not. On floor 5 the change between the worse (T1) to the better (T2) condition was significant which also was the change between the better condition (T2) to the worse (T3) (all p < 0.05; please see Fig. 3b). To conclude the second hypothesis was supported in that the better acoustical condition is related to lower reported nearby disturbances. Distant disturbances ~~~~~~~~~~~~~~~~~~~~ With the use of Wilks' criterion there was no significant main effect of time or floor. The interaction effects between time and the covariates were not significant. The time and floor interaction was significant between time and floor for the hypothesised quadratic function (F[1, 40] = 5.42, p = 0.025, partial η 2 = 0.12). The manipulations on each floor yielded symmetrically different U-shaped curves for distant disruption suggested lower disturbances in the better conditions in comparison to the worse. Contrast analyses comparing the conditions within each floor where carried out to test the first hypothesis. On floor 4 neither the change from the better (T1) to the worse (T2) condition nor the change from the worse (T2) to the better (T3) condition was significant. On floor 5 the change between the worse (T1) to the better (T2) condition was not significant but the change between the better condition (T2) to the worse (T3) was significant (all p < 0.05; please see Fig. 3c). To conclude, the third hypothesis was supported in that the better acoustical condition is related to less reported disturbances from distant sources. Cognitive stress ~~~~~~~~~~~~~~~~ With the use of Wilks' criterion there was no significant main effect of floor. However, the main effect of time was significant (F[2, 36] = 3.48, p = 0.042, partial η 2 = 0.16). The interaction effect between time and covariates were not significant. The time and floor interaction was significant between time and floor for the hypothesised quadratic function (F[1, 37] = 7.59 p = 0.009, partial η 2 = 0.17). The manipulations on each floor yielded symmetrically different U-shaped curves for cognitive stress, which suggested lower stress in the better conditions in comparison to the worse. Contrast analyses comparing the conditions within each floor where carried out to test the fourth hypothesis. On floor 4 neither the change from the better (T1) to the worse (T2) condition nor the change from the worse (T2) to the better (T3) condition were significant. On the other hand on floor 5 both the change between the worse (T1) to the better (T2) condition and the change between the better condition (T2) to the worse were significant (all p < 0.05; please see Fig. 3d). To conclude, the fourth hypothesis was supported in that the better acoustical condition is related to less cognitive stress. Professional efficiency ~~~~~~~~~~~~~~~~~~~~~~~ With the use of Wilks' criterion there was no significant main effects of floor or time. The interaction effects between time and the covariates were not significant. Further, the hypothesized quadratic function between time and floor was not significant (see Fig. 3e), meaning that the employees on each floor did not report significantly higher or lower efficiency depending on the different conditions. Given that the overall quadratic function of time and floor was not significant, no further analyses within each floor were carried out. Therefore the fifth hypothesis could not be supported (see Fig. 3e).","This study investigated if better and worse acoustic environments, created by less or more absorbing tiles and wall absorbents, affect employees' perception of disturbances, cognitive stress, and professional efficacy. In line with our expectations, the acoustical measures showed a lower overall noise level during the working day and also lower D2,s, Lp,A,S,4 m, and rc in conditions where tiles and wall panels that absorbed more sound energy had been installed. In addition, also supporting our expectations, the acoustical measures showed a higher overall noise level during the working day and higher D2,s, Lp,A,S,4 m, and rc in conditions where the more reflective tiles where installed and wall panels removed. Our results are in line with previous studies (Kaarlela-Tuomaala et al., 2009) and suggest that employees' perception of disturbances and health are affected negatively when exposed to increased noise levels. However, in contrast to previous research findings (Perham et al., 2007; Pierrette et al., 2014), the results from the present study showed that improved room acoustics was associated not only to lower objective noise levels, but also to lower perceived disturbances and lower cognitive stress. Consequently, the results imply that employees perceived better possibilities to make decisions, concentrate, and reported having lower amount of memory loss. These findings may be explained by that decreased general noise level in the office environment decrease the interference of noise on higher cognitive functions, which are important for knowledge workers ability to carry out their tasks (Diamond, 2013; Lavie et al., 2004; Seddigh et al., 2014, 2015). These findings can also be related to the findings of Leather et al. (2003) who reported that high noise levels in contrast to low noise levels interact with job strain and impact employees' job satisfaction, organisational commitment and symptoms of infectious diseases. Hence, the acoustical condition of the open-plan office seems to have a direct relationship to indicators of both health and performance of employees. Interestingly, these effects were evident despite the short exposure time to the new condition, suggesting that the effect of a change in room acoustics is quite immediate. However, the short exposure time might also explain why not all contrast analyses where significant, even if the effects went in the expected direction (i.e. better acoustics leading to less problems, for our measures of disturbances and health). As evident in Fig. 1, the manipulations between the two conditions had a larger impact on floor 5 in forms of differences in dBA-levels. Although the employees on both floors had similar tasks, one possible explanation to these differences was revealed during the feedback session to the employees. Employees on floor 5 had more conversations and meetings around their desks and in the open space compared to employees on floor 4 who either did not have as many meetings or conducted the meetings in separate meeting rooms. Therefore, the acoustical differences between the two acoustical conditions likely had greater consequence for floor 5 than for floor 4, as Fig. 1 illustrates. In fact it is possible to discern a similar pattern in employees' rating by looking at Fig. 3a–e. Also in these figures it is apparent that the differences between the better and worse condition is larger on floor 5 as compared to floor 4, meaning that manipulations on floor 5 had a larger effect on the employees. Nevertheless, also for the significant findings the survey responses have a low variation around the middle of the scale. Therefore it seems that even if noise has a significant impact it does not seem terribly disturbing to most respondents. Given the small variation in the objective measures between the conditions, our results might also indicate that quite large changes to the physical environment are needed to substantially improve the acoustical conditions in the office. In this study we conducted acoustical measurement according to the new ISO-standards for open-plan offices (ISO-3382-3, 2012). With these measurements and radius of comfort (rc) we could find differences which corresponded to the manipulations done to the acoustical environment. Although the acoustical measures showed quite small differences between our two conditions within each floor, a comparison of effect sizes for the sought quadratic function between time and floor reveal small effects approaching medium sized effects (η 2 between 0.15 and 0.29) (Cohen, 1988). This would suggest that even a minor improvement made to room acoustics could impact employees perceived health and disturbances. Strengths and limitations ~~~~~~~~~~~~~~~~~~~~~~~~~ One of the main strengths of this present study is that it was carried out in the field addressing regular office employees. Given that the social and other organisational structure within the organisation had not changed, we believe that our finding is highly relevant for the effect that noise has on employees' perception of cognitive stress and disturbances. Another strength of this study is its crossover design. By having two groups that constantly were exposed to the opposite condition than the other and by changing back and forth between the conditions, we created a highly controlled field experiment increasing the reliability of our findings. In addition, we also gathered objective data. The objective measurements ensured that the manipulations we carried out had an impact on the acoustical environment and further strengthened our findings by corresponding to the survey responds. By so doing we were able to show that improvements in acoustics have a direct impact on measures of both health and disturbances. The ceiling tiles of both conditions looked the same, but the wall absorbent installed during the better condition could have indicated to the respondents sitting near the wall absorbents that an improvement had been made. In turn this signal could have systematically affected the employees to respond more positively when the absorbents were present. Nevertheless, there are some aspects that speak against that our results mainly would be due to such a placebo effect. Before any manipulations were made employees on both floors were asked to answer the survey (T0). At T1 floor 5 had the worse condition – that is the condition without any wall absorbent and with reflective ceiling tiles. As we compare the result of T0 with T1 we see that these employees generally report more problems during T1 (compare Table 2a with Fig. 3a–e or with Table 2b). If our result would be due to placebo we should not have seen such pattern given that no visual manipulations had been made on floor 5 between T0 and T1. On the other hand on floor 4 and during T1 in comparison to T0, the employees rated only small improvements in disturbances and cognitive stress. This corresponds well with the minor improvements measured by the acoustical measurements conducted between T0 and T1. Therefore, all in all it seems that the employees' answers correspond well with the acoustical measurements conducted rather than the presence or absence of the wall absorbents. Another concern that could be raised as a limitation is the short exposure period for each condition before we collected the survey data. If people after some passage of time learn to adapt to an increased noise level then our findings might not be as relevant as they might suggest. However, a study by Banbury and Berry (2005) could not find any lasting habituation to office noise, which speaks against any major adaptation to increased noise levels taking place among employees.","By means of a crossover design, we investigated the effect of two different room acoustics on employees' perception of disturbances, cognitive stress, and professional efficacy. Although the acoustical measurements showed that the manipulations between our two conditions in general were quite small, the better acoustical condition nevertheless had a more positive effect on employees' perception of disturbances and cognitive stress. It was also shown that manipulations in the acoustical environment measured by measurements suggested in ISO 3382-3(ISO-3382-3, 2012) correspond well with employees' self-reported measures of health and disturbances. The study shows the importance of focussing on the acoustical conditions in open-plan offices in order to improve employees' well-being and through means of that also organisational efficiency.","AS designed and made the preparation for the study. AS anchored the project in the participating organisation and also collected the data. AS conducted the analysis and wrote the first and successive draft including introduction, method, result, and discussion sections. HW, FJ, EB, and CBD have contributed with substantial suggestions and comments to the manuscript. AS have changed the manuscript according to relevant suggestions and comments, and prepared the final manuscript. AS submitted the manuscript and is the corresponding author in the review process. HW, FJ, and EB have been supervising AS in every part of the study, contributed to the writing of the paper, and have also been involved in the design process. HW is overall project leader."],["Humans have a capacity to become aware of thoughts and behaviours known as metacognition. Metacognitive efficiency refers to the relationship between subjective reports and objective behaviour. Understanding how this efficiency changes as we age is important because poor metacognition can lead to negative consequences, such as believing one is a good driver despite a recent spate of accidents. We quantified metacognition in two cognitive domains, perception and memory, in healthy adults between 18 and 84. years old, employing measures that dissociate objective task performance from metacognitive efficiency. We identified a marked decrease in perceptual metacognitive efficiency with age and a non-significant decrease in memory metacognitive efficiency. No significant relationship was identified between executive function and metacognition in either domain. Annual decline in metacognitive efficiency after controlling for executive function was ~0.6%. Decreases in metacognitive efficiency may explain why dissociations between behaviour and beliefs become more marked as we age. © 2014 The Authors. --------------------------------------------------------------------------------","Metacognition refers to ‘thinking about thinking’ (Flavell, 1979), or the ability to become aware of thoughts and behaviours. Metacognition is a fundamental aspect of higher cognition in humans (Dunlosky & Metcalfe, 2009) and may support conscious awareness (Koriat, 2007), social interaction (Frith, 2012), and be impaired in neuropsychiatric disorders (David, Bedford, Wiffen, & Gilleen, 2012). Metacognition comprises both monitoring and control (Koriat & Goldsmith, 1996; Nelson, 1996). Intuitively, a person has good metacognitive monitoring if their subjective appraisals track their objective behaviour on a trial-by-trial basis. For example, an individual who is metacognitively efficient will report high confidence when they are objectively correct, and low confidence when they are objectively incorrect; conversely an individual with poor metacognitive efficiency has poor awareness, manifested by subjective reports that are unrelated to task performance. Different cognitive domains give rise to different types of metacognitive judgment. For example, in the perceptual domain one might have variable confidence about seeing or hearing a particular stimulus. Metacognitive judgments about memory may refer to whether one’s recall or recognition is likely to have been correct (a retrospective confidence rating) or whether a stimulus is likely to be remembered or recognised in the future (“judgments of learning”, JOLs; and “feelings of knowing”, FOK). Overall, metacognitive judgments are thought to be based on a combination of perceptual or mnemonic strength and additional analytic factors (Busey, Tunnicliff, Loftus, & Loftus, 2000; Koriat & Goldsmith, 1996; Vickers, 1979). Previous research has established that metacognitive efficiency in different domains can be isolated and studied independently of primary cognitive capacity (see Fleming & Dolan, 2012, for a review). However, there is some debate as to whether metacognition changes as we age. On the one hand, we might expect greater life experience leads to more accurate self-knowledge and greater metacognitive efficiency. On the other hand, convergent evidence has revealed a specific neural basis for metacognitive efficiency in human prefrontal and parietal cortex (Fleming, Huijgen, & Dolan, 2012; Fleming, Weil, Nagy, Dolan, & Rees, 2010; McCurdy et al., 2013; Rounis, Maniscalco, Rothwell, Passingham, & Lau, 2010; Yokoyama et al., 2010) regions which are also highly susceptible to aging-related atrophy (Raz et al., 2005; Resnick, Pham, Kraut, Zonderman, & Davatzikos, 2003) and therefore metacognitive efficiency may be expected to decrease as we age. Such a hypothesis is consistent with reports that lack of awareness of cognitive, physical and perceptual abilities in healthy older adults can be problematic in everyday life. (Hertzog & Hultsch, 2000) demonstrated that there are notable changes in self-appraisal as we age, and these tend to centre on inaccuracies regarding beliefs about cognitive ability and control over cognition. Older adults tend to demonstrate increased over-confidence compared to actual performance when compared to younger adults (Dodson, Bawa, & Krueger, 2007; Hansson, Rönnlund, Juslin, & Nilsson, 2008). For example when older adults between the ages of 65–91 years old were asked about their driving abilities, 85% of the drivers in this age range rated themselves as ‘good’ or ‘excellent’ drivers despite an increased frequency of accidents (Ross, Dodson, Edwards, Ackerman, & Ball, 2012). However, the literature on laboratory measures of metacognition such as confidence judgments and JOLs has shown mixed results. Some studies reveal stable or even improved accuracy of confidence ratings with age for general knowledge (Dodson et al., 2007; Pliske & Mutter, 1996), problem solving (Vukman, 2005), or memory recall tasks (Lachman, Lachman, & Thronesbery, 1979). Similarly, studies investigating JOLs, FOKs and “judgments of forgetting” have found that older adults’ predictions of recall or recognition were as good as those of younger adults (Eakin, Hertzog, & Harris, 2014; Haber, 2012; Halamish, McGillivray, & Castel, 2011). In contrast, other studies report significant age differences in the accuracy of confidence judgments about recall and recognition (Bender & Raz, 2012; Dodson et al., 2007; Huff, Meade, & Hutchison, 2011; Kelley & Sahakyan, 2003; Pansky, Goldsmith, Koriat, & Pearlman-Avnion, 2009; Perrotin, Isingrini, Souchay, Clarys, & Taconnat, 2006; Soderstrom, McCabe, & Rhodes, 2012; Souchay, Isingrini, & Espagnet, 2000; Souchay, Moulin, Clarys, Taconnat, & Isingrini, 2007; Toth, Daniels, & Solinger, 2011; Wong, Cramer, & Gallo, 2012) learning of emotional information (Tauber & Dunlosky, 2012), and study-time allocation (Froger, Sacher, Gaudouen, Isingrini, & Taconnat, 2011). In addition, the neural correlates of metacognitive judgments have been found to differ between younger and older adults (Chua, Schacter, & Sperling, 2009). In many of these studies it has proven difficult to decouple metacognitive accuracy from age-related changes in performance. Common measures of metacognitive accuracy such as the gamma correlation are affected by task performance (Masson & Rotello, 2009), potentially confounding changes in metacognition with age with changes in performance. For example, if two individuals A and B have identical metacognitive ability, but A performs better than B on the primary task, A’s metacognition score will appear higher than B’s due to this performance confound. Accordingly, Daniels, Toth, and Hertzog (2009) found that older adults had lower accuracy of immediate JOLs for predicting old/new item recognition, but reasoned that this may reflect age-related memory deficits as opposed to deficits in metacognition. In the present study we employ a recently developed signal detection theoretic measure, meta-d′/d′ (Maniscalco & Lau, 2012), to circumvent this problem. Meta-d′/d′ quantifies the efficiency with which confidence ratings discriminate between correct and incorrect trials in each task domain (perception and memory). Importantly, meta-d′/d′ is a relative measure: given a certain level of processing capacity (d′), meta-d′/d′ quantifies the extent to which a metacognitively optimal observer is aware of their performance. Previous literature has also drawn conceptual similarities between characteristics of memory metacognition and executive functions (Fernandez-Duque, Baird, & Posner, 2000; Pannu & Kaszniak, 2005; Shimamura, 1995; Souchay et al., 2000). In particular it has been suggested that any age- related decline in metacognition may be due to executive limitations associated with aging (Souchay & Isingrini, 2004). Again, however, results from initial studies examining this issue are mixed. FOK but not JOL accuracy has been shown to significantly correlate with executive function (Souchay, Isingrini, Clarys, Taconnat, & Eustache, 2004), indicating that perhaps only some forms of metacognitive judgement require intact executive function to be completed accurately. Perrotin, Belleville, and Isingrini (2007) compared patients with mild cognitive impairment (MCI) to healthy age-matched controls in their FOK abilities. FOK accuracy was primarily related to primary memory performance in MCI patients, whereas in control participants it was linked to executive function. In the present study we therefore examined the effects of both age and executive function on a performance-controlled measure of metacognitive efficiency (meta-d′/d′). We studied retrospective confidence judgments in two domains, perception and memory, following recent evidence that metacognition across different domains may draw upon dissociable neural and cognitive processes (Baird, Smallwood, Gorgolewski, & Margulies, 2013; McCurdy et al., 2013). Recruitment ~~~~~~~~~~~","were recruited via two distinct sampling methods as part of a larger clinical research study. A copy of the Royal Mail Postal Address File (PAF), containing all residential addresses in the boroughs of Lambeth and Southwark (London, UK), was obtained, and approach letters for participation in the study were sent out to a random selection of 184 addresses. A total of 28 participants were recruited using this method. Participants verbally confirmed an absence of mild cognitive impairment (MCI) or memory disorder. Participants were also recruited using the MindSearch database, available to researchers based at King’s College, London. Participants who had registered their details with the database and met the inclusion criteria (n = 585) for this study were selected. All potential participants were contacted via email twice in a 6 month period, with a total of 32 participants responding positively and being included in the study. Our inclusion criteria required participants to be over the age of 18 and have no previous or current presentation of a psychosis related disorder, MCI, memory disorders or Alzheimer’s dementia. We compared our healthy older adult sample to a sample of confirmed MCI patients who were screened as part of a larger clinical study (criteria = scoring > 20 on the MMSE and failing at least one section of the Consortium to Establish a Registry for Alzheimer’s Disease (CERAD) neuropsychological battery). The PAF and MindSearch participants performed significantly better on the Wechsler Memory Scale (immediate and long term recall, long term recognition) and Trails B-A executive function test than the clinical sample. Therefore we can be confident that the majority of the participants included were not suffering from undiagnosed MCI. There was no significant difference between the two recruitment methods regarding mean group age (t = .27, p = .79), years in education (t = −.49, p = .63) or BDI score (t = −.52, p = .61). The sample was recruited using opportunity sampling within the two methods with no set age/gender cells filled. Two subjects with Beck Depression Inventory (BDI; (Beck & Steer, 1984) scores greater than 20 were excluded from further analysis. All participants had normal or corrected to normal vision. Participants ~~~~~~~~~~~~ Our sample consisted of 60 participants, 33 of whom were female. The mean age was 40.28 years old, with a minimum of 18 and maximum of 84 years old, with no difference in age between males and females (t(58) = 0.76, p = 0.45). The mean depression score was 6.03 (SD = 4.32) out of a possible 64 indicating that the sample was not depressed. A subset of participants completed the memory metacognition and neuropsychological tests; therefore for each test the relevant N is stated (Table 1). Metacognition measures ~~~~~~~~~~~~~~~~~~~~~~ Perceptual metacognitive ability was investigated using a computerised visual perceptual task similar to that used previously (Fleming et al., 2010). Each trial required participants to perform a perceptual task. The stimuli used were Gabor patches: circular patches of alternating light and dark vertical bars (2.8 visual degrees in diameter, spatial frequency of 2.2 cycles per visual degree). The contrast between the vertical lines in each standard Gabor patch was 20%, where 0% indicates no difference between the light and dark bars and 100%, the maximum difference (black to white). Six such Gabor patches were arranged in a circle (eccentricity of 6.9 visual degrees) around a central fixation point set on a uniform grey background (see Fig. 1a). One of the six Gabor patches was made to pop-out from the others by increasing the contrast in that patch. The contrast of the pop-out Gabor patches varied from 23% (little effect of pop-out) to 80% (pop-out very clear). The task required participants to view two stimulus arrays, each presented for 200 ms, separated by an interval of 300 ms. The interval between stimuli was filled by a uniform grey screen without the Gabor patches. A single Gabor patch in one of the two intervals was designated as a pop-out. Which of the six Gabor patches popped-out varied randomly between trials. Participants were prompted by a computer display to respond ‘1’ or ‘2’ as to whether they thought the pop-out Gabor patch appeared during the first or second presentation. Participants responded by pressing the numerical keys on the top left-hand side of the laptop keyboard with their left hand. Participants had 2 s in which to make their decision, after which a red box surrounded their selection. No feedback was given as to whether they were right or wrong. Participants then indicated confidence in their decision on a scale of 1–6 (1: relatively low confidence; 6: relatively high confidence; see Fig. 1a.). Participants were encouraged to use the full range of the scale, thinking carefully about how confident they were after each decision. Participants responded by pressing numerically marked keys on the top left-hand side of the laptop keyboard with their left hand, with a red box again surrounding this selection. Participants had 3.5 s to complete this metacognitive judgement. Performance on the task was maintained at around 70% using a 2-down, 1-up staircase procedure (Levitt, 1971). Two consecutive correct visual judgments led to a one step (3%) decrease in contrast of the pop-out Gabor patch in the next trial, whereas one incorrect visual judgment led to a one step increase in contrast of the pop-out patch. This procedure ensures all participants perform with approximately the same accuracy on the primary perceptual task, allowing us to measure metacognitive ability independent of task performance. This was especially useful in the present study as the range of participant ages may otherwise have led to performance bias. Older participants who struggled to make manual responses gave verbal answers to a researcher who made manual responses. The task comprised 5 blocks of 8 min with short breaks between each block, taking approximately 50 min to complete. A standard task instruction sheet was read through by participants on their own, followed by the opportunity to ask the task administrator questions. Participants were seated in a darkened room approximately 60 cm from a laptop computer screen (Sony Vaio, PCG-71614 M laptop; 17 in display; 1280 × 800 pixels). Stimulus display and responses for the tasks were programmed in MATLAB 7.8 (Mathworks Inc., Natick, MA, USA) using the COGENT 2000 toolbox (http://www.vislab.ucl.ac.uk/cogent.php). A practice session of two blocks of eight trials was given at the start to familiarise participants with the task. Participants were tested individually in a quiet room. Memory metacognitive ability was investigated using the 2-alternative forced choice (AFC) memory confidence task devised by McCurdy et al. (2013) (see Fig. 1b). There were three learning and testing blocks, with a different learning time assigned to each block. At the beginning of each block, 50 English words (Calibri font, size 24) were presented simultaneously on the screen for either 0.5, 1, or 1.5 min to create three levels of difficulty in which participants performed at neither chance nor ceiling. English words were generated using the Medical Research Council Psycholinguistic Database (Wilson, 1988). These standard nouns were four to eight letters long, had one to three syllables, and had a familiarity, concreteness, and imagability rating of 400–700 each. Participants were instructed to memorize as many words on the list as possible during the study period. A small notice appeared at the bottom of the screen to inform them when there was 10 s left to study the list. After the study period, a series of trials probing memory for the word list was presented. In each trial, two words were presented to the left and right of fixation. One of these words had been presented on the study list (“old”), and the other word had not been presented previously (“new”). First, participants had 3s to provide a 2-AFC judgment with regard to which word was ‘old’, where pressing ‘1’ referred to the left hand word and ‘2’ referred to the right hand word. Participants then had 3 s to press one of four keys (“7,” “8,” “9,” or “0”) using their right hand to indicate their confidence in being correct on the 2-AFC judgment (signifying “not at all confident” to “very confident” respectively). Older participants who struggled to make manual responses gave verbal answers to a researcher who made manual responses. The task comprised 3 blocks of approximately 5 min each with short breaks in between each block. Calculating metacognitive efficiency Metacognitive efficiency was quantified using the meta-d′ measure developed by Maniscalco and Lau (2012). Meta-d′/d′ is a relative measure: given a certain level of processing capacity (d′), meta-d′/d′ quantifies the extent to which a metacognitively optimal observer is aware of their performance. A meta-d′/d′ value of 1 is equivalent to metacognitively “ideal”, whereas meta-d′/d′ < 1 indicates a failure of metacognitive awareness. Values of meta-d′/d′ greater than 1 indicate that awareness is more accurate than task performance, which may occur for instance if the initial judgment is made under time pressure (Charles, Opstal, Marti, & Dehaene, 2013). Meta-d′ was calculated using MATLAB code available at http://www.columbia.edu/~bsm2105/type2sdt (Maniscalco & Lau, 2012) for both perceptual and memory metacognition. Neuropsychological measures ~~~~~~~~~~~~~~~~~~~~~~~~~~~ IQ was measured using the shortened Wechsler Adult Intelligence Scale-III (WAIS-III; Wechsler, 1997a, 1997b) which comprises the Digit Symbol Coding, Arithmetic, Information and Block Design sub-components of the full WAIS assessment; 39 participants completed this assessment. Executive function (EF) was measured using the “Trail Making” test (Reitan, 1986), which measures ‘set-shifting’, or the ability to shift attention between one task to another. This test consists of two parts in which the subject is instructed to connect a set of 25 dots as fast as possible while still maintaining accuracy. The first task (A) requires participants to connect numbers in ascending order and the second task (B) requires participants to connect the dots, alternating between numbers and letters (in alphabetical order). Time taken to complete both tasks is recorded and a ‘B-A’ score is calculated by subtracting the time taken in task A from time taken in task B. The longer the composite time (B-A) the worse the set-shifting ability. 44 participants completed this assessment. The Wechsler Memory Scale (WMS) ‘logical memory’ sub-task (Wechsler, 1997a, 1997b) was used to gauge short-term, long-term and recognition memory. The logical memory task is a sub-test within the whole WMS battery, designed to detect attention and memory deficits, and asks participants to listen to and remember two spoken stories. After hearing each story participants are asked to immediately recount the story to the researcher (Immediate test), where a higher score is obtained by recalling more key facts. 30 min after the stories are initially heard participants are asked again to recount the stories to the researcher (Delayed test). Finally, a recognition task is then carried out, where participants are asked 15 questions about each story and required to give ‘Yes’ or ‘No’ answers (Recognition test). 34 participants completed the short-term recall (Immediate) test, and 33 participants completed the long-term recall (Delayed) and recognition (Recognition) tests.","Correlations between metacognitive efficiency (meta-d′/d′), age and neuropsychological measures were computed using Pearson’s product-moment correlations implemented in R 3.0.1. 5 subjects who performed lower than 65% in the perceptual task (indicating a failure of the staircase procedure to appropriately control performance) were excluded from analyses of perceptual task data. Regression analyses were carried out using the lm function in R and un-standardised regression coefficients are reported.","The mean percentage of trials answered correctly on the perceptual task was 70.8% (range 65.7–73.6%), with no significant relationship between age and percentage of trials answered correctly (r = −0.16, p = 0.25; n = 53). The mean percentage of trials answered correctly on the memory task was 67.8% (range 49–87%). Performance on the memory task was lower in older adults leading to a negative correlation between age and performance (r = −0.36, p = 0.03; n = 38). We calculated a metacognitive efficiency score (meta-d′/d′) for each participant on both perceptual and memory tasks. The average efficiency score was 1.08 (SD = 0.34) for the perceptual task and 0.73 (SD = 0.81) for the memory task. In the subset of subjects who completed both tasks (N = 32), metacognitive efficiency tended to be greater for perception than for memory (t(31) = 2.19, p = 0.04). Consistent with previous findings (McCurdy et al., 2013) we found a positive association between metacognitive efficiencies across domains (r = 0.40, p = 0.02; n = 32; see Fig. 2a.) However, when two memory metacognition outliers were removed (greater than 2 standard deviations beyond the group mean; see Fig. 2a) this relationship failed to reach significance (r = 0.25, p = 0.19; n = 30). No significant relationships were identified between IQ and metacognitive efficiency (perceptual: r = 0.12, p = 0.51, n = 33; memory: r = 0.11, p = 0.60, n = 27) or years of education and metacognitive efficiency (perceptual: r = 0.016, p = 0.25, n = 53; memory: r = 0.003, p = 0.98, n = 38), consistent with previous reports of a lack of correlation between IQ and metacognition in younger adults (Fleming et al., 2012; Weil et al., 2013). Significant relationships were identified between WMS scores and memory metacognitive efficiency (immediate, r = .37, p = .05, n = 28; delayed, r = .59, p = .01, n = 27; recognition, r = .37, p = .05, n = 27), but not perceptual metacognitive efficiency (immediate, r = .12, p = .54, n = 29; delayed, r = .05, p = .79, n = 28; recognition, r = .07, p = .73, n = 28). We next turned to the relationship between metacognition and age. We found a significant negative relationship between perceptual metacognitive efficiency (meta-d′/d′) and age (r = −0.38, p = 0.005; n = 53; see Fig. 2b). Thus, despite our measure of metacognition controlling for differences in task performance, older adults showed lower awareness of their perceptual task performance than younger adults. The relationship between age and memory metacognitive efficiency (meta-d′/d′) was negative but was non-significant (r = −0.064, p = 0.70; n = 38). We cannot however draw conclusions regarding a differential effect of age on perceptual compared to memory metacognition as the difference between the domain-specific metacognitive efficiency-age correlations in the subset of subjects who completed both perceptual and memory tasks was itself not significant (Hotelling’s t = 0.66, p = 0.51). As meta-d′/d′ is a ratio between two quantities, changes in either or both of its components may contribute to an overall decrease in metacognitive efficiency in the perceptual task. Examining each component separately, in the perceptual task we found a significant increase in d′ with age (r = 0.47, p < 0.001; n = 53) and a non-significant decrease in meta-d′ (r = −0.17, p = 0.22; n = 53). Under the ideal observer model, d′ should be equal to meta-d′. Thus both an increase in performance and a decrease in the ability to appraise this performance (given a particular level of d′) contributed to the observed decrease in overall efficiency (Fig. 2c). For the memory task, there were no significant changes in either d′ (r = 0.016, p = 0.92; n = 38) or meta-d′ (r = −0.15, p = 0.36) with age. Finally, we considered that the relationship between metacognitive efficiency and age may be mediated by changes in executive function. In order to assess this hypothesis we constructed a general linear model (GLM) that predicted metacognitive efficiency from age and executive function as measured by the Trail Making Test. This analysis was restricted to a subset of participants who conducted the Trail Making Test (N = 44). We found that the relationship between age and perceptual metacognitive efficiency remained significant after controlling for changes in executive function (β = −0.0058, p = 0.02), and increased executive function was not associated with better metacognitive efficiency (β = 0.0042, p = 0.26).","The primary aim of this study was to investigate the effects of age on metacognitive efficiency in healthy adults between the ages of 18 and 84. We found that perceptual metacognitive efficiency declined with age, despite task performance being controlled to ensure all participants performed with the same accuracy. In other words, older adults were less efficient at introspecting about whether they are performing well or badly on a perceptual task than younger adults. This result is consistent with previous observations of a weaker match between beliefs and abilities in older adults (Hultsch, MacDonald, Hunter, Levy-Bencheton, & Strauss, 2000; Ross et al., 2012) and age-related differences in the accuracy of confidence judgments (Bender & Raz, 2012; Dodson et al., 2007; Huff et al., 2011; Kelley & Sahakyan, 2003; Pansky et al., 2009; Audrey Perrotin et al., 2006; Soderstrom et al., 2012; Souchay et al., 2000, 2007; Toth et al., 2011; Wong et al., 2012). However, many previous studies did not control for the influence of task performance on measures of metacognition. This is particularly critical when studying aging as metacognitive ability may be difficult to distil from other age-related changes in cognitive abilities. In the current study we employ a measure of metacognition, meta-d′/d′, that controls for the influence of task performance and response bias (Maniscalco & Lau, 2012). Our results extend those of Weil et al. (2013) who took a similar approach to study the development of metacognitive efficiency during adolescence. In a sample of 28 adolescents and 28 adults it was found that perceptual metacognitive efficiency increased with age during adolescence (11–18 years old), with a non-significant decrease with age in adulthood. However the maximum age in Weil et al’s sample of older adults was 41, precluding the study of metacognitive efficiency in older age. Here we extend this age range to 84, finding that efficiency continues to decline despite task performance remaining stable. A subset of subjects additionally completed a recognition memory task that allowed us to calculate a metacognitive efficiency score in the memory domain. We found a positive correlation in efficiency scores across domains that was weakened after removal of two outliers. Our result provides some support for the notion that there is a global correlation in metacognitive ability across domains (McCurdy et al., 2013), but also is consistent with a large proportion of domain-specific variance that may be linked to separate brain systems underpinning perceptual and memory metacognition (Baird et al., 2013; McCurdy et al., 2013). There was no effect of age on memory metacognition, but we are cautious about over-interpreting this null result for two reasons: first, fewer subjects completed the memory task, reducing our power to detect an effect of age; and second, there was no statistical support for a differential effect of age on perceptual vs. memory metacognition. Further work is required to ascertain whether may be different trajectories for age-related changes in domain-specific metacognitive functions. Mixed results were obtained for the effect of age on measures of basic task performance (% correct and d′). In the perceptual task, effects of age on these two measures had opposite sign (negative for% correct and positive for d′). Dissociations between% correct and d′ are possible if the decision criterion is also changing (Macmillan & Creelman, 2004). However such differences in the current study are difficult to interpret, because performance was controlled in the perceptual task such that% correct varied over a narrow range, and poorly performing subjects were excluded prior to analysis. It is possible that increases in d′ with age reflect an overcompensation in difficulty adjustment that made the task slightly easier for older adults. In the memory task, % correct declined with age, whereas this effect was not seen in an analysis of d′. The meta-d′ approach estimates the subject’s metacognitive accuracy in units directly comparable to d′, a measure of primary task performance (Maniscalco & Lau, 2012). For a metacognitively ideal observer, meta-d′ = d′, and meta-d′/d′ = 1. Closer examination of the relationship between these quantities and age revealed an increase in perceptual d′ and a decrease in perceptual meta-d′, leading to an overall decrease in metacognitive efficiency (Fig. 2c). As noted above, the narrow range of performance levels in the perceptual task precludes strong interpretation of changes in d′. However, the relative values of d′ and meta-d′ are informative. In younger adults, meta-d′ is similar to or slightly above d′ (meta-d′/d′ ∼ 1), whereas in older adults, meta-d′ drops below d′, leading to metacognitive efficiency scores less than expected on an ideal observer model. One might expect that age-related changes in metacognitive processes would be related to neuropsychological measures of executive function (Fernandez-Duque et al., 2000), following evidence that perceptual metacognitive efficiency is linked to frontal lobe function (Fleming et al., 2010; Hoerold, Pender, & Robertson, 2013; Persaud et al., 2011; Rounis et al., 2010). Previous studies have found evidence for a link between the use of metacognitive information in the control of behaviour and executive function as measured by Wisconsin Card Sorting Test performance (Pansky et al., 2009; Souchay & Isingrini, 2004), but the relationship between executive function and metacognitive monitoring is less well understood. In particular, it is difficult to rule out an indirect effect of executive function on measures of metacognition via effects on primary task performance, highlighting the need to distil a measure of metacognitive efficiency that controls for differences in performance. In the current study, we did not find evidence for a relationship between perceptual metacognitive efficiency and performance on the Trail Making Test, a measure of set-shifting. Indeed, in a regression analysis explicitly controlling for changes in set-shifting performance, we observed an annual decline of 0.6% in metacognitive efficiency. Of course, executive function is a very broad construct, and the relationship between metacognitive efficiency and other components such as inhibition and updating (Miyake et al., 2000) remain to be determined. The perceptual and memory metacognition tasks were adapted from those used in recent structural and functional imaging studies of healthy participants (Baird et al., 2013; Fleming et al., 2010; McCurdy et al., 2013). However a limitation of this design is that there are differences between tasks that are potentially orthogonal to the domain in question. For example, our recognition memory task involved verbal stimuli; we cannot rule out the possibility that a different pattern of results would be obtained with memory tasks using nonverbal stimuli such as faces. Indeed, an important goal for future work is the development of perceptual and memory metacognition paradigms that are more closely matched for stimulus characteristics. Similarly it is unclear whether the observed decrease in metacognitive efficiency is specific to visual perceptual metacognition, or may extend to other perceptual modalities or other aspects of decision-making. An additional difference between domains is that the perceptual task used a staircase procedure to control task difficulty for each individual participant whereas this was not feasible with the current word recognition task. By adjusting stimulus selection online such performance control could be built into a memory paradigm in future studies. Finally, a small minority of older individuals struggled to input confidence ratings manually using the keyboard response, and their verbal ratings were instead entered by the experimenter. We cannot rule out effects of changes in response modality on our results, but we note that even in older individuals metacognitive efficiency did not decline to floor levels (meta-d′/d′ ∼ 0), indicating that confidence ratings tracked task performance in a meaningful way regardless of adjustments in response modality. In summary, we reveal an age-related decline in perceptual metacognitive efficiency whilst controlling for age-related differences in task performance and executive function. Our study is cross-sectional, and it is possible that other factors differing across the age range affected metacognitive ability. It will therefore be very informative in further work to investigate longitudinal changes in metacognitive efficiency with age. Quantifying changes in metacognition with age is critical for our understanding of higher-order cognitive functions in an aging population, especially as deficits in metacognitive monitoring may lead to impaired control of behaviour (Koriat & Goldsmith, 1996). Aging-associated diseases such as Alzheimer’s are accompanied by metacognitive deficits that may lead to non-adherence to treatment and impaired decision-making (Cosentino, 2014). Our results, when combined with previous research in adolescents (Weil et al., 2013), reveal a non-linear relationship between age and perceptual metacognitive efficiency, increasing during adolescence, plateauing in early adulthood, and declining in older age.","E.C.P, A.S.D and S.M.F. designed research; E.C.P performed research; E.C.P. and S.M.F. analysed data; E.C.P, A.S.D and S.M.F. wrote the paper."],["This article pays tribute to the seminal paper by Peter J. Lang (1977; this journal), “Imagery in Therapy: Information Processing Analysis of Fear.” We review research and clinical practice developments in the past five decades with reference to key insights from Lang's theory and experimental work on emotional mental imagery. First, we summarize and recontextualize Lang's bio-informational theory of emotional mental imagery (1977, 1979) within contemporary theoretical developments on the function of mental imagery. Second, Lang's proposal that mental imagery can evoke emotional responses is evaluated by reviewing empirical evidence that mental imagery has a powerful impact on negative as well as positive emotions at neurophysiological and subjective levels. Third, we review contemporary cognitive and behavioral therapeutic practices that use mental imagery, and consider points of extension and departure from Lang's original investigation of mental imagery in fear-extinction behavior change. Fourth, Lang's experimental work on emotional imagery is revisited in light of contemporary research on emotional psychopathology-linked individual differences in mental imagery. Finally, key insights from Lang's experiments on training emotional response during imagery are discussed in relation to how specific techniques may be harnessed to enhance adaptive emotional mental imagery training in future research. -------------------------------------------------------------------------------- Mental imagery refers to perceptual experience in the absence of sensory input, commonly described as seeing with the “mind’s eye,” hearing with the “mind’s ear,” and so on (Kosslyn, Ganis, & Thompson, 2001). In this paper we consider mental imagery both as an emotion-evoking stimulus that can be manipulated (e.g., during therapeutic techniques such as imaginal exposure; Foa, Hembree, & Rothbaum, 2007), and as a symptom of psychopathology (e.g., distressing intrusive memories/flashbacks in posttraumatic stress disorder; cf. Holmes & Mathews, 2010). In his bio-informational theory of emotional imagery, Lang (1977, 1979) postulated that a mental imagery representation of an emotionally charged stimulus (e.g., a spider) activates an associative network of stored information that overlaps with that activated during actual experience of the stimulus in reality (e.g., encountering a live spider). This associative network of information is said to consist of perceptual information about the stimulus (color, shape, size, texture of spider), semantic information about what it means (insect, danger, bite), somatovisceral response information about what it feels like to encounter the stimulus (fear, racing heart), and preparatory motor responses evoked by the encounter (e.g., muscles tensing to flee from the spider). According to this theory, mental imagery differs from verbal thought in that only mental imagery has the capacity to activate physiological and behavioral response systems (Lang, 1987). Research pertaining to this assertion will be discussed in Section III. Lang's (1977, 1979) associative-network information processing definition of mental imagery was influenced by Pylyshyn's (1973) propositional theory of mental imagery, which construed mental imagery as conceptual representations describing reality, rather than as pictorial representations depicting reality. Historically, the debate concerning whether mental imagery involves conceptual representations (Pylyshyn, 1973) or pictorial representations (Kosslyn, 1981) has been the subject of heated debate. Neuroimaging and psychophysics evidence gathered in the past 50 years has largely resolved the debate in favor of the latter view (see Pearson, Naselaris, Holmes, & Kosslyn, 2015, for a review). Although this appears to refute the grounds of Lang’s bio-informational theory, closer inspection reveals that the validity of this theory stands irrespective of whether mental imagery is conceptual or pictorial in nature. As Lang (1987) explicated, the bio- informational theory was developed not as a theory on the nature of mental imagery, but as a functional theory concerning the impact of mental imagery on emotional processing. Due to the overlap in perceptual information between imagined and real stimuli, Lang (1977, 1979) proposed that imagined interaction with a stimulus can evoke corresponding emotional responses associated with real interaction with that stimulus. As such, imagined interaction with stimuli can function as an “as-if real” template for rehearsing and modifying emotional and behavioral responses to the same stimuli in real life. Lang (1977, 1979) proposed that this function of mental imagery could be harnessed in clinical treatment to facilitate fear-extinction learning and habituation via the rehearsal and learning of new adaptive responses during imaginal exposure therapy. Indeed, numerous studies using a range of associative learning paradigms have shown that mental imagery can produce conditioned responses in the same way as real stimuli (cf. Dadds, Bovbjerg, Redd, & Cutmore, 1997; Lewis, O'Reilly, Khuu, & Pearson, 2013). Crucially, Lang proposed that the elicitation of a fear emotional response during mental imagery of the feared stimuli (simulation of both perceptual representations and autonomic and behavioral responses) is required for learning to occur, and is therefore necessary in order for imaginal exposure to be effective (Lang, 1977; Wolpe, 1958). Interestingly, Lang’s conception of mental imagery as an “as-if real” template parallels contemporary functional perspectives on mental imagery. These contemporary accounts view mental imagery as a core component of the \"prospective brain,\" which enables the simulation of hypothetical future events based on prior knowledge and memories of past experience for the purposes of prediction and planning (Moulton & Kosslyn, 2009; Schacter, Addis, & Buckner, 2008; Suddendorf & Corballis, 2007). Of particular relevance to Lang’s (1977, 1979) functional theory is Moulton and Kosslyn's (2009) theory of mental imagery as emulation. Emulation is defined as the episodic construction of a hypothetical scenario that simulates not only perceptual information about an event, but also rich semantic and affective information about plausible causes and consequences of the imagined scenarios (Moulton & Kosslyn, 2009). As such, both Lang (1977, 1979) and Moulton and Kosslyn (2009) postulate that mental imagery has the capacity to evoke cognitive and emotional responses, enabling the individual to not only “try out” one or more versions of what might happen (Schacter et al., 2008, p. 40), but to also “try out” the emotional consequences of alternative courses of action (Lang, 1987, p. 412; Moulton & Kosslyn, 2009, p. 1278). By providing a pivotal testable framework linking mental imagery to emotion, Lang’s (1977, 1979) bio-informational theory paved the way for much subsequent research concerning the functions of mental imagery. While Lang focused on the role of imagery in fear extinction learning, contemporary theorists point to the wider role of mental imagery in planning, problem solving, and self-regulation (Gilbert & Wilson, 2007; Suddendorf & Corballis, 2007; Taylor, Pham, Rivkin, & Armor, 1998). The first half of this review will evaluate evidence pertaining to Lang’s key proposal, that mental imagery has the capacity to evoke emotional responses, and explore how this capacity has been exploited therapeutically in clinical treatment. The second half of the review will consider how subsequent research has built on Lang’s initial work concerning individual differences to illuminate how variability in imagery may relate to emotional dysfunction, and highlight neglected insights from Lang’s research on emotional imagery training that could potentially contribute to future treatment innovation.","Developments in clinical psychology over the past 50 years have shown that unwanted and distressing mental imagery is a symptom present in a wide range of anxiety and mood disorders (Holmes & Mathews, 2010). In anxiety disorders, mental imagery phenomena include the intrusive flashbacks that define posttraumatic stress disorder (Brewin & Holmes, 2003; Ehlers et al., 2002), and imagery of feared stimuli or distorted images of one’s own physical appearance in the case of phobias or social phobia, respectively (Hirsch & Holmes, 2007). Depression and bipolar disorder symptomology also feature mental imagery of past failures and trauma (Kuyken & Brewin, 1994), as well as possible future events, such as suicidal acts (Holmes, Crane, Fennell, & Williams, 2007). While clinical presentation suggests that negative mental imagery symptomology can have emotionally distressing consequences for patients, researchers have noted a relative paucity of studies designed to directly evaluate this assumption (Holmes & Mathews, 2010; Watts, 1997). Lang (1977, 1979) was one of the first investigators to empirically examine the capacity for mental imagery to evoke emotional response. Much of Lang’s experimental work on mental imagery was devoted to assessing psychophysiological reactivity to mental imagery cued by verbal scripts depicting emotional scenarios and bodily responses (e.g., “your face is flushed as your muscles strain to continue the pace”). In this section, we review evidence resulting from the work of Lang, and others, that mental imagery does indeed have an impact on emotion, as indexed by physiological, neurological, and subjective measures. Furthermore, response information measured at these levels has also revealed biases in emotional processing in individuals suffering from emotional psychopathology, such as mood and anxiety disorders. The Impact of Emotional Mental Imagery on Physiological Activity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ It is widely accepted that emotional experience involves activation of the central and peripheral nervous systems (Gross & Barrett, 2011). Thus, if mental imagery representations of emotional scenarios are capable of eliciting emotional responses in an “as-if real” manner (Kosslyn et al., 2001; Kreiman, Koch, & Fried, 2000; Lang, 1979), this may be observable at a physiological level. Using verbal scripts depicting emotionally negative or neutral information, Lang’s (1977, 1979) experimental work measured peripheral nervous system responses while participants imagined the content of the scripts. The central hypothesis was that physiological effects would arise during emotional imagery as a result of activation of perceptual memories, triggering somatovisceral responses associated with experience of the stimuli in reality. In terms of autonomic arousal, Lang and colleagues observed greater heart rate acceleration during mental imagery of high versus low arousal scenarios. In healthy participants, imagery of fear and anger-related scenarios led to greater heart rate acceleration than did imagery of neutral scenarios (Cook, Hawk, Davis, & Stevenson, 1991; Vrana, 1995; Vrana, Cuthbert, & Lang, 1986; Witvliet & Vrana, 1995). Similarly, skin conductance response (SCR) levels also indicated elevated emotional arousal during mental imagery of emotional scenarios, such that scripts depicting highly arousing pleasant and unpleasant experiences produced larger increases in SCR relative to those depicting neutral experiences, in healthy populations (Lang, Levin, Miller, & Kozak, 1983; Weerts & Lang, 1978). Research has shown that respiratory responses are also influenced by mental imagery in healthy populations. Specifically, imagery evoked by scripts depicting high arousal scenes (fear and action-related) compared to low arousal scenes (relaxation and depression-related) produced greater drops in end-tidal fractional carbon dioxide concentration, likely reflecting hyperventilation (Van Diest et al., 2001). Importantly, this hyperventilation during emotional imagery was more pronounced in individuals with higher relative to lower imagery generation ability, as assessed using the Questionnaire Upon Mental Imagery (QMI; Sheehan, 1967). This individual differences dimension in imagery has also been explored in Lang’s work, and will be considered in more detail in Section III. In addition to autonomic indices, mental imagery has been shown to modulate other indices of physiological arousal, such as the startle blink reflex. Vrana and Lang (1990) required healthy participants to first learn and then recall six pairs of sentences depicting fear-related and neutral scenarios. During recall, participants were instructed either to relax and ignore the sentence (control condition), to silently articulate the sentence (verbal condition), or to imagine the sentence content as a personal experience (imagery condition). As expected, startle blink reflexes evoked by acoustic probes were found to be greater during recall of fear relative to neutral sentences. Importantly, this effect was greater in the imagery condition than in either the control condition or the verbal condition (Cuthbert et al., 2003). Finally, mental imagery of food stimuli has been shown to modulate the gustatory salivary reflex. Research in the field of brain computer interfaces (BCI) has found increases and decreases in salivary pH levels compared to baseline as a result of a healthy participant imagining consuming a lemon versus drinking milk, respectively (Vanhaudenhuyse, Bruno, Bredart, Plenevaux, & Laureys, 2007). This finding has also been harnessed to demonstrate conscious awareness in clinical patients with complete locked-in syndrome (LIS), individuals who cannot otherwise indicate the presence of conscious awareness (Wilhelm, Jordan, & Birbaumer, 2006). In addition to appetitive salivary responses, repetitive mental imagery of food consumption (eating M&Ms) has also been shown to lead to food item-specific satiation effects (Morewedge, Huh, & Vosgerau, 2010). Although no physiological measures were included in the study, consumption of that food item was reduced in the high- repetition imagery group (30 repetitions) relative to the low-repetition imagery group (three repetitions), indicating satiation effects (Morewedge et al., 2010). Together, the above evidence provides support for Lang’s (1977, 1979) contention that mental imagery has the capacity to activate the peripheral nervous system and evoke somatic responses, be it in the form of “fight or flight” sympathetic system response during fear- or anger-related imagery, or “rest and digest” parasympathetic response during gustatory imagery. However, one limitation associated with using physiological indicators of autonomic nervous system (ANS) activation is that these can be employed only to index high arousal emotions, such as anger, fear, disgust, or elation. In contrast, both high and low arousal emotional responses can be assessed using neurological measures, such as those provided by neuroimaging. Furthermore, the studies reviewed in this section have not explicitly contrasted mental imagery representations of emotional stimuli to an alternative mode of representation of the same information, and therefore one cannot conclude based on the evidence that the physiological effects observed reflect the impact of mental imagery per se, rather than the impact of emotional information processing in general. Research addressing this limitation will be discussed in Section III. The Impact of Emotional Mental Imagery on Neural Measures of Emotional Response ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ According to Lang's (1979) bio-informational theory, emotional imagery is hypothesized to be able to elicit similar physiological responses in both peripheral and central systems as would be evoked during actual experience (Lang & McTeague, 2009). Since the 1970s, developments in brain imaging technology have enabled assessment of the impact of emotional imagery on central nervous system (CNS) structures involved in coordinating peripheral nervous system (PNS) responses. Neural indices of emotion processing during emotional mental imagery can be evaluated in comparison to activity observed during veridical perception of emotional stimuli, or to emotionally neutral mental imagery. Meta- analytic reviews of neuroimaging studies on healthy participants show that the brain regions most consistently associated with emotional processing are the dorsomedial prefrontal cortex (mPFC), anterior cingulate cortex (ACC) and amygdala (Murphy, Nimmo- Smith, & Lawrence, 2003; Phan, Wager, Taylor, & Liberzon, 2002), and the insular cortex (Craig, 2009), all of which are involved in the coordination of ANS activity. Activation of these same emotion-processing regions has been observed during emotional mental imagery. Early hemodynamic neuroimaging studies using positron emission tomography (PET) showed that mental imagery of emotional relative to neutral information elicited increased regional cerebral blood flow (rCBF) to the mPFC, ACC and anterior insula (Kosslyn et al., 1996; Partiot, Grafman, Sadato, Wachs & Hallett, 1995; Schaefer et al., 2003). This effect has been found in PET studies not only for mental imagery evoked by scripts depicting negative scenarios (aggression and guilt-related) (Shin et al., 2000), but also for mental imagery evoked by scripts depicting positive scenarios (success and affection-related) (Schaefer et al., 2003). More recent studies have used script-driven emotional imagery to examine the specificity of neural circuits for the processing of emotional valence versus arousal. Costa, Lang, Sabatinelli, Versace, and Bradley (2010) used functional magnetic resonance imaging (fMRI) to examine negative, positive, and neutral script-driven imagery, and found that activation of the amygdala was enhanced during emotional relative to neutral imagery, irrespective of emotional valence. However, activation of the nucleus accumbens (NAc) and mPFC were selectively enhanced during positive relative to neutral imagery, not negative relative to neutral imagery (Costa et al., 2010). Results from this script-driven imagery study are also consistent with a previous fMRI study where participants viewed positive and neutral pictures (Sabatinelli, Bradley, Lang, Costa, & Versace, 2007). Studies using fMRI have also compared neural responses to real versus imagined emotional stimuli in the same participants. Kim et al. (2007) found comparable magnitudes of left hemisphere amygdala activity between when participants viewed faces with negative and positive emotional expressions and when they generated mental imagery of such faces. In fMRI studies using real-time neural activation feedback (“neurofeedback”), over successive trials participants were able to use self-generated visualization of positive and negative scenarios to regulate the activation of emotion processing regions such as the insula (S. Lee et al., 2011). In addition to face stimuli, mental imagery of positive and negative events in the past and future has been shown to activate the amygdala and ACC (Sharot, Riccardi, Raio, & Phelps, 2007). In another study examining negative and positive mental imagery in the same participants, Damasio et al. (2000) found different ACC subregions were differentially activated by fear, anger, happiness, and sadness imagery. Finally, direct evidence that visual imagery and visual perception share common neural systems implicated in emotional response comes from a single-cell recording study on human epileptic patients. Kreiman et al. (2000) found that 75% of the 89 recorded cells in the amygdala selectively altered their firing rates in comparable patterns during both visual perception and visual imagery recall of emotional faces, whereas 4% and 10% responded selectively to imagery or veridical perception, respectively (Kreiman et al., 2000). As such, results from neuroimaging and single cell recording studies provide converging evidence that mental imagery has the capacity to evoke an emotional response by activating neural networks involved in emotional processing and response coordination with the autonomic nervous system. However, as with the physiology studies reviewed earlier, the neuroimaging studies reviewed here have not explicitly contrasted mental imagery representations of emotional stimuli to an alternative mode of representation of the same information, and therefore one cannot conclude based on the evidence that the physiological effects observed reflect the impact of mental imagery per se, rather than the impact of emotional information processing in general. Disruption of Emotional Mental Imagery Reduces Emotional Impact ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ If mental imagery-based representation of emotional information serves to elicit emotional responses to this information, then it follows that disruption of such imagery should reduce the intensity of such emotional responding. This hypothesis has been investigated using dual task paradigms that reduce the availability of cognitive resources required for mental imagery generation. Typically, in such studies, participants attempt to generate mental imagery while undertaking a concurrent task that consumes working memory resources, such as sequenced finger tapping on a number pad, or mental arithmetic. Across several experiments, Andrade, Kavanagh, and Baddeley (1997) asked healthy participants to generate visual imagery from emotionally negative versus neutral cues (photographs and personal memories) while concurrently tapping a spatial pattern with a finger, engaging in lateral eye movements, or performing no concurrent task. Results showed lower vividness of mental imagery in the concurrent spatial tapping and lateral eye movement conditions, relative to the no-concurrent task control condition. Critically, this reduction in the imagery vividness was accompanied by a corresponding decline in the intensity of participants’ emotional responding to the negative, compared to neutral, cues. This result was replicated by Kavanagh, Freese, Andrade, and May (2001) using a within-subjects design in which participants generated both positive and negative autobiographical episodic memory imagery at three time points. At the first and third time points, imagery generation took place without a concurrent imagery interference task, whereas at the second time point, imagery interference was produced by concurrent performance of a lateral eye-movement task and viewing of dynamic visual noise, in counterbalanced order. Concurrent imagery interference was successful in reducing the vividness ratings for negative (but not positive) memories, as compared to the no-interference condition. Again, the critically important finding was that performance of the concurrent task also served to dampen the emotional impact of these memories. Clinical researchers have also begun to harness the potential for concurrent tasks to reduce the emotional impact of processing affectively toned mental imagery. The use of a concurrent mathematics task during recall of real-life collective trauma has been shown to reduce the emotional impact of recalling the traumatic event in the general population (Engelhard, van den Hout, & Smeets, 2011). Likewise, it has been shown that performance of a concurrent capacity-consuming task reduces the emotional impact of thinking about feared future events in healthy participants (Engelhard, van den Hout, Janssen, & van der Beek, 2010), in students who report high frequency of intrusive future fear imagery (Engelhard, van den Hout, Dek, et al., 2011), and in clinical patients with PTSD (Lilley, Andrade, Turpin, Sabin-Farrell, & Holmes, 2009). Lilley et al. (2009) asked a group of patients with PTSD to generate trauma-related mental imagery while engaging in concurrent lateral eye movements, counting out loud, or performing no concurrent task. Participants in the eye-movement condition reported the lowest levels of imagery vividness, followed by those in the phonological counting condition, while participants in the no-concurrent task control condition reported the highest levels of imagery vividness. The intensity of negative emotion experienced by participants during this procedure followed this exact same function, being lowest in the participants who performed the concurrent eye movement task, intermediate in those who performed the concurrent phonological counting task, and highest in those who performed no concurrent task. Thus, manipulating the vividness of mental imagery through the use of this concurrent task approach served to influence the emotional impact of processing trauma-relevant information. In addition to voluntarily generated imagery, researchers investigating intrusive flashbacks in PTSD have also begun to evaluate the possibility that capacity-consuming concurrent tasks may also be capable of reducing the frequency and emotional impact of involuntary mental imagery. There is growing evidence that playing the visuospatial computer game “Tetris” following exposure to analogue trauma reduces the frequency of subsequent involuntary memory imagery. Holmes, James, Coode-Bate, and Deeprose (2009) had healthy participants view film clips depicting traumatic scenes, then either play the visuospatial game Tetris after film viewing, or not. Participants used a diary to report intrusive imagery of the film clip content over the following week. Participants who played Tetris after film viewing reported experiencing fewer such imagery intrusions, relative to those who had not. This effect appears to be modality-specific, as playing a predominantly verbal “pub quiz” game after film viewing did reduce subsequent imagery intrusions to the same degree as playing Tetris (Holmes, James, Kilford, & Deeprose, 2010). Subsequent studies have also found Tetris game play to be effective in reducing intrusive memory frequency when memories were reactivated 24 hours after initial film viewing (James et al., 2015). Future research could profitably investigate whether Tetris alleviates imagery intrusions in PTSD symptoms specifically because of visuospatial disruption (James et al., 2015) or as a result of more general working memory taxation (Van den Hout & Engelhard, 2012), as investigators remain divided in their views concerning this issue.","Mental imagery-evoked emotional responses have been used to study emotional psychopathology, primarily in anxiety research. Evidence has revealed variations in neural and physiological responses across anxiety disorder subtypes. For example, while there is evidence that, overall, anxiety patients exhibit greater amygdala and insula activity during aversive imagery and picture viewing relative to healthy controls, this hyperactivation is more evident in social anxiety disorder and specific phobias compared to PTSD (cf. Etkin & Wager, 2007). Similarly, during fear relative to neutral memory imagery, specific and social phobics showed greater increases in heart rate and startle reflex response (but not skin conductance) compared to healthy participants (Cuthbert et al., 2003). However, those with panic disorder and PTSD showed hyporeactivity relative to healthy controls, even when baseline heart rate and startle potentiation were taken into account (Cuthbert et al., 2003). Using the same paradigm, McTeague et al. (2010) found that while PTSD patients generally exhibited greater increases in heart rate and startle reflex response (but not in skin conductance) compared to healthy controls, more severe multitrauma PTSD patients showed blunted heart rate, skin conductance, and startle reflex responses compared to single-trauma PTSD patients (McTeague et al., 2010). Indeed, reduced neural fear responding has been linked to the dissociative subtype of PTSD, characterized by symptoms of depersonalization and derealization, symptoms associated with higher trauma severity and chronicity (Lanius, Brand, Vermetten, Frewen, & Spiegel, 2012). Blunted neurophysiological reactivity to aversive imagery is hypothesized to be attributable to the effects of chronic stress and depression comorbidity, which appears to be more evident in individuals with pervasive anxiety-based disorders rather than focal fear-based phobias (for a review, see Lang & McTeague, 2009). Concordance Between Subjective and Physiological Response to Emotional Mental Imagery in Psychopathology ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Variations in the level of concordance between subjective emotional distress and physiological arousal during aversive mental imagery have also been investigated as an indicator of emotional psychopathology. This section will review studies that have assessed the level of concordance between self-report, physiological and neural measures. In healthy volunteers, there is evidence of concordance between self-report and physiological indicators of emotional responding. For example, in a study with healthy volunteers reporting high fear of snakes or public speaking, Lang et al. (1983, Experiment 2) found concordance between self-report and physiological indicators of arousal during fear-imagery, though only when participants had been trained beforehand to focus on their own somatic responses during imagery, and not when participants instead had been trained to focus on visual perceptual information about the stimuli. Similarly, Dalgleish et al. (2013) found concordant variation in subjective reports and physiological indicators (skin conductance and heart rate) of emotional responses during imagery of personal memory imagery of fear, anger, happiness and sadness, in 41 out of 53 participants. In anxiety disorders, Lang and colleagues found enhanced self-reported distress and heart rate and skin conductance reactivity during imagery of personal phobic stimuli relative to personal non-phobia-related stimuli and standardized stimuli (Cook, Melamed, Cuthbert, McNeil, & Lang, 1988). Interestingly, the congruence between levels of subjective distress and physiological fear response was greatest in participants with simple phobias, followed by those with social phobia, with the least congruence found in those with agoraphobia (Cook et al., 1988). Similarly, McNeil, Vrana, Melamed, Cuthbert, and Lang (1993) assessed aversive imagery in volunteers and clinically anxious patients with simple and social phobias, and found that subjective reports of imagery vividness and emotional distress were positively related to heart rate reactivity in those with simple phobias (e.g., dental surgery) but not in those with social (speech) phobia. As reviewed in Lang and McTeague (2009), subjective reports of emotional distress in anxiety disorders are not always accompanied by concordant levels of physiological fear response, at least for startle potentiated reflex responses. Whereas exaggerated startle reflex response is found in patients with specific and social phobia during phobic imagery and picture-viewing, blunted startle response is found in patients with more severe anxiety profiles, such as those with agoraphobia, panic disorder, generalized anxiety disorder (GAD), and comorbid depression (Lang & McTeague, 2009). Indeed, physiological response to aversive mental imagery appears to be negatively related to disorder chronicity, severity, and depression- comorbidity, possibly reflecting suppressed defensive reactivity due to long-term stress (McTeague & Lang, 2012). However, the view that chronic anxiety is associated with negative emotional suppression is inconsistent with the Contrast Avoidance Model of worry (Newman & Llera, 2011). This model postulates that individuals with GAD defensively recruit a perpetuated negative emotional state to emotionally prepare for unexpected negative events. Preliminary evidence in support of this model is provided by the finding that worry-induction in healthy volunteers and GAD patients increases negative emotional states at both subjective and physiological (skin conductance) levels, preventing further increases during exposure to negative imagery (Peasley-Miklus & Vrana, 2000) and film clips (Llera & Newman, 2014). The presence of ceiling effects in negative emotional reactivity to imaginal and real aversive stimuli should be further investigated in GAD and other anxiety-disorder subtypes. SECTION III: THE IMPACT OF MENTAL IMAGERY ON EMOTION—A COMPARISON WITH VERBAL REPRESENTATION -------------------------------------------------------------------------------- Section I reviewed evidence that mental-imagery-based representations of emotional information elicit emotional responses at physiological, neurological, and subjective levels. However, as noted earlier, such studies did not contrast mental imagery to an alternative mode of representation of the same information, and therefore it is unknown whether the effects observed during emotional imagery reflect an impact of emotional information processing in general or mental imagery per se. This section reviews studies that have attempted to answer this question by contrasting mental imagery to verbal representations of emotional information. In several studies, Lang and colleagues compared the impact of mental imagery versus verbal information processing of fear-related and neutral sentences (Cuthbert et al., 2003; Vrana et al., 1986). Participants were instructed either to generate mental imagery based on previously memorized fearful sentences, or to verbally repeat the sentences in silence without generating accompanying mental imagery. Results showed significant and sustained heart rate acceleration relative to baseline during mental imagery, but not during verbal repetition of fearful scenarios (Cuthbert et al., 2003; Vrana et al., 1986). This suggests that mental-imagery-based processing of emotional information has a greater impact on emotion than verbal processing. However, several experimental design issues constrain the interpretation of these results. First, self-reported emotional impact was elicited only following mental imagery of scenarios in Vrana et al. (1986), but not following verbal repetition of scenarios. Thus, it is unclear whether imagery-based processing led to greater subjective emotional response compared to verbal-based processing of the same information (Holmes & Mathews, 2005). Second, due to prior memorization of the cue sentences, participants necessarily engaged in verbal processing of the scenarios prior to completing both mental imagery and verbal repetition conditions. Thus, it is difficult to determine whether differences in emotional response between the imagery and verbal processing conditions simply reflect the recruitment of an additional representational system in the imagery condition (Holmes & Mathews, 2005). Finally, while participants in the imagery condition elaborated on sentence content by simulating them as a personal experience, those in the verbal condition simply repeated the sentences. Thus, it is possible that differences in the level of semantic elaboration between the conditions contributed to differences in observed physiological response. Following Vrana et al. (1986), a series of studies conducted more controlled comparisons of the subjective emotional impact of mental imagery versus verbal processing of emotionally negative information. A commonly used experimental paradigm has involved exposing participants to scenarios that, while initially ambiguous, ultimately resolved in a manner that produces an emotionally valenced representation of the described situation. Experimenters have contrasted the emotional impact of manipulating whether these representations are imagery-based or verbal in nature. Thus, for example, Holmes and Mathews (2005) instructed healthy participants to either generate imagery or focus on the semantic meaning of 100 such initially ambiguous auditory scenarios that always resolved negatively (“You are at work when you hear the fire alarms go off. You run to the exit to discover that it is . . . for real”). Holmes and Mathews (2005) found that participants in the imagery condition reported greater increases in self-reported anxiety as measured using the Spielberger State Anxiety Inventory (STAI-S; Spielberger, Gorsuch, Lushene, Vagg, & Jacobs, 1983) (Experiment 1: M = 7.3; Experiment 2: M = 5.67), relative to the verbal condition (Experiment 1: M = 1.9; Experiment 2: M = 1.55). Importantly, the enhanced emotional impact of mental imagery relative to verbal processing is not restricted to negative material. Using a similar experimental approach, but employing initially ambiguous auditory sentences that always resolved positively, Holmes, Mathews, Dalgleish, and Mackintosh (2006) found healthy participants in the mental imagery generation condition reported greater increases in positive affect (M = + 7.15, SD = 10.30) than those in the verbal condition (M = -6.08, SD = 9.63), as measured using the Positive and Negative Affect Schedule (PANAS-X; Watson & Clark, 1994). This effect of greater emotional impact of imagery relative to verbal processing has since been replicated in several studies (Holmes, Coughtrey, & Connor, 2008; Holmes, Lang, & Shah, 2009; Nelis, Vanbrabant, Holmes, & Raes, 2012; Pictet, Coughtrey, Mathews, & Holmes, 2011). In addition to transient effects on positive emotion, the impact of positive mental imagery has been shown to persist in the short term. Holmes, Lang, & Shah (2009) asked healthy participants to complete the task described above, and subsequently exposed these participants to a negative mood induction procedure. Participants in the imagery condition experienced a less severe decrement in positive mood following the negative mood induction procedure (PANAS M = -4.90, SD = 6.88) than did those in the verbal condition (PANAS M = -10.90, SD = 8.33). Consistent with research on the buffering effects of positive mood on negative mood (Fredrickson, 1998), Holmes, Lang, & Shah (2009) results suggest that positive mental imagery may have a greater potential to buffer the emotional impact of a negative mood induction than does positive verbal thought. The experimental approach adopted by Holmes and colleagues differs from that employed by Vrana et al. (1986) in two important ways. First, participants in the verbal representation condition were instructed to form meaningful representations of the auditory scenario, requiring them to engage in semantic processing rather than simply engage in verbal repetition of the information, as was the case in Vrana et al.’s approach. Second, subjective emotional response to the information was compared when this was represented in an imagery format and in a verbal format. Thus, while Vrana et al. (1986) were able only to conclude that imagery-based representations exert an impact on subjective emotional experience, Holmes and colleagues have been able to demonstrate that imagery-based representations exert a greater impact on subjective emotional experience than do verbal representations. In the studies described above, the information that participants were required to represent was always presented in a verbal format. Thus, it could be the case that the greater emotional impact of the imagery condition reflects the fact that the information is represented in two modalities, both verbal and imagery, while in the verbal condition the information is represented in only a single modality (Pictet & Holmes, 2013). To address this issue, an alternative experimental paradigm, the Picture-Word Task (PWT), was developed (Holmes, Mathews, Mackintosh, & Dalgleish, 2008). In the PWT, cue items consist of both a visual image and a negative or benign word caption (e.g., a picture of a person in a lake accompanied by the word sink or swim). Participants are asked to combine the picture and the word using either imagery (“form a mental image that combines the next picture and word”) or verbal statement (“form a sentence that combines the next picture and the word”). Thus, in this design, the stimulus materials are comprised of image-based and verbal information. Results from the PWT were consistent with the previous finding that imagery-based representation evoked a stronger impact on subjective emotional experience than verbal representation of the same information. Specifically, Holmes, Mathews, Mackintosh, & Dalgleish, 2008 found that using mental imagery to combine negative picture-word cues resulted in greater increases in anxiety (M = + 7.06, SD = 5.84) than did the use of verbal processing to combine these cues (M = + 2.43, SD = 3.01); and using mental imagery to combine benign picture-word cues resulted in a greater reduction in anxiety (M = -3.81, SD = 5.94) than did the use of verbal processing to combine these cues (M = + 1.0, SD = 2.22). Therefore, on the basis of the reviewed evidence, it is appropriate to conclude that imagery-based processing of emotional information has a greater impact on subjective emotional experience than does verbal processing of the same information. In relation to Lang’s original postulation that only mental imagery has the capacity to activate physiological and behavioral response systems, evidence from studies on anxiety suggests that more verbal forms of thinking, such as worry, can also impact cardiovascular (Brosschot, Pieper, & Thayer, 2005; Llera & Newman, 2010) and neural activity consistent with a negative emotional response (Oathes et al., 2008). We also note here that worry has been defined as “a chain of thoughts and images, negatively affect-laden and relatively uncontrollable” (Borkovec, Robinson, Pruzinsky, & DePree, 1983). SECTION IV: HARNESSING MENTAL IMAGERY’S POWERFUL EMOTIONAL IMPACT IN CLINICAL PRACTICE -------------------------------------------------------------------------------- When Lang's (1979) bio-informational theory of emotional imagery was developed, behaviorism was the dominant approach within clinical psychology. At this time, mental imagery was used to evoke emotion during imaginal exposure treatment of anxiety disorders to help clients learn new emotional and behavioral responses and train self-control (Goldfried, 1971) via graded imaginal desensitization (Wolpe, 1958) or rapid flooding to feared stimuli (Wolpe, 1973). Developments in cognitive and social psychology in subsequent decades has shown that an individual’s thoughts can exert a strong influence on his/her behavior and wellbeing. As a result, the content and phenomenology of an individual’s mental imagery-based thoughts have themselves become the subject of analysis and modification (A. T. Beck, 1970; Lazarus, 1968). Building upon Judith Beck’s cognitive therapy imagery formulations (J. S. Beck, 1995), new mental imagery techniques have been brought into the clinic. The following section will briefly review contemporary cognitive and behavioral (CBT) therapeutic techniques that have built upon the legacy of Lang’s early work, which include: (a) imaginal exposure, (b) the direct modification of the content of aversive imagery-based thoughts, (c) the promotion of adaptive imagery, (d) metacognitive reappraisal of imagery, and (e) imagery-based cognitive modification of maladaptive thinking habits. Imaginal exposure (IE) treatments are still widely used in contemporary clinical practice, with a strong evidence base for treating anxiety disorders ranging from specific phobias (Craske, Antony, & Barlow, 2006), obsessive-compulsive disorder (Abramowitz, Franklin, & Foa, 2002), generalized anxiety disorder (Zinbarg, Craske, & Barlow, 2006), to PTSD (Foa et al., 2007). Seminal advances in mechanistic understandings of fear extinction learning and habituation have been informed by the legacy of Lang’s bio-informational theory and experimental work (e.g., Lang, Melamed, & Hart, 1970). Specifically, it has been proposed that for extinction learning and habituation to occur in imaginal exposure therapy, the imagined feared stimulus itself must evoke fear processing structures in the brain, including both stimulus and response networks (Foa & Kozak, 1986). Evidence for this central mechanism of change comes from the numerous subsequent studies showing that patients who exhibit initial physiological fear responses and subsequent habituation of such responses benefit more from imaginal exposure therapy than those who do not (Craske, Sanderson, & Barlow, 1987; Jaycox, Foa, & Morral, 1998; Kozak, Foa, & Steketee, 1988; Mueser, Yarnold, & Foy, 1991). There is increasing evidence that the efficacy of IE, at times combined with other behavioral or cognitive treatments, is superior to psychopharmacology for OCD (Foa et al., 2005) and PTSD (Van Etten & Taylor, 1998). For GAD, IE-based CBT has been found to be more effective than nondirective and relaxation-based techniques (Borkovec & Costello, 1993). Interestingly, a growing body of research has used virtual reality simulations during exposure therapy to treat a range of anxiety disorders (for reviews see Opriş et al., 2012; Morina, Ijntema, Meyerbröker, & Emmelkamp, 2015). Preliminary evidence suggests that exposure treatments using mental imagery simulations (IE) versus virtual reality simulations of feared stimuli have comparable efficacy for fear of flying (Rus-Calafell, Gutiérrez-Maldonado, Botella, & Baños, 2013) and public speaking (Wallach, Safir, & Bar-Zvi, 2009). In addition to anxiety disorders, researchers have also successfully used IE to reduce intrusive mental imagery in major depression, with promising results in terms of therapeutic benefits (Kandris & Moulds, 2008). Imagery rescripting techniques designed to directly modify the content of emotion-inducing mental imagery have been developed, helping clients transform their unwanted distressing imagery into more benign forms (Hackmann, Bennett-Levy, & Holmes, 2011). Such techniques have been combined with exposure treatment for a range of disorders, from PTSD and personality disorders (e.g., Arntz & Weertman, 1999; Butler & Holmes, 2009; Long & Quevillon, 2009; Smucker & Dancu, 1999/2005), to snake and social phobia (Hunt & Fenton, 2007; Wild, Hackmann, & Clark, 2007). For PTSD, evidence suggests that imaginal exposure therapy combined with imagery rescripting is more effective than imaginal exposure therapy alone (Arntz, Tiesema, & Kindt, 2007; Grunert, Weis, Smucker, & Christianson, 2007). Imagery rescripting has also been used to reduce intrusive mental images in healthy populations (Rusch, Grunert, Mendelsohn, & Smucker, 2000) and in people with depression (Brewin et al., 2009). Rescripting of intrusive imagery associated with negative memories has shown initial promise as a stand-alone treatment for major depression (Brewin et al., 2009). For a detailed review of imagery rescripting treatments, see Arntz (2012) and Holmes, Arntz, and Smucker (2007)). In addition to the development of clinical techniques intended to attenuate negative imagery, clinical investigators have also sought to develop methods of fostering the generation of positive imagery, to help clients build more adaptive relationships between self, others, and the world (Hackmann et al., 2011). Symbolic positive imagery has been used to help clients access adaptive emotional states, in interventions such as Compassion Mind Training (CMT; P. Gilbert, 2009), which was inspired by work on the perfect nurturer image by D. A. Lee (2005). Other positive imagery techniques include guiding clients to construct and road test ideal ways of being in personality disorders (Mooney & Padesky, 2000) or to form mental images of positive scenarios to counter the types of negative beliefs observed in personality disorders, eating disorders, and obsessive-compulsive disorder (Korrelboom, de Jong, Huijbrechts, & Daansen, 2009). Positive imagery can also be used to attenuate or replace negative imagery in vivo. In social anxiety, generating and holding in mind positive images of one’s own appearance, to displace negative images of oneself appearing red and sweaty, has been shown to reduce anxiety and improve social performance (Hirsch, Mathews, Clark, Williams, & Morrison, 2003). Positive imagery has also been shown to function as an effective distractor for chronic pain patients (Fors, Sexton, & Gotestam, 2002). A third category of imagery-focused interventions aims not to directly alter potentially distressing negative mental imagery but rather to change how the patient thinks about these images, in ways that reduce their emotional impact. Metacognitive strategies aim to down regulate the impact of distressing mental imagery via reappraisal (Gross, 2002). Techniques include encouraging patients to perceive mental images as subjective phenomena (Wells, 2003) rather than meaningful premonitions or signs that one’s mind is out of control (Starr & Moulds, 2006; Williams & Moulds, 2007). Another approach is mindfulness- based cognitive therapy, in which clients are trained to regard their verbal thoughts and mental images as mere passing mental phenomena that do not require a response (Segal, Teasdale, & Williams, 2002). More recent developments in experimental clinical research have begun to develop mental imagery-focused cognitive training informed by cognitive bias modification (CBM) paradigms. CBM aims to reduce maladaptive imagery styles and boost adaptive imagery via computer-based training (Blackwell et al., 2015; Williams et al., 2015). Such mental imagery training approaches are informed by research that has illuminated psychopathology-linked individual differences in imagery, which will be discussed in more detail with the next section of this review. Hence, clinical practice is increasingly harnessing the power of mental imagery to facilitate adaptive emotional functioning. The efficacy of such approaches may be further enhanced by future research designed to illuminate the potentially critical role of variability in the strength of the emotional response, observed during imagery in the clinic, in determining treatment outcomes. Given that the elicitation of both subjective and physiological emotional responses during imaginal exposure is a central mechanism of change in the successful treatment of fear, the effectiveness of other treatments that use mental imagery to optimize emotional functioning may also depend on the client’s emotional response during imagery. In our view, research investigating differences between treatment responders compared to nonresponders in the emotional impact of mental imagery may prove fruitful in informing future translational research in this field. SECTION V: INDIVIDUAL DIFFERENCES IN MENTAL IMAGERY AND EMOTIONAL PSYCHOPATHOLOGY -------------------------------------------------------------------------------- So far, this review has discussed mental imagery in terms of its capacity to evoke emotional responses, the presence of disorder-linked variations in this emotional response, and the ways that clinical practice has harnessed mental imagery to facilitate emotional processing. However, one relatively neglected aspect of mental imagery that is of potential clinical significance is differences across individuals in the ability and tendency to experience emotional imagery. Here, “ability” is simply defined as the deliberate generation of imagery, and “tendency” as the nondeliberate (i.e., spontaneous) generation of imagery. This section will review research on individual differences in emotional mental imagery tendency and ability in both healthy and clinical populations. An interesting finding from Lang’s experimental work on the emotional impact of mental imagery is the observation that the level of emotional response experienced during imagery is related to how vividly an individual can generate mental imagery in general (imagery generation ability). Lang and colleagues found that, in healthy participants, heart rate acceleration during fearful imagery was more pronounced in individuals with high imagery ability, relative to those with low imagery ability, as assessed by a self-report questionnaire (Miller et al., 1987). However, such a relationship between self-reported imagery ability and physiological arousal during fear imagery is not always observed. McTeague, Bradley, and Lang (2002) found no association between imagery ability and physiological reactivity during emotional imagery, whether such reactivity was indexed by autonomic responses (heart rate, skin conductance), facial EMG (corrugator and orbicularis), or startle blink reflex. Relatively few studies have examined individual differences in mental imagery and its relationship to imagery-based treatment outcomes in clinical populations. McNeil et al. (1993) found that physiological reactivity and subjective reports of distress were greater in individuals with higher relative to lower self-reported imagery generation ability. However, this effect was only observed in individuals with specific phobias (e.g., dental), and not in those with social anxiety. In chronic PTSD patients undergoing prolonged imaginal exposure treatment, Rauch, Foa, Furr, and Filip (2004) found that participants’ self-reported anxiety and imagery vividness were correlated, and both decreased significantly during treatment. However, while subjective anxiety change was related to treatment outcome, imagery vividness change was not significantly related to treatment outcome (Rauch et al., 2004). Given these mixed results, it presently remains uncertain how robust a difference there exists in the intensity of emotional responding to imagery, between people who exhibit good and poor imagery ability, and how these differences relate to imagery-based treatment outcomes. The resolution of issue clearly requires further research. In recent years, researchers have begun to investigate whether bias in the relative ability to generate mental imagery of emotionally negative compared to emotionally positive scenarios may be associated with emotional psychopathology. Several studies have used the Prospective Imagery Task (PIT; Holmes, Lang, Moulds, & Steele, 2008; Stöber, 2000) to assess imagery vividness. In this task, participants are required to generate mental imagery cued by negative and positive sentences depicting self-referential future scenarios (e.g., negative scenario: “You will be the victim of crime”; positive scenario: “You will have lots of energy and enthusiasm”). Results from such studies indicate that, compared to healthy controls, individuals with depressed mood (Holmes, Lang, Moulds, & Steele, 2008), and those with a diagnosis of clinical depression or anxiety (Morina, Deeprose, Pusowski, Schmid, & Holmes, 2011) report lower imagery vividness (i.e., visual clarity) for positive scenarios. Participants with depressed mood (Holmes, Lang, Moulds, & Steele, 2008) and anxiety disorders (Morina et al., 2011) also report greater vividness of imagery for negative scenarios compared to healthy controls, though this latter effect is not observed in clinically depressed participants (Morina et al., 2011). In addition, researchers have begun to move beyond cross-sectional data and to examine temporal relationships between imagery vividness and emotional psychopathology. Blackwell et al. (2015) conducted a randomized controlled trial in which 150 depressed adults received either a positive imagery intervention involving repeated generation of mental imagery involving pleasant daily activities, or a nonimagery “sham training” control condition, both delivered via the Internet over 4 weeks. Results indicate that the more vividly participants in the positive imagery condition could imagine the positive training scenarios at baseline, the greater their reduction in depression symptoms over the course of the intervention (Blackwell et al., 2015). Future studies should more directly assess the relationship between changes in imagery vividness and changes in symptomology. Researchers have also started to explore disorder-linked individual differences in the tendency to spontaneously experience emotional mental imagery in daily life. Negative self-concept in unselected undergraduate students has been found to be positively associated with the frequency of retrospectively reported negative spontaneous cognitions, consisting predominantly of imagery-based memories of past negative experiences (Krans, de Bree, & Moulds, 2015). It has also been shown that, compared to formerly-depressed and never-depressed individuals, depressed individuals report experiencing spontaneous imagery-based memories, which were predominantly negative in content, as more vivid, evoking greater subjective distress, and causing greater disruption to daily life (Newby & Moulds, 2011). Future emotional psychopathology-linked individual differences research may benefit from obtaining convergent measures of emotional response during mental imagery, such as measures of physiology and hemodynamic neural activity. In addition to indices of emotional response during imagery, electroencephalography (EEG) measures of cortical brain electrical activity may provide indices of the level of cognitive resources expended during imagery as an additional measure of the impact of emotional imagery. One recent study from Lang and colleagues examining alpha-band activity during emotional mental imagery suggests EEG may provide an objective measure of neural activity involved in semantic, motor, and perceptual representations of emotional stimuli. Bartsch, Hamuni, Miskovic, Lang, and Keil (2015) found that while mental imagery elicited by pleasant and unpleasant verbs prompted equivalent levels of alpha amplitude, both were higher than that observed during imagery of emotionally neutral verbs. Alpha-band (8-12 Hz) activity is associated with internally directed attention and may be used to index individual differences in emotional imagery- related brain activity in translational and clinical research. Could it be that individual differences in imagery-based processing of verbal information causally underpins vulnerability to emotional psychopathology, and/or functionally contributes to its maintenance? So far, research linking individual differences in emotional mental imagery generation and emotional psychopathology has been correlational in nature. More experimental research is required to investigate if and how the ability and tendency to experience emotional mental imagery is causally related to emotional wellbeing and dysfunction. Several mechanisms have been proposed and require further investigation. One hypothesis is that mental imagery acts as an emotional amplifier and exacerbates states of mood instability, such as in bipolar disorder (Holmes, Geddes, Colom, & Goodwin, 2008). From a treatment perspective, enhancing the ability to mentally simulate positive scenarios may causally relate to improved emotional functioning via cognitive and behavioral mechanisms such as enhanced negative mood repair, and/or increased anticipatory pleasure and approach motivation for daily activities. The outcomes of such future research will inform how best to exploit our increasing understanding of individual differences in mental imagery in ways that can enhance the efficacy of our therapeutic interventions for clinical anxiety and depression. SECTION VI: NEGLECTED INSIGHTS FROM LANG’S WORK ON MENTAL IMAGERY EMOTIONAL RESPONSE TRAINING -------------------------------------------------------------------------------- A critical yet largely overlooked finding from Lang’s (1977, 1979) experimental work is the demonstration that the emotional impact of imagery can be modulated via training. Lang’s psychophysiology studies provided reliable evidence that physiological arousal during mental imagery of fearful scenes can be increased in two ways. One is via the inclusion of somatovisceral response information in the verbal scripts used to cue imagery (e.g., “your heart is pounding”). The other is via the positive reinforcement of participants’ verbal reports of mental imagery containing somatovisceral response during an initial training phase (e.g., “I felt myself running down the hill fast . . . my heart was beating”). Several studies by Lang and colleagues have demonstrated that such imagery training amplified participants’ ability to elicit physiological responses during imagery (Lang, Kozak, Miller, Levin, & McLean, 1980; Lang et al., 1983; Miller et al., 1987). As discussed earlier, the effectiveness of imaginal exposure therapy rests partly on the presence of initial emotional arousal to the imaginal feared stimuli. Thus, training that enhances emotional response during negative imagery may improve fear-extinction treatment outcomes. Furthermore, training that can increase the emotional impact of mental imagery may serve to enhance the effectiveness of positive imagery interventions for clinical patients who exhibit deficits in positive imagery. As described earlier, the cognitive bias modification approach has recently been combined with mental imagery training to target interpretation bias in both mood and anxiety disorders. Imagery-based interpretation bias modification (IBM) training involves repeatedly generating imagery in response to verbal sentences and picture-word pairs depicting hypothetical everyday scenarios that resolve positively. Experimental studies have shown that imagery-based IBM training of this nature improves mood to a greater degree than does verbal-based IBM training (Holmes, Lang, et al., 2009; Holmes et al., 2006; T. Lang, Blackwell, Harmer, Davison, & Holmes, 2012). To date, several randomized-controlled clinical trials have been conducted, with mixed preliminary findings. While results from two such trials indicate positive imagery IBM training to be a promising and accessible Internet intervention for improving mood and alleviating symptoms in depression (Williams, Blackwell, Mackenzie, Holmes, & Andrews, 2013; Williams et al., 2015), another trial found no support for the positive IBM imagery condition being better than a non-imagery-focused, noninterpretation training, control condition (Blackwell et al., 2015). It seems reasonable to suppose that the efficacy of this imagery-based IBM training would be further enhanced by increasing the positive emotional impact of the trained pattern of positive imagery. This objective may be served by integrating the somatovisceral response amplification techniques, developed by Lang and colleagues, into imagery-based IBM programs designed to increase the occurrence of positive imagery, which represents an important avenue for future research in this area.","Charting the progress of emotional mental imagery research across the past five decades since Lang’s (1977) seminal paper has clearly illustrated the major impact of his bio- informational theory and experimental research on this field. Following Lang’s footsteps, both experimental and translational researchers have continued to be fascinated by the special relationship that exists between mental imagery and emotion. In emphasizing mental imagery’s capacity to activate perceptual stimulus representations as well as physiological and behavioral responses, Lang’s theory and experimental work provided crucial insights into the field of behavior therapy. Specifically, physiological fear response during imagery of feared stimuli has been identified as the marker of successful learning during imaginal exposure and habituation. In light of this finding, evidence reviewed in Section I showing blunted physiological fear response in patients suffering more chronic anxiety and mood dysfunction represents a serious concern in the clinical field. Specifically, it suggests that such individuals may require additional support to enhance physiological responding in order to maximally benefit from the corrective effects of imaginal (or in vivo) exposure therapy. Given that the success of imaginal exposure therapy in facilitating fear-extinction hinges upon the patient exhibiting initial physiological arousal to the imagined fear stimuli, the efficacy of other CBT treatments using mental imagery to facilitate emotional processing may also be enhanced by taking steps to ensure that emotional responses to mental imagery are maximized. Future clinical translational and treatment efficacy research may well benefit from using training techniques pioneered by Lang to keep affect “hot” during critical moments in therapy. Recent individual differences research has begun to illuminate how bias in the relative ability to voluntarily generate emotional imagery, and bias in the tendency to spontaneously experience emotional imagery, differentiates those experiencing emotional disturbance from healthy individuals. This line of experimental research has important clinical implications, as an excess of unwanted emotional mental imagery requires therapeutic alteration. Further, the inability to generate positive imagery, or inability to respond to negative imagery, has treatment implications for depression and exposure therapy. Particularly promising is experimental research that investigates how the relationship between mental imagery and emotion can be exploited for clinical benefit via cognitive bias modification training. Future research on positive mental imagery training could be enhanced by incorporating Lang’s method of emotional imagery response training, using imagery-eliciting scripts that contain both stimulus and response information, combined with training procedures that positively reinforce emotional responding during mental imagery. Furthermore, measuring both subjective and neurophysiological indices of emotional response during positive mental imagery training may provide additional information in identifying treatment responders in future research. Nearly 40 years ago, Lang’s seminal work on emotional imagery laid the foundation for a new era of experimental and clinical research on this fascinating topic. While subsequent researchers have been able to build upon this firm foundation in ways that have greatly expanded knowledge and understanding, it is remarkable that Lang’s original questions, and the findings produced by Lang and his collaborators, have remained of such central importance within the burgeoning contemporary literature. We are confident that this will continue to be the case across the years that lie ahead, and we are honored to have been given this opportunity to celebrate, and pay tribute to, Lang’s extraordinary scientific and clinical legacy.","The authors declare that there are no conflicts of interest."],["Background Although characterised by motor impairments, children with Developmental Coordination Disorder (DCD) also show high rates of psychopathology (anxiety, depression, low self-esteem). Such findings have led to calls for the screening of mental health problems in this group. Aims To investigate patterns and profiles of emotional and behavioural problems in children with and without DCD, using the Strengths and Difficulties Questionnaire (SDQ). Methods and procedures Teachers and parents completed SDQs for 30 children with DCD (7–10 years). Teacher ratings on the SDQ were also obtained from two typically-developing (TD) groups: 35 children matched for chronological age, and 29 younger children (4–7 years) matched by motor ability. Outcomes and results Group and individual analyses compared parent and teacher SDQ scores for children with DCD. Teacher reports showed that children with DCD displayed higher rates of emotional and behavioural problems (overall, and on each subscale of the SDQ) relative to their TD peers. No differences were observed between the two TD groups. Inspection of individual data points highlighted variability in the SDQ scores of the DCD group (across both teacher and parent ratings), with suggestions of elevated hyperactivity but comparably lower levels of conduct problems across this sample. Modest agreement was found between teacher and parent ratings of children with DCD on the SDQ. Conclusions and implications There is a need to monitor levels of emotional and behavioural problems in children with DCD, from multiple informants. --------------------------------------------------------------------------------","In this study, we present a detailed investigation of emotional and behavioural problems in children with Developmental Coordination Disorder (DCD), using parent- and/or teacher- report versions of the Strengths and Difficulties Questionnaire (SDQ). We used both group and individual analysis, which enabled us to compare teacher-ratings of children with DCD to typically developing children (those who were matched for age, as well as younger children matched for motor ability), and to each other. Results demonstrated that there was variability in the SDQ scores of DCD children (across both parents and teacher ratings), but also some broad patterns; for example, individually, children with DCD tended to show high levels of hyperactivity, but comparably lower levels of conduct problems. For children with DCD, levels of agreement between parent and teacher ratings on the SDQ were modest. This suggests that information on emotional and behavioural problems in DCD should be collected from multiple informants.","Developmental Coordination Disorder (DCD, sometimes referred to as dyspraxia) affects between 2 and 6% of children (American Psychiatric Association [APA], 2013; Lingam, Hunt, Golding, Jongmans, & Emond, 2009) and is characterised by motor skills that are significantly below age-expected levels, persisting despite opportunities to acquire and develop these skills. These motor impairments must: have a significant impact on activities of daily living and academic achievement; occur early in development; and not be better accounted for by an alternative explanation (e.g., general medical conditions, intellectual disabilities, visual impairments) (APA, 2013). There are several reasons why children with DCD may present with emotional and behavioural difficulties. Despite being of average or above average intelligence (APA, 2013; Sumner, Pratt, & Hill, 2016), children with DCD often experience problems with school-related tasks (e.g., handwriting, organising their workload, completing tasks on time) (Zwicker, Missiuna, Harris, & Boyd, 2012). DCD also negatively affects leisure participation (Zwicker et al., 2012), meaning that children may become less likely to engage in group activities with peers (Chen & Cohn, 2003), potentially leading to social isolation and loneliness (Missiuna, Moll, King, Stewart, & MacDonald, 2008; Poulsen, Ziviani, Cuskelly, & Smith, 2007). Further, high rates of psychopathology – including anxiety (Pratt & Hill, 2011) as well as depression and low self-esteem (Lingam et al., 2012; Piek et al., 2007) – have been reported in children with DCD. DCD also commonly co-occurs with other conditions, such as attention- deficit-hyperactivity disorder (ADHD), which is often associated with emotional and behavioural problems (Missiuna et al., 2014). There have been calls for the screening of mental health problems in children with DCD (Rigoli & Piek, 2016), with the Strengths and Difficulties Questionnaire (Goodman, 1997) being suggested as a suitable tool for assessing possible psychosocial problems; both generally (Goodman, Ford, Simmons, Gatward, & Meltzer, 2000) and in the DCD population (Rigoli & Piek, 2016). Using the parent-report version of the SDQ in a sample of 47 children with DCD, Green, Baird, and Sugden (2006) found that 62% of children with DCD showed ‘clinical’ levels of emotional and behavioural difficulties (13% = ‘borderline’, 15% = ‘normal’).1 Further, 85% of the sample showed ‘significant’ problems in at least one of the five SDQ subscales (Emotional symptoms, Conduct problems, Hyperactivity, Peer problems, Prosocial behaviours). Using the teacher- report version of the SDQ, Van den Heuvel, Jansen, Reijneveld, Flapper, and Smits- Engelsman (2016) reported children with DCD (n = 23) to have significantly greater emotional and behavioural problems than typically developing (TD) (chronological age matched) children. However, the proportion of children showing ‘clinical’ levels of the Total difficulties scores (36%) was much lower than the 62% reported by Green et al. (2006). Indeed, mean scores across all subscales of the SDQ were lower in Van den Heuvel et al.’s (2016) sample, relative to Green et al.’s (2006) sample. This could be due to Green et al. (2006) recruiting their sample from a clinic, whereas Van den Heuvel et al. (2016) recruited their sample by screening large numbers of children and identifying those with significant motor impairments (from a community-based school sample). Alternatively, it could be due to the studies differing in their use of parent- versus teacher-report, with teachers potentially rating the children’s difficulties as less severe. This may be because teachers are less familiar with each child’s capabilities (relative to the parents), therefore underestimating the child’s difficulties. Or, it could be because teachers have a greater understanding of what typical performance is (due to working with a large range of children) and are, therefore, less likely to overestimate any difficulties. Indeed, a review of the psychometric properties of the SDQ highlighted only modest agreement between parent- and teacher-reported scores on the SDQ (Stone, Otten, Engels, Vermulst, & Janssens, 2010). The aim of the current investigation was to explore emotional and behavioural difficulties using the SDQ in a sample of children with a confirmed clinical diagnosis of DCD. First, we sought to confirm previous reports of high levels of emotional and behavioural difficulties amongst children with DCD by comparing teacher SDQ ratings of children with DCD to two groups of TD children: (1) a group matched by chronological age (hereafter ‘CA’ group); and (2) a group matched based on motor ability (motor-match, hereafter ‘MM’ group). The latter group was comparable to the DCD group in terms of performance on a motor task but was, inevitably, younger than the DCD group. Comparisons between these two groups provide an indication of whether the observed profile of children with DCD reflects a level of immaturity, to some extent. The second aim, focusing on the DCD group only, was to investigate levels of agreement between parent- and teacher-report on the SDQ (unfortunately, we were not able to collect parent- reported SDQ data from the TD children, to also explore this comparison in the CA and MM groups). A meta-analysis comprising 14,811 children between the ages of 3–17 years (from a range of typical and clinical populations), reported correlations between parent and teacher SDQ ratings to be between 0.26 and 0.47 (Stone et al., 2010). As such, only “modest” agreement was predicted in the current study. However, adopting group and individual analyses to explore this research question allowed more detailed analyses than has been undertaken in previous research. Further, it enabled us to explore individual profiles of emotional and behavioural problems across the DCD group.","As part of a broader study exploring the cognitive and behavioural profiles of children with DCD (see Sumner, Hutton, Kuhn, & Hill, 2016; Sumner, Leonard, & Hill, 2016), 30 children with a diagnosis of DCD (21 boys, 9 girls, all aged 7–10 years) were recruited through primary schools, as well as advertisements via a charitable organization (the Dyspraxia Foundation, UK). Prior to taking part in the study (and independent of the research study itself), children had received a diagnosis of DCD from a multi-disciplinary team of clinicians who were external to the research team. The second edition of the Movement Assessment Battery for Children (MABC-2; Henderson, Sugden, & Barnett, 2007) was used to confirm a DCD diagnosis, and all children scored at or below the 16th percentile on this measure. Additionally, on an initially screening questionnaire, parents confirmed that there was no history of additional diagnoses or medical conditions that might explain the child’s motor difficulties. The CA group comprised 35 children (26 boys, 9 girls, aged 7–10 years), whilst the MM group comprised 29 children (19 boys, 10 girls, aged 4–7 years), all recruited from primary schools in South London. In a screening questionnaire, their parents reported no identified diagnosis of a neurodevelopmental condition, including DCD. The MM group were screened based on the time taken to complete a peg placing task as part of the MABC-2 (in which they had to place 12 pegs into a board as quickly and accurately as possible, using both their preferred and non-preferred hands); as reported in Sumner, Hutton et al. (2016) and Sinani, Sugden, & Hill (2011), who have adopted similar approaches to matching. Raw scores were used for the motor matching process (i.e., number of seconds taken) of children in the MM and DCD groups. After screening peg placing time, all children in the MM group completed a standardised assessment of their fine and gross motor skills to determine age-appropriate motor skills (at or above the 25th percentile on the MABC-2), detailed below. Similarly, all children in the CA group had to score at or above the 25th percentile on the MABC-2. Background characteristics of the groups are presented in Table 1. This included a measure of parental education, which has been used as a measure of socio-economic status in similar studies (Fernald, Marchman, & Weisleder, 2013; LeBarton & Iverson, 2013), and was found to be comparable across the groups. Wechsler intelligence scale for children (WISC-IV) and wechsler preschool and primary scale of intelligence (WPPSI-IV) The WISC-IV (Wechsler, 2003) and WPPSI-IV (Wechsler, 2012) were used to determine inclusion in the study, which required a Full-Scale IQ (FSIQ) >70 for all groups. FSIQ calculated from the WISC-IV comprises four indices and ten subtests (items per index shown in brackets): verbal comprehension (3 items), perceptual reasoning (3 items), working memory (2 items), and processing speed (2 items). From the WPPSI-IV, FSIQ is calculated from five indices and 6 subtests: verbal comprehension (2 items), visual spatial (1 item), fluid reasoning (1 item), working memory (1 item) and processing speed (1 item). Participants completed all subtests. The DCD and CA groups completed the WISC-IV, while the younger MM group completed the WPPSI-IV. The psychometric properties of these tests have been established from large, representative samples which confirm good reliability (including internal consistency above 0.88 for the four indices of the WISC-IV, and above 0.89 for the WPPSI-IV indices; and test-retest stability above 0.86 for the WISC-IV, and above 0.84 for the WPPSI-IV). Movement Assessment Battery for Children (MABC-2) Children completed the age-appropriate assessments (age band 1: 4–6 years; age band 2: 7–10 years) of the second edition of the MABC-2 (UK norms; Henderson et al., 2007). Each age band comprises three components: manual dexterity (3 items), aiming and catching (2 items), and static and dynamic balance (3 items). Scores from the eight items are summed to provide a total standard score (Mean = 10, SD = 3) and percentile rank. The MABC-2 was used to confirm the diagnostic status of the DCD (i.e., at or below the 16th percentile) and to confirm age-appropriate abilities in both TD (i.e., at or above the 25th percentile) groups. Across studies that have addressed the psychometric properties of the MABC-2, reliability is considered good. Intra class correlations (ICCs) for test-retest reliability are reported between 0.77 and 0.95, and for inter-rater reliability ICCs have been reported at 0.95 and above (see Henderson et al., 2007, for details). Strengths and Difficulties Questionnaire (SDQ) The SDQ (Goodman, 1997) is a 25-item questionnaire that can be completed by parents or teachers (note: the questions asked are the same in both formats but the opening statement differs very slightly from ‘your child’ to ‘the child’, respectively). It comprises five scales (of five items each): Emotional symptoms (e.g., “Often unhappy, down-hearted or tearful”); Conduct problems (e.g., “Often has temper tantrums or hot tempers”); Hyperactivity (e.g., “Constantly fidgeting or squirming”); Peer problems (e.g., “Rather solitary, tends to play alone”); and Prosocial behaviours (e.g., “Considerate of other people’s feelings”). Ten of the questions are designed to tap the child’s strengths; 14 represent difficulties; and 1 item is neutral. Parents or teachers (depending on the informant) rate each question on a three-point scale (“Certainly True”, “Somewhat True”, “Not True”) and scores of 0, 1 or 2 are assigned (depending on whether the items are positively or negatively phrased). A ‘Total Difficulties’ score (ranging from 0 to 40) is generated by summing scores from all of the scales except the Prosocial behaviours scale (as this reflects positive behaviours). As well as using SDQ scores as continuous variables, scores can be classified into ‘normal’, ‘borderline’ and ‘clinical2’ categories, and risk factors can be determined regarding Emotional disorders, Behavioural disorders and Hyperactive disorders (‘low risk’, ‘medium risk’, and ‘high risk’) (see www.sdqinfo.org). Data from a large, representative sample has confirmed satisfactory reliability (mean Cronbach α = 0.73; mean cross-informant correlation = 0.34; mean retest stability after 4–6 months = 0.62) and validity (as assessed by comparing SDQ scores against independent psychiatric diagnoses) (Goodman, 2001).","Ethical approval was obtained from Goldsmiths, University of London. Written informed consent was provided by all schools and parents/carers, while verbal assent was obtained from the children that took part in the study. Children completed the cognitive and motor assessments in two separate sessions either in their school (for the TD groups) or during a visit to the research lab (for the DCD group). They were seen individually in a quiet room. Parents of children with DCD completed the questionnaire during the visit to the research lab, while teachers of children with DCD were sent a copy of the questionnaire and asked to post the completed copy back to the research team. Teachers of the TD groups completed the questionnaires during school time and returned them to a member of the research team. Statistical analyses Tests of normality and homogeneity were conducted prior to test selection. Parametric (paired samples t-test, ANOVA) and non-parametric equivalents (Kruskall-Wallis) were used to investigate differences at the group level and when comparing teacher and parent responses. Intra class correlations (ICCs) were also used, when comparing agreement between teacher and parent ratings on the SDQ.","Relating to the inclusion criteria for the study, Table 1 presents the background characteristics of the three groups. The DCD and CA groups were significantly older than the MM group. No significant differences were found between the CA and MM groups on FSIQ, but the DCD were shown to have slightly lower FSIQ scores (although still scoring within the average range for this test). FSIQ was not found to be correlated with any of the SDQ measures (ps > 0.18) and, therefore, is not included in subsequent analyses. Finally, the inclusion criteria for motor abilities were met for all three groups. Of note, 2 (7%) children with DCD scored on the 16th percentile, while the remaining DCD participants scored on the 9th (n = 6, 20%) or at or below the 5th percentile (n = 22, 73%). Teacher ratings of emotional and behavioural difficulties in children with and without DCD ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Data screening highlighted that SDQ Total Difficulties scores, as well as three SDQ subtest scores (Emotional symptoms, Conduct problems, and Peer problems), were not normally distributed (skewness values >1). Applying transformations to these data did not result in the assumptions of normality being met, so non-parametric (Kruskall-Wallis) tests were used. For the normally distributed variables (Hyperactivity, and Prosocial behaviour), one way ANOVAs were used. Results demonstrated that teachers rated the DCD group as showing greater levels of emotional and behavioural problems than their TD peers (CA and MM matched) on all subscales of the SDQ, and in their overall scores (see Table 2). As can be seen in Table 2, the TD average scores were low in each subtest, except for in the Prosocial behavior ratings where a higher score reflects more positive, prosocial behaviour. Notably, none of the children in the two TD groups (CA or MM) scored in the ‘borderline’ or ‘clinical’ ranges on the SDQ Total Difficulties component (all scoring in the ‘normal range’). Further, in relation to the subtest scores, teacher ratings demonstrated only one ‘borderline’ case from the CA group in each of the following categories: Emotional symptoms, Hyperactivity and Prosocial behaviour; in addition to, two ‘clinical’ scores on the Hyperactivity scale, three ‘borderline’ scores on the Prosocial behaviour and one on the Peer problems scale from the MM group. Analyses of borderline and clinical cases in the DCD group are discussed below. Comparing parent and teachers SDQ scores in children with DCD ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Focusing on the DCD group only, group analyses compared parent and teacher reports on the SDQ before examination of individual data points allowed us to: (a) look more closely at concordance between parent and teacher ratings on the SDQ; and (b) determine whether children with DCD presented with a specific profile of emotional and behavioural difficulties on the SDQ. As seen in Table 3, a paired samples t-test revealed no significant difference between parent and teacher scores when considering the Total difficulties scores of children with DCD on the SDQ. However, paired samples t-tests exploring SDQ subscale scores indicated two key differences between parent and teacher ratings of the children with DCD. First, parents reported their DCD children’s hyperactivity to be more problematic than teachers did. Second, parents rated the prosocial behaviours of their children with DCD more highly (i.e., less problematic) than teachers did. There was also a non-significant trend towards parents rating their DCD children’s conduct problems to be more severe than teachers did. Across the other subtests, there was general agreement between parent and teacher ratings. Significant correlations between parent and teacher ratings were found for three of the five subscales (Emotional symptoms, Conduct problems and Peer problems), with correlations approaching significance for the Total difficulties score (see Table 3). Classifying DCD children’s scores (overall, and for each subtest) into ‘normal’, ‘borderline’ or ‘clinical’ categories, we explored individual children’s scores on the SDQ for those who had both parent and teacher reported scores (n = 26; one teacher and three parents did not return the SDQs). Table 4 presents the classification for each individual child, on each subtest and overall, as rated by both their parent and teacher: unshaded cells reflect ‘normal’ scores, grey cells reflect ‘borderline’ scores, and black cells reflect ‘clinical’ scores. For Total difficulties scores, as well as scores on each subtest, a tick after the parent and teacher ratings reflects agreement between the informants’ classifications (with ‘agreement’ being classed as both parent and teacher reporting the child’s behaviour to be either ‘normal’ or in the ‘borderline/clinical’ range) and a cross denotes disagreement (i.e., a parent rating the child’s behaviour in the ‘normal’ range and the teacher reporting the child’s behaviour to be in the ‘borderline/clinical’ range, or vice versa). A tally at the end of each row or column indicates the number of parent-teacher classification agreements for each child or across the subscales (including the overall total). Inspection of Table 4 serves to further demonstrate the modest level of agreement between parent and teacher ratings of children with DCD on the SDQ. For example, on the Total difficulties SDQ scores, parents and teachers classified children similarly in 15 out of 26 cases (57.7%). Likewise, classification agreements across the other subtests ranged from 15 to 20 out of 26 (57.7%–76.9%), and classifications for risk of Emotional, Behavioural or Hyperactive disorders ranged from 15 to 19 out of 26 (57.7%–73.1%). A further feature of Table 4 is that it illustrates the variability in the scores of children with DCD across the different subscales of the SDQ. Visual inspection of Table highlights that no discernable pattern emerges amongst the group, aside from fairly severe levels of hyperactivity across the sample (with only 2/26 DCD children showing ‘normal’ levels of hyperactivity, as rated by both parents and teachers), as well as a relative lack of children with the presence of conduct problems (16/26 DCD children showed ‘normal’ levels of conduct problems, as rated by both parents and teachers). DCD children were also rated as having a slightly raised profile of peer problems and emotional difficulties; only 12 DCD children (46.1%) were rated as having ‘normal’ levels of peer problems by both parents and teachers, with this figure falling to just eight children (30.8%) in relation to emotional difficulties.","We used teacher and parent versions of the SDQ to explore reports of emotional and behavioural problems in children with DCD. Teachers reported that children with DCD displayed a higher number of emotional and behavioural difficulties than their TD peers (both those of a similar CA, and those with similar motor abilities). Further, exploring the profiles of emotional and behavioural problems in children with DCD more closely (using individual and group analyses), yielded two key findings. First, variability was noted amongst the SDQ profiles of individual children with DCD; there was a range of different combinations of typical and atypical presentations of emotional and behavioural problems, with some overarching themes (e.g., high levels of hyperactivity, comparatively lower levels of conduct problems). Second, agreement between parent and teacher reports of children’s emotional and behavioural difficulties was modest (as found by Stone et al., (2010), highlighting the need for clinicians to collect information from multiple informants when assessing a child with DCD. The results are consistent with other studies that have used the SDQ in this population; finding high rates of difficulties overall, as well as on the individual subscales (e.g., Green et al., 2006; Van den Heuvel et al., 2016). The present study was unique in that it compared teacher-report SDQ scores of children with DCD against two TD comparison groups – those matched on CA, as well as younger children matched on motor ability (MM). Consideration of the profiles of the two TD groups from teacher ratings demonstrated that all children fell in the ‘normal’ boundary for the SDQ Total Difficulties score, with very few TD children from these two groups showing ‘borderline’ or ‘clinical’ profiles in the subtests. This is in stark contrast to a high number of children with DCD showing the latter profiles. The finding that the CA and MM groups scored equivalently – and showed lower levels of emotional and behavioural problems than the DCD group – suggests that the emotional and behavioural profile of the DCD group is not simply due to motor immaturity. Instead, it may be a repercussion of the core diagnostic characteristics: difficulties experienced in the motor domain (e.g., problems with acquiring proficient fine and gross motor skills having consequences for the development of other aspects of functioning such as social and emotional development). Individual analyses (using SDQ data collected from both parents and teachers) highlighted variability between the SDQ scores of the DCD group, with a range of combinations (of typical and atypical) emotional and behavioural problems noted. Wide variability is perhaps not so surprising in a group often shown to be heterogeneous in nature (Biotteau et al., 2016). Therefore, this method of exploring individual differences is an important approach when interpreting the performance of the DCD group. Despite this variability, there was some suggestion that children with DCD have elevated levels of hyperactivity, with only 2/26 children (7.69%) showing ‘normal’ levels of hyperactivity (as rated by both parents and teachers). This is consistent with studies suggesting a high degree of overlap between DCD and ADHD (Kadesjo & Gillberg, 1998) and highlights that clinicians should be astute to hyperactivity in children with motor difficulties. Individual analyses also showed that nearly two thirds of children with DCD demonstrate no conduct problems – 16/26 children (61.54%) showed ‘normal’ levels of conduct problems (as rated by both parents and teachers), which is consistent with previous research (Gustafsson et al., 2014). Yet, it should be noted that conduct problems were reported in just over a third of this group. Finally, the slightly raised profile of emotional difficulties and peer problems in the children with DCD also confirms previous findings (e.g., King-Dowling, Missiuna, Rodriguez, Greenway, & Cairney, 2015; Wagner, Bös, Jascenoka, Jekauc, & Petermann, 2012). Group and individual analyses demonstrated modest agreement between parent and teacher scores on the SDQ for children with DCD. Whilst previous studies have reported high levels of emotional and behavioural problems in children with DCD, findings from SDQ studies using parent or teacher reports have not yielded equivalent results; for example, teacher-reported SDQ scores of children with DCD (Van den Heuvel et al., 2016) have been reported to be lower than those reported by parents (Green et al., 2006). These previous studies have each used parent or teacher reports in isolation, whereas the present research compared parent and teacher ratings across the same group of children. Adopting this approach provides further, more solid evidence for subtle disparities between parent and teacher ratings on the SDQ. They also confirm previous studies (outside the field of DCD) showing only modest agreement between parent- and teacher-reported scores on this measure (Stone et al., 2010). Analysis of parent and teacher reports also revealed that parents often rated their children with DCD to be more hyperactive than teachers did; however, the opposite was true for prosocial behaviours (with parents often reporting higher levels of prosocial behaviour than teachers). This could be accounted for by children behaving differently in school and at home, or could be related to teachers having a broader (and potentially more accurate) benchmark against which to compare the children. Irrespective of the reasons underlying these differences, it illustrates the importance of professionals obtaining reports from both teachers and parents when using the SDQ with children with DCD. In doing so, any bias ought to be removed, as well as ensuring no information is missing. This, in turn, will lead to a more comprehensive picture of the child’s difficulties. Although a strength of this study is the inclusion of a younger MM comparison group, it is noted that the procedure of matching on a single measure of fine motor skill has its limitations. Future research may consider extending this approach to a larger sample of participants and to consider matching on fine and gross motor skill. In selecting samples, care was taken to exclude children with known co-occurring difficulties (based on parent and clinical reports) so that conclusions related to the core motor difficulty could be made. However, it is recognised that further screening of possible co-occurring difficulties would help further categorise the profiles seen here. In taking this research forward, direct comparison of both parent and teacher SDQ ratings in the TD (CA and MM) groups would be fruitful, to align to the present findings. Interestingly, our data show that teachers do not rate the TD children particularly highly on any aspect of the SDQ; nevertheless, parent perspectives would be valuable here. Overall, the current findings highlight that a large proportion of children with DCD present with problems with attention (hyperactivity). In addition, albeit to a lesser extent, a number of DCD children had a raised profile of emotional difficulties and peer problems. As the sample did not have any confirmed co-occurring diagnoses (e.g., ADHD), it highlights the importance of exploring emotional and behavioural problems in a DCD population, to fully support these individuals. The findings also flag inconsistencies across parent and teacher ratings, stressing the importance of considering both perspectives. Moreover, the variability in SDQ scores across the DCD sample suggests that a tailored approach to intervention is necessary to support the emotional and behavioural needs of this group."],["Emotional intelligence (EI) is argued to predispose individuals to better apprehend and accomplish stressful tasks. Research produced to date has nevertheless mostly neglected the processes that render EI situationally advantageous. In this study, we used path analysis to explore how ability and trait EI relate to delivering a presentation under stress by distinguishing subjective and objective performance. We also proposed self-efficacy as a potential mediator of the trait EI-performance relationship. One hundred and twenty university students completed the STEU, STEM, and GERT (for ability EI) and the TEIQue (for trait EI); then they performed an oral presentation task in front of two evaluators. Students’ self-ratings and evaluators’ scores composed subjective and objective performance. Results indicated that: (a) self-efficacy fully mediated the relationship between trait EI and both subjective and objective performance; (b) ability EI, in particular emotion understanding (STEU), directly predicted objective performance. These findings highlight one way in which the two leading EI approaches may both contribute to performance under stress, but through distinctive paths. --------------------------------------------------------------------------------","Past results are mixed, and yet both trait EI and ability EI appear to predict various forms of performance such as cognitive task performance (e.g., Hui-Hua & Schutte, 2015; Lam & Kirby, 2002), academic performance (e.g., Perera, 2016; Song et al., 2009), workplace performance (e.g., O'Boyle et al., 2011), and performance under pressure (e.g., Laborde, Lautenbach, Allen, Herbert, & Achtzehn, 2014; Lyons & Schneider, 2005). Trait EI appears to predict performance when performance is assessed via both subjective and objective ratings (O'Boyle et al., 2011). One rationale advanced for the positive association between trait EI and objective performance/outcomes is the overall characterization of high trait EI individuals as being less implicitly prone to impulsiveness and more self-controlled (Petrides, 2011). Trait EI thus should cultivate attitudes and behaviors that are normally crucial for facilitating productive persistence in stressful situations. The role of ability EI in predicting performance has been investigated especially in the domain of workplace behavior, with evidence mainly supporting its utility for objective performance (O'Boyle et al., 2011). Being conceptualized as a ‘performance’ measure of EI and hence not linked to subjective perceptions of one's abilities—as also proven by the low correlation between trait and ability EI (e.g., Vesely Maillefer et al., 2018)—ability EI should be associated with objective, rather than subjective, performance. In reviewing the extant literature, we identified three major limitations. First, only few studies have investigated both trait and ability EI as predictors of the same final outcome (e.g., Davis & Humphrey, 2012). This empirical hole makes it hard to ascribe reliably a precise situational utility to each single construct. Second, most frequently, subjective and objective indicators of performance–with a disproportionate emphasis on job performance–have been counted as equivalent (Joseph et al., 2015). Third, while a rationale for a possible mediation effect of self-regulatory process standing between EI and performance outcomes has been theorized (Zeidner & Matthews, 2018), past research has largely neglected its testing (for an exception regarding coping as a potential mediating mechanism, see MacCann, Fogarty, Zeidner, & Roberts, 2011). Consistent with the above critiques, we considered: (a) both trait and ability EI as predictor of performance in a task expected to demand high emotional involvement; and (b) both subjective (self-reported) and objective (other-rated) ratings of performance; and (c) the mediating role of a particular self-regulatory process, namely self-efficacy, in the relationship between trait EI and performance in a stressful task. Self-efficacy, a potential mediator ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Self-efficacy refers to people's judgments of their own capabilities to organize and carry out actions and behaviors required to attain a certain performance (Bandura, 1986). The basic idea is that one's personal dispositions affect one's behavior by activating self- regulatory processes such as self-efficacy (Lent, Brown, & Hackett, 1994). Some EI experts have suggested that trait EI may positively impact performance because of its links to emotional self-efficacy (e.g., Dacre Pool & Qualter, 2012; Joseph et al., 2015). However, to our knowledge there exists no full investigation of the relationship between the three concepts. A few studies only have reported a positive correlation between trait EI and self-efficacy (e.g., Mouton, Hansenne, Delcour, & Cloes, 2013), and another has shown that trait EI predicted self-efficacy perceptions in a condition of stress as compared to a neutral one (Mikolajczak & Luminet, 2008). Yet, ‘subjective’ beliefs about one's capability to perform, we argue, is the most plausible mechanism through which trait EI may impact performance. In contrast, the ability EI-self-efficacy beliefs relationship appears less straightforward. Individuals high in ability EI use their emotional skills to effectively adapt to the situation and ultimately outperform. Performance seems to depend directly on one's actual ability to deal with an emotional situation. The present study ~~~~~~~~~~~~~~~~~ We aimed to test the EI-performance link and our assumption for self-efficacy to act as a unique mediator in two hypotheses Hypothesis 1. The relationship between trait EI and objective and subjective performance will be mediated by perceived self-efficacy in dealing with the task. Hypothesis 2. Ability EI will have a direct effect on objective performance, but not on subjective (self-rated) performance. We did not have specific hypotheses regarding which EI ability would better predict performance, although emotion understanding and emotion management would be more relevant for the task chosen involving high stress.","A mixed group of 120 students enrolled at a University in Switzerland (55 female; 80% of undergraduate, age range: 19-31; Mage = 21.63 and SD = 2.22) were pooled from a larger research sample collected a year earlier. All received 20 CHF in compensation for participation. Design and procedure ~~~~~~~~~~~~~~~~~~~~ Upon arrival in the lab, participants were instructed to prepare, in 10 min, a 3 min oral presentation that they would need to perform immediately after in front of two neutral evaluators. Participants randomly allocated to a high stress condition (n1 = 60, 26 females) were asked to synthetize a complex philosophical text, and those allocated to a low stress condition (n2 = 60, 29 females) to discuss professional experiences and career aspirations. In both conditions, they needed to speak until the end of the 3 min in front of a camera and were recorded. The evaluators seated behind a desk as they listened and rated independently the presentations using an online assessment form. Participants finally completed a post-questionnaire before being debriefed. The experiment lasted about 45 min.","The Trait Emotional Intelligence Questionnaire-Short Form (TEIQue; Cooper & Petrides, 2010) measured trait EI, while the Situational Test of Emotional Understanding-Short Form (STEU; MacCann & Roberts, 2008), the Situational Test of Emotion Management-Short Form (STEM; MacCann & Roberts, 2008), and the Geneva Emotion Recognition Test-short version (GERT; Schlegel & Scherer, 2016) assessed ability EI. A single-item indicator evaluated self-efficacy belief (SE) about coping with the task. Single item measures have been used to assess self-efficacy in previous studies (e.g., Hoeppner, Kelly, Urbanoski, & Slaymaker, 2011). As a measure of subjective performance (Perf-S), participants indicated how well they thought they performed on the task. For objective performance (Perf-O), the evaluators rated each presentation overall based on the criteria: clarity, organization, and coherence of the argument presented; inter-rater reliability calculated using the intra-class correlation coefficient (ICC) was good: .91. The Brief HEXACO Inventory (BHI; De Vries, 2013) and the Verbal Reasoning test from the Kit of Factor-Referenced Cognitive Tests (VR; Ekstrom, French, Harman, & Dermen, 1976) assessed personality and cognitive ability and were employed as controls1 (see supplementary materials 1 for details). Statistical analyses We used the Statistical Package for Social Sciences version 22 (IBM SPSS Statistics 22; SPSS Inc., Chicago, IL) to compute scores on the variables of interest, perform descriptive analyses, and calculate ICC and Cronbach reliabilities. Monte Carlo power analysis for indirect effects was conducted through a modern online application (Schoemann, Boulton, & Short, 2017). To test the fit of the data, path analysis with maximum likelihood estimation was performed with Stata 14 (StataCorp, 2015); path coefficients of direct and indirect effects were estimated using Z-tests for mediation analysis. A model testing both hypotheses 1 and 2 was fitted with the total score of the TEIQue, STEU, STEM, GERT as independent variables, confidence in the ability to cope with the situation as mediator (only for the TEIQue), and subjective and objective performance as outcomes (see Fig. 1). A trimmed model was then created removing the non-significant paths for the sake of parsimony. This model was compared with two full mediation models linking trait EI with each performance type. Only the control variables displaying significant correlations with the mediator or the outcomes were retained for the analyses. We let the error terms of the outcomes correlate in the model. We requested standardized solutions, thus the reported values are beta coefficients. Overall R2 was also computed. The fit indices used were: chi-square test statistic (χ2), the comparative fit index (CFI), the Tucker-Lewis index (TLI), the root mean square error of approximation (RMSEA), and the Standardized Root Mean Residual (SRMR). If the CFI value was .90 or above, the TLI values were above .95, the SRMR and RMSEA values were .08 or less, we considered the model to have a good fit.","In the two conditions, participants rated how stressed they felt before, during, and after performing. Participants assigned to the high stress condition felt more stressed after the task (M = 2.05 SD = 1.02) than those in the low stress condition (M = 1.68 SD = 0.93), F (1, 118) = 4.26, p = .041. Stress felt before and during performance did not differ in the two conditions, with participants reporting quite high levels of stress overall across the two conditions while performing (M = 3.54, SD = 0.99). Due to the similarity of effects of high and low stress conditions, and to low power to conduct mediation analyses with each condition treated separately, we dummy coded experimental condition (0-low stress condition, 1 high stress condition) and added it as a control variable in all models. Descriptive analyses ~~~~~~~~~~~~~~~~~~~~ Table 1 presents the means, standard deviations, and correlations among the variables. Trait EI positively correlated with self-efficacy and subjective performance but not with objective performance, only the STEU did. The STEU, STEM and GERT were correlated with each other, suggesting a common latent variable assessing ability EI. Self-efficacy in coping with the situation was positively correlated to subjective and objective performance. Verbal reasoning, gender, emotionality, agreeableness, and conscientiousness showed no significant relationship with either the mediator or the performance outcomes and were excluded from analysis. Mediation analyses ~~~~~~~~~~~~~~~~~~ The Monte-Carlo power analysis indicated power of 94% for detecting the indirect effect of trait EI on subjective and objective performance. Confirmatory factor analysis (CFA) using maximum likelihood estimation to assess the structural validity of ability EI using the STEU, STEM, and GERT further showed satisfactory fit indices and/or factor loadings >.45, confirming that each scale is indeed an indicator of ability EI.","We experimentally explored how ability and trait EI relate to subjective and objective performance under task-induced stress and proposed self-efficacy as a potential mediator of the trait EI-performance relationship. Self-efficacy mediated the relationship between trait EI and both subjective and objective performance (Hypothesis 1), and ability EI–albeit only in the case of emotion understanding–directly and strictly predicted objective performance (Hypothesis 2). While we did not hypothesize the exact nature of the mediation for trait EI, we note that a full mediation emerged for both objective and subjective performance. The mediating role of self-efficacy ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our findings suggest that the perceived capacity to understand and use emotions is cardinal in mobilizing self-regulatory processes that are powerful enough to influence both subjective and objective performance. In this respect, they corroborate the results of a recent study where trait EI was found to impact human behavior through a similar mechanism (Udayar, Fiori, Thalmayer, & Rossier, 2018). Our findings also reveal that, at least with respect to performance under stress, the contribution of trait EI on subjective and objective performance is fully accounted for by self-efficacy beliefs about being capable to deal with the situation's emotional and stressful components, further corroborating the path through which trait EI may exert positive outcomes. Utility of EI in predicting performance under stress ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Ability EI had a direct effect, and on objective performance only. Among the three ability EI components, only emotion understanding was a significant predictor. Although we expected that emotion management could have played a role given the stressful situation, the finding regarding the prominent role of emotion understanding within the ability EI framework was not particularly surprising based on a number of studies showing its relevance in predicting important outcomes (e.g., Qualter et al., 2019; Yip & Côté, 2013). Factor analysis approaches have also highlighted higher factor loadings for this ability EI component (e.g., MacCann, Joseph, Newman, & Roberts, 2014). Individuals high in emotion understanding received a better external score for their presentation but, as self-raters, they did not judge their individual performance significantly higher than their counterparts lower in emotion understanding. Perhaps this denotes the presence of a form of self-doubt in high ability EI scorers, and/or one of overconfidence in those less skilled in the construct. More certainly, the results indicate that actual abilities and subjective perceptions of performance have little to do with each other. This aligns in logic with the typical low correlation found between intelligence scores and self-reported measures of intelligence (Paulhus, Lysy, & Yik, 1998 cited in Brackett & Mayer, 2003). Our findings also show that like in Lyons and Schneider (2005), high ability EI individuals seem to have used their emotional skills to adapt to the presentation situation and to ultimately outperform. In contrast, trait EI seems to be a better predictor of subjective performance as ability EI had no effect on it. All in all, our results speak in favor of considering ability EI as an asset for elevating performance regardless of the subjective beliefs the person has about such performance. Still, EI training programs promoting either the trait or ability approach could benefit from incorporating the finding that self-mastery beliefs about one's ability to deal with a stressful situation (such as performing well at a job interview) are of practical help for thriving. Our findings, however, support the conclusions of previous studies for envisaging the utility of EI mostly in jobs high in emotional labor (e.g., Joseph & Newman, 2010) as, we speculate, it is in these where task-induced stress would be most exacerbated.","We assessed performance on one task and under the special state of stress. Our findings’ generalizability is thus limited to individuals subjected to similar performance challenges and conditions (Matthews et al., 2006). This being said, many school or work situations require to deliver presentations on subjects the person is not necessarily familiar with, and to articulate a position in front of others with often the quite stressful risk of being judged. Hence, we believe that the experimental condition employed in our study had good ecological validity. We also measured self-efficacy after the task was performed. In principle, how people performed on the task might have affected self- efficacy perceptions. Although the measurement of self-efficacy before performance would have been more effective for showing its mediational effect, we believe that performing the task did not substantially affect self-efficacy perceptions for two reasons: (a) the experimenters were instructed to remain neutral and to not provide any cue about how participants were performing and (b) both conditions were perceived to be rather stressful and therefore difficult (the average level of stress during the task across conditions was 3.54 out of 5); hence, participants did not have a clear sense of how well they performed after the task. Finally, we made the choice to use performance indicators explicitly tailored to our presentation task; an alternative for future research would be to use standardized questionnaires assessing subjective and objective performance, which might increase the validity of our findings.","As research on possible primary (e.g., inability to exploit/express one's full potential) and secondary (e.g., demotivation, poor decision-making) negative effects stress could have on performance continues to advance and gain importance in a plethora of domains (e.g., health care, sport; Lea, Davis, Mahoney, & Qualter, 2019), our findings bring reasonable optimism in revealing, through trait and ability EI, two possible paths that might aid to minimize such detrimental effects. Of course, moving forward, one big question is: how can knowledge about the virtues of each construct be most effectively diffused and strategically implemented in action?","The contribution of the authors was financed by the Swiss National Science Foundation (grant no 100014_165605 awarded to Marina Fiori). +"],["Drawing on discussions with Kenyan, Mexican and British teachers, this paper reports on emotional responses to international socio-economic inequality. Emotional regimes are explored to identify what ‘appropriate’ responses to inequality are in a variety of local and national contexts. These include rural and urban settings, and social milieus ranging from elite to deprived. Politeness, hand-wringing and humour can create a protective distance; while sadness, anger and hope for change connect with the issue of inequality and challenge the associated injustices. Distancing and connecting emerge as central themes in the analysis. The spatial patterns of emotions align with participants' socio-economic positions, in more disadvantaged settings unapologetic anger about inequality was expressed, as was humour in the face of group or national misfortune. These emotional regimes can be understood within the wider context of participants' socio-economic position; their senses of injustice; and their views on the possibility of social change. I argue that social norms surrounding justice and distribution can influence levels of inequality, and vice versa. This is of particular importance given the societal damage caused by inequality, which is now widely acknowledged. --------------------------------------------------------------------------------","Emotions are central to how people are positioned in relation to a topic or situation. Being emotionally engaged may amplify attitudes and provide an impetus for action. In contrast, denial of something being morally problematic may mean not feeling disturbed (Cohen, 2000). Connecting to or distancing from an issue is a key theme in this analysis of secondary school teachers' and trainee teachers’ attitudes towards socio-economic inequality. Emotions are important for understanding the interconnected yet unequal social world, to the extent that neglecting the vocabulary of emotions “leaves a gaping void in how to both know, and intervene in, the world.” (Anderson and Smith, 2001, p.7). Emotions are active elements of public debate on world issues. Emotions, such as fear of terrorism, may be provoked to justify political manoeuvres (Pain, 2009). Negative emotions surrounding inequality may be roused by unmet expectations. These expectations are based on experience of norms of remuneration, capacity to meet basic needs, and level of disposable income, amongst other factors (Hegtvedt et al., 2008). I am interested here in understanding how local socio-economic positioning and norms influence emotional responses to inequality. In particular, I consider the emotional regimes surrounding inequality in three countries that differ markedly in terms of national wealth. It is widely argued that current levels of world inequality, normally taken to imply income or wealth inequality, are unacceptable (Amin, 2006; Dorling, 2010; Ghosh, 2008; Roy, 1999; Sutcliffe, 2005). Economic inequality is closely associated with health, social and educational inequality. It has been argued that greater economic equality would enable a fuller use of human resources, create larger markets for goods, and reduce costs of managing society, such as policing costs (Sutcliffe, 2005; Wacquant, 2010). Many negative outcomes of national inequality in richer countries have been identified, which impact upon the wealthy as well as poorer groups (e.g., Wilkinson and Pickett, 2009). Being richer than others can even lead to feelings of vulnerability and depression. This may be partly due to searching for fulfillment in objects of social status (James, 2007). Thomas Pogge takes a Rawlsian approach to poverty, arguing that we have a responsibility not to cause harm (Pogge, 2008a). Bob Sutcliffe, on the other hand, emphasizes that redistribution is desirable for social justice independent of consequences (Sutcliffe, 2005). These authors provide some responses to inequality embedded in academic debate, a debate that provokes emotional and moral statements in addition to discussion about evidence and theory. This paper presents comparative research on attitudes to world socio-economic inequality. It focuses on how secondary school teachers and trainee teachers from Kenya, Mexico and the UK talk about such inequality. I emphasize the emotional stances imbued in discussions about inequality, and pay attention to how participants position themselves in relation to inequality. The three countries were selected to span a wide range of levels of international inequality, whilst having broadly comparable national inequality (the UK is the more equal society of the three, reflecting a trend of richer countries being more equal than poorer countries; Barford, 2010). The research locations also capture diversity in terms of the countries’ roles in the world economy, geographical location, and regional influences. Exploring how emotional regimes are interrelated with the local, national and international socio- economic positions of research participants offers insight into spatial patterns of emotions surrounding inequality. Characterising the emotional regimes concerning world inequality is important for understanding how people create connections and distances across social and physical divides.","Since the mid-1990s, there has been increased interest in emotions within the disciplines of sociology, psychology, philosophy, and geography (Reddy, 2005; Thien, 2005). More recent interest in emotions within the social sciences stems from the recognition of their political importance. The idea that emotions are regarded as separate from the public sphere and essentially private has been widely critiqued across feminist (or emotional) geography literature, since they are tied up with power relations. For example, social hierarchies are associated with psychosocial stress and status anxieties (Anderson and Smith, 2001; Wilkinson and Pickett, 2009; Routledge, 2012). Social constructionists view emotions as “culturally relevant, public performances, reflecting power relations and mediating between subjective experiences and social practices” (Zembylas, 2007, p.58). Understandings of politics, policies, experiences and attitudes can be improved by considering their emotional dimensions. Emotions can motivate people to act against injustices (Routledge, 2012). In recognition of the active role of emotions, William Reddy coined the term emotives, a word similar to performative, to express how emotions influence the world (Reddy, 2005). He describes how emotional information is conveyed in responses to others in words and facial expressions. Reddy argues that communities establish norms resulting in an ‘emotional regime’, where conformity to preferred emotions is endowed with authority. Social interactionist Arlie Hochschild (Hochschild, 2008) uses a similar concept of ‘feelings rules’. For Hochschild, emotions are based on cultural ‘prototypes’. Particular reactions are expected in response to certain events: one should be thrilled to win a prize, one should be furious when mistreated. As cultures are fluid and interconnected, feeling rules can be interpreted as having local, national, and international influences. These expectations vary between cultures and contexts due to local differences in general standards (Hegtvedt et al., 2008). A constellation of feeling rules contributes to emotional regimes, and both terms are employed in this paper. Attitudes to world inequality are likely to be influenced by spatial variations in emotional regimes and feeling rules. Emotional regimes, which make some responses acceptable and others distasteful or inappropriate, are partly influenced by material conditions. This is because material conditions underlie the procurement of essential goods and luxuries (and social norms of wealth influence what is deemed essential or luxurious). Economic position also influences the cultural norms and values to which people are exposed. Thus, geographies of inequality could bear some similarities to the spatial patterns of emotional responses to inequality. Reddy's and Hochschild's views of the social conditioning of emotions counter the common impression that emotions are involuntary (Anderson and Smith, 2001), and several researchers have documented how emotional expression is consciously controlled. For instance, protesters may avoid angry and violent responses to social injustices so as not to provide an excuse for others to delegitimize their objections (Routledge, 2012). Similarly, in Mexico and the USA, media accusations of being ‘crazy’ or ‘emotionally craven’ have undermined and silenced protest against injustices to women (Wright, 2008). In both cases, dominant groups have encouraged emotional control in an effort to maintain social control. Yet channeling rather than suppressing anger and aggression has enabled protests to persist, whilst conforming to wider emotional expectations (Thrift, 2004). However, a strong response is sometimes necessary to initiate social change. Roland Barthes distinguishes between punctum, as an emotionally charged response that ruptures complacency, and the more common studium, which is a general, polite interest in something (Barthes, 2000). Emotion demonstrates engagement, whereas polite interest suggests an emotional distance. Those carefully managing their emotional responses in order to maintain legitimacy, whilst acting upon strongly held views, negotiate conflicting demands. Norms for emotional performance have been identified as governing emotional labour. This ‘surface acting’ requires people to display emotions that they do not feel (Moore, 2008). For example, retail workers are expected to be cheerful and friendly, whereas judges should be emotionally neutral (Kiely and Sevastos, 2008). Emotional labour at work strains employees (e.g. Nylander et al., 2011) due to dissonance between actual and performed feelings. Some people resolve this dissonance by aligning their own feelings with expected behavior referred to as ‘deep acting’ (Moore, 2008). There is less need for such emotional labour for people of higher status. Generally, people who are powerful and of high status have more positive emotional experiences than people with lower status (Collins, 1990 in Moore, 2008). My interest in the socio-economic context of emotions resonates with a feminist approach that binds everyday emotions to networks of power and privilege within which they are located (Pain, 2009). Economic and other social inequalities are widely understood to be instances of injustice, and so have the potential to cause anger and frustration amongst those experiencing these injustices. Other researchers have demonstrated that those suffering disadvantages experience more distress and anger, and those whose advantages are associated with injustices are more likely to experience feelings of guilt (Hegtvedt et al., 2008). The type and extent of emotional response may reflect how much someone is influenced by an experience or observation, as well as by dominant feeling rules. Distancing is of particular interest given that people and places are now generally understood to be relational and connected to others, which affects their identities and capabilities. The uneven development of places is partly due to their interconnectedness: “The ‘gap’ between the ‘first’ world and the ‘third’ is not just a gap; it is also a connection.” (Massey and Jess, 1995, p.225). Through time, humans have empathised with increasingly large groups, from families to the nation state and beyond (Rifkin, 2010). At a smaller scale, a South African Xhosa proverb ‘umuntu gumuntu ngabantu’ (a person is a person through persons) acknowledges the importance of society to individuals' identities (Raghuram et al., 2009; Shutte, 1993 in Smith, 2000; Therborn, 2009). Given these interrelations and expanding geographical imaginations, distancing other people and places in our imaginations may constitute denial, whereas emphasizing connections may stress responsibility. Anglo-American culture has a propensity to suppress emotion and avoid discussion of responsibility. Emotional suppression was exemplified when ecologist Page Spencer, writing about her grief at the destruction caused by the 1989 Exxon-Valdez oil spill in Alaska, was criticized for her “unprofessional and embarrassingly emotional” accounts (Button, 2010). Similarly, geographer Melissa Wright (2008) received criticism for her protests against injustices to women. It seems more usual for the privileged to address injustice and inequality from an emotional distance. Judging their wealth to be deserved and in the national interest (Pogge, 2008a,b; Rowlingson and Connor, 2010), and denying the severity of poverty (Swaan, 2005) enables the wealthy to approach inequality in an apparently rational and unemotional manner. Misperceiving poverty, perhaps by constructing the poor as lazy and the rich as hardworking (Reis and Moore, 2005; Bamfield and Horton, 2009), also facilitates an emotionally disengaged approach to inequality. Income and life expectancy distances between people are increasing at the world and often country levels (Therborn, 2009). Yet the porous boundaries of places mean that these basic inequalities cannot be understood in isolation (Massey and Jess, 1995). Instead, contemporary inequalities are geographical expressions of the contradictions of capitalism (Smith, 1990). The feelings of the relatively wealthy towards the global poor are better documented than those who are more disadvantaged by inequality. This research aims to access multiple understandings of inequality to enhance appreciation of “the lives that others live partly because of us” (Cook, 2006, p.660). It is within this relational understanding of people and places that emotions towards world inequality are of particular interest. The emotions surrounding these international connections and injustices are central to this paper. Key questions are: How do people feel about their position in an unequal global system? How do physical distance and feelings of connectivity play out in the emotional regimes associated with increasing inequalities? Does one's position within this system influence the form of appropriate emotional responses, and how? What is the geography of the feeling rules concerning world inequality? Country context ~~~~~~~~~~~~~~~ The countries selected for this research are distinctly positioned in terms of national income, meaning that these three countries intersect with world inequality from divergent material situations. On the spectrum of Gross Domestic Product per capita, Kenya has a relatively low income, Mexico is towards the middle, and the UK has a relatively high income. The distribution of income within the three countries is comparatively unequal (poorer countries generally have higher income inequality than wealthier countries). See Table 1 for details. All three countries included in this study enacted neo-liberal policies since the late 1970s. In 1985, Mexico signed the General Agreement on Trade and Tariffs, which opened the economy and removed state subsidies. This was followed by further liberalisation or ‘Salinisation’ under President Carlos Salinas (Hamnet, 2006). This privatization in the 1990s enabled the Mexican Carlos Slim Helú to enter the ranks of the world's richest people (Harvey, 2007). In the 1980s and 1990s, private enterprise strategy and Kenyan capitalists acting as agents of foreign capital seemed to make Kenya an exception to sub-Saharan African poverty (Himbara, 1994). At the same time, under Margaret Thatcher, the UK transformed from a social democratic state, then comparable to Sweden, to one of increased service privatization (Harvey, 2007). Each country connects to a regional identity with distinctive cultures, politics and institutions. Kenyans share in a pan-African identity (Thiong’o, 2009); Mexico and the rest of Latin America are set in contradistinction to the United States along linguistic, economic and political lines (Paz, 2005); and the UK is half-heartedly engaged in European economic and political consolidation. Still, these countries also stand out from their region. Kenya is the East African hub; all three countries have close but dissimilar relationships with the United States. Trends in values and attitudes do vary between world regions, as shown by global values surveys (Pew Research Centre, 2015), so wider regional context is worth considering. Discussion groups ~~~~~~~~~~~~~~~~~ Discussion groups were used to shed light on how inequalities are addressed in social situations, and thereby access the social nature of human knowledge (Goss and Leinbach, 1996). The group setting invites discussion and accommodates open-ended questions that encourage people to talk in their own terms. Working with groups also emphasizes locality (Holbrook and Jackson, 1996), which is especially relevant to research into geographical variations. In order to capture the co-production of emotions in specific contexts, I have employed extended quotations, where appropriate, to show group interaction. Table 2 provides further detail on group composition and setting. The discussion groups were one- off meetings lasting roughly 90 min, a relatively low time-commitment compared to approaches that run multiple meetings with the same group (Burgess, 1996; Kneale, 2001). The discussion guide included seven questions, starting with what inequality means, in order to develop a working definition. Next participants were asked about their awareness of inequalities at the world scale, and then about the causes of inequality. They were asked about the importance of inequality as a world issue. Four visualizations of inequality were used to provoke discussion. Of these, two cartograms are shown in Figs. 1 and 2. Then, positive and negative aspects of inequality were discussed. To conclude the discussion participants were asked to comment on the frequency with which they discuss inequality, and with whom. This gave a sense of how much the ideas that arose during the groups extend into participants’ daily lives (Bedford and Burgess, 2001). Group sizes departed from the standard number of discussion group participants, which is often in double figures (Goss and Leinbach, 1996; Hennick, 2007). Instead, groups ranged from 2 to 8 people. Small group sizes allow more time for each person to speak, and reduce the likelihood of simultaneous conversations that are hard to facilitate and transcribe. With smaller groups it is also easier to recruit participants and gain head teachers’ approval. Discussion groups were recorded and subsequently transcribed. The words spoken (or their English translation), and how they were said, were noted. This included the stress on certain words (indicated by italicized text), tone of voice and other sounds such as sighs and laughter. The discussion groups in Mexico used Spanish. Those in the UK and Kenya were conducted in English (which is the medium of secondary education in Kenya). In Kenya, Kiswahili words were often inserted into English sentences. Learning languages for fieldwork can deepen understandings of different perspectives and increase cultural sensitivity more generally (Watson, 2004). Having studied both Spanish and Kiswahili in preparation for this research, I was able to conduct the groups and translate the recordings myself, checking meaning with others when appropriate.","School teachers were chosen for four main reasons. Firstly they have a wider influence on society, through their potential to expand pupils' sensitivities and awareness. Educational institutions also play a central role in reproducing social structures (Bourdieu, 1996). One teacher referred to her role as teaching children to be responsible citizens: “… You're also teaching them for a wider world in which that inequality will exist, and it will change if they have a different mindset” (urban private school, UK). Secondly, teachers often interact with pupils from diverse socio-economic backgrounds and may deal with inequalities between their students on a daily basis. Thirdly, although teachers are a heterogeneous group whose professional experiences vary considerably, working with participants with a shared occupation renders findings more comparable between countries. Lastly, as ‘global social dialogues’ of social movements and international institutions are already well documented (Yeates, 2009), new insights into global dialogues could come from research into teachers' perspectives given their unique positions in relation to education and schooling. In most cases, discussion groups were recruited from the same school or teacher training college, and most discussion groups took place within the school building. The professional context of the discussion, in terms of the social and physical setting, influenced the roles played by participants. School hierarchies and professional desirability influenced the views and emotions expressed, and probably minimized the number of disagreements that arose between participants, requiring emotional labour from some group members. Recruiting people who know one another can reduce anxiety about involvement and debate, as well as facilitate the telling of shared stories (Holbrook and Jackson, 1996). Table 2 details the participants in each group. I mainly recruited teachers of geography, history, languages and social sciences. These teachers had varying levels of experience: from trainees to retirees. They worked in towns and cities as well as rural areas; in schools serving richer and poorer students; in government and private schools. I conducted nine groups in Kenya, eight in Mexico and seven in the UK. The social position of participants ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Local context, such as being in a rural or urban area, or public or private schools, influences both teachers’ experiences and their socio-economic status. Particularly in rural areas of poorer countries there may be a lack of electricity, poor school buildings, minimal support and book shortages, amongst other challenges (Iredale, 1993). In Mexico, teachers are more respected in a rural setting than in cities, due to the status and remuneration compared to other skilled jobs available locally. Rural and urban teachers will be aware of divergent forms of inequality in their daily lives. A state-private divide exists in education systems in all three countries. Private school teachers are likely to accept dominant elites, given their choice of school. Many state school teachers will have some experience of poverty amongst pupils. Since the 1990s, Mexico opened its education system to the private sector due to World Bank pressure to reduce “non- productive expenditure”. Private schools taught 13.4% of lower secondary and 21.4% of upper secondary pupils by 1999 (Delgado-Ramos and Saxe-Fernández, 2009; Marchant, 2005). In the UK, the Conservative governments from 1979 to 1997 framed parents and pupils as consumers of education. ‘Consumer choice’ has led to prioritisation of good results over pupils' needs (Hill et al., 2009). Analysis ~~~~~~~~ To analyse discussion group data in a discourse analysis tradition I focused on what words allow the speaker to do, how some topics are prioritised, and considered the grammar and word choices (Wood and Kroger, 2000). Self-definition implicit in discussion was used to evaluate how interviewees positioned themselves and others in terms of socio-economic inequality. This is because there are multiple possible interpretations of status depending on reference group. It is almost always possible to find someone better or worse off than oneself, and to feel relatively privileged or disadvantaged as a result. Emotions were analysed according to how participants related to a privileged or disadvantaged self- categorisation. Status and emotions were identified by the words used and manner of expression, in order to identify the geography of the emotional regimes employed and the logic behind them. I worked with emic themes arising from the data, to make less expected findings possible (Goss and Leinbach, 1996). A reading, re-reading and cross-group comparison formed the basis for analysis. Codes were formed, altered and merged throughout. At times I mapped themes diagrammatically to create an overview of conceptual interconnections (Burgess, 1996; Kneale, 2001; Miles and Huberman, 1994). Ultimately, I collapsed emotional responses into those that emphasize connectivity between socio- economic groups, and those that create a distance. These responses infer something about spatiality and intimacy. Distance is physical, with apparent barriers, as well as a personal disentanglement from other people and places. By contrast, connection emphasizes political, economic, social or other forms of proximity, even over large physical distances. Descriptions of connection also express a level of personal engagement with the experiences of others. These approaches were analysed in relation to how respondents expressed their socio-economic status. Comparison between groups highlighted absences and presences and I identified several discursive repertoires and emotional treatments of inequality and considered how the groups relate to such repertoires (Jackson, 2001). These findings are best understood as case studies that illustrate variations and patterns of understandings of inequalities in particular locales, positioned differently in the same world order. Due to the small and unrepresentative sample, the findings cannot be considered representative of the countries in which the research was undertaken (Creswell, 2007). It is possible that cultural context or variations in sense of humour, amongst other things, could lead to my misinterpretation of some emotions. Nevertheless, multiple types of information from a discussion group can be used to aid a good interpretation of feelings. Respondents’ words, body language, tone, and other sounds such as giggling or sighing can be read in conjunction to triangulate the emotions expressed. Researcher positionality ~~~~~~~~~~~~~~~~~~~~~~~~ Postcolonial responsibility requires acknowledging the partial, embedded, political and messy nature of research and writing (Noxolo, 2009; Jazeel and McFarlane, 2010). This research was carried out within the context of the global inequalities that it seeks to challenge. On leaving the UK, my home country became a more significant part of my identity. Like others (Skelton, 2001), I tried to disassociate myself from British transnational politics, immigration laws and trade policies. Inequalities often existed between me as the researcher and the participants. Some teachers working in poorer contexts could well have found it insensitive for a privileged British student to ask about inequality, although this reaction was not openly communicated during the discussion groups. Some of the teachers working in richer contexts, where inequality usually is not openly critiqued, probably thought me to be radical due to the same questions. On first sight I was assumed to be a Gringo in Mexico, due to my physical appearance. I am white, female and undertook this research aged 26–27. Being British may have freed me from negative stereotypes of United States Americans whose holiday homes have proliferated in towns such as Chapala in Jalisco. On the other hand, the UK's foreign policies have often met with controversy. In Kenya I was a Muzungu, a racial and economic category. Typically, Muzungus receive a warm welcome, always pay for their Kenyan friends, and enjoy going on safari. In Kenya and Mexico there was surprise at my age and occupation: in both countries PhD students are usually over 30 years old and already have families (I had no children at the time). Within the UK, my accent gives away my middle-class background and often (mistakenly) leads people to think I come from the South of England. I attempted to manage facets of my identity whilst recognising Rose's (1997) point that we cannot fully understand, control or redistribute power. My identity influenced participants’ responses to my questions, as well as my analysis and interpretation of the data. Responses were influenced partly by a wish to maintain social desirability in front of me, the researcher, as well as other group members. Further, ideas and attitudes may have been expressed in particular ways in order that I would understand. Especially in Kenya and Mexico, it is likely that participants gave more explicit explanations to ensure my understanding. The themes that interested me, and the patterns that I discovered in the transcripts, were influenced by my own interest in and understanding of the causes of inequality. Thus aspects of my identity combined with my sensitivities as a researcher to influence the results produced by this research.","Below I demonstrate and discuss how emotional regimes relate to socio-economic inequality. My analysis shows how respondents tend either to connect or distance themselves from the issue of inequality. Connection and distancing have emotional and conceptual implications. Emotional distancing is being more cut off from an issue, whereas connecting is being emotionally engaged. Conceptually it means whether respondents describe themselves as being interconnected with others. These associated responses are influenced by whether the respondents consider themselves to be comparatively privileged or disadvantaged. Speaking of the global and local poor generally led the teachers to position themselves as privileged. Reference to wealthier nations, or the rich within their own society, resulted in participants self-defining as disadvantaged. UK groups rarely position themselves as disadvantaged. Certain emotions were associated with these positions and distancing approaches, and these intersecting themes structure the results section (see Table 3). Generally there was agreement within groups, illustrating the co-production of emotional responses. Sometimes emotional responses were led by a dominant group member with especially strong views. In Mexico Group 6 (rural government school) there was one dominant, politicized and critical male teacher. His discourse was followed with approval, e.g.: “He's taken the words from our mouths. We totally agree with our colleague.” Disadvantaged, connecting ~~~~~~~~~~~~~~~~~~~~~~~~~ Sometimes being disadvantaged whilst connecting with inequality took the form of outrage or disgust, and a sense of manipulation. Sometimes hope was expressed, at other times powerlessness and resentment were the dominant emotions. In many cases there was a strong emotional engagement with inequality, akin to Barthes’ concept of punctum (2000). “At the world level there are countries that are supposedly the powerful ones, and they are the ones that are almost moving the world. The smaller countries are those that are doing nothing more than depending on other countries. So, there is a lot of inequality.” “We're nothing more than their game of chess!” (Mexico 7, small rural school) The outrage at world inequality communicated is evident in word choice. The last quote here suggests disempowerment and resentment, but these words were spoken boldly and critically. The speaker challenges this situation by emphasizing the nature of this unequal relationship. The teachers in this Mexican village school saw few opportunities for themselves. One teacher suggested marriage to a foreigner as a way to progress socially. I found that rural groups within Mexico were particularly emotionally engaged with inequality. In small communities poverty and insecurity are not anonymous, but experienced by friends, family, neighbours and pupils. Thus socio-economic disparities may be more conspicuous and felt particularly deeply. The term ‘feeling inferior’ was mentioned three times by Kenya 2 (high-achieving urban government school). Kenya 2 and Mexico 6 (rural government school) use the terms “inferior”, “unequal”, and “terrible” to describe inequality. This differs from the qualified, apologetic tone of the later quotations from privileged groups. Kenya 2 positioned themselves as being disadvantaged at the world scale. “I think something else which also brings up world inequalities, is lack of finance. … other countries have a lot of finance to exploit all the resources that they have. So we are left behind and we feel as if there is no equality. But if we could have finance to exploit the resources at times we might be on a par with the others …” “and even issues of political ideologies, you see like we always adopt, if you look at the world, even like the developing countries the kind of political ideologies are foreign. They try to adopt them and try to use them to run their own affairs in those countries. And some of these things brings about a serious problem of inequality because some of the ideologies for instance, look at the way Kenya got its independence, we inherited a British kind of system, and this was a colonial system and so we had our own people come in and continue to perpetuate the system of colonialism and that creates inequality in the country.” (Kenya 2, high-achieving urban government school) The speakers object to poorer countries’ lack of financial and political control. The second speaker builds on the first. Both argue that uneven connections between places reinforce inequality and disempower poorer nations. Colonialism is blamed for causing national level inequality; this was a common theme in the Kenyan discussion groups. The explanation that lack of finance creates inequality paradoxically gives some hope. Through identifying the structural cause of the problem a possible solution becomes apparent. The first speaker is able to imagine a way in which greater equality could be achieved. Hope is the most positive emotion identified in relation to inequality: “I live in the slum, and I don't like that kind of life … I keep thinking, ‘how can I change things? How can I move out of this you see, and have that?’ What, I talk about it with people, ‘for how long shall we continue living in this situation?’ … ‘So what can I do to change this?’ Not just for me but for all of the other people who are living around me.” (Kenya 6, NGO-funded slum primary teachers) The speaker, in their mid- twenties, is dissatisfied but hopeful about the future. The speaker is a slum dweller who engages with inequality through talk of improving things for others as well. Solidarity and optimism are evident. The deprived setting meant that school meals might be the only food pupils receive. Following the focus group I was shown sacks of food donated by international organisations. This teacher's emotional expression, of hope for the future and dissatisfaction with the present, may be a well-practised interaction with international visitors (especially those involved in aid or charity work), a performance of emotional labour by a recipient of aid. This quotation contrasts with the politicized engagements cited earlier, yet still engages seriously with inequality whilst following a set of feeling rules. Anger and outrage at the structural causes of inequality are not appropriate responses to the donor community, not part of the role of being a grateful recipient. Emotional labour is needed to construct such politically acceptable responses, which perform the function of easing interactions between the privileged and disadvantaged, and facilitating a beneficial transfer of resources. Privileged, connecting ~~~~~~~~~~~~~~~~~~~~~~ Privileged participants also engaged with inequality as a serious and important issue. Emotional work was involved in acknowledging the difficulties of others, whilst not offending other group members in the privileged social context. This balancing is demonstrated below, where one female participant recovers her composure after an outspoken criticism of inequality. “But I do have moments when I think ‘God, this is, I cannot live with this, this is awful, how can this be’, you know. And then obviously you do, you're not actually affected by it, and I think maybe I'm just being a middle class white woman having a little bit of a worry, and then I'll buy something Fairtrade and it will be ok. But you know I do feel it personally to be quite difficult.” (UK 5, urban private school) This teacher has a strong emotional reaction against inequality. She emphasizes the necessity of change, describing inequality as awful. She then qualifies her own views using her white middle-class and female position to imply that she may be more emotional (female) and perhaps less in touch with local social conditions (white and middle class). Dismissing her feelings as “a little bit of a worry” tones down her response. This softens her response, bringing it closer in line with a more emotionally neutral approach. Perhaps this is a defensive move based on previous criticism for strong emotional responses. This speaker was not challenged for her feelings about inequality during the group discussion, instead her self-regulation appears to be a response to wider social norms and feeling rules governing how inequality is discussed. How some respondents express their critical and emotionally engaged approach, knowing that this is unconventional in their society, reminds us of the partial authority of feeling rules. Another British participant, with a strong interest in inequality, represents a small but vocal minority arguing against inequality within the UK. “If the problem is inequality, which I think it is, you can't just look at the third world and say, you know, ‘we've got to do something about the third world’. No, we've got to do something about the first world and the second world as well. Because they're, for me, they're equally problematic. … if I have to walk through Manchester and see somebody scrabbling around in the dustbin to find food, my life is worse. I mean I know that their life is worse (said quickly), but my life is worse. We've got to get to recognise that.” (UK 4, retired urban teacher) This participant emphasizes connections between people. This is achieved partly by identifying that the problems of inequality need to be addressed worldwide. Inequality, referred to here in terms of poverty, is engaged with emotionally. When speaking her voice rose, and she was insistent when stating “my life is worse”. By commenting that poverty is bad for all of society, this respondent conceptually positioned herself within inequality. That she did not qualify her criticisms or soften her opinions may be due to her lifelong commitment to tackling inequality desensitizing her to dominant feeling rules. Over the years she will have co-constructed an alternative set of emotional responses, building a critical culture to challenge inequality. In the following quotation a teacher describes how privileged pupils responded to the disadvantage of others. “… I took a group of students, year 11, 12 and 13 to, er, Mother Theresa Centre, missionaries of charities, she rescues the lowest of the low in society, people who have been abandoned by families, people that are paralysed, people that are deformed, mentally challenged and so on, and when the kids went there I assumed at that level they were psychologically ready to go in. But when we went in, they broke down and were totally disorientated …. So, yes, they know there is inequality, but they don't digest, but yeah, they don't get, they've got to be in contact with it to understand what it means.” (Kenya 7, urban British-system private school) The privileged pupils at this school are a mixture of international and home students. They are wealthy in the Kenyan context, and some will join a global elite. The initial obliviousness to the reality of poverty and disadvantage illustrates how wealthy people buffer themselves from other social groups. This participant recommends ‘being there’ to disrupt complacency, using a language of embodied understanding. Not understanding is described corporeally (or physically) as not “digesting” or incorporating new information, implying a superficial awareness. This relates to my earlier observation that the two rural Mexican groups were particularly emotionally engaged in the discussion of inequality. Nairobi has a broader spectrum of income groups than rural Mexico, yet this urban environment separates pupils from many other city dwellers. The teacher positioned herself as privileged. During the discussion she gave examples of how she connects across social distance. This included providing meals for a poor community and highlighting disadvantage to her pupils. The sense of privilege and connection resulted in tangible actions, a demonstration of Routledge's (2012) description of emotional engagement with injustices leading to action. Disadvantaged, distancing ~~~~~~~~~~~~~~~~~~~~~~~~~ The main way in which disadvantaged people emotionally distanced themselves from the inequality was through humour. Several participants in Kenya joked about poverty and the challenges experienced at the local or at national level. “And why are we living in this situation while the other people have enough so that they can even throw it? Like the politicians who come with the helicopter and just throws money [giggles]” Followed by group laughter (Kenya 6, teachers in poor urban area) The brazen behavior described was of some politicians flying over the Kibera slum, throwing money in an attempt to win votes. Group laughter confirms that these comments were funny, and demonstrates that humour is an appropriate emotional response in this context. Laughter releases tension, keeping the discussion light-hearted. This laughter may stem from two causes. Firstly, having such wealth and poverty together is so preposterous that it becomes funny. Secondly, participants feel powerless to change such entrenched wealth differences. The helplessness provokes a reaction and laughter is often easier than anger to handle socially. These teachers live and work in a particularly poor community, so have first-hand experiences of the injustices associated with inequality. Collective laughter creates an emotional buffer. Another example comes from comments about the world cartogram showing the distribution of people earning over US$ 200 per day (Fig. 1). On this map most of Africa shrinks into a thin black line. South Africa is visible due to the very high earners living there. The following dialogue is about this map: “[Chuckles] we're in real problems.” “There's a strip, a black one” [chuckles] “A black strip, of Africa” [laughs] Anna: “That's where Kenya is, in the black line” Anna: “Why are you laughing?” “Because it is not there [laugh], it is not seen.” (Kenya 8, rural government school) Like the previous quotation, this humour is not derisive of others, but is a response to the group's circumstances. Kenya is described using the personal pronoun ‘we’. Laughter builds throughout this dialogue, with almost every speaker chuckling or laughing as they describe the map. That Kenya is not visible on a map of high earnings could be funny due to a similar combination of ridiculousness and powerlessness. Firstly, seeing one's country missing from a world map is bizarre. Unexpected occurrences or actions are often used to provoke laughter, and this map had the same effect (albeit unintentionally). Secondly, that Kenya is ‘in real problems’ is well-established, the map authoritatively reinforces this point. Laughing is a protective response. Privileged, distancing ~~~~~~~~~~~~~~~~~~~~~~ Humour featured as a way of emotionally distancing the privileged from inequality: “Without inequality, I mean I, we would, we would all be the same, we'd all be the same [yes] who's going to do, you know, different types of jobs, [yes] you know, what it is that you aspire to.” “Actually that's a REALLY good point, like in Aldous Huxley's Brave New World, everyone's actually, so what they do is they actually engineer people so that they are equal [particularly] because if you have an entire society made up of incredibly bright, intelligent but nobody, nobody wants to clean the toilets.” “Well I think that's it really isn't it, and society, and economy need [Yeah.] variation.” “See in intelligence and in the things, in the things that, well let's face it, you want inequality in intelligence so that you can con some people into cleaning the toilets” “I mean my mum is a cleaner and she gains, she is honestly one of the people who is most satisfied with her job.” (UK 6, urban teachers at a private girls school) This conversation depoliticises inequality by using supposed differences in intelligence to justify some people doing undesirable jobs. These teachers distanced themselves from menial work, by asserting their own superior intelligence. The raucous laughter released tension in the discussion, which had grown whilst building an argument against equality. In the context of an elite school, intelligence justifying social disparities offers a socially acceptable explanation. One group member recovers from this laughter, by drawing on the trope of the happy poor. She states that cleaning is a fulfilling job. There were two other noteworthy instances of laughter distancing privileged British participants from inequality (UK 1, UK 3). In both instances the laughter followed statements of their own privilege, expressing uneasiness about this position. Awkwardness about privilege, rather than catalysing a social critique or expression of guilt, followed a pattern of diffusion by humour when group members were of a similar social position. It is likely that the emotional regime would have been altered by a different group composition. Distancing by those who consider themselves comparatively privileged often took the form of non-confrontational interactions. Apparently balanced responses, noting the disadvantages and benefits associated with inequality, emotionally and politically distanced groups from the poverty and injustice associated with inequality. The following responses are to a question about awareness of world inequality: “Obviously people are living in absolute poverty. They don't have anything, don't have enough food and other things they need such as water. And obviously that's not nice and that's what aid charities are trying to tackle I think.” “Saying good things to inequality is difficult because they're WRONG. But at the same time you've got things like cheap clothing that people, you know they want cheap clothing. And I know it's wrong to say, but due to inequality we do get cheaper clothes and things. And it feels wrong but then it's there, it's a fact.” (UK 1, trainee teachers, urban location) After acknowledging the existence of economic inequality and its serious consequences, both speakers distance themselves from the possibility of change. Both comments have similar narrative forms, expressing regret about inequality and distancing themselves from inequality. Locating the possibility for change away from themselves offers a neat solution to the problem, and presents it as self-contained. The second speaker focuses on the present, implying that inequality is immutable. Distancing the causes of inequality from the relatively privileged trainee teachers avoids confrontation, the associated feeling rule is: do not challenge those who benefit most from inequality. Emotions that others might find it hard to respond to, such as guilt, anger or sorrow, are not expressed. Acknowledging the problem and distancing themselves from it positions these trainees as both globally aware yet not accountable for how their lives intersect with those of others. Studying at a University in a wealthy British city could mean these respondents were rarely confronted with deep disparities. Inequality probably felt distant, making distancing easy. The trope of the ‘happy poor’ was used by groups in all three research countries. In Mexico and Kenya, this referred to the rural poor within that country. The argument is that rural subsistence lifestyles require little money. Kenya 1 (urban trainee teachers) suggested that $2 is too much in the countryside. UK 1 (urban trainee teachers) respondents referred to a generalised global poor, warning against patronizing pity when poor people might be happy. The quotation below shows how one Mexican group presented the idea of the happy poor within the district of Chiapas. “The geography of Mexico is one of internal differences. In Chiapas there are people who are happy with $2 per day.” (Mexico 1, urban teachers from different schools) The happy poor are constructed as distant by the urban, privileged groups cited above. The speakers live at a physical and social distance from the poor to whom they refer. This distancing is inconsistent with their claims to know about the emotional well-being of these distant poor. The distance allows poorer people to be imagined as having fundamentally different needs and desires from the research participants. The Mexican quote above refers to the relatively poor Chiapas district, which is home to many indigenous people. Emotions are contained and participants reason that poverty is not always a problem (see Barford, 2011 for further discussion). Following an emotionally neutral regime, research participants avoid challenging one another's socio-economic position and political stance by curtailing discussion of the difficulties associated with poverty, discussion which could otherwise prompt strong emotional reactions such as anger or shame. The effect of living amidst urban poverty was described by a British teacher working in Nairobi, Kenya. She had become accustomed to seeing slum living conditions: “I don't blink, so you do become a little bit immune to it, because it becomes so normal” (Kenya 7, urban British-system private school). She had adjusted to living surrounded by poverty by accepting it. This created an emotional buffer and protected her from thinking through the implications of poverty. Observing severe poverty on a daily basis and not intervening requires a form of distancing, to block the possible shame, pity or anger that might otherwise occur. Emotional regimes surrounding inequality play several roles: limiting discord, requiring politically correct responses, and protecting the group and speaker from emotional upset. Participants' socio-economic positions influence the way in which calls for social change are expressed, and the appropriate level of emotional commitment to such views. Those who distanced themselves from inequality also tended to justify it; conversely those who connected with the topic and connected with others across social divides were inclined to challenge inequality. In all three countries there were instances of connection and distancing. One's perceived socio-economic position depends upon choice of reference group, and more privileged groups in each country recognised this status. In the UK no groups considered themselves to be disadvantaged, whereas some groups in poorer and rural areas of Kenya and Mexico did identify with poorer segments of society. Perceived socio- economic position, combined with connecting or distancing, influenced whether and how research participants challenged global inequality.","Emotions are said to be “intensely political” (Anderson and Smith, 2001, p.7). Inequality itself is intensely political, because if inequality in resource distribution is understood as problematic, the logical response is redistribution. Politics arise because those with more than their equal share of resources may be unwilling to share what they, and others, consider to be deserved wealth (Rowlingson and Connor, 2010). Whilst some research participants defended and justified inequality, others felt it was unacceptable. As others have argued (e.g., Raghuram, Madge and Noxolo, 2009; Therborn, 2009), a more public and emotionally engaged appreciation of connectivity and relationality could diminish the conceptual, and ultimately socio-economic, distances between people. Future research might explore how changes in feeling rules associated with inequality come about, and how these changes relate to societal change. Self-defined socio-economic position (disadvantaged/privileged) and political approach to inequality (connecting/distancing) influence emotional regimes (see Table 3). An emotional regime guides whether anger, humour, or hope (amongst other options) is an appropriate response to the highly political and morally sensitive topic of inequality. Country context had some influence on whether research participants self-defined as being comparatively privileged or disadvantaged. It was more common for Kenyans and Mexicans to self-define as disadvantaged than UK groups, reflecting disparities in national wealth. However, it is not possible to distill these findings to a list of emotional regimes by country because of the importance of sub- national variations in socio-economic status. The more privileged participants softened or excused their confrontational views and laughter about inequality; in disadvantaged settings unapologetic anger and humour at group or national misfortune about inequality were acceptable responses. Whether research participants appeared connected to, or distanced from, inequality and its consequences is partly influenced by group dynamics; this is evident in the tendency towards consensus within the discussion groups which were composed pre-existing groups of colleagues. The emotional labour of some participants was observable, for example as they followed feeling rules of being concerned, whilst expressing their own anger or distance in a socially acceptable way. The geography of these distancing or confronting emotional techniques varies between people located at different points within the worldwide distribution of resources, and possibly relates to how much people feel they have to gain/lose from redistribution. At the beginning of the Millennium it was asked, “What possibilities are there for developing a geographical agenda sensitive to the emotional dimensions of living in the world?” (Anderson and Smith, 2001, p.8). The findings presented here offer an international perspective on the feeling rules that influence how teachers discuss inequality. Comparing emotional expressions about inequality from diverse social, economic, political and geographical settings offers insight into multiple perspectives on the same topic. Going beyond one-way analyses of how one socio-economic group makes sense of another, the work presented here helps to piece together a more holistic picture of how inequality is understood from multiple vantage points. This research has made space for both the similarities and differences between disparate groups to emerge. Paying attention to emotional expressions, whether people are connecting or distancing themselves from an issue, and how geographical distance and socio-economic inequality intersect with physical distance and relationality, offers purchase on the role that emotions play in our interconnected social, economic and political lives. This methodological approach could be productively applied to other global issues, such as the debate surrounding climate change."],["This study tested the habit discontinuity hypothesis, which states that behaviour change interventions are more effective when delivered in the context of life course changes. The assumption was that when habits are (temporarily) disturbed, people are more sensitive to new information and adopt a mind-set that is conducive to behaviour change. A field experiment was conducted among 800 participants, who received either an intervention promoting sustainable behaviours, or were in a no-intervention control condition. In both conditions half of the households had recently relocated, and were matched with households that had not relocated. Self-reported frequencies of twenty-five environment-related behaviours were assessed at baseline and eight weeks later. While controlling for past behaviour, habit strength, intentions, perceived control, biospheric values, personal norms, and personal involvement, the intervention was more effective among recently relocated participants. The results suggested that the duration of the 'window of opportunity' was three months after relocation. --------------------------------------------------------------------------------","Promoting environmentally friendly behaviours is arguably one of the most difficult behaviour change targets. When people are asked opinions on environmental issues such as global warming, many will express concerns and pro-environmental attitudes (e.g., Eurobarometer, 2014; Ipsos MORI, 2015; Verplanken & Roy, 2013). But when asked what the most important issues are today, the environment usually ends up low in these rankings (e.g., BBC, 2015; Gallup, 2015). Even if events such as hurricanes, flooding, or pollution, happen on the doorstep, the environment remains a distant and nebulous entity for most people (e.g., Lorenzoni, Nicholson-Cole, & Whitmarsh, 2007, Whitmarsh, 2008). Consequently, as is predicted by construal-level theory of psychological distance (Trope & Liberman, 2010), mental representations of “the environment” are not conducive to taking pro-environmental action (Spence, Poortinga, & Pidgeon, 2012). Environmental issues can also be framed as social dilemmas, that is, conflicts between immediate self-interest and longer-term collective interest, which also weaken an individual's motivation to act (e.g., Biel & Thøgersen, 2007; Lorenzoni et al., 2007). Some barriers to pro-environmental action are straightforward, in particular when people are restrained in their options, for instance due to inadequate public transport or limited financial resources. Gifford (2011) discussed a variety of psychological barriers to pro-environmental behaviour, such as judgemental biases, social comparison processes, psychological investments in current behaviours, and mistrust in authorities. In this article we focus on habit as a particular barrier to change. We will argue that while habits are hard to break, finding opportunities where existing habits are temporarily broken may make a behaviour change intervention more effective. Many behaviours that are considered as potential targets for behaviour change in a more sustainable direction, such as transportation, shopping, leisure activities, or water usage, are strongly habitual. Habits are learned dispositions to repeat past responses (Wood & Neal, 2007; Wood & Rünger, 2016). These behaviours are conducted frequently, usually at the same location and time, and are less guided by conscious intent (e.g., Danner, Aarts, & de Vries, 2008; Gardner, 2009; Ji & Wood, 2007; Neal, Wood, Labrecque, & Lally, 2012; Ouellette & Wood, 1998; Triandis, 1977; Verplanken, Aarts, van Knippenberg, & Moonen, 1998; Wood, Quinn, & Kashy, 2002). While the prevalent socio-cognitive models suggest that control of behaviour is anchored in an individual's motivation or willpower (e.g., Ajzen, 1991), when habits are forming some of that control shifts to the environment, that is, to the cues that elicit the habit (e.g., Neal, Wood, & Drolet, 2013; Neal, Wood, Wu, & Kurlander, 2011; Orbell & Verplanken, 2010; Wood & Neal, 2007; Wood, Tam, & Guerrero Witt, 2005). Habits are thus highly automatised behaviours (e.g., Aarts & Dijksterhuis, 2000; Verplanken & Orbell, 2003), or patterns of behaviour (e.g., Kurz, Gardner, Verplanken, & Abraham, 2015; Roy, Verplanken, & Griffin, 2015). This comes with a degree of ‘tunnel vision’, that is, a lack of choice awareness, superficial decision making, and little interest in new information, even if decision makers are explicitly asked to make deliberate choices (Aarts, Verplanken, & van Knippenberg, 1997; Verplanken, Aarts, & van Knippenberg, 1997). The features which thus characterise habit – lack of conscious intent, a shift of behavioural control from willpower to cues, and ‘tunnel vision’ – are making existing habits resistant to change and thus do not bode well for behaviour change interventions. However, it is not always possible to execute a habit. Circumstances may arise or contexts may change which limit or block a habit, perhaps temporarily, and thus require considering alternative courses of action (e.g., Jones & Ogilvie, 2012). For instance, Fujii, Gärling, and Kitamura (2001) studied the effects of a temporary freeway closure on commuters. While habitual car users were likely to take a longer route rather than switching to a more efficient public transport option, some car users did try public transport and, finding out they had overestimated the travel time, continued to do so during the freeway closure. Brown, Werner, and Kim (2003) observed how car users switched to a light-rail option due to temporary parking shortages, and for some this remained a long-run choice maintained by the positive experiences. Verplanken, Walker, Davis, and Jurasek (2008) found that university employees who had recently moved house and were concerned about the environment were commuting more sustainably than those who were equally concerned, but had not relocated, suggesting that the relocation might have temporarily activated important environmental values (cf., Gatersleben, Murtagh, & Abrahamse, 2014; Verplanken & Holland, 2002). These studies suggest that when habits are broken, this may create a “window of opportunity” for behaviour change. Change may occur spontaneously, for instance by discovering better options than the old habits, as supposedly was the case in the studies cited above. But this window may also be used strategically to promote behaviour change. Behaviour change interventions may thus be more effective when delivered in the context of major habit disruptions, such as those related to life course changes. This has been put forward as the habit discontinuity hypothesis (Bamberg, 2006; Verplanken et al., 2008; Walker, Thomas, & Verplanken, 2015). Major discontinuities may involve transitions to new phases in life (e.g., from education to a job), geographical or physical changes (e.g., residential or work-related relocations), or changes in the environment where habits are executed (e.g., infrastructural changes). Such discontinuities may force people to renegotiate ways of doing things, create a need for information to make the new choices, and a mind-set of being ‘in the mood for change’. Interventions that capitalise on these conditions may thus be more effective compared to interventions under default conditions. A number of studies have investigated the effects of behaviour change interventions that were intentionally delivered in the context of a discontinuity. Bamberg (2006) provided residents who recently had relocated with a 1-day free public transport ticket and information about the available public transport services. The intervention induced a significant increase in the use of public transport compared to a control group of relocated residents who did not receive an intervention. Thøgersen (2012), in a secondary analysis of an intervention study in which participants were given a free one-month public transport pass, found that the intervention was only effective among participants who had recently moved house or work place. Walker et al. (2015) followed workers of an organisation which had relocated and initiated a sustainable travel plan in the wake of it, and demonstrated how old habits decayed and new habits established. While the studies cited in the previous paragraph produced results that are in line with the habit discontinuity hypothesis, they did not provide a test whether the discontinuity itself had a distinct role. In other words, these studies demonstrated that interventions delivered in the wake of a discontinuity were effective, but did not contrast the effects with a default condition in which participants did not go through a discontinuity. The present study aimed to provide such a test in a field experiment in a middle-large city in the east of England. The study included participants who had, versus had not, recently relocated, as well as an intervention versus no-intervention control group in both segments. The intervention consisted of face-to-face interviews and the provision of information about sustainable choices. The outcome consisted of self-reported frequencies of twenty-five environmentally relevant behaviours, which were assessed at baseline and eight weeks later. The hypothesis was tested that higher frequencies of behaviour are reported in the intervention versus control group eight weeks later, but that this effect is stronger when participants had recently relocated. The effects were controlled for key determinants of environmental behaviour at baseline; past behaviour, habit strength, behavioural intention, perceived behavioural control, biospheric values, personal norms, and personal involvement (e.g., Steg, van den Berg, & de Groot, 2014; Steg & Vlek, 2009). Past behaviour obviously served as benchmark for change. Existing habit strength was included, as this might influence the resistance to change (Lewin, 2008/1946). Intention and perceived control represented the most proximal predictors of behaviour in the theory of planned behaviour (e.g., Ajzen, 1991), and thus covered the motivation to behave environmentally friendly and the perceived ability to do so, respectively. Biospheric values, personal norms, and personal involvement represented broader motivational, normative and identity-related factors which have been found related to related to pro-environmental behaviour and behaviour change (e.g., Göckeritz et al., 2010; Sparks & Shepherd, 1992; Stern, 1992; Thøgersen & Ölander, 2002; Verplanken & Holland, 2002; Whitmarsh & O'Neill, 2010). Participants and design ~~~~~~~~~~~~~~~~~~~~~~~ Participants were recruited among residents of Peterborough, a city in the east of England with approximately 186,000 citizens. A total of 1612 individuals were cold-contacted at the doorstep; 800 (49.6%) were willing to participate in the study.1 Half of the participants were known to have relocated within the previous 6 months (“movers”). These households had been identified through property websites and contacts with developers who had been active in the recruitment areas. The remaining 400 participants were recruited from the same areas (“non-movers”). Movers and non-movers were matched on house size (number of bedrooms), home ownership, recycling facilities, and access to public transport. Participants were assigned to an intervention or a control condition. In order to avoid neighbours being assigned to different conditions, a clustered randomisation procedure was applied, in which geographical units were designated as intervention and control areas, respectively. The study comprised two measurements (T1 and T2, respectively), which were approximately eight weeks apart. The measurements consisted of questionnaires, which were handed out upon recruitment and sent by post, respectively. From the original 800 participants at T1, a total of 521 (65%) submitted a completed questionnaire at T2. The final sample contained 330 females (63%) and 191 males (37%). Ages ranged from 19 to 85 years, M = 41 years. Participants received a £10 cash voucher and a lottery ticket for a £250 a prize draw for submitting the final T2 questionnaire. The study received approval from the Ethics Committee in the authors' department. Procedure and the intervention ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The field work (recruitment, delivery of the intervention and data collection) was conducted by the Peterborough Environment City Trust. This organisation had developed an intervention to promote sustainable behaviours among residents who had recently relocated. This intervention was adapted and delivered in the intervention condition. Participants in the control condition only completed the T1 and T2 questionnaires. The intervention consisted of a Personal Interview; a selection of sustainable items (“Sustainable Goodie Bag”); tailored and general information (“Green Directory”); a Newsletter. The intervention was targeted at a wide range of environmentally relevant behaviours, including water conservation, waste reduction, reducing car use, and saving gas and electricity. The intervention contained individualised as well as generic information. Multiple motives were addressed, which concerned ecological values such as preserving natural resources, as well as individual benefits such as financial savings (cf., Steg, Bolderdijk, Keizer, & Perlaviciute, 2014; Whitmarsh, 2009). The individualised information was tailored in the Personal Interview, and partly tailored in the Green Directory, which also contained generic information. The Sustainable Goodie Bag and the Newsletter conveyed generic information. Personal Interview Upon agreement to participate, participants were asked to fill out a questionnaire, which formed the T1 baseline assessment. The project assistant then conducted a personal interview. The responses provided in the questionnaire were used as the basis for this conversation. The purpose of the interview was to identify behaviours the interviewees were interested in changing, potential barriers to change, and possible solutions. The project assistants were trained to use one or more of the following intervention tools: (1) addressing perceived obstacles and barriers, such as lack of information or skills; (2) providing details of obtaining financial benefits, for example by saving water, electricity, gas or fuel; (3) emphasising long-term environmental benefits for humans and the environment; (4) setting and committing to behavioural goals; (5) emphasising a “green identity”; (6) emphasising pro-environmental injunctive and descriptive norms; (7) enhancing or maintaining engagement with and attention to the ecological agenda. The project assistants kept notes of which behaviours were addressed specifically during the interview. “Sustainable Goodie Bag” Participants were offered a free re-usable shopping bag containing sustainable products. The bag contained eco-washing liquid, vegetable and flower seeds, a bus timetable, a shower timer, and a set of brochures on environmentally friendly choices. “Green Directory” An information booklet was sent out shortly after the Personal Interview. Information for each participant was selected on the basis of their expressed interest and/or lack of awareness of issues during the Personal Interview. The Directory also provided generic information by referring to websites on how to live sustainably, and emphasised both environmental as well as the financial benefits of saving resources. Newsletter Participants received twice a Newsletter from the Peterborough Environment City Trust. The Newsletters contained a variety of information about sustainable solutions and provided links to relevant websites. Assessments ~~~~~~~~~~~ Behaviours were assessed both at T1 and T2, while all other measures were taken at T1. Relocation status Participants were asked how long ago they had moved to their current address. This was recorded in terms of weeks, months, and/or years. The time participants lived at the current address varied from one week to 32 years. A variable indicating relocation status was constructed by applying a log transformation on the number of weeks participants had lived at their current address. Behaviours Participants were asked how frequently they had performed twenty-five environmentally relevant behaviours during the last year (T1). These items were presented again eight weeks later (T2) with reference to the same time frame. The choice of behaviours was informed by behavioural goals formulated by the UK Department for Environment, Food, and Rural Affairs (Defra, 2008). The behaviours broadly covered the domains of water (e.g., taking less than 10 min in the shower; using the toilet dual flush), waste (e.g., using re-usable shopping bags; using leftover food for other meals), transportation (e.g., walking or cycling short journeys; ecologically friendly driving), and energy use (e.g., turning down the heating; washing clothes at cooler temperatures).2 Frequencies of performing these behaviours were reported on 5-point scales, which were labelled “never” (1), “seldom” (2), “sometimes” (3), “often” (4), “always” (5), respectively. Because the internal reliabilities of the four behavioural domains were unacceptable (i.e., Cronbach Alpha < 0.50), the twenty-five behaviours were aggregated. Cronbach Alpha for the collective behaviours was 0.77 and 0.84 for the T1 and T2 assessments, respectively. For each participant the behaviours were thus averaged into a T1 and T2 aggregated behaviour index, respectively. High scores indicate high frequencies. Habit strength Habit strength was assessed by a shortened version of the Self-Report Habit Index (SRHI; Verplanken & Orbell, 2003). In order to keep the response load within acceptable limits, the SRHI was applied to the four behavioural domains, which were labelled as, “Using less water”, “Producing less waste”, “Reducing the car less for short journeys”, and “Reducing gas and electricity use”, respectively. For each of these categories six items from the original twelve items contained in the SRHI were presented, Each set of items started with the stem “[Behaviour X] is something …”, which was followed by the six items; “… I do frequently”, “… I do automatically”, “… I do without thinking”, “… that is part of my daily routine”, “… is typically me”, and “… I have been doing for a long time”. The items were chosen such that the key features of habit (the experience of repetition and automaticity) were represented (Orbell & Verplanken, 2015). Responses were reported on 5-point scales labelled as “strongly disagree” (1), “disagree” (2), “undecided” (3), “agree” (4), “strongly agree” (5), respectively. Cronbach Alpha for the four behavioural domains varied between 0.96 and 0.97. Across all items Cronbach Alpha was 0.94. For each participant the SRHI responses were averaged into an aggregated habit index. High scores indicate strong habits. Behavioural intentions For each of the four main behavioural categories participants were presented with three intentions, e.g., “In the next six months I intend to conserve water”; “I expect that I will conserve water in the next six months; “I am not really intending to conserve water in the next six months” (reverse scored). Responses were reported on 5-point scales, which were labelled “strongly disagree” (1), “disagree” (2), “undecided” (3), “agree” (4), “strongly agree” (5), respectively. Cronbach Alpha for the four behavioural domains varied between 0.75 and 0.82. Across all items Cronbach Alpha was 0.86. For each participant the behavioural intention responses were averaged into an aggregated intention index. High scores indicate strong intentions. Perceived behavioural control For each behavioural category participants were presented with three items assessing perceived behavioural control (e.g., “I would find it easy to conserve water”; “Cutting back on my water consumption would not be hard to do”; “I don't really know how I could conserve water” – reverse scored). Responses were reported on 5-point scales, which were labelled “strongly disagree” (1), “disagree” (2), “undecided” (3), “agree” (4), “strongly agree” (5), respectively. Cronbach Alpha for the four behavioural domains varied between 0.63 and 0.77. Across all items Cronbach Alpha was 0.82. For each participant the perceived behavioural control responses were averaged into an aggregated perceived control index. High scores indicate strong perceptions of control. Personal norms For each behavioural category participants were presented with three items assessing personal norms with respect to the environment (e.g., “Conserving water is something that everyone should do”; “Because of my values and principles, I feel it is important to try and conserve water”; “I feel a moral obligation to save water for the sake of the environment”). Responses were reported on 5-point scales, which were labelled “strongly disagree” (1), “disagree” (2), “undecided” (3), “agree” (4), “strongly agree” (5), respectively. Cronbach Alpha for the four behavioural domains varied between 0.78 and 0.84. Across all items Cronbach Alpha was 0.91. For each participant the personal norm responses were averaged into an aggregated personal norm index. High scores indicate strong personal norms. Biospheric values Biospheric values were assessed by four items taken from De Groot and Steg (2008). Respondents rated the importance of “Preventing pollution: protecting natural resources”, “Respecting the earth: harmony with other species”, “Unity with nature: fitting into nature“, and “Protecting the environment: preserving nature” in terms of the extent to which they were “a guiding principle in their lives”. The response scale used was ‘‘not at all important” (1) to “of supreme importance” (5).’ Cronbach Alpha was 0.91. For each participant the responses were averaged. High scores indicate strong values. Personal involvement Personal involvement was assessed by eight items, which were developed for the present study. The items covered emotional involvement (e.g., “I feel anxious about what climate change will do to us”), interest in the environment (e.g., “There are more important things to worry about than the environment” - reverse scored), and empowerment (e.g., “I feel that I can really make a contribution to a better environment”). Responses were reported on 5-point scales, which were labelled “strongly disagree” (1), “disagree” (2), “undecided” (3), “agree” (4), “strongly agree” (5), respectively. Cronbach Alpha was 0.81. For each participant the responses were averaged into a personal involvement index. High scores indicate strong personal involvement.","All variables were screened on distribution normality, and were found satisfactory. Skewness and kurtosis values were between −0.05 and + 0.05 for all variables, except personal involvement, for which these values were −0.98 and 1.50, respectively, suggesting some degree of deviation from normality. In Table 1 means, standard deviations, and correlations of the behaviour indices at T1 and T2 and the determinants of behaviour at T1 are presented. All determinants assessed at T1 were statistically significantly correlated with the behaviour indices. The correlations were as can be expected on the basis of the literature, that is, in the range of 0.30–0.40. Testing the habit discontinuity hypothesis ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In order to test the main hypothesis, a multiple regression analysis was conducted. Behaviour at T2 was regressed on age, sex, behaviour at T1, the determinants of behaviour at T1 (habit, intention, perceived control, biospheric values, personal norm, and personal involvement), the intervention, relocation status, and the intervention x relocation status interaction. The intervention was coded as −1 (control group) and +1 (intervention group), and relocation status was z-transformed before calculating the interaction term. The adjusted R-square was 0.46, Cohen's f2 = 0.85. The variance inflation factors varied from 1.02 to 2.28, indicating that there were no multicollinearity problems. Details of the analysis are presented in Table 2. While all determinants at T1 correlated statistically significantly with behaviour at T2, only behaviour at T1, habit, and personal involvement retained statistically significant regression weights. Unsurprisingly, the main effect of relocation status was non-significant. The intervention effect was statistically significant, suggesting the intervention was effective in changing behaviour in a sustainable direction. Importantly, this effect was qualified by a statistically significant intervention x relocation status interaction, beta = −0.08, t = −2.37, p < .02. In order to inspect the nature of the interaction, simple slope analyses were conducted at the mean minus one standard deviation of relocation status, the mean, and the mean plus one standard deviation, respectively. The dependent variable was the behaviour index at T2, controlled for the behaviour index at T1 and the determinants. These analyses revealed that the intervention was most effective when participants had relocated relatively recently. The slopes showed a statistically significant effect of the intervention on behaviour change at the mean minus one standard deviation of relocation status, beta = 0.20, p < .001, 95% CI between 0.08 and 0.32, and at the mean, beta = 0.09, p < .04, 95% CI between 0.01 and 0.18, while there was no significant effect at the mean plus one standard deviation, beta = −0.01, 95% CI between −0.14 and 0.11. Because part of the information provided during the intervention was individualised, the multiple regression was re-run with the behavioural composites of participants in the intervention condition being replaced by composites of only the behaviours that were specifically addressed during the Personal Interview. Using these tailored scores, the results were very similar to those reported above, and included again the statistically significant intervention x relocation status interaction, beta = −0.08, t = −2.43, p < .02. Investigating the ‘window of opportunity’ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Tentatively, we investigated the actual time frame during which the intervention was most effective after participants had relocated. In other words, given that a habit discontinuity effect was present, how wide was this ‘window of opportunity’? While relocation status had been log-transformed in the previous analyses, in this case the raw number of weeks since relocation was used. As half of the total sample was recruited to have relocated within the previous 6 months, the sample was split as closely as possible into quartiles, i.e., participants who relocated less than 3 months previously, 3–6 months, 6–200 months and over 200 months, respectively. In Fig. 1 the mean behaviour composite scores at T1 and T2 in the intervention and control conditions are graphically presented for the four groups. The figure also provides the test of the intervention effect in each group.3 This breakdown suggests that the intervention was effective during the first three months after relocation, after which no effects could be detected.","This study tested the hypothesis that an intervention delivered in the wake of a major discontinuity (residential relocation) is more effective than if the intervention is delivered under default conditions. The rationale behind the hypothesis is that when old habits are temporarily disturbed, people may be more sensitive to new information and adopt a mind-set that is conducive to behaviour change. The results of the study gave support to this hypothesis. While controlling for baseline levels of past behaviour, habit strength, intentions, perceived control, biospheric values, personal norms and personal involvement, and pitted against a no-intervention control group, participants who received an intervention and had recently relocated reported more change eight weeks later on a composite of twenty-five environment-relevant behaviours compared to participants who had not recently relocated. Although other studies have produced results that were in line with the habit discontinuity hypothesis (Bamberg, 2006; Brown et al., 2003; Fujii et al., 2001; Jones & Ogilvie, 2012; Thøgersen, 2012; Verplanken et al., 2008; Walker et al., 2015), the field experimental design of the present study provided a more rigorous test of the hypothesis. Unlike the above mentioned studies, which were either correlational and/or included samples in which all participants had been subjected to a discontinuity of some sort, the present study was thus able to demonstrate the effect of the discontinuity per se. There are a number of caveats to consider. The effect size of the extent to which relocation boosted intervention effects was small. More than anything else, the results should be considered as ‘proof of concept’. Two conditions made the present test very conservative. The first is that for an individual participant not all behaviours were relevant, and only a selection of these were addressed in the Personal Interview. Secondly, the discontinuity effect was controlled for major determinants of behaviours, that is, past behaviour, habit, perceived behavioural control, and a set of motivation variables, which, as can be expected, explained most of the variance in T2 behaviour. The test of the discontinuity effect was thus confined to the mere additional contribution of relocation. The discontinuity effect was evident when using the highest level of aggregation of behaviour and the corresponding aggregates of habit, intention, perceived control and personal norm.4 First, the intervention was partly tailored, and thus focused on different behaviours for different participants, which provided a compelling argument for aggregation. Second, aggregation makes sense from a reliability point of view. In a seminal paper, Weigel and Newman (1976) showed that when single behavioural criteria were aggregated into an overall behavioural index, this measure correlated 0.62 with a general environmental attitude measure, compared to an average of 0.29 when single criteria were used. While that paper focused primarily on the issue of attitude-behaviour consistency, it demonstrated that, in line with the principles of test theory, combining multiple indicators provides a more reliable instrument. In order to provide the most rigorous test of the discontinuity effect, we thus also aggregated the behaviour-specific determinants (habit, intention, perceived control, personal norm) at the highest level. A fair question can be posed about the psychological meaning of these aggregates, as these do not have one-to-one connections to specific behaviours. Our view is that, if anything else, the aggregates might be considered as behavioural, motivational, and normative representations of higher order sustainability or ecological values. One of our reviewers suggested that the aggregated habit variable might capture variation in self-identity, in this case a “green” identity (e.g., Sparks & Shepherd, 1992; Stern, 1992; Thøgersen & Ölander, 2002; Verplanken & Roy, 2013; Whitmarsh & O'Neill, 2010), which would thus also elucidate why habit retained a significant regression weight (Table 2). The latter suggested that variations in existing habit strength modulated behaviour change over and above the discontinuity effect. While the results undoubtedly have theoretical significance, the practical implications are limited unless circumstances are found or created under which larger effect sizes can be realised. The latter may be accomplished in a variety of ways. For instance, larger effect sizes can be expected when interventions focus on single or small sets of behaviours (e.g., recycling, eco-driving, saving water; e.g., Abrahamse, Steg, Vlek, & Rothengatter, 2005). Effect sizes may increase by selecting and/or combining treatments. On the basis of a meta-analysis of 253 intervention studies in the domain of pro-environmental behaviours, Osbaldiston and Schott (2012) found that interventions that included cognitive dissonance, goal setting, social modelling and the use of prompts showed the largest effect sizes. Combining such tools with a discontinuity approach may thus lead to more powerful interventions. Finally, interventions may be more effective when these are carried forward by groups or communities which generate social support (e.g., Abrahamse et al., 2005; Staats, Harland, & Wilke, 2004; Weenig & Midden, 1991). Interventions and behaviour change cannot be seen in isolation from wider systems in which they occur (e.g., Hawe, Shiell, & Riley, 2009; Lewin, 2008/1946). A ‘system’ may be defined by geographical location, such as a residential area. This comes with an infrastructure and bundles of behaviours, such as driving children to school or shopping in that area. Relocation thus unfreezes such patterns. It is therefore interesting to focus on relatively large-scale discontinuities, which are confined to a specific location and time frame, such as when a new residential area is being built. These situations provide easy access to relatively large groups of residents, who are all undergoing the same life course change in the same time period. Other ‘systems’ may be culturally defined, such as social practices, for instance those involving hygiene or leisure activities (e.g., Kurz et al., 2015; Reckwitz, 2002; Shove, Pantzar, & Watson, 2012). Social practices also involve infrastructures and bundles of habits, and are empowered with a shared meaning (e.g., hygiene standards). When individuals move into a new phase, such as a transition from school to work, starting a family or entering retirement, the habits defined by a social practice are subject to change, and may thus be interesting targets for interventions. Habit discontinuities may also be considered from a stage model perspective (e.g., Bamberg, 2013; Dahlstrand & Biel, 1997; Gollwitzer, 1990; Prochaska & Velicer, 1997). Stage models of behaviour change distinguish a motivational phase and a volitional or execution phase. The motivational phase is characterised by deliberation, prioritising goals, and forming intentions. In the volitional phase goals and intentions are subsequently enacted. One of the issues in these models is why, how, and when an individual moves from a motivational to a volitional phase. A habit discontinuity may be conducive to instigating such a transition, and may thus facilitate the transition from contemplation to action (e.g., Gollwitzer, 1990; Holland, Aarts, & Langendam, 2006). A prerequisite is that a motivation to adopt the new behaviours is present in the first place, which needs to be genuine and self-related in order to have the potential to translate into action (e.g., Bamberg, 2006; Verplanken & Holland, 2002; Walker et al., 2015; Whitmarsh & O'Neill, 2010). The present study has a number of limitations. An obvious limitation is that behaviours were assessed by means of self-reports, which are vulnerable to social desirability and consistency biases. We opted for investigating a broad spectrum of behaviours. This put constraints on the types of measurements that were feasible. However, even if the self-reports contained a degree of bias, this cannot explain why a habit discontinuity effect was found. In other words, it is difficult to argue that biases would be stronger among those who recently relocated. The choice for breadth (many behaviours) also put restrictions on the assessments of behaviour-specific determinants (habit, perceived control, intention, personal norm), which could only be measured at the level of behaviour category (e.g., saving water). This introduced a lack of correspondence. Another limitation was that we had no a-priori foundation for defining what exactly a “recent” relocation is, and the chosen cut-off period of six months was arguably arbitrary. This points to a wider issue with respect of discontinuity hypothesis, namely the question what determines the size of what we referred to as the ‘window of opportunity’, and the time when discontinuity effects can be expected to occur (cf., Jones & Ogilvie, 2012). A disruption is a temporary condition, and once a person has settled into the new situation, old habits may easily be re-activated (e.g., Walker et al., 2015; Wood & Rünger, 2016; Wood et al., 2005). The present results suggest that in this case the ‘window’ was approximately three months wide. It can be argued that the dynamics of a discontinuity effect may not be confined to a period after the event itself. For instance, in the case of relocations, behaviours such as more efficient commuting may be the very reason for a relocation. The window of opportunity may thus open before the actual discontinuity takes place.","Prevalent models of behaviour and behaviour change may lead to the suggestion that people change attitudes and behaviour only if the alternatives offered are sufficiently convincing or beneficial. This leaves the context in which interventions are delivered out of the equation, and reflects a rather static view of human behaviour. The habit discontinuity approach focuses on contexts in which people are undergoing life course changes. Those moments of change, when minds and behaviours temporarily unfreeze, may provide precious opportunities for adopting healthier and more sustainable lifestyles."],["Latent inhibition refers to a retardation in learning about a stimulus that has been rendered familiar by non-reinforced preexposure, relative to a non-preexposed stimulus. Latent inhibition has been shown to be inversely correlated with schizotypy, and abnormal in people with schizophrenia, but these findings are inconsistent. One potential contributing factor to this inconsistency is that many tasks that purport to measure latent inhibition are confounded by alternative effects that also retard learning and co-vary with schizotypy (e.g. learned irrelevance and conditioned inhibition). Here, two within-participant experiments are reported that measure the effect of familiarity on learning without the confound of these alternative effects. Consistent with some of the clinical literature, a positive association was found between the rate of learning to the familiar, but not the novel, stimulus and the unusual-experiences dimension of schizotypy - implying abnormally persistent latent inhibition in high schizotypy individuals. --------------------------------------------------------------------------------","For many years, experimental designs have been translated from the study of animal learning to abnormal psychology in an attempt to understand the symptoms associated with schizophrenia. One of these symptoms is a disruption of attentional function (e.g. Hemsley, 1987; McGhie & Chapman, 1961), and latent inhibition (Hall & Honey, 1989; Lubow, 1973; Lubow & Moore, 1959) has been one of the most common designs used to model abnormal attention in schizophrenia. Here, a stimulus is rendered familiar by mere non-reinforced exposure, before being established as a cue for an outcome. Latent inhibition is observed when participants learn more slowly about the preexposed cue than a non-preexposed control cue during a subsequent test of learning (Lubow & Moore, 1959). Theoretical analyses of latent inhibition have focused upon an attentional explanation — proposing that during preexposure, attention diminishes to the preexposed stimulus so that, subsequently, participants take longer to learn the association between this stimulus and the outcome (Lubow & Gewirtz, 1995; Mackintosh, 1975; Pearce & Hall, 1980) than the non-preexposed cue. Consistent with the idea that individuals with schizophrenia have a deficit in attention is the observation of an attenuation of latent inhibition in these individuals, which is reflected as the absence of slower learning to the preexposed cue. This effect is typically seen in individuals with acute schizophrenia, rather than individuals with chronic schizophrenia (e.g. Baruch, Hemsley, & Gray, 1988a; Gray, Fernandez, Williams, Ruddle, & Snowden, 2002; Gray, Hemsley, & Gray, 1992; Rascle et al., 2001; Vaitl et al., 2002, but see also: Cohen et al., 2004; Swerdlow, Braff, Hartston, Perry, & Geyer, 1996; Williams et al., 1998), and this relationship has been suggested to account for the presence of spurious associations being formed between stimuli in the environment from which unusual thought patterns and positive symptoms may emerge (i.e., hallucinations, delusions) (Kapur, Mizrahi, & Li, 2005; Moran, Owen, Crookes, Al-Urzi, & Reveley, 2008). However, the relationship between attenuated latent inhibition and positive symptomatology has been challenged. Gray et al. (1992) suggested that a reduction in latent inhibition was associated with the acute stage of schizophrenia rather than the positive symptoms per se. When acute and chronic patients with schizophrenia were matched for their level of positive symptoms, an attenuation of latent inhibition was only observed in acute, not chronic patients. Later studies have provided mixed findings: normal latent inhibition has been observed in both acute medicated (Swerdlow et al., 1996) and un-medicated (Williams et al., 1998) patients. More recent studies have shown that acute patients with schizophrenia do show an attenuation of latent inhibition, but that this was correlated with their negative rather than their positive symptoms (Rascle et al., 2001), whereas Cohen et al. (2004) found latent inhibition in schizophrenia patients with high levels of positive symptoms did not differ from that of healthy controls (for a review see: Schmidt- Hansen & Le Pelley, 2012). One possible explanation for the inconsistencies in the literature may be because the effect has an additional pole of expression — an enhanced, or abnormally persistent, latent inhibition effect with the chronic stage of schizophrenia (Weiner, 2003). To the best of our knowledge, only three studies have shown that latent inhibition is abnormally persistent in chronic patients with schizophrenia. Rascle et al. (2001), Cohen et al. (2004) and Gal et al. (2009) all report enhanced latent inhibition in patients in a chronic stage of their illness. Although enhanced latent inhibition has been tentatively associated with negative symptoms, this effect appears more specific to illness chronicity (Gal et al., 2009). It thus seems accurate to suggest that schizophrenia is associated with an abnormal expression of latent inhibition. Whether an attenuation or enhancement of the effect is observed, depends on the stage of the illness. As has been noted elsewhere (e.g. Haselgrove & Evans, 2010) comparisons of the cognitive abilities of schizophrenic patients with controls can introduce a number of confounds, notably the medication state of the different groups. To overcome this issue, a dimensional approach can be adopted in which variations in schizotypal personality characteristics are measured in a normal population and correlated with performance on cognitive tasks. A number of studies have now indicated that attentional mechanisms are similarly disrupted in high psychometrically-defined schizotypal individuals and people with schizophrenia (e.g. Baruch, Hemsley, & Gray, 1988b; Evans, Gray, & Snowden, 2007; Granger, Prados, & Young, 2012; Gray et al., 2002; Le Pelley, Schmidt-Hansen, Harris, Lunter, & Morris, 2010; Schmidt-Hansen, Killcross, & Honey, 2009). However, like the schizophrenia literature (e.g. Baruch et al., 1988a; Gray et al., 1992; Rascle et al., 2001), previous studies that have investigated the relationship between schizotypy and latent inhibition have revealed mixed results. Baruch et al. (1988b) were the first to report a relationship between latent inhibition and schizotypy in the normal population; reporting reduced latent inhibition in participants who scored high, but not low (as determined by a median split) on the Psychoticism dimension of the Eysenck psychoticism questionnaire (EPQ; Eysenck & Eysenck, 1975 see also; Allan et al., 1995; Lubow, Ingberg- Sachs, Zalstein-Orda, & Gewirtz, 1992). Similarly, Gray, Snowden, Peoples, Hemsley, & Gray (2003) reported measures of schizotypy to be correlated with reduced latent inhibition, but only when using a between-participant latent inhibition task (see also: Braunstein- Bercovitz & Lubow, 1998; Burch, Hemsley, & Joseph, 2004). However, another between- participants latent inhibition task used by Lipp, Siddle, and Arnold (1994) reported no significant association of the effect with the EPQ (Eysenck & Eysenck, 1975), and an association between latent inhibition and the schizotypal personality questionnaire (Claridge & Broks, 1984) that only approached statistical significance (see also: Lipp & Vaitl, 1992). Furthermore, this trend was due to differences in the non-preexposed control group, with high scorers tending to learn faster than low scorers, rather than the theoretically more interesting, preexposed group. The between-participant tasks used by Baruch et al. (1988b) showed no association between latent inhibition and scores on the Launey–Slade Hallucination Scale (Launay & Slade, 1981). Other studies have shown that, given sufficient preexposure, individuals high in schizotypy can in fact demonstrate a latent facilitation effect (De la Casa, Ruiz, & Lubow, 1993) (but see Burch et al., 2004). Therefore, where some authors report a reduction in latent inhibition with higher levels of schizotypy, others do not, and with some authors suggesting a reversal of latent inhibition with schizotypy (see also: De la Casa & Lubow, 2002; Kaplan & Lubow, 2011; Lubow & Kaplan, 1997; Lubow, Kaplan, & De la Casa, 2001; Lubow & Weiner, 2010; Shira & Kaplan, 2009). More recent studies have tended to employ a within-participant procedure for detecting latent inhibition in which learning about a novel and familiar stimulus is measured in the same participant. Evans et al. (2007), Schmidt-Hansen et al. (2009) and Granger et al. (2012) argue their results support a deficit in latent inhibition, related to the positive dimension of the Oxford-Liverpool Inventory of Feelings and Experiences (O-LIFE) (Mason, Claridge, & Jackson, 1995). However, the attenuated latent inhibition effect with unusual experiences reported by Evans et al. and Schmidt-Hansen et al. did not reach the conventional cut-off point for statistical significance. A significant reduction in latent inhibition was attained by Granger et al., but this was a result of an association between the difference between the preexposed and non-preexposed stimuli and unusual experiences. This latter observation is problematic, because any correlation between schizotypy and a composite constructed from these two scores does not reveal which of its components is, or is not, contributing to the overall effect. As such it is entirely possible that it is a difference in performance to the non-preexposed stimulus, not the preexposed stimulus, that contributes to the co-variation of the composite measure with schizotypy. In support of this possibility, Granger et al. did not see any significant relationship between the unusual experiences dimension and learning about the preexposed stimulus alone. A number of studies of latent inhibition in humans have modified its basic procedure in order to ensure that participants engage with the experiment during preexposure. First, the outcome from the second stage of the experiment might also be included in the first stage of the experiment — unpaired with the cue (e.g. Cohen et al., 2004; De la Casa & Lubow, 2001; Gal et al., 2009; Lubow & De la Casa, 2002; Lubow & Kaplan, 1997; Swerdlow et al., 1996). Second, a secondary, masking, task may be presented concurrently with the preexposed cue. For example, a list of nonsense syllables may be presented and participants required to count the number of times one syllable appears during preexposure (e.g. Baruch et al., 1988a; Gray et al., 1992). The use of either of these modifications undermines the comparability of human latent inhibition to animal models that do not require such procedures to observe latent inhibition (Lubow, 2005). But, more importantly, they also generate procedures that align themselves with other learning phenomena, rather than latent inhibition. For example, by exposing the target outcome during the pre-exposure stage of the experiment in an uncorrelated (or unpaired) fashion with the pre-exposed cue, may result in the establishment of learned irrelevance or conditioned inhibition to the pre-exposed cue; both of these effects are known to retard the acquisition of later learning (e.g.: Baker & Mackintosh, 1977; Rescorla, 1969) and are known to co-vary with schizotypy (Le Pelley, Schmidt-Hansen, et al., 2010; Migo et al., 2006; Schmidt-Hansen et al., 2009). Evans et al. (2007) have described a within-participant latent inhibition procedure that, they suggest, circumvents the inclusion of a masking-task during preexposure. In this task participants were presented with a series of letters, presented one after the other in the centre of the screen and instructed to press the spacebar as quickly as possible when the letter X was presented. The letter X was either preceded on some trials by a letter (e.g., S) that had been preexposed amidst the filler letter earlier in the experiment or by a letter (e.g., H) that had not been preexposed. This task showed a latent inhibition effect — participants were slower to respond to presentations of X when it was cued by the preexposed letter rather than the non-preexposed letter, and a trend for a reduction in latent inhibition with the positive symptom dimension of schizotypy was observed. As this procedure did not include a concurrent masking task during the preexposure stage of the experiment, it is difficult to explain this result in terms of learned irrelevance. Furthermore, at first blush, it seems difficult to explain this result in terms of conditioned inhibition, as the target outcome was not presented to participants during the pre-exposure phase either. However, as Evans et al. note, an expectation of the target- stimulus was established prior to the preexposure phase through instruction. Thus, conditioned inhibition might be generated because the target outcome was expected to appear (but did not) at a time when the preexposed (but not the non-preexposed) stimulus was presented. Consequently, standard associative models of learning (e.g. Rescorla & Wagner, 1972) predict that during the preexposure stage an inhibitory association will form specifically between the preexposed stimulus and the target X, slowing later learning with this stimulus. Importantly, this slower learning is not a consequence of an attentional mechanism — such one that might generate latent inhibition. Here we introduce a procedure that examines variations in latent inhibition with schizotypy under conditions where the contribution of conditioned inhibition and learned irrelevance are minimised in order to provide a less ambiguous measure of the impact of learned variations in attention. However, removing the masking task altogether would result in an experimental paradigm that participants have no requirement to engage in. An alternative strategy then is to keep the masking task in place during preexposure but in such a way as to establish it as task-relevant. The two experiments reported here explored this possibility.","The first aim of Experiment 1 was to create a within-participant latent inhibition task that minimises the possibility of observing conditioned inhibition and learned irrelevance. The second aim was to examine how this task co-varies with schizotypy. Presented here are two variations of a task by Evans et al. (2007; itself modified from that designed by Young et al., 2005). The first version constituted a replication of the task described by Evans et al., to demonstrate latent inhibition, predominantly as a positive control. The second version constituted a modification of this task where no expectation of the target was established during the preexposure stage either through instruction or explicit exposure to the target outcome — thus removing the contribution of conditioned inhibition. Instead, as suggested by Evans et al., during the preexposure stage participants were simply asked to count the number of instances of one of the filler letters (M). This manipulation also establishes all of the stimuli in stage 1 as task relevant as participants must process each letter in order to determine whether it is a letter M or not. Consequently, this task is also less amenable to an explanation in terms of learned irrelevance. In the subsequent test stage of both versions of the task, participants continued to be presented with a series of letters, one after the other in the centre of the screen, but were now instructed to make a response as quickly as possible when the letter X appeared. On some occasions the letter X was preceded by a non- preexposed cue, whereas on other trials it was preceded by a cue that had been rendered familiar by being presented during the preexposure stage. Based on the results of Evans et al. it was expected that response-times would be shorter to X when it had been preceded by the non-preexposed, rather than the preexposed cue. We are interested in assessing whether the same effect was evident in the modified version of the task, as this would suggest the operation of a mechanism during stimulus preexposure that is not sensitive to learned irrelevance or conditioned inhibition.","Fifty-seven healthy Nottingham University participants and members of the general public (35 males and 22 females) took part, in exchange for course credit or a £4 inconvenience allowance. The age range was 18–54. Twenty-eight participants completed the replicated version of the Evans et al. (2007) LI task (‘replicated- task condition’), and twenty-nine completed a modified version of this task (‘modified-task condition’).","All experimental stimuli appeared on a standard desktop computer running Windows XP, and were programmed using Psychopy (Peirce, 2007; www.psychopy.org). Stimuli were white capital-letters in Arial-font (7 mm × 5 mm; h × w) presented for 1 s each on a computer-screen (28 cm × 35 cm; h × w) with a gray background. The stimulus-letters were S and H, one of the letters served as the preexposed stimulus and the other was the non-preexposed stimulus, counterbalanced across participants. The target was the letter X, with filler-letters D, M, T and V; see Fig. 1. Replicated-task condition The task had two stages: preexposure and test. After reading an information sheet and signing a consent-form, the following instructions were presented to participants on the computer monitor prior to the task: “In this task I want you to watch the sequence of letters appearing on the screen. Your task is to try and predict when a letter ‘X’ is going to appear. If you think you know when the ‘X’ will appear then you can press the space bar early in the sequence, that is before the ‘X’ appears on screen. Alternatively, if you are unable to do this please press the spacebar as quickly as possible when you see the letter ‘X.’ There may be more than one rule that predicts the ‘X.’ Please try to be as accurate as you can, but do not worry about making the occasional error. If you understand your task and are ready to start press the spacebar to begin.” During the preexposure stage the preexposed stimulus was presented 20 times, intermixed in a random order with presentations of filler letters each of which was presented 15 times; each stimulus was presented for 1000 ms separated by a 50 ms inter-stimulus interval. The non-preexposed stimulus and target letter X were not presented during the preexposure stage. The test stage followed the preexposure stage, without interruption. In the test stage, the preexposed stimulus and the non-preexposed stimulus were each presented 20 times followed by a 1000 ms presentation of the target stimulus X. There were also 20 non-cued presentations of X during which the target was preceded by one of the 4 filler letters, each of which preceding the target 5 times. In total there were 64 presentations of the filler letters throughout the test phase. The whole task lasted 7 min. Participants were required to press the space-bar, either when X appeared on screen, or if they could predict when the X would appear as the next letter in the sequence. Modified-task condition The procedure for the modified version of the task was as described for the replicated version of the Evans et al. (2007) latent inhibition task (Section 2.1.3.1), with the exception that participants received 2 sets of instructions, one set appeared on screen prior to the preexposure stage, instructing the following: “In this task I want you to watch the sequence of letters appearing on the screen. Your task is to count how many times the letter ‘M’ appears. This task will last about 3 mins. When this task ends, you will be given a new set of instructions. Press any key when you are ready to start the experiment.” Thus for the modified-task condition participants were not aware that the target stimulus would appear until after the preexposure phase. A second-set of instructions (identical to those administered at the outset of the replicated-task condition) were then presented prior to the test stage. A computer-based version of the O-Life (Mason et al., 1995) was administered to assess individual schizotypy. This questionnaire assesses four dimensions of schizotypy. The unusual experiences (UnEx) subscale measures auditory hallucinations, magical thinking and perceptual aberrations reflecting positive symptoms of schizophrenia (e.g., “Have you ever felt you have special, almost magical powers?”). The Introvertive Anhedonia (IntAn) subscale reflects anhedonia (inability to experience pleasure); analogous to the negative symptoms of schizophrenia (e.g., “Do you feel lonely most of the time, even when you're with people”). The Cognitive Disorganisation (CogDis) subscale assesses disruptions in attention/concentration; consistent with the disorganized symptoms of schizophrenia (e.g., “Do you ever feel that your speech is difficult to understand because the words are all mixed up and don't make sense?”). Lastly, Impulsive Nonconformity (ImpNon) measures recklessness, impulsivity and antisocial behavior (e.g., “Do you often have an urge to hit someone?”) ; similar to the Psychoticism scale of the Eysenck Personality Questionnaire (Eysenck & Eysenck, 1975).1 The OLIFE questionnaire has good validity as it maps on to the same multi-dimensional structure as schizophrenia; assessing positive, negative and disorganized symptoms (Mason et al., 1995). Scoring Reaction times (RT's) in stage 2 were recorded from the onset of the preexposed and non-preexposed stimulus that preceded the target (X) for each participant. As each stimulus was presented for 1000 ms separated by a 50 ms inter-stimulus interval, participants' RT could range from 0 to 2050 ms. If participants' RT was less than 1050 ms they predicted the X; whereas if their RT was between 1050 and 2050 ms, they responded to the X. Median RTs for responses to the PE stimulus and NPE stimulus were calculated for each participant as it is less biased by extreme values compared to the mean. RT's to both the PE stimulus and NPE stimulus across the 20 test trials were between 1050 and 2050 ms, excluding one NPE trial which had an RT less than 1050 ms for the modified-task version; thus the majority of responses to both stimuli were responses to the X. The scores derived for the four-schizotypy subtypes (complete for Experiment 1 and the subsequent Experiment 2) are presented in Table 2. Latent inhibition Fig. 2 shows the group mean of individual median reaction times to X across the 20 test trials2 with the PE and NPE stimuli. Both the replicated-task and the modified-task groups showed faster RTs to the non-preexposed stimulus than the preexposed stimulus — latent inhibition. A 2 (condition: replicated-task, modified-task) × 2 (stimulus: preexposed, non-preexposed) mixed analysis of variance (ANOVA) of individual median reaction times revealed a significant main effect of stimulus F(1,55) = 16.626, p < .001, partial η2 = .23, but no main effect of condition or interaction (Fs < 1), suggesting reaction times were similar for participants in both the replicated-task and the modified-task irrespective of target expectation during preexposure. On this basis, and to increase statistical power, the data were combined from the two test conditions for subsequent analyses. Latent inhibition and schizotypy A standard multiple regression analysis was carried out using the four schizotypy subscales taken from the O-Life: UnEx, IntAn, ImpNon and CogDis as the predictor variables, and individual median reaction times to the preexposed and non- preexposed stimuli as the dependent variables. If any of the predictor variables are associated with latent inhibition it would be expected that a relationship would be found with the preexposed stimulus, but not with the control non- preexposed stimulus. When reaction time to the preexposed stimulus was entered as the dependent variable, UnEx was a significant predictor of RTs (β = .36, p = .021), reflecting slower learning to the preexposed stimulus with individuals high in UnEx, i.e. enhanced latent inhibition. ImpNon was also a significant predictor of reaction time to the PE stimulus (β = −.36, p = .014), reflecting faster learning to the preexposed stimulus for individuals high in ImpNon, i.e. an attenuation of latent inhibition. Neither of the remaining schizotypy subscales (CogDis and IntAn) were significant predictors of reaction time to the preexposed stimulus (ps > .05). When reaction time to the non-preexposed stimulus was entered as the dependent variable, the only significant predictor of reaction time was ImpNon, which again was negatively correlated with RT (β = −.32, p = .035). None of the remaining schizotypy dimensions were significant predictors of reaction to the non-preexposed stimulus (ps > .05). All standardized regression coefficients and R2 values can be seen in Table 1. The results indicate that individuals high in UnEx are slower to learn the association between the preexposed stimulus and the target than individuals low in UnEx. This, in conjunction with the finding that UnEx was not a significant predictor of reaction time to the non-preexposed stimulus, indicates that individuals high in this subtype are exhibiting an enhancement of latent inhibition. A relationship between ImpNon and RTs to both the preexposed and non-preexposed stimuli was also found, showing that individuals high in ImpNon make faster responses irrespective of whether the stimulus is familiar or novel. The enhancement of latent inhibition with high UnEx, does not agree with a number of schizotypy studies (Evans et al., 2007; Granger et al., 2012; Schmidt-Hansen et al., 2009). However, the reported attenuation of latent inhibition with high UnEx, failed to reach the conventional level of significance in the studies reported by Evans et al. (2007) and Schmidt-Hansen et al. (2009). Furthermore, it cannot be ruled out that the latent inhibition task employed in each of these studies was not a consequence of alternative learning phenomena instead of latent inhibition, due to the limitations previously described. Before we can draw any further conclusions, however, it is important to acknowledge the possibility that we still might be observing a co-variation of schizotypy with learned irrelevance in the current study, as opposed to latent inhibition. Whilst the modified-task condition successfully minimised the contribution of conditioned inhibition, it still included a masking task (count the letter M). Although this procedure — which requires continuous monitoring of the experimental stimuli — establishes a situation in which all of the experimental stimuli are task relevant, it is conceivable that it still establishes learned irrelevance. In this task, participants are required to respond (albeit covertly) to the letter M, rather than any other stimulus. In this sense, then, the preexposed stimulus is irrelevant to the task in hand, thus learned irrelevance may still be the cause of the slower learning to the preexposed stimulus, rather than latent inhibition. As previously discussed, learned irrelevance is an effect which has been shown to influence human learning (Le Pelley & McLaren, 2003) and also co-vary with schizotypy (Le Pelley, Schmidt-Hansen, et al., 2010; Schmidt-Hansen et al., 2009). However, as previously outlined, it would be problematic to remove the masking task altogether as participants would have no requirement to engage in the task during the preexposure stage. Therefore, the aim of Experiment 2 was to design a procedure that examined latent inhibition under conditions where the contribution of both learned irrelevance and conditioned inhibition were minimised, but keep the masking task in place during preexposure but in such a way as to establish it as directly relevant (as opposed to irrelevant) to the preexposed stimulus. If latent inhibition is still observed under these circumstances, it would permit an evaluation of the effect in terms of models of attention that do not emphasise the importance of learned irrelevance (e.g. Esber & Haselgrove, 2011; Pearce & Hall, 1980).","To minimise the contribution of learned irrelevance (as well as conditioned inhibition), the purpose of Experiment 2 was to adjust the parameters of the modified-task condition from Experiment 1. In the preexposure stage, participants were now asked to say out loud each of the letters that appeared on the screen. This manipulation directly establishes all of the stimuli in stage 1 as task relevant as participants must process each letter by reading each of them aloud. Consequently, this version of the task rules out an explanation of any subsequent attenuation of learning to the preexposed stimulus with an appeal to learned irrelevance. Furthermore, as no expectation of the target stimulus (X) is established prior to, or during, preexposure the task is also not amenable to an explanation in terms of conditioned inhibition. The test stage of the task remained the same as the modified-task condition from Experiment 1: participants were required to make a response as quickly as possible when the letter X appeared on screen. We are first interested in assessing whether an effect of stimulus preexposure is still observed under these different circumstances and second, to assess whether the task co-varies with schizotypy. This being the case would suggest a relationship between schizotypy and of stimulus preexposure that goes beyond learned irrelevance.","Sixty healthy Nottingham University participants and members of the general public (10 males and 50 females) took part, in exchange for course credit or a £4 inconvenience allowance. The age range was 18–33 years.","The apparatus were the same as described in Experiment 1. Procedure The procedure for Experiment 2 was as described in the modified-task condition in Experiment 1 with the exception that the instructions received prior to the preexposure stage asked participants to say aloud each letter that appeared on the screen. A second-set of instructions (identical to those administered at the outset of the test stage of the modified condition from Experiment 1) were presented prior to the test-phase. As per the previous experiments, participants completed the O-Life (Mason et al., 1995) questionnaire. All scoring was performed in the same manner as described in Experiment 1. In keeping with Experiment 1, the majority of RT's to both the PE stimulus and NPE stimulus across the 20 test trials were between 1050 and 2050 ms, excluding 3 NPE trials which had RTs that were less than 1050 ms, indicating that participants were predicting the occurrence of the X on these trials.","The scores derived for the four schizotypy subtypes (complete for Experiments 1 and 2) are shown in Table 2. Unpaired t test analyses were carried out to assess if the reported schizotypy means differ from the population norms for each subscale. While the means for CogDis and IntAn do not differ significantly from the normative values, the means for UnEx and ImpNon are both significantly lower than the normative values for the modified-task version of Experiment 1, and for Experiment 2. Significant differences are highlighted in bold in Table 2. Previous studies have also obtained mean schizotypy scores that are below Mason et al.'s (1995) normative values, and similar to those reported here (e.g. Evans et al., 2007; Granger et al., 2012; Sellen, Oaksford, & Gray, 2005). Latent inhibition Fig. 3 shows the median reaction times to X across the test trials of Experiment 2 (shown in two-trial blocks) with the preexposed and non-preexposed stimuli. It can be seen that reaction times were faster during the non-preexposed than the preexposed stimulus. This impression was confirmed with a 2 (stimulus: non- preexposed, non-preexposed) × 10 (trial block: 1–10) ANOVA of individual reaction times, which revealed a significant main effect of stimulus, F(1,59) = 25.691, p < .001, partial η2 = .303 and a significant main effect of trial number, F(9,51) = 7.949, p < .001, partial η2 = .584, but no significant interaction between these variables, F < 1. In line with both conditions from Experiment 1, Experiment 2 successfully generated an effect of preexposure on reaction times during subsequent learning — latent inhibition. The task presented in Experiment 2 however, produced latent inhibition when the target was not expected during preexposure, and importantly, when using a masking-task that was not irrelevant to stimulus preexposure. These results encourage the suggestion that that an effect of exposure on learning is being observed here — that is to say latent inhibition rather than conditioned inhibition or learned irrelevance. Latent inhibition and schizotypy In keeping with Experiment 1, a standard multiple regression was carried out using the four schizotypy subscales from the O-Life (UnEx, IntAn, ImpNon and CogDis) as the predictor variables, and reaction time to the preexposed and non-preexposed stimuli as the dependent variables. Again, when reaction time to the preexposed stimulus was entered as the dependent variable, UnEx was a significant predictor of reaction times to the preexposed stimulus (β = .40, p = .021), reflecting slower learning to the preexposed stimulus with individuals high in UnEx — replicating the enhanced latent inhibition effect observed in Experiment 1. Unlike Experiment 1, however, ImpNon was not a significant predictor of reaction time to the preexposed stimulus, nor were the remaining schizotypy subtypes. When median reaction time to the non-preexposed stimulus was entered as the dependent variable, none of the schizotypy subtypes were significant predictors of reaction time to the non-preexposed stimulus (ps > .05). Standardized regression coefficients and R2 values can be seen in Table 3. In keeping with Experiment 1, the results of Experiment 2 show that individuals high in UnEx are slower to learn the association between the preexposed stimulus and the target than individuals low in UnEx. In both Experiments 1 and 2, we observed facilitation in RTs in individuals high in UnEx that was specific to the preexposed stimulus. These results encourage the suggestion that we are observing an enhancement of latent inhibition, rather than a more general effect of schizotypy on learning to both stimuli. Whilst the findings from both experiments presented here are comparable, the task employed in Experiment 2 is particularly notable as it comprises a relatively ‘pure’ demonstration of latent inhibition, as it minimises the contribution of both conditioned inhibition and learned irrelevance to stimulus preexposure.","Two experiments revealed slower learning of a stimulus-target association with a stimulus that had been rendered familiar through prior non-reinforced preexposure than a stimulus that had not — latent inhibition. In both experiments learning about the preexposed, but not the non-preexposed stimulus was related to the unusual experiences dimension of the O-LIFE — revealing an enhancement of latent inhibition in individuals scoring higher on the positive dimension of schizotypy. Experiment 2, in particular, arranged preexposure in a manner that resulted in the subsequent retardation of learning to be explicable in terms of the effects of mere exposure but not the confounding effects of conditioned inhibition or learned irrelevance. This is in contrast to other studies in the latent inhibition literature (e.g. De la Casa & Lubow, 2001, 2002; Evans et al., 2007; Granger et al., 2012; Lubow & De la Casa, 2002; Schmidt-Hansen et al., 2009; Swerdlow et al., 1996), which can be explained in terms of these alternative learning phenomena. To the best of our knowledge, the current data constitute the first observation of enhanced latent inhibition in sub-clinical high-schizotypy individuals. Three studies (Cohen et al., 2004; Gal et al., 2009; Rascle et al., 2001) have reported enhanced latent inhibition in schizophrenia patients. The first study by Rascle et al. (2001) used a between-participants design in which chronic schizophrenia patients in the preexposed group showed slower learning in comparison to controls, resulting in an enhancement of latent inhibition. The remaining studies, by Cohen et al. (2004) and Gal et al. (2009), like the current study, employed a within-subject manipulation of stimulus familiarity to demonstrate latent inhibition and were able to show an abnormality in learning that was specific to the preexposed stimuli. Both Cohen et al. and Gal et al. showed that latent inhibition enhancement was associated with the negative symptoms experienced by adolescents with schizophrenia. These results are what would be predicted based on Weiner's (2003) model that suggests enhanced latent inhibition is associated with depleted levels of glutamate (see Javiit, 2007; Javiit, 2010), which may be related to the prevalence of negative symptoms. On the other side of the coin, is the reported relationship between the positive symptoms of schizophrenia and attenuated latent inhibition (e.g. Baruch et al., 1988a; Gray et al., 1992, 2002; Rascle et al., 2001; Vaitl et al., 2002). This latter pattern of results is consistent with Gray et al.'s (1991) model for cognitive and neural associates of positive acute schizophrenia symptoms: that a loss of loss of latent inhibition is due to over-activity in the mesolimbic dopaminergic system. At first glance, the results presented here, an enhancement of latent inhibition with the positive UnEx dimension of schizotypy, conflict with these analyses. There has been considerable disagreement about the relationship between the attenuation of latent inhibition in schizophrenia and positive symptomatology: some authors have found a relationship between latent inhibition and positive symptoms (Baruch et al., 1988a; Gray et al., 1992, 2002; Rascle et al., 2001; Vaitl et al., 2002), others have not (Cohen et al., 2004; Gal et al., 2009; Rascle et al., 2001; Swerdlow et al., 1996; Williams et al., 1998; for a review see: (Schmidt-Hansen & Le Pelley, 2012). In particular, Rascle et al. (2001) reported an attenuation of latent inhibition was associated with low levels of negative symptoms in patients with schizophrenia, rather than with levels of positive symptoms. Whereas Cohen et al. (2004) reported no difference in the magnitude of latent inhibition between high levels of positive symptoms in schizophrenia patients, and healthy controls. These findings, along with the current results, do not support the relationship between latent inhibition attenuation and positive symptomatology. On the other hand, the proposition by Weiner (2003) — that enhanced latent inhibition is related to negative symptoms, refers mainly to chronic patients. However, the findings reported by Cohen et al. and Gal et al. (2009) were able to show an association between enhanced latent inhibition and clinical condition (chronic schizophrenia), but not with the level of negative symptoms per se. The discrepancy between these findings, and the results reported here are possibly due to the nature of the tasks employed by Cohen et al. and Gal et al.; as previously highlighted, these existing tasks confound learned irrelevance with latent inhibition itself. How the refined latent inhibition task reported here covaries with individuals with schizophrenia, is the focus of future research. One possible shortcoming of employing the multiple regression analysis that we have used in Experiments 1 and 2 is that the observed correlations between UnEx and RT to the non-preexposed stimulus could have been caused by any processes that impact upon the RTs to the preexposed stimulus, including those which also impact on RTs to the non-repexposed stimulus; that is to say, the common variance components affecting RTs to both preexposed and non-preexposed conditions. In order to evaluate this possibility, we pooled the data across Experiments 1 and 2 and conducted a hierarchical multiple regression in which RTs to the non-preexposed stimulus were added in the model in step 1 to act as a covariate, and examined the subsequent relationships between UnEx, CogDis, IntAn and ImpNon (as predictor variables), and RTs to the pre-exposed stimulus (as the dependent variable) in step 2. UnEx remained as a significant predictor of RT to the pre-exposed stimulus in step 2, β = .23, t = 2.61, p = .01, as did CogDis now, β = −.18, t = 2.07, p = .041. The remaining sub dimensions of the OLIFE were not significant however, βs < −.01, ts < 1.2, ps > .23. It therefore appears that the relationship that we observed between schizotypy and RT in the current studies is specific to the pre-exposed stimulus. For the purposes of completeness, we also repeated the previous regression but this time with RTs to the preexposed stimulus entered as a covariate in step 1, and examined the subsequent relationships between UnEx, CogDis, IntAn and ImpNon (as predictor variables), and RTs to the non-preexposed stimulus (as the dependent variable) in step 2. None of the beta coefficients were significant. βs < .03, ts < 1.0, ps > .39. In order to ensure that participants were engaged with the task during the preexposure stage of Experiment 2, a secondary task was employed in which participants were required to repeat, out loud, each stimulus that was presented on the screen. We have argued that immersing preexposure within such a procedure precludes the current results from being explained in terms of learned irrelevance — as the preexposed stimulus was established as task relevant. This raises the question, then, of whether the current results are a demonstration of latent inhibition or, instead, a circumstance in which establishing a stimulus as task relevant in stage 1 might hinder learning in stage 2 when the same stimulus is established as an explicit cue for a target stimulus. On balance, this possibility seems unlikely. A number of studies have now shown that when a stimulus is established as relevant to the solution of one task, the same stimulus is subsequently better, not worse, than a control stimulus at serving as a cue in a different task (e.g. Bonardi, Graham, Hall, & Mitchell, 2005; Le Pelley, Turnbull, Reimers, & Knipe, 2010). Furthermore, tasks of these sort have been shown to have a negative, not a positive, correlation with schizotypy (e.g. Le Pelley, Schmidt-Hansen, et al., 2010). To the best of our knowledge there is only one demonstration, in humans, of a stimulus being established as task relevant then going on to show a subsequent retardation in learning (Griffiths, Johnson, & Mitchell, 2011). However, this \"negative-transfer\" effect was demonstrated under circumstances in which the task type was the same between pre-exposure and learning (only the magnitude of the target outcome was changed). Furthermore, to date, there is no evidence of this effect having any relationship with schizotypy. The two experiments presented here show an effect of schizotypy on learning about a preexposed stimulus using a refined latent inhibition procedure. Both Experiments 1 and 2 show a comparable and novel effect of enhanced latent inhibition in individuals high in UnEx. We advocate the use of the task described in Experiment 2, as this task successfully minimised the contribution of both conditioned inhibition and learned irrelevance on the preexposure effect, and could be a useful tool for assessing attentional dysfunction in schizophrenia, as well as other clinical and sub- clinical populations.","No conflicts of interest declared."],["Although locus of control (LOC) has been the focus of thousands of studies we know little about how or if it changes over time and what is associated with change. Our lack of knowledge stems in part from the past use of cross-sectional and not longitudinal methodologies to study small numbers of participants from non-representative populations. The purpose of the present study was to use a longitudinal design with a large representative population to provide relevant information concerning the stability and change of adult LOC. Before the birth of their child, and again six years later, mothers and their partners participating in the Avon Longitudinal Study of Parents and Children (ALSPAC) completed LOC tests and structured stressful events surveys. Analyses revealed that stresses experienced in relationships with spouses, friends and family, financial stability and job security, and illness/smoking were associated with changes in LOC. Results suggest substantial variation of LOC within spousal/parent dyads and moderate stability of LOC over time for both men and women. Stressors associated with change in LOC may be possible candidates when considering interventions to modify LOC expectancies. --------------------------------------------------------------------------------","The purpose of this project was to examine the stability and change of locus of control (LOC) orientation in adult men and women over a 6 year period, and to identify events associated with LOC stability or change. LOC refers to individuals' generalized expectancy regarding the connection between their behavior and reinforcements received in a problem- solving context (Rotter, 1966). Individuals who fail to see a connection between what they do and what happens to them and view what happens as the result of luck, fate, chance, or powerful others are externally controlled. Conversely, those who tend to perceive a connection between their efforts and what happens to them are internally controlled. Rotter's article stimulated a remarkable amount of research. A search of PsychInfo resulted in 17,812 articles with a keyword “locus of control” as of summer 2015 and with 6600 of these appearing after 1996 (1425 dated 2010–2015). LOC has sustained itself as a concept for psychological study for more than a half century (Nowicki & Duke, 2016). As there are >100 different definitions of “locus of control” throughout the literature (Skinner, 1996), researchers need to clearly state and define which LOC concept and measure is being used (e.g. Reich & Infurna, 2016). Peterson and Stunkard (1992) have noted potential problems that could result from using similar appearing cognates, like efficacy (e.g. Infurna & Mayer, 2015; Lachman & Weaver, 1998) or attribution (Peterson & Seligman, 1983; Seligman, 1975) interchangeably with locus of control of reinforcement as described by Rotter (1966). Rotter defined LOC within his social learning theory (1954, 1966) emphasizing that it is an expectancy that has the capacity to affect behavior differently from situation to situation and has its greatest impact in circumstances that are novel, ambiguous or transitory. LOC has been related to an ever-growing number of important and significant aspects of human life including personality characteristics (e.g. Judge & Bono, 2001; Nowicki & Duke, 1974), social adjustment difficulties (e.g., Cheng, Cheung, Chio, & Chan, 2013), academic achievement (e.g. Flouri, 2006), health outcomes (e.g. Conell-Price & Jamison, 2015), and business success (e.g. Kormanik & Rocco, 2009). However, surprisingly little research has been completed concerning the origins of control orientations or their trajectories over time. Using data gathered from the Avon Longitudinal Study of Parents and Children (ALSPAC), Golding, Iles-Caven, Gregory, and Nowicki (2017), Golding, Gregory, Iles-Caven, and Nowicki (2017), and Golding, Ellis, Iles-Caven, Gregory, and Nowicki (2017), identified antecedent characteristics of the childhoods that were common to internality in both men and women; these included maternal warmth, being breast fed, having a stable home, and recollection of childhood as happy. Antecedents like these describe a home situation in which children feel comfortable and safe enough to explore their environments and learn more about the possible contingencies existing between their behavior and outcomes. Lefcourt (1976) had earlier theorized a similar set of circumstances underlying the development of appropriate internal control expectancies. Antecedent information may be helpful in suggesting what might underlie the learning of generalized internal or external control expectancies as well as to offer possible insights into what may be associated with changes in or maintenance of LOC. Unfortunately, since most of the antecedent data have been obtained from studies using cross-sectional methodologies they are inappropriate for objectively describing how individuals' LOC changes over time or, if changes do take place, with what they are associated. Longitudinal data would be valuable in supporting or refuting cross-sectional results describing the development of control expectancies. The purpose of this investigation was to identify events associated with LOC changes over time by analyzing longitudinal LOC data gathered over a 6 year period beginning when the women and their partners were expecting a baby. Some LOC studies with adults have used LOC pre-post scores as an outcome measure to evaluate the impact of interventions. Unfortunately, in these cases researchers rarely report correlations between pre-post testing of the group not receiving the intervention, as this would be valuable in determining the stability of LOC over time in typical participants. For example, Sørlie and Sexton (2004) studied the effects of psychosocial medical and treatment related factors upon change in the Multidimensional Health Locus of Control scale (MHLC) (Wallston, Wallston, & DeVellis, 1978). The MHLC was administered to surgical patients prior to surgery and again four months following discharge. They found a positive relationship with physicians' predicted internality, while severity of illness and subjective feelings of stress were associated with externality. Similar findings were obtained with patients who had undergone heart surgery (Sørlie & Sexton, 2004). They too became more internal with the passage of time and involvement with rehabilitation. Though researchers using the perceived control construct have produced results consistent with those found by Rotter, orientated LOC investigators have primarily studied older rather than younger adults. They used a variety of ways to assess “perceived” control, and not offered construct validity evidence in support of the perceived control measures they used or how they may have related to LOC. For example, in one study (Turiano, Chapman, Agrigoroaei, Infurna & Lachman, 2014) “perceived” control was measured using a 4-item test with a 7-point Likert Scale while in another (Infurna & Okun, 2015) it was assessed by a 6-item scale with a 4-point Likert scale. Because there is an absence of information of how the different perceived control measures correlate with one another or with Rotter's scale, it is difficult to determine if what they are measuring is similar to or different from what researchers have found using Rotter-based LOC scales. Rather than assessing “perceived” control in middle-aged and older adults, Schneewind (1997) was one of the few researchers to gather test–retest information from younger adults using a scale based on Rotter's definition. The initial test was administered to parents when their child was 10 years old; the retest took place 16 years later. Parents and their children completed German adaptations of LOC scales (Adult and Child Nowicki-Strickland Internal-External scales). LOC was part of a more comprehensive longitudinal study of personality in the context of family development (Schneewind, Ruppert, & Harrow, 1998). Initial data were gathered from a sample of mother–father-child triads recruited from six different German states. The mean age of the children was 12, mothers 39, and fathers 42. Sixteen years later Schneewind and his colleagues contacted former participants and had 197 triads volunteer to participate in a second assessment. In this paper we are primarily interested in what Schneewind found regarding parent LOC. At both testing times, mother and father LOC scores were positively correlated, but low (time one, r = 0.19; time two, r = 0.21). Pre-post LOC correlations for mothers and fathers across the 16 years were what Schneewind called “moderate” and ranged between 0.35 and 0.44. The present study ~~~~~~~~~~~~~~~~~ Little is known about how LOC expectancies develop and change in adulthood. We lack information about what the LOC association is within spousal or parental dyads, the trajectory of adult LOC over time or what life events are associated with LOC changes over time. Our goal here is to provide such LOC information. The ALSPAC data set is unique in that it contains LOC scores from both mothers and fathers before the child was born and six years later. Thus we can assess the stability of parents' LOC during a critical period of their lives. The absence of previous research made prediction difficult. Relationship theories offer competing predictions. Complementary theorists (e.g., Kiesler, 1982) suggest a negative association between mother and father LOC scores, with one individual being external and the other internal. Similarity perspectives (e.g., Byrne, 1969) suggests that the LOC scores would be similar; internals liking internals and externals liking externals. Based on Schneewind's findings, we predict parents' own pre-post LOC scores would be moderately related to one another over time, but relatively unrelated to one another at pre- and post-times. Although Schneewind obtained valuable data regarding adult LOC stability and change, he failed to provide any information about what might be associated with changes. The ALSPAC data set includes data concerning the stressors encountered by adults during the six years between the first and second LOC administration. Rotter (1966) and Lefcourt (1976) theorize that internality thrives when individuals are in warm, supportive, relatively stress free environments in which they can learn to perceive the connection between their behavior and outcomes. We predict that greater stress will be related to greater externality. The ALSPAC study ~~~~~~~~~~~~~~~~ This pre-birth cohort was designed to determine the environmental and genetic factors that were associated with health and development of children and their parents (Golding and ALSPAC Study Team, 2004; Boyd et al., 2013). As part of the study design, and in order to determine the parents' backgrounds prior to the birth of the child, there was a concerted effort to obtain details of their personalities, moods and attitudes, including a measure of their LOC, before the birth of the child. ALSPAC recruited 14,541 pregnant women resident in Avon, UK with expected dates of delivery 1st April 1991 to 31st December 1992. Enrolment strategies included encouragement through the local media, general practitioners, midwives, health services and obstetric hospitals; women then contacted the study center for further information; they were then sent a series of questionnaires to be completed at home. 14,541 is the initial number of pregnancies for which the mother enrolled in the ALSPAC study. Of these there were 14,062 livebirths, of which 13,988 survived to at least 12 months. For full details of all the data collected see the study website: www.bristol.ac.uk/alspac/researchers/data-access/data-dictionary/. Uniquely among the major UK cohort studies at the time it was decided to include the fathers of the children. To this end questionnaires were sent to the mother to pass to her partner if she was happy for him to take part. This strategy was approved by the ALSPAC Ethics and Law Committee (Birmingham, 2018). Consequently, there was no immediate way in which the study administrators knew the identity of the study fathers. Given the uncertainty of this approach, it is striking how many took part during the pregnancy (10,000 compared with 13,867 pregnant women (Fraser et al., 2013)). Follow-up used a variety of techniques including questionnaires, hands-on examinations, assays of biological samples and details of the environment at various stages of life. Questionnaires included detailed history of events, occupations and life styles of each parent. Loss to follow-up occurred when the child died, the mother refused, or moved away and could not be traced. As in all longitudinal studies, attrition increased as the study continued (Fraser et al., 2013). This is illustrated by the numbers of parents who completed the LOC questions at the two time points 6 years apart. For the women the numbers fell from 10,565 to 8378 (79% of the original), and for the partners the change was from 7365 to 3891 (53%). From Supplementary Table 1 it can be seen that the proportion of parents with LOC scores 6 years after the study child's birth had proportionately fewer families who (a) resided in public housing, (b) were young mothers, (c) were of manual social class or (d) who smoked. These factors need to be born in mind when generalizing the results. Follow-up used a variety of techniques including questionnaires, hands-on examinations, assays of biological samples and details of the environment at various stages of life. Questionnaires included detailed history of events, occupations and life styles of each parent. Loss to follow-up occurred when the child died, the mother refused, or moved away and could not be traced. As in all longitudinal studies, attrition increased as the study continued (Fraser et al., 2013). This is illustrated by the numbers of parents who completed the LOC questions at the two time points 6 years apart. For the women the numbers fell from 10,565 to 8378 (79% of the original), and for the partners the change was from 7365 to 3891 (53%). From Supplementary Table 1 it can be seen that the proportion of parents with LOC scores 6 years after the study child's birth had proportionately fewer families who (a) resided in public housing, (b) were young mothers, (c) were of manual social class or (d) who smoked. These factors need to be born in mind when generalizing the results. Measures of LOC ~~~~~~~~~~~~~~~ The LOC measure used in the present study is a shortened form of the adult version of the Nowicki-Strickland Internal-External locus of control scale (ANSIE) which comprises 40 items in a yes/no format to assess perceived control (Nowicki & Duke, 1974). This was chosen over other scales more specifically related to perceived control over health, as it was considered that this more generalized scale would relate to other factors in addition to health outcomes. Construct validity for the scale has been found in the results of over 1000 studies (Nowicki, 2016). The version used here comprises 12 of the original 40 items, which were chosen after factor analysis of the ANSIE in a pilot of 135 mothers in the USA. An Anglicized version of these 12 items was included in the questionnaires sent to the two parents in pregnancy (Golding, Iles-Caven, et al., 2017) and again six years after the child was born, with identical wording. From the responses LOC scores were derived for men as well as for women, the higher the score the more external the LOC. The scores ranged from 0 to 12. Classification into external and internal parents in pregnancy, using the definition for externality of greater than the median LOC score (>4 for women, >3 for men), identified 40.3% of mothers and 38.8% of their partners as externally-oriented in pregnancy. Because of increasing internality in the women over time, their median LOC score changed to 3, whereas their partners stayed at 3. Identification of events ~~~~~~~~~~~~~~~~~~~~~~~~ From 8 months post-birth, each parent was sent a questionnaire concerned with life events in the preceding period, at approximately yearly intervals. The life events inventory comprised 42 items, which were derived for the ALSPAC study using previous inventories as a basis for selection of items. Three main sources were used in this way: Brown and Harris (1978), Barnett, Hanna, and Parker (1983) and Stanley (1988). In order to develop a summary of variables to determine whether each specific event had occurred over the preschool period of the study child, we created a variable for each event that used all available questions (i.e. those sent to the parents at 8, 21, 33 and 47 months). We did not sum the number of times an event occurred as there was often overlap between ages. Rather than look at the number of the events occurring, we have deliberately looked at each separately using the hypothesis that some types of event will result in increasing internality, and some externality, a pattern that we found when considering the relationship between the events in childhood and the LOC of the adult (Golding, Ellis, et al., 2017). Other variables considered ~~~~~~~~~~~~~~~~~~~~~~~~~~ Other variables investigated here were chosen because they were associated with parental LOC as well as with differential response rates. They were included to determine whether they were associated with any change in LOC orientation across the 6 years. They include: (1) housing tenure – which is an accepted socio-economic indicator in the UK – divided into owner-occupied (includes having a mortgage), rented public housing, and other rented property; (2) age of the mother at the child's conception; (3) social class (based on occupation of the mother's partner classified into manual and non-manual (Standard Occupational Classification, 1990); (4) presence of mother's partner in the home (identified by questionnaire at 8 months); (5) maternal education level (based on the actual qualifications attained and divided into three levels of attainment); (6) the mother's parity (number of previous pregnancies resulting in either a live or stillbirth); (7) the smoking habit of each parent at 8 months of age (categorized in two ways – any regular smoking and regular smoking of ≥10 cigarettes/day). Statistical techniques ~~~~~~~~~~~~~~~~~~~~~~ The statistical strategy was to determine the pattern of factors that were associated with the change in LOC orientation in each parent, with a focus on events that occurred after the initial measure. The study was hypothesis free. Based on the findings from a study of the childhoods of these individuals we anticipated that some of the events might result in a change toward internality, and some toward externality. Results were descriptive, and compared individuals who changed from those who did not, using logistic regression. No account was taken of the number of tests undertaken, to avoid type I errors. Correlation of parents' LOC scores with one another and across time ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As demonstrated in Table 1, both mother and father LOC scores were positively and moderately correlated between the two-time points (mothers = 0.57; fathers = 0.56); cross- sectional correlations between parents at each of the two time points were positively correlated, but low (in pregnancy r = 0.33; 6 years later r = 0.29). Additional findings concerning parents' LOC ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The ‘change score’ was calculated as the difference between the LOC score in pregnancy and that obtained 6 years later. Mothers' and fathers' LOC change scores ranged from −9 to +8, with mode and median at 0 with a mean difference of −0.30 [SD 1.89; n = 8302] points for mothers and +0.037 [SD 1.98; n = 3968] points for fathers. Thus, mothers on average were getting more internal while their partners were tending to become more external. The correlation coefficient between the parents change scores was positive but low (0.11). Lifestyle and social conditions associated with changes in LOC over time ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In Tables 2a and 2b we consider certain lifestyles and social conditions at 8 months post- delivery: for life-style we consider the changes in LOC associated with the smoking habits of each parent; social conditions were demonstrated by the tenure of the home in which the family were living and whether the mothers' partner was part of the household. Smoking of the parents bore very strong associations with the likelihood of the parents changing their orientation (Table 2a). Of the women who stayed externally oriented 30.7% were smoking when the offspring was 8 months old, whereas 24.7% of those who had become internal had this history (P = .002); for women who stayed internal, only 11.9% smoked, but 20.0% of those who became external did so (P < .0001). Similar patterns were shown for women who were heavy smokers as well as when the partner smoked (Table 2a). Changes in the partners' orientation were equally associated with the smoking habits of the woman as well as of the partner himself (Table 2b). Mothers who lived in public housing were considerably more likely to stay external, and those in owner-occupied accommodation less likely to do so. Conversely mothers who were internally oriented in pregnancy were more likely to become external if in public housing (Table 2a). A similar pattern was shown for the mothers' partners (Table 2b). Most of the women who answered the relevant questionnaires had live-in partners (Table 2a), but those mothers who were externally oriented were more likely to become internal if the partner was present. Conversely the women who were internal but became external were less likely to have a live-in partner. Partners were less likely to be given (or complete) a questionnaire if they were not living with the mother 6 years post-birth; even so, among those who answered the questionnaire, there was some evidence that those who became internal were more likely to be living with the mother, and those who did not were more likely to become external (Table 2b). Stressful life events associated with changes in LOC over time ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 3a presents variations in the potentially stressful events occurring to parents and ways in which they were associated with maternal changes in the LOC orientation. The first two columns compare proportions of women who had a history of the stressor between those who stayed external and those who became internally-oriented. Of the 42 stressors occurring, just four showed a significant association: the women were more likely to become internal if (a) they went back to work and/or (b) they had problems at work; they were more likely to stay external if they had (i) separated from a partner, or (ii) argued with family or friends. The last two columns compare the occurrence of the events among those who stayed internal and those who were internal but became external. In this case, 25 of the 42 stressors were significant at the 5% level or higher, with 16 of the 25 at the .005 level and 9 of the 25 at the <.001 levels or better. We find 4 of the 16 significant events associated with staying internal: (i) having a partner who was very ill, (ii) having a partner who had problems at work, (iii) mother returning to work, and (iv) having problems at work. In contrast, three times as many stressful events, 12 in all, were associated with becoming external: (a) becoming divorced, (b) having a partner who rejected the child, (c) having a partner who was in trouble with the law, (d) separating from the partner, (e) arguing with family and friends, (f) becoming homeless, (g) having major financial problems, (h) getting married, (i) being physically abused by her partner, (j) being emotionally abused by her partner, (k) having a partner who emotionally abused her children, and (l) attempting suicide. As was found for mothers, partners had more stressful events associated with externality than internality (Table 3b). Only one event was significantly associated with changing from an external to an internal orientation: starting a new job. Four were associated with him being significantly less likely to become internal; (a) losing his job, (b) arguing with his partner, (c) arguing with family and friends, and (d) having major financial problems. Similar to the mothers, partners had more stressors, 11 in all, associated with becoming external. Four of the 11, were associated with being less likely to become external – (a) a friend or relative was ill, (b) he moved home, (c) he had a burglary, and (d) his partner became pregnant. Seven were associated with becoming external; (i) having been very ill himself, (ii) losing his job, (iii) reduction in his income, (iv) arguing with family and friends, (v) having major financial problems, (vi) being emotionally abused by his partner, and (vii) his partner starting a new job. LOC between spouses and over time ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This study is among the few that has obtained LOC scores from a large representative population composed of parental dyads at two time points, but is unique because the initial test took place prenatally, and the second six years later Hajek and Konig (2017) collected LOC information from a large representative cohort of adults over a five year period from 2005 to 2010, but did not delineate where participants were in regards to family or examined factors that were associated with changes in LOC. On the other hand, Schneewind (1997) did focus on LOC in parents of children who were first tested when children were aged 10 and tested again 16 years later. However his participants were a select group of nearly two hundred family members whose initial LOC test was not obtained before the child was born. By obtaining mother and father LOC scores prenatally we have the unique advantage of being able to describe parents' LOC perspectives before the arrival of the child and to evaluate the future impact of their prenatal orientations on their own and their children's behavior. Studies have already begun to show the benefit of this approach by revealing that the degree of externality in the parent dyad was associated with (1) negative outcomes in children's eating, sleeping and emotion regulation behaviors during their first 5 years of life (Nowicki, Iles-Caven, Gregory, Ellis, & Golding, 2017a) and (2) the number of teacher-rated emotional and behavioral difficulties when children were 9 and 11 years old (Nowicki, Iles-Caven, Gregory, Ellis, & Golding, 2017b). The design of the ALSPAC study and the data set it produced not only allowed us to evaluate the associations between parent prenatal LOC and child outcomes but also (1) to gather heretofore nonexistent normative data regarding the nature and stability of the spouses' LOC within their dyads prior to their child's birth and 6 years later, as well as (2) to obtain indicators of the stability of adult male and female LOC scores from before a child was born to a time 6 years later. We know of no other study that has obtained extensive normative LOC information from a large and representative population of adult women and men who were in a relationship with one another over time. Our results favor the similarity prediction (Byrne, 1969), but just barely. Correlations were positive but low at both testing times. While mothers and their partners tend to share a similar LOC perspective, it is only that, a tendency, and the reality is that for most parent dyads, various LOC combinations exist. To examine associations that might arise from different combinations with child outcomes, Nowicki et al. (2017a) created four combinations of prenatal LOC; internal/internal; mother internal/father external; mother external/father internal; and external/external. They found that as externality increased in the parental dyad so did associations with negative children's adjustment outcomes. Associations with stability and change in LOC over six years ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Moderate positive correlations between pre- and post-test scores indicate that many individuals were changing their LOC orientation over time. We predicted and found that LOC changes, especially toward externality, were associated with lifestyle, social, and stressful event variables. What is clear from the findings is that greater stress and less stability in personal relationships, health, and financial matters are related to remaining or becoming external for both mothers and their partners. In terms of relationships perhaps none are more important than the one between the parents themselves. Externality was associated with stressful spousal issues in the relationship including physical or emotional abuse of one another or other indicators of relationship failure such as divorce, trouble with the law or rejection of the child. Although mothers had more relationship stressors associated with externality than their partners, the direction of the association was the same. Besides their own relationship difficulties being related to externality, both parents indicated that “arguing with family and friends” added to the potential to remain or become external. If mothers and their partners didn't get along then an inability to relate to family and friends would probably increase the likelihood of being associated with feelings of externality. The importance of the spousal relationship was further substantiated by the finding that living together or living apart was related to stability and change in LOC. External mothers were more likely to move toward internality if they had a live-in partner in contrast to prenatally internal mothers who, if they did not have a live-in partner, moved toward externality; the same association was found for mothers' partners as well. Financial matters and “work” were also associated with externality for both mothers and their partners. For mothers, externality was associated with “having major financial problems” and “not going back to work” while for men it was “losing a job”, “reduction of income”, and “partner starting new job”. “Work” may be especially important for men's LOC as suggested by the finding that “starting a new job” was the only event associated with internality for them. Illness was associated with LOC differently for mothers and their partners. For men, “having been very ill himself” was significantly related to externality while for women “having a partner who was very ill” was associated with internality. In this regard, it is noteworthy that the four events endorsed by women who stayed internal (having had a partner who was very ill; having had a partner who had problems at work; having to go back to work herself; and having problems at work) involved problems in their spousal and work relationships. Perhaps meeting challenges and responding successfully to them facilitated internality in women. Unfortunately, we do not have access to what exactly the problems were or what the women thought, felt and did while experiencing them, but we might learn much from interviewing such women to obtain this information. Lifestyle differences were also related to LOC stability and change. Clearly, smoking was a significant marker of externality in the ALSPAC population. As smoking increased so did externality. While we cannot establish cause and effect with our design, the fact that smoking is associated with externality and externality in turn is associated with negative outcomes in personality, health, achievement, and adjustment suggests it would be worthwhile having a closer look at how those variables affect one another. Past research suggested a connection between smoking and LOC (Wallston & Wallston, 1978). Because of the association between smoking and externality, it might be helpful to see if stop-smoking programs had any effect on participants' LOC. Unfortunately, research has focused on determining who is more likely to successfully complete a stop-smoking program; not surprisingly most often it is the internals and not the externals (e.g. Kaplan & Cowles, 1978). One potential way around the fact that externals do not respond as well to more general smoking cessation programs is to use individualized interventions favoring those with external as opposed to internal control expectancies (e.g. Steffy, 1975) to see if externals who were successful changed toward internality. Having a choice about where we live is important, so it is not surprising that ALSPAC participants who had less choice of where to live and found their only choice was to live in public housing rather than renting or owning other dwellings are more external. To offset a lack of choice in where to live, it might be useful to give those in public housing opportunities to exercise some “control” over their living situation. Researchers have found that institutionalized individuals who were given control over small decisions in their lives responded by changing toward internality (e.g. Aasen, 1987). Perhaps pilot programs in which housing authority personnel and those living in public housing could work together to generate lists of ways in which public housing occupants would have the power to make more of their own choices; this might lead to changes toward internality over time. Strengths and limitations ~~~~~~~~~~~~~~~~~~~~~~~~~ Among the strengths of this study are the following: (i) it is based on a whole population, defined geographically and consequently not biased by particular selection (e.g. college students). Comparison with census data from 1991 show that the study participants are broadly representative (Fraser et al., 2013). (ii) The numbers considered are far larger than in any other study of LOC over time. (iii) Data on life events were obtained at times between the measures of the LOC orientation of the parents, and consequently cannot have been influenced by the LOC response. (iv) It is the only study to be able to compare the pairs of parents across time with respect to their LOC orientation. The disadvantages concern: (1) the attrition rate in this study, which was particularly high for the mothers' partners; (2) other modes of statistical analysis, choice of confounders and/or unobserved time-constant variables may provide different results; and (3) the fact that we are unable to confirm our results by comparing results with another study, since there is no (to our knowledge) study that has similar longitudinal data available. In consequence these results should be considered as hypothesis generators. Hopefully they will generate further studies over time.","Thousands of studies confirm the importance of LOC, yet we have learned little from them about true stability and/or change of LOC over time or what is associated with change when it occurs. The present study provides data to begin to better understand LOC, although it is important to note the paternal attrition over time. We found that: (1) LOC in parent dyads is correlated in a low, positive manner; (2) adult LOC is moderately correlated across time; (3) women become more internal from pregnancy to motherhood, whereas their partners become slightly more external; and (4) greater stress especially in relationships, finances, and health, is associated with increasing externality. Future designs need to build in techniques to gather information that would allow for the assessment of the short-term changes that may have been obscured by adaptation/habituation processes over a 6-year span. In this way we can gather additional information to help us to eventually identify variables that may be linked to change in LOC orientations. Hopefully our findings will help to elucidate the role of stressors in the development of internal and external LOC and bring us closer to being able to construct intervention programs to foster the development of appropriate internality, especially in expectant parents. The following is the supplementary data related to this article. Availability of parental LOC data in pregnancy for whole sample and for the parents 6 years after the baby was born, showing distributions of background factors. Supplementary data to this article can be found online at https://doi.org/10.1016/j.paid.2018.01.017."],["We investigated the relationship between semantic knowledge and word reading. A sample of 27 6-year-old children read words both in isolation and in context. Lexical knowledge was assessed using general and item-specific tasks. General semantic knowledge was measured using standardized tasks in which children defined words and made judgments about the relationships between words. Item-specific knowledge of to-be-read words was assessed using auditory lexical decision (lexical phonology) and definitions (semantic) tasks. Regressions and mixed-effects models indicated a close relationship between semantic knowledge (but not lexical phonology) and both regular and exception word reading. Thus, during the early stages of learning to read, semantic knowledge may support word reading irrespective of regularity. Contextual support particularly benefitted reading of exception words. We found evidence that lexical–semantic knowledge and context make separable contributions to word reading. --------------------------------------------------------------------------------","Knowledge of the meaning of words and phrases (semantic knowledge) has an important role to play in reading. Logically, a child needs to understand the meaning of the words and phrases contained within a text in order to fully understand it. The simple view of reading (e.g., Gough & Tunmer, 1986), an influential framework for understanding reading comprehension, posits that successful reading comprehension is underpinned by oral language comprehension (including semantic knowledge) as well as word reading abilities. Indeed, studies adopting longitudinal and experimental (randomized controlled trial) designs (e.g., Clarke, Snowling, Truelove, & Hulme, 2010; Nation & Snowling, 2004) have yielded convincing evidence that semantic knowledge is causally related to reading comprehension ability. There is also evidence that oral language ability contributes to the development of word reading in children, with influences from both phonology and semantics (e.g., Duff & Hulme, 2012; Nation & Cocksey, 2009; Nation & Snowling, 2004; Ouellette & Beers, 2010; Ricketts, Nation, & Bishop, 2007). We concentrated here on semantic influences. Nation and Snowling (2004) showed that semantic knowledge at age 8 years predicted later word reading at age 13 years after accounting for decoding ability, phonological skills, and the autoregressor (word reading at age 8 years). In an extension of this research, Ricketts and colleagues (2007) demonstrated a more specific relationship—that oral vocabulary knowledge was more closely associated with exception word reading than with regular word reading. Exception words are words with unusual mappings between spelling and sound (e.g., <yacht>, <pint>), whereas regular words contain only predictable spelling–sound mappings. Importantly, regular words can be readily decoded using knowledge of the usual relationships between spelling patterns (graphemes) and sounds (phonemes), whereas exception (or irregular) words cannot (e.g., using such a strategy would result in <yacht> being pronounced to rhyme with “matched” rather than “cot”). Regular words are usually read more accurately than exception words by typically developing children (e.g., Nation & Cocksey, 2009). In the literature outlined above, receptive and/or expressive oral vocabulary measures have typically been used to assess semantic knowledge. It is worth noting that the acquisition of oral vocabulary or lexical–semantic knowledge is incremental rather than an all-or-nothing process, with individuals adding to existing lexical–semantic representations, as well as acquiring new representations, throughout the lifespan. Studies conducted by Ouellette and colleagues (e.g., Ouellette, 2006; Ouellette & Beers, 2010) have acknowledged this by making a distinction between breadth (number of words known) and depth (what is known) in vocabulary knowledge. Ouellette and Beers (2010) found that for children aged 5 to 7 years a depth measure was a significant predictor of exception word reading, whereas a breadth measure was not; the reverse pattern was observed for older readers (11–12 years). Oral vocabulary is an important part of semantic knowledge. However, semantic knowledge also encompasses an understanding of the meaning-based relationships between words, the meaning of phrases, and so on. As far as we have ascertained, the study by Nation and Snowling (2004) is unique in investigating the relationship between semantic knowledge and word reading by using not only the usual measure of oral vocabulary (in this case an expressive measure) but also a measure that goes beyond such lexical–semantic knowledge—a composite of “semantic skills” comprising semantic fluency and synonym judgment. In regression analyses, Nation and Snowling found that their two measures of semantic knowledge made equivalent contributions to explaining variance in word reading, as measured concurrently and longitudinally by a well-established standardized test. However, their analysis of exception word reading, more specifically, showed that oral vocabulary at age 8 years was a significant predictor of exception word reading 4 years later, whereas the semantic composite was not. A number of mechanistic accounts for the relationship between semantic knowledge and word reading have been proposed. Walley, Metsala, and Garlock (2003) suggested that the relationship between semantic knowledge and word reading is indirect. According to their lexical restructuring hypothesis, oral vocabulary development serves to specify phonological representations, which in turn are critical for word reading development (e.g., Bishop & Snowling, 2004; Brady & Shankweiler, 1991; Goswami & Bryant, 1990). Computational models of word reading assume a more direct relationship. In the triangle model, words can be read aloud via two pathways, including one that maps indirectly from orthography to phonology via semantics (Harm & Seidenberg, 2004; Plaut, McClelland, Seidenberg, & Patterson, 1996). The dual route cascaded (DRC) model (Coltheart, Rastle, Perry, Langdon, & Ziegler, 2001) also makes reference to a semantic route; however, this route has not been implemented in its simulations, and the activation of semantics is not necessary for word reading. In the triangle model, semantic knowledge is necessary and has a particularly important role to play in the reading of exception words and for poor readers. Similarly, in his developmental account, Share (1995) argued that top-down support from semantic information helps readers to resolve decoding ambiguity (for similar proposals, see Bowey & Rutherford, 2007; Tunmer & Chapman, 2012). According to this view, when a word is encountered that cannot be readily decoded, either because it is an exception word or because the reader does not possess the requisite reading ability, semantic information relating to the context or the word can be combined with a partial decoding attempt to successfully read the word. In most studies, the relationship between semantic knowledge and word reading has been investigated by measuring both constructs and testing whether these constructs are correlated across participants, showing that there is a general relationship between some index of the semantic knowledge that individuals can access and the number of words that they can read on an unrelated measure. However, theoretical positions proposing a direct and necessary relationship between semantics and word reading (e.g., Harm & Seidenberg, 2004) motivate a more precise hypothesis of the relationship between these variables. Specifically, that knowledge of an individual word should aid reading of that particular word. This hypothesis is corroborated by evidence from semantic dementia patients, some of whom experience difficulty in reading exception words alongside their semantic impairments but who are more likely to successfully read exception words for which they know the meanings (Graham, Hodges, & Patterson, 1994; Woollams, Ralph, Plaut, & Patterson, 2007; but see Schwartz, Saffran, & Marin, 1980, for a contrasting case). In what follows, we summarize pertinent data from studies with children. Nation and Cocksey (2009) probed item-level relationships between semantic knowledge and word reading in children.","aged 7 years read lists of regular and exception words and completed auditory lexical decision and definitions tasks as indexes of phonological and semantic lexical knowledge, respectively. Nation and Cocksey found that children demonstrated phonological and semantic knowledge of the majority of words that they read correctly, and this relationship was stronger with exception words than with regular words, although a small percentage of words were read correctly without being recognized in the auditory lexical decision task or defined correctly. Across-items performance in both auditory lexical decision and definitions tasks showed equivalent correlations with word reading. In further analyses, both auditory lexical decision performance and definitions knowledge were entered into by-items regression analyses predicting exception word reading. Auditory lexical decision performance explained unique variance in exception word reading after accounting for the variance explained by definitions performance. However, definitions did not explain unique variance in exception word reading after accounting for the variance explained by auditory lexical decision. This led the authors to conclude that lexical phonology (familiarity with a word’s phonological form) is sufficient to support word reading and that possessing deeper semantic knowledge does not predict more successful reading. However, they interpreted their findings with caution due to the small sample size and the recognition that by-items performance on their auditory lexical decision task was skewed toward ceiling. Two training studies conducted by Duff and Hulme (2012, Experiment 2) and McKague, Pratt, and Johnston (2001) showed that pre-exposing children to the phonological forms of words facilitates learning to read those items, as does pre- exposure to phonology plus semantics (see also Ouellette & Fraser, 2009; Wang, Nickels, Nation, & Castles, 2013). In Duff and Hulme (2012) and McKague and colleagues (2001), pre- exposure to phonology plus semantics did not confer an additional advantage beyond pre- exposure to phonology alone, resonating with Nation and Cocksey’s (2009) claim that lexical phonology is sufficient to support word reading. In contrast, adult studies have indicated that semantic pre-exposure supports learning to read exception words over and above pre-exposure to phonology alone (McKay, Davis, Savage, & Castles, 2008; Taylor, Plunkett, & Nation, 2011), consistent with data from semantic dementia patients (for a review of relevant research, see Taylor, Duff, Woollams, Monaghan, & Ricketts, 2015). Taken together, findings are mixed. In relation to ideas put forward by Share (1995) and others (Bowey & Rutherford, 2007; Tunmer & Chapman, 2012), knowing a word’s phonological form may be sufficient to support partial decoding attempts, but knowledge of semantics may also be important. Resolving this issue was one motivation for our study. The current study ~~~~~~~~~~~~~~~~~ We investigated whether semantic knowledge predicts word reading in 6- and 7-year-old children, bringing together two approaches that have been used to explore this relationship. In the first approach, we measured semantic knowledge and word reading using standardized tests and also asked children to read lists of regular and exception words to assess whether there is a general relationship between semantic knowledge and word reading (cf. Nation & Snowling, 2004; Ouellette & Beers, 2010; Ricketts et al., 2007). As in Nation and Snowling (2004), we measured both lexical–semantic knowledge (expressive vocabulary) and broader semantic knowledge (semantic relations between words). Our measure of lexical–semantic knowledge was an expressive oral vocabulary measure that captured depth as well as breadth; such measures have been found to predict exception word reading more strongly than measures of breadth alone in children of this age (Ouellette & Beers, 2010). Our measure of broader semantic knowledge assessed awareness of meaning-based relationships between words. In our second approach, we investigated item-specific relationships between word knowledge and word reading (after Nation & Cocksey, 2009). We exposed children to lists of regular and exception words in tasks assessing word knowledge (auditory lexical decision and definitions) and reading (reading in isolation and reading in sentence context) to probe whether knowledge of a word’s phonological form or semantic attributes would predict the ability to read that particular word. Our study builds on previous work by assessing reading in a more naturalistic contextualized task in addition to the reading in isolation approach adopted by the majority of studies. Notably, children typically read words, particularly exception words, more accurately in context (Archer & Bryant, 2001; Nation & Snowling, 1998). We also extend previous research by using mixed- effects models to estimate item-specific relationships between word knowledge and word reading while accounting for error variance due to participants and items. In sum, we took a novel approach to probing the mechanisms underpinning the relationship between word knowledge and word reading by (a) investigating general and item-specific relationships in the same study with the same children, (b) measuring richer semantic knowledge using the semantic relationships task as well as oral vocabulary, and (c) measuring word reading in context as well as in isolation. Our hypotheses were as follows. First, we hypothesized a general relationship between semantic knowledge (both vocabulary and semantic relationships) and word reading (Nation & Snowling, 2004) that would be stronger for exception words than for regular words (Ricketts et al., 2007). Second, we predicted an item-specific relationship between word knowledge (as indexed by auditory lexical decision and definitions) and word reading, again expecting that this relationship would be stronger for exception words (Nation & Cocksey, 2009). We further predicted that auditory lexical decision might be an equivalent or stronger predictor of word reading compared with definitions (Nation & Cocksey, 2009). Finally, we expected that regular words would be read more accurately than exception words (Nation & Cocksey, 2009), words would be read more accurately in context than in isolation (Archer & Bryant, 2001), and this contextual facilitation effect would be more pronounced for exception words than for regular words (Nation & Snowling, 1998; Share, 1995). Participants ~~~~~~~~~~~~ A sample of 27 children (10 boys) aged 6 and 7 years participated in this study (M = 6.50 years, SD = 0.26). All children from one year group attending two schools serving socially mixed catchment areas in Birmingham, United Kingdom, were invited to take part provided that they spoke English as a first language and did not have any recognized special educational need. Data were collected and analyzed from all children for whom informed parental consent was received. Children had experienced 2 years of formal literacy instruction. Ethical approval was provided by the ethics committee at the Institute of Education, University of London. Standardized tasks Children completed standardized tasks in two sessions, each lasting approximately 30 min. Sessions were separated by approximately 1 week (mean amount of time between testing sessions = 5.26 days, SD = 1.58). All background measures were published standardized tasks and were administered according to manual instructions in a fixed order across the two sessions. Nonverbal reasoning was measured using the Matrix Reasoning subtest of the Wechsler Abbreviated Scale of Intelligence (WASI; Wechsler, 1999), which is a pattern completion task. Word- level reading was assessed using the Phonemic Decoding Efficiency (PDE) and Sight Word Efficiency (SWE) subtests of the Test of Word Reading Efficiency (TOWRE; Torgesen, Wagner, & Rashotte, 1999). In each subtest, children are asked to read a list of nonwords (PDE) or words (SWE) of increasing length and difficulty as quickly as they can. Efficiency was indexed by the number of nonwords or words read correctly in 45 s. Semantic knowledge was indexed by the Vocabulary and Similarities subtests of the WASI (Wechsler, 1999). The Vocabulary subtest is a measure of expressive vocabulary that requires children to verbally define words. The Similarities subtest measures knowledge of the semantic relationships between words; children are presented with two semantically related words and are asked to describe how these words are related in meaning. Experimental tasks Children were exposed to 40 words in the context of four tasks: two assessing reading (reading in isolation and reading in context) and two indexing lexical knowledge (auditory lexical decision and definitions). Tasks were completed in the following fixed order: auditory lexical decision, reading in isolation, definitions, and reading in context. Tasks were presented in this order to limit contamination across tasks. Nonetheless, repetition effects were possible and were confounded with the isolation versus context manipulation. However, the first three tasks were completed during the first session, and the final task was completed during the second session. Thus, the reading tasks were completed on separate days. All tasks were separated by time and interleaved with filler tasks to minimize children’s awareness of the repetition of items. The auditory lexical decision task was included to assess children’s familiarity with the phonological forms (lexical phonology), and the definitions task was administered to tap item- specific lexical–semantic knowledge (lexical semantics). Stimuli Stimuli are included in the Appendix and comprised 20 regular words and 20 exception words, taken from longer lists in the Diagnostic Test of Word Reading Processes (DTWRP; Forum for Research in Literacy & Language, 2012). Regular words included only graphemes that were pronounced according to grapheme–phoneme correspondence (GPC) rules (Rastle & Coltheart, 1999), whereas exception words included one or more graphemes with pronunciations that deviated from these rules (e.g., the <s> in <sugar> has an atypical pronunciation). The stimuli included monosyllabic and multisyllabic words. Because stress patterns affect pronunciation in multisyllabic words, during DTWRP design an expert panel of psychologists, linguists, and psycholinguists provided consensus that the regular multisyllabic words were pronounceable using usual grapheme–phoneme mappings. All words selected for the current study could be used as nouns. Regular and exception word lists were closely matched (all ps > .05) on length measured in phonemes, letters, or syllables and on printed word frequency, where available from the Children’s Printed Word Database (Masterson, Dixon, Stuart, & Lovejoy, 2003), otherwise from the CELEX Lexical Database (Baayen, Piepenbrock, & van Rijn, 1993). In addition, lists were matched (all Fs < 1) for bigram token frequency, bigram type frequency, trigram token frequency, trigram type frequency, and number of orthographic neighbors (data from N-Watch; Davis, 2005). See Table 1 for a summary of the stimulus characteristics of the regular and exception words. Reading tasks In the first reading task, children read each word aloud in isolation. In the second reading task, children read each word in a sentence context, with each word appearing at the end of a sentence stem ranging in length from four to nine words. In each trial of the contextualized reading task, a sentence stem was presented on the screen first. Following this, the target word was presented. Children were asked to read sentence stems and target words aloud. Sentence stems and target words were presented separately to minimize differences between the two reading tasks. In addition, the examiner corrected any errors made while reading sentence stems to maintain comprehension for the context. Errors made while reading target words were not corrected. To develop sentence stems, regular and exception words were paired according to difficulty (using the difficulty order from the DTWRP; Forum for Research in Literacy & Language, 2012) so that sentence stems could be matched in pairs for overall printed word frequency (Masterson et al., 2003) and for length in words, letters, and syllables (all Fs < 1). A series of cloze procedures was conducted with adults to develop contexts that were not overly constraining such that participants could not readily guess the target from the sentence stem and, therefore, would need to read it. For each cloze procedure, adults were asked to complete each sentence stem (with target words missing). For the sentence stems used in this study, a maximum of 2 of 25 adults inserted the target in any one case, showing that children were unlikely to guess the target word from the sentence stem. Within isolation and context reading tasks, trials were blocked by type (exception then regular). Stimuli were presented in random order within blocks using the E-Prime program (Schneider, Eschman, & Zuccolotto, 2002a, 2002b). Words and sentences were presented in Arial 25-point font, and the approximate viewing distance was 40 cm. Words subtended an approximate mean visual angle of 4.37° to 10.82° for 4- and 10-letter words, respectively. Accuracy was calculated for each child in each task (i.e., number of words read correctly). The maximum score was 20 for each list (regular or exception) within each task. Auditory lexical decision The auditory lexical decision task was administered to determine whether children were familiar with the phonological form of each word (lexical phonology). The 40 words were presented along with an equal number of nonwords from the ARC database (Rastle, Harrington, & Coltheart, 2002) that were matched to the words for number of letters and, in most cases (80%), for initial phoneme. Items were recorded by a native speaker of English. Stimuli were presented one at a time through headphones, and children were required to make a manual key-press response to indicate whether the item was a word or not. Children completed four practice trials at the beginning of the task to ensure that they understood the task demands. Stimuli were presented in random order, and response accuracy (max = 20 for each word list) and latencies were recorded using E-Prime. Definitions Children were asked to describe what each word meant, yielding a measure of lexical–semantic knowledge. All 40 words were administered in a single random order. Items were blocked such that children responded to items from the exception word list first and then items from the regular word list. The resulting definitions (N = 1080) were scored by two independent coders as 0 (no definition/incorrect definition), 1 (partial definition), or 2 (full definition). Criteria for scoring a 0, 1, or 2 for each word were agreed to by the first author and coders beforehand. The coders then scored each definition without any consultation. There was a high degree of inter-rater reliability, r(1080) = .96. Nonetheless, the coders discussed each discrepant score in turn (with advice from the first author), reaching consensus in all cases. A total definitions score (max = 40 for each list) was calculated for each child.","Mean normative scores were at or near the average range on standardized assessments of nonverbal reasoning, semantic knowledge, and word-level reading (see Table 2 for a summary). High reliability estimates are reported for all tasks. Table 3 summarizes performance by participants and by items, as well as reliability estimates (Cronbach’s alpha), for experimental word tasks. Reliability estimates were acceptably high for most tasks but were relatively low for auditory lexical decision. We next present findings on (a) correlation and regression analyses exploring general relationships between semantic knowledge and word-level reading (with scores calculated by participants in the more traditional way) and (b) mixed-effects models that probe effects of regularity (regular vs. exception) and reading task (isolation vs. context), as well as item-specific relationships between semantic knowledge and word-level reading (taking into account random effects due to participants or items). General relationships between semantic knowledge and word-level reading ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 4 presents bivariate parametric correlations (by participants) between raw scores on standardized measures of semantic knowledge (vocabulary, similarities) and all reading tasks (TOWRE PDE, TOWRE SWE, and regular and exception word reading in both isolation and context). Pertinent to our hypotheses, Table 4 indicates medium to large correlations between each measure of semantic knowledge and each measure of word-level reading. Contrary to our expectations, semantic variables were not more closely related to exception word reading than to regular word reading, and across word reading tasks performance was less highly correlated with scores on the vocabulary task than with scores on the similarities task. For nonword reading (measured by the TOWRE PDE), scores showed a higher correlation with vocabulary than with similarities, although coefficients were similar. A series of regression analyses (see Table 5) was then conducted to probe whether semantic knowledge explains additional variance in word reading after accounting for variance explained by phonological decoding ability (measured by TOWRE PDE score), which was entered at the first step. Separate analyses were conducted with performance on each word reading measure (number of words read correctly by each participant) as the outcome variable. Decoding was a significant independent predictor of each word reading measure in each analysis. After accounting for the variance that decoding explained, similarities but not expressive vocabulary explained additional variance in each outcome variable. These models explained between 62% and 79% of the variance in word reading. Models with similarities explained more variance (67%–79%), with similarities explaining approximately 5% to 10% of that variance. In summary, there was a clear relationship between semantic knowledge and word reading ability; this was more marked for the similarities task. Mixed-effects models ~~~~~~~~~~~~~~~~~~~~ In generalized linear mixed-effects models (GLMMs), we examined the factors that influenced the log odds of response accuracy, including fixed effects due to item regularity (regular vs. exception), experimental reading task (in isolation or in context), and word knowledge (auditory lexical decision or word definitions scores), as well as random effects due to variation in overall accuracy (random intercepts) or in the slopes of the fixed effects (random slopes) associated with differences between sampled participants or stimuli (Baayen, Davidson, & Bates, 2008). This approach allowed us to avoid the problems associated with analyzing dichotomous outcomes using linear models (discussed by, e.g., Baayen, 2008; Dixon, 2008; Jaeger, 2008). We analyzed 2160 observations—27 children reading 20 regular and 20 exception words, once in each of the isolated and context conditions—using the glmer function in the lme4 package (Bates, Maechler, Bolker, & Walker, 2014) in R (R Core Team, 2014). We tested the relative utility of including hypothesized fixed effects or potential random effects in our models by performing pairwise likelihood ratio test (LRT) comparisons (Barr, Levy, Scheepers, & Tily, 2013; Pinheiro & Bates, 2000) of simpler models with more complex models, where the former are nested within the latter. In the following, we outline the results of the model comparisons but report only estimates of fixed and random effects for the final model. Interested readers are invited to contact the first author for supplementary material, including the data, code used in analyses, and estimates associated with intermediate models. First, we tested our hypotheses by progressing through a series of models with varying fixed effects but the same random effects, starting with a model of the log odds of response accuracy with no fixed effects and just the random effects of participants and items on intercepts (average accuracy)—an “empty model.” Compared with the empty model, a model including terms corresponding to regularity, reading task, auditory lexical decision, and definitions significantly improved model fit: LRT, χ2 = 41.08, 4 df, p < .001. In this main-effects model, there were significant effects of reading task and definitions only (both ps < .001). Our remaining hypotheses were addressed by adding interaction terms. Compared with the main-effects model, a model also including the regularity by reading task interaction improved model fit: LRT, χ2 = 5.30, 1 df, p = .021. Adding regularity by auditory lexical decision and regularity by definitions terms did not further improve model fit: LRT, χ2 = 0.47, 2 df, p = .789. Thus, we adopted a final model that included the main-effects and regularity by reading task terms. Following Baayen (2008; see also Pinheiro & Bates, 2000), we examined whether both random intercepts terms were required by performing pairwise LRT comparisons of models with the same fixed effects as the final model but varying random effects as follows: (i) a model with both random effects of participants and items on intercepts, as in the models detailed in the forgoing; compared with (ii) a model with just the random effect of participants on intercepts; and compared with (iii) a model with just the random effect of items on intercepts. We found that both random intercepts terms were warranted by improved model fit to data (inclusion of a random effect of participants on intercepts: LRT, χ2 = 754.85, 1df, p < .001; inclusion of a random effect of items on intercepts: LRT, χ2 = 585.05, 1df, p < .001). In Models (ii) and (iii), the pattern of significant effects remained largely the same, with significant effects of reading task and definitions and a regularity by reading task term that was nearly significant (Model ii: p = .058; model iii: p = .094). However, in each model, the auditory lexical decision effect was also significant (both ps < .001). Thus, when variation relating to either participants or items alone was taken into account, both auditory lexical decision and definitions showed a significant relationship with word reading. However, after simultaneously accounting for variation relating to both, only the definitions effect remained. Following Barr and colleagues’ (2013) recommendations, we examined the importance of random slopes (random differences between participants or between items in the slopes) of the fixed effects due to reading task, word knowledge, or the regularity by reading task interaction. We did this by testing whether model fit was improved by the inclusion of terms corresponding to random effects of participants or items on the slopes of the fixed effects. We found that a model including terms corresponding to random effects of participant differences on the slopes of word regularity and reading task effects, and corresponding to random effects of item differences on the slopes of both word knowledge measures (definitions and auditory lexical decision), significantly fit the data better than a model including the same fixed effects and just random intercepts: LRT, χ2 = 34.46, 10 df, p < .01. Thus, including the observed variability between participants in the slopes of both regularity and reading task effects, and between responses to different items in the slope of the word knowledge effect, improved model fit. Table 6 summarizes the final model, with fixed effects due to regularity, reading task, the regularity by reading task interaction, and word knowledge (scores on auditory lexical decision and definitions tests) as well as random effects of participants and items on intercepts and on the slopes of the fixed effects. The estimated coefficients for the final model show that reading accuracy was higher for the context (vs. isolation) task and for words that had been defined more accurately. Furthermore, the model revealed a regularity by reading task interaction. Inspection of Table 3 indicates that a regularity effect was more evident when words were read in isolation rather than in context and that the influence of context was greater for exception words than for regular words. Contrary to our hypotheses, we found that (a) semantic knowledge showed equivalent relationships with regular and exception word reading and (b) auditory lexical decision performance was not associated with word reading.","The results support our primary hypothesis that variation in semantic knowledge is associated with variation in word reading performance. Indeed, we have provided robust evidence for this by observing this association, for the first time, across both general and item-specific analyses. We have extended previous findings on reading words in isolation by assessing word reading in context, which is more akin to how children encounter words naturally. We observed an interaction between context and word type such that sentence context particularly facilitated reading of exception words, in line with previous studies (Nation & Snowling, 1998). It is worth noting that the reading in isolation task was always administered before the reading in context task; thus, any contextual benefit must be interpreted with caution because some improvement might be attributable to practice effects. However, tasks were separated by approximately 1 week. Furthermore, the order of the tasks does not invalidate our finding of an interaction between context and word type, nor does it affect our key findings that semantic knowledge (as measured by the similarities task) predicted reading in context as well as reading in isolation in regression analyses and that lexical–semantic knowledge and contextual effects were independently predictive of word reading in our mixed-effects analyses. Theories of word reading focus almost exclusively on reading in isolation. Nonetheless, our findings are consistent with developmental theories that highlight the importance of contextual support for word reading (Share, 1995) and with the triangle model’s (yet to be implemented) assumption that semantics and context exert separable but interacting effects on reading aloud (Bishop & Snowling, 2004; Seidenberg & McClelland, 1989). We hope that the current study, along with other empirical studies of word reading in context (Martin- Chang & Levesque, 2013; Nation & Snowling, 1998), will pave the way for research that aims to probe the mechanisms that underpin word reading as it occurs naturally. An important first step will be to specify how context supports word reading, why this might be more beneficial for exception word reading than for regular word reading, and why this effect was separable from that of item-specific semantic knowledge in our analyses. In the current study, context was provided at the sentence level and may have conveyed useful semantic information along with other cues (e.g., grammar). Thus, one plausible interpretation of our findings would be that semantic information from the context supported word reading, and this was more effective for exception words than for regular words. However, this interpretation is premature; our data do not address whether this effect was driven by semantic information or other cues provided by context. We found that semantic knowledge showed equivalent relationships with regular and exception word reading, a finding that we replicated across by-participants regression analyses and mixed-effects models. To the extent that our measures of semantic knowledge map onto the way in which semantic representations are activated in the triangle model, this finding contrasts with the triangle model, where semantic knowledge is seen as more important for exception word reading than for regular word reading (e.g., Harm & Seidenberg, 2004; Strain, Patterson, & Seidenberg, 1995; but see Woollams et al., 2007, for effects of semantics on regular word reading within a triangle model framework). Notably, it is also at odds with pertinent developmental findings that semantic knowledge shows a closer relationship with exception word reading than with regular word reading in English- speaking children (Nation & Cocksey, 2009; Ricketts et al., 2007), whereas it is in accord with emergent findings from English-speaking children indicating relationships between semantic variables and both regular and exception word reading (Duff & Hulme, 2012; Mitchell & Brady, 2013; see also findings from Spanish-speaking adults reported by Davies, Barbón, & Cuetos, 2013, and from English-speaking adults reported by Strain & Herdman, 1999). It remains to be seen whether this finding is predicted by the DRC model (Coltheart et al., 2001) given that current instantiations have not yet simulated the role of semantics in word reading development (see Taylor, Rastle, & Davis, 2013). There are a number of possible explanations for discrepancies between our observations and previous findings. Marked ceiling effects on regular word reading could explain weaker relationships between semantic knowledge and regular word reading in previous studies (Nation & Cocksey, 2009; Ricketts et al., 2007). Another possibility concerns the age and reading ability of participants. Semantic knowledge may contribute more indiscriminately to word reading during the early stages of reading development when children have limited knowledge of orthography-to-phonology mappings (as in our study; for a similar argument, see Duff & Hulme, 2012). With reading experience, the role of semantics in regular word reading may decrease such that a closer relationship between semantic knowledge and exception word reading emerges. In addition, the impact of semantic knowledge on word reading may be influenced by item-level characteristics such as length, frequency, familiarity, and meaning (Mitchell & Brady, 2013). Indeed, our set of regular words were harder to define than our set of exception words. This could go some way to explaining the finding that semantic knowledge contributes to both regular and exception words. Future research should aim to explore the conditions under which semantic knowledge affects regular word reading, adopting developmental designs and varying stimulus characteristics. In correlation analyses (by participants), all standardized measures of semantic knowledge and word-level reading were inter-correlated. However, knowledge of semantic relationships (similarities) was consistently more highly correlated with word reading than with oral vocabulary knowledge. After controlling for decoding skill, regression analyses showed that similarities, but not expressive vocabulary, predicted word reading. One possible explanation for this finding is that scores on the similarities measure were more varied than scores on the vocabulary measure such that the similarities measure may have captured more fully the variability in semantic knowledge in our sample. The hypothesis that performance on the similarities measure was systematically more varied than performance on the oral vocabulary measure could be explored in future research. Previous studies that have investigated general relationships between more than one semantic measure and word reading have shown that the semantic predictors of word reading (after controlling for decoding) vary according to the age of the participants and the outcome measures used in analyses. In Ouellette and Beers (2010), both depth and breadth of vocabulary knowledge were measured. In younger participants (5–7 years) depth but not breadth predicted irregular word reading, whereas in older participants (11 and 12 years) the opposite pattern was observed. Nation and Snowling (2004) employed a measure of vocabulary and a “semantic composite” (semantic fluency and synonym judgment). Both measures predicted word reading concurrently, but only oral vocabulary was a longitudinal predictor of exception word reading. The finding that oral vocabulary did not predict word reading in our regression analyses also contrasts with the item-specific effects detected in our mixed- effects models. This seems surprising given that the definitions task was designed to parallel the standardized expressive vocabulary measure that we used by asking children to define words and adopting a three-point scoring approach. Plausibly, this discrepancy could be explained by differences in the variables included in the models. We controlled for decoding ability in our by-participants analyses so that we could examine the relationship between semantic knowledge and word reading after accounting for the substantial variance in word reading explained by decoding skill (this is a standard approach; see, e.g., Ouellette & Beers, 2010; Ricketts et al., 2007). However, we did not include decoding ability in our mixed-effects analyses because the models were specified as confirmatory analyses (following, e.g., Barr et al., 2013) of the effects of the following experimental factors: reading task, regularity, and word knowledge type. Nevertheless, the addition of decoding ability to the final model did not change the pattern of results. Different findings across our analytical approaches could instead reflect the way in which our mixed-effects models capture a specific relationship between knowledge of an item and reading that same item, whereas the by-participants regressions explore a more general relationship between a measure of children’s lexical–semantic knowledge, which could act as a proxy for their item-specific semantic knowledge or their ability to use context, and their ability to read a separate set of words. Arguably, this general relationship could be weaker. Taken together with the mixed findings discussed in the preceding paragraph, it is clear that although the relationship between semantic knowledge and word reading is robust, the precise pattern of findings observed varies across analyses and data sets. Notably, however, our observations show that semantic knowledge is predictive of word reading ability. Mixed-effects models demonstrated that correctly defining a word was a significant predictor of accurately reading that word, whereas accepting it as a word in our lexical decision task was not. This result was unexpected given that in Nation and Cocksey (2009) performance on definitions and auditory lexical decision tasks showed equivalent (significant) correlations with word reading and that auditory lexical decision was the stronger predictor in by-items regression analyses (for similar findings, see Duff & Hulme, 2012, Experiment 2; McKague et al., 2001). Thus, we did not replicate Nation and Cocksey’s (2009) finding that auditory lexical decision predicts word reading, nor did we provide support for their proposal that lexical phonology is enough to support word reading (i.e., lexical–semantic knowledge provides no additional benefit). Instead, our findings indicate that it is lexical–semantic rather than lexical–phonological knowledge that supports word reading. Other investigations of the relative importance of lexical phonology and semantics for word reading have indicated that semantic knowledge is a better predictor of reading success than phonological knowledge (Duff & Hulme, 2012, Experiment 1; McKay et al., 2008; Taylor et al., 2011), resonating with our findings. One plausible explanation for the discrepancy between our study and that of Nation and Cocksey (2009) relates to the different analytic approaches adopted in the studies. In Nation and Cocksey’s study, correlations and regressions were conducted across items (an F2 by-items analysis), thereby taking into account random error variance due to the items. In contrast, our final mixed-effects model incorporated random error variance due to both participants and items. Thus, accounting for both sources of error variance could have “washed out” the effect of auditory lexical decision. Indeed, when our model accounted for either error variance due to participants (akin to F1 by- participants analyses) or error variance due to items (akin to F2 by-items analyses), we replicated Nation and Cocksey’s finding; both definitions and auditory lexical decision performance predicted word reading. Analyses reported by Baayen and colleagues (2008) indicate that fixed effects are better estimated in repeated measures studies when both random participants and item effects are taken into account (see also Barr et al., 2013). Essentially, these models specify, rather than assume, the random variation in the data that is due to participants (in this case variation in children’s reading accuracy) and items (in this case variation in performance in response to individual words). It is possible that our findings would be replicated in Nation and Cocksey’s data if mixed- effects models were applied, supporting a conclusion that lexical semantics, but not lexical phonology, affects word reading. Caution is warranted in interpreting our auditory lexical decision results. Reliability for this task was low, and post hoc consideration of its stimuli has highlighted its limitations. Following Nation and Cocksey (2009), we selected nonwords that matched our words in terms of letter length and initial letter (or phoneme). However, we should have explicitly matched words and nonwords for number of syllables and phonemes. We checked this retrospectively, discovering that our words had approximately one more phoneme (M = 5.50, SD = 1.68 vs. M = 4.45, SD = 1.11) and one more syllable (M = 2.13, SD = 0.79 vs. M = 1.03, SD = 0.16). It is possible that this made the nonwords superficially distinctive from the words, making the task easier and reducing the extent to which lexical knowledge was used to make decisions (they could instead have been made on the basis of shallower processing). By participants, there is no indication of ceiling effects, and performance showed good variability. By items, performance again showed good variability, but scores were closer to ceiling (this is also the case in Nation & Cocksey, 2009), providing some evidence that discriminating between particular words and nonwords was fairly easy. Ceiling effects by items may also explain poor reliability (Cronbach’s alpha) on the auditory lexical decision task. Our choice of nonword distracters, therefore, may have restricted relationships between auditory lexical decision performance and reading because auditory lexical decision performance did not consistently reflect lexical knowledge or because scores on this task showed poor reliability (for further discussion of the impact of poor reliability on correlational analyses, see Vul, Harris, Winkielman, & Pashler, 2009). The nature of the nonwords used in the auditory lexical decision task has important implications for how performance on this task should be interpreted (e.g., Ernestus & Cutler, 2015). As mentioned above, superficial differences between our word and nonword stimuli may have reduced the use of lexical knowledge in making decisions. Equally, however, in tasks where nonwords are very word-like, lexical decisions are commonly assumed to reflect greater reliance on semantic processing (Binder et al., 2003). An important goal for future research will be to investigate the relative contributions of lexical phonology and semantic knowledge to word reading using more carefully controlled auditory lexical decision stimuli and/or other tasks designed to tap lexical phonology. In sum, our findings provide robust and novel support for the idea that semantic knowledge and sentence context independently support word reading (cf. Bishop & Snowling, 2004). In addition, they add to emergent evidence that lexical or semantic knowledge supports reading of regular words as well as exception words (Davies et al., 2013). If semantic knowledge is causally related to word reading success, then training knowledge of word meanings should benefit word reading. Findings from such training studies have so far been inconclusive, with some suggesting that training lexical-level phonological knowledge is sufficient to support word reading (Duff & Hulme, 2012, Experiment 2; McKague et al., 2001) and others indicating that semantic knowledge exerts an effect beyond phonology (McKay et al., 2008; Taylor et al., 2011). Future empirical and theoretical studies that adopt psychologically plausible approaches to learning and development should aim to advance our understanding of how the relationship between lexical knowledge and word reading changes with age and development and whether semantic knowledge is causally related to word reading."],["The ability to share and direct attention is a pre-requisite to later language development and has been predominantly studied through infant pointing. Precursors to pointing, such as showing and giving gestures, may display similar communication skills, yet these gestures are often overlooked. This may be due to difficulty in discerning these gestures in interaction. The current study had two aims; firstly, to identify the micro-behaviours associated with showing and giving gestures in infants under 12 months, in order to ascertain whether these form two discrete communicative behaviours. Secondly, to examine whether these micro-behaviours predicted caregiver responses to these gestures. Fine-grained coding of show and give gestures, their micro-behaviours and caregiver responses was conducted through secondary analysis of naturalistic, triadic interactions between 24 infants, caregivers and a selection of toys. Findings suggested that the micro-behaviours arm position, hand orientation and eye-gaze, were significant predictors of infant gesture type, however only arm positioning was a significant predictor of caregiver response. This suggests that early showing and giving gestures can be classified based on some associated micro-behaviours, however caregiver's responses may not be contingent on these same cues, potentially resulting in difficulty understanding infant gestures. Our findings enhance our understanding of infant communication before 12 months, provide guidance to both researchers and caregivers in the identification of infants' early shows and gives, and highlight the need for greater study of these early pre-linguistic behaviours. --------------------------------------------------------------------------------","Between 9–12 months, infants experience a transition in their interaction with the world. Systematic patterns emerge in their communicative behaviours as they begin to use deictic gestures combined with eye-gaze, vocalisations, body movements and facial expressions to engage in social interaction with a communicative partner (Bates, 1979; Igualada, Bosch & Prieto, 2015; Liszkowski, Brown, Callaghan, Takada, & De Vos, 2012). The presence of these multimodal communicative behaviours is believed to be an indicator of an infant’s joint attention abilities, and a considerable body of evidence links these skills to later language development (Kristen, Sodian, Thoermer, & Perst, 2011; Laakso, Poikkeus, Katajamäki, & Lyytinen, 1999). Generally, these early communicative skills have been studied mainly through the pointing gesture (Carpenter, Nagell, & Tomasello, 1998; Cochet and Vauclair, 2010; Tomasello, Carpenter & Liszkowski, 2007). Pointing is perceived as a tool used to initiate joint attention between the infant and adult and pointing declaratively (i.e. pointing with a motive to share or direct attention onto a specific object or event) is a good predictor of later language outcomes (Colonnesi, Stams, Koster, & Noom, 2010; Tomasello et al., 2007). Infants begin to use pointing with a communicative intent at around 11–12 months of age (Fusaroa, Vallotton, & Harris, 2014) and both experimental and non-experimental studies consistently highlight a relationship between this type of pointing and skills in both the production and comprehension of language, particularly verbal naming (see Colonnesi et al., 2010 for a review). Pointing is generally perceived as a landmark communication skill at around 12 months of age (Colonnesi et al., 2010; Liszkowski, 2010). There is evidence however to suggest that infants can engage in communicative behaviours prior to the emergence of pointing. Bates, Camaioni, and Volterra (1975) found that showing and giving behaviours emerged around 10 and 11 months respectively, whereas pointing with communicative intent did not appear until 12–13 months. The shift from showing and giving to pointing indicates the infant’s understanding of the difference between the self and objects. Pointing is believed to be more cognitively complex as it exists outside of the object context and draws attention to more distal referents (Bates, Thal, Whitesell, Fenson & Oakes, 1988). However, both showing and giving behaviours also reflect the ability to initiate joint attention and demonstrate an understanding that the adult is an agent separate from the environment and capable of engaging with an object. Support for the claim that shows and gives are precursors to pointing is presented in Cameron-Faulkner, Lieven, Theakson, & Tomasello (2015). In their study of 10–12 month old infants, shows and gives emerged prior to pointing behaviours and also had a strong association with the later use of points but not reaches (the latter of which are associated with imperative behaviours and are not deemed to be as cognitively complex in nature). Beuker, Rommelse, Donders & Buitelaar (2013) examined the developmental trajectory of specific joint attention skills and their interrelations with later vocabulary size. They found that infants who developed joint attention skills at an earlier age, specifically gestures which involved directing attention (such as showing, giving and pointing) displayed larger receptive and expressive vocabulary growth earlier in life. Thus, showing and giving gestures may be good candidates for studying the foundations of early communication and the skills that make us uniquely human. To date pre-linguistic showing and giving behaviours have been under- researched, particularly when compared to studies on pointing. A potential reason for this absence is the lack of a clear definition of the two constructs. Bates et al. (1975, 1976) highlighted the difficulty in distinguishing showing and giving behaviours, suggesting that their function is often ascertained by how others react to the social context and that often, infant intentions are misinterpreted by caregivers. They referred to shows and gives as an extension of the arm towards the adult and distinguished between the two behaviours based on whether the infant gave the toy to the adult or kept it for themselves. Clements and Chawarska (2010) built on these definitions in their study of shows, gives and points in 9 and 12 month olds with autism. Shows were defined as “a person’s arm extending toward another person’s face while holding an object” (p. 48) whereas giving behaviours were described as “placing an object in another person’s hand or pushing an object at least halfway toward another person” (p. 48). Even with these more detailed definitions, the authors noted that pointing gestures were more salient than showing gestures due to their specific hand form (i.e. an outstretched arm with the index finger extended). The lack of salience of many showing gestures creates problems from a methodological viewpoint. Typically, naturalistic research on gesture development is conducted through observation or parental diaries. Although this provides an ecologically valid measure of infants’ spontaneous gestures (Capirci, Iverson, Pizzuto & Volterra, 1996; Woodward, 2009, Crais, Douglas, and Campbell (2004) highlighted the concern researchers often have over parental report methods, (i.e. through parental diaries) as the reliability of their interpretations is questioned. Whilst researchers may be trained to recognise behaviours in infants, parents may find it difficult to recognise “researcher defined” gestures or their functions, potentially jeopardising the validity of communicative development research (Woodward, 2009). The difficulty in identifying these gestures extends to caregivers in the home environment too. Bates et al. (1975, 1979) highlighted the problem of caregivers misinterpreting these gestures as instrumental acts or overlooking this action completely. Early pointing studies have already established that children rely on verbal feedback to determine connections between their pointing gestures and intentions, and adults who respond promptly, contingently and appropriately to infant actions tend to improve infants’ subsequent production and comprehension of words (Colonnesi et al., 2010; Rowe & Goldin-Meadow, 2009). Furthermore, observation of the responses of others facilitates not only social learning, but enables the understanding of intentional communication (i.e. awareness of other people’s goals during interaction) which plays a fundamental role in language development (Elsner, Bakker, Rohlfing, & Gredebäck, 2014). Theoretically, being able to distinguish these gestures would provide greater insight into the emergence of intentional communication in pre- linguistic infants. Towards the end of their first year, infants’ communicative competencies increase and they begin to use gestures with a number of accompanying behavioural characteristics, such as systematic hand shapes and vocalisations, to help more directly express their social intentions. Exploration of these behaviours could provide insight into the different motives underlying early pre-linguistic gestures. It would also help determine whether these gestures are fully ambiguous and so interpretable only from the context of the shared interaction and preceding actions. The difficulty in pinpointing infant intentions outside of adults’ responses raises the question of whether infants formulate an intention before they hold out a toy, or if their behaviours are contingent on the adult’s response. If this were the case, it may be impossible to distinguish between early showing and giving gestures without relying on caregiver feedback. If, however, in a typical interactional context, shows and gives involved distinct behavioural cues (e.g. a particular hand position) it would allow for greater examination of the role of caregiver responses in forming and developing these gestures (Bates et al., 1975; Liszkowski, 2005). Identifying and understanding these behavioural cues is an important step in the understanding of the origins of language and social cognition in humans. Research on the pointing gesture has managed to overcome the issue of ambiguity to some extent by providing objective associated micro-behaviours to help study the development of this gesture in greater detail (Bates et al., 1975; Brooks & Meltzoff, 2008; Cochet & Vauclair, 2010; Gullberg, de Bot, & Volterra, 2008; Krause & Fouts, 1997; Liszkowski, Carpenter, & Tomasello, 2008; Woodward, 2009). Cochet et al. (2014) used a frame by frame video analysis to describe the differences between infant’s early pointing gestures versus reaching. Features such as arm extension, hand-shape and body posture were all found to demonstrate different functions of infants’ gestures. Arm extension was found to be greater with pointing gestures, and these were often accompanied by an outstretched index finger. In contrast, infants leaned further forward when reaching whereas pointing was characterised by a ‘sitting back’ posture (Lock et al., 1990). A number of other studies have looked at the impact of vocalisations and gaze alternation on early social interaction and language development (Franco & Butterworth, 1992; Grünloh & Liszkowski, 2015; Liszkowski & Tomasello, 2011). Infants tend to produce different types of vocalisations when they are engaged in social interaction compared to when they are alone (Goldstein, Schwade, Briesch, & Syal, 2010). Furthermore, vocalisations are more likely to accompany declarative pointing gestures compared to imperative reaching (Grünloh & Liszkowski, 2014). These differences suggest that declarative, communicative gestures are more closely interconnected to the vocal system and thus later language development (Cochet & Vauclair, 2010). The co-occurrence of gaze alternation with a gesture is often considered evidence of intentional communication (e.g. Cochet & Vauclair, 2010; Franco & Butterworth, 1996). Previous studies have found that gaze alternations with pointing were produced more frequently in a declarative context, particularly one that requested information (Cochet & Vauclair, 2010). All of these micro-behaviours provide a number of important developmental functions; they are a means of expressing infants’ intentions and feelings, they help engage a social partner in interaction and allow infants to display their affective experiences (Bates et al., 1975; Bates, 1979). They also provide quantitative measures to allow researchers to categorise different gestures and clarify differences based on various features. To our knowledge, these precise definitions for coding have not been looked at for early showing and giving gestures. Exploration of the associated micro-behaviours of showing and giving would allow researchers to examine these gestures to the same extent as pointing and reaching, and could provide parents with information to aid in their identification of these early gestures. The current study addresses this gap in the literature by conducting a fine-grained analysis of these early communicative gestures in naturally-occurring play, documenting early showing and giving, their associated micro-behaviours and caregiver responses to these gestures. The study has two specific research aims: Firstly, to examine whether shows and gives are two distinct behaviours by documenting associated micro-behaviours. From this we would hope to provide standardised behaviour codes to help identify and distinguish between these gestures. Secondly, to examine whether caregiver’s responses are predicted by the patterns of micro- behaviours associated with infants’ gestures. From this we hope to examine whether caregivers’ linguistic and non-linguistic responses to infants’ gestures are contingent on specific infant micro-behaviours displayed when they gesture. Dataset ~~~~~~~ The data from the current study was taken from pre-existing video corpus of pre-linguistic interaction between infants aged 10–13 months and their caregivers. The data was collected as part of a larger, longitudinal study on the emergence of proto-declarative gestures (Cameron-Faulkner et al., 2015).","24 infants (10 girls: mean age 313 days, range 267–356 days) and their mothers were recruited from the centre database at the University of Manchester Child Study Centre. All dyads were monolingual English-speakers from the north-west of the UK with no reported language delay. Families that participate at our Study Centre typically come from middle class backgrounds, though demographic information was not collected for the current study. The mothers were given travel expenses and the children were presented with a book for their participation.","Infants had to attend three monthly sessions which took place within a controlled environment in a child study lab. Two video cameras were used to capture the body movements and facial expressions of both the infants and caregivers. Infants engaged in 20–25 min of natural free-play sessions with their caregiver and a selection of age- appropriate toys (e.g. plastic cups, rattles, blocks, brushes etc.) with a variety of shapes, sizes and textures, supplied by the experimenter. All toys were picked with the aim of eliciting a declarative motive (i.e. a motive to share attention and interest) rather than an imperative. Specifically, all toys could be fully manipulated by the infants without requiring any assistance from the caregiver. The naturalistic play sessions were also established to encourage sharing interest in the toys and in terms of the target object behaviours. Infants and caregivers sat on the floor on a play rug without other distractions around them. Caregivers were asked to let their infants lead the play so that the maximum amount of infant initiated behaviours could be elicited. Sessions were divided into two 10 min phases. During phase 1, the experimenter left the room and the infant and caregiver were recorded playing together. At the end of phase 1, the experimenter returned to the room, collected the toys, and replaced them with a different selection of toys. Phase 2 repeated the social play session but with the new toys. Infants participated in these sessions once a month for 3 months. Each session was the same duration and had the same procedure. Gestures All coding was conducting using the video recordings of the original larger study (Cameron-Faulkner et al., 2015). For the original study, two trained research assistants coded the data for instances of infant showing and giving gestures based on the definitions in Table 1, as well as infant reaches and points. The coding criteria for shows and gives were broad and did not include any specific detail on the more fine-grained aspects of the gestures. The two research assistants were blind to the hypotheses of the original study. In addition to coding the infant gestures, the research assistants also categorised the interactional sequences following each instance of these communicative behaviours (i.e. shows, gives, reaches and points) with respect to eye gaze and object manipulation or maternal comment. For the current study, the first author began by recoding the data using the same broad definitions of show and gives as used in the larger study. These codes were then compared with those of the original study. We examined whether the current and previous coders identified the same behaviours as infant initiated show and give behaviours. Of the 112 instances identified by either coder, agreement was k = 0.70 (79%). Behaviours which could not be agreed upon were labelled ambiguous and removed from subsequent analyses. Of the 92 instances of communicative gestures identified by both coders, agreement was k = 0.76 (87%) for type of gesture (a show vs. a give) indicating good reliability between the categorisation of the original coders and the first author. Following the initial identification phase we established a coding scheme for the micro- behaviours by drawing on existing studies of prelinguistic communicative behaviour (see below). For each infant, only the first session to display shows and gives was coded, as we were interested in the earliest display of these gestures. The type of gesture was categorised using binary codes of 1 for a showing gesture and 0 for a giving gesture. Micro-behaviours The first author established a coding scheme for potentially relevant micro- behaviours. Micro-behaviours were selected based on previous codes used for the classification of more established gestures such as points and reaches (e.g. Brooks & Meltzoff, 2008; Cochet & Vauclair, 2010; Elsner et al., 2014). Table 2 displays the micro-behaviours coded for each show and give. The video-recordings were then re-examined and the micro-behaviours for each identified communicative gesture were coded. The micro-behaviour analysis was conducted on the data after a sufficient time lag from the identification phase and in addition the show/give categorisation was hidden. Also all show and gives which were identified during both the original and current study (as mentioned above) were included in the analysis. Vocalisation Vocalisation was scored as a binary code of 1 if the gesture was accompanied by vocalisation and 0 if it was not. These codes were then compared to the original coding. Agreement for this behaviour with the previous coding was good k = 0.73 (84%). Eye gaze Eye gaze was coded as 1 if the infant looked to the caregiver first when displaying a gesture and 0 if they primarily looked at the toy. Again these codes were compared to the coding from the original study. Coders agreement for the location of eye-gaze was k = 0.64 (66%) indicating moderate reliability. Similarly gestures that were accompanied by gaze alternation were scored as 1 if there was evidence of gaze alternation and 0 if the eye gaze was fixed, agreement between coders was again good k = 0.78 (81%). Morphological features Arm positioning was given a code of 1 if the arm was raised and 0 if it was straight out/lowered slightly. Body posture was given a code of 1 if the infant leaned forward and 0 if they did not move. Hand-orientation was split into “inverted” hand-shape (i.e. palm facing downwards) and “upwards” hand- shape (i.e. palm facing up or towards the caregiver) as these had been suggested in past literature as indicators of infant gives (Elsner et al., 2014). Codes which were difficult to define due to camera angles or the positioning of the dyads were coded as ambiguous. Caregiver responses Caregiver linguistic and non-linguistic responses to infants’ shows and gives were classified. These definitions were based on codes used in the larger study which in turn were based on pre-established categories used in the classification of caregiver contingent talk and follow-in behaviour towards infant’s gestures (Cameron-Faulkner et al., 2015; McGillion et al., 2013; Tomasello & Farrar, 1986). Table 3 displays the caregiver responses coded for each show and give gesture. These responses were split into ‘show’ responses (look and comment contingently on the object) and ‘give’ responses (reaches towards toy and asks for it). Caregiver responses were given a score of 1 for a perceived ‘show’ response and 0 if it involved for a perceived ‘give’ response. Agreement between coders for caregiver’s non-linguistic responses was k = 0.85 ‘Ignore’ responses from the caregivers were also coded whenever they failed to respond linguistically and non-linguistically to the infants' target gestures. Analysis ~~~~~~~~ To examine whether the observed micro-behaviours were linked to (1) categorisations of infant gesture type, and (2) caregivers’ interpretations of these gestures, a multiple correspondence analysis, followed by a mixed-effects logistic regression was performed in R (R Core Team, 2012). Results are organised as follows. Firstly, descriptive statistics are presented for the proportion of infant showing and giving behaviours, the associated behavioural cues and caregiver interpretations of these gestures across the selected sessions. Secondly, the graphical results of the multiple correspondence analyses are displayed and described. Finally, the mixed-effects logistic regression is reported for both infant gesture classification and caregiver response outcomes. Frequencies of shows, gives and associated micro-behaviours ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Overall, 20 participants displayed some form of showing or giving in at least one of their three sessions. Four participants were excluded due to a lack of shows and gives. There were a total of 92 episodes of gesturing across all participants. Of these 54% were classified as shows (n = 50) and 46% were labelled as gives (n = 42). There were instances of individuals who displayed only shows or only gives in their session. Infants produced an average of 2.5 shows (range: 0–8) and an average of 2 gives within their sessions (range: 0–10). Age of first use of gestures was not significantly correlated with overall number of showing and giving gestures produced (r = −0.1; p = 0.3) nor was it correlated with the type of gesture produced (r = −0.04; p = 0.7). This suggests that the age of onset for these gestures within these sessions did not predict the frequency of gestures or the type of gesture. Vocalisation accompanied just 17% of infants’ gestures whereas gaze alternation was present with gestures on average 64% of the time. Infant gaze focused primarily on the caregiver 46% of the time. Regarding the form of the gestures, 35% of gestures were characterised by the infant leaning forward and 45% of gestures displayed an ‘inverted’ hand-shape compared to a ‘palm up’ hand orientation. Gestures with a raised arm, as opposed to a straight out position also accompanied 45% of infants’ gestures. Fig. 1 displays the frequencies of the associated behaviours for both infant shows and gives. When broken down, there were clear differences in the numbers of these behaviours for infants’ showing vs. giving. Multiple correspondence analysis: associations between gesture categorisations, micro-behaviours and caregiver responses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To visually examine the pattern of relationships of the variables, a Multiple Correspondence Analysis (MCA) was conducted using the FactoMineR function in R (R Core Team, 2012). Firstly, we looked at associations between the micro-behaviours and the researcher classified shows and gives. Fig. 2 displays clusters for the micro-behaviours and the showing and giving gestures. Two fairly clear clusters of associations are apparent. In the top left quadrant, the behaviours classified as ‘show’ are highly associated with a ‘palm up’ hand position and a raised arm. In the bottom right quadrant, the behaviours classified as ‘give’ are highly associated with a ‘straight arm’ position and ‘inverted hand’ shape. We then looked at the associations between the micro-behaviours and caregivers’ responses to infant gestures (see Fig. 3). Again, there are two clusters, however these are not as distinct compared to the researcher’s classifications of infants’ showing and giving in Fig. 2. In the top left quadrant, both non-linguistic and linguistic responses classified as ‘show responses’ are clustered near to the ‘palm up’ hand position and a raised arm. In the bottom right quadrant, both non-linguistic and linguistic responses classified as ‘give responses’ are clustered near to the ‘straight arm’ position but actually are most closely associated with the infant’s lack of vocalisations. Furthermore, although these responses are positioned in the same quadrant as the researcher’s classifications of shows and gives, they are not as closely associated, suggesting a level of ambiguity when interpreting these gestures. Regression model 1: behavioural cues and infant gestures ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We then carried out two analyses to investigate whether these micro-behaviours significantly predicted (1) the classification of infants’ gestures and (2) caregivers’ linguistic and non-linguistic responses. Regression models were fitted using the lme4 function in R (Bates, Maechler & Bolker, 2012) with the two outcome variables; gesture (show vs. give) and caregiver response (show response vs. give response) used in two separate models. The following predictor variables; ‘vocalisation’, ‘gaze alternation’, ‘eye-gaze to caregiver, “arm raised”, ‘inverted hand-shape’ and ‘lean forward’ were entered as fixed effects and ‘participants’ and ‘age of first use’ as a random intercept effects. Overall fit of the model was tested using a likelihood ratio test, comparing the models with predictor behaviours entered to a baseline model (with no effects of interest) (Field et al., 2012). Before considering the main predictor variables of interest, we compared models with and without the random effect of age of first use. The addition of infant age did not provide a significantly better fit to the data (χ2 (1) = 0.92, p = 0.6) thus only ‘participant’ was retained to control for differences accounted for by individuals. Consequently, model 1 included all of the predictor variables to establish if any of the infant micro-behaviours were significant predictors of the categorisation type (show/give) assigned to infants’ gestures. Model 1 was then compared to a baseline model without the predictor variables to assess the fit of the data. Table 4 displays the coefficients and standard error results for all the micro-behaviours entered into model 1. The results showed that gaze alternation, infant posture and infant vocalisation were not significant predictors of the classification of infant gesture. In contrast, the location of infants’ eye-gaze significantly predicted gesture type classification, as eye-gaze primarily directed to the caregiver was associated with infant showing gestures. ‘Inverted’ hand-shape was also found to have a significant, negative association with infant showing, indicating that it was a significant predictor of infant gives. Likewise, in positioning was found to be a significant predictor of gesture type, as a raised arm was associated with infant showing. ‘Inverted’ hand-shape was found to have a significant, negative association with infant showing, indicating that it was a significant predictor of infant gives. When comparing model 1 to the baseline model, the predictive behaviours significantly improved the fit of the model (χ2 (9) = 53.0, p < 0.001). This suggests that there are specific micro-behaviours associated with each global behaviour type. Regression model 2: behavioural cues and caregiver responses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To test whether caregivers’ linguistic and non-linguistic responses to infant gestures were also predicted by the micro-behaviours present with the infants’ gestures, a further mixed effects logistic regression was conducted, with caregiver response (show response, characterised by shared orientation and comment vs. give response, characterised by reaching for and/or requesting the toy) as the outcome variable. Again, the inclusion of infant age as a random effect within the model did not provide a significantly better fit to the data (χ2 (1) = 0.05, p = 0.9) thus only ‘participant’ was retained to control for differences accounted for by individuals. Model 2 again included the following predictor variables; ‘vocalisation’, ‘gaze alternation’, ‘eye-gaze location’ (eye-gaze to hearer) ‘arm positioning’ (raised arm), ‘hand orientation’ (inverted hand-shape) and ‘body posture’ (lean forward) as fixed effects and ‘participants’ as random intercept effects. Again, fit of the model was obtained by likelihood ratio tests, comparing model 2 (with the predictors of interest) to a baseline model (with no effects of interest). Table 5 displays the coefficients and standard error results for vocalisation, gaze alternation and all the other behavioural cues entered into model 2. Like the previous model, vocalisation, gaze alternation and infants’ posture changes were not significant predictors of the type of caregiver response to infant gestures. However, infants’ eye- gaze was also not a significant predictor of caregiver responses. Infants’ hand- orientation also did not reach statistical significance indicating that infants’ body, hand-shape, vocal and visual behaviours were not predictors of caregiver’s responses. Arm positioning was again found to be a significant and indeed the only predictor of caregiver responses, with a raised arm displaying a significant and positive association with caregiver’s shared orientation and comment (show response). When comparing model 2 to the baseline model, the predictive behaviours significantly improved the fit of the model (χ2 (9) = 17.34, p = 0.008). Fig. 4 displays a summary of the variance explained by each predictor for both infant behaviour classifications (show vs. give) and caregiver responses (show response vs. give response). Positive values indicate a stronger association with behaviours classified as shows/caregivers’ shared orientation/commenting (show responses) whereas negative values imply a greater association with infant behaviours classified as gives/caregivers’ reach/request (gives responses) behaviour.","This study set out to address the gap in the literature concerning infants’ early showing and giving gestures, by conducting a fine-grained analysis of these gestures and their associated micro-behaviours. The aims of the study were (1) to investigate both showing and giving gestures in infants and to characterise them in terms of vocalisation, hand- shape, arm and body position, eye gaze location and gaze alternation (2) to investigate the predictive value of these micro-behaviours on caregivers’ responses to infant gestures. We used naturalistic video-data from a larger study to identify instances of showing and giving, the associated micro-behaviours and caregiver responses. Our results suggest a distinction between the micro-behaviours associated with showing and giving gestures and also a contrast in the predictive value of these behaviours for researcher classifications of these gestures compared to caregiver responses. Below we discuss some of the key findings and areas for future study. The type of gesture was not associated with age ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The current findings support the view that from 9 to 10 months, infants begin to use gestures as a way to communicate and interact with a social partner (Bates, 1979; Cameron- Faulkner et al., 2015; Cochet & Vauclair, 2010;). Overall, 83% of the infants demonstrated evidence of self-initiated shows and gives which displayed some evidence of communicative function (i.e. a motivation to share attentions and interest). Interestingly, the age of onset for these behaviours did not significantly predict the overall number of shows and gives observed or the type of gesture produced. A possible explanation for this is that before 12 months, infants are still learning how their gestures can achieve specific communicative goals and so do not yet use them consistently within social play (Carpenter et al., 1998). If this was the case, it is possible that the type of gesture produced at this stage may be contingent on the caregiver’s response rather than a previously formulated intention. Indeed, our findings revealed that individual infants often displayed a tendency to use one type of gesture over the other during the social play session. The identification of specific micro-behaviours for shows versus gives provides researchers with the opportunity to analyse the earliest instances of these gesture and examine how each is shaped by caregiver responses. Vocalisation and gaze alternation were not significant predictors of infant shows or gives, however eye-gaze location was ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Vocalisation only accompanied 17% of infants’ gestures and was not a significant predictor of gesture type. A possible explanation for this is that the infants in our study were still quite young to be consistently vocalising as a means of communication. Gros-Louis & Wu (2012) noted an increase in pointing gestures and vocalisation only after 12 months of age, and also found that infants produced vocalisations concurrently with gestures to orient the adults’ attention to the object. In our study, the adults were generally attentive to the infant throughout the interaction, which may have decreased the need for vocalisation when the adult was not attending. Further research is required to determine whether caregiver’s attentional focus dictates the frequency and type of vocalisation with showing and giving gestures, and to explore the developmental trajectory of this behaviour into the infants’ second year. Gaze alternation was also not a significant predictor of showing or giving, although it accompanied a higher proportion of gestures (64%). A possible explanation for this is that gaze alternation highlights the presence of general intentional communication rather than acting as a predictor of the type of gesture displayed. Research consistently highlights the presence of gaze alternation in conjunction with gestures such as giving, showing, reaching and pointing (Bates et al., 1975; Cochet & Vauclair, 2010; Gros-Louis, West & King, 2014). However, previous studies have noted that gaze alternation is often not a reliable predictor of infant communicative intentions and often can be influenced by the specific contexts in which the child was recorded (Cameron-Faulkner et al., 2015; Liszkowski & Tomasello, 2011). Clearly, the relation between the use of eye gaze to direct attention, and its relation to specific communicative behaviours and developments requires greater clarification, and so the conclusions we can draw from this behaviour should be interpreted with caution. Interestingly, although gaze alternation did not significantly predict the type of gesture used by infants, eye-gaze location did. Infant eye-gaze to the caregiver was associated with a showing gesture. Previous studies have suggested that early pointing in infants emerges from a shared practice of looking at things with a social partner (Liszkowski & Tomasello, 2011). Pointing with communicative intent is generally accompanied by visual checking, particularly eye gaze to the caregiver, to establish common ground (e.g. Dimitrova, Moro, & Mohr, 2015; Haynes et al., 2004). The significant association between eye-gaze to the caregiver and the showing gesture supports the idea that this gesture may be a precursor to communicative pointing, and reflect a similar motive to share that interest and attention with others (Liszkowski, Carpenter, Henning, Striano, & Tomasello, 2004; Tomasello et al., 2007). The presence of this behaviour alongside early infant showing gestures could provide greater insight into the developmental trajectory of pre- linguistic communication. Arm positioning and hand-shape were significant predictors of infants’ shows and gives ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Differences in arm position were a significant predictor of infant gesture categorisation; a raised arm was more frequently associated with a showing gesture than a giving gesture. This is an interesting finding as although arm positioning has been used to define pointing and reaching gestures (e.g. Franco & Butterworth, 1996) the relationship between this behaviour and other deictic gestures has yet to be studied in great detail. Another interesting observation was the significance of hand-shape variability in predicting infant gesture categorisation; an ‘inverted’ hand-shape was more frequently associated with a giving gesture. Previous studies suggested that hand-shape variability was a significant indicator of infants’ motives when pointing at an object or event (e.g. Brooks & Meltzoff, 2008; Franco & Butterworth, 1996) and a recent study reported that infants as young as 12 months displayed abilities to differentiate between goal-directed gestures based on their hand-shape (Elsner et al., 2014). Replication of the current findings (i.e. that a raised arm is a significant predictor of infant showing gestures whereas an inverted hand-shape helps predict the likelihood of a giving gesture) could establish these as salient indicators of shows and gives and aid in their identification in both research and social situations. Caregiver responses were not contingent on infants’ behaviours, with the exception of arm position ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ From our initial observations, it seemed that caregivers sometimes found it difficult to understand their infant’s gestures, for example responding to the infants with a question like “Is that for me?” when the infant held out a toy. This suggests that at this age, the intent of infants’ emerging gestures is sometimes unclear to the recipient. This assumption would support the recent finding of Dimitrova et al. (2015) who also found that the communicative status of prelinguistic gestures was often overlooked or misinterpreted by caregivers. Our second aim was therefore to investigate whether caregiver responses were contingent on infants’ micro-behaviours displayed when gesturing. Caregivers responses were divided into ‘show’ responses (looks and/or comments) and ‘give’ responses (extends hand and/or asks for toy) to the infants’ gestures (Cameron-Faulkner et al., 2015). With the exception of arm position, none of the behaviours produced by the infants significantly predicted the likelihood of particular caregiver responses. Rather than concluding that these behaviours have little relevance, one possibility is that there are complex multimodal interactions between the various behavioural cues that influence their interpretation. For example, it could be that infant vocalisations have particular significance that is contingent on the presence or absence of shared eye contact with the caregiver. To examine these more complex interactions, a much larger sample of infant show and give behaviours would be needed. Although further work is needed to investigate possible interactions between cues and caregivers’ abilities to utilise this information to interpret infant gestures, one additional cue was predictive of researcher categorisations of infant behaviours. Thus, there is still the possibility that the lack of significant associations between infant micro-behaviours and caregiver responses can be explained by caregiver failure to utilise all the available behavioural cues produced by the infant (Cochet & Vauclair, 2010). A number of studies indicate that when an infant uses clear communicative cues during a joint attention episode, the caregiver provides more contingent responses overall (Vallotton, 2009). This sharing of information helps the infant to gain an understanding of their social world and adults who respond contingently to infant gestures aid in the advancement of their social and language repertoire (Colonnesi et al., 2010). The current study provides a set of standardised behaviours to help identify showing and giving gestures not only for research purposes, but to inform caregivers in the home about the presence of infant behavioural cues during social interaction. This is a first step in the investigation of caregiver responsiveness to these early gestures and the potential factors that may influence this contingency.","Taken together, the results of this study provide valuable insight into infants’ communicative patterns before 12 months and question the assumption that infant pointing is the first communicative milestone achieved (Colonnesi et al., 2010; Liszkowski, 2010). Infants’ showing and giving gestures display evidence of differing behaviours and demonstrate infants’ ability to initiate these early gestures during triadic play, engaging the adult in the object of interest. This ability to adapt their behaviours may reflect earlier instances of communicative intentions, and the emergence of more sophisticated communicative abilities (Liszkowski, 2010). The research calls for future studies to explore these gestures and associated behavioural cues using more controlled experimental methods and longitudinal studies, to further determine their significance as a window onto infants’ intentional communication and to explore the relationship between early showing and giving gestures and later language acquisition. Furthermore, the current findings highlight the need for researchers to focus on a number of features in infant gesture, including arm positioning, hand shape, and eye-gaze reference when coding infants’ early shows and gives, in order to fully ascertain the functions of early communicative gestures. Clear documentation of behavioural cues associated with early shows and gives may not only aid in parental reports of infants’ gestures, but generally provide parents with clearer indications of their child’s intentions, potentially reducing the number of misinterpretations, which could be beneficial to both infant language development and social interaction (Dunham & Dunham 1992)."],["Background: Research suggests that autistic individuals may be more likely to come into contact with police and have more negative experiences in police custody. However, limited information about the difficulties they experience during the custody process is available. Aims: This study explores the experiences of autistic individuals and officers during a walkthrough of the custody process to identify specific difficulties in these encounters and what support is needed to overcome these. Methods and procedures: A participative walkthrough method was developed to provide autistic individuals and officers an interactive opportunity to identify areas where further support in the custody process was needed. Two autistic participants and three officers took part in the study. Outcomes and results: Autistic participants reported negative experiences due to: i) the emotional impact of the physical setting and custody process ii) communication barriers leading to increased anxiety and iii) exposure to sensory demands. Officers highlighted three factors which limit their ability to support autistic individuals effectively: i) the custody context ii) barriers to communication and iii) knowledge and understanding of autism. Conclusions and implications: Adjustments are needed to the custody process and environment to support interactions between autistic individuals and officers and improve the overall wellbeing of autistic individuals. --------------------------------------------------------------------------------","There has been little in-depth research into the lived experiences of autistic individuals throughout the custody process. Previous research has not explored the difficulties which autistic individuals might experience at specific points of the custody process or what adjustments might support the interactions between autistic individuals and officers to improve overall wellbeing. This study is one of the first to examine the experiences of autistic individuals during the custody process. A novel walkthrough technique using a combination of qualitative methods was employed to gain an insight into police and autistic interactions and experiences of the custody process. It builds upon the evidence that autistic individuals have negative experiences in police custody by identifying specific parts of the custody process that create difficulties for autistic detainees. The research also further explores the factors which may affect the ability of officers to support autistic detainees effectively.","Research suggests that autistic individuals are more likely to come into contact with police than the general population (Debbaudt & Rothman, 2001). For example, Tint, Palucka, Bradley, Weiss, & Lunsky (2017) reported that 16 % of 284 autistic participants had experienced police contact over a 12–18 month period (Tint et al., 2017). Some studies have suggested that this could be due to autistic individuals being more likely to commit a criminal offence or engage in certain types of criminal behaviour (see Allely & Creaby- Attwood, 2016 for discussion). However, other studies have disputed this, reporting that autistic individuals are in fact no more likely than the general population to commit an offence (see King & Murphy, 2014 for review). Alternative explanations for the increased risk of autistic individuals coming into contact with police have been reported by other studies. For instance, it has been suggested that autistic individuals are at greater risk of being victims of a crime (Brown-Lavoie, Viecili, & Weiss, 2014). In addition, another study has also reported that autistic individuals may be at greater risk of their behaviour being misinterpreted by police, resulting in an increased likelihood that they may be arrested for a criminal offence (Dickie, Reveley, & Dorrity, 2018). Police custody plays an integral role in the criminal justice system in responding to and dealing with individuals who have been arrested in connection with offending behaviour. In England and Wales, when an individual has been arrested by a police officer on reasonable suspicion that they have committed, are committing, or about to commit a criminal offence, they may be brought to police custody (see s 24 Police and Criminal Evidence Act 1984 ('PACE')) until a decision is made about whether they are to be charged with the offence. A suspect may be detained in police custody for up to 24 hours without charge for the purposes of investigating the offence (see s 37(1) and s 41 PACE). Many decisions made in the custody context can influence the personal and legal outcomes of individuals. For instance, the detainee can decide to exercise the right to legal representation or the right to silence, or they may choose to admit guilt and/or accept a police caution. The custody officer oversees many of these decisions as they are responsible for ensuring detainees are treated according to the main legislative safeguards (see s 39(1)(a) PACE). In particular, the custody officer is responsible for informing detainees of their rights and ensuring the welfare of individuals whilst in custody. To some extent, all individuals will experience difficulties in police custody as suspects will be subject to a series of intrusive processes including booking-in, fingerprinting, DNA swabbing, drugs testing and police interview (see Skinns, 2011 for discussion). They may also be detained in a cell for a significant period of time while the investigation takes place and they may experience a loss of privacy, isolation and a loss of control (see Skinns, 2011 for discussion). This can sometimes result in adverse outcomes (see Skinns, 2011 for discussion). However, the addition of sensory and communication differences in autism may exacerbate the impact of these factors further, making them more vulnerable to adverse outcomes (Chown, 2010; Crane, Maras, Hawken, Mulcahy, & Memon, 2016; Woodbury-Smith & Dien, 2014). Specifically, if important legal information is not conveyed in an accessible way, this could lead to autistic detainees making ill-informed decisions in custody. In addition, barriers to communication between neurotypical and autistic individuals may prevent adjustments to the custody process or environment from being made which could further impact negatively on their welfare. Despite the increased likelihood of coming into contact with the criminal justice system, police officers and other criminal justice professionals report feeling ill-equipped to adequately support autistic individuals (Crane et al., 2016; Dickie et al., 2018). But few studies have explored autistic individuals’ experiences in the system. Working in partnership with autistic individuals to gain their views on how to enhance service provision (Robertson, 2010), and to ensure research is both meaningful and valuable to the autistic community (Fletcher-Watson et al., 2019) can lead to better translation of the research findings to real-world settings, improving outcomes for autistic individuals. Of the few studies that have explored the criminal justice system from an autistic perspective, all report that autistic individuals have negative experiences in police custody and are generally dissatisfied with their interactions with police (Allen et al., 2008; Crane et al., 2016; Helverschou, Steindal, Nøttestad, & Howlin, 2018). As part of a larger study into the prevalence of offending behaviour among autistic individuals in South Wales, Allen et al. (2008) conducted interviews with six autistic individuals who had experience of the criminal justice system. They documented several difficulties including “not being able to take everything in; feeling in the spotlight; not knowing what was going to happen next, and feeling uncomfortable with the other people at the police station ...” (Allen et al., 2008: 754). Following an online survey of both police officers (n = 394), autistic individuals (n = 31) and their families (n = 49) in England and Wales, Crane et al. (2016) also reported that autistic individuals experienced emotional stress and breakdowns in communication in custody. In part, this was attributable to the inappropriateness of the physical setting of the interview room and custody suite and a lack of appropriate support. This study also found that some autistic individuals may not disclose their diagnosis due to a fear of being victimized or discriminated against. Helverschou et al. (2018) interviewed nine autistic individuals who had been convicted of a criminal offence. Although the majority of the research focused on reasons underlying the offending behaviour and experience in prison in Norway, some questions touched on their experiences during the custody process. While the processes in Norway may be different to England and Wales, it is interesting that the individuals reported similar negative experiences to those reported by participants in the UK. Specifically, they reported being confused about the reasons for their arrest or concerns about the lack of a legal representative that understood them. While these studies provide evidence for the need of greater support for autistic individuals within the criminal justice system, there are some limitations to the methodological approaches used. Interviews and surveys conducted after a period of time has passed may fail to capture important details about specific parts of police procedures that are significant barriers to participation in the criminal justice system. The sample of participants who complete online surveys may also not fully reflect the broader range of experiences as they may be disproportionately completed by those who are most dissatisfied by their experiences. Furthermore, these studies explored autistic experiences of the criminal justice system more broadly. They did not ask participants to reflect on specific aspects of the custody process or identify what support may be required to allow them to participate effectively. To address this, the current research will investigate the custody process to gain detailed information about potential changes that could make the process more accessible for autistic individuals. Crane et al. (2016) suggested it would be valuable to carry out direct observations of how police procedures are conducted with autistic individuals. This would allow the researcher to note specific areas of difficulty and when they occur, which might be missed in subsequent interviews with participants. This approach has further value in that it would allow researchers to obtain perspectives from both autistic individuals and the officers interacting with them. This allows us to address the double empathy problem (Milton, 2012; Milton, Heasman, & Sheppard, 2018; Sheppard, Pillai, Wong, Ropar, & Mitchell, 2016) which argues that breakdowns in communication are a two-way process. Specifically, just as neurotypical individuals may struggle to understand the intentions of non-autistic people, difficulties may also stem from an inability of neurotypical individuals to infer what autistic individuals are thinking (Sheppard et al., 2016). Therefore, it was the aim of this research to carry out a participative walkthrough which used a combination of direct observation and interview techniques to identify barriers that might affect an autistic individual’s participation in the custody process. Two autistic individuals and three officers took part in the study. The objectives were to: i) identify specific points in the custody process where adaptations are needed, ii) assess how autism impacts upon the custody process, and iii) identify recommendations for changes to the custody process/environment that would improve the experiences of autistic individuals. Participative walkthrough ~~~~~~~~~~~~~~~~~~~~~~~~~ Due to the practical, ethical and legal limitations associated with carrying out an observation of autistic individuals detained in police custody (see Demarée, Verwee, & Enhus, 2013), it was considered more appropriate to simulate this process using a walkthrough method. The participative walkthrough is an experimental method adopted for the purposes of this research. It builds on the basic principles of the ‘go-along’ method which is typically used by researchers to learn more about a neighbourhood or place, or experimentally, to explore new and unfamiliar situations (Carpiano, 2009; Kusenbach, 2003). This method combines two different qualitative methods: field observation and interviewing (Carpiano, 2009). During observations, people do not usually provide a verbal commentary on their activity, making it difficult to access their concurrent experiences and interpretations (Kusenbach, 2003). Similarly, interviews keep informants from engaging in their usual activities in the environments where they naturally occur meaning important aspects of that lived experience may be overlooked (Kusenbach, 2003). In contrast, by using a go-along approach researchers can observe an individual’s responses in context and simultaneously access their experiences at a specific time or location (Kusenbach, 2003). Researchers are therefore able to explore the informant’s experiences, interpretations, and practices using a combined approach of asking questions and observation (Carpiano, 2009). The participative walkthrough goes beyond the traditional ‘go-along’ to recreate an interactive process. This study was designed to provide officers and autistic individuals the opportunity to actively engage in the custody process and with each other, rather than simply being passive observers to what would take place in custody. Another layer was added by the researchers who observed the walkthroughs assessing the participants’ behavioural and emotional responses. These observations also helped guide the subsequent interviews and provide additional context to themes identified in the data, allowing the researchers to identify points to discuss in the interview and verify the accuracy of what was reported. The researcher also served a supportive role for the walkthrough and was there to guide the process and intervene, if needed.","Ethical approval for the study was granted by the ethics committee at the School of Law, University of Nottingham. Autistic participants were recruited on the basis they: i) were resident in the UK, ii) were over the age of 18, iii) had a diagnosis of an Autism Spectrum Condition (including Asperger’s Syndrome), iv) did not have previous experience of being detained in police custody, and v) had sufficient communication skills to take part in the research. Two autistic individuals, one female aged 24 (referred to as Anne) and one male aged 33 (referred to as Ben), took part. Both were British nationals with English as their first language. One participant lived independently and both had educational attainment through to University level. Nottinghamshire police were the focus of the study because of their proximity to the University and interest in developing their vulnerability strategy further. One custody sergeant and two detention officers were recruited via internal email sent by the custody and development officer to assist with the walkthrough. The officers were not required to have any training or experience of autism, although one reported having personal experience of interacting with autistic individuals outside custody. Informed consent was obtained from all participants prior to the study.","The walkthrough focused on three parts of the custody process: i) booking-in ii) processing and iii) cell detention. The walkthrough was carried out in an operating police station on an overflow floor which was not currently in use. Two walkthroughs were conducted. The walkthroughs lasted 51 and 53 min, respectively. These were audio recorded with the consent of participants. Two officers, the autistic participant and CH were present during each walkthrough. DR was also present during one walkthrough to provide additional support to the autistic participant. CH took notes on the quality of interactions between individuals (i.e. ease of communication), and recorded any visual emotional and behavioural cues (e.g. self-stimulatory behaviour) showed by the participants in response to what was happening, and provided support to both participants throughout the process. Prior to the walkthrough, participants read an information sheet and consent form which they signed and returned to the researcher. They were also given an overview of what would happen during the study and informed that each process was optional. Consent was verbally confirmed at each stage. Participants were allocated a scenario to contextualise the study (theft or criminal damage) which was explained to them. Booking-in Each autistic participant was asked to answer a series of background questions set out by standard police procedures. The officers answered any queries raised by the participants and explained the purpose of the booking-in process. Participants were then asked to undergo a personal search. The officers explained what would happen and showed them the items used in this process. Where the participants consented, the search was then carried out by an officer of the same gender. Processing The autistic participant was then taken to the processing room located in an operating part of the station. No other detainees were present at this time. The different processing procedures were explained accordingly: i) UV Smartwater test, ii) fingerprinting, iii) DNA test, iv) drug test and v) photographs. The custody sergeant also explained what safeguards would be in place and any legal implications. The officers simulated the Smartwater test, fingerprinting and photographs where the participant consented. Cell detention Finally, the autistic participant was taken to look around a holding cell, interview room and shower room. They were specifically asked to talk about how they felt about the cell. The officers demonstrated how the intercom in the cell operated, what it would be like if the door was closed and what would happen during a routine check. Interviews The participant interviews were semi-structured and conducted following an interview schedule to allow for consistency between individuals. All interviews were audio recorded and notes were taken throughout. DR conducted interviews with each autistic participant. One interview lasted 47 min and the other was 64 min long. Questions were divided into sections: i) overall study ii) booking-in iii) processing iv) cell detention and v) final comments. All questions were open-ended asking the participants to describe their experiences of each process and identify any difficulties with or ways of improving these processes (see Appendix A). CH also carried out a joint interview with the custody sergeant and one detention officer. The other detention officer was unavailable for interview after the study. The interview was 25 min long. General open-ended questions were asked about the officers’ experiences of taking part in the study and interacting with the autistic participants (see Appendix B). Upon completion, all participants were de-briefed and invited to discuss any concerns. Researchers maintained contact with the participants for a limited time after the study to check if they had further questions and ensure they were not experiencing any negative consequences from their participation. Data analysis The walkthrough and interviews were transcribed in full and analysed by CH using a combination of inductive and deductive approaches to thematic analysis (Braun & Clarke, 2006). This allowed the researchers to identify themes relating specifically to the officers or autistic participants and better compare their perspectives. Initially an inductive approach was taken due to the novelty of the walkthrough method and the exploratory nature of the research questions, which allowed us to code directly from the data to reflect the actual experiences of the participants. Each transcript was read line by line and colour coded according to identify specific themes. These initial codes were then reviewed, adapted and categorised into overall themes to better describe what was happening in the data. The second stage of analysis involved a more deductive approach as the transcripts and field notes were considered alongside each other to compare the experiences of the participants and allow the researcher to draw inferences from the data that were not explicitly articulated. Specifically, we drew upon the theoretical construct of “double empathy” (Milton, 2012; Milton et al., 2018) to help us better frame relationships between the data. To verify the accuracy of the final analysis, the data was subsequently reviewed and interpreted independently by NM and DR. A few discrepancies were resolved prior to the identification of final themes. Autistic perspective ~~~~~~~~~~~~~~~~~~~~ Three core themes were identified: i) impact of the custody environment ii) the contribution of communication difficulties to anxiety and iii) impact of sensory sensitivities on autistic individuals’ experiences in custody (see Table 1 for summary). Impact of the custody environment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Both participants were concerned about what it would be like to be detained longer in police custody. This was linked with the physical surroundings and the emotional impact of experiencing the custody process. Unfamiliarity The unfamiliarity of the custody environment was a concern: “(…) it’s mainly the novelty of the environment - the unfamiliarity of it” (Ben interview). This was partially attributable to the fact it was the participants first time in police custody. Small spaces The lack of space was also identified as a potential difficulty: “very claustrophobic” (Anne interview). Ben shared concerns about the small space particularly over a longer period of time with nothing to do: “I think if I was in there for a whole day… I think I’d be very claustrophobic… (…) I wouldn’t have been doing much constructive activity. (…) I’m worried about perhaps the space to move (…) perhaps, the exercising (..” (Ben interview). These concerns extended to the interview room: “a bit… boxy, a bit dingy in places, a bit (…) claustrophobic” (Ben walkthrough). Conversely, the processing room was viewed more positively: “(…) was quite spacious and I’m thinking that others might be tighter, so that helped in there (..” (Ben interview). Sleep The custody environment was also connected to the potential inability to sleep. This was due to the lighting in the cell which is always turned-on: “This doesn’t get any darker at night. This is what it’d be like at 2 a.m.” (Ben walkthrough). It was suggested this would cause difficulties sleeping which, combined with having nothing to do, would heighten the impact of detention: “if I’d been in here for… what, let’s say, a few hours… I think I’d probably start getting bored (…)” (Ben walkthrough). Privacy Both participants also discussed potential difficulties created by the lack of privacy in custody. The open toilet was a key issue: “(…) the toilets I think, I find most intimidating!” (Ben walkthrough). Both participants stressed they would not have used the toilet unless they were “desperate” (Ben interview). This was because they thought they would be visible on the CCTV camera. But this is not always the case: “(…) it’s not constantly being watched - but it’s constantly being filmed in case something happens and the toilet areas greyed out, so you can’t see people on the toilet.” (Anne walkthrough). Feeling overwhelmed Both autistic participants highlighted the emotional impact detention would have had on them if they had actually been taken into custody: “I think what realistically [would] happen is… probably start stimmi..”, “(…) hand flapping and things” (Anne interview). Anne reported how this would have been “[a] way of copi..” with detention, “trying to (…) calm the autism down..” (Anne interview). Concern about long-term consequences The long-term emotional impact was also discussed: “I’d probably felt self- destroyed, perhaps cause, you know, what could be the consequences of an arrest? (…) particularly if I was to lose my job as a result. And… if I was to go from being… doing well to, perhaps years of being unemployed (…) and I think I’d be very anxious - would I be worse off afterwards?” (Ben interview). Understanding of custody staff It was stressed that the manner and response of custody staff can help to minimise the emotional impact of custody: “(…) I felt a lot easier about it than I initially thought… and, the officers who we, dealt with today, were… very… understanding, (…) overall, I didn’t feel overly intimidated, but (…) specific points (…) could have been a little bit… intimidating.” (Ben interview). Importantly, it was suggested knowledge of autism was crucial: “(…) although they… seemed to… uh, try to be understanding they actually… appeared to have not that much knowledge. And so… probably being under more stress… if you were (…) in custody for real. (…) it (…) might be even harder to… feel comfortable to tell them, like, what was wrong. And especially some of the questions, which were key” (Anne interview). The contribution of communication difficulties to anxiety ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Both participants reported feeling less anxious when the officers explained what was happening as it helped them understand the custody process. In contrast, they reported feeling anxious when this was not fully explained, or the language was not clear. The need for more information about police procedures Not knowing what was going to happen was a key source of anxiety. This was illustrated during the walkthrough when the Smartwater test was conducted by switching off the normal lighting and turning on the ultraviolet light. It was observed by the researcher that the “sudden change startled [Anne]” (Field notes). This was because the detention officer did not explain what they were doing: “(…) I didn’t know what - was gonna happen. (…) all they said was ‘Stick your hands out’ (…) I didn’t know if they were gonna suddenly search me or what.” (Anne interview). It was mentioned that in some situations this might cause some autistic individuals to “lash out at someone” (Anne interview). Anne emphasised that it helped when the officers provided a context for what was happening by explaining what they were going to do and why they had to do each thing. Ben concurred that this helped reduce his anxiety: “(…) it just removed some of the fear about it. You know, when you know what’s actually being done… you don’t have the anxiety about something that’s… impending that isn’t.” (Ben interview). Anne also felt it might help to be kept informed about what is happening later on in the process and receive regularly updates of any changes if they were detained for longer periods of time. Ben agreed that additional information about what might happen would be useful: “I think the foresight of what would happen next and possibly when it’s likely to happen… (…), step-by-step what’s likely to happen and (…) if you’re charged, what happens if you’re not charged and um… you know, … perhaps explaining some of the (professionals) roles – you know have some literature, because reading material’s always good.” (Ben interview). It was also suggested that more accessible forms of information should be provided. For example, “visual prompts of what was gonna happen” or “a visual timetable” (Anne interview) and a booklet outlining the custody process. Ambiguous questions The autistic participants often found it difficult to understand some of the questions asked by the police officers. Open and ambiguous questions played an integral part of this. For example, it was noted by the observer that when Anne was asked “Where do you live?” in the booking-in stage, she only gave the name of the city where she lived. When further prompted by the officer with “Where in (city name)?”, Anne only gave the specific area rather than the address. Anne commented in interview “(…) they asked me (…) where I lived. (…) they could’ve easily just said ‘What’s your address? ’” (Anne interview). There was further evidence that the phrasing of other questions during the booking-in led to a failure to disclose important medical information or clinical diagnoses. For instance, when Ben was asked “do you have any injury or illness currently?” he did not disclose a medical condition due to the wording of the question. This was only picked up several questions later when he was asked about whether he was taking any current medication (Field notes). Both participants also reported that they found it difficult to know when it was appropriate to disclose they were autistic: “I wouldn’t have mentioned autism… (…) Had they not already known” (Anne interview). This was because there was no direct question on autism, only questions asking “do you suffer from any mental health problems or depression?” and “do you have any learning disabilities that we need to provide any help or support for?”. The phrasing of questions was also mentioned in relation to asking for food and drink. It was explained that standard protocol is to regularly check on the detainee in the cell to check they are okay. Anne stressed that: (…) ‘Are you okay?’ could be answered in general … I just say yes.” (Anne interview). She further commented in interview: “I think, being asked specifically if you want a drink or food [would help] … ‘cause it’s something I find difficult to ask for (…) in a strange situation” (Anne interview). It was suggested communication between the autistic individuals and officers could be better supported by making “open questions effectively into closed questions” (Anne interview). Prompt sheets would also have helped: “could have had like, prompt sheets (…) so when they say ‘How does autism affect you?’, be able to say (…) things that are quite common”; “like lighting, sound. “(Anne interview). This would make it easier for autistic individuals to communicate what specific support they might need and minimise the risk that they may not disclose this information. Legal jargon Technical language (i.e. solicitor) caused additional confusion. This was a specific concern for Anne: “added some anxiety because… (…) I had to re-ask… what it was”; “they didn’t seem to understand (…) why I didn’t understand” (Anne interview). It was observed that this encouraged her to waive her legal rights during the study (Field notes). Communicating in the cell There were also concerns about using the intercom to contact the officers: “I wouldn’t feel comfortable using it. ‘cause it’s like a phone… I don’t like phones...”;“‘cause you can’t see who you’re talking to. (…) it’s like the anxiety of waiting for someone to pick up”; “there’s no body language, (…) you have to rely on… words..”, “you can’t sign or anything” (Anne interview). The impact of sensory sensitivities on autistic individuals’ experiences in custody ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Several aspects of the custody environment were identified as creating sensory difficulties. Noise As the study was conducted on a quieter floor, there was minimal background noise. However, Ben commented on how he would have felt if the walkthrough had been conducted in busier surroundings: “(…) there wasn’t a lot of noise there, but if there was… if there was lots clattering, lots of loud noises, it could have been quite difficult.” (Ben interview). Anne also reported hearing “buzzing” coming from the lights and described the cell as “quite echoey” (Anne interview). She emphasised these noises would have been difficult to cope with: “you’re already anxious going in. And I would say all of those, sort of, aspects are heightened… and so it’s harder to deal with them”; “you can’t block them out as easily.” (Anne interview). To minimise exposure to noise, it was suggested that it would be better to use a quieter part of the station, if possible (Ben interview). Alternatively, it was thought that people might benefit from “ear-muffs” or ear defenders (Ben interview). Visual A potential source of visual sensory stress was the artificial lighting: ““it could throw a few (…) sensory issues after prolonged exposure” (Ben interview). Being able to dim the lighting would have minimised the impact of this. It was also suggested that the lighting should be replaced with “softer lights” or “different colours of lighting” (Ben interview). Colour was also a concern: “a lot of people with autism can be colour sensitive as well - so certain colours.” (Ben interview). Duller colours were preferred: “(…) you wouldn’t want a cell to be red for instance… because that could, you know, make people… quite angry? (…)” (Ben interview) In addition, it was felt that “softer floor colours” and “more matt” paint would be less demanding (Ben interview). Tactile Lastly, Anne and Ben referred to the tactile demands of the custody process. During the study, Ben took part in the search, fingerprinting and photographs. While the search did not create problems at the time, in different circumstances this could have raised sensory issues: “(…) if it wasn’t in a such a consensual nature it would have felt a little bit intrusive.” (Ben interview). In contrast, Anne opted for a description of this process because she was concerned about the sensory demands involved. Tactile demands were perhaps most evident during the processing stage for both participants who raised concerns about the invasiveness of the procedure. For example, they referred to the forcefulness of the fingerprinting: “(…) in a busy situation I’d imagine it might have been slightly more forceful, which might have been more irritating. (…)” (Ben interview). This was amplified by the fact that “they said that they could do it against your will (…) Feel sort of forced to do it.” (Anne interview). The DNA swab would raise similar issues: “Very uncomfortable. Having stuff in my mouth that I don’t want.” (Anne interview). It was noted that a further concern was having to wear different fabrics for the photographs, as this may be a “potential sensory trigger” (Field notes). Officer perspective Three main themes were identified in relation to how officers perceived and experienced the walkthrough process with autistic detainees: i) the custody context ii) barriers to communication and iii) training and experience (see Table 2). These factors also affected the identification of autistic individuals in police custody. Restrictions of the custody context ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The custody sergeant and detention officer described the demands of the custody environment and their role in custody. Time pressure The officers identified competing pressures in their day to day roles which could prevent them from spending more time with detainees than they would wish: “(…) some - like me - need to spend more time with some people as opposed to others. (…) ‘cause of custody and the beast that it is (…) the fact that we may want to spend some time with this particular person, but because, this other person’s come in and caused – (…) a bit of a rucus, it - it causes no end of problems. (.. ″ (Officer interview). The need to process detainees quickly especially during busier periods was also a factor: “it’s a time thing isn’t it as well? (…) It’s bang, bang, bang, bang, the next person. (…)” (Officer interview). Lack of alternatives to custody Another factor which influenced their ability to respond to autistic detainees was the limited alternatives to detention: “Doing this job of criminalising people when we don’t need to. Arresting them and being in custody when they don’t need to be here. And especially if it causes - like, I could see with [Ben] it was causing him that much hassle and grief (…)” (Officer interview). While the custody sergeant referred to the possibility of diversion, he described how this occurred at the end of the process. He also emphasised how it was important for arresting officers to consider the necessity of bringing individuals to custody in the first place. The impact of custody design on detention processes Finally, the custody environment itself was considered to play a role in the way officers were able to respond to detainees. In particular, the lack of privacy was highlighted as a factor that might discourage detainees from sharing important medical information: “(…) it’s difficult. ‘cause sometimes a lot of people hide things. Er, and a lot of people do not want to tell you (…) it’s not very private is it? (…)” (Officer interview). Barriers to communication ~~~~~~~~~~~~~~~~~~~~~~~~~ Communication between the custody staff and the autistic participants was difficult for both parties, not only for Anne and Ben: “(…) it was harder with [Anne] than it was with [Ben]. (…) much harder to get the answers and responses from her (…) I’ve booked a couple of people in who’ve, um, said they’re on the autistic spectrum – (…) and normally get more interaction with them (…) So it’s quite difficult… (…) I’d ask them… um… what affects them, and… how do you feel, and how can I make your stay better – (…) and they normally tell us (…)” (Officer interview) One key barrier to communication was the question structure, particularly with regard to disclosure. This is because there is no flexibility to adjust what is asked according to the needs of the detainee: “(…) all our questions, everything that I went over, are pre-set.” (Officer interview). In particular, the custody sergeant suggested it would help them if there was a set question about autism in the system: “it’s probably be better if it was a set question, because it does depend how you ask the question. (…) ‘cause [Ben] presents very well. (…) unless he told you, you wouldn’t think (…) - there was any issue at all? (…) ‘Do you require any help with reading or writing? Or have any learning disabilities the police can provide help or support with?’, I don’t know whether that’s enough (..” (Officer interview) Knowledge and understanding of autism ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ It was suggested that the response of officers was also influenced by their knowledge and understanding of autism. This was linked to two factors - prior experience with autistic individuals and training. Prior experience Having personal experience with autistic individuals was thought to influence officers’ responses to autistic detainees: “(…) I only ask certain questions [beyond the pre-set questions] because of my life experiences (…) if I’d not had that experience, I wouldn’t ask certain questions. (…) Or I wouldn’t be prepared to treat them any different (…)” (Officer interview). It was also observed that prior experience with autistic detainees in their role may improve their responses. During the walkthrough, it was noted that the officers may have adapted their interactions with Ben (the second participant) according to what they learned with Anne (Field notes). In contrast, their overfamiliarity with the role may prevent them from understanding an individual’s needs: “I think sometimes we take it for granted (…) the experiences of people coming to custody - especially when they’re brand new. (…) we don’t quite understand their anxieties. (…) ‘cause (…) we see people coming in all the time. (…)” (Officer interview). Training It was also implied that the training they receive may play a role. However, they referred to a lack of specific autism training: “We do get a little bit but (…) in email form, really. (…) we do have training - training is good [but] … it’s more mental health based (…)” (Officer interview). The custody officer suggested providing specific information to help them support autistic detainees better, such as “quick [reference] cards” (Officer interview).","This study was conducted to learn more about the experiences of autistic individuals and the officers interacting with them during the custody process. It aimed to identify potential barriers that might prevent autistic individuals from participating in the custody process and limit the ability of custody staff to support them. The data from the autistic participants showed three overarching themes: i) impact of the custody environment ii) the contribution of communication difficulties to anxiety and iii) impact of sensory sensitivities on autistic individuals’ experiences in custody. These themes partly overlapped with those from the officer data: i) restrictions of the custody context ii) barriers to communication and iii) knowledge and understanding of autism. However, the reasons underlying some of the common themes were different. This reflects the differing perspectives of the autistic individuals and the officers who participated in the walkthrough, illustrating how a “double empathy” problem (Milton, 2012; Milton et al., 2018; Sheppard et al., 2016) may present itself in a real-world context. Custody environment/context ~~~~~~~~~~~~~~~~~~~~~~~~~~~ A key issue identified from the data was how the custody environment impacted upon the well-being of the autistic participants. The physical layout of the custody suite increased anxiety due to its unfamiliarity, small/closed-off spaces, bright lights and lack of privacy. Interestingly, lack of privacy was also mentioned in the officer interview but in relation to disclosure of personal information, rather than the impact on detainee wellbeing. The emotional impact of the custody environment, although lessened by the conditions of the walkthrough, was clearly explained as a concern by the autistic participants. Feeling overwhelmed and worried about the consequences of arrest (e.g. losing job) were all highlighted as emotions they would feel quite intensely if they were taken into custody as a suspect. These findings are similar to those made by Skinns’ research which discussed the general experiences of non-autistic suspects in police custody (Skinns, 2011). However, the additional sensory sensitivities and communication barriers autistic individuals face are likely to amplify the intensity of these issues, making the custody environment particularly challenging. The officers highlighted other issues relating to the custody context which they felt restricted their ability to better support autistic individuals in custody. Specifically, time constraints and the lack of available alternatives to detention. These findings are consistent with Crane et al.’s (2016) study which surveyed a large number of police officers and highlighted time pressure and frustration in relation to other legislative restrictions. This emphasises the organisational and procedural constraints operating within police environments and suggests there may be a need for higher level organisational or policy changes. Communication / anxiety ~~~~~~~~~~~~~~~~~~~~~~~ Communication was perhaps the strongest theme. For the officers, the primary concern was their ability to obtain sufficient information from autistic individuals to help support them. However, the data suggests that the reliance on gathering this information solely through verbal exchange can be a substantial barrier to communication. The key issue for autistic participants was the ambiguity or vagueness of the questions which had implications for disclosure of important medical information or their autism diagnosis. Both the officers and the autistic individuals felt the pre-set questions did not adequately allow individuals to disclose their autism. This was due to the autistic individuals not considering their diagnosis to be an illness, injury, medical condition, learning disability, or mental health condition. While some autistic individuals may have associated learning difficulties (Emerson & Baines, 2010) or mental health conditions (Lai et al., 2019), autism is not classified as either. Furthermore, there has been a movement away from a medical model of autism towards the social model, specifically with regard to understanding autism as another layer of diversity within society (Robertson, 2010). The concept of neurodiversity has gained popularity among some autistic individuals who feel that autism is a part of their identity (see den Houting, 2019 for discussion). This indicates a risk that some autistic individuals may not be identified in police custody. Therefore, the issues in this research with disclosure suggest some changes may be needed at a procedural or policy level to incorporate a question about autism, or at least about needs arising from neurodiversity (i.e. alternative communication support). An important finding in relation to the theme of communication for the autistic participants was how strongly it was linked with anxiety. Specifically, being informed about what was going to happen was key to helping reduce feelings about uncertainty. It was suggested the addition of visual support (i.e. a toolkit) to convey information in a more accessible way could be valuable. Sensory sensitivities ~~~~~~~~~~~~~~~~~~~~~ Sensory sensitivities were highlighted as having the potential to negatively affect experiences of autistic individuals during the custody process. Tactile sensitivities were emphasised most, particularly in relation to the search and processing procedures. The visual glare from the lighting and glossy colours used in the station also added to their difficulties. Interestingly, smell was not mentioned during the walkthrough. Similarly, auditory sensitivities were not as strongly highlighted. This may have been due to using the training suite rather than the actual custody suite which was always in use. Thus, the impact of these factors may be areas for further investigation. Knowledge and understanding of autism ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The officers highlighted how previous experience with autistic or other vulnerable detainees prompted them to ask additional questions to better support them during the walkthrough. However, although officers have a good awareness and sensitivity to vulnerable detainees, they may be lacking in knowledge of autism specific support. This was clearly illustrated by the quote in section 3.1.2 (above) which conveyed that the officers showed compassion but lacked specific knowledge about how autism affected them. As mentioned by the officers, this may be due to the limited autism training they receive. This is consistent with previous research which raised concerns about the adequacy of autism training for police (Crane et al., 2016). Recommendations for support ~~~~~~~~~~~~~~~~~~~~~~~~~~~ In light of these findings, several recommendations can be made to improve the support of autistic individuals in police custody (see Table 3). These are divided into those that could be implemented by police officers on the ground immediately and those that require a change in policy or practice at a higher level. Notably, many of these recommendations would also benefit suspects more generally or those with other developmental conditions (i.e. sensory processing disorder). Therefore, it would be extremely valuable to consider how some of these changes could be incorporated as part of standard custody procedure, and not only for detainees who disclose a diagnosis of autism. This is particularly important as some individuals with autism may be undiagnosed or fail to disclose they have autism during the booking in process as evidenced in our current findings and in previous work (Crane et al., 2016). While this study provides valuable insight into the experiences of autistic individuals during the custody process, it also has limitations. The findings are limited to two autistic participants who both had attended University and only required moderate support in their daily lives. Given the heterogeneity of autism, it will be important to replicate this work with other autistic individuals representing a greater diversity of the population. Furthermore, although many of the custody processes will be the same in other police stations (at least in England and Wales), it is important to note that the facilities or environment may differ across police stations. The study is also an example of a ‘best case scenario’ and various forms of support were in place to help safeguard the welfare of the participants. As the findings do not reflect the heightened stress experienced by detainees in real life, they may not reflect the full range of difficulties some individuals may experience in police custody or the impact this may have on them. Moreover, this study did not include neurotypical individuals. This prevents further comparison of the experiences of autistic individuals with those from the general population. Nonetheless, the autistic participants did report a considerable number of issues that would likely only be amplified under real-world conditions. To gain further insight into these experiences, further research across different settings with a range of participants should be conducted.","The effects of communication barriers and sensory stress illustrated by this study indicate that autistic individuals are at greater risk of being adversely affected when detained in police custody. When combined with the typical difficulties suspects already experience, these additional factors are likely to affect how autistic individuals participate in the custody process. Adaptations to custody procedures and greater provision of accessible information are crucial as a lack of understanding can exacerbate anxiety levels. Taking steps to minimise sensory demands by adapting or better designing custody suites is also important to prevent sensory overload and reduce overall stress. In summary, these findings emphasise the need to implement a range of additional support in police custody and to make changes to the custody environment, police procedures and current policies.","This work was supported by the Economic and Social Research Council [grant number ES/J500100/1], by a PhD studentship to CH."],["The overall aim of the current study was to identify typical trajectory classes of externalising behaviour, and to identify predictors present already in infancy that discriminate the trajectory classes. 921 children from a community sample were followed over 13. years from the age of 18. months. In a simultaneously estimated model, latent class analyses and multinomial logit regression analyses suggested a five-class solution for developmental patterns of externalising problem behaviours: High stable (18% of the children), High childhood limited (5%), Medium childhood limited (31%), Adolescent onset (30%), and Low stable (16%). Six risk factors measured at 18. months significantly discriminated among the classes. Family stress and maternal age discriminated the High stable class from all the other classes. The results suggest that focusing on enduring problems in the relationship with the partner and partners' health may be important in preventive and early intervention efforts. © 2013. --------------------------------------------------------------------------------","We used data from the Tracking Opportunities and Problems Project (TOPP), a population- based prospective longitudinal study focusing on development of well-being, good mental health, and mental disorders in children, adolescents, and their families. More than 95% of Norwegian families with children attend public health services in infancy, which include 8–12 health screenings during the first 4 years of the child's life. Every family who visited a child health clinic within six select municipalities in eastern Norway (comprising 19 different health care regions) in 1993 for the scheduled 18 month vaccination visit, were invited to complete a questionnaire. Of the 1081 eligible families, the parents of 939 children (87%) participated at Time 1 (t1). These parents received a similar questionnaire when the children were 2.5 years of age (Time 2: n = 804, 86% of t1), 4.5 years (Time 3: n = 760, 81%), 8.5 years (Time 4: n = 535, 57%), 12.5 years (Time 5: n = 610, 65%) and 14.5 years (Time 6: n = 481, 51%). The questionnaires were administered by health-care workers at t1 to t3. In subsequent waves questionnaires were sent by mail. The parents chose whether the mother or father completed the questionnaire at t1–t4, at t5 the mothers were encouraged to answer, and at t6 separate maternal and paternal questionnaires were sent. The number of questionnaires completed by mothers at each wave included 921 (t1), 784 (t2), 737 (t3), 512 (t4), 594 (t5) and 481 (t6). Since so few fathers participated across time, the paternal questionnaires were not included in the current study. The 19 health care regions were chosen on the basis of their overall representativeness of the diversity of social environments in Norway: 28% of the families lived in large cities, 55% in small towns or other densely populated areas, and 17% in rural areas. The gender of the children in the sample was nearly evenly divided, with 48.9% (n = 450) boys. Maternal age ranged from 19 to 46 years at t1, with a mean of 30 years (SD = 4.7). At t1, 49% of the families had only one child, 37% had two, and 15% had three to ten children. The participating families were predominantly ethnic Norwegian. In 1993 only 2.3% of the Norwegian population came from non-Western cultures, therefore, this sample was largely representative of ethnicity in Norway at the time of data collection (Statistics Norway, 2013). Data from the child health clinics showed that nonparticipants at t1 did not differ significantly from the study participants with respect to maternal age, education, employment status, number of children, or marital status. Analyses of sample attrition from t1 to t7 (i.e., child age 16.5 years) showed that the families who had dropped out were not significantly different at t1 from the families who completed questionnaires at t7 in terms of child externalising behaviour, maternal depression, maternal age, financial status, number of children, negative life events, chronic stress, or social support. However, the dropout sample was significantly different from the remaining sample at t1, in that a greater proportion of mothers with low education had left the study. This is commonly found in longitudinal studies (Gustavson, Soest, Karevold, & Røysamb, 2012). Steps taken to minimize the impact on statistical analyses of this non-random attrition are addressed in the analyses section. Externalising behaviour problems Core aspects of mother-reported child and adolescent externalising behaviours were measured at all six waves with items rated on a three point scale: 0 (no difficulties), 1 (moderate difficulties), or 2 (substantial difficulties). At ages 18 months, 2.5 years, and 4.5 years the average of three items from the Behaviour Checklist (Richman & Graham, 1971) was used to measure temper tantrums, manageability, and irritability. Internal consistency (Cronbach's alpha) was .41, .46, and .49, at t1, t2, and t3, respectively, for the three-item scale. The average inter-item correlation was .21, .23, and .25 at the three time points, comparable to the average inter-item correlation of .25 for the 24-items of the Externalising syndrome grouping of the CBCL for 1.5–5 years (Achenbach & Rescorla, 2000) in a large study with a Norwegian sample of 4-year-olds — the Trondheim Early Secure Study (L. Wichstrøm, personal communication, June 10, 2011). This suggests that the low alphas reflect that few items are used to measure a broad construct rather than fundamental deficits in scale reliability. At age 8.5 years the Conduct Problem subscale from the Strengths and Difficulties Questionnaire (SDQ; Goodman, 1994) was used to measure tempers, obedience, fighting, lying, and stealing. The reliability and construct validity of the SDQ has been established in a Norwegian sample (Van Roy, Veenstra, & Clench-Aas, 2008). Internal consistency for the five item scale was .48. The alpha for the Conduct Problem subscale is similar to the findings from other studies (Van Roy et al., 2008). At age 12.5 and 14.5 years we used the TOPP Scale on Antisocial Behaviour (TSAB) as a measure of externalising behaviours in adolescence. The reason for the change of measures was the need for a broader and more comprehensive measure of externalising, covering a wider range of behaviours than the five-item SDQ subscale. The 18-item scale was constructed for the current project given the absence of an age and culture sensitive measure of problem behaviours ranging from relatively normative to serious (illegal) through adolescence. The TSAB is presented in Table 1. The specific behaviours are included with reference to Loeber and colleagues' model of three developmental pathways in child disruptive behaviours (Loeber et al., 1993). The items measuring inter-personal aggression refers to “overt behaviours” in the Loeber et al. model, stealing and vandalism to “covert behaviours”, and loitering to “authority conflict/avoidant behaviours”. The TSAB combines items from other Scandinavian scales (Bendixen & Olweus, 1999; Mahoney & Stattin, 2000; Rossow & Bø, 2003). The alpha coefficients were .69 and .77 at t5 and t6 respectively. Due to a change of wording for three items at t6 (excluding aggressive behaviours among siblings) the measure of physical aggression at t6 may be underestimated compared to t5. Time-to-time correlations (i.e., t1 to t2, t2 to t3, etc.) were .46, .50, .32, .29, and .43 for the externalising measures, with the lowest correlations corresponding to the longest intervals between waves. These time-to-time correlations between the six externalising measures across childhood are at about equal magnitude as the alphas for the BCL and the SDQ. Again, this supports the interpretation of a measure comprising few items covering a broad construct area as opposed to fundamental deficits in measure reliability. Temperament At age 18 months child temperament was assessed by the EAS Temperament Survey for Children: Parental Ratings (Buss & Plomin, 1984), which contains four dimensions: (a) Emotionality — the tendency to become aroused easily and intensely (often called Negative Emotionality); (b) Activity — preferred levels of activity and speed of action; (c) Sociability — the tendency to prefer the presence of others to being alone; and (d) Shyness — the tendency to be inhibited and awkward in new situations. The EAS for children aged 1–9 years was used. Because of ambiguity in translation, one item was deleted from each dimension. The items were scored on a Likert scale from 1 (very typical) to 5 (very untypical). Cronbach's alphas for the four items in each dimension were .66, .68, .52, and .75, respectively. Maternal mental health At child age 18 months maternal symptoms of anxiety and depression were measured by a 23-item version of the Hopkins Symptom Check List (HSCL-25; Hesbacher, Rickels, Morris, Newman, & Rosenfeld, 1980). The reliability of the HSCL has been well established in a Norwegian sample (Tambs & Moum, 1993). Two items, “thoughts of ending your life” and “loss of sexual interest or pleasure”, were excluded from the current version of the questionnaire because some mothers who participated in a pilot study had perceived the questions as offensive. The items were scored on a 4-point Likert scale, from 1 (not at all) to 4 (very much). The alpha coefficient was .90. Family stress At child age 18 months mothers were asked to indicate whether they had experienced enduring problems during the last 12 months in the following areas: housing, employment, financial status, their partner's health (somatic and mental), and their relationship with their partner, each scored 0 (no problem) or 1 (problem). The sum of the scores in the five stress areas formed the composite score of family stress, with a range of 0 to 5. The alpha coefficient was .56. Social support from partner At child age 18 months a social support from partner index was formed by taking the mean of three items, each on a Likert-scale from 1 (completely disagree) to 5 (completely agree), measuring closeness and contact, respect and responsibility, and a feeling of belonging (Dalgard, Bjork, & Tambs, 1995; Mathiesen, Tambs, & Dalgard, 1999). The alpha coefficient was .59. Social support from friends and family of origin Corresponding to the social support from partner index, this questionnaire targeted the same three qualities (closeness and contact, respect and responsibility, and a feeling of belonging) to describe the mothers' relationships to friends and members of her family of origin. This measure was also completed at child age 18 months. A social support from friends and family of origin index was computed by summing the mean value of the 6 items. The alpha coefficient was .72. Family demographics Maternal education at child age 18 months was measured using eight response categories, and was recoded to represent the approximate total years of education. Additional variables included: Maternal birth year; Mothers living without spouse or partner; Siblings, a dichotomous variable of 0 (no siblings) and 1 (one or more siblings); and Child gender, all values were reported by mothers at child age 18 months. See Table 2 for a description of sample characteristics.","Latent class analyses refer to modelling with categorical latent variables to represent subpopulations. The latent classes explain the relationships among the observed dependent variables, similar to factor analysis, but sort individuals into latent classes rather than producing continuous latent factor scores. Child externalising mean scores at each assessment (age 18 months and 2.5, 4.5, 8.5, 12.5 and 14.5 years), rescaled to have approximately equal variance at every time point to eliminate possible estimation problems due to different scales of measurement, were used in the latent class analyses. The rescaling was done by multiplying the variable by 10 (at t1, t2 and t3, respectively), 14.29 (at t4), and 26.67 (at t5 and t6). After rescaling, mean externalising was 4.2, 4.6, 4.7, 3.1, 1.9, and 1.9 at age 1.5, 2.5, 4.5, 8.5, 12.5 and 14.5 years respectively. Latent Class Growth Analysis (LCGA) is a popular latent class analyses method for longitudinal measures that produces classes that are similar in terms of development. Latent Profile Analysis (LPA) is similar to LCGA but it does not impose a parametric form to growth (e.g. linear growth) and is, therefore, more general than LCGA. LPA captures developmental change just as LCGA, but in the form of a profile of change rather than as slopes and intercepts. Because we did not want to restrict a priori the possible shape of developmental patterns to estimate, we considered this a sound choice. We allowed the residual time specific variances to be different across classes but forced them to be equal across time to minimize the number of variance parameters and potential convergence problems. Child and family factors measured at child age 18 months were used as predictors of the longitudinal latent classes, including emotionality, shyness, activity, sociability, maternal symptoms of anxiety and depression, family stress, social support from partner, social support from family of origin and friends, maternal education, maternal age, mothers living without spouse/partner, siblings, and child gender. Multinomial logit regression was used to test for group discrimination by the 18 month predictors one at a time. The predictors were then combined in one multi-predictor model to compare their relative strength in discriminating the latent classes; those that were not significant in the multi- predictor model were removed. Models were estimated by using the full information maximum likelihood estimator in Mplus, which allows for the inclusion of participants with partial data in the trajectory variables (externalising), but not participants with missing predictor data, under the assumption that missingness is at random, conditional on variables included in the model (MAR). Thus, the sample size varies somewhat across models depending on which t1 variables are included. The amount of missing data at t1 was minimal, however, with less than 2% for any particular predictor, and less than 3% for the multi- predictor models. It is not possible to test the MAR assumption unless the missing data can somehow be recovered, but even if the MAR assumption is not completely true, MAR based likelihood estimation performs well under most circumstances and is superior to obsolete methods based on including only subjects with complete data (Graham, 2009). Optimal number of latent classes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Externalising scores from all waves were first included in a series of latent class analyses in order to decide on the optimal number of latent classes. For models with 3 or more classes, the data drove the models toward solutions that included one or two classes with virtually no externalising in adolescence. This was due to skewness of the TSAB scores at t5 and t6, as 173 mothers reported that their adolescents had no externalising behaviours at both t5 and t6 combined. Since subgroups with means and variances of zero are known to produce estimation challenges and model nonconvergence (Hipp & Bauer, 2006), we thus fixed the model variance-parameters for these two groups at a small value near zero (0.22). (Another commonly used approach to the class variation problem, i.e., forcing the residual variance terms to be the same across all the classes, was not appropriate for these data, as the assumption that a low externalising class would have just as much variability as a high externalising class was not plausible and did not describe the observed data well.) Solutions with two to six classes were examined. We used the BIC and sample size adjusted BIC (SSA-BIC) fit statistics, and the Vuong-Lo-Mendell-Rubin (VLMR) test to decide on the optimal number of latent classes. The BIC's were 19,313, 18,983, 18,805, 18,743, and 18,767, respectively, for the two to six-class solutions, while the SSA-BIC's were 19,256, 18,895, 18,688, 18,594, and 18,590. The VLMR test indicated that three classes were the best solution but this solution did not include a low externalising class. There is an inherent uncertainty regarding the optimal class enumeration in exploratory LPA. However, BIC fit statistics are the most commonly used and accepted criterion for deciding on the optimal number of latent classes. We chose the model with five trajectories, as it had the lowest BIC fit statistic among models with two to six classes. Although the SSA-BIC did not show the same clear minimum, it is known that the BIC imposes a higher per parameter penalty then the SSA-BIC and the SSA-BIC usually indicates more classes than the BIC (Nylund, Asparouhov, & Muthen, 2007). Inspection of the profile forms, problem levels, and proportions of the various classes also supported the five-class solution. We tracked the (re)allocation of individuals as the number of classes increased, and the underlying heterogeneous data structure became apparent. We were thus able to inspect the meaningfulness of the various solutions in terms of the balance between model complexity and simplicity when depicting this heterogeneity. Both the BIC and the meaningfulness of the solution were supportive of the five-class solution. The five-class solution consists of a “High stable” class (a group with a stable high level of externalising through the study period that emerged in the solutions with three, four, and five classes, suggesting that the identification of the “High stable” class is a robust finding), a “High childhood limited” class (characterised by a very high level of externalising in the early epoch and a low level in adolescence), a “Medium childhood limited” class (characterised by high levels of externalising in early childhood), an “Adolescent onset” class (with the second lowest level of externalising early on and the second highest externalising level in adolescence), and finally a “Low stable” class. We present class proportions later when discussing results for models that included predictors. Single-predictor model results ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A second series of latent class analyses was conducted by including the 18 month predictors into the five-class model, one predictor at a time to test for the effect of separate predictor variables on the classification results. Including the early variables in the model increased the model stability without substantially changing the shape of the trajectories and the proportions of children in the classes. We used a nested chi-square test (likelihood ratio test or LRT) to compare all class contrasts (the H1 model) versus a model where all class contrasts were forced to zero (the H0 model). The nested chi- squares, resulting from a 4 df test of the null hypotheses that the predictors had no effect on class discrimination, were significant for all 18 month variables except three. Latent class membership was significantly predicted by child emotionality, LRT (4) = 169.4; p < .001, maternal depression, LRT (4) = 83.9; p < .001, family stress, LRT (4) = 54.8; p < .001, support from family and friends, LRT (4) = 23.6; p < .001, maternal education, LRT (4) = 23.4; p < .001, maternal age, LRT (4) = 23.4; p < .001, child gender, LRT (4) = 20.8; p < .001, and support from partner, LRT (4) = 17.2; p < .001. Child temperamental shyness, mothers living without partners, and having siblings were not significant predictors. Models with child temperamental activity and sociability had severe convergence problems, did not give meaningful results, and were excluded from further analyses. Multi-predictor model results ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Finally, in order to identify the most influential predictors differentiating the five trajectory classes, we included a multi-predictor, multinomial logit regression analysis in the latent class model. We entered all of the 18 month predictors, and successively eliminated variables that were not significant in the total model. Again, this significance was evaluated by comparing a 4 df nested chi-square test for a model with all class contrasts (the H1 model) versus a model where all class contrasts were forced to zero (the H0 model). Six predictors remained in the final model and had significant effects in discriminating among classes. The individual class contrasts, the four that were actually estimated (classes 1–4 vs. 5) and the other six that can be derived from those four along with the 4 df nested chi-squares for these six variables are presented in Table 3. Estimating all four class contrasts for each of six predictors, 24 estimated effects in all, with no restrictions on any of the effects, makes the solution difficult to interpret and tends to inflate the standard errors of the estimates. To make the model more interpretable, we constrained class contrasts that were about the same magnitude to be equal with possibly different signs and forced two class contrasts to zero that were very close in magnitude to zero. This constrained model had nine estimated class contrasts as opposed to 24, yet the fit of the model was not significantly worse (nested chi-square = 13.27 with 15 df, p = .58) and the standard errors for the nine effects were considerably better estimated, although the actual magnitudes of the effects were about the same as the unconstrained model. The class regression effects for the constrained model are shown in Table 4. In general, classification quality improves when more information, such as class predictors, are included in the model. Interestingly, the “High stable” class remained almost identical across models, while the other classes changed only moderately. Boxplots with pseudo-class and fitted trajectories for the five-class solution based on the multi-predictor model are presented in Fig. 1. As can be seen, there is a close correspondence between the two, indicating a reasonable solution. For three of the five groups there is a shift in externalising level between t3 and t4 and between t4 and t5, which may partly be a result of a change in measures. The five LCA classes varied both in level and pattern over time. The class with the highest level of externalising behaviours in the three first waves had a low level of externalising at the adolescent end-point. This “High childhood limited” (HCL) class was the smallest with 5% of the sample. The class with the second highest level of externalising in infancy continued to have high scores throughout the entire measurement period. This “High stable” (HS) class comprised 18% of the sample. The class with the median level of externalising in infancy declined to a near zero level at the adolescent end-point. This “Medium childhood limited” (MCL) class was the largest, including 31% of the sample. The class with the second lowest level of externalising in infancy was the second highest group at the adolescent end- point. This “Adolescent onset” (AO) group constituted 30% of the sample. Finally, 16% of the sample was assigned to a “Low stable” (LS) group. Due to change in measures and rescaled variables, note that only relative change across groups and not absolute (developmental) change can be interpreted (Fig. 1). As can be seen by the 4 df nested chi- squares at the bottom of Table 3, child negative emotionality made the strongest class discrimination, followed by maternal symptoms of depression, child gender, family stress, maternal age, and siblings. Turning now to the class regressions (Table 4), we start with the two variables, family stress and maternal age that had the hypothesized pattern of uniquely discriminating the high stable class (class 5) from the other classes (class 1–4). The fact that all four class regressions for family stress can be constrained to be equal without degrading model fit, − .29, and are significantly different from zero indicates that family stress uniquely discriminates the HS class from classes 1–4 but does not discriminate among classes 1–4. Higher levels of family stress were associated with lower log odds of membership in classes 1–4. The estimated class regressions can be exponentiated to determine the effect of a 1 unit shift in the predictor on the odds of being in one class vs. another, and for family stress this value is .75, indicating that each 1 unit increase in family stress lowers the odds of belonging to classes 1–4 by 25%. To appreciate the range of risk from family stress in the population we can consider the 10th percentile family stress score of 0 as low risk and the 90th percentile score of 3 as high risk. Going from low to high risk in the population would increase family stress by 3 units and thus increase the odds of belonging to the HS class by a multiplicative factor of 2.4 or 140%. Maternal age also had the same pattern of effects as family stress such that increasing birth year (i.e. younger age) was associated with lower log odds of being in classes 1–4 vs. the HS class. Going from low to high risk increased the birth year variable by 1.2 units (it was linearly rescaled to prevent convergence problems in the model, on the raw scale it was 12 years, age 36 compared to age 24), which corresponds to an increase in the odds of being in the HS class by a factor of 3.7. Because family stress and maternal age are the only two predictors that uniquely discriminated the HS class, it makes sense to consider the range of risk from both risk factors simultaneously. This entails comparing a family with both risk factors at the 10th percentile value to a family with both risk factors at the 90th percentile, all else being equal. Such a comparison yields an 8.8 fold increase in the odds of being in the HS class vs. classes 1–4, which is a very substantial increase in risk. The next two variables with relatively simple class discrimination patterns were female gender and children with siblings. Siblings increased the odds of being in the HS class vs. the MCL, HCL and AO classes by 2.1 but had no effect on discriminating the LS class and the HS class. Female gender increased the odds of being in the MCL, HCL, and in the LS classes vs. the HS class by 1.7 but decreased the odds of being in the AO class vs. the HS class by 0.6. Finally, maternal depression and child emotionality had a similar but more complicated pattern of effects on class discrimination. Both variables increased the odds of being in the HCL class vs. the HS class by very substantial amounts, but both variables also increased the odds of being in the HS class vs. the AO and LS class by substantial amounts. Maternal depression also increased the odds of being in the HS class vs. the MCL class, but child emotionality had no effect on this contrast. Child emotionality appears to discriminate largely based on early externalising in the first three time points. Maternal depression also fits this pattern, except for the significant discrimination of the HS and MCL classes, which have similar levels of early externalising, but quite different levels of adolescent externalising. In fact, for low vs. high risk on maternal depression, a shift of 0.8 units on the maternal depression scale, the odds of being in the HS class vs. the LS class increases by 2.0.","The overall aim of the current study was to identify typical trajectory classes of externalising behaviours, and to identify predictors already present in infancy that discriminate among the trajectory classes. A latent profile model with five classes was chosen as the best to capture the heterogeneity in the 13-year course of externalising behaviour development from infancy to mid-adolescence. The model identified a class of children following a developmental trajectory with a HS level of externalising behaviours. Identification of this HS class is a robust finding, in that the class emerged early in the analytic process and remained consistent across models. While most predictor variables discriminated between trajectory classes in the single-predictor models, six variables were the most influential (multi-predictor model). Two variables, family stress and maternal age, discriminated uniquely between the HS class and all other classes. Comparing families with both of these risk factors simultaneously at the 10th percentile value to families with both risk factors at the 90th percentile, all else being equal, yields an 8.8 increase in the odds of being in the high stable class versus the other classes. The other four significant variables in the multi-predictor model – child gender, siblings, maternal depression, and child emotionality – had less clear patterns of class discrimination. The finding that maternal age and family stress, measured in infancy, seems to have strong impact on externalising development from infancy to mid-adolescence has potential to inform preventive and early intervention efforts. There is an inherent uncertainty regarding class enumeration in exploratory LPA, and deciding on the optimal number of latent classes cannot be done based on test-based criteria alone. For our results, however, the BIC fit statistics clearly suggested that five classes were better than four or six. We fixed several model variance parameters at a small value near zero for two classes at t5 and t6 in order to avoid estimation difficulties due to no variance in externalising scores. In preliminary attempts at modelling the data without fixing the class variances, the BIC statistic never hit a clear minimum and just continued to decrease for four, five, and six class solutions. The longitudinal five-class solution that was chosen for the current study is in line with theory and earlier findings. As postulated by Moffitt's taxonomy (1993), an early-onset and an adolescent-onset type emerged. Childhood limited types were found as in other longitudinal studies (Moffitt, 2006). Among studies that have covered longer developmental periods starting from before the age of three years, the solutions vary between three, four, and five trajectory classes. Similar to our five-class solution, results from the NICHD study based on aggression ratings (Campbell et al., 2006; NICHD ECCRN, 2004) and on ratings based on a broader externalising construct (Fanti & Heinrich, 2010) suggested five trajectory classes. Further, two different classes with externalising behaviour problems limited to childhood were also found in studies starting at a very early age (Fanti & Heinrich, 2010; Côté et al., 2006; Shaw et al., 2005). For instance, Fanti & Heinrich (2010) identified a High Desister class as well as a Moderate Desister class, which seems to correspond well with our solution with two different childhood limited classes. The LS class consisted of 16% of the sample. In trajectory studies that begin from preschool or later, the LS externalising class tends to be the largest class (e.g., Campbell et al., 2010; Odgers et al., 2008). In infancy and toddlerhood, however, child misbehaviour is highly normative (Tremblay et al., 1999; Wakschlag et al., 2010). A small LS externalising class is in line with findings from other studies that have measures from infancy and onwards, like 10% in the Shaw et al., study (2005) and 5% in the Côté et al., study (2007). The HS externalising class comprised 18% of the sample, which is substantial in comparison to most previous studies. For example, 7% of the sample used by Shaw et al. (2005) was classified as having high stable externalising behaviour, and 3% fell into a similar class in the NICHD ECCRN (2004) study. The only study to identify a similar proportion is Côté et al. (2006), with 16.6%, although the Côté study identified three and not five groups as did the current study. Post-hoc analyses indicated that the HS class in the present study was characterized by a wide array of externalising behaviours. As the construct of externalising used in the TOPP cohort is broader, encompassing common oppositional and disruptive behaviour in early childhood and vandalism, status violations, and physical aggression from mid childhood onwards, the larger high stable class is an expected finding compared to the studies that have had a limited focus on overt conduct problems and physical aggression. The problem scores of the HS class are labelled “high” as they are compared to the other classes in the model, but are these scores exceptionally (deviant) high? The proportion of children in comparable populations expected to meet criteria for a disruptive disorder diagnosis is estimated at between 3% and 8% (e.g., Heiervang et al., 2007; Kessler et al., 2012; Wichstrom et al., 2012), thus for the majority of the children in the HS class (18%) the scores are probably not clinically high. When comparing the mean HS problem levels, however, to the means in comparable studies across childhood, it is clear that the HS class is substantially elevated in comparison to average children. First, we compared the BCL scores at age 4.5 to data from an epidemiological study of approximately 2000 Norwegian 4-year-olds in Vestfold county (Wefring, 2001). Means and SDs across the complete datasets were similar, and the mean problem score of the HS class corresponds to the 88th percentile value in the Vestfold data (K. K. Lie, personal communication, November 8, 2013). Further, we do not have Norwegian norms for SDQ, but we compared the age 8.5 score on SDQ Conduct Problems Scale to data from two large Norwegian population-based samples of 9-year-olds (Obel et al., 2004). Means and SDs across the three complete datasets are very similar, and the mean problem score of the HS class corresponds to a score between the 87th and 95th percentile value in the Akershus data (J. Clench-Aas, personal communication, November 8, 2013). Finally, to our knowledge, the only other project that has collected parent-reports at ages 12 and 14 with items identical to those in TSAB is the Swedish “10 to 18” study (Mahoney & Stattin, 2000). In that study, four items comparable to TSAB items no. 2, 3, 9, and 15 (see Table 1) were administered. It should be noted that the response format in that study differed from the one applied in our study, in that responses referred to lifetime involvement (“ever”) and not involvement during “the last 12 months”. These four items in the “10 to 18”-study, measured at ages 12 and 14 (H. Stattin, personal communication, November 1st, 2013) have close to identical means, SDs, and 95th percentile values as compared with the corresponding four items in the TOPP-study measured at ages 12.5 and 14.5 years. The problem score levels of the HS class correspond to a score between the 94th and 98th percentiles and between the 92nd and 98th percentiles in the “10 to 18”- data in early and mid-adolescence, respectively. More details on the comparisons are available by request from the first author. Concerning class-differentiating predictors measured at 18 months, one noteworthy result was that the LS class was uniquely different from all the other classes on low maternal depression and low child emotionality. The HS problem trajectory class was uniquely and strongly differentiated by young maternal age and higher levels of family stress from all the remaining developmental pathways identified in the study, including the HCL group, which started with an even higher externalising level in infancy. Given the need to extend knowledge about normative versus high-risk early externalising behaviour, the identification of these two family factors which can discriminate between the HS and other groups (odds ratio of 8.8 for high on both vs. low on both risk factors) is an important finding, with potential to inform prevention and early intervention efforts. Early family adversity is a well-established predictor of externalising development (Aguilar, Sroufe, Egeland, & Carlson, 2000; Shaw et al., 2001). Despite this, the current study is to our knowledge, the first trajectory study over the period from infancy to mid-adolescence to document this. The results suggest that support to young mothers and families under stress, especially those whose toddlers are exhibiting noncompliant acting-out behaviour, may interrupt a pathway of externalising behaviour which may otherwise continue into adolescence. The current study included a population based sample where the youngest mother gave birth at age 17. Young motherhood is identified as a risk for externalising development in both high-risk samples in the U.S. and Canada (Nagin & Tremblay, 2001; Shaw et al., 2005), and in nationally representative samples (Côté et al., 2007). Our results are in line with these previous findings, while we note that young motherhood was not a significant predictor in a U.S. general population sample (NICHD ECCRN, 2004). The measure of family stress that discriminated the HS class from all the other classes at 18 months covered problems experienced during the previous 12 months in areas including socioeconomic difficulties and relationship issues. Living condition stressors constituted important parts of this index, and were measured with three items covering enduring problems with housing, employment, and financial status. Our finding on living condition stressors is in line with results on low income from general population samples in the U.S., Canada, and the U.K. (Barker & Maughan, 2009; Côté et al., 2006; NICHD, 2004), underlining that living condition stressors are an important predictor of the development of chronic high externalising from early childhood onwards. The current findings suggest that other aspects of stress are also important. An exploratory post-hoc analysis suggested that the variables within the overall stress construct which were most closely related to class membership were problems in the relationships between mothers and their partners and partners' health problems. Regarding the mothers' experience of enduring problems in their relationship with their partners, studies are scarce, and we have located only one trajectory study from early childhood onwards with which to compare our results. The Avon Longitudinal Study (Barker & Maughan, 2009) measured grave parental relationship problems in the form of partner cruelty towards the mother (defined as any indication of emotional and/or physical abuse from the mother's partner), and found that partner cruelty between child ages 0 and 4 predicted chronic conduct problems from ages 4 to 13. Our findings expand upon the Avon study results by indicating that parental relationship problems do not need to include cruelty to have significance for early-onset externalising development. The last topic within the early family stress index is mothers' experience of enduring problems related to partners' somatic and/or mental health. Here, as well, we have located only one early starting trajectory study with a measure that was somewhat similar to ours. A Canadian population based study measured depressive symptoms in fathers at child age 5 months and found paternal depression to be a unique predictor of chronic disregard for rules between ages 29 to 74 months (Petitclerc et al., 2009). The Canadian finding, in combination with our results, point toward the importance of health issues in fathers (mothers' partners) for chronic high externalising development from early childhood onward. As these last issues within the family stress index were measured with only one item each, and given the overall scarcity of studies on these two topics, there is a need for more studies to shed light on the impact of enduring difficulties in mother's relationship to her partner and to partner's health, as early predictors of a chronic high externalising profile across childhood. Child negative emotionality measured in infancy had the strongest impact on the final model as a whole; this variable contributed to the discrimination of most of the classes. Child emotionality appears to discriminate largely based on early externalising in the first three time points. Since for all five trajectory classes the mean levels of negative emotionality and externalising were parallel to each other at t1, it is possible that parents did not differentiate clearly between externalising behaviour and temperamental negative emotionality at this early age. Furthermore, it may be that there is conceptual overlap between the items tapping negative emotionality and externalising at this age (both of which are related to child manageability), an issue that is common in the field (Sanson et al., 2011). However, studies suggest that confounding of measures does not fully account for the predictive role of emotionality, which has consistently emerged in the literature as a risk for externalising behaviour problems (Sanson, Hemphill, Yagmurlu, & McClowry, 2011). Parents raising a child high in negative emotionality may need support in helping the child to regulate a temperamental disposition, and in avoiding a punitive style of discipline and ‘coercive cycles’ of interaction with their child (Reid et al., 2002). Maternal anxiety and depressive symptoms when children were 18 months old had the second strongest impact on the final (multi-predictor) model as a whole, and similar to child emotionality, this predictor appears to discriminate largely based on early externalising in the first three time points. Unlike child emotionality, however, it discriminated the HS class from the MCL class (odds ratio of 2.0), two classes which have similar levels of externalising at the three earliest time points but quite different levels of externalising at adolescence. This finding points toward maternal symptoms of anxiety and depression as an early risk factor for all the pathways towards some level of externalising problems in adolescence. Once again, these findings suggest a need for early intervention for mothers showing signs of mental health problems. More of the children in the HS class had siblings compared to all the other classes except the LS class. Parenting two or more children simultaneously creates higher demands on parents and family resources. A mother who is exposed to the risk factors of family stress, young age, and the presence of siblings in the family has a 16 fold increase in the odds of having a child in the HS class compared to the MCL, HCL, and the AO classes. The substantial increase in risk due to siblings could be because multiple siblings may exacerbate each other's antisocial behaviour in a negative cycle of so-called deviancy training (Patterson, 1986). In support of this notion, sibling aggression has previously been found to have a unique contribution to externalising development in a genetic sensitive study (Natsuaki, Ge, Reiss, & Neiderhiser, 2009). Contrary to expectations, the HS profile had an even split between the genders, rather than mostly consisting of boys. The even split between the genders in a high externalising class is not in line with previous research (e.g. Côté et al., 2006). Previous findings have suggested that robust gender differences are typical for overt externalising behaviour types, as well as for a wider construct that includes overt types (Broidy et al., 2003; Moffitt, Caspi, Rutter, & Silva, 2001). To our knowledge, gender differences are not equally robust in other facets of externalising behaviours. For stealing and lying, the frequency may be equal between the genders (Tiet, Wasserman, Loeber, McReynolds, & Miller, 2001). The use of a broad measure of externalising in the current study, which includes oppositional and disruptive behaviour in early childhood and covert and authority avoidant behaviour from mid-childhood onwards, may have resulted in more girls included in the HS class. The post-hoc analyses of gender differences within the HS class suggest that the HS boys were on average more involved in overt externalising behaviours than the HS girls, while for the remaining externalising items there were no gender differences within the HS class. Furthermore, the size of the HS class identified in the current study is substantial in comparison to other studies (NICHD ECCRN, 2004; Shaw et al., 2005). Thus, the inclusion of externalising behaviours that were non- confrontational at all time points in addition to overt behaviours, may have resulted in a higher proportion of children in general, and also a higher proportion of girls, included in the HS profile. Maternal education level discriminated in the HS and LS groups in the single-predictor models, but failed to do so when all predictors were entered simultaneously. This is in contrast with findings from the U.S., Canada, and England, for example Nagin and Tremblay's (2001) study, in which low maternal education was one of two factors discriminating between “chronics” and “high decliners” in physical aggression. One possible explanation for why maternal education was not significant in the complete model of this study may be that education is not a strong marker of social class in Norway given that Norway is a relatively homogeneous society. Overall, this study produces substantially important findings, and adds to the literature in notable ways. To our knowledge, no latent class trajectory study has included such an array of simultaneously estimated early predictors of class membership, with the earliest externalising measure taken before child age of 2 years and continuing to mid-adolescence. The results suggest that preventive and early intervention efforts should have a broader focus and pay special attention to children's externalising behaviours in the context of young motherhood and higher levels of family stress, as well as child temperament, maternal distress, male gender, and presence of siblings in the family. It is possible that family support intervention programmes focusing on supporting a stable, low stress family environment would reduce the numbers of adolescents likely to engage in delinquency. What may the mechanisms be that link the identified factors in infancy with membership in a stable high trajectory pattern? A young mother may be less experienced as a caregiver and have poorer parenting skills, while at the same time she may be more vulnerable to the frustration of her own developmental needs. Additionally, when experiencing enduring strains in a relationship with a romantic partner, with health issues, with living condition related issues, and/or when having symptoms of depression and anxiety, these factors are likely to result in reduced sensitivity, contingency, and time available during day to day interactions with the child. While this study benefited from six waves of data with a community sample, it has some limitations. Attrition analyses showed that some of the less educated mothers had left the study by t7. This maternal factor is known to be associated with externalising problems in the child. Thus, the results may be underestimating the effects of low education on externalising trajectories. All analyses, however, were carried out using full information maximum likelihood estimation which includes subjects with partial data and minimizes biases due to attrition. Another limitation is the use of different measurement instruments to assess externalising behaviour in the trajectory model, meaning that the externalising construct is not identical through all developmental phases. Three out of five groups show a shift in trajectory shape with changing instruments, and the physical aggression component may be underestimated at t6. Our measures, however, are still appropriate for identifying different patterns of externalising development even if specific mean levels of externalising are not directly comparable across time. It has also been shown that children's development differs over the various domains within the broad construct of externalising (Bongers, Koot, van der Ende, & Verhulst, 2004). The current study focused on the broader construct, and not on its sub-domains. The current study's strength of using a developmentally appropriate broad measure to identify and compare trajectory groups should be weighed against the disadvantage of not considering sub-domain development. Furthermore, the early predictors measured in this study are likely to change over time in ways that are likely to continue to impact development. Future studies should address a broader view of potential predictors throughout development. Findings may have been weakened by the modest internal consistency of some predictor measures, which is likely to have attenuated some relationships. Internal consistency was also low for the outcome, but the latent class model accounts for imperfect reliability in the outcome and hence it is not likely that the modest internal consistency of the t1–t4 indicators has biassed the class solution, even though it may have contributed to weaker classification accuracy for some individuals, particularly in the first set of analyses before predictors were added to the model. In addition, because the mothers reported on themselves as well as their children, single-informant bias may have influenced the results. A reasonable case has been made, however, that maternal reports provide valid and useful information (Janson & Mathiesen, 2008; Rothbart & Bates, 2006). However, replication with multi-source data is needed. Also, although the study has tapped a range of important influences on children's development, there are other potential sources of influence such as children's peer relationships and genetic variation, on which more information is needed to fully understand the development of externalising problems. Finally, our findings may also be more easily generalised to Western European societies that are more similar to Norway with respect to social welfare systems including family and youth support systems as opposed to findings from for example the United States. This study contributes to the literature with data from a developing general population sample followed over a long time span. We have used a broad and developmentally appropriate measure of externalising, and taken advantage of a wide range of risk factors measured very early in development. The results add to the literature with unique and practically important findings that may be useful for prevention or early intervention efforts in minimizing the long term negative effects of early externalising behaviours. Although replication of these findings is necessary, they point towards the need for preventive interventions to start very early in life and to address multiple aspects of children's family life."],["During skilled music ensemble performance, a multi-layered network of interaction processes allows musicians to negotiate common interpretations of ambiguously-notated music in real-time. This study investigated the conditions that encourage visual interaction during duo performance. Duos recorded performances of a new piece before and after a period of rehearsal. Mobile eye tracking and motion capture were used in combination to map uni- and bidirectional eye gaze patterns. Musicians watched each other more during temporally-unstable passages than during regularly-timed passages. They also watched each other more after rehearsal than before. Duo musicians may seek visual interaction with each other primarily, but not exclusively, when coordination is threatened by temporal instability. Visual interaction increases as musicians become familiar with the piece, suggesting that they visually monitor each other once a shared interpretation of the piece is established. Visual monitoring of co-performers’ movements and attention may facilitate feelings of engagement and high-level creative collaboration. --------------------------------------------------------------------------------","When people interact, eye gaze serves as a means of giving, as well as obtaining information. The direction of a person’s gaze can indicate their focus of attention – which might be on another person, a salient environmental stimulus, or the direction they intend to move. Indeed, some researchers have hypothesized that the human eye evolved its current structure (specifically, the contrast between the coloured iris and white sclera) because people benefitted from the ability to track each other’s gaze direction (Tomasello, Hare, Lehmann, & Call, 2007). When eye gaze is directed towards another person, it often indicates an intention to interact. Imaging studies have shown that direct eye gaze activates the “social brain”, a network of structures involved in human communication and social interaction (Senju & Johnson, 2009). It follows that eye gaze can be a valuable means of communicating during joint action tasks, such as playing team sports or performing with a music ensemble. Our study assessed the communicative functions of eye gaze during music ensemble performance, a form of joint action that requires precise and multi-layered coordination between participants. Individual note onsets and offsets, as well as expressive nuances (e.g., changes in patterns of loudness and articulation) have to be aligned in time, even though the information provided by the score (if there is one) may be limited and subject to interpretation. During ensemble performance, visual communication between musicians is secondary in importance to auditory communication, but can become more relevant when performers are uncertain of each other’s interpretations (Bishop & Goebl, 2015). Prior research has shown that temporal instability, such as happens at piece entrances, long pauses, and sudden tempo changes (Bishop & Goebl, 2017; Davidson, 2012; Kawase, 2014a, 2014b), and disruptions to the clarity (e.g., loudness) of inter-performer audio feedback (Fulford, Hopkins, Seiffert, & Ginsborg, 2018) can prompt an exchange of visual cues. These are sometimes deliberately integrated into a performance plan across rehearsals (Williamon & Davidson, 2002). The current study was designed to build on these findings by examining patterns of uni- (one- way) and bidirectional (two-way or mutual, see Emery (2000)) eye gaze as they unfolded across the course of piano and clarinet duo performances. Our aim was to identify the conditions that prompt performers to interact visually. Based on theoretical understandings of the mechanisms that support musical interaction (discussed in the next section), we hypothesized that differences in playing conditions would encourage different degrees of visual interaction. Functions of eye gaze in musical interaction ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We defined two functions of interperformer eye gaze, both of which we expected to see evidence of during duo performance. Our definitions of these functions are theoretically- motivated, which we describe below, and demonstrate our attempt to reconcile the sometimes-contrasting predictions that derive from different perspectives on musical interaction. By defining these functions and seeking measurable behavioural evidence of them, we aimed to highlight the variable nature of musical interaction and shed some light on which processes support duo performance as playing conditions change. Our underlying argument was that, as a performance unfolds, musicians are at some moments in time drawn to engage deliberate planning and communication processes to ensure that coordination is achieved or maintained. They might exchange gestural cues, facial expressions, or make audible inhalations to facilitate synchronization following a long pause, for example. Outside of these moments, they allow coordination to “emerge” from the subtle, pre- reflective accommodations that they make to each other’s audiovisual feedback. Of course, successful emergent coordination at note and expressive levels is contingent on performers attending closely to each other and their combined musical output. Engagement-driven eye gaze Engagement-driven eye gaze is used by musicians to monitor each other’s attention and engagement in the joint performance task, as well as to indicate their own attention and engagement. This information can be provided by musicians’ gaze direction, facial expressions, and body movements. For example, eye gaze directed towards the observing co-performer might indicate a focus on the interperformer interaction, while a particular pattern of body sway might indicate focus towards an aspect of expression. It is important for musicians to be aware of each other’s focus of attention (e.g., whether it is directed to the score, or to the co- performer, signalling an intent to interact, or elsewhere, signalling potential distraction), because their attentiveness relates to how responsive they are likely to be to subtle fluctuations in each other’s audiovisual signals. In the context of the current study, evidence of engagement-driven gaze would include the occurrence of bidirectional gaze that does not occur exclusively at moments of temporal ambiguity (though perhaps at other structurally significant points; e.g., piece ending). Bidirectional gaze between duo performers allows visual signals to flow in both directions and signifies that both performers are monitoring each other. In particular, bidirectional gaze that occurs during periods of temporal stability, when visual communication is not necessary for performers to maintain coordination (Bishop & Goebl, 2015; Fulford et al., 2018; Kawase, 2014a), is suggestive of performers’ attempts to confirm each other’s attention. Intention-driven eye gaze Intention-driven eye gaze is used by musicians to communicate their individual intentions and learn their co-performers’ individual intentions. By “individual intentions”, we refer to action-based plans that are known to one performer, but not necessarily to others (e.g., a performer may intend to play a particular passage softly). These are to be distinguished from “shared intentions”, which are action-based plans that overlap between performers (e.g., all performers may intend to play the passage softly). In particular, musicians are known to communicate visually to align their intended timing, and their cueing gestures often take the form of an exaggerated nod (Bishop & Goebl, 2017). In the context of the current study, evidence of intention-driven gaze would include the occurrence of unidirectional gaze at moments of temporal ambiguity (e.g., piece onset), particularly when directed from follower to leader. Musical passages with ambiguously or imprecisely notated timing may be interpreted differently by different performers; thus, uncertainty about each other’s interpretation might prompt performers (especially those in a leading role) to communicate their own intentions or (in the case of followers) seek out information about their co- performer’s intentions. In such cases, the leader might watch the follower while performing a gestural cue to ensure that the follower is paying attention. While the follower’s gaze is intention-drive, the leader’s follower-directed gaze is engagement-driven. In the next section, we discuss the theoretical underpinnings of these definitions. Cognitivist and enactive perspectives on musical interaction ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Some forms of joint action involve deliberate collaboration between group members with the aim of achieving a particular goal. Tasks such as coordinating a pass between people on the soccer field, lifting a table into a truck, or playing a piano duet according to a particular style fall into this category. In the literature, this is referred to as planned coordination, and defined in contrast to emergent coordination, which occurs unintentionally as people make (largely) automatic accommodations to low-level features in each other’s audio and visual signals (e.g., people who are walking together may find themselves stepping in sync; van der Wel, Sebanz, & Knoblich, 2016). It is important to note that planned and emergent coordination are not mutually exclusive, and during ensemble performance, occur in parallel. According to the cognitivist account of joint action, during planned coordination, shared intentions and shared attention to the task are necessary for group members to achieve their desired goal (Keller, 2014; Knoblich & Sebanz, 2008). The term intentions refers broadly to a dynamic, action-based process of representing or anticipating upcoming actions or events. This process is flexible – intentions are constantly evolving as performers monitor their own and their co- performers’ output. Flexibility in action planning is important because it allows group members to compensate for variability in each other’s performance, which might arise because of errors, environmental disturbances, or the spontaneous introduction of new ideas (Loehr, Kourtis, Vesper, Sebanz, & Knoblich, 2013; Noy, Dekel, & Alon, 2011; for an overview see Bishop, 2018). Shared intentions comprise overlapping, but not identical, representations of how individual efforts should contribute to the overall outcome. In essence, each member of the group understands their own contribution in the terms of how it will combine with others’ contributions (Gallotti & Frith, 2013). Another perspective on the processes underlying interpersonal coordination derives from embodied, enactive, and distributed cognition and dynamical systems theories (Geeves & Sutton, 2014; Leman & Maes, 2014). By this account, entrainment between individuals occurs spontaneously as a result of low-level sensory (auditory and visual) couplings, or accommodation (Maes, 2016). For ensemble musicians, this process potentially enables coordination at both local (note) and global (expressive) levels. Some researchers argue that shared intentions, as defined by cognitivists, may not be needed for musicians to coordinate their performance in this way, because coordination can arise independently of such top-down control as musicians engage in cycles of small-scale responses to the joint output (Schiavio & Høffding, 2015). Evidence of coordinated improvised performance in the absence of preplanned structures has been observed in both music (Canonne & Garnier, 2015) and dance domains (Kimmel, Hristova, & Kussmaul, 2018), suggesting that emergent coordination can indeed occur in artistic contexts. MacRitchie, Varlet, and Keller (2017) propose that representational and entrainment processes may run in tandem, with one or the other exerting a dominant influence on behaviour depending on the performers’ familiarity with each other’s roles in the performance. For example, in the early stages of rehearsal, ensemble members may be unsure of how each other will interpret an unfamiliar piece, and representational processes may dominate as they try to establish a shared interpretation. Jointly rehearsing a new piece was the task given to participants in the current study. As they rehearsed, their familiarity with each other’s playing increased and they established shared intentions for how the piece should sound, potentially reducing the likelihood that they would have to communicate individual intentions in order to coordinate their performance. One of the specific hypotheses we tested was whether performers would spend less time watching each other as the rehearsal progressed. Such a finding would suggest that gaze is largely intention-driven during the early stages of rehearsal, and decreases once performers are more certain of how each other intends to play. The opposite finding – an increase in partner-directed gaze across the rehearsal (already observed in trios; see Vandemoortele et al., 2018) – would suggest that familiarity with the music and with a shared, practiced interpretation encourages the use of engagement-driven gaze. Influences on gaze behaviour ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ With these perspectives in mind, our study considered the possibility that eye gaze serves multiple communicative functions during ensemble performance – specifically, that it is primarily a means of affirming co-performers’ attention and engagement, and secondarily a means of ascertaining co-performers’ intentions. Engagement-driven gaze, which allows for the exchange of subtle attention, expressive, and movement-related signals, may help to support entrainment processes. Intention-driven gaze, in contrast, which reflects a drive to exchange intentions that would not otherwise be shared, may help to support representational processes. Much of the research on visual interaction during ensemble performance focuses on musicians’ body gestures. The questions addressed include how timing information might be encoded in these gestures, and how well other musicians are able to synchronize their own gestures with those they observe (Bishop & Goebl, 2017, 2018; Wöllner & Cañal-Bruland, 2010; Wöllner, Parkinson, Deconinck, Hove, & Keller, 2012). The assumption behind these studies is that musicians choose to attend to each other’s gestures to obtain information that might facilitate coordination. In particular, a performer who is assuming a “follower” role might monitor the performer who is assuming a “leader” role for gestural cues. The current study investigated whether patterns of partner-directed gaze relate to performers’ assumption of leader/follower roles. Research has already shown that leader-follower relationships emerge among ensemble members (Badino, D’Ausilio, Glowinski, Camurri, & Fadiga, 2014; Timmers, Endo, Bradbury, & Wing, 2014), and that these relationships are reflected in the amount of time that performers spend looking towards one another – when leader/follower roles are assigned, followers look more towards leaders than vice versa (Kawase, 2014a). Many factors contribute to the emergence of leader-follower relationships in musical duos, some relating to genre conventions (e.g., in Western classical music, the performer playing the solo or melody typically leads) and others relating to social factors (e.g., relative skill levels of the performers, personality traits, etc.). As a contrast to previous studies in which leader/follower roles were assigned, we tested specifically whether leader-follower relationships that are implied by piece structure (i.e., melody vs. accompaniment assignments) influence gaze patterns. If performers were to watch their partner more when following (playing accompaniment) than when leading (playing the melody), it would suggest the use of gaze as a means of obtaining information about the leader’s intended interpretation. Our hypothesis was that leader/follower relationships might be reflected in gaze behaviour during some parts of a performance (e.g., during periods of high temporal instability), but not during others (e.g., during periods with regular timing), indicating an interaction with musical structure. Musicians’ tendency to visually monitor each other’s attention and engagement was likewise expected to rise and fall across the course of a performance. To assess this tendency, we tested whether performers tended to look more towards their partners’ face than towards their partners’ bodies or instruments. In cases where instrument motion is likely to be informative (e.g., during clarinet performance), a preference for looking at the face would suggest that gaze is being used to communicate information about attention. We also considered how often partner-directed gaze was bidirectional, rather than unidirectional. During moments of bidirectional gaze, information flows both ways simultaneously – from primo to secondo and from secondo to primo – which, in addition to giving both performers visual access to each other’s body gestures and facial expressions, communicates to each performer that they are the focus of their partner’s attention. Direct eye gaze is treated specially by the human brain, affecting the way people perceive others and judge their mental capabilities. Faces with direct gaze are judged as belonging to people with more sophisticated mental faculties than faces with averted gaze, for instance, possibly because direct gaze acts as a cue to impending social interaction and prompts people to expend more effort in considering others’ internal experiences (Khalid, Deska, & Hugenberg, 2016). In certain types of interaction, gaze is drawn to others’ eyes and faces. For example, deaf viewers have been found to focus mostly on signers’ faces when watching sign language video clips, processing hand and arm movements largely with their peripheral vision (Muir & Richardson, 2005). Face-directed gazing during signing allows the viewer to detect facial expressions and lip movements, and it allows the signer see that the viewer is paying attention. Along the same lines, studies of spoken conversation have shown that speakers expect listeners to confirm their attention through auditory or visual backchannelling, and will periodically pause and look towards their listeners, seeking such a response (Bavelas, Coates, & Johnson, 2002). Direct eye gaze may likewise contribute to backchannelling during music ensemble performance. More broadly, in non-musical forms of interaction, perceived gaze direction has been shown to help people track each other’s attention. A study by Khoramshahi, Shukla, Raffard, Bardy, and Billard (2016) provides a clear example. In this study, human participants collaborated with an avatar on the “mirror game”, a task in which two participants, while facing each other directly, move a pair of sliders along horizontal tracks to produce creative but synchronized patterns of motion. Participants were told to follow the avatar’s lead. When the avatar provided anticipatory gaze cues (i.e., began looking in the direction of its next movement just before initiating the movement), leader-follower synchronization was better than when the avatar visually tracked its hand movements (i.e., always looked directly at its moving hand). Anticipatory gaze cues, therefore, allowed participants to account for the avatar’s intentions in their own action planning. We hypothesized that ensemble musicians monitor each other’s gaze direction when it is possible to do so, possibly because tracking fluctuations in each other’s attention – as well as receiving backchannelling signals – helps to support emergent coordination. Current study ~~~~~~~~~~~~~ This study investigated patterns of eye gaze during duo piano and clarinet performance. Motion capture and eye tracking were used to map musicians’ body gestures and gaze patterns as they rehearsed and performed a new duet piece. The piece was composed specifically for this study and contained a number of potential challenges for coordination (see Section 2.3). Duos recorded four full performances of the piece between periods of free joint rehearsal. The final performance was given with no visual contact between performers; the other performances and free rehearsal periods were completed under normal visual contact conditions. Eye gaze data were analysed to determine when and how much musicians looked towards their co-performers. Motion data are presented in Bishop and Goebl (submitted for publication). Patterns of partner-directed eye gaze were expected to distinguish conditions that encourage the use of visual interaction for confirming co- performers’ attention and engagement (engagement-driven gaze) from the use of visual interaction for communicating individual intentions (intention-driven gaze). To this end, we assessed the potential effects of rehearsal, leader/follower relations, and musical structure on the amount of time that musicians spent watching their partners. It should also be noted that glances from one performer to another can be multifunctional, enabling a performer to gain information about their co-performer’s attention and intentions simultaneously. For this reason, the manipulation of musical structure was particularly important, as we expected temporal uncertainty resulting from imprecisely notated timing to encourage intention-driven gaze in particular. Engagement-driven gaze, in contrast, could occur at any time, and was not expected to account for heightened tendencies for partner-directed gaze at moments of temporal instability. As evidence of engagement-driven gaze, periods of bidirectional gaze were expected to occur throughout performances. Musicians were expected to look at each other even during temporally stable passages and during performances that followed the rehearsal period. For clarinet duos, partner- directed gaze was also expected to target the face more often than the instrument/body, even though clarinet bell movements are known to carry information relevant to expression and timing (Wanderley, Vines, Middleton, McKay, & Hatch, 2005). As evidence of intention- driven gaze, musicians were expected to look towards their partner when uncertain of their partner’s intentions – for example, at piece onset, during periods of temporal instability in the music, and at the start of the rehearsal period. They were also expected to look more towards their partner when playing an accompaniment passage (i.e., when following) than when playing a melody passage (i.e., when leading). Design ~~~~~~ The effects of three main within-subject independent variables were tested: (1) piece structure, (2) rehearsal time, and (3) leader/follower roles. The structure of the duet piece is outlined below (see Section 2.3). As part of the analysis procedure, the piece was segmented into windows of interest, and performers’ behaviour was compared between windows (piece structure variable). The three performances recorded with normal visual contact were given before, partway through, and immediately following a rehearsal period (rehearsal time variable). Instrument pairing (piano-piano or clarinet-clarinet) was manipulated between subjects. Our primary dependent variable was the amount of time (i.e., percentage of recorded gaze samples) that performers spent looking at their partner. For clarinettists, as a secondary dependent variable, we also compared the percentage of time dedicated to face-directed versus body/instrument-directed gaze. Stimuli and equipment ~~~~~~~~~~~~~~~~~~~~~ A duet was composed for the experiment by the second author, a pianist and composer with 14 years of musical training. Some excerpts from the piece are given in Fig. 1. Though primo and secondo parts were not technically very difficult, the piece was meant to be challenging for a duo to coordinate. There were several changes in meter and tempo, some unusual meters (e.g., 5 + 7/8), one unmetered section, and sections with accent patterns that the primo and secondo were intended to synchronize. The piece was initially written as a piano duet, then arranged for two clarinets with the aid of a professional clarinettist. See the Appendix for the full piano score. Performers wore SMI ETG 2 wireless glasses, through which eye gaze was tracked at 120 Hz (Fig. 2). These glasses track gaze with up to 0.5° accuracy over all distances (SensoMotoric Instruments (SMI), 2017). The glasses cannot be worn over normal prescription glasses, which create too much interference; however, magnetic snap-on corrective lenses were available for participants requiring a distance correction (within the prescription range of −4.0 to +4.0). The glasses were fitted with markers for detection by the motion capture system. Two markers were placed on top of the glasses frame, at the corner of each lens. The third marker was placed on a small stick, which was attached to the glasses frame with dental putty and extended down from the left arm of the glasses. A 10-camera (Prime 13) OptiTrack motion capture system was used to track performers’ upper body movements, recording at a rate of 240 frames per second. Each performer was fitted with 25 reflective markers. Additional markers (3) were fixed to the eye tracking glasses, the corners of the music scores (4), and, in the case of clarinet duos, both performers’ instruments (4). Eye gaze data for the two performers were collected with separate laptops, both of which were connected to an OptiTrack eSync 2 device (see Fig. 2), which transmitted TTL triggers from the motion capture software to the eye tracking software at the start and end of each motion capture recording. These triggers were logged in the gaze data files as a numerical variable, allowing us to retrospectively align eye gaze and motion capture data by trimming data files to exclude observations outside the trigger range. Pianists performed on Yamaha Clavinovas, from which audio and MIDI data were collected via a Focusrite Scarlett 18i8 sound card and recorded on separate tracks in Ableton Live. Clarinettists performed on their own instruments, and their audio was likewise recorded in separate tracks, using DPA d:vote 4099 clip-on microphones. A clapboard was placed within range of an additional room microphone and within view of the OptiTrack and SMI cameras and struck once at the start and end of each recording. This provided a synchronization stimulus, recorded on all devices, that allowed us to align audio/MIDI with motion capture and eye gaze data.","At the start of the session, performers were presented with the piece and randomly assigned either the primo or secondo part. They were told that our aim was to investigate performer interaction during rehearsal of unfamiliar music. Duos were asked to practice the piece in preparation for making some high-quality recordings at the end of the session. They were not required to memorize the music. Performers were positioned so that they faced each other, roughly 1.5 m apart. Clarinettists played standing without any specific instruction on how to orient themselves, so they had more freedom to move around than did pianists (e.g., see how the clarinettists’ scores ended up at an angle to each other in Fig. 3b). An initial (sight-read) performance was recorded first, with performers encouraged to ignore errors and play through as much of the piece as possible without stopping. This performance was followed by up to 20 min of free joint rehearsal. (Some duos required more practice time than others. To prevent duos from reaching peak performance by the end of the first rehearsal, duos who were progressing quickly were asked to stop when the experimenters could hear that they still had some passages to work out.) A second full performance was then recorded, followed by up to 20 more minutes of rehearsal. Finally, two “polished” performances were recorded, one with visual contact between performers, and then one without (always in that order). Musicians also completed a short questionnaire on their musical background and answered a few debriefing questions about their perceptions of the experiment. Note onsets Note onsets in clarinet performances were identified manually using waveform and spectrogram information. For piano performances, MIDI data collected from the Clavinovas were matched to the score using the score-performance matcher developed by Flossmann, Goebl, Grachten, Niedermayer, and Widmer (2010). This system pairs performed pitches with score notes based on pitch sequence information, disregarding the timing. Incorrectly performed pitches (additions and substitutions) are omitted, so the resulting matched performance profile includes only correctly performed notes. Eye gaze Vectors indicating direction of gaze for the right eye of each participant were exported from the eye tracking software. For each sample of eye gaze data, a “gaze target” was identified. For pianists, gaze targets took one of three values, indicating whether the participant was looking towards the score (“score”), the other performer’s face (“face”), or elsewhere in the room, including towards the piano keys (“other”). For clarinettists, gaze targets could take an additional value, indicating whether the participant was looking at the other performer’s upper body/instrument (“instrument”) (Pianists were only able to see each other’s faces over their scores, and clarinettists’ views of each other’s lower bodies were largely occluded by the music stands). It was necessary to allow for small degrees of imprecision in the gaze vector coordinates obtained by the glasses resulting from calibration errors, which sometimes occurred over the course of the 1-h recording session (e.g., if participants bumped or shifted the glasses following the initial calibration step). To account for such errors, we expanded the region of each gaze target by a small amount so that gaze vectors that fell slightly outside a region would be counted as intersecting that region. For all duos, markers indicating the sides of the head were adjusted outwards by 1.4 times the length of the vector between them; the same was done for the markers indicating the left and right shoulders, and for the markers outlining the clarinets. The assumption here was that participants were unlikely to spend time watching anything but each other, the score, and their instruments, given their unfamiliarity with the music. Thus, a participant who appeared to be staring at a spot on the wall just to the left of their co-performer’s head was most likely actually looking at their co-performer’s face. We developed an automated procedure for identifying gaze targets, which involved remapping gaze vector coordinates into the motion capture space. As a first step, motion capture and eye tracking data were temporally aligned using the recorded trigger values (see Section 2.3), and linear interpolation was used to resample gaze vector profiles at 240 Hz, the sampling rate of the motion capture recordings. Finally, we tested whether the corrected gaze vector intersected with the score, the other performer’s face, or, for clarinettists, the other performer’s body or instrument. Fig. 3 provides visualizations of these tests. Scores were defined as four-sided polygons using markers that had been placed in the corners. Performers’ faces were defined as four adjacent triangles using markers located on the top and sides of the head, on the chest, and on the shoulders. Clarinettists’ upper bodies were designated by five adjacent triangles, using markers located on the shoulders, chest, upper arms, and waist. Clarinets were defined as four-sided polygons, using markers placed near the mouthpiece and near the bell. A sample video showing the gaze behaviour of two pianists during a final performance can be viewed in the supplementary files. Missing data Eye gaze data are not reported for two pianists and two clarinettists, for whom we were unable to get reliable recordings. (The glasses are sometimes unable to track pupil position; for example, if the participant is squinting, has swelling around the eyes, or has a flatter nose/narrower eyes than are typical for people of European descent.) Motion capture data was also lost for one clarinet duo and partially lost for a second clarinet duo, due to file corruption. Differences in temporal stability across piece sections ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ One of our primary hypotheses was that the percentage of time that performers spent watching each other would increase in passages where timing information was loosely- defined by the score (and the chances of performers’ interpretations diverging was greater) and decrease in passages with less temporal ambiguity. We refer to passages with loosely-defined timing information as “temporally unstable”, because variability in performance timing was expected to increase. To confirm, analyses of (1) tempo variability and (2) primo-secondo note asynchronies were conducted for piano duo performances. The analyses were only run on the MIDI data acquired from pianists since note onsets in clarinet performances were identified manually, and therefore less precise. The piece was divided a priori into sections, with boundaries placed where they made musicological sense (i.e., when meter and texture changed). The first 8 bars (“entrance”), the final 10 bars (“ending”), and the unmetered section from the middle of the piece (“unmetered”) were expected to invoke greater temporal instability than the other, more regularly-timed sections (“regular”). Interbeat intervals (IBIs) were taken as a measure of performance tempo using averaged primo-secondo note onsets. As a measure of within-performance tempo variability, the series of IBIs was differenced once (Fig. 4). The resulting series indicated how IBI durations changed across the course of the piece: high values reflected large note-to-note changes in tempo, while small values reflected a more consistent tempo over time. The increased tempo variability and note asynchrony that occurred in the unmetered section of the piece confirms our expectation that timing would be less stable during this passage. In the following sections, we test for effects of piece structure on eye gaze behaviour with the expectation that this section in particular will prompt visual interaction between performers. Effects of piece structure, rehearsal time, and instrument on unidirectional partner-directed eye gaze ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 6 shows the percentage of time per beat, averaged across individual performers, that gaze vectors intersected with a co-performer’s face, body, or instrument – that is, the average percentage of time for which unidirectional partner-directed gaze occurred. Results are given separately for clarinettists and pianists, as some differences were observed between instrument groups (see below). Thus, performers spent more time looking at each other during periods of temporal instability than during periods of regular timing, in line with the hypothesis that uncertainty about co-performers’ intended timing would prompt visual interaction. On the other hand, they spent more time watching each other after rehearsing than before, indicating an increase in visual interaction once they had established shared intentions regarding how the piece should sound. Differences between sections were stronger for clarinettists than for pianists. Effects of leader/follower roles on unidirectional partner-directed eye gaze ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Sections of the piece in which leader/follower roles were implied by a melody/accompaniment structure were identified a priori. Primos led a total of 27 bars, distributed across entrance (4 bars), unmetered (half of the section, counted as 1 bar), regular (12), and ending (10 bars) sections, while secondos led a total of 14 bars, distributed across entrance (4 bars), unmetered (half of the section, again counted as 1 bar), regular (1 bar), and ending (8 bars) sections. Effects of piece structure and rehearsal time on face- versus body/instrument-directed eye gaze among clarinettists ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Pianists in this experiment were only able to see each other’s faces over the top of their music stands, but clarinettists were standing and able (at least partially) to see each other’s upper bodies and instruments. So, for clarinettists, we tested for differences in how much time they spent looking at their partner’s face versus their partner’s body/instrument (Fig. 9). Effects of piece structure, rehearsal time, and instrument on bidirectional partner-directed eye gaze ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 10 shows the average percentage of time per beat that both performers’ gaze vectors simultaneously intersected each other’s face, body, or instrument – the average percentage of time for which bidirectional partner-directed gaze occurred. Thus, bidirectional gaze occurred rarely, but varied according to piece structure and rehearsal time similarly to unidirectional gaze – occurring more often during moments of temporal instability, and increasing as performers became more certain of how they wanted the piece to sound. As was the case for unidirectional gaze, differences between piece sections were greater for clarinettists than for pianists.","This study assessed the patterns of unidirectional and bidirectional eye gaze that occurred during duo performance of an unfamiliar piece, before and after brief periods of rehearsal. Piano and clarinet duos recorded four performances of a piece – one before, one partway through, and two at the end of a 20–40 min rehearsal session. Piece structure and rehearsal time were found to influence the percentage of performance time during which uni- and bidirectional gaze occurred. Specifically, musicians looked more often at their partners during periods of temporal instability than during periods of regular timing, suggesting the use of intention-driven gaze. They also spent more time watching each other after rehearsing than before, suggesting an increasing role for engagement-driven gaze as they learned the music. Clarinettists, finally, exhibited a preference for face-directed rather than body-directed gaze during the entrance section of the piece. Each of these observations is discussed in greater detail below. One of our central hypotheses regarding intention-driven gaze was that moments of temporal instability would prompt performers to interact visually, and this hypothesis was confirmed. Performers spent a greater percentage of their time watching each other during the unmetered section than during regularly-timed sections. This finding is in line with the idea that ensemble musicians exchange visual signals when temporal coordination is threatened by ambiguity in the timing information given by a musical score (Bishop & Goebl, 2015; Kawase, 2014a, 2014b). More broadly, it aligns with research suggesting that, on interpersonal coordination tasks, people draw flexibly on information from different modalities, depending on which modality seems most reliable at a given moment (Elliott, Wing, & Welchman, 2010; Hove, Iversen, Zhang, & Repp, 2013). During the unmetered section of the piece, temporal ambiguity came from both the absence of a notated meter – which encouraged high note-to- note variability in tempo – and the presence of long held notes, including the notes with fermatas shown in Fig. 1. As a side note, we should mention that several (though not all) duos deliberately integrated visual cues into their performance routine during this section of the piece, following verbal discussion (e.g., “watch me here, and I’ll give the cue”). This observation is in line with findings reported by Williamon and Davidson (2002), who also found musicians to deliberately integrate visual cues into a performance plan in this way. During the entrance section, the primary instance of temporal ambiguity came at the moment of piece onset, which the primo and secondo were intended to synchronize. Both uni- and bidirectional gaze peaked during the seconds just prior to piece onset. On one hand, it is unsurprising that performers watched each other closely as they prepared to coordinate the first chord of the piece, given the absence of prior audio. On the other hand, it is interesting to note the peak in bidirectional gaze at this point. A peak in uni- but not bidirectional gaze would have suggested that one performer visually monitored the other for information about when to begin. The peak in bidirectional gaze that we observed instead suggests that performers monitor each other. In doing so they may exchange a range of information about their shared focus of attention, their joint willingness to interact, and their intended timing of the first notes of the piece. Uni- and bidirectional gaze both also peaked in the beats leading up to the final chord of the piece. This was particularly the case for clarinettists. Many duos gradually slowed their tempo in the last notes of the piece, employing the well-known expressive device referred to as the “final ritard” (Friberg & Sundberg, 1999; Honing, 2003), but overall, there was less temporal ambiguity in this section of the score than, for example, in the unmetered section. We would therefore suggest that the increase in visual interaction at the end of the piece reflects performers’ attempts to monitor each other’s attention, rather than exchange specific information about their intended timing. For pianists, there was also an unexpected uptake in unidirectional gaze during one of the regularly timed sections of the piece. This uptake occurred in the third bar (53) of the line shown in Fig. 1b (around beat 380 in Figs. 6 and 10). It seems unlikely that performers would be prompted to look to each other for timing cues during this passage, as the timing is relatively unambiguous. Rather, we would speculate that the uptake in partner-directed gaze relates to the positioning of the passage on the score: this line appeared at the top of the page that stood in the center of the music stand, only a short vertical distance away from the co-performer’s face. It would have been easier for performers to glance at their partners while playing this line than while playing most other parts of the piece, which would have required the eyes to traverse a greater distance. The piano score spanned six pages, which we fixed together as two groups of three, so that there was only one page turn at the end of page 3. Beat 380 fell at the top center of page 5. A corresponding, though smaller, uptake can be seen at the start of the second regularly-timed section on score page 2 (see third dotted vertical line in Fig. 6), which likewise fell at the top center of the music stand, though this uptake could also be attributed to the change in meter that occurs at this spot. Uptakes in partner-directed gaze did not occur for clarinettists at either location, likely because they were free to reposition themselves (since they were standing), and often ended up at an angle rather than facing each other directly (as in Fig. 3b). The fact that pianists chose to glance at their partner when it became easier to do so is telling – it suggests a propensity to engage visually even when coordination does not depend on it. We also tested the hypothesis that performers would look less at their partners after rehearsing than before. This is what we would expect to see if visual interaction were primarily a means of communicating individual intentions. Performers’ familiarity with each other’s intentions would increase as they established a shared interpretation of the piece, reducing their reliance on visual cues. Challenging this hypothesis, however, the results showed an increase in partner-directed gaze with rehearsal. It seems that performers chose to engage visually with each other once they were familiar enough with the score to be able to glance away from it. Note synchronization was maintained in the final performance, even though performers could not see each other, indicating that duos do not require visual contact in order to coordinate temporally – in line with prior research (e.g., Bishop & Goebl, 2015; Fulford et al., 2018). The increase in partner-directed gaze that occurred across the first three performances thus indicates a desire to interact visually when doing so is not necessary to maintain coordination. Such a finding is in line with the hypothesis that gaze can be engagement-driven – that musicians interact visually in order to monitor each other’s attention and engagement. It might also be that performers seek additional means of interaction to ensure coordination once they have specific expectations for how the piece should sound, and what each performer should contribute. For clarinettists, we found a preference for face- over body-directed gaze only in the entrance section of the piece. We had hypothesized that clarinettists would spend more time looking at each other’s faces than each other’s bodies if their aim was to monitor each other’s facial expressions and gaze direction; in contrast, greater focus on the body/instrument might indicate an attempt to monitor gestures involved in sound- production, including fingering and breathing. Our observations suggest that clarinettists were compelled to monitor each other’s expressions and attention as they coordinated the entrance of the piece. The lack of meaningful results in other conditions could be partially attributable to imprecision in the identification of gaze vectors that intersected near the borders of the face and instrument regions. A possible consequence is that performers who were watching their partner’s instrument close to the mouthpiece might sometimes have been categorized as watching their partner’s face. The best fix to this issue would have been periodic recalibration of the glasses throughout the 1-h recording session. Another likely explanation is the way clarinettists were positioned in the recording space in relation to each other. Their view of each other was partially blocked by the music stand and score, and depending on how they chose to stand (e.g., close to the music stand vs. further away; facing their partner directly vs. at an angle), how high they had adjusted the music stand, how tall they were compared to their partner, and how much they and their partner moved, their views of each other were sometimes (variably, throughout the performance) partially occluded. We chose to prioritize naturalistic playing conditions by giving performers some freedom of movement, but a more controlled recording setting might yield a clearer identification of which cues (face or body) clarinettists attend to while performing. We had hypothesized that performers might spend more time watching their partner when playing accompaniment passages (following) than when playing melody passages (leading). This hypothesis was not confirmed; partner-directed gaze was not reliably influenced by the leader/follower roles implied by melody/accompaniment relationships in the score (though there was a nonsignificant tendency for performers to watch their partners more when following than when leading, in all but the ending section). This lack of effect is in contrast to the idea that followers visually monitor leaders for cues indicating how to play. Instead, it seems that both leaders and followers have reason to watch each other. The lack of effect here is at odds with prior research that has shown followers to look more towards leaders than vice versa (Kawase, 2014a, 2014b), possibly because of differences in experimental design. Kawase’s participants were assigned leader and follower roles, while our participants were not given any explicit instruction regarding who should lead or follow. We were instead interested in whether gaze patterns would be affected by the leader/follower relationships that emerge naturally during duo performance, with the acknowledgement that these relationships vary within and between performances and are influenced by numerous factors, including piece structure, performance conventions, and the social and personality dynamics that arise between the performers. Future research might investigate whether some of these other factors have a stronger influence on gaze behaviour. It is important to acknowledge limits to the generalizability of our results. We tested a specific type of performance situation: participants were rehearsing, not performing for an audience; the music was unfamiliar, not previously practiced; and participants were positioned in the recording space in a way that encouraged visual contact. Changes to any one of these factors would likely influence gaze behaviour. Skilled musicians are flexible in their abilities to draw cues from their co-performers’ audio/visual signals, and readily adapt to changes in playing conditions (Bishop & Goebl, 2015; Glowinski, Bracco, Chiorri, & Grandjean, 2016). As such, there is no single pattern of interactive or communicative behaviours that could be said to enable duo performance across different playing conditions. The patterns of gaze behaviour that were observed here are interesting primarily because of what they imply about how musical interaction unfolds. Our results imply that duo musicians who have full access to each other’s audio signals are prompted to supplement their audio interaction with interaction via a second modality (vision) when they want confirmation of piece timing or their partner’s focus. Future research might investigate how interaction techniques differ across performance conditions. The design of this experiment rests on the assumption that musicians who are learning a new piece from the score will have to focus their visual attention on the score most of the time. The fact that they have few opportunities to look elsewhere makes their partner-directed glances meaningful – they would not divert their gaze away from the music unless they had reason to do so. Without such a demand on their visual attention (e.g., if the music were memorized or the task was to improvise), it would not be possible for us as experimenters to deduce meaning from their partner-directed gaze. In this paper, we have attributed meaning to participants’ partner-directed glances according to the conditions in which they occur. First, we have argued that partner-directed glances occurring in moments of potential temporal variability likely reflect an effort to confirm shared expectations regarding performance timing. In such moments, performers may be watching for cues indicating their partner’s intended timing. In interpreting some of our other findings, we have suggested that performers keep track of each other’s attention and engagement by visually monitoring gaze direction, body movements, and facial expressions. The frequency of bidirectional gaze and the increase in partner-directed gaze that occurred with rehearsal were observations that we took as evidence that performers choose to visually monitor each other even when coordination is not threatened by uncertainty over timing.","The results of this study suggest that duo musicians seek visual interaction with their co-performer, particularly – but not exclusively – when coordination is threatened by temporal instability. Performers seem to monitor each other for timing cues when uncertain of whether temporal coordination will be successful. Visual interaction also increases as performers grow more familiar with the piece, which may reflect an attempt to engage with each other in a way that enables higher levels of collaboration. Gestural cues are an important source of information as performers interact visually, but a characterization of performers’ communicative gestures is still missing from the literature. Our ongoing investigation of the motion data collected as part of this study examines how performers move when they know that a co-performer is watching, where in the performance explicit cueing gestures are given, and how performance gestures evolve across rehearsals. In the literature, it has been hypothesized that emergent coordination between ensemble members is supported by low-level sensory couplings (e.g., Maes, 2016). Visually monitoring co- performers should facilitate such couplings by relaying information about body movement and facial expressions. Furthermore, performers who are able to visually confirm each other’s attention and engagement might feel more engaged in the collaborative task themselves, and perhaps more willing to take creative risks. Along these lines, an interesting continuation of this research would be to test the contribution of visual interaction to the quality of ensemble musicians’ performance experiences and, importantly, to the occurrence of group flow. The term group flow describes an intrinsically-rewarding state of intense focus, effortlessness, and perceived success that is shared among collaborating members of a group (Cochrane, 2017; Hart & Di Blasi, 2015; Sawyer, 2006). Reports from ensemble musicians suggest that a feeling of deep connectedness between co-performers is critical to group flow (Hart & Di Blasi, 2015). Visual interaction could help to support such feelings of connectedness."],["Background: Research evidence from studies in North America on the relationships between family-centered practices, parents’ self-efficacy beliefs, parenting confidence and competence beliefs, and parents’ psychological well-being was used to confirm or disconfirm the same relationships in two studies in Spain. Aims: The aim of Study 1 was to determine if results from studies in North America could be replicated in Spain and the aim of Study 2 was to determine if results from Study 1 could be replicated with a second sample of families in Spain. Methods and procedures: A survey including the study measures was used to obtain data needed to evaluate the relationships among the variables of interest. The participants were 105 family members in Study 1 and 310 family members in Study 2 recruited from nine early childhood intervention programs. Structural equation modeling was used to test the direct and indirect effects of the study variables on parents’ well-being. Outcomes and results: Results showed that family-centered practices were directly related to both self-efficacy beliefs and parenting beliefs, and indirectly related to parents’ psychological well-being mediated by belief appraisals. Conclusion and implications: The pattern of results was similar to those reported in other studies of family-centered practices. Results indicated that the use of family-centered practices can have positive effects on parent well-being beyond that associated with different types of belief appraisals. --------------------------------------------------------------------------------","This paper adds to our understanding of how the effects of family-centered practices can be traced to more positive and less negative parent psychological well-being. Specifically, the results showed that how early childhood intervention practitioners interact with families and provide support is directly related to both parents’ self- efficacy beliefs and parenting beliefs and indirectly related to parent well-being mediated by self-efficacy beliefs. In turn, positive parenting beliefs were related to parents' psychological well-being. This research illustrates how the complex relationships among the variables in the study could be identified as evidenced by the direct, indirect, and total effects of the predictor variables on both belief appraisals and parent well- being. This complexity has implications for both developmental disabilities researchers and ECI practitioners in order to understand how professional practices have positives consequences for improving family functioning.","The birth of a child with an identified disability or condition, and the rearing of a child with a developmental delay, is two life events that can and often do have deleterious effects on parental psychological health and well-being (Koehler, Fagnano, Montes, & Halterman, 2014; Schwarzer & Schulz, 2003; Weekes, 1999). Research indicates that parents of young children with disabilities or delays often experience increased stress (Innocenti, Huh, & Boyce, 1992; Smith, Oliver, & Innocenti, 2001) and attenuated psychological well-being (Barlow, Cullen-Powell, & Cheshire, 2006; Raina et al., 2005) in the absence of effective coping mechanisms and social support (Hassall, Rose, & McDonald, 2005; Krakovich, McGrew, Yu, & Ruble, 2016; Trute, Benzies, Worthington, Reddon, & Moore, 2010). In addition to the adverse effects on parental psychological health and well-being, raising a child with a disability or developmental delay can also have negative effects on parents’ beliefs about their child-rearing confidence and competence (Gowen, Johnson- Martin, Goldman, & Applebaum, 1989; Hassall et al., 2005). The more difficult child- rearing entails, the more parenting confidence and competence is likely to be compromised (Dempsey, Keen, Pennell, O’Reilly, & Neilands, 2009; Mas, Giné, & McWilliam, 2016). One of the coping mechanisms that lessen stress and bolsters well-being is self-efficacy beliefs (D’Amico, Marano, Geraci, & Legge, 2013; Hall, Neely-Barnes, Graff, Krcek, & Roberts, 2012). Self-efficacy refers to “people’s beliefs about their capabilities to exercise control over events that affect their lives” (Bandura, 1994, p. 71) including efficacy beliefs about rearing a child with a disability or developmental delay (Hassall & Rose, 2005; Maclnnes, 2009). The stronger perceived self-efficacy, the less negative are the consequences of adverse life events (Lightsey & Sweeney, 2008). The weaker perceived self- efficacy, the more negative are the consequences of adverse life events (D’Amico et al., 2013; Dunning & Giallo, 2012). In addition, the stronger perceived self-efficacy, the more positive the effects on parenting beliefs (Dunst & Dempsey, 2007) whereas the weaker perceived self-efficacy, the more negative are the effects on parenting beliefs (Dunning & Giallo, 2012; Kuhn & Carter, 2006). Two types of self-efficacy beliefs were the focus of investigation in the study described in this paper: (a) family member control over the resources and supports from ECI practitioners and (b) parenting beliefs of control over affecting changes in child behavior. A considerable amount of research indicates that self-efficacy is related to a host of different psychological and physical health benefits (DeVellis & DeVellis, 2001; Grob, 2000; O’Leary, 1985), including parenting belief appraisals and psychological well-being (Nelson, Kushlev, & Lyubomirsky, 2014). Bandura (1997), for example, noted that “Research…shows that a strong sense of parenting efficacy yields dividends in the emotional well-being of mothers raising children who present special difficulties” (p. 191, emphasis added). Parents of young children with disabilities and developmental delays often come in contact with many professionals as part of participation in early childhood intervention (ECI) (e.g., Affleck, Tennen, & Rowe, 1991; Bailey, Hebbeler, Scarborough, Spiker, & Mallik, 2004; Swick, 2004; Woods & Lindeman, 2008). The ways in which professionals interact with, treat, and provide support to parents and their children can influence self-efficacy and parenting beliefs in either positive or negative ways depending on how help is provided. Research indicates that professionals’ use of family-centered practices (FCPs) is positively related to both self- efficacy beliefs and parents’ sense of competence and confidence (Dunst, Trivette, & Hamby, 2007). Practitioners who employ FCPs treat families with dignity and respect; share information so parents can make informed decisions; acknowledge and build on family member strengths; actively engage family members in obtaining resources and support; and are responsive to each families’ changing life circumstances (Dunst & Espe-Sherwindt, 2016; Dunst, 2002). Dunst and Espe-Sherwindt (2016), in their review of FCPs measures, found that different measures tend to include two sets of indicators: relational practice indicators and participatory practice indicators. Relational practices emphasize the use of relationship-building strategies, active and reflective listening, and practitioner beliefs about family member strengths and capabilities. Participatory practices emphasize informed family choice and decision making, active family member involvement in achieving desired goals and outcomes, and practitioner use of capacity-building help-giving practices. Research reviews of FCPs studies indicate that this type of help-giving is related to a host of positive parent, family, and child outcomes, including self-efficacy beliefs, parents’ sense of confidence and competence, and parent and family psychological health and well-being (Dempsey & Keen, 2008; Dunst, Trivette, Trivette et al., 2007; Rosenbaum, King, Law, King, & Evans, 1998). The relationships between FCPs and self- efficacy beliefs, however, have been found to differ as a function of the targets of belief judgments (Bugental, Johnston, New, & Silvester, 1998). Family-centered practices are more highly related to parent belief appraisals of control over practitioner and program practices and are less strongly related to belief appraisals over life events not directly influenced by early childhood practitioners (Dunst, Trivette, Trivette et al., 2007; Dunst, Trivette, & Hamby, 2008). The same is the case for the relationships between FCPs and parents’ beliefs about their parenting confidence and competence (Dunst, Trivette, & Hamby, 2006, 2008). Results from these meta-analyses indicate that FCPs are indirectly related to parenting confidence and competence mediated by belief appraisals of control over program and practitioner responsiveness to family concerns and priorities. Based on the relationships among family-centered practices, self-efficacy beliefs, parenting confidence and competence beliefs, and parental psychological well-being described above, investigators have developed path analytic models for exploring the relationships among these variables where hypothesized relationships have been tested using structural equation modeling (Dunst & Trivette, 2009; Dunst, Hamby, & Brookfield, 2007; King, King, Rosenbaum, & Goffin, 1999; Raina et al., 2005; Trivette, Dunst, & Hamby, 2010). Close inspection of the results in these investigations indicates that FCPs have positive effects on self-efficacy beliefs where positive self-efficacy beliefs are associated with positive parenting belief appraisals (confidence and competence) and psychological well-being. Findings from structural equation modeling (SEM) studies by Dunst et al. (e.g., Dunst, Hamby et al., 2007, 2009, Dunst, Espe-Sherwindt, & Hamby, 2019; Dunst, Hamby, & Raab, 2019; Trivette et al., 2010) were used as the foundation for the studies described in this paper. We used the relationships in those studies to (1) test the fit of the relationships among the study measures to the model in Figs. 1 and 2 evaluate the pathways of influence in two structural equation modeling studies. Both studies were conducted in Spain where we determined whether the pattern of relationship found in structural modeling studies conducted mostly in North America could be replicated with families in another country. The study was conducted as part of a line of research and practice on factors facilitating practitioner use of FCPs and the outcomes of these practices (e.g., Costa, Serrano, Dunst, Mas Mestre, & Cañadas, 2017; Mas et al., 2018; Serrano, Mas, Canadas, & Gine, 2017). The study hypotheses, and prior research that are the basis for the pathways of influence, were: Family-centered practices would be directly related to both self-efficacy beliefs and parenting (confidence and competence) beliefs (Dempsey & Dunst, 2004; Dunst, Trivette, Trivette et al., 2007; Dunst, Hamby et al., 2007) and be indirectly related to parenting beliefs mediated by self-efficacy beliefs (2008, Dunst et al., 2006b). Self-efficacy beliefs of control over practitioner family-centered practices would be directly related to parenting beliefs (Dunst et al., 2006b) and be indirectly related to parental well-being mediated by parenting beliefs (Dunst, Hamby et al., 2007). Parenting confidence and competence beliefs would be directly related to parental well-being (Young, Karraker, & Lesley, 2006). Family-centered practices would be indirectly related to parental well-being mediated by both self-efficacy beliefs and parenting beliefs (Dunst & Trivette, 2009; Trivette et al., 2010). Results from the different sets of analyses were expected to determine if the pattern of relationships among the study variables could be replicated between countries and within countries. Our primary interests were the indirect effect of FCPs on parent well-being mediated by the different belief appraisals and the extent to which the direct and indirect effects of FCPs and efficacy beliefs on parent well-being could be replicated as the sine qua non of scientific research (Francis, 2012; Simons, 2014).","Study participants were recruited from different ECI centers in Spain: three ECI centers for Study 1 and six ECI centers for Study 2. Tables 1 and 2 show the characteristics of the study participants, their children and families, and the ECI the children and families received. Study 1. A convenience sample of 105 family members of children with identified disabilities and developmental delays receiving ECI and who responded to an invitation to complete a survey were the Study 1 participants. Most participants were mothers (70%) and married or living with a partner (87%). Sixty-five (65) percent were between 31 and 40 years of age. A majority of the participants (70%) completed at least a secondary education. The children’s mean chronological age was 3.73 years (Range = 1–8) where 73% were boys. The most frequently reported child diagnoses were developmental delays (30%), Autism Spectrum Disorders (18%), and intellectual disabilities (13%). Study 2. A convenience sample of 310 family members of children with identified disabilities and developmental delays receiving ECI and who responded to an invitation to complete a survey were the Study 2 participants. Most participants were the children’s mothers (80%) who were married or living with a partner (86%). Fifty-eight (58) percent were between 31 and 40 years of age. Eighty (80) percent of the respondents completed at least a secondary education. The children’s mean chronological age was 3.42 years (Range 1–6) and 66% were boys. The most frequently reported child diagnoses were speech and language disorders (23%), developmental delays (16%), and Autism Spectrum Disorders (13%).","A survey was used to obtain information needed to describe the study participants and obtain measures of the variables of interest. The survey included both investigator- developed items and scales previously validated in other studies (e.g., (Trivette & Dunst, 2004) including a study investigating the psychometric properties of Spanish versions of the scales (Mas et al., 2018). The latter was modeled after instruments used by other investigators interested in both the psychometric properties of the scale items the relationships among the variables that were the focus of investigation (e.g., Dunst, Trivette, & Hamby, 2006). Background information Investigator-developed questions were used to obtain information about the person completing the survey, and their child and family receiving ECI. This included child age, gender, and their primary diagnosis; the respondent’s age, gender, relationship to the child, level of education, employment status, and marital status; and family monthly income and number of adults and children in the nuclear family. The survey also included questions about the frequency of intervention and the number of months of receiving ECI. Each child’s primary caregiver completed the survey so that there was only one survey for each household. Family-centered practices We used the Spanish version (Mas et al., 2018) of the Family-Centered Practices Scale (Dunst & Trivette, 2003) to assess practitioners’ use of FCPs. This scale is a parent-completed instrument that assesses the degree to which professionals with whom they work employ FCPs. The scale includes six relational FCPs items and six participatory FCPs items. The relational FCPs items measure the interpersonal relationships between a family member and practitioner (e.g., “The practitioner really listens to my concerns and requests”). The participatory FCPs items measure practitioner use of capacity-building practices (e.g., “The practitioner helps me be an active part of obtaining desired resources and supports”). Participants indicate for each item the extent to which ECI staff interacts with and treats them and their families on a 5-point scale ranging from 1 = never to 5 = all the time. Coefficient alpha for the total scale score was .91 in Study 1 (relational subscale alpha = .81 and participatory subscale alpha = .83) and .89 in Study 2 (relational subscale alpha = .78 and participatory subscale alpha = .83). Self-efficacy beliefs Self-efficacy was measured in terms of the participants’ beliefs about control over the types of resources and supports provided by the ECI practitioners. Four items were used to measure the respondents’ self-efficacy beliefs where the participants were asked to indicate on a 5-point Likert scale the degree to which they agreed with different belief statements (e.g., “I am able to decide with staff which aspects I want to work for my children and family”). The items were obtained from Dunst et al. (2006a). Coefficient alpha was .72 in Study 1 and .75 in Study 2. Parenting competence and confidence beliefs Parenting beliefs were measured in terms of participants’ judgment of their ability to execute childrearing practices. Four items were used to measure the respondents’ parenting competence and confidence beliefs. Participants were asked to indicate on a 5-point Likert scale the degree to which they agreed with different parenting belief statements (e.g., “I am able to provide my children activities that help them learn”) (competence) (e.g., “I feel confident in my ability to help my son or daughter’s development”) (confidence). The items were obtained from Dunst et al. (2006a). Coefficient alpha was .64 in both Study 1 and 2. Psychological well-being Five items were used to measure the respondents’ psychological well-being. Participants were asked to indicate, on a 5-point Likert scale, the degree to which they agreed with different well-being statements. Three items asked about positive well-being (e.g., “I think that things are going well for my family”) and two items asked about negative well-being (e.g., “I feel anxious or upset”). The items were obtained from Dunst et al. (2006a). Coefficient alpha for positive well-being was .60 in Study 1 and .58 in Study 2. Coefficient alpha for negative well-being was .63 in Study 1 and .56 in Study 2. Procedure The researchers first contacted three ECI center directors in Study 1 and six ECI center directors in Study 2 to request their participation in this study and to obtain permission to contact the families of the children enrolled in this centers. ECI centers were selected via a criterion of convenience (i.e., centers that had already collaborated in previous or current projects developed by the same research group). After their approval, a meeting was held with the practitioners at the centers to explain the study, request their assistance in terms of explaining the importance of the study to the families, describe the aim and expected results, and to provide them the study measures to distribute to families receiving ECI on this services. The information provided to the parents included the following: A letter describing the study, an informed consent letter, and the survey described above. They were also provided a telephone number and email address they could use to contact the investigators if they had any questions or concerns about the study or if they wanted to withdraw their consent and participation. It was requested that only one family member (father, mother or other primary caregivers) complete the FCPs scale and other survey measures within 15 days and return completed survey and informed consent letter to the center in a sealed envelope to ensure confidentiality. Once the center had collected the sealed envelopes, the directors sent the completed scales and surveys to the researchers who prepared the responses for subsequent analysis. Method of analysis ~~~~~~~~~~~~~~~~~~ Structural equation modeling (SEM) was used to evaluate the effects of family-centered practices on parental well-being mediated by self-efficacy and parenting belief appraisals (Jöreskog & Sörbom, 2014). SEMs were used in both Study 1 and Study 2 with both measured and latent variables to identify the best fitting model. The measured variables were the sums of the scores for the (a) two family-centered practices measures (relational and participatory), (b) two parenting belief measures (confidence and competence), and (c) the two well-being measures (positive and negative). The latent variables included the two subscale measures for family-centered practices, parenting beliefs, and parental well- being. SEMs for all combinations of measured and latent variables were run. The fit of the models to the pattern of relationships among the variables in the models was evaluated using the root mean square error of approximation (RMSEA), standardized root mean residual (SRMR), comparative index (CFI), incremental fit index (IFI), and normed fit index (NFI). The closer RMSEA and SRMR are to zero, and the closer CFI, IFI, and NFI are to one, the better the fit of the model to the data (Hu & Bentler, 1995). A preponderance of indices reaching recommended levels was used as the criterion for establishing an acceptable model fit (Hooper, Coughlan, & Mullen, 2008). Both SEMs included tests for the direct, indirect, and total effects of the study variables on parental well-being. Effects decomposition (Kline, 2005) was used to identify the pathways of influence in the models. Standardized coefficients between contiguous variables were used to test for direct effects and the products of these coefficients were used to test for indirect effects (Bollen, 1987; Sobel, 1988). Total effects were determined by the sums of direct and indirect effects.","Table 3 shows the correlations among both the Study 1 and Study 2 measures. All of the correlation coefficients in both sets of data are statistically significant. The directions of the effects are also all as expected. Both family-centered practices measures were positively correlated with the three belief measures (self-efficacy beliefs, parenting competence beliefs, and parenting confidence beliefs) and positive parental well-being and negatively correlated with negative parental well-being. The pattern of relationships between the three belief measures and parental well-being was much the same. The correlation matrices in both studies were imputed in LISREL for the SEMs. Study 1 analyses ~~~~~~~~~~~~~~~~ The best fitting SEM is shown in Fig. 2. RMSEA was .00 (90% CI = .00–.10), SRMR was .03, CFI was .99, IFI was .99, and NFI was .95. All five indices met generally agreed upon thresholds and indicate a good fit of the model to the relationships among the variables in the model. Three of the five structural coefficients for the direct effects were statistically significant. The patterns of relationships among the variables in the model are generally consistent with the study hypotheses as evidenced by effects decomposition. Table 4 shows the effects decomposition results for the direct, indirect, and total effects of family-centered practices, self-efficacy beliefs, and parenting beliefs on parental well-being. As expected, FCPs were directly related to both self-efficacy beliefs and parenting beliefs. Family-centered practices were also indirectly related to parental well-being mediated by both self-efficacy beliefs and parenting beliefs. Self-efficacy beliefs were directly related to parenting beliefs, and indirectly related to parental well-being mediated by parenting beliefs. Parenting beliefs were directly related to parental well-being. Study 2 analyses ~~~~~~~~~~~~~~~~ Fig. 3 shows the SEM for the fit of the Study 1 model to the data. RMSEA was .11 (90% CI = .08–.14), SRMR was .05, CFI was .97, IFI was .97, and NFI was .96. The indices indicate a reasonably good fit of the model to the data. All of the path coefficients in the model were statistically significant except for the direct effect between self-efficacy beliefs and parental well-being. The patterns of relationships were consistent with expectations. The complete set of direct, indirect, and total effects are shown in Table 4. Family- centered practices were directly related to both self-efficacy beliefs and parenting beliefs, and indirectly related to parental well-being mediated by both the self-efficacy and parenting belief measures. Self-efficacy beliefs were directly related to parenting beliefs, and indirectly related to parental well-being mediated by parenting beliefs. Parenting beliefs were directly related to parental well-being. Model comparisons ~~~~~~~~~~~~~~~~~ A comparison of the results from the two sets of analyses found that the indices for the fit of the models to the data a were more similar than different, and that the indirect effects of FCPs on parental well-being mediated by both kinds of belief measures (self- efficacy beliefs and parenting beliefs) were almost identical (β = .37, p < .001, and β = .38, p < .001, respectively in Studies 1 and 2). Closer inspection of the pathways of influence, however, indicates a few differences in the two sets of analyses. In Study 1, the influence of FCPs on parenting beliefs was direct (β = .41, p < .01) but only marginally indirectly related to parenting beliefs mediated by self-efficacy beliefs (β = .73 × .21 = .15, p < .10). In contrast, the relationship between FCPs and parenting beliefs in Study 2 was both direct (β = .24, p < .005) and indirect mediated by self- efficacy beliefs (β = .70 × .54 = .38, p < .001). There were also differences in the pathway of influence of family-centered practices on parental well-being. Whereas the indirect effect of family-centered practices on parental well-being mediated by parenting beliefs was similar in both Study 1 (β = .41 × .58 = .24, p < .02), and Study 2 (β =.24 × .78 = .19, p < .005), the indirect effect of family-centered practices on parental well- being mediated by self-efficacy beliefs through parenting beliefs was significant in Study 2 (β = .70 × .54 x .78 = .29, p < .001) but not in Study 1 (β = .73 × .21 x .58 = .09, p > .10). The patterns of relationships among the variables in Study 1 indicate that the influence of FCPs on parental well-being is primarily through parenting beliefs. In Study 2, the patterns of relationships among the variables indicate two pathways of influence between FCPs and parental well-being; one through parenting beliefs, and one through self- efficacy beliefs and parenting beliefs. Despite the few differences in the sizes of effect for the relationships among the variables in the two SEMs, the results were more similar than different as evidenced by the fit indices and effects decomposition results.","The studies described in this paper examined the manner in which practitioners’ use of FCPs was related to parent psychological well-being. More specifically, we evaluated the relationships between family-centered relational and participatory practices, self- efficacy beliefs, parenting confidence and competence beliefs, and families’ positive and negative psychological well-being. Structural equation modeling was used in both studies to test the fit of a model based on results from previous research on the relationships among the variables in the model (e.g., King et al., 1999; Thompson et al., 1997; Trivette et al., 2010). The findings for families in Spain were the same or very similar to those reported in other studies mostly in North America (e.g., Dunst, Hamby et al., 2007; King et al., 1999; Thompson et al., 1997). The Study 1 model was used to evaluate the fit of the data to the SEM in Study 2 for replication purposes. The results provided support for both the hypothesized direct and indirect effects shown in both Fig. 3 and Table 5. The findings showed that how practitioners interact with families is directly related to parents’ self-efficacy beliefs and parenting beliefs and indirectly related to parenting beliefs mediated by self-efficacy beliefs (Bailey, Nelson, Hebbler, & Spiker, 2007; King et al., 1999; Trivette et al., 2010). More specifically, our results showed that the more parents judged the practices of the ECI practitioners as family-centered, the more control they reported over the supports and assistance received from the practitioners with whom they worked in a manner identical to that found in other studies (e.g., Dunst, Trivette, Trivette et al., 2007). Likewise, the more parents perceived control over FCPs, the more competent and confident was their parenting beliefs (Dunst, Trivette, Boyd, & Brookfield, 1994; Dunst & Dempsey, 2007; Dunst, Hamby et al., 2007). Positive parenting beliefs, in turn, were related to parents' psychological well-being in a manner similar to what has been found in other studies (Dunst & Trivette, 2009; Dunst, Hamby et al., 2007; Dunst, Trivette & Raab, 2013; Trivette et al., 2010). That is, the pathways of influence found in our study illustrate how the effects of family-centered practices could be traced to more positive and less negative parent well-being in a manner similar to that found in other studies (Dunst & Trivette, 2009; Dunst, Hamby et al., 2007; King et al., 1999). The indirect effects found in Study 2 are consistent with findings in previous research where the relationships between FCPs and child, parent, and family outcomes have been found to be mediated by different kinds of belief appraisals in a manner consistent with self- efficacy theory (Bandura, 1997). Parental attributions -like those investigated in the present research-have been found in other studies to mediate the relationship between FCPs and parent well-being (King et al., 1999), child well-being (Dunst & Trivette, 2009), parent-child interactions (Dunst, Espe-Sherwindt et al., 2019; Dunst, Hamby et al., 2019), and child behavior and development (Graves & Shelton, 2007; Trivette et al., 2010). The research described in this paper adds to this evidence base by demonstrating how the complex relationships among the variables in the study could be identified as evidenced by the direct, indirect, and total effects of the predictor variables on different mediator and outcome variables. The manner in which we found FCPs to be directly and indirectly related to parent psychological well-being illustrate how practitioner help-giving practices are related to belief appraisals and in turn are related to health-related outcomes similar to what has been found in other studies (Bailey et al., 2007; Dunst et al., 2008). This was expected since the use of both relational and participatory FCPs has been found to have capacity-building characteristics and consequences (Dunst & Espe- Sherwindt, 2016). The characteristics include, but are not limited to, practitioner beliefs about existing family member strengths and the capacity to become more competent; informed family choice and decision-making; active family member involvement in achieving desired goals and outcomes; family member use of existing capabilities; and the development of new skills for obtaining desired resources and supports and achieving desired goals (Dunst & Espe-Sherwindt, 2016). The consequences include, but are not limited to, parents’ beliefs about their ability to exercise control over important life events (e.g., Skinner & Greene, 2008), including parenting beliefs about their childrearing practices (Coleman, 1999; Newland, 2015). The results add to “the knowledge base regarding the role active help-receiver participation plays in people achieving desired life circumstances that in turn produce positive behavioral consequences in other domains of functioning” (Dunst et al., 2006a, p. 48). We therefore conclude that the results found in our studies provide evidence for the generalized effects of FCPs in samples in Spain in a manner similar to that found in other studies. The belief appraisals that were the focus of this investigation have been used in other studies as measures of individual empowerment (Woodall, Raine, South, & Warwick-Booth, 2010) and psychological empowerment (Zimmerman & Rappaport, 1988). According to Singh (1995), individual empowerment is defined as \"a process by which families access knowledge, skills, and resources that enable them to gain positive control over their own lives as well as improve the quality of their lifestyles” (p. 13, emphasis added). Self-efficacy beliefs have been used as a measure of control over practitioner help-giving practices (Trivette, Dunst, Boyd, & Hamby, 1995) and parenting beliefs have been used as a measure of control over executing parenting roles (Paczkowski & Baker, 2007). An empowerment perspective of efficacy beliefs helps explain how experiences afforded families (e.g., practitioner use of FCPs) “equip people with the requisite knowledge, skills, and resilient self-beliefs of efficacy to alter aspects of their lives over which they can exercise some control (Ozer & Bandura, 1990, p. 472, emphasis added). These types of control belief appraisals influence families’ abilities to cope, deal with, and affect life circumstances (Bandura, Caprara, Barbaranelli, Regalia, & Scabini, 2011; Skinner, 1995). Placed within the context of empowerment theory, FCPs include capacity-building experiences that provide family members opportunities to use existing abilities and acquire new abilities in ways that positively affect their beliefs about control over important life events in ways that influence health-related outcomes. As noted by both empowerment (Rappaport, 1987) and self-efficacy (Bandura, 1997) scholars, belief appraisals function as mechanisms for explaining how competency-enhancing experiences afforded people are manifested in terms of improved psychological functioning (Diener & Biswas-Diener, 2005; Woodall et al., 2010). According to Tensky (2005), this necessitates a paradigm shift in terms of how we conceptualize and operationalize the ways in which professionals go about their work with help seekers. Limitations ~~~~~~~~~~~ There are at least three limitations that must be considered as part of the data interpretation. The first limitation has to do with the use of self-report measurement instruments, which may account for the nature of the relationships among the variables in both studies, and the fact that the data are cross-sectional and not longitudinal. This limitation concerns the degree of confidence for asserting causal relationships between the study measures (Asher, 1983; Kenny, 2004). This limitation is partly mitigated by the fact that the pattern of results in both studies are nearly identical to those reported in studies where there was time precedence between the belief appraisals and well-being measures (Dunst & Trivette, 2009; Dunst, Hamby et al., 2007). A second limitation is the failure to include other measures in the analyses that might better explain the nature of the relationships among the variables in the SEMs (e.g., parenting styles of interaction; see especially Trivette et al., 2010). As noted by Tomarken and Waller (2005), such omissions may present a misleading depiction of how FCPs are directly and indirectly related to other variables of interest (see especially Trivette et al., 2010). A third limitation relates to the fact that only one model was the focus of investigation where the model was informed by results from previous investigations. It could be the case that there are other models that better explain the relationships among the variables in the model. However, as noted by MacCallum (1995), “In a strictly confirmatory [SEM] strategy, the researcher constructs one model of interest and evaluates that model by fitting it to appropriate data. If the model yields interpretable parameter estimates and fits the data well, it is supported and considered a plausible model (p. 31, emphasis added).","The pattern of results in the SEMs provided empirical support for the contention that the way in which help is provided can either enhance or impede the intended outcomes of that help (Dunst, Trivette, & Hamby, 2007). The findings showed how the effects of FCPs on parent psychological well-being are indirect and mediated by self-efficacy beliefs and parenting competence and confidence beliefs. These results have a number of implications for practice. The first implication has to do with an understanding of and the need to be clear about how FCPs are related to health-related outcomes. Parental well-being is recognized as an important family outcome of ECI (Bailey et al., 1998; Krauss & Jacobs, 1990). The effects of FCPs on parent psychological well-being, however, are primarily indirect and not direct, and it is important that practitioners have a clear understanding of how FCPs influence outcomes of interest, including, but not limited to, health-related outcomes (see e.g., Dunst et al., 2008). The second implication has to do with the importance of using both relational and participatory FCPs and especially the use of participatory help-giving practices, if FCPs are going to have capacity-building consequences. The latter is the case because practices that actively involve parents in ECI are more likely to promote a sense of competence and confidence as hypothesized by Bandura (1997). Accordingly, ECI programs and practitioners that focus on the empowerment of families will need to place primary emphasis on using participatory practices that strengthen existing capabilities and promote the acquisition of new family member competencies (Dunst, 2010). The third implication is not to assume that ECI programs and practitioners that claim to use FCPs, in fact, use these practices on a routine basis. Research has shown that ECI programs, organizations, and practitioners face many challenges and obstacles as part of adopting and using FCPs, and especially participatory practices (Dempsey & Keen, 2017; Dunst & Espe-Sherwindt, 2017). It is therefore hardly surprising that Dempsey and Keen (2017) noted that “effective implementation is the next major challenge for family-centered practices” (p. 65). Effective implementation includes, but is not limited to, the use of evidence-based professional development practices. Findings from a recently published study indicate that professional development specialist use of capacity-building professional development practices was related to practitioners’ use of capacity-building FCPs (Dunst, Espe-Sherwindt et al., 2019; Dunst, Hamby et al., 2019). This demonstrates an empirical relationship between evidence-based implementation practices and evidence-based FCPs.","This research has been carried out through funds from the Ministry of Universities and Research of the Department of Enterprise and Knowledge of the Generalitat de Catalunya."],["On average, marriage tends to lead to temporary increases in life satisfaction, which quickly return to pre-marital levels. This general pattern, however, does not consider the personality of individuals entering into marriage. We examine whether following marriage pre-marital personality predicts different changes to life satisfaction in a sample of initially single German adults (N = 2015), completing life satisfaction measures and indicating their marital status yearly for 8 years (during which 468 married). We find that conscientious women experience greater life satisfaction following marriage than less conscientious women. Our data also indicate that introverted women and extraverted men experience longer-term life satisfaction benefits following marriage. Our results refute the claim of limited life satisfaction effects from marriage and caution against relying on average effects when examining the influence of life events on well-being. --------------------------------------------------------------------------------","Considerable research has aimed at testing whether marriage leads to increases in life satisfaction. Married individuals robustly have higher average levels of life satisfaction than non-married individuals (Haring-Hidore, Stock, Okun, & Witter, 1985), but this relation is partially explained through social selection effects, whereby those with higher life satisfaction are more likely to marry (Mastekaasa, 1992). Nevertheless research that controls for selection effects suggests that any life satisfaction benefits of marriage are at best transitory. There are short-term life satisfaction increases following marriage but life satisfaction returns fairly rapidly to pre-marital levels (Yap, Anusic, & Lucas, 2012). However, this general pattern of results is unlikely to be true for everyone, with some people being more likely to experience greater life satisfaction benefits following marriage, whilst others may find the experience less beneficial. Here we explore whether a person's pre-marital personality predicts life satisfaction change following marriage. Personality represents basic individual tendencies and, as conceptualized by the Five Factor Model, (FFM; McCrae & Costa, 2008), comprises agreeableness, conscientiousness, extraversion, neuroticism, and openness-to-experience. Individuals can infer and express accurately what these basic tendencies are from their own behaviors and experiences (McCrae & Costa, 2008). The FFM traits relate to an individual's life satisfaction (Steel, Schmidt, & Shultz, 2008), which may be through a direct relation, capturing an individual's predisposition to experience positive or negative emotions (as with the positive or negative affective components of extraversion or neuroticism). Alternatively, the relationship between personality and life satisfaction may be indirect (as with agreeableness, conscientiousness, and openness) through orientating individuals toward positive situations (McCrae & Costa, 1991). However, evidence is emerging for a third pathway, in that there are differences in how personality influences response to life events. Specifically, personality has been shown to predict how life satisfaction is influenced following adverse life events such as disability (Boyce & Wood, 2011) and income loss (Boyce, Wood, & Ferguson, in press), as well as protecting against depression during widowhood (Pai & Carr, 2010). Importantly such studies have utilized personality measures before the events took place, thus preventing confounding any effects with the possibility that personality traits develop in response to these events (Boyce, Wood, Daly, & Sedikides, 2015). Only two studies have assesed whether personality moderates the extent to which individuals' life satisfaction changes following marriage (Anusic, Yap, & Lucas, 2014; Yap et al., 2012). However, owing potentially to limited statistical power Yap et al. (2012) obtained null effects, whilst Anusic et al. (2014) did not utilize personality traits measured before marriage. Since research in this area is limited we hypothesize that any of the FFM personality traits may be important. In accordance with our exploratory approach the literature on relationship satisfaction suggests an important role for agreeableness, conscientiousness, extraversion, and neuroticism (Malouff, Thorsteinsson, Schutte, Bhullar, & Rooke, 2010). Personality traits tend to influence relationship satisfaction via ongoing relationship dynamics (Solomon & Jackson, 2014), which may ultimately lead to the dissolution of the relationship (Roberts, Kuncel, Shiner, Caspi, & Goldberg, 2007). Since the attainment of a satisfying relationship is a near universal goal (Roberts & Robins, 2000) factors that enhance the quality of a relationship are also likely to influence life satisfaction. Given there are personality differences across men and women with regard to relationship satisfaction (Solomon & Jackson, 2014) we also explore personality differences across men and women. Since we make no specific predictions we consider statistical corrections for multiple comparisons (Nakagawa, 2004).","We used the German Socio-Economic Panel (SOEP) study, an ongoing longitudinal study of German households. The SOEP began in 1984 with a sample of adult members from private households in West Germany, initially over-representing immigrants. Since 1984, the SOEP has expanded to include East Germany and various sub-samples to ensure a broadly representative sample of the entire German population (Wagner, Frick, & Schupp, 2007). We focused on SOEP participants, regardless of their origin in the sample, who answered personality questions in 2005 and were single. Participants also responded to questions about their life satisfaction in every year from 2005 to 2012 and we ensured that their marital status was recorded in each of these years. We then observed the marital status across this period to determine whether individuals had married. Participants' current marital status is recorded in the SOEP as either married (living together with spouse), married (but permanently separated), single, divorced, or widowed. We concentrate only on those individuals that are initially single, got married and stayed married (remaining living together with spouse) in the study period. All individuals that marry in our sample therefore marry for the first time. We included a control group of individuals who remained single throughout the study period such that we could account for life satisfaction selection effects and to ensure life satisfaction changes were the result of marriage rather than some national event that affected the entire sample. Our final sample consisted of 2015 (986 females, 1029 males) participants of which 1547 remained single throughout the study period and 468 (248 females, 220 males) participants married for the first time at some point in the study and remained married. In 2005, when all individuals were single, age ranged from 17 to 88 (M = 30.99, SD = 12.53). Life satisfaction Life satisfaction was measured with one item each year for all 8 years. Participants responded to the question “How satisfied are you with your life, all things considered?” from 0 (completely dissatisfied) to 10 (completely satisfied). Participants' responses were standardized (M = 0, SD = 1) across the sample. Single item scales, although typical for large data sets, can have a low reliability resulting in an underestimation of the true effect size (inflating Type II, but not Type I, error). Lucas and Donnellan (2007) estimate the unstable state/error component of life satisfaction in the SOEP and show that approximately 33% of the variance in responses can be attributed to the unstable state/error component over a 1 year period. They infer that the life satisfaction has an acceptable reliability of at least r = .67. Although reliability diminishes with an increased time interval the reliability is approximately r = .45 across 7 years. This is higher than normally observed for single item measures. Big Five personality measures A 15-item (3 per trait) shortened version of the Big Five Inventory (Benet- Martínez & John, 1998) was administered in 2005. This version was developed specifically for use in the SOEP, where there is limited space for survey questions (Gerlitz & Schupp, 2005). Participants responded to 15 items (1 = “does not apply to me at all”, 7 = “applies to me perfectly” scale), with three items assessing each of the FFM domains of agreeableness (e.g., “has a forgiving nature”), conscientiousness (e.g., “does a thorough job”), extraversion (e.g., “is communicative, talkative”), neuroticism (e.g., “worries a lot”), and openness (e.g., “has an active imagination”). Across each personality dimension all three scores were aggregated after appropriate reverse coding and then standardized (M = 0, SD = 1). Life satisfaction and personality scores for the entire SOEP sample, as well as for each marriage category and by an individual's age group, are found in Tables A1 and A2 respectively in the Appendix A. These scores are broadly comparable to SOEP sample wide scores. The SOEP scale has comparable psychometric properties to longer FFM scales. For example, the short-item scale produces a robust five factor structure across all age groups (Lang, John, Lüdtke, Schupp, & Wagner, 2011). Donnellan and Lucas (2008) demonstrated that each of the scales in the SOEP correlates highly (r > .88) with the corresponding scale in the full Big Five Inventory. Although Lang (2005) illustrates that the retest reliability across 6 weeks is acceptable (r > .75) this reliability measure is insufficient as our study takes place over 7 years and may not apply to our specific marriage sub- sample. Since the shortened Big Five Inventory was administered 4 years later in the SOEP we estimate the retest reliability in our sample. It was at least r = .52 across this time period and similar for those that married and those that did not (see Table A3). These values are comparable to longer scales over this time frame (r = .55; see Roberts & DelVecchio, 2000). Table A4 shows the correlations between each of the FFM personality traits and life satisfaction in our sample. Neuroticism has a strong negative relationship with life satisfaction, whereas the remaining traits are less strongly positively related to life satisfaction, conforming with previous research (Steel et al., 2008). Covariates Marriage is correlated with a number other factors which may be associated with life satisfaction. We control for an individual's age, the presence of children in the marriage, education level, and an individual's satisfaction with family life. We also include time-period dummies to allow for time-period specific differences in life satisfaction. Since age and education also correlate with personality (Srivastava, John, Gosling, & Potter, 2003) any personality interactions may be driven by these factors. For example, older individuals (or analogously those more highly educated) may have a higher life satisfaction during marriage than those younger. Since age (or education) is also likely to be associated with personality, not appropriately controlling for the interaction of these variables with marriage may lead to a spurious interaction between personality and marriage. Thus we include interactions of both age and education (recorded in 2005) with our marriage variables. We dealt with missing data in education (15.9%) and family satisfaction (2.2%) using multiple imputation. We used multiple imputation chained equations (MICE; White, Royston, & Wood, 2011) using predictive mean matching and obtained 5 imputations (based on five sequential iterations using MICE). We also imputed the missing education–marriage interaction terms to ensure these variables had the correct means and covariances.","To examine whether personality predicts life satisfaction differences in how individuals respond to marriage we carried out an interaction analysis within a multilevel framework. We analyze the Level 1 effect of marriage on life satisfaction (LS) across all time points (t) from 2006 to 2012 for men and women separately. Since we were interested in life satisfaction over the course of the marriage we coded individuals at each time-point according to the number of years they had been married up to that time-point. Since the years before marriage are often associated with benefits to life satisfaction we include dummy variables to indicate that an individual will get married in the next year or alternatively that they will get married at some point in the study. At a given time-point participants were classified as either never experiencing marriage throughout the study, not yet married but would at some point during the study (Mt + > 1), experienced marriage in the following year (Mt + 1), or married for 1 to 7 years (Myrs). Our analysis allowed us to establish, and control for, any life satisfaction selection effects, and also determine the effect on life satisfaction at different years of marriage. To determine non-linear effects we included the square and cube of the number of years that the participant had been married (Myrs2, Myrs3). Person-specific slopes and intercept errors are captured by the σ terms and ε captures the overall model error. By controlling for life satisfaction in 2005, γ01 and γ02 are interpretable as marriage selection effects, and γ03, γ04, and γ05 signify changes in life satisfaction by year of marriage. The coefficients γ13, γ14, and γ15 represent the personality-marriage interaction effects. Since our analysis is largely exploratory we consider our results in light of possible Type 1 error through multiple comparisons (Nakagawa, 2004).","To test whether there is an interaction between personality and marriage in predicting life satisfaction we carried out multilevel regressions separately for both women and men, initially including no controls. Table 1 Regression 1 provides the results for women. The coefficients on the marriage main effect variables suggest that on average women that will marry during the study are 0.13 SD (coefficient on Mt + > 1) higher in life satisfaction than those who don't marry. In the year directly preceding marriage women on average have life satisfaction levels 0.29 SD (coefficient on Mt + 1) higher than those who never marry. The first year of marriage is then associated with a life satisfaction level of 0.21 SD, with each additional year of marriage changing life satisfaction according to 0.30 ∗ Years Married − 0.10 ∗ Years Married2 + 0.01 ∗ Years Married3. Thus the effect of marriage on life satisfaction is initially positive but eventually reduces. Regression 1 illustrates that the effect of marriage on life satisfaction is also dependent upon pre- marriage personality. There are significant interaction terms for both conscientiousness and extraversion. This suggests that women with basic underlying tendencies (see McCrae & Costa, 2008) that result in them endorsing behaviors reflective of conscientiousness or low extraversion experience higher life satisfaction during marriage. This is illustrated in Fig. 1. In the left-hand panel we observe that women who score themselves moderately high on conscientiousness (+ 1 SD) experience sustained life satisfaction benefits, whereas women who score themselves as moderately low on conscientiousness (− 1 SD) quickly experience falls in life satisfaction. After some years the life satisfaction levels of those moderately low in conscientious are similar to those that remained single throughout the study. The middle panel in Fig. 1 shows the effect of marriage on life satisfaction for women who score themselves moderately low (− 1 SD) and moderately high (+ 1 SD) on extraversion. There is significance only on the linear interaction but the curvature remains owing to the main effect coefficients and the interaction only changing the trajectory of the curvature. Initially there are no satisfaction differences by extraversion. However, after a few years of marriage women that endorse behaviors reflective of extraversion begin to experience reductions in life satisfaction, whilst those that don't endorse behaviors reflective of extraversion (i.e. intraversion) maintain their level of life satisfaction. Thus women who score low on extraversion appear to experience long-term life satisfaction benefits following marriage. We note the sudden spike in life satisfaction in years 6 and 7 but we suggest caution since only a small number of participants in our sample experienced 6 or 7 years of marriage. Table 1 Regression 3 provides the results for men. On average men that marry during the study experience life satisfaction increases in the year directly preceding the marriage, where life satisfaction rises to 0.18 SD. The first year of marriage is then associated with a life satisfaction of 0.17 SD higher than those who remain single, with each additional year of marriage changing life satisfaction according to 0.25 ∗ Years Married − 0.09 ∗ Years Married2 + 0.01 ∗ Years Married3. The effect of marriage on life satisfaction is on average positive but returns to pre-marital levels of life satisfaction quickly. Regression 3 suggests, however, that men who endorse behaviors reflective of extraversion experience higher life satisfaction during marriage. The right-hand panel in Fig. 1 shows the effect of marriage on life satisfaction for men that score themselves moderately low (− 1 SD) and moderately high (+ 1 SD) on extraversion. Whilst all men experience a pre- marital increase in their life satisfaction, men that are extraverted seem to experience longer-term benefits to their life satisfaction during marriage. Introverted men, however, experience significant drops in their life satisfaction that result in them being approximately 0.20 SD lower in life satisfaction than those who never marry. Due to the possibility of Type 1 errors owing to multiple comparisons we re-evaluate our results after making a Bonferroni-type correction (p = 0.05/α, where α represents the number of comparisons made which is 10 here). Only the conscientious interaction in women survives this correction (joint significance on both conscientiousness interaction terms; p = .004). Although Bonferroni-type corrections minimize the possibility of Type 1 errors they have been criticized for increasing the likelihood of Type 2 errors (Nakagawa, 2004). Thus we suggest that the other interactions, rather than being rejected, should be simply treated with caution. Further a number of other factors may correlate not only with marriage and life satisfaction but also with personality. We account for these by including additional controls including age, the presence of children, education level, and an individual's satisfaction with family life, as well as the interaction of the individual's pre-marital age and education with the marriage variables up to the quadratic term of years spent married. The conscientiousness interaction effects are still evident, whereas for extraversion these effects are significant only at the 10% level. There is now a linear effect on neuroticism. These changes are driven by the inclusion of family satisfaction, which is strongly correlated with life satisfaction. This suggests further caution for the extraversion result in both men and women.","Although individuals may experience initial life satisfaction increases following marriage these effects on average return quickly to pre-marital levels. However, we show that an individual's reaction depends on their pre-marital personality. Specifically, women who reported being conscientious experienced sustained increases in their life satisfaction following marriage whereas those less conscientious experienced small transitory life satisfaction increases. Such a result might be explained by the tendency for conscientious individuals to place more value on relationship goals (Roberts & Robins, 2000) and therefore conscientious individuals may strive harder to ensure success (Duckworth, Peterson, Matthews, & Kelly, 2007). This result is consistent with conscientious individuals being more satisfied with their relationships (Malouff et al., 2010). This conscientiousness effect, however, was not found in men. Although we expected some differences between men and women it is not clear why this was the case and we speculatively suggest that this could be due to differences in how men and women value life goals (Roberts & Robins, 2000). We also found effects that differed across men and women for extraversion, with introverted women but extraverted men experiencing long-term benefits to their life satisfaction. Extraversion generally predicts enhanced relationship satisfaction (Solomon & Jackson, 2014) and although it is not clear why we observed inconsistent effects it has been suggested that the importance of extraversion for relationship satisfaction may be cultural (Malouff et al., 2010). Some caution is, however, recommended with our results on extraversion since the effects were dependent on the inclusion of certain controls and further the effect did not pass the more stringent significance level to account for multiple comparisons. One reason for our limited power to detect some of the effects might be due to our scales for personality and life satisfaction being shorter than ideal. This may have also resulted in our effects being under-estimated. Our research nevertheless demonstrated a test–retest stability in a short-item personality scale over 4 years that was comparable to longer scales, adding to the literature on personality stability whilst being consistent with the literature on personality change (Boyce et al., 2015). Our exploratory approach to understanding how personality moderates the influence of marriage on life satisfaction was an attempt to establish initial research in this area. Now that the basic relationship has been established we hope that this opens up possibilities for future research to explore mechanistic pathways. We suggest that our results might be driven by specific personality types valuing or not valuing certain features of their new environment, such as different social opportunities that arise following marriage. Alternatively, life satisfaction may increase following marriage not due to the direct effect of marriage per se but via the indirect effect marriage has in protecting an individual when they encounter life stressors. Personality may both increase the likelihood of other life stressors occurring during marriage and/or moderate the impact of such life stressors. It is also possible that partner personality may have an important effect in explaining why some marriages yield more satisfaction than others (Solomon & Jackson, 2014). Since most of our sample did not include partners from the same marriage we were unable to examine the influence of partner personality. We recognize this limitation and future research should therefore explore the role of partner personality. Although more work is needed in understanding and testing precise mechanisms behind our results, our research is the first to demonstrate that personality moderates the effect of marriage on life satisfaction and adds to a growing literature illustrating the importance of personality traits for generating higher or lower well-being following commonly occurring life events."],["The use of image-based testing to assess individual differences has increased substantially in recent years, with proponents arguing that they offer a more engaging alternative to text-based psychometric tests. Yet research examining the validity of these tests is near to non-existent. Traditional image-based formats have been little more than an adaptation of self-reports, with images replacing questions but not response options. The current study develops a novel image-based creativity measure, where images replace conventional response scales, and scores on the measures are obtained using a linear regression scoring algorithm to predict three self-reported creativity measures. Using sequential forward selection on a set of 77 image-based items, an optimal solution of 14 items that were valid predictors of self-reported creativity scores were identified. The image-based measure had good test-retest reliability. Implications are discussed in terms of the usefulness of image-based testing for practitioners seeking engaging and short test formats. --------------------------------------------------------------------------------","The assessment of individual differences in psychological traits, such as personality, intelligence, and creativity, stretches back more than a century (Chamorro-Premuzic, 2007). The most common way of measuring differences between people is through psychometric tests (Ahmetoglu & Chamorro-Premuzic, 2013). Psychometric tests are used extensively in settings from selection (Rothstein & Goffin, 2006) to psychiatric diagnosis (Gilbody, Richards, Brealey, & Hewitt, 2007), and consumer profiling (Matz, Gladstone, & Stillwell, 2016). Despite their widespread use, psychometric tests are criticised for their inability to engage the test taker (Krosnick, 1991), the ease of faking responses (Morgenson et al., 2007), and adverse impact (Hough, Oswald, & Ployhart, 2001). Perhaps in response to these criticisms, and fuelled by technological advances, recent years have seen mounting interest in more engaging forms of assessment (Attali & Arieli-Attali, 2015), including gamification (Chamorro-Premuzic & Steinmetz, 2013; Landers & Callan, 2011; Reeves & Read, 2013) and social media analytics (Kosinski, Matz, & Gosling, 2015; Pennebaker, 2011). However, innovative assessment tools often serve entertainment purposes, with little indication to their validity (Naglieri et al., 2004). The increase in the quantity of these instruments has not been synonymous with an increase in research into their quality, that is, their reliability and validity. Indeed, the desire to use innovative assessment by professionals has outpaced the peer-reviewed literature (e.g., Roth, Bobko, Van Iddekinge, & Thatcher, 2013). This gap between research and practise is problematic if tests are used to make hiring decisions or provide clinical diagnosis. Consequently, developing scientific evidence for the validity and utility of image-based tests is critical, not only from an academic, but also an applied perspective. The current study takes a step in this direction. Specifically, an image-based creativity assessment and a predictive scoring algorithm are developed. The test-retest reliability, as well as its concurrent validity in relation to three text-based, self-report creativity measures are assessed, so that practitioners may better understand how such image-based tests compare to traditional tests. Advantages of image-based formats ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ One of the most common innovations in psychological assessment formats has been to replace item questions with visual representations, thereby increasing user engagement (Barrett & Ebbeling, 2003; Downes-Le Guin, Baker, Mechling, & Ruylea, 2012; Hamari, Koivisto, & Sarsa, 2014; Lugtigheid & Rathod, 2005). Beyond engagement, image-based formats could provide theoretical and practical advantages over text-based psychometric tests. First, they may be more suitable for culturally and linguistically diverse test takers, and remove misunderstanding of text items (Paunonen, Jackson, & Keinonen, 1990). Second, responding to image-based items may require less attention, reducing test taker fatigue. Finally, image stimuli evoke stronger preferences in respondents than verbal stimuli, providing for reduced length of image-based tests (Lugtigheid & Rathod, 2005; Meissner & Rothermund, 2015). Past research on image-based tests ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Despite being innovative, image-based formats in assessment are not new. Geist's (1959) Pictorial Interest Inventory pictures a person engaged in three activities, of which respondents pick the most appealing one. More recent image-based tests adapt text-based personality measures such that the question is replaced with an image: The Nonverbal Personality Questionnaire (Paunonen et al., 1990) measures Murray's (1938) psychological needs such that participants report the likelihood that they would engage in visually displayed behaviours. A version of the test measuring the Big Five also exists (Paunonen, Ashton, & Jackson, 2001). These adaptations of verbal personality tests have gained support in the academic literature for their internal reliabilities and validities (Hong, Paunonen, & Slade, 2008; Moore, Schermer, Paunonen, & Vernon, 2010; Paunonen, 2003; Paunonen, Jackson, Trzebinski, & Forsterling, 1992; Paunonen, Zeidner, Engvik, Oosterveld, & Maliphant, 2000). However, research examining the validity of image-based tests is scarce, and their use is mostly limited to special populations, such as children or illiterates. In addition, the use of response scales and scoring methodologies developed for verbal formats is not ideal: by using images to replace the question stem, questions are limited to those that can be visually represented. Assessment of creativity ~~~~~~~~~~~~~~~~~~~~~~~~ Creativity encompasses both personality and cognitive aspects related to the production of unique and useful ideas (Runco & Jaeger, 2012; Simonton, 2000). Three of the many components associated with creativity are: Cognitive Flexibility, the ability to switch cognitive sets to adapt to changing environmental stimuli (Scott, 1962); Curiosity, the recognition, pursuit, and intense desire to explore novel and uncertain events (Kashdan & Silvia, 2009) and Openness to Experience, the Big Five personality trait considered as a proxy of creativity (Feist, 1998; Furnham & Bachtiar, 2008; Martindale, 1989). Because of the broadness of the construct, multi-trait, multi-method approaches have been proposed as most suitable (Cropley, 2000; Plucker & Makel, 2010). An image-based measure of creativity may add to this array of measurement methodologies available for creativity testing. In addition, image-based response scales may be particularly effective in measuring creativity because images elicit aesthetic preferences, such as preferences for complexity, which in turn are indicative of self-reported creativity and aesthetic styles (Barron, 1953; Chamorro-Premuzic, Reimers, Hsu, & Ahmetoglu, 2009; Rawlings, 2003; Swami, Stieger, Pietschnig, & Voracek, 2010; Wiersema, van der Schalk, & van Kleef, 2012). A preference for complex polygons is associated with higher self-reported creativity, such that Eisenman and Robinson (1967, 1968) suggested the use of polygons varying in their level of complexity as measures of creativity. Accordingly, the present research aimed to a) develop a novel format image-based creativity measure, b) investigate its concurrent validity in relation to three text-based measures of creativity, and c) assess its test- retest reliability. Curiosity and Exploration Inventory-II (CEI-II; Kashdan et al., 2009) A 10-item, five-point Likert self-report scale. The CEI-II measures two traits: stretching (e.g., ‘I actively seek as much information as I can in new situations’) and embracing (e.g., ‘I am the type of person who really enjoys the uncertainty of everyday life’). The CEI-II demonstrates reliability estimates of 0.85, construct validity, discrimination, desirable breadth of difficulty (Kashdan et al., 2009), and predictive validity for task performance (Kashdan, Rose, & Fincham, 2004). Cognitive Flexibility Inventory (CFI; Dennis & Vander Wal, 2010) A 20-item, seven-point Likert scale, self-report measure of adaptive thinking in stressful situations. Thirteen items assess behaviours related to alternatives (e.g., ‘I consider multiple options before making a decision’), and seven items behaviours related to control (‘When I encounter difficult situations, I feel like I am losing control’). The CFI shows a reliable factor structure, internal consistency, test-retest reliability, and concurrent validity (Dennis & Vander Wal, 2010). Openness to experience (Goldberg, 1999) Measured on a five-point Likert scale (‘very inaccurate’ to ‘very accurate’) using the 10-item Openness scale from the International Personality Item Pool (e.g. ‘I enjoy hearing new ideas’). Item design ~~~~~~~~~~~ The question stem of image-based items retained its verbal format, but the response scale presented a range of images (see Fig. 1). Each item consisted of a text-based question and between two and eight image response options. The image response options took one of two forms: they either assessed varying levels of the same trait, or they represented different traits. Seventy-seven items were designed to reflect Cognitive Flexibility, Curiosity, and Openness. Scoring ~~~~~~~ The scoring algorithm was developed on a sample of 964 participants, recruited using a UK panel company, and compensated for their participation. The panel had an equal distribution of males and females, and participants were UK residents. Approximately half of the users were 18–25 and the other half 25–36 years old. Participants completed the three creativity measures as well as all 77 image-based items. Rather than stipulating which responses were indicative of which underlying trait, responses to image-based items were scored in relation to standard measures. This method is commonly used in measure validation procedures when testing concurrent validity between new and existing measures (Rust & Golombok, 2009), as well as for predictive personality measures (Bachrach, Kosinski, Graepel, Kohli, & Stillwell, 2012; Boyd et al., 2015; Lambiotte & Kosinski, 2014; Youyou, Kosinski, & Stillwell, 2015; Schwartz et al., 2013; Wang, Kosinski, Stillwell, & Rust, 2014). Responses to all 77 items were dummified, with each image response option being transformed into a binary variable. This resulted in 321 dummy variables. Dummified responses were used as the independent variables (predictors) and the creativity scores as the respective dependent (predicted) variables in linear regression models to estimate the creativity scores. Item selection As the large number of dimensions resulting from 321 dummy variables can cause over-fitting, two methods of feature selection were used. The first method, LASSO (Least Absolute Shrinkage and Selection Operator) regression with 10-fold cross validation was applied to reduce the number of image response options. LASSO is a regularized regression, which penalizes variables with large coefficients and discounts variables with inconsistent performance across the sample. Thereby LASSO selects response options that are most indicative of creativity. LASSO cannot take into account that some image response options were taken from the same question. In order to account for the contribution of single questions, a second feature selection method, Sequential Forward Selection, was used (Devijver & Kittler, 1982). Starting with an empty set of questions, LASSO regression with 10-fold cross validation was used to estimate the relevant scale. The predicted and measured scores were correlated, and additional questions added at each step until no new question improved the correlation by > 0.1. Questions with individual correlations higher than 0.2 were also retained. This resulted in a final set of 14 questions, or 64 dummy variables.","With the selected 14 questions as predictor variables, LASSO regression with 10-fold cross validation was performed to predict creativity scores. Coefficients for the models predicting Cognitive Flexibility, Curiosity, and Openness, are presented in Table 1. Validation ~~~~~~~~~~ 1071 participants (605 females) were recruited using Amazon's Mechanical Turk (MTurk). Participants completed the text-based creativity measures and the 14-item image-based measure, which is part of the Red Bull Wingfinder assessment. To assess test-retest reliability, a subset of 162 participants retook the test after 60 days. MTurk panellists were US citizens paid for their participation. 15% were aged 18–24, 47% aged 25–34, 24% aged 35–44, and 14% aged 45 to 59. Responses to the image-based measure were scored using the algorithm described in Section 2.3. Creativity scores were normally distributed on the text- and image-based measures (see Table 2). The three text-based creativity scores had moderate intercorrelations (average r = 0.47, with p < 0.001), as had the three image- based scores (average r = 0.5, with p < 0.001) (see Table 2). Correlations between text- and image-based scores were moderate to high. Concurrent validity was higher for Curiosity and Openness than for Cognitive Flexibility (see Fig. 2). The average test-retest reliability of the image-based measures was r = 0.63 (p < 0.001) (see Fig. 2).","The aim of this study was to examine the psychometric properties of a newly developed image-based creativity measure. Results obtained from two large samples provided preliminary evidence for the test-retest reliability and concurrent validity of the 14-item measure. The developed scoring algorithm accurately predicted creativity scores on two of the three existing scales. This finding is in line with studies demonstrating the use of predictive models for measuring personality (Chen, Hsieh, Mahmud, & Nichols, 2014; Lambiotte & Kosinski, 2014; Yarkoni, 2010) and indicates that predictive scoring algorithms are suitable for scoring image-based response scales. Moderate correlations between the image-based and the text-based measures for Curiosity and Openness demonstrated good concurrent validity of the image-based format (r = 0.5, p < 0.001). Furthermore, the measure exhibited good test retest reliability (r = 0.65, p < 0.001), indicating that the selected image-based items are able to reliably measure aspects of creativity. On the other hand, the concurrent validity for Cognitive Flexibility was relatively low (r = 0.35, p < 0.001), suggesting that the selected images may not assess this particular aspect of creativity equally well. The predictive scoring algorithm used fewer items than established text-based measures, assessing all three creativity aspects with 14 items, compared to 40 items on the text-based measures. Both the predictive scoring algorithm and stronger associations evoked by images may be reasons for achieving shorter length (Meissner & Rothermund, 2015). The image-based measure demonstrated good test-retest reliability (average r = 0.63, p < 0.001), in particular taking into account factors that might have reduced the correlation including the small number of items, long interval between test and retest (six weeks), and the small to moderate sample size. Indeed, the observed test-retest reliability for the image-based Openness measure was higher than that reported in other studies for the ten-item, text-based Openness measure (reported r = 0.55 in Kosinski, Stillwell, & Graepel, 2013). Participants were more likely to consistently select the same image than they were to consistently select the same point on a Likert scale. This could be due to the relatively broader construct of creativity as compared to Openness (i.e. broader constructs tend to display better reliability; Chamorro-Premuzic, 2011). In addition, some image-based items had only two images as response options, compared with five to seven response options on Likert scales, which could lower the probability of changing responses. Implications ~~~~~~~~~~~~ The current study has a number of implications for the development of image-based assessments, particularly those focusing on creativity, but also beyond. For researchers and practitioners interested in alternatives to traditional self-report Likert response scales, this study provides preliminary support for the validity and reliability of an image-based measure. Although additional research is needed to replicate and extend these findings, this study takes a step towards providing evidence for the utility of innovative psychological assessments. The image-based measure has a number of advantages in practice. The measure is shorter than both existing text- and image-based creativity measures. Image response scales may be less obvious in what they are measuring than Likert scales. As a consequence, image-based scales could be less prone to faking and appear less intrusive to the test taker. Limitations and future research ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The current study has a number of limitations. It assessed the concurrent validity of image-based measures in relation to text-based, self-report inventories only. Although this provides initial evidence of the validity of image-based measures, additional studies are needed to investigate its relationship to non-self-report measures of creativity, such as divergent thinking tests. In addition, reliability of the image-based measure should be improved by identifying additional image-based items, in particular if the measure is to be used in selection contexts. The incremental validity of the image-based measures in predicting creative performance, beyond existing tests, should be established. This research would be necessary for demonstrating the value of using image-based measures alongside, in addition to, or as a replacement of, current creativity measures. There has been a call for a multi-method, multi-trait approach to the study of creativity (Batey & Furnham, 2006; Cropley, 2000; Park, Chun, & Lee, 2016; Plucker & Makel, 2010). Image-based measures may tap into the same performance variance as other creativity measures. But provided that image-based measures are more engaging than self-reports and easier to administer than divergent thinking tests, they could provide an alternative to existing creativity measures. Otherwise, image-based tests may predict distinct performance variance from current measures. In this scenario, they could be used alongside traditional tests. Accordingly, the purpose of developing an image based assessment is not only to provide more engaging alternatives for established (self-report) methods, but to provide an optional methodology for a valid multi-method approach to assessing creativity (and perhaps other personality traits like the Big Five).","This study supports the proposition that creativity can be measured via preferences for image-based stimuli. It may encourage research into innovative assessment formats and help practitioners in applying alternative assessments in settings where evidence of validity and reliability is required. Image-based assessments may provide a solution to an evolving need for alternative assessments, and this study was one of the first to attempt to bridge the gap between practise and research."],["'Impulsivity' refers to a range of behaviours including preference for immediate reward (temporal-impulsivity) and the tendency to make premature decisions (reflection-impulsivity) and responses (motor-impulsivity). The current study aimed to examine how different behavioural and self-report measurements of impulsivity can be categorised into distinct subtypes.Exploratory factor analysis using full information maximum likelihood was conducted on 10 behavioural and 1 self-report measure of impulsivity.Four factors of impulsivity were indicated, with Factor 1 having a high loading of the Stop Signal Task, which measures motor-impulsivity, factor 2 representing reflection-impulsivity with loadings of the Information Sampling Task and Matching Familiar Figures Task, factor 3 representing the Immediate Memory Task, and finally factor 4 which represents the Delay Discounting Questionnaire and The Monetary Choice Questionnaire, measurements of temporal-impulsivity.These findings indicated that impulsivity is not a unitary construct, and instead represents a series of independent subtypes. There was evidence of a distinct reflection-impulsivity factor, providing the first factor analysis support for this subtype. There was also support for additional factors of motor- and temporal-impulsivity. The present findings indicated that a number of currently accepted tasks cannot be considered as indexing motor- and temporal-impulsivity suggesting that additional characterisations of impulsivity may be required. --------------------------------------------------------------------------------","Impulsivity encompasses a range of behaviours that include making premature decisions, preferring immediate gratification and having difficulties inhibiting motor responses. Impulsivity functions as a dimension of normal behaviour, and it is thought that it can be adaptive in certain situations (Dalley, Everitt, & Robbins, 2011). However, it is also well established that it is associated with a number of negative outcomes (Aichert et al., 2012; Schweizer, 2002; Vigil-Colet & Morales-Vives, 2005) and is elevated in many clinical populations (e.g. de Wit, 2009; Winstanley, Dalley, Theobald, & Robbins, 2004; Winstanley, Eagle, & Robbins, 2006). There is growing consensus that impulsivity is heterogeneous and should not be considered a unitary construct and should instead reflect a variety of behaviours and processes (Evenden, 1999). In laboratory-based research, investigators have focused on two subtypes of behavioural impulsivity: ‘motor’-impulsivity (MI), as a failure to inhibit a behavioural response (also termed inhibitory control) and the failure to delay gratification (which we will term ‘temporal’-impulsivity [TI], also referred to as delay discounting). A third subtype of ‘reflection’-impulsivity (RI), i.e. the tendency to make decisions without gathering or evaluating necessary information, has also been suggested although it received comparatively little attention. Multiple tasks have been designed to index each subtype including the Stop Signal Task (SST), Go/NoGo (GNG) and Immediate Memory Task (IMT) for MI, the Matching Familiar Figures (MFF20) and Information Sampling Task (IST) for RI, and pen-and-paper measures such as the Monetary Choice Questionnaire (MCQ) and experiential tasks including the Single Key Impulsivity (SKIP) and Two Choice Impulsivity Paradigm (TCIP) for TI. Impulsivity can also be indexed using self- report measures (e.g. Kirby & Finch, 2010; Whiteside & Lynam, 2001), including the Barratt Impulsiveness Scale (BIS-11, Patton, Stanford, & Barratt, 1995). However, despite agreement that impulsivity comprises of multiple subtypes, they are rarely investigated concurrently and multiple tasks are seldom simultaneously administered to the same participants. Researchers typically select a single measure and refer to it as ‘impulsivity’, disregarding the wide array of processes and subtypes contributing to impulsive behaviour. This practice has led to poor characterisation of the structure of impulsivity, and the relationship of the subtypes to one another. Of the small number of studies attempting to address this, investigators typically find correlations between dependent variables of a task and have also found relationships between tasks indexing the same subtype (e.g. Broos et al., 2012; Dougherty et al., 2009; Reynolds, Penfold, & Patak, 2008), suggesting that the subtypes may be well-defined. In contrast, relationships between subtypes are not uniformly found (e.g. Broos et al., 2012; de Wit, 2009; Messer, 1976; Reynolds et al., 2008) and investigators employing factor analysis procedures have found that measures of TI and MI load onto different factors of impulsivity, providing evidence that these two subtypes may be distinct (Broos et al., 2012; Lane, Cherek, Rhoades, Pietras, & Tcheremissine, 2003; Reynolds, Ortengren, Richards, & de Wit, 2006). Collectively these studies provide preliminary evidence that the subtypes of impulsivity may be well-defined and differentiated. However these studies are limited by including too few tasks (e.g. Broos et al., 2012; Lane et al., 2003; Reynolds et al., 2006) despite evidence that more detailed classifications of impulsivity are required. The SST, GNG and IMT are used interchangeably as measures of MI in spite of evidence that the tasks index distinct processes: ‘action cancellation’, i.e. the inhibition of a response during its execution, on the SST (Dalley et al., 2011; Winstanley, 2011) and ‘action restraint’, i.e. the inhibition of a response before it has started, on the GNG and perhaps the IMT (Dalley et al., 2011, 2009; Eagle, Bari, & Robbins, 2008; Reynolds et al., 2008; Winstanley, 2011; Winstanley, Olausson, Taylor, & Jentsch, 2010). There is evidence that different neurotransmitters may contribute to the two processes (Eagle et al., 2008; Winstanley et al., 2010) and factor analysis has indicated two distinct factors of MI (Dougherty et al., 2009; Reynolds et al., 2008). With regards to TI, participants respond differently to experiential versus pen-and-paper measures (Winstanley, 2011), hypothetical versus real rewards (Hinvest & Anderson, 2009; Madden, Bickel, & Jacobs, 1999), monetary versus point rewards (Frederick, Loewenstein, & O’Donoghue, 2002) and also to short versus longer delays (Odum, 2011). However, these paradigms are used interchangeably despite there being no evidence to validate the assumption that they all index the same underlying process. Research suggests that self-report measures of impulsivity are not analogous with behavioural tasks (Dick et al., 2010). The BIS-11 has been found not to correlate with measures of MI or TI (Lane et al., 2003; Lansbergen, Schutter, & Kenemans, 2007; Reynolds et al., 2006) and investigators predominantly find distinct factors of self-report and behavioural impulsivity (Broos et al., 2012; Havik et al., 2012; Lane et al., 2003; Malle & Neubauer, 1991; Meda et al., 2009). However, there is some evidence that self-report impulsiveness is related to GNG performance (Aichert et al., 2012; Reynolds et al., 2006). Importantly, no factor analysis studies have included measures of RI despite evidence that it is clinically significant and distinct from other subtypes (e.g. Caswell, Morgan, & Duka, 2013a, 2013b; Morgan, Impallomeni, Pirona, & Rogers, 2006; Morgan, McFie, Fleetwood, & Robinson, 2002). The IST was designed to minimise some of the potential shortcomings of the MFF20 that include confounding by other cognitive processes (Clark, Robbins, Ersche, & Sahakian, 2006; Messer, 1976; Zelniker & Jeffrey, 1976) but although both tasks have been proposed to be analogous measures of RI, there are no factor analysis studies to validate this. As such, current literature discusses three subtypes of behavioural impulsivity, RI, MI and TI. There is evidence that MI and TI are distinct, although available literature is hampered by limited selection of tasks. This limited selection of tasks is a cause for concern as there is evidence of task differences within proposed subtypes that may have implications for their factor loadings and call into question their validity as indexes of the subtypes. Previous studies have also failed to incorporate RI into factor analysis models of impulsivity, despite evidence of its importance (e.g. Caswell et al., 2013a, 2013b; Morgan et al., 2006; Morgan et al., 2002). The current study aims to address these issues by investigating the structure of impulsivity using exploratory factor analysis, including measures of RI, to confirm whether impulsivity can be categorised into distinct subtypes. We will include a greater number of putative measures of different subtypes of impulsivity than has been attempted previously, encompassing the three proposed behavioural subtypes of MI, TI and RI (the previously unexplored subtype). The tasks will also include the BIS-11 as a self-report index of impulsivity although it is expected that separate facets of self-report and behavioural impulsivity will be identified.","160 (80 m, 80f) student participants at the University of Sussex were recruited, providing informed consent. They were required to be 18–45 years of age, not suffering from any mental illness, not be a heavy smoker (<20 per day), not taking any medication (excluding the contraceptive pill). Participants were instructed to abstain from the use of illicit recreational drugs for at least 1 week prior to the experiment and from the use of alcohol for at least 12 h prior to the experiments.","Participants completed the BIS-11 and the National Adult Reading Task, Alcohol Use Questionnaire and Drug Use Questionnaire followed by a battery of behavioural impulsivity tasks. Tasks were computerised and completed in a random order. Self-report and demographic measures National Adult Reading Task (NART; Nelson & O’Connell, 1978): The NART gives an estimate measure of verbal IQ. Participants did not complete the NART if they were dyslexic or second language English (n = 23). Alcohol Use Questionnaire (AUQ; Townshend & Duka, 2002): Participants estimate the number of alcohol units they consume per week. Drug Use Questionnaire (see Townshend & Duka, 2005): Participants give details of use for main drug categories. Participants were given a score where 0 = no use; 1 = use of cannabis/hash/marijuana; 3 = use of ecstasy/other drugs. Self-report impulsivity Barratt Impulsiveness Scale, Version 11 (BIS-11; Patton et al., 1995): The BIS-11 is a 30-item checklist measuring impulsivity. The questionnaire gives a total impulsivity score as well as three subscales of motor-, attentional- and nonplanning-impulsivity. Behavioural impulsivity Information Sampling Task (IST; Clark et al., 2006): Participants open a matrix of boxes to reveal two colours underneath before selecting the colour in the majority. There are two conditions available, each consisting of 10 trials, treated as separate tasks: Fixed win (FW): Participants win/lose 100 points regardless of number of boxes opened. Reward conflict (RC): For every box opened, participants lose 10 points from a bank of 250. The task gives the probability of being correct that the participant tolerates at the point of decision-making [P(correct)]. Matching Familiar Figures Task (MFF20; Cairns & Cammock, 1978; Kagan, Rosman, Day, Albert, & Phillips, 1964): Participants select the one of six visually presented stimuli which is identical to an original image. Participants complete 20 trials. The task gives a composite Impulsivity score (I-score). Stop Signal Task (SST; Logan, 1994): Participants respond to the direction of visually presented green arrows withholding this response whenever the arrow turns red (the Stop Signal, occurs 25% of trials). Participants complete 120 trials. The task gives a measure of Stop Signal Reaction Time (SSRTi). Go/NoGo (GNG; adapted from Kim, Iwaki, Imashioya, Uno, & Fujita, 2007): Participants respond whenever a visually presented triangle is pointing upwards (Go trials, occur 60% of trials) withholding this response if a triangle is pointing in another direction (Stop trials, occur 40% of trials). Participants complete 120 trials. The task gives a measure of the percentage of commission errors to Stop signals. Immediate Memory Task (IMT; Dougherty, Marsh, & Mathias, 2002): Participants press the mouse-button if a 5-digit number string is identical to the preceding string. Participants complete two blocks of 180 s, with a 20 s rest period between blocks. The task gives a measure of commission errors occurring when a participant makes a premature Go response to a Catch trial (occur 33% of trials). Single Key Impulsivity Paradigm (SKIP; Dougherty, Mathias, Marsh, & Jagar, 2005): Participants press the mouse-button to obtain a point reward. The magnitude of the reward is dependent on the delay between consecutive responses. Participants complete a four-minute trial. The task gives a measure of average inter-response time. Two Choice Impulsivity Paradigm (TCIP; Dougherty et al., 2005): Participants choose between two shapes representing a smaller-sooner (3 points after 3 s) and larger-later (9 points after 9 s) point rewards. Participants complete 30 trials. The task gives a measure of the number of smaller-sooner choices. Monetary Choice Questionnaire (MCQ; Kirby, Petry, & Bickel, 1999): A pen-and-paper task on which participants choose between hypothetical large delayed rewards, and smaller more immediate rewards. Participants complete 27 items. The task gives a measure of discounting of delayed rewards (k). Delay Discounting Questionnaire (DDT): The DDT is a variation on the MCQ. The pen-and-paper procedure is identical to the MCQ; participants complete 212 items. Stimuli are presented in a fixed random order. The task gives a measure of discounting of delayed rewards (k). Statistical analysis Exploratory factor analysis was conducted using Mplus 7,2 (Muthén & Muthén, 1992–2012). Participant demographics and correlations were analysed using Statistical Package for Social Sciences (SPSS) version 22. Participant demographics characteristics Participant demographic information including age, average IQ and alcohol and drug consumption are reported. Gender differences on the tasks were calculated to ensure that there were no differences that may affect the factor structure. Pearson’s correlation coefficient was calculated to identify the relationship between age and impulsivity. Variable selection One primary dependent variable was selected per task for the factor analysis. This selection was made in part due to the comparatively small sample size; had the sample been larger, multiple measures from each task could have been included. The selection of one variable per task was also made to maintain consistency between tasks as each task contains a varying number of outcome variables and it is known that factor analysis depend heavily on the number of indicators included per expected factor– including multiple from a select number of measures would have caused imbalances in the factor structure by increasing shared variance (Russo, Leone, Lauriola, & Lucidi, 2008). Variables included in the factor analysis model were: DDT mean k value. MCQ mean k value. TCIP number of impulsive choices. IMT percentage commission errors. GNG percentage commission errors. Data were checked to ensure that any participant who did not understand the task, or displayed inconsistent responding, was excluded from that measure – see Section 3.1 for details of included participants. The SST, SKIP, DDT and MCQ were log 10 transformed to correct issues of non-normality. All variables were coded so that large values indicate increased impulsivity. Correlations between task Pearson’s correlation coefficient was calculated to identify the correlations between impulsivity measures selected for the factor analysis. In an additional exploratory analysis, as a variables of interest not included in the primary analysis, correlations between the BIS-11 subscales and the behavioural measures of impulsivity were calculated to identify the relationship between self-report and behavioural measures of impulsivity. Factor analysis Exploratory factor analysis (EFA) was conducted to identify the factor structure of impulsivity. The sample size of 160 participants for the 11 items exceeds the suggested minimum ration of 5 participants per item (Gorsuch, 1983). EFA was carried out using full information maximum likelihood with Geomin oblique rotation. χ2, comparative fit index (CFI), a root mean square error of approximation (RMSEA), and standardized root mean square residual (SRMR) were used to evaluate the fit between the model and the data. CFIs of ⩾0.90 indicate a good fit to the data (Browne & Cudeck, 1993). A RMSEA value <0.05 indicate a good fit to the data (Browne & Cudeck, 1993). Well-fitting models obtain SRMR values <0.05 (Cooke et al., 2013). As measures of appropriateness of factor analysis, a Kaiser–Meyer–Olkin (KMO) value >.5 indicates acceptable sampling adequacy. A significant result for Bartlett’s test of sphericity indicates that the null hypothesis that the correlation matrix is an identity matrix can be rejected. Missing data and exclusions ~~~~~~~~~~~~~~~~~~~~~~~~~~~ BIS-11 Data were missing for one participant. SST Data were missing for 3 participants; a further 9 were excluded for GoRTs > 1000msecs, or 100% Stop accuracy as it was assumed that participants had not understood task instructions. GNG Data were missing from 2 participants, one participant was excluded for failing to stop to any Stop signals as it was assumed that they had not understood task instructions. IMT Data were missing for 3 participants. SKIP Data for 3 participants were missing. DDT Data were missing for 1 participant, 14 participants were excluded according to the inclusion criteria (see Johnson & Bickel, 2008). MCQ Data for 3 participants were missing. Full information maximum likelihood (FIML) was used to handle missing data for the factor analysis and therefore all 160 participants were included in the factor analysis (Enders, 2010). Correlational analysis was performed on all included data. Participant characteristics ~~~~~~~~~~~~~~~~~~~~~~~~~~~ The age of participants ranged from 18 to 45 (M 20.85; S.D. 3.79). Estimated verbal IQ ranged from 90 to 124 (M 108; S.D. 7.19). Participants drank on average 17 units of alcohol/week (range 0–72, S.D. 14.39). 40% of participants reported no drug use, 31% reported marijuana use, 29% reported other drug use. There were no gender differences on any impulsivity measure, see Table 1. There were significant associations between age and the Two Choice Impulsivity Paradigm (r(160) = −.173, p = .029) and the Delay Discounting Task (r(145) = .182, p = .029). There were no other correlations between age and any other impulsivity measures [BIS-11, r(159) = .067, p = .40; SST, r(148) = .149, p = .07; GNG, r(157) = −.017, p = .83; IMT, r(157) = −.046, p = .57; ISTfw, r(160) = .135, p = .09; ISTrc, p(160) = .074, p = .35; MFF20, p(160) = .043, p = .59; SKIP, p(157) = −.036, p = .65; MCQ, p(157) = .082, p = .31]. Correlations between tasks ~~~~~~~~~~~~~~~~~~~~~~~~~~ The correlation matrix between primary variables of each task is presented in Table 2a. Correlations between the BIS-11 subscales and each of the behavioural impulsivity measures are presented in Table 2b. Factor analysis ~~~~~~~~~~~~~~~ A four-factor model was indicated and appeared to fit the data well [χ2(17) = 15.736, p = 0.5426; CFI = 1.000, RMSEA = 0.000, 90%CI = 0.000–0.660; SRMR = 0.031]. The KMO value was .511. Bartlett’s test of sphericity was significant, x2(55) = 156.06. The scree plot (Fig. 1) was uninformative and indicated no clear number of factors. Four factors were retained in the analysis. Factor 1 contained a high loading of the SST. Factor two represents RI with loadings of the IST and MFF20. Factor 3 represents performance on the IMT. Factor 4 represents performance on the DDT and the MCQ. The SKIP, TCIP, GNG and the BIS-11 did not load onto any factors. There were no significant correlations between factors. See Tables 3 and 4 for factor loadings and correlations.","The current study provides important new insights into the structure of impulsivity. The results indicate that impulsivity should not be considered a unitary construct and instead represents a series of independent subtypes. Importantly, the results provide the first factor analysis support for the suggestion of a distinct, well-defined factor of RI. There was also support for the characterisation of behavioural impulsivity into additional factors of MI and TI. The current findings indicated that a number of currently accepted tasks as measurements of MI and TI cannot be considered as indexing these two subtypes and therefore suggest that additional characterisations of impulsivity may be required. Overall, there does not appear to be a strong underlying factor structure; instead, measures purported to index impulsivity typically do not correlate other than in small independent clusters. The study is the first to implement RI in factor analysis protocols. The results indicated that all putative measures of RI loaded onto a single factor thereby validating these measures and suggesting that, in addition to its clinical significance, RI is distinct from other subtypes of impulsivity. The MFF20 has been criticised as being confounded by other cognitive processes (Block, Block, & Harrington, 1974; Clark et al., 2006; Southgate, Tchanturia, & Treasure, 2008) with the IST developed to circumvent these issues (Clark et al., 2006). However, despite their procedural differences the current study provides the first validation that the two measures index the same primary underlying process. The results provide evidence that MI can be considered independent from RI and TI; however, it appears that tasks purported to measure MI do not index the same underlying processes. Whilst the SST, IMT and GNG are often used interchangeably, the results indicate that they index different forms of inhibitory control. The SST loaded onto a distinct factor, providing evidence that ‘action cancellation’ (Winstanley, 2011) is dissociable from other forms of inhibitory control. While it has been proposed that the GNG and IMT both index ‘action restraint’ (Winstanley, 2011) the two tasks loaded separately suggesting that the tasks measure different processes. The IMT loaded onto the third factor; on the task, participants must refrain from responding until the correct cue is presented (Winstanley, 2011) and it has been noted that responding on the task is self- generated where participants regulate their behaviour in anticipation of a ‘go’ signal (Winstanley et al., 2010). The results suggest that this form of self-generated responding may be an independent facet of impulsivity, distinct from action cancellation on the SST. The GNG did not load onto any factor suggesting it should be treated with caution as a measure of impulsivity. Overall, the data indicate that types of motor-impulsivity are behaviourally characterisable and are dissociable. Investigators have developed pen-and- paper and experiential tasks to measure TI, however there has been little research validating the assumption that they index the same process. Our results indicate that pen- and-paper measures (the DDT and MCQ) are analogous and that participants respond consistently despite differences in reward and delay values. However, neither experiential task loaded onto the factor indicating that they do not index TI as currently understood. Ostensibly, the TCIP and pen-and-paper measures both require participants to select between smaller-sooner and larger-later rewards. However, the tasks differ in the magnitude and type of reward- and delay-values. The comparatively short delays on the TCIP may not have been sensitive to individual differences (Winstanley et al., 2006). The point rewards are received in the laboratory, removing expectations of inflation, future income and the probability of receiving the delayed reward (Frederick, Loewenstein, & O’Donoghue, 2002). The SKIP is methodologically distinct, utilising a free-operant procedure – the longer participants wait between consecutive responses, the more points they receive. The underlying processes are relatively unexplored; the task correlated with the GNG suggesting that the two may share underlying processes. The results provide evidence that the SKIP and TCIP do not index TI processes as they are currently understood, and that neither are analogous to pen-and-paper measures. Self-reported impulsivity on the BIS-11 did not load onto any factor. This supports evidence that self-report impulsivity loads separately from behavioural tasks (Broos et al., 2012; Lane et al., 2003; Malle & Neubauer, 1991; Meda et al., 2009) and suggests that the two are heterogeneous. Interestingly, although self-reported impulsivity did not load onto any factor, performance on the IMT was related to BIS-11 total score as well as the nonplanning subscale, the MCQ also correlated with the nonplanning subscale whilst the SST was the only task that correlated with the BIS-11 motor subscale. No behavioural task correlated with the attentional subscale. These correlations suggest that there may be some limited associations between self-reported impulsivity and performance on behavioural tasks. There were no gender differences in impulsivity indicating that the model applies to both genders, and age did not correlate with any of the measures except for the TCIP and DDT indicating that for the most part impulsivity does not differ with age amongst our sample. All participants were university students however we did not take a measure of income which may have been an important demographic factor of interest. There are further limitations to the analysis which should be discussed. Despite each of the impulsivity tasks providing multiple outcome measures, only the primary impulsivity index was selected from each. Including multiple measures may have provided a more nuanced profile of the constructs under study. For example, the BIS-11 can be categorised into three sub-scores of self-report impulsivity; as these were not included we cannot identify whether any would have loaded onto any of the identified factors, although the lack of consistent correlations between the subscales and the behavioural tasks suggest that they would not. There were two primary reasons for this selection – to address the limited sample size and to avoid imbalances in the number of variables included from each task. The number of subjects is adequate based on our reduced selection of only one variable per task, and selecting multiple variables from each task would have jeopardised this; a larger sample size would have permitted more refined analysis of multiple outcome variables and future studies with greater power are needed to evaluate the sub-scores for each task. The selection of one variable per task was also made to maintain consistency between tasks. Each of the tasks provide a differing number of outcome variables and had we selected multiple from one task (e.g. the BIS-11) we would also have had to select multiple from every other tasks to prevent imbalances in the factor structure (it is known that factor analysis depend heavily on the number of indicators included per expected factor). In light of this we made the decision to select only one per task. Unfortunately, in reality such imbalances are unavoidable, with the Fixed Win and Reward Conflict versions of the IST, and the two pen-and-paper measures of TI being very similar; these methodological overlaps may have implications for the observed factor structure by increasing shared variance. In summary, the results provide evidence that impulsivity should not be considered a unitary construct, instead consisting of a series of independent subtypes. The data provide compelling support for the suggestion of a distinct, well–defined factor of RI. There was also support for the categorisation of behavioural impulsivity into additional factors of MI and TI. However, the results suggest that a number of currently accepted tasks cannot be considered as indexing these two subtypes, instead indicating that additional characterisations of impulsivity may be important. The results indicate that the IMT represents an additional facet of impulsivity. The data suggest that a number of tasks purported to index impulsivity should be treated with caution, and the results should be used as a basis for investigators in selecting tasks. It is hoped that the results encourage more researchers to implement multiple tasks to index ‘impulsivity’, as opposed to tasks in isolation."],["Conflict between goals (inter-goal conflict) and conflicting feelings about attaining particular goals (ambivalence) are believed to be associated with depressive and anxious symptoms, but have rarely been investigated together. Kelly et al. (2011, Personality and Individual Differences, 50, 531-534) reported that inter-goal conflict interacted with ambivalence to predict concurrent depressive symptoms in undergraduates, with ambivalence being more strongly associated with depressive symptoms for persons reporting less inter-goal conflict. We sought to replicate and extend this finding in a larger sample, using separate measures of inter-goal conflict and facilitation, and a longitudinal follow-up. Undergraduates (N = 210) rated their goal strivings for ambivalence, inter-goal conflict and facilitation, and completed measures of depressive and anxious symptoms that were repeated after one month. Inter-goal conflict (but not facilitation) and ambivalence were both uniquely positively associated with depressive and anxious symptoms concurrently, but did not predict symptom change. Inter-goal conflict and ambivalence did not interact to predict concurrent symptoms, but inter-goal conflict was associated with greater reductions in anxious symptoms for people reporting low ambivalence. Findings suggest that different forms of motivational conflict across the goal hierarchy are associated with symptoms, but do not exacerbate symptoms over time. --------------------------------------------------------------------------------","Making progress on personal goals imbues life with meaning and contributes to well-being (Brunstein, 1993; Klinger, 1977; Klug & Maier, 2015), so it is unsurprising that goal conflict has long been considered to be associated with psychological distress (Higginson, Mansell, & Wood, 2011). This article examines how two different forms of conflict (inter- goal conflict and goal ambivalence) contribute to anxious and depressive symptoms. A person experiences inter-goal conflict when one of their goals makes it more difficult to pursue their other goals (Emmons, 1986; Riediger & Freund, 2004). For example, a person's goal to ‘spend more time with my family’ may conflict with their goal to “get promoted at work”. Conversely, a person may experience inter-goal facilitation if one of their goals makes it easier to pursue their other goals (e.g., “spend more time with family” may facilitate the goal to “deepen my relationships”). Inter-goal conflict is associated with negative affect and lower life satisfaction (Emmons, 1986) and more psychiatric symptoms among undergraduates (Perring, Oatley, & Smith, 1988) and adolescents (Dickson & Moberly, 2010). However, some studies using undergraduate samples have not found associations between inter-goal conflict and depressive (Emmons & King, 1988, Study 2; King, Richards, & Stemmerich, 1998; Segerstrom & Solberg Nes, 2006) or anxious symptoms (Emmons & King, 1988, Study 2). In community samples, no significant correlations emerged between inter- goal conflict and depressive symptoms (Wallenius, 2000) or negative affect (Kehr, 2003; Romero, Villar, Luengo, & Gómez-Fraguela, 2009). Equivocal results may reflect the use of bipolar measures that conflate inter-goal facilitation and conflict. Riediger and Freund (2004) found that unipolar measures of inter-goal conflict and facilitation loaded on distinct factors, with only inter-goal conflict being significantly associated with negative affect at the between- and within-person level. Boudreaux and Ozer (2013) found that inter-goal conflict, but not inter-goal facilitation, was positively correlated with anxiety and negative affect in undergraduates; the correlation with depressive symptoms was not significant. In their meta-analysis, Gray, Ozer, and Rosenthal (2017) revealed that goal conflict was positively associated with psychological distress (weighted effect size: r = 0.34), with studies using unipolar scales yielding larger effect sizes. Inter- goal conflict may be less distressing if it represents competition among goals for a shared limited resource (e.g., time or money) rather than inherently incompatible outcomes (Riediger & Freund, 2004; Segerstrom & Solberg Nes, 2006). However, conflicted motives about attaining specific goals, i.e., ambivalence (Bleuler, 1911; Sincoff, 1990), may illustrate more profound motivational conflict that is more strongly associated with psychological symptoms. Goal ambivalence has been conceptualised as an approach-avoidance conflict about the pursuit of a particular goal (Emmons, King, & Sheldon, 1993) that is generated by conflict between relevant goals at a higher level in the goal hierarchy (Kelly, Mansell, & Wood, 2015). For example, a student may feel ambivalent about an essay- writing goal because it is relevant to a higher-level goal conflict between excelling academically and maintaining interpersonal relationships. Higher-level goal conflict may be more irresolvable because such goals are self-defining (Powers, 1973). Goal ambivalence has indeed been found to be associated with anxious and depressive symptoms among undergraduates (Emmons, 1986; Emmons & King, 1988; King et al., 1998; but see Romero et al., 2009, for null results). Other research has examined the association between psychological symptoms and ambivalence about goals relevant to particular life stages. For pregnant women, ambivalence about childbirth was associated with concurrent depressive symptoms and increasing symptoms post-partum (Koletzko, La Marca-Ghaemmaghami, & Brandstätter, 2015). In another sample, daily fluctuation in ambivalence about having the child was associated with negative affect. In another study, ambivalence about attaining a degree was associated with lower life satisfaction both concurrently and longitudinally (Koletzko, Herrmann, & Brandstätter, 2015). Inter-goal conflict and ambivalence may overlap because people will often feel ambivalent about conflicting goals (Emmons & King, 1988). Indeed, modest positive correlations have been reported between goal ambivalence and inter-goal conflict at the within-person level (Emmons, 1986; King et al., 1998), if not at the between-person level. Few studies have examined whether inter-goal conflict and ambivalence have independent or interactive associations with symptoms (Kelly et al., 2015). Although Emmons (1986) found that ambivalence but not inter-goal conflict explained unique variance in psychological symptoms, this study was underpowered. Kelly et al. (2011) reported that goal ambivalence was positively associated with concurrent depressive and anxious symptoms, whereas inter-goal conflict did not predict significant additional variance. Moreover, these forms of conflict interacted such that ambivalence was more strongly associated with depressive symptoms for participants reporting less inter-goal conflict. The authors speculated that ambivalence may be more distressing if it is not attributable to the pursuit of conflicting lower-level goals, suggesting that the ambivalence is generated by higher-level goal conflict. A person who strives to run marathons and learn guitar may report inter-goal conflict due to limited leisure time, but may experience no ambivalence if these pursuits are consistent with higher-level goals (Kelly et al., 2015). Conversely, a person who strives to care for the vulnerable and provide childcare may report no inter-goal conflict, but may experience ambivalence if these pursuits conflict with a higher-order goal of being independent. A combination of low inter-goal conflict and high ambivalence may indicate a distressing lack of integration across levels of the goal hierarchy. However, Kelly et al.’s (2011) result requires replication, and it is unclear whether the relationship between ambivalence and depressive symptoms is moderated by lower levels of inter-goal facilitation and/or higher levels of inter-goal conflict. To further illuminate the unique and interactive relationship between inter-goal conflict, ambivalence and psychological distress, we extended Kelly et al.’s (2011) research using a larger sample and distinct measures of inter-goal conflict and facilitation (Riediger & Freund, 2004). We also examined whether inter-goal conflict, goal ambivalence and their interaction would predict symptom change over one month, consistent with the notion that inter-goal conflict actively contributes to psychological distress. Boudreaux and Ozer (2013) found that inter-goal conflict predicted increases in depressive and anxious symptoms over five weeks in undergraduates. Similarly, Koletzko, La Marca-Ghaemmaghami, and Brandstätter (2015) found that ambivalence about having a child in women was associated with worsening depressive symptoms after birth. Based on the notion that conflict is deleterious at all levels of the goal hierarchy (Powers, 1973), we hypothesised that inter-goal conflict and goal ambivalence would each predict unique variance in anxious and depressive symptoms. Inter-goal facilitation was included as a covariate, but was not expected to be associated with anxious or depressive symptoms (Riediger & Freund, 2004). We sought to replicate Kelly et al.’s (2011) interaction between ambivalence and inter-goal conflict, such that anxious and depressive symptoms would be highest for individuals reporting high level of goal ambivalence and low levels of inter-goal conflict. Prospectively, we expected that higher levels of ambivalence and inter-goal conflict would each predict increases in anxious and depressive symptoms. More tentatively, we predicted that the interaction between inter- goal conflict and ambivalence would explain additional variance in symptom change.","Two hundred and ten undergraduate students (169 women, 41 men; M = 20.0 years, SD = 2.5, range = 18–35) were recruited from the University of Exeter campus via online advertisements. Participants were remunerated with course credit or £15.","Participants attended an initial 1 h session in which they provided informed consent, before completing a personal strivings assessment, inter-goal conflict and facilitation matrices, and depressive and anxious symptom scales. Personal goal strivings (Emmons, 1986) Participants first read instructions asking them to list at least ten personal goals, defined as “things that you typically or characteristically are trying to do”, by completing the stem: “I typically try to…” Examples were provided (e.g., “Convince others that I am intelligent”) and participants were told that they should list goals that identified them as individuals, rather than goals that other people thought they should have. Participants who generated more than ten goals were asked to choose the ten that represented them most accurately. Allowing for minor wording changes, Emmons (1986) found that 82% of goals were consistent over one year. Goal ambivalence (Emmons, 1986) Participants rated their ambivalence about each of their goals on a 6-point scale from 0 (none at all) to 5 (extreme) in response to the following question: “Sometimes even though we successfully reach a goal, we are unhappy (e.g., if you're “trying to become more intimate with someone” and you succeed, you might also feel concern about being tied down). How much unhappiness do you or will you feel when you are successful in this striving?” Mean ambivalence scores across goals were calculated for each participant (α = 0.79). Goal ambivalence has previously shown a one-year stability correlation of 0.65 (Emmons & King, 1988). Inter-goal conflict and facilitation (Riediger & Freund, 2004) Participants next completed two 10 × 10 matrices to rate inter-goal conflict and facilitation respectively. In each matrix, each of the participant's ten goals was listed in both rows and columns. In the conflict matrix, participants rated the extent to which pursuing each of their goals in the rows “makes it more difficult to pursue” each of the other strivings across the columns. In the facilitation matrix, participants were asked to rate the extent to which pursuing each of the goals in the rows “makes it easier to pursue” each of the other goals across the columns, on a 6-point scale from 0 (not at all) to 5 (extremely). Thus, participants rated the extent to which each of their goals both conflicted with and facilitated each of their other goals (bidirectionally). Mean inter-goal conflict (α = 0.91) and facilitation ratings (α = 0.90) were calculated for each participant. Beck Depression Inventory–II (BDI-II; Beck, Steer, & Brown, 1996) Participants completed the BDI-II, a validated 21-item scale assessing depressive symptoms over the past two weeks. Each item is rated on a scale from 0 to 3, yielding a total score from 0 to 63 (α = 0.90). Generalized Anxiety Disorder–7 (GAD-7; Spitzer, Kroenke, Williams, & Löwe, 2006) Participants completed the Generalized Anxiety Disorder–7, a seven-item scale assessing the frequency of anxious symptoms over the past two weeks. Each item is rated on a four-point scale from 0 to 3, yielding a total score from 0 to 21 (α = 0.85). Participants completed further goal measures not relevant to the current study, before making an appointment for a follow-up session one month later (M = 35.0 days, SD = 5.4). Follow-up session One hundred and ninety-four (92.3%) participants returned for the follow-up, when they completed the BDI–II (α = 0.91) and the GAD–71 (α = 0.85), together with other measures irrelevant to the current study, before being remunerated. Cross-sectional analysis ~~~~~~~~~~~~~~~~~~~~~~~~ Table 1 presents correlations and (untransformed) descriptive statistics for all variables. Inter-goal facilitation scores clustered around the scale midpoint but two- thirds of the sample reported mean inter-goal conflict and ambivalence scores below 1 on the 0–5 scale. Sample means were below recommended cut-offs for mild depressive (Beck et al., 1996) and mild anxious symptoms (Spitzer et al., 2006), and these variables were highly correlated. Due to small means, depressive and anxious symptom scores, inter-goal conflict and goal ambivalence were positively skewed so were log-transformed to improve normality. Men reported less inter-goal facilitation than did women, t(208) = 1.99, p = 0.048, d = 0.35, but no other significant gender differences emerged. Inter-goal conflict and ambivalence were each significantly positively correlated with depressive and anxious symptoms at both time points. At Time 1, inter-goal facilitation was modestly positively associated with anxious symptoms. Inter-goal conflict and facilitation were positively correlated between persons: people reporting more conflict among their goals tended to report more facilitation, perhaps reflecting a general response tendency. However, multi- level models (accounting for clustering of goals within persons) revealed that inter-goal conflict and facilitation were negatively correlated at the within-person level: goals that conflicted more with other goals tended to be less mutually facilitative. Inter-goal conflict and ambivalence were positively correlated at both levels of analysis. Inter-goal facilitation and ambivalence were not significantly correlated at the between-person level but were modestly negatively correlated at the within-person level. To assess unique and interactive relationships, inter-goal conflict, facilitation and ambivalence were each standardised before entry into a multiple regression model as predictors of Time 1 depressive symptoms in the first step, followed by the interactions between (i) inter-goal conflict and goal ambivalence and (ii) inter-goal facilitation and goal ambivalence in the second step.2 When entered together in the first step, inter-goal conflict, facilitation and goal ambivalence jointly explained 12.1% of depressive symptom variance, F(3, 206) = 9.44, p < 0.001. Goal ambivalence was independently associated with depressive symptoms at Time 1, β = 0.22, p = 0.002, as was inter-striving conflict, β = 0.20, p = 0.007, but inter-striving facilitation was not a significant predictor, β = −0.00, p = 0.99. Critically, when entered in the second step, the interactions between (i) inter-goal conflict and goal ambivalence, and (ii) inter-striving facilitation and goal ambivalence, jointly failed to explain significant additional variance in depressive symptoms at Time 1, ∆F(2, 204) < 1, p = 0.84, with neither individual interaction reaching significance, ps > 0.54. An equivalent multiple regression analysis was conducted to predict Time 1 anxious symptoms. When standardised and entered simultaneously in the first step, inter-goal conflict, facilitation and ambivalence explained 14.0% of anxious symptom variance, F(3, 206) = 11.14, p < 0.001. Striving ambivalence was independently associated with anxious symptoms at Time 1, β = 0.23, p = 0.001, as was inter-striving conflict, β = 0.19, p = 0.009, but inter-striving facilitation was not a significant predictor, β = 0.10, p = 0.13. Critically, when entered in the second step, the interactions between (i) inter- striving conflict and goal ambivalence, and (ii) inter-striving facilitation and goal ambivalence, jointly failed to explain significant additional variance in Time 1 anxious symptoms, ∆F(2, 204) < 1, p = 0.81, with neither individual interaction reaching significance, ps > 0.51. Longitudinal analysis ~~~~~~~~~~~~~~~~~~~~~ Paired t-tests revealed a small, statistically significant decrease in depressive symptoms from baseline to follow-up, t(193) = 3.42, p = 0.001, d = 0.20, but no statistically significant difference in anxious symptoms over this period, t(192) = 1.45, p = 0.15, d = 0.05. Further multiple regressions investigated whether inter-goal conflict and ambivalence predicted change in depressive and anxious symptoms from Time 1 to Time 2 respectively. In the multiple regression predicting Time 2 depressive symptoms, Time 1 depressive symptoms was entered first, followed by standardised inter-goal conflict, inter-goal facilitation, and goal ambivalence in the second step. Finally, the interactions between (i) inter-goal conflict and goal ambivalence and (ii) inter-goal facilitation and goal ambivalence were entered in the third step. Controlling for Time 1 depressive symptoms, inter-goal conflict, inter-goal facilitation and goal ambivalence jointly failed to explain significant additional variance in Time 2 depressive symptoms, F(3, 189) = 1.06, p = 0.37, ∆R2 < 0.01. Goal ambivalence, β = 0.10, p = 0.09, inter-goal conflict, β = −0.05, p = 0.38, and inter-goal facilitation, β = −0.00, p = 0.97, were not significant predictors. The interactions entered in the third step failed to explain significant additional variance in Time 2 depressive symptoms, F(2, 187) = 1.35, p = 0.26, ∆R2 < 0.01, with neither interaction being significant, ps > 0.19. Thus, inter-goal conflict, facilitation and goal ambivalence did not predict change in depressive symptoms, independently or interactively. In the parallel multiple regression predicting Time 2 anxious symptoms, after entering Time 1 anxious symptoms, simultaneous entry of inter-goal conflict, facilitation and goal ambivalence failed to explain additional variance in anxious symptoms at Time 2, F(3, 188) < 1, p = 0.82, ∆R2 < 0.01. Goal ambivalence, β = 0.04, p = 0.43, striving conflict, β = −0.04, p = 0.45, and striving facilitation, β = 0.00, p = 0.97, were not significant predictors. The interactions entered in the third step failed to explain additional variance in depressive symptoms at Time 2, F(2, 186) = 2.47, p = 0.09, ∆R2 = 0.01. However, because the crucial interaction between inter-goal conflict and ambivalence was significant (β = 0.11, p = 0.03; the other interaction was not, p = 0.77), we proceeded to explicate it. Fig. 1 plots (log) T2 anxious symptoms for persons scoring one standard deviation above and below the mean on inter-goal conflict and ambivalence, calculated at mean levels of T1 anxious symptoms and inter-goal facilitation. Tests of simple slopes revealed that higher levels of inter-goal conflict at Time 1 were associated with reductions in anxious symptoms at Time 2 for persons with lower goal ambivalence (β = −0.15, p = 0.04). However, levels of inter-goal conflict at Time 1 were not significantly associated with levels of anxious symptoms at Time 2 for persons with higher goal ambivalence (β = 0.04, p = 0.53).","Our results support theoretical perspectives and empirical research suggesting that goal conflict is associated with psychological distress (Higginson et al., 2011). The positive association between inter-goal conflict and concurrent psychological symptoms mirrors the results of a recent meta-analysis (Gray et al., 2017). As hypothesised, inter-goal facilitation was not uniquely significantly associated with anxious or depressive symptoms, consistent with distinct relationships for inter-goal conflict and facilitation (Riediger & Freund, 2004). The relationship between inter-goal conflict and symptoms is unlikely to be due to a general tendency for distressed people to make more pessimistic goal ratings, because no negative correlation emerged between inter-goal facilitation and symptoms. Inter-goal facilitation may be more relevant to psychological well-being than to distress symptoms (Riediger & Freund, 2004). Consistent with previous research (Emmons & King, 1988; King et al., 1998), goal ambivalence was associated with greater anxious and depressive symptoms. Ambivalence was moderately positively associated with inter-goal conflict at both the between-person and within-person level of analysis, but was uniquely associated with both anxious and depressive symptoms, suggesting that they are not mutually redundant. Our study had greater statistical power (0.80 to detect a small-medium effect size f2 = 0.05) than those reported by Emmons (1986) and Kelly et al. (2011), which may explain why they did not find that goal ambivalence and inter-goal conflict had unique associations with symptoms. Goal ambivalence and inter-goal conflict may reflect motivational conflict at higher and lower levels of the goal hierarchy respectively (Kelly et al., 2015). Furthermore, whereas inter-goal conflict must be consciously reported, ambivalence towards goals could suggest higher-level goal conflict that is outside conscious awareness. The unique contributions of inter-goal conflict and ambivalence observed here support the utility of using distinct measures to capture motivational conflict associated with distress throughout the goal hierarchy. We found no evidence for an interaction between ambivalence and inter-goal conflict in predicting concurrent distress symptoms. Using a larger sample and a unipolar measure of inter-goal conflict, we did not replicate Kelly et al.’s (2011) finding that inter-goal conflict buffered the relationship between ambivalence and depressive symptoms. These authors speculated that ambivalence may be more closely associated with psychological distress at lower levels of inter-goal conflict because this combination implicates unconscious higher-order goal conflicts that are difficult to resolve. Instead, our results suggest that ambivalence is associated with psychological symptoms at high and low levels of inter-goal conflict. Thus, psychological distress is associated with both mid-level conscious conflict and higher-level goal conflict that generates ambivalence (Kelly et al., 2015). This combination is illustrative of a low level of motivational integration across the goal hierarchy, which may be consistent with goal blockage and negative affect (Emmons & King, 1988). It is noteworthy that people with more depressive (and to a lesser extent, anxious) symptoms report proportionately more abstract goals (Dickson & MacLeod, 2004; Dickson & Moberly, 2013; Emmons, 1992). Abstract goals are rated as more difficult (Emmons, 1992), and their centrality to the self may make goal conflicts at this level appear irresolvable. Counter to expectations, neither inter-goal conflict nor goal ambivalence predicted change in anxious or depressive symptoms over one month. Nevertheless, a significant interaction revealed that people with high levels of inter-goal conflict and low levels of ambivalence experienced reductions in anxious symptoms. Although it would be inappropriate to over-interpret this unexpected finding, which did not emerge for depressive symptoms, moderate conflict or differentiation among goal pursuits may protect against anxiety in the absence of ambivalence. Previous longitudinal studies have not examined the longitudinal interaction of inter-goal conflict and ambivalence, and should seek to replicate this result. Our longitudinal results are contrary to Boudreaux and Ozer's (2013) finding that goal ambivalence predicted change in depressive symptoms among undergraduates, and Koletzko, Herrmann, and Brandstätter’s (2015) finding that mothers' ambivalence about the specific goal of having a child predicted increased depressive symptoms post-partum. It could be concluded that our longitudinal results suggest that inter-goal conflict is a concomitant rather than a cause of distress. However, depressive and anxious symptoms were highly stable over the one month period, such that it was difficult for other predictors to predict change. Furthermore, we elicited goals as enduring strivings that are relatively stable (Emmons, 1986), while ambivalence and inter- goal conflict ratings demonstrate considerable stability (Emmons & King, 1988). Therefore, any long-established pattern of motivational conflict may not predict further increases in psychological distress. Consistent with this, Kehr (2003) found that emerging but not enduring inter-goal conflict predicted changes in affect over eight weeks in managers. Studies indicate that within-person fluctuations in motivational conflict are correlated with state affect (Koletzko, La Marca-Ghaemmaghami, & Brandstätter, 2015; Riediger & Freund, 2004), suggesting that changes in inter-goal conflict or ambivalence that are associated with the adoption of new strivings might predict increases in psychopathology. Although our findings illuminate the role of distinct forms of motivational conflict in contributing to psychological distress, this study has limitations. First, we used a single item to measure goal ambivalence that asked participants to what extent they would experience negative emotions after goal attainment. Koletzko, Herrmann, and Brandstätter (2015) argued that this measure does not capture the contradictory motives entailed in ambivalence, and developed a new scale for this purpose. Second, we were unable to determine whether the association between inter-goal conflict and symptoms was related to inherent incompatibility between goals or competition between goals for a limited resource, although these dimensions correlate positively (Riediger & Freund, 2004). Some studies (e.g., Segerstrom & Solberg Nes, 2006) have recruited independent judges to rate inter-goal conflict to obtain more objective judgements, which risks failing to capture idiosyncrasies relating to personal goal strivings. We used an undergraduate sample whose strivings may be more homogenous than older adults who have a more consolidated identity. Finally, it is unclear to what extent our findings are influenced by the probable inclusion of participants who would meet diagnostic criteria for mood disorders. Nevertheless, the association between goal conflict and well-being is relatively consistent across samples (Gray et al., 2017). In conclusion, our results suggest that both inter-goal and intra-goal conflict are uniquely associated with psychological distress, such that goal ambivalence is associated with anxious and depressive symptoms for individuals reporting high and low levels of inter-goal conflict. Although our results suggest that chronic goal conflict may not exacerbate symptoms, future research could usefully concentrate on examining cross-lagged relationships between different forms of motivational conflict and psychological distress in periods when goal strivings are adopted or discarded."],["The present research examined whether the environmental responsibility and actions attributed to large scale organizations, such as the government, can influence people's environmental efforts. In particular, we examined whether people increase or decrease their willingness to enact energy conservation behaviors (ECB) when there is a shortfall between others' actions and their responsibility. In Studies 1 and 2 we found that willingness to enact ECB was positively correlated with judgements about each of the organizations' eco-responsibility but not their eco-actions. Interestingly, each of the organizations' actions were perceived as falling short of their responsibility and this shortfall was positively associated with willingness to enact ECB. In Study 3, we found that manipulating respondents perceptions of government shortfall increased participants' willingness to enact ECB. Overall our findings provide support for social compensation theory as when others actions fall short of their responsibility people are prepared to \"go the extra green mile\". --------------------------------------------------------------------------------","Environmental campaigns and policy initiatives often attempt to influence people's behaviors (DEFRA, 2008; Owens, 2000). For example, a discussion paper from the UK cabinet office argued that in striving for green behaviors, “the eventual aim is to entrench a habit of personal responsibility” (2004, p.5). However, while the onus appears to be on individuals there are other key actors or agents who also have a role to play in energy conservation such as firms, communities, governments, and international organizations (see Stern, 1992). Yet, to date, this wider social context has typically been overlooked in psychological research. Consequently, it remains to be seen if people's willingness to enact Energy Conservation Behaviors (ECB) is influenced by (a) the responsibility ascribed to others to conserve energy, (b) the actions others are seen to be taking and, (c) incidences in which other agents' responsibility to conserve energy falls short of their perceived eco-actions. The influence of other organizations on individual environmental efforts ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We propose that people's actions are influenced by collective dynamics, such that individuals look to others (including larger organizations) when setting their own behavior standards. We suggest this on the basis that people do not operate in a social and political vacuum; rather they are aware that other organizations and entities have a role to play in energy conservation. Indeed, in several qualitative studies it has emerged that people consider a number of organizations to be responsible for environmental efforts (Barr, Gilg, & Shaw, 2011; Hargreaves, Nye, & Burgess., 2013; Hinchlifffe, 1996; Lorenzoni, Nicholson-Cole, & Whitmarsh, 2007). Interestingly, such findings emerged despite the fact that the majority of these qualitative studies did not seek to examine the role of other agents in environmental behaviors - which suggests that such perceptions may be pervasive. Moreover, it is likely that these perceptions are fostered by the media which frequently provides commentary on the environmental efforts of a variety of agents and institutions. For example, in April 2014 the UK was hit by high levels of air pollution caused by a combination of local emissions, light winds, pollution from the continent, and dust from the Sahara. News articles were quick to acknowledge that such pollution could bring further attention to the, “government's long-term failure to reduce air pollution” (BBC, 2014). As such it is clear that a person's environmental action is situated in a broader set of social relations that need to be taken into consideration (see also Catney et al., 2013). Responsibility ~~~~~~~~~~~~~~ The link between personal responsibility and willingness to enact or support ECB has been established in a multitude of research studies (e.g., Guagnano, Dietz, & Stern, 1994; Hines, Hungerford, & Tomara, 1987; Hunecke, Blobaum, Matthies, & Hoger, 2001; Jansson, Marell, & Nordlund, 2010; Kaiser, Ranney, Hartig, & Bowler, 1999; Kaiser & Shimoda, 1999; Nordlund & Garvill, 2002; Steg, Dreijerink, & Abrahamse, 2005). In contrast, far less is known about the relationship between ascriptions of environmental responsibility to other agents and personal willingness to enact ECB. Yet, it is apparent from both quantitative and qualitative studies, that individuals are aware that other agents, such as their neighbours, the government, corporate bodies (e.g., city council, offices) and multinationals, have a role to play in energy conservation (e.g., Hargreaves, Nye, & Burgess, 2010; Hinchliffe, 1996; Lorenzoni et al., 2007; Stern, Dietz, & Black, 1985). However, it remains to be seen how these perceptions of others' environmental obligations influence people's own environmental efforts. According to the bystander effect, we might expect a diffusion of responsibility to occur and individuals to be less inclined to help by enacting ECB when responsibility is distributed among several others (Darley & Latané, 1968; Latané & Darley, 1970; Latané & Nida, 1981). Yet, on the other hand, if individuals consider both themselves and others responsible for energy conservation this may foster a sense of shared responsibility, such that willingness to enact ECB is positively influenced by ascriptions of responsibility to others. Action ~~~~~~ Past research suggests that social norms play a pervasive role in an individual's willingness to enact ECB (e.g., Barr et al., 2011; Cialdini, Reno, & Kallgren, 1990, Goldstein, Cialdini, & Griskevicius, 2008; McDonald, Fielding, & Louis, 2013; Nolan, Schultz, Cialdini, Goldstein, & Griskevicius, 2008; Schultz, Nolan, Cialdini, Goldstein, & Griskevicius, 2007). Typically, marketing campaigns use social norms to try and influence people's behaviors by changing perceptions of what is considered normal (descriptive norms) or socially acceptable (injunctive norms). For instance, researchers found that hotel guests were significantly more likely to re-use their towels when presented with the following normative appeal, “Join your fellow guests in helping to save the environment”, than when presented with the message, “Help save the environment” (Goldstein et al., 2008). Social norms can also lead people to act in ways that are detrimental to the environment. For example, people are more likely to litter in littered environments, and this effect is even more pronounced if they have witnessed another person drop litter (Cialdini et al., 1990 Experiment 1). As such, there is substantial support for the idea that people may enact either more or less ECB depending on what others are (or are not) doing. However, typically norms have been examined at the individual level and, to the best of our knowledge; there is currently no research that examines if the norms of larger social organizations (e.g., the government, energy suppliers) influence personal environmental efforts. On the one hand, the environmental actions that an organization takes (or does not take) may set an important precedent (i.e., it may act as a norm), especially given the position of power these organizations may hold. Yet, on the other hand, people may not consider the actions of larger organizations as relevant if they perceive that they are operating on a substantially different level from themselves. Considering responsibility and action together ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We propose that in order to understand if the wider social context contributes to intentions to enact eco-behaviors it is necessary to consider both perceptions of others' eco-responsibilities, and others' eco-actions. This is because while responsibility and action are distinct and separable from one another they are also clearly related. This relation stems from their definition. Specifically, responsibility is defined as, “the state or fact of having a duty to deal with something …”, while action is defined as, “the fact or process of doing something”. In other words, responsibility is about “what we ought to be doing” whereas action is about “what we are actually doing”. Thus when people think about responsibility it is likely that they also consider action. Of course, this does not mean that the two inevitably co-occur in an applied setting. Rather, it is possible to be responsible for something but not to take action and vice versa. However, given the operational links between responsibility and action there are two good reasons for considering the dual influence of both factors on intentions to enact ECB. First, considering action without responsibility may render the influence of action irrelevant. If an agent is not considered responsible for conserving energy then their actions or inactions are irrelevant and may have little bearing on our own actions. Second, considering action alongside responsibility provides the basis for moral judgements to be made about whether other agents are meeting their environmental responsibilities. As such, considering both responsibility and action together enables us to address an important and hitherto unanswered research question: to what extent are others' actions seen as matching their responsibility and in cases where others' action are perceived as falling short of their responsibility how does this influence personal willingness to enact ECB? In the present paper we refrain from making specific predictions about whether perceptions of others' shortfall will lead to either an increase or decrease in willingness to enact ECB. We argue that to do so would be inappropriate given that there are psychological mechanisms that can be used to infer support for either possibility. Specifically, when confronted with others' shortfall, the sucker effect and feelings of personal inefficacy may explain why people will decrease their efforts; whereas social compensation theory may explain why people will increase their efforts. Doing less: running a mile The ‘sucker effect’ describes a phenomenon that occurs when individuals experience motivation loss when they suspect that capable others are not contributing (Kerr, 1983). There is some indication from qualitative studies that the sucker effect may occur in response to perceptions that powerful organizations are failing to meet their environmental responsibilities (Barr et al., 2011; Hinchliffe, 1996). For example, one interviewee observed, “But it is discouraging when you hear … that places like America won't sign up to the Kyoto agreement … That's just pushing us into thinking, ‘well, why should we bother? ’” (Barr et al., 2011, p.716), while another interviewee commented, “I am one person and you think, well why am I going to change my lifestyle if all these other people aren't? It's human nature” (Lorenzoni et al., 2007, p.451). Diminished feelings of personal efficacy or perceived helplessness may also lead individuals to do less when others' actions fall short of their responsibility. Personal efficacy refers to “the belief in one's capabilities to organize and execute the courses of action required to manage prospective situations.” (Bandura, 1995, p.2). We suggest that individuals' personal efficacy may be undermined in the face of powerful global entities failing to live up to their environmental responsibilities. Indeed, there is some suggestion of this in qualitative research: “I'm impotent in a way because America didn't sign up to the Kyoto agreement” (Lorenzoni et al., 2007, p.450) and in quantitative research findings that show perceived behavioral control is a strong predictor of environmental behaviors (e.g., Heath & Gifford, 2002; Kaiser & Gutscher, 2003). Doing more: going the extra green mile Social compensation theory (SCT) proposes that in a collective setting people may work harder or contribute more effort to group tasks to compensate for others when they expect their co-workers to perform poorly on a meaningful task (Williams & Karau, 1991). As such, if extrapolated from a small group context (where SCT has previously been studied) to a more societal context, SCT provides a theoretical basis for why people may be more likely to increase their own environmental efforts if they perceive a discrepancy between others' actions and their responsibility. Notably, for social compensation to occur in an environmental setting three conditions need to be met. First, individuals would need to attribute responsibility for energy conservation both to other agents and to oneself to ensure energy conservation is considered an important shared goal. Second, participants would need to perceive that the actions others were taking to conserve energy were minimal or ineffective. Third, as social compensation theory is grounded in the collective effort model (Karau & Williams, 1993), individuals would need to perceive that their environmental efforts could make at least some difference to the goal of energy conservation.","In Studies 1 and 2 we utilized a correlational design to examine if willingness to enact ECB were associated with perceptions of others environmental responsibilities, actions, and the discrepancies between others responsibility and actions. In Study 3, we employed an experimental design and manipulated the extent to which the government's actions were seen as falling short of their responsibility to examine the effects on willingness to enact ECB.","In both Studies 1 and 2, we administered a questionnaire measuring participants' willingness to enact ECB, perceptions of others responsibility, actions, and the discrepancies between others responsibility and actions (i.e., the shortfall). In Study 1, shortfall was calculated by subtracting actions scores from responsibility scores. In Study 2 we aimed to replicate the results obtained in Study 1 and to extend the findings from Study 1, by administering measures of shortfall to ensure that the results we obtained were not the product of measurement biases.","In Study 1, participants were presented with a list of six agents displayed in a randomized order and were asked to rate the extent to which each agent was responsible for conserving energy and was taking action to conserve energy. We presented these rating tasks to participants on two different pages because we aimed to avoid highlighting any discrepancies between others' eco-responsibility and eco-actions (i.e., others' shortfall) lest this affected ECB ratings. Having completed the rating tasks, participants then indicated their willingness to enact ECB. All measures are detailed below in the order they were presented. In Study 2, we followed the same procedure. However, this time participants were randomly assigned to complete either (a) measures where shortfall was computed by subtracting action ratings from responsibility ratings (as per Study 1) or (b) bipolar measures of shortfall (ranging from “action is less than responsibility” to “action is more than responsibility”).","then completed the Environmental Government Judgement Scale (EGJS). This is a 15 item scale that we developed to measure perceptions that the governments environmental actions were (i) less than, (ii) in line with, or (iii) exceeding their responsibility. We developed and included this measure to avoid relying on scales comprised of the difference score calculated between two single items. After completing the EGJS, participants then indicated their willingness to enact ECB before providing their responses to a scenario where governments' environmental actions were described as falling short of their responsibility. We focused on the government, as opposed to other organizations, because, in principle, the government has the most control via legislation. In asking participants explicitly for their responses to the shortfall scenario we were able to see if their self-reported responses to shortfall would be in line with the correlations between shortfall and willingness to enact ECB. Participants ~~~~~~~~~~~~ In Studies 1 and 2 we recruited participants using a convenience sampling method via Amazon's Mechanical Turk1 to respond to an online survey entitled, “Social Issues and Your Opinions”. All participants were USA citizens. In Study 1, a total of 197 participants (103 males, 94 females) aged from 18 to 73 (M = 33.32, SD = 12.99) completed the survey. In Study 2, 212 participants (128 males, 84 females), aged from 19 to 72 (M = 34.34, SD = 11.98) completed the survey. Of these 212 participants, 108 completed shortfall ratings by providing action and responsibility ratings and 104 completed judgements of the bipolar shortfall measures. Responsibility and action judgements Participants provided their own judgements about the extent to which they agreed/disagreed that each of the six agents were (a) responsible for conserving energy and (b) taking action to conserve energy. All judgements were completed using 7-point scales from 1 (‘Strongly Disagree’) to 7 (‘Strongly Agree’). The agents listed were: myself, other consumers apart from me, the government, the energy suppliers, the big countries that use the most energy (hereafter referred to as industrial countries), and industrial factories. We selected these 6 agents based on the findings from a pilot study in which 10 participants (3 males, 7 females aged between 18 and 60, Mage = 31, SD = 12.27) responded to the open-ended question: ‘In your opinion who does the responsibility lie with to conserve energy?’ We computed a mean score of ‘Other Agents’ Shortfall’ comprised of shortfall ratings of the following agents: the government, the energy suppliers, industrial factories and countries. We excluded “other consumers” from this mean score because we reason that participants will likely perceive “other consumers” as more similar to themselves than to the other agents. In Study 1, we always measured responsibility first and action second. In Study 2, we administered action ratings first and responsibility ratings second. In administering these measures in a different order we aimed to rule out the possibility that our results were influenced by ordering effects2 Shortfall judgements (subtracting action from responsibility) Shortfall judgements were calculated by subtracting action ratings from responsibility ratings. Hence, a positive figure indicates a perception that the actions that others are taking to conserve energy fall short of the responsibility that others have to conserve energy. Shortfall judgements (a bipolar scale, Study 2 only) In Study 2, half of the participants were randomly assigned to complete shortfall judgements using a bipolar scale. Specifically, they were asked to indicate for each agent, ‘if the action they take to conserve energy measures up to the responsibility they have to conserve energy’ using a 13 point bipolar scale ranging from −6 (“action is less than responsibility”) to 6 (“action is more than responsibility”). The midpoint of 0 was labelled “action is equal to responsibility”. We used a 13 point scale to ensure that results were comparable with the shortfall scores that we obtained by subtracting action ratings from responsibility ratings as this previous method allowed scores to range −6 and 6. In subsequent analyses, we reverse scored all responses so that a positive score equated to more shortfall. The Environmental Government Judgement Scale (EGJS, Study 2 only) In Study 2, we developed our own measure to assess the relationship between perceptions of government's responsibility to conserve energy and perceptions of the government's energy conservation actions. Logic dictates that there are only three possible relationships between these two variables. Accordingly, we generated a pool of statements intended to reflect beliefs that the government's actions (i) fell short of its responsibility, (ii) were in line with its responsibility, and (iii) exceeded its responsibility. In total we generated 15 statements, 5 for each of the possible three relationships. Participants were asked to indicate their agreement/disagreement with each statement using a 7-point scale ranging from 1 (‘Strongly Disagree’) to 7 (‘Strongly Agree’). All items are listed below in the results section in Table 3. Government shortfall scenario (Study 2 only) We asked participants to report what they would do if the government's action to conserve energy fell short of its environmental responsibility. Responses were made using a 5-point scale ranging from 1 (‘Decrease my environmental efforts a lot’) to 5 (‘Increase my environmental efforts a lot’). The midpoint of the scale (i.e., 3) represented ‘make no change to my environmental efforts’. Participants were then asked to provide qualitative explanations for the answer they had given. Willingness to enact ECB In both Studies 1 and 2, participants indicated the likelihood that they would make a concerted effort to enact each ECB using a 7-point scale ranging from 1 (very unlikely) to 7 (very likely). Six out of eight of these items were adapted from the Student Environmental Behavior Scale (Markowitz, Goldberg, Ashton, & Lee, 2012, p.99). We avoided selecting items that may be dependent on people's circumstances (e.g., the extent to which they have the opportunities to use public transport, carpool, replace CFL light bulbs, etc.) and instead focused on energy conservation behaviors that in theory most people could enact. Both exploratory and confirmatory factor analysis performed on Studies 1 and 2 revealed a two factor structure (see Supplementary Materials). Factor 1 consisted of items representing effortful ECB: ‘Donate money to research projects designed to reduce carbon emissions’, ‘Start a petition to support environmental protection efforts’ and ‘Tell others about ways in which they can be environmentally friendly’. Factor 2 consisted of items representing easier ECB: ‘Turn my thermostat down by one degree’, ‘Switch off lights in unoccupied rooms’, and, ‘Avoid leaving electronic appliances in stand by modes’. In both Studies 1 and 2 both subscales had good internal reliability (Study 1: Easy: α = .74, Effortful: α = .76. Study 2: Easy: α = .70, Effortful: α = .80) and were significantly positively correlated (Study 1: r = .44, p < .01, Study 2: r = .62, p < .01). Responsibility and action and shortfall judgements Means, standard deviations, and standard errors indicating respondents' perceptions of each agent's responsibility to conserve energy and the actions each agent is perceived to be taking to conserve energy are shown in Table 1. Responsibility ratings In both Studies 1 and 2, respondents' responsibility ratings indicate that each of the agents is perceived to be at least somewhat responsible for conserving energy (means ranged from 5 to 6). Overall, participants tended to rate their own responsibility to conserve energy as less than the other agents' responsibility. To examine if these differences were statistically significant, we conducted a series of pairwise comparisons where the significance level adopted for the results was adjusted to .01 to control for the Familywise error rate. In Study 1, the results showed that participants rated their own responsibility for conserving energy as significantly less than other agents' responsibility, although the effect sizes obtained were small (the government: t(196) = −2.77, p < .01, 95% CI = −.46 to −.08, d = −0.20, industrial countries t(196) = −4.25, p < .01, CI = −0.58 to −0.21, d = −0.31, energy suppliers: t(196) = −2.44, p < .01, 95% CI = −0.47 to −0.05, d = −0.20, industrial factories: t(196) = −3.87, p < .01, 95% CI = −0.60 to −0.19, d = −0.32, and other individual consumers: t(196) = 1.90, NS, 95% CI = 0.02 to 0.39, d = 0.11)3 In Study 2, participants again tended to rate their own responsibility as slightly less than other agents' responsibility, although these differences were not large enough to reach significance and the effect sizes were negligible (the government: t(107) = −79, NS, 95% CI = −0.42 to 0.18, d = −0.08, industrial countries t(107) = −1.91, NS, 95% CI = −0.58 to 0.01, d = −0.21, energy suppliers: t(107) = −0.71, NS, 95% CI = −0.39 to 0.18, d = −0.07, industrial factories: t(107) = −.24, NS, 95% CI = −0.34 to 0.27, d = −0.02 and other individual consumers: t(107) = 1.31, NS, 95% CI = −.04 to 0.21, d = 0.07). Perceptions of other agents' responsibility did not appear to diminish participants judgements of their own responsibility, in fact the more participants felt other agents were responsible the more they also felt they too were responsible (Study 1: r = .57 and Study 2: r = .55, both ps < .01). Action ratings In both Studies 1 and 2, the data show that participants were less inclined to agree that other agents are taking action to conserve energy with mean ratings for each agent typically falling between 3 and 4. This was with the exception of respondents own actions for which the mean score was higher at 5.21 in Study 1, and 5.17 in Study 2. This suggests that participants believe they are taking at least some action to conserve energy and that the action they are taking is more than that of both other individual consumers and other agents (Study 1: other individual consumers: t(196) = 7.84, p < .01, 95% CI = 0.56 to 0.94, d = .54, d = 0.75, the government: t(196) = 8.37, p < .01, 95% CI = 0.86 to 1.38, d = 0.75, industrial countries: (t(196) = 10.75, p < .01, 95% CI = 1.23 to 1.77, d = 0.96, energy suppliers: t(196) = 8.84, p < .01, 95% CI = 0.99 to 1.55, d = 0.84, and industrial factories: t(196) = 11.09, p < .01, 95% CI = 1.28 to 1.84, d = 1.00). Study 2: other individual consumers: t(107) = 3.72, p < .01, 95% CI = 0.24 to .79, d = 0.41, the government: t(107) = 5.43, p < .01, 95% CI = 0.69 to 1.49, d = 0.76, industrial countries: (t(107) = 7.95, p < .01, 95% CI = 1.25 to 2.08, d = 1.08, energy suppliers: t(107) = 5.99, p < .01, 95% CI = 0.79 to 1.57, d = 0.77 and, industrial factories: t(107) = 9.27, p < .01, 95% CI = 1.43 to 2.21, d = 1.23)). Shortfall ratings: responsibility minus action There were significant differences between each of the agents' eco responsibility ratings and their corresponding eco-actions, such that others' eco-actions fell short of their responsibility (other individual consumers: t(196) = 7.07, p < .01, 95% CI = .62 to .1.10, d = .65, the government: t(196) = 11.80, p < .01, 95% CI = 1.35 to 1.90, d = 0.65, industrial countries: (t(196) = 13.96, p < .01, 95% CI = 1.83 to 2.43, d = 1.40, energy suppliers: t(196) = 12.47, p < .01, 95% CI = 1.48 to 2.05, d = 1.19, and industrial factories: t(196) = 14.35, p < 01, 95% CI = 1.90 to 2.50, d = 1.45). Study 2: other individual consumers: t(107) = 5.67, p < .01, 95% CI = 0.56 to .1.18, d = 0.68, the government: t(107) = 9.89, p < .01, 95% CI = 1.32 to 1.98, d = 1.12, industrial countries: (t(107) = 11.07, p < .01, 95% CI = 1.97 to 2.82, d = 1.55, energy suppliers: t(107) = 9.17, p < .01, 95% CI = 1.35 to 2.09, d = 1.12, and industrial factories: t(107) = 11.55, p < 01, 95% CI = 1.91 to 2.69, d = 1.46)). Interestingly, participants perceived even their own eco-actions as falling short of their eco-responsibility (Study 1: t(196) = 2.54, p < .01, 95% CI = 0.05 to 0.42, d = 0.18, Study 2: (t(107) = 3.67, p < .01, 95% CI = 0.20 to 0.67, d = 0.32)). Nonetheless, participants own shortfall was still seen to be significantly smaller than others' shortfall (Study 1: other individual consumers: t(196) = −5.96, p < .01, 95% CI = −0.82 to −0.42, d = −0.41, the government: t(196) = −10.05, p < .01, 95% CI = −1.66 to −1.11, d = −0.84, industrial countries: (t(196) = −12.65, p < .01, 95% CI = −2.19 to −1.59, d = −1.06, energy suppliers: t(196) = −10.41, p < .01, 95% CI = −1.82 to −1.23, d = −0.91, and industrial factories: t(196) = −12.70, p < .01, 95% CI = −2.26 to −1.65, d = −1.10). Study 2: other individual consumers: t(107) = −3.09, p < .01, 95% CI = −.72 to −0.16, d = −0.31, the government: t(107) = −5.91, p < .01, 95% CI = −1.61 to −.81, d = −0.81, industrial countries: (t(107) = −8.48, p < .01, 95% CI = −2.41 to −1.49, d = −1.08, energy suppliers: t(107) = −6.04, p < .01, 95% CI = −1.71 to −.87, d = −0.89, and industrial factories: t(107) = −9.01, p < .01, 95% CI = −2.27 to −.87, d = −1.10)). Shortfall ratings (a bipolar scale, Study 2 only) The descriptive statistics for the shortfall ratings obtained using the bipolar measure are listed in Table 2 along with the results obtained from the series of one-sample t-tests which we conducted for each agent to examine if their shortfall scores were significantly different from 0. The findings show that each of the other agents' eco-actions is perceived to fall significantly short of their eco- responsibility (apart from Government that did not reach the Familywise corrected .01 significance level). In contrast, participants tended to report that their own eco-actions were either in line with, or marginally surpassed their eco- responsibility. Respondents own shortfall was rated as significantly smaller than each of the others' shortfall (other consumers: t(103) = −3.43, p < .01, 95% CI = −1.41 to −0.38, d = −0.41, government: t(103) = −2.50, p < .01, 95% CI = −1.83 to −0.21, d = −0.35, industrial countries: t(103) = −3.95, p < .01, 95% CI = −2.45 to −0.81, d = −0.54, energy suppliers: t(103) = −3.53, p < .01, 95% CI = −1.97 to −.0.55, d = −0.46, and industrial factories: t(103) = −4.78, p < .01, 95% CI = −2.65 to −1.10, d = −0.65). The Environmental Government Judgement Scale (Study 2 only) To assess the structural integrity of the EGJS we first conducted parallel analysis (PA; Horn, 1965) using the SPSS syntax developed by O'Connor (2000) to determine how many factors to extract.4 Previous studies have found that PA is one of the most accurate methods for deciding how many factors to retain (e.g., Zwick & Velicer, 1986). Only the first two eigenvalues were greater than the subsequent values, suggesting a two-factor structure solution. Accordingly, we conducted exploratory factor analysis specifying a two-factor solution. Given that we expected our subscales to be negatively correlated we selected an oblimin rotation. Bartlett's test (X02 (105) = 2718.04, p < .01) suggested that there was an adequate sample size for this analysis and the Kaiser-Meyer-Olkin (.95) test indicated that the data were suitable for factor analysis. The eigenvalue for the first factor was 9.31 and accounted for 62.07% of the variance. The eigenvalue for the second factor was 1.41 and this accounted for an additional 9.36% of the variance. All items loadings are shown in Table 3. Factor 1 is comprised of items representing the perception that the government's action either in line with or exceeds its responsibility. Factor 2 is comprised only of items representing perceptions that the government's environmental actions fall short of its responsibility. Both subscales had acceptable reliability (Actions in line with or exceeding responsibility: α = .95, Shortfall: α = .88). As predicted, the two factors were significantly negatively correlated (r = −.37, p < .01). Moreover, there was some indication of convergent validity as the government shortfall scale comprised of 5 items was significantly positively correlated with the government shortfall scale that we calculated by subtracting government action from government responsibility (r = .49, p < .01). Similarly, the subscale containing items indicating government actions were either in line with or exceeded its responsibility were significantly negatively correlated with the government shortfall scale comprised by subtracting government action from government responsibility (r = −.49, p < .01). We computed the means for each subscale and found that they were in line with our previous findings. Specifically, the mean score for government shortfall was 5.22 (SD = 1.16), indicating that respondents perceive that the government's actions fall short of its responsibility. Conversely, the mean score for the other subscale was 2.89 (SD = 1.16) indicating that the majority of participants do not tend to believe that the government's actions are either in line with or exceed its responsibility. Associations of willingness to enact ECB to responsibility, action, and shortfall The Pearson product–moment correlation coefficients between participants' willingness to enact ECB and eco-responsibility and eco-action ratings are shown in Table 4. Associations of willingness to enact ECB to responsibility In both Studies 1 and 2, willingness to enact both easy and effortful ECB was significantly positively correlated with perceptions of other agents as responsible for conserving energy (Study 1: rs ranged from .15 to .38, with rs >.20 being significant at p < .01. Study 2: rs ranged from .32 to .47, ps < .01). Associations of willingness to enact ECB to action The extent to which others agents were perceived to be taking action to conserve energy was not consistently associated with willingness to enact either easy or effortful ECB. Although, in Study 2, we found that both the ratings of government action and the ratings of industrial factories action were negatively correlated with willingness to enact easy ECB, albeit these findings did not reach statistical significance after applying the Familywise corrected significance level of .01). Associations of willingness to enact ECB to shortfall (responsibility minus action) Overall the mean score of other agents' shortfall (i.e., when their responsibility to conserve energy did not measure up to their actions taken to conserve energy) was significantly positively correlated with willingness to enact easy (Study 1: r = .26 and Study 2: r = .38, both ps < .01) but not effortful ECB (Study 1: r = .09 and Study 2: r = .16, both NS). In both Studies 1 and 2, the shortfall ratings for the government, industrial countries, energy suppliers, and factories were positively correlated with willingness to enact easy, but not effortful, ECB. Further analyses revealed positive correlations between each agent's shortfall and willingness to enact easy ECB (Study 1: rs ranged from .17 to .26, with rs >.20 being significant at p < .01, Study 2: rs ranged from .23 to .39, with rs >.25 being significant at p < .01). We then conducted further analyses excluding anyone with a score of 0 to see if this would influence the association of shortfall ratings to willingness to enact ECB. Notably, those with a score of 0 believe that the agents' actions are in line with their responsibility (i.e., they are doing what they should be). However, the obtained correlations were comparable to those displayed in Table 4.5 Associations of willingness to enact ECB to shortfall (bipolar scale) The mean score of other agents' shortfall obtained using the bipolar measure administered in Study 2 was also positively correlated with willingness to enact easy, but not effortful, ECB (r = .24, p = .016). Notably, this correlation was comparable to the correlation we obtained between other agents' shortfall as computed by subtracting action ratings from responsibility (Study 2: r = .24 vs. r = .38, Z = 1.12, NS). Further analyses revealed that the correlations between each of the agents' shortfall and willingness to enact easy ECB were positive (rs ranged from .14 to .27, albeit only those greater than .25 were significant at the after applying the Familywise corrected significance level of .01). Associations of willingness to enact ECB to environmental government judgement scale The 5-item government shortfall scale was significantly positively correlated with willingness to enact both easy and effortful ECB (respectively, rs = .43 and .45, both ps < .01). Conversely, the subscale comprised of items indicating that government's actions were either in line with or in excess of its responsibility was significantly negatively correlated with willingness to enact both easy and effortful ECB (respectively, rs = −.30 and −.21, both ps < .01). Controlling for personal responsibility did not substantially change the correlations between government shortfall and willingness to enact either easy or effortful behaviors (respectively, r(1056) = .35 and .31, both ps < .01). Government shortfall scenario (Study 2 only) The majority of participants reported that if the government's actions fell short of its responsibility they would make no changes to their environmental efforts (N = 75, 35.4%), increase their efforts a little (N = 98, 46.2%) or a lot (N = 37, 17.5%). Only 2 respondents reported that they would decrease their efforts a little (N = 1) or a lot (N = 1). Such findings replicate the positive correlations we obtained between government shortfall and willingness to enact ECB. We used the qualitative responses participants gave for their reactions to the government shortfall scenario to conduct thematic analysis using the five step process outlined by Braun and Clarke (2006) in which analysts (1) familiarize themselves with the data, (2) code it, (3) generate initial themes, (4) review these themes, and (5) define and name them. We employed an inductive approach consistent with an essentialist/realist method whereby the themes identified were strongly linked to the data. Hence, themes were largely identified at the semantic level. Table 5 shows the themes that emerged from our qualitative analysis for people that indicated they would increase, decrease or make no changes to their environmental efforts. The following themes emerged for participants who said that they would increase their efforts either a little or a lot: compensation, obligation (personal and collective), efficacy (personal and collective), importance of the environment, moral high ground, and setting an example. The following themes emerged for people who said they would make no change: limits to personal environmental contribution, belief that the environmental issues are not important, and disconnection between personal actions and the government. Of the two people who stated they would decrease their environmental efforts, both questioned why they should do more when the government was doing less. In other words, the sucker effect appeared to account for these two people's decreased environmental motivation. Perceptions of others' responsibility, actions, and shortfall Studies 1 and 2 present a first attempt to examine the wider contextual associations of ECB. The results indicate that both individuals and other agents are perceived to be at least somewhat responsible for conserving energy. Yet, other agents are seen to be taking significantly less action to conserve energy than individuals. Both other agents and individuals' eco-actions are seen as falling short of their responsibility to conserve energy and persisted irrespective of whether we measured responsibility first (as in Study 1) or action first (as in Study 2). Moreover, we consistently found that regardless of which measures of shortfall are employed there is still a perception that other agents actions are falling short of their responsibility. Notably, this discrepancy between eco-responsibility and eco-actions was significantly larger for other agents than for oneself. Associations of willingness to enact ECB to responsibility, action, and shortfall Across both Studies 1 and 2 the data show that the responsibility ascribed to others to conserve energy is associated with willingness to enact both easy and effortful ECB. In contrast, perceptions of others' eco-actions were not consistently and significantly related to either willingness to enact easy or effortful ECB. Interestingly, when other agents' actions to conserve energy were seen as falling short of their environmental responsibility, individuals were somewhat more likely to report willingness to enact easy ECB. Shortfall and willingness to enact ECB Importantly, the positive association between shortfall and willingness to enact ECB emerged consistently, despite our use of several different measures of shortfall which included (i) subtracting action ratings from responsibility ratings, (ii) using a bipolar scale (iii) measuring government shortfall using a 5 item scale, and (iv) asking participants to respond to a government shortfall scenario. As such, it seems that the findings are relatively robust and are not vulnerable to differences in measurement style. Moreover, a thematic analysis of the explanations revealed some insight into why individuals may increase their efforts when confronted with shortfall. Indeed, the themes that emerged were compatible with the conditions needed to prompt social compensation – i.e., that the goal to be obtained is both important and shared, and that any personal actions they undertake will contribute towards achieving the goal. Specifically, participants appeared to consider “caring for the environment” to be an important goal and one that they had a collective and personal obligation to help fulfill. Moreover, they believed that the environmental actions that they undertook could help make a difference. As such, the data provide some support for the idea that when others' actions are seen to be falling short of their responsibility people are more likely to be inclined to compensate for the discrepancy by increasing their own willingness to enact ECB. Of course, given that the magnitude of these correlations was generally small and only emerged between others' shortfall and willingness to enact easy (but not effortful) ECB, it seems that respondents may compensate for others' shortfall but only to a certain extent.","In both Studies 1 and 2, we found that participants appear willing to compensate for other agents' shortfall by personally enacting more easy to perform ECB. However, given the correlational design of Studies 1 and 2 we cannot infer causality. Hence in Study 3, we aimed to manipulate shortfall and examine the effect on willingness to enact ECB. We manipulated shortfall by first asking participants to provide their responsibility ratings, and then randomly assigning them to view either 10 government actions or 10 government inactions. Notably, we deliberately presented the materials in this order in an attempt to manipulate shortfall. Specifically, by presenting participants with the responsibility rating task first and then showing them the actions/inactions we intended to bring responsibility to the forefront of participants' minds. This was done to ensure that those participants who subsequently encountered the 10 government inactions perceived a shortfall to occur between ascribed responsibility to the government and its actions. We manipulated action on the basis that in Studies 1 and 2 participants typically rated the government as at least somewhat responsible for conserving energy (i.e., mean ratings were always between 5 and 6 out of a maximum of 7). Given the stability of this score we anticipated that it might be difficult to manipulate participants' responsibility perceptions. In contrast, we expected that it might be easier to manipulate participants' perceptions of government's action as we suspect that only those with a specialist interest in the environment would be familiar with environmental policies and legislation. To manipulate actions we randomly assigned participants to view either 10 environmental actions the government has taken or 10 environmental actions the government has not taken. We expected that participants in the inaction condition would score more on the government shortfall measures. Given the results of Studies 1 and 2 we hypothesized that participants that were induced to perceive high levels of government shortfall (i.e., in the inaction condition) would be significantly more willing to enact ECB than participants that were induced to perceive low levels of government shortfall (i.e., those in the action condition). While we predicted that our experimental procedure would manipulate participants' perceptions of government's shortfall we also considered it likely that in each condition there would be some people that were unaffected by the manipulation. In order to maximize the effect of our manipulation, we removed participants whose scores on our 5-item measure of government shortfall were discordant to the manipulation condition they had been allocated to, namely, (a) those participants who scored highly on government shortfall in the action condition and, (b) those participants who scored low on government shortfall in the inaction condition. Further details on this procedure are presented in the result section.","first rated the extent to which they believed the government is responsible for conserving energy. They were then randomly assigned to view either 10 statements emphasizing the action the government has taken to reduce carbon emissions/greenhouse gases or 10 statements emphasizing issues where the government has not taken action (see Appendix A). An example of a government action was, ‘The environmental actions the government has taken caused carbon pollution to fall to its lowest level in nearly 20 years’ and an example of a comparable government inaction was, ‘Recent data found that energy related carbon dioxide emissions in 2013 were 2% above the 2012 level, largely because the government has not implemented effective environmental policies.’ Each statement was displayed on a single page and the order in which the pages were presented was randomized. To ensure that participants read the statements we instructed them to ‘read this information carefully as you will be asked questions about it later’. To check if our manipulation had worked we asked participants to complete a measure of government action so that we could compute shortfall by subtracting action ratings from responsibility ratings (as in Studies 1 and 2) as well as our 5-item government shortfall judgement scale (as in Study 2). Finally, participants were asked to indicate how willing they would be to engage in ECB. Participants ~~~~~~~~~~~~ A total of 125 participants (71 males, 54 females) were recruited via Amazon's Mechanical Turk to complete the online questionnaire. All participants were USA citizens. Their ages ranged from 19 to 81 (M = 3.76, SD = 13.81). 63 participants were randomly allocated to the high shortfall (government inaction) condition and 62 participants to the low shortfall (government action) condition. Government shortfall We measured government shortfall using two methods. First, as in Studies 1 and 2, we subtracted action ratings from government ratings and second, as in Study 2, we administered the 5-item measure of government shortfall. The 5-item measure of government shortfall had excellent reliability (α = .90). Willingness to enact ECB As in Studies 1 and 2 participants completed a measure of their willingness to enact ECB. Each subscale had acceptable reliability (easy: α = .68, effortful: α = .79). Manipulation checks We expected that participants who were exposed to 10 government actions would rate government shortfall as low while participants who were exposed to 10 government inactions would rate government shortfall as high. To examine if this expectation was correct we used a median split to categorize participants' scores on the 5 item government shortfall scale as low (scores of 5 or less) or high (scores of 5.2 or more) to see whether the obtained scores were discordant to the condition participants had been allocated to (i.e. low shortfall scores to participants in the high shortfall condition and high shortfall scores to participants in the low shortfall condition). While this expectation held true for the majority of participants (N = 100), it seems that the manipulation did not work as intended for a fifth of our sample. Specifically, in the no-shortfall condition (i.e., where participants were shown 10 government actions) 12 participants still reported perceiving high levels of government shortfall. In the shortfall condition (i.e., where participants were shown 10 government inactions) 13 participants did not report perceiving high levels of shortfall. Hence, following the exclusion of these participants, the mean score of the 5-item government shortfall was 6.14 (SD = .55) in the shortfall (or inaction) condition and 3.42 (SD = .98) in the non-shortfall (or action) condition7 As an additional test of our manipulation we conducted a one way ANOVA to examine if the conditions had produced significantly different perceptions of government shortfall as assessed by subtracting government action ratings from government responsibility ratings. As intended, there was a significant effect of condition on the Shortfall difference score: F (1, 98) = 169.77, p < .01. η2 = .63) such that participants in the shortfall (or inaction) condition scored higher on government shortfall measures than participants in the non-shortfall (or action) condition (M = 3.37, SD = 1.73 vs. M = −.92, SD = 1.55). Effect of condition on willingness to enact ECB Findings from two one-way ANOVAs revealed that there was a significant effect of condition on willingness to enact both easy and effortful ECB (respectively, F (1, 98) = 3.98, p < .05. η2 = .04 and F (1, 98) = 5.29, p < .02. η2 = .05), such that those in the shortfall (or inaction) condition reported more willingness to enact ECBs than those in the non-shortfall (or action) condition (easy ECB: M = 5.97, SD = 0.96 vs. M = 5.55, SD = 1.15, effortful ECB: M = 4.44, SD = 1.46 vs. M = 3.84, SD = 1.11). These findings suggest that participants may be more willing to enact ECB when they are induced to believe that the government's environmental actions fall short of its responsibility.","The data from Study 3 demonstrate that it is possible to manipulate shortfall by highlighting discrepancies between responsibility and action by first having participants rate an agent's responsibility and then varying the information participants have about an agents actions. Our findings show that participants that are induced to perceive high levels of government shortfall are more willing to enact both easy and effortful ECB than participants that perceive lower levels of government shortfall. Importantly the data show the direction of causality between shortfall and willingness to enact ECB.","The present research presents a first attempt to examine some of the wider contextual variables that surround environmental issues, namely, how perceptions of other agents' responsibility and action contribute to individuals own environmental efforts. Moreover, by considering these variables in relation to one another we were able to make a novel contribution to the literature regarding how people react when others' eco-actions are seen as falling short of their responsibilities. In summary our findings show that several other agents are seen as responsible for conserving energy but are not perceived as taking action to do so. Respondents' willingness to enact ECB is positively correlated with other agents' responsibility ratings but not perception of their environmental actions. In circumstances where others' actions are seen as falling short of their responsibility people appear prepared to compensate for this shortfall by increasing their willingness to enact ECB. Willingness to enact ECB: why eco-responsibility matters ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We replicated findings from past research showing that (a) individuals consider themselves to be responsible for conserving energy, and (b) that this responsibility is positively associated with willingness to enact ECB (e.g., Hines et al., 1987; Kaiser & Shimoda, 1999). Our findings also complement existing quantitative and qualitative findings by showing that aside from themselves people attribute responsibility to several other agents (e.g., Lorenzoni et al., 2007; Hargreaves et al., 2010; Hinchliffe, 1996; Stern et al., 1985). Beyond replicating past research, we found that judgements about others eco- responsibilities are positively correlated with willingness to enact ECB. These findings present an important contribution to the literature, as they suggest that, when it comes to environmental issues, responsibility is shared but not diffused, leaving people willing to enact ECB. Indeed, the correlations of .57 and .55 between personal responsibility and others' responsibility obtained in both Studies 1 and 2 provide some support for the notion of shared environmental responsibility. Willingness to enact ECB: why others' eco-actions do not seem to matter ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In both Studies 1 and 2 our findings suggest that others are not perceived as taking much action to conserve energy, and that others' eco-actions are not significantly associated with willingness to enact ECB. Therefore, little support was provided for the effects of social norms, at least when considering other agents in the form of high-level institutions or organizations. Such results may be explained by the fact that the average person may have relatively little knowledge of the environmental actions' that organizations are taking/not taking. Equally, people may find it difficult to clearly envisage an entire organization's environmental actions rather than a singular person's actions. Indeed, on average participants neither agreed nor disagreed that others were taking action to conserve energy, thus suggesting that they were unsure about their responses. It is possible then that a positive correlation may have been found between willingness to enact ECB and other agents actions if respondents were more inclined to believe these other agents were taking action to conserve energy either because they had knowledge of such actions and/or could envisage them. It is also possible that we did not find significant correlations between other agents actions and willingness to enact ECBs because the others agents were not considered by participants to be similar to themselves. Indeed, appeals involving social norms appear to be more effective when their content involves a more similar referent group such as guests in this room rather than guests in this hotel (e.g., Goldstein et al., 2008; Reese, Lowe, & Steffgen, 2013; but see Bohner & Schlüter, 2014; Study 1 for conflicting results). In summary the null relations between other organizations' actions and willingness to enact ECB could be explained by at least three factors – knowledge of others environmental actions; difficulty of envisaging an organization's environmental actions; and perceived similarity of the organizations to oneself. Future research could independently manipulate these factors to establish if they affect the influence of an organization's actions on individuals own environmental efforts. Willingness to enact ECB: when others' actions fall short of their responsibility ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In Study 1, we found that other agents' responsibility to conserve energy was seen as falling short of their action and that this shortfall was significantly positively correlated with willingness to enact easy to perform ECB. In Study 2, we replicated these findings using several different measures of shortfall. Moreover, correlations between shortfall and willingness to enact easy to perform ECB were corroborated as almost two thirds of respondents reported that they would increase their environmental efforts should the government's actions fall short of its responsibility. Finally, in Study 3 participants who were manipulated to experience high perceptions of government shortfall were more willing to enact both easy and effortful ECB than participants who were manipulated to experience low perceptions of government shortfall. We suggest these findings are best explained by social compensation theory. Indeed, the three conditions necessary for compensation to occur were present in our data sets. First, respondents attributed responsibility to everyone including themselves, indicating that energy conservation is seen as a collective or shared goal. Second, respondents appeared unconvinced that other agents were taking sufficient action as evidenced by their relatively low action ratings. Third, participants appeared to perceive that their own efforts could go some way in compensating for other shortfall. Indeed, in Study 2, the qualitative explanations provided support for the conditions needed for social compensation to occur.","Past research has typically identified ways in which individuals can be encouraged to enact ECB. However, it has done so without taking into consideration individuals' perceptions of the “external” world. Our findings show that it is important to consider the broader social context as the extent to which other organizations' actions were seen to fall short of their environmental responsibility increased personal willingness to enact ECB. In our research, we focused on one organization in particular – the government. We consistently found that perceptions of the governments' shortfall were related to increased willingness to enact ECB. Notably, although significant, in both Studies 1 and 2 many of the correlations we obtained between others' shortfall and ECB were small in magnitude. Similarly, in Study 3, manipulating shortfall did influence ECB but the effect sizes obtained were medium rather than large. This suggests that perceptions of others' shortfall may only explain a small percentage of the variance in willingness to enact ECB. However, this does not mean that our findings are of little importance or interest. Rather, encouraging people to act in environmentally friendly ways constitutes a complex challenge involving a multitude of factors. Thus to identify an additional and previously unconsidered factor still presents an important step towards gaining a more complete understanding of people's willingness to enact ECB. Nonetheless, although the present research offers an important and novel contribution there is ample scope for future research that could improve upon and extend our findings. First, future correlational research should counterbalance the order in which the measures are presented to participants. While we varied the order in which participants completed responsibility and action judgements, in both Studies 1 and 2 we always measured ECB after participants had completed the responsibility and action judgements. While this enabled us to compare the findings from the correlational studies to Study 3 (in which the experimental design meant that we had no choice but to measure ECB after responsibility and action) it is important to establish whether this ordering had a bearing on our results. Second, to ensure that the results obtained here are not attributable to self-report biases, future research should utilize an objective measure of ECB that captures participants physically enacting a meaningful and personally costly behavior. Given the vague information participants received about each study it is unlikely that they guessed the aim of our research and influenced their answers accordingly. Nonetheless, indicating a willingness to enact ECB is different from actually enacting ECB and establishing a relationship between shortfall and an objective measure of ECB would demonstrate ecological validity. Third, while our results are in line with social compensation theory it is likely that there are additional theories that may also explain the relationship between shortfall and willingness to enact ECB. Indeed, in Study 2 participants mentioned several other reasons for increasing their efforts in response to government shortfall including taking the moral high ground or setting an example to others. Future research could examine the validity of these explanations by operationalizing them and establishing their role as mediator between others shortfall and willingness to enact ECB. Finally, there is ample scope for further research to explore the links between the public goods dilemma and the findings we obtained. Indeed, there are some parallels between both the public goods dilemma and the perception of others' actions as falling short of their responsibility. Namely, in both scenarios, in the absence of everyone involved working together (i.e., co-operating) it is still in an individual's best interest to contribute even when others do not (i.e., they defect) because the payoff is still greater than doing nothing at all. Considering our findings in this vein may help illuminate other relevant research questions. For instance, participants that initially co-operate in public good games tend to reduce their contribution in subsequent rounds when others repeatedly defect. Are the same findings likely to occur in the shortfall scenario or will people continue to compensate for “repeat offenders” shortfall? In summary, while as an explanation of the outcome we favor social compensation theory, undoubtedly further research may be required to discount alternative theories. Nevertheless, the present research indicates that perceived shortfall induces a willingness to enact energy conservation behaviors."],["Adults with cerebral palsy (CP) are known to participate in reduced levels of total physical activity. There is no information available however, regarding levels of moderate-to-vigorous physical activity (MVPA) in this population. Reduced participation in MVPA is associated with several cardiometabolic risk factors. The purpose of this study was firstly to compare levels of sedentary, light, MVPA and total activity in adults with CP to adults without CP. Secondly, the objective was to investigate the association between physical activity components, sedentary behavior and cardiometabolic risk factors in adults with CP. Adults with CP (n= 41) age 18-62. yr (mean. ±. SD. = 36.5. ±. 12.5. yr), classified in Gross Motor Function Classification System level I (n= 13), II (n= 18) and III (n= 10) participated in this study. Physical activity was measured by accelerometry in adults with CP and in age- and sex-matched adults without CP over 7 days. Anthropometric indicators of obesity, blood pressure and several biomarkers of cardiometabolic disease were also measured in adults with CP. Adults with CP spent less time in light, moderate, vigorous and total activity, and more time in sedentary activity than adults without CP (p<. 0.01 for all). Moderate physical activity was associated with waist-height ratio when adjusted for age and sex (β= -0.314, p<. 0.05). When further adjustment was made for total activity, moderate activity was associated with waist-height ratio (β= -0.538, p<. 0.05), waist circumference (β= -0.518, p<. 0.05), systolic blood pressure (β= -0.592, p<. 0.05) and diastolic blood pressure (β= -0.636, p<. 0.05). Sedentary activity was not associated with any risk factor. The findings provide evidence that relatively young adults with CP participate in reduced levels of MVPA and spend increased time in sedentary behavior, potentially increasing their risk of developing cardiometabolic disease. © 2014 The Authors. --------------------------------------------------------------------------------","Although cerebral palsy (CP) is a non-progressive disorder, it is well reported that adults with CP experience a number of secondary conditions with age. These include pain, fatigue, stiffness, and poor balance (Opheim, Jahnsen, Olsson, & Stanghelle, 2009; Van Der Slot et al., 2012), and can lead to a decline in physical functioning and loss of mobility from early adulthood. Between 30% and 52% of adults with CP reported experiencing deterioration in walking function (Bottos, Feliciangeli, Sciuto, Gericke, & Vianello, 2001; Opheim et al., 2009). Loss of mobility is most commonly observed between the age of 20 and 40 years (Bottos et al., 2001). Deterioration in physical functioning over time may lead to difficulties performing everyday activities and potentially an inactive lifestyle. Only two studies have objectively measured physical activity in adults with CP to date, with conflicting results (Nieuwenhuijsen et al., 2009; van der Slot et al., 2007). Adults with unilateral spastic CP were reported to be as active as their able-bodied peers (van der Slot et al., 2007), whereas adults with bilateral spastic CP were less active than their able-bodied peers (Nieuwenhuijsen et al., 2009). Gross motor function was a strong predictor of physical activity in adults with bilateral CP (Nieuwenhuijsen et al., 2009). Differences in the gross motor function between samples may therefore explain the discrepancy in results. Although these studies reported levels of total physical activity in adults with CP, information about time spent in individual domains of activity, such as sedentary behavior, light activity (LPA) and moderate-to-vigorous activity (MVPA), is not available. Decreased levels of MVPA and increased sedentary behavior are independently associated with risk factors for cardiovascular disease (CVD) and type II diabetes mellitus (T2DM), including obesity, dyslipidemia, hypertension, insulin resistance, hyperglycemia, and inflammatory markers (Healy, Matthews, Dunstan, Winkler, & Owen, 2011; Loprinzi et al., 2013; Luke, Dugas, Durazo-Arvizu, Cao, & Cooper, 2011; Nelson et al., 2013). The current American College of Sports Medicine (ACSM) guidelines recommend that adults accumulate 150 min of moderate activity (MPA) or 75 min of vigorous activity (VPA), in bouts of at least 10 min, per week to reduce their risk of CVD and premature mortality (Garber et al., 2011). The potentially increased risk of adults with CP developing T2DM or CVD, as a result of inactivity, has led to comparisons to people with spinal cord injury – a population known to have insulin resistance, dyslipidemia, and an elevated presence of T2DM (Bauman, 2009). Furthermore, a recent study reported that reduced MPA and increased sedentary behavior are associated with an increased risk of the metabolic syndrome in adults with impaired mobility (Peterson, Al Snih, Stoddard, Shekar, & Hurvitz, 2013). Despite this, no study has investigated habitual levels of MVPA and their association with cardiometabolic risk factors in adults with CP. The purpose of this study was two-fold. The first objective was to compare levels of physical activity and sedentary behavior between adults with and without CP. The second objective was to examine the relationship between physical activity components, sedentary behavior and cardiometabolic risk factors in adults with CP.","Adults with CP (n = 41), age 18–62 yr (mean ± SD = 36.5 ± 12.5 yr), classified in level I–III of the Gross Motor Function Classification System (GMFCS) participated in this study. This study was limited to ambulatory individuals only as accelerometers are unable to quantify physical activity in wheelchair users (Hiremath & Ding, 2011). Participants were recruited from a national centre that provides services to people with a disability and through general practitioners (GPs) nationwide. The database of the centre was searched for eligible adults, resulting in 263 letters and study invitations being sent to potential participants. Letters were sent to 1367 GPs asking them to pass on information leaflets and study invites to clients who were eligible to participate. Physical activity data from age- and sex-matched adults without CP was obtained from an institutional database of physical activity control data that was collected between 2009 and 2013. Adults with a severe intellectual disability and pregnant women were excluded from participating in this study. Participants were informed of the testing procedures before written informed consent was obtained. In the case of participants with a mild to moderate intellectual disability, their guardians also provided written informed consent. Ethical approval for this study was granted by the University of Dublin's Faculty of Health Sciences’ ethics committee and the Central Remedial Clinic's ethics committee.","Participants were classified according to the GMFCS and according to type of motor abnormality and anatomical distribution as defined by the Surveillance of Cerebral Palsy in Europe (Rosenbaum et al., 2007). Information was also obtained from participants regarding their history of CVD and T2DM and current use of medication. Height, body mass, body mass index (BMI), waist circumference (WC), waist-hip ratio (WHR), and waist-height ratio (WHtR) were measured in participants. WC was measured, on bare skin, to the nearest 0.1 cm midway between the lower rib margin and the iliac crest at the end of gentle expiration. HC was measured to the nearest 0.1 cm at the end of gentle expiration around the maximum circumference of the buttocks. The mean of two measurements was used for both WC and HC. Overweight and obesity were identified as a BMI ≥ 25 kg m−2 and ≥30 kg m−2, respectively. Central obesity was defined as WC ≥ 80 cm for women and ≥94 cm for men. Blood pressure was measured from the right arm or the least affected side, in the case of significant asymmetry, using the Omron 705 IT BP monitor. The Omron 705 IT has demonstrated excellent validity in adults under the British Hypertension Society criteria (El Assaad, Topouchian, & Asmar, 2003). The appropriate cuff size was selected for the participant based on their mid-arm circumference and placed so that the lower edge was 3 cm above the elbow crease and the bladder was centred over the brachial artery. Participants rested in a seated position with their back supported for at least 5 min before three measurements were taken at a 1–2 min interval. The average of the last two measurements was used in data analysis. Blood was drawn following an overnight (10 h) fast and processed according to standard procedures in the Biochemistry Department, St. James's Hospital, Dublin. Participants were allowed to drink water during the fast and medications for cardiovascular stability were permitted (i.e. anti-hypertensive medications). Insulin was measured by electrochemiluminescence immunoassay (Elecsys Insulin Assay, Roche Diagnostics GMBH). Enzymatic, colorimetric assays (Roche/Hitachi cobas c systems) were used to measure fasting glucose, total cholesterol (TC), high-density lipoprotein cholesterol (HDL-C) and triglycerides. TC/HDL-C ratio was calculated as TC divided by HDL-C. Low-density lipoprotein cholesterol (LDL-C) was calculated using the Friedewald equation (Friedewald, Levy, & Fredrickson, 1972). High performance liquid chromatography (Arkray/Adams A1c HA-8160 Analyser System) was used to measure glycated haemoglobin (HbA1c) and c-reactive protein (CRP) was measured by particle enhanced immunoturbidimetric assay (Roche/Hitachi cobas c systems). High risk CRP was categorized as >0.3 mg dL−1 (Pearson et al., 2003). 25-Hydroxy vitamin D (25OHD) was measured on the API 4000 LC/MS/MS system (Norwolk, Connecticut). The Homeostasis Model Assessment index (HOMA-IR) (Matthews et al., 1985) was used to evaluate insulin resistance. The metabolic syndrome was defined according to the most recent joint interim statement (Alberti et al., 2009), i.e. the presence of three or more of the following: (1) central obesity, (2) elevated triglycerides (≥150 mg/dL [1.7 mmol L−1]) or drug treatment for elevated triglycerides, (3) reduced HDL-C (<40 mg/dL [1.0 mmol L−1] in men; <50 mg/dL [1.3 mmol L−1] in women) or drug treatment for reduced HDL-C, (4) elevated blood pressure (systolic ≥ 130 and/or diastolic ≥ 85 mm Hg) or antihypertensive drug treatment, and (5) elevated fasting glucose (≥100 mg/dL) or drug treatment for elevated glucose. Physical activity was measured using the RT3 accelerometer (Stayhealthy, Inc.). The following procedure was used to collect physical activity data in adults with and without CP. All participants were asked to wear the RT3 for 7 days on their right hip (or least affected side in the case of significant asymmetry) in the midaxillary line. Participants were told to wear the RT3 for waking hours and to remove it only for bathing and swimming. Participants were asked to record the times that they removed the monitor and the activities they completed while not wearing the monitor. Vector magnitude count data was collected in 1-min epochs. Valid activity data was defined as having at least four days data, of at least 10 h wear time per day (Ward, Evenson, Vaughn, Rodgers, & Troiano, 2005). Sedentary activity was defined as <100 counts·min−1, light activity (LPA) was defined as 100–984 counts·min−1, moderate activity (MPA) was defined as 984–2341 counts·min−1, vigorous activity (VPA) was defined as >2341 counts·min−1 (Rowlands, Thomas, Eston, & Topping, 2004). Data is presented as time spent in light, moderate and vigorous activity accumulated in 1-min intervals, percentage time spent in sedentary activity (i.e. minutes spent in sedentary activity/total wear time), and mean activity counts per minute (counts·min−1). Time spent in MVPA accumulated in 10-min bouts was also calculated. One minute of activity below the moderate activity count threshold was allowed for before the bout was considered to be ended. Finally the percentage of adults with and without CP meeting the ACSM recommendation was calculated. Data analysis Statistical analysis was performed using IBM SPSS Statistics (version 19). The distribution of the data was checked for normality by the Kolmogorov–Smirnov test. The logarithm function was applied to HOMA-IR, TC/HDL-C ratio and 25OHD to transform this data to a normal distribution. Means and standard deviations were computed for each of the normally distributed continuous variables. Medians and interquartile ranges were computed for skewed data. Prevalence data is presented as percentages. Differences between continuous variables with a normal distribution were determined by independent t-tests and one-way analysis of variance (ANOVA). Differences between continuous variables with a skewed distribution were determined by Mann–Whitney U tests and Kruskal–Wallis one-way analysis of variance. Pearson's χ2 test was used for comparison of independent groups of categorical data. Multiple linear regression was used to investigate the association between cardiometabolic risk factors (i.e. BMI, WC, WHtR, WHR, systolic blood pressure, diastolic blood pressure, TC, HDL-C, TC/HDL-C ratio, LDL-C, triglycerides, plasma glucose, HbA1c, HOMA-IR, 25OHD) and physical activity components (percentage time in sedentary behavior, LPA, MPA, VPA, MVPA, mean counts·min−1). Collinearity was examined using variance inflation factors. All analyses were controlled for age and sex. To avoid multicollinearity, each independent variable (i.e. percentage time in sedentary behavior, LPA, MPA, VPA, MVPA, and mean counts·min−1) was entered in separate analyses. When systolic or diastolic blood pressure was the dependent variable of interest the analysis was additionally controlled for anti-hypertensive medication (i.e. self-report of taking any hypertension lowering medication coded as 1 if yes or 0 if no). When TC, HDL-C, LDL-C, or TC/HDL-C ratio was the dependent variable of interest the analysis was also additionally controlled for drug therapy (i.e. self-reported taking of any cholesterol medication coded as 1 if yes or 0 if no). These analyses were repeated additionally controlling for total physical activity (mean counts·min−1). Variance inflation factors <5 suggested that multicollinearity was not an issue. Logistic regression was conducted to investigate the association between the metabolic syndrome, high risk CRP (dependent variables) and each physical activity component (independent variables). Analyses were initially controlled for age and sex before additionally controlling for total activity. Statistical significance was set at p < 0.05.","The characteristics of participants with CP are presented in Table 1. Both groups (i.e. adults with and without CP) wore the RT3 for a median (IQR) of 7.0 (1.0) days. Adults with CP wore the RT3 for a median (IQR) time of 840.5 (88.5) min per day; adults without CP wore the RT3 for a mean time of 841.2 (59.3) min per day. There was no significant difference in wear time between groups. Physical activity and sedentary time in adults with and without CP ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Time spent in MPA, VPA, and MVPA, and mean counts·min−1 differed across GMFCS levels I, II and III, respectively (see Table 2). Post hoc pairwise comparisons revealed that adults in level I spent more time in VPA and MVPA than adults in level II, and more time in MPA, VPA, MVPA and total activity (mean counts·min−1) than adults in level III (p < 0.01 for all). Adults in level II spent more time in MPA than adults in level III (p < 0.05). Adults in level III spent more time in sedentary behavior than adults in level II (p < 0.05). There were no other differences in physical activity components between GMFCS levels. Adults with unilateral spastic CP spent less time in sedentary behavior and more time in MPA, VPA, MVPA and total activity than adults with bilateral CP. Data from adults with non-spastic forms of CP were excluded from analyses due to the small numbers of adults in this group (n = 4). Overall, adults with CP spent more time in sedentary behavior (p < 0.001) and less time in LPA (p < 0.001), MPA (p < 0.001), VPA (p < 0.01), MVPA (p < 0.001), and total activity (mean counts·min−1) (p < 0.001) than adults without CP (see Fig. 1). When analyzed according to GMFCS level, however, there was no difference in any physical activity outcome between adults in GMFCS level I and age- and sex-matched control participants. A trend towards a significant difference was observed for VPA (p = 0.057). Adults in level II spent more time in sedentary behavior (p < 0.001), and less time in LPA and total activity than control participants (p < 0.01). Time spent in MPA, VPA and MVPA did not differ between adults in GMFCS level II and control participants. Adults in level III spent more time in sedentary behavior and less time in LPA, MPA, VPA, MVPA, and total activity than control participants (p < 0.01 for all). Adherence to physical activity guidelines among adults with and without CP ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Ten adults with CP (24.4%) and 22 adults without CP (53.7%) met the ACSM guideline of 150 min of MVPA per week (χ2 = 7.38, p < 0.01). The number of adults meeting the guideline was significantly less across worsening gross motor function [GMFCS level I: n = 7 (53.8%); GMFCS level II: n = 3 (16.7%); GMFCS level III: n = 0 (0.0%); χ2 = 9.92, p < 0.01]. Adherence to the guideline for vigorous activity was not calculated because of the small quantity of VPA accumulated by adults with CP. Cardiometabolic risk factors in adults with CP ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Two adults (5%) were taking antihypertensive medication. Three adults (7.3%) were taking cholesterol medication. One person reported a pre-diagnosis of type I diabetes mellitus. This person was removed from all analyses of blood biomarkers of glucose metabolism (i.e. plasma glucose, HOMA-IR, HbA1c). A value for plasma glucose was missing for one person and the HbA1c assay was missing for one person due to processing errors. Cardiometabolic risk factors across GMFCS level are presented in Table 3. The prevalence of overweight/obesity was 41.5%, the prevalence of central obesity was 36.6%, the prevalence of the metabolic syndrome was 20.5% and the prevalence of high risk CRP was 14.6%. Regression models to show how physical activity components were associated with cardiometabolic risk factors are presented in Table 4. Regression analysis revealed that when age and sex were adjusted for MPA was associated with WHtR. When total activity was additionally controlled MPA remained associated with WHtR. Significant associations were also observed between MPA and WC, systolic blood pressure and diastolic blood pressure when adjusted for age, sex and total activity. Logistic regression analysis revealed that neither physical activity nor sedentary behavior was associated with the metabolic syndrome or high risk CRP.","The results of the current study indicate that adults with CP spend more time in sedentary behavior and less time in light, moderate, vigorous and total physical activity than adults without CP. Although previous studies have similarly reported reduced levels of total activity among adults and children with CP (Maher, Williams, Olds, & Lane, 2007; Nieuwenhuijsen et al., 2009) this is the first study to quantify the time that adults with CP spend in each component of physical activity. In addition, this is first study to demonstrate that a significantly smaller proportion of adults with CP meet physical activity guidelines compared to adults without CP. MPA was negatively associated with a number of cardiometabolic risk factors, suggesting that this population is at increased risk of developing chronic disease as a result of reduced levels of activity. Similar to previous studies of adults and children with CP (Bjornson, Belza, Kartin, Logsdon, & McLaughlin, 2007; Nieuwenhuijsen et al., 2009), adults in level I accumulated comparable levels of physical activity to able-bodied adults. In fact they accumulated a mean of 27.4 min of MVPA per day, resulting in over half of these adults meeting the guideline. Adults in GMFCS level II, however, spent more time in sedentary behavior, and less time in LPA and total activity than adults without CP. Interestingly, they accumulated similar levels of MPA, VPA and MVPA as control participants. It is possible that adults with mild to moderate impairments participate in MVPA while attempting to keep up with their able- bodied peers during everyday life. The combination of trying to balance the activity involved in daily life and the greater energetic cost of locomotion associated with CP (Brehm, Becher, & Harlaar, 2007) may, however, result in an energy imbalance and chronic fatigue, a primary complaint among adults with CP (Van Der Slot et al., 2012). Adults in level II may therefore reduce light activity and increase sedentary behavior in an attempt to preserve energy. Despite achieving similar levels of MVPA as able-bodied participants, it should be noted that, of concern, only 16.7% of adults in GMFCS level II met the physical activity guideline. As expected, adults in GMFCS level III participated in the least amount of physical activity; no adults in this group met the physical activity guideline. As well as accumulating little or no MVPA they spent a large proportion of the day in sedentary behavior. There is evidence in the general adult population that sedentary behavior is strongly associated with T2DM and cardiometabolic risk factors, independent of time in MVPA (Healy et al., 2011; Stamatakis, Hamer, Tilling, & Lawlor, 2012). In the current study however, sedentary behavior was not associated with any cardiometabolic risk factor. This may be because the relationship between objectively measured sedentary behavior and cardiometabolic risk factors is less consistent than that between self-reported sedentary behavior and risk factors (Stamatakis et al., 2012). Some research suggests that the type of sedentary behavior, in particular TV-viewing, rather than the volume of sedentary behavior is associated with cardiometabolic risk factors (Carson & Janssen, 2011; Stamatakis et al., 2012). This information was not captured by accelerometry in the current study. Although no association was observed between sedentary behavior and cardiometabolic risk factors, MPA was negatively associated with a number of risk factors. This is in contrast to a recent study in Dutch adults with CP that was unable to find an association between physical activity and cardiovascular risk factors (van der Slot et al., 2013). This is likely because the association between individual components of physical activity and risk factors was not investigated. The association between physical activity and risk factors may be mediated through the effect of excess adiposity on cardiometabolic disease. In the current study MPA was associated with measures of central adiposity but not BMI. Similarly, anthropometric measures of central obesity, but not BMI, are associated with cardiometabolic risk factors in adults with CP (Peterson, Haapala, & Hurvitz, 2012), probably because BMI is unable to identify excess adiposity in adults with impaired mobility (Peterson et al., 2013). Future studies investigating the effect of exercise interventions in adults with CP should evaluate changes in abdominal adiposity, which can occur in the absence of changes in BMI even in the general population (Kay & Fiatarone Singh, 2006), and subsequent changes in cardiometabolic risk. Limitations ~~~~~~~~~~~ There are a number of limitations to this study including the cross-sectional design, the inability to generalize the results to non-ambulatory adults with CP, and the lack of dietary assessment. The studied sample was also relatively small. There is currently no CP register in the Republic of Ireland and the majority of rehabilitative services are only provided up until age 18 years. International research reported that less than one-third of adults with CP are under the regular control of a rehabilitation physician (Hilberink et al., 2007). Despite every effort being made to recruit adults for this study the relatively low participation may have resulted in selection bias. In particular adults with an interest in preventive health and being physically active may have been more likely to participate. As information was not available on adults who did not respond to the study invite, comparisons cannot be made between responders and non-responders. The accelerometer used in this study was also not water-proof and therefore was unable to measure physical activity accumulated during water sports. Only 6 participants (14.6%) however, reported swimming for between 1 and 3 h per week.","The results of this study suggest that a large proportion of adults with CP do not meet physical activity guidelines. As well as spending less time in total physical activity adults with CP spent less time in MPA and VPA, and more time in sedentary behavior than their able-bodied peers. The negative association between MPA and cardiometabolic risk factors suggests that policy and intervention should be implemented to increase MPA in adults with CP in order to reduce cardiometabolic disease risk in this population.","No external funding was received for this work. An honorarium, grant, or other form of payment was not given to anyone to produce this manuscript. The authors declare no conflicts of interest."],["In this review, incorporating functional and structural MRI and DTI, with evidence gathered over the last 15 years, we examine the neural underpinnings of extraversion and neuroticism, the two major personality dimensions in Eysenck's (1967) biological model of personality. We present clear evidence that, as proposed by Eysenck nearly half-a-century ago, these traits relate meaningfully to the functioning and structure of various cortical and limbic brain regions. Specifically, there is a robust relationship between neuroticism and the functioning of several emotion processing networks in the brain, particularly during exposure to negative stimuli. The brain regions showing this association include a number of cortical regions implicated in emotion regulation, depression and anxiety, in addition to many sub-cortical/limbic regions. Currently, there are few studies directly assessing the relationship between extraversion and the cortical arousal system in the context of varying stimulations but data available so far are remarkably consistent with Eysenck's model. Future neuroimaging studies guided by relevant personality and cognitive theories, and with sufficient power to allow application of sophisticated analysis methods (for example, machine learning) are now needed to improve our understanding of the biological basis of individual differences and its application in the promotion of well-being and mental health. --------------------------------------------------------------------------------","Well before the advent of modern human brain imaging, Hans Eysenck, the visionary psychologist and the most influential personality researcher in recent history, proposed a theory (Eysenck, 1967) that went beyond description and measurement of personality and, for the first time, provided the neurophysiological causes of personality. It was unique in trying to explain extraversion and neuroticism, the two major personality dimensions in Eysenck's model (the third dimension, psychoticism, added formally later in 1975), in terms of individual differences in the functioning of aspects of the central nervous system (Eysenck, 1967). Here, we review neuroimaging evidence, gathered mainly over the last 15 years, examining the association between extraversion and/or neuroticism and brain activation/connectivity patterns elicited by a wide range of cognitive and affective tasks. We have included relevant functional magnetic resonance imaging (MRI), structural MRI and diffusion tensor imaging (DTI) studies generated in the context of Eysenck's three-factor factor model as well as Costa and McCrae's five-factor personality model (Neuroticism, Extraversion, Openness to Experience, Agreeableness & Conscientiousness). There is a reasonable correspondence between the two models for extraversion and neuroticism (Costa & McCrae, 1995). We have also considered findings relating to the remaining three factors of the five-factor model as well as those relating to psychoticism, the third dimension in Eysenck's revised model (Eysenck & Eysenck, 1975), were examined within the same study, for completeness. Cognitive processing ~~~~~~~~~~~~~~~~~~~~ Eysenck's theory proposed that the extraversion–introversion dimension (extraversion = positive affectivity, marked by pronounced engagement with the external world and characterized by high sociability, talkativeness, energy and assertiveness) is caused by variability in cortical arousal (Eysenck, 1967). Those who score low for extraversion (introverts) have lower response thresholds and are consequently more cortically aroused than those who score high for extraversion (extraverts). It further postulated an inverted U-shaped relation between cognitive performance and ‘level of arousal’, jointly determined by environmental arousal potential (defined in terms of a range of environmental manipulations and task parameters) and subject arousability as reflected in extraversion. These postulates jointly predict that, at low environmental arousal potential, extraverts' performance would be lower than that of introverts'. As environmental arousal increases, performance of extraverts should improve and they should catch up with introverts; and, at high levels of environmental arousal, extraverts should out-perform introverts with a decline in introverts' performance, until it becomes so arousing as to evoke transmarginal inhibition (TMI) (Eysenck, 1994; Gray, 1964). With evocation of TMI, introverts may experience lower arousal increments than extraverts. There is considerable support for these predictions from behavioral studies (Eysenck, 1981). Eysenck's model further postulated that level of arousal, resulting from a combination of environmental arousal and subject arousability, is mediated by activity in a ‘cortical arousal system’, modulated by reticulo-thalamic-cortical pathways (Eysenck, 1967, 1981). A circuit that seemingly corresponds to this cortical arousal system, including the dorsolateral prefrontal cortex (dlPFC) and anterior cingulate regions, has been identified in studies applying fMRI to a wide range of cognitive tasks (Duncan & Owen, 2000). Importantly, findings of an fMRI study (Kumari, ffytche, Williams, & Gray, 2004), the only one so far to test the predictions concerning extraversion and cortical activity at different cognitive loads (or stimulation levels), are remarkably consistent with Eysenck's model. Specifically, this study showed that the higher the extraversion score, the greater the change in fMRI signal in the dlPFC and anterior cingulate from rest (through 1- and 2-back) to the 3-back working memory load condition. Furthermore, also consistent with Eysenck's model, which treats neuroticism and psychoticism dimensions as independent of extraversion, the relationship between extraversion and dlPFC and anterior cingulate activity was not found for neuroticism or psychoticism in Kumari et al.’s (2004) study. Concerning neuroticism, Eysenck proposed that the neuroticism-stability dimension (neuroticism = negative affectivity, marked by emotional instability and low tolerance for stress or aversive stimuli, and characterized by anxiety, fear, moodiness, worry, envy, frustration, jealousy, and loneliness) is explained by differences in the level of activity primarily in the limbic system (Eysenck, 1967). Perhaps not surprisingly, most existing fMRI studies have examined the effects of neuroticism in implicit or explicit affect processing, emotion regulation, fear/anxiety stress induction paradigms (reviewed and discussed in the next section) rather than with pure cognitive paradigms. A very recent study, which examined the effects of personality using the five-factor model, found that decreased and increased effective connectivity within the working memory network, activated by a 3-back working memory task, were associated with high neuroticism and high conscientiousness, respectively (Dima, Friston, Stephan, & Frangou, 2015). Although these findings show a significant effect of personality in neuroplasticity, their interpretation is rather difficult because neuroticism and conscientiousness had opposite effects. The effects of conscientiousness, however, appear consistent with possible extraversion effects, since conscientiousness correlates positively with extraversion when assessed using the Eysenckian scales (Costa & McCrae, 1995). Notably, extraversion itself, as in the five-factor model, did not have any influence in this study. An important line of enquiry in relation to fMRI of neuroticism using cognitive (and other) paradigms is indicated by experimental evidence showing greater trial-to-trial variability in cognitive performance (particularly reaction time) of high neuroticism scorers, relative to low neuroticism scorers (Robinson & Tamir, 2005). This behavioral effect may reflect task- irrelevant cognitions such as worries and preoccupations in neurotic individuals. Moment- to-moment brain signal variability is also known to be present in neuroimaging studies and has important implications for fMRI activation and connectivity studies (Garrett et al., 2013). Interestingly, moment-to-moment brain signal variability correlates with less, rather than more, reaction time variability across various paradigms and samples (Garrett, Kovacevic, McIntosh, & Grady, 2011; McIntosh, Kovacevic, & Itier, 2008; Misic, Mills, Taylor, & McIntosh, 2010; Raja Beharelle, Kovacevic, McIntosh, & Levine, 2012). Despite a highly likely influence of neuroticism in this phenomenon, given its known association with reaction time variability, no published study has yet examined the effect of neuroticism in moment-to-moment/trial-to-trial variability in brain activations. Affect ~~~~~~ One of the great challenges faced by the human mind is the need to comprehend the content of other minds. Thus a rapidly increasing literature has sought to explore the psychological and neural mechanisms behind “the mental operations that underlie social interactions, including perceiving, interpreting, and generating responses to the intentions, dispositions, and behaviors of others” (Green et al., 2008), namely ‘social cognition’. One of the most fundamental means we have of making these inferences is the emotion cues that other people display. However, individual differences associated with personality traits are a key influence on the way we perceive and respond to emotion cues (Britton, Ho, Taylor, & Liberzon, 2007). Indeed, our personality, whether we tend to be shy or outgoing, anxious or contented, has a major influence on our lives and the way we interact with the world around us (Hamann & Harenski, 2004). For example, highly neurotic individuals preferentially respond to negative emotion cues, and highly extravert individuals preferentially respond to positive emotion cues (Canli et al., 2001). Personality has these effects because it comprises an integrated pattern of thinking, feeling and behaving that varies between individuals but is relatively stable within individuals over time (Suslow et al., 2010). These chronic affective styles associated with personality tune the affective system to be more sensitive towards one class of cues than to another (Cunningham, Arbuckle, Jahn, Mowrer, & Abduljalil, 2010; Fruhholz, Prinz, & Herrmann, 2010). Beyond their everyday implications for understanding normal socio- cognitive behavior, neuroticism and extraversion are of great importance as trait dimensions, because of the implications for individual vulnerability for emotion-related psychopathologies such as anxiety and mood disorders (Brandes & Bienvenu, 2006; Foster & MacQueen, 2008; Gale et al., 2011; Keller, 2004; Klein, Kotov, & Bufferd, 2011; Wright, Kelsall, Sim, Clarke, & Creamer, 2013). Adoption of cognitive neuroscience techniques undoubtedly facilitates a clearer understanding of how personality influences the way people react to emotion cues. It has even been said by some that neuroimaging might prove superior to behavioral or cognitive paradigms in characterising the effects of personality dimensions on reactivity to emotion cues (Harenski, Kim, & Hamann, 2009). The thinking here is that whereas behavioral and cognitive indices represent the combined effects of all brain activity components during a task, neuroimaging can isolate specific aspects of neural reactivity as being influenced by specific personality dimensions. Although some psychological determinants of individual variability in emotional reactivity have been determined at the behavioral level, research has only recently begun to explore the brain mechanisms that might enable this individual variability (Canli et al., 2001). This is because most prior neuroimaging studies have taken a group-based approach, in which the mechanism that determines emotional reactivity is studied in a group of healthy individuals not preselected for any specific criteria (Calder, Ewbank, & Passamonti, 2011; Hamann & Canli, 2004). Here the effect of individual differences on emotion perception at the neural level is frequently ignored or dismissed as statistical noise (Calder et al., 2011; Canli, 2004; Hamann & Canli, 2004). Yet these individual differences can exhibit remarkable stability within participants, suggesting that they are not random fluctuations, and that they relate to traits that are different between, but consistent within, individuals (Canli, 2004). With a correlational approach, individual differences in neural reactivity do not represent noise, rather they represent valuable signal that can reveal much about aspects of brain function of fundamental value to the study of social cognition (Canli & Amin, 2002; Hamann & Canli, 2004). Since the trait of neuroticism involves enhanced processing of negative emotion cues (Canli et al., 2001), one way of establishing the influence of neuroticism on patterns of brain activity is to study the neural response to negative emotions (Haas, Constable, & Canli, 2008). In the original study of neuroticism, extraversion and neural reactivity to emotion cues, Canli et al. observed that neuroticism correlated positively with neural reactivity to negative scenes in the middle frontal and temporal gyri, whilst extraversion correlated with reactivity to positive scenes in the inferior, middle and superior frontal gyri, the cingulate, the inferior and middle temporal gyri, the basal ganglia and amygdala (Canli et al., 2001). These correlations were both robust; and in the expected direction, that is greater neural reactivity to positive emotion cues was associated with high extraversion, whilst greater neural reactivity to negative emotion cues was associated with high neuroticism. Since this original study, subsequent research has confirmed and extended these findings in the visual modality, using symbols, faces or scenes. Thus, neuroticism has been associated with activity in the middle frontal gyrus (‘negative’ upsetting scenes) (Canli et al., 2001), medial PFC (sad faces) (Haas et al., 2008), anterior cingulate (‘negative’ upsetting scenes) (Haas, Omura, Constable, & Canli, 2007), temporal pole (sad faces) (Jimura, Konishi, & Miyashita, 2009), amygdala (‘negative’ upsetting scenes) (Cunningham et al., 2010; Harenski et al., 2009), and the basal ganglia (‘negative’ non-smiling/sad emotion symbols) (Bruhl, Viebke, Baumgartner, Kaffenberger, & Herwig, 2011), when perceiving facial emotion cues. Beyond passive viewing of visual emotion cues or forced-choice labelling tasks, other research on personality and affective reactivity has sought to actively direct participant's view towards or away from emotion cues, by manipulating their attention. In this line of research, neuroticism has been shown to influence activity in brain regions such as the caudate and supramarginal gyrus even before a negative emotion cue is visible, that is through directing attention towards its mere anticipation (Bruhl et al., 2011), activation of the former having also been demonstrated more broadly for the anticipation of emotional music (Salimpoor, Benovoy, Larcher, Dagher, & Zatorre, 2011). Increasing the evaluative attentional processing demands by introducing emotion-related conflict has demonstrated that in such circumstances, neuroticism is typically associated with increased amygdala activity (Fruhholz et al., 2010), which makes sense given the role of the amygdala in vigilance (Davis & Whalen, 2001), and the aforementioned association of neuroticism with a predisposition towards negative emotion cues. Indeed, according to Eysenck's own biological theory of personality, high levels of neuroticism were hypothesized to reflect increased reactivity of the limbic system (of which the amygdala is part), which then predisposes highly neurotic people to react strongly to emotionally arousing experiences and take longer to return to pre-arousal states (Eysenck, 1967, 1994). Elsewhere in the limbic system, neuroticism has also been observed to associate with degree of activity in the medial PFC in response to emotional arousal, the authors in this study suggesting that neuroticism influences emotional reactivity by enhancing neural sensitivity to high levels of emotional arousal, predisposing highly neurotic individuals to react strongly to arousing experiences (Kehoe, Toomey, Balsters, & Bokde, 2012). Some emotion-relevant brain regions show increased or decreased activity in association with neuroticism depending upon specific emotional response/processes involved. For example, highly neurotic individuals show increased activity during anticipation of painful stimulation (possibly reflecting higher vigilance, anticipatory anxiety and emotional over-arousal) and decreased activity during painful stimulation (possibly reflecting emotional blunting, avoidance/passive coping strategy, learned helplessness, etc.) in the anterior cingulate, thalamus, parahippocampal gyrus and thalamus regions (Coen et al., 2011; Kumari et al., 2007). In the temporal domain, evidence has shown that higher degrees of neuroticism are associated with a sustained haemodynamic response across time in the medial PFC to negative stimuli, not just with an instantaneous response (Haas et al., 2008). Such a link might thereby create a mechanism whereby rumination over time could facilitate the association of neuroticism with increased vulnerability for depression. The importance of taking temporal dynamics into account when using neuroimaging to study the influence of personality on affective reactivity is further underscored when distinguishing between initial reactivity to an emotional stimulus, and subsequent recovery once an emotion cue terminates or ceases to be relevant. Closer examination of the time course of neural response in the amygdala in one study demonstrated that whilst initial amygdala reactivity did not predict trait neuroticism, slower amygdala recovery from negative images after their offset did predict greater neuroticism (Schuyler et al., 2012). Because neuroticism is associated with greater perseveration of emotional events (Robinson, Wilkowski, Kirkeby, & Meier, 2006), this link between greater neuroticism and slower amygdala recovery might form the basis of the inability to alter negative mood state once established, thereby enabling another important element of the increased vulnerability for depression associated with neuroticism. Following the same logic as for neuroticism, one way of establishing the influence of extraversion on patterns of brain activity is to study the neural response to positive emotions. Thus a high degree of extraversion has been linked in this way with a greater response to positive visual emotion cues e.g., happiness, in the amygdala (Suslow et al., 2010), the basal ganglia (Canli et al., 2001), and anterior cingulate cortex (Canli, Amin, Haas, Omura, & Constable, 2004; Canli, Sivers, Whitfield, Gotlib, & Gabrieli, 2002; Haas, Omura, Amin, Constable, & Canli, 2006). These brain regions are all known for their generic roles in emotion perception (Etkin, Egner, & Kalisch, 2011; Kringelbach & Berridge, 2009; Phan, Wager, Taylor, & Liberzon, 2002; Phillips, Drevets, Rauch, & Lane, 2003; Sabatinelli et al., 2011). Functional connectivity between the anterior cingulate and inferior parietal lobule, middle frontal gyrus and right orbitofrontal gyrus may also increase with degree of extraversion during reactivity to positive emotion cues (Haas et al., 2006), which is consistent with the known functional neuroanatomy of this region (Bush, Luu, & Posner, 2000), and likely reflects increased attentional input to monitoring for positive emotion cues among other functions (Ptak, 2012; Vandenberghe & Gillebert, 2009). Subcortical regions such as the thalamus have also been implicated as enabling the influence of extraversion on the neural response during (anticipation of) positive emotion cues (Bruhl et al., 2011). In the context of the protective influence of high extraversion on psychological well-being (Abbott et al., 2008), a relationship with emotional reactivity in such a deep brain structure might be particularly difficult to modify in people with low extraversion, for example by psychotherapy (Wang et al., 2013). Beyond the confines of the supposed link between extraversion and positive emotional reactivity, other research suggests that extraversion in fact reflects a generic increase in social engagement irrespective of emotional valence (Ponari, Trojano, Grossi, & Conson, 2013). Along these lines, one study demonstrated a link between emotional reactivity in the fusiform gyrus to both positive and negative emotion cues in an attentional dot-probe task (Amin, Constable, & Canli, 2004), which would be consistent with that region's known role as an alerting mechanism to emotional content in social communications (Schindler, Wegrzyn, Steppacher, & Kissler, 2015). Historically, literature on the relationship between personality traits and emotion perception mechanisms has focused on visual cues, yet beyond facial cues, another crucial means of transmitting emotion cues is through prosody. By manipulating features of speech such as pitch, duration, loudness, voice quality and spectral properties (Ross, 1993), we can alter our tone of voice and thereby alter the emotion that we convey. It is an especially effective means of conveying emotional meaning in everyday contexts (Scherer, 1986), because the ability to produce, coordinate, and understand emotion signals in speech in this way is a prerequisite for fundamental attributes of social cognition including, ‘negotiating claims to power, respect, or equality, defining degrees of intimacy, showing affiliation or non-affiliation, avoiding threats, repairing interpersonal misunderstandings and so forth’ (Arndt & Janney, 1991). Indeed, in some situations, it may be the only means of expressing nonverbal emotion cues. Thus far, there has been only one prior report on the effect of individual differences in personality on the neural response to prosodic emotion cues (Bruck, Kreifelts, Kaza, Lotze, & Wildgruber, 2011). In this study, for explicit emotional prosody identification relative to control tasks (task driven activity), positive correlations were observed between neuroticism and neural responsivity in the amygdala and anterior cingulate. During prosody identification for emotional vs. neutral trials (stimulus-driven activity), a correlation was observed with the neural response in medial frontal cortex, to happy prosody. Perhaps reflecting the clearer trends for neuroticism-linked than for extraversion-linked patterns of emotional reactivity with visual emotion cues, no correlations were observed between brain activity patterns and extraversion for prosodic cues in this study. Overall, there is no doubt that existing studies provide strong support for a brain basis of individual differences in extraversion and neuroticism, and the findings of many of these studies, despite them not being specifically formulated to test (in fact many conducted in a theoretical vacuum), can be considered broadly in line with Eysenckian predictions.","That personality is biologically-based as proposed by Eysenck immediately becomes apparent when considering evidence from structural MRI. One particular overview study was able to demonstrate associations with the volume of different brain regions and four out of the ‘big five’ personality traits (DeYoung et al., 2010). Specifically, extraversion covaried with volume of medial orbitofrontal cortex, a brain region involved in processing reward information (Noonan, Kolling, Walton, & Rushworth, 2012). Neuroticism covaried with the volume of brain regions associated with threat, punishment, and negative affect, such as the dorsomedial PFC and portions of the left medial temporal lobe (Kalisch & Gerlicher, 2014; Strange & Dolan, 2006). However, greater levels of neuroticism have also been linked to reduced overall ratio of brain to intracranial volume across the brain as a whole, particularly the element of neuroticism relating to chronic experience of arousing negative emotions (Knutson, Momenan, Rawlings, Fong, & Hommer, 2001), which might possibly reflect the cumulative effects of stress reactivity (McEwen, 2006). Agreeableness (which correlates negatively with Psychoticism; Costa & McCrae, 1995) in the DeYoung study, covaried with volume in regions that process information about the intentions and mental states of other individuals, including the posterior cingulate cortex and superior temporal gyrus (Nummenmaa & Calder, 2009; Schlaffke et al., 2015). Finally, conscientiousness covaried with volume in lateral PFC in the De Young study, a region involved in planning and the voluntary control of behavior (Tanji & Hoshi, 2008). Often the seat of functional associations with personality (see Section 5), the PFC cortex may further exhibit separable patterns of association dependent on the particular personality trait in question. Specifically, extravert people may have a thinner layer of cortical gray matter ribbon in regions of the right inferior PFC and fusiform gyrus, compared to introvert people, whereas neurotic people may have a thinner cortex mantle in anterior regions of the left orbitofrontal cortex (Wright et al., 2006). The difference in hemispheric lateralization for these effects was significant in this study, perhaps reflecting literature that greater left-sided activity of the PFC has classically been associated with positive affect as found in extraversion, whereas greater right-sided PFC activity has classically been associated with negative affect as found in neuroticism (Davidson, 2003; Spielberg, Stewart, Levin, Miller, & Heller, 2008). Perhaps the greatest body of research on the structural bases of the five major personality traits has come from Voxel-Based Morphometry (VBM), a technique with two distinct advantages relative to traditional tracing methods. Firstly, it allows detection of subtle morphometric differences in brain structure that may not be discernible by visual inspection, and secondly, it allows investigation of the entire brain rather than a particular structure, in an automatic and objective manner (Scarpazza, Tognin, Frisciata, Sartori, & Mechelli, 2015). In the body of research using VBM to examine personality-related differences, one major area of focus has been the amygdala. In one of the early studies, neuroticism appeared to negatively correlate with gray matter density in the right amygdala, whereas extraversion was positively correlated with gray matter density in the left amygdala (Omura, Todd Constable, & Canli, 2005). Given what we know about the affective style of extraversion and neuroticism, this laterality pattern would again be consistent with a model of affect that posits left-lateralized hemispheric specialization for positive affect or approach-oriented behavior, and right-lateralized hemispheric specialization for negative affect or withdrawal-oriented behavior (Maxwell & Davidson, 2007; Shackman, McMenamin, Maxwell, Greischar, & Davidson, 2009). However, subsequent research has not confirmed these results for neuroticism. In later literature, correlations with neuroticism have either been positively rather than negatively correlated, and the foci of the volumetric correlation has concerned the left/bilateral rather than right amygdala (Koelsch, Skouras, & Jentschke, 2013; Mincic, 2015). Nevertheless, amygdala volume differences in one form or another might therefore be indicative of an individual's risk of depression (Whittle et al., 2014), which would certainly fit with this structure's role in detecting emotional saliency and social relevance (Adolphs, 2008, 2010; Murray, Brosch, & Sander, 2014). Other gray matter volume differences reportedly linked to degree of neuroticism have included increased volume in the cerebellum, and decreased volume in the superior frontal gyrus (Lu et al., 2014), which have been linked to negative affect coordination and regulation of negative emotions respectively (Baumann & Mattingley, 2012; Becker & Stoodley, 2013; Falquez et al., 2014; Mak, Hu, Zhang, Xiao, & Lee, 2009; Mothersill, Knee-Zaska, & Donohoe, 2016). As a whole, this body of evidence suggests that neuroticism is related to several brain regions involved in regulating negative emotions. Regarding extraversion, the major finding seems to concern personality-dependent volume differences in the orbitofrontal cortex and other PFC brain regions. Thus increasing extraversion appears to correlate positively with increased/decreased orbitofrontal cortex gray matter density/volume (Coutinho, Sampaio, Ferreira, Soares, & Goncalves, 2013; Cremers et al., 2010; Omura et al., 2005). As mentioned above, changes in amygdala volume or density have also been observed to depend on the degree of extraversion an individual exhibits (Cremers et al., 2010; Lu et al., 2014; Omura et al., 2005). Given that extraversion acts as a protective factor against development of anxiety disorders and depression, one possible explanation is that the reduced likelihood of highly extravert individuals developing an affective disorder relates to modulation of emotion processing through the orbitofrontal cortex and the amygdala, two structures often implicated in concert in patients with diagnoses of affective disorders (Blackmon et al., 2011; Kanske, Heissler, Schonfelder, & Wessa, 2012; Zald et al., 2014; Zhang et al., 2014). Other frontal lobe structures which have evidenced a relationship between extraversion and personality-dependent changes in gray matter density or volume have included the anterior cingulate cortex, inferior frontal gyrus, middle frontal gyrus, superior frontal gyrus, all of which are known to be involved in emotional and socio-cognitive processes (Coutinho et al., 2013; Cremers et al., 2010; Forsman, de Manzano, Karabanov, Madison, & Ullen, 2012; Lu et al., 2014). Taken as a whole, the structural evidence from VBM suggests that one aspect of personality differences may be individual variation in the organization of brain networks for evaluation of socially relevant stimuli, including regions involved in recognition of social-affective stimuli and social judgment.","Whilst the evidence presented in the previous sections has revealed important initial insights into the structural and functional neural correlates of personality, current knowledge of relationships between brain function and structures and personality traits is still rather limited. White matter mediates communications in the brain and is critical for the integrity of brain function. As noted above, neuroticism indexes the tendency to experience negative effect, and functional imaging has evidenced the importance of the relationship between functioning of the amygdala and PFC in response to negative emotion cues (Haas et al., 2007; Harenski et al., 2009; Hooker, Verosky, Miyakawa, Knight, & D'Esposito, 2008). It is perhaps not surprising therefore that subsequent research with Diffusion Tensor Imaging (DTI) has evidenced a positive correlation between neuroticism and a measure of loss of white matter integrity in the anterior cingulum and uncinate fasciculus, tracts that connect the PFC and amygdala (Xu & Potenza, 2012). This modulatory effect of neuroticism on the connectivity of the PFC and amygdala is in line with similar demonstrations with functional MRI measures in the context of affective reactivity (Cremers et al., 2010). However, other DTI evidence has suggested that the breakdown in white matter integrity associated with neuroticism may be more widespread, and include fibres connecting frontal, occipital, parietal and temporal lobes, tracts connecting orbitofrontal regions with limbic regions, fibre tracts connecting thalamic nuclei with the frontal lobes, and cross-hemispheric pathways including the corpus callosum (Bjornebekk et al., 2013). It was speculated by the authors that the more widespread relationship between neuroticism and white matter integrity observed in this second study might reflect its greater methodological sensitivity. The wide distribution of effects in the latter study further suggests that general rather than regionally specific processes might be driving the effects, and emphasizes the importance of collecting white matter integrity measures throughout the brain. Evidence for personality-linked patterns of functional connectivity independent of performance of a given task, comes from the analysis of resting state data, that is from the patterns of spontaneous activity observed, that are intrinsic and stable across time. Identifying brain correlates of a situation-independent personality structure requires evidence of a stable ‘default’ mode of brain functioning (Sampaio, Soares, Coutinho, Sousa, & Goncalves, 2014). Thus in terms of low-frequency fluctuations, neuroticism has been observed to correlate negatively with regional activity of the middle frontal gyrus and precuneus, whilst extraversion has been observed to correlate positively with regional activity of the striatum, precuneus and superior frontal gyrus (Kunisato et al., 2011). The former might reflect the association of neuroticism with the role of the precuneus in anticipatory fear (Kumari et al., 2007), and an inbuilt mechanism to attempt emotional regulation in order to deal with the predisposition to negative affectivity (Pan et al., 2014; Seo et al., 2014). For extraversion, the striatal correlation observed more likely reflects the importance of reward processing (Schultz, 2015), but the positive involvement of the precuneus may this time reflect a reduced need for emotion regulation. The link between low frequency oscillations in the precuneus in the resting state and degree of extraversion has subsequently been confirmed by another group (Wei et al., 2014). Seed-based correlation analysis has continued this theme of differences between neuroticism and extraversion in reward processing and need for socio-emotional regulation. Thus higher neuroticism scores have been observed to correlate with increased amygdala resting state functional connectivity with the precuneus, and decreased amygdala resting state functional connectivity with the temporal poles, insula, and superior temporal gyrus (Aghajani et al., 2014). Conversely, higher extraversion scores in this study correlated with increased amygdala resting state functional connectivity with the putamen, temporal pole, insula, and occipital cortex. These changes in amygdala resting state functional connectivity associated with neuroticism perhaps reflecting the less adaptive perception and processing of self-relevant and socio-emotional information frequently seen in neurotic individuals (Robinson, Moeller, & Fetterman, 2010), whereas the amygdala resting state functional connectivity pattern associated with extraversion perhaps reflecting the heightened reward sensitivity and enhanced socio-emotional functioning in extraverts. Placing the seed region in the precuneus or anterior cingulate has revealed that neuroticism predicts resting state functional connectivity with brain areas involved in self-evaluation and fear, whilst extraversion predicts resting state functional connectivity with brain areas involved in reward and motivation (Adelstein et al., 2011). Given that neuroticism is known to be associated with anxiety and self-consciousness (Montag, Reuter, Jurkiewicz, Markett, & Panksepp, 2013), and extraversion is implicated in gregariousness and excitement-seeking (Li et al., 2010), it seems that the major contribution of this literature is the clear consistency it evidences with known qualities about each personality domain. At a more general level, compared with less neurotic individuals, highly neurotic individuals may exhibit a whole-brain network structure resembling more of a random network, with weaker functional connections. In such highly neurotic individuals, Servaas et al. (2015) demonstrated that functional sub-networks could be delineated less clearly and the majority of these showed lower efficiency. The authors concluded that the ‘neurotic brain’ has a less than optimal functional network organization and that it shows signs of functional disconnectivity. Moreover, in high compared with low neurotic individuals, emotion and salience sub-networks seemed to have a more prominent role in the information exchange relative to the sensorimotor and cognitive control sub-networks (Servaas et al., 2015).","If we could go back to the 1990s, the easiest way to answer this question would have required going to the canteen of the Institute of Psychiatry (now known as the Institute of Psychiatry, Psychology and Neuroscience) around 10.30 am and approach Hans Eysenck. We believe he would have given insightful suggestions with a smile, just as he did when approached by the second author when she first arrived, an unknown hopeful post-doc from India, at the Institute of Psychiatry. Nonetheless, we identify some areas, which we believe will advance our understanding of the neural basis of individual differences. We believe that there is a great need to understand why some people are better at interpreting emotion cues than others, and we believe that personality effects are one such mechanism by which these individual differences occur. If we are able to demonstrate clear personality-linked differences in the way people's brains shape their response to emotion cues, a myriad of future applications will likely ensue (Bruck, Kreifelts, & Wildgruber, 2011). Can brain activation patterns be used to classify normal and dysfunctional emotional reactivity with machine learning classifiers? Can we distinguish the types of people (i.e., personality types) skilled in decoding nonverbal emotion cues from those with difficulties interpreting such cues based on typical patterns of brain activation? Pursuing an individual differences approach in the study of emotional reactivity not only holds the potential to further advance current models of social cognition; it may also guide the way to a better understanding of individual disturbances of emotion perception associated with psychiatric disorders. Here it is worth noting that in the latest incarnation of current diagnostic guidelines (Diagnostic and Statistical Manual of Mental Disorders; DSM-V), there is much greater emphasis on dimensional approaches to describing mental ill health than in previous editions (Kraemer, 2007). From a clinical perspective, a clear demonstration of a relationship between personality and emotional reactivity could be used to argue that changes in emotion perception could be harnessed as a reasonable alternative way of monitoring clinical improvement among individuals receiving treatment for personality disorders, as a ‘surrogate marker’. Pursuing other cognitive and affective functions which relate to individual differences in behavioral studies (but not considered so far by personality neuroscience researchers), and also found to be aberrant in certain psychiatric disorders, would be just as valuable. We suggest that future studies are also sufficiently powered to examine personality effects, in particular trait interactions (e.g., extraversion with/without high neuroticism), in brain responses at rest, to varying cognitive demands and emotional challenges, and rule out the influence of other factors such as gender, age and IQ (which also affect brain properties). Most existing studies of extraversion, neuroticism (reviewed earlier) and related traits (reviewed by McNaughton, Corr, & DeYoung, 2015) have been too small to examine such effects and mostly ignored them. It may be possible to achieve larger samples sizes with increased co-operation between research groups, without much additional cost (Mar, Spreng, & Deyoung, 2013). We further suggest that personality neuroscience researchers, wherever possible, utilize relevant personality theories (e.g., reinforcement sensitivity theory) (Corr, 2008) and cognitive psychology (e.g., to determine task properties, cognitive processes involved, time on task etc.) to guide their neuroimaging experiments and interpretations. Finally, we suggest that future studies should examine the influence of personality traits in trial-to-trial variability in brain signals. If confirmed, personality influences in these signals have implications for our understanding of brain basis of individual differences in personality but also for the neuroimaging community at large.","The findings reviewed above provide compelling support that, as proposed by Eysenck, individual differences in extraversion and neuroticism are related to the functioning and structure of various brain regions. Our review suggests a strong relationship between neuroticism and emotion processing neural networks, particularly during exposure to negative emotion cues, and suggests a relationship between some of the same emotion processing regions and extraversion in response to positive emotion cues. These regions include some cortical regions implicated in emotion regulation, depression and anxiety, in addition to most sub-cortical/limbic regions. At present, there are few data directly assessing the relationship between extraversion and cortical arousal system in the context of varying stimulation levels (e.g., by increasing cognitive load), but they are remarkably consistent with Eysenck's model."],["This article challenges the claim that young children's helping responses in Buttelmann, Carpenter, and Tomasello's (2009) task are based on ascribing a false belief to a mistaken agent. In our first Study 18- to 32-month old children (N = 28) were more likely to help find a toy in the false belief than in the true belief condition. In Study 2, with 54 children of the same age, we assessed the authors’ mentalist interpretation of this result against an alternative teleological interpretation that does not make the assumption of belief ascription. The data speak in favor of our alternative. Children's social competency is based more on inferences about what is likely to happen in a particular situation and on objective reasons for action than on inferences about agents’ mental states. We also discuss the need for testing serious alternative interpretations of claims about early belief understanding. --------------------------------------------------------------------------------","Buttelmann’s helping paradigm plays an important role in the discussion about when infants or children come to understand belief (Buttelmann, Carpenter, & Tomasello, 2009: BCT). In the standard false belief test children are asked to predict where an agent, who is mistaken about an object’s location, will look for it. Quite reliably only by about 4 years children answer this question correctly (Wimmer & Perner, 1983; Wellman, Cross, & Watson, 2001). In contrast, children’s looking behavior that indicates their anticipations about where the agent will search for the object provides evidence for sensitivity to false beliefs in infants as young as 18 months (Clements & Perner, 1994; Southgate, Senju, & Csibra, 2007; Thoermer, Sodian, Vuori, Perst, & Kristen, 2012). In violation of expectation paradigms evidence was found at an even younger age around 14–16 months (Onishi & Baillargeon, 2005; Surian, Caldi, & Sperber, 2007). Prolonged looking when an agent’s belief does not match the child’s own belief showed sensitivity to the agent’s belief as young as 7 months and similar ages are reported for neural signatures of representing belief (Kampis, Parise, Csibra, & Kovacs, 2015; Kovacs, Teglas, & Endress, 2010; Southgate & Vernetti, 2014). BCT provide an importantly different kind of evidence for early understanding of belief because they used helping behavior, an intentional action, as an indicator of understanding. Before expanding on the ongoing debate about the nature of young children’s false belief understanding and on our alternative interpretation of BCT’s findings we start by describing their procedure and interpretation in detail. Two experimenters E1 and E2 (E2 being the agent to be helped) engage with the child C. After a short warm up E2 discovers two boxes of different color (A and B) and opens and closes the lids of both with interest. In E2’s absence E1 shows C how the boxes can be locked and opened with a pin. E1 leaves the boxes unlocked and E2 returns excitedly with a caterpillar toy. She plays for a while with it, introducing it also to E1 and C, and puts it into box A. In the false belief (FB) condition E2 leaves again to get her keys. E1 and C “play a trick” on E2 and E1 sneakily moves the toy to box B, continuously checking on the door, and then locks both boxes with the pin. E2 returns and tries to open box A (where she thinks the toy still is). In the true belief (TB) condition E2 stays in the room and watches how E1 moves the toy to box B with mutual eye contact with E2 and C. E2 looks briefly away when E1 locks the boxes with the pin. After going to check whether the door was properly shut, E2 approaches box A and unsuccessfully tries to open it. If the child does not respond immediately, E1 suggests to C to help E2 and if needed, several prompts follow. BCT tested one group of 18-month- and another group of 30–32-month-old children and found that in each age group children in the FB condition tended to go to box B to retrieve the toy for E2, while in the TB condition they tended to help E2 to open box A. The mentalistic interpretation preferred by BCT is as follows: When E2 tries to open the locked box A the child has to figure out the reason for E2’s action. Since in the FB condition E2 thinks that the toy is still in box A she most likely is looking for her toy. Since she does not know where the toy really is the children help her find it in box B. In the TB condition E2 knows that the toy is in box B, therefore she cannot be after the toy when trying to open box A. She must be trying to open A for some other (unknown) reason. BCT’s results play a central role for theories about the cognitive basis of early theory of mind competences. Two lines of explanation are particularly prominent. The first one distinguishes between an implicit and explicit understanding (Clements & Perner, 2001; Onishi & Baillargeon,2005). Children’s looking is an indirect measure of their knowledge. They look as a consequence of their expectations and not in order to serve a purpose (They look there because the agent will go there, not in order that the agent will go there, nor in order to tell the experimenter what they are thinking). Whereas, they answer the test question posed in the traditional false belief task in order to answer the question. Appropriate responses on an indirect measure (looking) in the absence of correct responding to a direct measure (answer to question) is taken as a sign of implicit knowledge in the consciousness literature (Reingold & Merikle, 1993; applied to false belief studies: Clements & Perner, 2001). Children’s helping behavior in BCT’s study is highly relevant evidence, since it is a direct measure: children do it in order to help the experimenter. So their data speak against the implicit knowledge explanation (Carruthers, 2013, p. 145). The second main explanation of the discrepancy in when children take belief into account assumes that children have explicit knowledge from early on but cannot show it in the standard task due to processing limitations. Baillargeon, Scott, and He (2010) distinguish between spontaneous and elicited test responses. Looking time and gaze direction are spontaneous responses, while the answers to the traditional test questions are elicited responses. In particular the additional processing required by the test question is supposed to exceed younger children’s processing capacity. BCT’s finding thus poses a problem for this theory since children’s helping is elicited by E1’s verbal suggestion to help E2 (the relevant agent to whom a belief is supposed to be attributed). Hence it should be as difficult as the standard test, which it is not. To account for BCT’s data Carruthers (2013, p. 152), for instance, saw the need to amend Baillargeon et al.’s theory with assumptions from language pragmatics as proposed by Helming, Strickland, and Jacob (2014) and Helming, Strickland, and Jacob (2016). When being asked a question by the experimenter in the traditional test, children have to coordinate their third-person perspective as a listener to the story with their second person perspective when interacting with the experimenter. It is this coordination of perspectives that makes the traditional task so difficult and helping in BCT’s procedure easy since this does not necessitate such coordination. In contrast, Setoh, Scott, and Baillargeon (2016, also Scott, 2017) argue that despite requiring elicited responses BCT’s task is easier than the standard false belief task because it lacks the need for inhibiting a prevalent (reality oriented) response. Evidently, BCT’s results are of great theoretical importance for the field. For this reason we decided to have a closer look at their replicability and interpretation. Although there are quite a few demonstrations of early sensitivity to belief, hardly any of them have been replicated by different laboratories (e.g., for BCT’s task Fizke, Butterfill, & Rakoczy, 2013 found similar results but they used a noticeably different procedure). Another question, of course, concerns the interpretation of the results. Again, although there are many demonstrations of infants’ sensitivity to belief in different situations, no study that we are aware of, has yet specifically tested the more recently suggested alternative interpretations for several of the VOE and AL studies (e.g., Ruffman, 2014; Wellman, 2014, chp 8). To our knowledge the only study that tested an alternative explanation for BCT’s findings was run by Allen (2015). Allen (2015) contrasted the FB condition with a new clairvoyance condition, which was the same as the FB condition except that E2 tried to open box B, where the toy is but E2 thinks it is empty. If, so Allen reasoned, children take E2’s false belief into account E2 must be intending to open an empty box, so they should direct E2 to box A, which is empty. Children did not do this. They helped E2 find the toy as often in the clairvoyance as in the FB condition. However, Allen’s argument is not very persuasive, since trying to open a box of which one thinks that it is empty, does not mean one wants to open box A because one thinks it is empty. One might have plenty other reasons for opening it. To gain experience with BCT’s paradigm we started with a straight replication of the original TB and FB conditions.2 We soon noticed that with the FB-TB manipulation not only E2’s belief changed but the conditions had a quite different feel. One very obvious difference is the trick which E1 plays on E2 only in the FB condition. As already pointed out by Allen (2015, p.66):“Hiding the toy in the context of playing a trick not only makes the toy particularly salient but it also creates an expectation that the adult is going to return and look for the toy”. Two further potentially confounding factors pertain to ownership of the toy and to E2’s projected interest in the toy. The procedure of the FB condition (1) confirms the initial impression that the toy belongs to E2 because, for example, E1 only dares move it secretively in E2’s absence. Children this age have already an acute sense of ownership (Nancekivell, Van de Vondervoort, & Friedman, 2013). (2) The condition also leaves unquestioned that E2 takes great interest in her toy. In contrast, E2’s behavior in the TB condition signals that (a) it may not be E2’s toy since he watches E1 move the toy without complaining that they have not asked for permission and (b) that E2’s enthusiasm and interest for the toy must have waned towards the end since E2 watches as a bystander while E1 moves the toy. Ownership and interest therefore are further possible common sense reasons which could explain children’s different behavior in the two conditions. Taken together we proposed three possible factors (i.e., ownership, interest and playing a trick, but there could be many more) that provide teleological reasons for children to show a distinct helping pattern in the two conditions that are not based on belief reasoning. However, the purpose of the study is neither to investigate the specific reasons that might motivate children’s behavior nor the general nature of early pro-sociality. What we want to test is whether their responses in this task are based on belief reasoning or on some kind of teleological reasoning (as understood by Perner & Roessler, 2010). Perner and colleagues (Perner & Esken, 2015; Perner & Roessler, 2010) argue that by 9–18 months children become ‘teleologists’ able to derive an agent’s reason for an action without concern for the subjective views provided by mental states. The important feature of teleology is to see objective facts as providing the reasons for action. For instance: it starts to rain at the birthday party. The teleologist naturally perceives the need for a shelter for the birthday cake (the cake being under a shelter is preferable/more-desirable/better than it being left in the rain, which makes it a potential action goal) and uses her knowledge of where to find a shelter to bring the cake there. Although the evaluation of the cake being under a shelter as ‘better’ (or desirable) and the ‘facts’ of the shelter’s location are based on the teleologist’s subjective view, the teleologist treats these ‘subjective facts’ as objective. For the teleologist it is a simple fact that it is better to shelter the cake than leave it in the rain; it is not a matter of being considered better by some people depending on their subjective views of what is good or bad. This subjectivity is inherent in the mentalist concepts of desire (subjective view of what would be a better state, i.e., a goal) and belief (subjective view of what the facts are). To illustrate how teleology applies to BCT’s scenario we use the suggestion from above that children perceive the agent’s (E2’s) continued interest in playing with the toy. In the FB condition E2 is continuously playing with the toy (which makes her happy and thus for a pleasant, desirable situation). That E2 has to interrupt her play in order to fetch a key suggests that she will continue to play with the toy on her return. Hence, enabling her to do so will make for a better situation (she’ll be happy; a goal to be achieved) than preventing her from doing so (she’ll be nervous and grumpy). Consequently, when she is looking for the toy in the wrong box children have good reason to help her find the toy in box B to achieve a better situation (a goal). In the TB condition the agent interrupts her play with the toy and watches the child and E1 play with it and then briefly moves away to close a door. This gives less clear indication that the agent is likely to resume her play with the toy. So when she tries to open the empty box the teleologist child perceives no compelling reason to direct her to the box with the toy. To note: this explanation of the results by BCT does not depend on children making use of E2’s belief about the content of the boxes to infer what E2 wants, as BCT suggested. There is no need for mentalizing; a basic concern about what purpose (goal) E2 pursues is directly indicated by her different behavior in the early stages of the TB and FB conditions. With these assumptions outlined above BCT’s data can be explained without children having to concern themselves with E2’s beliefs or knowledge. When in the FB condition E2 returns it is very likely that she is coming back for her toy. When she is trying to open box A children recognize her error and correct her by redirecting her to her toy in box B. In the TB condition it is not so clear what E2 wants. When she tries to open box A she is likely to want to open this box – for whatever reasons – and thus children help her do so. We tested our teleological alternative against BCT’s mentalistic explanation of their data using three boxes, A, B, and C in four conditions. Two conditions, the old-FB and old-TB condition, corresponded to BCT’s FB and TB conditions, except for the presence of the third box that was not used in the test procedure, only in the training phase. The new-FB condition conformed the original FB condition except that E2 tried to open the third and empty box C. In the new- TB condition E2 behaved as in the original TB condition except for trying to open the empty box C rather than the empty box A, into which E2 had originally placed the toy. We expect that the presence of the third box will not influence children’s helping behavior and that we will replicate BCT’s original results in the old-FB and old-TB conditions. For the new-FB condition the two theories make different predictions. If children use E2’s belief and knowledge to infer what E2 wants (BCT’s theory) then children should behave in the new-FB condition in the same way as in the original TB condition: they should help open box C, since E2 knows that box C never contained her toy and she cannot be looking for it. Our teleological alternative predicts that since E2’s behavior in the early phases of the new-FB condition was the same as in the original FB condition, this should signal directly that E2 is interested in getting her toy. So children should help her find it by directing her to box B that contains the toy. For the new-TB condition the two theories make the same behavioral prediction. BCT predict that children help open C because E2 knows that it is empty. The teleological alternative predicts that children help open C because E2, who lost interest in the toy, is now interested in opening box C. One should point out, though, that teleology by itself does not yield a strict prediction of how interested children should be in opening box C. Only because we know from the results of BCT that they are under these conditions that we make this prediction. Table 1 provides an overview of the four conditions and each theory’s predictions.","The objective of this study was not primarily to see whether BCT’s result can be replicated but whether we would manage to replicate it using – within our possibilities – the same materials and procedure before we proceeded to test an alternative interpretation of the data in Experiment 2.","Overall, 45 children between 18 and 32 months (Mage = 24.47, SD = 4.08, 20 girls) participated in the study. The age range was chosen to cover the range between BCT’s youngest in their sample of 18 month olds in Study 2 and their oldest in Study 1. These two groups showed about the same size of effect, which we try to replicate. A third group of 16 month olds did not show a significant effect and so we did not include this age in our study. Data were collected in the Theory of mind Child Lab of the University of Salzburg (n = 20), the Parent-Toddler Group of the University of Stirling, (n = 17) and in the Little Stars Nursery (n = 8). Seventeen children had to be excluded because of parental/teacher error (3), fussiness (10), unclear responses (2), or because they did not respond to any helping request by opening or at least touching one of the boxes (2). Compared to BCT, who excluded 23% of the 2,5-year-olds and 54% of 18–16 month-olds, our overall dropout rate was similar (35%). For child centered reasons only (fussiness and no response) BCT excluded 6.5% of the 2,5 year olds and 26% of the 16–18 months olds while we excluded 13% of the older children (26–32 months) and 41% of the younger children (18–25 months olds) of our sample. The 28 children of the final sample had a mean age of Mage = 25.64 months, (SD = 3.64, range = 20–32 months, 12 girls).","Materials were produced according to the description in Buttelmann et al. (2009). Boxes were identical in size, virtually identical in locking mechanism, handlebars and color; for the stuffed toy we also used a caterpillar which was roughly the same size (48 cm). Following David Buttelmann’s advice, the locking mechanisms of the boxes were loosened so that they would make a noise when E2 tries to open them and the caterpillar was stuffed especially with plastic foil so that it made sizzling noises when being moved. Both features may be important for grabbing children’s attention during the procedure. For the warm up and instruction phase a wooden pegboard game was used. Procedure For a detailed description of the procedure see Buttelmann et al. (2009). To replicate the original procedure we prepared transcripts (English and German) of a video provided by David Buttelmann. We then videotaped our procedure and received written feedback from him for both, the English and the German version. We implemented his detailed feedback on the protocol as well as on our realization of the procedure. As in BCT children participated either in a TB or in a FB condition. Position and color of the boxes as well as where the toy was placed first were counterbalanced. E2’s eye gaze before attempting to open a box was also counterbalanced. Test-sessions were videotaped and coded by two independent raters. As in the original study, it was coded which box a child opened or touched first. Our coding and exclusion criteria followed closely that of BCT. However, in the FB condition two children showed an interesting response pattern that was not reported by BCT but also occurred in Allen’s (2015) study with 3–5-year olds; they responded by opening the old toy box first but not to help E2 to open it but only to gleefully show E2 that the box is now empty. Immediately after doing so they opened the current toy box to show E2 that the caterpillar had been hidden there. Both children additionally verbalized “Here it is!” Although these children touched the old box first, we interpret this behavior in BCT’s favor as an understanding that E2 did not know the new location of the caterpillar. In order to avoid false negatives a “re-coding” according to this interpretation was added to the “original coding” by BCT and we will report results for both. Overall, two raters disagreed only on two trials. As a result, one child was excluded due to an unclear response. For the other child agreement could be reached through discussion with a third rater. Results & Discussion ~~~~~~~~~~~~~~~~~~~~ The 28 children needed different amounts of prompting to either touch or open a box. Eleven children spontaneously responded to E2’s nonverbal request, nine children responded to one of E1’s prompts, four children responded to one of E2’s verbal prompts and four children helped only when their parent or teacher prompted (3) or assisted (1) them. The center panel of Table 2 shows the number of children choosing either box B with the toy (two of them re-coded in the FB condition) or the empty box A. The right panel shows the original results by BCT in comparison. As one can see the proportion of children helping to open box B that contained the toy in the FB condition in our study is very high (recoded: 93%, original coding: 79%) and comparable to the original study (76%) and differs from a uniform distribution (Binomial test: re-coded: p = .002; original: p = .057). In contrast, the distribution of responses in the TB condition looked different. The proportion of the expected empty-box response was not even half (43%) in our study as opposed to the original study (81%) and the responses did not differ significantly from a uniform distribution (Binomial test: p = .79 in our case). If we compare our results in the TB condition directly with BCT’s the difference is highly significant (χ2 = 7.15, p < .01). We have no good explanation to offer for this difference except for the observation that most children appeared to be at a loss of what is being asked of them in this condition. Despite this different performance in the TB condition, we nevertheless find a trend of a different proportion of B (box with toy) responses in the two conditions: Fisher’s exact test p = .038 (relying on the recoded responses, with original coding: p = .21, both one-tailed) but two-tailed p = .077 (relying on the recoded responses, with original coding: p = .42). Hence, the data do not clearly speak against the null- hypothesis. We will take up this point again in the General Discussion. Since our study supported children’s strong tendency to help the agent in the FB condition of the original study we thought to be in a good enough position to venture a test of our alternative explanation for that condition.","The prime objective of this study was to assess how children in the FB condition infer from E2’s attempt to open an empty box that E2 must be looking for her toy. Is it critical for their inference that the agent is trying to open that box, in which she mistakenly thinks her toy is located or is their inference based on some other information? We have suggested at least three other factors that may lead to that inference: in the FB condition – in contrast to the TB condition – children are primed by the initial part that (1) the toy belongs to E2, (2) E2 has great interest in playing with her toy, and is likely to engage in a hide and seek routine (Allen 2015). Any one of these factors make children expect that E2 will look for her toy on her return. When E2 returns and tries to open the now empty box children help her find the toy, without any concern for her false belief about where the toy now is. Children do not show the same helping behavior in the TB condition since the TB condition strongly suggests that the toy does not belong to E2, that she has lost interest in the toy, and because no hide and seek routine is indicated. Hence when she tries to open the empty box fewer children assume that she is still looking for the toy and that she tries to open the empty box for some unknown reason. Consequently, fewer children direct her to the toy and, instead, help open the empty box. A new version of the FB task (new-FB) with three boxes can distinguish between the different explanations. Instead of trying to open the box that she believes contains her toy, E2 tries to open the third box which is also empty and has never contained the toy. If children rely on E2’s false belief to figure out that she is likely looking for her toy – as BCT suggested – then they should come to the conclusion that she must have some other reason for trying to open this box since she knows that it does not contain her toy. If, however, children assume that the agent is keen to play with her toy again, as we surmised, then they should also help her to find the toy in this new-FB condition. In order to tie in the results with the 2-boxes version, there will be an old-FB and an old- TB condition identical to the two boxes experiment in Study 1, except for the presence of a third box (which is involved in the training but not in the transfer in the test phase). For these two conditions we anticipate to replicate the results from Study 1. What differs in the new-TB condition is that instead of attempting to open box A the agent tries to open the third box C (as in the new-FB condition). Both theories predict the same behavior, that is, children will show the same response pattern in the new-TB as in the old-TB condition. Its purpose is to check whether the active involvement of the third box creates any deviation from what is expected.","Overall, 126 children between 18 and 32 months were tested either in the Theory of mind Child Lab of the University of Salzburg (n = 20), in different childcare institutes in the city of Salzburg (n = 87) and in Scotland (n = 19). Testing in institutes took place in a separate room and in the presence of the child’s teacher or parent. Thirty-six children (28%) had to be excluded due to parental/teacher (4) or experimenter error (4), fussiness (20), unclear responses (3) or because they did not respond to any helping request (5). Overall, 29.1% of children (20,6% of the older (28–32 months) and 37.5% of the younger (18–27 months) ones) were excluded. The final sample consists of 90 children between 18.04 and 32.82 months (M = 27.15 months, SD = 3.65, 40 girls). Thirty-seven children participated in the replication conditions (Mage = 27.17 months, SD = 3.69). Six children spontaneously responded to E2’s nonverbal request, 13 children responded to E1’s prompts, five children responded to E2’s verbal prompts, 11 children responded to their parents/teachers prompt and one child needed parental/teacher assistance. Fifty-three children participated in the new conditions (Mage = 27.13 months, SD = 3.66). Seven children spontaneously responded to E2’s nonverbal request, 26 children responded to one of E1’s prompts, eight children responded to one of E2’s verbal prompts, 10 children responded to their parents/teachers prompting and one child needed parental/teacher assistance (unfortunately, for one child video recording is missing and therefore amount of prompts cannot be reported). Materials & Procedure In this study we added a third, differently colored but otherwise identical, box. All remaining materials were the same as in Study 1. The three boxes were set up in a semi-circle in equal distance (80 cm) to each other. The distance between the child and each box was 1 m. As in Study 1 the experiment started with E2 discovering the boxes. She opened and closed the lids of all three boxes several times before leaving the room so E1 could explain the locking mechanism to the child. By doing that E1 treated all boxes identically and according to the protocol of Study 1. The procedure in the test-phase was also identical to Study 1, with the exception that E1 locked the third box with the pin without attracting the child’s attention to it and that E2 was sitting in some distance behind the three boxes in the decision-phase. Further details can be found in the supplementary materials. Design & Coding Children participated in one of the four conditions. We had two TB conditions and two FB conditions that differed according to which of the three boxes E2 tried to open (see Table 1). In neither condition the third box C was involved in the transfer of the toy from box A (old) to box B (current). Two of the conditions (old-FB and old-TB) were a replication of the original study, with the only exception that the third box was also present in the setting. In these conditions, E2 always tried to open the now empty box A. In both of the new conditions (new-FB and new-TB), E2 always tried to open the third, non-involved box C. Colours of boxes were assigned to position according to a Latin Square Design within each condition. So the location of the boxes was fully counterbalanced, as was where the toy was put first. This makes for 36 different combinations to which children were randomly assigned without using a combination twice. The direction of transfer was varied such that the toy was always transferred from box A to the box to the right, if A was the box on the right the toy was transferred to the box on the left. Eye gaze was varied such that E2 always looked at box B first, then at the left empty box (which could either be box A or C), then at the right empty box (which could again either be box C or A). Test-sessions were videotaped and coded by two raters. Again, it was coded which box a child opened or touched first. Two raters agreed on 87 of 90 trials and the disagreements were resolved through discussion with a third rater. Again, three children in the FB conditions showed a response as described in Study 1: after opening the box that E2 tried to open they showed her that it was empty and immediately proceeded to show her where the caterpillar was hidden. We re-coded these children as directing their response (primarily) at box B. We will again report results according to this re-coding as well as for the original coding. Results & Discussion ~~~~~~~~~~~~~~~~~~~~ Table 3 shows children’s responses in the four conditions. Children practically ignored the third box C (with the exception of one child) and showed a preference for box B, containing the toy, in both conditions but not significant in either (Binomial Test both ps > .11). However, the pattern of results for A-responses and B-responses is not significantly different from the results of Study 1 (χ2 (3) = 3.212, p > .36). The results of the old FB and TB conditions, which we hoped would replicate the results of the two box conditions in Study 1, indicate that the use of three boxes flattened the distribution of chosen boxes. This suggests increased error responding. Closer inspection of the data showed that children had a strong preference for approaching the center box, the one closest to E2. The choices of boxes for all four conditions were 26% left, 53% center, and 21% right box. This differed from the expected uniform distribution (χ2 (2) = 9.8, p > .007), which should have occurred since assignment of A, B, and C to box location was counterbalanced. Evidently the center box attracted children, which increased error trials. So we excluded this source of error by looking only at responses directed at the left or right box, shown in Table 4. Now the picture becomes much more accentuated. The response frequencies for the old conditions now resemble closely those of Study 1. Without the distraction of the center box Study 2 replicates Study 1, which is reassuring as it shows that the use of the third box does increase error but does not distort the results. We now turn to our core concern, children’s responses in the new-FB condition with our focus on children’s response directed at the box with the toy (B) in relation to the box, which E2 tries to open (C). According to BCT’s theory children will know that E2 believes that the toy is in box A and should therefore show the same preference for the box E2 is trying to open (C), as they do in the original TB condition. Our theory predicts, however, that children assume that E2 is looking for her toy and, therefore, direct her to the box with the toy (B), as they do in the original FB condition. The data confirm that more responses were directed at B than at C. To see whether this refutes one and supports the other theory we calculated a Bayes Factor (BF). Each theory is tested against the null hypothesis H0 of no preference between the two boxes. We use Rouder’s Bayes calculator for binomial data: http://pcl.missouri.edu/bf-binomial. For specifying the model of H1 for BCT’s theory we use the observed proportions from BCT’s original TB condition (see Table 2) since their theory lets us expect similar results to their TB condition in our new-FB condition. The model thus specifies the proportion of 30/37 for C (which the agent tries to open) vs 7/37 for B (where the toy is). The corresponding observed proportions in the new-FB condition of 6/24 for C and 18/24 for B result in BF = 0.00076, very strong evidence against BCT’s theory3 (Dienes, 2014; Lee & Wagenmakers, 2014). In contrast, our explanation predicts that children should behave in the new-FB condition as in BCT’s FB condition. Thus we use the data from their FB condition with 28/37 for B (where the toy is) vs. 9/37 for C (which the agent tries to open) to specify H1. Given the observed proportions this yields BF = 17.86 substantial evidence for our hypothesis. The same result is obtained for the reduced data set excluding middle box responses (see Table 4) with BF = .005 for BCT’s theory and BF = 9.09 for our theory. Moreover, there were significantly more B responses in the new-FB condition than in the new-TB condition in relation to C responses (Fisher’s Exact p = .031). One can argue that what matters for BCT’s mentalist claim is that the original FB condition should lead to more B-directed responses because E2 is trying to open the box she mistakenly thinks her toy is in. Whether children direct their helping at the box E2 is trying to open in the TB condition is not really of essence. In fact our data suggest that children do not have much idea of what they are supposed to do in the TB condition. Hence a fairer test might be to contrast the number of B-directed responses with responses directed at any one of the other boxes (A or C). In this case we get a highly significant difference between conditions: Fisher’s exact test p = .006.4 4 General discussion ~~~~~~~~~~~~~~~~~~~~ The overarching aim of this paper was to test an alternative interpretation of Buttelmann et al. (2009). Study 1 was a direct replication attempt to build a fundament for testing this alternative hypothesis. Like in the original study, children’s responses in the FB condition were much more often directed at the box that contained the toy than at the empty box. We could also show that this tendency was stronger in the FB than in the TB condition, albeit children not showing the expected preference for the empty box in the TB condition in our study. The theoretical thrust of our data comes from our new-FB condition in Study 2 designed to distinguish BCT’s original mentalistic explanation from our teleological alternative. BCT claimed that in order to be able to help the agent children must infer in the FB condition that the agent wants her toy as she is trying to open the box where she believes her toy to be. Whereas in the TB condition, the agent tries to open a box, which she knows is empty, she probably wants to open that box for unknown reasons. In our new-FB condition the agent tries to open box C, which she knows to be empty. Therefore, on BCT’s reasoning, she cannot be looking for her toy but must intend to open that box for some other reason. Consequently children should direct their response to this box the agent is trying to open, in analogy to the original TB condition. Our data speak strongly against this explanation. Our suggestion is that children infer what the agent wants without concern for her mental states. The procedure in the FB condition signals that she is still highly interested in her toy or is likely to engage in a hide and seek routine when she returns, therefore children will help her find her toy when she makes an error of looking into box A instead of box B. For our new-FB condition this theory, as opposed to BCT’s theory, predicts the same behavior as for the old-FB condition since the pre-test procedure is exactly the same. Children – according to our hypothesis – are fairly sure that the agent is coming back to look for her toy. However, she goes to the wrong box (A in the original and C in the new-FB condition), which provides children with a good reason to help her find the toy in box B. The data speak for this explanation much more strongly than for the one advanced by BCT. Our explanation has potential relevance for how we should look at infants’ impressive social competence. An important issue concerns the question of why such young children are so keen to help others. Paulus (2014) outlined several different classes of models of what motivates pro-social behavior in very young children. (1) Emotion-sharing-models assume that helping behavior arises as a result of emotional contagion in combination with the development of self-other-differentiation and the arising ability to respond to others’ negative emotions in a solution oriented manner, for example comforting (Hoffman, 2000; Preston & De Waal, 2002). (2) Social- interaction-models propose that children act pro-socially merely because they enjoy interacting with other people without a specific motivation to be of benefit to others (Over & Carpenter, 2009; Reingold & Merikle, 1993). (3) The social-normative-model emphasizes the role of the social environment and characterizes the emergence of helping behavior as a process of internalizing the rules of their environment (in support of this model see e.g., Hammond & Carpendale, 2012). (4) Proponents of goal-alignment-models (e.g., Kärtner, Keller, & Chaudhary, 2010; Kenward & Gredebäck, 2013) agree upon the idea of a goal contagion process by which children take over the other’s goal and consequently act as if it was their own. Goal-alignment-models are the most similar to the (5) teleological account as both propose that an understanding of goals, and not an understanding of others’ mental states, is the driving factor for helping. In Perner and Roessler’s account, however, there is no need for a process of goal contagion because teleological reasoning is based on objective facts and seeing the possibility of a desirable state should give anyone reason to make this the goal of their action. Although these models differ about children’s motivation for engaging in helping behavior, they all presuppose that children in our study know what the purpose/goal of the other person’s action is. On all five accounts children need to know this in order to (1) respond adequately to other’s emotional state, to (2) engage in behavior apt to promote good interaction, to (3) show that one is willing to help, or to (4) take on the other’s goal. The central question tested in our new conditions is whether children need a belief-desire theory to do so, or whether they can do it on the basis of what they observe the agent doing. This question is most explicitly addressed in teleology which is particularly explicit that no mental states are needed. This stands in stark contrast to how Tomasello (2014) describes the cognitive basis of cooperation and helping. Tomasello (2014) argues that an early inclination for helping stems from a specific human genetic trait, a cooperative turn in human evolution, which consists of the ability to engage in higher order mentalizing. For instance in the object choice task, where children or apes are faced with several up-side down buckets, one of which is baited, when the experimenter marks or points to one of them, very young children spontaneously look for the bait under this bucket while chimpanzees need excessive training to learn this. According to Tomasello (2014, p. 57) “the key point is that the inferences used in cooperative communication are socially recursive… In the object choice task…, the recipient infers that the communicator intends that she knows that the food is in that bucket – a socially recursive inference that great apes apparently do not make.” Our finding put into question whether such intricate, recursive mentalizing abilities are needed to explain young children’s cooperative inclinations. But, if not recursive mentalizing, what does give children this cooperative knack? Roessler and Perner (2015) and Perner and Esken (2015) have argued that an understanding of the reasons for acting (teleology) provides a more direct, hence less vulnerable, basis for cooperation than recursive mentalizing. Reasons for action consist of non-mental, objective facts including value facts, that is, what is good or bad to have. Our explanation of BCT’s results implements this approach even though we cannot be sure exactly how the children interpret the interactions. We have mentioned several possibilities. For instance, we proposed that children notice E2’s interest and emotional engagement with the toy and therefore see her playing with it as desirable (a goal). Since it is a ‘good’ thing it provides reason – not just for the agent – for everyone, who can contribute to help bring it about. So when the agent does something inadequate for achieving this goal the child teleologist has a natural inclination to help—without any need to engage in reasoning about each other’s desires and beliefs. Allen (2015) has pointed out another plausible possibility. The secretive cue in the FB condition may give children the impression that this is a game of hide and seek, in which the seeker is to find the hidden target. We know that children at this age tend to help by directing the seeker to the target. From an adult’s point of view this helping is counterproductive and misses the point of the game; yet that is what children tend to do (Gratch, 1964; Theo Wimmer anecdote in Perner, 1991, p. 153). From a teleologist’s point of view this counterproductive help makes perfect sense. The overall goal, when everybody is happy, thus a desirable, good state to be in, is for the seeker to find the target. So the teleologist child chips in to get to that goal by helping the seeker. Our result has also implications for how we should view other supposed demonstrations of early belief understanding. Of these – there are many by now – only few, if any, have been replicated in a strict sense and some, as this journal issue attests, are not easily or perhaps not at all replicable A recent number of studies report severe replication issues (see e.g., Dörrenberg, Liszkowski, & Rakoczy, 2017; Powell, Hobbs, Bardis, & Carey, 2017; Schuwerk, Priewasser, Sodian, & Perner, 2017; Yott & Poulin-Dubois, 2016). Even the data of our Study 1 did not reach the same significance criterion (two-sided p-value below .05) as the original study. But one should not conclude from that that our data provide evidence against the existence of BCT’s effect, that is, the difference between the two conditions. To see this we carried out a Bayes analysis which shows a BF = 2.81 in favor of BCT’s finding, but a BF below 3.0 is regarded as inconclusive evidence (Dienes, 2014; Lee & Wagenmakers, 2014). Nevertheless, by the sheer weight of numbers of other kinds of demonstrations, early belief understanding is widely accepted as a fact. For many of the findings alternative explanations have been proposed (e.g., Apperly & Butterfill, 2009; Fenici, 2015; Perner & Roessler, 2010; Ruffman 2014; Wellman, 2014, chp 8), but few of the pithier ones have been tested.5 Peter Carruthers (2013, p. 150), in a recent evaluation of the evidence, concluded with soothing caution: “…at present we seem warranted in tentatively endorsing the infant-mindreading hypothesis, based on its record so far,” and he added wisely: “But if it should turn out that these existing studies cannot be replicated, or if additional control experiments provide evidence of non-mentalizing mechanisms underlying the results, then the situation may yet reverse itself.” The present study heralds this reversal."],["The recent interest in epigenetics within mental health research, from a developmental perspective, stems from the potential of DNA methylation to index both exposure to adversity and vulnerability for mental health problems. Genome-wide technology has facilitated epigenome-wide association studies (EWAS), permitting ‘hypothesis-free’ examinations in relation to adversity and/or mental health problems. In EWAS, rather than focusing on a priori established candidate genes, the genome is screened for DNA methylation, thereby enabling a more comprehensive representation of variation associated with complex disease. Despite their ‘hypothesis-free’ label, however, results of EWAS are in fact conditional on several a priori hypotheses, dictated by the design of EWAS platforms as well as assumptions regarding the relevance of the biological tissue for mental health phenotypes. In this short report, we review three hidden hypotheses — and provide recommendations — that combined will be useful in designing and interpreting EWAS projects. -------------------------------------------------------------------------------- HIDDEN HYPOTHESIS 1: EWAS COVERAGE IS SUFFICIENT FOR COMPLEX PSYCHIATRIC PROBLEMS -------------------------------------------------------------------------------- Array-based platforms have become widespread in psychology research, largely due to their ease of use, relatively high through-put, and well standardised and validated pipelines for processing, quality control, and analysis techniques. In particular, the Illumina 450k and EPIC arrays feature 480 000–850 000 probes targeting nearly 99% of RefSeq genes, as well as a range of other genomic categories, such as CpG islands, shores and shelves, miRNA promoters and enhancers, where DNAm can be influenced by and/or impact transcription in distal genomic regions [11••]. Compared with the Ilumina 450k, the newer Illumina EPIC 850k array provides much greater coverage of ENCODE and FANTOM5 enhancers [12••], and shows higher genetic influence underlying DNAm probes [13]. Nevertheless, these microarrays are limited in the number of sites they can assess, and thus lack true genome- wide measurements [14]. Furthermore, during the design process of the 450k and EPIC arrays, CpG sites were chosen as potentially biologically informative based on consultation with a consortium of DNA methylation experts [15]. Whilst the coverage of genes and CpG islands on these microarrays are comprehensive, it does not represent a complete picture of methylated cytosines across the genome. Selection was, in part, based on data from a number of phenotypes (some medical in nature such as cancer), and thus is not specifically targeted to brain-based, stress-related complex mental health phenotypes. This is an important point: if a sizeable proportion of the CpG sites tested are not relevant to the phenotype of interest, the likelihood of detecting relevant results is reduced. HIDDEN HYPOTHESIS 2: PERIPHERAL TISSUE IS MEANINGFUL FOR MENTAL HEALTH PROBLEM(S) -------------------------------------------------------------------------------- The second hidden hypothesis relates to the tissue that is used to quantify DNAm. The majority of mental health research is based on DNAm profiles obtained from peripheral tissues from living persons, such as blood and saliva. When investigating outcomes such as conduct disorder or depression, however, the brain is often the main tissue of interest when it comes to mechanistic interpretations of results [16••]. To this end, research suggests that the correspondence of methylation profiles from blood and saliva to the brain is in fact quite limited, but can be higher with cross-tissue genetic influence [13,17]. This presents a critical disadvantage if the investigator would like to use the peripheral tissue as a surrogate of the central nervous system (CNS; the brain). One promising avenue is to establish DNAm as a biomarker for mental illness. A biomarker does not have to be mechanistic (i.e. CNS surrogate). Indeed, blood-based biomarkers have been used for diagnostics, predictive risk, disease monitoring and/or treatment response in cancer, cardiovascular and infectious disease [18,19]. However, even within a biomarker framework, the assumption is often that distinct peripheral tissues are interchangeable and equally suited for biomarker detection, when in fact it is highly probable that peripheral tissues themselves correspond differently to environmental adversity and/or disease state [14]. For instance, biomarkers for mental health traits (e.g. depression) may be more detected in blood than saliva, as blood is more central to inflammatory processes related to stress and disease [16••].","The last hypothesis relates to the assumption that biology can be informative to the phenotype itself. Focal phenotypes (e.g. oppositional defiant disorder, anxiety) in mental health research are often complex and multiply determined [20]. The lack of established robust biomarkers for mental health problems (e.g. [19]) may suggest that some of these traits might not strongly associate with detectable biological processes. Furthermore, effect size associations in EWASes are often very small suggesting that — while significant — distinguishing the importance of DNAm in the aetiology of the mental health phenotype may prove difficult [21••]. Perhaps unsurprisingly then, most EWAS in mental health include some form of gene ontology analysis, which queries the role of larger biological systems based on existing databases [22]. These analyses result in general statements such as ‘neurodevelopment’ or the ‘immune system’ being involved in the aetiology of a given phenotype. Whether these broad categories play indeed a substantial role in the aetiology of the mental health problem is often hard to determine given the post hoc nature of the interpretation. Relatedly, many EWASes have tried to infer downstream effects of observed variation in DNAm such as differences in gene expression. Many of these studies find very little in terms of functional relationships, but a small number do report downstream biological associations (e.g. [21••]). In general, it has proven difficult to pinpoint EWAS-related biological relevance of observed DNAm changes, even if they are in genes which seem ‘plausible’ based on reported functionality and previous literature.","An alternative to using arrays with limited coverage is to use next-generation sequencing- based approaches to interrogate the whole methylome [21••]. However, these methodologies are high in cost and time intensive. Despite the limitations described above, pragmatic and strategic study design can maximise utility and interpretation of results of the Illumina 450k and EPIC arrays. For example, for researchers interested in targeting CpGs likely to associate with ‘brain-based’ mental illnesses, an a priori set of CpGs (e.g. a ‘systems approach’) could be isolated from the array data, which could still span thousands of loci. The suggestion is to prioritize CpGs within biological systems that are known to associate with variation in post-mortem brain samples [23••] or even structural or functioning brain imaging [24] if this is of primary interest to the investigator. The second recommendation for optimising the use of EWAS CpGs is to target those probes with underlying genetic influence — methylation quantitative trait loci (mQTLs). This approach may have the advantage that cross-tissue concordance (e.g. blood, saliva, post-mortem brain) appears higher for CpGs that show cross-tissue genetic influence [13]. Another advantage of mQTLs is that CpGs under considerable genetic influence are less affected by confounds [11••,13,25]. However, while mQTLs are a worthwhile approach, it is a relatively new area and at present, there is a small proportion of methylation sites with consistently reported mQTLs [11••,13]. Furthermore, large-scale and detailed information on tissue-specific mQTLs is still sparse.","One strategy to maximise the interpretability of EWAS projects is to examine DNAm as a biomarker for mental health problems that have mechanistic underpinning in tissues other than the brain, such as blood. A wide-range of psychiatric disorders have been associated with immune function as measured by peripheral inflammation [26]. Furthermore, there is good evidence from animal studies, and increasing evidence in humans, that peripheral inflammatory markers can affect brain areas implicated in certain psychiatric disorders [27]. Consequently, adversity-related immune processes and DNAm may be well measured in blood samples (see [28••]). For biomarkers to be useful, they must be cost effective, drawn from accessible tissue and predictive of future risk [29]. Biomarkers for brain- based disorders (e.g. depression) have proven more difficult to establish [19]. Liu et al. [30] performed an EWAS on blood tissues across 13 population-based cohorts and reported that a composite biomarker (consisting of 144 CpGs) discriminated drinkers from non- drinkers. It was thus suggested that a blood-based DNAm diagnostic test could be developed. It is important to note, however, that in addition to methodological considerations [31], the Liu et al. study was cross-sectional, thus it may prove difficult to use this specific biomarker as a predictor of future alcohol use, as the variation in DNAm may be the result of chronic drinking (i.e. reverse causality [32]). Importantly, large-scale meta-analyses based on new and growing consortia (e.g. PACE [33•]) are beginning to report consistent epigenetic effects on traits such as schizophrenia or smoking behaviour (e.g. [34,35]) which suggests that we may begin to be able to utilise this information to further optimize DNAm biomarker approaches.","Several suggestions have been put forward to address the complex nature of the biology that may underlie mental health problems. Most notably, the Research Domain Criteria (RDoC) initiative has proposed alternative approaches to study mental illness by integrating many different levels of information including genetics, neurocircuits and behaviour [36]. Methylation-based research can be integrated into an RDoC perspective. Here, researchers could employ a two-stage analysis, first investigating epigenetic effects on intermediate dimensions of mental health and then, using the results as biomarkers to query the more complex phenotypes. For, example, if externalising difficulties (e.g. ADHD, aggression) are the focal phenotype, rather than performing an EWAS directly on the disorder(s), the researchers could instead, as the first step, perform an EWAS on brain imaging endophenotypes of the externalising phenotype (e.g. [24]). In the second step, the results of the EWAS could be used to create poly-epigenetic genetic biomarker score (e.g. [28••]) to be (potentially) associated with the disorder. This type of two stage of EWAS may examine the epigenetic changes associated with antecedents of diagnosable mental health conditions, which would be could be more useful as a risk biomarker than a biomarker of the actual diagnosis.","The recent interest in epigenetics, from a developmental perspective, stems from the potential of DNA methylation to index both exposure to adversity and vulnerability for mental health problems [2••]. To this end, there has been substantial activity in examining EWASes of adversity-related disorders, such as conduct disorder [37] and psychosis [38]. Of interest, from these EWAS, DNAm in genes that underlie stress response, neurotransmitter activity and immune regulation have been identified. These preliminary findings may provide a useful framework for more in-depth investigations — potentially as CNS surrogates or biomarkers — of the biological pathogenesis of a mental health problem. However, we argue that understanding hidden hypotheses within the EWAS is an important first step in interpreting the results in relation to mental health phenotypes."],["Verbal hallucinations are often associated with pronounced feelings of anxiety, and it has also been suggested that anxiety somehow triggers them. In this paper, we offer a phenomenological or 'personal-level' account of how it does so. We show how anxious anticipation of one's own thought contents can generate an experience of their being 'alien'. It does so by making an experience of thinking more like one of perceiving, resulting in an unfamiliar kind of intentional state. This accounts for a substantial subset of verbal hallucinations, which are experienced as falling within one's psychological boundaries and lacking in auditory qualities. --------------------------------------------------------------------------------","In this paper, we offer an account of the relationship between a substantial subset of verbal hallucinations (VHs) and feelings of anxiety.1 It is widely acknowledged that VHs are heterogeneous. Variables include volume, auditory quality, number of voices, degree of personification, emotional tone, thematic content, mode of address (second- or third- person), level of control over voices, and level of distress associated with them (Larøi, 2006; McCarthy-Jones et al., 2014; Nayani & David, 1996). We focus specifically on VHs that have repeated insults, threats and terms of abuse as their thematic contents. Many studies report that the majority of ‘voice hearers’ report such contents, and often only such contents.2 Although frequently associated with schizophrenia diagnoses, these experiences also arise in several other psychiatric conditions, including post-traumatic stress disorder, psychotic depression, bipolar disorder and borderline personality disorder, as well as in non-clinical subjects (Johns et al., 2014). First-person reports of abusive, insulting or threatening voices in these populations have much in common (Aleman & Larøi, 2008, p. 78), and we offer an account of VHs that is consistent with their diagnostic non-specificity. We propose that certain VHs arise due to pronounced and pervasive social anxiety, of a kind that is common to several psychiatric conditions. We begin by presenting the view that anxiety is not merely a consequence of VHs: it both triggers them and shapes their content. Then we examine a model of how this happens, according to which VHs result from an anxiety-induced failure to anticipate thoughts. We argue that lack of anticipation is neither necessary nor sufficient for VHs. Instead, we introduce the notion of an ‘emotional style’ of anticipation and focus on one such style: anxious anticipation. Our central claim is that anxious anticipation of one’s own thought contents generates VHs by making an experience of thinking that p more like one of perceiving that p.3 We will show how this serves to clarify what Stephens and Graham (2000) call the “alien quality” of VHs, something that is not always attributable to their seeming to originate in the external environment or to their having sensory properties much like those of veridical perceptions. Our account captures those VHs that are experienced as internal in origin and as lacking in auditory properties. Others, we concede, require different explanations. We conclude by briefly examining how our personal-level account ties in with subpersonal theories. Our case throughout is principally philosophical. We develop a phenomenological account of anxious anticipation, one that is both independently plausible and compatible with various empirical findings concerning VHs. However, we also draw upon first-person testimonies, in order to illustrate and provide further support for some of our claims. The main source of testimony is an Internet questionnaire study, which we conducted in collaboration with colleagues.4 Participants were invited to provide open-ended, free-text responses to questions that included “Please try to describe your voice(s) and/or voice-like experiences”, “How, if at all, are these experiences different from hearing the voice of someone who is present in the room?” and “What kinds of moods or emotions are associated with your voices?” All respondents quoted here had psychiatric diagnoses.5 Our intention is not to suggest that the phenomenology can simply be ‘read off’ first-person reports, thus comprising straightforward empirical evidence for our account. Rather, these are testimonies to be interpreted.6 The account that we offer is not only consistent with their content but aids in their interpretation, helping to make sense of otherwise puzzling experiences that people struggle to describe. It is corroborated by its ability to do so.7","VHs, especially those with unpleasant contents, are generally accompanied by depression and anxiety. It is easy to see how such an experience might cause anxious distress. However, anxiety and depression frequently arise before the onset of VHs, and anxiety is particularly prevalent among ‘voice-hearers’ in both clinical and non-clinical populations (Allen et al., 2005; Kuipers et al., 2006; Paulik, Badcock, & Maybery, 2006). It is important to distinguish two findings: (a) generalised anxiety is present before voices arise, and (b) there is heightened anxiety immediately before and during the VH experience. We focus on (b), but will also suggest a role for (a). According to Delespaul, de Vries, and van Os (2002, p. 97), anxiety is the “most prominent emotion during hallucinations and reports of anxiety intensity exceeded baseline levels before the first report of auditory hallucinations”. It has also been hypothesised that increased anxiety both triggers VHs and shapes their content, although the mechanism remains unclear (Freeman & Garety, 2003, p. 923). This view is consistent with first-person descriptions of the emotional states that immediately precede VHs: “It’s worse when I’m stressed, anxious or scared.” (#3) “When I am feeling anxious they grow stronger. When I am alone as the day goes on they get stronger.” (#6) “Fear, unsafety, scare, not knowing” (#19) “Loneliness, depression, anxiety, feeling unloved, deserted, uncared for” (#28) What is described is not simply ‘anxiety’ but ‘social anxiety’. Hence the individual often withdraws from others and may complain of feeling socially isolated and estranged (Hoffman, 2007).8 Romme, Escher, Dillon, Corstens, and Morris (2009) provide fifty detailed first-person accounts by voice hearers, which include numerous references to social vulnerability, anxiety, fear, social isolation, shame and feeling lost. Many add that their feelings of anxiety, depression and estrangement originated in distressing social relationships and traumatic events, including neglect and abuse during childhood. All of this is consistent with the prevalence of abusive and insulting voices; the content of VHs is usually mood-congruent (Larøi, 2006, p. 165). That many VHs arise in a context of social anxiety and are immediately preceded by heightened anxiety does not, in itself, imply a causal relationship. However, the fact that treating the anxiety often leads to a reduction in VH frequency and severity lends further support to the hypothesis that anxiety is causally implicated in VHs (Kuipers et al., 2006, p. 28). Even so, how anxiety might cause VHs remains thoroughly unclear. In what follows, we will offer a personal- level account of how this happens.","We think that the notion of ‘anticipation’ is central to understanding how anxiety induces VHs. However, contrary to the received view, we will argue that VHs are not attributable to a lack of anticipation but to how one anticipates. A recent phenomenological account that attributes VHs to anxiety and to a consequent lack of anticipation is offered by Gallagher (2005), who seeks to accommodate both VHs and thought insertion. (He does not say whether or how the two differ, an issue we will return to later.) We take this account as a starting point from which to develop our own position. Gallagher’s approach is premised on the view that acts of thinking incorporate experiences of anticipation. It is not that one anticipates thinking something before one thinks it. That would fall foul of an infinite regress objection: anticipating the thought that p is itself a thought with the content ‘the thought that p’, which would be anticipated by a further thought with the content ‘the thought that the thought that p’, and so on.9 What one anticipates is less determinate in content. Gallagher offers the analogy of listening to a melody, where one might not anticipate hearing a particular note before one hears it, but one has at least some sense of what will come next, as illustrated by the surprise one feels when a note is out of tune.10 By analogy, one might anticipate the thematic “gist” of a thought, the content of which is more determinate (Hoffman, 1986). According to Gallagher, if a thought were not anticipated at all, it would arrive fully formed, rather than crystallizing out of something that is congruent with it but less determinate in content. This would amount to a sense of its coming from elsewhere, like the unanticipated and fully-formed communications we receive from other people. He further proposes that such experiences could occur due to “unruly emotions such as anxiety”. In brief, anxiety disrupts anticipation and, were it to immediately precede a particular thought, that thought would “appear as if from nowhere”; it would be “sudden and unexpected” (Gallagher, 2005, pp. 194–200). This is consistent with a widespread emphasis in the VH literature on prediction failure (e.g. Frith, 1992). Of course, the breakdown of a subpersonal mechanism that predicts (or, more specifically, ‘monitors’) the generation of thoughts is distinct from a personal-level prediction failure. However, the personal-level correlate of sub-personal prediction failure is often taken to be an experience of one’s thoughts as unanticipated. As Fletcher and Frith (2009, p. 56) put it, “an inner voice is unpredictable and therefore feels alien”.11 However, Gallagher’s account faces a serious problem, as does any other account that appeals to a lack of conscious anticipation. First of all, it is arguable that lack of conscious anticipation is not sufficient for VHs. Many thoughts appear to arise unanticipated, such as a song that suddenly starts ‘playing in one’s head’ or a seemingly random thought that does not cohere with the gist of one’s thinking and may also disrupt one’s train of thought. In response, perhaps even these thoughts are anticipated to at least some degree and therefore differ from VHs. It is difficult to arbitrate between conflicting phenomenological claims here. But, whatever the case, it can be added that lack of anticipation is clearly not necessary for VHs. In short, many voice-hearers do anticipate their voices, to the extent that they may be able to solicit a voice, dialogue with it and predict the thematic content of what it will ‘say’ next. In fact, it has been claimed that a majority of voice-hearers are able to converse with their voices (e.g. Garrett & Silva, 2003, p. 449). In one influential study, 51% reported that they had at least some control over their voices, 38% that they could initiate a voice, and 21% that they could stop a voice (Nayani & David, 1996, p. 183). Some also describe a feeling of ‘personal presence’ preceding a voice. Hence we suggest that the simple contrast between anticipating and failing to anticipate is an unhelpful one; VHs do not arise due to lack of anticipation. Indeed, having a thought that is completely unanticipated might well be a fairly mundane experience, one that does not involve the relevant sense of externality. Instead, VHs arise when thoughts are anticipated in a distinctive way. All instances where one anticipates that p can be qualified in terms of the following: Determinacy of content: p can be more or less specific. Mode of anticipation: p can be anticipated as certain, uncertain, probable, improbable, doubtful, and so forth. Emotional style of anticipation: the prospect of p can be the object of a range of different emotions, such as excitement, curiosity, hope or fear. By appealing to a combination of [1] and [3], we will argue that VHs arise due to a distinctive emotional style, that of anxious anticipation. We will assume that the mode of anticipation, [2], is usually that of certainty or high probability.","We accept, as a premise, that many emotions either are intentional states or at least incorporate intentional states: R is afraid of p; S is guilty about q. This is not to imply that the intentionality of emotion is a matter of cognitive ‘judgment’ or ‘appraisal’ rather than ‘affect’ or ‘feeling’. There are various ways of arguing that some or all emotional feelings are themselves intentional, and that their objects are not restricted to one’s own bodily states (see, for example, Goldie, 2000; Prinz, 2004; Ratcliffe, 2008). We further maintain that types of emotion are not just commonly but properly associated with only certain other types of intentional state. For instance, feeling guilty about something is properly associated with remembering it but not with imagining it or anticipating its occurrence. And fearing something is properly associated with anticipating its occurrence but not with remembering that it has already occurred. One might feel guilty about something that has not happened, in a situation where one has already set the wheels in motion such that it almost certainly will happen. Even so, the guilt remains past-directed: one feels guilty about p in virtue of the fact that one remembers doing q, where q is likely to cause p. It is not psychologically impossible to fear what has already happened or to feel guilty about a merely imagined state of affairs. Nevertheless, the experience would be a strange one. In the case of guilt, one might then think ‘I am wrong to feel guilty about this’ or, alternatively, ‘maybe I am not just imagining doing it; maybe I actually did do it’. When remembering that p is associated with feeling guilty about p, it is debatable how the two relate. Perhaps there is a singular kind of intentional state, that of ‘guiltily remembering that p’. Alternatively, a distinction might be drawn between two distinct and simultaneous intentional states with the content p. Or it could be that one first remembers p and then feels guilty about p; so guilt borrows its content from a preceding intentional state. The same applies to fearing and perceiving that q. However, even if the two experiences are – according to some criterion – distinct, it is plausible to maintain that they affect each other. Were one to feel persistent, intense, recalcitrant guilt about something merely imagined, one’s imagining having done p might take on some of the qualities of memory; p would start to feel like something one had actually done. More generally, we propose the following: where ‘emotion x with content p’ is properly associated with ‘intentional state y with content p’ but not with ‘intentional state z with content p’, its association with z can result in z’s taking on some of the phenomenological characteristics of y. In extreme cases, the result is a novel kind of experience, one that is of neither y nor z. This, we will argue, is how anxiety induces VHs. It is not properly associated with one’s own thought contents and, when it is associated with them, they are experienced as the contents of a perception-like intentional state that also retains some of the features of thought. Our proposal requires further clarification. It is commonplace to think something and also feel anxious about it. Hence one might object that thoughts clearly are proper objects of anxiety. However, what we are ordinarily anxious about is p, not having the thought that p. For example, where p is ‘I might lose my job’, I am anxious about actually losing my job, not about having the thought that I might. Feeling anxious about ‘the thought that p’ is a more unusual experience. But how could this account for VHs? There is a sense in which anxiety is intrinsically ‘alienating’ or ‘externalising’. It presents its object – however determinate – as something unpleasant that one is confronted with. The object of anxiety is something that threatens, something that one feels helpless in the face of. It is important to distinguish different senses of ‘externality’ here. To say that an object of anxiety is essentially external is not to insist that it be experienced as physically external to one’s bodily boundaries. Our own bodily experiences can be objects of anxiety. A person who fears she has a serious medical condition may become increasingly anxious about certain persistent bodily sensations. And chronic illness can involve more widespread feelings of alienation from one’s body. It is encountered in a way that is strange and previously unfamiliar, as an actual or potential impediment to one’s activities and an object of anxiety, rather than something in which one has implicit ‘trust’ (Carel, 2013; van den Berg, 1966). Where an object of anxiety is physically external to oneself, one feels ‘alienated’ from it in a similar way. One could feel comfortably immersed in or uncomfortably separate from one’s physically external, interpersonal surroundings. When suffering from pronounced social anxiety, one does not simply feel physically separate from others. One experiences a different kind of separation from them; one is estranged from them, threatened by them, vulnerable and helpless. In this respect, anxiety is comparable to some experiences of pain. Consider an intense, lingering pain in one’s hand that persists independently of any external stimulus. The painful hand is not experienced as something ‘external’ to one’s body. But one feel alienated from it all the same, in the sense that one is confronted by something unpleasant, something one seeks to avoid but can do nothing about. Hence something can be experienced as external and alien or, alternatively, as internal and alien. The sense of alienation that we are concerned with here has nothing to do with perceived physical location.12 And neither does that which many voice-hearers describe, since the voices are often reported as internally located, and yet alien. We suggest that certain VHs arise when one’s own thought contents become objects of anxiety and are thus experienced as ‘alien’. This is consistent with the observation that many voice-hearers dread their voices and, more specifically, what it is that the voices ‘say’. The distressing content is something the voice-hearer is confronted with, something she might try unsuccessfully to resist, to avoid: “it’s mocking me, I hate that one […] I am left in a state of fear […]. They don’t sound like me. They are angry most of the time. I don’t like to think of mean things, I try hard not to, but the more I try not to think the more the voices get nasty” (#22). Of course, it could be that the person experiences p and is subsequently anxious about it, due to the unpleasantness of p and also to p’s seeming to originate from elsewhere. We acknowledge this, but propose that causation goes both ways: anxious anticipation of content p can also generate an experience of p as alien. To illustrate how this happens, consider various familiar experiences that involve an indeterminate, affectively charged thought content coalescing into something more determinate. Take the realisation that you have left your bag on the train. As you depart from the station, this might begin as a surge of anxiety, the content of which can be roughly characterised as ‘something is wrong’; ‘I’ve not done something’ or ‘something important is missing’. This becomes ‘I’ve left something on the train’ and then ‘I’ve left my bag on the train’, after which the repercussions of what has happened becomes progressively clearer. Indeterminate content p arises, eliciting anxiety, and one anxiously anticipates the dawning of q, where q is a more determinate form of p. (The initial experience can also arise when nothing is wrong, in which case the content sometimes remains indeterminate, and the feeling fades upon recognition that all is well.) The view that thought contents can increase in determinacy as they form is consistent with various proposed explanations of VHs. For instance, Fernyhough (2004) suggests that inner speech is more usually condensed and fragmented, and that the experience of externality is attributable to its anomalous re- expansion. And, according to the influential theory proposed by Hoffman (1986, p. 503), VHs are generated by disruption of a discourse planning process, which involves “abstract planning representations that are linked to goals and beliefs”. These give one a broad sense of what is coming next and precede more determinate contents. More generally, the emphasis that much of the VH literature places on ‘inner speech’ suggests a process of some kind whereby thoughts are converted into inner speech (Stephens & Graham, 2000, p. 81). The received view is that VHs involve experiencing one’s own ‘inner speech’ as non- self-produced, rather than one’s thoughts per se, where inner speech is construed as a medium in which only some of our thoughts appear. Hence thought content could provoke anxiety, the object of which is the subsequent content of inner speech. For current purposes, we do not need to endorse a specific theory of what happens or how it happens. All we need commit ourselves to is the claim that thought content p precedes thought content q, where q is a more determinate form of p. Now, it could be argued that the content of the thought remains the same throughout such a process, that it is simply ‘translated’ into inner speech. However, whatever the process we are referring to might turn out to consist of, we suggest that it does at least involve differing degrees of content determinacy. That this is so becomes clearer once we emphasise the emotional content of VHs. Many emotions are intentional states with contents that can be conveyed in linguistic form. However, even if one were to insist that they incorporate some kind of ‘propositional content’ from the outset, this is not the same as their incorporating inner speech. It has been argued that spoken language does not just serve to convey pre-formed emotional states but also to individuate or even partly constitute them, at least in some instances (e.g. Campbell, 1997; Colombetti, 2009). Amongst other things, language can gives an emotion a more specific content. Similar points are made by the phenomenologist Maurice Merleau-Ponty (1962, pp. 177–182), for whom speech increases the determinacy of thought (and emotion): “the most familiar thing appears indeterminate as long as we have not recalled its name”; “the clearness of language stands out from an obscure background”.13 It is informative to revisit Hoffman (1986) in the light of these reflections. His approach has a cognitive emphasis throughout. The abstract plans and goals that enable discourse planning have – it appears – a propositional structure, although their content is less specific than that of inner speech. But consider an example Hoffman uses to illustrate his point. When asked to describe where she lives, a patient says the following: Yes, I live in Connecticut. We live in a 50-year-old Tudor house. It’s a house that’s very much a home…. ah…. I live there with my husband and son. It’s a home where people are drawn to feel comfortable, walk in, let’s see…. a home that is furnished comfortably – not expensive – a home that shows very much my personality. (Hoffman, 1986, p. 506) The overall theme that constitutes a sense of ‘where things are heading’ is a consistently emotional one.14 Utterances are not just ‘mood-congruent’; what we have is the articulation of an emotion or mood, something that renders its content increasingly specific. Now, let us assume that ‘inner speech’ can play a similar role to spoken language. Thus, in the case of an abusive ‘voice’, there is an unpleasant emotional content p, which provokes anxious anticipation of a more determinate linguistic content q, one that is elicited by p and consistent with p. Anxiety is intrinsically alienating and so its object, the thought that q, is experienced as alien, as something unpleasant that one faces and is unable to avoid. Whatever forms of anticipation our thinking more usually involves, anxious anticipation of thought content is not one of them. That style of anticipation is more typical of certain affectively charged perceptual experiences. So an unfamiliar, perception-like experience of thought content arises. It might be objected that the thought content ‘I’ve left my bag on the train’ is not experienced as alien and that the process sketched here therefore fails to account for the ‘alien quality’ of VHs. But there is a crucial difference between the two. In the train case, one is anxious about the fact that one has left one’s bag on the train, not about the thought content ‘I’ve left my bag on the train’. However, in the case of a VH with the content ‘you’re a worthless piece of filth’, the thought content is itself an object of anxiety. The way it is anticipated as it coalesces thus renders it alien, something unpleasant before which one feels helpless. Hence an emphasis on lack of anticipation is misleading.15 To revisit the melody analogy, consider listening to a piece of music that involves a build-up of tension (an opera by Wagner, perhaps). One senses that something intrusive will blast in; it is on its way. And yet, when it arrives and conforms to the indeterminate expectation one had of it, it is still encountered as disruptive, as arising from elsewhere, set apart from the music that preceded it.","The approach we have outlined is consistent with empirical findings concerning (i) the thematic content of many VHs, (ii) the prevalence of anxiety, and (iii) the occurrence of similar kinds of VH experience in several psychiatric conditions and in non-clinical populations. In addition, it is consistent with many first-person testimonies and serves to further illuminate those testimonies. VH contents are sometimes explicitly described as linguistic manifestations of negative, self-directed emotional appraisals that are themselves sources of anxiety. For example: “It’s hard to describe how I could ‘hear’ a voice that wasn’t auditory; but the words used and the emotions they contained (hatred and disgust) were completely clear, distinct and unmistakeable, maybe even more so that if I had heard them aurally. [ …]. I heard the voices of demons screaming at me, telling that I was damned, that God hated me, and that I was going to hell”. (#9) This person further describes how the contents of his VHs “reflected all the judgmental attitudes I had heard from my family and church”. The emotional judgments are themselves feared; they are met with “anxiety/panic and incapacitating depression” (#9). Another questionnaire respondent remarks, “I hate everything about myself. [ …] I hear a voice that confirms everything I think about myself and sometimes it feels as if it is the only one that will tell me the real truth about myself”. She adds, “I can never concentrate on anything but how I am feeling and the voice I hear”. Thus, there is a combination of distressing, self-directed emotional appraisals and heightened attentiveness towards them; she waits for them to arrive.16 The content of her ‘voice’ is specifically associated with that of negative, self-directed emotions: “I took an anti-psychotic to stop the voice and it helped a bit as I don’t have something that seems so real confirming my feelings so outrightly”17. Reference to a voice ‘confirming’ one’s feelings indicates that it not only expresses pre- formed feelings but also adds to them in some way. This is plausibly construed in terms of its giving them a more determinate content, one that is an object of anxious anticipation. Feelings of inadequacy and the like are at first indeterminate, but can take on a more determinate linguistic guise that ‘confirms’ the emotional appraisal it is congruent with and out of which it arises. Some first-person accounts further indicate a process of exactly the kind that we have described. An experience induces anxious anticipation, which then proceeds to shape the experience in question: “Due to the murmuring voice experiences being so distressing with each successive occurrence however, I grew to dread ever more either whenever another experience would appear to possibly be forthcoming or, once in the midst of an actual ongoing experience, what would come next; waiting for the next shoe to drop.” (#31) “It’s very difficult to describe the experience. Words seem to come into my mind from another source than through my own conscious effort. I find myself straining sometimes to make out the word or words, and my own anxiety about what I hear or many have heard makes it a fearful experience. I seem pulled into the experience and fear itself may shape some of the words I hear.” (#32) “I have come to recognise the voices as expressions of anxiety, perhaps even a recognition of a fear I have about myself that I am not prepared to entertain as being part of my personality.” (#34) There is anxious anticipation of what is coming next, rather than a lack of conscious anticipation (#31), which affects what is then experienced (#32). And, as indicated by (#34), anxiety about one’s own thought contents is also associated with a sense of their being alien; the anxiety is constitutive of one’s ‘disowning’ something distressing. One might object that inner speech does not have auditory properties, regardless of whether or not its content is experienced as alien, whereas VHs are auditory experiences (e.g. Wu, 2012). Furthermore, VHs are often experienced as originating in the external environment. Nothing we have said explains their auditory properties or the fact that they are physically ‘external’, as well as ‘alien’ in our sense. We acknowledge that some VH experiences most likely conform to orthodox definitions of hallucination, according to which verbal hallucinations are auditory experiences that arise in the absence of appropriate external stimuli (e.g. Frith, 1992, p. 68; Halligan & Marshall, 1996, p. 242). Several accounts of VHs emphasise such an experience (e.g. Garrett & Silva, 2003, p. 445; Leudar et al., 1997, p. 888; Wu, 2012, p. 90). However, others suggest that VHs are generally lacking in auditory qualities (e.g. Moritz & Larøi, 2008; Stephens & Graham, 2000, p. 104). As Frith (1992, p. 73) suggests, a VH can involve something more abstract than hearing a voice, “an experience of receiving a communication without any sensory component”. The two views can be reconciled by acknowledging that VHs come in both guises. David (1994) states that most but not all subjects experience voices as arising “inside the head”, while Nayani and David (1996) report that 49% of their subjects heard voices through their ears, 38% internally and 12% in both ways. Leudar et al. (1997, p. 889) state that 71% of their subjects heard only internal voices, 18% heard voices “through their ears”, and 11% heard both. Internal VHs are not always described as bereft of auditory properties but, whatever auditory properties they might have, there is a substantial phenomenological difference between these two types. This is readily apparent when we turn to first-person descriptions by individuals who experience both: “The voice inside my head sounds nothing like a real person talking to me, but rather like another person’s thoughts in my head. The other voices are to me indistinguishable from actual people talking in the same room as me.” (#1) “There are two kinds – one indistinguishable from actual voices or noises (I hear them like physical noises, and only the point of origin (for voices) or checking with other people who are present (for sounds) lets me know when they aren’t actually real. The second is like hearing someone else’s voice in my head, generally saying something that doesn’t ‘sound’ like my own thoughts or interior monologue.” (#17) Hence we suggest that it is fruitful to draw a broad, over-arching distinction between two subsets of VHs: (i) those that are experienced as external in origin and auditory in character; (ii) those that are experienced as internal in origin and lacking in auditory properties. Our account applies specifically to (ii). One might worry that what we have proposed conflicts with the observation that even internal VHs are usually described in terms of audition, rather than other kinds of perceptual experience. However, information of the relevant kind is usually received through auditory channels, at least in the absence of visual stimuli such as reading materials. So, even when it is bereft of the relevant sensory qualities, it lends itself to description in those terms. Furthermore, talk of hearing is often qualified. For example, references to the ‘sound’ of a voice and to ‘hearing’ might appear in scare quotes. Some of these internal ‘voices’ may not have any auditory qualities at all, a view that is consistent with reports of VHs in congenitally deaf subjects (e.g. Aleman & Larøi, 2008, pp. 48–9). Nevertheless, it is plausible to suggest that some of them do have audition-like properties. The claim that inner speech is sometimes or always wholly bereft of auditory properties is by no means uncontroversial. For example, Hoffman (1986) takes it to involve ‘auditory imagery’, and there may be considerable interpersonal variation too.","Even though anxiety about one’s own thought contents is unusual, one might object that it is plausibly more widespread than the kind of VH experience we seek to account for. There could well be many people with disruptive, intrusive and self-directed thought contents that provoke anxiety but are not experienced as VHs. However, up to this point we have only emphasised one of two roles played by social anxiety, that of an immediate trigger for VHs. Generalised anxiety and social isolation are also disposing factors (which is not to rule out others). One experiences heightened anxiety about something specific in a context of already feeling more generally anxious and estranged. Depression and anxiety are both associated with what we might call ‘diminished agency’. One’s personal and interpersonal surroundings are globally oppressive and no longer invite effortless responses to meaningful possibilities in the way they once did. So there is a pervasive sense of being incapable of action, and even thoughts may seem sluggish, effortful, bereft of the kind of active anticipation that is involved when one is drawn into and absorbed in a theme (Benson, Gibson, & Brand, 2013; Ratcliffe, 2013). If the person feels more generally passive, helpless and incapable in the face of a threating world, the phenomenological distance between actively initiating something and receiving it from elsewhere may already be lessened. This could render her more vulnerable to a blurring of the phenomenological boundary between thinking that p and perceiving that p. This is consistent with the observation that voices are generally “perceived as being extraordinarily powerful” (Birchwood, Meaden, Trower, Gilbert, & Plaistow, 2000; Chadwick & Birchwood, 1994, p.191). The individual’s relationship with her voices corresponds to her relationship with the social world; she feels passive and vulnerable in the face of interpersonal threat. There are also reports of inner dialogue becoming “more pronounced” before the onset of voices in some cases, with “subtle pre-psychotic distortions of the stream of consciousness – such as abnormal sonorization of inner dialogue and/or perceptualization of thought” (Raballo & Larøi, 2011, p. 163). This suggests a more general blurring of the experienced difference between kinds of intentional state, which would render one more prone to experiencing anxiety-inducing thought contents as alien. It also points to a gradual process, whereby the person becomes anxious about certain thematic contents, with raised anxiety leading to their progressive alienation, sometimes culminating in an experience of the voices as “almost personified” (Raballo & Larøi, 2011, p. 165).18 What we have outlined here also complements an approach to delusions proposed by Currie (2000) and Currie and Jureidini (2001), according to which a delusion is not a recalcitrant false belief but an imagining that is mistaken for a belief. In the case of VHs, there is similarly confusion between two kinds of intentional state: perceiving and thinking. Currie and Jureidini (2001) construe this as an epistemic problem, where one actually imagines something but mistakes one’s imagining that p for the belief that p. However, they later reject a categorical distinction between imagination and belief, allowing for the possibility of intentional states that fall somewhere between the two (Currie & Jureidini, 2004).19 Whether or not our account of the blurring between thought and perception is an epistemic or constitutive one depends on which definitions of ‘perception’ and ‘thought’ one is working with. Our emphasis is on the phenomenology of VHs. We have suggested that an experience of perceiving differs from one of thinking, and that anxious anticipation can lead to a perception-like experience of thought content.20 We grant that non-phenomenological conceptions of thought and perception could be adopted, which would re-cast the situation in epistemic terms. For example, if perception is defined as necessarily involving receipt of information from a sensory source, then a VH of the kind we have described is non-perceptual, pure and simple, and any first-person impression to the contrary is mistaken. However, in purely phenomenological terms, one does not mistake an experience of type x for one of type y; one has an intentional state that is neither x nor y: “it definitely sounds like it is from inside my head. It’s at some kind of border between thinking and hearing” (#18). The view that VHs involve an unfamiliar kind of experience, one that falls between thinking and perceiving, is supported by the observation that people frequently struggle to describe them. VHs are often described as ‘almost like’ something; it is ‘as though’ something were the case. For example, they might be described as ‘like’ telepathy: “The commentary and the violent voices I heard as though someone was talking to me inside my brain, but not my own thoughts. Almost like how telepathy would sound if it were real. I don’t know how else to explain it.” (#4) “.. there are things I ‘hear’ that aren’t as much like truly hearing a voice or voices. [ ….] Instead, these are more like telepathy or hearing without hearing exactly, but knowing that content has been exchanged and feeling that happen.” (#7) Given this, it is interesting to consider the possibility that certain kinds of VHs and what is called ‘thought insertion’ – which are traditionally treated as separate kinds of symptom – are in fact the same phenomenon (Ratcliffe and Wilkinson, 2015). It could be that there are varying degrees of ‘perceptualisation’, which lend themselves to description in terms of either VH or thought insertion. And it could also be that much the same experiences are described in different ways. A quasi-perceptual experience of thought content could be related in terms of (a) a perception with an anomalous content or (b) a thought content that is not one’s own.21 Acknowledging that VHs involve an unfamiliar kind of intentional state also casts light on the phenomenon of ‘double bookkeeping’. Many who voice delusional beliefs and describe hallucinatory experiences also speak and act in ways that distinguish their delusions from other beliefs and their hallucinations from veridical perceptions, thus suggesting different kinds of experience. Sass (1994, p. 3) describes this as follows: Many schizophrenic patients seem to experience their delusions and hallucinations as having a special quality or feel that sets these apart from their ‘real’ beliefs and perceptions. [ …] Indeed, such patients often seem to have a surprising, and rather disconcerting, kind of insight into their own condition. With specific reference to VHs, van den van den Berg (1982, p. 105) observes that voices are often given a “special name” to set them apart from perceptual experiences, due to their having a “recognizable character of their own which distinguishes them from perception and also from imagination”. This is complemented by our view that VHs are not quite like perceptions or thoughts, a view that also explains why the majority of clinical and non-clinical voice- hearers are readily able to distinguish their ‘voices’ from veridical auditory perceptions (Moritz & Larøi, 2008).","As we have emphasised, our account is intended to accommodate only some of those experiences that are labelled as ‘VHs’. Furthermore, it could be that ‘internal VHs that are bereft of at least some auditory properties’ are themselves heterogeneous. While we have emphasised anxious anticipation of inner speech, Michie, Badcock, Waters, and Maybery (2005) propose that VHs involve memory intrusions. Now, McCarthy-Jones et al. (2014) report that only 39% of their subjects acknowledged VH contents resembling memories and fewer still said that their VHs were memories. Even so, it could be that some internal VHs are like this. Indeed, internal VHs could encompass experiences of inner speech, memories and imaginings, as well as some contents that blend memories with imaginings. And the predominance of one form or another may reflect individual differences, different life histories and different diagnostic categories. For instance, we might find a predominance of alienated memory contents in cases where there is past trauma. However, inner speech VHs with less pronounced auditory qualities may be more common in schizophrenia, thus accounting for more frequent reports of ‘thought insertion’ in schizophrenia. However, this is not a problem for our account, given that what we have proposed need not be specific to the experienced boundaries between perception and inner speech. The alienating role of anxiety could be easily extended to the anticipation of distressing memories and imaginings, both of which may have more pronounced auditory qualities. It is also worth keeping in mind that ‘social anxiety’ is not a singular phenomenon but something that would benefit from further analysis. Some first-person accounts emphasise shame and humiliation, others interpersonal threat and helplessness, and others guilt and self-hate. Different variants of social anxiety are likely to be associated with different thematic contents, given that VH contents are mood-congruent. Again, some of these differences may correspond to different psychiatric categories. For example, a person with a diagnosis of psychotic depression might hear voices that mock her and criticise her for her failures (Larøi, 2006). Hence the account we have offered has the potential to accommodate considerable phenomenological diversity, and to distinguish VH characteristics that are more typical of one or another diagnosis. What about external VHs? If first-person accounts are to be taken at face value, they are sometimes much like veridical auditory perceptions. Hence they are unlike what we have described and also arise in a different way. However, social anxiety is implicated here too. Extreme social anxiety could dispose a person towards the anticipation of interpersonal communications with negative, self- directed contents. Delespaul, deVries and Van Os (2002) found that VHs are most likely to occur either when one is in the presence of lots of people or when one is alone. Building on this, Dodgson and Gordon (2009) propose, on the basis of clinical case-studies, a kind of VH called a ‘hypervigilance hallucination’ which occurs especially in ‘noisy’ environments where stimuli are susceptible to multiple interpretations. This, they suggest, accounts for a “substantial subset of externally located voices”.22 The existence of hypervigilance hallucinations as a separate subtype was subsequently supported by Garwood et al. (2013) who, based on a cluster analysis, showed that VHs tend to occur when (i) attention is directed inward in quiet contexts, and (ii) attention is directed outward in noisy contexts. This finding fits nicely with our account. In both internal (inner speech or memory-based) VHs and external, hypervigilance VHs, anxious anticipation shapes and distorts the experience. In instances of anxious hypervigilance, out of external stimuli (such as a ticking clock or the muffled sound of neighbours talking) will emerge the experience of a voice telling the subject exactly what the subject is afraid of hearing. Thus anxiety could also be the underlying cause of at least some external VHs, even though internal and external VHs are generated in different ways. Our account also allows for in-between cases. For example, a perceptual stimulus might trigger an imagining, which is then experienced as an object of pronounced anxiety, and therefore as alien and perception-like. The same applies to inner speech: an external stimulus with auditory qualities could trigger an increasingly determinate linguistic content that is experienced as alien. Unlike an internal VH, this would seem to originate in the external environment, given its association with a perceived environmental cause. In such a case, auditory properties might also be ‘interpreted’ in such a way that they are consistent with the content of the communication. Even so, our emphasis on anxiety cannot do justice to all VH experiences. Some VHs do not have distressing contents (Copolov, Mackinnon, & Trauer, 2004). Indeed, some voice-hearers obtain consolation, support and/or guidance from their voices. This applies to many of those VHs that arise in the context of grief. In a study of nearly 300 widows and widowers in Wales, Rees (1971) found that nearly half had hallucinations of the deceased spouse, sometimes lasting many years. Feelings of presence were most common, but VHs were also reported by 13%. Most of these people found their hallucinations comforting and helpful. We do not claim to have dealt with such cases, and we concede that the phenomenology of grief – in all its complexity – needs to be addressed separately. The same goes for various other kinds of VH experience. However, it may be possible to extend our general approach to VHs without the specific emphasis on anxiety. We have not claimed that anxious anticipation is the only way of anticipating one’s own thought contents that blurs the boundaries between intentional state types. A further limitation of our account is that we have addressed only the ‘phenomenological’ or ‘personal’ level of description and have not postulated any associated mechanisms. What we have supplied here can, however, operate as an explanandum for neurobiological approaches. If one wants to provide a subpersonal account of how x is generated, where x is phenomenological in character, it helps to have a good account of what x consists of. That is what we have tried to provide, and there are clearly implications for accounts of the subpersonal mechanisms involved in VHs. According to a popular family of approaches, our nervous systems distinguish endogenous from externally produced stimuli through a process of self-monitoring. In particular, when a motor command is sent, a copy of that motor command is used to predict the sensory consequences of the action. When the actual sensory consequences match the predicted sensory consequences, the nervous system ‘judges’ that it is self-produced and sensory attenuation occurs (as if it were ‘saying’: ‘don’t worry about this: it’s only you’). When monitoring goes awry, sensory attenuation fails to occur and endogenous stimuli are erroneously attributed to an external cause (e.g. Campbell, 1999; Frith, 1992; Frith, Blakemore, & Wolpert, 2000; Jones & Fernyhough, 2007; Seal, Aleman, & McGuire, 2004). It is not clear how anxiety might fit into a self-monitoring story, as a cause rather than an effect of VHs. Furthermore, our anxiety-based account does not appeal to motor processes. Indeed, one might sense a tension between our account and the self-monitoring approaches, given that we emphasise a change in the style of anticipation, rather than a lack of anticipation. Of course, personal- and subpersonal- level explanations are to be distinguished from each other. And it could be that our anticipating p in style x rather than y involves the breakdown of certain subpersonal prediction mechanisms, while others continue to operate or perhaps operate differently. But, until a more specific account along such lines is developed, the claim that personal level anticipation in style x rather than y involves a breakdown of subpersonal prediction lacks explanatory power.23","We have argued that anxiety induces VHs in the following way: anxious anticipation of thought contents as they become increasingly determinate results in a quasi-perceptual experience of thought content. This is because anxiety is not properly associated with ‘the thought that p’; anxiety alienates us from its objects in a way that we are not ordinarily alienated from our own thought contents. Insofar as the person faces something that she seeks to avoid and is confronted by something that she feels helpless to resist, the resulting experience resembles an affectively charged perception more so than a mundane episode of thought. This account fits in well with subjective reports to the effect that anxiety triggers or aggravates VHs, that the voices confirm negative self- evaluations, and that the voices fall somewhere in between experiences of hearing and thinking. However, we also made clear that our account applies only to a subset of VHs, those with negative content, and most clearly to those that are experienced as internal in origin and lacking in auditory qualities."],["Temporal binding refers to the compression of the perceived time interval between voluntary actions and their sensory consequences. Research suggests that the emotional content of an action outcome can modulate the effects of temporal binding. We attempted to conceptually replicate these findings using a time interval estimation task and different emotionally-valenced action outcomes (Experiments 1 and 2) than used in previous research. Contrary to previous findings, we found no evidence that temporal binding was affected by the emotional valence of action outcomes. After validating our stimuli for equivalence of perceived emotional valence and arousal (Experiment 3), in Experiment 4 we directly replicated Yoshie and Haggard's (2013) original experiment using sound vocalizations as action outcomes and failed to detect a significant effect of emotion on temporal binding. These studies suggest that the emotional valence of action outcomes exerts little influence on temporal binding. The potential implications of these findings are discussed. --------------------------------------------------------------------------------","Temporal binding refers to the compression of the perceived time interval between voluntary actions and their sensory consequences (Haggard, Clark, & Kalogeras, 2002). More specifically, an outcome (e.g., a tone) is experienced earlier when it is triggered by a voluntary action compared to when it occurs in isolation or is triggered by an involuntary movement. Similarly, actions that trigger an event are experienced later than actions with no discernible outcome (see Moore & Obhi, 2012, for a review). For example, Haggard et al. (2002) examined judgements of the onset time of both a voluntary action and a resulting tone using the Libet clock method (Libet, Gleason, Wright, & Pearl, 1983), where one estimates the time of onset of an action or outcome via the position of a rotating clock- hand around a clock-face. These judgements were compared to those made when only the action was performed (i.e., with no outcome) and when a sound was heard in isolation (i.e., without a prior cause). Haggard et al. found that the perceived time of an action was later when the action produced a tone compared to when there was no outcome. Moreover, the perceived time of a sound was earlier when the sound had been produced by an action compared to when it was heard in isolation. In other words, temporal binding means that the time interval between an action and its outcome becomes perceptually compressed when we think there is a causal relationship between action and outcome. Temporal binding has also been observed with methods other than the Libet task, such as verbal or numerical estimates of the interval between action and outcome (Buehner & Humphreys, 2009; Humphreys & Buehner, 2010). Temporal binding has been shown to occur for both self- and other- generated actions (Moore, Teufel, Subramaniam, Davis, & Fletcher, 2013; Poonian & Cunnington, 2013) and may be a general phenomenon linking causally related events (Buehner, 2012). To date, researchers have mostly investigated the conditions required for temporal binding and the mechanisms that underpin it (Hughes, Desantis, & Waszak, 2013), and they have done so using experimental tasks that often involve basic actions, such as a button press, producing sensory feedback, such as an auditory tone (David, Newen, & Vogeley, 2008; Sato & Yasuda, 2005). These temporal binding tasks arguably lack any real- world complexity with which humans perform goal-directed actions to produce meaningful outcomes in everyday life (Moretto, Walsh, & Haggard, 2011). Researchers have started to examine the generalizability of temporal binding effects to stimuli beyond simple and arbitrary outcomes, such as priming social cues (Aarts et al., 2012), authorship of action cues (Desantis, Weiss, Schütz-Bosbach, & Waszak, 2012), leader-follower cues (Pfister, Obhi, Rieger, & Wenke, 2015) and economic and pain cues (Caspar, Christensen, Cleeremans, & Haggard, 2016). For example, Aarts et al. (2012) found that, when primed with a positive picture (taken from the International Affective Picture System; Lang, Bradley, & Cuthbert, 1999) that indicated a reward, temporal binding during the Libet clock task increased compared to neutral primes. Takahata et al. (2012) trained participants to associate two tones with either financial gain or loss. Using the Libet task, they found that the temporal interval between judgements of onsets for actions and outcomes of financial loss was significantly larger than for judgements of financial gain. In other words, negative outcomes reduced the effect of temporal binding. This points towards the possibility that the effect of valence on temporal binding might be driven by self-serving biases, where one is more inclined to associate positive events with the self compared to negative events (Mezulis, Abramson, Hyde, & Hankin, 2004; Miller & Ross, 1975). Yoshie and Haggard (2013) directly tested this idea by investigating whether temporal binding differed between outcomes that varied in terms of their intrinsic emotionality. They asked participants to make voluntary actions (a key-press) that produced auditory sounds that were either of positive or negative emotional vocalizations (e.g., laughter or disgust). Participants made temporal estimations of their actions and the ensuing sound via the Libet clock method. They found that positive sounds produced shorter estimations of onset- time between the action and sound compared to negative sounds (Experiment 1), with this effect being mostly driven by decreased binding to negative outcomes (Experiment 2). Yoshie and Haggard’s (2013) research provided promising evidence that negative emotional outcomes reduce temporal binding, which occurs presumably because people are less inclined to attribute negative outcomes to themselves. However, despite the potential importance of Yoshie and Haggard’s (2013) findings, they have yet to be replicated using other temporal binding tasks and different emotionally-valenced action outcomes. Thus, answering Christensen, Yoshie, Di Costa, and Haggard’s (2016) call for more research exploring the emotional modulation of temporal binding using alternative methods, the goal of the current research was to conceptually replicate Yoshie and Haggard’s (2013) temporal binding effects using an interval estimation procedure (vs. the Libet task; Moore & Obhi, 2012) and images of faces conveying positive and negative emotions (vs. emotional vocalizations; experiments 1 and 2). Moreover, we conducted a separate study to validate the perceived valence of the face stimuli we used in Experiments 1 and 2 (Experiment 3), and we conducted a highly-powered direct replication of Yoshie and Haggard’s first experiment (Experiment 4). On the basis of Yoshie and Haggard’s findings, we expected that temporal binding would be smaller for negative outcomes (faces or vocalizations conveying negative emotions) than for positive outcomes (faces or vocalizations conveying positive emotions).","We used an interval estimation procedure to gauge temporal binding (Ebert & Wegner, 2010; Engbert, Wohlschläger, & Haggard, 2008; Moore, Wegner, & Haggard, 2009). In this procedure, participants are asked to judge the time interval between an action and its sensory outcome (e.g., a button press and a sound). Using this procedure, Engbert et al. (2008) found that the interval between voluntary actions and visual, auditory, and somatic outcomes were compressed compared to the interval between passive actions and similar outcomes. For our task, participants were asked to press the space bar, which was followed by emotionally valenced action-outcomes—namely, emoticons depicting positive, neutral, or negative emotions (see Fig. 1). Emoticons are prevalent throughout modern technological communication, and frequently used to convey emotion (Derks, Bos, & Von Grumbkow, 2008; Hudson et al., 2015). Research has shown that emoticons elicit similar cortical responses to real faces (Churches, Nicholls, Thiessen, Kohler, & Keage, 2014) and that emotions conveyed in emoticons are subject to similar behavioural biases (Öhman, Lundqvist, & Esteves, 2001) and neural processing disruptions (Jolij & Lamme, 2005) as real faces.","We recruited 80 native English-speaking participants (51 males, Mage = 33.91, SDage = 11.27) through prolific.ac, an online crowdsourcing platform. Participants received monetary compensation. We screened participants for the following inclusion criteria: an approval rating of above 90% on prolific.ac (based on prior experiment performance/approval scores) and aged between 18 and 65. The required sample size was fixed ahead of data collection, and a power analysis showed we had 90% power to detect a small effect (Cohen’s f = 0.10) of emotional valence on temporal binding (α = 0.05).","Experiment 1 consisted of 100 trials: 10 practice and 90 experimental trials. We used an interval estimation procedure to measure temporal binding (see Moore & Obhi, 2012). For each trial, participants saw a fixation cross on the screen, and in their own time, pressed the spacebar. In the practice block participant actions produced a neutral stimulus, which was a green circle with a diameter equal to the emoticon images. During practice trials, the green circle appeared after a randomly selected time interval from either 0 ms or a multiple of 100 ms up to 900 ms. We used all intervals in the practice block, to encourage participants to expect the full range of durations in the experimental block. During the practice block, feedback was provided to participants after they made their time estimations. Feedback consisted of both the participant’s estimated time and the actual time of stimulus onset to enhance familiarity with estimating time in milliseconds. In the experimental condition, an emoticon appeared after either 100, 400 or 700 ms (Moore et al., 2009), which remained on the screen for a further 400 ms. We varied the delay intervals to increase participants’ uncertainty regarding the interval between action and outcome to allow for variation in judgement times (cf. Ebert & Wegner, 2010). The emotional expressions of the emoticons were manipulated by orienting the lines representing the mouth: curved upwards for positive, curved downwards for negative, and a straight line for neutral. The emoticons were genderless, varied only in the shape of the mouth, and were presented on a white background in the center of the screen (see Fig. 1). Participants underwent two blocks of 45 trials, allowing for 30 presentations of each emoticon image in total. Participants were instructed that they would not receive feedback for their time estimations during the experimental trials. A schematic display of the sequence of trial events is shown in Fig. 2. Both the time intervals and emoticons (either positive, negative or neutral) were pseudo- randomised across trials, such that there was the same number of trials in each condition at each time interval. A blank screen then followed the emoticon for 400 ms, replaced by a horizontal time estimation scale in the center of the screen (see Fig. 3). The scale ranged from 0-1000 ms, with demarcation lines every 100 ms. Participants were instructed to scroll the slider along the bar to the time that they believed it took the image to appear since their action (in multiples of 100 ms). Once selected, participants confirmed their selections by clicking on a ‘finish’ button, and proceeded to the next trial.","Participants’ mean time estimations for each of the three onset times (100, 400 and 700 ms) and the three emoticons (positive, neutral and negative) were subjected to a 3 (emotional valence: positive, neutral, and negative) × 3 (temporal delay: 100, 400 and 700 ms) fully within-subjects ANOVA (see Fig. 4). Analysis revealed a significant main effect of Temporal Delay, F(2, 158) = 56.54, p < 0.001, ηp2 = 0.77, showing that even under less controlled experimental contexts (i.e., within an online testing platform), participants perceived distinct time intervals corresponding to their actual length (see Dewey & Knoblich, 2014, for comparable findings within a laboratory context). There was no statistically significant effect of emotional valence on time estimations, F(2, 158) = 0.22, p = 0.80, ηp2 = 0.003, nor was there an interaction between temporal delay and emotion, F(4, 316) = 1.47, p = 0.21, ηp2 = 0.018.","In Experiment 1, the emotional valence of action outcomes did not affect temporal binding. One potential limitation of Experiment 1 is that although previous research has shown that emoticons can have the same affective consequences as real faces do (Öhman et al., 2001), the emoticons we used might not have elicited enough of an emotional response to modulate temporal binding. Thus, rather than using emoticons for action outcomes, in Experiment 2 we replicated our Experiment 1 procedure using images of real human faces expressing either negative or positive emotions.","In Experiment 2, we used real-face images as the outcomes to participants’ actions. Real face images have been well-documented to elicit electrocortical responses, and emotional expressions are typically rated along the dimensions of valence and arousal: Smith, Weinberg, Moran, and Hajcak (2013), using the NimStim collection of face-images (NimStim, Tottenham et al., 2009), found that emotional expressions (e.g., happy, fearful, sad), elicited greater cortical responses than neutral face images. Generally, both negative and positive emotions invoke stronger emotional responses than faces with neutral expressions (Ito, Cacioppo, & Lang, 1998), however the current literature suggests negative emotions elicit stronger cortical responses than positive emotions (Leppänen, Kauppinen, Peltola, & Hietanen, 2007; Smith, Cacioppo, Larsen, & Chartrand, 2003).","Participants We recruited 89 participants (55 males: Mage = 33.73, SDage = 10.74) through prolific.ac.uk. An additional participant was excluded due to a technical problem. Participants received monetary compensation. A power analysis showed that we had 95% power to detect a small effect (Cohen’s f = 0.10) of emotional valence on temporal binding (α = 0.05).","Materials and procedures Experiment 2 consisted of 110 trials: 30 practice trials, and 80 experimental trials. To prepare participants for the experimental procedure, we asked participants to initially perform a practice task consisting of 10 trials where their actions produced a neutral stimulus (the green circle). Similar to Experiment 1, during practice trials the time interval for the stimuli to appear was randomly selected from either 0 ms, or a multiple of 100 ms, up to 900 ms. Participants were provided with feedback about the accuracy of their time estimations. Outcome stimuli consisted of 80 face images of young adults either portraying positive or negative expressions, taken from a widely used and validated set of face stimuli (NimStim, Tottenham et al., 2009). The facial images were balanced for gender, such that 10 males and 10 females were randomly chosen from the set (see Fig. 5). Four facial images per male/female were chosen: two depicting positive facial emotions, and two depicting negative facial emotions (80 images in total, 4 × 20). The positive facial emotions included 40 images of a happy expression comprised the positive facial emotions, and 36 images of disgust and 4 images of fear expressions for the negative. Images were presented on a white background in the center of the screen. For the initial practice trials, we used the same neutral stimulus (green circle) as Experiment 1. Participants underwent two experimental task blocks of 40 trials each, with a break between blocks. Each block was dedicated to either solely positive expressions or negative expressions, and the order of task blocks was counterbalanced between participants. Therefore, action-effects were predictable within their own blocks. Furthermore, participants were instructed that they would not receive feedback for their time estimations. The time interval for face images to appear was randomised at 100 ms, 400 ms, or 700 ms (Moore et al., 2009), with the same number of trials in each condition at each time interval. A practice block of 10 trials that contained stimuli of the related task block preceded each experimental block. Upon block completion, participants were instructed that they would be asked to complete another practice task where they would see a different set of images, receiving feedback with their time estimations. To incentivize participant to attend to the face stimuli, we also implemented catch-trials by informing participants that they would also be occasionally asked a question about the image they had just seen (specifically, “Was the previous face male or female?”). If they were correct, then they would be awarded an extra 10 pence per correct question. There were six catch trials in total – three trials per experimental condition. Seventy-six participants (84%) scored correctly on all catch trials, 8 participants (9%) scored correctly on 5 catch trials, and the remaining 6 participants scored correctly on 4 catch trials.","Results We averaged time estimations for each of the three onset times (100, 400 and 700 ms) and for each of the two levels for face-expressions (happy and disgust). We conduced a 2 (emotional valence: positive and negative) × 3 (temporal delay: 100, 400 and 700 ms) fully within-subjects ANOVA. Analysis revealed a significant main effect of temporal delay, F(2, 176) = 225.75, p < 0.001, ηp2 = 0.72 (see Fig. 6). Consistent with Experiment 1, there was no statistically significant effect of emotional valance on time estimation, F(1, 88) = 0.092, p = 0.76, ηp2 = 0.001. There was also no significant interaction between temporal delay and emotion, F(2, 176) = 0.63, p = 0.53, ηp2 = 0.007.","Discussion Similar to Experiment 1, the findings from our second experiment indicated no modulation of negative versus positive emotions on temporal binding. This is despite the use of real facial images depicting emotional expressions (as opposed to emoticons), and the predictability of which emotion-expression (either positive or negative) would result from the participant’s action. For both Experiments 1 and 2, we failed to find any meaningful effect of emotion on temporal binding, which seems inconsistent with earlier findings. One potential issue with our first two experiments, however, is that the stimuli we used for the positive and negative action outcomes (emoticons and real faces) might be perceived as less positively and/or negatively valenced than the sound vocalizations that Yoshie and Haggard (2013) used and therefore produced weaker temporal binding effects. To validate our stimuli, in Experiment 3 participants rated the emotional valence and arousal of the emoticons and faces we used in Experiments 1 and 2 and the positive and negative sound vocalizations that Yoshie and Haggard used. Participants Forty-nine participants were recruited via Amazon’s Mechanical Turk (25 males, Mage = 34.80, SDage = 11.56). To ensure data independence, one additional participant was not included in the analyses because they had a duplicate IP address. Materials and procedures Participants were informed that they would rate several images of faces and sound vocalizations in terms of how negative-to-positive and emotional arousing they appeared or sounded, respectively. Participants first performed a sound check that asked them to identify three different sounds (e.g., a cow mooing) from three choices (e.g., a pig’s oink, a cow’s moo, or a chicken’s cluck) in order to ensure participants both could hear the sounds properly and were paying attention. All respondents saw the emoticons used in Experiment 1, all 80-face expressions used in Experiment 2, and heard 24 sounds (three repetitions of the 8 different sounds). The sounds were the same as those used by Yoshie and Haggard (2013), which were a selection of 8 different non-verbal emotional vocalizations: four negative vocalizations (screams expressing fear or retches expressing disgust, each with both male and female voices) and four positive vocalizations (cheers expressing achievement or laughs expressing amusement, each with both male and female voices). The block order of which type of stimulus the participants rated was randomly determined, and the stimuli presented within those blocks was randomised. Using the same rating scales as Yoshie and Haggard, after seeing/hearing the stimulus, participants judged the extent to which each stimulus looked (for the images) or sounded (for the vocalizations) negative-to-positive, on a 7-point scale ranging from 1 (highly negative) to 7 (highly positive). Participants also rated the extent to which they believed each stimulus sounded or looked emotionally arousing (1 = not arousing at all to 7 = highly arousing). Results ~~~~~~~ Ratings of valence and emotional arousal were averaged across the different positive and negative faces and sounds. Because we were primarily interested in determining whether the different stimuli were perceived to be of equivalent valence, we conducted a one-way ANOVA with stimulus type on three levels (emoticons, faces, and vocalizations) separately for positive and negative stimuli. Shown in Table 1, there was a significant main effect of stimulus type in terms of perceived valence for both positive stimuli, F(2, 96) = 15.21, p < 0.001, ηp2 = 0.24, and negative stimuli F(2, 96) = 22.44, p < 0.001, ηp2 = 0.32. Paired sample t-tests revealed that the happy emoticon was rated as significantly more positive than the positive vocalizations, t(48) = 3.52, p = 0.001; there was no significant mean difference between the positive faces and positive vocalizations in terms of perceived valence, t(48) = 1.95, p = 0.057. For the negative stimuli, the negative vocalizations were rated as more positive (less negative) than both the sad emoticon, t(48) = 5.76, p < 0.001, and the negative faces, t(48) = 2.70, p = 0.01. Thus, the emoticon and face stimuli we used in Experiments 1 and 2 were perceived as either the same or more emotionally- valenced than the sound vocalizations used by Yoshie and Haggard (2013). We also conducted a one-way ANOVA with stimulus type on three levels (emoticons, faces, and vocalizations) separately for positive and negative stimuli for perceived emotional arousal. There were no significant differences among the types of positive stimuli for the ratings of emotional arousal, F(2, 96) = 1.44, p = 0.24, ηp2 = 0.03. For the negative stimuli, F(2, 96) = 3.70, p = 0.028, ηp2 = 0.07, the negative vocalizations were rated as more arousing than the sad emoticon, t(48) = 2.31, p = 0.025, but were no more arousing than the negative faces, t(48) = 0.79, p = 0.43. Discussion ~~~~~~~~~~ The findings of Experiment 3 indicate that the visual stimuli used within Experiments 1 and 2 and the audio stimuli of Yoshie and Haggard (2013) were by and large rated similarly across dimensions of perceived valence and emotional arousal. More specifically, the positive emoticon was rated as more positive and more emotionally arousing than those of real faces and emotionally valenced vocalizations. Similarly, the negative emoticons and the negative faces were rated as more negative than the vocalizations. As such, the failure to find the predicted modulation of temporal binding by emotion in Experiments 1 and 2 does not seem to be driven by differences in the emotional appraisal of the stimuli.","Because we did not find an effect of emotion on temporal binding in Experiments 1 and 2, we conducted a direct replication of Yoshie and Haggard (2013) to investigate the replicability of their findings.","We recruited 24 participants to achieve 95% power to detect Yoshie and Haggard's reported effect size for their Experiment 1 (dz = 0.77): 12 males and 12 females (aged 18–23: Mage = 21.75, SDage = 3.11), one for each of the 8 (2 × 2 × 2) possible orders of conditions (agency/baseline, action/sound and positive/negative vocalizations), counterbalanced between participants. Participants were paid for their time. Following Yoshie and Haggard (2013), we screened for the following exclusion criteria: native language other than English, left handedness, recent use of illicit drugs, uncorrected visual or auditory impairment, and history of psychiatric or neurological illness.","Experiment 4 used the exact same auditory stimuli as Yoshie and Haggard (2013). The stimuli were a selection of non-verbal emotional vocalizations, previously validated in the native English population to significantly differ in perceived valence, but not in perceived arousal (Sauter, Eisner, Calder, & Scott, 2010). In the negative condition, each participant’s keypress was followed by one of four negative vocalizations (screams expressing fear or retches expressing disgust). In the positive condition, these were replaced by positive vocalizations (cheers expressing achievement or laughs expressing amusement). The auditory stimuli in each condition were carefully matched for pitch (peak frequency) and duration. This experiment faithfully replicated the same procedure used by Yoshie and Haggard (2013). We presented the experiment via Macintosh computers (OS X 10.9.5), and used a customised program running in Inquisit v4.01 (Draine, 1998; Millisecond Software) to present participants with the temporal binding task on a 27-inch flat screen. We used the Libet clock task to measure the perceived timing of actions and sounds. During the experiment, participants viewed a Libet clock. In agency conditions, the participant was instructed to press a key on a computer keyboard with the right index finger at a time of his/her choosing, which caused a sound to appear 250 ms later. The participant was then prompted to report where the clock hand was at the onset of their key-press or (agency action condition), in a separate block, at the onset of the sound (agency sound condition). In the single- event baseline action condition, the participant pressed a key at a time of his/her choosing. This keypress did not cause a sound, and the participant was asked to judge the time of his/her keypress. In the single-event baseline sound condition, the participant heard sounds at random intervals, which mimicked time intervals of participant key-presses, and judged the times of sound onsets. To make sure that participants understood the task, we asked participants to perform 5 practice trials before each condition. Participants underwent four task blocks of 32 trials each (baseline action, baseline sound, agency action, and agency sound) for both the negative and positive conditions, or 256 (32 trials × 8 blocks) trials in total. In each block four different sounds of an emotional condition were presented in a randomised order (4 sounds × 8 repetitions). Since each block contained only positive or negative sounds, the four different vocalizations consisted of either the disgust and fear sounds, or the achievement and amusement sounds (each in both male and female voices). Each block was further divided into two sub-blocks of 16 trials each, with the stimuli randomised across the two sub-blocks, such that each sub-block could contain an uneven distribution of sounds. To ensure attention to the auditory stimuli, at the end of every sub- block we asked participants which of the four sounds they heard most frequently during that sub-block. Participants gained a reward of 25 pence for each correct answer to this question. The whole experiment was divided into two sessions of four blocks each. Each session was devoted to action judgments (baseline action and agency action) or sound judgments (baseline sound and agency sound) only. Half of participants (n = 8) judged the times of action in the first session and of sound in the second session, while in the other half (n = 8) the order was reversed. A 10-min break was inserted between the two sessions. To maximize the effects of emotional valence, within each session the baseline and agency blocks of one emotional condition (e.g., negative) were presented successively, and after a 5-min break the blocks of another emotional condition (e.g., positive). Thus, there was an additional 5-min break within each session. Both the order of emotional conditions (negative first or positive first) and the order of task types (baseline first or agency first) were consistent across the two sessions for each participant, and counterbalanced between participants (see Yoshie & Haggard, 2013).","We used Yoshie and Haggard’s (2013) protocol for extracting binding scores. Judgement errors were calculated individually for each block by subtracting the actual onset of the event with the perceived onset. Positive values reflect a delayed judgement, and negative outcomes reflect an anticipatory (early) judgement. Action binding (shift) was calculated by subtracting the mean judgement error of the action in the baseline condition from the mean judgement error in the agency condition. Similarly, sound binding (shift) was calculated by subtracting the mean judgement error of the sound in the baseline condition from the mean judgement of the sound in the agency condition. Composite binding was calculated by subtracting the mean shift in sound judgements from the mean shift in action judgements. Per Yoshie and Haggard (2013), paired t-tests (negative vs. positive) were used to assess the effects of emotional valence on temporal binding. We performed a Grubbs test for outliers (Grubbs, 1950), and no participant met the criteria for exclusion (all ps > 0.05). Additionally, we compared scores between positive and negative vocalizations on an attention task asking participants to state the most frequent sound within the preceding sub-block. A paired-samples t-test revealed no difference in participants’ attention to sounds between negative (M = 3.83, SD = 1.34) and positive (M = 3.63, SD = 1.21) vocalizations, t(23) = 0.96, p = 0.35; dz = 0.20. Table 2 shows the mean judgment errors and shifts relative to baseline conditions for different emotional conditions. The presence of action binding was confirmed by a shift in judgement errors that was significantly different from zero for action judgements in both the negative, t(23) = 2.94, p = 0.007, dz = 0.60, and positive conditions, t(23) = 3.78, p = 0.001, dz = 0.77. Similarly, sound binding was also significant for both negative, t(23) = 5.47, p < 0.001, dz = 1.12, and positive vocalizations, t(23) = 5.27, p < 0.001, dz = 1.08. Composite binding did not differ significantly between the negative (M = −234.68, SD = 174.71) and positive conditions (M = −280.38, SD = 134.20), t(23) = 1.20, p = 0.24, dz = 0.24. Similarly, paired t-tests revealed no significant difference in sound binding, t(23) = 0.64, p = 0.53; dz = 0.13, or action binding, t(23) = 1.16, p = 0.26, dz = 0.24, between the positive and negative conditions.","The findings of Experiment 4 suggest that temporal binding, as measured using the Libet clock method, was not significantly modulated by positive versus negative sound vocalizations as action outcomes. It is worth noting that although we did not find significant modulation of temporal binding by emotional valence, the effect we observed was nonetheless in the same direction as Yoshie and Haggard’s (2013) effect. Thus, if there is an effect of emotional valence on temporal binding using the Libet task and sound vocalizations, it is smaller than previously thought. Moreover, given the results of our Experiments 1 and 2, the effect of emotional valence of action outcomes on temporal binding does not seem to generalize using emotionally valenced visual stimuli and time interval estimation tasks.","The objective of this series of experiments was to investigate the degree to which temporal binding is modulated by emotional valence. Studies 1 and 2 found no significant difference in temporal binding between positive and negative emoticons (Study 1) or positive and negative real facial expressions (Study 2). Study 3 revealed that the stimuli used in Studies 1 and 2 were equivalent in valence and arousal to stimuli that have previously been observed to modulate temporal binding (Yoshie & Haggard, 2013). Furthermore, in a highly powered replication study (Study 4), we observed no significant modulation of temporal binding by emotionally valenced vocalizations (Yoshie & Haggard, 2013). Taken together, these finding cast doubt on whether temporal binding is influenced by outcome valence. Despite showing no significant modulation by valence, temporal binding itself was clearly present in Study 4. Indeed the binding scores were overall somewhat larger than Yoshie and Haggard’s (2013). This suggests that the absence of a valence effect in our study was not due to reduced sensitivity to detect emotional modulation. Although not significant, the effect of valence on binding was in the predicted direction in the current study. However, it is worth noting that this was largely driven by greater action binding to positive tones, whereas Yoshie and Haggard’s (2013) effect was more strongly localised on outcome binding. More recently Christensen et al. (2016) investigated the effect of outcome valence on prospective and retrospective components of action binding (see Moore & Obhi, 2012) with the same vocalizations used here and in Yoshie and Haggard (2013). They observed significantly increased retrospective action binding only when the valence of the outcome was unpredictable. However, for predictable outcomes (as used in the current study) there was reduced action binding for both positive and negative outcomes compared to neutral outcomes. Taken together with the current findings, a complex picture emerges whereby the precise effect of emotion on temporal binding cannot be clearly attributed to a simple self-serving bias such that positive outcomes increase binding. This may reflect a genuine complexity in the precise mechanisms driving the emotional modulation of binding, or it might reflect the fact that the underlying effect is small or unreliable. The absence of an effect of valence in Experiments 1 and 2 suggest that any effect, if present in the population, does not generalize to other measures of binding. Future work should attempt to replicate and extend other examples of self-serving bias in temporal binding (Aarts et al., 2012; Takahata et al., 2012) and sensory attenuation (Gentsch, Weiss, Spengler, Synofzik, & Schütz-Bosbach, 2015; Hughes, 2015) to further advance our understanding of how (or if) outcome valence influences implicit agency. Assessing the degree to which binding is modulated by factors that also modulate explicit agency reports is important to determine the relationship between implicit and explicit agency. Recent evidence suggestions that neither sensory attenuation (Dewey & Knoblich, 2014) nor temporal binding (Dewey & Knoblich, 2014; Saito, Takahata, Murai, & Takahashi, 2015) correlate with explicit reports of agency. While explicit and implicit measures will never show total convergence, positive evidence of covariation is important to argue that conscious reports and unconscious biases are indeed measuring the same underlying process. The current studies provide new evidence that questions the degree to which temporal binding is modulated by self-serving biases.","This research was supported by studentship ES/J500045/1 from the Economic and Social Research Council."],["Deception research has been criticized for its common practice of randomly allocating senders to truth-telling and lying conditions. In this study, we directly compared receivers’ lie-detection accuracy when judging randomly assigned versus self-selected truth-tellers and liars. In a trust-game setting, senders were instructed to lie or tell the truth (random assignment; n = 16) or were allowed to choose to lie or tell the truth of their own accord (self-selection; n = 16). In a sample of receivers (N = 200), we tested two alternative hypotheses, predicting opposite effects of random assignment (vs. self-selection) on receivers’ lie-detection accuracy. Accuracy rates did not differ significantly as a function of veracity assignment, failing to support the claim that random assignment of liars and truth-tellers alters the detectability of deception. Equivalence tests indicated that, while a small effect of random assignment cannot be ruled out, moderate (or larger) effect sizes are unlikely. --------------------------------------------------------------------------------","General Audience Summary In everyday communication, people typically decide whether to lie or to tell the truth of their own accord. In most studies on lie detection, however, researchers instruct individuals to lie or tell the truth on a random basis. This approach has received critique from experts in the field, because it does not reflect what happens in real life. Since self-selected and instructed liars and truth-tellers differ in several ways (e.g., motivation, proficiency of lying), the two modes of veracity assignment may give rise to different cues to truth and deception. In the current study, we tested whether random assignment, as compared with self-selection, improves or impairs people's ability to detect deception. Liars and truth-tellers (senders) tried to convince participants (receivers) to trust them with their money, promising cooperation and financial gain in return. Half of the senders had been randomly assigned to lie or tell the truth, whereas the other half had chosen to lie or tell the truth of their own accord. We tested two competing hypotheses: First, on the assumption that it prevents good liars from choosing to lie (and poor liars from choosing not to lie), random assignment would improve receivers’ ability to detect lies. Second, on the assumption that there are detectable differences between senders who are likely to lie when given the opportunity and those unlikely to lie, random assignment would make such differences uninformative and impair receivers’ ability to detect lies. Our results did not support any of the hypotheses (lie-detection accuracy was near chance level in all experimental conditions), thus failing to support the claim that random assignment of liars and truth-tellers alters the detectability of deception. Instead, they indicate that the widely documented poor ability of humans to detect lies holds for both self-selected and instructed liars.","Seventy-two male economy students at the University of Gothenburg, Sweden, were recruited for a study on economic decision making. Video messages from 32 of these (age range: 19–32 years, Mage = 24.03, Mdnage = 23, SD = 3.19) were used as stimulus material in Phase 2. The decision to specifically recruit male economy students was based on research showing that economists are more likely to act self-interestedly compared to non-economists (Carter & Irons, 1991; Frank, Gilovich, & Regan, 1993). Of particular interest, male economy students have been found to act more self-interestedly than female economy students in a trust-game dilemma (James, Soroka, & Benjafield, 2001). Considering our aim to elicit voluntary lies in a trust game dilemma, only male economy students were recruited.","Procedure Upon arrival, participants were seated in front of a computer and introduced to a trust-game dilemma. The scenario was an amended version of that used by Holm and Nystedt (2008). Participants were told they had been randomly paired with another participant, referred to as “X”, and that they had been given 50 SEK each to a total sum of 100 SEK. Participants were further told that this amount could be doubled to 200 SEK on the condition that X trusted the participant with the decision how to distribute the money between them. If they did not convince X they would, however, only receive the 50 SEK. Participants were told that they were going to send a video message to X, expressing why they are altruistic and can be trusted, in order to convince him/her to pass the decision over to them. Participants were presented with two distribution options: (a) X gets 0 SEK and the participant gets 200 SEK (egoistic option), and (b) X gets 80 SEK and the participant gets 120 SEK (altruistic option). The instructions stated that both parties would be better off if X chose to hand over the decision to the participant, provided that the participant would choose the altruistic option (b). Participants in the random-assignment condition were instructed, through randomization, either to select the egoistic alternative (a) or to select the altruistic alternative (b). Thus, they needed to lie about having altruistic intentions if they were assigned the egoistic alternative (n = 20), and tell the truth if they were assigned the altruistic alternative (n = 17). Participants in the self-selection condition were instead instructed to choose the distribution option they personally preferred. Thus, they needed to lie if they chose the egoistic alternative (n = 8), and to tell the truth if they chose the altruistic alternative (n = 27). When participants had been informed about their randomly assigned choice or had made their own choice, they were given 5 min alone to prepare their message. They were told that they could say whatever they wanted in order to convince X to pass the decision over to them, and that they had to keep within the upper and lower time limits of 30 s and 1 min, respectively. After having prepared what to say, an assistant entered the room and asked the participant to take a seat in a chair in front of the camera and to look straight into the camera, whereupon the assistant started the video recording. The camera was placed straight in front of the participants and was slightly higher than eye- level. The image was cropped slightly above their heads and just below their knees. If participants talked for longer than 1 min the assistant, who was seated in the corner of the room, signaled by standing up that they should end the message. When the video recording was finished, participants were shown to a second room while the experimenter allegedly brought the message to X to watch. In reality, the experimenter waited for 5 min before re-entering the room. All participants were then told that X trusted them after having viewed the message. After this, participants were asked to divide the 200 SEK in accordance to their distribution choice and to place the money in each of two envelopes marked “for you” and “X.” Finally, participants brought the envelopes back to the first room, after which they left with the compensation that they had put in their own envelope. The experimenter subsequently verified that all participants had left the correct amount of money in the “X” envelope (i.e., corresponding to the chosen distribution option). Participants and design Sample size was determined before data collection, based on an a priori power analysis using the G*Power 3.1.9.2 software (Faul, Erdfelder, Lang, & Buchner, 2007). As no previous tests of the current hypotheses are available in the literature, we computed the necessary sample size to achieve 80% power, with a 5% Type-I error rate, to detect the average effect size of social cognition studies (r = .20) reported by Richard, Bond, and Stokes-Zoota (2003). The analysis indicated a sample size of 199 participants. To allow for attrition, 215 participants were recruited for a study on social judgments. They were recruited from a research participant pool, consisting mainly of university students and people from the general public, at the Department of Psychology, University of Gothenburg, Sweden. They received 50 SEK (∼6 USD) as compensation for participating in the study. Fifteen participants were excluded from analyses—five knew someone in the videos, four did not follow the instructions, two had language difficulties, two experienced technical problems, one was disturbed while performing the test, and one had impaired hearing—which resulted in a final sample size of 200 participants. One hundred and thirty-seven participants were female, 60 were male, and three did not specify their gender. Ages ranged between 18 and 72 years (M = 30.03, Mdn = 26, SD = 11.26). Participants were randomly assigned to one of four conditions in a 2 (veracity assignment: random vs. self-selection) × 2 (detection strategy: feeling focus vs. detail focus) factorial design. The exact number of participants per design cell is reported in Table 1. Sender veracity was varied within participants (see Procedure section below). Procedure Participants were seated in front of a computer where they read and signed an informed consent form. They were informed about the procedure of the trust game in Phase 1. Specifically, they were told that participants in a previous study had made a choice about how to distribute 200 SEK between themselves and another participant (referred to as “X”). There had been one egoistic decision alternative (take 200 SEK and give 0 SEK to X) and one altruistic decision alternative (take 120 SEK and give 80 SEK to X). Importantly, participants in the previous study had been instructed to try to convince X to trust them with the decision, because otherwise they would get only 50 SEK each. To this end, the participants in the previous study had recorded a persuasive video message to X. Participants were further told that they were going to watch a series of these video messages, while assuming the role of X, imagining that each message was directed toward themselves. They were also informed that some of the persons in the videos lied about choosing the altruistic distribution option, while some told the truth about choosing the altruistic distribution option. Participants were told that their task was to judge the veracity of each of the video messages. Four video sets were created, each containing eight videos (four true and four deceptive messages). Two of the video sets contained only senders who had been randomly assigned to lie or tell the truth (random assignment) and the other two contained only senders who had voluntarily chosen to lie or tell the truth (self-selection).1 Participants were randomly assigned to one of the two sets in their respective veracity- assignment condition. Each participant thus watched and judged eight video messages, the order of which was randomized for each participant. To manipulate detection strategy, participants in the feeling-focus condition were instructed to base their veracity judgments on how they felt about the sender and the message: “When watching the following video clips and making your judgments, we want you to focus on how you feel about the person and what he is saying. Does your gut feeling tell you that the person is lying or telling the truth?” Participants in the detail-focus condition were instructed to base their judgments on details in the verbal content of the message: “When watching the following video clips and making your judgments, we want you to focus on analyzing the details in what the person is saying. Do these details indicate that the person is lying or telling the truth?” A reminder of the detection strategy was presented on the screen immediately before the presentation of each video message. Dependent measures Immediately after watching each video clip, participants made a dichotomous veracity judgment by answering the following question: “Do you think that the person told the truth or lied (i.e., did he choose to share the money or not)?” For the purpose of statistical analysis, participants’ veracity judgments were converted, using the procedures described in Stanislaw and Todorov (1999), into signal-detection measures of discrimination accuracy (d′) and response bias (c).2 Participants also indicated how confident they were that their veracity judgment was correct (1 = not at all confident, 7 = very confident). Participants in the feeling-focus condition further indicated how strongly they felt that the person was lying or telling the truth (1 = strongly feel that the person is lying, 7 = strongly feel that the person is telling the truth). Participants in the detail- focus condition instead indicated how strongly they perceived that the details in the person's statement suggested that he was lying or telling the truth (1 = strongly suggest that the person is lying, 7 = strongly suggest that the person is telling the truth). The latter two measures were included to strengthen the manipulation of detection strategy, by repeatedly drawing participants’ attention to their affective impression of the sender or to the verbal details of the messages. Hence, they were not included in any statistical analyses. Manipulation checks and control variables In order to check the effectiveness of the detection-strategy manipulation, participants were asked to rate on 7-point scales the extent to which they (a) based their veracity judgments on their gut feeling and (b) based their veracity judgments on the details of what the individual was saying (1 = not at all, 7 = completely). Participants’ accuracy motivation was also assessed by asking how important they thought it was to make correct veracity judgements (1 = not at all important, 7 = very important). In addition, participants’ general trust was measured using the General Trust Scale (Yamagishi & Yamagishi, 1994). The scale consists of six items measuring peoples’ propensity to trust others (e.g., “most people are basically honest”), using 5-point rating scales (1 = strongly disagree, 5 = strongly agree). A general trust index was created by averaging participants’ ratings on the six items (Cronbachs's α = .64). Given a previous finding that higher scores on the General Trust Scale are associated with better lie-detection performance (Carter & Weber, 2010), we included the general trust index as a covariate in the analyses of the lie-detection measures to account for variance due to individual differences. Preliminary Analyses ~~~~~~~~~~~~~~~~~~~~ Participants in all four conditions were highly motivated to make correct veracity judgments, with means varying between 5.38 and 5.69 on 7-point Likert-scales. A 2 (veracity assignment: random vs. self-selection) × 2 (detection strategy: feeling focus vs. detail focus) ANOVA showed no significant main or interaction effects (Fs < 0.94, ps > .335), indicating that the four groups did not differ significantly from each other in level of motivation. A second ANOVA on participants’ general trust also did not reveal any main or interaction effects (Fs < 3.05, ps > .082). Main Analysis ~~~~~~~~~~~~~ Means and standard deviations for the variables related to lie detection are reported in Table 1. Participants’ average detection accuracy was 48.3%, which did not differ significantly from chance performance (50%) as indicated by a one-sample t-test, t(199) = −1.33, p = .186, Hedges’ g = −0.09, 95% CI [−0.23, 0.04]. On average, participants judged 51.1% of the messages as truthful. There was not a significant truth or lie bias, as the proportion truth judgments did not differ significantly from 50%, t(199) = 1.10, p = .271, Hedges’ g = 0.08, 95% CI [−0.06, 0.22]. Additional Analyses ~~~~~~~~~~~~~~~~~~~ Given the null finding regarding the effect of veracity assignment on discrimination accuracy (d′), we ran tests of statistical equivalence to examine the informativeness of the result. The procedure recommended by Lakens (2017) involves running two one-sided tests (TOST) to test whether the observed effect size differs significantly from the smallest effect size of interest, defined by a lower (e.g., d = − 0.20) and an upper (e.g., d = 0.20) equivalence bound. If the observed effect is significantly higher than the lower bound and significantly lower than the upper bound, it is deemed equivalent to the absence of an effect that is worth examining (Lakens, 2017). A first TOST with equivalence bounds set at d = ±0.20 was non-significant, t(203.95) = 0.06, p = .522 (one- tailed), indicating that our result is not equivalent to the absence of a small effect (by conventional standards; Cohen, 1988). In the current experiment, an effect of d = 0.20 translates into a difference in lie-detection accuracy of 3.73% (0.20 × 18.63 [SDpooled]). A second TOST, with equivalence bounds set at d = ±0.50 was significant, t(203.95) = −2.10, p = .019 (one-tailed), indicating that our results are equivalent to the absence of a medium-sized or larger effect (by conventional standards). In the current experiment, an effect of d = 0.50 translates into a difference in lie-detection accuracy of 9.32% (0.50 × 18.63 [SDpooled]). In sum, these analyses show that while our findings do not provide evidence for the absence of a minor effect, they do indicate that a moderate or large effect of veracity assignment on discrimination accuracy is unlikely.","In the current study, we pitted two competing hypotheses against each other, predicting that receivers’ lie-detection accuracy would increase (H1) or decrease (H2) as a result of the random assignment (vs. self-selection) of lying and truth-telling senders. The observed results did not provide support for any of the hypotheses. In fact, participants across the experimental conditions performed similarly to what would be predicted by chance. Moreover, we found no indication that the effect of random assignment on detection accuracy differs as a function of receivers’ basing their veracity judgments on verbal content details or on intuitive feelings. When interpreting null results, it is important to consider the possibility of a false negative finding. Our study was powered to detect effects as small as r = .20 (d = 0.41) at 80% power. Moreover, follow-up tests indicated that our results are statistically equivalent to the absence of a medium-sized or larger (d ≥ 0.50) effect of veracity assignment. The current findings, thus, indicate that medium or large effects of our manipulation, as studied under the current paradigm, likely do not exist. However, a few caveats should be noted before generalizing our findings further. First, our results cannot rule out the existence of a small effect. The meta-analysis by Bond and DePaulo (2006) indicates that most moderators of lie-detection accuracy are, indeed, quite weak, and such effects are difficult to detect in an individual experiment, unless very large samples are used. Second, our relatively modest stimulus material (32 senders) may have restricted the range of detectable differences between self-selected liars and truth-tellers. Provided that such differences exist, they should emerge more reliably when larger samples of stimuli are used (Wells & Windschitl, 1999). Third, our study produced null findings in a trust-game paradigm, but their replicability in other deception contexts is unknown. Therefore, the role of veracity assignment procedures in other domains (e.g., criminal, relational) should be examined. Finally, manipulation checks indicated that the manipulation of detection strategy was only moderately effective. It is thus questionable whether we achieved the intended separation of participants who focused predominantly on verbal details and those who focused predominantly on feelings. We also do not know whether the manipulation shifted receivers’ relative attention to verbal and nonverbal cues, as we did not measure this directly. These limitations would have been particularly problematic had the effect of detection strategy been the main focus of the study. However, given that we explored detection strategy only as a potential moderator of any effect of veracity assignment, the weak manipulation of detection strategy is less consequential. These caveats notwithstanding, however, the current findings indicate that veracity assignment and detection strategy, as operationalized in the current experiment, have at most a modest influence on human lie- detection performance. We unexpectedly found an interaction effect between veracity assignment and detection strategy on participants’ response bias. Specifically, participants relying on a feeling-focused (vs. detail-focused) strategy were significantly more likely to judge senders as truthful, but only when senders had self-selected to lie or tell the truth. Given the effect's modest size, and the fact that it emerged from exploratory analyses, we refrain from a lengthy attempt at an explanation. The result may be an indication, however, that senders attuned to feeling-based intuitive impressions (vs. verbal content details) are more receptive to cues that signal high believability, which may in turn be more prevalent among self-selected (vs. randomly assigned) senders. This possibility could be tested in future confirmatory research. Another notable finding is that we were unable to replicate Carter and Weber's (2010) finding on the relationship between receivers’ general trust and lie-detection ability. They concluded, based on a very small sample (N = 29), that “high trusters” are better lie detectors than “low trusters.” In the current study, however, we found no correlation between general trust and lie-detection accuracy in a sample almost seven times as large (N = 200). Instead, we found the more intuitive result that high trusters displayed a stronger truth bias than did low trusters. The current finding casts serious doubt on the replicability of the original finding reported by Carter and Weber (2010) and calls for further replication attempts. Moreover, compared with Carter and Weber's result, ours is more compatible with the established finding that individual differences in lie-detection ability are very small, and that there is considerably more individual variation in people's inclination to regard statements as truthful (Bond & DePaulo, 2008). In conclusion, the current study was the first to fully compare the detectability of self-selected versus randomly assigned truth-tellers and liars. We found no evidence that the mode of veracity assignment, separately or in conjunction with receivers’ detection strategy, influences lie-detection accuracy. This finding fails to support, and provides some evidence against, the concern that random assignment would substantially alter the detectability of truths and lies (Levine, 2018). This should not be taken as evidence that the method of veracity assignment is of little consequence to deception researchers. Rather, the choice of veracity assignment should be informed by the focal research question. If the aim is to examine differences between truthful and deceptive messages, independent of senders’ propensity to lie, then random assignment is appropriate. If, however, the researcher is interested in differences that correlate with senders’ propensity to lie, then allowing liars and truth-tellers to self-select is necessary.","K.A. conceived and designed the studies and collected the data. K.A., S.C., and E.M.G. all contributed to analyzing and interpreting the data, writing the manuscript, and approving the final version of the manuscript for submission.","The authors declare no conflict of interest."],["Are individuals responsible for behaviour that is implicitly biased? Implicitly biased actions are those which manifest the distorting influence of implicit associations. That they express these 'implicit' features of our cognitive and motivational make up has been appealed to in support of the claim that, because individuals lack the relevant awareness of their morally problematic discriminatory behaviour, they are not responsible for behaving in ways that manifest implicit bias. However, the claim that such influences are implicit is, in fact, not straightforwardly related to the claim that individuals lack awareness of the morally problematic dimensions of their behaviour. Nor is it clear that lack of awareness does absolve from responsibility. This may depend on whether individuals culpably fail to know something that they should know. I propose that an answer to this question, in turn, depends on whether other imperfect cognitions are implicated in any lack of the relevant kind of awareness.In this paper I clarify our understanding of 'implicitly biased actions' and then argue that there are three different dimensions of awareness that might be at issue in the claim that individuals lack awareness of implicit bias. Having identified the relevant sense of awareness I argue that only one of these senses is defensibly incorporated into a condition for responsibility, rejecting recent arguments from Washington & Kelly for an 'externalist' epistemic condition. Having identified what individuals should - and can - know about their implicitly biased actions, I turn to the question of whether failures to know this are culpable. This brings us to consider the role of implicit biases in relation to other imperfect cognitions. I conclude that responsibility for implicitly biased actions may depend on answers to further questions about their relationship to other imperfect cognitions. --------------------------------------------------------------------------------","Are individuals responsible for behaviour that is implicitly biased? Implicitly biased actions are those which manifest the distorting influence of implicit associations. That they express these ‘implicit’ features of our cognitive and motivational make up has been appealed to in support of the claim that, because individuals lack the relevant awareness of their morally problematic discriminatory behaviour, they are not responsible for behaving in ways that manifest implicit bias. However, the claim that such influences are implicit is, in fact, not straightforwardly related to the claim that individuals lack awareness of the morally problematic dimensions of their behaviour. Nor is it clear that lack of awareness does absolve from responsibility. This may depend on whether individuals culpably fail to know something that they should know. I propose that an answer to this question, in turn, depends on whether other imperfect cognitions are implicated in any lack of the relevant kind of awareness. In Section 2 clarify our understanding of ‘implicitly biased actions’ and then argue that there are three different dimensions of awareness that might be at issue in the claim that individuals lack awareness of implicit bias. Having identified the relevant sense of awareness, in Section 3, I argue that only one of these senses is defensibly incorporated into a condition for responsibility, rejecting recent arguments from Washington & Kelly for an ‘externalist’ epistemic condition. Having identified what individuals should – and can – know about their implicitly biased actions, I turn in Section 4 to the question of whether failures to know this are culpable. This brings us to consider the role of implicit biases in relation to other imperfect cognitions. I conclude that responsibility for implicitly biased actions may depend on answers to further questions about their relationship to other imperfect cognitions.","What are the phenomena at issue when we talk of implicit biases? Amodio and Mendoza (2010) describe them as ‘associations stored in memory’ (364).1 These associations can influence behaviours and judgements. For example, implicit associations have been posited as explaining differential evaluations of the same CV whose only difference was race, indicated by the name at the top (Dovidio & Gaertner, 2000); as implicated in shooter bias, whereby in a computer simulation individuals were more likely to ‘shoot’ black men with weapons than white men with weapons (Glaser & Knowles, 2008); and as playing a role in the seating distances between experimental participants and stigmatised group members (Tidswell, Sheeran, & Webb, in preparation). But what is it about the associations involved in producing such varied behaviour that makes them implicit, and how do we delineate which of the many associations of this sort constitute biases? ‘Implicit’ ~~~~~~~~~~ Some have suggested that associations are implicit simply because the measure used to access them is an implicit one; namely, one that does not rely on self-report measures, nor the voluntary offering of information about one’s attitudes. An implicit measure might involve a prime of which the agent is not aware, then a measure of how being so primed influences behaviour or judgement. But why use an implicit measure to access these associations? One reason is that individuals may not be forthcoming or frank about associations they would rather they did not have. Another reason is that the associations are characterised by features of automatic processes which render them difficult for the agent to identify and report on. DeHouwer, Teige-Mocigemba, Spruyt, and Moors (2009) pick out the following features as ones taken to be characteristic of implicit associations: operation without the guidance of proximal goals (that would enable the agent to initiate, intervene or stop the processes); operation without substantial cognitive resources (such as when one’s attention is occupied with some other task); and operation with very limited time (such as when one is required to respond very quickly); or without awareness. Notably, some philosophers and psychologists have taken this latter feature as characteristic or even definitional of implicit bias (see e.g. Kelly & Roedder, 2008; Washington & Kelly, in press; Greenwald & Banaji, 1995; Saul, 2013). I will say much more about this characteristic in the following. What is important is that these features are not specified as necessary for an association to be implicit, which leaves considerable scope for variation in the properties of the associations being measured in such studies, and discussed in subsequent philosophical literatures.2 ‘Bias’ ~~~~~~ What is it about some implicit associations that should lead us to characterise them as ‘biases’? We can think of such associations as biases when they are disposed to exert a distorting influence on judgement. The influence at issue can be characterised as distorting in that it leads to a judgement which departs from the norms of rationality. This can be most clearly seen in the CV studies mentioned above: the name at the top of a CV does not provide a reason for judging it to be better or worse than an otherwise identical CV. It is less clear how this analysis explains the behavioural outputs, such as increased seating distance from stigmatised groups. One possibility would be to extend the definition to include not only distortions of judgement but also undesired or undesirable influences on action. Another would be to suppose that these behaviours are preceded by (tacit) judgements, which are distorted (judgements about the suitable place to arrange the seat say; or about the level of danger posed by an individual). Perhaps either way of proceeding is adequate, but for the sake of providing a simple contrast with explicit bias, I will work with the model that sees all implicitly biased action as involving a distortion of judgement. This judgement is sometimes the output measured; in other cases it informs the behavioural output which is measured.3 We can proceed, then, with the following understanding of implicit bias: it is operative when implicit associations produce a distorting influence on judgement and hence behaviour informed by that judgement (this leaves room for some implicit associations which are not implicit biases). Responsibility ~~~~~~~~~~~~~~ What are we asking when considering whether or not an agent is responsible for implicit bias? One dimension of responsibility is forward-looking: is this something an agent can be asked to take responsibility for, and bring about changes in her cognition? This is an important sense of responsibility when considering implicit biases, where one of the primary aims is to bring it about that individuals act in ways that are less biased.4 But this is not the only sense of responsibility I am concerned with here. The question here is whether an individual can be held responsible, in the sense of liable to praise or blame, for the manifestation of implicit bias. If an individual acts in ways that express bias, is this something that they are blameworthy for (there is, of course, a further question about whether expressing blame would be appropriate)? To say that the agent is blameworthy, then, is to say that they have intentionally done something that violated a moral standard that we expected them to maintain, and as a result certain responses would be warranted: disapprobation or other forms of informal sanction on the part of others; resentment on the part of the wronged party; guilt on the part of the wrong-doer, and resolution to avoid such behaviours or actions in future (indeed, to take responsibility for that).5 Whether individuals are responsible in this sense is the question with which I am concerned. Awareness ~~~~~~~~~ Authors who have addressed the question of responsibility for implicit bias have argued that to the extent that individuals are not (Saul, 2013) or could not reasonably be expected to be (Washington & Kelly, in press) aware of their implicit biases, they are not responsible for action influenced by these biases. However, there are different senses of ‘awareness’ at work in this debate. In the next sub-section I articulate three different senses of awareness that are circulating in the literatures (philosophy and psychology) about implicit biases. This will enable us to identify more precisely the sense in which individuals have or lack awareness, such that we can evaluate whether being in such an epistemic situation exculpates from moral responsibility. Three kinds of awareness ~~~~~~~~~~~~~~~~~~~~~~~~ What do individuals lack awareness of, in the case of implicit bias? In the literature from empirical psychology and from philosophy, we find different views on this. For example, in making the claim that individuals should not be blamed for implicit biases, Saul writes that ‘a person should not be blamed for an implicit bias of which they are completely unaware’ (Saul, 2013 p. 55), where what is at issue is that individuals are not aware of the operation of the bias in the production of action. (We might also interpret Saul as claiming that individuals are not aware of the presence of the bias, but let us restrict our focus to the operation of the bias, in relation to which our concern with responsibility arises most pressingly). The sense of awareness at issue here seems to be introspective awareness; awareness that might yield knowledge of one’s cognitive processes simply by reflecting on one’s internal states and processes. This sense of awareness is also in play in Kelly & Roedder’s description of implicit measures as accessing aspects of cognition ‘not easily accessible or readily available to introspection’ (Kelly & Roedder, 2008, p. 524), and in Anderson’s discussion of ‘unconscious stereotypes’; representations of which the agent is introspectively unaware (2010, p.74, 48).6 One might have introspective awareness with respect to whether certain beliefs or feelings are playing a role in one’s decisions: one can ask oneself, and on reflection give an answer. But, the claim goes, one cannot simply introspect and discern if an implicit bias is operating in the production of action. Alternatively, we might be concerned with whether individuals are aware of a set of propositions about implicit bias, which are likely to be true of themselves. This is a second sense of awareness at issue in Saul’s claims: when she writes that individuals may ‘become aware that they are likely to have implicit biases’ (2013, p.55), the awareness at issue is of the body of knowledge concerning the disposition of individuals to be biased. Similarly, Washington and Kelly (in press) focus on whether some individuals in fact know, or should know, certain empirical facts about their probable susceptibility to implicit biases. In attempting to explain their divergent intuitions about the responsibility of an egalitarian on a hiring committee who manifests implicit bias in the 1980s, and a similarly placed contemporaneous egalitarian, they observe that ‘in 1980, no one knew the creepy psychological facts about implicit biases; the psychological research had not yet been done, and so today’s wealth of empirical evidence simply did not exist’ (in press). This fact (about what individuals can reasonably be expected to know) figures in their explanation of why they seek to exculpate the 1980s discriminator, but not the contemporaneous one. It is unreasonable to expect the 1980s discriminator to be aware of facts about implicit bias as yielded by empirical psychology: those facts were not part of our epistemic milieu then. But now, those on hiring committees have epistemic responsibilities, which include familiarising themselves with that body of knowledge. At issue in these claims, then, is not some particular aspect of one’s cognition, nor some observed behavioural effects, but rather some body of knowledge pertaining to individuals’ general tendencies to manifest implicit biases. Knowledge of this would be achieved through what we might call inferential awareness; awareness reached through inferences made about this body of empirical knowledge, and one’s own behavioural dispositions in light of that. Finally, other authors are concerned with whether individuals have awareness of the manifestation in behaviour of implicit bias: in particular, their focus is on an individual’s awareness of discrepant behavioural responses on tests for biases, and their willingness to attribute such responses to implicit prejudice. For example, Monteith, Voils, and Ashburn-Nardo (2001) undertook studies on implicit race associations, in which they measured different response times to pairing tasks (black names with unpleasant [congruent] or pleasant [incongruent] terms, and white names with unpleasant [incongruent] or pleasant [congruent] terms. The congruent pairings are those that individuals are expected to respond more quickly on, insofar as they are informed by stronger, more accessible, associations. So, a faster ‘black/unpleasant’ response than ‘black/pleasant’ response indicates a stronger association between the former, negative, construct than the latter. Following the study, Monteith et al. ‘identified participants who recognized that they were slower on incongruent [contra-implicit association] than on congruent [consistent with implicit association] IAT trials’ (2001, p.405), and found that a significant portion of these individuals were able to attribute their differential response times to implicit prejudice. The striking claim here is that a considerable number of individuals (64% of the study participants) were able to recognise, on the basis of observations of their own behavioural responses, that they were responding differently to the different stimuli. At issue here, then, is whether individuals are aware of their differential or discriminatory behavioural outputs. As a matter of fact, it turns out that at least some individuals are. Individuals at least sometimes have observational awareness of the extent to which their own actions are biased; and sometimes they additionally have what we can refer to as attributional awareness: awareness that some feature of their cognition – implicit prejudice – can be attributed as causally producing that influence.7,8 We have three candidate senses of awareness, then: introspective awareness of the implicit association itself, or its operation; inferential awareness of the body of knowledge about people’s tendencies to harbour, and display, implicit bias; observational awareness of the effects of the implicit associations on behaviour (sometimes alongside attributional awareness of the cause of these effects). Because these distinct senses of awareness have not been distinguished in either the empirical or philosophical literatures, we should be cautious in our claims about the relationship between these different senses of awareness; considerably more conceptual (and empirical) work is needed to understand their relationship than is possible here. However, some preliminary remarks can be made. Firstly, if one denies that introspective awareness of implicit biases (or indeed any mental state, cf. Levy, 2014) is possible, then acquiring knowledge of the other kinds (inferential or observational) will not garner that sort of awareness. Secondly, gaining inferential awareness (of the fact one is likely to be biased) may well aid observational awareness, if it can prompt reflection on one’s behaviours that might yield evidence of subtle discrimination. Finally, inferential and observational awareness of facts about implicit bias may also generate attributional awareness, in that one may be able to attribute one’s discriminatory behaviour to the probable presence of implicit bias, even if one is unable to introspect on such biased processes. A normative epistemic condition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Each of these senses of awareness of implicit bias has been used in the literature regarding whether individuals meet the epistemic conditions for responsibility. How should we think of these different senses in relation to the epistemic conditions for responsibility? When considering whether lack of awareness, or ignorance, exculpates, we have to ask not merely what an individual does or does not have awareness of, in some relevant sense of awareness.9 Some failures of awareness are themselves culpable. For example, ignorance of the harm one is perpetrating may not exculpate if one can reasonably be expected to check that one’s actions are not harmful in that way. Similarly, forgetting – being unaware of – a meeting one had promised to keep does not exculpate, and may itself be culpable, in the usual course of things (setting aside excessive pressures or stresses or distractions). Not being aware of a motive of cruelty or jealousy that shapes one’s interactions with friends does not excuse actions that are cruel or express jealousy. This is because, in all these cases (bar exceptional circumstances) we think that the agent should be aware of the morally relevant facts – whether they are causing harm, or are forgetting a commitment, or expressing a cruel motive – and their failures of awareness or knowledge are themselves culpable. (See Sher (2009) for a recent articulation and defence of the epistemic conditions for responsibility in these terms. The reasons for this culpability need further unpacking; something we return to in Section 4.). So rather, we should proceed by asking not whether individuals are in fact aware, in the (to be determined) relevant sense, of their implicitly biased actions; but rather, whether an individual should be aware in this sense, and whether their failures of awareness are culpable.10 We are asking what individuals should know qua responsible agents.11 This approach is consistent with the idea, defended in legal philosophy, that negligence does not require that an individual in fact be aware of the harm caused by her action; only that a reasonable person would have been.12 So our question is now more well focused: do the epistemic requirements that apply to responsible individuals include awareness of the kinds identified above? In the next section, I consider each requirement in turn in order to identify the relevant epistemic conditions for responsibility for implicitly biased actions.","We are interested, then, in whether the following claims are true as conditions for responsibility: individuals should be aware, introspectively, of the operation of implicit associations individuals should be aware of, or know about, the body of knowledge about people’s tendencies to harbour and display implicit biases, and make the relevant inferences about their own tendencies to express implicit biases. individuals should have observational awareness, or knowledge, of the effects of implicit associations on their behaviour. How might we proceed in evaluating which kinds of knowledge individuals should have? There is a methodological difficulty here, in that we have quite a few moving parts. The conditions for moral responsibility are in question, but the mental phenomena, processes, and actions they influence are in some respects unfamiliar – they are states about which we are learning ever more from the findings of empirical psychology. What should be held fixed, in trying to understand how we should think about the role of these unfamiliar processes in our agency? In question is not whether we are ever morally responsible for anything; that we are, at least sometimes, is a starting assumption of the argument.13 Those who argue that we are not responsible for actions influenced by implicit biases (or for the implicit associations themselves) are not seeking to vindicate general scepticism about responsible agency. At issue, rather, is whether certain aspects of our agency fall into the remit of those things for which we are responsible. One strategy we can employ is to consider other more familiar aspects of agency, and our judgements about responsibility regarding these. Should the same be said of implicit biases? Another strategy is to consider where a certain condition for moral responsibility would set the bar if it ruled out holding individuals responsible for a certain kind of state or behaviour: would it set the bar implausibly high, and lead us to general scepticism about responsibility? I will deploy each of these strategies in proceeding. In asking whether we should consider each epistemic condition as specifying a requirement for responsibility, we should ask whether it is a) desirable, and if so, b) possible, to meet the norm. For if it is impossible to gain awareness of some sort, then it would be unreasonable to require that individuals have such knowledge as a condition of responsibility. So our evaluation will require attention both to philosophical questions about the defensibility of certain conditions, and empirical questions about what sorts of awareness are possible, as far as the findings of empirical psychology reveal. Individuals should be aware, introspectively, of the operation of implicit associations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Is this a requirement on responsible agency that we should endorse? If so, then the non- culpable failure of individuals to have introspective awareness of the activation and operation in their cognitions of implicit biases would indeed mitigate from responsibility. My inability to detect the activation of an implicit association, or the extent to which it is influencing my cognitive processes, would suffice to absolve me from responsibility for any influence that association exerts on behavioural outputs. It should be noted that any failure of this kind would, seemingly, be non-culpable. Many have claimed that the operation of implicit associations – their activation and role in our cognitive processes – is not something to which we have access simply by reflecting, or by introspectively checking (see e.g. Saul, Washington & Kelly). And these assertions are borne out by empirical studies also. Indeed, various studies rely on the opacity of various aspects of our cognition. For example, the ‘evaluative priming test’ operates by presenting a prime (e.g. a black or white face, a picture of a prominent Republican or Democrat) and then measuring how long it takes individuals to recognise and categorise a negative or positive word. The idea is that if an individual is faster to categorise the negative terms as bad, than positive words as good, this reveals a negative association with the prime (because negative constructs were made more accessible by the prime).14 The idea is that individuals do not have introspective access to the ways in which the prime (of which they may or may not be aware) activates certain associations which influence their ease of categorisation. Such implicit processes in general do not appear to be ‘operationally transparent’ to us, such that it is not possible to have introspective awareness of their operation. If so then any failure to have such introspective awareness would be non-culpable. But should we endorse this requirement as a condition for responsible agency, such that non-culpable failure to meet it does exculpate? This requirement does not seem to me to be defensible. Firstly, it is not a standard we apply to other aspects of our cognition. Secondly, were such a standard to be applied, it would lead to radical scepticism about the possibility of responsible agency. On the first point: we should note that with respect to other aspects of cognition, it is not a condition on responsibility that individuals be aware of the cognitive processes that produce action. Here are two examples that help us to see this. Consider a case discussed by Nancy Snow (2006), of an individual instinctively making an intervention when she observes an elderly woman being cheated by a sales clerk (556–557). A central feature of Snow’s example is that the agent does not recognise that her sense of justice is activated (her justice related goals, in Snow’s terms); there are important aspects of her cognitions, then, in relation to which she lacks introspective awareness. But that she lacks awareness of this aspect of her cognition does not mean that she cannot be held responsible and praised for her actions. Likewise with blameworthy actions. Cases of forgetting are clear candidates of instances in which individuals might be blamed, despite lack of awareness of whatever processes led them to forget (indeed, awareness of this might alleviate the forgetting!). Consider Sher’s discussion of forgetting for which the agent is responsible: a distracted parent forgets that a dog is languishing in an overheating car. The parent lacks awareness of the cognitive processes whereby various competing demands crowd out the relevant belief; this leads her to act in a negligent and harmful way. But Sher asks us to share the intuition that the agent is still responsible for this failure, despite the lack of awareness of the processes (the failures of attention) that produced the action. Here are two cases, then, in which it is plausible that we hold an agent responsible (for creditable action, and for blameworthy action) even whilst they fail to be aware of the cognitive processes that play a role in producing the actions for which we hold them accountable. Thus, it seems that we do not require introspective awareness of the processes involved in the production of action as a condition on responsibility. My contention is not, of course, that these examples are exactly the same as cases of implicitly biased actions. There might ultimately be different judgements to be made about the two kinds of cases. However, what these cases show is that merely lacking introspective awareness of the processes involved in deliberation and action does not suffice to exculpate, and is consistent with praiseworthy and blameworthy action. Nonetheless, the more familiar processes described above are similar in some important respects (whilst of course dissimilar in others) to those involving implicit associations that produce implicitly biased actions: they are fast, automatic, not readily under the agent’s deliberative control, unreflective, and (in the latter case) processes the agent would not endorse, and productive of morally undesirable outcomes. Unless a case can be made for implicit associations being treated differently, then the fact that an agent lacks awareness of the operation of implicit associations would not be grounds for exemption from responsibility.15 For in general, we do not maintain that individuals should have this kind of knowledge of their cognitive processes; we do not require ‘operational transparency’ for individuals to be held responsible for actions that result from these processes. Indeed, if we did, very many of our actions – perhaps all of them? – would be exempt from responsibility, insofar as we are never aware of all of the processes that input into the production of action. Lack of awareness of this kind, then, does not exempt from responsibility. Even if it is not possible to secure this kind of awareness, that is irrelevant: because introspective awareness is not a requirement for morally responsible action. Individuals should be aware of, or know about, the body of knowledge about people’s tendencies to harbour and display implicit biases, and make the relevant inferences about their own tendencies to express implicit biases ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Let us turn, then, to the second formulation of the epistemic condition for responsibility in relation to implicit biases: that individuals should be aware of the relevant facts about the tendencies to express implicit biases in action. This is the sort of condition that Washington & Kelly argue for, in developing their ‘externalist’ account of the epistemic conditions for responsibility. The condition on responsibility here is one which maintains that individuals should, in some contexts, have an awareness of the facts uncovered by empirical studies about tendencies to discriminate as a result of implicit bias; inferences can then be made about one’s own propensity to do so. (Ultimately, steps to avoid biased actions can then be taken). There are two distinctive moves made by Washington & Kelly in advancing this condition: first, Washington & Kelly argue that what individuals should know (with respect to bodies of knowledge such as that about implicit bias) is role dependent in a significant sense. For example, a football commentator is required to keep up with changes to team composition in the seasonal transfer market; a heart surgeon is required to have up to date knowledge of aortic valve technologies; I am not required to have knowledge of either. The second key move is to point out that the possibility of meeting this requirement is dependent upon one’s epistemic environment. Accordingly, individuals could not meet the requirement to know about implicit biases in 1980, when the findings of empirical psychology about implicit cognition were less well established and less readily available. In this way the epistemic conditions for responsibility are ‘externalist’ in a significant sense – depending not on facts about the agent, but on facts about her epistemic environment – what is known, and what is available to be known. This leads Washington & Kelly to conclude that we should endorse this condition for moral responsibility, (b), in the case of individuals who (i) need to have information about implicit biases because of the social role they occupy; and (ii) have that information available in their epistemic environment, such that they are responsible for not availing themselves of that knowledge. Thus they maintain that responsibility ‘accrues first, or at least more quickly and disproportionately, to occupiers of specific social roles. These include those involved in hiring decisions, obviously, but also teachers, social workers, and those in other “gate keeper” positions whose activities can have the most amplified effects on various institutions and population level outcomes’ (in press). This issue speaks to the extent to which failure to meet this epistemic condition is culpable – is such ignorance itself a culpable failure to fulfil one’s epistemic responsibilities? For the contemporaneous hiring committee member, according to Washington & Kelly, it is; they can be held responsible (and perhaps blamed) for not knowing what they should – given the availability of the relevant information, and the role they occupy – be aware of. An implication of this is that for other individuals – the proverbial ‘person on the street’ – the lack of awareness of this knowledge is not culpable; her failure to meet this epistemic requirement does not mean she fails to grasp something she should grasp – so she, unlike the ‘gatekeepers’, can be exculpated from responsibility. This, Washington and Kelly maintain, is due to the fact that our epistemic environment is one in which knowledge about implicit bias ‘has still not risen to the level of common knowledge, and it probably will not any time in the immediate future. … and the percentage of people who have not heard of implicit bias at all is probably still quite high’ (in press). Accordingly, the extent of this exculpation depends on the ‘availability’ in the epistemic environment of the relevant knowledge. When it is pervasively known, and not just within the remit of academic researchers, this failure will be more widely culpable.16 Should we agree with Washington & Kelly’s statement of this epistemic condition for responsibility? To start, we should ask ourselves what motivates the move to restrict the normative requirement (should know) to those in ‘gatekeeper’ roles. Why should such individuals know about implicit biases? The answer is presumably that individuals in such roles are more likely to manifest implicit biases in ways that we can reasonably foresee to have deleterious effects on those who are stigmatised by or discriminated against by those biases, or in ways that affect population level distributions of benefits and burdens. But this motivating assumption seems to me implausible. Many people make decisions about who to hire, who to fire – literal ‘gatekeeper’ decisions – but also about who to grant a loan to, where to live, who to stop and search, who to give a lift to, what news stories to report (and how), who to write prescriptions for, who to sit by on a train, how to evaluate co-workers, who to smile at, what grades to assign or references to write, who to cross the road to avoid, who to believe, who to befriend …and so on. These kinds of interactions can all be affected by implicit biases (see Jost et al., 2009 for an overview).17 And it seems to me to be difficult to substantiate the claim that the reasonably foreseeable cumulative effects of these interpersonal interactions are not greater than those of the gatekeepers who make decisions with population level effects. The widespread impact of ‘informal’ discrimination and segregation in sustaining patterns of disadvantage is described in detail in Anderson’s work on racial inequality (2010). Valian (1999) also emphasises that population level discrepancies in the distribution of advantages can often be traced to the accumulation of small instances of differential treatment. Accordingly, it seems to me implausible to restrict the realm of ‘responsibility to know’ to some few individuals charged with hiring decisions or other population level outcomes, especially given what we know about the pervasiveness of implicit bias and its effects. If this is right, then the epistemic requirement to be familiar with the findings of empirical psychology will apply very widely indeed. Almost everyone will be subject to the requirement that they are aware of the relevant body of empirical findings about the nature of implicit biases, which of their actions are likely to be susceptible to it (and ultimately, how to avoid this). This knowledge is available in their epistemic environment, and it is knowledge that they should acquire. This seems like a very demanding epistemic condition! Washington & Kelly might appeal to the idea that such knowledge, whilst available, is not readily accessible as ‘common knowledge’, so failure to grasp it is not culpable.18 But the general epistemic requirement cannot be that the knowledge is ‘common’ in this way – otherwise the hiring committee members would also fail to be culpable (these gatekeepers, too, are making decisions where knowledge of implicit bias is not common knowledge). Moreover, if the potential and cumulative damage foreseeable as a result of the biased actions of, for example, loan-makers (and prescription-writers, testimony-takers, stop-and-searchers, and so on) is commensurate with that of hiring committee ‘gatekeepers’, there is no reason to subject them to different epistemic standards. But if the same epistemic standards apply more broadly it looks like we are committed to maintaining that very many individuals in fact are responsible for acting in an implicitly biased way, because they are culpably blameworthy for not knowing what they should know (namely, about a body of knowledge in empirical psychology). This seems rather implausible: not because it involves maintaining that very many individuals might be responsible for acting in implicitly biased ways (I think we may be); but that the grounds for this are the failures to engage with the findings of empirical psychology.19 To summarise: plausibly, very many of us play a role in sustaining and perpetuating patterns of disadvantage by acting in ways that are implicitly biased; also plausibly we have a responsibility to avoid doing so. But Washington & Kelly seem to be committed either to denying this, or to the claim that, therefore, almost everyone has a responsibility to engage with the findings of academic research.20 This is an implausibly demanding epistemic requirement, and one that accordingly, we should reject.21 One of the reasons we should reject this condition is because it supposes that the only access individuals might have to knowledge about implicit bias is via (some perhaps mediated) academic research. But there is reason to believe that this is not the case. This brings us to the third epistemic condition for moral responsibility. Individuals should have observational awareness, or knowledge, of the ways their behaviours are influenced by implicit associations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Let us now turn, then, to the third sense of awareness identified earlier. I have suggested that individuals cannot, and ought not to, have awareness of all of their (or the relevant) cognitive processes prior to acting; and that it is unreasonable to demand of almost everyone that they engage with complex bodies of academic knowledge about implicit biases. But there remains the question of whether individuals can, and should be expected to, have knowledge of the morally relevant features of the actions they perform. In this context, the morally important features include the property of being discriminatory, or treating differentially on the basis of some arbitrary feature (such as race, gender, or age). Do we in general require that individuals have awareness of these properties as a condition of responsibility? Do we require that individuals should have this knowledge? And to what extent is failure to have that knowledge culpable? Let us again consider two cases that will help us to think about this condition in general, before we turn to consider it in relation to our concern with implicit bias. Let us consider a variant of a case presented by Adams (1985), in which an individual acts in a way that expresses ingratitude; perhaps she is insufficiently warm in accepting a favour, say. This individual has not confronted the fact that she harbours this attitude of ingratitude, Adams tells us. And this is because she cares more about having a good opinion of herself than confronting unsavoury truths about herself. In being unaware of the presence of this attitude in her cognitions, the agent will therefore be unaware of her action manifesting this attitude; so unaware that her behaviours manifest the morally undesirable characteristic of being ungrateful. But we should nonetheless hold the agent responsible for the actions that are inflected by this attitude, Adams claims, because she should be aware of the fact that her actions are inflected by this attitude. If we share Adams’ judgements, then we should hold that it is not a condition for responsibility that the agent is in fact aware of the morally relevant features of their behaviour. A failure to know that one’s action is expressive of ingratitude does not absolve one of responsibility for acting ungratefully. This is something the agent should be aware of. In Adams’ version of the story, the agent’s failure of awareness is due to some further agential fault (the desire to have a good opinion of herself). Moreover, in this case, it seems clear that this is something that, with a bit of reflective work, the agent could become aware of. Observing that her action manifests ingratitude, then, is something we can reasonably expect the agent to do. Let us consider a second case, this time from Sher (2009). Sher presents an individual who tells an anecdote, failing to notice that in doing so she is being insensitive to her audience (an anecdote about a financial failure that is not well received by the individual who has recently experienced financial failure, say). Her behaviour has the property of insensitivity, but this is something of which she is not, at the time of her remarks, aware. Nonetheless, Sher urges us to share his intuition that this individual is responsible and blameworthy for her insensitivity. This is because she should be aware of her insensitivity, he claims, even if she is not.22 If we agree with Sher, then, again, we should not hold that it is a condition for responsibility that the agent is in fact aware of the morally relevant features of her behaviour (in this case, insensitivity). The agent should be aware of this feature, and we can suppose that this is something that, with sufficient reflection, it would have been possible for her to notice about her behaviour. Given this, it is a reasonable expectation to hold the agent to. These cases seem far from controversial, and are instances of common parts of our practice of holding each other responsible. So far, then, the epistemic condition I have described in this section is far from revisionary. Finally, then, let us consider whether the same considerations apply to the case of actions that express implicit bias – such as the differential evaluation of CV whose only difference is the racialised name at the top. These actions have the morally undesirable characteristic of being discriminatory. If we treat this case as a direct analogue to those considered above, then the mere fact that an agent is not aware that her action is discriminatory will not suffice to exculpate. Ignorance alone does not excuse. So we turn to our next question: is awareness of the discriminatory nature of her action something that we think the agent should have knowledge of? Certainly we should strongly affirm the desirability of this: insofar as agents are expected to bring their behaviour into conformity with moral norms, we should insist that they ought to self-monitor, and be aware of ways in which their behaviour may depart from these moral standards. But whether this expectation is reasonable in the case of implicit bias depends on whether it is possible for agents to have this awareness. Given the implicit nature of the biases at issue, shouldn’t we suppose that this is not something that individuals can reasonably be expected to be aware of? As we have already seen (1.5 above), the empirical evidence does not provide support for that claim, and rather indicates the contrary. First, it is worth noting that whilst some authors have characterised implicit associations as attitudes of which the agent is unaware (Saul, Washington & Kelly, see above), DeHouwer et al. (2009, 357–358) point out that there is in fact little reason to endorse this claim. Whilst in some implicit measures the way in which the association is activated is something that the agent is not aware of (because the association is primed) this does not mean that the effects of the bias on behaviour are not something that the agent can be aware of. Likewise, that an attitude is measured with an implicit measure (a measure that does not require self-report or reflective articulation of one’s attitude) does not mean that the behavioural manifestation of the attitude cannot be reported on. Recall the studies by Monteith et al. (2001), which I mentioned earlier in introducing the notion of observational awareness. In this study, participants undertook a race IAT (pairing white or black names with pleasant or unpleasant terms). The participants were then asked to evaluate their performance on the IAT, and the interesting finding for present purposes was that a significant number of participants (64%) were able to report on the basis of observational awareness of their own responses that they responded more slowly when pairing black names with pleasant terms (than with unpleasant terms, and than white names with pleasant terms). These findings garner support from more recent experimental studies.23 Hahn, Judd, Hirsh, and Blair (2013) examined individuals’ accuracy of predictions regarding the expression of implicit biases. They found that individuals were, when asked to carefully reflect, able to accurately predict this, both in experimental terms: ‘My sorting of [the congruent pairings] will be very/moderately/slightly easier...’ (p.5) and in conceptual terms: ‘My true implicit attitude is a lot/moderately/slightly more positive towards white’ (p.8). This was so even where individuals showed discrepancies between recorded implicit attitudes, and reported explicit attitudes, such that predictions were not being made on the basis of explicit attitudes of which the participants were aware and alert to the possibility of their subtle influence on behaviour.24 We might think that the participants made their predictions on the basis of general knowledge of social context rather than on the basis of awareness of their own behavioural dispositions. But the experimenters ruled this out by asking participants to predict both their own, and the average responses. There was divergence between these predictions, with predictions about their own behaviour more closely matching biases measured (p. 10). Or, we might think that reports of explicit attitudes are unreliable, such that individuals in fact had explicitly biased attitudes, and made reports in anticipation of these attitudes influencing their behaviour. This possibility seems unlikely, given that the effect was found in the condition in which participants were told that the results did not reveal their ‘true attitudes’, but merely cultural associations (manipulation checks revealed that these participants accepted this story) (p.7). This study is important, because it indicates that individuals are not only able to detect morally relevant features of their actions post hoc; they were also able to predict morally undesirable features ex ante.25 In a way, this should not seem surprising given the remarks above: there is nothing about the manifestation of an implicit association or attitude in behaviour that prevents it from being accessible to report on. But it is surprising in the context of philosophical discussions that have supposed that implicitly biased behaviour is something of which individuals are not (in some sense) aware, and have elided the different notions of awareness at issue. But these assumptions are not supported: there is evidence that supports the claims that, with reflection, individuals are at least sometimes able to detect and predict discrepant responses. These are tentative findings: we might wonder whether the possibility of detecting and predicting biased responses extends to the full range of behaviours that might be influenced by bias. This is particularly so where the biased actions are identified in studies that observe statistical tendencies across groups, rather than in intrapersonal differences in responses. And we should want to know more about the functioning of this awareness ‘outside the lab’. Nonetheless, the present findings provide reason to examine how widespread observational awareness of ones actions as biased is, given that it is at least possible sometimes to have such awareness.26 The key points, then, are as follows: firstly, individuals should have knowledge of the morally relevant features of their actions, such as that they are discriminatory. It is only if individuals non-culpably lack knowledge of this sort, then, that they should be absolved of responsibility for discriminatory, implicitly biased behaviour. One way of non-culpably lacking this knowledge would be if it were not possible to have such awareness. But the evidence suggests that this sort of awareness is not ruled out; and moreover, sometimes individuals are able to gain this sort of awareness.27 This raises the question, then, of whether there are other grounds for supposing that individuals who lack awareness of the ways in which their actions are inflected with bias are culpable for this lack. I suggest two possible (non-exhaustive, non-exclusive) explanations that might implicate other imperfect cognitions in our responsibility for implicitly biased behaviour.","We have arrived at the following question: when individuals lack the awareness that it appears possible for them to have, with respect to the morally relevant properties of their implicitly biased behaviour, is this lack a culpable one? Not all failures to know what should be known are culpable failures. What might be said about the failure to have the relevant kind of awareness of one’s own actions as implicitly biased? Here are two possible answers that yield different judgements about the extent of an individual’s blameworthiness for their actions. Failures of attentiveness ~~~~~~~~~~~~~~~~~~~~~~~~~ It might be tempting to suppose that the kind of awareness at issue requires some revisionary moral understanding, such that the culpability of any failure to meet the epistemic condition is considerably mitigated. Not knowing what one should know may be less culpable if it is harder to gain that knowledge, because it is not yet normalised and nor accessible to all as moral knowledge (it is rather at a ‘frontier’ of moral knowledge) (Calhoun, 1989, cf. Washtington & Kelly, above); or if the pervasive tendency is not even to think of an issue as ‘morally charged’, so that moral reflection is not focused upon that action or its consequences at all (Isaacs, 1997). Examples of this include the failure of many to know that the supposedly gender-neutral use of the pronoun ‘he’ perpetuates sexism. Such sexist language use may be less blameworthy if the knowledge of this that we all should have is not yet normalised, or requires some imaginative leap to access; or if it is not yet clear that this domain of our activity is even morally scrutable.28 Analogously, we might think that the common sense view of discrimination involves the explicit intention to treat differently, or the explicit intention to harm others or manifest ill will (Garcia, 1996). Given this, the failures of awareness with respect to implicitly biased actions may be candidates for a kind of moral ignorance that is mitigated in the ways described above. Perhaps the knowledge that one’s actions can be discriminatory, even if not intentionally so, is not yet ‘normalised’ so as to make this knowledge readily accessible; or perhaps such behaviour does not prompt reflection, as it is not on the radar as a ‘moral issue’ at all. In this respect we can see how inferential awareness of the findings about implicit bias may make it considerably easier to possess observational awareness of one’s actions as biased. Knowing that one is likely to be biased, given the body of empirical evidence we now have, can introduce pressures to reflect that might generate observational awareness. To the extent that one lacks inferential awareness, then, it may be harder to know what one should know. Moreover, we might think culpability is mitigated further if this moral ignorance is compounded by misleading introspective evidence, which seems to suggest to us that our motives are good, and without discriminatory content. However, these remarks do not seem quite right as a full diagnosis of the failing, because what is at issue is not simply whether a certain behaviour is morally problematic or would fall under the rubric of ‘discrimination’; but rather whether the fact that a behaviour displays differential treatment is noted at all. The lack of awareness at issue here is not whether the situation is a ‘morally charged’ one; it is surely accepted that one’s behaviour would be morally questionable if it were known to involve such differential treatment. Rather, the lack of awareness pertains to the behaviour manifesting differential treatment itself. There is some lack of attentiveness, a failure to notice, that means the morally relevant feature of behaviour is not something of which the agent is aware (irrespective of how it would, once noticed, be labelled). It might be that this failure of attention is driven by the mistaken beliefs that, since one is not intentionally discriminatory, it is simply not possible for one’s actions to have discriminatory effects. But on the other hand, some failures to notice morally important things can be indicative or expressive of an agent’s evaluative stance – what they care about, and how much. Lacking the motive to reflect can indicate what one takes to be worthy of moral scrutiny. Such lacks are ones that it is feasible to hold individuals responsible, and blameworthy, for (Sher, 2009; Smith, 2005). The extent to which an individual is culpable for the failures of awareness with respect to implicitly biased actions, then, will depend on answers to further questions about the role of implicit associations in our broader agential structures, and how their expression is related to the values we hold.29 Self-deception ~~~~~~~~~~~~~~ It is one thing to lack a motive to reflect; another to be motivated not to reflect, as Calhoun points out: ‘self-interest can motivate the suppression of reflection ... self- deception is a matter of not being motivated to examine one’s actions or reasoning too carefully, lest something unpleasant turn up’ (399).30 Whilst it might be possible to detect in one’s actions differential treatment – or to predict it – when specifically prompted to do so, these are difficult truths to confront. Acknowledging not only one’s complicity in but perpetuation of patterns of discrimination, albeit in subtle and unintended ways, is something that no doubt many of us find hard to accept. The belief that one’s actions are implicitly biased, and other implied beliefs about one’s role in sustaining patterns of discrimination, are clearly beliefs that, for a range of reasons, agents might be motivated not to confront. Conversely, the belief that one’s actions are consistent with one’s moral ideals (of non-discrimination, of being evidence sensitive and unbiased) is one that agents are motivated to maintain. That we are motivated to avoid the sort of moral reflection that might overturn those desirable beliefs, and turn up these undesirable facts about our propensity to bias, is a plausible explanation of this lack of awareness. After all, there is ample evidence from empirical psychology that we are motivated to maintain a positive self-concept (Brown, 1986; Suls, Lemos, & Stewart, 2002). Indeed the motive to present a positive view of ourselves is what raises concerns about the validity of self-report measures (we do not want to see ourselves, or let others see us, as prejudiced), and makes access to attitudes via implicit measures so valuable.31 Moreover, recent studies indicate that such a motivation has a role in sustaining our view of ourselves as immune to bias: Pronin and Kugler (2007) found individuals to over-rely on on misleading introspective evidence of propensity to bias (introspection revealing – surprise! – no bias). Meanwhile participants ignored behavioural evidence of their own bias that they were willing to take as evidence of bias in others: ‘actors … preferred to see themselves as bias free’ where it was possible to ignore evidence to the contrary (576). This form of self-deception could explain the failure of individuals to have awareness that their actions manifest implicit bias.32 Would this explanation yield the judgement that such ignorance is culpable? Whilst pervasive, such self-deceptions are not so overwhelming as to be insurmountable: in Pronin and Kugler’s study, when experimenters reminded subjects of the unreliability of introspection (in contrast to observed behaviour) as a guide to their own bias, the participants no longer denied their susceptibility to bias. Insofar as the motivated lack of reflection is serving to bolster a misleadingly positive view of oneself, and serving to cover up morally undesirable aspects of one’s actions that are not otherwise impossible to detect, a case can be made that failures of awareness resulting from such mechanisms are culpable. In this case, such ignorance would be culpable, and the lack of awareness of one’s actions as implicitly biased would not serve to exculpate from responsibility.","I have argued that there three different claims that need teasing apart when thinking about whether individuals have awareness in relation to implicit bias: whether we have introspective awareness of its operation in influencing behaviour; whether we are aware of the relevant bodies of knowledge about implicit bias that enable us to infer our likely susceptibility to such biases; whether we are observationally aware of the fact that our behaviours have the morally undesirable property of being discriminatory. In each case, the relevant question is not whether an individual has this awareness, but whether they should have such awareness, and whether lacking it is culpable. I argued that lacking introspective awareness of the operation of implicit associations, or lacking inferential awareness of the propensity to display implicit bias, does not in itself exculpate. Rather, I argued that we should have observational awareness of the morally relevant features of our behaviour, namely, their discriminatory nature. And indeed the empirical studies on implicit bias do not entail that such knowledge is impossible to gain; in fact, some show that sometimes individuals do have this kind of awareness. When we lack it, are we culpably ignorant? Perhaps so, if this lack is due to failures of attention that express our values, or self-deception that narrowly serves our interests. In this way, our responsibility for implicitly biased actions may be bound up with imperfect cognitions of other kinds."],["In this paper I explore the nature of confabulatory explanations of action guided by implicit bias. I claim that such explanations can have significant epistemic benefits in spite of their obvious epistemic costs, and that such benefits are not otherwise obtainable by the subject at the time at which the explanation is offered. I start by outlining the kinds of cases I have in mind, before characterising the phenomenon of confabulation by focusing on a few common features. Then I introduce the notion of epistemic innocence to capture the epistemic status of those cognitions which have both obvious epistemic faults and some significant epistemic benefit. A cognition is epistemically innocent if it delivers some epistemic benefit to the subject which would not be attainable otherwise because alternative (less epistemically faulty) cognitions that could deliver the same benefit are unavailable to the subject at that time. I ask whether confabulatory explanations of actions guided by implicit bias have epistemic benefits and whether there are genuine alternatives to forming a confabulatory explanation in the circumstances in which subjects confabulate. On the basis of my analysis of confabulatory explanations of actions guided by implicit bias, I argue that such explanations have the potential for epistemic innocence. I conclude that epistemic evaluation of confabulatory explanations of action guided by implicit bias ought to tell a richer story, one which takes into account the context in which the explanation occurs. --------------------------------------------------------------------------------","In this paper I explore the nature of confabulatory explanations of action guided by implicit bias in the non-clinical population. My aim is to highlight the potential epistemic benefits of some confabulatory explanations and tell a richer story about the overall epistemic status of such explanations. Although confabulation is characterised and often even defined on the basis of its epistemic costs, I argue that some confabulations can play a positive role, not just because they act as a psychological defence by enhancing coherence, stability, self-confidence, and well-being (Ramachandran, 1996: 351; Fotopoulou, 2008: 542), but also because they bestow epistemic benefits which are otherwise unavailable. In section one I describe two imagined cases of confabulatory explanations of actions guided by implicit bias.1 In section two, I characterise non- clinical confabulation by identifying some common features of the phenomenon, features shared by my two imagined cases. In section three, I introduce the notion of epistemic innocence to capture the status of those cognitions that have obvious epistemic faults, but also have some significant epistemic benefits that could not be otherwise obtained. In section four, I argue that confabulatory explanations of actions guided by implicit bias have the potential to deliver epistemic benefits, insofar as they maximise the acquisition of true beliefs in the long run by filling an explanatory gap, and they help the agent maintain consistency among her cognitions. I also argue that, at the time of the confabulatory explanation, no alternative explanations are available—in a sense to be explained—to the subject. I conclude that epistemic evaluation of confabulatory explanations should be indexed to context, taking into account the (un)availability of alternatives and potential epistemic benefits. This allows us to resist a kind of trade- off view about the epistemic status of confabulatory explanations: the view that pragmatic benefits come at the expense of epistemic ones. A closer focus on the potential epistemic benefits of these cognitions, as well as the context in which they occur, can result in a more careful epistemic evaluation of them.","In this section I describe two imaginary cases of confabulatory explanations that could occur in the non-clinical population. I will assume Jules Holroyd’s definition of implicit bias in the discussion which follows, according to which: ‘[an] individual harbors an implicit bias against some stigmatized group (G), when she has automatic cognitive or affective associations between (her concept of) G and some negative property (P) or stereotypic trait (T), which are accessible and can be operative in influencing judgment and behavior without the conscious awareness of the agent’ (Holroyd, 2012: 275). Implicit biases then can be understood as ‘largely unconscious tendencies to automatically associate concepts with one another’, such tendencies can result in judging ‘members of stigmatized groups more negatively’ (Saul, 2012b: 244). Care is needed when using the term ‘unconscious’ with respect to implicit bias. Gawronski, Hofmann, and Wilbur (2006) distinguish three types of awareness: source, content, and impact awareness. If a subject has source awareness of some attitude of hers, she has awareness of the origin of that attitude. If a subject has content awareness of some attitude of hers, she has awareness of the attitude itself. Finally, if a subject has impact awareness of some attitude of hers, she has awareness of the influence of that attitude on other psychological processes (Gawronski et al., 2006: 486). Gawronski and colleagues’ review of empirical evidence suggests that implicit attitudes only differ from explicit attitudes with respect to impact awareness. They conclude that the term “unconscious” is adequate for indirectly assessed attitudes only with regard to one particular aspect: impact awareness. However, the term “unconscious” is inadequate when it is assumed to imply a lack of source awareness or content awareness. Note that though the evidence suggests that source and content awareness of implicit attitudes is possible, that is not to say that subjects always have such awareness in ordinary settings. We can be cautious here and take a lesson from the nearby literature: in her discussion of moderating automatic stereotypes, Irene Blair claims that though ‘the evidence is compelling with regard to the possibility of moderating automatic stereotypes, the likelihood of such moderation in everyday social encounters is not yet known’ (Blair, 2002: 249). Similarly then, though the work Gawronski and colleagues reviewed showed evidence that content and source awareness is possible, that is not to say that ‘unconscious’ used in this way is always inappropriate. Empirical work has shown that implicit biases are held by most people, even those who avow egalitarian positions, or are members of the targeted group. Results from Implicit Association Tests (IATs) support this. IATs work by measuring the speed at which subjects pair two categories of object with, for example, pleasant and unpleasant stimuli (e.g. the words ‘wonderful’ and ‘awful’) or stereotypical and unstereotypical stimuli. The idea behind the tests is that we can discover which categories a subject associates with one another. This is done by measuring the categorisation performance of combinations of categories (De Houwer, Teige-Mocigemba, Spruyt, & Moors, 2009: 347). In a review of over 2.5 million IAT results across seventeen topics, Brian Nosek and colleagues report that ‘[i]mplicit and explicit comparative preferences and stereotypes were widespread across gender, ethnicity, age, political orientation, and region’ (Nosek et al., 2007: 40). In the rest of the paper, I will follow Holroyd in understanding implicit biases as operative when they ‘produce a distorting influence on judgement and hence behaviour informed by that judgement’ (Holroyd, 2015). The cases I introduce in the next two sections are ones in which implicit biases are operative in this sense. Implicit gender bias and the case of Roger ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In a study on the influence of the gender of an applicant on reviewing CVs, both male and female participants were more likely to vote in favour of hiring a male applicant than a female applicant when the CVs presented were otherwise identical (Steinpreis, Anders, & Ritzke, 1999: 509). The authors of the study suggest that the results ‘indicate a gender bias for both men and women in preference for male job applicants’ (Steinpreis et al., 1999: 510). Both male and female participants were ‘significantly more likely to hire a potential male colleague than an equally qualified potential female colleague’, and were ‘more likely to positively evaluate the research, teaching, and service contributions of a male job applicant than a female job applicant with an identical record’ (Steinpreis et al., 1999: 522). Corinne A. Moss-Racusin and colleagues found that when assessing application materials of a student applying for a Laboratory Manager position, ‘[f]aculty participants rated the male applicant as significantly more competent and hireable than the (identical) female applicant’, as well as ‘select[ing] a higher starting salary’ (Moss-Racusin, Dovidio, Brescoll, Graham, & Handelsman, 2012: 16,474). Monica Biernat and Diane Kobrynowicz predicted that ‘it would be more difficult for women than men […] to document their ability in a competence-related domain’ (Biernat & Kobrynowicz, 1997: 554). They ran two studies which confirmed this prediction showing that women were required by participants to ‘jump through more hoops’ in order to show that they were able to fill a position (Biernat & Kobrynowicz, 1997: 554). These studies suggest that ‘[w]ithout any intention of bias, once we have categorised someone as male or female, activated gender stereotypes can colour our perception’ (Fine, 2010: 56). Having pointed to empirical work showing the presence and effects of gender bias, here is the case of Roger: Roger is on a hiring panel deciding from a stack of CVs which candidates to invite to interview. Roger thinks of himself as an egalitarian, and not as somebody who is sexist. The CVs are not anonymous with respect to gender. Roger chooses not to invite any female applicants to interview. Katie is one of the female candidates who Roger chooses not to invite to interview. Katie’s CV is of equal or better quality than at least some of her male competitors who did get invited to interview, and had Katie’s CV been headed with a male name, Katie would have been invited to interview. The case might be made more plausible if the job were a traditionally male one, say, an ‘executive chief of staff’ position (Biernat & Fuegen, 2001: 711). As Madeline Heilman points out, ‘research has repeatedly demonstrated sex bias in employee selection processes […] with male applicants generally recommended for hire and seen as more likely to succeed than female applicants with the identical credentials when jobs are male in sex-type’ (Heilman, 2001: 660). Let us say that Roger’s decision not to invite Katie to interview was guided by his implicit bias against women. I will not worry too much about how exactly to cash out the guided by locution here, the thought is only that Roger’s implicit bias was efficacious in his decision not to invite Katie to interview, and the appropriate counterfactuals are true (if Roger did not have an implicit bias against women, Katie would have been invited to interview, and if Katie’s CV was headed with a typically male name, Katie would have been invited to interview). Let us ask Roger to explain his decision. He claims that Katie’s CV indicates that she is not good enough for the job, and so she ought not to be interviewed. Later I will argue that Roger’s explanation of his decision not to invite Katie to interview is a confabulatory one (Section 2.6). Implicit race bias and the case of Sylvia ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Empirical work has shown that implicit race bias is exhibited by many people. Such biases may affect not only our beliefs, but also the way we perceive the world. A study by Keith B. Payne (2001) showed that participants were able to more quickly identify guns (as opposed to non-gun tools) when they were primed with a Black face, as compared to when they were primed with a White face. In this experimental set up, the ‘presence of Black faces facilitated the identification of guns relative to the presence of White faces’ (Payne, 2001: 185). When a time condition constraint was introduced into an otherwise identical experiment, participants more often mistakenly identified tools as guns when they had been primed with a Black face as compared to when they had been primed with a White face. Payne identifies as the ‘critical finding’ that ‘priming participants with a Black rather than a White face was sufficient to make them call a harmless item a gun’ (Payne, 2001: 188). In a similar experiment in which participants were instructed to shoot at subjects holding a gun, and not shoot at subjects not holding a gun, participants were found to shoot an armed subject quicker if that subject was African American than if he was White, and the decision to not shoot an unarmed subject was made more quickly when the target was White and not African American (Correll, Park, Judd, & Wittenbrink, 2002: 1317). Finally, behaviour which is ambiguously aggressive committed by a black person is more likely to be perceived as hostile than the same behaviour committed by a white person (see for example Duncan, 1976; Sagar & Schofield, 1980). Implicit bias then is ‘getting to us even before we get to the point of reflecting upon the world—it affects our very perceptions of that world’ (Saul, 2012b: 246). In summarising a body of work on this Tamar Gendler claims that ‘[h]undreds of studies conducted in dozens of laboratories over nearly two decades have shown that—for most twentieth and twenty-first-century American participants in most circumstances’ the White-or-positive and Black-or-negative pairing ‘is more natural’ than the White-or-negative and Black-or-positive pairing. Gendler claims that this suggests ‘that the former categories are, for most participants, more readily constructed or more easily accessed, and hence experienced as more “natural” than the latter’ (Gendler, 2011: 52). Here is the case of Sylvia: Sylvia is walking down the road on her way to work. Sylvia thinks of herself as an egalitarian, and not as somebody who is racist. She sees a black man walking towards her. Sylvia crosses the road. The man is not acting in a threatening manner. Had the man not been black, Sylvia would not have crossed the road. Let us say that Sylvia’s decision to cross the road was guided by her implicit bias against black people. Again, what I mean by the guided by locution here is just that Sylvia’s implicit bias against black people was efficacious in the production of her behaviour, and that the appropriate counterfactuals hold (if Sylvia did not have an implicit bias against black people, she would not have crossed the road, had the man not been black, she would not have crossed the road).2 Let us ask Sylvia to explain her behaviour. She claims that the man was behaving in a threatening way, and so she crossed the road out of fear for her safety. I have built into the case that the black man is not behaving in a threatening way, and let us suppose that if one did not hold an implicit bias towards black people, no threat would be felt. In the next section I will argue that Sylvia’s explanation of why she crossed the road is a confabulatory one (Section 2.6). Next I will look at the explanations offered by Roger and Sylvia in the light of five features of confabulatory explanations, to support the claim that at least sometimes, explanations of decisions or actions guided by implicit bias are confabulatory.","Here I do not aim to solve the thorny issue regarding how best to define confabulation, rather, I will just highlight what the two cases described in the previous section have in common, focusing in particular on those features that are relevant to their epistemic status. As we will see, these features appear in popular accounts of confabulation. I will consider five common features of confabulatory explanations, they (1) are false or ill- grounded; (2) are offered as the answer to a question; (3) have a motivational component; (4) fill a gap, and (5) are reported without any intention to deceive. These are not necessary and sufficient conditions, rather, they have been thought to be common features of confabulatory explanations, and they are features which characterise the cases of Roger and Sylvia. False or ill-grounded ~~~~~~~~~~~~~~~~~~~~~ Confabulatory explanations are epistemically faulty. Generally speaking, they are false explanations, their being so has been identified as a key or defining feature of them (see for instance Berrios, 2000: 348; McKay & Kinsbourne, 2010: 289). However, someone might confabulate and hit on something true (see for example William Hirstein’s discussion of subjects with Korsakoff’s syndrome—a form of amnesia caused by alcohol abuse or severe malnutrition—who happen to confabulate correctly (2009: 3), or Ryan McKay and Marcel Kinsbourne’s example of a subject who lacks ‘access to his biographical information, yet by chance may confabulate the correct answer when asked his age’ (2010: 289)). We cannot then, rule out by definition the possibility of true confabulatory explanations. But even when confabulatory explanations are not false, they are epistemically poor in other respects. A key epistemic feature of confabulations then is that they are ill-grounded or poorly supported by evidence (Hirstein, 2005: 33–4). In confabulation, the explanation offered for a decision or action does not reflect what caused that decision or action and also poorly matches the details of the situation at hand. For example, the participants in Johnathan Haidt’s experiment were presented with a scenario in which two siblings engage in incest. Asked how they felt about the scenario, most of the participants claimed that it was wrong, and justified that attitude on the basis of the risks of inbreeding, even when the case of incest they were presented with explicitly excluded the possibility of reproduction after sexual intercourse (Haidt, 2001). Explanations of actions guided by implicit bias look to share the same features as the classical cases of non-clinical confabulation reported by Haidt. They are false explanations, since they do not map on to what was efficacious in guiding the subjects’ decisions or actions. They are ill-grounded, insofar as they misrepresent important features of the situations. Roger chooses not to invite Katie to interview, when her CV is of equal or better quality than at least some of her male competitors. Sylvia crosses the road upon seeing a black man, when the black man is not acting in a threatening way. Roger offers an explanation of his decision by claiming that Katie’s CV was not as strong as the CVs of the (male) candidates who were invited to interview. Sylvia explains her action by claiming that the black man was behaving in a threatening way. The reasons given in these explanations were not the reasons which were actually in play with respect to Roger’s decision and Sylvia’s action, and they are not supported by evidence, and therefore the explanations are both false and ill-grounded. Provoked ~~~~~~~~ In the psychological literature there is a recurrent distinction between two types of confabulation. A confabulation is spontaneous when it is not elicited by questioning, and it is provoked when it is offered as a response: the person is asked to offer an explanation for something, and she provides a confabulation (Kopelman, 1999: 197–8; Hirstein, 2005: 20). Provoked confabulations can occur in people ‘who are fully in possession of most of their cognitive faculties, and able to respond correctly to all sorts of requests and questions’ (Hirstein, 2005: 21). In the two cases of confabulation I considered in the previous section and in many other instances of non-clinical confabulation, the explanation is likely to be generated as a response to a specific question, such as ‘Why did you not invite this candidate to interview?’ or ‘Why did you cross the road?’ The second feature of confabulation then, which is shared by the cases of Roger and Sylvia, is that they can be provoked as responses to questions. Motivated ~~~~~~~~~ A third feature common to many instances of confabulation is that they involve a motivational element, which means that they are goal-directed states or processes (Bayne & Fernandez, 2009). This is something ‘many accounts of confabulation’ have recognised (Bortolotti & Cox, 2009: 954, see also Zangwill, 1953: 700). This motivational element might be a causal factor for the explanation’s being provided at all (for example, to cover a gap in memory, Bonhoeffer, 1901, cited in McKay & Kinsbourne, 2010: 291), or as a factor with respect to the content an explanation has (Fotopoulou, Conway, & Solms, 2007; Fotopoulou et al., 2008; Metcalf, Langdon, & Coltheart, 2010). So disagreement with respect to the role of motivational factors in confabulation concerns whether they play a role in the very existence of a confabulation, or whether they play a role additionally in the content of a confabulation. Confabulatory explanations are then, at least in part, motivated explanations, that is, some pro-attitude plays a role in either the very formation of a confabulation, or more specifically, the content of a confabulation. We can distinguish between a motivation to offer an explanation as opposed to no explanation (type-a), and a motivation to offer an explanation with a specific content as opposed to an explanation with a different content (type-b). In the non-clinical context of confabulatory explanations of decisions or actions guided by implicit bias, confabulatory explanations may be motivated in both senses. In some cases, they allow the person to avoid the appearance of ignorance or incompetence. In response to a question to which a person does not know or cannot know the answer, she may confabulate when admitting that she does not know the answer would be costly. This typically occurs when ‘the provoking question touches on something people are normally expected to know’ (Hirstein, 2005: 30), (we might think that someone on a hiring panel ought to be able to give reasons for rejecting candidates, and someone crossing a road ought to be able to give reasons for so doing). Thus, confabulatory explanations may protect a person from claiming not to know why they decided or acted as they did. Equally, there are cases in which the motivational component plays a role in the very content of the explanation. Perhaps the most obvious motivational component guiding the content of a confabulation is the desire a subject might have to maintain coherent beliefs about the self or a good self-concept. It looks like both kinds of causal work are being done by confabulatory explanations offered in the cases of Roger and Sylvia. In the first place, the explanation might be motivated by a desire not to be dumbfounded, or the desire not to fail to offer an explanation at all for one’s decision or action. In the second place, the content of the explanation might be motivated by one’s desire not to appear sexist or racist. Filling a gap ~~~~~~~~~~~~~ Related to the last point, and especially with regards to type-a motivation, confabulatory explanations have been thought to, in some sense, fill a gap. According to the classic conception of clinical confabulation as a consequence of a memory impairment—that best applies to the phenomenon when it occurs in the context of amnesia and dementia—confabulations are ‘stories produced to cover gaps in memory’ (Hirstein, 2005: 32). But confabulation plays a gap-filling role in more cases than those in which there is a memory deficit (I do not think that Roger and Sylvia offer the explanations they do because they fail to remember why they acted as they did3). The gap-filling claim needs to be construed more broadly and not couched only with respect to a gap in memory. We might instead think of confabulatory explanations as filling gaps ‘at a certain level in the cognitive system’, insofar as they help produce ‘complete, coherent representations of the world’ (Hirstein, 2005: 30). In the cases I have given of explanations of actions guided by implicit bias, the person may not be aware of having a bias, or at the very least, may not be aware of that bias’s influence, (recall that this is what Gawronski and colleagues term impact awareness (2006: 486)). This is the gap that might be filled by the confabulatory explanation. Arguably, explanations of decisions or actions guided by implicit bias fill a gap which could not be otherwise filled—I will discuss the idea of alternative explanations being in some sense unavailable to subjects who confabulate later (Section 4.2). No intention to deceive ~~~~~~~~~~~~~~~~~~~~~~~ My talk of a motivational component and of gap-filling may be suggestive of an intention to deceive (even if it is an intention to deceive oneself) when subjects offer confabulatory explanations. There is an interesting literature on the potential overlap between confabulation and self-deception which I cannot address here (see, for example Bayne & Fernandez, 2009; Hirstein, 2000, 2005; Ramachandran, 1996). For my purposes it suffices to say that the ‘orthodox position’ is that people who confabulate should not be understood as lying (Hirstein, 2005: 28), and this follows from there being no intention to deceive in confabulation.4 This is made very clear in the epistemic concept of confabulation, according to which a confabulation is ‘a certain type of epistemically ill- grounded claim that the subject does not know is ill-grounded’ (Hirstein, 2005: 33). According to this conception, when a person confabulates she offers an ill-grounded claim and she is not aware of that claim as being ill-grounded, and she does not believe contrary to it. If a necessary condition on intending to deceive another is that the deceiver believes contrary to what she avows, then she who confabulates does not so intend, because she does not recognise her confabulation as epistemically ill-grounded. People offering confabulatory explanations of decisions or behaviour guided by implicit bias have no intention to deceive: subjects can lack impact awareness of the operations of implicit biases, and so it would be strange to understand the subjects in my examples as seeking to deceive in order to cover up their implicit sexism or racism which is operative when they act. Given that their attitudes are implicit ones and are not recognised by the subjects as attitudes they have (or at least, not recognised as attitudes which are influencing their decisions or behaviour), subjects should not be construed as intentionally seeking to cover up those attitudes (see Holroyd, 2015, for more discussion on this point). Do Roger and Sylvia confabulate? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Are the explanations given by Roger and Sylvia genuine cases of confabulation? I have outlined five features common to confabulation, these are their being false or ill- grounded, provoked in response to questioning, motivated, filling a gap, and there being no intention to deceive on the part of the confabulator. In my outline of each of these features, I have suggested that they are ones which characterise the cases of Roger and Sylvia. Care needs to be taken though with respect to how the cases of Roger and Sylvia are described, since under a certain kind of description, Roger and Sylvia are not confabulating. In the cases as I have described them Roger genuinely (and falsely) believes that Katie’s CV is of inferior quality, and Sylvia genuinely (and falsely) believes that the black man is behaving in a threatening way. If this is the interpretation of the circumstances—that the subjects have false beliefs and base their decisions and actions upon those beliefs—we would not be in the presence of confabulatory explanations. Our subjects would be explaining their behaviour by appealing to a belief of theirs for which the evidence was poor. Under the hypothesis I am currently discussing, Roger believes that Katie has a worse CV than her male competitors and he offers a true explanation of his action when he explains why he did not choose to invite her to interview. His belief about the quality of the CV is not supported by the evidence, but instead is guided by an implicit bias, the belief is false but the explanation of the choice is not confabulatory. The explanation is one which is both true and well-grounded. This is not how I want to understand these cases. Rather, let us ask Roger why he has the belief that Katie’s CV is of inferior quality. Now Roger claims that he has the belief because the CV is inferior. Now we are in the presence of confabulatory explanations since the reason Roger has the belief that the CV is inferior is not because it is, it is rather because of some implicit bias he has. My cases then ought to be understood as ones in which the subject is asked to explain why they acted in the way that they did, and they give an explanation in which they cite what they take to be a fact about the world which both informs the belief and explains the decision or action. When Roger says that he did not invite Katie to interview because her CV was inferior he confabulates, because he cites something which was not efficacious in the making of his decision. This is not to say that it is sufficient for confabulation that one is wrong about the cause of one’s belief, and it is for this reason only that Roger’s explanation is a confabulatory one. As we have seen, Roger’s explanation exhibits many other features typical of confabulatory explanations. Note also the way the cases were described: these are not cases of explicit sexism or racism, rather, they are cases of subjects who explicitly take egalitarian positions, but have implicit biases against certain groups, which guide their decisions or actions. It is possible (indeed common) for a subject to have no explicit sexist or racist beliefs and yet still have implicit biases towards these groups. According to the most common reading of the IAT results, subjects have implicit biases, even though such biases are not always reflected in explicit judgements. EPISTEMIC INNOCENCE5THANK YOU TO LISA BORTOLOTTI, WITH WHOM I DEVELOPED THE NOTION OF EPISTEMIC INNOCENCE AS I AM UNDERSTANDING IT HERE (SEE BORTOLOTTI, 2015, FOR AN EXPLANATION OF WHY WE USE THE TERM ‘INNOCENCE’ IN THIS WAY).5 -------------------------------------------------------------------------------- I am interested in the epistemic status of confabulatory explanations of decisions or actions guided by implicit bias, and whether they have the potential for epistemic innocence. I will understand the notion of epistemic innocence in the following way: an epistemically faulty cognition is epistemically innocent if, at a given time, it endows some significant epistemic benefit (Epistemic Benefit) onto the subject, which could not be otherwise had, because alternative, less epistemically faulty cognitions are in some sense unavailable to her at that time (No Alternatives). I will further elucidate the notion of epistemic innocence in the rest of this section. Which benefits? ~~~~~~~~~~~~~~~ A cognition being epistemically innocent does not imply that that cognition is free from epistemic faults, perhaps very few cognitions are innocent in that sense. The claim is only that such cognitions can confer epistemic benefits which are otherwise unavailable, in some contexts. If confabulatory explanations of the kind I have in mind here have epistemic benefits, might we be better off to call them epistemically good rather than epistemically innocent? No, since saying this would be to ignore or deny the obvious epistemic faults of confabulation. My aim here is to highlight the potential epistemic benefits of some confabulatory explanations. One way of doing this without polarising the debate is to claim some sort of inbetween status for confabulations. An ideal agent would not need to confabulate explanations, but human agents have significant limitations that lead to confabulatory explanations. Now, this is not always or not entirely a bad thing. Confabulatory explanations can sometimes play a positive epistemic role. Which unavailability? ~~~~~~~~~~~~~~~~~~~~~ What I mean by unavailable in the No Alternatives condition on epistemic innocence will differ depending on the kind of cognition under investigation. Here is a first pass at three ways we might think about the notion of unavailability: alternative cognitions might be strictly unavailable, motivationally unavailable, or explanatorily unavailable. An alternative explanation is strictly unavailable if it is based on information that is opaque to introspection, or otherwise irretrievable. We might understand this sense of unavailability in terms of alternative cognitions being inaccessible to the subject. For example, consider the case of a subject with dementia who suffers from severe memory impairment. She claims to remember going to the beach with her parents that morning, but the trip she recalls occurred when she was a teenager, sixty years ago. A memory of the trip which included the correct time at which it took place, or information which would suggest to the subject that she had made a mistake with respect to the time of the trip, is inaccessible to her, and so strictly unavailable, due to the severe memory impairment she suffers as a result of her dementia. An alternative explanation is motivationally unavailable if it is inhibited or not accessed due to motivational factors. It is generally agreed upon in the literature that self-deception includes a motivational element, which makes it a good case to refer to in explicating the notion of motivational unavailability.6 This motivational element might make less epistemically faulty cognitions unavailable to the subject. Take the case of the cuckholded husband who self-deceptively believes that his wife is faithful. Evidence that she is unfaithful may be available to him (insofar as it is perceptually available—he sees that his wife returns home late, dishevelled, and uninterested in him), but an alternative cognition, such as the belief that his wife is having an affair, is motivationally unavailable, due to the husband’s very strong motivation for it to be the case that his wife is faithful (wishful self- deception) or for it to be the case that he believes that his wife is faithful (willful self-deception) (see Van Leeuwen, 2007: 331–2, for further elucidation of the distinction between wishful and wilful self-deception). An alternative explanation might be unavailable in a weaker sense, such that it is strictly speaking available for consideration by the subject, but it is not regarded as a genuine contender. It is explanatorily unavailable to the subject insofar as it is dismissed due to its apparent implausibility. For example, a subject may come to have a cognition which explains some experience she has. If alternative cognitions which might also be candidate explanations for her experiences are such that they strike her as seriously implausible or explanatorily inadequate, these alternative cognitions are explanatorily unavailable. Suppose there are bite marks in my cheese, I hear scratching at night, and my cat is agitated. I come to the conclusion that I have mice in my house. An alternative explanation might be that a cheese-eating, cat-irritating, noisy fairy is infiltrating my home at night. This explanation is not available to me in the sense I have in mind here due to the incredulity I would feel towards it. It is either not considered by me, or it is such that I rule it out on grounds of implausibility or poor explanatory power, relative to the preferred and adopted cognition.","In order for confabulatory explanations of decisions or actions guided by implicit bias to be epistemically innocent in the way described above, it would need to be the case that they are epistemically beneficial as per the Epistemic Benefit condition, and that less epistemically faulty alternative explanations delivering the same epistemic benefit are in some sense unavailable to the subject at the time, as per the No Alternatives condition. In this section, I will argue that at least in some cases, confabulatory explanations of this sort meet both conditions on epistemic innocence. Epistemic benefit and confabulatory explanations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As in my earlier characterisation, confabulatory explanations are epistemically poor insofar as they are false, or at the very least ill-grounded. Here I want to focus on whether confabulatory explanations have any epistemic benefits notwithstanding their epistemic faults. I claim that there are two potential epistemic benefits: in the context of epistemically imperfect agents, confabulatory explanations of decisions or actions guided by implicit bias may maximise the acquisition of true beliefs in the long run by filling an explanatory gap, and they may help the epistemic agent maintain consistency among her cognitions. Sometimes confabulations will be epistemically innocent and overall good, for example, viewed from a consequentialist framework this might be when the benefits outweigh the costs. In other cases, they will be epistemically innocent, and epistemically bad, from the same framework, as when such benefits are counteracted by the epistemic costs of confabulation. Importantly, epistemic innocence does not stand or fall with epistemic goodness. Maximising the future acquisition of true beliefs by filling explanatory gaps Here I will argue that there are epistemic benefits in providing a confabulatory explanation to the question ‘Why did you do this?’ This is because filling an explanatory gap allows people to consider their reasons for their actions which may otherwise remain unexplored and unchallenged. Earlier I suggested that one of the characteristics of confabulatory explanations was their filling a gap, where this was to be understood not just in terms of memory gaps, but rather, more broadly. The performance of the function of gap filling may be pragmatically beneficial in different ways depending on the case. In the non-clinical case, explanations of decisions or actions guided by implicit bias may fill a gap which gives the illusion of competence, and prevents people from having to claim that they do not know or cannot remember why they formed an attitude, made a choice, or performed an action (see Dalla Barba, 1993, cited in Van Damme & d’Ydewalle, 2010: 221). McKay and Kinsbourne identify two types of motivational account which apply to confabulations. According to the first, subjects who confabulate offer the explanations they do ‘in order to conceal embarrassing gaps in their memories’ (McKay & Kinsbourne, 2010: 291). McKay and Kinsbourne reject this account as one having ‘very little evidence’ in support of it, and I have already said that confabulations are doing more than filling gaps in memory. We might understand this kind of account more broadly though, as one which claims that people who confabulate do so in order to conceal gaps more generally (for example, explanatory gaps regarding why they decided or acted in a certain way). If this kind of account were correct, at least for confabulatory explanations in the non- clinical population, such explanations could be seen as conferring a pragmatic benefit with respect to the prevention of embarrassment which may have indirect positive epistemic consequences (perhaps my not suffering the discomfort of embarrassment might mean I am more willing and able to investigate my environment and participate in the exchange of information). According to the second type of motivational account, confabulations are ‘purposive constructions that function to embellish the situation of the patient’, meaning that confabulatory explanations are compensatory in virtue of their content (not merely their very existence) (McKay & Kinsbourne, 2010: 291). If the second type of account were right, there would be additional pragmatic benefits, due to the confabulation enhancing the concept of the self or the situation of the person, and potentially enhancing self-confidence and wellbeing. This matches with my description of the cases Roger and Sylvia. But is there room for epistemic, as well as pragmatic, benefits? My claim here is that by having an explanation and endorsing a position, one can receive feedback and the problematic things that are believed (I did not invite Katie to interview because her CV was of poorer quality or I crossed the road because the black man was behaving in a threatening way) become available to introspective reflection, and open to feedback and revision. Jeanette Kennett and Cordelia Fine argue that when we become aware of our biases and are committed to not being prejudiced, then we can counteract or compensate for our biases and achieve better consistency between explicit attitudes and behaviour: research demonstrates that when people become aware that they have a tendency to make certain types of judgments in a biased way (for example, due to the activation of negative stereotypes about a racial group) then, if they are motivated to be unprejudiced, they will effortfully over-ride their intuitively-based judgments, so long as they have the cognitive resources to do so. If becoming aware of our biases and being motivated to not be prejudiced can lead to trying to overcome the judgements we make on the basis of bias, confabulatory explanations may be indirectly epistemically beneficial, insofar as they might contribute to our becoming aware of our biases. We might think that a much better way to become aware of our biases would be to fail to offer an explanation at all. In Roger’s case for example, if he claimed not to know why he did not invite Katie to interview, this might be a better way of encouraging him to reflect on what was guiding that decision. However, I will argue later (Section 4.2) that this dumbfounding response is unavailable to Roger, and so the confabulatory explanation here confers a benefit which is not otherwise obtainable. Related to this, research has shown that individuals who avow low prejudice but recognise that they are prone to behaviour which is inconsistent with this—so-called ‘low- prejudice discrepancy-prone’ subjects—are more likely to feel guilt when expressing biased behaviour (as measured by the IAT for race), and more likely to interpret such behaviour as being related to race factors (Monteith, Voils, & Ashburn-Nardo, 2001: 411). Speculatively, if subjects become aware that they are discrepancy-prone, the guilt felt when expressing biased behaviour, and their interpreting it in this way, might be instrumental in seeking to achieve better consistency between their explicit attitudes and their behaviour. There are reasons to think such awareness is possible. For example, in a study looking at the relationship between self-report scores on the Modern Racism Scale (MRS) and implicit racial attitudes, Jason Nier found that: when participants believed their ‘true attitudes’ were being accurately assessed, there was a significant relationship between an implicit measure of racial attitudes (the IAT) and an explicit measure of racial attitudes (the MRS). When participants did not believe that their self-reported explicit attitudes could be accurately corroborated with an implicit measure, there was no association between implicit and explicit attitudes. One way to become aware of our biases is to observe our own behaviour (see Holroyd, 2015) compare them with our explanation of our actions, and identify discrepancies. Giving reasons (even confabulatory ones) contributes to becoming aware of one’s own attitudes (and conflicts in attitudes), priorities and values and offers people the opportunity to construct a coherent narrative of themselves where behaviour aligns with explicit commitments as opposed to implicit biases (Bortolotti, 2009). By reporting a false explanation, we make some (confabulatory) explanation for our decisions or actions available, and we may initiate a process by which we acquire a new true belief, whereas if we did not have any explanation to offer we may have been stuck with an implicit bias that is not detected and does not get challenged. The flip side to this benefit is that offering a confabulatory explanation might serve to mask the discriminatory nature of the behaviour, which might otherwise be observed if the explanation were an accurate one or one indicating dumbfounding. I will argue later though that alternative explanations such as these are unavailable to the person confabulating, given certain other conditions (Section 4.2). I do not deny the epistemic costs of the confabulatory explanation—that it might mask the discriminatory nature of the behaviour being one such cost—the idea is only that there may also be some epistemic benefits, including the confabulatory explanation initiating a process by which the subject acquires a new true belief. So confabulatory explanations are epistemically beneficial insofar as in offering them as explanations they become open to attack. This might then start a thinking process, where I might end up with a less epistemically faulty explanation. So the confabulation is epistemically good insofar as it acts as an enabling condition for further reflection. Discussing and justifying one’s decisions or actions, specifically, reflecting on the potential role played by stereotypes can reduce the effects of such stereotypes (Saul, 2012a: 259). This does not require the subject to have any kind of conscious control over her implicit biases, the idea is rather that when she engages in this kind of discussion, it ‘may help to flag up at least some cases in which there is no defensible reason for a judgment. And some of these may be cases of bias’ (Saul, 2012a: 259). This can be epistemically beneficial insofar as if one discovers that the decision or action is not carried out for a defensible reason this might initiate a process whereby the subject comes closer to a true explanation of their decision or action. The subject may come to have impact awareness of their bias, and seek to reduce its influence which might have positive epistemic consequences. One worry is that the benefit I have outlined is one had by a lot of things, and that now my just saying anything false is epistemically beneficial insofar as it might prompt someone to suggest I think again, which might in turn bring about my doing so, and then I might come to have a more epistemically worthy belief. There are two things to say in response to this kind of worry. Firstly, this is an epistemic benefit which is had in a context in which alternatives are not available, whereas not all cases of avowed false belief are going to be cases where less epistemically faulty cognitions are unavailable. So though avowing false beliefs might be epistemically beneficial, the context in which this benefit comes is important for judgements of epistemic innocence. Secondly, it is not on this benefit alone that I suggest that confabulatory explanations of decisions or actions guided by implicit bias are epistemically beneficial. Maintaining consistency Here I will argue that there are epistemic benefits in providing a confabulatory explanation in response to a question such as ‘Why did you do this?’ because it allows one to maintain a coherent self-concept. The epistemic benefit here comes from the explanation’s having a particular content, and so this maintaining of consistency is only had by explanations with certain contents. Not having unexplained gaps with respect to one’s decisions or actions may be instrumental to maintaining a coherent set of beliefs about oneself. In the case of implicit bias, Roger has to make consistent his belief that he is egalitarian, and his belief that he did not invite Katie to interview. Similarly, Sylvia has to make consistent her belief that she is egalitarian, and her belief that she crossed the road when she saw a black man. Given our subjects’ lack of impact awareness that an implicit bias is playing a role here, their confabulatory explanations maintain consistency between their beliefs about the values they are committed to, and their beliefs that they made some decision or performed some action.7 In addition to this, the explanation may fill a gap in the sense that it provides a reason for the action that is more desirable than the (true) alternative. In the confabulatory explanations of actions driven by implicit bias, Roger’s decision not to invite Katie to interview, or Sylvia’s action of crossing the road to avoid a black man, is explained by appeal to something external (the inferior quality of Katie’s CV, and the threatening behaviour of the black man, respectively). The subjects make an attribution error but in doing so they maintain their positive self-concept. In the true explanations, the action would be caused by something internal, a bias against women, and a bias against black people, respectively. The limitations of this potential epistemic benefit lie in the fact that the confabulation achieves consistency at the expense of truth. This may lead to a revision of the self-image such that in confabulation, ‘coherence trumps correspondence’ (Bortolotti & Cox, 2009: 962).8 We might wonder then why consistency here should be considered as epistemically beneficial, might it make the acquisition of true beliefs less likely in the long run? If we conceive of consistency being a benefit only when it tends towards truth or when it adds to the overall coherence of belief, perhaps the maintaining of consistency had here is not epistemically beneficial. For those who feel the force of this worry I say this: at the very least the apparent maintaining of consistency might be indirectly epistemically beneficial. Not experiencing an inconsistency might have indirect epistemic consequences. Just like the prevention of embarrassment (see Section 4.1.1), perhaps by not suffering the discomfort which an inconsistent set of cognitions might bring, I might be more willing and able to investigate my environment and participate in the exchange of information. No alternatives and confabulatory explanations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ I have argued that confabulatory explanations can have epistemic benefits. Here I will argue that such benefits are only attainable via the confabulatory explanation, because less epistemically faulty cognitions which deliver the same epistemic benefits are in some sense unavailable to the subject who confabulates. We have seen that confabulatory explanations of decisions or actions guided by implicit bias can be sincerely and confidently avowed, and offered with no intention to deceive on the part of the subject confabulating. Given this, as well as the third-person implausibility of the explanations offered, it looks like subjects who confabulate do not recognise the poor epistemic status of the explanations they give. This claim is supported further by accounts of confabulation which postulate a deficit in the confabulator’s ability to evaluate hypotheses. On Hirstein’s view, two errors are involved in confabulation. The first is the creation of a false response, and the second is a failure by the subject to ‘check, examine it and recognize its falsity’ (Hirstein, 2005: 15). The thought is that a non- confabulator might have the first error, the generation of a false response, but they would recognise the falsity or absurdity of it. So the person who confabulates ‘should either not have created the false response or, having created it, should have censored or corrected it’ (Hirstein, 2005: 15). The mechanisms by which this failure is achieved is suggested by Hirstein to result from the fact that the ability to construct plausible responses and the ability to verify such responses are separate in the brain, and confabulators retain the first ability, but the second has been compromised by brain damage (Hirstein, 2005: 16–7). My focus has been on confabulation as it occurs in the non- clinical population, specifically, in explanations of actions guided by implicit bias. Hirstein states that young children have been reported to confabulate when reporting on their (false) memories (see for example Ackil & Zaragoza, 1998), and subjects of hypnosis can confabulate (Hirstein, 2005: 16), as well as ordinary people in experimental settings (for example, in choice blindness experiments, see Hall & Johansson, 2008). In these cases, people do not have the kind of brain damage which would make the particulars of Hirstein’s account plausible in such a way that it has general application across all cases of confabulation. The etiology of confabulatory explanations may not be shared across the board; we cannot give an account which cites brain damage when we are explaining why non-clinical subjects offer confabulatory explanations in certain contexts. What we can say in general though is that—for whatever reason—when people confabulate, they do not experience doubt, which results in a failure to check on the plausibility of hypotheses when they are occur. Perhaps this failure is realised in different ways across cases. In a case of explanations of decisions or actions guided by implicit bias, for example, the failure might be caused by some motivational component of the agent’s cognitive economy. The motivational component of explanations of actions guided by implicit bias is straightforward to spell out. In terms of type-a motivation (that is, motivation to offer an explanation as opposed to no explanation), subjects are asked to justify their decision or action but because their decision or action was guided by an implicit bias, the true causal explanation of their decision or action is not strictly available to them (unless they are well-read in psychology and they suspect that their decision or action may be influenced at least to some extent by biases they are not aware of). As we have seen, this lack of impact awareness with respect to implicit attitudes is a distinguishing feature of them (Gawronski et al., 2006: 486). Given that failing to give reasons for one’s decision or action may be costly, such subjects have a reason to provide an explanation. In terms of type-b motivation (that is, motivation to offer an explanation with a particular content), our subjects are motivated to think of themselves as persons of egalitarian persuasion, who do not have gender or racial prejudices. Thus, it is easy to see that the explanations they offer for their decisions and actions are influenced by motivational factors. We might think that motivational factors are playing a role in the content of the explanation offered by running the appropriate counterfactual: if Roger were explicitly sexist or Sylvia were explicitly racist, and the subjects were not ashamed by this or did not have the desire to conceal such prejudices, and they decided or acted in the above ways and were asked for explanations of their decisions or actions, we might expect Roger to claim that he did not invite Katie to interview because women make for bad colleagues, and Sylvia to claim that she crossed the road to avoid the black man because black men are dangerous. If a subject’s decision or action is guided by an implicit bias pertaining to some group, there is a sense in which explanations citing that bias are unavailable to her, given certain other conditions. Such other conditions may include the subject considering herself a person who does not hold sexist or racist prejudices, or at the very least having the desire to be perceived as such by her peers: The belief that one’s actions are implicitly biased, and other implied beliefs about one’s role in sustaining patterns of discrimination, are clearly beliefs that, for a range of reasons, agents might be motivated not to confront. Conversely, the belief that one’s actions are consistent with one’s moral ideals (of non-discrimination, of being evidence sensitive and unbiased) is one that agents are motivated to maintain. This kind of story is supported empirically. In their review article, Gawronski and colleagues found evidence that correlations between self-reported attitudes and indirect attitudes are ‘often higher when the impact of motivational factors is controlled’ (Gawronski et al., 2006: 489). Confabulatory explanations of actions driven by implicit bias are sometimes such that less epistemically faulty alternative cognitions are motivationally unavailable to the subject insofar as they cite implicit biases against certain groups, biases which she is motivated not to recognise in herself or have others recognise in her.9 Thinking in terms of one’s attitudes being regulated for truth, we might say that the truth regulation present in the cases when a motivational component is in play is sufficiently different from the truth regulation present in cases where motivational factors are not playing a role, such that in the former case, the content of the explanation offered is distorted away from a true explanation. Once the subject is motivated to think of herself as someone who does not have sexist or racist attitudes, or motivated for other people to not think of her in that way, the regulation for truth in her belief formation with respect to her explanation for her decision or action is weakened. We might think then, that without the motivational component, the explanation offered for decisions and actions guided by implicit biases, would not be confabulatory, since studies have shown that controlling for the motivational components results in a higher correlation of explicit and implicit attitudes (Gawronski et al., 2006: 489). What cases of confabulation share is that something is going on such that where doubt should be cast on the confabulatory story, it is not. My claim is that whatever this component is—whether it has a neurological basis (as in many clinical cases), or a motivational one—it indicates that other, less epistemically faulty cognitions are unavailable to the subject. When an explanation is presented to the subject upon which doubt is not cast, alternative cognitions are unavailable insofar as the confidence which comes with the confabulation, and the motivational factors which are driving both the presence and content of it, close off alternatives such that they are not sought or entertained.","In this paper I characterised confabulation in terms of five features which are common to them: their being false, their being offered in response to a question, their involving a motivational component, their filling a gap, and their being produced and avowed without any intention to deceive. After outlining two possible cases of confabulatory explanations of decisions or actions guided by implicit bias, I introduced the notion of epistemic innocence. This notion was intended to capture those cognitions which deliver some epistemic benefit which could not have been otherwise had in virtue of alternative cognitions being in some sense unavailable. I then argued that confabulatory explanations of this sort have the potential for epistemic innocence. So what should we conclude from these considerations about the epistemic status of confabulatory explanations in the non- clinical population? What might my analysis mean for the epistemic evaluation of non- clinical confabulation, specifically confabulatory explanations of decisions and actions guided by implicit bias? As I noted earlier, epistemic innocence does not track epistemic goodness. The benefits I identified though should not be neglected. These were filling an explanatory gap which cannot be otherwise filled, which may lead to the acquisition and retention of true beliefs or knowledge, as well as helping to maintain consistency between a subject’s other beliefs (which might be directly or indirectly epistemically beneficial). In some cases, less epistemically faulty cognitions with the same epistemic benefit are in some sense unavailable to the subject, and so the confabulatory explanation delivers some epistemic benefit which is otherwise unobtainable. I conclude then that at least in some cases, confabulatory explanations of decisions or actions guided by implicit bias can be epistemically innocent, and that epistemic evaluation of confabulatory explanations ought to take into account the context in which the cognition occurs. If we do this we are able to provide a richer account of the epistemic status of confabulatory explanations, and we can resist the view that pragmatic benefits come at the expense of epistemic ones. If we focus on the context in which confabulatory explanations occur, and their potential epistemic benefits, we can give a richer epistemic evaluation of them."],["Habits may develop when meaningful action patterns are frequently repeated in a stable environment. We measured the differing tendencies of people to form habits in a population sample of n = 533 using the Creature of Habit Scale (COHS). We confirmed the high reliability of the two latent factors measured by the COHS, automaticity and routines. Whilst automatic behaviours are triggered by context and do not serve a particular purpose or goal, routines often have purpose, and because they have been performed so often in a given context, they become automatic only after their action sequence has been activated. We found that both types of habitual behaviours are influenced by the frequency of their occurrence and they are differentially influenced by personality traits. Compulsive personality is associated with an increase in both aspects of habitual tendency, whereas impulsivity is linked with increased automaticity, but reduced routine behaviours. Our findings provide further evidence that the COHS is a useful tool for understanding habitual tendencies in the general population and may inform the development of therapeutic strategies that capitalise on functional habits and help to treat dysfunctional ones. --------------------------------------------------------------------------------","Habits are repetitive, meaningful actions in our daily lives, which often go unnoticed due to their automatic nature (Robbins & Costa, 2017). The scientific interest in habits has increased in recent years, not least because habits can make behaviour either highly efficient or severely dysfunctional and distressing. Understanding the factors that influence habit formation for the better and for the worse may thus have implications both for enhancing performance when habits are to our benefit, and developing treatments for habits that have become maladaptive. People also differ in their readiness to form and engage in habits in their daily lives, which we refer to as habitual tendencies. There is widespread agreement that habits form over time when behaviour is repeated regularly in the same context. Meanwhile, control over this behaviour gradually shifts from being guided by intentions to being automatically triggered by cues in the environment (see Wood & Runger, 2016). An example would be a person's tendency to take their shoes off automatically by the front door when returning from work. Importantly, habits are not restricted to single actions but may also involve sequences of actions, as exemplified by a night time routine to always prepare one's clothes for the next day before going to bed. Although this deferral of control to environmental stimuli can make habits highly functional by providing structure, reducing uncertainty and freeing up cognitive resources, it also makes behaviour less flexible, since much more effort is required to break or adjust a habit (Wood & Runger, 2016). In people with problems of regulatory control, habits may run the risk of spiralling out of control. A prime example is obsessive-compulsive disorder (OCD), where rather mundane habitual behaviour patterns such as washing one's hands or locking the front door when leaving the house become problematic for the afflicted individual. Patients with OCD find themselves unable to stop performing certain routines, even when they become dysfunctional and negatively affect their lives (American Psychiatric Association, 2013). Another example is drug addiction, where individuals lose control over the initiation, amount and duration of their habitual drug use, and pursue their drug-taking routines even in the face of extremely adverse consequences (American Psychiatric Association, 2013). Experimental research in both animals and humans has been trying to elucidate the mechanisms underlying habit formation, and to identify factors that may precipitate the development of habits. External factors such as exposure to stimulant drugs (Corbit, Chieng, & Balleine, 2014; Ersche et al., 2016; Gourley, Olevska, Gordon, & Taylor, 2013; Nelson & Killcross, 2006), excessive training (Colwill & Triola, 2002; Holland, 2004; Thrailkill, Trask, Vidal, Alcala, & Bouton, 2018) and stress (Dias-Ferreira et al., 2009; Schwabe & Wolf, 2009) have all been shown to facilitate the formation of habits. Thus, behaviour that is regularly performed in a state of either acute or chronic stress is more prone to become habitual and controlled by environmental stimuli (Dias-Ferreira et al., 2009). For example, individuals with a history of stressful life events who also use drugs recreationally, are highly likely to develop a drug-taking habit when they use drugs regularly in the same context (e.g. always in a club on a Friday night). Even if these individuals do not have the intention to use drugs, going to a club on a Friday night increases the likelihood that they will be using drugs that night. As habitual behaviour is triggered by the context, it is the environment that prompts their drug use, overriding their intentions. Habitual drug use does not, however, equate to addiction, but internal factors such as impulsive personality traits have been shown to precipitate – at least in animal models of addiction – the transition of initially adaptive habits into maladaptive ones (Belin, Mar, Dalley, Robbins, & Everitt, 2008). Impulsive individuals are thus at risk of their drug-taking habits spiralling out of control and becoming compulsive. This means that they may continue using drugs even if this poses an acute threat to their health, professional or social life, such as by risking myocardial infarction, job loss, or relationship breakdown. The exact mechanism underlying the transition from functional to dysfunctional habits is, however, still unclear, but as there is a great need for effective treatments for individuals who have lost control over their habits, as well as for strategies to promote habit formation in individuals affected by cognitive decline or dementia who struggle to manage their daily lives, the interest in understanding habitual behaviours will continue to increase. Healthy individuals in the general population who describe themselves as ‘creatures of habit’ are therefore of particular interest for research as they show an increased propensity to develop habits without them becoming pathological. We recently developed the Creature of Habit Scale (COHS, Ersche, Lim, Ward, Robbins, & Stochl, 2017) to measure individual variation in habitual tendencies in the general population. We confirmed modulatory effects of stimulant drug use and life adversity on habitual tendencies, as assessed by the COHS. We also identified positive relationships between compulsivity and habitual tendencies, but at the time we did not examine relationships with impulsivity (Ersche et al., 2017). Impulsivity and compulsivity are two personality traits associated with a lack of control over behaviour, which might explain why in some people automatic habits risk spiralling out of control. Whilst impulsivity reflects a failure to inhibit the initiation of behaviour, compulsivity characterizes a failure to stop an ongoing behaviour that is becoming inappropriate to the situation (see Robbins, Gillan, Smith, de Wit, & Ersche, 2012). Although both constructs reflect distinctly different deficiencies in the regulatory process, they may occur together. We also did not investigate at the time the influence of participants' prior experience with each situation referred to in the questionnaire items. This is a fair concern given that habits develop gradually through associative learning and repetition (Wood & Neal, 2007), suggesting that the more often action patterns are performed in a given context, the greater the likelihood of the context triggering the actions automatically irrespective of the goal (Verplanken & Orbell, 2003). In the aforementioned example of drug-taking, it would be important to know whether it really matters how often people take drugs before their drug use becomes habitual, or whether predisposing personality factors that precipitate the development of habits are more important. In experimental settings, overtraining of an instrumental action is commonly used to induce stimulus-response habits (Adams & Dickinson, 1981; de Wit & Dickinson, 2009; Tricomi, Balleine, & O'Doherty, 2009). However, whilst in animal models the extent of training seems to be directly related to the predominance of the stimulus-response habit (Dickinson, 1985), such a relationship does not seem to hold for humans, according to the findings of experimental work (for review de Wit et al., 2018). It is therefore conceivable that the strength of human habits is not proportional to the extent of prior practice. The aim of the present study was therefore to assess the extent to which the self-reported frequencies with which habitual actions have been performed accounts for habitual tendencies, as assessed by the COHS. We also aimed to investigate the effects of impulsivity and compulsivity on habitual tendencies in daily life. As the COHS is a relatively new measure, we used the opportunity to re-evaluate its psychometric properties in this new sample.","We recruited study participants through the online platform Amazon MTurk to examine variations in regular behaviours, since this platform has been regarded as suitable for obtaining data from the general population (Mortensen & Hughes, 2018). Participation requirements were a minimum age of 18 years and current residency within the United States of America. There were no restrictions with respect to gender, ethnicity or employment status. We also collected background information, including ethnicity, native language, education level, and employment status. A total of 565 participants completed the study, but data of 32 participants (6%) had to be excluded post hoc due to invalid responses or inattentive responding [assessed through recommendations by Meade & Craig, 2012]. Excluded individuals did not differ from the remaining sample on any demographic variable. The final sample included 533 participants with a mean age of 36.8 years [±10.8 standard deviation (SD), age range 18–71 years]. The sample was almost evenly split between male (50.8%) and female (49.2%) participants, of whom 81% identified as Caucasian, 9% as African American, 4% as Asian, 4% as Hispanic, and 2% as Multiracial. The overwhelming majority of participants were native English speakers (99%), who were at the time of the study in full-time employment (70%) [15% in part-time employment, 13% not in paid work, and 2% studying].","The study was approved by the School of Biological Sciences Research Ethics Committee (PRE.2015.124; PI: KD Ersche). Participants received $3.00 for the completion of the study, which included the assessment of functional habits using the COHS questionnaire (Ersche et al., 2017), which includes 27 statements to which participants indicate their level of agreement on a 5-point Likert scale, ranging from strongly disagree (1) to strongly agree (5). Once participants completed the COHS, they were asked to indicate how often they engage in the behaviour described by each item by selecting one of the following response options: never, once a month, twice a month, three times a month, once a week, twice a week, three times a week, four times a week, five times a week, six times a week, once a day, twice a day, three times a day, four times a day, five times a day, or more than five times a day. These responses were then coded respectively from 0 (never) to 15 (five times a day). This fine-grained scale was deliberately selected to capture the wide spread of individual lifestyles as it had been previously used to quantify habitual behaviours (Verplanken & Orbell, 2003). In order to determine participants' levels of trait impulsivity, we administered the Barratt Impulsiveness Scale (BIS-11, Patton, Stanford, & Barratt, 1995). The BIS-11 is a widely-used 30-item questionnaire that measures impulsive personality traits in three dimensions: attention (inattention and cognitive instability), motor behaviour (spontaneous actions), and non-planning (lack of forethought). As a personality trait, impulsivity covers the spectrum from normal to maladaptive behaviour, and it can be assessed using the same tool in both healthy people and patients. The following cut-off scores have been suggested to differentiate variation in trait-impulsivity: BIS-11 total scores between 52 and 71 are indicative of the normal range of impulsivity (n = 296, 56%), scores below 52 reflect individuals who are extremely over-controlled (n = 176, 33%), and scores above 72 signify highly impulsive individuals (n = 61, 11%) (see Stanford et al., 2009). For the assessment of compulsive tendencies, we administered the Obsessive-Compulsive Inventory–Revised (OCI-R, Foa et al., 2002), which requires participants to rate 18 common obsessive-compulsive symptoms in terms of the degree to which they have been bothered or distressed by them in the past month on a 5-point scale, ranging from not at all (0) to extremely (4). Whilst subclinical levels of obsessive-compulsive symptoms are common in the general population (Sher, Martin, Raskin, & Perrigo, 1991), an OCI-R score of 21 or more suggests obsessive-compulsive symptoms of clinical severity (Foa et al., 2002). In the present sample, the vast majority of participants (n = 425, 80%) scored within the normal range, whilst the scores of 20% of participants (n = 108) pointed toward particularly high levels of compulsivity. Factor structure of the Creature of Habit Scale (COHS) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We used confirmatory factor analysis to verify that the COHS structure consists of two latent factors (automaticity and routine) as reported in our previous study (Ersche et al., 2017). Means and variances adjusted Weighted Least Squares (WLSMV) were used as estimators in all presented models. Fit was evaluated using traditional indices such as Root Mean Square Error of Approximation (RMSEA), Comparative Fit Index (CFI) and Tucker- Lewis Index (TLI). Reliability of subscales was assessed by McDonald's omega (McDonald, 1999), and for reasons of convention, we also computed Cronbach's alpha (Cronbach, 1951). We extended our factor analytic model by regressing each COHS item onto the corresponding frequency item. The conceptual path diagram of this model is shown in Fig. 1. The rationale for this approach was two-fold: firstly, it allowed us to quantify the influence of frequency on COHS items (by using difference in R-squares of COHS items between frequency-adjusted and non-adjusted factor analytic models). Secondly, we evaluated whether the constructs of automaticity and routine hold whilst taking into account the frequency with which participants previously engaged in these behaviours. Relationship between impulsivity and compulsivity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We further extended the CFA model by including self-reported impulsivity (BIS-11) and compulsivity (OCI-R) levels as predictors of the latent factors for automaticity and routine. We used participants' individual responses on the COHS without any adjustments for frequency. Our aim was to investigate how levels of impulsivity and compulsivity relate to participants' responses with respect to automaticity and routine. We estimated the models using the statistical software Mplus, version 8 (Muthén & Muthén, 2018) and standardized the results. Solely for descriptive purposes, we divided the sample into three subgroups reflecting the three categories of impulsivity, as measured by the BIS-11. For illustrative purposes only, we compared these subgroups with respect to compulsivity (OCI-R) and habitual tendencies (COHS automaticity and routine) using the Kruskal-Wallis and the Jonckheere's trend tests to identify differences and trends respectively (Fig. 2). Factor structure of the Creature of Habit Scale (COHS) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The two-factor structure of the COHS fitted the data well (Chi-square(df) = 1169(323), p ≤0.001, RMSEA = 0.070, 90%CI for RMSEA = (0.066, 0.075), CFI = 0.977, TLI = 0.975), supporting the notion of automaticity and routine being two subscales of COHS. Both factors were also moderately correlated with each other (r = 0.296, p < 0.001). The loading of both factors was high, indicating high factorial validity of items. Likewise, the reliability of the coefficients for both factors was high as well, i.e. COHS routine (Cronbach's alpha: 0.90; McDonald's omega: 0.94) and COHS automaticity (Cronbach's alpha: 0.87; McDonald's omega: 0.92), providing support for satisfactory measurement precision of both subscales (Fig. 3). The influence of frequency on the COHS scores ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The inclusion of the frequency ratings in the model also fitted the data well (RMSEA = 0.032, CFI = 0.93, TLI = 0.93). In this model, factor loadings were adjusted for the influence of frequency ratings. Model estimates are shown in Table 1. In brief, the loadings were smaller compared to the model without frequency, but remained high enough to support the existence of two underlying factors. This suggests that habits, as conceptualised in the COHS, cannot be fully explained by how frequently the corresponding activities have previously been carried out. Standardized regression coefficients (adjusted for relationships between items and the corresponding factors) between the COHS items and the corresponding frequency ratings ranged from 0.09 to 0.65. Except item 13 (I rely on what is tried and tested rather than exploring something new.), all COHS items were statistically significant, suggesting that participants' responses can be partially, but not fully, explained by the self-reported frequency with which the behaviour in question has been repeated. The average difference between R-squares for frequency- adjusted and non-adjusted factor models was 0.15 for automaticity items and 0.11 for routine items, suggesting that frequency explains slightly more variance for automaticity than for routine. The effects of deficient regulatory control on habitual tendencies ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 4 depicts the model we used to evaluate the effects of insufficient regulatory control, as reflected by self-reported measures of impulsivity and compulsivity. All factor loadings, regressions and correlations included in this model were statistically significant (all p-values <0.001). Automaticity was positively related to both impulsivity and compulsivity, suggesting that participants with higher levels of impulsivity and/or compulsivity are more likely to report increased tendencies for automaticity. It is noteworthy that the effect of impulsivity (β = 0.345) on automaticity was larger than the effect of compulsivity (β = 0.155). For routine behaviours, we observed relationships of approximately the same magnitude but of different directions with respect to impulsivity and compulsivity (βimpulsivity = −0.270 versus βcompulsivity = 0.273). This may indicate that more compulsive individuals are more prone to routine behaviours. This effect is, however, attenuated in individuals who are also impulsive, as impulsivity prevents the occurrence of routines.","Habitual responses are part of everyday life, but the propensity to form habits differs substantially across individuals. Here we provide evidence for the validity of the two- factor model underlying the COHS for assessing habitual tendencies in the general population. We further confirm that the COHS is consistent with the theoretical concept of habits, which explains the formation of habits through a process of context-dependent repetition (Robbins & Costa, 2017; Wood & Runger, 2016). Both scales of the COHS were significantly influenced by the frequency of past behaviour but, importantly, they were not fully explained by it. Our findings thus not only concur with prior experimental work in humans (de Wit et al., 2018), they also have important implications for the use of the COHS, making it a potentially useful psychometric instrument for investigating habitual tendencies in the general population. Our data further suggest that habitual tendencies are related to personality traits, specifically those that characterise an individual's disposition in regulating behaviour. Whereas both impulsivity and compulsivity were positively related to automaticity, they had conflicting associations with the tendency to routine. Impulsivity was negatively associated with routine behaviours, perhaps because of impaired behavioural regulation with respect to timing. By contrast, compulsivity, perhaps unsurprisingly, was positively associated with both routine and automaticity. Frequency of past behaviour and habitual tendencies ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There is widespread agreement that habit formation is a process, rather than an event, which implies that the repetition of behaviour is to some degree necessary for it to become habitual. This notion is also reflected in our data suggesting that habits are dependent on repetition but are not fully explained by it. What may seem self-evident has, however, not always been seen in this way. A decade ago, habits were understood as a combination of both frequency (as defined by repetition) and automaticity (as defined by a lack of control and awareness) (Verplanken & Orbell, 2003). Whilst repetition was deemed critical in the early stages of habit formation, automaticity was thought to play a more prominent role in the final stages of habit formation by implicitly ensuring that behaviour will be repeated once the habitual response has been established (Gardner, 2012; Verplanken, 2006). Clearly, repetition of behaviour through practice is needed for habit formation, but well-practised behaviours are not necessarily habits. Our two-step approach of composing the COHS questionnaire, first to select items of sufficiently general nature, and subsequently to collect frequency information for each of them, has allowed us to control statistically for the effects of repetition. As this study shows, both COHS scales clearly survived the corrections for different frequencies between individuals, thereby further supporting the notion that habitual tendencies are more than merely the frequency of their occurrence. Therefore, as recommended by Ajzen (2002), we should not solely rely on behavioural frequencies as the defining criterion for habits, but rather focus on the qualitative characteristics of habits, such as the declining influence of cognitive factors and the increasing control of stimulus cues over behaviour. Personality and proneness to habit As habit formation is characterised by a devolution of control from intentions to the contextual cues, individuals with problems in self-regulation run the risk of automatic habits getting out of control. Our data suggest that trait impulsivity promotes habitual behaviour by facilitating this regulatory imbalance between the goal-directed and the habit system. Impulsivity seems to selectively enhance automatic stimulus-driven, goal-independent actions whilst diminishing the occurrence of routine behaviours. Our findings seem to complement previous findings in regular smokers, whose levels of trait impulsivity predicted how likely they would pick up a cigarette if they were craving for one (Hogarth, 2011). Conscious cravings predicted smoking only in less impulsive smokers, whereas in highly impulsive individuals smoking was more habitual since it was decoupled from the conscious desire for a cigarette. This subtle difference between goal-independence, as reflected by the COHS automaticity scale, and some relatedness of routines to a goal, appears to be critical in defining the influence of trait impulsivity on habit formation. Given that impulsive actions are spontaneous, premature, lack forethought, and have previously been described as an ‘inability to control automatic reactions to stimuli’ (Kopetz, Woerner, & Briskin, 2018), the positive relationship with stimulus-driven behaviours captured by the COHS automaticity scale may thus seem intuitive. The negative influence of impulsivity on routine behaviours, however, might be less obvious. Clearly, routines such as going to church on Sunday morning or brushing your teeth before going to bed are not fully independent from a goal, but are meaningful familiar action patterns, which have been performed many times in the same context so that they become automatic once the action sequence has been activated by the cognitive representation of the goal (Aarts & Dijksterhuis, 2000). Consequently, the difficulty that impulsive individuals have in adjusting the optimal timing for an action (as they initiate actions prematurely) is opposed to adhering to regularity and developing routines. In fact, the implementation of family routines is explicitly recommended to parents of impulsive children as routines provide them with a stable context in which consequences become more predictable, encouraging them to act less impulsively (Lanza & Drabick, 2011). Compulsivity, by contrast, was positively associated with both aspects of habits, routine behaviours and automaticity, albeit to a lesser degree than impulsivity. The positive relationship is not surprising given that routines, at least those routines that have an almost ritualistic nature, may run the risk of developing into compulsive behaviours, as exemplified in obsessive-compulsive disorder (Mell et al., 2005). Likewise, the unconscious, stimulus-driven nature of compulsions has often been described by patients with compulsive disorders (McCusker & Gettings, 1997). It is thus conceivable that compulsive personality traits may enhance habitual tendencies under conditions of insufficient inhibitory control. Habit formation from a neuroscientific perspective ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our findings also agree well with the substantial work in both animals and humans that has been defining the neural circuit architecture underlying the neuroadaptive processes of habit formation (Balleine & Dickinson, 1998; Tricomi et al., 2009; Yin, Knowlton, & Balleine, 2006). There is now broad consensus that habitual responding develops through the progressive engagement of different striatal subsystems along a ventral to dorsal gradient (Burton, Nakamura, & Roesch, 2015). During the early stages of habit formation, behavioural control relies heavily on dopaminergic pathways modulating interactions of the associative striatum and the lateral and medial prefrontal cortices, which encode the value of expected outcomes (Haber & Knutson, 2009). As during this stage behaviour is mainly goal-directed and performed on purpose, actions are selected on the basis of their anticipated consequences. With prolonged practice, however, the regulation of these actions is reorganised as behaviour becomes increasingly automatic and less dependent on dopaminergic neurotransmission in the aforementioned regions (Ashby, Turner, & Horvitz, 2010). This change is hypothetically underpinned by devolution of behavioural control to the sensorimotor striatum (Ashby et al., 2010). During this transition, information about action sequences is clustered together to form units of behavioural repertoires (or ‘chunks’), which can be quickly and more efficiently implemented (Graybiel, 1998). In other words, the sensorimotor system acts on the signals it receives from the prefrontal cortex and does not further evaluate how appropriate they are. Consequently, behaviour is executed irrespective of its potential consequences. Experimental evidence further shows that the extent of training (or repetition) is directly linked to changes in dopamine signalling, which induces the transition of control from the associative system to the sensorimotor system (Choi, Balsam, & Horvitz, 2005). Thus, from a neuroscientific point of view, the repetition of behaviour is not equivalent to habit, but represents an essential ingredient for its development. With respect to the influence of impulsive and compulsive traits on habit formation, there is also growing evidence of impulsivity being associated with deficits in goal-directed control (Gillan, Kosinski, Whelan, Phelps, & Daw, 2016; Hogarth, Chase, & Baess, 2012). Consequently, reduced prefrontal involvement during goal- directed choices may facilitate habitual responding (Deserno et al., 2015). The neural substrates of compulsivity have been associated with dysregulation of control functions implemented by frontostriatal and corticostriatal circuitries (van den Heuvel et al., 2016). Subclinical levels of compulsivity in healthy volunteers seem to be associated with a volume enlargement of the putamen (Kubota et al., 2016), a structure that is significantly enlarged in patients with compulsive disorders (Chamberlain et al., 2008; Ersche et al., 2011, 2012; Pujol et al., 2004), suggesting a highly active sensorimotor system (Kwon et al., 2003) that facilitates habit formation. Strengths, weaknesses and outlook for further research ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Strengths of the study include the relatively large sample size, the re-evaluation of the COHS's psychometric properties (given that it is a relatively new tool), and our aim to address the theoretically important question about the influence of frequency on habit formation. The latter is particularly pressing for assessing habitual tendencies by questionnaire given the discrepancies between animal and human experimental work. Due to practical constraints of this online study, the occurrence of habitual behaviours and their frequencies could only be assessed by self-report and not verified objectively, which should be addressed by future studies. We used a 15-point frequency scale to allow for the variation in lifestyles in our sample. We acknowledge that for some COHS items, frequency data are more difficult to collect than for others. For example, items such as ‘I like routines’, may not spark an obvious answer but our pilot work suggests that participants understood this item in terms of the frequency with which they generally obtain gratification when engaging in their routines, which is consistent with the item's meaning. Our finding that frequency is important but not sufficient for the formation of habits is critical for the assessment of habits by questionnaire more generally, given that habits are dependent on people's lifestyles, which determine how frequently they engage in a behaviour. The COHS items were therefore carefully chosen, and as suggested by the present study, are not compromised by how frequently respondents report having previously engaged in the behaviour in question. Our findings are thus in keeping with both the theoretical concept of habits and the neuroscientific evidence of their development. The COHS has not been designed to measure habit strength in terms of how persistently an individual pursues a particular habit, but rather to assess individual variation in readiness to engage in habitual behaviours in daily life. It should therefore not be mistaken with the Self-Reported Habit Index (Gardner, De Bruijn, & Lally, 2011; Verplanken & Orbell, 2003), which requires individuals to rate their experience with the regular behaviour in question. Further work is, however, warranted to investigate how proneness to habits expresses itself during habit learning in an experimental paradigm. It would be of particular interest to test whether high COHS scores are predictive of enhanced acquisition of stimulus-response habits and the (in)ability to adjust them according to changing circumstances. The present study does not necessarily indicate that the proneness to form functional habits is a risk marker for the transition to dysfunctional habits; however, we would welcome research in patients with dysfunctional habits to investigate whether the COHS might inform preventative or therapeutic strategies for these patients.","This present study examined the inter-relationships between three potentially related traits: impulsivity, compulsivity and the proneness to habits, defined according to two distinct criteria, automaticity (goal-independent) and routine (goal-related), and not dependent on their frequency of occurrence. Whereas both impulsivity and compulsivity are associated with increased automaticity of behaviour, they may thus exert opposite effects on routine habitual behaviour, impulsivity being negatively correlated. These findings may be relevant for identifying vulnerability to certain mental health disorders and for understanding how learning may become aberrant and produce psychopathology."],["Evidence obtained with new experimental paradigms has renewed the debate on the development of theory of mind in general and false belief ascription in particular. Namely, several studies contend to prove that infants already have the capacity to attribute false beliefs. The aim of the current meta-analysis is to review and summarize the empirical evidence about spontaneous-response false belief tasks in infants younger than 2 years old. Fifty-six false belief conditions using the violation-of-expectation, the anticipatory looking and interactive paradigms were included in this meta-analysis, including 1469 infants. The role of several moderators was examined, following Wellman et al.ös meta-analysis (2001). Results show that correct performance on spontaneous-response false belief tasks was about 1.76 times more likely than incorrect performance (β = 0.57, 95% CI 0.33; 0.80, p <.0001). Mediator analyses revealed that (i) year of publication had a significant influence on performance, reducing the average log odds of successful performance (β = -0.11, 95% CI: -0.16; -0.06, p <.0001); and (ii) correct performance was more likely than incorrect performance when the task was conducted in the violation-of-expectation paradigm (β = 0.75, 95% CI: 0.25; 1.26, p =.003). However, heterogeneity was high across the studies and the funnel plot revealed an asymmetric distribution suggesting that studies with small effect sizes were not published. These results cast doubt on the alleged robustness of the phenomenon: its effect size decreases as time passes, it seems to depend on the type of paradigm employed, and the variance across studies is not well understood yet. --------------------------------------------------------------------------------","In our everyday social interactions, we frequently attribute mental states (e.g., beliefs, desires and intentions) to other people in order to predict or explain their behavior. For example, if you know that Ben wants to drink orange juice, and you know that Ben believes that the orange juice is in the fridge, you can predict that he will search for the orange juice in the fridge. This ability to impute mental states to others (and also to oneself) to make predictions about the future behavior of other agents is called theory of mind (Premack & Woodruff, 1978, p. 515). Competence in theory of mind consists of understanding agents as holders of beliefs and desires, which jointly give rise to intentions and goals, carried out in action. False-belief ascription was chosen as diagnostic of theory of mind capacity (Bennett, 1978; Dennett, 1978; Harman, 1978) because an agent with a false belief holds an informational state incongruent with reality. Then, this ability requires that an individual understands the situation in the terms of the agent, which are necessarily different from how the individual himself conceives of it. As a result, developmental psychologists devised the false belief task (FBT) to investigate children’s ability to represent another person’s mistaken belief (Wimmer & Perner, 1983). There are two main versions of the classical FBT: the “change of location” or the “unexpected transfer” version (Baron-Cohen, Leslie, & Frith, 1985; Wimmer & Perner, 1983), and the “unexpected contents” variant (Gopnik & Astington, 1988; Hogrefe, Wimmer, & Perner, 1986; Perner, Leekam, & Wimmer, 1987). In the former, the child is told (or shown) a story in which the main character puts an object in location x and then, while the character is absent, the participant witnesses how the object is transferred from location x to location y. Upon the character’s return, the participant is asked, “Where the character will look for the object?” In the unexpected contents task, the child is shown a box whose content is not what the cues depicted on its surface suggest, and is asked, “What do you believe is in the box?” After revealing what is in fact inside, the participant is asked “What did you believe was in the box?” and “What another child will think is in the box when he first encounters it?” To pass the test, the child has to distinguish between her current and former perspective of the situation as well as the character’s perspective. Correct performance on any version of the FBT indicates that the child understands that the agent will behave according to her beliefs about the reality (i.e., her perspective) which diverge from the reality itself (i.e., these beliefs are false) and from the child’s true beliefs about the situation. Using multiple variations of this basic test, a huge amount of converging evidence supported the conclusion that children consistently pass the FBT by the age of 4 (Wellman, Cross, & Watson, 2001). Interestingly, younger children consistently give the wrong answer in every version of the FBT, responding from their own point of view, not from the agent’s (Apperly, 2011). However, doubts remained that passing the classical FBT involved additional abilities than just theory of mind, both in terms of linguistic abilities and executive control requirements (Ozonoff & McEvoy, 1994; Russell, Mauthner, Sharpe, & Tidswell, 1991). Further support for this suspicion came from an experiment in which 35-month-old children, who were allowed to answer by looking instead of verbally answering a question, passed the test (Clements & Perner, 1994). More recently, Rubio-Fernández and Geurts (2013), who also used the classical change of location task but required the participants to move the puppet character themselves instead of verbally answering questions, found that 36 months old passed the test. In addition, Setoh, Scott, and Baillargeon (2016) revealed that 30 months old also succeeded in the change of location task if they received two previous practice trials resembling the structure of the test question and if the target object was removed from the scene, so the child was uncertain about the object’s last location. As the structure of the classical FBT is made simpler, in terms of verbal competence and executive control requirements, attribution of false beliefs is found earlier than previously thought. During the last decade, this suspicion gave rise to a new set of simpler, spontaneous- response FBT, developed in order to find out how early attribution of false beliefs, and therefore, theory of mind, can be said to appear in development (Scott, Baillargeon, Song, & Leslie, 2010). These new experimental paradigms involve violation-of-expectation (VOE) (Onishi & Baillargeon, 2005), and anticipatory looking (AL) measures (Southgate, Senju, & Csibra, 2007), as well as a variety of interactive designs (Buttelmann, Carpenter, & Tomasello, 2009). The VOE paradigm exploits the fact that when an individual’s expectations are violated, she is surprised and surprise involves looking longer at the spot where the relevant event is taking place. Accordingly, it involves two steps: first, a habituation or familiarization phase, in which the infants are experimentally induced to form context-dependent expectations by being repeatedly exposed to the same event. Secondly, in the test trials, they are presented with either an expected or an unexpected event, given the induced expectation. The prediction is that infants will look longer to unexpected events as a sign of surprise. By experimentally modifying the event, authors try to determine the content of the infants’ expectations formed during the habituation trials (Jacob, 2013). Onishi and Baillargeon (2005) designed the first spontaneous- response study applying this paradigm in false belief understanding as a variation of the classical change of location task. In the familiarization phase, infants first watched as an actor hid a slice of watermelon in a green box and retrieved it from there twice. Next, a sort of magic change occurred. In one false belief (FB) condition, the watermelon moved by itself (in fact, the agent moved it using a magnet under the table) from the green to a yellow box while the agent was absent (FB-green condition). In another FB condition, the agent was present while the watermelon moved to the yellow box but then it returned to its original position as soon as the agent disappeared from the scene (FB-yellow condition)1 . In the experimental trials, infants watched the actor searching in one of the two boxes. The authors predicted that infants would look for a longer period of time when the agent acted unexpectedly (according to her false belief): if she went to the yellow box in the FB-green condition or to the green box in the FB-yellow condition. After this seminal paper, the VOE paradigm has spread over the field as a measure of belief understanding2 . On the other hand, other researchers have chosen infants’ first look in the test trial as the dependent variable. This first look generally takes place before the display is shown. This anticipatory looking method is a measure of what an infant expects to happen, given what happened before in his/her presence (Heyes, 2014). Clements and Perner (1994) were the first in using the AL measure in the context of a standard change of location FBT. Instead of directly asking where the agent would look for her toy upon her return, participants heard the experimenter say to himself “I wonder where she’s going to look?”. Most 35 month-old children correctly anticipated the spot where the agent would search according to her belief looking first at that location. Southgate et al. (2007) updated the AL methodology using an eye-tracker and removing any verbal interaction in order to test 25-month-old infants in a spontaneous-response FBT. They predicted that infants would look in anticipation where an agent with a false belief will search for an object, on the grounds of the visual information previously available to both. Interestingly, the object was not just changed of location but removed from the scene. After Southgate and colleagues’ positive results, other authors started to use infants’ first look as a dependent measure of false belief understanding. In addition, a new set of spontaneous- response FBT emerged that took advantage of infants’ intentional interaction abilities, such as active helping and referential communication. These tasks typically involve simple language and introduce participants to a scene in which an agent does or does not witness an event. Children are then given a verbal prompt –but one that only indirectly taps their representation of the agent’s epistemic state. Authors measure young children spontaneous behaviors in response to the prompt, such as helping and pointing, as a proxy for belief understanding. In the first FBT using the interactive paradigm, infants watched as an agent placed a toy inside a box, and later as another person switched the toy from one box to a second one while the agent either witnessed the switch (true belief condition) or not (false belief condition). Then, the agent unsuccessfully attempted to open the box that originally contained the toy and infants, who were previously instructed on how to open the boxes, spontaneously helped her. Buttelmann et al. (2009) expected that infants helped the agent to open the other box, where the toy is actually hidden, only in the false belief condition. So far, many other studies have employed infants’ helping or pointing behavior to measure false belief ascription. Each of the previous studies on early false belief understanding provided positive results showing that infants, from 15 months old, can attribute false beliefs to another agent. This new picture of young children tracking another agent’s false beliefs gave rise to a paradox in development: on the one hand, infants succeed in different non-verbal and spontaneous versions of the FBT but, on the other hand, children cannot reliably pass the verbal, explicit versions of the FBT until they are 4 years old. Different theoretical approaches tried to explain this developmental paradox. For instance, advocates of a nativist or an early, full mentalistic capacity account take the results of spontaneous-response FBT at face value, claiming that they prove genuine false belief understanding (Carruthers, 2013; Jacob, 2013; Leslie, 1994; Leslie, Friedman, & German, 2004; Scott & Baillargeon, 2017). They argue that 3 year-olds’ difficulties with the explicit FBT are related to processing demands but not to their understanding of beliefs, thus applying the strategy first proposed by Fodor (1992) of an innate, full competence but limited performance due to the task’s requirements. Working memory and related executive control abilities change throughout children’s development but not their ability to understand propositional attitudes in general, which is assumed to be complete, universal, and innate or it appears very early in ontogeny. In other words, 2 year-old children do attribute full-blown beliefs –beliefs with propositional contents– as part of their understanding of others as intentional agents. Empiricists, on the other side, deny that infants’ performance on spontaneous-response FBT show that they attribute beliefs in first place. From their point of view, infants’ behavior can be explained by low-level processes such as a novelty preference (Heyes, 2014) or simpler behavioral rules (Perner & Ruffman, 2005). For example, Heyes (2014) has forcefully argued that the novelty generating the surprise in the VOE and the AL procedures appears at a lower representational level than the true or false belief representations. Infants' looking behavior is a function of the degree to which the observed (perceptual novelty) and remembered or expected (imaginal novelty) low-level properties of the test stimuli –their colors, shapes, and movements– are novel with respect to the earlier events encoded by the infants in the experiment. On Perner’s view, on the other hand, infants’ performance in these tasks can be accounted in terms of learned or innate behavioral rules, such as that “people tend to look for an object where they last saw it and not necessarily where the object actually is” (Perner & Roessler, 2012). More recently, though, a two-systems account of belief-ascription has been put forward (Apperly & Butterfill, 2009; Butterfill & Apperly, 2013; Low, Apperly, Butterfill, & Rakoczy, 2016). According to this dual-system view, overcoming the spontaneous-response FBT depends on a fast and implicit but inflexible system. Such system reasons about “belief-like” states, called “registrations”, and develops early in ontogeny. A registration is a non- propositional, extensional state that holds relations to objects and properties and has its effects directly on action. Someone who reasons about registrations would track beliefs, true and false, but in a limited range of situations. For example, infants can attribute registrations regarding the location of an object and thus succeed in many spontaneous-response FBT (Apperly & Butterfill, 2009). However, they cannot track beliefs that involve quantifiers, indefinitely complex combinations of properties, or aspectuality, that is, situations in which a protagonist refers to an object under one aspect in contrast to another (Apperly & Butterfill, 2009; Rakoczy, Bergfeld, Schwarz, & Fizke, 2015; Rakoczy, 2017). The fact that registrations have signature limits explains why young children fail spontaneous tasks in which the protagonist was mistaken about the identity of an object (Fizke, Butterfill, van de Loo, Reindl, & Rakoczy, 2017). Since registrations are relational states, they fail to capture the intensionality characteristic of propositional attitudes and the aspectuality of the agent’s mental state (Rakoczy, 2017). On the other hand, a late-developing and flexible system supports the attribution of genuine beliefs. This system is effortful and explicit, but also inefficient, and supports children’s success in explicit FBT. Beliefs, unlike registrations, are propositional, intentional states, which engage in holistic and complex interactions with other propositional attitudes. Such a flexible system allows children to overcome the limitations posed by the early system and track more complex beliefs. Lately, Tomasello (2018) has argued that insofar as the previous theories are based on individual cognition, the developmental puzzle will not be solved. The key to understanding beliefs requires an account of the processes of social and mental coordination with other agents and their perspectives. According to Tomasello (2018), the origins of this coordination lie in infants’ join attentional activities around their first birthdays, in which they relate two perspectives. To pass the spontaneous-response FBT, infants just need to display their ability to triangulate, the same one that they use in joint attention, which involves tracking what the agent sees or has seen in the test, and how this information will affect the agent´s behavior. But such epistemic tracking does not imply that the child should understand that the agent´s belief is incorrect; in other words, the infant does not need to compare the agent’s perspective with his own view of the situation or with the objective situation (Tomasello, 2018). While three-year-old children answer the false belief question focusing on the objective perspective, 4- to 5-year-old children pass the FBT because they come to coordinate the three perspectives involved: the child’s, the agent’s and the objective perspective. Much social and communicative interaction with others is required before such achievement, which is mainly characterized by communicative exchanges involving joint attention to mental content (Tomasello, 2018). Almost fifteen years after the seminal study from Onishi and Baillargeon, there is no consensus about the capacities displayed by the infants in the non-verbal tasks. After the positive evidence obtained in the pioneer studies, new tasks either confirmed a similar pattern of results or failed to replicate the original findings (most of them were not published; see Kulke & Rakoczy, 2018). Given the mixed pattern of the accumulated evidence, it is time for a meta-analysis of the literature on implicit theory of mind. Specifically, this meta- analysis explores how correct performance in spontaneous-response FBT has revealed and whether (and how) other variables mediate this effect. Inclusion and exclusion criteria ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We used the following criteria to select studies to include in the meta-analysis. We only included studies testing participants’ capacity to attribute false beliefs in which their mean age was 26 months old or younger. Thus, we excluded all false belief conditions testing older children. We only included tasks that employed one of these three paradigms for implicit false belief ascription: the violation of expectation methodology, the anticipatory looking and the interactive paradigms (grouping both tasks that require infants’ helping and pointing behaviors). We only included implicit false belief conditions; thus, we excluded true belief, ignorance and control conditions that might be present in the studies as well. We only included false belief conditions in which typical- developmental infants were tested. Then, we excluded false belief tests carried out on samples with atypical development. We included all the tasks published from Onishi and Baillargeon seminal paper (2005) until 4th February 2019. Although we checked for unpublished data in the overview of studies done by Kulke and Rakoczy (2018), most of these data are now published and we have already included in the current meta-analysis3 . Moderator and mediator analyses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Each condition included in the meta-analysis was coded for the following dependent variables, most of them following Wellman et al. (2001)4 : Paradigm: the paradigm employed in each condition, that is, the violation of expectation (VOE), the anticipatory looking (AL) or the interactive paradigm. Age: the mean age of participants in each condition (expressed in months). In some cases, the authors reported the mean age of participants from the whole group (grouping together either true and false belief conditions or two false belief conditions). Sample size: the number of participants in a condition. Dropped participants: the number of participants excluded from each condition. In many cases, the authors reported the total amount of excluded participants in the whole study, grouping together false and true belief conditions, or two false belief conditions. Familiarization trials: the number of familiarization trials performed in each condition. When the authors employed warm-up trials, instead of familiarization trials, we coded the familiarization trials as 0. Belief: the type of belief attributed to the agent. We coded four levels, distinguishing beliefs about: an object location (L); a non-obvious property of an object (NP); the identity of an object (I); or Agent: three levels that described the nature of the agent as: a real present person; a videotaped person; or a virtual agent (including cases of a human-like individual, a geometric shape or an animal). Real presence of the target object: whether, at the test trial, the target object was: real and present (i.e., the object was inside one of the boxes); or absent (i.e., the object had been removed from the scene). Object movements: the number of displacements of the object that the agent did not see before the object arrived at its final position. Motive of the transformation: two levels capturing whether the key transformation (e.g., the change of location or the substitution of unexpected contents) was done: to trick: to explicitly trick the agent, or for other reason: for some other reason including no explicit reason at all. Salience of the agent’s mental state: we distinguished four levels: absence: the false belief state had to be tracked from the agent’s absence during the key events; back turned: the false belief state had to be tracked due to the agent’s back turned during the key transformation; first-person experience: the false belief experience was demonstrated initially on the children themselves; and blindfold: the agent used a blindfold during the transformation. Interaction: we coded whether there was interaction between the child and the agent during the test or not. Design: indicates the design employed in each condition. We distinguished between: a within-subjects design (WS): participants were tested in both test events (congruent and incongruent events in the VOE case) or in both conditions (false and true belief conditions); and a between-subjects design (BS): participants were tested in only one test event (in the VOE case) or in only one false belief condition. Test trials: the number of test trials that were conducted with each child. Depending on the paradigm employed, we coded the dependent variable in each case. In the VOE paradigm, we coded the mean looking time (in seconds) that infants spent looking to (a) the expected event, that is, when the agent acted according to her false belief; and (b) the unexpected event, that is, when the agent acted in a way that was not according to her false belief. In the AL and the interactive paradigms, we coded the percentage of passers, that is, the number of infants that showed the correct anticipatory behavior (in the AL paradigm)5 or the correct helping or pointing behavior (in the interactive paradigms). Search strategies & coding procedures ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The literature on implicit false belief tasks is relatively recent (the first study was published in 2005), so new studies are known in the area and are cited as new articles appear. However, we checked the databases EbscoHost – PsycINFO, SCOPUS, and Web of Science (WoS) to ensure we were taking into account all the studies published up to date. The search terms employed were false belief, implicit false belief tasks, false belief + infants, spontaneous response false belief tasks, violation of expectation + false beliefs, looking paradigm + false beliefs, theory of mind. We addressed studies published in English and we only included published studies. All the studies were coded by the first author, after agreement on the codes with the third author. Table 1 shows the final list of references included in the current meta-analysis. Statistical methods ~~~~~~~~~~~~~~~~~~~ We calculated the effect sizes based on the data described in the articles and on the responses from the authors to requests for further information6 . For within-subjects designs, in which infants were tested on both congruent and incongruent test trials, the standardized mean change was calculated. Its formula is equal to (m1i - m2i)/Sdif, where m1i and m2i refer to the observed means at the two measurement occasions. To calculate the standard deviation of the paired differences, the corresponding standard deviations of the two measurement occasions and the correlation coefficient between both should be reported. No condition that employed a within-subject design reported the correlation between participants’ two measures. Then, we fixed a medium correlation of 0.5 for these cases7 . On the other hand, studies conducted with the AL and the interactive paradigms provide data for single groups with respect to a dichotomous dependent variable: correct anticipatory look (it means, looking in anticipation to the place where the agent should go according to her false belief), or incorrect anticipatory look; or correct helping or pointing behavior (i.e., opening the correct box, giving the appropriate object, or pointing to the correct location), or incorrect helping or pointing behavior. Therefore, we used the logit transformed proportion (log odds) as an outcome measure for the studies from both paradigms. The log odds for single groups is equal to the log((xi/ni)/(1-(xi/ni))) where xi denotes the number of individuals experiencing the event of interest and ni refers to the total number of individuals in the group. The variance of log odds is equal to 1/xi + 1/(ni-xi). Finally, we converted d scores into log odds by applying the formula d · (π/√3); and the variance of d into the variance of log odds following the formula v(d) · (π2/√3). Then, we calculated the effect size standard error (SE) which is equal to the square root of the variance. 95% confidence intervals were calculated in the statistical program R Core Team (2013). The lower bound of the confidence interval results from applying the formula log odds – (1.96 · SE), while the upper bound of the interval is determined by log odds + (1.96 · SE). We employed a random effects model in which we assume that the true effects are normally distributed (Borenstein, Hedges, Higins, & Rothstein, 2009). Under this model, two sources of variability are assumed: intra-study variability (due to sampling error) and inter-study variability (each study estimates its own parametric effect) (Ausina & Meca, 2015). Heterogeneity was assessed by calculating different statistics. First, we used the Q statistic, which is the weighted sum of squares (WSS) on a standardized scale, and its p-value as a test of significance. As a standard score, it can be compared with the expected WSS (on the assumption that all studies share a common effect) to yield a test of the null and also an estimate of the excess variance (Borenstein et al., 2009). We also quantify T2 that reflects the variance of the true effects; and I2 that reflects the proportion of observed dispersion that is due to true heterogeneity and is expressed as a ratio. Higgins, Thompson, Deeks, and Altman (2003) suggest that values on the order of 25%, 50%, and 75% might be considered as low, moderate, and high, respectively. Finally, when we add moderators to the model, we report R2 which reflects the amount of heterogeneity accounted for the variables included in the model. Possible publication bias was assessed with a funnel plot. The funnel plot is a graphical display of effect sizes against their standard errors and shows that, in the absence of publication bias, studies will be symmetrically distributed around the mean effect size, since the sampling error is random. In the presence of publication bias, the studies are expected to follow the model, with symmetry at the top, a few studies missing in the middle, and more studies missing near the bottom (Borenstein et al., 2009, p. 284). We used the metafor package (Viechtbauer, 2010) in R Core Team (2013) to calculate effect sizes, heterogeneity estimates, to run the meta-analysis, and to perform the forest, the funnel, and the meta- regression plots.8 Description of studies ~~~~~~~~~~~~~~~~~~~~~~ A PRISMA flow diagram illustrates the search strategy and study selection (see Fig. 1). The search procedure yielded 44 references for which additional information was obtained. Two of these references were duplicates (the VOE task with 18-months-old from Poulin- Dubois & Yott, 2018 with the one from Yott & Poulin-Dubois, 2016; and the AL study of Sodian et al., 2016 with the one of Thoermer, Sodian, Vuori, Perst, & Kristen, 2012), thus we removed them. Of these 42 papers, 2 studies were carried out with another paradigm (Kovács, Téglás, & Endress, 2010; Southgate & Vernetti, 2014) and, in 5 of them, the target group was constituted by older children (Burnside, Ruel, Azar, & Poulin-Dubois, 2017; Burnside, Wright, & Poulin-Dubois, 2018; Grosse Wiesmann, Friederici, Singer, & Steinbeis, 2017; Oktay-Gür, Schulz, & Rakoczy, 2018; Wang & Leslie, 2016), thus we excluded them. Out of them, 2 additional articles were eliminated because of insufficient reported data to calculate the effect sizes (Song & Baillargeon, 2008; Träuble, Marinovi, & Pauen, 2010). Finally, the data from 33 articles were included in the meta-analysis. From these definitive 33 papers, we additionally excluded some conditions in which authors tested older children (Fizke et al., 2017; Grosse Wiesmann, Friederici, Disla, Steinbeis, & Singer, 2018; Kulke, Reiß, Krist, & Rakoczy, 2018; Priewasser, Rafetseder, Gargitter, & Perner, 2018; Schuwerk, Priewasser, Sodian, & Perner, 2018) or deaf participants (one condition from Meristo et al., 2012), and other conditions due to insufficient data to calculate the effect sizes (two VOE conditions from Powell, Hobbs, Bardis, Carey, & Saxe, 2018). The final selection of articles amounted to 56 false belief conditions in which a total of 1469 infants were tested. Table 2 reports the descriptive information of each false belief condition included in the meta-analysis like the year of publication, the sample size, infants’ mean age, the paradigm employed, the effect size and the 95% confidence intervals. Table 3 provides a descriptive summary of the database grouped by the variables coded. The mean age of all the participants of the database was 19.55 months and, from the 56 false belief conditions, 19 employed the VOE paradigm, 15 used the AL procedure and 22 the interactive methodology. Model 1 We built a random effects meta-analytic model with maximum likelihood estimation on infants’ performance in spontaneous-response false belief tasks. This first model (model 1) included no covariate. The results showed that the estimated average effect size (calculated in log odds) is equal to β = 0.57, 95% CI [0.33; 0.80] (see Fig. 2). It, therefore, suggests that the average log odds corresponds to an odds of 1.76, meaning that correct performance on spontaneous-response false belief tasks was about 1.76 times more likely than incorrect performance9 . In terms of probabilities, the probability of correct performance was 64%10 . The null hypothesis can be rejected (z = 4.69, p < .0001). The total set of studies was heterogeneous as evidenced by significant Q-values (Q[1,55] = 163.81, p < .001). T2, which reflects the amount of true heterogeneity (i.e., the variance of the true effects), is 0.50 (SE = 0.15). The proportion of observed dispersion in the data that is due to heterogeneity is specified by I² = 71.5%, which indicates a medium-high proportion of variance due to heterogeneity rather than chance. The funnel plot showed that all the effect sizes falling within the grey area would be statistically non-significant in a two-tailed test (Fig. 3). We can see an asymmetrical distribution of the effect sizes, with more studies on the right part of the plot -showing larger effect sizes- that have less precision (smaller samples). This graphical display suggests a publication bias: some studies with non-significant results were not published. Model 2 Part of the heterogeneity found may be due to the influence of moderators (Viechtbauer, 2010). We examined this possibility by fitting a mixed-effects model including theoretically relevant predictors as the year of publication, infants’ age, the paradigm employed, the type of belief, the type of agent, the motive of the transformation, the design of the study, and the presence of the target object as moderators. This second model 2 includes 8 predictors. We performed a multiple meta-regression on model 2 using the restricted maximum likelihood method. Table 4 shows the results. In this model, the estimated amount of residual heterogeneity is equal to T2 = 0.23, suggesting that 53.80% (R2) of the total amount of heterogeneity can be accounted for the eight moderators included in the model. The proportion of observed dispersion due to heterogeneity (I²) is reduced to 52.14%. We can reject the null hypothesis based on the omnibus test (QM = 49.67, df = 12, p < .0001). Before analyzing the implication of each moderator, a statement of caution is crucial. Given the large number of moderators considered and their confounds with one another, we need to be circumspect in the interpretation of the results about the influence of these variables on children’s performance. Examining the predictors of model 2, we find that year of publication (z = -4.54, p < .0001), infants’ age (z = 2.26, p = .02) and the motive of transformation (to trick the agent; z = 0.32, p = .04) have a significant influence on performance. The test for residual heterogeneity is significant (QE = 87.20, df = 43, p < .0001), indicating that other moderators, not considered in the model, are influencing children’s performance. These results indicate that as time passes, the probability of passing the test diminishes -0.15 units (95% CI: -0.22; -0.09) in terms of the average log odds of successful performance. This means that incorrect performance on spontaneous-response false belief tasks is 0.86 times more likely than correct performance as the year of publication is higher, suggesting that null results or children’s incorrect performance on the task are currently being published. On the other hand, when older infants are tested, the log odds of successful performance increase 0.08 units (95% CI: 0.01; 0.15), meaning that correct performance is 1.08 times more likely than incorrect performance as infants get older. When the motive of the transformation is to trick the agent, it results in an increase of 0.65 (95% CI: 0.02; 1.29) units in the log odds of correct performance. This indicates that correct performance is 1.91 times more likely than incorrect performance. Anyway, the confidence intervals of the effect sizes of these two predictors almost touch zero. Then, we should cautiously interpret them. The effects of all other predictor were not significant (all p values > .05). Model 3 Lastly, we examined a simpler model by fitting a mixed-effects model only including the year of publication and the paradigm as moderators. The results of the multiple meta-regression for model 3, that includes two predictors, are presented in Table 5. The estimated amount of residual heterogeneity is equal to T2 = 0.19, suggesting that 61.30% (R2) of the total amount of heterogeneity can be accounted for the inclusion of these two moderators in the model. The proportion of observed dispersion due to heterogeneity (I²) is 48.19%. We can reject the null hypothesis based on the omnibus test (QM = 40.12, df = 3, p < .0001). A closer inspection of the predictors shows that year of publication appears to have a significant influence on performance (z = -4.32, p < .0001) as well as the VOE paradigm (z = 2.92, p = .003). Results indicate that as the year of publication is higher, there is a reduction of -0.11 (95% CI: -0.16; -0.06) in the average log odds of successful performance. This result indicates that incorrect performance on spontaneous-response false belief tasks is 0.90 times more likely than correct performance as time passes, in the same vein as the result of model 2. On the other hand, performing the task in the VOE paradigm implies a change of 0.75 (95% CI: 0.25; 1.26) in increasing the log odds of successful performance in the task. This means that correct performance is 2.12 times more likely than incorrect performance when the test is run with the VOE paradigm. Models’ comparison The models were compared using the Akaike information criterion (AIC) and the Bayesian information criterion (BIC). Table 6 contains the values of each criterion for the three models. Comparing model 2 with model 1, both AIC and BIC decrease: AIC is reduced from 160.26 to 123.74, and BIC decreases from 164.28 to 148.40. This shows that the fit of the second model is improving. If we then consider model 3, in comparison with model 2, AIC increases (AIC = 129.09) while BIC considerably decreases (BIC = 138.84). BIC penalizes for each additional variable included in the model, thus, the fit of model 3 –with only two parameters– improves in comparison to model 2 that possesses eight predictors. As the AIC for model 2 and model 3 are not substantially different, the fit of the third model is better than the fit of the first one and therefore it is preferable over the other two models. A meta-analysis per paradigm ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We also run separate meta-analysis per each of the paradigms employed (see Fig. 2). Results show that the VOE paradigm had a significant mean effect size β = 1.28, 95% CI [0.86; 1.69], implying that correct performance on the task was about 3.58 times more likely than incorrect performance (see Table 7). The results from the interactive paradigms also indicate a significant mean effect size β = 0.31, 95% CI [0.01; 0.61], suggesting that correct performance on the task was about 1.36 times more likely than incorrect performance. However, this is a lower probability and the confidence interval almost reach zero. On the other hand, the meta-analysis for the AL paradigm provided no significant mean effect size (β = 0.16, 95% CI: -0.24; 0.57). Results from the VOE and the interactive paradigm suggest that the average effect size found when all paradigms were included as moderators come from these methodologies. Main effects analyses ~~~~~~~~~~~~~~~~~~~~~ We also inspected the main effects of all the variables we have coded. We will first consider the effect of the year of publication and we will then analyze the other methodological moderators separately. Year of publication might have a special status, likely indexing publication bias rather than methodological differences11 . Year of publication The year of publication, on its own, has a significant influence on performance (z = -5.32, p < .0001). Insofar as the year of publication is higher, the probability of passing the test falls -0.14 units (95% CI: -0.19; -0.09) in terms of the average log odds of successful performance (see Fig. 4). This result suggests that children’s incorrect performance on spontaneous-response false belief tasks is 0.87 times more likely than correct performance as the year of publication increases, in the same direction as we found in models 2 and 3. Methodological moderators We found that the following variables produced a significant main effect: the paradigm, the type of agent, the object’s movements, the salience of the agent’s mental state and the sample size. Each of them has, on its own, a significant influence on young children’s performance. All the remaining methodological moderators produced non-significant main effects. Table 8 presents the estimates, the p values and the confidence intervals for all the significant predictors. The effect of the paradigm was already analyzed in model 3. On the other hand, when the agent was a videotaped person, there was a change of -0.62 (95% CI: -1.15; -0.08) in reducing the average log odds of successful performance.12 Regarding the object’s movements, when it was moved twice outside the view of the agent, there was a change of -1.16 (95% CI: -1.90; -0.42) in reducing the average log odds of successful performance. Then, when the agent’s false belief was induced because she was turned back during the key events, there was a change of -0.77 (95% CI: -1.31; -0.24)13 . The sample size also seemed to influence performance: when bigger the sample size, lower the probability of finding a correct performance in the task (β = -0.01, p = .02, 95% CI: -0.02; 0.00). However, the confidence interval for this effect size includes zero. Finally, the following Figs. 5–7 show how the variables are distributed in the conditions recruited for the meta- analysis. Fig. 5a shows how all the studies that found big effect sizes were run with small samples. When bigger sample sizes were employed, the effect sizes tended to decrease. Fig. 5b, on the other hand, shows the relation between dropped participants and the effect sizes found in each condition: only a few studies excluded more than 20 participants. In addition, when all infants were included in the sample, the resulting log odds of correct performance was low (it only happened in two studies). Effect size as a function of infants’ age is shown in Fig. 6. We can see that, as infants become older, correct performance tended to decrease although it was a non-significant trend (p = .06)14 . Interestingly, most of the tests were carried out with 18-month-old infants. This age group, however, showed an uneven performance on the task: they correctly passed the VOE tests but showed all possible results in the other two paradigms considered (above, at chance and below chance performance). Interestingly, there were no studies carried out in the specific age range from 19 to 23 months old. Also, there was only one study employing the VOE paradigm with older infants (24 months old) and interactive tests were consistently conducted from 18-month-old. Fig. 7 shows the dispersion of the data in each paradigm according to (a) the design employed, (b) the interaction between the child and the experiment, and (c) the motive of the transformation. Interactive paradigms mostly employed between- subjects designs (by testing the child only in a FB condition), while within- subjects designs were used in the other paradigms (although not as often as between-subjects ones; Fig. 7a). Fig. 7b shows how interaction overlaps with the paradigm: AL and VOE paradigms are non-interactive measures and interaction appears in those tasks in which infants should help or point to the agent. Finally, Fig. 7c shows that the transformation of the scene (i.e., change of location) was mainly done to trick the agent in the interactive tasks. All the other changes that provoked a false belief in the agent were done for other (including no) reason.","This meta-analysis reveals that, on average, correct performance on spontaneous-response false belief tasks is more likely than incorrect performance, suggesting that the tasks are tapping a real phenomenon. However, the meta-analysis simultaneously uncovers important issues: (i) an increase in year of publication implies a decrease in correct performance; (ii) a considerable heterogeneity is present among studies; (iii) the asymmetrical distribution of effects sizes in the funnel plot suggests publication bias; and (iv) correct performance depends on the type of paradigm used. Taken together, these results indicate that the developmental puzzle is still unsolved. First of all, we saw that large effect sizes were found at the onset of the research in implicit false belief attribution, but this outcome could not be replicated in more recent experiments. Relatedly, those large effect sizes were found with small samples, whereas later studies used bigger samples. Secondly, the high percentage of heterogeneity found across studies suggests that, despite the significant effect, we still do not know which factors really explain the variance across studies. Since two of the parameters (i.e., the year of publication and the paradigm) considerably reduced such variance among studies -from a high ratio to a medium ratio of heterogeneity-, there was still much variance in the dataset that the model could not explain. Thirdly, the skewed distribution of the effect sizes hints at the presence of publication bias in this literature. Current initiatives to look for unpublished papers and to replicate previous studies imply that many researchers become aware of this bias. For example, some journals have invited to submit replications of the tasks (for instance, the special issue Understanding theory of mind in infancy and toddlerhood published by the journal Cognitive Development) and the scientific community is doing an effort to conduct direct as well as conceptual replications (Kulke & Rakoczy, 2018). Especially relevant for the future of the field is the ManyBabies framework (Frank et al., 2017) which is currently running a large size replication on infants’ theory of mind. The project ManyBabies 2 aims to run three waves of multi-lab studies developing tasks to measure VOE, AL and interactive behaviors. In this effort, researchers from different theoretical point of views will collaborate to devise the experiments and run them in their own labs. The results of the present meta-analysis, then, constitute a first step on that direction insofar as it provides an overall picture of the field that would be comparable with a future overview of the area as new results are accumulated. Fourth, the type of task employed seems to be relevant. When the year of publication and the paradigm were included in the model, only the VOE paradigm –among the other two– produced a significant effect. This consequence implies that infants reliably performed above chance only within this methodology. But if false belief attribution were achieved at these ages, then it should take place regardless of the task used, just as it was found for elicited-response FBT. Studies employing within-subject design are critical to shed light on this lack of convergence across paradigms (Poulin-Dubois & Yott, 2018; Powell et al., 2018). In addition, many of the VOE effects are due to a single lab; further, analytic decision-making varies substantially from paper to paper in this body of work. These studies combine three looking time measures in different ways. Disruptive looking and continuous-looking measures (used to end a trial), and a cumulative minimum looking time (before which a trial could not end) are systematically varied across conditions and studies without independent justification (Rubio-Fernández, 2019b). This fact certainly obscures the interpretation of the results in the VOE paradigm. On the other hand, there is a great variability in the confirmatory looking times across the FB conditions in the VOE paradigm that it is surprising in itself, because it could not be predicted a priori (all the looking times for expected and unexpected trials are reported in Table 9). Looking times considerably varied from one experimental condition to another: they stretched from 4.52 to 29.5 s when unexpected outcomes are presented and from 3.27 to 18.8 s in the case of expected events. At least part of this variability could be due to the different materials employed in each task, the setting in which studies were embedded, and the cultural context in which the studies were run. However, the existing variability should be pointed out and make researchers to raise the question of how to set and justify a criterion for expected versus unexpected looking times rather than stipulate them ad hoc as it seems it is done depending on the times found.15 Anyway, what matters for the VOE methodology is that longer looking times, in fact, indicate a surprise reaction rather than duration per se. As it was well established long ago, longer looking may also indicate a preference for novelty (Fantz, 1967): without surprise, longer looks might indicate a preference for an event that is seen as new rather than unexpected. This point brings us to a second consideration, common in the debate about the use of this methodology to assess conceptual development in first year-olds: the familiarization phase involved in the VOE paradigm might just induce perceptual expectations instead of the assumed one in terms of belief and intention attribution. The fact that a long and constant familiarization phase is required reinforces this possibility, providing the participants with some anticipation of where the elements involved in the scene are likely to move around. In other words, the familiarization phase might induce the infant to attend to a particular location (Falck, Brinck, & Lindgren, 2014) or to a particular perceptual configuration (Heyes, 2014)16 . However, the studies did not include a measure of familiarization or habituation. An independent way to assess the generated expectations, if any, during the familiarization phase would be welcome. The design should also involve greater perceptual variability in terms of colors, distances, objects and all the perceptual details of the situation. It could easily be tested in (i) a control condition in which perceptual changes were introduced in the familiarization phase to check whether the VOE response also appeared, thereby indicating conceptual understanding; and (ii) a control condition in which the familiarization phase involved perceptually different instances of the same intentional event. Additionally, if infants were already able to ascribe beliefs in the familiarization phase, a very brief familiarization period would suffice. Main effect analyses showed how multiple aspects, on their own and in combination, decreased the probability of correct performance in these tasks. We may be unaware of many other variables that moderate infants’ performance on these tasks. Although we attempted to codify more parameters, like the separation between locations and the delay between hiding the object and testing, most authors did not report such information. For instance, the distance between the two locations (i.e., boxes) was reported in only 13 out of 56 conditions. In those studies, such distance considerably varied from 5 cm to 100 cm. On the other hand, the delay between hiding the object on its last position and testing was not specified. In some cases, it could be estimated but it was a difficult task since infants controlled the end of the trials in the VOE studies. Any estimation, then, would be imprecise. The analysis of both parameters could give us a hint for the explanation of the variability in the data. Study limitations ~~~~~~~~~~~~~~~~~ We remark some potential limitations to the current meta-analysis that simultaneously constitute venues for new research. First, we only focused on peer-reviewed journal publications; hence we did not have access to reports that failed to be published. The present meta-analysis, as we claimed before, is a necessary first step to have a panoramic view of the evidence published so far. It is also an open invitation to submit non- published results in order to increase the number of false belief conditions considered and rerun the analyses. Second, other moderators that we did not code could have influenced the phenomenon (like nationality, the distance between locations, the delay between hiding and testing, etc.). Third, although we chose to work with the more prolific paradigms in the field, new ways to measure implicit false belief ascription are being developed, like the use of electroencephalography (EEG) to record brain wave patterns during the task (Southgate & Vernetti, 2014). Four, we selected sample sizes in which infants were younger than 26 months old. In future work, it would also be compelling to study the developmental trend of the capacity to attribute beliefs and see what happens in the transition from two to three years old. But this decision would also imply broadening the inclusion criteria in order to select verbal spontaneous-response paradigms, like the verbal AL and the verbal VOE tasks (He, Bolz, & Baillargeon, 2012; Scott, He, Baillargeon, & Cummins, 2012). Finally, we only included conditions in which infants were required to attribute false beliefs to another. However, almost all studies include additional conditions, like true belief as well as ignorance and other control conditions. Then, infants’ above chance performance on implicit FBT would be doubtful if they, for example, performed at chance or below chance level in the true belief condition. As Rubio-Fernández (2019a) correctly pointed out, different kind of criteria have been proposed for passing the FBT in infants and older children. While a differential performance between true and false conditions is expected for infants, older children also have to perform above chance level in both conditions in the traditional task. So, a careful analysis of infants’ performance in spontaneous-response true belief tasks is also needed to get a more comprehensive view of infants’ mentalistic abilities.","In conclusion, this meta-analysis is the first one in proposing an integrative view of all the published studies on spontaneous-response false belief tasks in children younger than 26 months-old. Infancy is a key developmental period in which implicit methodologies are being used to test the origins of many psychological constructs, such as theory of mind and the specific capacity to attribute false beliefs. The emerging picture from this systematic review, however, is puzzling. Although correct performance on spontaneous- response FBT was on average more likely than incorrect performance, the analyses revealed publication bias and high heterogeneity across the data. While several authors have shown that preschoolers display a robust developmental pattern from a below-chance to an above- chance performance in the standard FBT (Wellman et al., 2001), infants’ performance on the nonverbal FBT is equivocal and does not show any clear developmental pattern. Understanding the behavior of another agent according to her false belief would minimally require robust attributions. We mean that the attributer, who ascribes a false belief to an agent, should correctly help or inform her when necessary, look in anticipation to the location where she will go and look longer when she acts incongruently with her epistemic state. Experimental settings should move forward in that direction and test the same children across a battery of implicit tasks that employ different paradigms. In this way, we could examine individual differences in performance (e.g., infants that show understanding across all the implicit measures and infants that fail all of them) and convergent validity of the tests (like Poulin-Dubois & Yott, 2018 who failed to find convergent validity). Similarly, tasks that concurrently combine two implicit measures are welcomed, like measuring infants’ anticipatory look and (un)expected looks (e.g., Dörrenberg, Rakoczy, & Liszkowski, 2018) or participant’s anticipatory look and interactive behavior (work in progress)."],["Influential developmental theories claim that infants rely on goals when visually anticipating actions. A widely noticed study suggested that 11-month-olds anticipate that a hand continues to grasp the same object even when it swapped position with another object (Cannon, E., & Woodward, A. L. (2012). Infants generate goal-based action predictions. Developmental Science, 15, 292–298.). Yet, other studies found such flexible goal-directed anticipations only from later ages on. Given the theoretical relevance of this phenomenon and given these contradicting findings, the current work investigated in two different studies and labs, whether infants indeed flexibly anticipate an action goal. Study 1 (N = 144) investigated by means of five experiments, under which circumstances (e.g., animated agent, human agent) 12-month-olds show flexible goal anticipation abilities. Study 2 (N = 104) presented 11-, 32-month-olds and adults both a human grasping action as well as a non-human action. In none of the experiments did infants flexibly anticipate the action based on the goal, but rather on the movement path, irrespective of the type of agent. Although one experiment contained a direct replication of Cannon and Woodward (2012), we were not able to replicate their findings. Overall our work challenges the view that infants are able to flexibly anticipate action goals from early on, but rather rely on movement patterns when processing other's actions. --------------------------------------------------------------------------------","During the first year of life, infants start to visually anticipate other people’s actions (Adam et al., 2016; Ambrosini et al., 2013; Cannon & Woodward, 2012; Daum, Gampe, Wronski, & Attig, 2016; Falck-Ytter, Gredebäck, & von Hofsten, 2006). For example, 12-month-olds anticipate the goal of a simple manual reach-and-transport action (Falck-Ytter et al., 2006). However, infants show difficulties in anticipating actions when situations become more complex (Gredebäck, Stasiewicz, Falck-Ytter, Rosander, & von Hofsten, 2009). In addition, it has been suggested that action anticipation depends on whether a human or a non-human agent is performing the action (Cannon & Woodward, 2012; Daum, Attig, Gunawan, Prinz, & Gredebäck, 2012; Kanakogi & Itakura, 2011) and that movement characteristics of actions such as distances, durations and velocities have a strong impact on the anticipation of the goal of observed actions (Daum, Gampe et al., 2016). In this paper we present a series of studies conducted in several laboratories that investigated whether infants and adults are able to visually anticipate an action goal when two goals are available and whether they differentiate between human and non-human (i.e. animated) agents. Being able to understand that other people have goals is essential for processing social information. Understanding the goal-directedness of human actions has been related to the ability of perspective taking (Krogh-Jespersen, Liberman, & Woodward, 2015) and of coordinating one’s own actions with others (Sebanz, Bekkering, & Knoblich, 2006). Influential developmental theories have therefore stressed the role of goal encoding and anticipation for early social-cognitive development (e.g., Woodward, 2009). In the study by Falck-Ytter et al. (2006), infants anticipated the action of a hand that placed objects into a container (for a related setup see also Brandone, Horwitz, Aslin, & Wellman, 2014). However, movement path and goal were confounded in this paradigm, and for this reason, no conclusion is possible whether infants’ anticipations were based on the information provided by the movement or the information about the action goal. The encoding of an action as goal directed goes beyond the representation of pure physical movements. Accordingly, a study by Cannon and Woodward (2012) presented a manual reaching action with not only one but two possible action goals to infants at the age of 11 months. After being familiarized with the hand always grasping the same of the two objects, the objects’ position changed place for the following test trial. During the test trials, the infants observed an uncompleted manual movement where the hand stopped before indicating a clear movement direction. The infants showed anticipations towards the familiarized object in the now new location. This indicates that infants did not just anticipate the mere movement pattern but that they encoded the action as being directed towards a specific goal. This goal attribution served as the basis for the subsequent goal anticipation. As striking as these results are, different results were obtained by Daum et al. (2012). They used a similar methodological approach with a slightly modified paradigm. In their study, participants were familiarized with an animated fish that moved behind an occluder to one of two goal objects. In the test phase, with the location of the two goal objects being swapped, only 3-year-olds and adults anticipated the correct goal. Two-year-olds anticipated both the path and the goal, indicating that they seem to be in a transition phase. In contrast, 9- and 12-month-olds expected the fish continue to move on the movement path as in the familiarization phase. These contradictory findings represent a puzzle, particularly for developmental theories that capitalize on the role of goal understanding in early development. Notably, a number of studies that used a similar paradigm as Cannon and Woodward (2012) do not give a clear picture either. These studies are different from Cannon and Woodward as the human agent sat at a table and was fully visible to the participants (Krogh-Jespersen & Woodward, 2014, 2018; Krogh-Jespersen et al., 2015; Krogh-Jespersen, Kaldy, Valadez, Carter, & Woodward, 2018; Paulus, 2011). Participants observed an agent grasping one of two objects for once, followed by two consecutive test trials in which the agent performed an uncompleted reaching action. Again, the objects’ position was swapped for test trials. Krogh-Jespersen and Woodward (2014) demonstrated with this paradigm that goal-directed fixations of 15-month-olds needed more time to be initialized, indicating additional cognitive effort when taking an action goal into account instead of when the mere movement pattern was anticipated. Interestingly, goal-anticipations were not demonstrated consistently in this paradigm for older children. Two-year-olds anticipated neither the goal nor the previous location systematically (Krogh-Jespersen et al., 2018) and 21-month-olds only made goal-directed anticipations in the first, but demonstrated chance performance in the second test trial (Krogh-Jespersen et al., 2015). Overall, there is a heterogeneous pattern of results on whether or not young children show flexible goal anticipations for a human actor. Given this evidence it is on the one hand unclear from which age on infants anticipate other’s actions as goal-directed and on the other hand, whether they differentiate between non- human and human agents. Developmental theories claim that one’s own experiences are fundamental for understanding others actions. This is also known as the human first-view (Luo & Baillargeon, 2005; Meltzoff, 1995; Woodward, 2005). Because a human hand performed the action in the study by Cannon and Woodward (2012), this could have facilitated infants’ goal encoding. Some suggest that infants use their own motor abilities when anticipating other’s actions (e.g., Kilner, Friston, & Frith, 2007; Paulus, 2012). For example Krogh-Jespersen and Woodward (2018) demonstrated that already 8-month-olds anticipate an action goal, but only after they practiced reaching for an object themselves. Others suggest that infants are simply more familiar with human hands than with a non-human agent (Ruffman, Taumoepeau, & Perkins, 2012). This account would imply that infants are more experienced with hands grasping objects than with moving fish, as they have probably not often seen an animated fish before. Indeed, previous studies demonstrated earlier anticipations within infants when actions were more familiar (Cannon & Woodward, 2012; Filippi & Woodward, 2016; Gredebäck & Kochukhova, 2010; Gredebäck & Melinder, 2010; Kanakogi & Itakura, 2011). Cannon and Woodward (2012) showed that infants did not anticipate an action as goal-directed when performed by a mechanical claw (see also Adam et al., 2016; Kanakogi & Itakura, 2011). In contrast, the so called all agents- view (Luo & Baillargeon, 2005) proposes that infants attribute goals to any individual that can be identified as an agent (Leslie, 1995). It has been argued that humans attribute goals to non-human agents as long as they show specific characteristics, like self-propelledness (Luo & Baillargeon, 2005), equifinal variations (Kamewari, Kato, Kanda, Ishiguro, & Hiraki, 2005), or an action outcome produced by the agent (Adam, Reitenbach, & Elsner, 2017). However, most empirical evidence concerning the all agents-view comes from looking-time studies (e.g. Gergely, Nádasdy, Csibra, & Bíró, 1995; Luo & Baillargeon, 2005). Thus, it is still an open question whether infants are able to generate online anticipations when perceiving non-human instead of human actions. However, as it is neither empirically nor theoretically clear when and to what extent infants are able to anticipate an action based on the action goal and not on the movement pattern, we report two studies, assessed in two different labs, that address this question in further detail. Study 1 contains five experiments, which disentangle the precise methodological issues between Cannon and Woodward (2012) and Daum et al. (2012). The goal of the first study was to investigate which aspects are fundamental for infants’ goal encoding abilities within 12 months of age. Two major questions guided this research: First, whether infants show goal-directed anticipations for non-human as well as for human actions equally, as proposed by the all agents-view, or whether infants are better in anticipating the goal of a human action than the goal of a non-human action, as suggested by the human first-view. Second, whether the seemingly contradictory findings of Cannon and Woodward (2012) and Daum et al. (2012) are a result of further methodological differences. In the study by Daum et al. (2012), the agent shortly disappeared behind an opaque occluder. This paradigm was used to trigger participant’s eye-movements to the position where they expect the agent to reappear (adapted from Kochukhova & Gredebäck, 2007). However, infants have to maintain the association between the agent and the object when the agent is not visible, which might require additional cognitive capacities, such as memory or attention (Hespos, Gredebäck, von Hofsten, & Spelke, 2009; Jonsson & von Hofsten, 2003). This additional requirement of cognitive resources could be another reason why infants in the Daum et al. study (2012) showed goal-directed anticipations only from later on. Further, the two studies used different designs in providing the stimuli. Cannon and Woodward (2012) presented infants four blocks, which each contained three familiarization trials, one swap trial (in which the objects’ position changed place) and one test trial. Daum et al. (2012) presented participants eight familiarization trials, one swap trial and two test trials. The more frequent presentation of learning trials in the Daum et al. study (2012) could have increased the attentional focus to the location of the object. As was demonstrated by Paulus et al. (2011), already 9-month-olds based their anticipations of an agent’s choice for a path on the agent’s previous choices. Further, it is also possible that infants encoded both goal and path of the action simultaneously, thus it is unclear which aspect dominates infants visual anticipations. The following five experiments of Study 1 present a step by step approximation of the paradigm from Daum et al. (2012) to the paradigm of Cannon and Woodward (2012). The last experiment represents a replication as close as possible to the study by Cannon and Woodward (2012). Initially, we hypothesized that the type of agent and the presence of an occluder influence infants’ goal anticipations. Following the human first-view, we expected more goal anticipations in the experiments that contain a human agent. In contrast, the all agents-view predicts no differences in the goal anticipations between all five experiments, because human and non- human agents are processed equally from early on. We would further assume that infants show more goal-directed anticipations in the experiments without an occluder. The second study, conducted in a different lab, directly compared the two paradigms in a within- subjects design. That is, Study 2 assessed whether children and adults show systematic differences in their anticipations when observing human and non-human goal-directed actions. A second goal of Study 2 was to answer the question whether the two paradigms assess the same underlying ability regarding their goal anticipations, as proposed by the all agents-view. Since non-human agents are widely used in studies on social perception within children (e.g., Hamlin, Wynn, & Bloom, 2007; Kuhlmeier, Wynn, & Bloom, 2003), it is crucial to find out whether they actually perceive animated stimuli in the same way as human stimuli. We tested 11-month-olds as our youngest age group, since this age group showed goal-directed anticipations in the study by Cannon and Woodward (2012). We also included 32-month-olds, because developmental changes of children’s goal anticipations for an animated agent between the age of 24 and 36 months were observed by Daum et al. (2012). Developmental changes were also observed by Krogh-Jespersen et al. (2015, 2018), although they found a decrease of goal-directed anticipations for a human agent in toddlers. Given this puzzle, the inclusion of 32-month-olds in Study 2 seems informative. We additionally wanted to clarify whether and how adults differ in their perceptions of the stimulus material presented in the two paradigms. Accordingly, stimulus material was presented in a within design with a human- and a non-human animated agent. One is based on Cannon and Woodward (2012) and contained a human hand grasping one of two objects; the other is based on Paulus, Schuwerk, Sodian, and Ganglmayer (2017) and contained an animated agent walking along a path towards one of two targets (similar to Daum et al., 2012, where an opaque occluder was used). The human first-view proposes goal encoding for 11- and 32-month-olds to be more likely in the hand- than in the path-paradigm (Cannon & Woodward, 2012; Daum et al., 2012). In contrast, the all agents-view proposes goal anticipations in both paradigms. Either of the theories predicts adults to visually anticipate an action goal for both human- and non-human agents (Daum et al., 2012; Pfundmair, Zwarg, Paulus, & Rimpel, 2017). The current effort from two labs is a valuable approach with the aim to conceptually and partly even directly replicate a finding that is central in a heated debate in developmental psychology on the early origins of social cognition. It is essential to know in greater detail to which extent the findings by Cannon and Woodward (2012) are replicable before drawing strong theoretical conclusions. Thus, one central point of this endeavor was to examine whether or not we could (conceptually or directly) replicate Cannon and Woodward (2012) and contribute thus to the theoretical debate by examining the robustness of a key finding.","Study 1 investigates whether 12-month-olds are able to make goal-directed anticipations. Given the contradicting findings of Cannon and Woodward (2012) and Daum et al. (2012), the aim was to test which aspects are relevant for infants’ ability to anticipate an action goal. Therefore, the stimuli of Daum et al. (2012) were assimilated step-by-step over five experiments to the stimulus material used by Cannon and Woodward (2012). In Experiment 1 the animated stimuli of Daum et al. (2012) were used but displayed in the same presentation order as in Cannon and Woodward’s study (2012). In Experiment 2, the occluder was removed and the action direction was changed from vertical to horizontal, whereas the fish still remained as the agent. In Experiment 3, the fish was replaced by a human hand as the agent. Additional adaptions regarding timing were made in Experiment 4. Finally, Experiment 5 used newly filmed videos, which were designed to be as comparable to the stimuli of Cannon and Woodward (2012) because the original stimulus material was not available. According to the human first-view, one would hypothesize to find anticipations towards the previously observed movement path or random gaze behavior in conditions using a non-human agent (Experiment 1 and Experiment 2). In contrast, one would expect anticipations towards the previously observed goal in the conditions in which a human hand served as the agent (Experiment 3, 4 and 5). According to the all agents-view infants should demonstrate in all five experiments goal directed anticipations. We further expected that experiments without an occluder would facilitate infants processing of the action, thus we expected to find an increase of goal-directed anticipations in the conditions that did not make use of an occlusion paradigm. Experiment 1 ~~~~~~~~~~~~ In the first experiment, the influence of the design of stimuli presentation on infants’ encoding of the action was tested. Infants were shown the stimulus material as used by Daum et al. (2012), an animated fish that moved towards one of two goal objects and was briefly occluded. We combined the stimulus material with the procedure used by Cannon and Woodward (2012) where the stimuli were presented in four blocks; each block contained three familiarization trials, one swap trial, and one test trial. Further some criteria for inclusion (three of four blocks with usable data) and analysis of gaze shifts (gaze shift from the start-AOI to one of the goal-AOIs with 200 ms fixation) were the same as in the study by Cannon and Woodward (2012).","The preprocessed eye-gaze data of both studies is available at https://osf.io/bucrv/?view_only=fa9e929fe4524755b38383fd223378f5. To protect participants’ data privacy, demographic information is not shared in this data set.","Participants Participants Participants Participants The sample included 24 healthy 12-month-olds (12 girls, mean age = 12 months and 4 days; 11;21–12;15). Ten additional infants had to be excluded due to inattention and restlessness (n = 1), crying (n = 4), technical problems (n = 1) and failure to provide enough eye-tracking data (n = 4; see","Measures section for details).","Participants were presented videos of a red-blue fish, that moved by itself (self-propelledness) on a blue background. At the beginning of the videos, the fish was situated at the bottom in the middle of the screen. The targets were a yellow duck and a colored ball placed on the left and right corner at the top of the screen (see Fig. 1). In the middle of the screen was a round occluder in the color of wooden grain. Infants saw four blocks and each block consisted of three familiarization trials, one swap trial and one test trial. Before each block an attention getter was presented to direct infants’ attention to the screen. In the familiarization trial (total duration was 15.12 s) the fish first jumped up and down (accompanied by a sound) for 3 s and then moved towards the occluder (2.44 s). The agent disappeared behind the occluder for 0.92 s and reappeared to aim for one target. At the goal object (after 2.08 s) the fish poked the target for three times (3 s) and the target reacted with small movements, which was combined with a sound. During the swap trial, the two targets changed place (4.96 s) and were shown for another second after the changeover to the infant. In the test trials, the fish again jumped up and down (3 s) before approaching the occluder (2.44 s). The agent stayed behind the occluder for the rest of the trial (another 10.12 s). The total duration of the whole presentation (all four blocks) was 5 min and 24 s. Target object as well as the position of the target object was counterbalanced between participants. Setting and procedure For testing, infants were seated in a car safety seat (Maxi Cosi Cabrio) with a distance of 60 cm between the eye tracker and the child and stimuli were presented on a 17”-monitor (25° × 21°). Gaze was measured through a Tobii 1750 eyetracker (precision: 1°, accuracy: 0.5° and sampling rate: 50 Hz) and a nine-point infant calibration was used. The stimuli-presentation was conducted via the software ClearView (version 2.7.1., Tobii). Measures To analyze infants’ eye-movements, three areas of interest were defined based on Daum et al. (2012, see Fig. 1). The lower area was the starting area of the agent, the other two included the two goal objects. For all measures, we analyzed the first fixation participants performed from the start area to one of the goal-AOIs (first look analysis). Infants had to fixate the goal-AOI for 200 ms, within a radius of 50 pixels (based on Cannon & Woodward, 2012). Gaze shifts were categorized as anticipatory, when the first fixation was directed in one of the goal-AOIs before the agent reappeared from the occluder during familiarization trials (occlusion-time plus 200 ms). In the test trials, this time interval was extended for another 1000 ms (see Cannon & Woodward, 2012), because the agent did not reappear from behind the occluder (occlusion-time plus 200 ms plus 1000 ms). Gaze shifts were categorized as reactive, when fixations to one of the goal-AOIs occurred after the agent reappeared in the familiarization trials. For the test trials, a fixation to one of the goal AOIs after the 2120 ms was categorized as reactive. First, we calculated the anticipation rate, which is the relation of all anticipations to all gaze shifts (anticipatory as well as reactive). This measure indicates how much participants generally perform anticipations; it is not including information to which specific location infants anticipated. To further analyze the type of anticipations, the accuracy rate was calculated. For this, the number of anticipations towards the specific target AOI was divided by the total number of gaze shifts (anticipatory and reactive). For this analysis, the number of anticipations was averaged over all familiarization trials and test trials. For familiarization trials the goal-related accuracy rate (ratio of goal anticipations and all gaze shifts) and the non-goal-related accuracy rate (ratio of anticipations to the other object and all gaze shifts) were defined. To analyze infants’ learning performance in the familiarization phase, two scores were compared (Daum et al., 2012): The accuracy score of the averaged anticipations of the first familiarization trials of all four blocks and the accuracy score of the averaged anticipations of the last familiarization trials of all four blocks. For test trials the identity-related accuracy rate (ratio of anticipations to the goal object in the new location and all gaze shifts) and the location-related accuracy rate (ratio of anticipations to the other object in the old location and all gaze shifts) were generated. Again, the anticipations were averaged over all four test trials for each score. In sum, for both the familiarization as well as the test phase, each accuracy score consists of four trials. To be included in analysis, infants had to watch the screen at least 200 ms from the start of the movie until the agent disappeared behind the occluder; and 200 ms after disappearance until the end of the movie. Infants were included for final analysis if they had at least three of four test trials that fulfilled these criteria (Cannon & Woodward, 2012). Further they had to look at the swap trial for at least 2000 ms to be included. In all experiments of Study 1, we controlled for the possible influence of the type of target and position of target on the number of anticipations for the first four familiarization trials and test trials. As no significant influence could be found in none of the five experiments, the following analysis was averaged over these factors. Further, the anticipation rate of the first four and last four familiarization trials, as well as test trials, were averaged across the four blocks. Anticipation rate The anticipation rate for the whole experiment was 0.79 (SD = 0.15) and 0.74 (SD = 0.20) for the familiarization phase only. We compared the anticipation rate of the last familiarization trials with anticipation rate of the test trials with a Wilcoxon signed-rank test and found a significant difference. Participants anticipated more in the test trials (M = 0.91, SD = 0.17) than in the last familiarization trials (M = 0.72, SD = 0.30), with z = −2.92, p = .003, r = .42. Familiarization phase A Wilcoxon signed-rank test was calculated to compare the goal-related accuracy rate with the non-goal-related accuracy rate in the last familiarization trials. Indeed, children anticipated more to the goal (M = 0.58, SD = 0.34) than to the non-goal (M = 0.14, SD = 0.25), with z = −3.25, p = .001, r = −.47. This indicates that the children learned to correctly anticipate the reappearance of the agent from behind the occluder during the familiarization phase. A comparison of the first familiarization trials averaged across the four blocks (M = 0.51, SD = 0.34) and the last familiarization trials did not show a significant increase of goal-directed anticipations over time, z = −1.41, p = .16, r = −.20. Test phase The Wilcoxon signed-rank test demonstrated a higher location-related accuracy rate (M = 0.60, SD = 0.32) than an identity-related accuracy rate (M = 0.31, SD = 0.31), z = −2.14, p = .03, r = −.31. Further, a comparison of the goal- related accuracy rate of the last familiarization trials with the identity- related accuracy rate in the test trials showed a significant difference, z = −2.56, p = .01, r = −.37. In contrast, there was no significant difference between the goal-related accuracy rate of the last familiarization trials and the location-related accuracy rate of the test phase, z = −0.34, p = .73, r = −.05. For the following analysis, only anticipations (and not reactions) were used. A Chi-Square-Test was calculated over the number of identity- and location-related anticipations in test trials. Infants anticipated the reappearance of the agent based on location (n = 51) than on identity of the goal object (n = 28), χ²(1) = 6.70, p = .01. Additional analysis Finally, when interpreting these findings, and comparing them to the original study, one has to consider that the data was differently analyzed than the original study of Cannon and Woodward (2012). The inclusion criteria used are stricter than in the original study by Cannon and Woodward (2012). For example, infants had to look for a specific time at the swap trial or fixate the start area for a certain time before the agent moved behind the occluder, etc. Also, gaze shifts that occurred after a certain time were no longer defined as anticipatory, but as reactive. The resulting scores were calculated different to the original study, which used a proportion score and did not include non-anticipations. Although our use of stricter criteria should result in a more reliable assessment of true goal anticipation, one could argue that the different results are caused by these stricter criteria. To exclude this possibility, we additionally analyzed our data as closely as possible to the approach by Cannon and Woodward (2012; details can be seen in the supplementary material). This additional analysis did not change the pattern of results; the mean proportion score of 0.34 (SD = 0.30) was significantly different from chance with t(23) = −2.64, p = .015, Cohen’s d = 0.54, indicating a significant looking bias towards the location and not the goal.","Discussion Discussion Discussion Discussion Experiment 1 aimed to examine whether the different findings reported in Cannon and Woodward (2012) and Daum et al. (2012) are the result of differences in the procedure of the stimulus presentation. The findings show that 12-month-olds learned to correctly anticipate the reappearance of the agent by the end of the familiarization phase. However, in the test trials, infants anticipated the action based on the location of the goal object and not on its identity. Therefore, it doesn’t seem that the more frequent presentation of the action in Daum et al.’s study (2012) highlighted the path of the action and caused children’s location- related anticipations. Ultimately, the divergent findings are not caused by the different presentation order and amount of learning and test trials. The next experiment will test whether the occluder has a significant effect on infants’ anticipations. Experiment 2 ~~~~~~~~~~~~ In Experiment 2, the animated stimuli were used without an occluder; the agent was visible the whole time. Additionally, the direction of the movement was changed from a vertical to a horizontally movement (as in Cannon & Woodward, 2012). Given the claim that horizontal movements are easier to anticipate for infants (Gredebäck, von Hofsten, & Boudreau, 2002), we intended to facilitate anticipations and to test whether a change of the movement direction increases infants’ identity-related anticipations. Also to draw infants’ attention to the screen at the beginning of the action, an ostensive cue was integrated (a voice stated “Look”). Participants Again, the sample included 24 healthy 12-month-olds (12 girls, mean age = 12 months and 4 days, 11;20–12;10). Nine additional infants had to be excluded due to inattention and restlessness (n = 2), crying (n = 1) or failure to provide enough eye-tracking data (n = 6).","Stimuli and procedure Stimuli and procedure The experimental setup was exactly the same as in Experiment 1, only that the targets (duck 3.9° × 4.1°, penguin 3.5° × 4.0°) were now situated in the two corners at the right side of the monitor (see Fig. 2). Further the starting point of the agent (1.9° × 4.0°) was close to the left monitor side. No occluder was visible and the background was light blue. Stimuli where similar to Experiment 1 with the exception that in the first two seconds of each movie a voice stated “Look”, to catch infants’ attention. In the familiarization trial the agent jumped up and down (1.96 s, combined with a whistle sound) and started to move towards the target objects. After 4.72 s the agent took a turn to one of the targets, which he reached after another 2.60 s and poked it (see Experiment 1). The whole familiarization trial lasted for 14.92 s. In the swap trial the two objects changed position within 5.96 s and the whole trial lasted for another 5 s. Test trials were similar to Experiment 1, except that the agent, when approaching the target, stopped after 4.72 s just after the middle of the scene and remained there for the rest of the trial (another 4.96 s). The whole presentation time for the movies was 5 min and 2 s. Again goal object and position of goal object were counterbalanced between participants. Measures For Experiment 2 the calculation of the measures followed Experiment 1. Areas of interest can be seen in Fig. 2. Gaze shifts were defined as anticipatory, if anticipations took place from the beginning of the movement towards the objects until the agent made a turn to one of the targets (for test trials: time of standstill of the agent plus 1000 ms). This results in an anticipatory period of 5720 ms for test trials in total. All gaze shifts that occurred after the turn of the agent were coded as reactive (in test trial after the additional 1000 ms). Anticipation rate For this experiment the anticipation rate was 0.81 (SD = 0.14) for the whole experiment, and 0.76 (SD = 0.19) for the familiarization phase. As in Experiment 1, we found an increase in anticipations in the test trials (M = 0.94, SD = 0.13) compared to the last familiarization trials of all four blocks (M = 0.74, SD = 0.30), with z = −2.62, p = .01, r = −.38. Familiarization phase Infants showed in the last familiarization trials more anticipations to the goal (M = 0.60, SD = 0.35) than to the non-goal (M = 0.14, SD = 0.21), z = −3.43, p = .001, r = −.50. The goal-related accuracy rate in the first familiarization trials (M = 0.52, SD = 0.33) was not significantly different from the goal-related accuracy rate in the last familiarization trials (M = 0.60, SD = 0.35), z = −0.85, p = .40, r = −.12. The infants quickly learned to correctly anticipate the target. Test phase The analysis showed no significant difference between the identity-related accuracy rate (M = 0.41, SD = 0.35) and the location-related accuracy rate (M = 0.54, SD = 0.33), z = −1.12, p = .26, r = −.16. Further, a comparison between the goal-related accuracy rate in the last familiarization trials (M = 0.60, SD = 0.35) with the identity-related accuracy rate (M = 0.41, SD = 0.35), z = −1.58, p = .12, r = −.23, and the location-related accuracy rate (M = 0.54, SD = 0.33), z = −0.91, p = .36, r = −.13, was not significant. Infants anticipated in the test phase towards the goal as well as to the original location of the object. Following Experiment 1, only anticipations were analyzed. The Chi-Square test over the four test trials between identity- (n = 37) and location-related gaze-shifts (n = 47) was not significant, χ²(1) = 1.19, p = .28. The additional analysis, with the same measure and inclusion criteria as Cannon and Woodward (2012, see Experiment 1 for details) revealed chance performance of the infants’ looking behavior, with t(23) = −0.72, p = .482, Cohen’s d = 0.15, M = 0.45, SD = 0.33. Discussion Experiment 2 investigated whether the absence of an occluder and the change in movement direction (from vertical to horizontal) had a facilitating effect on infants’ goal encoding abilities. Results of the test trials demonstrated that anticipations of the 12-month-olds were at chance level. They neither showed a significant looking bias towards the goal object, nor to the old location. As they demonstrated goal-directed anticipations at the end of the familiarization phase, problems in learning the association between the agent and the target are not the case. It seems likely that the absence of the occluder facilitated infants’ goal anticipations (Hespos et al., 2009; Jonsson & von Hofsten, 2003). It seems easier for infants to encode the goal of an action, when the agent is visible for the whole time. The additional change from a vertical to a horizontal position and the use of a verbal cue at the beginning of the action could have facilitated the task as well. For the next experiment we wanted to see whether the type of agent influences infants’ anticipations. Experiment 3 ~~~~~~~~~~~~ To test whether infants are more likely to flexibly attribute goal-directed behavior to a human agent, the animated fish was replaced by a human hand that moved to and grasped one of two goal objects. Timing and procedure of the action, as well as size of the targets differed from Cannon and Woodward (2012), as they remained the same as in Experiment 1 and 2. Participants The final sample included 32 healthy 12-month-olds (16 girls, mean age = 12 months and 1 day, 11;15–12;15). The sample size is larger for this and the following experiments of Study 2, due to more conditions than in the previous experiments (see section Stimuli and Procedure). Eleven additional infants were tested but excluded due to inattention and restlessness (n = 4), crying (n = 2), or not enough eye-gaze data (n = 5). Stimuli and procedure Stimuli in Experiment 3 are the same as in Experiment 2, except for the following differences: Instead of an animated fish, a human hand (4.6° × 7.6°) was filmed. The two targets (duck 4.8° × 4.9°, penguin 4.5° × 4.8°) still have been animated with CINEMA 4D (Maxon, Version R10). The human hand was inserted in the video via a blue screen method (see Fig. 3). For the familiarization trials the human hand started to move its fingers and a whistle sound occurred (1.92 s). Then the hand moved from the left side in the direction of the targets. After 4000 ms the hand crossed the middle and made a turn to one of the targets. The hand reached the target after 2 s, grasped it (1.12 s) and moved it further to the right (action effect plus sound lasted for 0.92 s). The whole familiarization trial took 14 s. The test trials started exactly like the familiarization trials. After the hand started to move towards the middle, it stopped after 4 s for another 5 s. Duration of the test trial was 14.12 s in total. The whole presentation lasted 4 min and 48 s. Type of goal object and position of the object were counterbalanced between participants. Further, the orientation of the hand, that is, thumb pointing towards the goal object, as well as left and right hand were counterbalanced within participants. This led to four different combinations for the blocks, namely 1) familiarization: right hand, test trial: right hand, 2) familiarization: right hand, test trial: left hand, 3) familiarization: left hand, test trial: left hand, 4) familiarization: left hand, test trial: right hand. The order of combination was also counterbalanced. For analysis the same scores were calculated as in Experiment 1 and 2. Gaze shifts for test trials were defined as anticipatory if they occurred within a time period of 5000 ms (from the beginning of the reaching action onwards). Also inclusion criteria remained the same. AOIs were identical to Experiment 2. To control for an influence of hand orientation on the scores, Kruskal–Wallis tests were performed for Experiment 3, 4 and 5 and turned out not significant. Therefore, the following analysis was averaged over this factor. Anticipation rate For the whole experiment, the anticipation rate was 0.91 (SD = 0.10) and for the familiarization phase 0.89 (SD = 0.12). Although participants performed more anticipations in the test trials (M = 0.95, SD = 0.11) than in the last familiarization trials of all four blocks (M = 0.88, SD = 0.19), the Wilcoxon- signed rank test was not significant for this experiment, with z = −1.86, p = .06, r = −.23. Familiarization phase The Wilcoxon signed-rank test showed that the goal-related accuracy rate in the last familiarization trials (M = 0.60, SD = 0.34) was significantly higher than the non-goal related accuracy rate of the last familiarization trials (M = 0.28, SD = 0.31), z = −2.58, p = .01, r = −.32. There was no significant difference between the goal-related accuracy rate of the first familiarization trials (M = 0.53, SD = 0.27) and the goal-related accuracy rate of the last familiarization trials (M = 0.60, SD = 0.34), z = −1.21, p = .23, r = −.15. Test phase There was no significant difference between infants’ identity-related (M = 0.40, SD = 0.31) and location-related (M = 0.55, SD = 0.28) accuracy rate, z = −1.41, p = .16, r = −.18. Infants anticipated more to the goal in the familiarization phase (M = 0.60, SD = 0.34) than to the same goal in the test trials (M = 0.40, SD = 0.31), z = −2.26, p = .02, r = −.28. In contrast they did not differ in their anticipations to the goal in the last familiarization phase (M = 0.60, SD = 0.34) and their anticipations to the old location in the test phase (M = 0.55, SD = 0.28), z = −0.74, p = .46, r = −.09. A further Chi- Square-Test with the number of anticipations in all four test trials between goal- and identity-related anticipations turned out to be not significant, χ²(1) = 3.64, p = .057, although a tendency towards more location-related (n = 65) than identity-related anticipations (n = 45) could be observed. Again, also the additional analysis according to Cannon and Woodward (2012) resulted in chance performance of the infants, with t(31) = −1.53, p = .137, Cohen’s d = 0.27, M = 0.43, SD = 0.27. Discussion Experiment 3 examined whether the type of agent has an influence on infants’ goal anticipations. Therefore, a human hand reaching for one of two objects was presented to 12-month-olds. Results showed that infants did not increase their goal-related anticipations in this experiment. In contrary, they showed the tendency to anticipate the grasping action based on the movement path and not the goal object. Theoretically this finding speaks against the claim that experience with an action improves infants’ anticipations (Cannon & Woodward, 2012; Gredebäck & Melinder, 2010; Kanakogi & Itakura, 2011; Kochukhova & Gredebäck, 2010). Still, it is not clear which factors caused the differences between our experiments with Daum et al’s findings (2012) and Cannon and Woodward’s (2012). Some differences (timing of the action, start of the presentation, effect of the action, type of goal objects) have remained in the last three experiments. For example, in Experiment 1, 2 and 3, the animated objects visibly swap places in the swap trial and are not presented already in their new position like in Cannon and Woodward’s (2012) stimuli. It is not clear in what way this could have affected their anticipations. Hence changes in the stimuli material have been made on these factors for the next experiment. Experiment 4 ~~~~~~~~~~~~ Because in Experiment 2 and 3, the anticipations were ambivalent, this could be an indication for an increase of identity-related gaze shifts caused by the nature of the agent. To further assess this potential shift of processing caused by the stimulus material, the stimuli have been changed further in Experiment 4 to increase similarity to the stimuli of Cannon and Woodward (2012). Therefore, the timing of the action and the action effect were adapted. Additionally, movements of the hand at the beginning of the trials were removed and the swapping procedure of the target objects was no longer presented. Participants only observed a still frame of the already swapped objects. The rest of the factors, such as blocked design, no occluder, direction of movement, and type of agent stayed the same as in Experiment 3. Differences to Cannon and Woodward (2012) remained regarding the target objects (a duck and a penguin here, a green frog and a red ball in Cannon & Woodward, 2012). Also following Experiment 3, the agent was inserted via a blue screen method and targets were still animated. Further, type and size of targets were different than in the original study. Participants The sample included 32 healthy 12-month-olds (16 girls, mean age = 11 months and 26 days, 11;12-12;11). Thirteen additional infants were tested but excluded due to inattentiveness and restlessness (n = 1), crying (n = 1), not enough eye-tracking data (n = 10) or technical problems (n = 1). Stimuli and procedure Movies were similar to Experiment 3 and adapted to the videos of Cannon and Woodward (2012). Each movie started with a black still frame for 0.36 s. After 0.04 s the agent moved into the picture. After another 1.56 s the hand made a turn into the direction of a target. The agent reached the target after 1.04 s. He grasped the object and did not move thereafter (no action effect, only sound; 0.4 s). After 0.52 s the black screen was presented for 0.48 s. One familiarization trial lasted 4.36 s. For the swap trial a still frame with the objects in changed position was presented for 3.56 s. For test trials the agent again moved into the picture after 0.04 s. Then the hand moved in the direction of the targets and stopped after 1.52 s right after the middle of the screen. Movies of test trials lasted each 2.88 s. Presentation time of the whole stimuli material over all four blocks was 1 min and 33 s. Goal object, position of goal object, orientation of hand, and the order of presentation was again counterbalanced across the infants. The analysis remained the same as in the previous experiments. However, since the test trial is very short in this experiment, the criteria to treat gaze shifts after the standstill of the hand plus another 1000 ms as reactive is no longer applicable. Therefore, all gaze shifts were treated as anticipatory. Inclusion criteria were identical to the other three experiments, except that infants had to watch each movie at least for 100 ms from the start until the turn and again from the turn until the end of the movie. Also they had to look at the swap trial at least for 700 ms. Anticipation rate The anticipation rate for the whole experiment was 0.74 (SD = 0.18) and 0.65 (SD = 0.24) for all of the 12 familiarization trials. Again, the anticipation rate of the test trials was higher (M = 1.00, SD = 0.00) than in the last familiarization trials of all four blocks (M = 0.61, SD = 0.30), z = −4.32, p < .001, r = −.54. Familiarization phase Goal-related accuracy rate of the last familiarization trials (M = 0.49, SD = 0.29) was significantly higher than the non-goal-related accuracy rate of the last familiarization trials (M = 0.12, SD = 0.19), z = −3.98, p < .001, r = −.50. Further the goal-related accuracy rate of the first familiarization trials (M = 0.47, SD = 0.37) was not significantly different from the goal- related accuracy rate of the last familiarization trials (M = 0.49, SD = 0.29), z = −0.53, p = .60, r = −.07. Test phase Results revealed that the identity-related accuracy rate (M = 0.30, SD = 0.29) was significantly lower than the location-related accuracy rate (M = 0.70, SD = 0.29) in the test phase, z = −3.23, p = .001, r = −.40. There was also a significant difference between the goal-related accuracy rate in the last familiarization trials (M = 0.49, SD = 0.29) and the identity-related accuracy rate in the test phase (M = 0.30, SD = 0.29), z = −2.14, p = .03, r = −.27. Infants showed more anticipations towards the goal in the familiarization phase than in the test phase. Moreover, another Wilcoxon signed-rank test showed that the location-related accuracy rate in the test phase (M = 0.70, SD = 0.29) was significantly higher than the goal-related accuracy rate in the last familiarization trials (M = 0.49, SD = 0.29), z = −3.04, p = .002, r = −.38. Again, when only the number of anticipations was included, a Chi-Square test demonstrated that infants anticipated the action more in relation to the location (n = 83) than to the identity of the goal (n = 36), χ²(1) = 18.56, p < .001. The same pattern was observed in the additional analysis according to Cannon and Woodward (2012), in which infants demonstrated a preference for the previous location, with t(31) = −4.05, p < .001, Cohen’s d = 0.72, M = 0.29, SD = 0.29. Discussion In the fourth experiment we wanted to discover, whether the changes of the stimuli in comparison to Experiment 3 would now enable a replication of Cannon and Woodward’s results (2012). Findings indicate that 12-month-olds show more anticipations directed towards the location and not the goal object. Even after using stimuli that are highly similar to Cannon and Woodward (2012), we could not replicate their findings. However, our stimuli still contained subtle differences regarding the targets and construction of the stimuli, such as that a human hand was overlaid on the animated background with a blue screen method. Hence, we performed a last experiment that contains a replication as close as possible to the Cannon and Woodward (2012) study. Experiment 5 ~~~~~~~~~~~~ To replicate the results of Cannon and Woodward (2012), the stimuli were newly filmed. Direction, timing, and procedure of the action were as close as possible to the original study. The two targets, now also a red ball and a green frog, were at the same position of the screen and had the same size. There should be no to only few differences in the method used between this experiment and Cannon and Woodward (2012). Participants Thirty-two healthy 12-month-olds were included for the final sample (16 girls, mean-age = 11 months; 27 days, 11;16-12;13). Additionally, 22 infants were tested but excluded due to inattentiveness and restlessness (n = 7), crying (n = 1) or not enough eye-gaze data (n = 14). Stimuli, procedure and analysis Stimuli were identical to Experiment 4 except that the goal objects were now a green frog (3.4° × 4.4°) and a red ball (4.3° × 4.1°; see Fig. 4). For the action, the human hand (4.9° × 9.4°) reached from the left to the right side of the screen. No animations were used anymore. Material was edited and cut with Final Cut Pro (Version 7.0.3). To ensure that the action was timed accurately, one movie of Cannon and Woodward (2012) was used as a basis for cutting the stimuli. As already Experiment 4 followed the timing of Cannon and Woodward (2012), there were no other changes made (see Experiment 4 for details). Goal-object, position of goal-object, hand orientation and order of hand orientation was counterbalanced throughout participants, leading to 16 different combinations. Analysis of the data and inclusion criteria were carried out as in Experiment 4. Anticipation rate For the whole experiment, the anticipation rate amounts to 0.74 (SD = 0.14). For the familiarization trials the anticipation rate is 0.65 (SD = 0.18). A Wilcoxon signed-rank test revealed a significant difference between the anticipation rate of the test trials (M = 1.00, SD = 0.00) and the last familiarization trials (M = 0.56, SD = 0.31), with z = −4.40, p < .001, r = −.56. Familiarization phase Goal-related accuracy rate in the last familiarization trials (M = 0.39, SD = 0.31) was significantly higher than the non-goal-related accuracy rate of the last familiarization trials (M = 0.18, SD = 0.27), z = −2.03, p = .04, r = −.25. There was no significant difference between the goal-related accuracy rate in the first familiarization trials (M = 0.50, SD = 0.30) and the goal- related accuracy rate in the last familiarization trials (M = 0.39, SD = 0.31), z = −1.51, p = .13, r = −.19. Test phase A Wilcoxon signed-rank test demonstrated no significant difference between the identity-related (M = 0.40, SD = 0.35) and location-related (M = 0.60, SD = 0.35) accuracy rate in the test phase, z = −1.53, p = .13, r = −.19. Also the goal-related accuracy rate in the last familiarization trials (M = 0.39, SD = 0.31) did not differ from the identity-related accuracy rate in the test phase (M = 0.60, SD = 0.35), z = −2.62, p = .009, r = −.23. Nevertheless, the goal- related anticipation rate in the last familiarization trials (M = 0.39, SD = 0.31) was significantly lower than the location-related accuracy rate in the test trials (M = 0.60, SD = 0.35), z = −2.62, p = .009, r = −.23. A Chi-Square test over the number of anticipations of the test trials showed a significant difference, χ²(1) = 4.40, p = 0.036. Infants anticipated more often to the location (n = 66) than to the identity of the goal (n = 44). The analysis according to Cannon and Woodward (2012) with the same measure and inclusion criteria demonstrated chance level of infants’ looking behavior, with t(30) = −0.92, p = .363, Cohen’s d = 0.17, M = 0.44, SD = 0.39. Discussion While the previous experiments 1–4 represent conceptual replications of the study by Cannon and Woodward (2012), Experiment 5 represents a direct replication. Because the original stimuli were not available, the stimuli were newly filmed for Experiment 5 and video edited to make them as similar to the original stimuli as possible. Nevertheless, the current findings show that 12-month-olds demonstrated more anticipations to the location than to the goal. The results of this direct replication are in line with the findings of Experiment 1–4 and of Daum et al. (2012). The differing findings between the two paradigms can neither be explained by the presence of an occluder nor by the type of agent. Detailed implications of these findings are discussed in the General Discussion. Discussion Study 1 ~~~~~~~~~~~~~~~~~~ The aim of Study 1 was to examine which factors (type of agent, occlusion, order of trials, timing and procedure of the action, movement direction) effect infants’ goal encoding and caused the contradictory findings of two previous studies focusing on the same research question (Cannon & Woodward, 2012; Daum et al., 2012). Based on previous findings (Cannon & Woodward, 2012; Hespos et al., 2009; Kanakogi & Itakura, 2011), the hypothesis of Study 1 was that the different results could primarily be attributed to two factors, namely type of agent (human vs. non-human) and differences in the requirement of cognitive resources caused by the use of an occlusion paradigm. For this purpose, these differences were consecutively aligned in Study 1. In 5 experiments, an agent was presented to 12-month-old infants who moved to one of two goals. Before the test phase, positions of the target objects were swapped. Results over all 5 experiments showed that 12-month-olds learned the association of the agent and the target during familiarization phase already after three trials, thus demonstrating fast learning within infants (Krogh- Jespersen & Woodward, 2014). For test trials, infants showed in four of the five experiments more location- than identity-related anticipations. In one experiment, infants anticipated equally often towards the identity and the location of the target. Over all five experiments, infants demonstrated 311 location-related and 189 identity-related anticipations. The hypothesis that the different results of the two studies (Cannon & Woodward, 2012; Daum et al., 2012) could be attributed to the type of agent and the occluder was not confirmed. Further, our results are neither in line with the human first- nor with the all agents-view, since we did not find goal anticipations in any of the experiments. This finding is not in line with previous studies, which highlight the role of experience (human first-view) for understanding an action as goal-directed in the first year of life (Cannon & Woodward, 2012; Gredebäck & Melinder, 2010; Kanakogi & Itakura, 2011; Krogh-Jespersen & Woodward, 2018). They are in line with findings demonstrating that infants anticipate others’ actions based on their previous path (Paulus et al., 2011). However, results are not in line with looking time studies that demonstrated understanding of goal-directed actions of non-human agents already within infants, as stated by the all agents-view (e.g. Kamewari et al., 2005; Luo, 2011; Luo & Baillargeon, 2005; Schlottmann & Ray, 2010). Nonetheless, it remains an open question, why infants expected the agent to move to the familiarized location and not to the previous goal. The assumption, that the occluder could have an influence on infants’ anticipations, was not confirmed either in Study 1. Infants showed location-related anticipations, independent of the presence of an occluder. On the one hand, studies have suggested that the use of an occluder requires extra skills for infants (Hespos et al., 2009; Jonsson & von Hofsten, 2003). On the other hand, previous findings showed that infants in their first year of life are able to anticipate the reappearance of a temporarily occluded object (Gredebäck & von Hofsten, 2004; Gredebäck et al., 2002; Rosander & von Hofsten, 2004). Our results are in line with the latter set of findings. In sum, the current findings suggest that at the age of 12 months, the ability to flexibly anticipate the actions of others based on goal identity has not yet developed, irrespective of whether the agent was human or not, and therefore, infants have also shown anticipations based on the location. In Study 2 this issue is further addressed in another lab and with different stimuli. The inclusion of older age groups (32-month-olds and adults) in Study 2 will also reveal how stable this effect is over the course of development.","The second study analyzed whether children and adults show goal encoding for two different paradigms, one using a human hand as an agent (e.g., Cannon & Woodward, 2012) and one a non-human animated animal (e.g., Daum et al., 2012; Paulus et al., 2011). We intended to find out whether infants use similar processing strategies for both paradigms. Thus, both tasks were presented to 11-month-olds, 32-month-olds and adults in a within-subject design. In both paradigms, the agent walked to one of two objects for several times. For test trials the objects’ position was swapped and participants observed an uncompleted action. The hand-paradigm is based on the study by Cannon and Woodward (2012), while the stimuli differed in a few manners. Most notably, the number of trials and the presentation order of learning and test trials were different. Further details are described below in the method section. Moreover, to facilitate infants’ encoding of the scenario, we also decided to familiarize participants with the actor and the targets. Goal-directed anticipations from 11 months onward in the hand-paradigm would replicate the findings of Cannon and Woodward (2012). In the path-paradigm, we expected 11-month-olds and 32-month- olds to anticipate the old path, thus the novel goal, as the action is performed by a non- human agent (Daum et al., 2012). We hypothesized to find anticipations towards the familiarized goal within the adult sample in both paradigms (Daum et al., 2012; Pfundmair et al., 2017). A correlational analysis of the anticipatory looking behavior between the two paradigms should further clarify whether the different processing mechanisms assessed by the two paradigms are related to each other. We expected to find at least in the adult sample a positive correlation, because they showed goal anticipations for human and non- human agents in previous studies (Daum et al., 2012; Pfundmair et al., 2017). In case we would find a correlation between the two paradigms within children, as predicted by the all agents-view, we wanted to make sure that this relation is not mediated by individual factors of the child (Licata et al., 2014). Hence, we included a measure for temperament (Infant Behavior Questionnaire revised very short form; Putnam, Helbig, Gartstein, Rothbart, & Leerkes, 2014; Early Childhood Behavior Questionnaire very short form; Putnam & Rothbart, 2006) as well as a cognitive measure for pattern recognition. Therefore, items were adapted from the Bayley-Scales (Bayley, 2006; German version by Reuner & Rosenkranz, 2014; see supplemental material) for the 32-month-olds. As there are, to our knowledge, no comparable items for 11-month-olds, a task for working memory (Pelphrey et al., 2004; Reznick, Morrow, Goldman, & Snyder, 2004) was used as an indicator for pattern recognition. In this task infants had to remember in several trials at which of two windows the experimenter appeared beforehand.","The final sample included 34 11-month-olds (mean age = 11.44 months; SE = 0.15; 15 girls), 35 32-month-olds (mean age = 32.11 months; SE = 0.09; 23 girls) and 35 adults (mean age = 23.03 years; SE = 1.04; 30 women). Additionally, 8 children were tested but excluded, as they did not want to watch the second movie (n = 4), were inattentive (n = 2), or because of measurement failure (n = 2). Three additional adults had to be excluded due to technical problems. Participants came from a larger city in Germany. The infant population was recruited from local birth records and adult participants were recruited from a student population. Participants or their caregivers gave informed written consent. The study was approved by the local ethics board. Stimuli and procedure Participants were presented with two different paradigms. Both paradigms were shown on a 23-inch monitor, which was attached to a Tobii TX300 corneal reflection eye tracker with a sampling rate of 120 Hz (Tobii Technology, Sweden). Children were either seated on their parent’s lap or in a car-safety seat (Chicco) about 60 cm away from the screen. For 11- and 32-month-olds, a 5-point calibration was performed. Adults were calibrated with a 9-point procedure. Data collection and analyzation was carried out with Tobii Studio (Tobii Technology, Sweden). All movies had the size of 1920 × 1080 pixels. In the subsequent section, the stimuli for the hand-paradigm are described first, the animated path-paradigm second. Hand-paradigm The procedure started with a movie, which introduced the agent. The female actor was sitting on a table and waving at the participant (6 s). To familiarize participants with the targets, two pictures were shown (each 2 s). Each showed one of the two targets, a green ball and a blue cube, accompanied by a sound. Next, the learning trials started. Similar to Cannon and Woodward (2012), the movies of the learning trials presented the two targets at the right side of the screen. The ball was situated in the upper position, the cube in the lower position (see Fig. 5). The table was light brown. After 0.10 s, the hand reached into the picture (supplemented by a subtle, short sound) from the left side of the screen towards the targets, until just past midline (2.65 s). It then made a curvilinear path towards the ball and grasped it (after another 0.85 s) for 0.40 s combined with a squeak-sound. The whole learning trial lasted for circa 4 s, and was shown for 5 times in a row. Next, participants were presented with the swap trial, which consisted of a picture of the objects in swapped position accompanied by a rattle sound (4 s). Finally, in the test trials the hand reached in from the left (with the same sound as in the familiarization trials) and stopped just past midline (2.80 s). A still-image of the hand in this position was presented for another 6.10 s. A test trial lasted for 8.90 s and was presented for three times in a row. Between each test trial a black screen with an attention-getting sound was presented, to redirect infants’ attention to the screen. Path-paradigm First, participants were familiarized with the setup using an introductory movie. It showed a horizontal path that led from the right to the left side of the screen. A rabbit was sitting at the right side of the path and a transparent occluder was located in the middle of the path. After 0.2 s the rabbit started to jump up and down, supplemented by a sound. At the same time the occluder turned opaque and the rabbit started to move towards the end of the path through the occluder (6.88 s), turned around and went back towards the starting point. The whole movie lasted for 12.54 s. Afterwards five learning trials were presented. The learning trials contained a path that was leading to two different goals (similar to Paulus et al., 2017); a house that was situated on the upper path and a wood, situated on the lower path (see Fig. 5). The occluder was overlaid at the crossroad where the path divided into two options leading to the different goals. The agent was a pig that was located at the left side of the path. At the beginning of the movie the transparent occluder turned opaque (0.44 s). Afterwards the pig started to jump for two times (accompanied by a sound) and moved towards the occluder, until it disappeared for 2.37 s. The pig then reappeared on the upper path and moved towards the house. When it reached the house a bell sound was played. The whole movie took 10.52 s. Following the learning trials, a swap trial was shown (total duration 10.03 s). First a frame of the objects in the old position was presented; after 3 s both targets disappeared with a sound and reappeared in changed position after another 2 s with the same sound. Targets in changed position were presented for 5 more seconds. The pig was situated at the beginning of the path during the whole trial. The following test trials started completely identically to the learning trials, except that the targets were now in changed position and the pig did not reappear from the occluder for 6 s. The duration of one test trial was 11.52 s and test trials were presented three times in a row. Between the three test trials the same black screen with the attention-getting sound was inserted, as was done in the hand- paradigm. Both paradigms were shown to each participant. Order of paradigms was counterbalanced. Between the two eye-tracking movies 11-month-olds and 32-month-olds did a task for a control measure on the table. A detailed description of control measures can be found in the files of the supplemental material.","The Tobii standard fixation filter with a velocity threshold of 35 pixels/window and a distance threshold of 35 pixels was used to define fixations. For the hand- paradigm three AOIs were defined: One AOI covered the area of the hand (“starting area”, 31.94%), the other two AOIs covered the targets (each 13.67%). Participants gaze behavior was measured for the whole test trial. AOIs for the path-paradigm were implemented as followed: The “start-AOI” covered the starting point of the action, namely the first part of the path, which led to the occluder (5.46%, see Fig. 5). The other two AOIs were situated on the area where the paths reappeared from the occluder (each 14.53%). The “start-AOI” was active for 1.79 s before the agent disappeared behind the occluder. Once the agent disappeared, fixations to the other AOIs were measured. For analysis of participant’s gaze behavior, three different measures were used for both paradigms, following previous research (Paulus et al., 2011; Schuwerk & Paulus, 2015). First we assessed whether participants generally anticipated to one of the two targets, irrespective to which one. The second measure assesses participants’ expectations based on their first fixations to either of the targets. The third measure assessed looking time durations to both targets. This measure was included to control for corrective eye movements. Frequency Score This score assesses whether all three age groups showed an equal amount of general anticipations (irrespective to which of the two targets they fixated) in all three test trials for each paradigm. In the hand-paradigm, a fixation was counted with 1, when participants first fixated the hand and fixated one of the targets after (no matter which of the targets). For the path-paradigm, a fixation was counted with 1, when participants first looked at the path leading to the occluder before the agent disappeared and fixated on one of the target-AOIs after the agent disappeared. A fixation somewhere else or no fixation was coded with 0 (Schuwerk, Sodian, & Paulus, 2016). First Fixation Score This score was calculated similar to the Frequency Score, with the only difference that now the type of first fixation, namely to which specific target was fixated, was considered. To be counted as a first fixation for the hand-paradigm, participants had to fixate the hand first and fixate on one of the targets after. In the path-paradigm, first fixations were only included when participants first looked at the path leading to the occluder before the agent disappeared and fixated on one of the other AOIs after the agent disappeared. A fixation from the beginning of the path to the correct path was coded with 1, a fixation to the incorrect, that is, the familiarized path, was coded with −1 and a fixation somewhere else on the screen or no fixation was coded with 0 (Schuwerk & Paulus, 2015). The score was averaged over the three test trials for analysis. If participants have two or more missing values for test trials in each paradigm they were excluded for that score. One 32-month- old in the hand-paradigm did not show gaze data for at least two test trials and one 11-month-old in the path-paradigm. Differential Looking Score (DLS) A score was calculated that represents the relative amount of time on one AOI in relation to the other. This score was additionally included to control for corrective eye-movements, as participants could fixate first on one AOI but direct most of the following fixations to the other AOI (Schuwerk & Paulus, 2015; Senju, Southgate, White, & Frith, 2009). Therefore the total looking time to the incorrect goal-AOI was subtracted from the total looking time to the correct goal-AOI and divided by the sum of total looking time to both goal-AOIs in that time. This results in scores between −1 and 1; a value towards −1 would indicate a preference for the novel goal, a value towards 1 a preference for the old goal. Frequency of anticipations For the hand-paradigm, participants anticipated in 280 out of 312 trials (9.74%). For the path-paradigm, participants anticipated in 270 out of 312 trials (6.54%). A generalized estimating equations model (GEE; Zeger & Liang, 1986) was calculated for each paradigm separately (unstructured working correlation matrix, logit link function, binomial distribution) to see whether age group or test trial (first, second or third trial) or the interaction of age group and test trial had an effect on participants’ frequency of anticipatory first fixations. Results can be found in Tables 1 and 2. For both paradigms, neither of the predictors had a significant influence on the frequency of anticipations. This means that all age groups showed an equal amount of anticipations in all three test trials for both the hand- and path-paradigm. First Fixation Score A repeated measures ANOVA with the First Fixation Score was performed with the within-subject factor paradigm (hand- vs. path-paradigm) and the between-subject factor age group (11-months, 32-months, and adults). No main effect of paradigm was found, F(1, 99) = 0.64, p = .427; ηp2 = 0.01, but a significant main effect of age group, F(2, 99) = 6.55, p = .002, ηp2 = 0.12. The interaction of paradigm and age group was also significant, F(2, 99) = 9.65, p < .001, ηp2 = 0.16. Consequently, a one-way ANOVA with the between subject factor age group was performed for each paradigm separately. Analysis for the hand-paradigm did not reveal significant differences between the age groups, F(2, 100) = 1.93, p = .151, ηp2 = 0.04, whereas the ANOVA for the path-paradigm turned out significant, F(2, 100) = 13.35 p < .001, ηp2 = 0.21. Bonferroni’ post-hoc tests showed a significant difference between 11-month-olds (M = −0.34, SE = 0.10) and adults (M = 0.24, SE = 0.10) with p < .001, Cohen’s d = 0.95, and 32-month-olds (M = −0.45, SE = 0.10) and adults, p < .001, Cohen’s d = 1.15. The difference between the two infant groups was not significant, p = 1.00, Cohen’s d = 0.19. DLS Results for the DLS showed a similar pattern. The repeated measures ANOVA with the between subject factor age group and the within subject factor paradigm demonstrated no significant effect of paradigm, F(1, 100) = 0.16, p = .215, ηp2 = 0.02, and age group, F(2, 100) = 1.30, p = .277, ηp2 = 0.18. The interaction effect of paradigm and age group turned out significant, with F(2, 100) = 10.61, p < .001, ηp2 = 0.18. One-way ANOVAs were performed for each paradigm separately. The ANOVA for the DLS for the hand-paradigm was significant, F(2, 100) = 3.35, p = .04, ηp2 = 0.06. Bonferroni’ post-hoc tests showed no significant difference between 11-month-olds (M = −0.05, SE = 0.07) and 32-month-olds (M = −0.23, SE = 0.07), with p = .23, Cohen’s d = −0.44, and between adults (M = −0.30, SE = 0.07) and 32-month-olds, p = 1.00, Cohen’s d = −0.17. However, comparison between adults and 11-month-olds turned out significant, p = .04, Cohen’s d = −0.59. Similarly, analysis for the path-paradigm showed a significant effect with F(2, 101) = 6.60, p = .002, ηp2 = 0.12. Bonferroni’ post-hoc tests demonstrated a significant difference between 11-month-olds (M = −0.27, SE = 0.08) and adults (M = 0.12, SE = 0.08), p = 0.003, Cohen’s d = 0.80, and between 32-month-olds (M = −0.20, SE = 0.08) and adults, p = .020, Cohen’s d = 0.68. Again the difference between the 11-month-olds and 32-month-olds was not significant, p = 1.0, Cohen’s d = 0.15. Comparisons across paradigms per age group Repeated measures ANOVAs were performed for each age group separately with the within-subject factor paradigm. Analysis of the First Fixation Score for the 11-month-olds revealed no difference in performance for the two paradigms, F(1, 32) = 3.41, p = .07, ηp2 = 0.10, just as for the 32-month-olds, F(1, 33) = 0.45, p = .51, ηp2 = 0.01. In contrast, adults performed in both paradigms differently, with F(1, 34) = 16.64, p < .001, ηp2 = 0.33 (see also Fig. 6 for descriptives). Adults showed a looking bias towards the novel object in the old location in the hand-paradigm, but a looking bias towards the goal in the path-paradigm. The same pattern was demonstrated for the DLS: No difference was found for the 11-month- olds, F(1, 33) = 3.37, p = .08, ηp2 = 0.09, and 32-month-olds, F(1, 33) = 0.04, p = .84, ηp2 = 0.001. However adults showed a significant difference between the two paradigms, with F(1, 34) = 3.01, p < .001, ηp2 = 0.43, with the same pattern as for the First Fixation Score. Type of anticipated action Further to check whether participants showed a significantly different looking bias from chance towards one or the other AOI, one sample t-tests against chance level were calculated for each paradigm and each age group separately (indicated by the asterisks in Fig. 6). 11-month-olds showed chance performance in the hand- paradigm for the First Fixation Score, t(33) = −0.88, p = .39, Cohen’s d = −0.15, and the DLS, t(33) = −0.73, p = .47, Cohen’s d = −0.13. For the path-paradigm they showed a significant looking bias towards the previous path, for the First Fixation Score, t(32) = −3.36, p = .002, Cohen’s d = −0.58, and the DLS, t(33) = −3.30, p = .002, Cohen’s d = −0.57. 32-month-olds performed in every paradigm above chance. They anticipated that the hand would grasp the novel object in the old location, indicated by the First Fixation Score, t(33) = −4.46, p < .001, Cohen’s d = −0.77, and DLS, t(33) = −3.44, p = .002, Cohen’s d = −0.59; as well as that the agent would reappear on the old path aiming for the novel goal, for First Fixation Score, t(34) = −4.63, p < .001, Cohen’s d = 0.78, and DLS, t(34) = −2.59, p = .014, Cohen’s d = 0.44. Similarly adults anticipated significantly above chance that the hand would grasp the novel object in the old location, with t(34) = −2.57, p = .02, Cohen’s d = −0.44 for the First Fixation Score, and t(34) = −4.33, p < .001, Cohen’s d = −0.73 for the DLS. In contrast for the path-paradigm, adults anticipated the agent’s reappearance on the correct path, with t(34) = 2.24, p = .032, Cohen’s d = 0.38 for the First Fixation Score. Results of the DLS turned out not significant, t(34) = 1.41, p = .167, Cohen’s d = 0.24. To be able to compare our results better with Cannon and Woodward (2012) we further analyzed only the first test trial of the hand-paradigm. As Cannon and Woodward (2012) presented infants four blocks in the design of three learning trials and one test trial per block, and given that we presented participants five learning- and three consecutive test trials, an analyzation of only the first test trial is closer to a replication of the original study (with the difference being that we have five instead of three learning trials). One-sample t-test revealed chance performance for the 11-month-olds for the DLS (M = 0.05, SE = 0.11), t(32) = 0.49, p = .63, Cohen’s d = 0.09, and a First Fixation Score of M = −0.12, SE = 0.17. The 32-month-olds showed a significant looking bias towards the old location for the DLS (M = −0.26, SE = 0.08), t(33) = −3.18, p = .003, Cohen’s d = −0.55, with also a negative First Fixation Score (M = −0.53, SE = 0.15). Similarly the t-test was also significant for the adults with t(34) = −4.05, p < .001, Cohen’s d = −0.69, for the DLS (M = −0.36, SE = 0.09), and a First Fixation Score of M = −0.37, SE = 0.15, indicating a looking bias towards the location. In sum, even when we only analyzed the first test trial, we did not find goal directed anticipations over all age groups for the hand-paradigm. Correlation analysis For each age group the Spearmen’s Roh correlation between scores of the hand- and path-paradigm were calculated (see Table 3). All correlations of the control measures can be found in the supplemental material. Analysis of the 11-month-olds and 32-month-olds revealed no significant correlations between the two paradigms. Interestingly, for adults the First Fixation Score did not correlate with each other, whereas the DLS turned out significant. Additional analysis Also Study 2 was additionally analyzed as similar as possible to the study by Cannon and Woodward (2012). Details of analysis and results can be seen in the supplementary material. Results suggest that neither for the hand- nor for the path-paradigm, infants anticipated the action goal-directed. Even when looking at the three test trials separately for the hand-paradigm (see Fig. 7), participants’ performance was never above chance level.","The second study investigated two questions: (1) Do children and adults anticipate an action of a non-human agent and human agent based on the previously observed goal instead of the previously observed movement path? (2) Do children and adults show similar processing strategies regarding goal encoding for two different paradigms? To this end 11-month-olds, 32-month-olds and adults were presented with two different paradigms. One presented a human hand reaching for one of two objects and the other presented an animated animal walking to one of two goals. Both followed a similar paradigm as described in Study 1 with a previous familiarization phase and a subsequent test phase in which the position of the goal objects were swapped. Results revealed that when the location of targets had changed, 11-month-olds and 32-month-olds did not perform goal-directed anticipations for both types of actions. This is surprising, since we expected goal anticipations for human actions from early on (Cannon & Woodward, 2012; Meltzoff, 1995; Woodward, 2005). Instead we observed for 11-month-olds chance performance when presenting the action performed by the hand. Considering the path-paradigm with the animated agent, 11-month-olds looked longer towards the familiarized path leading to the novel goal. Hence results are similar to Study 1 and Daum et al. (2012). Together with the experiments of Study 1, our results are not in line with those by Cannon and Woodward (2012) and do not support theoretical claims (e.g., Woodward, 2009) that the ability to flexibly anticipate other’s action goals emerges in infancy. Rather, they support approaches that assume that early action understanding is a multi-faceted construct that involve different kinds of processes and mechanisms (Uithol & Paulus, 2014). 32-month-olds showed a looking bias in both paradigms towards the old path. On the one hand, this is noteworthy, as not even older children visually anticipated the goal of a human action. Yet, even previous research reported a mixed pattern of goal anticipations for human actions in toddlers (Krogh-Jespersen et al., 2015, 2018). The performance in the path-paradigm is in accordance to previous findings (Daum et al., 2012) and our hypothesis. Our results imply that 11-month-olds and 32-month- olds are not able to anticipate goals of non-human actions. This does not support the theoretical claim that infants understand goal-directed actions of non-human actions from early on (Luo & Baillargeon, 2005). The correlational analysis confirmed that children do not use the same processing strategies when observing a human or animated animal, as we didn’t find any significant correlations between the paradigms. In sum, our results cannot confirm the widely made assumption that children perceive human actions based on their goals (Ambrosini et al., 2013; Cannon & Woodward, 2012; Falck-Ytter et al., 2006; Woodward, 1998). Quite the contrary, the findings suggest that infants process actions based on visuo-spatial information and represent actions as movement patterns. Thus, our findings support low-level accounts of social understanding, indicating that infants make use of simple information when they process actions, such as statistical regularities (Daum, Wronski, Harms, & Gredebäck, 2016; Ruffman, 2014; Uithol & Paulus, 2014). Further we did not find any correlations within children between the two paradigms, which speaks against the claim that children process human und non-human actions similarly from early on (Leslie, 1995; Luo & Baillargeon, 2005). Adults reacted according to our expectations in the path-paradigm: They fixated the novel path leading to the familiarized goal first. Surprisingly adults showed contradicting anticipations in the hand-paradigm; they expected the hand to grasp the novel object. Thus, they encoded the movement trajectory instead of the action goal. This is interesting, since the goal of the actor should be clear at least for adults. However, so far there is only one study that tested the paradigm of Cannon and Woodward (2012) in adults. Pfundmair et al. (2017) found that adults anticipated the goal of a grasping hand as well as of a grasping claw, which is in contrast to our findings. Given this contradiction and lack of prior studies within adults, further investigation is needed. Maybe this paradigm is not equally suitable for adults as it might be for children (i.e. measuring the same underlying ability). Interestingly, in our study we found a positive correlation between adults’ looking times in the two paradigms, which indicates that they demonstrated related processing strategies for the two types of stimuli. Ramsey and de C. Hamilton (2010) showed in a fMRI-study that adults process goal-directed actions of a triangle similar to goal-directed grasping actions of a human hand. They concluded that the fronto-parietal network, which is often referred to as the human mirror neuron system, actually encodes goals rather than biological motion (Ramsey & de C. Hamilton, 2010). Similar assumptions have been made by Schubotz and Von Cramon (2004), who found activity in motor areas for abstract, non-human stimuli. One assumption would be that the role of experience influences participants’ performance in the two tasks (Ruffman et al., 2012). Adults probably gained more experience with animations and cartoons throughout their life, whereas infants are still not well familiarized with them.","In the previously described studies, we conducted six similar experiments that measured whether infants between 11- and 12-months of age anticipate the goal of an action instead of the movement path. Infants performed in none of the six experiments visual anticipations towards the goal, questioning the theoretical claim that infants selectively encode and anticipate other’s action goals (Cannon & Woodward, 2012; Woodward, 1998). Instead we observed the tendency of infants to anticipate the movement path. In order to produce a more reliable estimate of the observed looking bias towards the location, a meta-analysis was conducted. We wanted to find out whether this effect is significant over all of our six experiments. Effect sizes Effect sizes for the meta-analysis were expressed as correlation coefficients, r. This metric was chosen, as correlation coefficients are easier to interpret and compared with other metrics (Field, 2001; Rosenthal, 1991). The experiments of Study 1 were treated as single studies for the meta-analysis. As we wanted to see whether the effect of infants making location- instead of goal-directed anticipations is significant over all experiments, we used for each experiment of Study 1 the effect size of the comparison between location and goal-directed anticipations in test trials. This resulted in five effect sizes for each experiment of Study 1 (see Table 2 for effect- and sample sizes). For Study 2 we used the mean effect size of both paradigms of the First Fixation Score, following Rosenthal’s (1991) suggestion for cases when there are more effect sizes within a study. As the effect sizes of Study 1 are based on infants’ first fixations, we did not include the Differential Looking Score of Study 2 in this analysis. In Study 2, the direction of infants’ looking bias was measured via one-sample t-tests, which is why we used the effect-sizes thereof (see Table 2). Method of meta-analysis We assume that our sample sizes come from the same population and expect our sample effect sizes to be homogenous, since all our experiments are similar and collected from similar populations. According to this we decided on a fixed- effects model instead of a random-effects model, as suggested for these circumstances by Field and Gillett (2010). For calculating the fixed-effects model, the approach by Hedges et al. (Hedges & Olkin, 1985; Hedges & Vevea, 1998) was used. The analysis was calculated via written syntax for SPSS (see Field & Gillett, 2010).","The mean effect size based on the model was −0.24 with a 95% confidence interval (CI95) of’ −0.38 (lower) to −0.09 (upper) and a highly significant z score (z = 3.07, p = .002). According to Cohen’s criterion (1988) this is a small to medium effect. A chi-square test of homogeneity of effect sizes was not significant, χ²(5) = 1.51, p = .912. This supports our previous assumption of a low between-study variance indicating that our participants are sampled from the same population. To illustrate infants’ looking bias towards the location instead of the goal over all experiments, the number of anticipations towards the location versus the goal was summed up over all six experiments. From a total of 681 anticipations, infants directed 423 anticipations towards the location whereas 258 anticipations were performed towards the goal (Table 4).","The meta-analysis aggregated data from 178 participants over all six experiments and demonstrated that samples of 11- to 12-month old infants show anticipations directed towards the path of an action and not to its goal, with a small to moderate magnitude. This result underlines the previous findings of Study 1 and 2, and highlights the conclusion of our findings: Through all our 6 experiments we were not able to find a goal- directed looking bias within 11- to 12-month-olds. Rather, we observed anticipations of infants towards the location indicating that infants encode the movement path instead of the goal of an action. This is further discussed in the next section.","The present work addresses the question of whether and under which circumstances children and adults flexibly anticipate actions as goal-directed. It has been claimed that infants show visual goal anticipations earlier for human than for non-human actions (Meltzoff, 1995; Paulus, 2012; Ruffman et al., 2012; Woodward, 2005), whereas others proposed that infants are able to anticipate goals for non-humans equally well as for humans (Leslie, 1995; Luo & Baillargeon, 2005). Cannon and Woodward (2012) demonstrated that 11-month-old infants are able to encode an action of a human hand, but not of a mechanical claw as goal-directed. In contrast, Daum et al. (2012) could only find goal-directed anticipations within children from 3 years of age and adults, when using an animated agent instead of a human hand. With the current set of studies, we aimed to contribute to the current debate on infants’ goal anticipations by exploring whether infants indeed anticipate actions flexibly based on goals rather than based on movement paths and patterns. Two different labs collaborated over the attempt to replicate both conceptually and directly the findings reported by Cannon and Woodward (2012). We hypothesized that infants understand human actions earlier than actions of a non-human animated agent. Therefore, we expected to replicate the findings of Cannon and Woodward (2012) in our experiments that used a human hand as agent but not necessarily in the experiments that used non-human animated animals as agents. However, despite our systematic variation of any other factors that could have had an influence on infants’ action perception (such as type of agent, presence of an occluder or number of learning- and test trials), we failed to replicate the findings of Cannon and Woodward (2012) across all our experiments and labs. Infants between 11 and 12 months of age did not anticipate an action in relation to its goal, but rather to its movement path. This was observed, irrespective of whether the agent was a human or an animated animal. We even included an experiment that was as close as possible based on the stimuli of Cannon and Woodward (2012) and were not able to replicate their finding. A meta-analysis aggregated over all of our data emphasizes our result and illustrates a consistent effect for a looking bias towards the location and not the goal within 11- and 12-month-olds. The finding, that infants and young children base their anticipations on the movement path relates to previous eye-tracking studies (e.g., Paulus et al., 2011; Schuwerk & Paulus, 2015) and further stresses the role of frequency information for infants’ action processing. Our results underlie the claim that already infants are able to detect contingencies and regularities in their environment, as supposed to be an important learning mechanism (Kirkham, Slemmer, & Johnson, 2002; Ruffman et al., 2012; Smith & Yu, 2008). Does this mean that infants do never show flexible goal- directed anticipations? It should be noted that the paradigm by Cannon and Woodward (2012) presented participants with a scene of two objects placed on a white background and a hand that appears from the left side to grasp one of them. This scene is filmed from a birds’-eye view and does not include the whole situation, namely an agent sitting on a table in front of two targets. Theoretically, it has been proposed by predictive coding accounts that active perception is informed by environmental features (e.g., Clark, 2013). Also studies with infants could demonstrate the informative influence of context information on infants’ action anticipations (Fawcett & Gredebäck, 2015; Stapel, Hunnius, & Bekkering, 2012). Usual situations in our environment provide us with lots of information that is already there before an action takes place, and thus open room for top-down influence before an action is performed (Clark, 2013). Thus, it is possible that the situation in the Cannon and Woodward-paradigm (2012), as well as in our experiments, might be too abstract and out of context. The fact that we also couldn’t find goal- directed anticipations for adults adds additional concern towards this paradigm. However, adding more ecological cues might help infants. Other studies reported goal anticipations within infants when they were presented with a movie depicting the entire human agent grasping for an object (Kim & Song, 2015; Krogh-Jespersen & Woodward, 2014, 2018; Krogh- Jespersen et al., 2015). Furthermore, there are studies demonstrating goal anticipations in infants and toddlers when they had the possibility to learn about the agent’s goal (Paulus, 2011; Paulus et al., 2017). For example, in Paulus et al.’s study (2017), children observed an agent walking to a goal for several times whereas the goal’s position changed from time to time. This highlighted the goal object on the expense of the movement path (see also Ganglmayer, Schuwerk, Sodian, & Paulus, 2019). In sum these findings suggest that infants might need additional information in order to visually anticipate the goal and not the movement path. Another possibility relates to cultural factors. While the study of Daum et al. (2012) and the current studies were implemented in Europe, the study of Cannon and Woodward (2012) was conducted in the U.S. A comparison of various studies on children’s TV consumption suggests that infants watch more TV in the U.S. than in Europe (Feierabend & Mohr, 2004; Zimmerman, Christakis, & Meltzoff, 2007). In an additional pilot study, we asked 22 parents of 12-month-olds about the average time their infants watch TV. Results showed that only 23% of the 12-month-olds watch TV and for an average time of 4 min. In comparison, Zimmerman et al. (2007) reported that infants in the US spend on average an hour per day in front of a TV screen. Given the fact that prior experience with technology and real objects increases learning through TV (Hauf, Aschersleben, & Prinz, 2007; Troseth, Saylor, & Archer, 2006), it is possible that different experiences with TV and screens could be a reason for our diverging findings. Yet, on the other hand, since we did not find goal-related anticipations in 32-month-olds or even adults (Study 2) who arguably have ample experiences with TV, this factor is rather unlikely to be the central cause underlying our non-replication. Our work highlights the importance for replication studies in developmental psychology, as recently researchers warned of false positive findings and publication biases. Recent attempts to replicate findings of implicit Theory of Mind tasks turned out difficult as well (e.g., Burnside, Ruel, Azar, & Poulin-Dubois, 2018; Kammermeier & Paulus, 2018; Kulke, Reiß, Krist, & Rakoczy, 2017; Powell, Hobbs, Bardis, Carey, & Saxe, 2017). Some specifically tried and failed to replicate an implicit Theory of Mind task based on anticipatory looking measures, questioning the robustness of these tasks (Kulke, von Duhn, Schneider, & Rakoczy, 2018; Schuwerk, Priewasser, Sodian, & Perner, 2018). One strength of the present manuscript concerns the integration of work conducted at two different labs. Such a collaborative effort is not only in line with suggestions for future best practices (e.g., Frank et al., 2017), but also increases the conclusiveness of our results. In sum, our findings over various types of stimuli and across two labs suggest that infants around 11 and 12 months of age do not flexibly anticipate an action based on its goal, but primarily use the information about the movement path, when presented with two possible targets that change location for test trials (following Cannon & Woodward, 2012). This indicates that infants base their anticipations on frequency information (similar to Henrichs, Elsner, Elsner, Wilkinson, & Gredebäck, 2014). When they observe an action performed in a certain way, they use this information to predict that action in the future (Ruffman, 2014). However, infants in their first years of life seem to process this information in relation to movement paths (Paulus et al., 2011). Our findings also imply that we need to reflect on what we exactly mean when we say that infants understand goals, as a goal, in its general meaning, refers to a desire (Ruffman et al., 2012; Uithol & Paulus, 2014). In conclusion, according to our six experiments, infants do not flexibly anticipate action goals around 12 months. Rather, their anticipation of others’ actions might rely on more simple heuristics, such as movement trajectories.","This work was partly supported by the German Research Foundation (DFG, Nr. PA 2302/6-1)."],["Intentional motor actions and their effects are bound together in temporal perception, resulting in the so-called intentional binding effect. In the current study, we address an alternative explanatory mechanism for the emergence of temporal binding by excluding the role of motor action. Employing a sensory-based Libet clock paradigm, we examined temporal perception of two different auditory stimuli, and tested the influence of beliefs about the causal relationship between the two auditory stimuli, thus simulating a crucial feature of intentional action. In two experiments, we found a robust temporal repulsion effect, indicating that instead of being attracted to each other, the auditory stimuli were shifted away from each other in temporal perception. Interestingly, repulsion was attenuated by causal beliefs, but this effect was fragile. Furthermore, temporal repulsion was unaffected by the intensity of prior learning. Findings are discussed in the context of intentional action awareness research and multisensory integration. --------------------------------------------------------------------------------","The ability to integrate sensory information into a coherent picture of the world is an important human trait. It enables people to build a conscious representation of complex visual scenes and auditory events that are otherwise disconnected in time or space (Serences & Yantis, 2006; Shamma, Elhilali, & Micheyl, 2011). Apart from integrating perceptual input, humans also form coherent representations of their own behavior (Aarts & Custers, 2009; Vallacher & Wegner, 1987). Self-produced actions that predict and result in sensory feedback, such as pressing keys that result in a sound or light, are perceived as coherent events of meaningful behavior (e.g., playing a tone on the piano or illuminating a room), even though action and effect are spatially and temporally separated and do not form a single entity. This ability to construct a coherent conscious experience of one’s own behavior has important implications for the understanding of the self and the sense of agency, i.e. the feeling of causing one’s own action and effects in the external world. Specifically, people perceive their intentional actions and resulting outcomes as if they are temporally bound together in conscious awareness; a phenomenon that is referred to as intentional binding (Haggard, Clark, & Kalogeras, 2002b). It is assumed that this temporal binding effect is rooted in internal prediction processes resulting from intentionally selected motor actions, causing the compression of the time interval between action and effect and the formation of a coherent representation of the action-effect event (Frith, Blakemore, & Wolpert, 2000). Recent research, however, suggests that this temporal binding effect might not be the result of intentionality per se (e.g., Desantis, Hughes, & Waszak, 2012; Dogge, Schaap, Custers, Wegner, & Aarts, 2012; Hughes, Desantis, & Waszak, 2013), but rather a product of other factors that are associated with intentional action. In the present study, we examine this subject in more detail and raise the question whether the temporal binding effect is similarly sensitive to events that do not require any form of action. Building on the notion that intentional actions are associated with beliefs in causality that might bias the perceptual relation between two events (Buehner & Humphreys, 2010; Desantis, Roussel, & Waszak, 2011; Eagleman & Holcombe, 2002), we tested whether causality beliefs cause two successive auditory stimuli to be perceptually bound together in the Libet clock paradigm. The emergence of such a binding effect as a function of causality beliefs would undermine the validity of the temporal binding task as a read-out of intentionality and a measure of sense of agency. One central argument that has been put forward in favor of the central role of intentionality in action concerns the impact of externally induced action on temporal binding. Some studies employed transcranial magnetic stimulation (TMS) to induce an involuntary action followed by a tone (Haggard & Clark, 2003; Haggard et al., 2002b). Subjects judged the time of movement onset and tone onset by using the Libet clock. The time judgments did not reveal a temporal binding effect but a clear and strong repulsion effect instead, indicating that subjects perceived the motor movement and subsequent tone as two separate events. While these TMS experiments suggest that the temporal binding task can distinguish involuntary from voluntary action in producing a coherent representation of one’s own action, it is important to note that the magnetic stimulation of motor areas results in an abrupt, reflexive and oddly felt movement of a limb (here a finger). Clearly, human agents have the capacity to evoke actions by direct brain stimulation, but such action onset does not fit with everyday externally induced motor movement where intentionality is ruled out, for example when someone else pushes or pulls one’s hand into a specific direction. Building on this notion, recent studies show that temporal binding might be reduced but does not vanish for actions that are not self-chosen or driven by internal motivation (Antusch, Aarts, & Custers, 2019; Caspar, Christensen, Cleeremans, & Haggard, 2016) or passively induced in the agent by another experimenter or a mechanical lever press (Borhani, Beck, & Haggard, 2017; Dogge et al., 2012; Engbert & Wohlschläger, 2007; Moore, Wegner, & Haggard, 2009). These findings suggest that it is not intentionality per se that drives the temporal binding effect, as externally induced movement does not cancel out temporal binding in settings that include clear causal relations between action and outcome. One explanation for the observed binding between externally induced actions and a tone might be that participants were able to predict the outcome and believed to have a causal role in producing it. Causal beliefs impact cognitive processes at several stages, e.g., learning, judgment and decision-making (Harris, Fiedler, Marien, & Custers, 2019; Lucas & Griffiths, 2010). Moreover, higher-order causal beliefs have been shown to affect low-level processes of perception (Buehner & Humphreys, 2009, 2010; Cravo, Claessens, & Baldo, 2009). A study by Buehner and Humphreys (2009), for example, suggests that mentally linking one’s key press to a subsequent tone following a presumed causal association enhances the anticipation of the tone. However, not only effects stemming from motor actions are subject to temporal binding, also external (visual) events can be spatially bound together when they share a common causal association (Buehner & Humphreys, 2010). Accordingly, holding a higher-order belief about the causal relationship between two events increases binding between them, even when action is not actively involved. In a further exploration of the role of causality beliefs in temporal binding of involuntary movement and effect, Dogge et al. (2012) subjected their participants to a temporal binding task with the Libet clock. Participants either executed a voluntary key press or an involuntary key press that was induced by a lever pulling down their index finger. In the involuntary conditions, some participants were asked to consider themselves as the agent, construing the resulting effect (tone) as being caused by their key press. The magnitude of the binding effect increased between the simple involuntary and involuntary self-causation condition, with mean judgment errors (indicating a stronger binding effect) being larger in the involuntary self-causation condition (Dogge et al., 2012). Similarly, it has been shown that explicit beliefs of being the causal agent resulted in stronger intentional binding in participants in a joint action task (Desantis et al., 2011). Whereas informative, these studies do not entirely rule out the role of motor prediction and intentionality in temporal binding between action and effect. It is for example conceivable that participants started up their motor system once they noticed a slight movement of the mechanical lever, causing them to experience a motor prediction signal and/or intentionality over the key press. Because of the tight link between perception and action (e.g., Decety & Grèzes, 2006; James, 1890; Prinz, 1997), a similar reasoning might account for findings that show binding effects as a result of observing others’ action (e.g., Desantis et al., 2012; Kirsch, Kunde, & Herbort, 2018). A more stringent test of the role of causality beliefs in temporal binding is provided by sensory paradigms that do not employ motor action. In one such study by Haggard and colleagues (2002a), participants judged the temporal onset of two identical successive auditory stimuli that were separated by a 250 ms time interval. Results revealed that especially the second stimulus was temporally repulsed from the first stimulus (Haggard, Aschersleben, Gehrke, & Prinz, 2002a). Similar results were obtained by Desantis et al. (2012) who compared the temporal perception of two non-motor stimuli separated by a 400 ms time interval against an operant action condition in a Libet clock paradigm. The findings were such that both, congruent and incongruent (with precedingly learned associations), tone pairs resulted in temporal repulsion instead of temporal binding, suggesting that temporal contiguity and identity prediction were not sufficient for binding to emerge (Desantis et al., 2012). The first tone was perceived as occurring earlier than in actuality, and the second stimuli was perceived as occurring later than in actuality, leading to a repulsion instead of a binding effect. Thus, instead of integrating the two stimuli such as in intentional binding, participants seemed to detach the two from each other by experiencing them as two separate entities and shifting them away from each other in temporal perception. It is important to note that these studies did not investigate the role of causal beliefs and might have suffered from flaws that reduced the probability of observing a temporal binding effect. For example, Haggard et al. (2002a) used successive tones of the same frequency and length. Hence, there might have been little inclination for participants to integrate the tones into a coherent event and to perceive them as causally related. Thus, participants perceptually shifted the tones away from each other. Using tones of different frequencies, Desantis et al. (2012) paired the tones in congruent and incongruent within- subject blocks and employed an unconventional 400 ms time interval. Because the tones were sometimes congruent and sometimes incongruent, a stable predictive link could not be formed. Furthermore, it is unclear how the extended time interval affected the results. Finally, while suggesting temporal repulsion, comparisons with the baseline tone condition in the Haggard et al. (2002a) study only suggest temporal repulsion of the second tone, while the first tone seems to even be bound to the subsequent tone. Desantis et al. (2012) only tested the temporal perception of the second tone while a comparison baseline condition was lacking completely. Hence, the interpretation of the findings as demonstrating temporal repulsion was neither statistically tested nor can it be claimed in the absence of a comparison baseline condition. Thus, it remains subject to empirical scrutiny whether external stimuli in successive contexts are repulsed or bound to each other in temporal perception. The present study therefore aimed to replicate the temporal binding (repulsion) effect with two different auditory stimuli by employing a unimodal sensory paradigm based on the Libet-clock. In our task, participants acquired the associations between two auditory stimuli pairs, such that the first stimulus predicted the onset of the second one. Each pair consisted of a noise sound (such as a scratch sound) followed by a sinusoidal tone 250 ms later. The pairings remained constant for each participant, allowing for participants to perceive the stimuli as being linked and represented as one event. Furthermore, as an important addition to earlier work, we examined the role of causality belief in the perceived temporal attraction between the two auditory stimuli. To that end, we manipulated explicit beliefs about the causal relationship between the two successively presented auditory stimuli to more stringently test their influence on temporal binding of the stimuli. The first experiment addresses the basic role of perceived causality in modulating the temporal binding between two different auditory stimuli. Experiment 2 was designed to further explore the role of perceived causality in a context where the association between the first and second auditory stimulus was strengthened by practice. Participants and design ~~~~~~~~~~~~~~~~~~~~~~~ In total, 641 individuals (49 females), with a mean age of 23 years (M = 23.02, SD = 6.15) participated in the experiment in exchange for course credit or monetary reimbursement. Participants were randomly assigned to one of two between-subject conditions, resulting in 33 participants in the control condition and 31 participants in the causal belief condition. The study was carried out in line with the guidelines of the Declaration of Helsinki and was approved by the Ethics committee of the Faculty of Social and Behavioral Sciences at Utrecht University as part of a project line (ethics approval code: FETC17-124). All participants gave written informed consent. A 2 (target: noise vs. tone) × 2 (type of trial: baseline vs. succession trials) × 2 (condition, between: control vs. causal belief) mixed design was employed.","The experiment consisted of an altered Libet clock paradigm as used by Haggard et al. (2002) and was programmed using Eprime Software version 2.0. Four different auditory stimuli of 100 ms length were used: two noise sounds and two sinusoidal tones. The noise sounds resembled scratch sounds while the sinusoidal tones were of either 300 Hz or 1000 Hz frequency.2 All stimuli were played via a Sennheiser headphone. Acquisition phase The experiment began with an acquisition phase in which participants were familiarized with the auditory stimuli and their combinations. At the start, participants were told that four different auditory stimuli would be used in the experiment and that they would begin with experiencing each stimulus once. While participants experienced a stimulus, an on-screen message stated which stimulus they were experiencing. By pressing F1, they could advance to hearing the next stimulus. Subsequently, to facilitate the distinguishability of the stimuli, participants experienced each stimulus four additional times. These trials advanced automatically and like in the preceding trials, it was stated on screen which stimulus participants were experiencing. After participants were familiarized with the different auditory stimuli, they learned that each noise sound would be paired with one of the sinusoidal tones during the experiment. In the control condition, participants were told that the computer would determine the combinations of the noise and tone, and that they would therefore be independent of each other. Put differently, the pairs of auditory stimuli would not be causally related. They then heard the independent noise – tone combinations. In the causal belief condition, participants learned that the computer would choose the noise that would cause the second tone to occur. The second tone would thus be dependent on the selection of the noise. Put differently, the pairs of auditory stimuli would be causally related. They then heard the dependent noise – tone combinations. The noise – tone combinations were counterbalanced across participants but remained constant for a given participant. In both conditions, the noise – tone combinations were presented five times each (i.e., ten trials in total). Participants were asked to pay close attention. They then completed eight trials in which they heard noise – tone combinations that were either congruent or incongruent with the previously learned pairs. In the control condition, they were asked to indicate which combinations of tones were presented before, whereas in the causal belief condition, they were asked to indicate which combinations of tones were causally related. This was done to further emphasize and boost the independent (control condition) and the dependent nature (causal belief condition) of the stimuli pairs in the respective conditions before the start of the actual experiment. Experimental trials (altered Libet clock task) Next, participants commenced with the experimental trials. In both conditions, participants completed four blocks in randomized order. The blocks consisted of two baseline blocks (baseline noise, baseline tone) and two succession blocks (succession noise, succession tone). Each block comprised 24 trials, including four practice trials and 20 experimental trials. The practice trials preceding the experimental trials of each block were identical to and indistinguishable from the experimental trials but were not analyzed. Each trial began with the presentation of a dotted clock face consisting of forty grey dots arranged in a circle around the center of the screen, with a diameter of six centimeters. A black dot served as a clock hand and was moving alongside the grey dots with a speed of 2560 ms per rotation. The starting position of the clock hand varied between trials and was random. Participants were asked to attend to the clock face and the rotating clock hand. Depending on the trial, they then either perceived a single auditory stimulus or a stimuli pair in succession. The onset of the stimuli was programmed to always occur during the second rotation of the clock hand. In baseline blocks, participants either perceived one of the two noise sounds (baseline noise) or either of the two sinusoidal tones (baseline tone). In the succession blocks, participants always perceived an initial noise sound followed by the paired sinusoidal tone at 250 ms. After the presentation of the last stimulus, the clock hand continued moving for a duration of 1000 ms before it disappeared. Subsequently, the mouse cursor as well as a prompt to make a time judgment appeared in the center of the clock face. Participants were required to judge either the temporal onset of the noise sound (succession noise trials) or the temporal onset of the sinusoidal tone (succession tone trials) by clicking on one of the dots to indicate the position of the clock hand at the time of the event. In baseline trials, participants always judged the onset of the single presented stimulus (noise in baseline noise and sinusoidal tone in baseline tone trials). While blocks and trials were identical in both between-subject conditions, instructions differed. In the causal belief condition, instructions for the individual blocks emphasized the causal and dependent relationship of the two stimuli on each other. For example, participants in the causal belief condition read that they would hear two stimuli in succession, meaning that the tones were dependent on each other and that the first noise would cause the following sinusoidal tone. In the control condition, in contrast, the independence of the stimuli was emphasized. At the end of the experimental trials, participants filled in their age and gender. Moreover, participants also answered a question concerning their belief in the causal dependence of the second stimulus on the first stimulus (as indicated on a 9-point Likert scale ranging from 1 - absolutely not to 9 - very much). This question served as a manipulation check. Finally, participants were debriefed and thanked for their participation as well as reimbursed.","Prior to the main statistical analyses, extreme temporal judgments were excluded on a trial basis. Following Aarts et al. (2012), temporal judgments, which were more than ±640 ms away from the actual temporal onset, were excluded as it can be assumed that participants either rated the onset of the wrong stimulus or were inattentive in general on those trials. Overall, the number of judgments excluded based on this criterion was very small (0.14% of all judgments). Based on earlier research suggesting a temporal repulsion effect between auditory stimuli (Desantis et al., 2012; Haggard et al., 2002a), we first aimed to test whether there was a significant overall temporal repulsion effect independent of condition. First, for each trial a judgment error (in milliseconds) was calculated as the difference between the judged time of an event and its actual time of occurrence. A positive judgment error corresponds to delayed awareness of the event, and a negative judgment error corresponds to anticipatory awareness. Next, for each participant separate shifts for the first (noise) and second (sinusoidal tone) stimulus were computed by subtracting the mean baseline judgment errors from the mean judgment errors in succession trials (noise shift = mean judgment error in succession noise trials – mean judgment error in baseline noise trials; tone shift = mean judgment error in succession tone trials – mean judgment error in baseline tone trials). Finally, overall binding scores were computed by subtracting the shift of the second stimulus from the shift of the first stimulus (i.e., overall score = noise shift – tone shift). Positive overall binding scores thus indicated temporal compression of the interval between the stimuli (temporal binding) while negative scores were indicative of temporal repulsion. See Fig. 1 for the distribution of the raw data in the full sample. First, we examined the direction of the overall temporal shift. To that end, we conducted a one-sample t-test of the overall binding score against zero. As we had no hypothesis pertaining to either a temporal binding or a temporal repulsion effect, we conducted a two-tailed test. Next, we aimed to test whether temporal perception of the stimuli in the two conditions was statistically different. Specifically, we assumed that stimuli in the causal belief condition would be bound more strongly than in the control condition. In line with that, we conducted a one- tailed independent samples t-test on the overall binding scores in both conditions. Finally, we tested the direction and significance of the temporal shift in the two conditions. That is, we tested whether the perceptual shift was equally present in both conditions. For that purpose, two one-tailed one-sample t-tests of the overall binding scores in each condition against zero were conducted. Additionally, to illustrate the strength of the evidence and confirm the Frequentist analyses, equivalent Bayesian analyses were performed. For the convenience of the reader, complete Bayesian analyses are reported. Manipulation check ~~~~~~~~~~~~~~~~~~ Prior to the main analysis, we conducted an independent samples t-test, testing the belief in causality rating between control and causal belief condition. Confirming our manipulations, the results revealed that participants in the causal belief condition (M = 6.68, SD = 2.15) believed more strongly in the causal dependency of the second stimulus on the first stimulus than participants in the control belief condition (M = 5.24, SD = 2.48), t(62) = −2.4688, p = .016, two-tailed, d = 0.61, 95% CI [−1.12, −0.11]. Main analyses ~~~~~~~~~~~~~ Next, we proceeded with the main analysis as outlined in the data analysis plan. Overall temporal repulsion To examine the direction of the perceptual shift and its statistical significance, a one-sample t-test of the overall binding scores against zero was conducted. Results revealed a highly significant repulsion effect equaling 77 ms (M = −77.20, SD = 65.94), t(63) = −9.366, p < .001, two-tailed, d = −1.17, 95% CI [−1.49, −0.85]. This temporal repulsion effect occurred bi-directionally. That is, noises were significantly more anticipated on succession trials (M = 13.12, SD = 46.97) as compared to baseline trials (M = 50.97, SD = 47.94), t(63) = 5.949, p < 001, two-tailed, dz = 0.74, 95% CI [0.36, 0.80]. Similarly, tones were perceived significantly later on succession trials (M = 90.17, SD = 49.34) than on baseline trials (M = 40.83, SD = 50.82), t(63) = −6.932, p < .001, two-tailed, dz = −0.87, 95% CI [−1.31, −0.66]. Difference in temporal repulsion between conditions In addition, to test whether temporal perception differed for participants in the control and the causal belief condition, overall binding scores were subjected to an independent samples t-test based on condition. The results revealed that participants in the causal belief condition (M = -52.50, SD = 67.24, N = 31) perceived the stimuli as significantly closer to each other in time than participants in the non-causal control condition (M = −100.40, SD = 56.35, N = 33), t(62) = −3.096, p = .0015, one-tailed. The difference of −47.9 ms was of moderate effect size, d = −0.77, 95% CI [−1.28, −0.26].3 See Fig. 2 for a visualization of the overall shifts for the separate between-subject conditions. Temporal repulsion per condition Following, two one-sample t-tests were conducted to test whether the judgments were significantly different from zero. In the control condition (M = −100.40, SD = 56.35), judgments showed a statistically significant temporal repulsion effect, t(32) = −10.24, p < .001, one-tailed, d = −1.78, 95% CI [−2.33, −1.22]. Similarly, also in the causal belief condition (M = −52.50, SD = 67.24), individuals showed a significant temporal repulsion effect, t(30) = −4.35, p < .001, one-tailed, d = −0.78, 95% CI [−1.18, −0.37]. The separate shifts for the noise and the tone in the between-subject conditions are visualized in Fig. 3. Bayesian analyses ~~~~~~~~~~~~~~~~~ To assess the strength of the evidence, we performed Bayesian analyses in JASP (JASP Team, 2018, Version 0.9). In the absence of sufficient information regarding a reasonable effect size and given the novelty of the paradigm, the default Cauchy prior of 0.707 instead of an informed prior was used for the Bayesian t-tests (Bartlett, 2018; Quintana & Williams, 2018). To assure the robustness of Bayes Factors, robustness checks were conducted. The results indicated that the chosen prior did not inflate the Bayes Factors. Hence, the results of all tests can be assumed to be robust. The results of all robustness checks are reported in the Supplementary Materials. Bayesian t-test on overall repulsion To test the strength of the temporal repulsion effect, we conducted a Bayesian two-tailed one-sample t-test on the overall temporal repulsion score. The alternative hypothesis stated that the temporal repulsion effect differed from zero (zero meaning no perceptual shift). The results revealed a BF10 = 5.199e + 10 (=51,990,000,000), 95% CrI [−1.46, −0.84], indicating extreme evidence for the alternative hypothesis. Bayesian t-test on temporal repulsion between conditions Next, to test the strength of the difference between the two conditions, a one- tailed Bayesian independent-samples t-test was conducted. The alternative hypothesis assumed that the temporal judgments in the causal belief condition would be smaller than in the control condition, indicating less repulsion. The BF+0 was 25.01, 95% CrI [0.19, 1.19], suggesting that the data was 25 times more likely to occur under the alternative than under the null hypothesis. Thus, there was strong evidence for the idea that temporal repulsion was smaller in the causal belief condition than in the control condition. Bayesian t-tests on temporal repulsion in the separate conditions Consecutively, two Bayesian one-sample t-tests were conducted. Both overall binding scores were tested against a population mean of zero. For the causal belief condition, a BF10 = 183.4, 95% CrI [−1.13, −0.34], was found, indicating very strong evidence for the alternative hypothesis. For the control condition, a BF10 = 7.136e + 8 (=713,600,000), 95% CrI [−2.36, −1.13], was found, indicating even extreme evidence for the alternative hypothesis.","In this first experiment, we aimed to replicate the temporal repulsion effect between two auditory stimuli suggested by existing literature (Desantis et al., 2012; Haggard et al., 2002a). Employing a sensory based Libet clock paradigm, strictly excluding the possibility for motor action, for the first time, we empirically tested and statistically demonstrated this temporal repulsion effect (i.e., corroborated by Frequentist and Bayesian analyses). Additionally, we found that participants who held an explicit belief about the causal dependency of the stimuli on each other, perceived the stimuli as being less strongly repulsed than participants who did not hold such belief. This finding suggests an increased perceived coherence between the stimuli in the causal belief condition. However, temporal binding was still absent, indicating that the temporal integration of the two auditory stimuli was weak and thus did not produce a full coherence experience.","To test the reliability of the temporal repulsion effect in the perception of successive auditory stimuli and the attenuation thereof in the causal belief condition, a second experiment was set up. It is conceivable that temporal repulsion persisted in the light of causal belief because of the arbitrary nature of the stimuli and their associations. For example, recent research using a different paradigm suggests that stimuli that have a naturalistic link (such as the picture of a hand and a clapping sound) are more easily bound in temporal perception (Thanopoulos, Psarou, & Vatakis, 2018). Research suggests that these associative links between stimuli are acquired through learning and repeated exposure (Turk-Browne, Scholl, & Chun, 2008). Especially in the absence of motor action such learning might be pivotal for forming a stable predictive link. Hence, in an attempt to strengthen these associations, we increased the amount of learning trials from ten to sixty trials in the additional learning conditions. Participants and design The sample of the second experiment consisted of 1004 (66 females) volunteers with a mean age of 22 years (M = 22.22, SD = 2.03) who participated in the experiment in exchange for course credit or a monetary reimbursement. The study was carried out in line with the guidelines of the Declaration of Helsinki and was approved by the Ethics committee of the Faculty of Social and Behavioral Sciences at Utrecht University (ethics approval code: FETC17-124). All participants gave written informed consent. The design of the second experiment resembled the design of the first experiment with the addition of another between-subject factor. The complete design was a 2 (target: first vs. second tone) × 2 (type of trial: baseline vs. succession) × 2 (condition, between: control vs. causal belief) × 2 (additional learning, between: no vs. yes) mixed design. The addition of the learning factor resulted in two additional between-subject conditions. That is, participants could be randomly assigned to either one of the following conditions: control condition or causal belief condition (replication of the conditions of Experiment 1), control learning condition or causal belief learning condition (resulting from the addition of the learning factor to the design).","The general procedure of the experiment resembled that of the first experiment. Also the stimuli used were identical to the stimuli in the first experiment. The experiment was programmed using Eprime 2.0. The procedural differences between the first and the second experiment are outlined below. Acquisition phase The experiment began with an acquisition phase. In the control and causal belief condition, the same acquisition phases as in the first experiment were employed (for details see Experiment 1). In the additional learning condition, after having acquired the stimuli, participants completed extensive learning phases of the stimuli pairs. In total, participants completed 60 learning trials (2 × 30 trials, separated by a 30 s break). The repeated exposure to the stimuli pairs thus would strengthen the association between the stimuli. Like in the first experiment, in the causal belief condition, the causal relationship between the two stimuli of a pair was emphasized whereas in the control condition their independence was emphasized. Experimental trials (altered Libet clock trials) The control and causal belief condition were identical to the first experiment. The instructions for experimental blocks in the learning conditions were borrowed from the respective control conditions. Data analysis plan As in Experiment 1, we excluded temporal judgments that were more than ±640 ms away from the actual temporal onset of the stimulus on a trial basis (Aarts et al., 2012). Overall, the amount of excluded trials was very small (0.15% of the total amount of judgments). Furthermore, we calculated separate shift scores as well as overall binding scores per participant using the same formula as in Experiment 1. Upon inspection of the overall binding scores, one participant seemed to be an extreme outlier with an average mean judgment score of −532.38 ms. As this score was more than six standard deviations away from the average overall binding score of the remaining participants (M = −67.28, SD = 74.34), we decided to remove this person from further analysis. A visualization of the raw data including the outlier is provided in Fig. 4. We proceeded with testing the significance of the overall temporal perceptual shift. Since we had no directional hypothesis regarding the perceptual shift in the overall sample (either temporal binding or temporal repulsion), we conducted a two-tailed one-sample t-test on the overall temporal binding scores. After establishing the significance of the temporal repulsion effect in the sample, we subjected the overall binding scores to an ANOVA with causal belief and learning as between-subject factors to assess their influence on temporal perception. Following, four separate one-tailed one- sample t-tests were conducted to test the significance of the temporal repulsion shift in the separate between-subject conditions. Finally, equivalent Bayesian analyses are reported to give an impression of the strength of the evidence. Manipulation check ~~~~~~~~~~~~~~~~~~ To check that the belief in the causal relationship between paired stimuli was indeed stronger in the causal conditions than in the control conditions, we conducted an ANOVA with causal belief condition and learning as between-subject factors on causal belief. The results revealed that, on average, participants in the causal belief conditions (M = 5.6, SD = 2.24 , assessed on a 9-point Likert scale) believed more strongly in the causal dependency of the second stimulus on the first stimulus than participants in the control belief conditions (M = 3.67, SD = 2.17), F(1, 95) = 19.615, p < .001, η2 = 0.17, 95% CI [0.05, 0.30]. Learning was not found to have an influence on the causal dependency belief, F(1, 95) = 0.101, p = .751, η2 = 0.00, 95% CI [0.00, 0.04]. Similarly, also the interaction between the causal belief manipulation and learning was not significant, F(1,95) = 1.631, p = .205, η2 = 0.01, 95% CI [0.00, 0.10]. Overall temporal repulsion A one-sample t-test on the overall binding score across conditions was conducted to assess the direction of the perceptual shift and to test its statistical significance. The results revealed a significant repulsion effect of around 67 ms (M = −67.28, SD = 74.34), t(98) = −9.005, p < .001, two-tailed, d = −0.91, 95% CI [−1.14, −0.67]. Simple main effects confirmed that temporal repulsion occurred on both sides – the noise and the tone side, meaning that they were both perceptually repulsed from each other. While noises were perceived as occurring earlier on succession trials (M = 3.49, SD = 39.06) than on baseline trials (M = 35.44, SD = 41.05), t(98) = 7.805, p < .001, two-tailed, dz = 0.78, 95% CI [0.31, 0.73], tones were perceived significantly later on succession trials (M = 64.58, SD = 60.34) than on baseline trials (M = 29.25, SD = 44.54), t(98) = −6.018, p < .001, two- tailed, dz = -0.60, 95% CI [-0.90, -0.43]. Influence of causal belief and learning on temporal repulsion The overall binding scores were then subjected to an ANOVA with condition and learning as between-subject factors. Neither the main effect for condition, F(1, 95) = 0.203, p = .653, η2 = 0.00, 95% CI [0.00, 0.05], nor the main effect for learning, F(1, 95) = .061 p = .806, η2 = 0.00, 95% CI [0.00, 0.03], was significant. Moreover, also the interaction between condition and learning was not significant, F(1, 95) = 0.117, p = .733, η2 = 0.00, 95% CI [0.00, 0.05]. Also the planned comparison between the control and the causal belief condition did not yield a significant effect, F(1, 95) = 0.331, p = .566, η2 = 0.00, 95% CI [0.00, 0.06]. The causal belief effect that we found in the first experiment was thus not replicated. See Fig. 5 for a visualization of the overall binding scores in the four between-subject conditions. Temporal repulsion per condition The scores of all four between-subject conditions were negative, indicating temporal repulsion of the two stimuli from each other: control condition (M = −74.94, SD = 57.46, N = 27), causal condition (M = −62.90, SD = 68.02, N = 25), control learning condition (M = −66.01, SD = 98.73, N = 24) and causal learning condition (M = −64.37, SD = 73.30, N = 23). Separate one-sample t-tests were conducted to test if temporal repulsion in the conditions was significantly different from zero. The results were as follows: t(26) = −6.777, p < .001, one- tailed, d = −1.30, 95% CI [−1.81, −0.78] (control), t(24) = −4.624, p < .001, one- tailed, d = -0.92, 95% CI [−1.39, −0.45] (causal), t(23) = −3.275, p = .0015, one- tailed, d = −0.67, 95% CI [−1.11, −0.22] (control, learning), t(22) = −4.211, p < .001, one-tailed, d = −0.88, 95% CI [−1.35, −0.39] (causal, learning). Fig. 6 provides an overview of the separate temporal shifts for the perception of the noise and the tone as compared to baseline. Bayesian analyses ~~~~~~~~~~~~~~~~~ Furthermore, Bayesian analyses were conducted using JASP software (JASP Team, 2018, Version 0.9). Like in Experiment 1, the default Cauchy prior of 0.707 was used for the Bayesian t-tests as a lack of existing information regarding a realistic effect size did not allow for the determination of an informed prior (Bartlett, 2018; Quintana & Williams, 2018). To assure the robustness of the Bayes Factors, robustness checks were conducted. The results indicated that the chosen prior did not inflate the Bayes Factors and that all test can be assumed to be robust. The results of all robustness checks are reported in the Supplementary Materials. Bayesian t-test on overall repulsion score Like in the first experiment, to test the strength of the temporal repulsion effect, we conducted a Bayesian two-tailed one-sample t-test on the overall temporal repulsion score. The alternative hypothesis stated that the temporal repulsion effect differed from a population mean equal to zero. The BF10 = 3.979e + 11(=397,900,000,000), 95% CrI [−1.46, −0.84] indicated extreme evidence for the alternative hypothesis. Bayesian ANOVA Subsequently, we conducted a Bayesian ANOVA on the overall binding scores with causal belief and learning as between-subject factors. The results showed that, given the data, there was no evidence that the separate effects of causal belief (BF10 = 0.234) and learning (BF10 = 0.219) were a good addition to the null model explaining the extent of temporal binding or repulsion. Similarly, neither a model including both factors (BF10 = 0.051) nor a model including both factors and the interaction term (BF10 = 0.015) was favorable given the data. Bayesian t-test on the difference between the control and causal belief condition Next, to test the likelihood of decreased temporal repulsion in the normal causal belief condition as compared to in the normal control condition (replication of Experiment 1) given the data, a one-tailed Bayesian independent-samples t-test was carried out. The null hypothesis that there was no difference between the two conditions was tested against the alternative hypothesis that the temporal repulsion in the causal condition was smaller than in the control condition. The BF+0 = 0.497, 95% CrI [0.01, 0.7], suggested no evidence for the alternative hypothesis. Bayesian t-tests on temporal repulsion per condition Finally, we conducted separate Bayesian one-sample t-tests to test each group judgment against zero. The results were as follows: control condition BF10 = 47998.34, 95% CrI [−1.75, −0.75], causal condition BF10 = 250.6, 95% CrI [−1.34, −0.39], control learning condition BF10 = 12.36, 95% CrI [−1.05, −0.19], and causal learning condition BF10 = 87.08, 95% CrI [−1.23, −0.30]. That is, while there was extreme evidence in favor of the alternative hypothesis in the control and causal condition, the evidence for the alternative hypothesis was (very) strong in the two learning conditions. This suggests that the evidence for temporal repulsion was strong in all conditions but slightly less strong in the learning conditions. In short, we replicated the temporal repulsion effect of the first experiment, but we could not replicate the moderating influence of causal belief, even when participants received more extensive practice in causally linking the two auditory stimuli.","Research on action awareness consistently finds voluntary actions and their effects to be attracted to each other in temporal perception. This temporal binding effect is said to be rooted in factors inherent to intentional action – hence the name, intentional binding. Here we explored a potential confound of intentional action in producing temporal binding. Specifically, we explored the role of beliefs of causality, and cancelled out any role of action by examining temporal perception of auditory stimuli. Employing an adapted sensory- based version of the Libet clock paradigm, in two experiments we found a temporal repulsion effect for the perception of auditory stimuli in the absence of motor action. This finding supports past research that observed a similar effect (Desantis et al., 2012; Haggard et al., 2002a). However, we were able to statistically test it with a sufficiently large sample, and unambiguously demonstrated the existence and robustness of such an effect. That is, both - the noise and the sinusoidal tone - were significantly shifted away in comparison to baseline. We also found evidence for a moderating effect of causal belief on temporal binding of auditory stimuli. Specifically, in Experiment 1 we showed that making participants believe that the first auditory stimulus is causally related to the second auditory stimulus reduced the repulsion effect, suggesting an increased form of temporal integration and perceived coherence. The repulsion effect was not fully reduced, and we could not replicate the moderating influence of causal belief in a second experiment, even though we exposed participants to more training in applying causal beliefs to the auditory stimulus combinations. Hence, while we demonstrated a robust and clear temporal repulsion effect for sensory stimuli in the absence of motor action, the influence of higher order beliefs was less robust and not sufficient to cause actual temporal binding, as would be the case in intentional action. One possible explanation for the absence of a causal belief effect on temporal binding of auditory stimuli might lie in the naturalistic character of the stimuli combinations. Recent research on voluntary action, for example, suggests that naturalistic as compared to non-naturalistic links between action-events carry an implicit (learned) causality experience which makes them likely to bind to an effect (Dogge et al., 2012; Dogge, Hofman, Custers, & Aarts, 2019; Thanopoulos et al., 2018). Following this line of reasoning, the absence of temporal binding may be due to the lack of a natural association between the two auditory stimuli. Although previous research showed that spatial binding of visual stimuli could emerge rather rapidly (Buehner & Humphreys, 2010), auditory sequences might require practice for temporal binding and perception of coherence to occur (Hsu, Le Bars, Hämäläinen, & Waszak, 2015; Lange, 2009; Turk-Browne et al., 2008). We found that the number of trials that we used was not enough and that it is unclear how much prior experience is necessary to induce temporal binding between non-naturalistic stimuli links. While there is research suggesting that a few trials might suffice (Shimi & Logie, 2018), other research suggests that more elaborate prior learning is necessary (Ernst, 2007). The failure to replicate the causal belief effect on temporal binding thus indicates the fragility of building a coherent representation of a non-naturalistic combination of sounds (Grotheer & Kovács, 2014; Jacobsen & Schröger, 2001), which seems less problematic for intentional action where one produces sounds by key-presses that might be more naturalistic and more easily to acquire. All in all, the present finding might give reason to assume that intentionality has a special status in binding action and effect. As far as the evidence in the current literature can tell, temporal binding mainly occurs when people perform, observe and simulate actions (such as key-presses) to produce effects. Despite potential flaws in experimental designs and testing (Hughes et al., 2013), these overall findings strongly suggest that sensorimotor processes are vital to conscious experiences of action coherence and agency. It is unclear, though, whether the predictions involved in this sensorimotor mechanism stem from intentional action, actual body movement or action simulation, or whether temporal binding arises from multisensory causal integration that links physical sensations to predicted sensory effects (Ehrsson, Spence, & Passingham, 2004; Prikken et al., 2019). Future research might therefore design temporal binding studies that more indirectly operate on the motor-sensory system, possibly ruling out action simulation and using conditions that resemble sensory experiences (e.g., tactile stimulation of the finger) that often accompany intentional motor action during interaction with the world. Finally, the present experiments employed the Libet clock paradigm to examine how two successive auditory stimuli are temporally bound together in conscious awareness. A key feature of the Libet clock in a full temporal binding within- subject design is that one can specifically examine whether repulsion (or attraction) is caused by a shift away from the first stimulus or from the second stimuli, or both. Our data suggest that the temporal repulsion effect occurred bi-directionally, such that subjects perceived the first stimulus earlier and the second stimulus later in time than actually was the case. Repulsion in the temporal binding task thus represents an instance in which two successive stimuli are disconnected and separated in conscious experience. It is important to note that recent research uses alternatives for the Libet clock, in particular time interval estimations. Whereas these time estimations seem to yield findings that correspond with a temporal binding effect, the measure does not offer a clear inspection of the effect: time interval estimations are compared with the actual time intervals, leading to under- or overestimations of time intervals. Notably, in some studies underestimations are considered to represent temporal binding, whereas in other studies overestimations are treated as temporal binding when comparing conditions that are hypothesized to modulate binding effects (Damen, Van Baaren, Brass, Aarts, & Dijksterhuis, 2015; Kühn, Brass, & Haggard, 2013; Wen, Yamashita, & Asama, 2015). Thus, although the time interval estimation measure is practical and easy to administer, it is ambiguous in showing whether attraction or repulsion effects are at stake (for studies on time interval estimations with respect to auditory and visual stimuli links, see e.g., Humphreys & Buehner, 2009; Imaizumi & Tanno, 2019). Therefore, whereas the Libet clock paradigm has downsides on its own (Pockett & Miller, 2007), it allows for reliable assessments of the relative temporal shifts compared to baseline time perception, which may be preferred over time interval estimation to examine whether, how and when a positive temporal binding (attraction) binding or negative temporal binding (repulsion) effect will occur. To conclude, the present research clearly shows that two auditory stimuli occurring in the absence of motor actions are not subject to temporal binding but temporal repulsion, even when people believe that the stimuli are causally related. Understanding the role of causal beliefs in binding action and effect has received some attention lately, but most studies still involve motor movement, and hence, it remains unclear whether and how intentional action shapes coherent representations and a sense of agency of our own behavior. We believe that ruling out any role of motor movement (albeit overtly or covertly) might be crucial in examining alternative mechanisms for temporal binding between action and effect, such as multisensory causal binding that integrates associated information from different sensory modalities. The current research might serve as a first step in this enterprise.","S. A., R. C. and H. A. conceived the idea and planned the experiment. S. A. programmed the experiment, supervised data collection and analyzed the data. S. A. and H. M. interpreted the data. S. A., H. A. and R. C. wrote the manuscript. H. M. provided useful comments."],["Idling engines contribute significantly to air pollution and health problems. In a field study at a busy railway crossing we used the Theory of Planned Behavior to design persuasive messages to convince car drivers (N = 442) to turn off their engines during long wait stops. We compared the effects of three different messages (focusing on outcome efficacy, normative reputation, or reflection on one's intentions) against a baseline condition. With differing effectiveness, all three messages had a positive effect compared with the baseline. Drivers were most likely to turn off their engines when the message focused on outcome efficacy (49%) or reflection (43%), as compared to the baseline (29%). The increased compliance in the normative reputation condition (38%) was not significantly different from baseline. Thus, stimulating self-regulatory processes, particularly outcome efficacy, is demonstrated to have a positive effect on pro-environmental driving behavior. Theoretical and practical implications are discussed. --------------------------------------------------------------------------------","Individuals’ energy use greatly contributes to greenhouse gas emissions (e.g., Druckman & Jackson, 2016). The behavior of drivers (e.g., engine idling) has been identified as an important contributor to unnecessary emissions, and hence is a key target for intervention (Dietz, Gardner, Gilligan, Stern, & Vandenbergh, 2009). These emissions also pose a health threat. For example, in the Netherlands, car, bus, and bicycle passengers along high intensity traffic routes were all exposed to significantly more particulate air pollution than those on low intensity routes (Zuurbier et al., 2010). To address these environmental and health hazards we developed a field study to test the effectiveness of messages designed to persuade drivers to switch off their engines to reduce engine idling during long waits at a railway level crossing. Using Theory of Planned Behavior to encourage pro-environmental behavior ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Theory of Planned Behavior (TPB) asserts that attitudes, norms, and perceived behavioral control lead to intentions which proximally predict action (Ajzen, 1991, 2002). This TPB framework is useful for developing interventions aimed at pro-environmental behavior (PEB). Traffic-related behavior, despite being considered an important category of research for TPB, is only rarely investigated (Steinmetz, Knappstein, Ajzen, Schmidt, & Kabst, 2016). We focused here on norms, perceived behavioral control (via outcome efficacy), and intentions as the basis for messages aimed at encouraging a pro- environmental behavior (engine switch-off) at a long-wait stop. With respect to Abraham and Michie's taxonomy (2008), our intervention falls in between a “persuasion” and a “motivation” behavior change method, which were identified as quite effective methods when paired with TPB framework (see Steinmetz et al., 2016). Norms and reputation To invoke norms, we designed a persuasive message appealing to social reputation (Emler, 1990). Indeed, norms are often invoked by signaling the reputational relevance of behavior (Abrams & Hogg, 1990), which plays a key role in determining whether the behavior will be adopted by the individual. It has been suggested that people are more willing to act in a prosocial (including pro-environmental) manner when they can acquire a reputation for doing so (Roberts, 2012). For example, participants invested greater sums in a climate protection program when their investment was made public, i.e. when they gained social reputation (Milinski, Semmann, Krambeck, & Marotzke, 2006). In a similar vein, university students reduced their energy consumption on campus when a delegate provided them with, and commented on, their consumption feedback, but not when the feedback was provided electronically and privately (Emeakaroha, Ang, Yan, & Hopthrow, 2014). Finally, avoidance of publicly deviating from ingroup PEB norms can also increase PEB (Player et al., 2018; see also; Goldstein, Cialdini, & Griskevicius, 2008). For the “normative reputation” condition, we therefore developed a sign that asked drivers: “When barriers are down, turn off your engine to show others you care”. This designed to subtly induce a reputational motivation (to be caring) to engage in PEB. Perceived control and outcome efficacy Our second persuasive message relied on the attitude component of TPB, specifically, evaluation of behavioral outcomes. Attitude toward the behavior is driven by beliefs about the consequences of performing the behavior, such as whether their behavior will have a positive outcome (Hausenblas, Carron, & Mack, 1997; McEachan, Conner, Taylor, & Lawton, 2011). PEB can often be regarded as a ‘drop in the ocean’ (Kerr, 1996; Lorenzoni, Nicholson-Cole, & Whitmarsh, 2007), whereby an individual may feel that their single contribution will not have any impact, and therefore there is no point in acting. In contrast, outcome expectancy highlights the importance of an individual behavior in leading to a particular outcome (Doherty & Webler, 2016). Efficacy over the outcome can have direct, positive, and significant influences on public behavior about climate change (Doherty & Webler, 2016) and recycling behavior (Lindsay & Strathman, 1997). Assuming that all drivers are capable of turning off their engines, the relevant goal for perceived control is their belief that they can affect air quality (i.e., outcome efficacy), rather than the lower level ability to turn off their engine. Using outcome efficacy as the foundation, we therefore developed a message that asked: “Please switch off your engine when barriers are down. You will improve air quality in this area” (hereinafter labelled ‘outcome efficacy’). This message aimed at showing the positive consequence of a simple action that was within the drivers' control. Intention to act and self-reflection The third persuasive message addressed intention to act. TPB asserts that intention is the closest predictor to behavior, with attitudes, norms, and perceived behavioral control as antecedents. However, an intention-behavior gap persists (see Sheeran & Webb, 2016), in part because people often forget to perform the behavior (Sheeran & Orbell, 1999). In cognitively demanding situations that involve multiple goals (such as driving), depleted cognitive resources can disrupt the link between an intention and its behavioral implementation, meaning that despite pro-environmental attitudes (and intention), the behavior is not actioned (Steg, Bolderdijk, Keizer, & Perlaviciute, 2014). Reminding people to consider their intentions could foster the enactment of consistent behavior. Indeed, studies found that simply questioning people about their behavioral intentions increased subsequent adoption of the behavior (mere measurement effect; see Godin et al., 2010; Todd & Mullan, 2011). Using a question to focus driver attention on their own behavioral intention about a salient action (switch off engine) should increase action. For this “reflection on intention” condition, the message therefore asked drivers: “When barriers are down do you intend to turn off your engine?”.","The study complied with ethical standards of the British Psychological Society and American Psychological Association, and was approved by the School of Psychology Ethical Review Board (# 20122491). We collected data on engine switch-off rates at a busy long- wait stop at the time of displaying one of three intervention signs (compared to baseline). We hypothesized that all messages would improve behavior relative to baseline and additionally explored differences in effectiveness between messages.","Data collection took place on 14 different days spread on a 6-month period (October 2012 to March 2013). Collection lasted for 1 h at the time, between 8am and 6pm, Mondays to Saturdays. We randomly varied time of collection in order to reduce the chances that the same driver would be sampled repeatedly (e.g., while commuting to work), and so that intervention conditions would not be confounded with time of testing. Overall, 565 vehicles were sampled, a large majority of which were cars (n = 442).1 Location and setting Engine idling was observed at a busy level crossing in a medium-size city in the UK. The local council had erected a permanent sign asking people to switch off their engines that was visible before, and throughout the duration of the study (see supplementary material). The intervention messages were printed on a placard (420 × 594 mm; font type = Franklin Gothic medium, font size = 100 pt), 2 m above ground level. To ensure that all vehicles would pass the sign on their approach to the level crossing, the placards, held by stationary research assistants on the sidewalk, were positioned facing traffic on each side of the crossing, 75 m before the barrier and approximately 5 m from the existing council signs (futher detail on the methodology can be found in supplementary materials). While the barrier was down and the vehicles were stationary, another research assistant walked along the sidewalk from the barrier toward the queuing traffic as far as the sign, discreetly recording whether each vehicle's engine was on (coded 0) or off (coded 1). This was assessed by viewing exhaust activity and listening for engine noise emitted from each vehicle (see supplementary materials for interrater reliability). Intervention messages We used three intervention messages: Normative reputation, Outcome efficacy, and Reflection on intentions, which were compared to a Baseline condition where no message was present (see Table 1). Control variables It is possible that drivers' behavior is influenced by the weather (e.g., less likely to turn off the engine on a hot day where the car AC is on). To account for this, we coded for each observation period the weather (rainy = −1, sunny or cloudy but dry = 1) and the visibility (dark or foggy = −1, visible = 1). Since a person's behavior can be impacted by the presence of others, we additionally recorded the number of people in the vehicle (M = 1.55, SD = 0.76). Most drivers were alone (58%) or accompanied by one passenger (32%), and the remaining vehicles had three or more people aboard.","We conducted a hierarchical logistic regression to analyse the impact of the intervention messages on drivers’ idling behavior. In a first step we included only the type of intervention as a predictor. To test more accurately the effect of the three interventions against the baseline, we computed the following contrast (hereinafter C1): Baseline = −3, Outcome efficacy = 1, Reflection = 1, Reputation = 1. We also entered the orthogonal contrasts in order to better estimate residuals (C2: Baseline = 0, Outcome efficacy = −2, Reflection = 1, Reputation = 1; C3 = Baseline = 0, Outcome efficacy = 0, Reflection = −1, Reputation = 1). In a second step, we additionally entered the three covariates: weather, visibility, and number of people in the car. The analyses revealed a significant effect of the intervention messages. Relative to the baseline, drivers turned off their engine significantly more after exposure to any of the three messages. This effect held in the absence and presence of the covariates, which did not impact engine idling (see Table 2). Percentages of drivers that turned off their engine per experimental condition are illustrated in Fig. 1.2 To complement these results, we additionally tested the difference between each message and the baseline condition. The difference was significant for both Outcome efficacy, OR = 2.30, 95% CI [1.26, 4.18], Wald's χ2 = 7.42, p = .006, and Reflection on intention, OR = 1.83, 95% CI [1.02, 3.30], Wald's χ2 = 4.04, p = .045. It failed to reach significance for Normative reputation, OR = 1.49, 95% CI [0.81, 2.73], Wald's χ2 = 1.62, p = .20. The present results ~~~~~~~~~~~~~~~~~~~ In this experiment we tested the effectiveness of interventions aimed at encouraging PEB in a context where they are commonly rare. Drawing from TPB (see Steinmetz et al., 2016, for a recent meta-analysis), we hence tested the effectiveness of three messages (normative reputation: “… show others you care”; outcome efficacy: “… you will improve air quality”; reflection on intention: “… do you intend to switch off? “) to encourage drivers to turn off their engines at a long wait stop. Our aim was to test whether focusing on different aspects of behavioral regulation could be effective in triggering more environmentally positive behavior. Initial results showed that all messages led to higher compliance (higher rates of engine switch-off) than the baseline. However, analyses also revealed that not all messages had the same impact: despite descriptively increasing compliance (by 8.8 points), the reputation message did not significantly differ from baseline, whereas reflection and outcome efficacy did (by 13.8 and 19.4 points, respectively). Hence, the study shows that the mere presence of a person carrying a persuasive sign on the sidewalk is not enough to impact drivers’ behavior: the content of the message is essential. The normative reputation message aimed to motivate drivers based on opportunity to enhance their reputation by showing others they care (Steg et al., 2014). However, our results suggest that this was not enough to elicit behavior change. It is possible that drivers were not convinced that others would be able to detect their PEB, and hence that there was in fact little reputational gain to be achieved. Moreover, the message may have caused some dissonance for drivers with strong self-interest or hedonic values, and thus reduced the likelihood of action since the messaging opposed their existing belief system (Kollmus & Ageyman, 2002). This seems consistent with evidence that PEB messages can be effective when they target self-interest motives (Van de Vyver et al., 2018) or appeal to drivers with stronger hedonic or egoistic values (Steg et al., 2014). The reflection on intention message aimed to target drivers with preexisting intentions to behave pro-environmentally who might not do so in certain settings because they are distracted or forget to enact their intentions. This message elicited significantly higher behavioral compliance than baseline, which is consistent with the mere measurement effect (Godin et al., 2010; Todd & Mullan, 2011). However, the nature of the message does restrict its impact to receivers who already hold pro-environmental intentions. Drivers who are relatively unconcerned about the environment would be less likely to harbor intentions to turn their engines off, and therefore would not be influenced by the message. This might be especially important in the particular context of this study, since factors such as status and comfort have been found to influence behavioral choice (Gatersleben, Steg, & Vlek, 2002), so switching off the engine and losing other functions (such as heating/AC and radio) may outweigh the PEB. Finally, the outcome efficacy message aimed to increase drivers' sense of efficacy and remove uncertainty regarding the effectiveness of their (non-idling) behavior. This message elicited significantly higher behavioral compliance than baseline and was, descriptively at least, the most effective message, increasing compliance by 66%. This remarkable effectiveness could be explained by the fact that the message went beyond the heightening of values and normative behavior by delivering a specific and positive message that the driver's personal contribution would make a difference to the environment (Doherty & Webler, 2016; Hine & Gifford, 1996; Meyerowitz & Chaiken, 1987; Morton, Rabinovich, Marshall, & Bretschneider, 2011; Webb & Eves, 2007). Limitations, future research and conclusions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This study has some limitations that must be acknowledged. First, observations were limited to one location. Future studies will need to ensure that the results can be replicated in different places. Moreover, the procedure relied on the presence of a research assistant to hold the placard. Even if past studies have ensured that the mere presence of a person on the sidewalk (holding a blank sign) did not significantly affect the rated of engine idling (Meleady et al., 2017), it may have somewhat altered the drivers’ usual behavior. Future studies should therefore investigate the effect of messages attached to fixed poles. Finally, given the non-intrusive nature of the study, we assessed behavior but could not measure its psychological antecedents. Hence, we cannot know of sure whether the messages impacted the targeted cognitions and/or others. It will be valuable to conduct additional experimental studies that test the impact of the messages on the relevant cognitions – but this goes beyond the scope of the present field study. In sum, this study makes a substantial contribution by showing that behavior change towards PEB is sensitive to the way that persuasive messages tap into particular routes for self-regulation. It also raises interesting questions for future research about how different approaches might work in particular contexts and about how different aspects of self-regulation might feed into one another. The most effective message, outcome efficacy, led to a decrease of engine idling by 19.4% (corresponding to an increase in compliance by 66%). Relative to the usual traffic at the level crossing where the experiment took place, this ismpact would prevent the emission of 3538 tons of CO2 per year.3 We believe that this can still be increased and the message made even more effective, notably by giving additional, specific details on the efficacy of the behavior (Steg & Vlek, 2009; Webb & Eves, 2007). For example, adding a specific outcome (e.g. engine switch off would improve air quality by X%) or by providing some sort of feedback to drivers (e.g. each car that switches off reduces pollution by X amount) could boost the effectiveness of the efficacy message still further."],["The formation of urban social order has been studied since the classic accounts of the modern city (e.g. Park, Simmel, Weber, Goffman; see also Ocejo and Tonnelat, 2014). In studies inspired by the ‘mobility turn’ in sociology (Brighenti, 2012, 404), public transportation has been analyzed as a space for conviviality as well as a symbolic platform for struggles and protests over space and rights, particularly for racialized minorities (Fleetwood, 2004, 36–37; Mitchell, 2003). This article studies public transport from the perspective of 15- to 17-year-olds, who are often categorized as teenagers. Their views have rarely been taken seriously when analyzing urban social order (Ocejo and Tonnelat, 2014, 494; Tironi and Palacios, 2016), hence leaving us blind to young people's everyday struggles with their entitlement to belong in the city (Maira and Soep, 2005). With regard to young people's urban engagements, it is necessary to acknowledge the recent development and accessibility of mobile technologies and the expansion of social media platforms. Young people aged 15 to 17, the so-called Post-Millennials (also known as Generation Z or iGen, who have come of age in the first two decades of this millennium), are growing up amid ongoing, hectic change shaped by media and information technology advances. Finland, which is in focus in this article, ranks second only to Japan in terms of mobile broadband penetration (OSCE, 2017). More than 95 percent of Finns aged 16 to 24 own a smartphone (Statistics Finland, 2016a), and 92 percent of young people regularly use smartphones to play games online, listen to music and watch videos, while 89 percent use them to access social networks (compared to 9% of 65- to 74-year-olds; Statistics Finland, 2016b). Despite the widespread use of media technology, very few studies on new media have focused on young people's media consumption in relation to their mundane engagements in cities or other youth cultural contexts (Buckingham and Kehily, 2014, 10–11). In general, the implications of mobile technologies for social interaction in public places have not yet been thoroughly explored (e.g. Hampton et al., 2010). Considering the above, this article examines the dynamics of social order in the contemporary digitalized ‘media city’ (Georgiou, 2013) from the perspective of young generations. Social order is formed interactively in the everydayness of the city and through the processes of social control. By social control, we refer to forms of social interaction through which norms regulating human conduct are created, monitored and sanctioned. We conceptualize social control as being inherently interwoven into young people's intra- and intergenerational relations, having both formal and informal features (Honkatukia and Keskinen, 2018). In the media city in particular, social order involves technological mediation and is constituted via mundane encounters ‘with social diversity in a situation of spatio-temporal proximity’ (Brighenti, 2012, 46). Empirically, we have focused on the Helsinki metro and have interviewed young people who use it daily. The metro represents for us a highly technologically driven site where Internet, digital technology and mobile devices are deeply embedded in the everyday life of the passengers (Hatuka and Toch, 2016). Moreover, the metro is a site of both formal and informal surveillance and control (CCTV, gates, security lighting; Gardner et al., 2017; Lianos, 2012). Even if formal surveillance has been intensified recently in many cities due to the securitization and militarization of social control (e.g. Hussain and Bagguley, 2012; Franko Aas, 2007, 63–65), the Helsinki metro, perhaps due to its small size and its globally remote position, lacks many of the features of formal control that characterize larger cities. In the Helsinki metro, there are no gates with mechanical barriers as there are in most other metro systems, no specially dedicated police department (unlike in Washington, D.C. or throughout the Russian metro network) and no airport-style security with metal detectors and X-ray scanners as in St. Petersburg or Moscow. In the technological and automated Helsinki metro system, social control is often manifested through informal and non-verbalized means as well as voluntary self-control of passengers, deriving either from their willingness to conform, imitation, learned routines or contagiousness of habits. Some omnipresent tensions concerning social control become visible in the Helsinki metro due to its concentration of mechanical and digital technology in the closed spaces of stations and metro carriages (also Lianos, 2012; Zaporozhets, 2014). A specific feature of this context is that during transit, social control and norms are often negotiated in interactions with unknown people situationally and often in short-term encounters. Moreover, in the closed space of metro carriages moving through tunnels underground, human interaction with various technological objects (travel card readers and mechanical doors) becomes an important part of the metro experience (Lianos, 2012). Hence, metro carriages and stations are anthropological spaces for complex emotional and sensory experiences, social engagement and symbolic self-expression of which mechanical and digital devices are essential parts (e.g. Hayward, 2004). We are interested in young people's experiences of (mainly informal) social control in this technologized context. Our inquiry assumes that young people's encounters with adults, including those encounters related to social control, bear importance for creating a sense of belonging, citizenship and being legitimate members of the city. The article will also shed light on the dynamics behind earlier research findings from Finland claiming that public transport is seen as an ambiguous space for young people. Young people's experiences of threatening situations on public transport are rather common in larger cities (Kivivuori et al., 2014; Ojanen, 2018, 26–27), and it is also common to regard the metro as the least-safe option compared to other forms of public transport in Helsinki (Tuominen et al., 2014, 36–38). In the following section we will explore more closely the specificities of the metro as a technologized form of public transport, and present the theoretical underpinnings of the study. In the analysis of social order and control, we were inspired by Pierre Bourdieu's (1977) idea of symbolic violence (as applied by e.g. Bourgois, 2001 and Listerborn, 2015), as well as by conceptualizations of the digitalization of urban life (e.g. Hatuka and Toch, 2016; Georgiou, 2013; Brighenti, 2012; Hampton et al., 2010). This framework enables us to reflect on how everyday life experiences relate both to intergenerational power structures and the digitalization of urban spaces, and how these two intertwine. Thereafter, and based on thematic interviews conducted with a sample of young people aged 15 to 17, we will analyze their narration on encounters with adults on the metro, especially those related to social control and processes of symbolic violence. In the concluding section we reflect on the nature of the intergenerational social order in a digitalized media city based on our empirical inquiry.","On the metro, one is usually unable to choose one's company or the people one is surrounded by. Events change rapidly, and interaction with others is usually brief. Compared to trams and buses, a greater number of people share the same space on the metro, multiplying the possibilities of coming across people with diverse backgrounds in a physically restricted space and often crowded conditions (Ocejo and Tonnelot, 2014; Gardner et al., 2017). With larger numbers of passengers on the metro, one is less likely to meet the same people on the same routes at the same times or to be acquainted with their habits beforehand. Due to the high level of automation and digitalization, passengers rarely encounter metro personnel or guards in Helsinki. On buses and trams, by contrast, contact between passengers and drivers is almost unavoidable. On most buses, for instance, the drivers check that every passenger who boards has paid the fare, and they can directly intervene in any situations that may occur. On the metro, formal control is mediated by technical means such as video surveillance, and there are only random checks by security workers and ticket inspectors. This setting distinguishes the metro from other forms of public transport and determines the significance of informal control, which largely rests on conventions, shared norms, expectations and negotiation. Age, gender and cultural background exert significant impacts on social order. Accordingly, we approach the metro not only as a mechanical or technological mode of transport but also as a social, experienced and affective urban space in which digital devices, and through them virtual realities, play a part. The metro is not merely a vehicle that takes passengers from A to B, nor is a metro carriage or station merely a ‘non-place’ (Augé, 1995). On the contrary, in these spaces, people gain experiences, learn about and take part in public social order and are socially, virtually and emotionally engaged (see Zaporozhets, 2014; Thibaud, 2015; Hatuka and Toch, 2016). Their experiences and emotions are deeply interwoven with interactions with other people of different ages and from diverse backgrounds.1 We concur with Carina Listerborn (2015, 97), who studied veiled women's experiences in Malmö, Sweden, and who suggests that ‘encounters between strangers, as different users of public spaces, are one of the core subjects for discussion in relation to orders in the public space’. Regarding the metro in particular, we align ourselves with Richard E. Ocejo and Stéphanie Tonnelat (2014, 494), who regard the metro as a unique public space in which social order is negotiated with strangers, namely those who are located spatially close but socially distant from each other (following Simmel's (1964) notion of ‘the stranger’ as an inherent aspect of city life). The presentation of strangeness is an essential feature in this context, involving informal norms and their guidance through social control (Ocejo and Tonnelat, 2014, 495). We study social order and control in the highly-digitalized metro environment, which is equipped with wireless infrastructure and provides broadband Internet access to mobile devices. Tali Hatuka and Eran Toch (2016, 2195–2196) referred to the use of mobile devices as ‘a portable private- personal territory’, which is characterized by permeable boundaries between the private and the public or the material and the virtual in daily life places, as smartphones and devices connect their owners to multiple webs of relations. In the neo-liberal discourse, the appearance of mobile devices and new media is often celebrated as a means of making a city more accessible, flexible and smooth for its citizens, providing them with increased safety through electronic connectivity (Brighenti, 2012, 408; Hampton et al., 2010, 701–702). As alleged digital natives, young people's competencies are often believed to empower them as they are able to teach the older generations how to use the newest apps or play popular games (Giddings, 2017; Mäyrä, 2017). Not exclusively adhering to these discourses, we are interested in young people's subjective understandings of the nature of social order on the metro, which serves for us as an example of a highly technologized and digitalized social system of the media city. We understand social control as a necessary feature of social life that makes social interaction smooth and predictable. Yet at the same time, it involves power relations and possibilities for oppression as well as conveys messages on the conditions of belonging versus exclusion (Von Scheve and Von Luede, 2005). Through addressing social control, we examine what young people's accounts about travelling on the metro reveal about their entitlements to using urban spaces (Listerborn, 2015; Maira and Soep, 2005). From this perspective, we will focus on young people's stories about their encounters with adults, primarily those encounters characterized by ambivalence, and analyze the meanings created within the framework of symbolic violence (Bourdieu, 1977; Bourgois, 2001). Integral to the notion of symbolic violence is the premise that dominant views are perceived from a limiting perspective, which portrays social order as natural and self-evident, thereby concealing the structural power relations behind it. Our aim is to analyze personal narratives to understand the formation of the intergenerational order, to determine the role of digital devices in this context and to provide insights into young people's entitlements in the contemporary media city.","This study focuses on 15- to 17-year-olds, who lack many formal rights (in Finland, 18 is the age of majority and the voting age, while 15 is the minimum age of criminal responsibility and for other legal actions like concluding an employment contract) but for whom the metro provides the possibility of exploring the city independently of parents. In public discourse, the ‘underaged’ are commonly depicted with ambivalent perceptions as either victims or perpetrators of disruptive acts, especially in public places (e.g. Fleetwood, 2004). We are distancing the study from these depictions by carefully analyzing young people's subjective experiences of social control. The empirical material originates from interviews conducted during the Helsinki City Youth Department project in 2016 and 2017. The City of Helsinki recruited young people for a period of one month each summer to work on an art project, Ole hyvä Helsinki! (You're Welcome, Helsinki!), which is part of the city's branding campaign, Brand New Helsinki (http://brandnewhelsinki.fi/). Young people produced art for metro stations and worked in workshops on city premises (youth centres, art centres) near several metro stations (in 2016, four groups worked on four metro stations, and in 2017, the city expanded the project to 10 stations). A total of 31 interviews were conducted in 2016–2017 with 57 young people aged 15 to 17 using semi- structured individual or group interviews. The interviews were voluntary, and participants could choose whether they wanted to participate in a group (2–4 participants) or individually. Only a couple of the young people who had engaged in the city workshops in the centres we visited declined to participate in the study. The main themes of the interviews concerned everyday metro experiences, the use of apps and devices in transit, mobility in the city and neighbourhood identities. The participants represented various social strata and lived in neighbourhoods near the workshops where they were working. The majority (80 percent) of workshop participants were young women, with one participant describing their gender as non-binary, while around 10 percent represented ethnic or racial minorities. By and large, groups in Helsinki city centre mostly consisted of people from middle-class backgrounds. We do not, however, regard this as a problem, considering that many previous studies on youth engagements in the urban sphere have focused on young people from working-class backgrounds (e.g. Kehily and Nayak, 2014; Skelton, 2001). The interviewees from eastern neighbourhoods, however, had more diverse socio-economic backgrounds. Narratives describing encounters with adults were extracted from the transcribed interview data, and through a process of close reading, we identified themes including stories and reflections on norms and expectations concerning young people's conduct on the metro. We also paid attention to young people's descriptions of their use of digital devices while in transit. The studied accounts represented actual examples of the formation of social order in the digitalized and technological metro setting as articulated in the participants' own words (Ocejo and Tannelot, 2014, 499–500).","In almost every interview, the participants talked about metro carriages as mundane, convivial spaces, which hosted very limited social interaction. According to the interviewees, metro carriages are silent spaces where one is not expected to make a loud noise, stare at other people or intrude into other people's personal space. This was termed ‘the norm of silence’ or ‘light sociality’ by our participants, and it has been theorized by Goffman (1963), for example, as ‘civil inattention’, in which the presence of others is visually acknowledged without focused interaction or verbal communication (Ocejo and Tonnelat, 2014, 495, 497). Sometimes it can be like that. For instance, you bump into a stranger, apologize, everyone smiles and then ‘okay, goodbye’. Or, okay, someone dropped something, forgot something and then you kind of ‘hey, you dropped your umbrella’, ‘okay, thanks’ … [you] don't necessarily go to talk to a stranger because you don't know whether [he/she]2 wants to talk, what kind of person [he/she] is. You don't know whether [he/she] will look at you strangely if you start talking and everything … This kind of norm of silence is very hard to break. Even if you are not afraid of embarrassing situations, somehow it does not feel natural. (Neo, 17, male) The norm of silence, as described by Neo above, allows the passengers to be ‘in their own bubble’; to lose themselves in thought, look through the window or daydream. Most of the young people we interviewed admitted to routinely entertaining themselves in their ‘portable private-personal territory’ (Hatuka and Toch, 2016) by scrolling social media, chatting with their friends online or watching their favourite series via streaming services. Some even related to the metro carriage as ‘a space of productivity’ (Hampton et al., 2010, 714), as they described how they used the metro ride for doing their homework with their devices. Even though some young people were very critical of the extensive use of smartphones and social media in public places, many could not imagine a tedious metro journey without engaging in private activities on- screen, making visible the nature of the metro space as a media city. Sixteen-year-old Riina, for example, referred to this somewhat self-ironically as she admitted to having been ‘terrified’ every time her phone ran out of power on the metro. Furthermore, immersing oneself in virtual realities while travelling was sometimes portrayed as an actively chosen means of shutting out the ‘background noise’ in order to ‘be somewhere else’ – as explained by 17-year-old Aino. Hence, the young people's accounts illustrate the ways in which digital technology is explicitly used to reduce attentiveness to one's surroundings (Hatuka and Toch, 2016; Ocejo and Tonnelat, 2014, 504). Despite the strong norm of silence, most of the young people admitted that they had sometimes been approached by adults or engaged in brief exchanges with them while travelling on public transport. In the interview accounts, the adults who approached young people were either referred to as ‘oldies’ or ‘different’ in some way, such as alcoholics. By contrast, ‘people like mothers with small children’ or those going to work were, according to the interviewees, very task-oriented within public spaces and too preoccupied with their own matters (including their private digital bubbles) to engage in conversations. The exchanges with ‘grannies’ or ‘grandpas’, as the older passengers were called, were often described neutrally or in positive terms as amusing encounters or pleasant chats, emphasizing the convivial nature of travelling. Similarly, listening to intoxicated adults' stories was often seen as funny and entertaining, and being drunk in public was not regarded as wrongful for the most part. By contrast, many saw substance users as an inevitable and normal feature of urban life. In general, the young people felt that they were treated well or at least ‘neutrally’ on the metro compared to everyone else. Some stated that any possibly negative reactions were caused by the bad behaviour of young people themselves. Notwithstanding individualized tensions, the interviewees also challenged the strict division between their generation and adults into two, easily identifiable and inherently oppositional groups (Kallio, 2017; Fleetwood, 2004, 35; Ocejo and Tonnelat, 2014, 500). Instead, a distinction was made between well- and badly-behaved people, with the latter including noisy and unruly people of different ages. This division is visible in the following account, in which the young man positions himself as an ‘invisible’ rider (Ocejo and Tonnelat, 2014, 498), hence conforming to the atmosphere shaped by civil inattention discussed above: There are young people who behave well and there are noisy young people. [ People] always look angrily at those who are very noisy. I even do it myself. As for me, no one ever looks at me. I am kind of invisible; I never cause any kind of disorder, so no one ever looks at me. (Alex, 17, male)","Aside from the emphasized neutrality of contact with adults, some more ambiguous situations with adults were mentioned, including experiences of sexual harassment, racism and being labelled negatively according to stereotypes. The civil inattention discussed above called for flexibility from passengers as well as the need to overlook norm violations on occasion (Ocejo and Tonnelat, 2014, 497), which our interviewees were aware of. At the same time, the problematic nature of this kind of overlooking was perhaps most vividly apparent when some of the young women discussed their experiences of sexual harassment on the metro. Several of them admitted to having experienced it, as explained by the following interviewee: It was precisely on the metro where the incident occurred, and it was very distressing when someone started to follow me. He followed me for 15 minutes or so, then I kind of … very nice [ironical tone], he shouted after me that I was really beautiful and everything. Yeah, those are very distressing experiences for me at least. (Minna, 17, female) These incidents were not regarded as violence, nor were they seen as particularly scary situations – rather, they were viewed as occasions during which ‘nothing really happened’ (Stanko, 1990). As such, they were acknowledged as an inevitable part of everyday life for many girls and younger women in the city. One of the young women, for example, recalled a recent and, according to her, ‘minor’ yet irritating incident, and commented that it came to mind only because sexual harassment was discussed in the interview. Some of the young women related how they had grown accustomed to this kind of conduct as early as age 12 or 13. The self-evident nature of the young women's experiences of sexual harassment had come as a surprise to some at first, especially when being among other people in a metro carriage is not perceived as a situation in which one would expect intrusions into bodily integrity (see also Listerborn, 2015, 110). The perpetrators were often described as middle-aged men (or those in their 30s) who were sometimes – but not always – under the influence of alcohol. They were described as looking for some sort of sexual contact with young girls. Unwanted advances consisted not only of verbalized messages or touching, but also took the form of intense staring or sexual gestures. The men's comments about beauty or good looks were generally regarded as unacceptable and unwanted as opposed to flattering. Once I was going to school, and there [on the metro] was someone, a sort of, some guy was staring at me and smiling. And when I was about to leave, he winked at me. I just looked at him in astonishment. I was going to school! It was kind of amazing [said ironically], yes. (Olivia, 17, female) Reactions to harassment and boundaries of tolerance are decidedly contextualized. Fifteen-to sixteen- year-olds are highly sensitive to harassment in urban public places – particularly on public transport and by older males (Aaltonen, 2017). However, in the social order of the metro, shaped by civil inattention and passengers' immersion in their private digital bubbles, these experiences become either invisible or tolerated. According to the interviewees, nobody intervened as a rule, and unless something very exceptional occurred, no help could be expected from fellow passengers. Hence, they emphasized that they could only rely on themselves and develop skills to cope with such situations. The following incident was recounted by 16-year-old Ruska, who described their gender as neutral (not adhering to the strict binary division of male or female) and recalled being harassed by ‘lads’. The feeling of being left to cope alone during the incident is aptly summed up by Ruska below: Ruska: Once, some guy just put his hand on my knee and started to move his hand, and then I kind of really hit his hand and left. Ruska: And then people sort of noticed that something had happened, but they did nothing; they kind of said nothing to this person. Interviewer: So, no one intervened? In addition to sexual harassment, quite a few interviewees, regardless of gender, mentioned being accused of various issues related to their age while on the metro, such as loitering, being too noisy or too preoccupied with their gadgets, or of alleged bad behaviour in general. A frequently mentioned example was when the young people were watched to ensure that they did not occupy the priority seats (i.e. seats reserved for the disabled or elderly). These encounters were not always described as explicitly confrontational but generally consisted of angry gestures or scowls, thus making visible the norms and general attitudes towards young people in these situations (Fleetwood, 2004, 36). One of the young women said that she was always a bit paranoid about this and was constantly prepared to encounter negative comments or ‘shouting’. Hence, her strategy was to avoid carriages occupied by older passengers. A male respondent regretted not being able to respond with a witty retort to a ‘grandpa’ accusing him of sitting in the ‘wrong seat’. This was because he had been surprised by the elderly man's reaction, particularly seeing as the carriage had been half empty at the time. An old man came and tapped me on the back and said, ‘When you're young, it's good that you can sit [on public transport], so when you get old you'll be able to stand’. Then I asked, ‘Do you want to sit here?’ He said that he didn't. He told me to go ahead and sit if it was so important for me. In retrospect, I should have said that yes, I actually wanted to sit down. But in that kind of situation, witty replies seldom come to mind. Besides, it's probably better not to start arguing, even if it sometimes crosses your mind to answer an inappropriate comment in an equally inappropriate way. (Satoshi, 17, male) Racist insults and comments about young people's personal style or outward appearance were also mentioned as examples of unpleasant experiences on the metro. In the following excerpt, the young woman described how she had been upset and angry about the impolite attitude of an older woman she called ‘granny’, who commented on her looks: Once there was a granny who started to look at me in a kind way and smiled, so I smiled back. Then she suddenly said, ‘You have very ugly eyebrows; why did you paint them like that?’ Just like that; smiles politely and then simply tells me that I look ugly and to go and wipe some of it off. (Lotta, 17, female) Outward appearance, clothes and makeup are important means of identity-building for young people (McRobbie, 1997), and while often influenced by the fashion industry, they are highly personal. Hence, comments concerning one's looks can be particularly insulting. What was even more upsetting for Lotta above was that the comments were uttered by a kind-looking person who was smiling at the same time (cf. Listerborn, 2015, 109). Later in the interview, she emphasized her attempts to make a good impression by being polite and avoiding swearing in public, and regarded the comments by the older woman as unfair in this respect, too. Like the young man above, Lotta felt unable to talk back. Responding in an impolite way would have reaffirmed the negative stereotype of young people as troublesome, as she explained: So, I don't really understand; you can't say ‘mind your own business’ or anything like that because there is still some kind of authority-politeness kind of position. I was expected to apologize for my ugly eyebrows and then leave. (Lotta, 17, female) Based on the respondents' accounts, it can be said that young people are expected to play a part as ambassadors of youth (cf. Listerborn, 2015), which prevents them from defending themselves, as the following story by Olivia reveals: I was walking out of the metro with friends when a very rough woman just pushed me for walking too slowly because I didn't want to push the person in front of me. I made a comment of some sort, and then she began complaining that young people are blah blah blah. That's what young people get for being nice and giving up their seats to everyone on the bus, metro and so on but still get blamed for not respecting older people. (Olivia, 17, female) The above examples are revealing when it comes to young people's ambivalent positions in the social order of the contemporary city (see Ocejo and Tonnelat, 2014, 498). Some participants described using the metro as uncomfortable and stressful because of the possibility of being treated unfairly, condescendingly or rudely by older passengers. The respondents seemed to think that a double standard exists in intergenerational relations: older people are entitled to comment on or even touch young people, but young people are expected to be well-behaved and polite towards adults (also Listerborn, 2015, 108). For this reason, the participants felt that their ability to negotiate was limited, particularly with older people they encountered on the metro.","Based on the interviews, avoiding conflicts with adults seemed to be a major concern for some young people on the metro. In the context of a media city, the interviews revealed several significant patterns in this respect. First, young women in particular seemed to use their portable devices (smartphones, headphones, tablets) not only for entertainment and communication but also to secure privacy and to claim and safeguard their personal space (see Fleetwood, 2004, 38). Some referred to their gadgets or music as ‘a getaway’, meaning both the ability to distance oneself from one's surroundings and to protect oneself from unwanted advances, especially from adult males. The following excerpt exemplifies this: Well, usually I travel on the metro with a phone in my hands. I listen to music or watch something, check Twitter or apps like that. So, if someone comes and actually speaks to you, some random dude, then you can just pretend that you aren't listening to him or haven't noticed him. That's what I always do. It's a kind of safety issue for me, meaning that I don't necessarily need to encounter people. (Minna, 17, female) Second, the respondents reported that they resorted to their smartphones or other gadgets (i.e. headphones to listen to music) especially if they felt that there were potentially intrusive passengers around or if they were travelling alone. Hence, they used digital devices as part of their personal safety routines (Stanko, 1990) or ‘street wisdom’ (Anderson, 1990) on and around public transport. Communicating with someone (on the other end of the line), for example, was portrayed as one of the tricks young female travellers used to overcome stressful situations. This is described below by Olivia. Talking on the phone after leaving the metro station was described in her account as a safety routine that increased her confidence about travelling in public. I use a bus line that starts outside the metro station. One Saturday evening at the bus stop, I realized that a man was staring at me. I thought, okay, now [I] should do something, and I decided to call a friend. I talked to her during the entire bus trip home. I felt safer since I could not have been attacked while on the phone, especially seeing as I told my friend where I was and some other general information just to be on the safe side. (Olivia, 17, female) Considering the data, it seems that the young women's digitalized coping mechanisms were individualistic, as they were adopted in an atmosphere shaped by civil inattention and the travellers' general immersion in their private digital bubbles. However, in some accounts, digital devices were connected to safety in a more communitarian way, as some young women emphasized their habits of looking after one another and making sure that friends arrived safely at their destinations when travelling later in the evening. Third, digital devices were used by young people as shields against the adult gaze and for minimizing the opportunities for interaction with them out of fear of negative comments related to their youth. This is described by 17-year-old Lotta and Veera in the following exchange. For them, engagement in their private digital bubbles gave an impression of business as opposed to loitering, which, they claim, young people are often accused of. Lotta: Yes, it is kind of strange; when you are my age, whatever you do is just loitering. Like, if I'm waiting for public transport or waiting for a friend downtown, I always try to look terribly busy so that people would not get an impression that I'm just hanging around there. Veera: In that situation I usually grab my phone and pretend to be scrolling something […] Interviewer: […] You said that you take your mobile out and then what do you do? Veera: Well, I go to Instagram. Even if I have already seen all the feeds many times, I go there again in the desperate hope that there will be something new, please, anything new. Fourth, regarding the technologized metro system, many of the young people we interviewed had expectations of safety from the representatives of formal control, such as security guards. Despite sometimes being viewed as possessing stereotypical views of young people (see also Saarikkomäki, 2017), they were nevertheless regarded as safe adults, sources of help and advice and professionals with official rights to maintain public safety, thereby representing protective and caring forms of social control for young people. However, in the digitalized metro system, which is in itself effective in following passengers everywhere (Brighenti, 2012, 410), one quite rarely encounters such safe adults and cannot expect help from CCTV cameras, as was pointed out by our interviewees. Also, some respondents felt that security personnel had disregarded or overlooked their safety needs on the metro. When asked what they would like to change in the Helsinki metro system, interviewees most commonly mentioned that they wished to see more security guards in metro carriages and stations, especially in the evening and at night.","As previously mentioned, it was common for the young people we interviewed to portray metro carriages and stations as neutral spaces with limited social interaction with unknown people (Goffman, 1963; Augé, 1995), and the tendency for passengers to be in their private digital bubbles during the metro ride. We have, however, also documented another side to this story, describing young people's recollections of mundane but ambivalent encounters in which they felt that their presence on the metro was monitored or challenged by older generations. Depending on the young person in question, these experiences were either seen as isolated incidents or as common features of public transport. Hence, these experiences were not unitary, and their meanings varied depending on contextual factors. Despite this variability, the ways in which the interviewees made sense of their experiences can be interpreted in line with conceptualizations of symbolic violence (Bourdieu, 1977; Bourgois, 2001). Accordingly, sexual harassment, for example in the form of gestures or staring, makes young women in particular conscious of their visibility and vulnerability on the metro. They learn to view themselves as self-evident and natural objects of harassment, especially when performing idealized femininity. They also learn that, as young women, they are expected to cope with harassment by themselves and to put up with it without making a big issue of it. What they do not recognize, however, are the unobserved structural or cultural mechanisms behind the objectifying practices (Bourgois, 2001). At worst, if something happens, the young women might end up blaming themselves for not having been cautious enough (Pedersen, 2009). While some seemed confident about their coping strategies and cleverly used technological means to facilitate them, others talked openly about their fears and worries concerning their personal safety. The possibility of sexual violence was discussed in some interviews, illustrating the continuum of sexual violence in the consciousness of at least some young women (Kelly, 1987). Similarly, experiences of being labelled as misbehaving youngsters can be interpreted as a form of symbolic violence. The participants seemed to have internalized the fact that they could be confronted as representatives of the ‘troublesome’ young generation when they travel in the city. They did not usually portray themselves as victims of these utterances but were nevertheless irritated that they did not have sensible means of talking back. Experiencing negative reactions can negatively affect their trust in adults and their confidence while travelling in the public sphere (Listerborn, 2015, 109). Based on the interviews, it can be said that the mundane nature of these experiences disguises the intergenerational power relations that are reproduced in them and that portray young people as immature, deficient and of less worth than adults. The acts giving rise to ambivalent emotions are not overt violence but often mundane experiences – situations which may not even be noticed by other passengers often engaged in their own virtual networks. The rude comments are based on the negative imaginaries of youth, claiming that young people are in need of education, discipline or adult guidance. The older people's conduct towards young people can therefore reflect their genuine fear of youth, which they feel safe uttering only in certain situations. Despite this, many young people regard this kind of treatment as insulting as it affronts their dignity as people and categorizes them as deviant youth. When allowed to continue without intervention, the negative imageries are reproduced, and both the older people who intrude into young people's privacy and the adults who remain passive bystanders operate in ways that fail to recognize the harmful consequences of their actions for young people as users of public transport (Cooper, 2012, 66). As such, they take part in forming a social order of discriminating patterns that might restrict young people's movements in the city (Listerborn, 2015, 106–108).","Intergenerational interaction in the Helsinki metro seems to be characterized by young people's expectations of conviviality and travelling in their private digital ‘bubbles’ on the one hand, and confusion concerning their position on the other. They feel that they are required to act like adults, yet they must be prepared to be treated like children without rights, targets of sexual harassment or intrusive comments, or problematic kids who should be disciplined or given advice on how to behave or look. In other words, young people feel that they are expected to behave well, which does not, however, preclude adversarial encounters with adults. These experiences may arouse feelings of alienation in the urban public realm, as this condition is often viewed as a self-evident reality that young people feel they must adjust to without many opportunities for change. It is notable that the characteristics of social order that we have analyzed here as symbolic violence are formed in the metro system of Helsinki, which lacks many features of formal surveillance that exist in larger cities, and where social control mainly entails informal interaction between strangers. Moreover, the Helsinki metro sociality is characterized by widespread use of digital devices. Based on our analysis, we claim that the neoliberal discourses emphasizing the positive aspects of digitalization on urban sociability (see e.g. Hampton et al., 2010) disregard other, more complex impacts related to intensified use of digital technology in public spaces. Digital technology is documented to have created, in addition to new opportunities for city dwellers, new automated forms of control involving discriminatory restrictions of freedom of movement (Brighenti, 2012, 408; Graham and Wood, 2003). In addition, as we argue here, the contemporary mediatized and digitalized metro sociality might contribute to the constitution of an oppressive intergenerational order in the public sphere in which young people are relegated to subordinated positions and treated according to stereotypical depictions. This is made possible by the young and middle-aged passengers' extensive immersion in their digital bubbles when travelling on the metro. Aside from the narratives of our respondents, it is easy to observe that this is a general mode of conduct on public transport, especially on the metro. Based on the narration of the participants in this study and on observations in earlier studies (e.g. Hatuka and Toch, 2016; Hampton et al., 2010), it can be claimed that the travellers' extensive immersion in their ‘portable private-personal territories’ decreases their attentiveness to social surroundings and less familiar people. It also seems to reduce the general preparedness of passengers to intervene in problematic situations between generations, which, according to the interviewees, have occurred between them and adults who are ‘older’ or somehow ‘different’ from busy, task-oriented adults. Secondly, young women in particular have adopted mobile technology as part of their personal safety routines. In that sense, it can be argued that digitalization has increased their feelings of safety in their personal space amid strangers (Hatuka and Toch, 2016, 2203–2204). Yet simultaneously, as they expect no support from their fellow passengers, despite the existence of mobile technology, they still feel vulnerable and talk about coping mainly as their own responsibility. The implications of mobile technology for social relations in the urban sphere are hence ambiguous. The digital revolution has not only liberated teenagers to explore the media city on their own terms, but internet connectivity and portable digital devices are also involved in processes based on stereotypical imaginaries of youth, which can have harmful consequences for young people's capacities to use city spaces. The persistence of these perceptions is of course only one aspect of the lived realities of young people in urban spheres. Nonetheless, if not problematized, their continued use will restrict young people's opportunities to relate to the media city as a safe place to explore."],["An extensive body of literature indicates that people differ in the extent to which they attend to, process, and regulate emotions. The present research sought to build on this knowledge by examining whether general self-determination (GSD) could account for individual variation in emotional intelligence (EI) and psychological well-being (PWB). A simple and multiple mediation model using bootstrap analyses tested these relationships in a sample of students (Study 1, N= 283) and workers (Study 2, N= 265). Results supported the hypothesized mediating role of EI in the relationship between GSD and PWB across both studies. When the inter-related facets of EI were considerately separately, indirect effects emerged for mood regulation/optimism and social skills across both studies as well as for utilization of emotions, albeit negatively, in Study 2. Our findings support and extend past work on the antecedents of EI and have important implications for human functioning across a variety of settings. © 2014 The Authors. --------------------------------------------------------------------------------","A wealth of scientific evidence indicates that people vary in the extent to which they use emotion-related information in their day to day lives. Mayer, Salovey, and Caruso (2000) refer to this capacity as emotional intelligence (EI) which they formally define as “the ability to perceive and express emotion, assimilate emotion in thought, understand and reason with emotion, and regulate emotion in the self and in others” (p. 396). To these authors, EI is therefore a set of abilities and should be assessed with maximum performance measures much like traditional intelligence tests (e.g., Mayer, Salovey, & Caruso, 2002; Petrides, 2011; Petrides & Furnham, 2000a). A distinct but complementary conceptualization of this construct (Schutte, Malouff, & Bhullar, 2009) defines EI as a set self-perceptions, dispositions, and motivations that are affective in nature and that share some common variance with major personality traits (Petrides, Pita, & Kokkinaki, 2007; Petrides, Pérez-Gonzalez, & Furnham, 2007). Unlike the ability-model, this trait model of EI captures the inherent subjectivity underlying one’s emotional experience and should therefore be assessed via self-report measures (e.g., Petrides & Furnham, 2000a; Petrides & Furnham, 2000b; Schutte et al., 1998). Notwithstanding these divergent operationalizations, EI has emerged as a viable and important construct in the literature evidenced by the accumulation of handbooks, book chapters, review papers, and meta- analyses on the subject. For instance, those who score high on measures of EI perform better at work (e.g., O’Boyle, Humphrey, Pollack, Hawver, & Story, 2011) and in school (e.g., Petrides, Frederickson, & Furnham, 2004); they also report more positive relationships (e.g., Mavroveli, Petrides, Rieffe, & Bakker, 2007) and better physical health (e.g. Costa, Petrides, & Tillmann, 2014). However, it’s the enhancement of emotional health and well-being wherein lies the construct’s greatest potentiality and interest. For instance, EI is negatively related to several indices of psychopathology (Malterer, Glass, & Newman, 2008) such as personality disorders (Petrides, Pérez-González, et al., 2007) and anxiety disorders (Summerfeldt, Kloosterman, Antony, McCabe, & Parker, 2011) as well as self-harm (Mikolajczak, Petrides, & Hurry, 2009) and externalizing behaviors in adolescents (Downey, Johnston, Hansen, Birney, & Stough, 2010). In non- clinical samples, EI correlates positively with a variety of well-being indices such as life satisfaction, happiness, optimism, self-esteem, and decreased negative affect (for reviews see Brackett, Rivers, & Salovey, 2011; Petrides, 2011) with a meta-analytic correlation of .34 (Martins, Ramalho, & Morin, 2010). But why do some people attend to, process, and regulate their emotions with greater ease than others? In other words, what accounts for the individual variation in EI? Consistent with the trait-model of EI (e.g., Petrides, Pita, et al., 2007), higher-order personality factors are purported to shape people’s affective self-perceptions. For instance, trait EI mediated the relationship between each of the Big Five personality traits and self-reported mental health and well- being (e.g., Johnson, Batey, & Holdsworth, 2009). Other research suggests that trait EI may stem from dispositional differences in quality of attention. Schutte and Malouff (2011) observed that the relationship between mindfulness and various indicators of subjective well-being (i.e., positive affect, negative affect, and life satisfaction) were mediated by trait EI. Significant indirect effects for a specific subcomponent of EI, namely mood regulation have also been documented. For example, Kämpfe and Mitte (2010) observed that mood repair accounted for the relationship between extraversion and life satisfaction as well as between extraversion and happiness. Other research found cognitive reappraisal of emotion to partially explain the relationship between secure attachment and well-being (Karreman & Vingerhoets, 2012). Together, these findings suggest that the ability to perceive and manage one’s emotions is partly due to stable individual differences such as one’s personality, attachment style (i.e., secure attachment), and mindfulness. In the present research, we investigated self-determination as a plausible antecedent of EI that contributes to psychological well-being (Bhullar, Schutte, & Malouff, 2013). At the core of self-determination theory (Deci & Ryan, 1985) lays a motivational perspective of the self which is endowed with integrative capacities toward increasing organization and coherence (Ryan, 1993). The expression of this coalescence is reflected in the degree of perceived autonomy or self-determination underlying the regulation of action. For instance, behaviors which are initiated out of inherent interest and enjoyment for their own sake (intrinsic regulation) are experienced as the most self- determined followed by reasons to act in accordance with one’s deepest values (integrated regulation), and then by personal identification with the activity (identified regulation). However, not all behaviors are experienced as authentic and freely chosen; many are initiated out of pressure and obligation to bolster or protect one’s sense of self-worth (introjected regulation), to comply with external demands (external regulation) or without any intention (amotivation). These behaviors are experienced as controlling and coercive because the underlying self operates in a fragmented and compartmentalized manner. These six styles of behavior regulation can be combined into a single index, whereby higher scores reflect greater self-determination which is linked to healthier functioning and well-being (e.g., see Deci & Ryan, 2008 for a review). The integrative capacity for effective and adaptive self-regulation of action is also reflected in the manner with which one meets their moment to moment experiences. According to Hodgins and Knee (2002), greater self-determination endows a person with more openness and less defensiveness toward potentially threatening and difficult events. For instance, when primed with self-determination, people report less desire to escape and engage in fewer self-serving attributions in response to failure (Hodgins, Yacko, & Gottlieb, 2006). Autonomously-oriented individuals also exhibit better emotional regulation and integration of negative affect after viewing a traumatic film clip (Weinstein & Hodgins, 2009) and retrospectively recalling negative life events and identities (Weinstein, Deci, & Ryan, 2011). However, little is known on the skills utilized by those with greater self- determination which promote effective assimilation of emotionally-laden experiences into a more unified and cohesive self. We propose that these skills are attributed in part to the inter-related abilities of EI. The objective of the present research was to investigate individual variation in EI by examining the determining role of self-determination which was assessed at the dispositional or general level indicative of a more enduring motivational orientation toward the environment (Guay, Mageau, & Vallerand, 2003). To this end, the inter-related abilities of EI were hypothesized to mediate the relationship between general self-determination (GSD) and psychological well-being (PWB). These relationships were initially tested with a sample of undergraduate students (Study 1) and then replicated with a sample of working adults (Study 2).","A sample of 283 undergraduate students of which the majority were female (n = 226) took part voluntarily in this two-phase study (Mage = 18.95 years, SDage = 1.75). Participants were recruited from a campus subject pool and received course credit in exchange for their participation.","of GSD and EI were completed at the beginning of the semester (Phase 1) while a measure of PWB was completed three months later (Phase 2). Measures GSD was assessed with the 18-item General Motivation Scale (GMS; Guay et al., 2003). The six subtypes of motivation proposed by Deci and Ryan (1985) are each represented by three items. Respondents rated the extent to which each item (e.g., “…because I like making interesting discoveries”; intrinsic regulation) corresponded to their reasons as to “why they do things in general” on a scale from 1 (does not correspond to my reasons at all) to 7 (corresponds exactly to my reasons). Internal consistency estimates ranged from .68 to .84 across subscales. Mean subscale ratings were combined to form a GSD index whereby higher scores indicate greater GSD: +3 * (intrinsic) + 2 * (integrated) + 1 * (identified) − 1 * (introjected) − 2 * (external) − 3 * (amotivation). Cronbach’s alpha for the entire scale was .81. EI was measured using the Assessing Emotions Scale (AES: Schutte et al., 1998) where responses were rated from 1 (strongly disagree) to 7 (strongly agree). As to its structure, some suggest the existence of a single global EI factor (e.g., Schutte, Malouff, Simunek, McKenley, & Hollander, 2002) while others propose the existence of four sub-factors (e.g., Petrides & Furnham, 2000a; Saklofske, Austin, & Minski, 2003). Cognizant of this debate, EI was represented by a global EI factor derived by averaging scores across all 33 items as well as by four sub-factors derived by averaging scores across each subscale’s respective items. The subscales were derived from the work of Petrides and Furnham (2000a). Internal consistency estimates ranged from .72 to .84 across subscales (α = .91 for the entire scale). PWB was assessed using Ryff’s (1989) short form Scales of Psychological well-being (SPWB) which tap six different facets of positive psychological functioning. Responses were rated on a scale from 1 (strongly disagree) to 7 (strongly agree) and then averaged across all 18 items to represent PWB (α = .84).","Descriptive statistics Descriptive statistics are reported in Table 1. As predicted, positive relationships emerged between GSD, EI, and PWB. On the bivariate level, age did not correlate with any variable. However, gender differences did emerge for certain facets of EI with women scoring higher than men on ‘appraisal of emotions’ and ‘social skills’. Regardless of these observations, both gender and age were controlled for in subsequent analyses for theoretical reasons (e.g., Mavroveli et al., 2007; Petrides & Furnham, 2000b). Mediation analyses The hypothesized mediating role of EI in the relationship between GSD and PWB was tested using the SPSS macro, INDIRECT (Preacher & Hayes, 2008). The macro relies on the resampling method of bootstrapping; a procedure that provides an estimate of the indirect effect in the population by resampling the dataset k times in order to obtain the indirect effect’s sampling distribution and confidence intervals (CI). An estimate is considered statistically significant if its 95% CI does not include zero. Testing a simple mediation model, a direct effect emerged for GSD on PWB (B = .019, p < .001) as did an indirect effect through global EI with a point estimate of .0177, 95% CI [.0111, .0254]. The magnitude of this indirect effect was represented by an index of mediation (Preacher & Kelley, 2011) which was equal to .171. Thus, for every 1 SD increase in GSD, PWB would increase by .171 SDs through global EI. Effect size was estimated using Kappa2; a ratio of the obtained estimate over the maximum possible estimate of the indirect effect (Preacher & Kelley, 2011). In this case, Kappa2 = .18. Overall, GSD accounted for 33% of the variance in PWB both directly and indirectly through EI, B = .037, p < .001, F(4,278) = 34,86, p < .001. To understand which facet of EI accounts for the relationship between GSD and PWB, a multiple mediation model (MMM) was tested with each of the four sub-factors of EI. These results are reported in Table 2. Indirect effects emerged for ‘mood regulation-optimism’ and ‘social skills’. Pairwise contrasts revealed that the former was greater than the latter. The index of mediation for each indirect effect indicates that for every 1 SD increase in GSD, PWB would increase by.188 and .043 SDs through ‘mood regulation-optimism’ (Kappa2 = .20) and ‘social skills’ (Kappa2 = .05), respectively. In the MMM, GSD accounted for 40% of the variation in PWB both directly and indirectly through the four facets of EI, F(7,275) = 25.80, p < .001. Participants, measures and procedure This community sample was comprised of 265 working adults recruited from multiple work environments using a snowballing strategy, of which 46% were employed in the private sector and the remaining 54 % were employed in the public sector. The majority were women (n = 184) and the sample ranged from 18 to 58 years (M = 32.93, SD = 12.32). Participants took part in this cross-sectional study voluntarily and were invited to complete a questionnaire comprised of the same self-report measures described in Study 1 with the exception of the SPWB for which a mid-length version of the scale was used whereby each factor was represented by nine items instead of three (GMS: αs ranged from .66 to .84 across subscales, with a Cronbach’s alpha for the entire scale of .82; AES: αs ranged from .69 to .81 across subscales and was equal to .88 for the entire scale; SPWB: α = .92 for the entire scale). Descriptive statistics Descriptive statistics are reported in Table 1. Bivariate relationships between all constructs were consistent with predictions with the exception of ‘utilization of emotions’ which was not significantly related to GSD (r = .11, p > .05) nor to PWB (r = .02, p > .05). Age and gender emerged as significant correlates of GSD and EI. Consistent with the results of Study 1, women’s EI scores were higher than men’s on the global factor as well as on all facets of EI except ‘mood regulation- optimism’. Age and gender were controlled in subsequent analyses. Mediation analyses Testing a simple mediation model, a direct (B = .027, p < .001) and indirect effect with a point estimate of .0100, 95% CI [.0065, .0144] emerged for the global factor of EI with an index of mediation of .131 (Kappa2 = .14). GSD accounted for 36% of the variance in PWB both directly and indirectly through EI, (B = .037, p < .001), F(4,260) = 37.59, p < .001. Next, a MMM was tested, the results of which are reported in Table 2. Indirect effects emerged for all facets of EI with the exception of appraisal of emotions. Pairwise contrasts revealed that the indirect effect through ‘mood regulation-optimism’ was the greatest, followed by ‘social skills’ and then by ‘utilization of emotions’. The index of mediation for these indirect effects indicates that for every 1 SD increase in GSD, PWB would increase respectively by .146 and .043 SDs through ‘mood regulation-optimism’ (Kappa2 = .16) and ‘social skills’ (Kappa2 = .05). Contrarily, for every 1 SD increase in GSD, PWB would decrease by .031 SDs through ‘utilization of emotions’ (Kappa2 = .04). In sum, GSD accounted for 46% of the variation in PWB both directly and indirectly through the four factors of EI, F(7,257) = 33.19, p < .001.","The efficiency and effectiveness with which people can identify, process, and manage their emotions has important implications for their health and well-being. Part of this capacity is attributed to structural factors such as one’s personality (e.g., Johnson et al., 2009). Grounded in the framework of self-determination theory (Deci & Ryan, 1985), we propose that part of this capacity stems from an underlying self that is motivational in nature, oriented toward greater organization and unity. At the behavioral level, this integrative propensity is manifested as general self-determination (GSD) which endows a person with greater openness and receptivity to their environment (Hodgins & Knee, 2002). The present research tested this proposition by investigating GSD as a plausible antecedent that may account for individual variation in emotional intelligence (EI) thereby resulting in differing levels of psychological well-being (PWB). Data obtained from two different samples (students vs. workers) support these hypothesized relationships. First, we examined the mediating role of global EI in the relationship between GSD and PWB. All paths in the model were significant and positive across both studies suggesting that greater GSD is associated with greater PWB, directly and indirectly through increased global EI. Therefore, the more people undertake their daily activities with a sense of volition and autonomy, the more skilled they become in responding to and using emotion-laden information in their day to day decision making processes thereby experiencing greater PWB. Consistent with the trait-model of EI, these findings imply that part of its variability is motivational in nature (Petrides, Pita, et al., 2007) and therefore suitable for intervention work. Indeed, people who underwent an ‘emotional competence’ training program significantly improved their employability, their subjective well-being and the quality of their relationships post intervention (Kotsou, Nelis, Grégoire, & Mikolajczak, 2011; Nelis et al., 2011). Our findings suggest that intervention efforts might also benefit from targeting people’s GSD. For instance, participants primed with subtle reminders of GSD (i.e., choice, opportunity, freedom) evidenced better emotional integration following exposure to a traumatic film designed to induce negative affect (Weinstein & Hodgins, 2009) and after recalling difficult life events (Weinstein et al., 2011). Thus, participants undergoing an ‘emotional competence’ training program might experience accrued benefits if primed with GSD beforehand. Second, we sought to better understand which of the four (if not all) inter-related abilities of EI accounted for the relationship between GSD and PWB. Both ‘mood regulation-optimism’ and ‘social skills’ emerged as significant mediators suggesting that effective and adaptive regulation of action is linked to effective and adaptive regulation of emotions, “within the self and in relation to other people” (Vesely, Siegling, & Saklofske, 2013, p. 222). These results emerged in both studies, lending strength to their effect and are in line with the work of Spence, Oades, and Caputi (2004) who noted that mood regulation-optimism was the strongest predictor of emotional well-being in a sample of students. Our analyses also yielded one surprising result that warrants further discussion. Contrary to expectations, the indirect effect of ‘utilization of emotions’ (UE) was negative and significant in the worker sample but not in the in the student sample. To be specific, this result suggests that greater GSD is linked to greater UE which in turn is associated with lower levels of PWB. This particular subscale was designed to tap the extent to which one is capable of using emotional information in generating ideas and solving problems. Yet, a closer examination of its individual items (n = 4) suggests some bias toward neutral and positive emotions (e.g., “When I am in a positive mood, solving problems is easy for me”). Theoretically, this seems incongruent with the open and non-defensive disposition of someone with greater GSD who is equipped at meeting and internalizing a broad spectrum of emotions, even difficult ones. Stated differently, all emotional inputs (positive and negative) represent potential sources of information in making decisions for someone who initiates their actions based on well-integrated values. Statistically, there’s also the possibility that UE acted as a suppressor in the model. When we tested a simple mediation model linking GSD to PWB through UE, the mediator was not statistically significant. However, when several MMM were tested that included UE, its indirect path became significant. This finding may also be inherent to the subscale itself as other studies noted similar problems (e.g. Gignac, Palmer, Manocha, & Stough, 2005). A few limitations of the present research are worth noting. First, responses were limited to self-report data. Future work could examine these relationships using a motivational priming procedure and behavioral evidence of emotional integration and well-being (Hodgins et al., 2006). Second, Cronbach’s alphas were low (<.70) for some of the subscales of the GMS in Study 1 and 2, an issue that has been reported by other researchers (Julien, Guay, Senécal, & Poitras, 2009). However, the entire scale was used in the creation of GMS indexes whereby the reliability statistics were respectively .81 and .82. Third, the four- factor solution documented by Petrides and Furnham (2000a) for the AES did not emerge in our data. Attempts were made to keep our factor structure consistent with findings in the literature, but some of the items did not provide the same fit across our samples suggesting potential instability with respect to a four-factor solution. Future research is needed to help elucidate sample differences when examining sub-facets of EI with the AES. Fourth, the longitudinal nature of Study 1 implies that factors other than GSD and EI assessed during Phase 1 influenced the reporting of PWB assessed at Phase 2. Future studies should include potential covariates of this relationship (e.g., personality factors). Despite these limitations, findings from the present research contribute to a burgeoning literature on the antecedents of EI given the importance of EI for adaptive and healthy functioning. By investigating the motivational underpinning of EI, our findings lend credence to the growing interest in programs and workshops aimed at increasing EI. Moreover, the present research was grounded in self-determination theory; one of the most validated and comprehensive frameworks of human needs and motivation. Our results support the predictions of Hodgins and Knee (2002) by linking a motivational orientation with specific socio-emotional competencies which are conducive to better emotional integration and therefore enhanced PWB. By uncovering the emotional pathways by which GSD leads to better adjustment, future work might garner a better understanding of how people with varying motivational profiles cope with adversity (e.g., Amiot, Blanchard, & Gaudreau, 2008) across a variety of settings (e.g., school, work, sports)."],["Mindfulness-based interventions have proven effective for the transdiagnostic treatment of heterogeneous anxiety disorders. So far, no study has investigated the potential of mindfulness-based treatments when delivered remotely via the Internet. The current trial aims at evaluating the efficacy of a stand-alone, unguided, Internet-based mindfulness treatment program for anxiety.Ninety-one participants diagnosed with social anxiety disorder, generalized anxiety disorder, panic disorder, or anxiety disorder not otherwise specified were randomly assigned to a mindfulness treatment group (MTG) or to an online discussion forum control group (CG). Mindfulness treatment consisted of 96 audio files with instructions for various mindfulness meditation exercises. Primary and secondary outcome measures were assessed at pre-, posttreatment, and at 6-months follow-up.Participants of the MTG showed a larger decrease of symptoms of anxiety, depression, and insomnia from pre- to postassessment than participants of the CG (Cohen's dbetween= 0.36-0.99). Within effect sizes were large in the MTG (d= 0.82-1.58) and small to moderate in the CG (d= 0.45-0.76). In contrast to participants of the CG, participants of the MTG also achieved a moderate improvement in their quality of life.The study provided encouraging results for an Internet-based mindfulness protocol in the treatment of primary anxiety disorders. Future replications of these results will show whether Web-based mindfulness meditation can constitute a valid alternative to existing, evidence-based cognitive-behavioural Internet treatments.The trial was registered at ClinicalTrials.gov (NCT01577290). © 2013 The Authors. --------------------------------------------------------------------------------","The trial was registered at ClinicalTrials.gov (NCT01577290). The regional ethics committee of Umeå University approved the study protocol. Participants were recruited via advertisements in regional and national newspapers and on the project’s study website (www.studie.nu). After registering with their e-mail address, participants obtained detailed information about the theoretical background, the goals and the design of the study, and were asked to give written informed consent. They were informed that the study aimed to compare a mindfulness-based treatment to a control condition and that participants randomized to the control group would receive access to the active treatment after post-assessment. The selection of participants followed two steps. Participants were asked to fill out the outcome questionnaires. These included, among others, the Beck Anxiety Inventory (BAI; Beck, Epstein, Brown, & Steer, 1988), the Beck Depression Inventory-II (BDI-II; Beck, Steer, & Brown, 1996), and additional questions regarding current and past psychological or medical treatment for mental problems. Participants who indicated at least mild anxiety on the BAI (cutoff > 8) and who did not indicate severe depression according to the BDI-II (cutoff < 29) or suicidal ideation as assessed by the suicide item of the BDI-II (item 9 < 2) were then invited to take part in a diagnostic interview. The interview was conducted via telephone, a procedure with adequate psychometric properties (Crippa et al., 2008). Four advanced MSc clinical psychology students conducted the depression and anxiety disorders sections of the Structured Clinical Interview for DSM-IV Axis I Disorders (First & Gibbon, 2004). The interviewers had received training in using the interview. The SCID training included sample videos, role-plays, and supervised training interviews. We applied the following inclusion criteria: (a) at least 18 years old, (b) access to the Internet, (c) meeting diagnostic criteria for a primary diagnosis of social anxiety disorder, panic disorder with or without agoraphobia, generalized anxiety disorder, or anxiety disorder not otherwise specified, (d) not participating in any other psychological treatment for the duration of the study, (e) no extensive prior experience with mindfulness meditation, and (f) if on prescribed medication for anxiety/depression, dosage had to be constant for 3 months prior to the start of the treatment.","After pre-assessment, participants were randomly allocated to the mindfulness treatment group (MTG) or the discussion forum control group (CG) by an online true random-number service independent of the investigators. After randomization, participants of the MTG received access to a website where the mindfulness program was presented. They were asked to work with the mindfulness program daily for 6 days of the week for 8 weeks. Participants in the control group received access to an online discussion forum and were invited to take part in online discussions during 8 weeks. Participants in both groups received an automated e-mail at the end of Week 4, encouraging them to carry on with the assigned treatment. Primary and secondary outcome measures were administered over the Internet prior to the treatment, after the treatment at the end of Week 8, and, for participants of the MTG, at 6 months follow-up. Participants of the CG were offered the mindfulness program after post-assessment. Internet-Based Mindfulness Treatment At the core of the unguided, Internet-based mindfulness treatment program were brief, instructive audio files presenting mindfulness exercises (Schenström, 2010). Mindfulness exercises included instructions for sitting meditation, mindfulness movement, three different types of body scan, and four different forms of breathing anchor. The program was organized into eight modules. At the beginning, participants were presented with a 20-minute video that explained the concept of mindfulness and its relevance for anxiety disorders and introduced the eight modules of the program. The modules were as follows: Stopping and Getting Started Managing Your Thought Noise SOAL: Stop, Observe, Accept, and Let Go It Is What It Is Sitting With Whatever Comes Up In each module, mindfulness exercises were combined with brief psychoeducation and written instructions to apply the concept of mindfulness in daily life. Each module included 12 mindfulness exercises, each of which lasted 10 minutes, resulting in a total of 960 minutes (16 hours) of mindfulness exercises for the 8-week treatment period. Before each mindfulness exercise, participants were instructed to reflect on the purpose for the exercise. Participants were asked to complete one module each week and to conduct mindfulness exercises twice a day on 6 days of the week. Participants only gained access to the next module once they had completed the previous one. Online discussion forum Participants in the control group received access to a closed, anonymous, and supervised online discussion forum. Each week, a new topic was presented for discussion. All topics were related to anxiety or panic but were not therapeutic in nature. For example, topics included how participants perceived the health care provided for anxiety disorders, how they discussed anxiety problems with others, and how they perceived seasonal changes of mental health problems. These online dialogues were supervised but the investigators did not take active part in the discussions. Outcome measures Our primary outcome measure was the BAI, a 21-item self-report questionnaire that assesses the severity of somatic and cognitive anxiety symptoms. Items are scored on a 0–3 Likert scale and the total score ranges between 0–63 points. In addition, as secondary outcome measures, we administered the BDI- II, the Quality of Life Inventory (QOLI; Frisch, Cornell, Villanueva, & Retzlaff, 1992), and the Insomnia Severity Index (ISI; Morin, 1993). The BDI- II measures depression on 21 items, with a total score ranging between 0 to 63 points (0–3 Likert scale). The QOLI assesses the importance of (0–2 Likert scale) and satisfaction with (-3 to 3 Likert scale) 16 life domains on 32 items (total score -6 to 6). The ISI is a brief questionnaire that assesses insomnia on 7 items (0–4 Likert scale; total score 0–28). All outcome measures were administered online, a procedure that has demonstrated adequate psychometric properties for all applied instruments (Hedman et al., 2010; Lindner, Andersson, Öst, Carlbring, 2013; Thorndike et al., 2009, 2011). In the current sample, reliability estimates at pre-assessment were as follows: BAI: α = 0.82, BDI-II: α = 0.82, QOLI: α = 0.79, and ISI: α = 0.86. Statistical analyses All analyses on change in primary and secondary outcome measures were conducted as intention-to-treat analyses using a mixed models approach. Analyses were carried out in R Version 2.15 (R Development Core Team, 2010), and mixed models were fitted with NLME (Jose, Douglas, Saikat, Deepayan, & R Development Core Team, 2012). In this approach, main and interaction effects are evaluated on the basis of their contribution to an increase of goodness of model fit (Field, Miles, & Field, 2012). The increase of fit is χ2 -distributed. Within- and between-group effect sizes were calculated using Cohen’s formula based on pooled standard deviations (Cohen, 1988). Clinically significant change was calculated for the BAI for the completer sample according to the criteria suggested by Jacobson and Truax (1991). In order to facilitate comparison of the present findings with those of other studies, we adopted criteria for improvement and recovery from two previous trials (Vøllestad et al., 2011; Westbrook & Kirk, 2005). Reliable improvement or deterioration was defined as a pre-post change score of 10 points or more and recovery was defined as a post BAI score of 10 or less.","A total of 91 participants met all inclusion criteria and were randomized to one of the two groups (see flow chart in Figure 1). Seven participants (7.7%) did not complete the outcome measures at posttreatment. Dropout rates did not differ between the two groups, χ2(1) = 1.47, p = .267. Ten participants in the mindfulness group (11%) failed to fill out self-report measures at 6-months follow-up-assessment. Table 1 displays sociodemographic characteristics and Table 2 depicts pretreatment scores of the outcome measures for the two groups. There were no significant group differences at pretreatment on any demographic variable or outcome measure: all χ2(1-3) < 2.33, all p > .31; all t(89) < 1.31, all p > .19. The computer automatically registered the amount of completed mindfulness exercises for the MTG. Participants in the mindfulness group completed on average 44 (SD = 33.7) out of 96 mindfulness exercises, which corresponds to an average of 7.3 hours of mindfulness practice during the 8-week intervention period. These time specifications can only be estimators of the real practice time. The program only recorded the amount of started exercises. It remains unknown whether the mindfulness exercises were not only started but also conducted for the full intended 10 minutes. In both groups, participants were asked at post-assessment how satisfied they were with the received treatment. Answers ranged from 1 (not at all satisfied) to 5 (very satisfied). Participants of the mindfulness group were on average “satisfied” with the treatment (M = 3.7, SD = 1.0) whereas participants of the online forum control group were only “somewhat satisfied” (M = 3.0, SD = 1.2). This group difference in satisfaction was significant, t(80) = 2.88, p < .01. Change in primary and secondary outcomes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Means, standard deviations, and effect sizes for all four outcome measures in both groups are displayed in Table 2. The mixed models analysis on the BAI revealed that participants in the mindfulness group showed a larger decrease of anxiety from pre- to post-assessment than participants of the control group (Group × Time: χ2[1] = 9.71, p = .002). Pre-post effect sizes were large (d = 1.33) for the mindfulness group and moderate (d = 0.76) for the control group. The between-group effect size at post-assessment indicated a large group difference (d = 0.99). Results on the depression scale were similar. The mixed model analysis using the BDI as dependent variable showed a significant Time × Group interaction, χ2(1) = 15.60, p < .001. Participants of the active group indicated more improvement on depression scores from pre- to post-assessment than participants of the control group (between-group effect d = 0.84). Pre-post effect sizes were large (d = 1.58) in the mindfulness group and small (d = 0.49) in the control group. The mindfulness group also improved significantly more than the control group on the ISI. The mixed model analysis revealed that participants of the active group showed a larger decrease in insomnia from pre- to post-assessment than did participants of the control group (Time × Group: χ2[1] = 5.77, p = .016). Effect sizes indicated large improvements for the active group (d = 0.82) and small improvements for the control group (d = 0.45), as well as a small group difference at post-assessment (d = 0.36). The mixed model analysis on the QOLI also revealed a significant interaction effect of Time × Group, χ2(1) = 6.68, p = .009. Participants in the mindfulness group reported a moderate increase of life satisfaction (d = 0.64), whereas participants of the control group showed no change (d = 0.04). Group differences at post-assessment were small (d = 0.37). Differences Between Diagnostic Groups To examine whether the different diagnostic groups predicted or moderated primary treatment outcome, we entered diagnostic group as an additional independent variable in the mixed model on the BAI. Results indicated that although the different diagnostic groups showed different levels of anxiety across both assessment points and both treatment groups (main effect diagnostic group: χ2[3] = 19.84, p < .001), the specific anxiety diagnosis did not predict change in anxiety (Diagnostic Group × Time: χ2[3] = 1.89, p = .596) nor did it moderate the treatment effect (Diagnostic Group × Time × Treatment Condition: χ2[6] = 6.83, p = .337). In other words, in both treatment conditions, participants with PD, GAD, SAD, and ADNOS showed similar rates of anxiety change. Psychotherapy Experience To investigate the impact of former psychotherapy experience on treatment outcome, we included experience with psychotherapy (yes/no) as an additional independent variable into the mixed model analysis on the BAI. Results showed that the experience with psychological treatment did not influence change in anxiety scores across both treatment groups (Psychotherapy Experience × Time: χ2[1] = 2.28, p = .131). Participants with former experiences with psychotherapy showed similar rates of improvement as participants without such experiences. Differences in previous psychotherapy experience also did not lead to differential change rates of anxiety symptoms in the two treatment conditions (Psychotherapy Experience × Time × Treatment Condition: χ2[2] = 2.12, p = .345). Amount of Exercises We also examined whether the amount of completed mindfulness exercises predicted change in anxiety in the mindfulness treatment group (MTG). Amount of mindfulness exercises was entered as an independent variable into a mixed model within the MTG using the BAI at pre- and post-assessment as dependent variable. Results showed that the amount of exercises did not predict treatment outcome (Time × Exercises Completed: χ2[1] = 2.54, p = .111). Clinical significance of change ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 3 depicts the rates of clinical significant change on the primary outcome measure BAI at post-assessment for the two groups (completer sample). Sixteen participants (40%) of the mindfulness group met the criteria of improvement and recovery compared to 4 participants (9%) in the control group. This difference in response rates was significant, χ2(1) = 11.04, p = .002. Maintenance of treatment effects ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ At 6-month follow-up, the control group had received access to the mindfulness treatment program, so analyses are based on the MTG only. Means and standard deviations at 6-month follow-up are included in Table 2. Paired t-tests from pre- to follow-up assessment showed that there was a significant decline in anxiety scores (BAI: t[34] = 8.54, p < .001, d = 1.44). Similarly, on secondary outcome measures, participants of the MTG showed a significant reduction of symptoms of depression and insomnia (BDI-II: t[34] = 5.89, p < .001, d = 1.00; ISI: t[34] = 4.77, p < .001, d = 0.82) and a significant improvement in quality of life from pre- to follow-up assessment (t[34] = -3.12, p = .004, d = 0.53). Differences between post and 6-month follow-up indicated stable treatment results for three out of the four outcome measures. There were no significant post- to follow-up differences for anxiety, t(34) = -0.95, p = .347, insomnia, t(34) = -1.97, p = .057, and quality of life, t(34) = 1.67, p = .104. Results showed a significant increase of depressive symptoms from post- to follow-up assessment, t(34) = -2.18, p = .036.","The current trial aimed to evaluate the efficacy of an Internet-based mindfulness treatment program for persons with anxiety disorders. Pre-post change scores as well as the comparison to an active control condition indicated that participants of the mindfulness program benefitted substantially from the treatment and experienced a significant decrease of anxiety symptoms. Participants of this group also achieved substantial reductions in symptoms of depression and insomnia, which are very common comorbid conditions among individuals with anxiety disorders. Most improvements were stable at 6-month follow-up with the exception of changes in depressive symptoms. Overall, the present study supports previous results achieved in face-to-face settings on mindfulness as an effective transdiagnostic treatment approach in heterogeneous anxiety disorders (Arch et al., 2013; Vøllestad et al., 2011). In contrast to the present study, Arch and colleagues treated more severely disturbed patients in a more clinically representative setting and reported only moderate changes on self-report measures. The selection and recruitment of participants in the study of Vøllestad et al. (2011), on the other hand, was very similar to the current randomized controlled trial and effects on primary and secondary measures, as well as clinical change rates, were comparable. Results of the present study are also in line with findings of recent meta-analyses on mindfulness-based treatments in anxiety disorders (Hofmann et al., 2010; Vøllestad et al., 2012). Similar to the current results, Vøllestad and colleagues (2012) reported large average reductions in depression and anxiety and moderate improvements in quality of life through stand-alone and combined face-to-face mindfulness treatments. In conclusion, mindfulness exercises delivered in face-to-face settings or remotely via the Internet seem to yield similar changes in symptoms. The remote delivery does not seem to lessen the efficacy of mindfulness interventions. This is surprising as some differences in the setting are prone to affect common mechanisms of change. The most salient difference constitutes the lack of contact to a clinician and to other patients in the Internet setting. As Baer, Carmody, and Hunsinger (2012) point out, the contact with a warm and empathetic group leader as well as the sharing with fellow participants very likely stimulate therapeutic changes in face-to-face mindfulness programs, above and beyond the effects elicited by more specific mechanisms of change. Accordingly, Malpass and colleagues (2011), who describe the therapeutic process in mindfulness treatments from the participants’ point of view, highlight the perceived importance of the group as a therapeutic factor. The proposed specific therapeutic factor in mindfulness-based interventions is an increase of mindfulness. An increase of mindfulness has repeatedly been associated with the reduction of symptoms (e.g., Bränström, Kvillemo, Brandberg, & Moskowitz, 2010; Carmody & Baer, 2008; Vøllestad et al., 2011). Baer and colleagues demonstrated that an increase in facets of mindfulness, such as observing, nonreactivity, acting with awareness, and nonjudging, preceded and mediated changes in stress in an MBSR program. While the beneficial effect of an increase of mindfulness very likely also applies to Internet-based mindfulness treatments, the effects of more common mechanisms of change, such as the therapeutic relationship or group therapeutic factors, do not apply to an Internet-based intervention. In the present study, participants practiced mindfulness without any contact or support from clinicians or fellow participants. This also applies to two previous Internet-based mindfulness studies in nonclinical samples (Glück & Maercker, 2011; Krusche et al., 2012). In effect, the good outcomes of unguided Internet- based mindfulness studies at least partly question the necessity of interpersonal common factors in mindfulness treatments. A second important difference between the current Internet trial and previous trials constitutes the intensity of treatments. The applied Internet-based program was restricted to 20 minutes of daily mindfulness exercise. The MBSR treatment program encompasses 30 hours of group treatment paired with instructions to train in mindfulness daily for 45 to 60 minutes (Kabat-Zinn, 1990). Similarly, in mindfulness-based cognitive therapy (Segal, Williams, & Teasdale, 2012), patients spend 24 hours in group treatment and train an additional 45 to 60 minutes per day at home. Unfortunately, most clinical trials on mindfulness-based interventions failed to document to which extent patients actually adhered to these extensive homework assignments. As an exception, Vøllestad et al. (2011) reported that their participants practiced mindfulness for, on average, 34 minutes a day. This constitutes a vast difference to the 7 minutes of mindfulness practice per day found in the current trial. As both treatment protocols yielded comparable outcomes in similar patient populations, these findings suggest that treatment intensity does not crucially affect treatment outcome in mindfulness-based interventions. Indeed, reviews on mindfulness interventions found no or equivocal results on the relationship between number of treatment sessions/amount of homework exercise and treatment outcome (Toneatto & Nguyen, 2007; Vettese, Toneatto, Stea, Nguyen, & Wang, 2009; Vøllestad et al., 2012). In the present trial, the association between completed mindfulness exercises and change in anxiety was weak (r = .26, p = .114). Also, the good results of the present study were paired with a rather low adherence. Participants completed on average only half of the treatment protocol. There is yet no empirical data on the necessary and sufficient amount of mindfulness practice. Future studies should investigate dose-response relations in Internet as well as in face-to-face mindfulness- based interventions. Previous research on transdiagnostic Internet-based treatments for anxiety disorders found that tailored or unified cognitive-behavioral programs can be effective in the reduction of anxiety symptoms (Berger et al., 2013; Carlbring et al., 2011; Johnston et al., 2011). The reported between- and within-group effect sizes of the current trial are comparable to those achieved through these CBT trials. Also, the attrition rate of 8% in the current trial ranges well within the proportions reported in online CBT trials (4%–10%; Berger et al., 2013; Carlbring et al., 2011; Johnston et al., 2011). When comparing ratings of satisfaction with the received treatments, participants in the current mindfulness trial seem a bit less satisfied (average of 4 on a 1–5 scale) than participants in the two previous transdiagnostic CBT trials that reported satisfaction ratings (average of 3–4 on a 1–4 scale; Berger et al., 2013; Johnston et al., 2011). Overall, the Internet-based mindfulness treatment of anxiety disorders seems equally effective and acceptable as Internet-based CBT approaches. Although totally different in content, the applied Web-based mindfulness program and online CBT programs share some common features that have been found to be associated with good therapeutic outcome. For example, participants in both types of Internet treatments underwent the same extensive diagnostic process that has been found to promote adherence and therapeutic change (Barak, Hen, Boniel-Nissim, & Shapira, 2008; Boettcher, Berger, & Renneberg, 2012). Furthermore, participants in both Internet treatments completed their therapy within a clear deadline, an additional characteristic of Internet-based treatments that has been found to be associated with good outcome (Nordin, Carlbring, Cuijpers, & Andersson, 2010). These shared features and the comparable good results suggest that some characteristics of the Internet setting per se contribute to therapeutic change, possibly by the stimulation of positive outcome expectations (Boettcher, Renneberg, & Berger, 2013). Limitations and future research ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The current trial is, to our knowledge, the first to evaluate an Internet-based mindfulness treatment program for anxiety disorders. As such, it concentrated on the examination of efficacy. Nonetheless, the lack of information regarding the proposed mechanism of change in mindfulness treatments, the increase of mindfulness, is the central limitation of the present study. Future Internet-based studies should assess likely common and specific agents of change on a regular basis and relate these to changes in different outcome domains. A second limitation of the present design constitutes the lack of a diagnostic interview at post-assessment. Even though rates of clinical change are an indicator of how many participants benefitted from the treatment, a clinician-rated reevaluation of diagnostic status would have been helpful to estimate the effects of this treatment on remission and comorbidity (Johnston et al., 2013). To this end, the administration of disorder-specific self-report questionnaires also would have been helpful. The restriction to use the BAI as primary outcome measure makes it harder to compare the current trial to disorder-specific studies. Furthermore, as the BAI mainly focuses on the assessment of somatic anxiety symptoms, it may not adequately reflect changes in cognitive symptoms of anxiety, such as worrying in GAD (Leyfer, Ruberg, & Woodruff-Borden, 2006). A further limitation of the present study lies in the assessment of treatment adherence. We only assessed how many exercises were initiated and were unable to verify whether and for how long these exercises were performed. Clearly, the present study needs replications and can only be considered as a first step towards the establishment of online mindfulness programs in the treatment of anxiety disorder. In order to estimate the comparative efficacy of mindfulness online treatments, future studies should apply more carefully controlled comparison groups. One limitation of the current trial is the failure to assess the engagement of the participants in the online discussion forum control group. We do not know how much time the participants spent in the forum and how actively they participated in the discussions. Moreover, participants of the control group were informed beforehand that they would receive the mindfulness treatment after post-assessment, perhaps inadvertently producing a wait-list quality to the control group as some participants may have ignored the offer of the online discussion group and waited for the active treatment. This makes it hard to interpret the between-group effect sizes. Nevertheless, the results of the current trial suggest that an unguided mindfulness program can be as effective as established Internet-based CBT programs. Mindfulness-based treatments could form an alternative to existing online programs. They could offer a valid choice for persons seeking treatment for anxiety disorders in general and for patients who do not respond to CBT in particular. Mindfulness and CBT treatments differ substantially in the demands they pose on participants. Both treatment approaches are participatory and ask the patient to engage actively in the therapeutic process. However, whereas mindfulness treatments request the patient to follow repeated meditation instructions, CBT protocols ask the patient to actively engage in varying exercises (e.g., dysfunctional thoughts protocol, behavioral experiments). In direct comparisons of mindfulness and CBT treatments, future studies should investigate patient characteristics and preferences that are potentially associated with differential treatment outcome. In accordance with face- to-face multicomponent mindfulness interventions, mindfulness exercises could also complement cognitive-behavioral treatment protocols. Future studies should seek to explore reasonable ways to combine both treatment approaches in the Internet-based setting and empirically evaluate the potential benefits of combined treatments.","This study was made possible in part by a generous grant from the Swedish Council for Working Life and Social Research (FAS 2008-1145). The funding body was not involved in the study design, in the collection, analysis and interpretation of data, in the writing of the report, or in the decision to submit the article for publication.","Five of the six authors have no competing interests to report. Dr. Ola Schenström has founded a company that, among other things, markets online mindfulness products."],["Objectives: This experiment examined whether electroencephalographic (EEG)-based neurofeedback could be used to train recreational golfers to regulate their brain activity, expedite skill acquisition, and promote robust performance under pressure. Design: We adopted a mixed-multifactorial design, with group (neurofeedback, control) as a between-subjects factor, and pressure (low, high), session (pre-test, acquisition 1, acquisition 2, acquisition 3, post-test), block (putts within each training session), and epoch (cortical activity in the seconds around movement initiation) as within-subject factors. Methods: Recreational golfers received three hours of either true (to reduce frontal EEG high-alpha power, N=12) or false (control, N=12) neurofeedback training sandwiched between pre-test and post-test sessions during which we collected measures of cortical activity (EEG) and putting performance under both low and high pressure conditions. Results: Individuals in the neurofeedback group learned to reduce their frontal high-alpha power before striking putts. Despite causing this more \"expert-like\" pattern of cortical activity, neurofeedback training failed to selectively enhance performance, as both groups improved their putting performance similarly from the pre-test to the post-test. Finally, both groups performed robustly under pressure. Conclusions: Performers can learn to regulate their brain activity using neurofeedback training. However, research identifying the cortical correlates of expertise is required to refine neurofeedback interventions if this training method is to expedite learning. Suggestions for future neurofeedback interventions are discussed. --------------------------------------------------------------------------------","Neurofeedback training provides individuals with real-time information about their level of cortical activity via sounds or visual displays (Hammond, 2007). Based on principles derived from operant learning theory (e.g., Skinner, 1963), rewarding positive reinforcement, such as a change in the pitch of a tone, is provided when a desired level of cortical activity is achieved. Electroencephalography (EEG) is perhaps the most common brain imaging method that is used to provide neurofeedback training (e.g., Vernon, 2005). In brief, EEG involves the recording of electrical activity on the scalp to detect voltages generated in the brain. EEG offers exquisite temporal resolution, whereby changes in activation are detected more or less instantaneously (e.g., Harmon-Jones & Peterson, 2009). Moreover, EEG can be measured while participants stand and perform a range of movements, which makes the method particularly well suited for providing neurofeedback in sport (e.g., Thompson, Steffert, Ros, Leach, & Gruzelier, 2008). To this end, there have been a handful of studies investigating whether EEG neurofeedback training can facilitate performance in sport, and while the evidence concerning the effectiveness of EEG neurofeedback is not conclusive, it is certainly encouraging (e.g., Arns, Kleinnijenhuis, Fallahpour, & Breteler, 2007; Kavussanu, Crews, & Gill, 1998; Landers et al., 1991; Rostami, Sadeghi, Karami, Abadi, & Salamati, 2012). For instance, the seminal study of neurofeedback in sport was conducted by Landers et al. (1991), and investigated the effects of neurofeedback in sixteen experienced archers. Landers et al. (1991) reasoned that archery performance should be associated with activation of the right-hemisphere of the brain, which is associated with visual-spatial processing, and deactivation of the left-hemisphere of the brain, which is associated with verbal-analytic processing (e.g., Hatfield, Landers, & Ray, 1984; Landers et al., 1994). Accordingly, they measured EEG activity and archery performance in pre- and post-test sessions, separated by approximately 60 min of neurofeedback training during which the archers watched their relative left- and right-hemisphere activity on a visual display. Results revealed that performance improved from the pre-test to the post-test in eight archers who were rewarded when they reduced cortical activity over their left-hemisphere. In contrast, performance deteriorated in the remaining eight archers, who were rewarded when they reduced cortical activity over their right-hemisphere. Although this finding implies that neurofeedback training could be used to expedite learning in archery, it is important to note that left- hemisphere cortical activity in the pre- and post-test sessions was the same for members of both neurofeedback groups. This indicates that the neurofeedback training protocol did not cause members of the left-hemisphere neurofeedback group to suppress left-hemisphere function. Consequently, the improvement in performance that was achieved by this group may not be directly attributable to the neurofeedback that they received. A more recent study of neurofeedback training by Rostami et al. (2012) also adopted a pre- and post-test design to increase the power of the sensory motor rhythm (i.e., cortical activity between 13 and 15 Hz) over central motor areas (i.e., C3 electrode site) of the brain. Specifically, twelve experienced marksmen attended 15 h of laboratory sessions spread over five weeks, and were trained to control their cortical activity by sitting and watching this activity on a screen. Results revealed that neurofeedback training led to marginal improvements in shooting accuracy from the pre-test to the post-test, whereas the performance of a control group who received no neurofeedback training was unchanged. While this finding is also supportive of neurofeedback as a tool to aid the development of expertise and excellence in sport, the study was subject to two principal limitations. First, the choice to train participants to increase the sensory motor rhythm over central motor areas was somewhat arbitrary, with the authors providing no theoretical or empirical rationale to support this key methodological feature. Second, no measures of EEG activity were obtained during the pre- and post-test sessions. Accordingly, it was impossible to evaluate whether the beneficial effects of neurofeedback training were attributable to participants having learned to control their patterns of cortical activation. Finally, perhaps the most informative neurofeedback study in sport was conducted by Arns et al. (2007). They adopted a crossover design in which six amateur golfers completed 12 blocks of putts. The golfers putted as normal in the odd-numbered blocks, and putted while receiving auditory neurofeedback training in the even-numbered blocks. Importantly, the element of cortical activity that was fed back to participants was partly customised to the task. Specifically, a comparison of cortical activity associated with the best (i.e., holed) and worst (i.e., missed) putts during a baseline session was conducted to customise the neurofeedback for each participant. This resulted in participants being trained to reduce a combination of theta (4–8 Hz), alpha (8–12 Hz), sensory motor rhythm (13–15 Hz) and/or beta (15–30 Hz) power in the final moments preceding putts. It is also important to note that the auditory neurofeedback tone was played to participants while they stood over the ball and prepared to execute putts. By adopting these innovative design features, Arns et al. (2007) were the first researchers to provide customised, concurrent neurofeedback training during task performance. Their results revealed that participants holed more putts during the blocks in which they received neurofeedback compared to those in which they did not. Although this study provides arguably the strongest support for the efficacy of neurofeedback training as a tool to foster expertise and excellence in sport, it nonetheless suffers from key limitations, including low sample size and no control group. Thus, the results of the Arns et al. (2007) study may simply reflect a placebo effect whereby improved performance was elicited by the presence of the neurofeedback system and auditory tone, rather than by changes in cortical activity per se.","Since the work of Arns et al. (2007), two studies have systematically examined the patterns of cortical activity that underpin successful golf putts, and the results of these studies could form the empirical grounding for new neurofeedback interventions. Specifically, Babiloni et al. (2008) compared patterns of EEG activity characterising holed putts and missed putts in a sample of expert golfers. They found a widespread reduction in EEG alpha power during the four seconds preceding putts. This is consistent with the well-established finding that voluntary self-paced movements are preceded by a reduction (i.e., desynchronisation) in EEG alpha power (around 8–12 Hz) in both hemispheres of the brain during bimanual tasks (e.g., Leocani, Toro, Manganotti, Zhuang, & Hallett, 1997; Pfurtscheller & Aranibar, 1979). Crucially, Babiloni et al. (2008) also found that compared to missed putts, holed putts were characterised by a greater reduction in high-alpha power (10–12 Hz) at sites overlying the premotor and motor cortex (e.g., Fz, Cz, C4), indicating that these sites and the high-alpha frequency band are ideal candidates to be targeted by neurofeedback interventions. Second, the study by Babiloni et al. (2008) was recently replicated and extended by Cooke et al. (2014). Specifically, Cooke et al. (2014) compared cortical activity preceding holed versus missed putts, in both experts and novices. The results showed that expert golfers displayed a greater reduction in high-alpha power than novices, and that holed putts were characterised by less high-alpha power than missed putts, at frontal and central sites (e.g., Fz, F3, F4, Cz) in the two seconds preceding movement. In sum, the results of Babiloni et al. (2008) and Cooke et al. (2014) provide consistent evidence that expertise and optimal golf putting performance is characterised by a suppression of EEG high-alpha power in the final moments preceding movement initiation.","To address the limitations of previous studies, the present experiment was designed to be the largest neurofeedback investigation to date, while also being the first to employ concurrent, empirically grounded neurofeedback along with an active control group. Our empirically guided neurofeedback was based on the results of Babiloni et al. (2008), and Cooke et al. (2014), and aimed to teach participants to reduce their frontal high-alpha power before striking putts. Importantly, we also investigated the effects of neurofeedback training on performance under pressure. Previous research has demonstrated that increases in psychological pressure can disrupt patterns of physiological activity during motor preparation and cause impaired performance (e.g., Cooke, Kavussanu, McIntyre, & Ring, 2010; Weinberg & Hunt, 1976). Consequently, our experiment is the first to examine whether neurofeedback training can help performers produce consistent physiological responses which could ensure robust performances across both low and high-pressure conditions. We hypothesised a series of interactions in which the neurofeedback group would display greater reductions in pre-movement high-alpha power, and greater improvements in performance, from a pre-intervention test to a post-intervention test, when compared to a control group. We also expected individuals in the neurofeedback group to produce consistent patterns of cortical activity and robust performances across both low- and high-pressure conditions during the post-intervention test.","Twenty-four right-handed male golfers volunteered to participate in the experiment, and were randomly assigned to a neurofeedback training group (N = 12, M age = 23.00, SD = 5.83 years, M golf handicap = 23.00, SD = 6.62) or a control group (N = 12, M age = 21.00, SD = 2.52 years, M golf handicap = 23.33, SD = 4.62). All participants provided informed consent before taking part. Task ~~~~ Participants used a standard length (90 cm) blade style golf putter (Titleist Scotty Cameron Circa 62) to putt regular-size golf balls (diameter 4.7 cm) towards a standard- size hole (diameter 10.8 cm) from a distance of 2.4 m. The hole was located 1.25 m from the end and 0.75 m from the sides of a flat artificial putting surface (Turftiles), which had a stimpmeter reading of 3.05 m (i.e., a medium-to-fast paced green). Design ~~~~~~ We adopted a mixed-multifactorial design. Participants attended a pre-test and a post-test (hereafter referred to as the test phase of the experiment), which included both low- and high-pressure conditions. They also attended three training sessions following the pre- test and preceding the post-test (hereafter referred to as the acquisition phase of the experiment); each session consisted of twelve 5-min blocks of putts. Our design therefore included group (neurofeedback, control) as a between-subject factor, as well as pressure (low, high), session (pre-test, acquisition 1, acquisition 2, acquisition 3, post-test), block (referring to each 5-min block of putts within the training sessions), and epoch as within-subject factors. Epoch refers to the time-windows around movement in which cortical activity was assessed (i.e., −4 to −3 s, −3 to −2 s, −2 to −1 s, −1 s to 0 s, 0 s to +1 s). The inclusion of this factor was in keeping with previous studies (e.g., Babiloni et al., 2008; Cooke et al., 2014). Further details are provided in the data reduction and statistical analyses sections below. Neurofeedback training protocol ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ During the acquisition phase of the experiment, participants in the neurofeedback group received three 1-h sessions of neurofeedback training. Cortical activity was recorded from the Fz site on the scalp using an active electrode connected a DC amplifier (Brainquiry PET-4), with reference and ground electrodes attached to the right and left mastoids, respectively. We decided to provide feedback at the Fz site because this site has been shown to capture the strongest differences in high-alpha power between experts and novices and successful and unsuccessful outcomes in the moments preceding golf putts (Cooke et al., 2014). In tandem with cortical recordings, a computer running Bioexplorer software (Cyberevolution) extracted high-alpha power from the EEG signal, and fed this back to participants in the form of an auditory tone. Importantly, the tone was programmed to vary in pitch based on the level of high-alpha power. Moreover, the tone was set to turn off completely (i.e., be silenced) when the following criteria were met: a) high-alpha power was reduced (relative to each participant's individual baseline) by 26.8%, 53.6%, and 80.4% in the first, second, and third training sessions, respectively; b) theta power was reduced (relative to each participant's individual baseline) by 18.2%, 36.4%, and 54.6%, in the first, second, and third training sessions, respectively; c) impedance was low, as reflected by <10 μV of 50 Hz activity in the signal (Arns et al., 2007); d) eye-blinks were absent, as detected by an active electrode that was placed over the orbicularis oculi muscle of the right eye. Participants were required to reduce theta power alongside high- alpha power because theta power at the Fz site has been revealed as an additional marker of expertise in the moments preceding golf putts (Cooke et al., 2014). The high-alpha and theta power threshold values initially were based on the performance of experts (Cooke et al., 2014) but subsequently were refined in pilot testing. They were designed to encourage gradual development of participants' ability to reduce high-alpha power, with the thresholds in the final training session roughly corresponding to the reductions in power that characterised the successful putts of experts in previous studies (i.e., Babiloni et al., 2008; Cooke et al., 2014). In line with the study by Arns et al. (2007), the neurofeedback tone was turned off completely for 1.5 s once the neurofeedback thresholds had been reached, except for when high-impedance or an eye-blink was detected within this dwell-period, in which case the tone re-started immediately. Auditory neurofeedback was provided to participants in twelve 5-min blocks per training session while participants practiced putting (i.e., concurrent). Specifically, participants were instructed to execute putts when they felt ready, but only when the tone was turned off (i.e., indicating that high-alpha power had been reduced). Participants in the control group followed an identical procedure to those in the neurofeedback group, except the tone that they heard was not based on their brain activity. Instead, participants in this group were played a recording of the tone from a matched (i.e., yoked) participant in the neurofeedback group. This feature ensured that the tone was turned off on the same number of occasions and for the same duration in both groups, and guaranteed that a similar number of putts were completed by both groups of participants during the acquisition phase of the experiment. Pressure manipulation ~~~~~~~~~~~~~~~~~~~~~ Two pressure conditions (each consisting of 50 putts) were manipulated in the test phase of the experiment using social evaluation, competition, and rewards (Baumeister & Showers, 1986), with the order of the conditions counterbalanced across participants. The two conditions are described below. Low-pressure condition This condition was non-competitive, contained no rewards, and was designed to minimise any pressure that may have been elicited by social evaluation. Specifically, participants were informed that performance would be assessed by the average distance of putts from the hole, with holed putts counting 0 cm in the calculation. Crucially, participants were also told that although the accuracy of each putt would be recorded, their individual performance would not be analysed. Instead, it was explained that the performance of all participants would be pooled to generate one accuracy score for the sample as a whole. High-pressure condition The high-pressure condition was set up as a competition that offered rewards, and placed an explicit emphasis on social evaluation. Specifically, participants were informed that they would be individually evaluated in this block of putts. To this end, they were told that all participants would be ranked on a leaderboard based on their average distance from the hole in this condition. Moreover, they were informed that the leaderboard would be emailed to all participants at the end of the study, and that cash prizes of £50, £25, £10 and £5 would be awarded to the top four performers (e.g., Wilson, Smith & Holmes, 2007). Finally, they were told that twenty-four participants were to be recruited for the study, allowing each individual to evaluate their chances of winning a prize (e.g., Cooke et al., 2010). Number of putts struck We recorded the number of putts struck during each block in the acquisition phase of the experiment. The purpose of this measure was twofold. First, it allowed us to identify and control for possible group differences in the number of putts (i.e., practice) during skill acquisition. Second, we were able to indirectly assess the extent to which participants in the neurofeedback group learned to produce the desired high-alpha power profile, with better control of cortical activity being indicated by more putts being struck per 5-min block. Performance Mean radial error (i.e., the mean distance of the balls from the hole) was recorded as our measure of performance (Cooke et al., 2010; Cooke, Kavussanu, McIntyre, Boardley, & Ring, 2011). Zero was used in the calculation of mean radial error on trials where the putt was holed (Hancock, Butler, & Fischman, 1995). Mean radial error scores were obtained by taking a photograph of the final position of each putt using a digital camera (Sony Handycam) suspended above the hole. Photographs were then analysed offline using custom developed “ScorePutting” software (see Neumann & Thomas, 2008). Pressure manipulation check The 5-item pressure/tension subscale of the Intrinsic Motivation Inventory (Ryan, 1982) was used to test the effectiveness of our pressure manipulation. Items, including “I felt pressured”, were rated on a 7-point Likert scale, with labels of 1 (not at all true), 4 (somewhat true), and 7 (very true). The item responses were averaged to provide one score for the subscale. Cooke et al. (2011) reported reliability coefficients ranging from .66 to .90 for this subscale. In this experiment, alpha coefficients in the low and high-pressure conditions during the pre-test and post-test sessions ranged from .66 to .85, demonstrating acceptable internal consistency. Cortical activity We recorded EEG activity during the test phase of the experiment from an array of 16 silver/silver chloride pin electrodes on the scalp (Fp1, Fp2, F4, Fz, F3, T7, C3, Cz, C4, T8, P4, Pz, P3, O1, Oz, O2) positioned in accordance with the 10-20 system (Jasper, 1958). The BioSemi EEG system replaces the conventional “ground” electrode with two separate electrodes placed on either side of the vertex: a common mode sense active electrode and a driven right leg passive electrode; the “reference” at the time of recording lies somewhere between these two electrodes. These electrodes form a feedback loop which serves to reduce common mode voltage and thereby increase the signal-to-noise ratio above what would be achieved by systems using conventional ground electrodes with the same impedance (e.g., Metting van Rijn, Peper, & Grimbergen, 1990; see also Biosemi website: http://www.biosemi.com/faq/cms&drl.htm). Electrodes were also placed at the left and right mastoids, to permit offline referencing. All signals were amplified and digitized at 512 Hz with 24-bit resolution (Biosemi ActiveTwo) using ActiView software (Biosemi). Conductive gel (ECI Electro-gel) was applied to all recording electrodes, and all sites were abraded using a blunt needle (for sites on the scalp) and a combination of abrasive paste (Nuprep) and alcohol wipes (Mediswab) (for the mastoids) prior to electrodes being attached. Performance Mean radial error was used as our measure of performance, as described above.","The protocol was approved by the local research ethics committee. Participants attended five 2-h testing sessions (i.e., pre-test, acquisition 1, acquisition 2, acquisition 3, post-test) on separate days (to prevent any confounding effects of fatigue), with each session being an average of 2.3 (SD = 2.7) days apart. In the pre-test session, participants were briefed, instrumented to allow the recording of cortical activity and provided with instructions about the golf-putting task. Specifically, they were asked to try to get all putts “ideally in the hole, but if unsuccessful, to make them finish as close to the hole as possible.” Next, they performed 20 familiarization putts to become accustomed to the putting surface and to putting while instrumented for EEG recordings. After the familiarization block, participants performed two blocks of 50 putts, which represented the low- and high-pressure conditions. Each block was preceded by its respective pressure manipulation as described above. After each putt, a photograph was taken to record the terminal location of the ball, and then the ball was replaced at the start position by the experimenter. This ensured that participants did not need to move between trials, thereby keeping movement artefacts to a minimum, while also regulating the time between putts, which approximately ranged from 15 to 30 s. The manipulation check (i.e., pressure/tension subscale of Intrinsic Motivation Inventory, as described above) was administered immediately after the final putt in each of the pressure conditions, while cortical activity was recorded continuously during each block. On completion of the pre-test, participants were thanked and reminded of the time and date of their acquisition sessions. In the acquisition phase of the experiment, participants were welcomed to the laboratory and instrumented with the neurofeedback system. During the first acquisition session, they were asked to address and fixate on a golf ball for five seconds. This procedure was repeated five times in order to calculate their average baseline high-alpha power. Customized computer scripts were then prepared for each participant in the neurofeedback group in order to set the neurofeedback tone to silence when high-alpha power was reduced from their individual baseline by 26.8%, 53.6%, and 80.4%, in acquisition sessions one, two, and three, respectively. Participants were then issued with the following instructions: “The computer will play a tone that is linked to your brain activity. When you reach a prescribed level of brain activity, the tone will turn off. We would like you to produce this level of brain activity before you putt, so we want you to address the ball and get ready to putt, and then wait for the tone to turn off before executing your stroke. You will receive 1-h of neurofeedback training today in the form of 12 blocks of the tone being played for 5-min at a time. During each 5-min block, you will putt as many balls as is permitted in the time, which will depend on the number of times that you manage to turn the tone off. It does not matter how many balls you putt.” Participants then completed the twelve 5-min blocks of putts, interspersed with 2-min breaks between each block. After each putt a photograph was taken to record the terminal location of the ball (e.g., Neumann & Thomas, 2008). In acquisition sessions two and three, participants were instrumented, reminded of their instructions, and immediately began the 12 blocks of putts (i.e., the baseline high-alpha measure only occurred during acquisition session 1). At the end of each session participants were thanked and reminded of the time and date of their next visit. We decided to fix the amount of exposure to the tone (rather than fix the number of putts) because we wished to standardize the amount of experience with the feedback stimulus. By standardizing exposure our neurofeedback training protocol was consistent with other recent neurofeedback studies (e.g., Rostami et al., 2012). After completing the acquisition phase of the experiment, participants attended the post-test session, which was identical to the pre-test. On completion of the post-test, participants were thanked and debriefed. Upon completion of the study, leaderboards were emailed to participants and competition winners were contacted and paid their prize money. Data reduction ~~~~~~~~~~~~~~ Individual trials within the continuous cortical recordings were identified using an optical sensor (Datasensor S51-PA 2-C10PK), which detected the initiation of putts, and a microphone (Rode NT1) connected to a mixing desk (Studiomaster Club 2000), which detected the putter-to-ball contacts. These signals were recorded using both Actiview (Biosemi) and Spike2 (Cambridge Electronic Design) software. Processing of EEG data recorded during the test phase of the experiment was conducted with EEGLab software (Delorme & Makeig, 2004) using the following procedure: First, datafiles were resampled (256 Hz), filtered (1–50 Hz) and referenced to the average mastoid. Next, a neutral EEG baseline was identified (cf., Babiloni et al., 2008). Specifically, we performed a fast Fourier transform (1 Hz bins, Hanning window taper) spanning 7 s before until 1 s after initiation of each putt. We then performed exploratory analyses to identify a period within this 8 s window during which cortical activity was similar across both between (i.e., group) and within-subjects factors (i.e., session and pressure). Specifically, potential baselines were initially identified by eye and then subjected to a series of 2 Group × 2 Session × 2 Pressure ANOVAs to verify the absence of main or interaction effects which, if present in the baseline, would confound our interpretation of subsequent results. ANOVAs confirmed no main or interaction effects for high-alpha power in the period from 0 ms to 200 ms around movement, so we selected this window as a neutral baseline, before proceeding with the remaining data processing steps as follows: First, we created new epochs spanning 5 s before until 1 s after each putt (e.g., Babiloni et al., 2008), and performed baseline removal (i.e., subtracted power during the baseline period from power in the other epochs). Next, we screened the data to reject any artefacts. Gross artefacts were removed by rejecting any large deviations in the signal (>100 μV) from the baseline level. This important pre-processing step (see Onton, Westerfield, Townsend, & Makeig, 2006) was followed by independent component analyses, which, in combination with the ADJUST algorithm (Mognon, Jovicich, Bruzzone, & Buiatti, 2011), were used to identify and remove remaining artefacts including eye blinks, eye movements, and the blood pressure pulse. Next, we performed a fast Fourier transform (1 Hz bins, Hanning window taper) on the artefact-free epochs, and averaged the data in successive 1 s epochs from 4 s before (i.e., preparatory period) until 1 s after (i.e., movement period) the initiation of putts. Finally, we computed power in the high-alpha (10–12 Hz) frequency band. For brevity of reporting, only the results from the key Fz electrode, and those in its immediate surroundings (i.e., F3, F4, Cz) are presented. We selected these electrodes because they roughly overlie the primary motor cortex, the premotor cortex, and the supplementary motor areas that are related to movement control (e.g., Ashe, Lungu, Basford, & Lu, 2006), and which have been implicated in previous EEG-based golf-putting research (Babiloni et al., 2008). Moreover, these electrode sites captured the strongest effects, and were largely representative of the other sites. Statistical analyses Performance and number of putts struck in the acquisition phase of the experiment were subjected to 2 Group × 3 Session × 12 Block Analyses of Variance (ANOVAs). Measures of cortical activity during the test phase were subjected to 2 Group × 2 Session × 2 Pressure × 5 Epoch (i.e., −4 to −3 s, −3 to −2 s, −2 to −1 s, −1 s to 0 s, 0 s to +1 s) ANOVAs. Finally, the pressure manipulation check and performance recorded during the test phase were subjected to 2 Group × 2 Session × 2 Pressure ANOVAs. These analyses were designed to assess the three aims of our experiment as explained in the results section. Significant effects were probed by polynomial trend analyses and planned post hoc comparisons. The results of univariate tests are reported, with the Huynh-Feldt correction procedure applied to analyses, which violated the sphericity of variance assumption. Partial eta- squared is reported as a measure of effect size, with values of .02, .12, and .26 indicating relatively small, medium, and large effect sizes, respectively (Cohen, 1992). Acquisition phase ~~~~~~~~~~~~~~~~~ Analyses of the measures obtained in the acquisition phase of the experiment allowed us to indirectly assess whether neurofeedback taught participants to produce expert like cortical activity (i.e., number of putts struck) and promoted expedited learning. Test phase ~~~~~~~~~~ Analyses of the measures obtained in the test phase of the experiment allowed us to directly assess whether neurofeedback taught participants to produce expert like cortical activity and promoted expedited learning. They also allowed us to assess whether neurofeedback encouraged consistent patterns of cortical activity and robust performance across both low and high levels of pressure.","Neurofeedback training offers the exciting possibility of expedited motor skill acquisition, and, thus, more efficient development of excellence and expertise in sport (Gruzelier, Egner, & Vernon, 2006). This experiment is one of the first to empirically investigate the effectiveness of neurofeedback training as a tool to accelerate learning in sport. Our aims were to: a) examine whether neurofeedback training could teach recreational golfers to produce the patterns of brain activity that are exhibited by experts in the moments preceding successful putts; b) evaluate whether neurofeedback training could improve performance and thereby accelerate the novice to expert evolution; and c) examine whether neurofeedback training could help produce patterns of cortical activity and performance levels that would be robust to the potential deleterious effects of increased psychological pressure. Our results are discussed in the sections that follow. Effects of neurofeedback training on cortical activity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We supported our hypothesis that neurofeedback training would be successful in teaching recreational golfers to suppress frontal high-alpha power during preparation for putting. First, in the acquisition phase of the experiment, we revealed an effect of block, indicating that participants were able to complete more putts in five minutes in those blocks towards the end of each training session compared to those at the start. This provides indirect evidence that the neurofeedback group learned to reduce high-alpha power and thereby turn off the neurofeedback tone more frequently as the blocks in each training session progressed. Second, we provided the first direct evidence of neurofeedback training leading to selective changes in cortical activity, as implied by the group by session interactions for EEG high-alpha power at frontal sites. Specifically, members of the neurofeedback group reduced their pre-movement high-alpha power from the pre-test to the post-test, whereas members of the control group did not. This key finding clearly reveals for the first time that relatively brief neurofeedback training (i.e., 3 h over three separate days) can teach recreational golfers to regulate their cortical activity even when the source of neurofeedback (i.e., the auditory tone) has been removed. Effects of neurofeedback training on performance ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our hypothesis that neurofeedback training would facilitate performance and thereby expedite the development of expertise was not supported. Although members of the neurofeedback group learned how to suppress their high-alpha power in the moments preceding putts, their performance during acquisition and the improvement in their mean radial error score from the pre-test to the post-test were similar to the improvements achieved by members of the control group. This separation between cortical activity and behavioural outcome could indicate that reduced pre-movement high-alpha power is not a key determinant of golf-putting success. However, based on the evidence provided by Babiloni et al. (2008) and Cooke et al. (2014), this conclusion seems unlikely. There are three alternative explanations that could reconcile this discrepancy. First, it must be conceded that the pattern of high-alpha power that our neurofeedback group learned to produce was not entirely representative of the activity that is produced before the successful putts of experts. Specifically, our participants produced a general suppression of high-alpha power spanning the entire four second preparatory period, whereas experts have been shown to produce a sharp decline in high-alpha power that is most evident during the final two seconds before and during movement (Cooke et al., 2014). It is thus possible that a dynamic reduction in high-alpha power (i.e., event-related desynchronisation) is required in order to deliver putting success (e.g., Babiloni et al., 2008). However, to counter this claim, our neurofeedback group produced a dynamic reduction in pre-movement high- alpha in the pre-test, yet they performed worse than in the post-test, where a dynamic reduction was absent (Fig. 2). Similarly, the control group produced a dynamic reduction in high-alpha power in both pre-test and post-test sessions, but this did not yield better performances than the neurofeedback group. A second alternative is that the method employed to elicit the pre-movement reduction in high-alpha power is important for putting success, rather than the reduction in high-alpha power per se. For example, it is well- established that high-alpha power has an inverse relationship with cortical activity, such that a decrease in high-alpha power reflects heightened activation (e.g., Pfurtscheller, 1992). Accordingly, it would have been possible for the recreational golfers in our study to learn that they could reduce their high-alpha power and thereby turn off the neurofeedback tone by engaging in a number of cognitive activities that are irrelevant to golf putts (e.g., Nowlis & Kamiya, 1970). In contrast, one may speculate that expert golfers produce their reduction in high-alpha power by engaging specific working memory processes in order to programme movement parameters such as direction and force (e.g., Cooke, 2013). To counter this suggestion, one could argue that if members of our neurofeedback group were engaging in irrelevant activities to turn off the neurofeedback tone, their performance is likely to have been impaired. Nevertheless, it would be useful for future studies to examine the precise strategies that expert golfers use to produce reductions in pre-movement high-alpha power. The outcomes of such studies could provide cues to help ensure that participants regulate their cortical activity by engaging in the same processes as experts in future neurofeedback interventions. Finally, it is possible that it is a suppressed absolute level of high-alpha power in those seconds immediately preceding and during movement (i.e., –2 s to +1 s) that is the key to expertise and putting success (e.g., Babiloni et al., 2008; Cooke et al., 2014). In the present study, neurofeedback training led to suppressed high-alpha power early during movement preparation, but in the final second before and during movement the high-alpha power of the neurofeedback group and the control group was the same (Fig. 2). This seems the most likely explanation for the null effect of neurofeedback training on performance in our experiment. It could be attributed to our neurofeedback tone being silenced for 1.5 s once the prescribed high-alpha power threshold had been breached (cf. Arns et al., 2007). This afforded the possibility for participants to begin to increase high-alpha power in the final moments before and during movement, while the neurofeedback tone remained turned off. This potential artefact could be remedied in future studies by removing a pre-set silence duration once the neurofeedback threshold is achieved. Effects of neurofeedback on cortical activity and performance under pressure ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our final hypothesis was that neurofeedback training would help promote consistent preparatory cortical activity and performances under both low and high levels of pressure. We found that performance levels remained robust for both the neurofeedback and the control groups, so neurofeedback training failed to yield any selective benefits for performance during pressured conditions. The null effects of pressure on performance could be attributed to the high number of trials that were required to generate meaningful EEG data (e.g., Luck, 2005). Specifically, it is known that multiple trials dilute the strength of pressure manipulations, in this case providing participants with several chances to redeem bad putts, and likely eliciting levels of pressure far below those experienced in real-life (Cooke et al., 2010; Woodman & Davis, 2008). Future studies should afford special consideration to methods of maximising the impact of the pressure manipulations, especially when large numbers of trials are planned. Finally, in contrast to our hypothesis, we found that neurofeedback training failed to inoculate participants against pressure-induced changes in cortical activity. Elevated pressure served to increase the change in high-alpha power over time at sites overlying frontal regions of the cortex (i.e., Fz, F3) in both groups of participants. Given that high-alpha power is inversely related to cortical activity (e.g., Pfurtscheller, 1992), increased high-alpha power during the early phases of motor preparation in the high-pressure condition could reflect worrisome thoughts diverting attentional resources away from the motor planning (i.e., frontal) regions of the brain (Fig. 3). It is possible that more training sessions were required in order for the volitional control of cortical activity to outweigh the involuntary changes that appear to be induced by increases in psychological stress. Limitations and future directions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In addition to the potential limitations of the neurofeedback protocol outlined above, we concede that the sample size was relatively small. However, the sample size was larger than those adopted in relevant previous studies (e.g., Arns et al., 2007 N = 6; Landers et al., 1991 N = 16). Moreover, our study was sufficiently powered to detect a number of main and interaction effects as detailed above. Notwithstanding, it may be beneficial for future studies to replicate and extend our experiment with a larger cohort. Future research would also do well to investigate other aspects of cortical activity, such as EEG coherence neurofeedback training. In brief, EEG coherence analyses assess the degree of linear interrelatedness between signals recorded at two sites, with high coherence commonly thought to represent functional communication between the two sites (e.g., Thatcher, Krause, & Hrybyk, 1986). To this end, it has recently been suggested that a reduction in coherence between the left-hemispheric verbal-analytic regions (e.g., electrode site T3) and the motor planning regions (e.g., electrode site Fz) of the brain represent a reduction in verbal-analytic information processing, and characterise successful performance across golf putting, rifle shooting and surgical laparoscopy tasks (e.g., Deeny, Hillman, Janelle, & Hatfield, 2003; Zhu, Poolton, Wilson, Hu, et al., 2011; Zhu, Poolton, Wilson, Maxwell, & Masters, 2011). This understanding of what T3-Fz coherence reflects presents researchers with a first opportunity to supplement neurofeedback with coaching strategies (e.g., implicit learning techniques – see Masters & Poolton, 2012; Zhu, Poolton, Wilson, Maxwell, & Masters, 2011; Zhu, Poolton, Wilson, Hu, et al., 2011) to help ensure that trainees produce the desired “expert-like” coherence patterns through relevant methods. Such a cross-pollination of training methods could increase the likelihood of neurofeedback yielding expedited learning and robust performances under stress (e.g., Zhu, Poolton, Wilson, Hu, et al., 2011; Zhu, Poolton, Wilson, Maxwell, & Masters, 2011). Finally, future research should also carefully investigate the mechanisms which underpin the expected beneficial effects of neurofeedback training on performance. For instance, researchers could conduct source localisation analyses to more precisely identify the individual brain structures that are responsible for generating the EEG activity measured at the scalp. This would offer insight into the physiological mechanisms of excellence. Alternatively, researchers could conduct statistical mediation analyses to verify proposed causal relationships among cortical, cardiac, somatic and motor systems. We were unable to perform such analyses here because neurofeedback failed to produce selective benefits for performance.","In conclusion, this is the first experiment to show that brief neurofeedback training (i.e., 3 h) can teach recreational golfers to regulate their cortical activity during a retention test. Importantly, and in line with the themes of this special issue, these results provide proof of principle that neurofeedback training can be used to alter cortical activation, thereby providing further support for the notion that this technique can help expedite the development of expertise in athletes. Neurofeedback training failed to yield any selective benefits for performance in the present study. This could be explained by limitations of our neurofeedback protocol as outlined above. By refining neurofeedback protocols and furthering our understanding of the cortical correlates of expertise, future studies could expedite development of expertise and excellence, and neurofeedback can increasingly feature in motor learning protocols in the years and decades to come."],["Trait emotional intelligence (TEI) is emerging as a useful and promising individual difference in predicting vocational behavior (e.g., Di Fabio & Saklofske, 2014). Little is yet known about the underlying processes that may lead TEI to associate with career related outcomes. This study investigates the role of career adaptability in mediating the association between TEI and career decision-making difficulties and self-perceived employability, in a sample of Swiss university students (N = 400). The results of a series of path analysis in which we controlled for intelligence, sex and personality showed that career adaptability fully mediated the effect of TEI on self-perceived employability and career decision-making difficulties, in particular the subscales of lack of information and inconsistent information. Our findings shed light on the role of regulatory processes in shaping the effects of TEI on career-related outcomes. --------------------------------------------------------------------------------","Our contemporary globalized world involves managing increasingly uncertain professional trajectories (Guichard, 2015), and career uncertainty is known to be associated with higher anxiety (Fuqua, Seaworth, & Newman, 1987). In this context, a better understanding of one's emotional experience has been found to play an important role in career-related issues (e.g., Di Fabio & Saklofske, 2014). However, little is still known about the affective components that may impact the process of career exploration and career development. The aim of this study is to understand the pathway through which emotional self-perceptions representing the affective aspects of personality, i.e., trait emotional intelligence (TEI; Petrides, Pita, & Kokkinaki, 2007), may affect career related outcomes, such as career indecision and self-perceived employability. TEI, defined as the emotional traits that reflect self-perceptions regarding one's ability to deal with emotions, is associated with important career related outcomes including career indecision, career indecisiveness and career decision-making self-efficacy (Di Fabio & Saklofske, 2014), with self-perceived employability (Di Fabio & Kenny, 2015), and in predicting career success (De Haro García & Castejón Costa, 2014). The path through which TEI and career indecision and self-perceived employability are related, however, remains unexplored in the literature, with no study to our knowledge yet investigating the mediating paths between these variables. Social Cognitive Career Theory (SCCT; Lent, Brown, & Hackett, 1994) posits that the relationship between dispositions and career choice is not direct, but mediated by internal processes, such as self-efficacy. Similarly, Rossier (2015) highlights the role of regulatory processes, including career adaptability, in mediating the relationship between individual dispositions and career behaviors. Interestingly, regulatory processes have a strong adaptive function in allowing dispositions to fit the characteristics of the environment. Career adaptability, career indecision and self-perceived employability as career-related outcomes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Studies indicate the potential for TEI to have an indirect effect on career indecision and self-perceived employability through career adaptability (Coetzee & Harry, 2014; Harry, 2017). Career adaptability refers to a set of personal resources that help individuals manage career transitions (Savickas & Porfeli, 2012). Some studies have highlighted the association between career adaptability and career decision-making difficulties (e.g., Hirschi, Herrmann, & Keller, 2015). A recent meta-analysis (Rudolph, Lavigne, & Zacher, 2017) showed career adaptability to be significantly associated with employability. Of note, no studies explored the possible mediator effect of career adaptability on the relationship between TEI and career related outcomes. We posit that the psychosocial resources of career adaptability may account for the effect of trait emotional intelligence on career-related outcomes. To explore the mediating role of career adaptability we chose two career-related constructs to capture the more negative and more positive aspects of transition to employment: career decision-making difficulties (Gati, Krausz, & Osipow, 1996), more generally termed career indecision, and self-perceived employability. Previous work has primarily emphasized the cognitive aspects of career indecision (e.g., Gati & Tal, 2008). Some authors have also identified its affective aspect: Betz and Sterling (1993) found that chronic career indecisiveness was strongly correlated with fear of commitment to a decision, an affective disposition. Saka and Gati (2007) found emotional and personality-related factor affecting severe career decision- making difficulties. Self-perceived employability is defined as the characteristics needed to secure a job that corresponds to one's interests and goals (Rothwell & Arnold, 2007). Indeed, self-perceived employability appears to be related to higher job satisfaction, higher work engagement (Ngo, Liu, & Cheung, 2017), and higher perceived marketability (De Vos, De Hauw, & Van der Heijden, 2011). Overall, both career indecision and self-perceived employability capture fundamental aspects of career-related outcomes, and may help elucidate the role of TEI and career adaptability in contributing to career development. Personality, intelligence and sex ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The Big Five personality traits have been shown to associate with TEI (Saklofske, Austin, & Minski, 2003) and with career adaptability, especially conscientiousness (Rudolph et al., 2017). Although previous studies have established no significant effects of fluid intelligence on career decision (Di Fabio & Saklofske, 2014), this variable logically could affect career adaptability. Sex also may impact TEI (Petrides, Furnham, & Martin, 2004), although previous studies have shown no effects of sex on career adaptability (Rudolph et al., 2017). The present study ~~~~~~~~~~~~~~~~~ Based on the framework delineated above, the present study examined the indirect effect of TEI on career indecision making difficulties and self-perceived employability through career adaptability, controlling for personality traits, fluid intelligence, and sex. The strong association between career indecision and emotional intelligence is widely acknowledged in the literature, although no study to our knowledge has looked at the process that may explain this association. Furthermore, self-perceived employability is a key element for career development and success that has not been much studied in relation to TEI. We hypothesized that career adaptability would fully mediate the relationship between TEI and career indecision, including its 3 subscales (H1) and between TEI and self-perceived employability (H2).","We recruited 400 participants (46% female), ranging in age from 17 to 48 (Mean = 21.39 and SD = 3.27) from several universities in the Lausanne area (Switzerland) through the university subject-pool. Participants were bachelor (69,8%) and master or advance studies students (30,2%) in various disciplines. They gave written consent to participate in the study and were compensated for their participation. Trait Emotional Intelligence Questionnaire-Short Form (TEIQue-SF) The Trait Emotional Intelligence Questionnaire-Short Form (Cooper & Petrides, 2010) is a 30-item self-report questionnaire that measures TEI using a Likert scale ranging from 1 = ‘strongly disagree’ to 7 = ‘strongly agree’. Cronbach reliability for the total score in the current sample was 0.83. Career Adapt-Abilities Scale (CAAS) The Career Adapt-Abilities Scale (Savickas & Porfeli, 2012) includes 24 items equally divided into 4 subscales measuring resources of concern, control, curiosity, and confidence. Participants rate how strongly they have developed these resources on a Likert scale ranging from 1 = ‘I don't have this ability’, to 5 = ‘I have a very strong ability’. We employed the total score for the analysis as the literature shows that the 4 subscales load into a single second-order factor (Savickas & Porfeli, 2012). Cronbach reliability for the total score in the current sample was 0.91. Career Decision-making Difficulties Questionnaire (CDDQ) The Career Decision-making Difficulties Questionnaire (Gati et al., 1996) includes 34 items, but in this study only 32 were used due to a technical problem with the administration of items 27 and 29, which were part of the ‘inconsistent information’ subscale. The two missing items did not seem to significantly affect the subscale score and reliability. We employed the total score and the scores of the three subscales (lack of readiness, lack of information, inconsistent information) in the statistical analysis. The Likert scale ranges from 1 = ‘Does not describe me’ to 9 = ‘Describes me well’. Cronbach reliability in the current sample for the total score was 0.92, 0.61 for lack of readiness, 0.93 for lack of information, and 0.83 for inconsistent information. Self-Perceived Employability Scale (SPES) The Self-Perceived Employability Scale for university students (Rothwell, Jewell, & Hardie, 2009) is a 16-item scale used to evaluate expectations and self- perceptions of employability in university students. An example item is: “The skills and abilities that I possess are what employers are looking for”. The Likert scale ranges from 1 = ‘strongly disagree’ to 7 = ‘strongly agree’. Cronbach reliability in this sample for the total score was 0.87. Raven's Standard Progressive Matrices (RPM) Raven's Standard Progressive Matrices (Raven, 1938) were used to assess cognitive abilities, especially to evaluate fluid intelligence. The test is composed of five sets (A to E) of 12 multiple-choice, progressively more difficult items (60 total). The test administration was time-limited to 20 min. Cronbach reliability of the total score in the current sample was 0.88. Brief HEXACO Inventory (BHI) The Brief HEXACO Inventory is a 24-item questionnaire that assesses six personality dimensions: Honesty, emotionality, extraversion, agreeableness, conscientiousness, and openness (De Vries, 2013). The HEXACO model is an acknowledged measure of personality that proved to be a valid and reliable indicator of the major personality traits (e.g., Ashton & Lee, 2009). The brief version of the questionnaire showed adequate levels of test-retest reliability and high convergence with the longer HEXACO measure (De Vries, 2013). Participants respond to self-reflective items using a Likert scale ranging from 1 = strongly disagree to 5 = strongly agree. Alpha reliabilities of the dimensions in the brief version of the questionnaire range from 0.43 and 0.72 (De Vries, 2013). The Honesty scale was not included in the analytic model as we did not have any hypothesis regarding its relationship with TEI.","The data presented here were collected as part of a larger project on the investigation of emotional competencies, intelligence, and performance. Students from several French- speaking Swiss universities participated by filling out questionnaires both online and in a lab session. The scores of trait EI and personality were collected through an online survey filled out up to seven days before the lab session. The scores of career adaptability, fluid intelligence, career indecision and self-perceived employability were collected during the lab session. Only 308 participants completed this questionnaire online, as it was added at a second stage of the data collection. Statistical analysis To test the fit of the data, structural equation modeling for path analyses were conducted using Stata 14 (StataCorp, 2015) with maximum likelihood estimation; path coefficients of direct and indirect effects were estimated using Z-tests. Five separate path models testing full mediation were fitted: the first four with trait emotional intelligence as the predictor, career adaptability as the mediator, and career decision-making difficulties (total score and scores on the 3 subscales) as the outcome, and the fifth with the same predictor and mediator, but with the outcome of self-perceived employability. Each model was compared with a partial mediation model for comparative purposes. Sex, personality traits, and general intelligence were added as control variables in both models. Because the mediator was measured and not manipulated in our study, the mediator and the dependent variables might share some common causes; we therefore employed the instrumental variable method to account for potential endogeneity problems (Antonakis, Bendahan, Jacquart, & Lalive, 2010), allowing covariance between the standard error of the mediator and of the outcome. We requested standardized solutions, thus the reported values are beta coefficients. Overall R2 was also computed for each model. The following fit indices were used: chi-square test statistic (χ2), the comparative fit index (CFI), the Tucker-Lewis index (TLI), and the root mean square error of approximation (RMSEA). If the CFI value was 0.90 or above, the TLI values were above 0.95, the SRMR and the RMSEA values were 0.08 or less, the model was considered to have an acceptable fit (Cheung & Rensvold, 2002; Hu & Bentler, 1999).","The means, standard deviations, and correlations of the variables used are shown in Table 1. As scores on agreeableness were not correlated with emotional intelligence, we did not include it as a control variable in the final model. The Shapiro-Wilk normality test showed that the scores on CAAS, CDDQ (total score and lack of readiness), and SPES were normally distributed. Results regarding career indecision ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We tested a full mediation model in which career adaptability fully mediated the relationship between trait emotional intelligence and career decision-making difficulties, after controlling for the effect of sex, intelligence, and personality traits on TEI and career adaptability. Overall, the model showed satisfactory fit, χ2(6) = 12.29, p = .056; RMSEA = 0.058; SRMR = 0.020; CFI = 0.975; TLI = 0.938, R2 = 0.50. Because we observed that the effects of almost all control variables on career adaptability except conscientiousness were not significant, we created an alternative full mediation trimmed model (see Table 2) where only the path from conscientiousness to career adaptability was retained. A Chi-square test for model comparison indicated that the trimmed model had the same predictive power as the model including the regression paths from all control variables, ∆χ2(5) = 6.73, p = .241. We therefore retained the most parsimonious model with the regression path from conscientiousness to career adaptability only, for the rest of our analyses and for all tested models (Fig. 1). We tested a full mediation model in which career adaptability fully mediated the relationship between trait emotional intelligence and career decision-making difficulties, after controlling for the effect of sex, intelligence, and personality traits on TEI and career adaptability. Overall, the model showed satisfactory fit, χ2(6) = 12.29, p = .056; RMSEA = 0.058; SRMR = 0.020; CFI = 0.975; TLI = 0.938, R2 = 0.50. Because we observed that the effects of almost all control variables on career adaptability except conscientiousness were not significant, we created an alternative full mediation trimmed model (see Table 2) where only the path from conscientiousness to career adaptability was retained. A Chi-square test for model comparison indicated that the trimmed model had the same predictive power as the model including the regression paths from all control variables, ∆χ2(5) = 6.73, p = .241. We therefore retained the most parsimonious model with the regression path from conscientiousness to career adaptability only, for the rest of our analyses and for all tested models (Fig. 1). The comparison of the Chi-square test between the full mediation (trimmed) model and the partial mediation model (with a direct path from trait emotional intelligence and career decision-making difficulties) indicated no significant differences, ∆χ2(1) = 3.53, p = .060 (see Table 2), and supported a full mediation. This result highlights the importance of career adaptability in totally accounting for the relationship between TEI and career decision-making difficulties. TEI had a significant indirect effect on career decision-making difficulties through career adaptability, β = −0.39, Z = −8.19, p < .001. Results regarding career indecision's subscales We tested three different models in which career adaptability fully mediated the relationship between TEI and each of the CDDQ's subscales (lack of readiness, lack of information, and inconsistent information). Overall, the model with lack of readiness as dependent variable showed close to acceptable fit, χ2(11) = 28.66, p = .003; RMSEA = 0.072; SRMR = 0.032; CFI = 0.919; TLI = 0.890, R2 = 0.46. A Chi- square test for model comparison between the full mediation model and the partial mediation model (with a direct path from TEI to lack of readiness) indicated a significant difference, ∆χ2(1) = 5.49, p = .019—the partial mediation model had better predictive power and showed better fit indices. TEI had a direct effect on lack of readiness (β = −0.33, Z = −2.46, p = .014), but the indirect effect was not anymore significant. The model with lack of information as a dependent variable showed a good fit to the data, ∆χ2(11) = 15.19, p = .174; RMSEA = 0.035; SRMR = 0.021; CFI = 0.983; TLI = 0.977, with an R2 of 0.47. A Chi-square test for model comparison between the full and the partial mediation model (with a direct path from trait emotional intelligence to lack of information) showed no significant difference between the two, ∆χ2(1) = 3.54, p = .060, supporting a full mediation. TEI also had a significant indirect effect on lack of information through career adaptability (β = −0.38, Z = −8.03, p < .001). The model with inconsistent information as dependent variable showed a good fit to the data, χ2(11) = 14.74, p = .195; RMSEA = 0.033; SRMR = 0.022; CFI = 0.983; TLI = 0.977, with a R2 of 0.48. A Chi-square test for model comparison between the full and the partial mediation model (with a direct path from TEI to inconsistent information) showed no significant difference between models, ∆χ2(1) = 0.20, p = .658—the full mediation model has the same predictive power as the partial mediation model. TEI also had, a significant indirect effect on inconsistent information through career adaptability (β = −0.31, Z = −6.69, p < .001). Results regarding self-perceived employability ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Similar to the previous model, we tested a model in which career adaptability fully mediated the relationship between TEI and self-perceived employability, after controlling for the effect of conscientiousness on career adaptability. Overall, this model showed satisfactory fit, χ2(11) = 29.30, p = .002; RMSEA = 0.064; SRMR = 0.026; CFI = 0.945; TLI = 0.924, with an R2 of 0.44. A Chi-square test for model comparison between the full and the partial mediation model (with a direct path from trait emotional intelligence to self- perceived employability) showed no significant difference between the models, ∆χ2(1) = 0.14, p = .709. This result further highlights the importance of career adaptability in the relationship between TEI and self-perceived employability, confirming the full mediation effect of career adaptability (Fig. 2). A Chi-square test for model comparison between the full and the partial mediation model (with a direct path from trait emotional intelligence to self-perceived employability) showed no significant difference between the models, ∆χ2(1) = 0.14, p = .709. This result further highlights the importance of career adaptability in the relationship between TEI and self-perceived employability, confirming the full mediation effect of career adaptability (Fig. 2). TEI also had a significant indirect effect on self-perceived employability through career adaptability (β = 0.30, Z = 7.56, p < .001).","The aim of the current study was to examine the mediating role of career adaptability on the relationship between: a) TEI and career indecision, and b) TEI and self-perceived employability. We hypothesized a full mediation for both paths. Our hypotheses were supported in all cases except for the career indecision aspect of lack of readiness. Career adaptability acted as a mediator of the relationship between TEI and career indecision. Indeed, individuals scoring high on TEI obtained higher scores on career adaptability, which in turn had a positive impact on career indecision scores. These results suggest that the perceived capacity to understand and use emotions in different contexts may mobilize individuals' self-regulatory resources, which in turn reduce difficulties in making career decisions. Career adaptability was shown to be a mediator of the relationship between personality dispositions and career exploration and engagement (Li et al., 2015; Nilforooshan & Salimi, 2016). Our study indicates that career adaptability may also regulate the impact of more emotional dispositions, in particular trait emotional intelligence, on career decision-making difficulties. Indeed, when including career adaptability as a mediator, TEI has no anymore a direct effect on career indecision. This finding highlights the importance of considering the mechanisms through which a predictor may affect an outcome, and not simply testing the direct effect of the predictor. Concerning the three categories of career decision-making difficulties, career adaptability fully mediated the relationship between TEI and lack of information and inconsistent information, but TEI had only a direct positive effect on lack of readiness to enter the career decision-making process, the indirect effect through career adaptability being non-significant. Overall these results suggest that TEI has an important role in affecting perceived difficulties arising before making career-related decisions. Thus, individuals with high TEI would have fewer difficulties related to a lack of willingness to make a decision, to a general difficulty in making decision, or to dysfunctional beliefs about career decision-making process. Career adaptability seems to regulate the relationship between TEI and career decision-making difficulties that involve a lack of information and difficulties utilizing information due to inconsistency encountered during the decision-making process. TEI had once again no direct effect on both outcomes. Career adaptability also acted as a full mediator of the relationship between trait emotional intelligence and self-perceived employability. Indeed, individuals with high trait emotional intelligence through an activation of their career adaptability resources showed a better self-perceived employability, indicating that the perceived capacity to deal with emotions may activate the psychosocial self-regulatory resources, which in turn raise the confidence in finding a future employment. As for previous models, TEI had no direct effect on employability. The established relationship between conscientiousness and career adaptability was also confirmed by our study. Indeed, when the effect of conscientiousness on career adaptability was not controlled for, all the models showed poor fit. The meta-analysis conducted by Rudolph et al. (2017) showed that career adaptability and conscientiousness are strongly associated, which explains the better model fit once accounting for this link in the analysis. Limitations and future directions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Limitations of the current study include the use of a cross-sectional analysis to test the mediation hypotheses. It would be interesting to test them with a longitudinal design to see if the activation of career adaptability and its impact on career-related outcomes occurs across different time-lags. An additional limitation to consider is that this study used a student sample, and the results are not based on real-life career experiences. Previous research has shown that career adaptability influences job-related outcomes by affecting positive and negative emotional responses in the workplace (Fiori, Bollmann, & Rossier, 2015). The current findings add to our understanding of career adaptability by showing the dispositional affective antecedents that may mobilize the cognitive resources of career adaptability to support more effective career decision-making and more positive perceived employability. Ultimately, career adaptability emerges as a self-regulatory skill that shapes affective reactions (Fiori et al., 2015) and is also shaped by the perceived capacity to manage emotional experience. Based on the results of the current investigation, fruitful future directions could involve the development of interventions for career orientation that address potential deficits in both affective dispositions (e.g., trait emotional intelligence) and emotional states (e.g., negative affect) and cognitive self-regulatory resources (e.g., career adaptability). Important to note is that the affective and cognitive determinants of career-related outcomes may reciprocally influence each other, thus it would be advisable to intervene on both components at the same time. Because our results show that both conscientiousness and TEI affected career adaptability and career-related outcomes, future research might investigate potential interaction effects of these two personality dispositions. That trait emotional intelligence exerts positive effects on career indecision and employability is well known in the literature. Our contribution adds evidence for the path through which such effects may occur: individuals who are more confident in their capacity to handle emotional situations, including those generated by the search for an employment, mobilize more cognitive self-regulatory resources. These in turn positively affect individuals' self- perceived employability and lower their career indecision. Ultimately, this study helps provide a more comprehensive picture of the factors affecting career related outcomes. The following are the supplementary data related to this article. Final models with career indecision, lack of readiness, lack of information, and inconsistent information. Final model with self-perceived employability. Supplementary data to this article can be found online at https://doi.org/10.1016/j.paid.2018.06.046."],["Objectives: We examined the impact of judgement utility on the use of explicit contextual priors and visual information during action anticipation in soccer. Design: We employed a repeated measures design, in which expert soccer players had to perform a video-based anticipation task under various conditions. Methods: The task required the players to predict the direction (left or right) of an oncoming opponent's imminent actions. Performance and verbal reports of thoughts from players were compared across three conditions. In two of the conditions, contextual priors pertaining to the opponent's action tendencies (dribble = 70%; pass = 30%) were explicitly provided. In one of these experimental conditions, players were told that an incorrect ‘right’ response would result in conceding a goal, which created imbalanced judgement utility (left = high utility; right = low utility). In the third control condition, no explicit contextual priors or additional instructions were provided. Results: The explicit provision of contextual priors changed players’ processing priorities, biased their anticipatory judgements in accordance with the opponent's action tendencies, and enhanced anticipation performance. These effects were suppressed under conditions in which the explicit contextual priors were accompanied by imbalanced judgement utility. Under these conditions, the players were more concerned about the consequences of their judgements and were more inclined to opt for the direction with the higher utility. Conclusions: It appears that judgement utility disrupts the integration of contextual priors and visual information, which results in decreased impact of explicit contextual priors during action anticipation. --------------------------------------------------------------------------------","In dynamic and rapidly evolving environments, fast and accurate anticipation of an opponent’s actions underpins expert performance (Williams, 2000). Researchers have tried to elucidate the processes by which athletes combine prior knowledge and beliefs relevant to the specific performance context (i.e., contextual priors) with evolving visual information during anticipation (see Cañal-Bruland & Mann, 2015; Williams & Jackson, 2018). It has been proposed that athletes may employ Bayesian reliability-based strategies to weigh up, and integrate, contextual priors with relevant visual information: the reliance on contextual priors and visual information is modulated by the comparative reliability of the information at hand, where more reliance is assigned to information of relatively higher certainty with regard to the to-be-anticipated event (Gredin, Bishop, Broadbent, Tucker, & Williams, 2018). However, when examining anticipation in sport, a key, but often overlooked, component of the Bayesian framework is the comparative utility associated with possible judgements. Judgement utility refers to the potential costs and rewards associated with the consequences of one’s predictions: high judgement utility is associated with high rewards (if accurate) and low costs (if inaccurate), and vice versa for low judgement utility. According to Bayesian theory, people attempt to maximise the probability of their judgements being accurate while simultaneously maximise the expected utility of their judgments (Geisler & Diehl, 2003). In the current study, we examined the impact of judgement utility on the integration of explicit contextual priors and visual information during anticipation in soccer. It is well-established that expert athletes use advance visual information, such as opponent kinematics, to predict an opponent’s next move (e.g., Farrow, Abernethy, & Jackson, 2005; Loffing & Hagemann, 2014; Wright, Bishop, Jackson, & Abernethy, 2013). However, in recent years, researchers have taken a broader focus when examining the nature of anticipation in sport, by exploring the importance of non-visual information, such as contextual priors (Broadbent, Gredin, Rye, Williams, & Bishop, 2018; Gray & Cañal-Bruland, 2018; Gredin et al., 2018; Navia, Van der Kamp, & Ruiz, 2013; Runswick et al., 2018a). Due to advances in technology that enable sophisticated analyses of opponents, an increasingly prevalent component of elite sport preparation is to provide athletes with contextual priors pertaining to the behaviours of forthcoming opponents (Memmert, Lemmink, & Sampaio, 2017). Gredin et al. (2018) recorded gaze patterns and anticipatory judgements while soccer players predicted an opponent’s imminent actions with the ball in 2-versus-2 defensive soccer scenarios. The players performed the task both with and without explicitly provided contextual priors pertaining to the action tendencies of the attacker on the ball – specifically, whether he tended to dribble or pass the ball to a teammate. When given priors, the players allocated their visual attention toward the attacking teammate off the ball over the first half of each trial, as their determination of the teammate’s run trajectory enabled them to inform their judgements according to the action tendencies of the attacker on the ball. Over the second half of each trial, closer to the key point of action, the players relied more on evolving kinematic information of the attacker on the ball to inform their judgements. The provision of explicit priors biased the players’ judgements toward the most likely action, given the action tendencies, which resulted in enhanced task performance (see also Broadbent et al., 2018; Navia et al., 2013). Similarly, the beneficial performance effect of contextual priors was reported by Runswick et al. (2018a) using a task that required cricket batters to predict the location of forthcoming deliveries from bowlers. The batters exhibited superior anticipation when facing the same bowler six times in a row and when explicitly provided with information about the game state and field setting, compared to under conditions in which they faced six different bowlers and were not explicitly presented information about the game state and field setting. In the former condition, retrospective verbal reports demonstrated that the batters relied more on contextual priors to inform their judgements when compared to the latter condition. In the quest to develop an overarching framework that might explain the processes by which athletes combine contextual priors and visual information during anticipation, Loffing and Cañal- Bruland (2017) suggested that Bayesian theory may provide a suitable framework to elucidate these processes. Bayesian models for probabilistic inference assume that people base their judgements on probabilistic if-then relationships between known informational variables and unknown to-be-anticipated variables. That is, if ‘X’ (a known informational variable) occurs, then there is a certain probability that ‘Y’ (an unknown to-be- anticipated variable) will occur. This process suggests that, if one informational variable is associated with greater reliability (i.e., higher certainty) than another, then the individual’s joint estimate should be biased toward the ‘more reliable’ informational variable (Knill & Pouget, 2004). In line with the suggestion by Loffing and Cañal-Bruland (2017), Gredin et al. (2018) evaluated their findings using a Bayesian framework for probabilistic inference. In keeping with Bayesian theory, the authors proposed that players integrated explicit contextual priors and evolving visual information according to the comparative levels of reliability associated with the different sources of information (see Vilares & Körding, 2011). That is, the players were more dependent on priors over the first half of each trial where the reliability of kinematic information was low, and vice versa (see also Gray & Cañal-Bruland, 2018). A fundamental aspect of Bayesian theory is that our ultimate judgements are affected by both the reliability of available information and the potential costs and rewards associated with inaccurate and accurate judgements. Put in Bayesian terms, the weighted average of the reliability conveyed by prior and current sources of information is convolved with the utility values assigned to possible judgements (Geisler & Diehl, 2003). The biasing effect of judgement utility reflects not only to our desire to gain rewards and avoid costs, but also our tendency to use current information to confirm the viability of the high-utility option. As such, we tend to overestimate the likelihood of the outcome in accordance with its utility (see DeKay, Patiño-Echeverri, & Fischbeck, 2009; Russo & Yong, 2011). The biasing effect of judgement utility was, for example, shown in a study by Wallsten (1981) in which physicians judged it to be more likely that a patient had a malignant tumour than a cyst, despite the higher objective likelihood that the patient had a cyst. It was proposed that physicians overestimated the chances of a tumour due to its more severe consequences, relative to a cyst. In other words, a tumour diagnosis would yield greater rewards (if correct) and lower costs (if incorrect) than would diagnosing a cyst. In line with this finding, Canãl-Bruland, Filius, and Oudejans (2015) reported that skilled baseball batters tended to predict that fastballs would be pitched to a greater extent than change-ups. It was suggested that this strategy enabled them to handle the high speed of a fastball and, due to the slower nature of change-ups, to adapt their swing if confronted with this latter pitch type. If expecting a change-up, on the other hand, batters would not be able to catch up with the speed of a fastball. This finding suggests that judgement utility (e.g., comparatively higher utility of predicting a fastball than a change-up in baseball) may influence anticipation in sport. However, the manner in which judgement utility modulates the impact of contextual priors and processing priorities during anticipation has yet to be examined empirically. Greater awareness of the effects of judgement utility on athletes’ use of contextual priors will strengthen our understanding of action anticipation processes in sport – which will ultimately enhance coaches’ and performance analysts’ ability to maximise the effectiveness of explicit contextual priors (see Cañal-Bruland & Mann, 2015). In the current study, we examined the impact of judgement utility on the integration of explicit contextual priors and visual information as expert soccer players anticipated an oncoming opponent’s imminent actions. We employed the same task as used by Gredin et al. (2018): a video-based anticipation task simulating 2-versus-2 defensive soccer scenarios, in which the players had to predict the direction (left or right) of the attacker on the ball’s (termed ‘the opponent’ from hereon) action at the end of each sequence. In addition to conditions with and without explicitly provided contextual priors pertaining to the opponent’s action tendencies (dribble = 70%; pass = 30%), an additional condition was added, in which contextual priors were explicitly provided and judgement utility was explicitly manipulated. Judgement utility was manipulated by telling the players that an incorrect ‘right’ response would result in conceding a goal, which created imbalanced judgement utility (left = high utility; right = low utility). We recorded anticipation judgements and collected retrospective verbal reports of the thoughts the players engaged in when solving the task. We predicted that the explicit provision of contextual priors would bias anticipatory judgements toward the most likely action given the opponent’s action tendencies (i.e., dribble) and consequently enhance their anticipation performance (Broadbent et al., 2018; Gredin et al., 2018). Furthermore, we predicted that the explicit provision of contextual priors would result in players engaging in more thoughts related to the positioning of the attacker off the ball (see the Method section for detailed explanation) and the opponent’s action tendencies, whereas fewer thoughts would be related to the opponent’s kinematics (cf., Runswick et al., 2018a). However, in keeping with Bayesian theory (see Geisler & Diehl, 2003), we predicted that these effects would be supressed when the priors were accompanied with manipulated judgement utility, as players would be more inclined to opt for the direction associated with the higher utility value (i.e., left). In the condition where judgement utility was manipulated, it was hypothesised that the players would engage in more thoughts related to the costs and/or rewards that their responses could bring about. Based on Bayesian theories of weighted integration of information, we predicted that this would result in fewer thoughts related to the opponent’s action tendencies and relevant visual information.","A total of 18 (10 male and 8 female) expert soccer players (Mage = 23 years, SD = 3) participated. On average, the players had 14 years (SD = 2) of competitive experience in soccer at university or semi-professional level. At the time of the study, they participated in an average of 9 h (SD = 4) of soccer practice or match play per week. A spreadsheet for estimating sample size for magnitude-based inferences (Hopkins, 2006) was used to calculate the number of participants needed to find a clear performance effects (i.e., chances of the true effect to be substantially positive and negative < 5%; Batterham & Hopkins, 2006; see the Data Analysis section for further details). We used data from the study by Gredin et al. (2018) to calculate the minimum required sample size (n = 12; Hopkins, Marshall, Batterham, & Hanin, 2009). The study was approved by the Research Ethics Committee of the lead institution and conformed to the recommendations of the Declaration of Helsinki. All participants gave their written informed consent before taking part. Test stimuli ~~~~~~~~~~~~ The test stimuli comprised 30 video sequences of 2-versus-2 counter attacking scenarios in soccer, viewed from the perspective of one of the defenders. The test stimuli were filmed on an artificial turf soccer pitch using a high-definition digital video camera (Canon XF100, Tokyo, Japan) with a wide-angle converter lens (Canon WD-H72 0.8x, Tokyo, Japan). The camera was attached to a moving trolley, at a height of 1.7 m, to replicate the perspective of a central defender moving backwards. The final test footage was edited using Pinnacle Studio software (v15; Pinnacle, Ottawa, Canada). In total, 130 video simulations were created, but only clips that were independently selected by two qualified soccer coaches (UEFA A Licence holders), based on their representativeness of actual game play, were included in the final test footage of 30 clips. In each sequence, there was one attacking player in possession of the ball (‘the opponent’), a second attacker off the ball, and one marking defender who was following the second attacker throughout the sequence (see Figure 1). The videos were projected onto a 3.3 × 1.9 m projection screen using an Optoma HD20 DLP projector (Optoma, New Taipei City, Taiwan). The participant viewed the scenarios from a first-person perspective, as if they were the second defender. At the start of the sequence, the opponent was positioned approximately 7 m in front of the participant, 10 m to the right of the centre of the pitch, and 3 m inside the defensive half. The attacker off the ball and the second defender started approximately 3 m behind, either on the left, or the right, side of the opponent. When the sequence started to unfold, the players approached the participant and, after approximately 1.5 s, the attacker off the ball made a direction change towards either the left or the right. The sequences lasted for 5 s and, at the end of each sequence, the opponent could either pass the ball to his teammate who was positioned either to the left or right of the opponent (30% of trials) or dribble the ball in the opposite direction of the teammate (70% of trials). Task design ~~~~~~~~~~~ The participant was required to predict the direction (left or right) of the ball from the opponent’s final action. At the start of each trial, the participant was standing 3.5 m in front of the projection screen holding a response device in each hand − one for ‘left’ responses and one for ‘right’ responses. The participant was instructed that they should respond as soon as they were certain enough to carry out an action based on their prediction and that they could not change this response. Immediately after the participant’s response, the trial was occluded and feedback with regard to their response time and accuracy was displayed on-screen. Feedback for response time and accuracy was given as pilot testing suggested that receiving this information after each trial motivated the participant to stay engaged with the test task throughout the test session. A full 5-s trial was occluded 120 m s after the foot-ball contact of the opponent’s final action and if the participant responded after this point, then that trial was counted as incorrect. Since players in a soccer match are normally aware of the position of the ball and other players each trial started with a frozen frame for 1 s to allow the participant to detect this information. During the course of a trial, the participant was free to move as preferred in order to maximise the real-world representativeness of the task (cf. Gredin et al., 2018).","Before testing, the participant undertook 25–40 min of training in how to provide retrospective think-aloud reports. This training consisted of instructions on how to report thoughts retrospectively, including practice on a number of generic tasks. The participant was given feedback on their verbal reports, along with good and bad examples for these practice tasks (see Eccles, 2012). Throughout the training, the participant was encouraged to ask the researcher questions if they were unsure about how to articulate their reports. The aim of the verbal report training was to ensure that the participant learned how to verbalise only the thoughts that they used during the preceding task performance and to report them as they were naturally experienced during performance. This procedure is important to follow, as elaborating or providing commentary on the reported thoughts, or reporting thoughts that were not used during task performance may influence performance and violate the natural cognitive processes that occur during the task (Ericsson & Simon, 1980). Following this training, the participant was fitted with a lapel microphone and a body-pack transmitter that was wirelessly connected to a compact diversity receiver (ew112-p G3; Sennheiser, Wedemark, Germany) and a recording device (Zoom H5; Zoom Corporation, Tokyo, Japan), so that verbal reports could be recorded. Thereafter, the participant was given an overview of the experimental protocol and performed eight familiarisation trials to become accustomed to the experimental setup and response requirements. Verbal reports of thoughts were collected after four of the familiarisation trials. Following the familiarisation trials, the 30 test trials were presented in three blocks of ten trials, and under three different conditions (i.e., 90 test trials in total). In the control condition (Control), the participant performed the task without any additional information. In one of the experimental conditions (EXP), the participant performed the same task, but before each block, the participant was explicitly primed (verbally and on-screen) with contextual priors pertaining to the opponent’s action tendencies (i.e., dribble = 70%; pass = 30%). Due to the nature of the task, the positioning of the attacker off the ball revealed information that enabled the participant to use information about the opponent’s action tendencies (i.e., if the attacker off the ball was on the left, 70% of the opponent’s final actions were to the right, and vice versa). This meant that the participant had to incorporate the explicit contextual priors (i.e., the opponent’s action tendencies) with evolving visual information (i.e., the positioning of the attacker off the ball) in order to use the priors to inform their judgement. In complex and highly dynamic performance environments, such as in soccer, this interdependency between contextual priors and visual information may be an important component when seeking to elucidate the integration of different sources of information during anticipation (Gredin et al., 2018). In the other experimental condition (EXPJU), the same task was performed, and the same contextual priors were explicitly provided, but in this condition, the participant was instructed that if they were incorrect (i.e., responded ‘right’) when the opponent passed or dribbled the ball toward the left, their team would concede a goal. This instruction was given in order to increase the comparative utility associated with responding ‘left’. In other words, correct and incorrect ‘left’ responses came with greater rewards (stopping a goal) and costs (conceding a goal), respectively, than ‘right’ responses. This manipulation was based on the fact that in soccer, possession of the ball in a more central position near the penalty area (i.e., to the participant’s left in the present task) is more frequently associated with positive attacking outcomes than possession in a wider position (i.e., to the participant’s right; Brooks, Kerr, & Guttag, 2016). The participant received this instruction before each block, both verbally and on-screen, and was informed as to the number of goals they had conceded after each block. Each condition started with a condition-specific familiarisation trial, for which retrospective verbal reports were collected. To eliminate the influence of trial-specific characteristics, the same test trials were used in all three conditions. Thus, the distribution of trials where the opponent dribbled (70%) and passed (30%) the ball was identical across all three conditions. Furthermore, these actions were equally distributed across left and right outcome directions. The order in which the conditions were presented was randomised and counterbalanced across participants, which is a commonly used design in order to mitigate potential learning and carryover effects across various experimental conditions in sport (e.g., Gray, 2009; Jackson, Ashford, & Norsworthy, 2006; Runswick, Roca, Williams, Bezodis, & North, 2017). To further avoid any potential familiarity between conditions, the trial order in each condition was randomised (cf. Gredin et al., 2018). Response time and accuracy were recorded for each trial and verbal reports of thoughts were collected after six trials in each condition (cf. Runswick et al., 2018a). The selection of trials for which verbal reports were given was pseudorandomised in order to counterbalance the number of trials where the opponent dribbled (n = 3) and passed (n = 3) the ball, as well as the number of trials where the direction of the final actions was left (n = 3) and right (n = 3). The same trials were selected for all conditions and participants in order to avoid trial- specific characteristics from violating the verbal report data. The whole test session was completed within 90 min. Dependent measures ~~~~~~~~~~~~~~~~~~ Anticipatory judgements. To ascertain whether any systematic speed–accuracy trade-off effects were evident in the response data, Pearson correlations between response time and accuracy were determined for each condition. A positive correlation between the two variables was found in each of the three conditions (Control, r = 0.59 ± 0.28 [90% CI]; EXP, r = .66 ± 0.24; EXPJU, r = 0.68) ±0.23). To account for this speed-accuracy trade off, anticipation performance was expressed as an anticipation efficiency score, which was calculated by multiplying the average response time by the proportion of inaccurate responses in each condition (note: lower efficiency score indicates superior anticipation performance; cf. Gredin et al., 2018). In addition to the anticipation efficiency score, anticipation performance was expressed by the change in response accuracy between each condition, where the change in response time was used as a covariate in order to account for the covariation in the speed and accuracy measures (cf. Abernethy, Schorer, Jackson, & Hagemann, 2012). As we predicted that explicit contextual priors would bias the players’ judgements toward the most likely action, given the opponent’s action tendencies, we calculated the proportion of ‘dribble’ responses in each condition (note: a ‘dribble’ response corresponded to when the participant responded ‘right’ and the attacker off the ball was on the left side of the opponent or when the participant responded ‘left’ and the attacker off the ball was on the right side of the opponent). As we predicted that judgement utility would bias the players’ judgement toward the direction with the comparatively higher utility, the proportion of ‘left’ responses in each condition was assessed. Verbal reports. The verbal reports of thoughts were first transcribed verbatim, and the statements conveyed by each report were then coded into different categories (cf. Murphy et al., 2016; Roca, Ford, McRobert, & Williams, 2013; Runswick et al., 2018a). Statements were coded into three categories of visual information: positioning of the attacker off the ball, statements referring to the horizontal position (e.g., left, right, inside, outside) of the attacker off the ball relative to opponent; kinematic information, statements referring to the kinematic cues of the oncoming opponent; other visual information, statements referring to other kind of visual information not captured by the previous two categories. Furthermore, statements were coded into two categories of non- visual information: action tendencies, statements referring to the opponent’s tendency to pass or dribble the ball; judgement utility, statements referring to the costs and/or rewards their responses could bring about. Verbal report data from 3 participants, and 9 reports (<4%) from the remaining 15 participants, were excluded from the analyses due to participants failing to follow the procedure required for providing retrospective think- aloud reports (see Eccles, 2012). Once statements within each eligible report had been categorised, the proportion of reports in each condition that contained statements of each category was assessed (note: each report could contain references to multiple sources of information; cf. Murphy et al., 2016). Statistical analysis Descriptive statistics are reported as means and SDs. Magnitudes of observed effects along with their 90% CIs are reported as standardised (d) and unstandardised units. The effects were standardised by dividing the mean difference between conditions by the combined SD and then interpreted against the following scale: 0.2 > |d|, trivial; 0.2 ≤ |d| < 0.5, small; 0.5 ≤ |d| < 0.8, moderate; 0.8 < |d|, large (Cohen, 1988; Cumming, 2012). Cohen’s standardised unit for the smallest substantial effect (0.2) was used as a threshold value when estimating the uncertainty in true effects. The following scale was used to convert the quantitative chances to qualitative descriptors: 25–75%, possible; 75–95%, likely; 95–99.5%, very likely; > 99.5%, most likely (Hopkins, 2002). If the lower and upper bounds of the CI exceeded the thresholds for the smallest substantial negative and positive effect, respectively, meaning that there is ≥ 5% that the true effect could be substantially negative and ≥5% that it could be substantially positive, then the effect was deemed unclear. All other effects were deemed clear and evaluated as per the description above (Batterham & Hopkins, 2006). The confidence level of the CIs was not adjusted for multiple comparisons (Hopkins, Marchall, Batterham, & Hanin, 2009; Perneger, 1998; Rothman, 1990). Due to the low number of participants making references to the sources of non-visual information in their verbal reports, leading to high variability across participants and skewed data, inferential analyses were not conducted on the verbal reports of action tendencies and judgement utility. For these reports, descriptive statistics, including the number of participants referring to each information category, are presented. Anticipatory judgements ~~~~~~~~~~~~~~~~~~~~~~~ Table 1 shows the descriptive statistics for the response time and accuracy scores in each condition1,2. As shown in Figure 2a, the participants exhibited superior performance, manifested in a lower anticipation efficiency score, in EXP than in Control (d = 0.54 ± 0.35) and EXPJU (d = 0.43 ± 0.35), whereas no substantial effect was obtained when the anticipation efficiency score in Control was compared to that in EXPJU (d = 0.14 ± 0.34). The analysis of the adjusted effects for response accuracy between conditions revealed that response accuracy was higher in EXP than in Control (d = 0.35 ± 0.29) and EXPJU (d = 0.25 ± 0.30), whereas no substantial effect was found between Control and EXPJU (d = 0.12 ± 0.27; note: changes in response time were included as a covariate in these analyses). The proportion of trials in which the participants predicted that the opponent would dribble was higher in EXP than in Control (d = 0.66 ± 0.47) and EXPJU (d = 0.89 ± 0.46), whereas no clear effect was found when the proportion of ‘dribble’ responses in Control was compared to that in EXPJU (see Figure 2a). Figure 2c shows that the proportion of trials where the participants predicted that the opponent would dribble or pass the ball toward the left was higher in Control than in EXP (d = 0.33 ± 0.42) and higher in EXPJU, both compared to Control (d = 0.65 ± 0.44) and compared to EXP (d = 1.03 ± 0.47). Unstandardised effects for anticipation efficiency, percentage ‘dribble’ responses, and percentage ‘left’ responses across conditions are presented in Table 2. Verbal reports ~~~~~~~~~~~~~~ The percentage of verbal reports referring to each category of visual information in each condition is presented in Figure 3. The participants reported a higher proportion of statements referring to the positioning of the attacker off the ball in EXP than in Control (d = 0.37 ± 0.34) and EXPJU (d = 0.22 ± 0.39), whereas no substantial difference was found between Control and EXPJU (d = 0.12 ± 0.28). The percentage of statements referring to the kinematic information of the opponent was higher in Control than in EXP (d = 0.59 ± 0.27) and EXPJU (d = 1.03 ± 0.43), and higher in EXP than in EXPJU (d = 0.47 ± 0.27). Regarding other sources of visual information, the participants reported a higher proportion of statements within this category in Control than in EXP (d = 0.32 ± 0.37) and EXPJU (d = 0.40 ± 0.27), while no clear difference was obtained between EXP and EXPJU. Unstandardised effects for the percentage of verbal reports referring to each category of visual information across conditions are presented in Table 3. In reports containing statements referring to non-visual information, the proportion containing statements relating to the opponent’s action tendencies were 1.1% (SD = 4.3) in Control, 32.2% (SD = 28.5) in EXP, and 22.2% (SD = 29.5) in EXPJU. Only one participant mentioned this information source in Control, whereas eleven and nine participants referred to the opponent’s action tendencies in EXP and EXPJU, respectively. Statements relating to judgement utility were reported in 34.2% (SD = 32.5) of the reports in EXPJU, while the corresponding proportions in Control and EXP were 14.2% (SD = 23.3) and 3.3% (SD = 12.9), respectively. Statements referring to this category was reported by five participants in Control, one participant in EXP, and twelve participants in EXPJU.","We examined the impact of judgement utility on the integration of explicit contextual priors and visual information as expert soccer players predicted the direction (left or right) of an oncoming opponent’s imminent actions. The players performed the anticipation task under three different conditions: one condition without any explicit contextual priors; another condition in which contextual priors pertaining to the opponent’s action tendencies (dribble = 70%; pass = 30%) were explicitly provided; and a third condition in which these contextual priors were explicitly provided, and judgement utility was explicitly manipulated (left judgements = high utility; right judgements = low utility). We recorded anticipatory judgements and collected retrospective think-aloud reports in all three conditions. In line with our predictions, the explicit provision of contextual priors biased the players’ anticipatory judgements toward the most likely action, given the opponent’s action tendencies. This effect was manifested in the fact that the players predicted the opponent would dribble the ball to a greater extent in the condition with explicit contextual priors, compared to the condition in which no priors were provided. This finding supports previous research suggesting that expert soccer players use explicit contextual priors to inform their judgements when predicting an oncoming opponent’s next move (Broadbent et al., 2018; Gredin et al., 2018). As we predicted, this biasing effect of explicit contextual priors resulted in enhanced performance, which was evidenced by superior anticipation efficiency in the condition with, compared to without, explicit priors. These findings support the growing body of research demonstrating the performance- enhancing effect of contextual priors on anticipation in sport (Broadbent et al., 2018; Gredin et al., 2018; Navia et al., 2013; Runswick et al., 2018a). An important objective of this study was to examine the impact of judgement utility on the players’ anticipation. As we hypothesised, judgement utility supressed the impact of explicit contextual priors: the proportion of ‘dribble’ responses did not increase when priors were accompanied by the judgement utility manipulation. In line with our predictions, this resulted in anticipation performance comparable to that in the control condition. As prescribed by Bayesian models for informational integration, this finding suggests that expert soccer players base their judgements both on the reliability of the information at hand and the costs and rewards their responses could generate (see Geisler & Diehl, 2003). In other words, it appears that judgement utility reduced the impact of contextual priors, as the players were more inclined to opt for the direction with the comparatively higher utility. Further support for the impact of judgement utility on anticipatory judgements is apparent in the comparison made between the proportion of ‘left’ responses across conditions. As predicted, the proportion of responses where the players opted for a leftward outcome was higher in the condition where judgement utility was manipulated, compared to the other two conditions. Such a biasing effect of judgement utility has been demonstrated across various domains (see Canãl-Bruland et al., 2015; DeKay et al., 2009; Russo & Yong, 2011). It is noteworthy though, that the players responded ‘left’ more often than ‘right’ in all three conditions. It may be the case that the players associated the left side with greater goal threat, due to the position on the pitch (Brooks et al., 2016) even though this information was not explicitly provided, and hence there were more ‘left’ responses in all three conditions. This suggestion aligns with the findings reported by Canãl- Bruland et al. (2015) demonstrating that baseball batters predicted fastballs to a greater extent than change-ups, even if they were not explicitly instructed that predicting fastballs came with greater utility. Interestingly, in the two conditions without manipulated judgement utility, the players were more inclined to respond ‘left’ in the condition without explicit contextual priors. In keeping with Bayesian theory (see Geisler & Diehl, 2003; Vilares & Körding, 2011), the perceived imbalance in threat between left and right outcomes may have been weighted higher in the former condition where less reliance was placed on the opponent’s action tendencies. To explore the processing priorities employed by the players during task performance, we collected immediate retrospective verbal reports of their thought processes. The data provide tentative support for our prediction that the players would refer to the opponent’s action tendencies to a greater extent when this information was explicitly provided, relative to when it was not. Furthermore, when contextual priors were explicitly provided, the players reported a higher proportion of thoughts relating to the attacker off the ball, compared to the latter condition. This finding aligns with the gaze data in Gredin and colleagues' (2018) study, in which expert soccer players shifted their overt visual attention away from the opponent, and toward the attacker off the ball, when explicit contextual priors were provided. In the case of that study, and the current one, it is proposed that the players prioritised the positioning of the attacker off the ball more so when contextual priors were explicitly provided compared to when they were not. This information enabled the players to inform their anticipatory judgements in accordance with the opponent’s action tendencies (i.e., when the attacker off the ball was on the left, 70% of the opponent’s final actions were to the right, and vice versa). The verbal report data in the current study also showed that the proportion of thoughts relating to the opponent’s kinematic information was lower in the condition in which contextual priors were explicitly provided, relative to in the condition where they were not. This finding mirrors that of Runswick et al. (2018a), who demonstrated that an increased reliance on contextual priors came with a decreased reliance on kinematic information, when cricket batters predicted the location of forthcoming deliveries from a bowler. The previous research, and the current study, provide support for the Bayesian notion that, when the reliability of one informational variable increases (e.g., increased reliability of contextual priors via explicit guidance or sequential pickup), people’s judgements become less contingent upon other informational variables (e.g., kinematic information; see Vilares & Körding, 2011). In the condition where judgement utility was manipulated, the players engaged less in thoughts relating to the positioning of the attacker off the ball and fewer players referred to the opponent’s action tendencies, compared to when explicit contextual priors were provided, but judgement utility was not manipulated. Furthermore, in the former condition, the players engaged less in thoughts relating to the opponent’s kinematic information, and more players referred to the costs and/or rewards their responses could bring about, relative to in the other two conditions. In line with our predictions, these data provide tentative support that the biasing effect of judgement utility on anticipation was underpinned by changes in the thought processes that the players employed during task performance; namely, an increased concern about the costs and/or rewards their responses could bring about (see also DeKay et al., 2009; Russo & Yong, 2011). The current study lends support to the idea that Bayesian theory may provide a suitable framework to elucidate the processes by which athletes inform their judgments during action anticipation (Loffing & Cañal-Bruland, 2017). However, it is important to note that the experimental design in this study did not allow us to test the predictions prescribed by Bayesian theory in relation to the impact of contextual priors, visual information and judgement utility in a fine-grained quantitative manner. Thus, we encourage researchers to further explore the merits of Bayesian theory in this regard; for example, by adopting temporal occlusion paradigms to more tightly control the availability of visual information (e.g., Runswick, Roca, Williams, McRobert, & North, 2018b), by altering the reliability of priors in a more fine-grained manner (e.g., Gray & Cañal- Bruland, 2018), and/or by using continuous outcome possibilities, rather than a binary- choice task (e.g., Tassinari, Hudson, & Landy, 2006). Our findings may prove highly informative to coaches and performance analysts, who typically use contextual priors to guide their players’ on-pitch decision making. In situations where inaccurate and accurate judgements incur varying costs and rewards, it seems like judgement utility disrupts players’ processing priorities and decreases the effectiveness of explicit contextual priors. However, it is worth noting that soccer matches are inherently complex, comprising multiple crossmodal sources of environmental information, priors that relate to multiple players and/or situations, and a multitude of potential outcomes. Thus, it is possible that the task in the current study did not evoke behaviours that we might observe under natural performance conditions (see Cañal-Bruland, Müller, Lach, & Spence, 2018; Müller & Abernethy, 2012; Vaeyens, Lenoir, Williams, Mazyn, & Philippaerts, 2007). Clearly, the impact of judgement utility should be further explored, using more naturalistic test tasks. Another potential limitation of this study was that we did not include a condition in which we manipulated judgement utility and provided no explicit contextual priors. Adding such a condition would allow us to assess the moderating effects of judgement utility with greater certainty. However, pilot work suggested that the effects that would have been obtained in such a condition would not have been substantially different to those obtained in the EXPJU. Also, it is worth highlighting that it is possible that the players in the current study used information acquired in preceding conditions to inform their task performance in subsequent conditions. Thus, in order to mitigate potential confounding order effects, we employed a randomised and counterbalanced within-participant design, which is a well-established approach when comparing the effects of various informational conditions in sport (cf. Gray, 2009; Jackson et al., 2006; Runswick et al., 2017). To further avoid any potential familiarity effects across conditions, the trial order in each condition was randomised (cf. Gredin et al., 2018). Furthermore, it is possible that the feedback for response time and accuracy provided after each trial created another source of information that could have facilitated trial-to-trial learning. However, self-reported pilot data suggest that, without this immediate feedback about their performance, the players’ motivation to engage with the task would decline over the course of testing. In the study by Gredin et al. (2018), in which the players performed a similar task and received feedback for response time and accuracy after each trial, the authors did not find any substantial performance effect when the initial 24 trials of a conditions were compared to the final 24 trials of that condition. This finding suggests that task experience, including feedback on performance after each trial, did not change the players’ performance on the task. A related issue is the potential risk that the retrospective verbal reports of thoughts may have disrupted the natural thought processes in which the players engaged during task performance and, as such, influenced their behaviours on the task. In order to mitigate this risk, the players undertook 25–40 min of training on how to only report heeded thoughts in a way that they were naturally experienced during the task, rather than trying to explain, justify, or qualify their thoughts or behaviours during the task (see Eccles, 2012; Ericsson & Simon, 1980). In summary, our findings suggest that the explicit provision of contextual priors biased expert soccer players’ processing priorities during action anticipation: greater reliance was placed on the priors and context-relevant visual information, while less reliance was placed on evolving kinematic information. This, in turn, biased anticipatory judgements toward the most likely outcome, given the contextual priors, and enhanced anticipation performance. However, the biasing impact of explicit contextual priors, in regard to both anticipation and associated thought processes, was supressed when the comparative utilities associated with potential judgements differed. Under these conditions, the players became less reliant on the contextual priors and unfolding visual information, and more inclined to opt for the outcome associated with the highest rewards and the lowest costs."],["Background: Government policy and national practice guidelines have created an increasing need for autism services to adopt an evidence-based practice approach. However, a gap continues to exist between research evidence and its application. This study investigated the difference between autism researchers and practitioners in their methods of acquiring knowledge. Methods: In a questionnaire study, 261 practitioners and 422 researchers reported on the methods they use and perceive to be beneficial for increasing research access and knowledge. They also reported on their level of engagement with members of the other professional community. Results: Researchers and practitioners reported different methods used to access information. Each group, however, had similar overall priorities regarding access to research information. While researchers endorsed the use of academic journals significantly more often than practitioners, both groups included academic journals in their top three choices. The groups differed in the levels of engagement they reported; researchers indicated they were more engaged with practitioners than vice versa. Conclusions: Comparison of researcher and practitioner preferences led to several recommendations to improve knowledge sharing and translation, including enhancing access to original research publications, facilitating informal networking opportunities and the development of proposals for the inclusion of practitioners throughout the research process. --------------------------------------------------------------------------------","Government policy and national practice guidelines have highlighted an increasing need for professionals working in autism services to adopt an evidence-based approach in the delivery of diagnostic methods and clinical and educational interventions. However, a gap continues to exist between research knowledge and its application in practice (Parsons et al., 2013; Reichow, Volkmar, & Cicchetti, 2008). One factor that may contribute to this gap is a difference between academic and non-academic professional groups in their approach to acquiring knowledge of autism research. Practitioners’ views about what counts as a credible knowledge source is historically influenced by their training and experience (Rycroft-Malone et al., 2004). They may, therefore, routinely use different methods from those used by researchers when updating their specialist professional knowledge and have different views about how they could potentially benefit from research evidence in the future. Greater understanding of these perspectives is therefore important for researchers who are aiming to adapt scientific evidence to meet the needs of the wider, non-academic community (Lemay & Sá, 2014). Effective knowledge translation into practice depends on effective facilitation by researchers (Kitson, Harvey, & McKormack, 1998; Rycroft-Malone et al., 2004), a goal that has been heightened by government policy in recent years by the impact assessment of academic research (e.g. Research Excellence Framework, 2014; http://www.ref.ac.uk/). It has been argued that attempts to bridge the research-practice gap need to involve greater collaboration between autism researchers and research-users, such that both communities are engaged in the research process from the beginning (Parsons et al., 2013). Such collaborative activity, or engagement, can facilitate co-participation in the development of research design and method through reciprocal exchange of knowledge. Engagement between researchers and the people who use research is a central component of interactive models of knowledge translation in health policy (Jacobson, Butterill, & Goering, 2003) and a key facilitator of effective knowledge translation (Huberman, 1990). It enables the researcher to orient towards the needs of the user group, provides opportunities for discussion about the values and interpretation of evidence, and helps to facilitate trust and collaboration between researchers and research-users (Milton, 2014; Parsons et al., 2013). Recent research studies in the field of autism have highlighted the importance of developing a research agenda that is oriented towards the research-user. This work has identified topic areas that research-users prioritize as important areas for future research (Pellicano, Dinsmore, & Charman, 2013; Pellicano, Dinsmore, & Charman, 2014a). Results showed that although researchers and research-users agreed on some of the priorities for future research in autism, there was also a mismatch in priorities for other areas. This research also reported a mismatch in the level of engagement reported by researchers and research-users in the field of autism (Pellicano et al., 2013; Pellicano, Dinsmore, & Charman, 2014b). While academic researchers perceive themselves to be engaging with non-researchers, the same view is not held by non-research users of research. The findings above emphasize that research should focus on priority areas that meet the needs of the research-user community, a goal more likely to be fulfilled by improved engagement between researchers and non-researchers. The focus of the current study was not on priorities for what should be researched, as previously studied, but on the process of knowledge acquisition. To investigate this, similarities and differences in the methods and preferences for acquiring knowledge used by researchers and research-users were examined, targeting individuals from one sector of the autism research-user community: professional practitioners working in clinical, educational and policy settings. In this respect our definition of ‘research-user’ is consistent with the definition used by Lemay and Sá (2014) in a study of the translation of research evidence into professional practice. If methods of knowledge acquisition differ, and researchers present evidence in ways that are incompatible with the preferences of practitioners, communication will be inhibited and the translation of research evidence impeded. Therefore, greater understanding of the gap between research and practitioner groups in both their current practice and future preferences will facilitate conditions for collaborative engagement that starts from a more common ground and enable reciprocal exchange. The current questionnaire study formed part of the development work for an online web-based initiative being designed for the purpose of connecting research, practice, and policy communities. In a set of three questions, researchers and professionals working in practice communities were asked about how they use research information. First, both groups were asked how they currently keep up to date with information in the field of autism. Second, both groups were asked which methods they thought would be beneficial to increase research access and research knowledge for research-users. This question necessarily focuses on the translation of information from researchers to research-users, to help explore issues about the inadequacy of the communication of research evidence to non-research professionals. Finally, researchers and practitioners were asked to report on their current level of engagement with the other community. This question was intended as a general measure of the extent to which each group felt that they had any form of active involvement with the other community and was intentionally non-specific, to capture any form of interaction or engagement. It was expected that the two groups would use different sources of information to keep up to date with professional knowledge, with researchers relying more on primary evidence sources and practitioner research-users relying on alternative evidence sources. However, it was not known if views would differ about the best methods for increasing research knowledge in the non-academic community. Better understanding of potential differences between these groups may identify mechanisms by which knowledge sharing could be improved. Finally, based on previous findings, it was expected that practitioners would report lower levels of engagement with researchers than vice versa.","A total number of 683 respondents answered at least one of the three questions. Two hundred and sixty one respondents described themselves as working in the area of practice or policy. When asked to specify their occupation, of these 261, the two largest occupational groups represented were professionals working in schools (teachers or teaching/learning assistants; n = 55), and psychologists, including educational, clinical and occupational psychologists (n = 47). Other occupational groups included speech and language therapists (n = 21), psychiatrists (n = 5), nurses (n = 3), support workers (n = 16) and social workers (n = 6). Nine respondents stated that their occupation included both “policy work” and practice roles, with a further two indicating that their primary occupation was “policy work”. Eighty-two respondents selected “Other” and a further 26 did not respond to this question. The title of ‘practitioner’ group was therefore assigned, to include a range of professionals working in clinical, educational, and policy-based practice. In the research professional group (n = 422), the majority described themselves as either established career researchers (including completed a PhD, working in research; n = 256) or early career researchers (including post-graduate, research assistant or associate, or PhD student; n = 143). A minority (n = 17) selected “Other”, which included retired or freelance/independent researcher, and six gave no response. Researchers were also asked in which country they worked.1 There were 147 researcher responses from North America, 137 from Australia/New Zealand, and a further 108 from elsewhere in the UK. The remaining respondents (n = 30) came from a range of countries including Serbia, Japan, South Korea, Argentina, South Africa, Columbia and Singapore. Design and procedure ~~~~~~~~~~~~~~~~~~~~ The survey questions were first designed for practitioners and trialled on a first ‘wave’ set of 54 local professionals from clinical, educational and policy sectors recruited through professional contacts in two neighbouring regions of Wales (UK), in order to identify potential problems. Concurrent with this initial trial of the survey, a separate sample of eight local autism professionals from teaching, occupational therapy, speech and language therapy, care work and educational psychology professions were recruited through professional contacts to take part in 40 minute telephone/face-to-face interviews using a semi-structured, open-ended interview format. The interviews focused on the topics of evidence-based practice, research colleagues, and collaboration. Responses from these interviews revealed broad themes that were consistent with the online survey questions and response options. The survey was then distributed to the researcher group in addition to the geographically wider practitioner community without alterations. The survey was written and distributed only in English. The data for the main survey were collected via Google Survey over a period of three weeks in August 2013. The online survey was also sent to UK professionals and two international email lists made up of both practitioners and researchers who attended face-to-face or online professional conferences. Consistent with previous online surveys of this type (e.g. Pellicano et al., 2014a, 2014b), a snowball method of sampling was used. The link was shared extensively via national and international email contact lists and social media, and recipients were encouraged to forward it to colleagues, resulting in a ‘cascade’ distribution. Emails were sent out three times inviting people to respond. The survey was shared on Twitter and Facebook weekly. The written introduction to the survey explained that the questions formed part of the planning and design stage for a new knowledge hub initiative aiming to improve connections between autism researchers and non-academic professionals, and that the survey aimed to compare the needs and views of different professional groups. Ethical approval was obtained from the University's School of Psychology Ethics Committee. All participants gave informed consent on the first online page of the questionnaire prior to participating. Participants were first asked to indicate if they were a researcher or a professional working in practice or policy. Both groups were asked about the methods they use to access information: “At the moment, how do you keep up to date with current information in the area of autism?” Respondents were asked to select their top three options out of a possible ten (see Fig. 1 for summarized response options). In a separate question, respondents were asked to identify methods that they perceived would be beneficial for increasing research access and research knowledge (see Fig. 2 for response options). Non-academic professionals were asked “Of the options below, which would help you benefit most from research knowledge and evidence in your specialist area of autism?” Respondents were asked to select their top three options out of a possible nine. The researcher version of the questionnaire asked a similar question, which aimed to elicit researchers’ views of what would help most to promote engagement, research awareness and knowledge translation in practice and policy communities: “The [online initiative] aims to engage researchers with non-academic professionals, promote awareness of research and create opportunities to translate research knowledge. Which options below do you think would help achieve this aim? Tick your top three”. The identical response options for both groups are shown in Fig. 2. Both questions included an ‘other’ option. Finally, both groups were also asked about their level of engagement with research or non-research professionals: either “As a practitioner or working in policy in autism, do you currently engage with researchers?” or “As a researcher in autism, do you currently engage with practitioners or those working in policy?” To reduce the potential carry-over effect of responses, the questionnaire design presented the following fixed order: (1) methods for improving research knowledge question, (2) methods for updating of current information question, (3) engagement with the other professional group question. Updating of current information ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Each participant was asked to select their three preferred methods of keeping up to date with current information. Six participants, all researchers, did not endorse any items and were excluded. Of the remaining participants, 233 (89%) practitioners and 336 (81%) researchers selected three options. The distribution of choices for the sample selecting three options is shown in Fig. 1. Two-by-two Chi square analyses compared group differences for each option, with alpha level set at .005 following Bonferroni correction for multiple comparisons. The following information options were selected significantly more frequently by practitioners than researchers: campaigns: χ2(1) = 28.52, p < .001 (20% practitioners, 5% researchers); non-academic journals: χ2(1) = 15.10, p < .001 (23% practitioners, 11% researchers); conferences and Continuing Professional Development (CPD): χ2(1) = 14.26, p < .001 (64% practitioners, 48% researchers); newspapers, online news, TV and radio: χ2(1) = 7.93, p = .005 (29% practitioners, 19% researchers). In contrast, the following options were selected more frequently by researchers than practitioners; colleagues: χ2(1) = 20.94, p < .001 (74% researchers, 55% practitioners); academic journals: χ2(1) = 142.77, p < .001 (91% researchers, 45% practitioners). The top three options for each group overlapped, although the order of these differed. The most frequently endorsed option for practitioners was conferences/CPD courses (64%) followed by colleagues (55%) and academic journals (45%), while the most frequently endorsed option for researchers was academic journals (91%), followed by colleagues (74%) and conferences (48%). No group differences (p > .005) were found for Google searches (31% practitioners, 25% researchers), membership of voluntary organizations (19% practitioners, 11% researchers), or social media (8% practitioners, 9% researchers) and there was no difference in the “other” category (7% practitioners, 6% researchers). Analyses were re- run including participants who failed to endorse all three choices. Results were identical with the exception of one additional group difference for voluntary organizations, which was favoured by practitioners (χ2(1) = 9.28, p < .002; 18% practitioners, 10% researchers). Increasing research knowledge ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Each participant was asked to select the three methods that they considered most beneficial for increasing research access and knowledge (their top three). The question to practitioners asked which methods would most benefit their own research knowledge and evidence in their specialist area of autism (top three choices), while the question to researchers asked which would help promote research awareness and opportunities for knowledge translation in non-academic professionals. Eleven respondents (all of whom were researchers) did not endorse any items and were excluded. Of the remaining participants 228 practitioners (87%) and 389 researchers (95%) endorsed three choices. Fig. 2 compares the choices of the respondents who selected three choices. As above, 2×2 Chi square analyses were computed to compare groups for each option, with alpha level set at .006 using Bonferroni adjustment for multiple comparisons. The following options were selected significantly more frequently by practitioners than researchers: connect directly to research articles to read original research: χ2(1) = 29.09, p < .001 (58% practitioners, 36% researchers); access to practice based articles that have been based on reliable research: χ2(1) = 11.47, p < .001 (61% practitioners, 47% researchers). In contrast, the following options were selected more frequently by researchers: speak to a researcher to ask questions about specific research findings: χ2(1) = 27.59, p < .001 (35% researchers, 15% practitioners); access a large directory of researchers to find out about research and/or develop opportunities for collaborating: χ2(1) = 8.72, p < .005 (45% researchers, 33% practitioners). The most frequently endorsed option for practitioners was access to practice-based articles based on research (61%), followed by connect directly to research articles (58%) and then learn to apply evidence-based research methods (45%). For researchers the most frequently endorsed option was non-technical one-page lay summaries (50%), followed by access to practice-based articles based on research (47%) and access a large directory of researchers (45%). No group differences were found for researchers’ blogs (16% practitioners, 23% researchers), non-technical one-page summaries (43% practitioners, 50% researchers), Twitter or news updates (26% practitioners, 23% researchers) or apply evidence-based research methods to use in practice and policy (45% practitioners, 40% researchers). Analyses were re-run to include all those participants who did not endorse the full three choices; the results were identical. Engagement with the other professional group ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Eight practitioners and one researcher did not respond to this question. Table 1 shows the current level of inter-group engagement reported by the remaining practitioners (n = 253) and researchers (n = 421). Researchers reported significantly higher levels of engagement with practitioners than practitioners reported for their level of engagement with researchers (χ2(3) = 87.03 p < 0.001); the majority of researchers (56.3%) indicated that they had either “a lot” or “quite a bit” of engagement with practitioners, while the majority of practitioners (79.1%) reported that they only “occasionally” had engagement with researchers, if at all.","This study presents the first evidence of similarities and differences between researchers and practitioners working in autism in their preferred methods for acquiring up-to-date information and gaining research knowledge. The focus on a professional practitioner research-user group in the current study specifically aimed to facilitate greater understanding of the incompatibility in preferred methods used by researchers and professional practitioners. By identifying gaps in the way that researchers communicate research evidence it may be possibly to highlight potential targets for evidence-based approaches in clinical, educational and policy-based practice. Multi-dimensional approaches to knowledge translation (Kitson et al., 1998; Rycroft-Malone et al., 2004) have emphasized the importance of the nature and accessibility of the research evidence to be translated. In the current study we examined the evidence sources that practitioners prefer to use when updating their knowledge and the methods that they consider would help them benefit from research knowledge and evidence. Previous work on knowledge translation in areas of health and education indicates that practitioners can be resistant to learning about and accepting research evidence (Parsons et al., 2013; Russell, Greenhalgh, Boynton, & Rigby, 2004; Rycroft-Malone et al., 2004), and indicate preferences for accessible, practitioner-relevant and non-technical sources of evidence information (Graham, Tetroe, & Gagnon, 2013). Specific work in the area of autism has also emphasized practitioners’ priorities for accessing practitioner-oriented methods in the area of intervention (Reichow et al., 2008). When both researchers and practitioners were asked about how they keep up to date with current information, their responses showed that the preferred current sources of information for practitioners were conferences and continuing professional development (CPD). There was also evidence that practitioners relied more than researchers on publicly accessible sources of knowledge such as news media and campaigns, as well as on non-academic journals. In general, however, more accessible methods of updating current information, such as social media, Google searches and news/TV, were of relatively low priority for practitioners. Researchers differed from practitioners in their methods for keeping up to date with current information due to their highly frequent use of academic journals compared with practitioners. However they were similar in their low priority for social media and other accessible evidence sources. When researchers and practitioners considered what would be beneficial for increasing their research knowledge, the most notable finding was the high value given by practitioners to original articles from academic journal articles. Another notable difference was that researchers gave higher scores than practitioners to the options relating to direct contact with researchers (access to researcher directory to find out about research or develop opportunities for collaborating and speak to a researcher). Finally, both groups were in agreement about their priorities for learning to apply evidence-based research methods and accessing non-technical research summaries, both of which were highly rated. Again, more accessible methods such as social media were of relatively low priority for both groups and showed no difference between the groups. Finally, when asked about level of engagement with the other group, the results for the current study supported Pellicano et al.’s (2014b) earlier finding. We found that practitioners reported relatively lower levels of engagement in comparison to the level of engagement reported by researchers; that is, researchers identified themselves as more highly engaged with the practitioner community than vice versa. One potential explanation for this discrepancy could be the perceived impact of engagement in the two fields; while researchers may more directly perceive the benefits of involving the practitioner community in designing and conducting research, the outcome of research and the implications for practice may be felt only over a longer time-frame. Practitioners may, therefore, feel less engaged by the research community as the impacts of any involvement may be less immediately tangible. Alternatively, this discrepancy could reflect differences in the perceived value of research evidence and practitioner expertise; if practitioners felt that research evidence was valued more highly than evidence coming from their own knowledge and expertise, they may be less likely to engage with the research community. The questionnaire in the current study did not allow us to explore these issues, but these findings highlight that an important area for further investigation will be to clarify differences in the definition and value of what counts as evidence and to facilitate more reciprocal recognition of both research and practitioner expertise. Taking the results from all three questions together, the findings show a pattern in which researchers and practitioners not only have mismatched perceptions of engagement with each other, but also have different views about how to best facilitate the communication and sharing of research knowledge. For practitioners, direct access to original research articles in academic journals is prioritized very highly as a beneficial method for increasing research knowledge, while access to researchers themselves (speaking to a researcher or accessing a directory of researchers) is prioritized more highly by researchers. There are also important similarities between the two groups; for example, both assign relatively low ratings to social media as a mechanism for accessing and gaining research knowledge. Moreover, the top three options for updating current information in both groups were conferences or CPD, academic journals, and colleagues. Although the order of these options differed in the two groups, this overlap suggests that practitioners and researchers use the same mechanisms to update their current research knowledge, with a desire by the practitioner group in this study to access original research articles from academic journals. These findings have important implications for the effective translation of research knowledge, indicating the ways in which research evidence may best be made accessible to meet the preferences of practitioners. The findings also have some limitations and raise issues for both for the design of future studies and for initiatives directed at researcher-practitioner engagement. With respect to limitations, the two samples were not completely representative. The practitioner group, who had responded to a request to complete an online survey may already have been motivated towards research. Moreover, as recruitment included the use of email contact lists, some of which originated from conference delegate lists, the sample may have had an interest in research to some degree. Similarly, researchers who completed the survey were likely to have an interest in communicating research findings beyond the academic community. It could also be argued that as respondents were aware that the purpose of the survey was to inform the development of a new knowledge hub aiming to improve connections between autism researchers and practitioners, their responses may have been biased, reflecting a desire to demonstrate good practice both in terms of accessing research information and engaging with other communities. This potential bias highlights the limitations of self-report measures, which are dependent not only on the honesty of respondents but also on the accuracy of their evaluation of their own behaviour. A further potential limitation of the design of the survey could be that the response options for all three questions were presented as a list and in a fixed order. However, this did not appear to affect the results, as when asked which methods would be most beneficial in increasing access to research knowledge, the top three for both groups included options from the second half of the list. Moreover, some aspects of the question wording might suggest that caution should be applied when interpreting some of the responses options. For example, the relevant question posed to the researcher and non-academic professional groups regarding methods beneficial for increasing knowledge, differed slightly; however, this difference was necessary, as the aim in the current study was specifically to identify methods that would be beneficial for increasing practitioners’ research access and knowledge. Finally, some missing demographic information limited the potential for a fuller interpretation of the results. Although researchers were asked to indicate their current status as either early or established career researcher (or other), thus giving some indication of their level of experience, the practitioner group were not asked about their level of experience. Nor was either group asked to indicate their age or gender. It seems likely that some or all of these characteristics could potentially influence respondents’ perceptions and use of social media and technology. However, there were no significant differences in how the research and practitioner groups rated their use, or the perceived benefits, of social media and technology, suggesting that if there were differences in age, gender, or experience between the two groups, these differences did not unduly affect the findings. The relatively low priority placed on social media and other more accessible methods, however, is perhaps surprising given the prevalence of blogs and twitter postings from well-established researchers, research organisations, and charities. Moreover, the survey was advertised using e-mail and social media, and as such, some role for social media could have been expected. The reported lack of importance placed on social media and Google searches may have reflected recognition that some media sources are not necessarily direct scientific sources; that is, they may be brief, simplified interpretations of the evidence. This issue could be further explored by more specifically probing dependence on different types of social media, distinguishing between postings from credible research organisations or charities (e.g. the National Institute for Health and Care Excellence) and more general postings or comments on research findings in the mainstream media. In their ethnographic study of professionals working within Public Health Units, Lemay and Sá (2014) found that practitioners primarily relied on academic literature accessed through journals or professional training and communications to update their research knowledge. The group of professionals were actively engaged in using research to develop and implement evidence-based practice. Given the recruitment of a research-user group in the current study that was similarly focused on professional practice, a similar priority placed on more formal sources of research information was perhaps to be expected. Finally, it might be argued that that asking participants to indicate only three sources of information in the current questionnaire may have underestimated the role of more accessible sources such as social media. However, by constraining participants’ choice in this way, priorities were more clearly identified. The results from the current study indicate that more accessible methods such as social media were not given high priority by either group, suggesting that other sources were considered more informative or reliable. Further research is needed to identify the role played of social media in knowledge acquisition in this field. Other issues may be addressed in the design of future studies. For example, the design of the questionnaire did not allow for identification of differences between distinct practitioner groups working in either health or education practice who might hold different views. Further exploration of the preferences and priorities of the groups would also be helpful. For example, when selecting the top three options that would be beneficial for facilitating access to research knowledge, ranking of preferences would provide richer data. Another important future direction for further research is in understanding how both researchers and practitioners assess the quality of the research that they access. Improving practitioners’ confidence in their evaluation of the available evidence could empower them to apply the findings from research in their practice. Practitioners’ preferences regarding access to academic journals were a key finding. It is possible that very recent policy changes in open access publishing of academic research articles may be already reducing barriers to knowledge access. However, practitioners’ needs in this respect should, be continually reviewed and we recommend that researchers are attentive to the importance of making their published findings easily accessible for practitioner use. Another way to address the desire to access and use research is through a scientific training ‘master-class’ method that has proved to be effective for increasing research competencies in public health professionals (Jansen & Hoeijmakers, 2013), which also gives opportunity for direct contact and engagement. These ‘master-classes’ could be further extended to include training in the application of evidence-based practice, which was one of the top three options selected by practitioners as being beneficial in translating research knowledge. Previous commentary on the gap between research knowledge and practice in autism emphasizes the need to build new research evidence based on collaboration between research and non-research professionals (Parsons et al., 2013), a view that is also consistent with interactive models of knowledge translation in health sectors (Jacobson et al., 2003). Research evidence in the area of autism, however, indicates that research-users, including practitioners, family members and adults with autism, may experience low levels of engagement with researchers (Pellicano et al., 2014b). Differences in reported experiences of engagement by researchers and practitioner groups reported in this study point to the need to provide direct face-to-face contact and real world shared activities to improve mutual understanding with stakeholders. Mottron (2011) has highlighted the considerable insight that can be gained by including individuals with autism in all aspects of the research process – including the design, implementation, and interpretation of studies. Similar engagement between academic researchers and practitioners in clinical, educational and social care professions is also vital. We recommend that opportunities are increased for informal contact and equitable collaboration in the design of research studies, from the beginning of the research process (Parsons et al., 2013). In so doing, it will be possible to build stronger communication channels, more trusting relationships and a shared interpretation of research with practice professionals and other users (Milton, 2014; Parsons et al., 2013). This study investigated just one aspect of the complex pattern of knowledge sharing relevant for effective translation of research evidence into practice. Its focus was to explore views about access to research information and levels of engagement experienced. As such it might be considered as the tip of an iceberg, with a number of other barriers to evidence based practice still unexplored. The translation of research evidence into practice is a complex process and the relationship between researchers and research-users is not simply one of knowledge provider and user. Models that focus on the interactive nature of knowledge sharing between members of different academic and non-academic communities, stress the importance of active engagement between members. Yet it is important to acknowledge the notable lack of work in this area and the need to build initial foundations for such models. Building these foundations may first require the study of knowledge translation from researcher to researcher-user, and then trace the complementary influence of research-user expertise on research as a result of engagement and collaboration over time. Furthermore, as Mitton, Adair, McKenzie, Patton, and Waye- Perry (2007) points out, there is remarkably little formal evaluation of the actual effectiveness of the application of knowledge translation strategies in context. We recommend a new focus within basic research to ensure better understanding of “what works for whom and in what circumstances”. The results from the current study support a recommendation to provide access to a range of research-based information in the development of an online knowledge platform, from non-technical summaries to access to original research articles. Our findings also point to the need for improved engagement and communication, and continual review of user needs and the barriers to research implementation in practice. However, it is essential that future work should more fully characterize a range of potential factors that hinder knowledge sharing and translation in order to best integrate research and practice in autism."],["How do infants’ emerging language abilities affect their organization of objects into categories? The question of whether labels can shape the early perceptual categories formed by young infants has received considerable attention, but evidence has remained inconclusive. Here, 10-month-old infants (N = 80) were familiarized with a series of morphed stimuli along a continuum that can be seen as either one category or two categories. Infants formed one category when the stimuli were presented in silence or paired with the same label, but they divided the stimulus set into two categories when half of the stimuli were paired with one label and half with another label. Pairing the stimuli with two different nonlinguistic sounds did not lead to the same result. In this case, infants showed evidence for the formation of a single category, indicating that nonlinguistic sounds do not cause infants to divide a category. These results suggest that labels and visual perceptual information interact in category formation, with labels having the potential to constructively shape category structures already in preverbal infants, and that nonlinguistic sounds do not have the same effect. --------------------------------------------------------------------------------","One of the central questions in early language and cognitive development is how the emergence of language affects infants’ preverbal understanding of the world. During the past 20 years, many studies have shown that even very young infants can rapidly form categories on the basis of the static perceptual features of objects (Behl-Chadha, 1996; French, Mareschal, Mermillod, & Quinn, 2004; Mareschal & Quinn, 2001; Oakes, Coppage, & Dingel, 1997; Quinn, 2002, 2004; Quinn & Eimas, 1996) and that during the first year of life these early categories are gradually enhanced with more sophisticated knowledge such as feature correlations, sounds, motion, function, and animacy cues (Baumgartner & Oakes, 2011; Burnham, Vignes, & Ihsen, 1988; Pauen & Träuble, 2009; Perone & Oakes, 2006; Younger & Cohen, 1986). Other research has shown that in older children language can shape object categories by enabling the grouping together of perceptually dissimilar objects (e.g., dogs and whales as mammals) and the separation of similar objects (e.g., bats and birds) into different categories as well as forming the basis for inferences about hidden object properties (Graham, Kilbreath, & Welder, 2004; Welder & Graham, 2001). There is also evidence that labeling aids object individuation by 12 months of age (Xu, Cote, & Baker, 2005), and even younger infants at 9 or 10 months expect different labels to refer to different kinds of object (Dewar & Xu, 2007, 2009). However, although several studies have investigated the emerging effect of language on categorization during the first year of life, they have yielded little agreement and it is not yet clear how linguistic information interacts with object representations that have developed preverbally. Two separate questions have been asked about the role of language in infants’ object categorization during the first year of life. The first is whether labels can facilitate object categorization, that is, whether the labeling of objects enables infants to form categories that they would not form without labels. The second question is whether, like in older children, labels can override nonverbal perceptual information and change the structure of perceptual categories when visual similarity and category labels are in conflict with each other. Much of the work exploring early categorization is based on the familiarization/novelty preference procedure (Fantz, 1964), which relies on the fact that infants tend to spend more time looking toward novel objects than toward familiar objects. In a typical study, infants are familiarized with a sequence of objects from one category and are then tested with two new objects—one a novel member of the familiarized category and one from a different category. If infants show a looking preference to the object from the new category, it can be concluded that they have formed a category that includes the novel within-category object but excludes the object from the different category. The emerging ability of language to shape categories has been studied in variants of this paradigm in which the objects presented during familiarization are accompanied by auditory labels or other sounds and the effect on infants’ looking preferences during test are investigated. The question of whether labels facilitate category formation has been addressed in several studies by Waxman and colleagues (Balaban & Waxman, 1997; Booth & Waxman, 2002; Ferry, Hespos, & Waxman, 2010; Fulkerson & Waxman, 2007; Waxman & Braun, 2005; Waxman & Markow, 1995). For example, in a seminal study, Balaban and Waxman (1997) showed that infants who were familiarized with a sequence of pig drawings exhibited at test a looking preference for a rabbit over a novel pig only when the familiarization items were accompanied by a labeling phrase (“A pig!”) but not a tone sequence. Unsystematic labeling with different novel words did not lead to category formation in 12-month-olds (Waxman & Braun, 2005). Other studies have aimed to provide evidence that consistent novel labels facilitate categorization in infants at 6 months of age and even at 3 or 4 months (Ferry et al., 2010; Fulkerson & Waxman, 2007). This and other work has led to the claim that words serve as “invitations to form categories,” enabling category formation by highlighting commonalities between objects with the same label (Waxman & Markow, 1995). However, two considerations appear to weaken this claim. First, as described, many studies have shown that preverbal infants as young as 3 months can form perceptual categories in the absence of any auditory input, and these early categories can be at the basic or superordinate levels (Behl-Chadha, 1996; Quinn & Eimas, 1996). On this basis, it seems plausible that, in silence, 9-month-olds may be able to form a perceptual category of the stimuli used in the studies described above (e.g., a category of pigs that is different from rabbits). Because the described labeling studies did not involve a control condition in which the objects were presented in silence, it is difficult to assess whether labels had an effect above and beyond purely visual information (Plunkett, Hu, & Cohen, 2008; Robinson & Sloutsky, 2007a). Sloutsky and Napolitano (2003) and Sloutsky and Robinson (2008) have aimed to reconcile the results from categorization-in- silence studies and categorization-with-labels studies by arguing that, far from facilitating categorization, auditory information instead interferes with visual processing and, therefore, disrupts category formation when categories would have been formed in silence. According to this “auditory overshadowing hypothesis,” labels disrupt category formation less than other auditory stimuli because they are highly familiar auditory signals. Therefore, visual categories that are learned in silence can still be formed in the presence of labels, but tone sequences usually disrupt visual category formation (Robinson & Sloutsky, 2007a, 2007b). A second limitation of the described labeling studies is that they familiarized infants with only a single category. Even if we assume that labels have a facilitative effect above and beyond presentation in silence, such a paradigm cannot answer the question of whether labels constructively interact with preverbal representations to shape category structures. It is a well-known finding that infants’ attention is maintained longer in the presence of labels compared with silence (Althaus & Mareschal, 2014; Baldwin & Markman, 1989; Robinson & Sloutsky, 2007a). The processes involved in this longer looking period may well have to do with more intensive visual processing, but not necessarily in the sense of using the information given by the label to adjust category boundaries. In familiarization with a single category, an increase in novelty preference might simply be the outcome of an effectively prolonged stimulus exposure. The question of a constructive role of category labels can be addressed by familiarizing infants with two separate but visually similar categories with two distinct labels simultaneously. This is similar to word learning studies that test whether infants can associate words and (individual) objects (Mani & Plunkett, 2008; Schafer & Plunkett, 1998; Werker, Cohen, Lloyd, Casasola, & Stager, 1998). This paradigm can then be used to answer the two questions posed above. If infants do not form categories in silence but are successful under labeling, then labels facilitate category formation. If infants successfully categorize in silence but form different categories under labeling, then labels can serve to override visual information. Plunkett and colleagues (2008) reported a two-category study, using a set of stimuli introduced by Younger (1985). These stimuli consisted of two sets of line drawings of schematic animals based on the same set of four distinctive features. In the “broad” set, features were combined such that there were no discernable clusters of stimuli; the eight stimuli formed a single large category. In the “narrow” set, the feature values were correlated (e.g., long-necked animals always had thick tails), and based on these correlations two subcategories of four stimuli could be formed. In Younger’s study, 10-month-old infants were familiarized with the animal drawings in silence. Younger found that the infants were sensitive to the feature correlations and formed a single category for the broad set but formed two distinct categories for the narrow set. In a separate study, Younger and Cohen (1986) showed that 4-month-old infants were not able to make use of feature correlations, indicating that this ability develops between 4 and 10 months of age. Plunkett and colleagues (2008) introduced labeling into Younger’s (1985) paradigm. When 10-month-old infants were familiarized using the broad stimulus set and were provided with the same label (“Look! Dax!”) for every stimulus, they formed one category, just like in silence. Likewise, when infants were familiarized with the narrow stimulus set and heard two different labels corresponding to the two subcategories, they formed two categories, again like in the silent case. However, importantly, when infants saw the narrow stimulus set but heard the same label for each stimulus, they formed a single category. Plunkett and colleagues argued that in this case the labels served to override visual categorization, leading infants to merge two perceptually distinct categories when they shared a common label. However, these results leave open an alternative explanation, namely that forming two categories in the narrow condition depended on infants’ ability to detect feature correlations, an ability that develops between 4 and 10 months of age. It is possible that the added complexity of the condition in which visual information (two categories) and labels (one label) were in conflict led to overloading the infants, resulting in a more shallow processing of the visual information. This explanation is compatible with results showing that when complex stimuli (e.g., conflicting visual and auditory information) overload the learning system, infants can regress to an earlier level of processing—here, that of 4-month-olds (Cohen, Chaput, & Cashon, 2002). In this case, instead of infants constructively using label information to merge two distinct categories, they could merely have failed to detect the feature correlations that define the two categories in the first place and, like 4-month-olds, formed a single category of the visual stimuli. One way in which to exclude the possibility that labels merely lead to shallower object processing is to conduct a study with a similarly ambiguous visual stimulus set that is in the absence of labels perceived as a single large category but that is parsed into two subsets when two distinct labels are presented. In other words, distinct labels cause infants to divide a large category rather than a single label causing them to merge two sets of objects. In this case, deeper processing is a prerequisite for the change in category boundary because a more detailed representation is necessary for subcategory formation. Such a result would constitute evidence for a constructive effect of labels. This is the approach taken in our study. In sum, previous work with infants at the transition to language has not provided a coherent picture of how nonlinguistic and linguistic information interact in category formation. In particular, the central question of whether labels can guide the structure of object categories in preverbal infants remains open. Demonstrating such a constructive role requires the following methodological considerations. First, a labeling condition should be contrasted with a silent condition to ensure that labels have an effect beyond the visual information alone. Second, infants should be familiarized with a stimulus set that allows parsing into one or two categories to show that labels can shape the encoded category structure in a way that goes beyond merely enhancing attention. Third, the task should be designed so that category formation in the presence of labels cannot be explained by the processing of less visual detail than when familiarized in silence. A fourth point concerns the difficulty of ensuring that looking preferences represent novelty preference rather than familiarity preference. Previous work has demonstrated that infants’ mode of preference can change as a function of task complexity or duration of exposure (Hunter, Ames, & Koopman, 1983; Roder et al., 2000). Because the addition of (one or more) auditory stimuli arguably changes task complexity, a test trial including a visually novel test item is necessary to establish whether infants display a novelty preference (longer looking at the novel test item) or a familiarity preference (shorter looking at the novel stimulus) (Cohen, 2004; Oakes, 2010). Note that Plunkett and colleagues (2008) did not include a novel test stimulus, which complicates interpretation of their results. Here, we describe a study based on these considerations. We presented infants with a set of objects that we expected to be perceived as a single category when viewed in silence, and we asked whether pairing half of the objects with one label and half with another label would lead infants to split the objects into two categories according to their labels. This would require infants to use more detail for categorization with labeling than in the silent condition. We constructed a set of stimuli by morphing a drawing of one novel animal into a drawing of a distinct second animal in a fixed number of steps, thereby obtaining a set of perceptually equidistant stimuli over which it was possible to define categories as well as subcategories (at either end of the morphing continuum) and their averages. The overall logic of our study followed that of Younger (1985), but here differences between objects were based not on the variation in distinct features but rather on a holistic difference achieved by the morphing process. Infants were familiarized with eight objects either in silence, with one label, or with two labels. If labels serve to constructively shape category structure, infants should form a single category comprising all stimuli when familiarized in silence or when accompanied by a single label but should form two categories when two distinct labels are provided.","A total of 63 full-term 10-month-old infants participated in this study (mean age = 302 days; 28 girls). Infants were randomly assigned to either a silent condition (n = 20), a two-label condition (n = 20), or a one-label condition (n = 23). One additional infant was tested but excluded due to technical problems. Infants were recruited on a voluntary basis via local advertisements. Informed consent was received from the caregivers, and infants received a small gift for their participation. Visual stimuli The familiarization and test stimuli are depicted in Fig. 1. All stimuli were obtained by morphing Stimulus 1 into Stimulus 19 in 18 steps using the software MorphMagic. To ensure discriminability of the individual exemplars, only every second morphed image was used in the familiarization set. The gap between Stimuli 7 and 13 ensured that there was a visual grounding for a division into two separate subcategories (Stimuli 1, 3, 5, and 7 and Stimuli 13, 15, 17, and 19, respectively). Choice of test stimuli followed the design of Younger’s (1985) study. It was assumed that infants forming one broad category would find the overall average stimulus (here, Stimulus 10) less novel than more peripheral stimuli (here, Stimuli 4 and 16). Conversely, infants who form two categories should find the overall average more novel because it falls between the categories but should not show preference for the peripheral stimuli because they now fall into the center of each of the two categories. A novel test animal was constructed to establish whether infants were showing novelty or familiarity preference at test. The familiarization stimuli were depicted against a white background and presented on either the left or right half of the screen (counterbalanced such that no regular structure emerged). Stimuli subtended approximately 16 degrees visual angle on the screen. Roughly half of the infants saw the stimuli facing left, and the other half saw them facing right. The test stimuli were depicted pairwise facing in the same direction as the familiarization stimuli, either depicting the overall average (10) next to one of the subcategory averages (4 or 16) or depicting one of the average objects (10, 4, or 16) alongside the novel animal. This novel animal had been constructed to have the same configuration as the overall average (10) but different individual features. Features were chosen to correspond to those that were relevant to discriminate original animal drawings that served as the basis for morphing (e.g., tail, feet, ornament on tummy, wings, head shape) but looked different enough to be easily discriminable. Auditory stimuli The auditory stimuli were the phrases “Look, a geepee!” and “Look, a boota!” recorded digitally from a female native speaker of British English in enthusiastic, infant-directed speech. The label (“A geepee!” /“A boota!”) was repeated twice, thereby occurring three times per stimulus. The onset of the first auditory phrase was at 2000 ms after the onset of the visual stimulus, with the labeling phrase (“A geepee!” /”A boota!”) repeated at 5000 and 8000 ms after stimulus onset. In the two-label condition, half of the infants heard the label “geepee” with Stimuli 1, 3, 5, and 7 and heard the label “boota” with Stimuli 13, 15, 17, and 19—and vice versa for the other half of the infants. In the one-label condition, half (n = 12) of the infants heard “geepee” paired with all eight stimuli, and the other half heard “boota” paired with all eight stimuli.","After a warm-up phase during which the experimenter explained the procedure and obtained written consent, infants were seated on the parent’s lap at a distance of approximately 65 cm from a 22-inch screen. A 9-point calibration sequence was used to calibrate a Tobii X120 remote eye tracker (sampling frequency = 120 Hz, system accuracy = 0.5 degrees). During this procedure, a small animated object with accompanying sound was displayed at nine locations (left, center, and right for top, middle, and bottom row) on the screen. Calibration was repeated up to three times or until all nine points had been calibrated successfully. After the calibration procedure, infants were presented with the eight familiarization stimuli, shown in a pseudorandomized order for on average 9750 ms (SE = 22) each. Individual sequences were obtained by using three Latin squares of three pseudorandom sequences, ensuring that no more than three consecutive stimuli were from the same subcategory. In the label conditions, infants heard the labels as described above. Prior to the first trial, and in between all subsequent trials, small animated video clips were shown at the center of the screen to direct infants’ attention to this location. Their duration varied between 800 and 1800 ms. The familiarization phase was followed by a test phase consisting of six trials of 10 s each. The first and second, third and fourth, and fifth and sixth test trials were identical apart from left/right positioning of the test stimuli. The first two test trials always paired the overall average (Stimulus 10) with one of the subcategory averages (Stimulus 4 or 16, counterbalanced). The order of Test Trials 3/4 and 5/6 was counterbalanced; these presented infants with (a) the novel item paired with the overall average and (b) the novel item paired with the same subcategory average that was used on Test Trials 1 and 2 (i.e., each child either saw the subcategory average, 4, on Test Trials 1, 2, and either 3 and 4 or 5 and 6, or they saw the subcategory average, 16, on these four trials). The reason for always presenting the pair of overall average and subcategory average at the start of the test phase was that this is the most sensitive contrast and needs to be presented while infants are at the peak of familiarity with the target stimuli. In particular, there is the possibility that once the novel stimulus has been presented, the overall average and subcategory averages will appear equally uninteresting to participants regardless of the category structure established by the end of familiarization. The location (left/right) of the items on Test Trials 1, 3, and 5 was counterbalanced across infants. The test trials were silent in all conditions. Looking time during familiarization To assess whether familiarization had occurred, we compared average looking times during the first four and last four familiarization trials in all three conditions. The data were submitted to a mixed two-way analysis of variance (ANOVA) with the between-participants factor condition (silent, two-label, or one- label) and the within-participants factor block (Block 1 [Trials 1–4] or Block 2 [Trials 5–8]). There were significant main effects of condition, F(2, 60) = 15.357, p < .001, and block, F(1, 60) = 14.563, p < .001. The interaction between condition and block was not significant (F = 0.087, p = .917). Post hoc tests confirmed that infants in both label conditions had longer looking times than those in the silent condition (two-label: M = 8275 ms, SE = 280; one-label: M = 8274 ms, SE = 271; silent: M = 6249 ms, SE = 271; both ps < .001, Bonferroni- adjusted). Looking times between one- and two-label conditions did not differ significantly (p = .59). This result is consistent with previous studies showing longer looking under auditory input than in silence (Robinson & Sloutsky, 2007a). Furthermore, average looking times during Block 1 (M = 7763 ms, SE = 160) were higher than those during Block 2 (M = 7113 ms, SE = 189; p < .001, Bonferroni- adjusted), as demonstrated by a paired t-test on the collapsed data. This indicates that infants across all conditions had become familiarized by the end of the familiarization phase. Preferential looking during test trials Test trial results are given in Table 1 (proportions and test results) and Table 2 (absolute looking times). For the first two test trials, a preference score for the overall average object was obtained for each infant by dividing the total looking time at the overall average object (Stimulus 10) by the total looking time on that trial. For the remaining four test trials, a novelty preference score was obtained for each infant by dividing the looking time at the novel object by the total looking time on that trial. Preference scores for each pair of test trials (1 and 2, 3 and 4, and 5 and 6) were calculated as the average between the two. A mixed effects ANOVA with the between-participants factor condition (silent, two- label, or one-label) and the within-participants factor test trial (novel vs. overall average or novel vs. subcategory average) revealed a significant interaction of condition and test trial, F(2, 59) = 4.96, p = .010.1 The main effects of condition, F(2, 59) = 0.467, p = .629, and test trial, F(1, 59) = 0.013, p = .91, were not significant. Planned comparisons against chance (.50) were conducted for each test trial in order to test whether infants preferred one stimulus over the other. In the silent condition, infants showed no preference for either stimulus when the subcategory averages and overall average stimuli were paired, t(19) = 0.67, p = .512. Similarly, there was no looking preference in the pairing of the novel stimulus and the subcategory average, t(18) = 1.10, p = .27. However, infants exhibited a preference for the novel stimulus when it was paired with the overall average stimulus, t(18) = 4.36, p < .001, d = 1.00). This looking pattern indicates that the overall average, but not the subcategory averages, was more familiar than the novel exemplar. These results, therefore, indicate that in the silent condition infants formed a single category. In the two-label condition, results for Test Trials 1 and 2 again showed no preference for either item, t(19) = 0.279, p = .78. For the remaining test trials, the results from the silent and one-label conditions were reversed; infants showed no preference when the overall average and novel item were paired, t(19) = 0.79, p = .43, but had a preference for the novel stimulus when paired with the subcategory averages, t(19) = 4.17, p < .001, d = 0.94. These results indicate the formation of two categories in the two-label condition, resulting in the subcategory averages being perceived as familiar and the overall average stimulus being perceived as novel. Results in the one-label condition resembled those in the silent condition; with no preference on the first two test trials, t(22) = 0.97, p = .34, infants exhibited a significant novelty preference for the novel test item when it was paired with the overall average, t(22) = 2.57, p = .017, d = 0.54, but not when it was paired with the subcategory averages, t(22) = 0.438, p = .66. As in the silent condition, this pattern of results indicates that the infants formed one category when all objects were paired with the same label. Taken together, these results indicate that infants in the silent condition and in the one-label condition formed a single category comprising all eight familiarization items, but they formed two categories when the objects were paired with two labels.","Experiment 1 showed that infants formed a single category when objects were presented in silence and when they were paired with a single label. In contrast, infants split this category in two when the objects were paired with two distinct labels. Infants who were familiarized with the visual stimuli in silence appeared to accept a large variety of head shapes, leg types, and body posture as possible realizations of the target category. Infants’ preferences during the test phase indicated that in the silent condition they integrated all of the different familiarization exemplars into a single category representation whose prototype was the overall average (see also Younger & Gotlieb, 1988, for evidence that infants form prototype representation for perceptual categories). In contrast, infants distinguished between the two subsets at either end of the morphing continuum when the same objects were accompanied by two different labels, forming two more restrictive exclusive categories. If this result were due to labels merely enhancing visual processing overall (i.e., the trials simply being more interesting in the presence of speech), we would expect two categories to be formed regardless of how many different labels accompanied the visual stimuli. This is not the case; when a single label was presented with all stimuli, infants formed just one large category. Our results, therefore, suggest that infants tracked the correspondences between labels and object features. Furthermore, the current pattern of results is incompatible with the auditory overshadowing hypothesis, which predicts that visual processing in the presence of labels is less detailed than in silence; to form two categories, infants needed to process more detailed visual information to enhance the perceptual difference between the objects in the two subcategories. In particular, our findings indicate that the condition with the largest amount of novel auditory stimulation (two novel labels) leads to the most detailed representation (two separate subcategories). Together, therefore, these results provide unequivocal evidence that labels can have a constructive role in category formation in infants prior to their first productive use of language. One open question with regard to infants’ performance on the test trials is why infants in all three conditions did not exhibit any preference on the first two test trials contrasting the overall and subcategory averages. A possible explanation is that novelty preference depends on a sort of “feature pop-out.” Because the stimuli were obtained using morphing, there were no clearly identifiable feature differences between stimuli presented side by side, that is, between the overall average stimulus (10) and the two subcategory averages (4 and 16). It seems plausible that the lack of clear novelty could have prompted infants to compare stimuli, resulting in overall similar looking times to both test items. Specifically, none of the stimuli used here (the overall average, 10, and the subcategory averages, 4 and 16) violates the familiarized category structure. Even if an infant has formed two categories during familiarization and, therefore, the overall average (10) is perceived as relatively novel, the infant could attempt to integrate this stimulus into the previously constructed category structure by merely extending the category boundaries without needing to completely re-form the categories. Similarly, if the infant has formed one category, the subcategory average is not implausible as a category member but merely further away from the overall category average. Infants’ looking patterns on test trials containing a novel stimulus suggested competition effects between the two test stimuli presented side by side; pop-out of the features of the novel stimulus creates a novelty effect, and if the other stimulus is perceived as highly familiar, it presents no competition to the novel item. In contrast, if the other item is perceived as less familiar, infants will spend time exploring the novel stimulus but also scan the other item because more processing is necessary to incorporate it into the established category. In other words, competition between the test items here, or lack thereof, reveals how different category boundaries have been formed. One remaining open question is whether the observed constructive effect for labels in category formation is specific to language or whether it extends to other forms of auditory input. In the work by Waxman and colleagues (Balaban & Waxman, 1997; Ferry et al., 2010; Fulkerson & Waxman, 2007), labels were contrasted with tone sequences and differential effects were found. However, as discussed above, these differences were interpreted by Sloutsky and colleagues as being due to the different familiarity of speech sounds and nonspeech sounds. Even in studies that did not find positive effects for labels, nonlinguistic sounds had a more detrimental impact than labels (Robinson & Sloutsky, 2007a). Finally, the study by Plunkett and colleagues (2008), which is so far the best controlled study, did not involve a nonspeech sound condition. Therefore, to establish whether labels take a special role in shaping visual categories, we tested an additional group of infants on a nonlinguistic sound condition.","The design and procedure for this experiment were identical to the two-label condition in Experiment 1 except that instead of auditory labels, nonlanguage sounds were played.","A total of 17 10-month-old infants participated in this experiment (mean age = 300 days; 8 girls). A further 6 infants were excluded due to fussiness (n = 1) or technical problems (n = 5).","The visual stimuli were the same as in Experiment 1. Two sounds were used: a tingling bell sound and a wooden xylophone tone sequence. Each sound lasted 1 s. Each sound was played three times during the presentation of a picture, starting at 2000, 5000, and 8000 ms after stimulus onset. Half of the infants heard the bell sound with Stimuli 1, 3, 5, and 7 and the xylophone sound with Stimuli 13, 15, 17, and 19—and vice versa for the other half of the infants. Procedure The procedure was identical to that in Experiment 1. Familiarization A paired t-test on average looking times for Blocks 1 and 2 of familiarization (Block 1 [Trials 1–4] or Block 2 [Trials 5–8]) revealed that there was a trend for infants to begin to habituate by the second block of familiarization (Block 1: M = 6937 ms, SE = 446; Block 2: M = 5999 ms, SE = 573), t(16) = 2.033, p = .059. An ANOVA on the average looking time per trial was conducted to compare looking time during familiarization in Experiment 2 with the conditions of Experiment 1. This revealed a significant effect of condition, F(3, 76) = 10.13, p < .0005. Post hoc comparisons showed that infants in Experiment 2 spent approximately equal amounts of time gazing at the stimuli (M = 6468 ms, SE = 336) as those in the silent condition of Experiment 1 (p > .99); that is, they exhibited shorter familiarization looking times than infants in the labeling conditions (one-label: p = .023; two-label: p = .001). Test trials Looking preference scores for all test trials are shown in Table 1. Test trials and calculation of looking preferences were the same as in Experiment 1. For Test Trials 1 and 2 (overall average vs. subcategory average), the preference score did not differ from chance, t(16) = 1.48, p = .16, two-tailed. In test trials pairing the novel stimulus with the overall average stimulus, there was a marginally significant preference for the familiar item (i.e., the overall average) over the novel item, t(16) = 2.01, p = .061, d = −0.49. On the test trials pairing the novel stimulus with a subcategory average, infants did not reliably prefer either stimulus on these two test trials (t‐test on average preference scores: t(16) = 1.13, p = .28). In summary, infants’ results differed from those found in Experiment 1. The marginal familiarity preference exhibited when the novel stimulus was paired with the overall overage stimulus indicates that infants formed a single category in the two-tone condition of Experiment 2.","The two-label condition in Experiment 1 revealed that pairing objects from a single category with two labels led infants to split the category into two separate categories. Pairing the visual stimuli with two nonlinguistic tones did not have the same effect. The only (marginal) preference shown by the infants was one for the overall average when it was paired with the novel item. This looking pattern can only be interpreted as a familiarity preference. As described above, familiarity preferences can occur when the familiarization stimuli are complex and/or infants have not become sufficiently familiarized during training. In line with our interpretation of the results in Experiment 1, we need to interpret the familiarity preference here as indicative of infants forming a single category over the familiarization stimuli. This means that despite the presence of two distinct auditory stimuli, infants are not persuaded that there are two visual categories—in contrast to the label condition, where their category representation was guided by the structure provided by the labels. This result is compatible with the auditory overshadowing hypothesis in the sense that infants (a) seem to struggle with category formation, as indicated by the marginal test preference, and (b) do not use the auditory stimuli to group the visual objects. Therefore, nonlinguistic tone stimuli, unlike labels, appear to interfere with infants’ processing of visual material. Although category formation per se does not appear to be disrupted in Experiment 2, the finding is consistent with results from previous studies showing that nonlinguistic sounds do not facilitate category formation (Balaban & Waxman, 1997; Ferry et al., 2010; Fulkerson & Waxman, 2007; Robinson & Sloutsky, 2007a; Sloutsky & Robinson, 2008).","Our aim in this work was to investigate how object labels interact with visual information during category formation in preverbal 10-month-old infants. Although a body of literature that has addressed this question exists, we argued that methodological considerations left the results of these experiments inconclusive. We argued that to investigate whether labels can constructively guide category formation and override visual similarities in young infants it is necessary to (a) familiarize infants on a set of stimuli that could be parsed into one or two categories, (b) test infants under varying labeling conditions (silence, one label, and two labels), (c) include a novel test stimulus to ascertain whether at test infants show a novelty or familiarity preference, and (d) construct the stimulus set so that the effects of labels cannot be explained as processing less visual detail than without labels. We presented a paradigm that incorporates these considerations, familiarizing infants on a stimulus set that could be grouped into one large category or two smaller categories. We found that on presentation in silence or when all objects were paired with the same label, infants formed a single category. In contrast, when half of the objects were paired with one label and the other half were paired with another label, infants formed two separate categories in line with how the objects had been labeled. Conversely, when the objects were paired with nonlinguistic sounds, infants formed a single category but exhibited familiarity preference. In infants at the transition to language, these results present clear evidence that labels interact with nonlinguistic information to constructively shape categories and that these effects are specific to linguistic labels and do not extend to nonspeech sounds. The results of our experiments reconcile arguments that sounds act to overshadow visual processing and that labels have a constructive effect on category formation. The results of Experiment 2 suggest that nonlinguistic tones indeed interfere with visual processing so that infants did not form a category in the tone condition, whereas in silence they did. However, whereas Sloutsky and Napolitano (2003) and Sloutsky and Robinson (2008) argued that labels merely interfere less with visual processing than nonlinguistic sounds but contribute nothing to category formation, our Experiment 1 clearly shows that labels do in fact constructively affect categorization. Therefore, depending on their nature, auditory signals can both interfere with and shape visual categories. Furthermore, recent results also indicate that the temporal relationship between presentation of visual and auditory stimuli matters so that interference is especially strong when visual and auditory stimuli have a common onset (Althaus & Plunkett, 2015). Some authors have discussed the role of labels in terms of whether they function as features that increase similarity between objects in a bottom-up manner (Gliozzi et al., 2009; Sloutsky et al., 2001) or whether they act in a referential manner, serving as “names” (Waxman & Gelman, 2009). Waxman and Markow (1995) described the role of labels as “invitations to form categories” by highlighting common visual features, suggesting a top-down role for labels. Likewise, Waxman and Gelman (2009) argued that words are “referential” in nature and, as such, are more than associates or perceptual features. Gliozzi and colleagues (2009), in contrast, demonstrated that even results such as those presented by Plunkett and colleagues (2008), which show a change in category formation based on label–object correlations, can be simulated with a connectionist model that treats labels as features. In our study, labeling caused infants to use more restrictive criteria for a classification of two items as “similar,” effectively producing two smaller categories. The label, therefore, serves not so much as a mere additional feature but instead as a feature that modulates the way in which visual features are used. As such, our interpretation of the labels’ role is somewhere between the extremes of “labels as features” and “labels as names”; labeling has a supervisory function in the sense that it modulates the processing of similarity, but whether this is enough to establish a “referential” relationship between the labels and objects in these infants at just 10 months of age remains open. This interpretation is in line with recent computational models (Althaus & Mareschal, 2013; Westermann & Mareschal, 2014) that simulate the role of words in the context of categorization. Althaus and Mareschal (2013) demonstrated how early interactions between word learning and learning about objects led to improved category representations compared with isolated learning, without labels explicitly acting as supervisory signals. The model by Westermann and Mareschal (2014) shows how labels can affect the similarity relations between objects in a top-down manner by warping the perceptual similarity space to increase perceptual distance between objects that have different labels. Such a mechanism could also account for the results presented here. Our findings fit well with recent theories about interactions of language with other domains. Although the strong Whorfian hypothesis of linguistic determinism (Whorf, 1956) has been largely discounted over the last few decades, newer research is gradually establishing links between language processing and perception or cognition in adult processing (Boroditsky, 2001; Lupyan, 2008a, 2008b; Lupyan, Rakison, & McClelland, 2007). In particular, the “label feedback hypothesis” (Lupyan, 2012) suggests that language modulates cognitive processes, for example, benefitting visual search or recognition memory (Lupyan, 2008a, 2008b). It seems likely that the effects we observe in the experiments presented here reflect early stages of these interactions between language and cognition. To summarize, the experiments presented in this article show that labels have a constructive impact on category formation in preverbal infants; when a set of objects is accompanied by labels, 10-month-old infants adapt their category structure in correspondence to how the objects are labeled. This ability necessitates a deeper processing of objects when they are paired with labels than when they are presented in silence, with different labels highlighting visual differences between objects. Our results further indicate that the effect of labels is not merely to enhance attention but that it is also a constructive process in the sense that infants’ categorization behavior was contingent on the number of labels they heard and that infants aligned their categories with the labels. Object labels, thus, are powerful stimuli that are able to direct infants’ attention to different visual dimensions in the construction of visual categories."],["This study examined the influence of younger siblings on children's understanding of second-order false belief. In a representative community sample of firstborn children (N = 229) with a mean age of 7 years (SD = 4.58), false belief was assessed during a home visit using an adaptation of a well-established second-order false belief narrative enacted with Playmobil figures. Children's responses were coded to establish performance on second-order false belief questions. When controlling for verbal IQ and age, the existence of a younger sibling predicted a twofold advantage in children's second-order false belief performance, yet this was the case only for firstborns who experienced the arrival of a sibling after their second birthday. These findings provide a foundation for future research on family influences on social cognition. --------------------------------------------------------------------------------","Individual differences in children’s development of theory of mind (ToM), defined as the “understanding of mental states, what we know or believe about thoughts, desires, emotions, and other psychological entities both in ourselves and in others” (Miller, 2009, p. 749), have traditionally been explored using the false belief task (Perner & Wimmer, 1983). Researchers have noted various sources of individual differences on this task, including number of siblings (Lewis, Freeman, Kyriakidou, Maridaki-Kassotaki, & Berridge, 1996; McAlister & Peterson, 2007; Perner, Ruffman, & Leekam, 1994; Ruffman, Perner, Naito, Parkin, & Clements, 1998), family sociodemographic status (Cole & Mitchell, 2000; Cutting & Dunn, 1999), and maternal education level (Pears & Moses, 2003). Passing false belief tasks has also been found to be related to children’s language (Astington & Jenkins, 1999) and executive function (Carlson & Moses, 2001). Although it is well established that preschoolers with older siblings outperform those without siblings on ToM tasks (Lewis et al., 1996; Ruffman et al., 1998), the influence of younger siblings on ToM remains unclear. Piaget (1959) suggested that both younger and older siblings facilitate social understanding through discussion and reflection. Dunn (1994) claimed that siblings influence social understanding through talk about causality and internal states, management of conflict by parents, joint play, shared jokes, and reasoning about moral issues; both younger and older siblings may equally facilitate ToM (Jenkins & Astington, 1996; Perner et al., 1994; Peterson, 2000). Indeed, in a recent meta-analysis by Devine and Hughes (2016), the number of child-aged siblings, regardless of birth order, predicted false belief understanding during early childhood. Despite these findings, the evidence is mixed. Some studies found no effect of younger siblings on ToM tasks (Calero, Semelman, Salles, & Sigman, 2013; Farhadian et al., 2010; Ruffman et al., 1998; Shahaeian, 2015), and in one case younger siblings had a negative effect on ToM (Wright & Mahford, 2012). Younger siblings may influence ToM development negatively by placing increased demands on parents’ time, resulting in a decrease in mother–firstborn positive interactions (Baydar, Greek, & Brooks-Gunn, 1997), including play and conversation with the firstborn child (Dunn & Kendrick, 1980a, 1980b). It is also possible that parents’ explanations to their firstborn children are frequently interrupted due to younger siblings’ demands (Wright & Mahford, 2012). The age threshold model proposes that younger siblings may need to reach a certain threshold in age before providing a positive influence on ToM (Kennedy, Lagattuta, & Sayfan, 2015). If so, it is possible that some null findings may be due to the younger siblings in those studies being too young to provide any advantage. Younger siblings may become more important in fostering children’s more advanced understanding of minds during middle childhood, but research on sibling influences on the later development of ToM remains limited (Devine & Hughes, 2016; Hughes, 2016; Miller, 2009). Most studies examining younger sibling influence on ToM focused on first-order false belief tasks (Miller, 2009). However, during middle childhood, second-order false belief tasks are thought to be more age appropriate (Perner & Wimmer, 1985). Whereas first-order false belief tasks typically assess children’s understanding that someone may have beliefs that differ from their own, a second-order task assesses whether children understand that one story character can have a mistaken belief about another character’s belief. Some children pass this higher-order test of ToM between 6 and 7 years of age (Perner & Wimmer, 1985). Findings about sibling influence on ToM in older children are mixed; in some cases both younger and older siblings facilitated higher-order ToM performance (Kennedy et al., 2015; McAlister & Peterson, 2007), but in other studies younger siblings had no effect (Calero et al., 2013; Miller, 2013). It is possible that older siblings begin to benefit from younger siblings as the latter become more proficient playmates (Lagattuta et al., 2015). Alternatively, as firstborn children start school and spend less time with family members, the initial sibling advantage may disappear. Before a more definitive conclusion can be drawn, larger-scale studies are required to tease apart the benefits of particular kinds of sibling constellations (Cassidy, Fineberg, Brown & Perkins, 2005). Studies finding no effect of younger siblings may have lacked sufficient statistical power to detect smaller effects once samples are separated into sibling constellation groups (i.e., sibling presence, birth order, age spacing, and gender composition) (see Miller, 2013). “Only child” subsamples typically are small (Miller, 2013). This not only leads to a decrease in power to detect an advantage in having a sibling over none but also results in samples with a very high proportion of children who have siblings—in some studies more than 90%, which exceeds the estimate that 80% of children in Western families have a sibling (Volling, 2012). Although previous research on ToM has highlighted covariates that need to be accounted for in studies of sibling influence, rarely have these all been controlled in a single study, which may also explain the mixed findings. These covariates include child age (Wellman, Cross, & Watson, 2001), family context (Lewis et al., 1996), sociodemographic risk factors (Cutting & Dunn, 1999; Cole & Mitchell, 2000), language ability (Astington & Jenkins, 1999), and executive function, specifically working memory and inhibition (Carlson & Moses, 2001; Lagattuta, Sayfan, & Harvey, 2014). Children’s understanding of second-order false belief is positively associated with their language and executive function (Astington, Pelletier, & Homer, 2002; Lagattuta, Sayfan, & Blattman, 2010; Lagattuta et al., 2014; Perner, Kain, & Barchfield, 2002). However, in studies of sibling influences on ToM, rarely are age, sociodemographic risk, language, and executive function all controlled (see Kennedy et al., 2015; Miller, 2013); when examined together in one study, executive function was positively associated with second-order false belief when age was controlled but not when language ability was partialed out (Hasselhorn, Mahler, & Grube, 2005). Although correlates of first-order false belief may also be relevant for second-order false belief, this has not yet been fully established (Miller, 2012). Some of these correlates, such as executive function, may be most important during early development of ToM; after children reach a certain threshold of ToM skills during middle childhood, these relationships may attenuate or disappear (Lagattuta et al., 2015). To address these issues, we explored the ways in which younger siblings might influence 7-year-olds’ performance on a second-order false belief task while controlling for known correlates of ToM in a study of a nationally representative community sample of firstborn children and their families. Our moderately sized dataset of firstborn children and their families provided a unique opportunity to examine the effect of younger sibling constellation factors, including sibling presence, gender composition, and age spacing. Design ~~~~~~ The Cardiff Child Development Study (CCDS) is a prospective longitudinal study of a nationally representative sample of mothers and their firstborn children. Data collection took place during pregnancy and at means of 6, 12, 21, 33 and 84 months postpartum. The current study focuses on the home visit that took place at a mean age of 84 months. The CCDS is funded by the Medical Research Council (MRC), and ethical approval was obtained for the procedures from the National Health Service (NHS) Multi-Centre Research Ethics Committee and the Cardiff University School of Psychology Research Ethics Committee.","A total of 332 primiparous women and their partners were recruited between November 2005 and June 2008 from NHS antenatal clinics in hospitals and general practitioner (GP) surgeries in two Health Care Trusts in Wales, United Kingdom. The CCDS is nationally representative in terms of sociodemographic factors; it did not significantly differ from families with firstborn children in the large, nationally representative sample in the Millennium Cohort Study (see Hay et al., 2014). A total of 321 families were seen after the birth of the first child, with 286 (89.01%) assessed at 7 years; of these, 272 (95%) were directly observed at home. The current sample comprises 229 of these families (Fig. 1). The participants’ mean age at the time of testing was 83.20 months (range = 67–104). The demographic characteristics of the children included in this subsample (69.0% of the original sample) are summarized in Table 1. A child’s exposure to socioeconomic adversity was indexed by (a) the mother not having achieved basic educational attainments (i.e., having no qualifications or fewer than five general certificates of secondary education (GCSEs) or equivalent attainments), (b) the mother being 19 years of age or younger at the time of the child’s birth, (c) the mother not being legally married during the pregnancy, (d) the mother not being in a stable couple relationship during the pregnancy, and (e) the mother’s occupation being classified as working class according to the Standard Occupational Classification 2000 (SOC2000; Elias, McKnight, & Kinshott, 1999). A principal components analysis (PCA) based on the polychoric correlation matrix confirmed that all of these items contributed to a single component, which explained approximately 77% of the shared variance in these risk indicators. Summary scores derived from this PCA measured the family’s exposure to socioeconomic adversity (Perra, Phillips, Fyfield, Waters, & Hay, 2015). In terms of sibling constellation, 172 children (75.1%) had at least one younger sibling living in the home; of these, 133 (58.1%) had one sibling, 32 (14.0%) had two siblings, and 7 (3.1%) had three siblings. Gender composition and age spacing were examined with the younger sibling closest in age to the firstborn. In total, 91 children (52.9%) were in a same-gender sibling dyad and 81 children (47.1%) were in an opposite- gender sibling dyad; of these, there were 47 (27.3%) older boy–younger boy dyads, 44 (25.6%) older girl–younger girl dyads; 46 (26.7%) older boy–younger girl dyads, and 35 (20.3%) older girl–younger boy dyads. The firstborn children entered siblinghood at a mean age of 35.7 months (SD = 16.8). To investigate the influence of sibling birth interval, children were grouped according to the interval between the firstborn and secondborn sibling births. Children who entered siblinghood at or below the first quartile (≤24 months) were categorized as having an early arrival sibling (n = 45, 19.7%), children who entered siblinghood at or above the third quartile (≥43 months) were categorized as having a later arrival sibling (n = 44, 19.2%), and children with a sibling arriving between these quartiles were categorized as having an average arrival sibling (n = 83, 36.2%).","Research assistants visited each family at home for two 2-h sessions. The caregiver (typically the mother) was given questionnaires and interviewed by a trained research assistant to gather information on the caregiver and firstborn’s well-being as well as family lifestyle arrangements and social network. Where possible, these interviews would take place in a separate room from the child. During these interviews, the child completed various cognitive, social, and emotional assessments in a quiet space with a second trained research assistant. A third research assistant attended to keep any younger siblings occupied while the assessments took place. A remuneration of £20 was given to the caregiver, and a book voucher of £10 was given to the child, at the end of the session. Second-order false belief task This task was adapted from second-order belief paradigms (Coull, Leekam, & Bennett, 2006; Perner & Wimmer, 1985). Each child was told a story enacted with plastic Playmobil figures by the experimenter. The protagonist was gender matched to the participant, and the sibling was gender matched to the participant’s closest-in-age younger sibling. In cases where the focal child had no siblings, the sibling character’s gender was randomly selected. The narrative is shown in Fig. 2. Pathways to passing or not passing this task are shown in Fig. 3. Children were classified as minimally passing second-order false belief if they correctly answered the first location question with an appropriate justification, and they were classified as passing second-order false belief with full comprehension if they also correctly answered the additional probe questions. An independent observer coded transcripts for 32.9% of the participants and established excellent agreement for passing second-order false belief (kappa coefficient = 1.00) and for appropriate or inappropriate justifications (kappa coefficient = 1.00). There was also very good agreement within appropriate and inappropriate justification codes, where the kappa coefficients were .89 and .79, respectively. Verbal IQ Each child’s vocabulary knowledge was assessed using the British Picture Vocabulary Scale (BPVS; Dunn, Dunn, Whetton, & Pintillie, 1982). In this task, the experimenter spoke a word to the child, who was asked to point or say the number of the picture that corresponded to the word. Each child’s verbal IQ was calculated by age normalizing the data to produce a standardized score. The mean score for verbal IQ was 99.54 (SD = 11.99), and the average age children in the sample were equivalent to was 84.14 months (SD = 14.66) and ranged from 49 to 150 months. Executive function Cognitive function was assessed using tasks from the Amsterdam Neuropsychological Tasks (ANT) (de Sonneville, 1999). The ANT is a well-validated and sensitive instrument to evaluate executive functioning in population-based samples (Brunnekreef et al., 2007) and clinical samples (Rommelse et al., 2008). The tasks were presented on a laptop computer, and children made responses using a mouse. For each task, the experimenter gave verbal instructions while showing examples. Following this, children were given a practice trial before starting the test trials. The Response Organization Objects (ROO) task was used to measure response inhibition via children’s reaction times to stimuli. Children were asked to hold the mouse with a forefinger of each hand on each button of the mouse. In Part 1 (compatible condition), children were presented with a fixation cross in the middle of the screen and were asked to respond to a red ball appearing on either side of the cross by clicking the same side of the mouse on which the ball appeared. In Part 2 (incompatible condition), children were presented with a white ball on the screen. Children were instructed to click the opposite side of the mouse according to the position of the ball. Response inhibition was operationalized as the difference between children’s mean reaction speed times in milliseconds (M = 314.32 ms, SD = 195.65) between the incompatible (Part 2) and compatible (Part 1) tasks. The Visuo-Spatial Sequencing (VSS) task was used to measure visuo-spatial working memory. In this task, children were presented with a gray square containing 9 circles symmetrically positioned in a 3 × 3 matrix on a computer screen. After a beep, a sequence of circles was pointed at by a computer animated hand, and after the sequence children took control of the mouse to replicate the sequence of circles. The test consisted of 24 trials and gradually increased in difficulty in the number of targets and complexity of the sequence. Working memory was assessed using the total number of correct targets in the correct order, with a total of 100 possible correct targets. The mean score for correct targets in the correct order was 67.24 (SD = 17.94). Children’s understanding of second-order false belief ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Correlations, means, and standard deviations for all variables of interest are presented in Table 2. In total, 95 children (42.8%) passed the minimal second-order false belief questions, and 67 children (30.2%) passed the second-order false belief questions with full comprehension (Fig. 3). Minimal second-order false belief and second-order false belief with full comprehension were positively associated (Table 2). A Guttman scaling analysis using the Goodenough–Edwards method revealed a developmental progression from minimal second-order false belief to second-order false belief with full comprehension (CR = .99). Despite these two levels of passing second-order false belief being a part of a single continuum, younger sibling constellation factors were only associated with passing second-order false belief with full comprehension (Table 2). Therefore, the subsequent analyses focused on children’s full comprehension of second-order false belief. Prior to investigating the influence of siblings on this measure of false belief understanding, a preliminary investigation of its correlates was conducted. Correlates of second-order false belief understanding ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Examination of the correlation matrix (Table 2) and the collinearity statistics revealed no issues with collinearity among predictor variables: firstborn age, firstborn gender, sociodemographic risk, verbal IQ, response inhibition ANT, and working memory ANT (variance inflation factor < 10, tolerance > .20) (Menard, 1995; Myers, 1990). Verbal IQ and sociodemographic risk were significantly associated with passing the second-order false belief questions with full comprehension, with higher verbal IQ scores associated with better performance and higher sociodemographic risk scores associated with lower performance. No relationship was detected between the ANT measures of response inhibition and working memory and children’s passing second-order false belief with full comprehension, nor was a relationship detected between age at the time of testing and second-order false belief (all ps > .19) (Table 2). However, in view of earlier research suggesting that individual differences exist in performance on false belief tasks across different ages (Wellman et al., 2001), age was included in the subsequent logistic regression. In the logistic regression, these potential confounds accounted for 11% of the variance in second-order false belief with full comprehension, χ2(3) = 18.45, p < .001, Nagelkerke R2 = .11. Children who were older at the time of testing, Wald statistic = 4.21, p < .05, odds ratio (OR) = 1.08, 95% confidence interval (CI) = 1.00–1.16, and those who had higher verbal IQ scores, Wald statistic = 7.17, p < .01, OR = 1.04, 95% CI = 1.01–1.07, performed significantly better on second-order false belief; therefore, age and verbal IQ were used as covariates in the subsequent analysis. Do younger sibling constellation factors influence the firstborn’s second-order false belief performance? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There was no significant association between number of siblings living in the home and second-order false belief performance (r = .10, p = .15) (Table 2). Therefore, all subsequent analyses explored sibling constellation factors related to the closest-in-age sibling. Presence of a sibling in the home To test for variations in second-order false belief as a function of presence or absence of siblings in the home, the sample was divided into two groups. Preliminary analyses showed no differences between the groups in ratio of boys to girls, firstborn mean age, sociodemographic risk, verbal IQ, and working memory (all ps > .15) (see Table 3). Children with siblings performed better on the response inhibition task, t(76.83) = 2.03, p < .05. Children with a sibling had a twofold advantage in passing the second-order false belief task with full comprehension, χ2(1) = 5.00, p < .05, OR = 2.33, 95% CI = 1.10–4.97 (see Fig. 4). In a subsequent logistic regression analysis (Table 4), the covariates were entered into the first step of the model, which accounted for 9% of the variance in second-order false belief understanding, χ2(2) = 15.07, p < .001, Nagelkerke R2 = .09. At the second step, the presence of a younger sibling accounted for significant additional variance in understanding second-order false belief, χ2(1) = 4.97, p < .05, and the overall model remained significant, χ2(3) = 19.98, p < .001, Nagelkerke R2 = .12. Within this model, verbal IQ remained a significant predictor of second-order false belief performance. Children with a younger sibling were twice as likely as children without siblings to pass second-order false belief with full comprehension, Wald statistic = 4.53, p < .05, OR = 2.35, 95% CI = 1.07–5.15. Gender composition Gender composition was examined in two ways; after same-gender and opposite-gender dyads were compared, all four possible gender compositions—older girl–younger girl, older girl–younger boy, older boy–younger boy, and older boy–younger girl—were explored. Preliminary analyses showed no differences between the groups in ratio of boys to girls, firstborn mean age, sociodemographic risk, verbal IQ, and working memory or in inhibition across all of the groups (all ps > .10). No associations were detected between gender compositions of sibling dyads and second-order false belief (ps > .20). Birth interval Although no association was detected between timing of sibling arrival and second- order false belief understanding (r = .02, p = .81) (see Table 2), the four sibling arrival groups (no-sibling group and early-, average-, and late-arriving sibling groups) were investigated. Preliminary analyses showed no differences among these groups in ratio of boys to girls, verbal IQ, and working memory (ps > .15). Significant differences were detected among groups in sibling age, F(3, 224) = 3.16, p < .05, sociodemographic risk, F(3, 225) = 4.13, p < .01, and ANT inhibition scores, F(3, 220) = 3.12, p < .05. Post hoc tests were selected in accordance with results from tests for homogeneity of variances. Games–Howell post hoc tests indicated that children in the late-arriving sibling group were older than those in the no-sibling group and that children with an average-arriving sibling performed better on the inhibition task than those without a sibling (ps < .05). A Tukey post hoc test indicated that children with an average-arriving younger sibling had lower sociodemographic risk than those with a late-arriving sibling (p < .01). A significant difference was detected among the four sibling groups in their passing of the second-order false belief task with full comprehension, χ2(3) = 8.97, p < .05 (Fig. 5). This finding was explored further while controlling for covariates of second-order false belief. Because late- arriving siblings did not significantly differ from the average-arriving sibling group in performance on passing second-order false belief with full comprehension, these were collapsed into one “average to late”-arriving sibling group. There were no significant differences among the groups in ratio of boys to girls, firstborn mean age, sociodemographic risk, verbal IQ, and working memory or in inhibition when these groups were collapsed (all ps > .06). The three remaining sibling status groups were dummy coded, with the no-sibling group assigned as the reference category in a logistic regression. Covariates (age and verbal IQ) were entered into the first step of the logistic regression model. When entered into the model at the second step, early arrival of a younger sibling and average to later arrival of a younger sibling accounted for a significant step when entered into the model, accounting for an additional 4% of the variance in second-order false belief with full comprehension, χ2(2) = 6.57, p < .05. The overall model remained significant, χ2(2) = 21.57, p < .001, Nagelkerke R2 = .13. The early arrival of younger siblings did not predict firstborns’ passing of second-order false belief with full comprehension; however, “average to late”-arriving siblings conveyed a significant advantage, Wald statistic = 5.63, p < .05, OR = 2.66, 95% CI = 1.19–5.96 (Table 5).","When predictors of second-order false belief understanding were controlled, children with a younger sibling living in the home were twice as likely to succeed on a second-order false belief task. It was established that this sibling advantage occurred only for firstborns who did not experience the early arrival of a sibling. Our finding stands in contrast to the first study of sibling effects on second-order false belief tasks, which found no effect (Miller, 2013), but is consistent with previous research showing that presence of a younger sibling in the home is advantageous for ToM (Lewis et al., 1996; Perner et al., 1994; Peterson, 2000). In contrast to earlier work (Kennedy et al., 2015), the younger sibling’s influence on a higher-order ToM task in our sample was not limited to same-sex siblings. There are various mechanisms by which younger siblings could facilitate their siblings’ social understanding; these might include engaging in joint pretense (Youngblade & Dunn, 1995), sharing knowledge through teaching (Azmitia & Hesser, 1993; Zajonc & Markus, 1975), or engaging in conflict and resolution (Dunn, 1994; Foote & Holmes-Lonergan, 2003). Our focus on younger siblings, however, revealed that the firstborn’s experience of the arrival of a younger sibling before the second birthday did not provide a similar advantage. The first 2 years of life represent an important time in ToM development, when evidence for consciousness, pretense, and the use of lexical terms for mental states emerges (Astington, Harris, & Olson, 1988; Bartsch & Wellman, 1995). Transition to siblinghood during this time may disrupt this process. Future work should examine multiple factors in family interactions that explain differential effects of early- and late-arriving siblings on the oldest child’s social cognitive development. Children who experienced socioeconomic adversity performed less well on the second-order task; however, this association did not remain significant when accounting for age and verbal IQ. This finding stands in contrast to previous research (Cole & Mitchell, 2000; Cutting & Dunn, 1999), perhaps because our study took into account a number of dimensions of sociodemographic risk beyond occupational class or income. Although a number of sociodemographic risk factors have been found to be associated with ToM, such as income, maternal education (Andersson, Sommerfelt, Sonnander, & Ahlsten, 1996), and parental occupational class (Cutting & Dunn, 1999), rarely are these factors all controlled in a single study (Pears & Moses, 2003). Although the effects reported in this study were not large, it is important to note that the sample size used in the current study provided sufficient power to enable detection of such small to moderate effects. Thus, the absence of an association with children’s executive function abilities in this sample is noteworthy given that there was sufficient power to detect such an effect. Although executive function abilities and first-order ToM have been found to be positively related (Carlson, Moses, & Breton, 2002), a finding replicated in the current study with respect to working memory in particular, there has not been consistent evidence for a correlation between executive function and second-order ToM (for a review, see Miller, 2009). Indeed, executive function has been found to be positively associated with second-order false belief when age was controlled, but not when language ability was controlled (Hasselhorn et al., 2005). Alternatively, it is possible that the nonverbal measures used in this study to assess executive function might not be comparable to other verbal measures of inhibition and working memory such as Bear/Dragon, “Simon Says”–type inhibition tasks or word/digit span working memory tasks (Carlson et al., 2002). Before a more definitive conclusion can be made, replication of this finding using other executive function tasks is warranted. In light of previous research suggesting that some 6-year-olds and the majority of 7-year-olds are successful at attributing second-order beliefs (Perner & Wimmer, 1985), it is noteworthy that only a minority of children in this community sample passed the second-order task. This finding must be interpreted with some caution in view of the limitations of our study procedures. Data collection took place in the family homes; therefore the assessment may have been influenced by distractions within the home environment. However, evidence from this representative community sample may provide a more accurate estimate of the number of children at this age who understand second-order false belief. Finally, given our sampling strategy where we recruited firstborn children, we are unable to determine whether our findings were driven by a general sibling effect, not just the influence of younger siblings. Therefore, more work is needed to determine whether older siblings, as well as younger siblings, continue to foster children’s understanding of minds into middle childhood. In conclusion, the finding that the presence of a younger sibling in the home facilitated the firstborn’s false belief understanding draws attention to the unique contribution of the sibling relationship to social cognitive development during middle childhood. Taken together with evidence from the vast literature on first-order false belief understanding, our findings contribute to knowledge about the influence of both younger and older siblings on a child’s development of a ToM during the middle childhood years."],["Research shows that people infer the time of their actions and decisions from their consequences. We asked how people know how much time to subtract from consequences in order to infer their actions and decisions. They could either subtract a fixed, default, time from consequences, or learn from experience how much time to subtract in each situation. In two experiments, participants’ actions were followed by a tone, which was presented either immediately or after a delay. In Experiment 1, participants estimated the time of their actions; in Experiment 2, the time of their decisions to act. Both actions and decisions were judged to occur sooner or later as a function of whether consequences were immediate or delayed. Estimations tended to be shifted toward their consequences, but in some cases they were shifted away from them. Most importantly, in all cases participants learned progressively to adjust their estimations with experience. --------------------------------------------------------------------------------","A famous experiment by Benjamin Libet and his colleagues (Libet, Gleason, Wright, & Pearl, 1983) used an oscilloscope clock with a dot rotating around a sphere and asked their experimental participants to perform a quick movement of their finger or wrist at any time they felt they wished to. This procedure was repeated for a number of trials, and at the end of some of those trials participants were asked to indicate on the clock the position of the dot at the exact moment of their “conscious awareness of wanting to move”. In addition, electroencephalographic (EEG) activity was recorded in order to assess the exact time at which the readiness potential (an electroencephalographic component that precedes the initiation of voluntary actions) occurred. Contrary to what intuition would suggest, the sequence for self-initiated actions that was recorded did not start with the participants’ reporting their conscious will to perform the action. Instead, the readiness potential was recorded first in the temporal chain (around 550 ms before the action), with participants reporting their conscious will as occurring much later in time (around only 200 ms before the initiation of the action). This seems to suggest that voluntary actions are initiated unconsciously at a neuronal level, and that the conscious decision to act occurs much later in time (about 350 ms later). This experiment, along with several others that confirmed, extended, or discussed the initial findings (Banks & Isham, 2009; Guggisberg & Mottaz, 2013; Haggard & Eimer, 1999; Rigoni, Brass, Roger, Vidal, & Sartori, 2013; Rigoni, Brass, & Sartori, 2010; Soon, Brass, Heinze, & Haynes, 2008) have fueled some rather heated debates about consciousness and free will in psychology, law, and related disciplines (Alquist, Ainsworth, & Baumeister, 2013; Baumeister, Masicampo, & DeWall, 2009; Guggisberg & Mottaz, 2013; Moore, 2016; Rigoni, Kühn, Gaudino, Sartori, & Brass, 2012; Vohs & Schooler, 2008). Libet’s clock procedure has also been used in many experiments studying not only how people become aware of their own decisions but also of their own actions (Banks & Isham, 2009; Caspar & Cleeremans, 2015; Haggard, Clark, & Kalogeras, 2002; Moore, Lagnado, Deal, & Haggard, 2009; Rigoni et al., 2010; Tobias-Webb et al., 2017). A standardized, open source version of Libet’s clock has recently been published (Garaizar, Cubillas, & Matute, 2016), which should facilitate reproducibility of these findings. In spite of the many experiments that have pursued this line of research, there are still many questions that remain unanswered. Perhaps the most obvious problem is that it is impossible to assess directly the moment at which conscious will actually occurs. In particular, experimental participants do not have direct access to the time of their conscious decisions (Banks & Isham, 2009; Fahle, Stemmler, & Spang, 2011; Guggisberg & Mottaz, 2013; Rigoni et al., 2010). Moreover, we do not even know whether the temporal order in which the events are recorded in those experiments is the temporal order in which they actually occur. It might well be that people are poor at estimating the time of their conscious decision, so that perhaps they simply misjudge this time when asked to report it. Indeed, the only information that Libet et al.’s (1983) experiment (or, to our knowledge, any other experiment) can reveal in relation to the time of conscious will is the subjective report of when participants estimate that it took place. Therefore, one crucial question is how people estimate the time of their decision. One possibility is that they directly perceive the moment at which they decide to act, whilst another is that they infer retrospectively the time at which they decided to act. Preliminary evidence for this latter possibility exists in the literature. For example, Banks and Isham (2009) added auditory feedback (i.e. a tone) to the participants’ action and manipulated the time intervals between action and feedback. If people perceived the time of their decisions directly, then the time at which the tone occurred after the action should be irrelevant. However, their participants reported that their conscious decision to act occurred sooner or later as a function of whether the tone occurred sooner or later after the action. In other words, participants appeared to retrospectively infer, rather than perceive, the time at which they decided to act. Banks and Isham concluded that the generation of the action was largely unconscious, with participants inferring the time of their decision from observable cues, particularly from the apparent time of their own actions (see also Lau, Rogers, & Passingham, 2007; Rigoni et al., 2010). But assuming that the time of decision is inferred from the estimated time of action leads us to another crucial question: How do people estimate the moment of their actions? Again, they might directly perceive their own actions or they might infer the occurrence of their actions from their consequences. This has been a central question in Experimental Psychology since the nineteenth century. According to the Ideomotor Theory proposed by James (1890), people use the sensorimotor properties of the action in order to estimate when they performed the act. It is easy to extend this logic to assume that people use not only the sensorimotor properties of their action, but also any other consequences of their action, such as a tone — which an experimenter might introduce into the situation — to infer when they have acted. Interestingly, this line of research has shown that when the consequences of the action (e.g. a tone) are delayed, then people estimate that their action occurred later in time (Banks & Isham, 2011; Haering & Kiesel, 2014; Haering & Kiesel, 2015; Wegner, 2002). Moreover, a very similar finding has been consistently reported using the so-called “action-effect binding” paradigm, which shows that actions and their consequences tend to be perceived as closer to each other in time than they actually are (Haggard et al., 2002). According to Haggard and his colleagues, the intentionality of the action is the crucial factor in producing this effect, but it has also been suggested that intentionality is not always necessary (e.g., Isham, Banks, Ekstrom, & Stern, 2011), and that it might be the causal relationship between the action and the outcome, rather than its intentionality, that is critical (Bechlivanidis & Lagnado, 2013; Buehner, 2012; Buehner & Humphreys, 2009). In any case, the finding that the estimated time of action is affected by the time of its consequences and it is shifted toward them is a robust result that has been reported in many different experiments (Buehner, 2012; Cravo, Haddad, Claessens, & Baldo, 2013; Garaizar et al., 2016; Haering & Kiesel, 2015; Haggard et al., 2002; Moore & Haggard, 2008; Nolden, Haering, & Kiesel, 2012). Thus, it seems that it is not only the case that people cannot directly perceive the time of their decisions, they also seem to have difficulties in estimating the time of their own actions. Indeed, the estimation of the time of decisions and the time of actions appears to vary in the same manner and they both seem to be strongly influenced by the time at which the consequences occur (see also Banks & Isham, 2011). The most basic example of such consequences are sensorimotor consequences (such as visual and haptic cues that signal, for example, that a button has been pressed), but additional consequences such as added auditory stimuli can also be signals that the action (and decision) has occurred. In order to infer the time of their actions, people appear to discount some time from the time at which the consequences occur, and then, they possibly discount some extra time from the apparent time of action in order to infer the time of decisions. But then the question is how they know how much time they need to subtract in each situation. Do people simply subtract some fixed, default, time, which is identical in all cases, or do they learn from experience how much time they need to subtract in each particular case? In daily life, it seems reasonable to expect that people will constantly need to learn how much time to discount in each particular situation. People may easily learn that a light of a room will turn on immediately after they press the correct switch, but they should as readily learn to predict a long delay for tomatoes to grow after they have planted the seeds. Indeed, it has been shown that humans and other animals have expectations of when consequences should occur after the potential cause has occurred, and these expectations affect their learning and behavior (e.g., Arcediano, Escobar, & Miller, 2003; Buehner & May 2002; Buehner & May 2003; Buehner & May 2004; Matzel, Held, & Miller, 1988; Miller & Barnet, 1993; Savastano & Miller, 1998). Quite possibly, these expectations can be learned through experience and are continuously adjusted for each action-outcome pair and context. By the same reasoning, people will possibly learn to assume that in some cases their action occurred immediately before its consequences whilst in other cases they might learn that their actions (and decisions) took place a long time before they observed the consequences. The present study will test whether people learn to infer the time of their actions and decisions or if, in contrast, they just subtract a fixed time upon observing their consequences. Although the majority of research in the area of temporal binding implicitly assumes that learning plays an important role, we are aware of very few investigations that explicitly examine such learning. Indeed, most experiments report the dependent variable (e.g., subjective judgments or estimations of when the action occurred) averaged over the total number of trials (Haggard et al., 2002; Moore et al., 2009), and so there is typically a lack of information on what happens on a trial-by-trial basis during the course of learning. Thus, the possibility exists that action-effect binding might take place during the early trials due to pre-experimental biases, after which people might progressively learn to correct their error and improve the accuracy of their estimated time of actions (and perhaps also their decisions). Alternatively, it is possible that estimations do not improve with experience and instead show a fixed binding effect throughout the experimental session. This might occur, for instance, if actions were always, by default, subjectively shifted toward their effects. This is what most reports on temporal binding seem to suggest, given that they typically provide only one (averaged) binding value for each group in each experimental session. We are aware of only two studies looking at learning in the binding tradition. One of them was reported by Cravo et al. (2013). They used three different intervals (250, 300, 350 ms) between the action and the outcome (i.e. Action condition), and between an external event and the outcome (i.e., No Action condition). They observed, first, that the mean reported interval between the action and the outcome was shorter than between the external event and the outcome, an effect that was evident from the very early trials. According to the authors, this implies that action-outcome binding effects are due to people expecting short intervals between actions and outcomes from the very first trials, suggesting that pre-experimental biases must be important. In addition, they observed no differences between the early and later trials when considering the whole experimental session. Thus, they concluded that previous biases are probably more important than learning in producing the action-outcome binding effect. Importantly, however, they also noted that some learning was taking place within each of the consecutive and independent training blocks of trials in which they had subdivided the training phase. Thus, although their experiment represents a good starting point, there are a number of reasons why we are unable to extract any firm conclusions about learning from this study. First, the fact that learning seemed to take place within each block but no differences were observed between early and later trials throughout the session suggests that learning might have occurred but was eroded each time a new training block started. If this were the case, participants might not have benefited from the entire learning session and might have started anew each time a new block of trials was presented. Thus, it is possible that the learning phase might have just been too short to observe any appreciable differences. In addition, there were two different conditions and three different intervals in each block (with a total of 20 trials for each interval in each block), which might also have made the problem too difficult to be learned in the short time allocated to each block. A potential way to solve these problems might be to make the task easier and to provide a larger number of learning trials, that is, to ensure a more effective learning experience. This is one strategy that we will follow to examine the course of learning in the present research. The other study we are aware of was reported by Moore et al. (2009). They tested whether contingency learning played a role in the binding effect. Unfortunately, however, they used a complex design which did not allow them to observe the course of learning. As they put it: “we averaged our estimates across several trials because of the high variability of human timing performance. Therefore, we could not measure the time-course of the learning process but we can infer that causal learning occurs based on our contingency effects” (p. 282–283). This study also suggests that reducing the difficulty of the task might possibly be the first step needed to study the learning curve. The present research will take such steps in order to test whether people progressively learn to estimate the time of their actions and decisions with experience.","In order to facilitate the observation of a learning curve, we will use only one continuous training session (i.e. blocks of trials will constitute a continuous and incremental training session, rather than independent learning experiences). In addition, we will use only two different conditions during training (i.e. immediate feedback vs delayed feedback). A third condition (i.e. no-feedback) will be added as a separate baseline phase to assess judgments in the absence of feedback, but it will be presented only after the training phase has already been completed, so that baseline assessment could not interfere with the critical learning experience. Ethics statement ~~~~~~~~~~~~~~~~ The computer program informed participants that their involvement was voluntary and anonymous. We did not ask participants for any data that could compromise their privacy, nor did we use cookies or software in order to obtain such data. The stimuli and materials were harmless and emotionally neutral, the goal of the study was transparent, and the task involved no deception. The ethical review board of the University of Deusto examined and approved the procedure used in this research, and the two experiments were conducted in accordance with the approved guidelines. Data selection criterion ~~~~~~~~~~~~~~~~~~~~~~~~ Data from participants who did not press the key more than 25% of trials on either one of the three feedback conditions described above (i.e., immediate, delayed, and no-feedback) were discarded.","The purpose of this experiment was twofold. First, we tested whether people estimate the time of their actions accurately or whether their estimate is affected by the time of consequences. We expected that a delay in the consequences of the action should produce a delay in the estimated time of action, thereby reproducing previous reports (e.g., Banks & Isham, 2011; Garaizar et al., 2016; Haggard et al., 2002). Most importantly, this experiment aimed to extend previous findings by assessing whether the estimation of the time of action was subject to learning. That is, we tested whether people will then progressively learn to correct their estimation errors and improve their reported time of action with experience.","The sample was comprised of 60 Psychology students from the University of Deusto who volunteered to take part in the experiment in exchange for academic credit. Data from four of the participants were discarded according to the data selection criterion described above.","Participants performed the task on personal computers in a large computer room. They were seated approximately 1 m apart from each other. We developed a Visual Basic computer program based on Garaizar et al.’s (2016) open source HTML version of Libet’s clock. The computers run the program on Microsoft Windows 7. A screen capture of this version is shown in Fig. 1. Its accuracy and settings were validated using the identical methodology published by Garaizar et al. Auditory stimuli were presented via headphones connected through the audio output of the PCs. Procedure and design The procedure and design were identical to those described by Garaizar et al. (2016). On each trial, participants observed a clock face on the computer screen with a dot rotating around the sphere at a constant speed of 2560 ms per cycle. There were two cycles of the dot per trial. Participants were asked to look at the center of the clock during the first cycle within each trial, and to press the space bar anytime they wished to during the second cycle on each trial. Pressing the bar did not stop the rotating dot, but was followed by auditory feedback at different delays as described below. We used a within-subject design with two types of trials during training. On half of the trials (i.e., immediate condition) a 1000 Hz 200 ms tone was programmed to occur immediately following the registration of the action, using Microsoft Windows Multimedia Timers,1 which have a resolution of 1 ms as assessed by popular experimental software packages such as SuperLab (Abboud, Schultz, & Zeitlin, 2006) and E-Prime (Schneider, Eschman, & Zuccolotto, 2002). On the other half of trials (i.e., delayed condition), the same tone was programmed to occur with a 500 ms delay after the registration of the action. There were 40 trials for each condition, and they were presented in pseudorandom order. Immediately after each one of the 80 trials was completed, participants observed an empty clock face and were asked to indicate the location at which the dot was when they pressed the space bar during that trial (see Fig. 1). This was their subjective judgment of their time of action. Participants answered this question using the mouse, after which the empty clock face was presented again and the inter-trial interval (ITI) began. The ITI lasted between 1000 and 3000 ms. At the beginning of the ITI, a text prompting participants to be prepared for the next trial was presented along with a complex 1000 ms tone (500 ms at 250 Hz followed by 500 ms at 440 Hz). Then, after a variable interval of 0–2000 ms the next trial began. Finally, after all 80 trials had been completed a screen instructed participants that on the following trials no auditory feedback would be presented. Participants were then presented with 20 additional trials without auditory feedback in order to assess their baseline judgments without the provision of feedback (i.e., baseline condition).","For comparison purposes, we will first describe our results as they have typically been described in previous research. Thus, we will first report the subjective estimates for the time of action averaged across the training session. Mean estimated time of action is shown in Fig. 2 against the actual recorded time for action (which is represented as zero on the Y axis). A positive value means that the action was estimated on average after it already had taken place, whereas a negative value means that the action was estimated on average before it had taken place. In addition, it should be noted that baseline judgments are often subtracted from target judgments in the literature and the data are reported after this transformation has already been applied (e.g., Moore et al., 2009; Tobias-Webb et al., 2017). However, in order to facilitate potential comparisons, we decided to depict the experimental trials and the baseline trials separately (rather than subtract baseline data from target action-outcome trials). Thus, Fig. 2 shows the raw average estimates for the time of action (in ms from the actual time of action) for both the action-outcome trials and the baseline trials separately. Means (and standard errors of the mean) are −11.16 (4.97), 81.09 (17.98), and −8.45 (8.15) for the immediate, delayed and baseline conditions, respectively. The subjective estimation of the time of action in the delayed condition was also significantly longer than the time at which the action actually occurred (i.e. zero on the Y axis of Fig. 2), t(55) = 4.508, p < 0.001, dz = 0.602. Interestingly, there was also a discrepancy between the estimated time of the action and the occurrence of the action in the immediate condition, t(55) = 2.246, p = 0.029, dz = 0.3. In those cases, however, participants reported that their action occurred before, rather than after, the actual time of the action. We will use the term antibinding to refer to this effect that occurs in the immediate condition, since the estimated time of action was shifted away from the outcome. No significant differences were observed during baseline trials between the estimated time of action and actual time of action, t(55) = 1.037, p = 0.304, dz = 0.138. Subsequent planned comparisons showed that in the delayed condition there was a significant difference between the first and the last block of trials, which suggests that learning occurred and participants improved the accuracy of their estimations through the training session, t(55) = 5.054, p < 0.001, dz = 0.675. Also, and even though learning and improvement are evident in the delayed condition, the participants’ estimations are significantly different from the actual time of action, even on the fourth block of trials, t(55) = 6.511, p < 0.001, dz = 0.87, for the first block; t(55) = 2.637, p = 0.011, dz = 0.352, for the fourth block. The learning curve in Fig. 3 suggests that the accuracy of judgments could still improve if a more extended learning experience were provided. The immediate condition also shows the significant difference between the first and the fourth block which reflects the occurrence of learning, t(55) = 2.409, p = 0.019, dz = 0.321. In this condition, however, the estimation was at first quite accurate and did not differ from zero (i.e. the actual time of action) on the first block of trials, t(55) = 0.247, p = 0.785, dz = 0.033, but as learning proceeded, estimations began to shift away from the outcome (and thus from the action as well) so that the antibinding effect become evident, and by the fourth block participants came to believe that they had acted significantly earlier than they had, t(55) = 3.035, p = 0.004, dz = 0.405. Thus, the observed antibinding effect seems to be a learning effect. The estimations of the time of action during the baseline phase did not significantly differ between the first and the last block of trials, t(55) = 0.294, p = 0.77. Moreover, no significant differences were observed between the estimated time of action and the actual time of action during baseline, t(55) = 0.896, p = 0.374, dz = 0.119, for the first baseline block of trials, t(55) = 0.7, p = 0.487, dz = 0.093 for the last block. That is, the time of action appears to be accurately judged throughout the baseline phase.","Experiment 1 showed that participants inferred the time of their actions from the consequences of those actions, so that when the consequences were delayed participants estimated that they had acted later than they had. Most importantly, participants learned progressively to improve their estimations of their action on the delayed trials, possibly by subtracting some time from the time at which the auditory feedback was presented. However, as they did so, they also began to subtract some time from the time of feedback in the immediate condition, so they also began to assume that their action occurred earlier than it did when feedback was immediate. Our next purpose was to explore whether those effects would also occur if we asked participants to estimate the time of their conscious decision to act rather than the time of their actions. If the subjective time of action can be so easily mislead by the time in which the consequences of the action occur, then the subjective time of decisions should also be affected. As has been shown previously, participants should believe that their decision took place sooner or later as a function of whether the auditory feedback for their actions occurred sooner or later (Banks & Isham, 2009). This implies that people cannot be certain of when they made a decision. Instead, they possibly infer their time of decision upon observing the consequences of their behavior. Investigating whether people infer a fixed time for decisions as a function of when the outcome occurs or whether, in contrast, this time is learned and modifiable through experience is the aim of Experiment 2.","The sample was comprised of 56 Psychology students from the University of Deusto, who volunteered to take part in the experiment in exchange for academic credit. Data from seven of them were discarded according to the data selection criterion described for Experiment 1. Design, apparatus, and procedure All aspects of the experimental design and procedure are identical to those of Experiment 1 except that we asked participants to report the time of their decision to act rather than the time of their action.","As in Experiment 1, we will first report the subjective judgments averaged throughout the experimental session. In the present experiment, these judgments reflect the estimated time of decision rather than action. Again, in order to facilitate potential comparisons, we decided to depict the raw data from the experimental trials and the baseline trials separately. Fig. 4 shows the raw average estimates for the time of decision in ms from the recorded time of action (which is shown as point zero on the Y axis) for both the action- outcome trials and the baseline trials. Means (and standard error of the mean) are −11.33 (6.98), 27.92 (9.25), and −24.06 (6.90) for the immediate, delayed and baseline conditions, respectively. As can be observed in Fig. 4, when the action was immediately followed by feedback participants estimated that their decision had occurred slightly before their action, which seems to be a sensible inference. However, when feedback following the action was delayed, participants then moved their estimated time of decision toward the tone. This replicates previous findings (Banks & Isham, 2009) and can also be regarded as an extension of the standard binding effect, as it shows that the shift of the subjective estimation toward the outcome occurs not only for actions but also for decisions. It is interesting to note that the estimated time of decision in the delayed condition was also significantly delayed when compared with the time at which the action actually occurred (i.e. zero on the Y axis of Fig. 4), t(48) = 3.017, p = 0.004, dz = 0.431. We refer to this finding as a superbinding effect, as participants came to believe that their decision occurred after they had already acted, rather than before they acted. Thus, the mean estimation of the time of decision was shifted toward the tone, and beyond the action. This suggests that when estimating the time of their decisions, participants were more strongly influenced by the observation of consequences, which probably signaled the apparent time of action, that than by their actual decisions (or their actual actions). There was also a significant difference between the estimated time of decision and the actual occurrence of the action on the baseline trials, t(48) = 3.486, p = 0.001, dz = 0.498 (see Fig. 4). In those cases, participants tended to report that their decision occurred before the action, which is a sensible inference. This temporal priority of the reported conscious decision over the actual action was not observed in the immediate condition, t(48) = 1.623, p = 0.111, dz = 0.231, perhaps due to the fact that the delayed condition was presented during the same block of trials. (Nevertheless, recall that these are averaged data for the entire session; the analyses of the learning curve are described below and show a different and interesting pattern). The most critical and newest results in this experiment are shown in Fig. 5. It shows the learning curves for the subjective timing of conscious decision throughout the experimental session, depicted in blocks of 10 trials. The learning curves are very similar to those observed in Experiment 1, which suggests that the course of learning is very similar for actions and decisions. Although the binding effect for decisions in this experiment appears to be less pronounced than the effect for actions found in Experiment 1, it should be noted that the estimated time of decision should be located, at least in principle, before the action. Thus, the fact that the binding effect is still present, that is, decisions are still shifted toward the outcome and beyond the action, suggests that the binding effect for decisions is actually strong. As already mentioned, we are referring to it as a superbinding effect. Particularly during the early trials, participants seem to infer that their decisions occurred near the tone, and after their actions had already occurred. Quite possibly, they misjudged the timing of their actions, as was evident in Experiment 1, and therefore they also misjudged the location of their decisions. Then, as training proceeds, the participants appear to progressively learn that their decisions must have occurred earlier than they thought, so they gradually reduce the superbinding effect. Interestingly, in the delayed condition decisions are judged to occur much later than the actual actions during the first block of trials, t(48) = 4.089, p < 0.001, dz = 0.584, but this superbinding effect is then gradually corrected and by the fourth block of trials decisions are no longer judged to occur after their actions, t(48) = 0.186, p = 0.853, dz = 0.026. Nevertheless, by the end of training decisions are not yet judged as preceding the action. It is quite possible that the estimation of the action is still suffering from some degree of binding by the end of the training session, which suggests that, like in Experiment 1, the accuracy of judgments could still improve if a more extended learning experience were provided. In the immediate condition, however, the estimated time of decisions did not differ from the actual time of action in the first block of trials, t(48) = 0.625, p = 0.535, dz = 0.089, but by the fourth block participants had learned to locate their decisions significantly before their actions, t(48) = 5.27, p < 0.001, dz = 0.752. This suggests that, as should be expected, the immediate condition was easier for them to learn. The estimations of the time of decision during the baseline trials did not significantly differ between the first and the last block of trials, t(48) = 0.752, p = 0.456, dz = 0.107, and thus no learning was observed during baseline. Estimated time of decision was significantly lower than zero (see Fig. 5) from the very first block of trials, t(48) = 3.381, p = 0.001, dz = 0.483, and remained so through the last block, t(48) = 3.097, p = 0.003, dz = 0.442. That is, participants inferred that their decisions occurred before their actions throughout the baseline phase.","The two experiments reported here suggest that participants infer the time of their actions and decisions from their consequences. By manipulating the time of those consequences we were able to change both the time at which they believed they had acted (Experiment 1) and the time they believed they had decided to act (Experiment 2). The results of Experiment 1 extend previous reports in the action-outcome binding literature (Garaizar et al., 2016; Haering & Kiesel, 2012; Haering & Kiesel, 2015; Haggard et al., 2002; Moore, 2016); and the results of Experiment 2 extend previous reports in the conscious will literature (Banks & Isham, 2009; Libet et al., 1983; Wegner, 2002). Taken together, the two experiments suggest that the subjective estimates of the time of action and decision are governed by similar processes. That is, both the subjective estimates of the time of actions and decisions were affected by their consequences and tended to be shifted toward them. Most importantly, we observed that estimations of the time of action and decision were subject to learning and were progressively adjusted as learning proceeded. We believe this is the first time that a learning curve has been reported to show how people estimate the time of their actions and decisions. Below we elaborate on these findings. Learning, antibinding, and superbinding ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The two experiments showed that inferring the timing of decisions and actions is subject to learning, since the participants were continuously adjusting their timing inferences. We believe this is the first time that a clear learning curve has been shown for the estimation of actions and decisions. Action-outcome binding effects found in previous studies were replicated in the current experiments for both actions and decisions, and not only when judgments were averaged across the experimental session (i.e. as they are typically reported; see Figs. 2 and 4), but were also observed during the early training trials (see Figs. 3 and 5). Importantly, however, those binding effects were progressively corrected as learning proceeded (see Figs. 3 and 5). In some cases the binding effects vanished for both actions and decisions as learning proceeded whilst in other cases this learning and adjusting process yielded what we have called an antibinding effect: The reported time of action tended to be progressively shifted away from the outcome, rather than toward it, and in the immediate condition the action ended up being judged as occurring significantly earlier than the actual action. Interestingly, as participants learned to improve their estimations in the delayed feedback condition, they also overcompensated (erroneously) their subjective inferences in the immediate condition. Thus, this antibinding effect is also the result of learning. Seemingly, participants learned that some time interval needed to be subtracted from the apparent time of action (which was strongly influenced by the time of feedback), and so they subtracted this time not only in the delayed condition but also in the immediate condition, thereby producing the antibinding effect on the immediate trials. Thus, this antibinding effect has possibly been favored by the fact that various intervals were presented in the same block, so that what was learned for one of them implied an error when applied to the other one. It could be argued that the antibinding effect might simply be reflecting that the computer program is registering the time in which the action is completed whereas the participants might be reporting the time at which they initiated the action. It is true that the time at which an action is initiated should differ slightly from the time at which it is completed and registered, but even if participants were using the initiation of the action as the assessment point, we believe that this could not explain the observed effects. During the early trials the participants’ estimate of their time of action in the immediate condition does not differ from the actual time of action recorded by the computer. It is only with time and experience that participants start to locate their actions away from the outcome (and from the actual action), and this happens only in the immediate condition. In the delayed condition the estimated time of action is shifted toward the outcome. If antibinding were simply due to participants reporting the time of the initiation of their action, then the learning curves should be identical (and flat) in the immediate and the delayed conditions. We believe the interpretation in terms of learning is more plausible. Indeed, effects similar to these ones have also been reported in cases in which, instead of being trained on several delays simultaneously, participants are trained with one delay, so that they learn to expect a certain action-outcome interval, and then they are tested with different delays in a subsequent test phase. When tested with shorter delays, participants feel that the outcome occurs before their action, and sometimes they even feel that they have not caused the outcome, because it occurred too early for this to be the case (Haering & Kiesel, 2012; Haering & Kiesel, 2015; Heron, Hanson, & Whitaker, 2009; Stetson, Cui, Montague, & Eagleman, 2006). Moreover, in some cases participants may even distort their perception of temporal order in order to make it consistent with their causal beliefs. If the outcome occurs slightly before the action, participants will not realize that it does. They may even reorder the sequence of events so that they come to believe that the events occurred in the “correct” order, that is, action before effect (Bechlivanidis & Lagnado, 2013; Bechlivanidis & Lagnado, 2016; Desantis, Roussel, & Waszak, 2011; Heron et al., 2009). One potential problem of any experiment using Libet’s clock and related procedures is that participants could preplan their actions. Then the delayed feedback might be taken as indication of their inaccuracy to press the spacebar exactly at the planned time during the early trials, which might be the reason why they typically believe they acted later than they did. Note, however, that even though preplanning is certainly possible, participants would still be taking the tone as a signal of when the action (and decision) was performed, so we believe that the possibility of preplanning does not undermine our main proposal that the time of actions and decisions is inferred from consequences. Moreover, the role of learning is still evident in the learning curve and shows that participants learn to gradually correct their estimation error (regardless of whether their action was planned), and they gradually learn to give more weigh to cues that are better signals of the action (such as, for instance, haptic cues) while they progressively give less weight to cues such as auditory feedback, which in this case is a totally unreliable cue. Also of interest is the effect observed in Experiment 2 with the estimated time of decision as the dependent variable (Experiment 2). In this case, participants shifted their estimated time of decision toward the outcome, as we expected, but quite surprisingly, they even shifted it beyond the action. That is, they reported that their decision had occurred after they had already acted. Given that the subjective time of decision should reside, at least in principle, before the action, we are using the term superbinding to refer to this effect. Importantly, just as participants learned to reduce their binding effect for the time of actions in Experiment 1, they also learned to reduce this superbinding effect for the time of decisions in Experiment 2. Thus, superbinding is also subject to learning. In the immediate condition participants learned in just a few trials to locate their decisions before their actions. In the delayed condition, however, they needed more time to progressively locate their decisions away from the outcomes and closer to their actions. Even so, by the end of the session they were still estimating that they had decided after they had already acted. This finding adds support to the idea that the time of decision is inferred from the apparent time of action, which in turn is strongly affected by the time of its consequences. Nevertheless, the shape of the learning curve suggests that if training were extended, then participants would probably learn to estimate their decisions as occurring before their actions, just as they do in the immediate and in the baseline conditions. In other words, this superbinding effect should probably vanish with a more prolonged learning experience. What this experiment clearly shows is that decisions are not directly perceived, nor are they inferred to occur at a fixed interval prior to actions. Instead, participants progressively learn to locate their decisions before their actions. How do people learn to adjust their inferences for actions and decisions? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ It is worth noting that with respect to decisions (i.e. Experiment 2) there are no objective parameters to learn from, that is, no parameters that require adjustment. In other words, there is nothing that can be taken as an objective cue for learning about the actual time of decisions. In principle, if no cues exist that can facilitate learning, then learning should not take place, and adjustments would not make sense. Even so, participants in this experiment are learning to locate their decisions before their actions, and they are modifying their estimates with experience. We believe that they are quite possibly learning to improve their estimations of the time of actions, as shown in Experiment 1, and then they progressively learn to also subtract some time from the estimated time of action, probably assuming that decisions must precede actions. As already noted, the action is possibly the most reliable cue that can be used to infer that a decision has been made and therefore the question of how people learn to infer their decisions can be substituted, at least in part, by the question of how people learn to infer their actions. James (1890) Ideomotor Theory posits that associations are formed between actions and effects, and once they are formed the action will prime the mental representation of the effect, just as the observation of the effect will also activate the mental representation of the action. Consistent with this view, Ebbinghaus (1885) had also suggested that associations between mental representations of events were bidirectional, so that each event could activate the representation of the other one in either direction. Although early 20th century research was not conclusive about the bi-directionality of associations and the idea was generally abandoned (see Asch & Ebenholtz, 1962; Ekstrand, 1966, for review), more recent research has shown that associations can indeed activate the mental representation of either event upon observation of the other one (Arcediano et al., 2003; Arcediano, Escobar, & Miller, 2005; Dignath, Pfister, Eder, Kiesel, & Kunde, 2014; Gerolin & Matute, 1999). Recent extensions of James’ Ideomotor Theory (Dignath et al., 2014; Elsner & Hommel, 2001), as well as the Temporal Coding Hypothesis developed by Miller and his colleagues in the domain of associative learning with animals (Matzel et al., 1988; Savastano & Miller, 1998) and humans (Arcediano et al., 2003; Arcediano et al., 2005), suggest that associations encode not only the representation of the associated events but also their timing relationship. That is, if a given temporal relationship exists between the two associated events, each of these events will prompt the mental representation of the other at the time when it should occur or should have occurred (Dignath et al., 2014; Matzel et al., 1988; Savastano & Miller, 1998). This view predicts that whenever an action is performed its effect will be predicted based on previous learning; and that whenever the effect is observed, the occurrence of the action will be inferred at its usual time. This proposal is also consistent with the idea that action- effect binding might be the result of people using the consequences of their actions as temporal markers in order to reconstruct the time of their actions (and thus their decisions as well) (Banks & Isham, 2011; Friedman, 1990). These two processes (predicting the second event from the first and inferring the first event from the second) have often been regarded in the literature as two alternative hypotheses, but they are perfectly compatible. Indeed, the two of them have been shown to play a critical role, and it has been suggested that there is both a predictive and a postdictive (inferential) component in the subjective estimation of actions and in the action-outcome binding effect (Chambon, Sidarus, & Haggard, 2014; Moore & Haggard, 2008; Moore et al., 2009; Wenke, Fleming, & Haggard, 2010). We do not want to mean that these predictions and inferences need be perfectly explicit, conscious, and verbalized. However, because the action is associated with its consequence (which includes timing), performing the action involves the prediction (whether explicit or incidental) that its consequence will occur at its usual time. Additionally, because of the bidirectionality of the association, the representation of the action (including its time of occurrence) can also be backwardly activated upon observation of its consequence. All of the above suggests that continuous adaptation (i.e. learning) takes place when people infer the occurrence of their actions and decisions from consequences, a view that is strongly supported by the present findings. According to the most widely accepted theories of associative learning, learning consists of reducing prediction errors (Rescorla & Wagner, 1972). Combining traditional learning theories with the proposals described above concerning temporal coding and bi-directionality of associations (Dignath et al., 2014; Matzel et al., 1988; Savastano & Miller, 1998), it seems reasonable to expect that when participants perform an action they will automatically predict (whether explicitly or not) that a consequence will occur at a given time (as a function of their previous experience). We should also expect that if the consequence does not occur at the expected time, then learning will consist of adjusting those prediction errors until they are minimized. Moreover, from the occurrence of the consequence, people should also adjust the time at which they believe the action should have occurred, so that, if the outcome is moved forward, then the action should also have been moved forward. It is currently a matter of debate as to whether the subjective action-outcome interval becomes shortened in the action-outcome binding paradigm or whether the whole subjective interval is not shortened but instead displaced. In other words, it might be that action-outcome binding effects are due to a reduction in the subjective interval between actions and outcomes, but it could also be that when the outcome is delayed on a given trial, then the perception of the entire interval is shifted in that direction so that the estimated time of action will preserve the length and synchrony of the original interval (Heron et al., 2009). We believe our results are more consistent with this second view. The estimation of the action (and decision) tended to be shifted in the direction of the outcome or in the opposite direction as a function of whether the outcome occurred later or sooner than the time it was expected to occur. The estimated time of action and decision were also moved in the forward direction or in the backward direction as a function of whether learning was in its earlier or more advanced stages. Initially, the estimations tended to be displaced toward the outcome (which occurred surprisingly late), but then they moved away from the outcome as participants learned that the tone was not a reliable cue for the occurrence of action. Thus, it was not the case that the interval became shorter, it just moved in one or the other direction. A similar proposal has been advanced by different versions of the Comparator Theory in the area of sensorimotor research (Frith, 2005; Frith, Blakemore, & Wolpert, 2000; Miall & Wolpert, 1996; Synofzik, Vosgerau, & Newen, 2008; Wolpert, Ghahramani, & Jordan, 1995), according to which, people compute the difference between the sensory input that they expect when performing the action and the one they receive, and they use the output of this sensorimotor comparator process to realign the temporal interval (see also Heron et al., 2009). Because in the present experiments we presented the consequences of the action at two different intervals rather than just one, it was impossible for participants to accurately predict whether a tone would follow the action (and decision) immediately or whether it would be delayed. Under those uncertain circumstances, we suggest that participants are likely to have minimized those errors by encoding some averaged prototype of the intervals rather than (or in addition to) the correct ones. Participants could then use this averaged interval, both when they predict the outcome of individual trials and when they infer the time at which their action (or decision) occurs. In other words, they probably predicted the consequences of their actions at some intermediate point (e.g., 200 ms) and then, when they observed the consequences of their actions, they possibly inferred that the actions had occurred at some intermediate time before the consequences (e.g., 200 ms). We believe this might be important in the development of the antibinding effect; on trials in which feedback occurred immediately after the action, participants would still infer that their action should have occurred earlier, thereby producing the antibinding effect. This view assumes that the whole interval moves forward (or backward) rather than becoming shortened in binding effects (see Heron et al., 2009). Learning is, by definition, continuous, and should always correct errors. This means that one should expect that if changes are introduced in an already learned situation, then new errors could occur, and then relearning will need to take place again. Our results suggest that before starting the experiment, participants already have some expectations about the standard interval between their action and their consequences, due to their previous experience with the experimental context (typically a computer). This is consistent with Cravo et al.’s (2013) proposal that pre-experimental biases play a crucial role in the observation of binding effects with delayed outcomes, an effect that we replicated here during early trials. If those outcomes are presented later (or sooner) than expected, then participants assume that the action must have also occurred later (or sooner). Our results show that those expectancies, or pre-experimental biases, need to be learned and are continuously adjusted for each action-outcome pair. By the same reasoning, this adjustment does not always consist of shifting the estimation of the action toward the outcome (binding); sometimes it consists of shifting the estimation of the action away from the outcome (antibinding), and sometimes decisions can even be reported to occur after their actions (superbinding). In any case, and as learning proceeds, those errors tend to be corrected and people gradually learn to adjust the estimation of their actions to the occurrence of the actual actions. As they learn to adjust the estimation of their actions, they also learn to adjust the estimation of their decisions so that they appear to occur before their actions."],["This study examined the effect of word level phonological knowledge on learning to read new words in Down syndrome compared to typical development. Children were taught to read 12 nonwords, 6 of which were pre-trained on their phonology. The 16 individuals with Down syndrome aged 8-17 years were compared first to a group of 30 typically developing children aged 5-7 years matched for word reading and then to a subgroup of these children matched for decoding. There was a marginally significant effect for individuals with Down syndrome to benefit more from phonological pre-training than typically developing children matched for word reading but when compared to the decoding-matched subgroup, the two groups benefitted equally. We explain these findings in terms of partial decoding attempts being resolved by word level phonological knowledge and conclude that being familiar with the spoken form of a new word may help children when they attempt to read it. This may be particularly important for children with Down syndrome and other groups of children with weak decoding skills. © 2014 Elsevier Ltd. --------------------------------------------------------------------------------","Although there is variability between individuals, reading accuracy has been identified as a relative strength in Down syndrome in comparison to general ability and reading comprehension (e.g. Boudreau, 2002; Buckley, 1985; Cardoso-Martins & Frith, 2001; Hulme et al., 2012; Laws & Gunn, 2002; Nash & Heath, 2011). Typically individuals with Down syndrome show better reading of words than nonwords (Roch & Jarrold, 2008). Further, they tend to have poor phonological awareness (Lemons & Fuchs, 2010; Næss, Melby-Lervåg, Hulme, & Lyster, 2012), a skill that is important for decoding in typical development (e.g. Byrne & Fielding-Barnsley, 1989; Hulme, Goetz, Gooch, Adams, & Snowling, 2007; Wagner et al., 1997). This relationship is also present in Down syndrome (Fowler, Doherty, & Boynton, 1995; Roch & Jarrold, 2008); indeed a recent meta-analysis found that the deficit in nonword reading in individuals with Down syndrome is moderated by their performance on phoneme deletion tasks (Næss et al., 2012). The relative difficulties in nonword reading and phonological awareness among individuals with Down syndrome suggests that they may be recruiting compensatory strategies to support the development of word reading, such as visual skills or vocabulary knowledge (Boudreau, 2002; Buckley, 1985; Kay-Raining Bird, Cleave, & McConnell, 2000). A recent longitudinal study suggested that vocabulary is a stronger predictor of reading development in Down syndrome than in typical development (Hulme et al., 2012). Related to this, Nation and Cocksey (2009) have suggested that phonological (sound-based) aspects of vocabulary knowledge may be particularly important for learning to read when decoding is compromised. The aim of the current study was to investigate the contribution of word level phonological knowledge to reading and orthographic learning in individuals with Down syndrome. In order to study the mechanisms behind reading development in Down syndrome, it is useful to consider models of reading in typical development. The triangle model of reading (Plaut, McClelland, Seidenberg, & Patterson, 1996) proposes two ways in which the phonology of words can be activated from their orthography: directly or indirectly via semantics. Semantic information (meaning- based knowledge) is argued to be most important when the direct phonological pathway is impaired or compromised such as when reading irregular words (McKay, Davis, Savage, & Castles, 2008; Taylor, Plunkett, & Nation, 2011) or in disorders such as Down syndrome or dyslexia. Semantic knowledge is not well-specified in the triangle model and was conceptualised as additional input to the phonological representation (Plaut et al., 1996). In studies with children, semantics is typically assessed by vocabulary tasks. However, these tasks often require both phonological and semantic knowledge about a word (Johnson, Paivio, & Clark, 1996; Levelt, Roelofs, & Meyer, 1999). Nation and Cocksey (2009) devised separate measures of phonological knowledge and semantic knowledge and found that at the item-level, phonological, but not semantic, knowledge uniquely predicted variations in children's reading of irregular words. They suggested that irregular words may result in partial decoding attempts, and possession of a whole-word phonological representation allows children to resolve these attempts and produce the correct response. Such a process may well be important for learning to read in individuals who have decoding difficulties, such as individuals with Down syndrome. One method to determine how existing phonological knowledge affects reading is to conduct training studies with novel words. This controls what phonological information about a word children have been exposed to. In such studies, children are taught to read unfamiliar words or nonwords with or without pre-training. Pre-training can be used to train phonological knowledge of words or nonwords and may involve individuals hearing and saying the items (McKague, Pratt, & Johnston, 2001) or discriminating between the target items and phonologically similar distracters (Duff & Hulme, 2012). It has been found that typically developing children are more accurate when attempting to read words that have received such phonological pre- training (i.e. those they have phonological knowledge about), than those words which have not (Duff & Hulme, 2012; McKague et al., 2001). This supports the idea that knowing the phonological form of a word is causally related to children's ability to learn to read it. Training studies can also explore how children establish orthographic (written) representations of new words. According to Share's self-teaching hypothesis (Jorm & Share, 1983; Share, 1995) when a child encounters an unknown word, they independently convert the letters into sounds, a process termed phonological recoding. Multiple successful encounters with a word result in the formation of an orthographic representation. Orthographic knowledge can be tested using orthographic choice tasks that assess whether children can discriminate the correct written form of the word from a nonword with the same pronunciation but a different spelling. If familiarising the child with the phonological form of a word supports decoding, this should also promote the creation of an accurate orthographic representation. The aim of the present study was to examine the effect of phonological knowledge on written nonword learning in individuals with Down syndrome. A group of typically developing children was matched to the individuals with Down syndrome on word reading. A subgroup of the typically developing children was also individually matched to the individuals with Down syndrome on decoding. This two-stage matching procedure was used to investigate whether individuals with Down syndrome learnt to read new words at a rate commensurate with their reading or decoding level. It was predicted that individuals with Down syndrome would show written nonword learning in line with their decoding skill but below the level expected from their word reading ability. Phonological pre-training was designed to familiarise individuals with the phonological form of the nonwords prior to encountering them in print. It was expected that this would benefit the typically developing children (Duff & Hulme, 2012; McKague et al., 2001). The individuals with Down syndrome would likely have poorer decoding skills than typically developing children at the same level of reading and so make more partial decoding attempts that could be resolved through phonological knowledge. Therefore it was also predicted that individuals with Down syndrome would benefit more from phonological pre- training than the typically developing children matched for word reading but that when the groups were matched for decoding, they would benefit equally from phonological pre- training.","Sixteen children and adolescents with Down syndrome (mean age of 13;08 (SD = 2;11; range 8–17 years); five males) were recruited through families who were also taking part in a longitudinal study and all individuals completed the study. All individuals also participated in a spoken word learning study approximately 12 months previously (Mengoni, Nash & Hulme, 2013), which trained different nonwords using different methods and none of the background data were re-used in the present study. All individuals had Trisomy 21 according to parental report and were known from previous testing to have a reading age of at least five years. Parental consent was obtained for the individuals to participate in the study. Of the individuals with Down syndrome, ten were in mainstream education: five in primary schools and five in secondary schools. Four individuals were in special education: two in secondary schools and two in college. Two individuals had joint attendance at special and mainstream secondary schools. Thirty typically developing children (mean age of 6;01 (SD = 0;06; range 5–7 years); 16 males) were recruited from Year 1 and 2 classes in two primary schools. These year groups were chosen to correspond to the reading ability of the group of individuals with Down syndrome, and indeed there was no difference between the two groups on a test of single word reading (see Table 1). This matching procedure means that any differences in word learning cannot reflect overall differences in word reading ability, but it means that the typically developing group are likely to have had more limited reading experience. Consent for the typically developing children to participate was obtained from the headteachers of the schools and from their parents. Children who had been identified with special educational needs were excluded. All participants in both groups were monolingual English speakers. Assessment battery ~~~~~~~~~~~~~~~~~~ Testing for the longitudinal study in which the individuals with Down syndrome were participating took place at the same time as this study and the tasks reported below were part of the longitudinal study. It was decided to administer the same tasks to the typically developing group to keep the procedure the same and to act as filler tasks before the post-tests. The results of standardised tests are reported later to give an indication of the relative profiles of the two groups. Nonverbal reasoning The Matrices subtest from the Wechsler Pre-School and Primary Scale of Intelligence IIIUK (WPPSI-IIIUK; Wechsler, 2003; reliability coefficient .90) was administered to measure nonverbal reasoning skills. In this task, children were asked to look at an incomplete matrix and choose the missing section from four or five options. Testing was discontinued after four incorrect answers on either four or five consecutive items. Word reading The Single Word Reading Test from the York Assessment of Reading for Comprehension (YARC) Passage Reading battery (Snowling et al., 2009; reliability coefficient .98) was administered. The test consists of 60 words that increase in complexity from simple words such as ‘see’ to more complex words such as ‘pseudonym’. Individuals were shown all words and asked to read as many as they could. Nonword reading The Graded Nonword Reading Test (Snowling, Stothard, & McLean, 1996; reliability coefficient .96) was used to test children's decoding skills. Individuals were presented with nonwords that increased in difficulty from ‘hast’ to ‘sloskon’, and were asked to read them aloud. There were five practice trials and 20 test items, and the task was discontinued after six consecutive errors. Individuals were awarded one point for each nonword read correctly. Expressive vocabulary The WPPSI-IIIUK Picture Naming subtest (reliability coefficient .88) was administered to test expressive vocabulary ability. Individuals were asked to name a series of 30 pictures ranging from car to thermometer, and the test was discontinued if five consecutive incorrect responses were made. Phonological awareness To assess phoneme awareness individuals were given an Alliteration Matching task, adapted from Carroll (2004). All stimuli were presented to children as spoken words and colour pictures. Individuals were asked which word out of a choice of two started with the same sound as a target word. The distracters were matched to the correct answer for global similarity to the target word. There were two practice items and 10 test items, and individuals were administered all items. The Sound Deletion subtest from the YARC Early Reading battery (Hulme et al., 2009; reliability coefficient .93) was also administered to test phonological awareness. Individuals were presented with spoken words and corresponding colour pictures, asked to repeat the word and then asked to delete a sound. Some of the items resulted in nonwords, e.g. say sheep without the /ʃ/, whereas some items resulted in real words, e.g. say boat without the /t/. There were 12 items, which tapped deletion of syllables and phonemes in initial, medial and final positions and individuals completed all items. Verbal short-term memory The Word Recall subtest from the Working Memory Test Battery for Children (WMTB-C; Pickering & Gathercole, 2001; reliability coefficient .72) was used to measure verbal short-term memory skills. The individuals heard a sequence of words and had to repeat them in the same order. The sequence of words increased in length, across trials. The test was discontinued when individuals scored less than four out of six items correct at a given list length. The number of correct trials, rather than span score, was used in analyses as this score had a greater range and therefore was more sensitive to individual differences. Design and procedure ~~~~~~~~~~~~~~~~~~~~ All individuals were taught to read 12 nonwords, and there was phonological pre-training for six of these nonwords. Individuals were then tested on their phonological and orthographic knowledge of the taught nonwords. The phonological pre-training condition and control condition (no pre-training) took place in two separate sessions on different days. All children were seen individually and the testing sessions lasted 20–30 min each. All typically developing children were seen at school and the individuals with Down syndrome were either seen at home or school. In the pre-training condition, the phonological pre- training occurred first, followed immediately by written nonword learning. In the control condition, the written nonword learning was the first activity. The background tasks were then administered, followed by the phonological choice post-test and the orthographic choice post-test. Training materials Twelve pairs of nonwords were created that had two different but homophonic vowel digraphs (e.g. ‘nirp’ and ‘nurp’). All nonwords had four graphemes and three phonemes and are shown in Appendix. Within the nonword pairs, there were no significant differences between the number and frequency of orthographic neighbours of the nonwords according to the ARC database (Rastle, Harrington, & Coltheart, 2002). The nonword pairs were separated into two lists ensuring that similar vowel patterns were not all in the same list. Within each pair, nonwords were randomly assigned to be either a target nonword or pseudohomophone distracter, which would be used in the orthographic choice post-test. These assignments were the same for all individuals. Learning procedure The individuals were introduced to the training procedure by being told they were going to learn some words from an alien planet. The two conditions (phonological pre-training vs. control) occurred on different days and the training procedure is depicted in Fig. 1. The two lists of target nonwords were counterbalanced across the two conditions, as was the order of conditions resulting in four versions of the experiment. The learning procedure lasted approximately 5–10 min and the children generally enjoyed it and attended well. Phonological pre-training There were four trials in the phonological pre-training procedure, each of which consisted of a repetition activity and a phonological consolidation activity. These activities were spoken and individuals did not see any written material. First, the individuals heard the nonword and repeated it. This was the same across each trial. They then did a phonological consolidation activity that was designed to familiarise individuals with the sound of the nonword. This increased in difficulty throughout the training. For the first trial the individuals heard the nonword segmented into its constituent phonemes and repeated this and they also heard the first and last sound isolated. The second trial was the same as the first but the individuals sounded out the nonword independently. For the third and fourth trial, individuals independently isolated the initial sound and the final sound, respectively. The six nonwords appeared in a fixed random order within each trial and corrective feedback was given, which was either “well done” or “that's not quite right” followed by the correct answer. Written nonword learning (reading aloud) There were four trials of written nonword learning, in which the individuals saw the written nonwords individually and read them aloud. Individuals were permitted one attempt at the answer and were then given corrective feedback, which was either “well done” or “that's not quite right” followed by the correct answer. The nonwords were printed in size 36 Century Gothic lower-case font-type on sheets of A4 paper. All six nonwords appeared in a fixed random order within each trial. Individuals were awarded one point for each nonword they read correctly. Problems with articulation are common in children with Down syndrome (Kumin, Councill, & Goodman, 1994; Roberts et al., 2005) and if a phoneme was pronounced incorrectly but with a consistent realisation then this pronunciation was accepted as correct in the nonword learning task. Experimental post-tests ~~~~~~~~~~~~~~~~~~~~~~~ Two computerised post-tests, a phonological choice task and an orthographic choice task, were administered approximately 10–15 min after the training procedure. The phonological choice task assessed whether individuals could recognise the phonological forms of the trained nonwords. The orthographic choice post-test assessed individuals’ recognition of the nonwords’ spelling. Phonological choice task The phonological choice task was presented on a laptop computer using e-Prime version 1.0 (Schneider, Eschman, & Zuccolotto, 2002). Pictures of three children were shown sequentially, each accompanied by a spoken nonword that was either the target nonword or a distracter. The three pictures of the children then appeared simultaneously on the computer screen and the individual was asked who had said an alien word that they had learnt earlier. They could respond verbally or non- verbally by pointing. For each target nonword, two distracters were devised (see Appendix). The first differed by one phoneme; for half of the nonwords, the initial phoneme was changed and for the remaining half, the final phoneme was changed. The second distracter differed by two phonemes and was based on the first distracter but the remaining consonant was also changed, for example the distracters for ‘nirp’ were ‘nirt’ and ‘mirt’. Prior to the experimental trials, two practice trials with real target words were administered to familiarise the individuals with the task demands. These presented the real word ‘bird’ with the distracters ‘mird’ and ‘mirg’ and the real word ‘ball’ with the distracters ‘borz’ and ‘morz’. Orthographic choice task The orthographic choice task was also presented using e-Prime. Three written nonwords appeared on the computer screen simultaneously and individuals were asked to pick the alien word that they had learnt that day. They could respond verbally or non-verbally by pointing. As outlined in Section 2.4.1, all target nonwords had a pseudohomophone distracter. A visual distracter was also created by changing the last consonant of the target nonword to a visually similar one. Therefore, the target nonword ‘nirp’ had the pseudohomophone distracter ‘nurp’ and the visual distracter ‘nirg’. All distracter items are shown in Appendix. Prior to the experimental trials, two practice trials with real words and pictures were administered. These presented the real word ‘bird’ with the distracters ‘berd’ and ‘birg’ and the real word ‘door’ with the distracters ‘doar’ and ‘doof’.","Raw scores and a significance level of .05 were used for all analyses. One child with Down syndrome did not score on the alliteration matching and sound deletion tasks for behavioural reasons therefore N was reduced for these tasks. Performance on background measures ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To examine the distribution of scores within the groups the Shapiro–Wilk test was used and histograms were examined. If the distribution of scores for a task were non-normal for either or both groups then a Mann–Whitney U test was used to test the differences between the groups. If the distributions were normal in both groups, then an independent samples t-test was used. The mean scores, standard deviations and between-group test results are reported in Table 1. Where possible age-equivalent scores are also reported to indicate the developmental level of the two groups. As can be seen the two groups were matched on word reading and from previous research it would be expected that the typically developing children would perform significantly better than the individuals with Down syndrome on all other tasks. This was the case for matrices, alliteration matching, sound deletion and word recall, and there was also a trend for a difference on nonword reading. However the two groups performed similarly on the picture naming task. Phonological learning ~~~~~~~~~~~~~~~~~~~~~ In order to investigate whether the phonological pre-training improved written nonword learning, it must first be established that it resulted in increased phonological knowledge in both groups. The mean scores for the correct response in the phonological choice post-test are shown in Table 2. For both conditions and both groups, the correct answer was chosen above chance levels (score of two). The two distracters were chosen at similarly low levels across groups and conditions. Written nonword learning ~~~~~~~~~~~~~~~~~~~~~~~~ The written nonword learning scores for both groups in the control and phonological pre- training conditions across the four learning trials are shown in Fig. 2 and the overall means are in Table 2. Individuals with Down syndrome generally performed at a lower level than the typically developing children, with both groups doing better in the phonological pre-training condition than the control condition. Learning increased across trials, although this appeared to be greater in the control condition. Summary of results ~~~~~~~~~~~~~~~~~~ The phonological pre-training resulted in better performance on the phonological choice post-test and improved ability to read the taught nonwords. In comparison to the typically developing group, the individuals with Down syndrome performed more poorly when reading the nonwords and there was a marginally significant trend for them to benefit more from phonological pre-training. However there was no effect of phonological pre-training or group on orthographic knowledge. Matching for target nonword decoding skill ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To test the prediction that individuals with Down syndrome would show written nonword learning performance and benefit from phonological pre-training at a level equivalent to their decoding skill, a subgroup of children from the typically developing group were individually matched to the individuals with Down syndrome on their scores on the first written nonword learning trial (first reading aloud attempt) in the control condition. This criterion was chosen as it reflects ability to decode the nonwords with no previous exposure to the spoken or written form. This resulted in 16 typically developing children (seven males) being selected who had a mean age of 6;02 matched to the 16 individuals with Down syndrome. Performance on background measures The performance of the typically developing subgroup on the background cognitive measures is shown in Table 3 along with the individuals with Down syndrome for comparison purposes. Matching on target decoding also resulted in the groups having similar scores on the standardised nonword reading test. Due to the uneven profile in individuals with Down syndrome, with word reading generally being a strength compared to nonword reading, the individuals with Down syndrome performed significantly better on the word reading measure than the typically developing subgroup. The two groups had similar scores on the picture naming task, but the typically developing subgroup performed significantly better on the remaining tasks of alliteration matching, sound deletion, matrices and word recall. Written nonword learning The scores in the control and phonological pre-training conditions during written nonword learning are shown in Fig. 3 for both groups and the overall means can be seen in Table 2. Levels of learning were similar in the two groups in both conditions and there was a clear benefit from pre-training in both groups. As with the data from the whole sample of typically developing children, learning occurred over the training procedure but this appeared to increase at a faster rate in the control condition. The significant interaction between condition and trial was followed up with a Tukey's HSD test, with means collapsed across groups. The only significant difference either across trials within conditions or between conditions on each trial was the difference in favour of the phonological pre- training condition compared to the control condition on trial 1 (HSD = 0.93). Summary of results: matching for target decoding Phonological pre-training led to similar increases in phonological knowledge in the individuals with Down syndrome and the subgroup of typically developing children matched on decoding. The two groups were also equally accurate when reading the nonwords aloud and benefitted from the phonological pre-training to a similar extent. There was no significant difference between the two conditions or groups when individuals were required to identify the correct spelling pattern of nonwords they had learnt.","Vocabulary knowledge has been proposed to play an important role in reading development for individuals with Down syndrome (Hulme et al., 2012). The present study built on these findings and findings from studies with typically developing children, which suggest that the phonological aspects of vocabulary knowledge are related to word reading skill, particularly when decoding is compromised (Duff & Hulme, 2012; McKague et al., 2001; Nation & Cocksey, 2009). Specifically this study investigated whether word level phonological knowledge provides a benefit to individuals with Down syndrome when reading new words. Compared to typically developing children matched on word reading, the individuals with Down syndrome performed more poorly during written nonword learning. Although both groups benefitted from phonological pre-training, there was a trend for this benefit to be greater for individuals with Down syndrome. When the individuals with Down syndrome were matched to a subgroup of the typically developing children on target decoding skill, written nonword learning was equated in the two groups and both groups benefitted equally from phonological pre-training. However, there was no effect of phonological pre-training on a measure of orthographic knowledge in either set of analyses. During written nonword learning, individuals with Down syndrome were less accurate than typically developing children matched on word reading. However, as there was a trend for the typically developing children to perform better on the standardised measure of nonword reading, the two groups were not equated for baseline performance. Therefore 16 of the typically developing children were pairwise matched to the individuals with Down syndrome for target decoding. This also resulted in the two groups being well matched on the standardised measure of nonword reading. There were no significant differences between this subgroup and the individuals with Down syndrome during written nonword learning. Thus as predicted, on the task of written nonword learning the individuals with Down syndrome performed in line with their existing decoding skill but more poorly than would be expected from their existing word reading level. This is also in line with studies with children with dyslexia (Bailey, Manis, Pedersen, & Seidenberg, 2004; Ehri & Saltmarsh, 1995). The phonological pre-training exposed individuals to the spoken nonwords and required them to engage in activities that emphasised the individual phonemes within the nonwords. Familiarised nonwords were subsequently better recognised in a phonological choice post-test by both typically developing children and individuals with Down syndrome and as predicted, they were also read more accurately. Nonwords in the control condition were also recognised at above chance levels in the phonological choice post-test; this is likely to be because individuals had also been exposed to the phonological form of the nonwords during the written nonword learning. Phonological pre- training is argued to support reading by providing individuals with a whole-word phonological representation of the new word. This can resolve competition between the correct pronunciation and pronunciations derived from decoding the letter-sound mappings and therefore provide a pronunciation match for partial decoding attempts (Nation & Cocksey, 2009). There was a marginally significant trend for the individuals with Down syndrome to benefit more from phonological pre-training than the typically developing children matched for word reading. In the context of the triangle model, when the phonological pathway is compromised, such as with irregular words that cannot be decoded, existing phonological knowledge is considered to have a stronger effect on reading (Nation & Cocksey, 2009). This may also be the case for children who have an impaired phonological pathway and thus may apply to all words regardless of regularity for individuals with Down syndrome. Specifically, as the individuals with Down syndrome were less able to decode the words, they would be more likely to produce partial decoding attempts that phonological knowledge could then help resolve. This between-group difference was no longer present when the individuals with Down syndrome were matched to the typically developing subgroup for target decoding. This is argued to be because the two groups now made an equal number of partial decoding attempts, therefore resulting in similar opportunities for phonological knowledge to aid performance. Our findings support the suggestion of Taylor et al. (2011) that existing phonological knowledge about a word could be incorporated within the triangle model as item-specific phonological representations. To assess whether semantic knowledge also contributes to reading when decoding is compromised, such as with individuals with Down syndrome, the effect of semantic pre-training could be compared to phonological pre-training. Both typically developing children and individuals with Down syndrome showed increased scores across the trials during the written nonword learning. However this improvement was greater in the control condition than the phonological pre- training condition, particularly at the beginning of the training procedure. The nature of the corrective feedback during written nonword learning may have contributed to this difference between the two conditions. For the feedback, individuals heard the correct pronunciation of the target nonword and, as suggested by the lexical quality hypothesis, this may have provided more benefit to the control nonwords as it added new phonological information to the lexical representation (Perfetti & Hart, 2002). Previous studies that have examined the effect of phonological pre-training on written nonword learning in typically developing children have not included post-tests that specifically tested orthographic learning (Duff & Hulme, 2012; McKague et al., 2001). In the present experiment the target item was selected at relatively high levels of accuracy in an orthographic choice post-test by both groups in both conditions. Therefore the same level of orthographic knowledge was demonstrated for nonwords that received phonological pre- training and those that did not. Furthermore despite showing poorer learning, the individuals with Down syndrome showed the same level of orthographic knowledge as the typically developing sample matched on word reading. The self-teaching hypothesis posits a central role for phonological recoding, or decoding skill, in forming orthographic representations. Therefore it would predict that better performance during training (i.e. in the pre-training condition and for the typically developing children) would result in a stronger orthographic representation that would enhance performance in an orthographic choice task. It could be that the task was not sensitive enough to capture this; alternatively, repeated exposure to the orthography of the target word may have resulted in individuals creating and storing an orthographic representation of the nonword, regardless of whether they read it correctly. To test this, the orthographic choice task could be administered after the first trial of written nonword learning, when in the control condition participants will have had minimal exposure to the phonology of the word. It should be noted that the present study assessed written nonword learning immediately after training and not in the longer term. It is clearly important both theoretically and practically to know if learning is maintained, particularly as there is some evidence that points to impairments in consolidating newly learnt skills and on verbal, visual and spatial long-term memory tasks in Down syndrome (Carlesimo, Marotta, & Vicari, 1997; Pennington, Moon, Edgin, Stedron, & Nadel, 2003; Visu-Petra, Benga, Ţincaş, & Miclea, 2007; Wishart, 1993). In a written nonword learning study with typically developing children by Bowey and Muller (2005), an orthographic choice task was administered immediately after learning and six days later, and there was significantly better performance on the immediate post-test than the delayed post-test. In future, there is a need for studies with individuals with Down syndrome to include a delayed test of written nonword learning. There are a growing number of studies evaluating reading accuracy interventions for children with Down syndrome. Training phonological awareness, letter-sound knowledge and new vocabulary results in improvements of the directly taught skills and the reading of untaught words, although there is less evidence that this training extends to nonwords (Cologon, Cupples, & Wyver, 2011; Goetz et al., 2008; Lemons & Fuchs, 2010). In a randomised controlled trial, Burgoyne et al. (2012) found that an intervention combining both phoneme awareness and broader oral language skills, with a large emphasis on vocabulary instruction, had a significant effect on word reading. The present study provides evidence that the phonological aspect of vocabulary knowledge may help promote reading achievement and therefore being presented with the spoken form of a new word before seeing it in print could be a valuable part of reading instruction for children with Down syndrome."],["In this paper we reflect on the kind of listening that happens in research whilst taking part in a keep fit group and getting sweaty, that pushes us to ask an interviewee 'Are you alright?' and haunts us when the project is over. This is the kind of listening that weaves through, around and beyond what is immediately heard, including the unspoken, the articulateness of objects and the listening that comes through participating. The paper stems from a project concerned with how people live, experience and manage cultural diversity and ethnic difference in their everyday lives in urban England. Divided into two sections, the first part introduces our methods that included participant observation, interviews and repeat in-depth discussion group meetings. The second section reflects on our experiences of listening whilst doing, explores feelings that mediate listening and considers the time involved in listening. --------------------------------------------------------------------------------","Omar was eighteen years old and a student at Grafton College in Milton Keynes. He was one of the young people taking part in a discussion group for our research concerned with multiculture and how people live ethnic diversity. Outside of college Omar told us that he enjoyed spending time with his girlfriend, working to earn money and working out at the gym. His uncle liked body building too. He pulled his phone out of his pocket, found an image of his uncle's muscly, naked torso and passed it around for us to look at. Omar talked quickly, positively and hopefully, but how he talked sat uneasily with what he said about his life, the transience of his childhood, his absent mother, and how we experienced the group discussion. The phone was returned to Omar and placed on the table, contributing, somehow, to the group. Crisps, biscuits and drink fed into the unfolding discussion, adding something substantial, providing us with something to hold onto, bite into as we talked and listened to one another. We return to Omar later, but we begin with him to introduce the focus of this paper concerning the kind of listening that weaves through, around and beyond what is immediately heard, including the unspoken, the articulateness of objects and the listening that comes through participating. Whilst what Omar said was important to our research, there was more than this to listen to. Textbooks on research methods are helpful regarding talk - who to talk to, how to ask questions, turning talk into transcripts and how to analyse them (Valentine, 2005; Longhurst, 2010; Crang and Cook, 2007), but are less helpful on listening (Bennett, 2002). This paper aims to fill that gap a little and is for students and researchers who want to think about their listening and how this shapes their understanding and research. In this paper we draw upon our experiences of listening in our research on ‘Living Multiculture’1, one of our aims was to devise a methodology that involved trying to listen better to everyday, mundane practices and experiences of multiculture. Without marginalising everyday racism, exclusion and inequalities we wanted to explore the quieter micro narratives and routine encounters that are part of the lives of a growing majority of people living in urban England. In our ‘Living Multiculture’ project we took a different approach to ‘the segregation-distrust-conflict model’ that has underpinned and shaped UK public and policy debates about cultural difference (see Neal et al., 2013, 2015a, 2015b). Instead we were influenced by an alternative narrative of convivial encounter across difference (Back, 1996; Amin, 2002; Gilroy, 2004; Wise, 2009; Gidley, 2013; Thrift, 2005; Swanton, 2010; Wise and Velayutham, 2014; Wessendorf, 2014) that does not ignore tensions (Clayton, 2008; Rogaly and Qureshi, 2013) but recognises unpanicked, everyday, routine experiences of multiculture (Noble, 2009) and the developing skills and competencies that shape these (Wise, 2009; Neal and Vincent, 2013; Sennett, 2012; Wilson, 2011). This means that our listening involved methods that asked questions and we were concerned about developing good research practices for listening to complex stories of hardship, loss, disorientation and exclusion. These practices involved repeated and sustained connection with people and places, reflection on our experiences of the research and recognizing that we were stitched into the possibilities and limits of our listening and understanding (Pratt, 2010; Kanngieser, 2012; Dreher, 2009). Our methods also involved attending to mundane practices that shape everyday lives and experiences of multiculture; the sometimes wordless encounters, small gestures, ambiguity and atmospheres that embody getting about, amongst and along with others. As we detail later, our research involved interviews, repeat in-depth group discussions and participant observation (Neal et al., 2013, 2015a, 2015b; Jones et al., 2015). In this paper listening is broadly conceived. This is in part a reflection of our different disciplinary backgrounds and how individuals listen in unique ways (Forsey, 2010). It also reflects trends regarding ethnographic research, which involves a range of methods often including participant observation alongside other methods such as interviews (Crang and Cook, 2007). Some lament the demise of participant observation as a stand alone method and there is concern regarding the crowding of this method with others which demand talk under the label ‘ethnographic’ (Gans, 1999; see also Forsey, 2010). Interesting questions have been asked regarding what gets lost in ‘verbal methodologies’ (Crang, 2005; Back, 2003, 2007, 2012). Whilst this paper does not prioritise one method over another, it is interested in the kind of listening required in ethnographic research that involves not only attending to, and analysing talk, but, for example, acknowledging the context in which stories are told and how research interviews and group discussions are experienced. For us listening has a relational, intersubjective dynamic, decentring the researcher and illuminating the role of participants in the creative process of research and understanding. Participants include the non-human and in the paper we consider some of this vital matter, such as Omar's phone and a crisp packet, that substantiated our listening (Thrift, 1999; Conradson, 2005; Simpson, 2013). Finally, we consider the time involved in listening. Whilst projects and research contracts have end dates, listening does not and we live with, return to and listen to moments, conversations and memories of participants that won't leave us alone. The paper is divided into two sections. In the first section we briefly introduce our methods, detailing our planning for listening. In the second we reflect on our listening, focusing on three issues in particular: listening whilst doing, the feelings that mediate listening and the time involved in listening. The paper draws upon a diverse literature to think about listening. To help us explore ‘listening whilst doing’ we explore an embodied approach to listening evoked in methods that include participant observation (Laurier and Philo, 2006; Swanton, 2010; Rogaly and Qureshi, 2013; Wise and Velayutham, 2014; Wessendorf, 2014; Crang, 1994), participatory action research (Askins and Pain, 2011) and non-verbal research techniques (Macpherson and Fox, 2014; Fox and Macpherson, 2015; Bingley, 2003). Different theoretical, disciplinary and political agendas motivate this various research, but common to all is not only listening to what participants might say, but listening whilst doing and involved in happenings, evoking something of the experience of being there. To consider feelings that mediate listening, the paper is indebted to the work of geographers with expertise in psychoanalysis and psychotherapy (Kingsbury and Pile 2014, Cullen et al 2014, Davidson and Parr, 2014; Kingsbury, 2009; Bondi, 2003, 2005, 2014a, 2014b). Especially helpful has been writing on empathy (Bondi, 2003, 2014a) and listening that involves ‘tuning in’ (Paterson, 2014) to our experiences and how they might connect to the experiences of participants and what they are trying to ‘tell’ us, whilst also mindful of our different subject positions that require reflection and shape what is heard. Finally, a body of work in the field of education and arts based research concerning acousmatic texts is a source of inspiration regarding our reflection on the time involved in listening (Daignault, 2005; Aoki and Aoki, 2003; Leggo, 1999). Sometimes texts that make us tingle, the really inspirational ones that leave us buzzing, seem to be listening to us too as we read them. In this paper we explore work on acousmatic texts to consider the impact of time on our listening, research and transcripts.","Our attempt to listen better in our research did not involve inventing new methods, but did involve weaving the work and ideas of researchers and writers regarding participating and embodied listening, empathy and acousmatic texts outlined above into our approach and planning. As this section begins to introduce, our methods involved participating and joining groups in case study areas that were home to at least one of us, repeat group meetings with time in between to reflect on our listening and a team of researchers shaped by individuals who brought different readings and interpretations to transcripts and field notes. We situated our research project in three case study areas, chosen because of their particularly dynamic populations, urban geographies of diversity and multicultural formation. Our case study areas were the London Borough of Hackney, Milton Keynes, a new city near London established in the 1960s, and Oadby, a small suburban town on the edge of Leicester in the English midlands. These case study areas represent some of England's most dynamic and (super) diverse populations shaped by people with a wide range of ethnic backgrounds. Between 2001 and 2011 Hackney and Milton Keynes were amongst the UK's top ten fastest growing places with their populations increasing by 20% and 17% respectively (Milton Keynes Council, 2014; Hackney Borough Council, 2013). The ethnic composition of both places also changed between 2001 and 2011 with Hackney's long history of ethnic diversity intensifying (Neal et al., 2015a) and Milton Keynes' black and ethnic minority group doubling. Although Oadby's population growth was nothing like as dramatic, between 2001 and 2011 it was amongst England's fastest changing places in terms of its ethnic composition (Leicestershire County Council, 2013). Each of our three case study areas was also home or very close to home for at least one of us and we saw ourselves as part of the social world we were studying. Our connections with these places also played a role in shaping what we did and knew and added an extra layer of responsibility regarding these places and our research participants. We imagined the research through specific sites that included a park, café, library, social and leisure clubs and a college in each of our case study areas where different and differently positioned people (around age, education, residency, gender, class and ethnicity) encounter and interact with one another in ways that feel comfortable and less so. Our way of getting involved in these sites included taking part in activities already happening inside them, joining, for example, a running or writing group. Whilst our attempts to listen better involved decentring ourselves and taking part in selected sites, we are a group of mostly white social scientists which raises ethical issues around privilege (Skelton, 2001; Dwyer, 1999). In this paper ‘we’ is complicated. One layer of ‘we’ involves the paper's authors, but another layer of ‘we’ refers to our fieldwork which involved not only us, but also researchers and research consultants who were involved in different stages of the project. We are a diverse group of people in all sorts of ways but especially with regard to the ways in which people are typically identified around class, gender, age and ethnicity. Whilst ‘we’ describes a collective and collaborative approach, some of what we write about in this paper involves individual experiences fuelled by particular contributions to the project. So our use of ‘we’ is somewhat nuanced and we break out into individual experiences when this is necessary. Undoubtedly how others identified and related to us shapes our particular findings and how they are read. Who we are and how people see us affect how we listen, what people tell us and what we hear (Dreher, 2009). That said, we attempted a biographical approach to identity, one that recognises its complexity and ways in which it is shaped through interaction with people and places through the course of lives (Swanton, 2010; Nayak, 2006). For example, one of us – Giles – is mostly identified as white and usually identifies himself as white, but is mixed ethnicity with an Indian father and a White British mother, living his infancy in Ghana and later moving to Sheffield where he tells people he is from. Many of our respondents had similarly complex and interesting stories to tell regarding their biography and processes of identification. As we have discussed elsewhere (Neal et al., 2015b), though, we found that when we were writing about our observations in our field notes we sometimes slid into seeing through an essentialising lens that highlighted people's physical and cultural characteristics to identify them and evoke difference while paradoxically writing about how difference might have been disrupted. We attempted to counter this objectifying way of seeing through the layering of methods that comprised joining in and taking part and listening to interviewees' experiences of everyday multiculture that evoked both multi-textured understandings of identity and belonging that reckon with essentialism, but also their experiences of objectification and racism (Clayton, 2012). Our key methods included participant observation, interviews and repeat in-depth group discussions. The aim of our participant observation (Laurier, 2010; Cook, 2005; Crang and Cook, 2007) was to listen, watch, feel and be able to describe site worlds and the context, people, practices, etiquette, uses, rhythms, things and atmospheres that shaped them. We recorded our observations through writing, attempting to capture the minutiae of encounters and interactions happening around us and that we ourselves were involved in. We wrote about what happened in college canteens, parks, libraries and cafes and at events like world picnics in parks (Crang and Cook, 2007). We wrote about the keep fit group and writers group meetings that we joined. We returned to our sites at different times of the day and at different points of the year to get a sense of their daily and seasonal rhythms. A second key method of the research was repeat in-depth discussion groups (Burgess et al., 1988a,b). We set up groups of people who used the park, attended the college or were a member of the selected social club. We used a number of strategies to meet potential participants. These involved going along to and taking part in group activities, such as a keep fit group, running club or a writing group, and events, for example a community fun day. Before we set up the first group meeting, we interviewed participants one-to-one to get a sense of their biography (Valentine, 2005; Longhurst, 2010). In total we interviewed 88 people. Our repeat in-depth discussion groups were influenced by the psychoanalytically informed group work of the geographers Burgess, Limb and Harrison who, in the 1980s, considered not only what people said, but also listened to group dynamics and how they themselves experienced the groups (Burgess et al., 1988a,b). We met with 12 groups three times over a six month period, stretching from Autumn 2012 to Spring 2013. Our groups ranged in size with between 5 and 11 members. Some of our groups (such as our park user groups) were made up of people who did not know each other whilst others were comprised of people who did. Our in-depth discussion group meetings often involved two researchers for practical reasons around hosting and organising groups, but also to explore how they experienced group meetings. Repeat meetings gave us time to step back and reflect on our listening and understanding with others (members of the research team and locally based advisory groups2) (Bondi, 2013, 2014a). The final group meeting embraced the intersubjective nature of research (Jervis, 2014), allowing us to explore issues that seemed pertinent to that particular group, our understanding of those issues and their reactions to what we had picked up on. A final stage of the research involved iterative interviews with local and national policy makers. In each of our case study areas we shared some of our emerging findings with policy makers and community activists, prompting discussion and reflection around these. This stage of the research involved 22 interviewees working at a local level and four at a national level concerned with, for example, issues of race equality. Our aim was to listen to policy maker and activist responses to our research, whist attempting to connect our research to their work. Listening whilst doing ~~~~~~~~~~~~~~~~~~~~~~ There are some methods in particular that involve listening whilst doing and these include participant observation, participatory action research and non-verbal research techniques (Phillips and Johns, 2012). Different research agendas and ambitions often underpin these various methods, but they do not usually involve asking questions, rather sliding into the worlds of others and creating space in which the research can unfold (Crang and Cook, 2007). Listening whist doing is valuable because it concerns an active, sensuous, embodied approach to listening bringing experiences and feelings to the foreground of research as they weave through words said and happenings observed. Research practices at the interface of social science and psychotherapy involve embodied, or ‘whole body listening’ (Macpherson and Fox, 2014; Fox and Macpherson, 2015), that entails attending to gestures, textures, atmospheres, things and the context of happenings (Back, 2003, 2007, 2012; Mitchell and Back, 2006; Bissel, 2010; Kanngieser, 2012). Whole body listening involves ears, eyes, beating hearts, feelings, skin, pores, tingly, hair raising moments and more besides (Paterson, 2015). Some of our methods involved listening whilst doing as we joined groups to take part in their various activities, such as writing, running and keeping fit. None of these activities were set up for the purposes of our research, but concerned an involved, embodied and agile approach to listening. One of us, Katy, joined a keep fit group in Knighton Park, on the edge of Oadby: It's 9.30, damp and cold and I'm taking part in a keep fit group. We're a mixed group of young to middle aged, white and British Asian women. .......We set off. Jogging, sprinting, dropping onto the wet ground to do press ups, crawling. Some of the exercises require us to work in pairs and so we're working with partners, using them as support to do leg exercises before we swap partners again. It feels a bit odd holding people I've only just met. The group runs into a woman doing her morning walk, we're difficult to avoid and Jake yells at us to give the walker space. I work with Maggie, Paula, Kay, Amita, Shivani and others, we introduce ourselves, laugh a lot because it's all a bit awkward and hard work. I'm crawling down a bank being yelled at by Jake for holding my bum too high. We're slipping and sliding in the mud (Knighton Park 15/10/12). Three times a week a group of mostly women meet in the car park just before 9.30am to attend a keep fit group. Participation involves running and jogging through the park on paths and off track, using the ground and park furniture, such as park benches, for various exercises. The keep fit group is run by Jake, a black, male fitness instructor. Katy went to group sessions on Mondays, and sometimes Wednesdays, over a six month period. The size of the group varies depending on the weather, time of week or year. Group members are generally middle class, although not always, young to middle aged, ethnically mixed, comprising women who work part-time, are self-employed or full time mothers. Surprising to Katy was the physicality of her listening whilst sweating, sliding around in the mud and holding, touching and pressing upon others she had just met. The group involves people who wouldn't ordinarily cross paths, and involves encounters with others also using the park, for other activities, at the same time (Neal et al., 2015b). Dog walkers are the other main users of the park at 9.30am on a weekday. Dog walkers are generally white and older than the group of women and not always at ease with the bubble of chatter and Lycra clad bodies taking up paths accompanied by Jake loudly yelling instructions. There is little in the way of verbal exchange across the groups, but listening involved running past unsmiling faces and cold shoulders from the dog walkers, a change in pace and an altered atmosphere amongst us. We're sliding around in the muddy wet grass, walking like bears, hands and feet on the ground, going up to each other, lifting one hand off the ground, then another, for a ‘Hi 5’. We're grunting, groaning and giggling as we approach each other. The ground is slippery and Jake warns us about the wet leaves ….. (Knighton Park 29/10/12). Notable are the non-human things and beings that mediate social relations and substantiate listening (Askins and Pain, 2011). Things sparked and facilitated interactions between participants in ways barely noticed such as the wet leaves described in the extract above, but also water bottles, park benches, dog shit and dogs amongst many things and non-human beings that formed our field notes. In our field notebooks we wrote about things that mattered to us and were pointed out to us by others. Loose dogs unsettled the group, causing screams if they got too close, dog shit was irritably pointed out by a group member to others, reigniting tensions between dog walkers and keep fit group members, the offer of a water bottle soothing atmospheres experienced in the park, creating friendly relations amongst group members. All of this vital matter substantiated listening, shaping experiences of others and this place, making an impression, having an affect (Thrift, 1999; Conradson, 2005; Simpson, 2013). Interesting to the research project was the dynamic, simultaneous interplay of tense and easy relations amongst ethnically diverse individuals and groups and the impact of these experiences on relationships with sites and places. The kind of ‘data’ that listening whilst doing elicits is somewhat different compared to, say, interviewing in that it emerges around and through the shared activity, being with, but not generally sat opposite, and being amongst others – human and non-human. Listening whilst interviewing often involves building up information pixel by pixel, but listening whilst doing involves being confronted by a big, blurry screen and trying to bring this into focus. This analogy is too reliant on the visual, but explains the contrasting processes of understanding. Bringing a research site into focus is nicely exemplified in a researcher's field notes as he had lunch in a college six form area in Oadby: ‘The mood in the canteen was generally playful – as I filled my bottle up at the water tank, for example, an Asian group of boys and girls were messing around with a packet of crisps, with one trying to crunch the packet up and the other running away. Towards the end of our time in the canteen, a boy in a turban and another Asian girl sat next to us. The boy quietly got on with studying, but the girl was attracting the attention of a group of Asian lads, and some girls, on the comfy chairs next to us. I heard one of them lean over and say to the girl, ‘sorry, we were just taking the piss out of you’. The girl smiled and the lad said, ‘sorry, what was your name again?’ (Uplands College, Oadby, 10/10/12) Understanding produced through listening whilst doing is created through words said but also whilst, as exemplified above, experiencing a six form area unfold through an atmospheric mix of people and objects involving a crisp packet, messing around, laughter and wipe clean leather effect comfy chairs supporting more bodies than intended as students squeezed into them. Listening generated understanding about a college that, although very ethnically diverse, involves space and groups often organised around ethnicity (and gender), but also mixing on the edges of, and across, groups in ways that staff and students (less) consciously recognise. Experiences and feelings that mediate listening ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Throughout the previous section dealing with listening whilst doing there was reference to experience and feelings. In this section we bring feelings centre stage to consider how these mediated our listening. In our field notes we wrote about our experiences of interviews, group discussions and fieldwork, grappling with feelings and happenings that weaved around and through talk. A common thread in emotional geographies is the relationality of feelings and how what we experience is, in some way, connected to the feelings of others (Bondi, 2005). We drew upon this to listen to what was being communicated beyond and around what was immediately heard. To start this section we return to the college in Oadby to explore a third and final group discussion. In an earlier meeting this group had talked about ‘mixing’ and in our final meeting we asked what this involved. Ayo, who usually sat at a table in the six form area identified as where black African women sat, responded: ‘It's like if I'm just talking randomly in my perspective like, mixing is just us interacting with other people of course, but it's not to an extent where I can engage with them on that level because obviously I haven't experience of what – how they see things in like in their perspective, so it's mixing – you understand them to an extent but it's not like, you know diluting, if you know what I mean?’ (Ayo, Group Meeting 3, Uplands College, Oadby 22/4/13) Amelia, a white student with a British father and German mother, picked up the conversation, emphasising that she mixed with people with different ethnic backgrounds and that she and her friends did not have a regular table. Ayo responded that although she was associated with the Black African table that she also mixed with others, but said, ‘I'm not comfortable with everyone that I sit with, so there's different levels of comfortability’. Amelia picked up on the diluting idea, describing it as ‘weird’ at which point Ayo withdrew from discussion. ‘Are you alright?’ one of the researchers asked her. ‘Yes’, but she was much quieter. Silent for the time being. Whilst this conversation was on-going a researcher detailed in her field notes how a blue plastic cup was passed back and forth between two students, Amira and Tahir, who were silently and amusingly communicating their horror of lipstick marks on what should have been a clean cup. This was another exchange happening between students, who could have been opting out of what was being discussed and/or busy with their own relationship. Another researcher focused on her experiences of the group not feeling very relaxed, ‘a bit of tension’, ‘giggly’ talk and the ‘unhappy’ response of Ayo: ‘We're a group of nine now (….). The group doesn't feel as relaxed as I expected (….) They help themselves to drinks and snacks, seem to enjoy us reiterating, repeating and pushing issues and phrases they've introduced in previous meetings. But there is definitely a bit of tension. Perhaps this has something to do with impending exams that all students seem to be sitting soon (…) But there are other things happening in the room, especially between Amira and Tahir. Amira is giggly and constantly looking across me to Tahir for reassurance, for him to agree with what she has said, to tease him about something he's been involved in, winding him up. Tahir is defensive, a little less open than in previous meetings. He's definitely much quieter. Amelia and Ayo keep the conversation going, keep the group on track, ease the tension or whatever's going on between Tahir and Amira. Ayo and Amelia are two of the chattier, more confident students and are forthright regarding their experiences and opinions. At one point Amelia says something that Ayo is unhappy about. I try to encourage Ayo to have her say, but she withdraws, still managing to make some kind of point in the process.’ (Field notes, Uplands College, Oadby, 22/4/13) On the one hand we have the recording and transcript of the group meeting that we listen to, read and analyse. Shaping our listening are our experiences of that group, which take us by surprise because this group has worked well in the previous two meetings. The dirty plastic cup is amusing the students, annoying us, with Amira seeking the attention of Tahir and getting nothing back apart from a dirty cup. What do participants mean by mixing? Ayo's response perhaps unravels the carefully constructed consensus that this is a very ethnically diverse school comfortable with its diversity. Amelia, a white student, holds Ayo to account, uses the word weird and the atmosphere feels worse. The thread of conversation between Ayo and Amelia had kept us (researchers) on track whilst others were distracted and amused by Tahir and Ayo. The word weird is said, ouch, oh hell. Ayo is silenced. We empathise with Ayo. Are you alright? Should we have asked this? She withdraws from the conversation, her silence, our experience of that silence, is now a rather different kind of silence, as punchy and informative as what is being said as others move the conversation on, smoothing over what has happened. Is this how students cope with tensions? How would Ayo and the group have coped had we not acknowledged our feelings? Liz Bondi (2003) described empathy as “a process in which one person imaginatively enters the experiential world of another” (Bondi, 2003: 71). For us this involved working with our feelings and a kind of imagining in (to the experiences of participants - Are you alright (Ayo)?) in an attempt to appreciate where they were coming from whilst mindful of differences (and connections) around ethnicity, age, gender, residency, class and more besides (Sennett, 2012; Cochrane, 2014). The importance of sustaining a sense of difference or ‘alterity’ in empathy is emphasised by Liz Bondi (2014a) through her use of the ‘third position’ which involves shifting between participating in a relationship and observing it, encouraging a ‘stepping back’ rather than stepping into the (metaphorical) shoes of the participant. Empathic listening involves engaging fully with the unfolding story being recounted, but does not involve shared experience or greater knowledge (Bondi, 2014a; see also Gair, 2012; Watson, 2009; Dreher, 2009; O’Donnell et al., 2009). Sometimes listening to feelings seemed to involve more than us and our participants, but others – parents, uncles, grandparents and children. Writing about transference (and countertransference) Bondi (2005) wrote that ‘we carry the affective impress of our earliest patterns of relating into all of our subsequent relationships’ (2005, 440). We sometimes felt and listened to the presence of these relationships in what was said and not said in our discussion groups, in the ways in which we experienced individuals and groups that we did not fully understand and are difficult to articulate. Relationships with significant others hovered on the edges of meetings, taking shape around a plate of homemade biscuits brought to a meeting or the image of an uncle's naked, muscly torso shown around. These relationships were sometimes far from straight forward, occasionally painful to listen to and difficult to follow. Mother-child relationships contributed to, we think, (our experiences of) the conversation with Omar, introduced at the start of the paper. Omar was 18, born in Afghanistan, moving to Milton Keynes in 2009. Also involved in the conversation was Salima, who was 24, born in Somalia, moving to Kenya and then Milton Keynes in 2006. Talk flowed; Omar was confident and jaunty, Salima poised and friendly. How they talked sat uncomfortably with our experiences of what they actually said. Omar remembered a loving mother who taught him at home, but his mother slid away from view when he talked about his life in Milton Keynes. When Salima talked about her mother, it turned out she was talking about her step mother, her own mother had ‘left’ when she was three, ‘during the wars’. Omar focused on the transience of his life, travelling with his father, his impressive body building uncle, living in foster care when he moved to Milton Keynes, leaving that care and living in temporary accommodation, a room in a house with strangers. His talk was cheerful, fast, buoyant, fuelled by working out at the gym, working hard at College, working in a shop, a girlfriend he enjoyed spending time with, coping despite everything. He came across as self-reliant, hopeful and upbeat, although this is difficult to convey when confronted with the words (Back, 2012, 2007): ‘Because the thing is, like, when I was living with my family I wasn't – I wasn't there with my family, to be honest. I was like all the time travelling around, with my dad. And being brought up like that. And then, I mean, at the moment for me – ‘cause everybody's different. Everybody has different life situations. For me, it doesn't change anything. I'm just the only person I was. And I mean, I'm getting better but not worse. I mean, the life is being with my family. Sometimes I just remember I have a mum – I had a mum, yeah, and just don't go into it. I mean, I missed it but not much. Probably sometimes when I'm alone or when I see an old lady, oh yeah, I have a mum. It reminds me of, like yeah, somebody that. But I don't remember anybody. Or I don't miss anybody. Because I mean, you miss people when you're free, you don't have a job, you're bored and suddenly like everything comes up and you're thinking of things. I have a very tight timetable, I get up at six o'clock in the morning (….) Go to – go to the college, come back from college, go to work, finish work. Go to straight away to the gym, do exercise, come back home, study, cook, eat, clean and sleep 12 or 1am in the morning, five, seven hours sleep. And that's what I do. I'm forever busy, like I don't really remember anybody or talk to anybody. If somebody wants to call me, call me and if I want to talk to them I will like try and speak with them. If they don't…’ (Omar) To some extent listening involves participants allowing researchers to listen and at the time it felt like Omar opened and closed down multiple channels of listening making our understanding messy, stilted and broken. At times he felt playful. He was hard to follow, opening up and then withdrawing, whilst talking all the while. He introduced us to the image of his uncle on his phone, and then placed this to one side, on the table. The screen went black, but the uncle somehow remained. The glossy, upbeat talk of survival and success reiterating pervasive neo-liberal talk in Britain was painful to listen to when juxtaposed with ‘I have a Mum’, ‘I had a Mum’, ‘don't go into it’, ‘I don't miss anybody’, ‘you miss people when you're free’. What we experienced was caught up in Omar's incomplete stories, growing up quickly, maturity, coping, girlfriend, work and college alongside transience, loneliness, isolation, separation and missing his mother, amongst others. We were listening to fragments and a ‘jumble of emotion’ that defied conventional genres of narrative (Pratt, 2010). Perhaps our experiences were also caught up in what we brought to our listening, such as, for example, Katy's relationship with her son and her recently born daughter and an imagining in around what their loss might involve (Aitken, 2001). She found herself not only empathising with Omar and Salima, but with their mothers too. Katy was miserable after the meeting, packing up and going home. We have focused on this moment in our research to show that listening is an intersubjective experience, involving us, our participants and our relationships with others who shape our experience of research which can sit differently to participant expressions of emotion. Listening involves recognising the creative, fractured, intersubjective process of listening that underpins understanding that is ‘imperfect’ and ‘faltering’. As Bondi (2014a,b) wrote: ‘(E)mpathy does not generate direct or perfect apprehension of the subjective experience of another. Rather it requires effort and is always imperfect and faltering. However much of the experience of the other is accurately recognised, empathy also entails acknowledging that the effort to understand can only ever yield an imperfect grasp of what the other feels’ (Bondi, 2014a, 50). The time involved in listening…… (to) transcripts ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our final reflection is that listening takes time. Listening often continues long after research has supposedly finished. Feelings and conversations stay with us, live with us and we mull over research moments involving, for example, Omar and Salima (Bondi, 2014a; Bennett, 2009). Although this might be stretching his take on ‘intimate distance’ which unsettles the drive for intimacy and rapport, as Steve Pile (2010) wrote ‘to see the emotional dynamics underlying any research situation, you'll need appropriate distance: a (conditional) balance of intimacy and distance’. We have certainly found that we need time (and space) that extends beyond end points of research contracts to allow our research to sink in. Listening to (our experiences of) the transcripts also takes time. We are surprised by how different the transcripts can feel compared to what we experienced at the time. Some of this is not simply because of what they lack in terms of the buzz of the context (Back, 2012, 2007) but what they add. Transcripts can be provocative. What surprises us is how we are moved at different points and in unexpected parts of the transcript. Rachel Thomson et al. (2012) wrote evocatively about letting the text speak, listening after the event, allowing ourselves to be provoked by the text. When we return to our transcripts what we hear is shaped by identity, subjectivity and where we are in our lives. As Steve Pile (1991) wrote: ‘reflection is profoundly intersubjective, operating not only within the relationship between the researcher and researched but also within wider social and personal relationships’ (1991, 465). Listening over time involves the accumulation of (partial) understandings as relationships develop, texts, diaries and recordings materialise, as we provoke those texts and they us and we remember, reflect on and relive encounters. Provocative transcripts and field notes stem not only from us listening to them, but them listening to us too. Explaining acousmatic texts, Jacques Daignault (2005) wrote: ‘The text, if we need to remind ourselves, is not the printed thing that the author publishes but each of the readings of the printed page and even the sum of all these readings which always contain a part added by the reader.’ (Daignault, 2005, 6). (Acousmatic) texts are the sum of our readings, dynamic and thought provoking, open to unexpected thoughts that mediate our reading of them, created through our reading, listening to what we bring to them. Although the printed words in our transcripts do not change, just the paper yellows, we bring to them developing, maturing selves and new experiences which the transcripts absorb as they listen to us whilst we read them. They absorb our individual reading and response to them and what we read is shaped by how other members of the team have read them. Finally, although research projects and contracts have end points, some of the relationships we develop with participants and groups do not. Our field notes are read through the lens of on-going connections with people and places. One of us, for example, became a Friend of Knighton Park and involved in its activities, another joined a writing group long(er) term, writing the following in her field notes during the research: ‘Larry said he'd really like for me to continue attending, as it was nice to have me in the group, but I should consider this before bringing in an object (for a community project) [….] I openly admitted to him that this was a funny position to be in – you get attached […] and it's not easy or even appropriate to divide your engagement as a researcher from your engagement as a person’ (28/5/13 Writing Group, Hackney). ‘After we went around reading out our pieces, during which there was a lot of laughter and banter, the woman from the museum came by with a camera to take photos of us with our objects. I ended up bringing something in after all – a book of poetry entitled 'hackney, my hackney' from 2004 which includes one of my poems. I found it and it felt like it fit so perfectly into the whole thing that I had to bring it along, although I know perhaps committing to attending the group long-term is foolish’. (30/5/13 Writing Group, Hackney) Sometimes listening can be complicated over time by our developing relationship with groups and places that extend well beyond the end dates of projects. Although projects end, listening never quite does.","In this paper we have attempted to reflect on our listening in a way that we hope researchers and students find useful, drawing in particular upon three elements of listening that have especially resonated with us: listening whilst doing, the feelings that mediate it and the time involved in listening. In trying to listen better we did not invent new methods or prioritise one method over another, but we did draw upon the writing and ideas of others on embodied listening, empathy and acousmatic texts to plan and reflect on our listening. In this concluding section we finish with thoughts on the process of listening and the kind of insights that our listening might provide. Our first strand of concluding thoughts broadly concerns the process of listening that shaped our participant observation, interviews and repeat, in-depth discussion groups. Our participant observation involved sliding into the worlds of others and taking part in their activities. Sometimes these activities involved little in the way of chat or conversation and so listening involved attending to much more than this. Our field notes and reflections in this paper reveal a pattern and process to our listening. We often drew upon the mechanics of an activity in which we were involved to initially anchor us in the field, detailing what we were doing, where, with whom and how this made us feel. We used our eyes, ears, feelings, beating hearts - and more besides – to gauge moods and atmospheres as we grappled with often unfamiliar sites and settings and people who we did not know that well. We described people, encounters, interactions and contexts and explored how we experienced these. Our methods involved channelling our listening through ourselves, feeling our way into situations and conversation, working with our feelings, imagining in, empathizing whilst mindful of differences around ethnicity, age, gender and class (Bondi, 2003). Sometimes we shared our feelings with our participants, emoting, for example, whilst running or in a less direct way – ‘Are you alright (Ayo)?’ Objects mattered even more than we expected, supporting our listening as we gripped our pens whilst jotting notes during group discussions or filled water bottles in six form canteens. Objects contributed texture to our listening but were sparky in their own right too, hot-wiring social relations and interactions as, for example, a crisp packet was thrown around prompting movement, laughter and unsettling groups of students and the disciplined space of a six form canteen. As individual members of a research team, we do listen differently and perhaps, as Forsey (2010) suggests, some of us are more aural whilst others are more visual in our listening, but this is too much of a simplification. Our listening was also shaped by our individual life experiences, our relationships and where we were in our lives, but also by more immediate experiences, such as our journey to the fieldwork site. Our listening involved our feelings as we tried to make sense of what we were experiencing, what was being communicated in the context, contours and conversation in which we were immersed. The process of listening involves time that stretches well beyond the end dates of projects and research contracts. Our listening matures and evolves alongside our feelings and experiences of fieldwork. Transcripts listen to us and develop through our listening to them over the course of time and around our memories of fieldwork and the unfolding of our lives. Listening to transcripts involves attending to the understanding and experiences that other members of the research team bring to them. Our second strand of concluding thoughts concern the insights that our listening brought to our research (see also MacKian, 2009). Sometimes these insights are more obvious, when, for example, we listen to feelings that are surprising because they do not seem to correspond with how someone is speaking. Returning to Omar introduced at the start of the paper, his buoyant, happy talk, sat differently to the content of what he said which sat differently again to how one of us was feeling while he was talking. There was more than one story being told and listening to feelings was a way into the complexity of what was being communicated. Focus on feelings and a sensuous, embodied approach to listening decenters talk, without undermining it, knitting it into a mix of atmosphere, interaction, stuff, happenings, context, sensations – and more besides – requiring attention. For us the point of this kind of listening was to engage with the complexity of places, interactions and encounters that shape multiculture which research participants might not (want to) put into words in interviews and group discussions. Our listening might also bring insights about which we are less aware, but emerge in how we evoke people and places (see, for example, Rogaly and Qureshi, 2013). The kind of listening discussed in this paper encompassing doing, feelings and time creates insights that get under the skin of researchers shaping a less conscious, embodied, evolving knowing."],["Reading and listening to stories fosters vocabulary development. Studies of single word learning suggest that new words are more likely to be learned when both their oral and written forms are provided, compared with when only one form is given. This study explored children's learning of phonological, orthographic, and semantic information about words encountered in a story context. A total of 71 children (8- and 9-year-olds) were exposed to a story containing novel words in one of three conditions: (a) listening, (b) reading, or (c) simultaneous listening and reading (“combined” condition). Half of the novel words were presented with a definition, and half were presented without a definition. Both phonological and orthographic learning were assessed through recognition tasks. Semantic learning was measured using three tasks assessing recognition of each word's category, subcategory, and definition. Phonological learning was observed in all conditions, showing that phonological recoding supported the acquisition of phonological forms when children were not exposed to phonology (the reading condition). In contrast, children showed orthographic learning of the novel words only when they were exposed to orthographic forms, indicating that exposure to phonological forms alone did not prompt the establishment of orthographic representations. Semantic learning was greater in the combined condition than in the listening and reading conditions. The presence of the definition was associated with better performance on the semantic subcategory and definition posttests but not on the phonological, orthographic, or category posttests. Findings are discussed in relation to the lexical quality hypothesis and the availability of attentional resources. --------------------------------------------------------------------------------","Vocabulary development starts during infancy and is a lifelong endeavor; children and adults acquire new words, and specify existing lexical representations, throughout the lifespan (Nagy, Anderson, & Herman, 1987). The majority of words are not acquired through direct instruction but rather incidentally from conversations, television, and texts (Akhtar, 2004; Elley, 1989; Henderson, Devine, Weighall, & Gaskell, 2015; Houston-Price, Howe, & Lintern, 2014). The current study explored how children learn new words when they are exposed to them incidentally in stories. To our knowledge, this is the first investigation of whether children show greater word learning from listening to, reading, or both listening to and reading stories. Many studies have shown that exposure to stories fosters vocabulary development in children (Henderson et al., 2015; Nagy et al., 1987; Ricketts, Bishop, Pimperton, & Nation, 2011; Wilkinson & Houston-Price, 2013; Williams & Horst, 2014). Suggate, Lenhard, Neudecker, and Schneider (2013) compared the word learning shown in three story presentation conditions: independent reading, listening to an adult reading the story, and listening to an adult telling the story in his own words. In the listening conditions children (8- to 10-year-olds) were exposed to the spoken forms (phonology) and meanings (semantics) of new words but not their written forms (orthography), whereas in the reading condition they encountered the words’ written forms (orthography) and meanings. The children who listened to the stories were more likely to demonstrate knowledge of the new words’ meanings than the children who read the stories, suggesting that oral presentation is more beneficial for vocabulary learning in school- aged children than written presentation. In contrast to this result, studies of adults learning English as a second language tend to show that participants acquire new words more easily if presented with material in written form rather than oral form (Brown, Waring, & Donkaewbua, 2008; Sydorenko, 2010). Similarly, studies exploring memory for word lists or verse show better performance for written material than for oral material in both adults and children (Hartman, 1961; Menne & Menne, 1972). Because the primary medium for vocabulary acquisition is oral language, it makes sense that oral presentation may be the preferred and easiest method for acquiring vocabulary early on, but as reading ability improves, children become better at learning from written texts. There is reason to suppose that both listening to and reading a story at the same time will be maximally beneficial for word learning. Studies using e-book presentations show that presentation in both modalities is more beneficial for vocabulary acquisition than simply listening to the book read by an adult (Shamir, Korat, & Fellah, 2012). Other work has also found that access to orthographic forms promotes oral vocabulary learning, an effect referred to as “orthographic facilitation” (e.g., Hu, 2008; Ricketts, Bishop, & Nation, 2009; Ricketts, Dockrell, Patel, Charman, & Lindsay, 2015; Rosenthal & Ehri, 2008). In these studies, children were taught phonological forms and semantics either with or without orthography; greater learning of phonology, semantics, and orthography was seen for items where orthography was provided. To date, studies investigating orthographic facilitation all have employed a direct instruction approach to teaching new words. Whether these findings generalize to an incidental learning context, where children’s attention is not explicitly drawn to the new words, needs to be explored. Further evidence that simultaneously listening to and reading stories leads to better learning than a single modality of presentation comes from a study by Rosenthal and Ehri (2011). In that study, children read stories that contained novel words silently, pronouncing half of the new words aloud when they encountered them. Semantic and orthographic learning was greater for words that had been pronounced, demonstrating “phonological facilitation.” Although in this study words were embedded in passages, thereby simulating incidental learning from reading, attention was drawn to the target words (through underlining), and only these words were pronounced aloud. In the current study, to provide a more realistic learning context, new words were not highlighted and presentation modality was manipulated for the story as a whole rather than for individual words. To date, listening and reading conditions have not been compared with a combined listening and reading condition in relation to school-aged children’s learning of vocabulary in their first language. However, studies of adult second-language learners typically find superior learning when material is presented in written or dual modality format rather than in spoken format (Brown et al., 2008; Sydorenko, 2010). For example, Brown et al. (2008) presented three stories in three different modalities (listening, reading, and combined) to 35 Japanese adult students and found greater semantic learning in the reading and combined conditions compared with the listening condition. In line with these results, studies investigating the impact of multimodal presentation on memory for lists of words or verse found superior performance in the combined and reading conditions compared with the listening condition in children aged 8 or 9 years and adults (Hartman, 1961; Menne & Menne, 1972). In summary, evidence of orthographic and phonological facilitation for vocabulary acquisition has so far come from studies that present words in isolation, rather than in context (Ricketts et al., 2009; Rosenthal & Ehri, 2008), or from studies that confine the dual modality of presentation to the words of interest, rather than to the narrative as a whole (Rosenthal & Ehri, 2011). Evidence for the beneficial effect of dual presentation modality compared with single modality presentation has also come from studies with adult second-language learners (Brown et al., 2008). Therefore, questions remain about whether such facilitation effects are seen when school-aged children are learning new words in their first language and when the presentation modality is extended to the entire story in which words are embedded. Why might there be an advantage for a combined presentation modality over a single presentation modality? One possibility is that providing both oral and written forms frees up attentional resources during encoding, meaning that resources can be allocated to story comprehension and to the encoding of new word meanings. This idea is related to Ehri’s (2014) views on how good decoders use their strong knowledge of the link between orthography and phonology to generate phonology automatically while reading, leaving more resources for text comprehension and the learning of words’ meanings. It also resonates with cognitive load theory in multimedia learning (Mayer, Moreno, Boire, & Vagge, 1999), which states that situations that reduce cognitive load are more conducive to learning. According to cognitive load theory, presenting information in two modalities frees resources by increasing working memory capacity. This process occurs online as new information is encountered and while a representation is being created (Mousavi, Low, & Sweller, 1995). On this view, then, presenting both phonological and orthographic forms may free cognitive resources so that more attention can be paid to comprehension and to encoding word meanings. Notably, this framework focuses on the conditions that facilitate learning, suggesting that presenting information in a dual modality creates connections between two representations. It does not, however, probe the nature of the representations created (Mayer et al., 1999). An alternative framework for interpreting the benefits of simultaneous bimodal presentation is provided by the lexical quality hypothesis (Perfetti & Hart, 2002). This hypothesis posits that lexical representations that include both phonological and orthographic information are of “higher quality” than representations that contain only phonological or orthographic information. Although the lexical quality hypothesis focuses on existing lexical representations rather than their acquisition, it is consistent with the idea that a combined presentation will support the building of better quality lexical representations that, therefore, can be more readily accessed when they are encountered at a later stage. In addition to story presentation modality, the content of the story and the kinds of information provided within the story are likely to play a role in children’s learning of new words. While reading, children use contextual cues to infer semantic information about new words even when these cues convey minimal information (Nagy et al., 1987). Whereas general context supports learning of the broad category of a new word, more constraining contexts enhance learning of a word’s specific features (Ricketts et al., 2011). In addition, when words are encountered aurally, presenting a definition alongside the new word is beneficial for semantic learning (Dickinson, 1984; Penno, Wilkinson, & Moore, 2002; Wilkinson & Houston-Price, 2013). There are important differences in the information conveyed by contextual cues and definitions. Contextual cues typically provide general information about a word’s meaning and are spread through the text (Nagy, 1995). Definitions provide detailed semantic information about the word and typically occur alongside its form. If definitions are considered a highly constraining context, they should allow more specific information about a word to be extracted than the more general categorical information revealed by less constrained contexts (Ricketts et al., 2011). Although previous studies have shown that definitions elicit greater semantic learning than general contextual support in both recognition tasks (Wilkinson & Houston-Price, 2013) and production tasks (Justice, Meier, & Walpole, 2005), previous research has not explored the nature of the semantic learning shown in such conditions, a further aim of the current study. The current study ~~~~~~~~~~~~~~~~~ Children aged 8 or 9 years were exposed to a story containing eight new words. This age range was chosen to ensure a range of reading abilities in a population of children used to reading and understanding written texts. Children were divided into three groups, with one group listening to the story (listening group), one group reading the story (reading group), and one group reading and listening to the story simultaneously (combined group). Half of the words were accompanied by a definition the first time they were presented, and the other half of the words were not. The story was presented twice over 2 weeks to promote learning via repetition and allow for sleep-related consolidation (e.g., Henderson, Weighall, & Gaskell, 2013) and to fit in with the school timetable. After the second story presentation, children’s knowledge of phonological forms, orthographic forms, and meanings was assessed in a series of posttests. In phonological and orthographic posttests, children were required to recognize correct spoken or written forms from two alternatives. In three semantic posttests, children were asked to identify the words’ categories, subcategories, and definitions. By using these tasks, we were able to capture acquisition of different aspects of semantic information about the given words. From a practical perspective, this investigation serves to establish how best to promote vocabulary acquisition in school-aged children and whether children with different ability levels are facilitated by different presentation modalities. From a theoretical perspective, the results assess the hypothesis that different presentation modalities facilitate the acquisition of words of higher or lower lexical quality. The framework provided by the lexical quality hypothesis (Perfetti & Hart, 2002) has previously been applied to evidence that better readers create higher quality representations of new words (Perfetti, 2007), but it has not been used to make clear predictions about the conditions under which words are optimally encoded or retained. This study explored whether word representations created through different presentation modalities vary in their lexical quality. The study addressed several hypotheses regarding orthographic, phonological, and semantic learning. In relation to dual versus single modality presentations, it was hypothesized that the combined condition would elicit superior orthographic, phonological and semantic learning (Ricketts et al., 2009; Rosenthal & Ehri, 2008; Rosenthal & Ehri, 2011) compared with the other two conditions. Regarding semantic learning, given the conflicting evidence, clear predictions could not be made as to whether the listening group would outperform the reading group in learning words’ meanings, as suggested by the first language literature (Suggate et al., 2013), or whether the reading group would show an advantage, as suggested by the second language and memory literatures (Brown et al., 2008; Menne & Menne, 1972). In line with previous research on word learning from stories, it was expected that definitions would promote semantic learning (Penno et al., 2002; Wilkinson & Houston-Price, 2013), particularly in tasks assessing learning of specific features of the words’ meanings. Individual differences might also constrain children’s learning of new words from spoken and written story exposure. Therefore, measures of children’s abilities were collected. Based on previous research, it was expected that reading accuracy would predict learning of orthographic forms in the reading condition (Ricketts et al., 2011) and that vocabulary knowledge, reading accuracy, and reading comprehension would predict semantic learning in the reading condition (Cain, Oakhill, & Lemmon, 2004; Ricketts et al., 2011), whereas oral vocabulary ability should predict semantic learning in the listening condition (Lin, 2014; Penno et al., 2002; Senechal, Thomas, & Monker, 1995). Given that previous studies have not investigated monolingual children’s word learning when children are simultaneously reading and listening, it was unclear which abilities might influence learning in the combined condition. We anticipated that because children could rely on both oral and written modalities, oral vocabulary, reading accuracy, and reading comprehension all would influence performance in the combined condition.","A total of 71 children aged 8 or 9 years participated in the study (Mage = 9.03 years, SD = 0.31; 28 boys). Participants were recruited from four primary schools in England. Ethical approval was obtained from the first author’s institution, and informed parental consent was received for all participants. All children had normal or corrected-to-normal vision, and teachers confirmed an absence of learning or neurological disabilities. All children were monolingual native English speakers. Background measures Children completed background measures in one session prior to the word learning task. All were standardized assessments and were administered according to test manual instructions. Nonverbal abilities were measured using the Raven’s Coloured Progressive Matrices (CPM; Rust, 2008), a pattern completion task in which participants solve visual diagrammatic puzzles using analogies or inferences (split-half reliability reported in the test manual = .97). Oral language abilities were assessed using the British Picture Vocabulary Scale–Third Edition (BPVS; Dunn, Dunn, & National Foundation for Educational Research, 2009), a receptive vocabulary measure in which children need to choose the correct picture for a given word among four alternatives, and the Understanding of Spoken Paragraphs (USP) subtest of the Clinical Evaluation of Language Fundamentals–Fourth Edition (CELF-4; Semel, Wiig, & Secord, 2006), a test in which children listen to several short passages and answer questions about them. This test assesses both oral language comprehension and sustained oral attention (test–retest reliability reported in the test manual = .80). Nonword reading was assessed using the Phonemic Decoding Efficiency subtest of the Test of Word and Reading Efficiency (TOWRE; Torgesen, Wagner, & Rashotte, 1999), in which children read as many nonwords as they can in 45 s (test–retest reliability reported in the test manual = .90), and word reading was assessed using the Single Word Reading Test (SWRT 6–16; Foster, 2007), which is an untimed word reading task with words in sets of increasing difficulty. The York Assessment of Reading for Comprehension (YARC; Snowling et al., 2009) was used to assess text reading accuracy and reading comprehension. In this task, children read two passages aloud and reading errors are noted. Following reading, they answer comprehension questions indexing knowledge of literal content and inferential processes (parallel form reliability reported in the test manual for reading accuracy: all rs > .70; Cronbach’s alpha for reading comprehension scores from two passages: all αs > .70). Design Story presentation modality was manipulated between participants such that children were assigned to either the listening, reading, or combined condition in order to form three comparable groups matched on key background measures. The groups did not differ on gender, χ2(2) = 0.09, p = .957, reading, oral language, or nonverbal abilities (see Table 1). Details of the background measures used to match groups are included above. Children from each school were equally distributed across conditions. The presence of a definition was manipulated within participants, with all children encountering half of the words with a definition and half without a definition. Stimuli See Appendix A for stimuli. Words Target items were 8 low-frequency English concrete nouns (e.g., destrier, hauberk), each belonging to one of eight categories (e.g., animal, object). From an initial set of 36 words related to historical periods (Roman Empire, Norman Period, and British Empire in India), we selected 8 words from the Norman Period and embedded these target items into a meaningful narrative. In addition, 8 control words were selected from other historical periods and were included in the pretest and all posttests to control for prior knowledge (see below for details of these tasks). The pretest confirmed that knowledge of target and control words was negligible (see below). Given that control words were not presented in the story, it was expected that any difference in performance between target and control words could be ascribed to story presentation. Target and control lists were composed of words of the same category, matched for length and frequency using the SUBTLEX–UK database (Van Heuven, Mandera, Keuleers, & Brysbaert, 2014) (Mann–Whitney tests, all ps > .200). Target and control words were also matched on pilot data collected for a previous study showing the proportion of adults who spelled and pronounced each word correctly and the proportion of children who correctly categorized each word at its first encounter (Mann–Whitney tests, all ps > .200). Definitions The definition for each target and control word comprised information about the word’s subcategory and a further phrase to specify it. For example, for destrier, the definition “a horse used for fighting” comprises both the subcategory information “a horse” and the specific characteristic “used for fighting.” Words were divided into two lists of four items (Word Lists A and B) for counterbalancing purposes. The words in the two lists were paired for length, frequency (Van Heuven et al., 2014), the aforementioned results from pilot studies, the length of the definition, and the distance between the first and second mentions of the word in the story. Mann–Whitney tests revealed no difference between the words in the two lists on these measures (all ps > .200). Stories A story set during the Norman Period was written for this study to include all target words. Each target word was repeated three times in the story. Contextual references to the meaning of each word were included to ensure that the children could learn the meaning of the words from context alone. Two different versions of the story were created so that the definitions of either Word List A or Word List B were included as part of the text following the first mention of the word. The two versions were matched for length (1346 and 1347 words, respectively), Flesh Reading Ease (84.3% and 84.5%, respectively) and Flesh–Kinkaid grade level (5.2). Pilot data ensured that the stories were written at an appropriate level for children in this age range. Recordings of a female native English speaker reading the stories were created, and booklets of seven pages that contained only the text of the stories, written in Calibri 14-point font, were prepared. Pilot studies established how long, on average, children took to read the stories independently (M = 548 s, SD = 172); the recordings of the adult storyteller were controlled to match this (M = 507 s, SD = 2).","Fig. 1 summarizes the study procedure. Assessments were completed in a quiet room within schools. Prior to the word learning task, children completed the word knowledge pretest and were administered the background measures. The word learning task was completed in two sessions lasting between 30 min and 1 h. In the first session, children were exposed to the story for the first time and completed a story comprehension task. In the second session, they were exposed to the story a second time and completed the same story comprehension task as well as the phonological, orthographic, and semantic posttests. The first and second sessions were completed 1 week apart. The story comprehension task and all posttests were delivered through a laptop computer using E-Prime software (Schneider, Eschman, & Zuccolotto, 2002). The phonological and orthographic posttests were completed first, in counterbalanced order, followed by the semantic posttests. Given that the semantic posttests involved presenting the novel words in spoken and written forms, these were completed last to avoid any contamination from these to the phonological and orthographic tasks. We did not expect that there would be cross-contamination between orthographic and phonological posttests and the semantic posttest because the former provided no semantic information. The semantic posttests were multiple-choice format and presented in a fixed order, with the category recognition task first, followed by the subcategory recognition task and finally the definition recognition task. This order minimized any impact of previous semantic tasks on later semantic tasks. For each posttest, on-screen instructions followed by four practice trials ensured that children understood the demands of the task. Target and control words were presented with item order randomized, and accuracy was recorded for each trial. Word knowledge pretest Children were asked to define the target and control words. Here, 3 participants showed preexisting knowledge of one target word (2 children knew pottage and 1 child knew motte); the remaining 68 children had no knowledge of any target words. In addition, 3 children showed preexisting knowledge of one control word (2 children knew catacomb and 1 child knew verandah); again, 68 children had no knowledge of any control words. Thus, preexisting knowledge of control and target words was similarly scarce. Analyses that excluded these participants yielded the same pattern of results as those reported. Thus, all participants were retained in the analysis. Story exposure Participants were asked to listen to (listening group), read silently at their own pace (reading group), or listen to and read (combined group) the story, after which they were told they would be assessed on their comprehension of the story. No mention was made of the presence of the target words. Children were presented with a version of the story that contained definitions for either Word List A or Word List B. The oral presentation of the story was delivered through headphones connected to a laptop with a blank screen. For the written presentation, children read the story from a booklet. In the combined condition, children received both presentations simultaneously. Story comprehension task After each story exposure, children were asked four multiple-choice questions to ensure that they had paid attention to the story (e.g., “Fred wanted to reach the king’s castle. How long did he think the journey would last?” ; correct response: A month; foils: A day/A week/One hour). Each question, along with an array of four potential answers, was presented both orally and visually via a laptop. Pilot data provided by 12 children of the same age as participants confirmed that, without exposure to the story, children were unlikely to guess the answers to questions at above chance levels. Phonological posttest In this task, two dinosaurs were presented sequentially on a computer screen: one providing the correct spoken form of a word and the other providing an incorrect word form (distractor). Pilot data collected for an earlier study were used to generate distractors for this task. Adults were asked to pronounce each written target and control word, and the most frequent mispronunciations were used as distractors for most words. If no mispronunciations were produced (for 4 words), alternative pronunciations were created. For example, the distractor for “hauberk” (hɔːbək, first vowel as in horn) was “hɑ:bək” (first vowel as in heart). Children were instructed to choose the dinosaur that “said the word best” by pressing the corresponding button on the keyboard. The dinosaurs appeared in turn for 2 s in an alternating loop until an answer was provided. The association between the two dinosaurs and the correct answer was randomized. Both target and control words were presented in this task (16 items). Test–retest reliability for these items was obtained from pilot data collected from a separate group of 57 children of the same age (8–9 years) who completed the tasks twice, with a 1-week interval between test sessions. Pilot data were available for 13 items used in the current study (all 8 target words and 5 control words). Test–retest reliability for these items was r(57) = .71. See Appendix B for stimuli. Orthographic posttest As in the phonological task, children saw two dinosaurs sequentially, this time accompanied by a letter string, and were asked to choose the correct spelling of the given target or control word (by choosing the dinosaur who spelled the word best). The dinosaurs, alongside their spelling option, appeared in turn for 3.5 s in an alternating loop until an answer was provided. Pilot data provided by adults in an earlier study were used to generate distractor spellings. The most common misspellings of orally presented words were used; when no misspelling was produced (for 2 words), alternative spellings were created. For example, the distractor for “hauberk” was “horberk.” The association between the two dinosaurs and the correct answer was randomized. Both target and control words were presented in this task (16 items). As for the phonological task, test–retest reliability was computed from pilot data available for 13 of the items: r(57) = .57. Although this value highlights a less than excellent relationship, we deemed the measure to be sufficiently reliable because an amount of variability is to be expected in representations of newly learned words over time. See Appendix B for stimuli. Semantic posttests The three semantic subtests followed the same format. In each task, the children were asked to choose the correct alternative from an array of four choices pressing the corresponding button on the keyboard. First, a category recognition task assessed recognition of the category of the new word (e.g., clothing) among three of the categories of other target words (e.g., part of a house, job, animal). At the beginning of each trial, a target or control word was presented in spoken and written forms; the written form appeared at the top of the screen. The alternative category labels appeared one at a time both orally and in written form in randomized positions on the screen (see Fig. 2). The second and third tasks followed the same procedure. The second task assessed recognition of the subcategory of the word (e.g., for an item of clothing, the four alternatives were different kinds of clothing), and the third task assessed recognition of the word’s definition when presented with three distracter definitions (each differing from the correct choice by one feature). The definitions were not identical to those presented in the stories but were rephrased. Both target and control words were presented in all three semantic posttests (16 items). As for the other posttests, test–retest reliability for category recognition was computed from pilot data available for 13 of the items: r(57) = .76. Table 8 lists the correct responses for each semantic task. Story comprehension ~~~~~~~~~~~~~~~~~~~ A story comprehension task was used to confirm that children paid attention to the story. Participants completed it twice, once after each story exposure, giving a maximum score of 8 and chance-level performance of 2. One-sample Wilcoxon signed-rank tests indicated that all groups performed better than chance on this task (see Table 2). A Mann–Whitney test showed that children in the combined group outperformed children in the reading group (U = 140.50, p = .003) but not children in the listening group (U = 253.50, p = .461), whereas there was a nonsignificant trend for the listening group to outperform the reading group (U = 192.00, p = .069). This suggests that hearing the story promoted comprehension. Approach to data analysis ~~~~~~~~~~~~~~~~~~~~~~~~~ For each posttest, two sets of analyses were carried out. The first set (see Table 3; see also Table 6 below) compared the recognition of target and control words and compared each of these with chance (50% for phonological and orthographic tasks and 25% for semantic tasks). The second set of analyses used a mixed-effects modeling approach to explore our hypotheses relating to (a) presentation modality (listening, reading, or combined group), (b) definitions, and (c) individual differences within each of the three experimental tasks in turn. Because the groups were matched on age and all background measures, we included in the analyses only measures for which specific hypotheses were considered: vocabulary knowledge (BPVS raw score) and a composite reading accuracy score in all analyses, and reading comprehension (YARC reading comprehension ability score) in semantic task analyses. The composite reading accuracy score was formed by merging TOWRE, SWRT, and YARC text reading accuracy scores into a factor using the regression method (M = 0.00, SD = 1.00, range = −2.89 to 2.24). The three groups did not differ significantly on this measure, H(2) = 0.02, p = .992, and the creation of this factor was supported by high correlations between its constituent measures (all rs > .80, all ps < .001). See Appendix C for correlations between background measures. Because the data collected on each trial were binomial (a child could choose either the correct or incorrect alternative, obtaining a score of 1 or 0), mixed-effects models were conducted using generalized linear mixed models for binomial data (Jaeger, 2008), using the function “glmer” from the package lme4 (Bates, Maechler M., & Walker S., 2014), the function “mixed” from the package afex (Singmann, Bolker, Westfall, & Aust, 2016), and the function “lsmeans” from the package lsmeans (Lenth, 2016), computed with the software R (R Core Team., 2014). Each of the 71 children provided eight responses to target words on each task; this score was the dependent variable in each analysis. For each dependent variable, an initial model included a maximal random effects structure that captured our experimental design (Barr, Levy, Scheepers, & Tily, 2013). This entailed the random intercept terms for both participants and items and the random slopes terms for participants and items that relate to our repeated-measures manipulation: presence of a definition. However, models including random slopes were prone to nonconvergence; therefore, the simpler and convergent models are reported (Bates, Kliegl, Vasishth, & Baayen, 2015). We then compared these “empty” models (using pairwise likelihood ratio test comparisons; Barr et al., 2013) with models that also included performance on control words as a control variable and the hypothesized fixed effects: group (combined, listening, or reading), presence of definition (definition present or definition absent), and specific background measures. Scores on control word trials were included in models to control for general task effects, such as the ability to recognize word-like phonological forms, irrespective of any in-task learning. Further analyses were performed to explore whether results differed for the two sets of target words by entering word set as a further fixed factor. The pattern of results was identical, and set was not a significant predictor within the models; thus, these models have not been reported. All continuous factors were centered around the mean for analysis. Hypothesized interactions were included one at a time in the model with all fixed effects and were retained only if significant. The interactions between group and the background measures were separately introduced to test whether any background measure had a differential effect on performance in each task, depending on presentation modality. Estimates of fixed effects and interactions for the final models are reported in Tables 4, 5, and 7 below. Phonological task ~~~~~~~~~~~~~~~~~ The mean number of phonological forms correctly recognized by the children in each group is presented in Table 3 along with analyses comparing target and control word performance with each other and with chance. Target word performance was greater than control word performance and was above chance for all groups. Control word performance was also greater than chance for the combined group, but this was not the case for the listening and reading groups. The fixed-factor model for the phonological task significantly improved fit compared with the empty model, χ2(6) = 48.53, p < .001. Interaction terms did not significantly improve model fit. The final model (Table 4) indicates reading accuracy and control word scores as significant predictors of performance on the phonological task; phonological learning was greater for better readers and those better able to identify the phonological forms of control words. Most important, phonological learning did not differ across groups. Orthographic task ~~~~~~~~~~~~~~~~~ The mean number of orthographic forms correctly recognized by the children in each group is presented in along with analyses comparing target and control word performance with each other and with chance. In the Table 3orthographic task, target word performance was significantly greater than control word performance and was above chance for the reading and combined groups. For the listening group, target word performance was not significantly greater than control word performance or above chance. Control word performance was not above chance for any group. The fixed-factor model for the orthographic task significantly improved fit compared with the empty model, χ2(6) = 31.53, p < .001. Interaction terms did not significantly improve model fit. The final model (Table 5) indicates group and reading accuracy as significant predictors of orthographic performance, with children in the combined and reading groups showing better performance than children in the listening group (who did not show significant orthographic learning) and greater orthographic learning associated with greater reading accuracy. Semantic tasks ~~~~~~~~~~~~~~ The mean number of words correctly recognized by each group in each semantic task is presented in Table 6 along with analyses comparing target and control word performance with each other and with chance. In all three semantic tasks, all groups performed significantly above chance on target words and significantly better on target word trials than on control word trials. Control word performance was not significantly above chance in any task for group. The final model for each semantic task is presented in Table 7. Category recognition The fixed-factor model significantly improved fit compared with the empty model, χ2(7) = 55.34, p < .001. Interaction terms did not significantly improve model fit. Group, presence of definition, reading accuracy, vocabulary, and control word scores were significant predictors in the final model. The combined group performed better than the other two groups, and greater reading accuracy, oral vocabulary knowledge, and performance on control words was associated with better performance. Category recognition was surprisingly better for words presented without definitions than for those presented with definitions. Subcategory recognition The fixed-factor model significantly improved fit compared with the empty model, χ2(7) = 21.15, p < .001. Interaction terms did not significantly improve model fit. Presence of definition and vocabulary level significantly predicted performance, with the presence of a definition and greater existing oral vocabulary knowledge associated with better performance. No group effect was evident; the three groups performed similarly on this task. Definition recognition The fixed-factor model significantly improved fit compared with the empty model, χ2(7) = 42.61, p < .001. Addition of the group by reading comprehension interaction also improved model fit. As for subcategory recognition, the presence of a definition and oral vocabulary knowledge were significant predictors, but group was not a significant predictor. To further explore the group by reading comprehension interaction in the definition recognition task, a separate model was computed for each group. For the combined group, the presence of a definition predicted performance, β = .97, χ2(1) = 9.63, p = .002. The performance of the listening group was positively influenced by vocabulary, β = .42, χ2(1) = 4.61, p = .030, and reading comprehension, β = .47, χ2(1) = 5.32, p = .020, whereas the performance of the reading group was positively influenced by vocabulary, β = .42, χ2(1) = 4.39, p = .040, and control word scores, β = .79, χ2(1) = 10.60, p = .003. These supplementary analyses suggest that the group by reading comprehension interaction reflects a positive association between reading comprehension and performance for the listening group but not for the other two groups. In addition, the presence of definitions may have particularly supported later definition recognition by children in the combined group, whereas existing vocabulary knowledge particularly influenced performance in the listening and reading groups. However, because the group by definition and group by vocabulary interactions were not significant, no strong conclusions can be drawn in relation to these findings.","This study investigated children’s incidental word learning from stories, comparing listening, reading, and combined conditions for the first time. Children learned information about the phonology, orthography, and semantics of new words, but the extent and nature of their learning depended on the presentation modality. Learning was also modulated by the presence of a definition and by children’s existing abilities. The influence of presentation modality, presence of a definition, and individual differences are discussed in turn. In relation to presentation modality, the orthographic facilitation reported in previous studies (Rosenthal & Ehri, 2008) motivated our hypothesis that the combined group would outperform the other two groups on the phonological task. However, in our paradigm, where children learned new words incidentally rather than being taught them, children learned the phonological forms of the new words with equal proficiency across conditions. Thus, children who read the story appear to have formed a phonological representation of the new words that not only was better than chance and better than for control words but also was equivalent to the learning shown by children who had been directly provided with the phonological form. This is in line with Share’s (1995) self- teaching hypothesis, which states that phonological recoding occurs while reading. Children were able to use their knowledge of the relationship between orthography and phonology to learn the phonological form of words that they saw in written format only. To be certain that children’s phonological representations were equivalent across the three conditions, however, we must be confident that our phonological task provided a robust measure of phonological learning. Three issues are worthy of mention in relation to this point. First, performance on this task involved distinguishing between targets and plausible foils, which were generated, where possible, from adult mispronunciations of the written forms. Consequently, it is possible that the task probed abilities other than children’s in-task learning such as their general sensitivity to an oral form’s word- likeness. This idea is supported by the finding that control word scores significantly predicted target word scores on this task. To account for such general effects, scores on control word trials, therefore, were included in all analytical models, enabling us to have confidence that our results reflect children’s phonological learning within the task. A second consideration is whether the use of plausible foils made the task particularly challenging for children in the reading group, who were not provided with a phonological form and who, therefore, may have generated an alternative form while reading that was aligned more closely with the foil than with the target. However, this does not seem to have been the case; the reading group performed just as well as the other groups on the phonological task. Finally, it is possible that the recognition task was insufficiently sensitive to detect subtle differences in the quality of children’s phonological representations across conditions due to either the nature of the task itself or the number of items tested. Had we used a production task, for example (cf. Rosenthal & Ehri, 2008), group differences might have been detected. However, an identical recognition task format with the same numbers of items elicited group differences on the orthographic task in the current study (see below), and similar tasks and trial numbers have been found to be sufficiently sensitive in previous studies (e.g., Ricketts et al., 2009; Ricketts et al., 2011; Rosenthal & Ehri, 2008; Rosenthal & Ehri, 2011). Nevertheless, future studies might seek to corroborate our conclusions using alternative measures of phonological learning. In contrast to the findings relating to phonological learning, the three groups did not show equivalent orthographic learning. Here, children in the combined and reading groups outperformed those in the listening group, whose performance was at chance. This indicates that a word’s orthographic form is not automatically extrapolated from its phonology; rather, the presentation of written text prompts orthographic learning. Contrary to previous findings (Rosenthal & Ehri, 2011), the additional presentation of the oral form did not enhance orthographic learning relative to the presentation of the written form alone. As for the phonological task, it is possible that a more sensitive measure of orthographic knowledge might have revealed a difference in the orthographic representation of the two reading groups. Nevertheless, there is no evidence in our study for phonological facilitation of orthographic learning. Given the relatively low reliability for the orthographic task, these findings warrant replication in future studies. In summary, the results of the phonological and orthographic recognition tasks suggest that listening to stories promotes learning of new phonological but not orthographic forms. Reading, in contrast, appears to support the acquisition of both phonological and orthographic information about new words, presumably creating higher quality lexical representations (cf. Perfetti & Hart, 2002). The asymmetry in our results resonates with observations that children tend to perform better in reading tasks that require them to produce an oral form from a written one than in spelling tasks that require the reverse (Cossu, Gugliotta, & Marshall, 1995). Turning to the semantic tasks, we explored whether semantic learning would be greater in the reading group than in the listening group (Brown et al., 2008) or vice versa (Suggate et al., 2013). Our results indicated no support for the superiority of either oral or written presentation. It appears that, at 8 or 9 years of age, children learn as much about the meanings of words from reading as they do from listening to stories. It was also hypothesized that children in the combined condition would acquire more semantic information about new words than children in the listening and reading conditions (Ricketts et al., 2009; Rosenthal & Ehri, 2008; Rosenthal & Ehri, 2011). This prediction was upheld, but only for the category recognition task, which required the abstraction of word knowledge to categorize the word correctly. In this task, we observed both an orthographic facilitation effect (i.e., better performance in the combined group than in the listening group) and a phonological facilitation effect (i.e., better performance in the combined group than in the reading group). Before we consider why the benefit of the combined condition was found in only one semantic task, we first reflect on the causes of children’s better performance in the combined condition. The lexical quality hypothesis (cf. Perfetti & Hart, 2002) suggests that words with higher quality representations are more easily retrieved from memory. This theoretical perspective focuses on existing lexical representations rather than their acquisition. Nonetheless, it is consistent with the proposal that the combined condition promotes the building of better specified representations, facilitating access to stored word knowledge at test. However, this framework would appear to predict consistent differences in the quality of the phonological, orthographic, and semantic representations formed in the combined and single modality presentations, as reported previously in studies finding orthographic facilitation effects (Ricketts et al., 2009; Rosenthal & Ehri, 2008) and phonological facilitation effects (Rosenthal & Ehri, 2011). Such consistent effects across tasks were not found in the current study. As discussed above, it is possible that the insensitivity of our phonological and orthographic measures masked subtle differences between the groups. Although we find no evidence for this, it remains possible that the phonological or orthographic forms acquired in the combined condition were of a higher quality than those in the other two conditions. An alternative possibility is that the combined condition reduced cognitive load during word learning, freeing resources for comprehension and word meaning extraction (Mayer et al., 1999). Compared with the combined group, the reading and listening groups were charged with additional processing demands at the point of encountering the new words—the reading group in the form of spontaneous phonological recoding (as evidenced by the children’s performance on the phonological task) and the listening group due to the attentional demands associated with continuously monitoring the oral story presentation (without any “backup” support from the written text). Children in the combined group, therefore, may have had more resources available to allocate to processing the contextual support (including definitions) that immediately followed a new word, thereby encoding the word meanings better. This perspective would predict word form representations to be of similar quality across conditions but would predict representations of words’ meanings to be better in the combined condition, the pattern found in the current study. Although both approaches could, in theory, explain the greater semantic learning of the combined group, it remains to be discussed why this effect was seen in only one of our semantic tasks, the category recognition task. Two potential explanations occur to us. First, only the category recognition task required children to abstract category-level knowledge about the target words from the information provided in the story: the category label for each new word was never directly provided in the narrative. In contrast, recognizing the correct subcategory or definition in the other semantic tasks required participants to choose a subcategory or definition that was very similar to those provided in the story. Second, choosing the correct response on the subcategory and definition recognition subtests is likely to have been easier than choosing the correct form in the category recognition task. In the former tasks, only one of the four alternative response options had been encountered in the story (e.g., among the four clothing options offered in the definition task for the word hauberk, only a soldier’s shirt made of chain mail had been mentioned in the story, making the other alternatives less likely regardless of any learning of the label for this item). In contrast, all four of the alternatives in the category recognition task correctly described one of the new words presented in the story (i.e., to recognize hauberk as a piece of clothing in the category recognition task, children needed to know that this new word did not identify a new animal, job, or part of a house, each of which had been encountered in the story). These factors could have made the category recognition task the most challenging measure, and therefore the most sensitive measure, of children’s semantic learning. By freeing resources, the combined condition might have especially facilitated this task due to the level of abstraction and/or precision of the mapping required, a hypothesis that warrants further investigation. The study also investigated the impact of definitions on word learning. It was hypothesized that the presence of an accompanying definition alongside each new word would foster semantic learning (Wilkinson & Houston-Price, 2013). Although definitions did not affect children’s phonological or orthographic learning in this study, they facilitated the learning of subcategories and definitions of the words but hindered category recognition. These findings indicate that definitions help children to learn detailed information regarding new words but do not help to extract more general categorical information. This finding is perhaps not surprising considering that in the current study the definitions directly provided most of the information needed in the definition recognition task along with the subcategory (or a synonym of this) required in the subcategory recognition task. When definitions were provided, the level of abstraction required by these two tasks, therefore, was quite low. When definitions were not provided, children needed to extract the details of the word’s semantics from the text, synthesizing various cues to construct a coherent representation of its meaning, in order to succeed on these tasks. In contrast, the abstraction required to identify words’ categories, necessary to succeed on the category recognition task, would be similar whether or not a definition was provided because this information was not supplied within definitions. When provided with a definition, children appear to have learned the specific information in the definition at the expense of abstracting a more general representation of the word’s meaning. We also explored the effects of children’s reading and language abilities on their learning in different presentation conditions. Reading accuracy predicted performance in the phonological, orthographic, and category recognition tasks irrespective of presentation modality. Children with better baseline knowledge of orthography–phonology mappings were better able to learn new orthographic and phonological forms and extract semantic information (Ricketts et al., 2011; Rosenthal & Ehri, 2008). By facilitating the learning of new word forms, better decoding skills could potentially reduce the cognitive load in a similar way to that discussed earlier in relation to the dual presentation modality, enabling the learning of more complex semantic information (category learning). As hypothesized, oral vocabulary level predicted performance in all three semantic tasks. It is likely that children with smaller vocabularies face a more difficult task when reading or listening to stories because their reduced familiarity with the vocabulary in the story increases the size of the word learning challenge (Shefelbine, 1990). In addition, children with larger vocabularies may employ better word learning strategies (Cain et al., 2004); for example, they may be more proficient at linking new words with contextual information (McKeown, 1985). Written text comprehension ability was also positively related to definition recognition, but only in the listening group. This finding is surprising, because children in the listening group were not required to comprehend written text. It is possible that a third factor—an unmeasured one—is involved in both reading comprehension and word learning while listening to stories, such as the ability to build a representation of the incoming story or other discourse online without the need to revisit the text (e.g., by rereading). Despite the broad variation in children’s abilities, it is worth noting that background measures did not interact with the impact of presentation modality on category recognition. It is particularly interesting that, irrespective of reading ability, children’s category learning benefited from combined oral and written presentation. Thus, even good readers can benefit from oral presentation while reading, despite their ability to read efficiently by themselves. Similarly, poor readers can benefit from the presence of the written text alongside an oral presentation despite their poorer reading abilities. Given the growing use of e-books, these results are important in suggesting that listening to a narration of the story while reading enhances vocabulary acquisition. Having said that, given the narrow age range of the children in our study, it is possible that individual differences in ability might interact with presentation modality to affect learning in a younger or older population of school-aged children. Older, more efficient readers might, for example, learn equally well from reading and reading and listening, as adults do (Sydorenko, 2010), whereas less experienced readers might not be able to take advantage of a dual presentation modality compared with listening. Future studies that explore the relationship between the effect of presentation modality and individual differences in ability in younger and older readers would elucidate whether the pattern found in the current study changes with development.","This study shows that children aged 8 or 9 years are able to extract information about new words from a story and that the modality of the story presentation influences the nature and extent of their learning. Children learned the new words’ phonology similarly well from the three presentation modalities (listening, reading, and combined conditions), but orthography was learned only when it was directly presented (reading and combined conditions). Children showed the strongest learning for words’ semantic categories when words were presented both orally and in written form (combined modality). The practical implications of these findings for the classroom are that children are able to learn information about the phonological forms, orthographic forms, and meanings of new words when they are listening to and/or reading stories. Importantly, both listening to and reading stories can support the learning of new phonological forms, whereas opportunities to see the new words written down are crucial for building representations of their orthographic forms. Finally, allowing children to hear stories while they are reading along may be optimal for learning, and especially for the extraction of semantic information, supporting teachers’ practice of reading aloud in the classroom."],["Objectives: This study is the first examination of the longitudinal associations between behavioural regulation and accelerometer-assessed physical activity in parents of primary-school aged children. Design: A cohort design using data from the B-Proact1v project. Method: There were three measurement phases over five years. Exercise motivation was measured using the BREQ-2 and mean minutes of moderate-to-vigorous physical activity (MVPA) were derived from ActiGraph accelerometers worn for a minimum of 3 days. Cross-sectional associations were explored via linear regression models using parent data from the final two phases of the B-Proact1v cohort, when children were 8–9 years-old (925 parents, 72.3% mothers) and 10 to 11 years-old (891 parents, 72.6% mothers). Longitudinal associations across all three phases were explored using multi-level models on data from all parents who provided information on at least one occasion (2374 parents). All models were adjusted for gender, number of children, deprivation indices and school-based clustering. Results: Cross-sectionally, identified regulation was associated with 5.43 (95% CI [2.56, 8.32]) and 4.88 (95% CI [1.94, 7.83]) minutes more MVPA per day at times 2 and 3 respectively. In the longitudinal model, a one-unit increase in introjected regulation was associated with a decline in mean daily MVPA of 0.52 (95% CI [-0.88, −0.16]) minutes per year. Conclusions: Interventions to promote the internalisation of personally meaningful rationales for being active, whilst ensuring that feelings of guilt are not fostered, may offer promise for facilitating greater long-term physical activity engagement in parents of primary school age children. --------------------------------------------------------------------------------","Physical activity is associated with reduced risk of a variety of health outcomes, including heart disease, stroke, type 2 diabetes, several forms of cancer, and depression (Kyu et al., 2016; Rebar et al., 2015). Similarly, physical inactivity has been shown to be detrimental for health and well-being and has been identified as a source of great economic cost globally (Ding et al., 2016). As such, global physical activity guidelines recommend that adults undertake at least 150 min per week of moderate intensity activity, including additional muscle-strengthening activities at least twice a week, alongside the general aim of reducing their sedentary time (Australian Government, 2017; Department for Department of Health, 2011; Health Council of the Netherlands, 2017; Tremblay et al., 2017; World Health Organisation, 2010). However, evidence suggests that between 15% and 43% of adults in western countries do not meet physical activity recommendations (Colley et al., 2011; Craig, Mindell, & Hirani, 2011; Hallal et al., 2012; Kapteyn et al., 2018). From a public health perspective, it is clear that efforts need to be made to encourage more adults to be regularly active. Thirty-five percent of adults (aged 18–65) in the UK have dependent children (Office for National Statistics, 2017). Parents of young children have been shown to engage in less moderate-to-vigorous physical activity (MVPA) than similar aged adults without children (Bellows-Riecken & Rhodes, 2008; Berge, Larson, Bauer, & Neumark-Sztainer, 2011), with a noticeable decrease in physical activity at the point of transition to parenthood (Hull et al., 2010; McIntyre & Rhodes, 2009). Engaging in regular physical activity may be particularly challenging for parents with dependent children due to increased demands on their time, financial burdens, and a change in priorities compared to before parenthood. Yet, promoting parental engagement in physical activity could be beneficial for both parents and children in terms of health benefits, parenting behaviour, and energy levels (Hamilton & White, 2010; Lewis & Ridge, 2005). Additionally, if parents are active then they model active behaviour to their children (Sebire et al., 2016), with some evidence suggesting a weak positive association between parent physical activity and child activity (Jago, Solomon-Moore, Macdonald-Wallis, Thompson, Lawlor & Sebire, 2017a, b; Shutz, Browning, Smith, Lohse, & Cunningham-Sabo, 2018; Yao & Rhodes, 2015). A recent review of reviews highlighted that individual level variables (e.g., motivation, age, and health intentions) are the most consistent correlates of physical activity (Choi, Lee, Lee, Kang, & Choi, 2017) therefore indicating that interventions should be either tailored to specific populations or should encompass ways of manipulating these variables to increase physical activity. Self-determination theory (SDT; Ryan & Deci, 2017) is a framework through which the motivational processes that underpin physical activity can be investigated. Within SDT, quality of motivation is placed upon a continuum whereby different types of motivation differ in the extent to which they are autonomous or controlled (Deci & Ryan, 2000; Howard, Gagne, & Bureau, 2017). Three types of motivation are said to be more autonomous in nature: Intrinsic motivation, the most autonomous form of motivation characterised by an individual inherently enjoying or gaining satisfaction from the activity; integrated regulation, when the behaviour aligns with an individual’s identity; and identified regulation, when an individual consciously values the behaviour (Ryan & Deci, 2017). More controlled types of motivation are introjected regulation, when behaviour is controlled by self-imposed sanctions, such as shame, pride, ego, or guilt (Deci & Ryan, 2002) and external regulation, the most controlled form of motivation when behaviour is driven by external factors such as rewards, compliance and punishments (Deci & Ryan, 1987). Additionally, a lack of either autonomous or controlled forms of motivation is classed as amotivation (Ryan & Deci, 2000). In the context of physical activity, effortful and persistent behaviour is more likely to occur when an individual’s motivation is autonomous as opposed to controlled (Standage & Ryan, 2012). Cross-sectional evidence consistently shows autonomous motivation for exercise (i.e., intrinsic, integrated, and identified regulation) to be positively associated with self-reported and accelerometer-assessed physical activity in healthy adults (see Teixeira, Carraca, Markland, Silva, & Ryan, 2012 for a review). Controlled motivation is generally shown to have little cross-sectional association with self-reported and accelerometer-assessed physical activity behaviour (Teixeira et al., 2012). However, when analysed separately, introjected regulation more frequently shows a positive cross-sectional association with physical activity whereas external regulation is more commonly negatively associated with physical activity (e.g., Edmunds, Ntoumanis, & Duda, 2006; Wilson, Rodgers, & Fraser, 2002). In addition to the cross-sectional evidence, there are a small number of studies that have examined, and provided evidence, for a small to moderate positive association between autonomous motivation and self-reported physical activity over periods of time ranging from 1 to 6 months (Barbeau, Sweet, & Fortier, 2009; Fortier, Kowal, Lemyre, & Orpana, 2009; Gunnell, Crocker, Mack, Wilson, & Zumbo, 2014). In support of these longitudinal associations, qualitative evidence aligns with the theoretical tenet that movement through the behavioural regulation continuum towards more autonomous motivation is central to physical activity adherence (Kinnafick, Thogerson-Ntoumani, & Duda, 2014). To date, studies have shown no evidence for a longitudinal association between controlled motivation and physical activity (Barbeau et al., 2009; Gunnell et al., 2014). The limited number of studies assessing the associations between motivation and physical activity over time have used autonomous and controlled composites and not disaggregated the types of behavioural regulation in statistical models, thus limiting the study of the roles of qualitatively different types of motivation. Additionally, these longitudinal studies have also relied on self-reported measures of physical activity behaviour which are prone to bias (Sallis & Saelens, 2000). Therefore, further investigation of the longitudinal associations between behavioural regulation and physical activity, using more reliable behavioural estimates (e.g., through accelerometers) is warranted. Despite evidence indicating lower levels of physical activity in parents compared to the wider adult population, theoretical models have been seldom used to understand physical activity during parenthood (Bellows-Riecken & Rhodes, 2008). The quality of motivation for exercise may be particularly pertinent to parents’ physical activity due to extensive competing demands (e.g., time, fatigue, childcare) that may make converting some forms of motivation to behaviour more challenging for parents than non-parents (McIntyre & Rhodes, 2009; Solomon-Moore et al., 2016.) In support of this, in a study involving 1067 parents of children aged 5–6 years old, only identified regulation showed evidence of a cross-sectional association with MVPA after adjustment (Solomon-Moore et al., 2016), suggesting that, for parents of younger children, identifying with personally meaningful and valuable benefits of exercise may be the strongest motivational driver. However, there are no studies of the motivation-physical activity associations amongst parents of older children or evidence for any longitudinal associations between motivation and physical activity behaviour during parenthood. The aims of this research were to 1) examine the cross-sectional associations between the behavioural regulations set forward in SDT and objectively-estimated physical activity in parents of children aged 8–9 years old and then two years later when the same child was 10–11 years old and 2) assess the longitudinal associations between behavioural regulation type and accelerometer-assessed physical activity in parents over a five-year period.","The current analyses used data from the B-Proact1v project. The broader project is a longitudinal study exploring the factors associated with physical activity and sedentary behaviour in children and their parents throughout primary school. Briefly, data collection was conducted at three timepoints: between January 2012 and July 2013 when all participants had a child aged five to six years (Year 1, time 1), between March 2015 and July 2016 when the same child was aged eight to nine years (Year 4, time 2) and between March 2017 and May 2018 when the same child was in year 6 (aged 10–11). A total of 57 schools consented to participate at time 1 and were subsequently invited to participate at times 2 and 3. Forty-seven schools participated at time 2 and 50 participated at time 3. Across all three timepoints, data were collected from 2555 parents from 2132 families: 1195 parents were involved at time 1, 1140 at time 2, and 1233 at time 3. A total of 546 parents took part across two timepoints and 246 across three timepoints. The study received ethical approval from the University of Bristol ethics committee, and written consent was received from all participants at each phase of data collection.","Characteristics. Parents completed a questionnaire which included information about their date of birth, gender, ethnicity, height, weight, education level, and number of children. BMI was calculated from their self-reported height and weight. Parents also reported their home postcode, and this was used to derive Indices of Multiple Deprivation (IMD scores), based upon the English Indices of Deprivation (http://data.gov.uk/dataset/index-of- multiple-deprivation). Higher scores indicate areas of higher deprivation. Motivation to exercise. The Behavioural Regulation in Exercise Questionnaire (BREQ-2) was used to assess motivation to exercise (Markland & Tobin, 2004). The BREQ-2 consists of 19-items each assessing one of five forms of behavioural regulations: intrinsic (4 items e.g., ‘I exercise because it’s fun’), identified (4 items e.g., ‘It’s important to me to exercise regularly’), introjected (3 items e.g., ‘I feel like a failure when I haven’t exercise in a while’), external (4 items e.g., ‘I exercise because other people say I should’), and amotivation (4 items e.g., ‘I don’t see the point in exercising’). Due to difficulties in empirically distinguishing between identified and integrated regulation, the BREQ-2 does not assess integrated regulation (Markland & Tobin, 2004). Participants rated each item on a 5-point Likert scale ranging from 0 (not true for me) to 4 (very true for me). In the current study, the BREQ-2 subscales had good internal consistency in both the cross sectional and longitudinal samples (α > 0.7; see Table 1). Physical activity. Participants were asked to wear a waist-worn ActiGraph wGT3X-BT accelerometer for five days, including two weekend days. Accelerometer data were processed using Kinesoft (v3.3.75; Kinesoft, Saskatchewan, Canada) in 60-s epochs. In line with recommendations for monitoring habitual physical activity in adults, analysis was restricted to participants who provided at least three days of valid data including at least 1 weekend day (Aadland & Ylvisaker, 2015; Trost, McIver, & Pate, 2005). A valid day was defined as at least 500 min of data, after excluding intervals of ≥60 min of zero counts allowing up to 2 min of interruptions. The average number of MVPA minutes per day were derived for each participant using population- specific cut points for adults (≥2020 counts per minute (Troiano et al., 2008)). Data analysis The data analysis consisted of cross-sectional analyses of time 2 and time 3 data (a cross-sectional analysis of baseline data is published elsewhere; Solomon-Moore et al., 2016) and a longitudinal analysis including data from all three timepoints. All analyses were conducted at the parent level, so where two or more parents/guardians from the same family were included in the project, each parent was treated as a separate participant. Participants were included in the cross-sectional analysis if they had valid accelerometer data, and BREQ-2 responses with not more than 1 missing item per subscale. Participants were included in the longitudinal analysis if they met the above criteria for at least one timepoint. Following recommendations for dealing with missing data, and in order to reduce bias and increase statistical power, multiple imputation using chained equations was used to impute missing data for participating parents cross-sectionally (Rubin, 1996; Tabachnick & Fidell, 2012). The imputation models included parent gender, age, ethnicity, BMI, education level, IMD score, number of children, MVPA and the five subscales of behavioural regulation. For each, 20 imputed datasets were created using 20 cycles of regression switching and estimates were combined across datasets using Rubin’s rules (Rubin, 1996). Five independent variables, reflecting the five motivation types, were treated as continuous variables. In the cross-sectional analyses, linear regression models were used to examine associations between behavioural regulation and mean MVPA minutes per day. To identify longitudinal associations between the motivation variables and physical activity, we used a multi-level model to capture how MVPA changes over time. We also included interaction terms between behavioural regulation variables and age to explore whether change in MVPA over time differs with changes in behavioural regulation. In line with evidence for their influence on MVPA in adults (Cerin, Leslie, & Owen, 2009; Hull et al., 2010; Trost, Owen, Bauman, Sallis, & Brown, 2002), all analyses were adjusted for age, gender, number of children in the household and IMD score. In all models, robust standard errors were used to account for school clustering in the study design, and parents were clustered within families, to account for family-level similarities. All analyses were performed in Stata version 15 (StataCorp, 2017). Preliminary analysis ~~~~~~~~~~~~~~~~~~~~ At both time 2 and time 3, missing data among participating parents was minimal, and the distributions of observed and imputed characteristics were similar (Table 1). At time 2, the sample consisted of 925 participants of whom 72% were female, with a mean age of 41.34 years (SD = 6.20), and mean BMI of 25.83 kg/m2 (SD = 4.84). At time 3, the sample consisted of 891 participants, of whom 73% were female, with a mean age of 43.32 years (SD = 5.84), and mean BMI of 25.83 kg/m2 (SD = 4.77). Mean IMD was consistent across time (14.92 at time 2 and 14.44 at time 3). Average daily MVPA increased slightly from 50.03 min (SD = 23.91) at time 2–52.28 min (SD = 24.31) at time 3. At both timepoints, means and standard deviations for all motivation variables were similar, with participants reporting higher levels of autonomous motivation than controlled motivation for exercise, with levels of identified regulation being higher than intrinsic regulation. Full descriptives for time 1 are reported elsewhere (see Solomon-Moore et al., 2016). A total of 2374 parents had valid accelerometer data for at least one timepoint (93% of parents involved across the whole study). Of these, 463 had valid accelerometer data for two timepoints (85% of those parents involved at two timepoints) and 185 for three timepoints (75% of parents involved at three timepoints). Twenty-five parents provided no identifiable information at any timepoint and so were excluded from the analyses. Due to the study design, school attrition accounted for 244 families not taking part at time 2 and 167 families at time 3. At the family level, children moving to schools not involved in the project accounts for the drop out of 253 families. Further, as the same parent was not required to participate at every timepoint, in 227 families who had been involved at time 1, a different parent participated at time 2. A total of 302 families had a different parent participating at each timepoint. In these cases, all parents involved at any timepoint are included in the analysis. Cross-sectional associations between motivation and physical activity Consistent with baseline findings (Solomon-Moore et al., 2016), fully adjusted cross-sectional regression models (Table 2) showed a positive association between identified regulation and MVPA, with a one-unit increase in identified regulation associated with a 5.4-min (95% CI [2.6, 8.3]) and 4.9 min (95% CI [1.9, 7.8]) increase in MVPA per day at time 2 and time 3, respectively. There was no evidence for an association between any other type of behavioural regulation and MVPA at time 2. At time 3, introjected regulation was negatively associated with MVPA, with a one-unit increase in introjected regulation associated with a 3.2-min decrease in MVPA per day (95% CI [-5.0, −1.4]). There was no evidence for an association between the other types of behavioural regulation and MVPA at time 3. Longitudinal associations between motivation and physical activity The full multi-level model (Model 1, Table 3) explores how parent MVPA changes across the three timepoints in relation to behavioural regulation. Daily MVPA increased by an average of 0.60 min per year (95% CI [0.17, 1.03]). Identified regulation was positively associated with MVPA, with a one-unit increase in identified regulation being associated with an average of 3.96-min more MVPA per day per year (95% CI [2.42, 5.50]). Due to the association between time and MVPA in model 1, in model 2 we investigated whether change in MVPA over time was associated with behavioural regulation (Table 3). At time 1, identified regulation remained positively associated with MVPA with a one unit increase in identified regulation being associated with an 8.73-min increase in average daily MVPA per year (95% CI [3.41, 14.05]), however there was no association with change in MVPA over time. Introjected regulation was not associated with MVPA at time 1 but had a small negative association with change in MVPA over time, with a one unit increase in introjected regulation being associated with an average decline in daily MVPA of 0.52 min per year (95% CI [ −0.88, −0.16]).","This study presents the first longitudinal analysis of the association between parents’ exercise motivation and accelerometer-estimated physical activity. The analyses indicate that high levels of introjected regulation can lead to a small decrease in MVPA over time. Substantiating the baseline findings from the B-Proact1v cohort, the cross-sectional analyses also show that identified regulation is consistently associated with higher levels of MVPA but was not associated with change in MVPA over time. Our cross-sectional findings corroborate previous evidence showing that identified regulation is the type of behavioural regulation most strongly associated with physical activity behaviours in adults (Standage, Sebire, & Loney, 2008; Teixeira et al., 2012; Wilson, Sabiston, Mack, & Blanchard, 2012). These also extend the baseline findings of the B-Proact1v cohort to show that, across all three phases of the project, parents who are motivated to be physically active due to personal value engage in higher levels of MVPA than parents who are motivated in other ways. However, despite being consistently associated with higher levels of physical activity, the multi-level models suggest that identified regulation is not associated with change in physical activity over time. This is consistent with the theoretical assumption that more autonomous motivation is associated with long-term behavioural engagement (Ryan & Deci, 2017) and could indicate that identified regulation is more pertinent to behavioural maintenance, as behaviours that align with an individual’s personal values are more likely to be sustained (Kwasnicka, Dombrowski, White, & Sniehotta, 2016). Whilst the motivational continuum proposed within SDT suggests that intrinsic motivation (i.e., enjoyment of the behaviour itself) is the strongest motivational driver of behaviour, it has been recognised in the wider literature that behaviours such as exercise or being active might be more strongly driven by what can be achieved through doing it (e.g., physical or mental health benefits, social contact with others, development of skills and feelings of competence) rather than inherent enjoyment of the activity itself (Standage & Ryan, 2012). In the present study, parents’ endorsement of intrinsic and identified types of motivation were similar, emphasising that enjoyment and satisfaction are important sources of motivation. However, the lack of association between intrinsic motivation and physical activity suggests that, for parents, enjoyment of physical activity is not sufficient to lead to action, and that personally valuing activity is a more stable motivation factor for underpinning behaviour. One explanation for this could be related to the parental role, where activities that do not directly align, and potentially compete, with core parenting duties (e.g., taking time out to be active or exercise for fun or enjoyment) are not prioritised and may even result in feelings of guilt and selfishness (Hamilton & White, 2010). The present cross-sectional findings show that when the benefits of physical activity are perceived to be personally relevant and valuable, engagement in physical activity is greater. Therefore, valuing the benefits of exercise as an individual (e.g., physical & mental health, better sleep, more energy) and/or as a parent (e.g., modelling healthy behaviours, being physically fit to “keep up” with active children; Hamilton & White, 2010) may be central to being a more physically active parent. The cross-sectional findings showed a differential change in the association between introjected regulation and parents’ MVPA across the three timepoints, moving from a small positive association at time 1 (Solomon-Moore et al., 2016) to an increasingly negative association across times 2 and 3. The longitudinal results further highlight the impact of introjected regulation on behavioural outcomes, with a negative association between introjected regulation and change in MVPA over time. Within SDT, it is proposed that more controlled forms of motivation are detrimental to both behavioural and well-being outcomes (Ryan & Deci, 2017), yet previous longitudinal studies have found no evidence for this association (Barbeau et al., 2009; Gunnell et al., 2014) and cross- sectional research has indicated that introjected regulation may be associated with higher levels of physical activity (e.g., Brunet & Sabiston, 2011; Edmunds et al., 2006). These findings offer the first evidence of a long-term negative impact of introjected regulation on physical activity behaviour. Although further longitudinal research is warranted, given the universality of SDT we would anticipate that similar associations in the wider adult population with previous studies finding no association due to aggregating introjected and external regulations into a controlled motivation variable (e.g., Barbeau et al., 2009; Gunnell et al., 2014). However, there may also factors unique to parents that mean motivations grounded in feelings of guilt and shame have an increasing negative effect on physical activity as your child gets older. Extending previous studies, we explored the individual association of each type of behavioural regulation with change in physical activity behaviour. A potential explanation for this could be the association between introjected regulation and maladaptive social comparisons (Thogersen-Ntoumani & Ntoumanis, 2006). For example, parents may feel envy and resentment towards other parents who appear to be managing to be physically active (Hamilton & White, 2010), with such feelings having a more detrimental effect on behavioural engagement once the child is older and the parent might have more discretionary time. In terms of internal comparisons, parents of young children may accept that they cannot be physically active in the short term, hoping that they will be more active as their child gets older. If parents do not meet these self- imposed expectations when their child is older, this could result in a cyclical relationship between failure to be active and guilt. For parents these effects might be particularly salient due to the competing demands on their time. Evidence suggests that, for adults, increasing daily MVPA by 5–10 min can have clinically meaningful health benefits (Dohrn, Kwak, Oja, Sjostrom & Hagstromer, 2018; Greaves et al., 2011). With this in mind, the data presented in this paper indicate that strategies to encourage parents to find personal value, relevance and importance in being physically active whilst ensuring that feelings of guilt or shame are not induced, may be particularly important for promoting increased physical activity engagement with the potential for meaningful impact on health outcomes. This requires environments that support rather than thwart the three psychological needs of autonomy, competence and relatedness (Ryan & Deci, 2017). Such environments are characterised by the provision of choice (e.g., when and how to be active/exercise), optimal challenge (e.g., one’s exercise goals/plans are neither too easy nor too hard to achieve), and strong connections with others (e.g., being active is normal/accepted amongst one’s friends, family, or colleagues). Additionally, and particularly relevant for the promotion of identified regulation, may be the provision of a personally relevant and valuable rationale for being active, which has been shown to promote long-term behavioural engagement (Samdal, Eide, Barth, Williams, & Meland, 2017). Qualitative evidence suggests that long-term goals can be abstract and undermine physical activity and, as such, physical activity messages should emphasise that being active can also be immediately gratifying and help people cope with or manage their daily goals (Segar, Taber, Patrick, Thai, & Oh, 2017). Research on the content of exercise goals (essentially benefits of exercise) shows that goals such as health, social affiliation and development of exercise competence or skills are associated with greater autonomous motivation and psychological well-being (Sebire, Standage & Vansteenkiste, 2009). For parents, effective messages could include promoting physical activity as an important family activity that provides the opportunity to spend time together and interact with others, increase energy levels, and to relax and escape daily pressures (Segar et al., 2017). However, further qualitative work is required to inform the development of specific health messaging that aligns with identified regulation.","The strengths of the study lie in the theoretically-grounded approach to motivation, longitudinal data collected from the same cohort on three occasions over a five-year period and the use of accelerometers to estimate levels of MVPA which when combined have made a novel contribution to the literature. However, limitations should be acknowledged. First, the sample was largely female, and therefore is more representative of mothers than parents in general. However, given the universality of SDT, and supporting evidence indicating no gender differences in behavioural regulation as measured through BREQ-2, we would not anticipate significant differences in findings in a more male-dominated sample (Guerin, Bales, Sweet, & Fortier, 2012). Additionally, participating parents were generally more active than average in the UK (Craig, Mindell, & Hirani, 2011) and had high levels of self-determined motivation. This may limit the generalisability of the findings to the wider parent population and may also have limited the extent to which we can observe change over time. Whilst we did not have the power in the present study, in future researchers may want to look at the differences in motivation between parents with low physical activity levels, and those with high physical activity levels. A further limitation is that the BREQ-2 questionnaire used to measure motivation does not assess integrated regulation which sits on the SDT continuum of motivation between identified and intrinsic motivation types and represents the assimilation of the behaviour with one’s values, goals, and sense of self (e.g., “I am active”; Ryan & Deci, 2017).","This study is the first to assess the longitudinal associations between parent’s behavioural regulation to exercise in relation to accelerometer-estimated physical activity. The results indicate that motivation grounded in the personal meaning and value of exercise is associated with higher levels of MVPA. Additionally, the results suggest that being motivated by feelings of guilt and shame impact negatively on MVPA over time. Therefore, interventions that promote greater enjoyment, personal relevance and value, whilst also ensuring that guilt is not promoted, may offer promise for the facilitation of greater long-term physical activity engagement in parents.","We have no competing interests to report.","This work was funded by the British Heart Foundation (Grant numbers: PG/11/51/28986 and SP 14/4/31123)."],["Visual working memory (VWM) is the ability to hold in mind visual information for brief periods of time. The current study investigated VWM precision development longitudinally. Participants (N = 40, aged 7-11 years) completed delayed reproduction sequential VWM tasks at baseline and two years later. Results show age-related improvement in recall precision on both 1-item and 3-item VWM tasks, suggesting development during childhood and early adolescence in the resolution with which both single and multiple items are stored in VWM. Probabilistic modelling of response distribution data suggests age-related improvement in precision is attributable to a specific decrease in the variability (noisiness) of stored feature representations. This highlights a novel developmental mechanism which may underlie longitudinal improvement in VWM performance, crucially without invoking improvement in the number of items that can be stored. VWM precision provides a sensitive metric with which to track developmental changes longitudinally, shedding light on underlying cognitive mechanisms. --------------------------------------------------------------------------------","Visual working memory (VWM) provides a temporary storage mechanism for the retention and manipulation of visual information to support other cognitive processes (Luck & Vogel, 1997). This ability is considered to be a critical contributor to many essential cognitive functions such as decision-making, complex reasoning and goal-directed action (Baddeley, 2003). Numerous developmental studies have shown that performance on established neuropsychological tests of VWM improves during childhood (Alloway, Gathercole, & Pickering, 2006; Gathercole, Pickering, Ambridge, & Wearing, 2004). Typically in these tests, participants view a static visual array (e.g. coloured shapes) or spatiotemporal sequence of visual events (e.g. block tapping) which are held in memory and then reproduced following a delay. These studies have demonstrated, within large cross- sectional datasets, robust age trajectories and evidence for developmental stability in the relationship of VWM to other cognitive components (Gathercole et al., 2004). The mechanisms underlying VWM development and its relationship with other cognitive measures remain debated (Astle & Scerif, 2011). That is, what are the fundamental cognitive mechanisms underlying developmental improvements in VWM? Most traditional measures of VWM rely on indices of the number of items that can be remembered, e.g. using tasks that measure recall span or change detection. These reports rely on the classical concept that VWM is limited to holding only a very small number of items, as little as 3 or 4 in adults, arguing in favour of a ‘quantal’ memory mechanism limited to 3–4 memory ‘slots’(Luck & Vogel, 1997). In effect, this is a digital view of how items are stored in memory, with both span and change detection tasks providing a discrete estimate of the number of items that can be retained in VWM. Recently, an innovative empirical approach has shown potential to accurately and sensitively characterise VWM performance in a completely different way. Unlike span or change detection tasks, this approach relies on participants reproducing the exact qualities of the retained information on a continuous response scale, e.g. the orientation of a bar stimulus. Such analogue responses provide a measure of working memory precision, where precision reflects the resolution with which items are held in VWM (Bays, Catalao, & Husain, 2009; Fougnie, Asplund, & Marois, 2010; Wilken & Ma, 2004; Zhang & Luck, 2008). Recent investigations suggest that continuous recall measures may be more sensitive indices of changes in VWM than discrete capacity measures, e.g. in adult neuropsychological populations (Zokaei, Burnett Heyes, Gorgoraptis, Budhdeo, & Husain, 2014). Findings from such delayed reproduction tasks have also challenged the ‘quantal’ view of VWM. Results from such studies are not consistent with the view that there is a fixed upper limit to the number of items that can be stored in VWM (Bays et al., 2009; Gorgoraptis, Catalao, Bays, & Husain, 2011; Ma, Husain, & Bays, 2014; Zokaei, Gorgoraptis, Bahrami, Bays, & Husain, 2011). Instead, these investigations have shown that response data can be modelled by considering VWM to consist of a limited cognitive resource (Bays & Husain, 2008), with recall error arising due to noise in tuned populations of neurons (Bays, 2014, 2015; Ma et al., 2014). As the number of items stored increases, the precision with which each item can be retained decreases, purportedly because the same pool of tuned neurons must represent more items, and thus the representation of each item becomes noisier (Bays, 2014; Bays, 2015; Ma et al., 2014). Using the statistical procedure of mixture modelling of response data, continuous VWM tasks also provide a means to dissect out sources of error contributing to the pattern of overall VWM performance (Bays et al., 2009). In other words, what basic, fundamental cognitive processes contribute to working memory failure and success? Firstly, response error can theoretically arise due to variability in memory – ‘noise’ – associated with the remembered feature; that is, how accurately information (e.g. orientation) is stored. Alternatively, error can arise due to random guessing, for example, due to failures at encoding or at retrieval leading to a completely random response. Lastly, error may arise due to systematic interference or corruption of information by the other items encoded into VWM. These latter misbinding responses occur when one feature of an object (e.g. colour) is erroneously linked with the feature (e.g. orientation) of a different object stored in memory (Bays et al., 2009). Crucially, characterising VWM development in terms of these distinct sources of error may provide further insights into underlying developmental cognitive mechanisms. The current study investigated development of VWM precision longitudinally between age 7 and 13 years. Forty participants completed an identical task battery twice, two years apart. Tasks consisted of 1-item and 3-item sequential VWM precision tasks and a sensorimotor control task. Precision was calculated as 1/SD of error (reciprocal of variability) in response. After controlling for developmental changes in sensorimotor performance, we predicted longitudinal increases in recall precision on both the 1-item and 3-item VWM tasks. In addition, we predicted mixture modelling of recall data from the 3-item VWM task would provide evidence that effects of age are attributable specifically to a decrease in variability in the representation of target stimuli, and not to any changes in random guessing or misbinding. That is, the mixture modelling would shed further light on more fundamental cognitive mechanisms underlying memory development—whether changes in VWM performance are due to developmental increases in continuous memory resource, or to some other process such as reduced frequency of misbinding or guessing.","40 participants were recruited from a single-sex (male) preparatory school in Oxfordshire and tested on the same protocol twice, two years apart (t1 and t2). Participant age at t1 ranged from 7.9 to 11.7 years, with mean 10.2 and standard deviation (SD) 1.02 (see Table 1). Participants were from a larger (N = 90) cross-sectional cohort (Burnett Heyes, Zokaei, van der Staaij, Bays, & Husain, 2012). The current longitudinal cohort of 40 participants consisted of all t1 participants who were still attending the school at t2 for whom parent/guardian consent could be obtained at t2 and for whom timetabling constraints permitted attendance of the testing session at t2. FSIQe ~~~~~ Standardized yearly test scores (CAT-3; www.gl-assessment.co.uk) were provided by the school for each participant at t1 and transformed into estimated full-scale IQ (FSIQe; see Table 1), as per Wright, Strand and Wonders (2005):FSIQe = 1.1 × CAT-Av − 12.0where FSIQe is estimated FSIQ and CAT-Av is average CAT-3 score calculated by combining standardised scores on verbal, non-verbal and quantitative reasoning subtests (Wright et al., 2005). Colour-naming task At the start of the experiment, each participant completed a colour-naming task in which they were shown five screenshots from the VWM task, each containing a bar in one of the following stimulus colours: red, yellow, green, blue and purple/pink. Participants were asked to name aloud each of these five colours. All participants successfully completed this task. Sensorimotor control task Directly after completion of the colour-naming task, participants completed 25 trials of a sensorimotor control task (Fig. 1A). Stimuli were presented on a laptop monitor (32° × 19°) at a viewing distance of ∼52 cm. On each trial a coloured oriented bar (2 × 0.2° of visual angle) was presented against a grey background. After 500 ms following the presentation of the bar, a probe bar of the same colour surrounded by a black circle appeared below the target. Participants were asked to adjust the orientation of the probe bar using a rotating dial to match the orientation of the target which remained on screen until response. The black circle surrounding the probe item disappeared upon rotating the dial. Bar colour was selected so as to be easily distinguishable—as in the colour-naming task. The orientation of the target and probe were independently randomized. The inter-trial interval (ITI) was 500 ms. Visual working memory task: 1-item Directly after completion of the sensorimotor control task, participants completed 30 trials of a 1-item VWM task. On each trial, participants were presented with a single coloured oriented bar at the centre of the screen for 500 ms. Following a blank 500 ms delay, a probe bar of the same colour appeared at the centre of the screen surrounded by a black circle. Participants were asked to adjust the probe bar’s orientation to match the orientation of the probed (‘target’) bar using the dial (Fig. 1B). Participants completed 30 trials of this task divided into 2 blocks of 15 trials each with a short break in between blocks. During the break participants were encouraged to focus on a far point in the room, to minimise ocular fatigue. The ITI was 500 ms. Visual working memory task: 3-items Directly after completion of the 1-item VWM task, participants completed 90 trials of a 3-item VWM task. A schematic presentation of this task is shown in Fig. 1C. On each trial, a sequence of 3 coloured bars was presented at the centre of the screen against a grey background. Each bar was presented for 500 ms followed by a 500 ms blank interval prior to the presentation of the next bar. At the end of each trial, a randomly oriented probe bar of the same colour as one of the bars in the previous sequence was presented within a black circle. Using the dial, participants were asked to match the probe’s orientation to the orientation of the similarly-coloured bar in the preceding sequence (the ‘target’). All colours/serial positions in the sequence were probed with equal probability. Participants completed 90 trials of this task, with a break every 15 trials. The ITI was 500 ms. Developmental changes in precision We calculated recall precision as the reciprocal of the circular standard deviation of response error (where response error is the difference between response and target angles) (Fisher, 1993). Precision is a measure of response variability: less variability corresponds to more precise recall. Recall precision was calculated for the sensorimotor, 1-item and 3-item VWM tasks. To evaluate the effect of age on VWM precision controlling for any changes in sensorimotor performance, we corrected performance on the VWM task by subtracting sensorimotor error from VWM recall error and recalculating precision accordingly (i.e. 1/sqrt(difference in error variance)). To evaluate developmental changes in precision, Wilcoxon signed-rank tests (the non-parametric equivalent of paired samples t-tests) were conducted comparing t1 vs. t2 corrected 1- and 3-item VWM precision values (overall and at each serial position) as well as sensorimotor precision. Statistical significance was p < 0.05. Non-parametric statistics were used since in our data precision is not normally distributed. Finally, to investigate differential trajectories of development for 3-item relative to 1-item VWM tasks, Wilcoxon signed-rank tests were conducted comparing t1 vs. t2 3-item VWM precision values, corrected for 1-item VWM performance by subtracting 1-item VWM recall error from overall 3-item VWM recall error and recalculating precision accordingly (i.e. 1/sqrt(difference in error variance)). Mixture modelling of error in response Several sources of error could contribute to developmental changes in performance on VWM tasks such as those employed in the present study. Firstly, changes in VWM performance could result from a change in variability in memory for target features – here orientation – captured by the model concentration parameter (κ). κ is a measure of variability, where higher κ corresponds to lower variability in memory representations (Fig. 2A). Successful performance of the 3-item VWM task also requires memory for the correct combination of orientation and colour. Therefore, changes in performance on the 3-item VWM task could arise as a result of changes in the proportion of responses arising as a result of an incorrect conjunction of colour and orientation (misbinding errors). In such trials participants make an error centred on the orientation of other (non-probed) items in the sequence (Fig. 2B). In clearer terms: if the probed item is red but participants respond with the orientation of one of the other coloured bars in the sequence, this would be classified by the model as a misbinding error. In the model, the probability of reporting a non-target item is given by β, with {ϕ1, ϕ2 … ϕm} the orientations of the m non-target items. Alternatively, change in recall precision could occur due to changes in the number of guesses/random responses—captured in the model by γ (Fig. 2C), where γ=1-α-β. Maximum likelihood parameters of κ, α, β and γ were obtained for each task using an expectation maximization procedure for each participant (Myung, 2003). Using this probabilistic model, we were able to determine the underlying sources of developmental change in VWM. Similar to our analysis of recall precision, we used paired t-tests to compare t1 vs. t2 measures of κ and the proportion of target, non-target and random responses. Sensorimotor precision during development ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Sensorimotor precision improved significantly between t1 and t2 (Z = 3.19, p = 0.01; Table 2). Therefore, in the remaining analyses of performance in the 1-item and 3-items VWM tasks, we corrected performance for changes in sensorimotor precision (for details, see Section 3). Working memory precision improves with age on the 1- and 3-item VWM tasks ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Mean precision on the 1-item VWM task improved significantly after 2 years after correcting for developmental change in sensorimotor precision (Z = 2.87, p = 0.004; one outlier >2.5 SD > mean excluded). Thus, precision of recall for even a single item maintained in memory increased after 2 years in childhood and early adolescence. Recall precision on the 3-item VWM task also improved significantly with age. Participants performed significantly better at t2 compared to t1 (Z = 2.39, p = 0.017; Fig. 3A). Wilcoxon signed-rank tests comparing t1 vs. t2 corrected precision at each serial position of the target demonstrated a significant improvement in WM precision for items presented first and second in the sequence (Z = 2.05, p = 0.04 and Z = 3.06, p = 0.002 respectively) with no significant difference for the third item (Z = 1.57 p = 0.12; Fig. 3A). Improvement on 3-item WM task greater than improvement on 1-item VWM task ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Comparison of t1 vs. t2 3-item VWM precision corrected for 1-item VWM precision showed evidence for a steeper developmental trajectory of precision in the 3-item VWM task relative to the 1-item VWM task (Z = 2.79, p = 0.005; see Fig. 4 for individual participants’ change in precision). Mixture modelling of error in response ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To visualize the distribution of responses in the 3-item VWM task, we plotted the distribution of responses around the target (i.e. probed) orientation. As shown in Fig. 3B, the proportion of responses falling close to the target orientation increased from t1 to t2. This is specifically illustrated in the peak of response distribution around zero, i.e. the greater proportion of responses at the target orientation and lower proportion of responses at the tail of the distribution. However, overall performance does not inform us as of the sources of error, and how these alter with age. That is, overall performance does not shed light on whether improved VWM performance is due to decreased variability in response for the target feature (kappa), changes in proportion of target (p(T)) or non- target responses (p(NT)), or increased random responses (p(U)). We therefore fit a probabilistic model (Fig. 2; Bays et al., 2009) to each participant’s 3-item VWM dataset at each time point to examine the effect of age on each of the possible sources of error. Results show that kappa (inverse of variability in response around probed or target orientation) increased significantly with age (t(39) = 3.3, p = 0.002), crucially with no change in other model parameters: t(39) = 1.6, p = 0.12 for p(T), t(39) = 1.4, p = 0.18 for p(NT and t(39) = 0.3, p = 0.73 for p(U) (Fig. 3C). Thus variability around the probed target orientation improved significantly with age, without other sources of error changing.","The current study investigated longitudinal development of VWM precision during childhood and early adolescence. Forty participants aged 7–11 years at t1 completed a VWM precision task battery twice, two years apart. Results demonstrate firstly that recall precision increased with age on both 1-item and 3-item sequential VWM precision tasks. These increases remained significant after controlling for age-related improvement in performance on a sensorimotor control task. Second, the longitudinal effects of age were attributable to a specific decrease in variability in the representation of target stimuli, and not due to changes in random responding or to misbinding, i.e. corruption by features of other items retained in memory. We discuss implications in terms of developmental mechanisms underlying observed improvements in VWM performance with age. Development of working memory precision ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A number of studies have shown improvement across childhood in performance on standard tests of VWM (Alloway et al., 2006; Gathercole et al., 2004). In the current study, we contribute to understanding the mechanistic underpinnings of VWM development. We show that the precision with which items are recalled from VWM, whether presented individually or in sequences of three, increased after a time interval of two years in participants aged 7–11 years at study entry. Because participants were sampled longitudinally, we can rule out interpretations based wholly or partially on inter-individual variability and cohort effects, enabling conclusions to be drawn regarding the developmental trajectory of VWM precision. Importantly, longitudinal increases in precision withstood correction for improvements in performance on a control task requiring fine hand-eye co-ordination, and hence are not readily explicable on the basis of improvement in sensorimotor factors. Instead, the results suggest that the resolution of items recalled from VWM increases during middle childhood and early adolescence. These conclusions align with those from cross-sectional studies showing age-associated development during childhood and adolescence in VWM performance (Burnett Heyes et al., 2012; De Luca et al., 2003; Luciana, Conklin, Hooper, & Yarger, 2005; Swanson, 1999; Zald & Iacono, 1998). The current results are also consistent with studies showing developmental improvements in performance on executive tasks that have a VWM component (Brocki & Bohlin, 2004; Luciana et al., 2005; Luna, Garver, Urban, Lazar, & Sweeney, 2004). Using continuous recall measures, rather than discrete measures of capacity, may offer enhanced sensitivity for tracking behavioural changes in VWM, as has recently been demonstrated using even relatively small samples (N = 12) of adult patients with neurodegenerative conditions (Zokaei et al., 2014). Importantly, the approach taken here also enables investigation of the underlying cognitive mechanisms associated with these developmental changes. Modelling the distribution of responses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Continuous VWM tasks provide a means to examine the sources of error in recall. We applied a probabilistic model (Bays et al., 2009) to response data from the 3-item VWM task to decompose the contributions of different sources of error. Within this model, error in memory can arise from three sources. Firstly, it can be due to variability – ‘noise’ – in memory for the remembered feature. Alternatively, error can arise due to random guessing, for example due to failures at encoding or retrieval. Lastly, response error in the 3-item VWM task may also theoretically arise due to systematic interference or biasing of information by features belonging to other items encoded into VWM (i.e. misbinding or non- target responses). Here, we show that improvement with age in the 3-item VWM task was attributable to a specific decrease in variability in the representation of target stimuli, and could not be attributed to any changes in the frequency of guesses or misbinding errors. As such, our modelling analysis sheds new light on potential specific mechanisms underlying longitudinal development in VWM performance. In future, delayed reproduction VWM tasks accompanied by probabilistic modelling of response data could also provide a useful tool for characterising more fully neural changes associated with VWM across the lifespan. Several studies have shown that it is possible to track development of the functional neural substrates of VWM during childhood and adolescence (Bunge & Wright, 2007; Dumontheil & Klingberg, 2012; Geier, Garver, Terwilliger, & Luna, 2009; Klingberg, Forssberg, & Westerberg, 2002). Using precision, rather than traditional measures of capacity, may offer enhanced sensitivity for detecting variance associated with age. Differential development on 1-item and 3-item VWM tasks ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The current study showed effects of age on the difference in recall precision between 1-item and 3-item tasks, conceptually equivalent to an age by task interaction. That is, whereas we demonstrate significant improvement with age on both VWM tasks, improvement on the harder 3-item task was more substantial. Further empirical studies are needed to explore the developmental cognitive mechanisms underlying this result. Possibilities include development of attention (Astle, Nobre, & Scerif, 2012) or even emerging metacognitive capability (Weil et al., 2013).","We present longitudinal evidence that the precision of VWM for individually and multiply- encoded items develops throughout middle childhood and early adolescence. This development is attributable to a specific decrease in variability – noisiness – of stored feature representations. This highlights a mechanism that may underlie longitudinal improvement in VWM performance, without invoking improvement in the number of items that can be stored. These findings demonstrate protracted development in VWM performance and shed light on the potential mechanisms which could underpin this development longitudinally. Currently, there is much interest in modifying cognition as a means of addressing suboptimal developmental trajectories (Klingberg et al., 2005). VWM recall precision might provide a sensitive metric for assessing the impact of behavioural training and neurofeedback interventions (Wass, Scerif, & Johnson, 2012)."],["Objectives: This study investigated different patterns of physical activity (PA; frequency, intensity, and duration) among employees during and after participating in a worksite health-promotion intervention over a period of one year. The study aimed to assess whether different patterns of PA were associated with perceived competence and motivational regulations for PA. Design: A cluster randomized controlled trial with a delayed-intervention control group. The design of the group-based intervention was based on the tenets of Self-determination theory (SDT). Method: The study consisted of employees (N = 202, M age = 42.5) working with manual labor in an (Anonymized) transport and logistics company. A person-centered approach was applied in order to explore if there were different latent trajectories within the sample related to PA. The data was analyzed with latent class growth analysis (LCGA) and the modified BCH method. Results: The LCGA identified three PA trajectories: (1) employees high at baseline who declined significantly (n = 16), (2) employees who remained stable at a moderate level (n = 55), and (3) the majority of employees who reported low levels at baseline and increased significantly (n = 128). High levels of PA were associated with higher levels of perceived competence and autonomous forms of motivation for, which is in line with the tenets of SDT. Contrary to study hypothesis, controlled forms of motivation increased in all three trajectories after the intervention. Conclusions: Different trajectories of PA were found, and the intervention was able to attract employees with low levels of PA. --------------------------------------------------------------------------------","Despite great media attention and increased public awareness, people struggle to be physically active at the level required to maintain their health and well-being, and reduce their risk of chronic diseases. A national survey among (Anonymized nationality) adults revealed that only 35% reported being sufficiently physically active as recommended by the (Anonymized nationality) health authorities (150 min of moderate physical activity [PA], or 75 min of vigorous PA, per week; Hansen et al., 2015). Participating in a health promotion program can provide the necessary structure and support to initiate changes related to regular PA or objective measures of PA effects such as cardiorespiratory fitness (CRF). Composite interactive interventions that apply self-management and motivational enhancement approaches have demonstrated the most promising results in terms of effectiveness (Hutchinson & Wilson, 2011; Michie, Abraham, Whittington, McAteer, & Gupta, 2009). However, program participation is typically limited in time, particularly in non-treatment contexts. In order to produce changes in health and well-being of clinical relevance to the individual and to society in general, participants must be able to persist with lifestyle changes over a longer period of time – and on their own. Consequently, it is of importance that they develop a sense of competence and an autonomous motivation to persist with PA. Over the last four decades, the field of health promotion research has called attention to the worksite context because programs here have the potential to reach a large number of people, usually before they develop health problems (Abraham & Graham-Rowe, 2009; Rongen, Robroek, van Lenthe, & Burdof, 2013). Employers are willing to invest financial resources in programs because they appreciate the potential benefits of increased PA for health and well-being, such as decreased sickness absence (Cancelliere, Cassidy, Ammendolia, & Côté, 2011) and improved work productivity (Pronk & Kottke, 2009). Moreover, the presence of natural and lasting social networks offers a source of social support that can be incorporated into programs, and may persist after the program has finished (Linnan, Fisher, & Hood, 2012). Meta-analyses have demonstrated that worksite PA interventions can offer important albeit variable changes in health, well-being, and certain worksite outcome measures such as reduced job stress (Conn, Hafdahl, Cooper, Brown, & Lusk, 2009). There is a growing number of systematic reviews and meta-analyses of worksite PA promotion studies. Overall, they report positive effects albeit small effect sizes (Cohen’s d = 0.10–0.27) for self-reported measures of PA (Abraham & Graham-Rowe, 2009; Conn et al., 2009; Dishman, DeJoy, Wilson, & Vandenberg, 2009; Malik, Blake, & Suggs, 2014; Proper et al., 2003). The research evidence related to objective measures of fitness, such as CRF or muscle strength, is inconclusive and divergent (Proper et al., 2003). For instance, meta-analyses have reported positive effect sizes for CRF ranging from d = 0.29 (Abraham & Graham-Rowe, 2009) to d = 0.57 (Conn et al., 2009). High-quality randomized controlled trials tended to report lower effect sizes or non-significant effects compared to quasi-experimental and pre-post studies, and to studies with less rigorous methodology (e.g., randomization procedure poorly implemented or described, lack of intention-to-treat analysis, lack of control for confounders, lack of objectively measured outcome variables, and short follow-up assessments; Rongen, Robroek, van Lenthe, & Burdorf, 2013; To, Chen, Magnussen, & Kien, 2013). Moreover, Taylor, Conner, and Lawton (2012) reported findings indicating that worksite PA intervention studies using theory in an explicit and systematic manner were considerably more effective (Cohen's d = 0.34) compared to studies which did not (d = 0.21). However, a systematic review concluded that worksite programs had relatively low participation rates (M = 33%, the majority below 50%), and males, blue-collar workers, and smokers were less likely to participate (Robroek, van Lenthe, van Empelen, & Burdorf, 2009). Studies have revealed that employees have mixed feelings towards worksite health-related PA programs. For example, Fletcher, Behrens, and Domina (2008) found that employees perceived social support and their own levels of PA self-regulation to be the most important enabling factors for participating in worksite PA programs. The most frequently reported barriers, apart from lack of time, were increased self-consciousness and a lack of belief in their own ability to perform PA. Rossing and Jones (2015) found that employees were sensitive to the possible loss of credibility and stigmatization from colleagues if they appeared less competent or fit during collective exercise sessions at work. We argue that in order to attract employees broadly, and particularly those who will benefit the most due to their low levels of PA and an unfavorable health risk profile, PA promotion programs must offer support in a manner that makes employees comfortable, and increases their competence, and self-regulation regarding PA. The theoretical foundation of the present worksite PA intervention was based on the tenets of Self-determination theory (SDT; Deci & Ryan, 1985; 2000) in combination with elements from Motivational interviewing (MI; Markland, Ryan, Tolbin, & Rollnick, 2005; Miller & Rollnick, 2013). SDT is a theory of motivation that emphasizes the importance of the quality of motivation towards a specific behavior, or what people hope to obtain by doing the behavior. SDT presents a multidimensional approach to motivation, distinguishing between three types of motivational qualities: autonomous, controlled, and amotivation (Deci & Ryan, 2000). Autonomous motivation is characterized by a sense of choice and freedom from external pressure. Here, people engage in a behavior because they find it inherently satisfying (intrinsic regulation) or because they identify with the behavior and find it personally meaningful (identified regulation). When motivation for a specific behavior is contingent on the presence of external factors, such as a reward or the expectations or demands of others, it is termed extrinsic regulation. Once the external control is partially assimilated, people will typically experience a sense of guilt or shame if they fail to perform the behavior in question. This is termed introjected regulation. Both extrinsic and introjected regulations are controlled forms of motivation, and are characterized by a low level of internalization (Deci & Ryan, 1985; 2000). Amotivation is characterized by a lack of motivation for a behavior, and hence a lack of intention to act (Markland & Tobin, 2004). According to SDT, these different forms of motivation are not mutually exclusive, and people can simultaneously endorse controlled and autonomous motives for a behavior (Deci & Ryan, 1985; 2000). However, review studies have demonstrated that autonomous motivation has a consistent and positive effect on outcome variables related to health and well-being (Ng et al., 2012). The SDT based health model of behavior change postulates that in order to make lifestyle changes, such as increased PA, people need to perceive themselves as sufficiently competent as well as motivated (Williams, Gagné, Ryan, & Deci, 2002). When people feel unfit, unskilled, inexperienced, or restricted by health limitations or lifestyle situations that they struggle to overcome, their sense of competence will be affected (Ryan, Williams, Patrick, & Deci, 2009). A review of 53 PA studies demonstrated a consistent and rather strong association between autonomous forms of motivation for PA and prolonged PA (Teixeira, Carraça, Markland, Silva, & Ryan, 2012). However, the presence of a strong association between controlled forms of motivation and PA has not received consistent empirical support. The majority studies (57%) found no significant association, whereas the remainder (43%) reported a negative relation (Teixeira, Carraca, Markland, Silva, & Ryan, 2012). The same was found for the association between amotivation and PA. Teixeira and colleagues revealed that studies reporting negative associations between extrinsic regulation and PA considered regular exercisers, as opposed to non-exercisers. In line with SDT, regular exercisers have been found to report higher levels of autonomous motivation for PA compared to exercise initiates and non-exercisers (Thøgersen-Ntoumani & Ntoumanis, 2006). A study of four longitudinal datasets investigated the process of becoming a regular exerciser and how this is related to motivational regulation (Rodgers, Hall, Duncan, Pearson, & Milne, 2010). Participants who had completed a PA intervention program reported increases in intrinsic and identified regulation eight weeks after baseline. Despite a steady increase, autonomous motivation remained significantly lower among exercise initiates compared to regular exercisers six months after baseline. Changes in controlled motivation were non-significant during the study period. To date, there are a limited number of SDT-based intervention studies that incorporate follow-up assessments of motivational regulation and PA several months or years after the intervention. Silva and colleagues found that autonomous motivation predicted enhanced maintenance of behavioral change two years after the intervention (Silva et al., 2011). Duda et al. (2014) reported improvements in PA of clinical relevance immediately after an intervention, and the changes were sustained at six months follow-up. Moreover, changes in motivational processes during the intervention were significantly related to PA levels at follow-up. Sweet, Fortier, and Blanchard (2014) investigated the longitudinal effects of a PA intervention on sedentary patients by means of hierarchical linear modeling of growth trajectories. The study found a curvilinear trend for PA (increased at 13 weeks, and decreased post-intervention between 13 and 25 weeks). In line with the tenets of SDT, both intrinsic and identified regulation demonstrated a pattern of linear increase during the 25-week study period. However, the changes in extrinsic regulation followed the same curvilinear trend as PA. No fluctuation was found for introjected regulation, and the two controlled forms of motivation were not significantly related to changes in PA. We need more knowledge about the nuances of longitudinal fluctuations in PA and their relationships with motivational regulations, particularly in the period after PA interventions when participants are expected to persist with their PA habits without the support of a program. The majority of SDT-based PA intervention studies have been carried out in the context of health care, often incorporating patients in need of treatment, such as those who are overweight or obese (Silva et al., 2011), have type 2 diabetes (Sweet et al., 2009), or require cardiac rehabilitation (Mildestvedt, Meland, & Eide, 2008). Fortier and colleagues recommend that future studies assess intervention effects on groups that are more diverse in terms of demographic characteristics, such as age and gender, motivational regulation, and physical characteristic related to health (Fortier, Duda, Guerin, & Teixeira, 2012). The number of SDT-based PA intervention studies has grown in numbers, particularly studies with children (Owen et al., 2016) and adolescents (e.g., Lonsdale et al., 2016). However, PA intervention studies directed at adults are limited in numbers and participants are by and large patients in treatment contexts. There is a need for SDT-based intervention studies in non-treatment contexts targeting a more heterogeneous population in terms of PA levels, health risk profile, and motivation for PA. It can be argued that this will contribute to the applicability and effectiveness of SDT based intervention principles. Two SDT-based PA interventions have been carried out in the worksite context, and they both reported increases in PA in addition to positive associations between adherence, autonomous motivation for PA, and increases in cardiorespiratory fitness (Thøgersen-Ntoumani, Loughren, Duda, Fox, & Kinnafick, 2010; Thøgersen-Ntoumani, Ntoumanis, Shepherd, Wagenmakers, & Shaw, 2016). Both studies incorporated university administrative personnel. This study is the first to include a sample of employees working with manual labor. In the present study, we applied latent class growth analysis (LCGA) in order to explore if there were meaningful subpopulations related to PA behavior over a period of one year. Latent growth modeling techniques, such as LCGA, are person-centered methods suited for the estimation of between-person differences in within-person change, often referred to as trajectories (Bollen & Curran, 2006). These techniques offer the possibility to “model unobserved heterogeneity in a population by identifying different latent classes of individuals based on their observed response pattern” (Clark & Muthén, 2009, p. 3). Latent growth modeling has become increasingly popular because it is highly flexible and able to incorporate complexity such as partially missing data, nonlinear change, unequal time-points, and heterogeneous growth processes (Curran, Obeidat, & Losardo, 2010). Studies using LCGA related to PA are growing in numbers. A large cohort study explored the long-term patterns of PA involvement over a period of 22 year (Barnett, Gauvin, Craig, & Katzmarzyk, 2008). The analyses identified four distinct patterns (“inactive”, “increasers”, “active”, and “decreasers”). The results indicated that people with lower income levels and educational levels were less likely to follow the “active” trajectory and more likely to follow the “decreasing” trajectory. A recent study applied person-centered analysis to the PA levels of senior citizens resident in assisted living facilities (Park et al., 2018). Three distinct profiles were found in relation to autonomous motivation and perceived support for PA. The profile characterized as “high in both” also reported significantly higher levels of PA and more favorable impressions of exercise facilities in their physical neighborhood. The study indicates that person-centered approaches are suitable for detecting and analyzing differences in PA and their relationship to SDT based constructs, such as motivational regulations. Several studies have explored the associations between individual motivational profiles and PA among adult exercisers and athletes applying a more traditional person-centered approach; cluster analysis (e.g., Gillet, Vallerand, & Paty, 2013; Guérin & Fortier, 2012; Matsumoto & Takenaka, 2004). The studies based their clustering on motivational regulation, and all studies reported between two and five-cluster solutions related to PA (Friederichs, Bolman, Oenema, & Lechner, 2015). Friederichs et al. (2015) carried out a study on adults who did not comply with the PA recommendations, applying cluster analysis and one-way ANOVA to assess differences between clusters with regard to PA. Three clusters were found: (1) “autonomous motivation” (high on autonomous and low on controlled forms of motivational regulation), (2) “controlled motivation” (high on controlled and moderate on autonomous forms of motivational regulation), and (3) “low motivation” (moderate on controlled and low on autonomous forms of motivational regulation). Cluster (1) reported the highest levers of PA, and cluster (3) the lowest. The results indicate that low levels of autonomous motivation is more predictive of inactivity than high levels of controlled motivation. Moreover, the motivational profiles reported were in fact similar to those found in other studies, both among non-exercisers (Guérin & Fortier, 2012) and regular exercisers (Matsumoto & Takenaka, 2004). Cluster (1) accounted for 52.9% of the sample, and the sample could possibly be biased by the fact that the participants were recruited among individuals who had agreed to participate in a web-based PA intervention. The present study applied PA levels as the basis for a person-centered approach and included the perceived competence and motivational regulations for PA as distal outcome variables. We used a rather modern approach, the BCH method, to assess how the distal outcome variables related to underlying patterns of PA. This is a tree-step approach recommended by Asparouhov and Muthén (2014) because it avoids the undesirable shift in the latent class variables caused by the direct inclusion of the distal outcome variable in the analyses. Moreover, the above-mentioned studies applied cross-sectional data, whereas the present study explored if there were latent classes explaining different patterns of PA change during the course of an intervention. Study aim and research questions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ First, we aimed to explore whether there were latent classes or trajectories in the sample related to PA over a period of one year (prior to, during, and after a worksite PA promotion intervention). Second, we aimed to investigate whether the intervention was able to recruit employees with different levels of PA, particularly those with low levels. Moreover, we examined whether the possible changes during the study period were clinically relevant according to the recommendations of the Norwegian health authorities. Third, we aimed to assess whether demographic characteristics of the participants could predict differences in PA patterns. Fourth, we aimed to explore whether different patterns of PA were associated with the participants' perceived competence and motivational regulation for PA at baseline and follow-up. Regarding the fourth aim, the following hypotheses were tested: Employees reporting higher levels of PA were expected to have higher levels of perceived competence for PA compared to those employees reporting lower levels of PA. Employees reporting higher levels of PA were expected to have higher levels of autonomous motivation for PA (intrinsic and identified regulation) compared to those employees reporting lower levels of PA. Employees reporting higher levels of PA were expected to have lower levels of controlled motivation (introjected and extrinsic regulation) for PA compared to employees reporting lower levels of PA. Employees reporting higher levels of PA were expected to have lower levels of amotivation for PA compared to employees reporting lower levels of PA.","The study sample consisted of employees participating in a worksite health promotion program designed to support them in increase their PA. They were employed in the logistics sector working as drivers, mail carriers, and terminal workers. The study sample consisted of predominantly male participants (76.2%). The participants were between 19 and 68 years, and mean age was 42.5 years (SD = 11.65). Education levels were relatively low, and only 14.3% had a college degree. Participants were recruited during team-based information meetings at six worksites. A total of N = 320 were defined as eligible (working more than 20%), and n = 202 (68%) agreed to participate (written informed consent). Participants were assessed at three time-points; at baseline, post-test (5 months), and at follow-up (12 months). Questionnaires were applied to measure regular PA in addition to the motivational and demographic variables at all three time-points. The baseline and post- test assessments were in the form of health screenings that assessed their cardiorespiratory fitness, biomedical health markers (e.g., blood pressure, waist circumference, and cholesterol levels), and their lifestyle. A health practitioner presented them with the results (health status and risk factors), recommended lifestyle changes, and in some cases advised them to consult their physician for further testing and medical treatment. They received an individual, written report of their health profile. After baseline assessments in January, participants were randomized by means of six clusters (worksites) into an intervention condition (n = 113, 56%) and a control condition (n = 89, 44%). The former received a group-based intervention consisting of six sessions (two workshops and four exercise support-group meetings) and a booklet. The sessions were dialogue-based, and PA was expected to be self-organized, primarily during leisure time due to shift work and a lack of onsite exercise facilities. The design of the intervention elements were based on a model that combined the tenets of SDT with techniques from MI, which had previously been applied in PA intervention studies (Fortier et al., 2012). The workshops were facilitated by two health and exercise advisors (physiotherapists) who were trained to provide the workshops in a manner that supported the basic psychological needs according to study protocol. Pre-post intervention effects related to cardiorespiratory fitness, PA, cholesterol, blood pressure, and waist circumference have previously been published together with statistical power calculations and detailed intervention protocol descriptions (Anonymized). Participants in the control condition were offered a delayed group-based intervention eight months after baseline. The delayed intervention consisted of standard group-based sessions offered by the worksite health promotion program. Both conditions were presented with a follow-up assessment 12 months after baseline, and a total of n = 114 (55%) agreed to participate, of these n = 62 (55%) were from the intervention condition and n = 52 (45%) were from the control (delayed intervention) condition. A total of n = 195 participants completed the assessments at baseline, n = 155 completed at post-test, n = 114 completed at follow-up, and n = 101 (50%) completed all three assessments. The study was approved by the Data Protection Official for Research in (Anonymized). Educational level Education level was assessed with a one-item questionnaire applying the following scale: (1) primary and secondary school (10 years), (2) high school (13 years), (3) college/university degree (1–4 years), and (4) college/university degree (more than 4 years). Physical activity PA was measured with the three-item questionnaire International Physical Activity Index (IPAI), which was previously applied and validated on a large sample in (the HUNT study; Kurtze, Rangul, Hustvedt, & Flanders, 2008). The questionnaire assesses the frequency (number of sessions per week), the duration (the length of the PA sessions in minutes), and the intensity (the amount of energy expended during the sessions). Frequency was assessed with the item: “How frequently do you exercise?” using the following scale: 0 (Never), 1 (Less than once a week), 2 (Once a week), 3 (2–3 times per week), and 4 (Almost every day). Intensity was assessed with the item: “How hard do you push yourself?” using the following scale: 0 (I do not exercise), 1 (I take it easy without breaking into a sweat or losing my breath), 2 (I push myself so hard that I lose my breath and break into a sweat), and 3 (I push myself to near-exhaustion), Duration was measured with the item: “How long does each session last?” using the following scale: 0 (I do not exercise), 1 (Less than 15 min), 2 (16–30 min), 3 (30 min to 1 h), and 4 (More than 1 h). According to protocol, each item’s score was multiplied with a weighing factor listen in parenthesis: frequency 0 (0), 1 (0,5), 2 (1), 3 (2.5), and 4 (5), intensity 0 (0), 1 (1), 2 (2), and 3 (3), and duration 0 (0.1), 1 (0.38), 2 (0.75), and 3 (1). The weighted scores were then multiplied to calculate a summary index (Kurtze et al., 2008). Perceived competence for PA Participants rated their sense of perceived competence regarding PA by means of the Perceived Competence in Exercise Scale (PCES; Williams & Deci, 1996). The questionnaire consists of four items (e.g., “I feel confident in my ability to exercise on a regular basis”, Cronbach's αtime1 = 0.90; αtime3 = 0.94), and was answered on a seven-points Likert-scale ranging from 1 (strongly disagree) to 7 (strongly agree). Motivational regulations for PA The quality of motivational regulations was measured with the Behavioral Regulation in Exercise Questionnaire (BREQ-2; Markland & Tobin, 2004). The questionnaire consists of five subscales: intrinsic regulation for PA by four items (e.g., “I exercise because it’s fun”, αtime1 = 0.86; αtime3 = 0.89); identified regulation for PA by four items (e.g., “I value the benefits of exercise”, αtime1 = 0.76; αtime3 = 0.73); introjected regulation for PA by three items (e.g., “I feel guilty when I don't exercise”, αtime1 = 0.64; αtime3 = 0.77); extrinsic regulation for PA four items (e.g., “I exercise because other people say I should”, αtime1 = 0.80; time3 = 0.83; and amotivation for PA by four items (“I don't see the point in exercising”, αtime1 = 0.78; αtime3 = 0.80). Integrated regulation was not included in the present study because BREQ-2 does not contain the subscale. Participants responded according to a 5-point Likert scale, ranging from 0 (not true for me) to 4 (very true for me).","Preliminary analyses were performed to identify possible patterns of missing data. Dropout rates were n = 7 (3.5%) at baseline, n = 47 (23%) at post-test (5 months), and n = 88 (44%) at follow-up (12 months). Little's test of missing completely at random (MCAR) indicated that the data were not missing completely at random (x2 = 1036, df = 917, p = .004). One-way ANOVA, performed using IBM SPSS Statistics 21 (IBM Corp., Boston, Mass, USA), tested whether there were significant differences regarding the study variables between those participants who completed all three assessments and those who completed one or two. No significant differences were found, and data was assumed to be missing at random (MAR). We decided to include all N = 202 participants in the subsequent analyses applying Mplus version 8 (Muthén & Muthén, 1998–2012), and the missing data were handled by means of full information maximum likelihood estimation (FIML; Enders & Bandalos, 2001). Prior to the main analyses, the distal outcome variables (perceived competence and motivational regulation for PA) were assessed to evaluate the scale factor structure and measurement invariance (see Supplementary Material, Appendix A). LCGA were conducted on data collected at all three time-points to explore the different trajectories. The estimates of variance and covariance for the growth factor, PA, were fixed to zero assuming that all growth trajectories within each class were identical. An exploratory approach was chosen, and we did not hypothesize an expected number of classes. A stepwise model comparison approach was conducted to compare a one-class model to models with successively more classes (Nylund, Asparouhov, & Muthén, 2007). According to recommendations, a combination of goodness of fit indices (GOF) should be considered together with class sizes (>5%), theoretical justification, and interpretability in order to decide on the appropriate model (Jung & Wickrama, 2008). These following GOF indices were considered: the smallest Bayesian information criteria (BIC) and Aikaike's information criterion (AIC) to assess model fit, followed by the highest possible entropy to assess precision/quality of classification, and finally a significant p-value on the bootstrap likelihood ratio test (BLRT) and the Lo-Mendell-Rubin adjusted likelihood ratio test (L-M-R). The latter tests indicate whether the k-1 class model is rejected in favor of the k class model (Jung & Wickrama, 2008; Nylund, Asparoutiov, & Muthén, 2007). We examined the plot of each model to consider whether the differences between trajectories were logical. Finally, we considered the sample size of the trajectories. Based on the following procedure, we decided on the best model. Because PA was measured with a summary index, a manifest variable was applied as a continuous indicator of a latent class variable. We proceeded to consider the clinical relevance of the reported PA levels and changes related to the recommendations of the Norwegian health authorities (PA ≥ 150 min of MVPA per week). PA frequency scores were multiplied with duration scores (both weighted) to obtain minutes per week. The amount of participants with ≥150 m/w of moderate-to-vigorous intensity was calculated. Next, we tested whether there were differences regarding the probability of class-membership in relation to the covariates in the study. The automatic BCH approach was used for the continuous covariate (age and educational levels), while for the categorical covariates (onset of intervention and gender) the DCAT approach was used (Asparouhov & Muthén, 2014). Finally, we conducted a series of analyses to explore whether there were differences between the trajectories related to distal outcome variables (perceived competence and motivational regulations for PA). We applied the three-step BCH approach in Mplus, which offers an omnibus test that includes differences between the three classes on each distal outcome variable (Bolck, Croon, & Hagenaars, 2004). According to a comparative analysis of different approaches, the findings indicated that BCH was the most robust and flexible approach, yielding the least biased estimates (Bakk & Vermunt, 2016). Effect sizes were calculated for the differences between trajectories using Cohen's d for continuous variables and Cramer's v for categorical variables. Latent trajectories in the sample ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The stepwise comparisons of the LCGA favored a solution with three classes (Table 1). The entropy values of 0.96 indicated that both a three-class and a four-class model were able to accurately place subjects into classes. Both the AIC and the BIC decreased consistently for one-class to four-class models. However, the four-class model did not obtain a significant p-value on the L-M-R test, favoring the three-class model. In addition, the four-class model contained a class with a sample size of 4.2%, which is less than the recommended level of 5% (Jung & Wickrama, 2008). The identified classes represented three distinctly different and meaningful course trajectories (Figure 1): Trajectory 1 (prevalence: n = 16, 8% of the total sample) is labelled “Decrease from high”, and refers to subjects with the highest levels of PA and with scores significantly decreasing over a period of one year (intercept: M = 8.269, SE = 0.294, p < .001; slope: M = −1.433, SE = 0.579, p = .013). Trajectory 2 (prevalence: n = 55, 27.5% of the total sample) is labelled “Stable moderate”, and refers to subjects with moderate levels of PA and no significant change over a period of one year (intercept: M = 4.288, SE = 0.115, p < .001; slope: M = 0.090, SE = 0.227, p < .691). Trajectory 3 (prevalence: n = 128, 64.5% of the total sample) is labelled “Increase from low”, and refers to subjects with the lowest levels of PA and with scores significantly increasing over a period of one year (intercept: M = 0.700, SE = 0.070, p < .001; slope: M = 0.882, SE = 0.126, p < .001). Clinical relevance of physical activity levels and changes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ At baseline, 40% of the total sample reported PA levels consistent with or above the recommendations of the Norwegian health authorities (PA ≥ 150 min of moderate-to-vigorous physical activity per week). This percentage increased to 50.4% at post-test and to 55.7% at follow-up (Table 2). The large majority of participants in trajectories (1) “Decrease from high” and (2) “Stable moderate” remained within or close to the recommended levels of PA during all three time-points. Participants in trajectory (3) “Increase from low” reported low levels of regular PA at baseline, and 93.5% did not meet the PA levels recommended. However, they reported considerably higher levels of PA during the study period. At post-test, 31.3% met PA recommendations, increasing to 40.5% at follow-up. Controlling for the onset of the intervention ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We controlled for the onset of the intervention period (primary intervention group and delayed-intervention control group). The differences were non-significant and Cramer's v effect size was small between trajectories (1) “Decrease from high” and (2) “Stable moderate” (X2 = 0.03, p = .859, ES = 0.19), between trajectories (1) “Decrease from high” and (3) “Increase from low” (X 2 = 0.76, p = .385, ES = 0.19), and between trajectories 2 and 3 (X 2 = 2.97, p = .085, ES = 0.13). Sociodemographic covariates ~~~~~~~~~~~~~~~~~~~~~~~~~~~ We proceeded to test whether sociodemographic variables differed according to class membership. There were no significant differences between the trajectories (1) “Decrease from high” and (2) “Stable moderate” related to age, gender, or level of education. However, Cramer's v effect sizes were moderate for gender and age, and small for educational levels. There were more men in trajectory (2) “Stable moderate”, and they were somewhat older. The same pattern was found for the difference between (1) “Decrease from high” and (3) “Increase from low” (Table 3). The difference between (2) “Stable moderate” and (3) “Increase from low” were non-significant and effect sizes were small. Distal outcome variables related to competence and motivational regulation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The trajectories demonstrated a linear pattern of PA across three time-points. Hence, distal outcome variables were analyzed at baseline (T1) and at follow-up 12 months after baseline (T3). Several sets of analyses, which applied the BCH method, were carried out in order to assess whether the distal outcome variables (perceived competence for PA and motivations for PA) differed across the three trajectories. At baseline, five of the six omnibus tests were significant (p < .05), with the exception of extrinsic regulation for PA (Table 4). Employees in trajectory (3) “Increase from low” were significantly lower in perceived competence for PA compared to employees in trajectory (1) “Decrease from high” and (2) “Stable moderate”, and Cohen's d effect sizes (ES) were very large (1.20–1.57). Moreover, employees in trajectory (3) “Increase from low” reported significantly lower levels of autonomous motivation compared to the other two, and the ES were large to very large (intrinsic: 1.13–1.37; identified: 0.98–1.05; Sawilowsky, 2009). Regarding the more controlled forms of motivation, the differences were not as consistent. Employees in trajectory (3) “Increase from low” demonstrated significantly lower levels of introjected motivation, and ES were moderate (0.46–0.51). However, there were no significant differences between the trajectories related to extrinsic regulation. Employees in trajectory (3) “Increase from low” were considerably higher on amotivation compared to the two others, and ES were large (0.80–0.83). None of the distal outcome variables demonstrated a significant difference between trajectory (1) “Decrease from high” and trajectory (2) “Stable moderate”, and ES were small to very small (0.00–0.22). At follow- up, the pattern of significant differences between trajectories related to autonomous motivation for PA remained the same. Employees in trajectory (3) “Increase from low” reported considerably higher levels of autonomous motivation, and ES were moderate compared to baseline (intrinsic: 0.46–0.74; identified regulation: 0.65–0.71). The same pattern was found for perceived competence for PA, but ES were still moderate to large (0.49–0.87). Considering introjected regulation, the difference between trajectories (1) “Decrease from high” and (3) “Increase from low” was no longer significant. All the differences between trajectories that were related to extrinsic regulation were still non- significant. At follow-up, employees in trajectory (3) “Increase from low” reported lower levels of amotivation for PA compared to baseline. The difference between trajectories (1) “Decrease from high” and (3) “Increase from low” was no longer significant, and ES were small (0.00–0.35). Differences between employees in trajectories (1) “Decrease from high” and (2) “Stable moderate” remained non-significant on all distal outcome variables, and ES were very small to small (0.00–0.44).","In the present study, a person-centered approach was able to distinguish between three distinct and linear trajectories. The trajectories differed considerably in sample size, and one trajectory accounted for 65.5% of the participants. Other studies with a person- centered approach, such as Barnett et al. (2008), have found support for a four-trajectory model with a higher degree of stability. The latter study was a 22-year cohort study whereas the present study was an intervention study over a shorter period of time; just one year. Despite the considerable increase in PA among participants in trajectory (3) “Increase from low” across one year, they could possibly have returned to a stable “inactive” state after the study period. The effect of PA interventions reduces with time, and especially when the structural and social support are removed. We also explored whether the program was able to recruit employees with different levels of PA, particularly low levels. The findings indicate that the present intervention was able to attract employees who initially did not comply with the PA recommendations (70%). Moreover, they belonged to a population considered to be underrepresented in health promotion interventions, particularly in the worksite context; male employees with low educational levels and low occupational prestige (Marshall, 2004; Wong, Gilson, Van Uffelen, & Brown, 2012). Moreover, the findings indicate that the program was able to recruit a diverse sample, including a number of employees with moderate levels of PA, as represented by (2) “Stable moderate”. However, we question whether the intervention appealed to employees who were already highly active, as represented by (1) “Decrease from high”. This group could possibly have been underrepresented in the present context of eligible employees. However, this population of highly active employees was not the primary target of the program. Employees in (3) “Increase from low” initially reported considerably lower levels of PA, compared to the other employees. However, the mean value of perceived competence for PA at baseline could be characterized as moderate. According to Standage and Ryan (2012, p. 263) “feelings of competence are essential for any intentional behavior, irrespective of whether the action is motivated by extrinsic, introjected, identified, integrated, or intrinsic regulations”. The finding could indicate that the present intervention was unable to attract employees who felt inexperienced, incompetent, or unable to exercise on a regular basis. Employees in (3) “Increase from low” reported what we would describe as low-to-moderate levels of autonomous motivation, albeit significantly lower than the rest. These findings are in line with other SDT-based PA promotion intervention studies in the context of health care, which mainly attracted participants with elevated levels of autonomous motivation (Fortier et al., 2012). Third, the study aimed to test whether the associations between perceived competence, motivational regulation, and PA were in line with the tenets of SDT. Employees in trajectory (3) “Increase form low” exhibited a motivational profile and development comparable to exercise initiates previously found in a study comparing exercise initiates to regular exercisers (Rodgers et al., 2010). Both samples reported moderate levels of intrinsic and identified regulation at baseline. A review of worksite health-promotion programs reported that positive effects were mainly found in samples of motivated employees who volunteered to participate (Marshall, 2004). Employees in (1) “Decrease from high” reported significantly lower levels of PA at follow-up compared to baseline. We find it somewhat surprising that their levels of perceived competence and autonomous motivation for PA remained moderate-to-high and consistent throughout the whole period of one year. The results indicate that employees moderate-to-high on perceived competence and autonomous motivation for PA seem less vulnerable to fluctuations in PA and remain self- endorsed and confident that they are able to be physically active on a regular basis. This is in line with the findings of Sweet et al. (2014). The participation rate (68%) was considerably higher than mean values previously reported for worksite intervention programs (33%; Robroek et al., 2009). This could indicate that employees felt obligated to take part, possibly because the whole team was invited. If this was the case, we would expect the participants to exhibit relatively high levels of controlled motivation and amotivation for PA at baseline, particularly among employees in (3) “Increase from low”, in line with the tenets of SDT. However, all three trajectories reported what we considered to be low levels of amotivation and extrinsic regulation at baseline, although (3) “Increase from low” did exhibit somewhat higher levels of extrinsic regulation compared to the rest. Their initial level of introjected regulation was more apparent: the participants reported low-to-moderate levels, particularly employees in trajectories (1) “Decrease from high” and (2) “Stable moderate”. These findings indicate that participants were sensitive to and partially recognized the importance of taking part in the program and making lifestyle changes. Employees in (1) “Decrease from high” and (2) “Stable moderate” reported moderate-to-high levels of autonomous motivation for PA, particularly intrinsic regulation. This could possibly counteract their low-to-moderate levels of introjected regulation, reflecting a wish to participate in the program for their own reasons. We question whether the fact that they were not expected to participate in collective PA sessions during the intervention could have made them more comfortable since they may have felt less exposed to social comparison and loss of credibility from co- workers (Rossing & Jones, 2015). Given their moderate levels of PA at follow-up, it is not surprising that employees in the sample reported low levels of amotivation at all three time-points. However, they reported considerably higher levels of controlled motivation at follow-up compared to baseline, particularly extrinsic regulation. Employees in (2) “Stable moderate” and (3) “Increase from low” demonstrated the same pattern with regard to extrinsic regulation with an increase in mean values of around 1. Given their diverse development in PA over a period of one year, the findings did not support hypothesis 3, which was related specifically to controlled forms of motivation. Furthermore, the findings are not in line with other PA intervention studies in the health care context, which found non-significant changes in controlled forms of motivation (Rodgers et al., 2010; Sweet et al., 2014). We question whether participating in the program could actually have enhanced their controlled motivation for PA, even though their autonomous motivation remained moderate-to-high. The health screening results and recommendations together with the information, discussions and response they received during the intervention sessions could possibly have increased their awareness of the opinions and expectations of important others in their environment (e.g., family, co-workers, health practitioners, and health and exercise advisors). Participating in the program is likely to make them more sensitive to the fact that their employer invested time and money on the program in order to obtain organizational benefits, such as reduced sickness absence and increased work productivity. Although the intervention was designed to support the basic psychological needs and thereby increase autonomous motivation, it appears that aspects of the context were perceived as controlling. This is not surprising given the element of professionalism and mutual dependency between employer and employee. We argue that this is a challenge inherent in the worksite context, particularly at follow-up after the intervention period. This must be taken into consideration when designing worksite health promotion programs. Limitations and future direction ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The present study has methodological limitations. First, we completed statistical power calculations prior to recruitment in order to specify the sample size required to detect a between-groups effect size. We chose to include all participants in the LCGA, but separate power analyses were not completed. Having a too small sample size increases the risk of choosing an inadequate model with too few classes (Dziak, Lanza, & Tan, 2014). Despite of acceptable model fit indices and distinct differences between the trajectories, the three- class model could be subject to underextraction. Given the moderate sample size, we chose to explore the relationship between PA trajectories and distal outcome variables separately using manifest variables rather than latent (Andruff, Carraro, Thompson, & Gaudreau, 2009). This approach reduces the complexity but also the strength of the findings. Second, the dropout rates were considerable, particularly at follow-up, and data were not missing completely at random. Information not included in the study, such as general health condition or reasons given for not attending, may have provided a better understanding of what caused employees to drop out and what characterized those who were able or willing to participate at all three time-points. Third, in the present study, we were not able to assess PA using objective methods (e.g., accelerometers). A large number of studies have reported low-to-moderate agreement between self-reported and objectively measured PA levels (Prince, Adamo, Hamel, Hardt, Gober, & Tremblay, 2008). Both sources of PA information have the potential for under and over-estimation, and a combination is recommended (Steene-Johannessen et al., 2016). For instance, self-reported measures are biased by subjective interpretation and social desirability, whereas accelerometers do not capture all activities precisely depending on the placement on the body. Hence, the results of the present study should be interpreted cautiously. Finally, the present study also has limitations related to study design. We question whether the recruitment process could have been altered to better attend to the needs of employees with low levels of perceived competence and autonomous motivation for PA. For example, the information meetings, during which participants were recruited, could have been more dialogue-based, inviting participants to express their doubts and ambivalence more explicitly. Participants and co-workers may perceive such dialogue as being supportive of the basic psychological needs, and it may encourage them to reflect on their motivation toward PA before a decision to participate is made (Markland et al., 2005).","To our knowledge, this is the first study to apply LCGA to the investigations of the associations between longitudinal developmental trajectories of PA and SDT-based concepts of motivational regulation and perceived competence for PA using data from a PA intervention in the worksite context. The findings indicate that LCGA is a useful approach for detecting longitudinal trajectories in heterogeneous samples of both exercise initiates and regular exercisers. The present findings emphasize the effectiveness of the SDT-based intervention design and the generalizability of the results to non-treatment populations.","We would like to inform the Editor of the fact that the following PhD research project was financed by the Industrial Ph.D. scheme, administered by The Research Council of Norway: The Industrial Ph.D. scheme for funding for industry-oriented doctoral research fellowships was established to facilitate the recruitment of researchers to Norwegian industry. Funding for industry-oriented doctoral research fellowships will help many companies to step up their research efforts. The Industrial Ph.D. scheme does not represent a new type of doctoral degree, but is designed to support long-term, industry- oriented research that has the same level of scientific merit as the general doctoral degree education. The Industrial Ph.D. scheme is designed to enhance interaction between companies and research institutions, increase research activity in industry, and equip newly-educated researchers with knowledge of relevance to their company. Under the Industrial Ph.D. scheme, companies receive an annual grant equal to maximum 50 per cent of the applicable rate for doctoral research fellowships for a three-year period. The candidate must be an employee of the company and be formally admitted to an ordinary doctoral degree programme. The present PhD project was a cooperation between the Post Norway and The Norwegian School of Sport Sciences (NSSS) as the degree-conferring institution, and the candidate followed the PhD program at NSSS. For more information, please visit the web-page: https://www.forskningsradet.no/prognett- naeringsphd/About_the_PhD_scheme/1253952592832."],["Eyewitness memory for the perpetrator or circumstances of a crime is generally worse for scenarios involving weapons compared to those involving non-weapon objects—a pattern known for decades as the weapon focus effect. But despite ample support from laboratory experiments and recognition by experts, testimony concerning weapon focus is rarely admissible in court. The present article summarizes a selection of key findings within the weapon focus literature and considers whether the effect warrants consideration by the criminal justice system at this time. We conclude that weapon focus is sufficiently robust and uncontroversial to guide practice so long as consideration is given to the circumstances surrounding the criminal event with a particular emphasis on witness expectation. --------------------------------------------------------------------------------","Substantial laboratory evidence now exists that the presence of an unexpected weapon reduces the accuracy of subsequent suspect identification attempts or witness accounts of a crime (for reviews, Fawcett, Russell, Peace, & Christie, 2013; Kocab & Sporer, 2016; Pickel, 2015; Steblay, 1992). This finding is depicted schematically in Figure 1, but should not be applied indiscriminately. The magnitude (or even presence) of the weapon focus effect has been found to vary according to the characteristics of the eyewitness, the scenario in which the weapon is embedded (e.g., the perpetrator, surroundings), and the procedure through which memory is tested. As a result, each of these factors (witness, scenario, testing procedure) must be considered prior to evaluating the risk of weapon focus for any given situation. Each of these factors is discussed in turn. The Eyewitness ~~~~~~~~~~~~~~ With respect to the eyewitness, the roles of expectation and perceived threat deserve special emphasis given that the weapon focus effect has been attributed historically to the weapon drawing attention away from other details by virtue of its unexpected or threatening nature (Loftus et al., 1987). Concerning the former, a great deal of evidence has emerged over the past decade showing that witness expectation mediates the magnitude of the effect (for a quantitative model, see Erickson, Lampinen, & Leding, 2014). For example, weapon focus is diminished when the weapon is anticipated on the basis of the environment (e.g., a gun in a shooting gallery; see Figure 1) or due to prior knowledge of the individual holding the weapon (e.g., a gun held by a police officer; Pickel, 1999). Conversely, the effect is larger when the weapon violates cultural stereotypes (e.g., a woman holding a gun; Pickel, 2009, see also Sneyd, 2016). In contrast to the link between witness expectation and weapon focus, the contribution of perceived threat is less clear. Attentional narrowing due to threat was thought to play a central role by early theorists (e.g., Maass & Kohnken, 1989) and is a common feature of anecdotal reports provided by victims of weapon crime (as in our example; see also Loftus, 1979). Although meta-analyses have shown greater weapon focus in situations judged to be threatening or arousing (Fawcett et al., 2013; Steblay, 1992) or involving criminal compared to non-criminal events (Kocab & Sporer, 2016), laboratory efforts to quantify their contributions have often failed to reveal a reliable relationship (e.g., Erickson et al., 2014; Mitchell, Livosky, & Mather, 1998; Pickel, 1998; although see Peters, 1988). Importantly, weapon focus also emerges in austere scenes where threat is unlikely (e.g., Kramer, Buckhout, & Eugenio, 1990), challenging the view that threat is a necessary condition for weapon focus to occur. Convincing evidence for the role of expectation over threat also comes from studies showing that unexpected non-weapon objects elicit an effect analogous to weapon focus (e.g., someone brandishing a stalk of celery; Mitchell et al., 1998). In this respect, Pickel (1998) manipulated in her experiment whether an object carried by the perpetrator was perceived as threatening and whether it was expected given the context in which it appeared. Her study revealed worse memory for the perpetrator when an unexpected weapon (i.e., gun) or unexpected non-weapon object was held (i.e., whole raw chicken or chef doll in a barber shop) but no overall effect of threat and no interaction between threat and expectation (see also Figure 1).2 Thus, although it is possible that laboratory studies fail to elicit threat similar in nature or magnitude to that experienced during actual criminal events (although see Maass & Kohnken, 1989; Peters, 1988), there is little empirical evidence that threat or arousal are a key factor—at least under typical laboratory conditions. We will return to this idea again later, when we consider whether weapon focus research should be used to guide practice. The Scenario ~~~~~~~~~~~~ Both the scenario in which the weapon is presented and the individual holding the weapon can moderate the weapon focus effect independently of witness expectation. For example, factors diverting attention away from the weapon, such as a perpetrator with some unusual feature, will diminish the effect (Carlson & Carlson, 2012, 2014). The duration of exposure to a weapon is thought to be critical, typically eliciting a smaller effect in experiments involving either brief (i.e., <10 s) or extended (i.e., >60 s) exposure to the weapon (Fawcett et al., 2013; Steblay, 1992). However, memory for items or individuals experienced prior or subsequent to the encounter of a weapon must be considered with care, as exposure to the weapon may not influence memory for those details to the same extent (Erickson et al., 2014). For example, viewing an old acquaintance holding a weapon could result in relatively poor memory for novel features of the scene, but would be unlikely to affect memory for features that were already highly familiar. The same would apply to a stranger who draws a weapon, but then proceeds to interact with the witness for an extended period of time, either with or without the weapon. The Testing Procedure ~~~~~~~~~~~~~~~~~~~~~ Less is known about how the testing procedure interacts with weapon focus. Evidence now suggests an extended retention interval (e.g., testing after a delay of 24+ hours) diminishes the effect (Fawcett et al., 2013). However, given that retrieval accuracy is low in the weapon relative to the non-weapon condition to begin with, the decrease in weapon focus across longer delays is likely an artifact reflecting declining memory in the non-weapon condition rather than better memory in the weapon condition. Testing procedures in which participants are required to recall details relevant to the event have shown larger weapon focus effects relative to selecting a suspect from a police line-up (Fawcett et al., 2013; Kocab & Sporer, 2016; Steblay, 1992). That said, recent research indicates that the influence of weapon focus on suspect lineups may be more pronounced than originally considered. For example, witnesses appear to have difficulty discriminating a perpetrator from innocent foils in line-ups, and are more likely to commit false identifications when a weapon was involved in a mock crime (Carlson & Carlson, 2012; Erickson et al., 2014; but see, Kocab & Sporer, 2016). This finding is troubling and warrants careful consideration, particularly in light of the evidence that individuals exposed to weapons are also more susceptible to false information, introduced for instance through leading questions or exposure to police suspects (Saunders, 2009). Luckily, witnesses may be sensitive to the fact that exposure to a weapon reduces memory accuracy, shown by a greater correspondence between suspect identification accuracy in a mock police line-up and reported confidence concerning those identifications (Carlson, Dias, Weatherford, & Carlson, 2016).","Having reviewed the key findings within the literature, it is our view that the weapon focus effect is sufficiently robust to warrant consideration by the judicial system so long as the circumstances surrounding the crime are considered in tandem. Although some applications are certainly in need of further attention and development (see Table 1 for some examples), we are compelled in this evaluation by three converging sources.3 First, laboratory experiments have consistently demonstrated impaired memory for scenarios involving unexpected weapons. This provides a solid empirical basis for the weapon focus effect and its boundary conditions. Second, the weapon focus effect has been observed in a variety of simulated events (e.g., Maass & Kohnken, 1989; Peters, 1988; Pickel, Ross, & Truelove, 2006) and virtual scenarios (e.g., Kim, Park, & Lee, 2014), providing varying degrees of experimental control and ecological validity. Hence, weapon focus is observable under realistic conditions emulating natural behaviour while retaining experimental control. Finally, weapon focus is a common feature of witness accounts following exposure to a weapon (for discussion see, Loftus, 1979), suggesting that it also emerges under realistic conditions. Therefore, it is our view that the question should not be whether weapon focus influences eyewitness memory, but under what circumstances and to what degree. The circumstances under which weapon focus is most likely to occur vary in their empirical support. The most robust factor highlighted here is that of witness expectation, wherein weapon focus is most profound when the weapon is unexpected. It is therefore important to adopt a multifaceted approach to understanding the mindset of the witness when the weapon was encountered. Were there environmental- or perpetrator-specific cues that the weapon might appear? Was the perpetrator considered likely to have a weapon? Many of these considerations may be idiosyncratic and several individuals could witness the same scenario but experience weapon focus to varying degrees (if at all) dependent upon their expectations, background, and understanding of the event. Given that expectations unfold over time, it is also important to consider cues that build to the appearance of a weapon. For example, a store clerk who observes a customer acting suspiciously might expect a weapon to be present before a weapon is actually drawn. Although this scenario has not been investigated as such, the available evidence suggests that forewarning could mitigate the effect (e.g., Pickel, 2009; Pickel et al., 2006). In summary, each witness must be considered individually and in context to determine whether weapon focus is a probable concern. At least two major factors—the effects of test delay and exposure duration—have been identified largely through meta-analysis rather than primary investigations (although, see Kramer et al., 1990). These findings emerge from comparisons across studies and hence should be viewed as promising but not yet fully established, pending further investigation. Their discussion in court is consequently advocated only with suitable caution. The finding that weapon focus is larger for witness accounts than suspect identification technically falls within the same category, though this distinction is far more consistent, as demonstrated by difficulties obtaining the effect for suspect identification (for discussion, see Kocab & Sporer, 2016). The precise cause of this difference across measures is unclear. One possibility is that weapon focus simply has a larger effect on measures of recall (on which most witness accounts are based) than measures of recognition (which is the basis of suspect identification). Others have argued that the difference emerges because weapon focus has a larger impact on suspect absent line-ups whereas most investigations of suspect identification use only suspect present line-ups (Carlson & Carlson, 2012; although, see Kocab & Sporer, 2016). The etiology of this difference remains to be discovered, but the key fact is that current evidence supports a moderate effect of weapon presence on measures of feature accuracy and only a small effect of weapon presence on measures of suspect identification (Fawcett et al., 2013; Kocab & Sporer, 2016; Steblay, 1992). Thus, whereas weapon focus should still be considered in the context of a suspect line-up, the effect seems to have a greater impact on a witness's description of the perpetrator. Similarly, laboratory evidence for the role of threat has been sparse and, at present, can only be viewed to moderate (rather than elicit) the weapon focus effect. However, as touched upon earlier, we are unable to rule out the possibility that threat elicited in the laboratory fails to encompass the range of emotions experienced during actual criminal events. Different emotions, including threat, are known to cause attentional narrowing predominantly in scenarios involving an attention magnet (e.g., a dead body; Laney, Campbell, Heuer, & Reisberg, 2004) or an active personal goal capable of capturing attention (e.g., escape; Levine & Edelstein, 2009). Enhancing the salience of such magnets and personal goals using an immersive virtual reality simulation of a weapon crime has revealed complementary—but additive—effects of threat and expectation (Kim et al., 2014). This preliminary finding suggests that weapon focus could involve independent effects arising from threat and expectation that collectively influence eyewitness memory. That said, the perception of threat by witnesses in real criminal events could differ from the physiological threat or arousal experienced by those witnesses. Such differences may help to reconcile the discrepancy between frequent witness reports of feeling threatened, which is in contrast to the yet limited laboratory support for threat as a causal factor in weapon focus. This is a key area in need of greater development using further simulated events (e.g., Maass & Kohnken, 1989; Peters, 1988; Pickel et al., 2006) and applied techniques (e.g., Hulse & Memon, 2006; Kim et al., 2014) capable of balancing experimental control with the immersion necessary to emulate the threat experienced during a crime. Although considering the boundary conditions summarized above provides useful guidance, we caution that crimes rarely conform to these parameters. Weapon focus is presently expected to be largest for a brief crime involving an unexpected, threatening weapon that is committed by a previously unknown perpetrator with no pre- or post-weapon exposure and for which the witness's statement is taken within 24 h of the event. Seldom do criminal events fit this description perfectly, which may explain why most archival or field studies of actual crimes have reported little or no evidence of the weapon focus effect (e.g., Behrman & Davey, 2001; Valentine, Pickering, & Darling, 2003; but, see Tollestrup, Turtle, & Yuille, 1994). A recent meta-analysis has demonstrated an aggregate weapon focus effect when these studies are combined, but the magnitude of that effect is smaller compared to laboratory studies. Interestingly, laboratory studies that closely match typical criminal events (i.e., long exposure duration, long retention interval, high perceived threat, unexpected weapon) also show a reduced weapon focus effect comparable to the aggregate effect for archival or field studies. Difficulties observing a consistent weapon focus effect within archival or field studies therefore may reflect the influence of moderators typical of real-world crime (Fawcett et al., 2013; although, see Footnote 3). Overall, we regard the weapon focus effect to be sufficiently robust to be considered in court, but we emphasize that each witness account must be scrutinized to determine whether weapon focus is relevant, rather than presuming it to be applicable simply because a weapon was present.","Finally, in light of the evidence discussed thus far it is important to consider how the criminal justice system presently evaluates weapon focus. While criminal justice professionals acknowledge that weapon focus may negatively impact memory, they view their everyday practice as divergent from research outcomes. For example, police officers often report that weapon focus is inconsistent across cases, and memory deficits can be circumvented by enhanced interviewing skills. Further, weapon-related offences such as robbery (i.e., the “typical” weapon focus scenario) are less common than a weapon being present in domestic violence scenarios (often charged as a level 2 or aggravated assault in Canada; Statistics Canada, 2016; The Daily, 2015). The evidence summarized above should enlighten the perceived inconsistencies in this effect: whereas weapon focus may occur for a typical robbery (so long as it adheres to the specified boundaries), domestic violence cases might instead involve a variation of this effect with preserved memory for the perpetrator. In fact, in some instances police officers have described apparently greater weapon focus (and better recall of weapon-related details) without impaired perpetrator identification, owing to the victim's familiarity with the perpetrator (Edmonton Police Service, personal communication, 2016). Overall, from a policing standpoint, weapon focus is commonly acknowledged to exist, but is an issue for the courts to address. While it is clear that many judges realize the potential impact of a weapon on eyewitness memory (see R. v. Turner, 2012), few courts have recognized weapon focus as warranting expert testimony (e.g., Jordan v. State, 1996; United States v. Smith, 1984) despite a wealth of research establishing this effect (e.g., Fawcett et al., 2013; Kocab & Sporer, 2016). Even in recent cases, the judicial response has been twofold: (1) denial of expert testimony concerning weapon focus (e.g., Benton, McDonnell, Ross, Thomas, & Bradshaw, 2007), and (2) non-specific instructions to juries regarding eyewitness phenomena. Further adding to the problem is the prevalent view that “…the problems of identification are clearly within the general knowledge and comprehension of judges and properly instructed juries” (R. v. Fengstad, 1994, Supra. 74).4 However, counter to this widespread belief, the weapon focus effect is often associated with the lowest scores on surveys of lay knowledge concerning eyewitness memory for both judges and jurors, suggesting that the effect is not well understood by either population (Magnussen et al., 2008; Magnussen, Melinder, Stridbeck, & Raja, 2009). Given that weapon focus is not considered one of the five generally accepted eyewitness principles with consensus in the scientific literature (see Commonwealth v. Gomes, 2015), we would endorse the two-pronged solution proposed by Wise, Dauphinais, and Safer (2007): (a) permit expert testimony when case evidence relies heavily on eyewitness reports, and, (b) educate the principal participants in the criminal justice system. Although expert testimony on weapon focus continues to be excluded, there has been some progress on the educational front. In the Report and Recommendations of the Supreme Judicial Court Study on Eyewitness Evidence (2013), it was recommended that model jury instructions should be adjusted to incorporate acknowledgement of the weapon focus effect.5 Further, the Canadian Judicial Council (2012) recommended that judges use a series of instructions to guide jurors on how to assess witness testimony, including whether anything may have interfered with or distracted the witness from observing the details of the event and whether anything unusual happened that would enhance memory. While these recommendations are a step in the right direction, there remains no consensus concerning the evaluation of weapon focus in court.","The present article evaluated the merit of the weapon focus effect with respect to applications within the criminal justice system. We have summarized a selection of key findings within the literature and identified areas of further development. It is our conclusion that converging evidence from laboratory studies, simulated events, and actual witness accounts support the relevance of expert testimony concerning weapon focus in cases involving a weapon. In fact, such testimony could prove vital in addressing misconceptions concerning the effect amongst both judges and jurors (e.g., Magnussen et al., 2008). However, we caution that the particulars of the crime are important, and that they must reside within the boundaries for which weapon focus is thought to occur in order to promote accurate translation of empirical research to courtroom testimony.","All authors contributed to the writing of this manuscript."],["Humans have a highly developed visual system, yet we spend a high proportion of our time awake ignoring the visual world and attending to our own thoughts. The present study examined eye movement characteristics of goal-directed internally focused cognition. Deliberate internally focused cognition was induced by an idea generation task. A letter-by-letter reading task served as external task. Idea generation (vs. reading) was associated with more and longer blinks and fewer microsaccades indicating an attenuation of visual input. Idea generation was further associated with more and shorter fixations, more saccades and saccades with higher amplitudes as well as heightened stimulus-independent variation of eye vergence. The latter results suggest a coupling of eye behavior to internally generated information and associated cognitive processes, i.e. searching for ideas. Our results support eye behavior patterns as indicators of goal-directed internally focused cognition through mechanisms of attenuation of visual input and coupling of eye behavior to internally generated information. --------------------------------------------------------------------------------","The peculiar gaze of someone deeply absorbed in thought or “staring into space” suggests that eye behavior during internally focused attention might be different from states of externally focused attention. In fact, recent research provides evidence that mind wandering, the involuntary slipping away of attention from an external task to an unrelated internal train of thought, is associated with characteristic changes in eye behavior such as spontaneous pupil activity (Franklin, Broadway, Mrazek, Smallwood, & Schooler, 2013; Smallwood et al., 2011, 2012). Besides this spontaneous form of internally focused cognition, there are also many goal-directed cognitive activities which require a voluntary shift of attention to an internal focus such as planning or idea generation. For such activities external visual information is irrelevant and can even be distracting. Eye behavior is assumed to differ between states of goal-directed internally and externally focused cognition. This study aims to examine the characteristic eye behavior during goal- directed internally focused cognition. “Internally directed cognition” or stimulus- independent thought (Smallwood & Schooler, 2006) refers to cognition with an attentional focus on internally generated information rather than the external environment. Internally generated information results from memory retrieval of internal representations and combinations and modifications of those representations (mental simulations, ideas, dreams; see Chun, Golomb, & Turk-Browne, 2011; Dixon, Fox, & Christoff, 2014). The literature suggests that eye behavior may respond to and even support internal cognition in at least two ways: first, by means of attenuation of potentially distracting perceptual input, and second, by means of a coupling of eye behavior to internally generated information and processes. We will briefly review evidence in support of these two hypothesized mechanisms. Attenuation of visual input ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Due to limited information processing capacity, processing external and internal information simultaneously impairs performance in one domain or both (Chun et al., 2011; Dixon et al., 2014). On the one hand, external tasks are impaired by internal events, like daydreaming or trying to solve a puzzle whilst driving a car (He, Becic, Lee, & McCarley, 2011; Savage, Potter, & Tatler, 2013). On the other hand, tasks requiring internally focused cognition, like thinking about an amazing new idea and how to implement it, can get interrupted by more or less irrelevant external events, such as an e-mail popping up on one’s desktop. External (perceptual input) and internal (self-generated) information are competing for limited processing resources (Chun et al., 2011). Therefore, being able to focus on one’s internal stream of thought and effectively shielding it from distracting external information has a high influence on the resulting performance. Active attenuation or shut out of external information can occur at different levels of visual perception. At the neural level, research using EEG alpha frequency activity and functional MRI suggests an active suppression of visual input when participants have to solve a divergent thinking task with high internal processing demands. For example, EEG alpha activity is assumed to reflect a top-down mechanism for suppression of external visual information processing (Benedek, Bergner, Könen, Fink, & Neubauer, 2011; Benedek, Schickel, Jauk, Fink, & Neubauer, 2014; Fink & Benedek, 2013). Moreover, the right inferior parietal cortex was implicated in the down-modulation of visual information processing during high internal attention demands (Benedek et al., 2016; for an overview, see Benedek, 2017). At the level of eye behavior, the most straightforward mechanism would be to simply look away from distracting stimuli (e.g. avoid eye contact: gaze aversion) or close the eyes. Indeed, gaze aversion and eye closure were frequently observed during demanding internally focused cognition and are thought to reduce cognitive load and thereby free up cognitive resources (Doherty-Sneddon & Phelps, 2005; Mastroberardino & Vredeveldt, 2014; Vredeveldt, Hitch, & Baddeley, 2011). The frequency of gaze aversion depends on cognitive load. The more resources an internally focused task demands, the more resources need to be recruited/withdrawn elsewhere, e.g. via gaze aversion or eye closure, to prevent performance impairments (Doherty-Sneddon & Phelps, 2005). Very short eye closures – blinks – may also reduce the processing of visual input and enhance internally focused cognition. Solving problems through spontaneous insight was associated with higher blink rates as compared to analytical problem solving (Salvi, Bricolo, Franconeri, Kounios, & Beeman, 2015). The relationship between mind wandering and blink rate is currently not clear (Mooneyham & Schooler, 2016; Smilek, Carriere, & Cheyne, 2010; Uzzaman & Joordens, 2011), but demanding internally focused cognition tasks like idea generation and insight problem solving are typically accompanied by higher blink rates (Akbari Chermahini & Hommel, 2012; Salvi et al., 2015; Ueda, Tominaga, Kajimura, & Nomura, 2015), supporting the hypothesis of an active decoupling strategy. Besides actual interruption of the visual input via blinks or gaze aversion, the processing of external visual information can also be attenuated. One potential oculomotor mechanism for attenuation would be disaccommodation. To focus on an object in three-dimensional space the eyes need to be aligned (eye vergence) and the lens adjusted to the object’s distance. These changes are associated with changes in pupil diameter. The three processes – eye vergence, lens constriction and pupil size changes – are coupled in the so-called convergence reflex or near response triad (Delank & Gehlen, 2006; McLin & Schor, 1988; Myers & Stark, 1990). Defocusing an object results in blurring and double vision, impairing further processing of visual information. Therefore, eye vergence is an important variable in research on covert attention - when attention is shifted to another location without moving the eyes (Solé Puig, Pérez Zapata, Aznar-Casanova, & Supèr, 2013) - and dyslexia (e.g. Kapoula et al., 2007), and it could also play a role in goal-directed internally focused cognition as suggested by the phenomenon of “staring into space”. The phenomenon of “staring into space” refers to the peculiar gaze people have when they are deeply absorbed in thought. “Staring into space” seems to be characterized by a strong stare (lack of eye movements) and the eyes seem to focus at a distance far away. In a similar way, attenuation of visual perception could also be achieved by means of reduced microsaccade activity. When fixating static stimuli, neuronal adaptation leads to perceptual fading within seconds. To counteract perceptual fading, our eyes perform fixational eye movements (microsaccades, drift and tremor), of which microsaccades are considered the most important ones (Martinez-Conde, Macknik, Troncoso, & Dyar, 2006; Martinez-Conde, Otero-Millan, & Macknik, 2013; McCamy et al., 2012). Microsaccades cannot be generated voluntarily, but they can be suppressed voluntarily for a few seconds (Bridgeman & Palca, 1980; Winterson & Collewun, 1976). As external visual information is not needed during periods of internally focused cognition, there is no need to counteract fading through microsaccades. Fading may even facilitate internally focused cognition by reducing distracting visual input. So far, the role of eye vergence and microsaccades for internally focused cognition has not been addressed by research. Coupling to internally generated information ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ When we remember a special occasion or think about the upcoming holiday, we often seem to actually look at a mental picture in our mind’s eye. This notion has been supported by research on mental imagery and memory retrieval, showing that eye movements are commonly coupled to cognitive processes during internally focused cognition (Ferreira, Apel, & Henderson, 2008; Johansson, Holsanova, Dewhurst, & Holmqvist, 2012; Johansson, Holsanova, & Holmqvist, 2005). When retrieving information about objects from memory, eye behavior patterns reflect those made while actually looking at the scene (Johansson et al., 2005). Even when eye movements during encoding are prohibited through continuous fixation of a certain point, eye movement patterns during retrieval reflect spatial characteristics of the remembered material, as one was looking at it. Those eye movements during retrieval seem to play an important role for retrieval itself, as prohibiting them impairs retrieval (Johansson et al., 2012). Furthermore, irrelevant eye movements impair spatial imagery (e.g. navigating through a matrix) but not visual mental imagery (e.g. imagining features of objects, for instance whether the letter A contains a curve or not), suggesting that eye movements are more crucial for spatial than depictive aspects of content (de Vito, Buonocore, Bonnefon, & Della Sala, 2014). It is even possible to determine whether a visual or a linguistic task was performed using eye movement patterns although the presented stimuli were the same pictures (Coco & Keller, 2014). Also searching for and focusing on information in long-term memory is reflected in different eye movement patterns. These differences have a counterpart in viewing external visual information: there are more eye movements during search and only a few when focusing on an object (Ehrlichman & Micic, 2012). Cognitive processes do not only affect fixations and saccades, but also pupil diameter. Pupil diameter adjusts to imagined luminance changes (Laeng & Sulutvedt, 2014), and even responds to other properties of the imagined content such as valence, cognitive load or target detection (Palinko, Kun, Shyrokov, & Heeman, 2010; Piquado, Isaacowitz, & Wingfield, 2010; Privitera, Renninger, Carney, Klein, & Aguilar, 2010). Further evidence comes from studies on mind wandering, showing that moving attention away from external events results in spontaneous, more variable pupil activity (Franklin et al., 2013; Smallwood et al., 2011). Moreover, mind wandering episodes while reading reflected deviations from typical reading behavior including longer fixation duration and reduced within-word regressions (Reichle, Reineberg, & Schooler, 2010; Uzzaman & Joordens, 2011). The present study ~~~~~~~~~~~~~~~~~ The presented literature suggests that internally focused cognition is accompanied by characteristic eye behavior changes supporting the decoupling from external information and coupling to internally generated information. Much of this evidence comes from studies examining differences between goal-directed externally focused cognition (e.g., reading) and spontaneous, internally focused cognition (i.e., mind wandering). Additionally, previous studies often focused on subsets of specific oculometric parameters (e.g., blinks or pupil diameter) and their role for either decoupling or coupling of eye behavior to internal processes. Therefore, the interplay of various oculometric parameters and their integrative functional role for perceptual decoupling and coupling mechanisms during internally focused cognition still need to be investigated. The present study aims to extend available research in at least two important ways: first, we aim to identify specific eye behavior patterns associated with deliberate internally focused cognition by contrasting two goal-directed thinking tasks, one with an external focus (letter reading) and one with an internal focus of attention (idea generation). Second, we consider a large range of oculometric parameters including less common measures like eye vergence and microsaccade rate that appear particularly relevant to examine potential mechanisms of perceptual decoupling. The broad range of oculometric parameters allows us to get an integrated picture of eye behavior during internally focused cognition and the interplay of both mechanisms: attenuation of visual input and coupling to internal processes. Finally, we explore the robustness of all effects by means of an experimental variation of luminance across trials. The attenuation of visual input hypothesis predicts that goal- directed internally focused cognition should be associated with higher blink rates and blink durations to reduce visual input (Salvi et al., 2015). Additionally, we explore the role of eye vergence and microsaccade activity as a more nuanced mechanism to attenuate visual input. Internally focused cognition may be related to a smaller angle of eye vergence (indicating a focus behind the screen) and to fewer microsaccades, counteracting perceptual fading. Predictions regarding the coupling to internally generated information hypothesis are less straightforward. Our external task requires only little eye movements for reading single letters in the center of the screen. Therefore, we expect that searching for an idea – similar to searching for an object in a visual environment (Ehrlichman & Micic, 2012) – is reflected in a more active eye behavior. This eye behavior could manifest itself in more and shorter fixations, more saccades and saccades with greater amplitude (Ehrlichman & Micic, 2012). Pupil diameter (Laeng & Sulutvedt, 2014) and eye vergence (McLin & Schor, 1988) also respond to imagined differences in luminance and distance, respectively. Imagined objects and their uses during the idea generation may vary in luminance and distance within one trial. Therefore, we expect more within-trial variation of pupil diameter (Smallwood et al., 2011) and eye vergence.","Forty-eight young adults (23 ± 4 years old, 32 females, 15 males, 1 non-binary gender identity), mostly university students, participated in the experiment for course credit and/or the possibility to win a skiing holiday. All participants had normal or corrected- to-normal (soft contact lenses) vision, reported no strabismus or other medical condition affecting vision. All participants gave written informed consent. Two additional participants had to be excluded from analyses. One had excessive missing data (>50%) due to partial occlusion of the pupil by the eye lids, the other had slight strabismus distorting gaze data. The study was approved by the local ethics committee. Stimuli and apparatus As stimuli, a stream of letters was presented in the center of the screen. All letters were lowercase, and presented in black Arial font of size 15 pt resulting in a character width of 0.32° visual angle (ca. 2.5 mm or 8.5 pixels). Umlauts were written-out (e.g. “ä” to “a – e”) and words were separated by a dash (“-”) instead of a space so the screen was never empty (see Fig. 1A). Letters were presented for 500 ms each, directly followed by the next letter. To test possible luminance effects on eye parameters, background luminance was brighter (RGB color code: 204,204,204) in one half of the trials and darker (102,102,102) in the other half. Participants were placed in a sound attenuated room with the lights turned on and sat at a distance of 50 cm from the screen. Their heads were stabilized using chin rest and forehead rest of the EyeLink Tower Mount (SR Research, Ontario, Canada). Stimuli were presented on a 19-in LG flatroon L1920P monitor run at 60 Hz and a 1240 × 1024 pixels resolution, subtending 29.4 pixels per degree visual angle. Binocular eye data were recorded using an EyeLink 1000 Plus Tower Mount eye tracker (SR Research, Ontario, Canada) with a temporal resolution of 500 Hz. For stimulus presentation and response recording, the EyeLink Experiment Builder software (SR Research, Ontario, Canada) was used. For calibration, validation, drift correction and computation of the eye movement parameters (blinks, fixations, saccades), we used the manufacturer’s software (SR Research, Ontario, Canada). Online velocity threshold for saccade detection was set to 35°/s and acceleration threshold to 9500°/s2. There was a 9-point calibration procedure before each block and a drift correction before each trial. Spatial resolution was typically better than 0.30°. Participants’ answers were recorded with a microphone to monitor and record task performance. Procedure and task ~~~~~~~~~~~~~~~~~~ Participants performed eight experimental trials of a reading task and eight trials of an idea generation task (see Sections 2.3.1 and 2.3.2), preceded by one practice trial of each task. The experimental trials were organized in two blocks with a break in between to minimize effects of visual fatigue. Each block comprised four consecutive reading trials and four consecutive idea generation trials (half with dark and half with bright background, respectively). Half of the participants started with the reading task, the other half with the idea generation task (see Fig. 1B). After each block, participants filled out a short questionnaire regarding aspects of task performance, including questions about their concentration, the perceived demands of tasks and the amount of distraction by the letter stream. Between blocks participants also filled out personality questionnaires (part of another study) in the antechamber and were instructed to take as much time for the break as they needed to rest their eyes (at least 10 min). The experiment took about 45 min in total. Externally focused cognition: Reading task A reading task was used as externally focused cognition task. Participants were required to read a German text of 120-character length. To reduce saccades, the text was presented as a stream of consecutive letters, each presented for 500 ms at the middle of the screen, resulting in a presentation duration of 60 s for the entire text (see Fig. 1A). After text reading, two comprehension questions on the content of the message were presented on the screen to test if participants had focused on the letters throughout the task (e.g., text: “Albert and I are going to a concert on Tuesday. I asked Franz if he wants to join us, but he visits his grandmother that day” Questions: “What day are we going to the concert?”, “Who is Franz visiting?”). Short mind-wandering episodes during this task would lead to problems in understanding the text and in answering the questions. In a third question, participants had to indicate the direction of attentional focus during the task, from 0 = “totally absorbed in thought” to 5 = “totally focused on external events”. Participants were informed that if they were focusing on the letters of the text during the whole trial they should indicate 5. If they had trouble to concentrate on the letters and were mind wandering during the whole trial, they should indicate 0. Participants answered the questions aloud. At the end of each trial participants pressed the spacebar to continue with the next trial. Internally focused cognition: Idea generation task As an internally focused cognition task, we used the alternate uses task (Guilford, 1967), a popular and well-established idea generation task commonly employed in research on divergent thinking and creativity (Kaufman, Plucker, & Baer, 2008). The alternate uses task asks to generate creative uses for a common household object (e.g., paper cup). This task relies on imagination based on an initially presented stimulus word and was found to be independent of further sensory processing (Benedek et al., 2014). The procedure of the idea generation task was essentially the same as in the reading task (see Fig. 1A): at the beginning of each trial participants received the instruction (e.g., “find as many and as creative alternative uses for this object: credit card”). Then they had 60 s to think of creative uses but without verbalizing them. During this period, the same letters as in the reading task were presented one character at a time, but in reverse order to make it unreadable. At the end of each task, participants were prompted to tell their two best ideas. Finally, participants had to indicate again their attentional focus during the task. They were informed that if they were focusing on idea generation during the whole trial, they should indicate 0 = “totally absorbed in thought”. If they were constantly distracted by the letters on the screen or any other external events and could not focus on their ideas, they should indicate 5 = “totally focused on external events”.","Creativity of the generated ideas was rated by four experienced raters (overall inter- rater reliability of 0.67) on a four-point scale ranging from 0 = “not creative” to 3 = “very creative” (Diedrich, Benedek, Jauk, & Neubauer, 2015; Silvia et al., 2008). Answers to the comprehension questions after each reading task trial were rated as correct or incorrect and percentage of correct answers was analyzed. Regarding eye tracking data, fixation durations and counts of fixations, blinks and saccades and saccade amplitudes per trial were calculated offline using Data Viewer software (SR Research, Ontario, Canada). Based on the default settings of the Data Viewer (SR Research, Ontario, Canada) saccades were defined as eye movements that exceeded 30°/sec velocity, 8000°/sec2 acceleration and/or 0.15° motion. Blinks were defined as a period with the pupil data missing for three or more samples in a sequence (= at least 6 ms) and fixations were defined as any period that is not a blink or a saccade. Blinks as well as additional 200 ms periods before and after each blink were removed from gaze position and pupil diameter data to eliminate parts where the pupil was partially occluded (McCamy et al., 2012). Only data for which the eye tracker had recorded both eyes were analyzed. Further data analyses were performed using R (www.r-project.org). For calculation of pupil diameter and eye vergence, eye tracking data were down-sampled from 500 Hz to 50 Hz by averaging across 10 data points (20 ms). Pupil diameter data were transformed from arbitrary units (measured by the eye tracking software) to z-values to make them comparable across subjects. For the mean pupil distance a length of 60 mm was used for all participants. For microsaccade calculation, original 500 Hz gaze position data were used. Microsaccades (count and amplitude) were determined using the Microsaccade Toolbox for R (Engbert, Sinn, Mergenthaler, & Trukenbrod, 2015) with λ = 4 and a minimum microsaccade duration of 6 ms. Microsaccades were defined as saccades with an amplitude smaller than 1.0° (McCamy et al., 2012). Only binocular microsaccades with a minimum overlap of one data sample were considered. Microsaccade amplitude was defined as the mean across both eyes. For all eye parameters, means or within-trial variability was computed for each trial. Trials with more than 50% missing data and trials with values beyond three standard deviations from the individual’s mean were discarded (percent trials discarded: M = 3.78, SD = 5.58, Max = 25.0). Remaining trials were averaged for the reading and the idea generation trials and for the bright and darker background luminance, respectively. For raw and processed data as well as additional analysis see Walcher, Körner, and Benedek (2017). Task performance ~~~~~~~~~~~~~~~~ On average, participants were able to correctly answer 90.36% of the comprehension questions in the reading task. In the idea generation task, participants reported two ideas in 98.31% of trials. All trials were included in further analysis. Participants reported being more focused on external events during the reading task than during the idea generation task (reading: M = 3.20, SD = 0.84, idea generation: M = 2.27, SD = 1.02, t(47) = 5.23, p < .001, r = .61). There were no task differences between ratings on how well they could concentrate (reading: M = 3.03, SD = 1.20, idea generation: M = 3.01, SD = 1.21, t(47) = 0.10, p = .925, r = .01). The letters were perceived as relatively little distracting during the idea generation task (M = 1.72, SD = 1.39, on a scale from 0 = not distracting at all to 5 = totally distracting). Participants reported that they perceived the idea generation task as more demanding than the reading task (reading: M = 1.91, SD = 1.24, idea generation: M = 2.58, SD = 1.13, t(47) = 3.95, p < .001, r = .50). Eye tracking data ~~~~~~~~~~~~~~~~~ A total of 12 eye movement variables were analyzed, which are not necessarily independent from each other (e.g. blink count correlated with fixation count up to rS = .58, p < .001 and with saccade count up to rS = .58, p < .001). Most of the eye movement variables violated the assumption of a normal distribution on the sample level (see Table 1, asterisk indicates a significant Shapiro-Wilk test). Medians for all tasks and conditions are reported in Table 1. Univariate ANOVAs indicated no interaction effects of task type (idea generation, reading) and background luminance on eye parameters (bright, dark; all ps > .05). Hence, effects of task type (idea generation versus reading) and background luminance were analyzed separately using non-parametric Wilcoxon rank-sum tests (for calculation of effect size r see Field, 2013). Effects of reading vs. idea generation Effects of idea generation versus reading on eye parameters are visualized with effect size r in Fig. 2. In the idea generation task participants produced more blinks (Z = 5.97, p < .001, r = .61), longer blink duration (Z = 4.97, p < .001, r = .51) and a lower microsaccade count (Z = 5.70, p < .001, r = .58) than in the reading task. But we observed no task difference in the angle of eye vergence (AoEV; Z = 0.14, p = .890, r = .01) and within-trial variability in the AoEV (Z = 1.87, p = .060, r = .19). Pupil diameter was larger during idea generation compared to reading (Z = 5.20, p < .001, r = .53) but within-trial variability of pupil diameter did not differ between task type (Z = 0.01, p = .999, r < .01). More and shorter fixations were made during the idea generation task compared to the reading task (fixation count: Z = 5.82, p < .001, r = .59; fixation duration: Z = 5.77, p < .001, r = .59). Closely associated with the number of fixations, more saccades (Z = 5.82, p < .001, r = .59) and saccades with larger amplitudes (Z = 5.40, p < .001, r = .55) were made during the idea generation compared to the reading task. Like saccade amplitude, microsaccade amplitude was higher during idea generation compared to reading (Z = 3.51, p < .001, r = .36). In the idea generation task, originality of the two ideas was not correlated with any oculometric parameter (rS < .23, p > .12). Background luminance Effect sizes for the influence of background luminance are visualized in Fig. 3. Pupil diameter was smaller with bright background (Z = 6.03, p < .001, r = .62). Angle of eye vergence was larger with bright background compared to the dark background (Z = 3.87, p < .001, r = .41). Both pupil diameter and angle of eye vergence showed less within-trial variability with bright background than with the dark background (pupil: Z = 4.48, p < .001, r = .46, angle of eye vergence: Z = 4.47, p < .001, r = .46). Fixations were longer with bright background than with dark background (Z = 2.84, p = .004, r = .29). Background luminance had no significant effect on fixation count, blink count, blink duration, saccade count, saccade amplitude, microsaccade count and microsaccade amplitude (Z < 1.93, p > .05, r < .28).","The aim of the present study was to examine whether and how characteristics of internally focused cognition are reflected in oculometric parameters. In the following, we discuss our findings with respect to the hypotheses of attenuation of visual input and of coupling of eye behavior to cognitive processes, respectively. Attenuation of visual input ~~~~~~~~~~~~~~~~~~~~~~~~~~~ In support of an active attenuation of visual input through eye behavior during internally focused cognition, participants blinked more often (21 versus 8 blinks per minute) and longer (120 versus 100 ms), and greatly reduced the number of microsaccades (15 versus 4 per minute) during the internal compared to the external task. No effect of attentional direction was found for the mean angle of eye vergence. More and longer blinks induce a longer disruption of the visual input, which may serve as a basic mechanism to reduce the processing of irrelevant external visual information. Salvi et al. (2015) also reported longer blink duration right before insight solutions as compared to analytic solutions. Findings on blink rate have been less consistent for spontaneous internally focused cognition. For example, during mind wandering episodes decreased blink rate was observed only in comparison to a low load task like breath counting (Grandchamp, Braboszcz, & Delorme, 2014) but not in contrast to a higher load task like reading (Uzzaman & Joordens, 2011). Goal-directed internally focused cognition (as in our task of timed idea generation) may require more cognitive resources and therefore a stronger shielding of ongoing internal processes from intruding irrelevant external input than mind wandering. In further support of attenuation of visual input through eye behavior, participants performed fewer microsaccades during idea generation than during reading (see also Benedek, Stoiser, Walcher, & Körner, 2017). Given the close linkage of microsaccade suppression and perceptual fading (Martinez-Conde et al., 2006), reduced microsaccade activity may favor fading of visual perception and thereby serve as another mechanism to gate out irrelevant external information during internally focused cognition. Given the higher average frequency of microsaccades compared to blinks, microsaccades could be seen as a more time-sensitive indicator of internally focused cognition. Contrary to our exploratory expectations, we found no effect of attention direction on eye vergence. This means that participants mean distance of focus was similar in both tasks. Overall, the eye behavior pattern of heightened blink rates and reduced microsaccade rates favors the hypotheses of an active attenuation of visual input. In the present study, defocusing was not part of that eye behavior pattern (but see Benedek et al., 2017). Further studies should investigate whether this eye behavior pattern emerges generally during internally focused cognition or if it is specific for certain study conditions. Distractor saliency could be such a critical study condition. People avert gaze more often if a distractor is salient (Doherty-Sneddon & Phelps, 2005). In the present study, external events (appearance of a single letter every 500 ms) during the internal task were perceived as relatively little distracting (1.72 on a scale from 0, not at all distracting, to 5, very distracting). Some eye behavior patterns related to attenuation of visual input, like eye vergence thus could only become necessary under conditions of higher distractor saliency. Coupling to internally generated information ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Goal-directed internally (vs. externally) focused cognition was associated with larger average pupil diameter. This is in line with recent research on mind wandering (Franklin et al., 2013; Smallwood et al., 2011). But pupil diameter is an intricate eye parameter because it is affected by different variables such as imagined luminance (Laeng & Sulutvedt, 2014) and cognitive load (Piquado et al., 2010), leading to ambiguous results, as demonstrated by Grandchamp et al. (2014) in the domain of mind wandering. While mind wandering seems to be associated with a larger pupil diameter (Franklin et al., 2013; Smallwood et al., 2011), so is increasing cognitive load (Palinko et al., 2010; Piquado et al., 2010). In the present study, idea generation (internally focused cognition) was perceived as more demanding than reading (externally focused cognition). In the internal task, participants had to generate ideas and keep the best two ideas in memory for the remaining time. We tried to account for the working memory load by including some working memory load in our external task. In the external task, participants had to read a letter stream. That also included working memory to a certain point, as participants had to retrieve the previous letters and append the present letter to be able to read the resulting message. Furthermore, they had to remember the whole sentence until the questions appeared. So in both, the external and internal task, participants had to modify material in mind and had to remember something during the trial. Therefore the main difference between the two tasks was the focus of cognition: during the external task participants had to continuously focus on external events (externally focused cognition) and during the internal task they had to focus on internally generated information, namely their ideas and could ignore external events (internally directed cognition). Yet, one cannot fully rule out some differential effects of working memory load on pupil diameter. This point underscores the need to include a wide range of oculometric parameters and to look at the overall pattern instead of only investigating a small subset thereof (e.g. only pupil diameter). Supporting our hypotheses that eye behavior is linked to the cognitive processes, idea generation and reading differed strongly in their eye behavior. Consistent with the external task’s presentation characteristics (only one letter at a time in the center of the screen), pupil diameter and angle of eye vergence were relatively stable. Moreover, only a few fixations and saccades and only saccades with small amplitudes were made allowing for effective reading of the letter stream. This eye movement pattern was in strong contrast to that of idea generation. Like searching in a real visual environment (Ehrlichman & Micic, 2012), searching for an idea in the mind’s eye evoked a more active eye movement pattern than in reading. Internally focused cognition induced a higher rate of fixations and saccades and a slight but not significant increase in variability of angle of eye vergence. Given that angle of eye vergence can be influenced by imagined changes in depth (McLin & Schor, 1988), those variations could reflect features of imagined ideas during the generation task. Eye behavior patterns strongly depend on the cognitive processes active and therefore on task characteristics. This implies that the observed differences (e.g., higher fixation counts for internal cognition) may not be intrinsic to internally focused cognition, but can only be interpreted in the context of the different cognitive processes involved in internal and external cognition tasks. It hence is not clear whether task differences in these parameters would also have been observed when tasks had been more similar (e.g., idea generation for a visible object as compared to idea generation for an imagined object; but see Benedek et al., 2017). As a consequence, coupling of eye behavior to cognitive processes may only serve as a reliable indicator of internally versus externally focused cognition when cognitive processes are well understood and differentiable. Effects of background luminance ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Background luminance was varied within tasks in order to test the robustness of attention effects and the sensitivity of parameters to luminance. We observed no interaction between attention direction and luminance supporting the idea that effects of attention direction are robust across different conditions of luminance. As expected, pupil diameter was smaller for bright compared to dark background, and we observed further effects on the mean angle of eye vergence and within-trial variability of pupil diameter and eye vergence. The latter findings may in part be due to eye tracking characteristics itself. Recent studies found a systematic distortion of gaze measurements due to pupil dynamics (Drewes, Zhu, Hu, & Hu, 2014; Eberhardt & Huckauf, 2016). Implications and future directions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ An objective discrimination between internally and externally focused cognition has important implications for cognitive research that often relies on self-report for assessing the focus of attention. Further investigation of the conditions under which attenuation of visual input occurs might help to refine theories of distractor inhibition during internal cognition and of switches between external and internal cognition. For example, load theory (for an overview see Murphy, Groeger, & Greene, 2016) postulates different effects of cognitive and perceptual load on distractor interference during externally focused cognition. Both cognitive and perceptual load also have a strong impact on distractibility during internally focused cognition. Adding to that, Buetti and Lleras (2016) showed that increased cognitive engagement reduces sensitivity to visual events. In a next step, one should investigate if the attenuation of visual input under higher cognitive load and cognitive engagement could be related to changes in the eye behavior pattern (blinks, microsaccade suppression) as found in the present study. Regarding the properties of the used tasks, one can also interpret the results in the framework of the predictive and reactive control system theory (PARCS, see Tops, Boksem, Quirin, IJzerman, & Koole, 2014). During the reading task, participants had to integrate new incoming visual information (letters) continuously, which would involve mostly the reactive control system or switches between the reactive and predictive control systems of the PARCS. Compared to that, the idea generation task requires a focus on internally generated information and a decoupling from incoming signals, therefore mainly the predictive system should be involved. Attenuation of visual input could indicate the involvement of the predictive control system. Establishing an eye behavior pattern that can indicate this decoupling from the external world reliably not only allows for a more accurate categorization of internally and externally focused cognition between tasks but also within tasks. The latter would be valuable for the investigation of the brain mechanisms associated with switches between the reactive and predictive control systems. Further, one should investigate individual differences in attenuation and coupling related to eye behavior in the face of distracting visual input. This could bring us closer to the answer why some people are so good at shielding their internal train of thought while others have serious trouble doing so, especially in a world where we are confronted with an immense amount of external stimulation (e.g. adds on webpages, new emails popping up on the desktop).","The present study established that goal-directed internally and externally focused cognition can be reliably discriminated based on eye parameters. Increasing blink rates and durations as well as reduced microsaccade activity may facilitate perceptual attenuation via shut-out and fading of visual information. This eye behavior pattern thus may represent an important mechanism to shield ongoing internal trains of thought from irrelevant, potentially distracting external stimulation, and future research should attempt to test this notion experimentally. Further differences in eye parameters between internally and externally focused cognition may be interpreted in terms of a coupling of differing cognitive demands between idea generation and reading. We conclude that the eyes do not idle during cognitive activities that are independent from sensory information but rather seem to play an active supporting role by shielding internal thought processes as well as reflecting them."],["This article is concerned with what it means to think of online spaces as emotionally safe or safer. It does this by looking at the sharing of emotional distress online and the role of organisations in identifying and proactively engaging with such distress. This latter type of digital engagement is analytically interesting and rendered increasingly feasible by algorithmic developments, but its implications are relatively unexplored. Such interventions tend to be understood dualistically: as a form of supportive digital outreach or as emotional surveillance. Through an analysis of blog data about a Twitter-based suicide prevention app, this article attempts to understand the tensions and also the potential points of connection between these two meanings of ‘looking out for each other’ online. From an avowedly sociological, relational and emotional perspective, it tries to offer a more nuanced account of what it might mean to share emotional distress ‘safely’ online. --------------------------------------------------------------------------------","While it is not always possible to clearly distinguish between the two (Tucker and Goodings, 2017), there is increasing recognition of the significance of both informal emotional support on social media (Eggertson, 2015) and digital outreach, a proactive form of support facilitated by organisations such as voluntary or government bodies. The challenges of this latter type of ‘digital professionalism’ (Ellaway et al., 2015) are debated but, as Facebook's recent announcement about the use of algorithms to identify suicidal users makes clear (Kelion, 2017), it is increasingly presumed that social networking platforms have the potential to ‘promote positive change’ in relation to mental health (Inkster et al., 2016) and to offer digital ‘safe spaces’. At the same time, however, there are concerns that such outreach can become a form of ‘health surveillance’ (French and Smith, 2013; Lupton, 2012). Regardless of whether or not it is framed as emotional surveillance, the complexities of how selves are presented, on and offline (Murthy, 2012), make the interpretation of online content by those engaged in online support difficult. This is particularly evident in relation to suicide (Mok et al., 2016). Reflecting wider dichotomous views of the Internet, it's role in relation to suicide is often perceived as Manichean: a struggle between dark/death and light/life (Robert et al., 2015). Social media is increasingly part of this context (Christensen, 2014). There is evidence, on the one hand, that informal online suicide communities can be valued by participants for their mutuality (Baker and Fortune, 2008); on the other, of suicidal postings not being acknowledged or of forums maintaining or even amplifying suicidal feelings (Mok et al., 2016). Similar concerns about how online sharing may reinforce or exacerbate risky behaviours have also been noted in relation to other practices (Cantó- Milà and Seebach, 2011). Reflecting the ambiguity of emotional expression, the task of distinguishing between harmful and helpful content in these contexts is complex: apparently ‘dangerous’ text can act as a deterrent, be life-affirming or empowering, and be read differently over time (Mars et al., 2015). Suicide prevention apps (resources designed to support a person in distress) are increasingly a part of this complex digital landscape. They range from social media interventions2 to smart phone self-help apps3 (Aguirre et al., 2013). Facebook has been the platform most closely associated with these developments4 but, in 2014, Samaritans, a charity that provides emotional support to people experiencing distress, launched a Twitter-based app, Radar,5 under the tagline, ‘turn your social net into a safety net'6. This allowed registered users to be alerted when a Twitter account they followed included messages that might suggest depressed or suicidal thoughts. Although media and other responses to the app were initially positive, it proved increasingly controversial and was withdrawn nine days after its launch. To understand this, we need to look beyond digital outreach to what it means to share emotions in online spaces and how such spaces come to be thought of as safe or safer.","Notions of ‘space’, including ‘safe spaces’, are key to the sharing of emotion online. An understanding of space as not innate but constituted and reconstituted through our actions and interpretations is longstanding (Bondi, 2005; Cronin, 2014) and has been applied to online settings (Marino, 2015). However, despite initial framings that digital space might involve the collapsing of time and the overcoming of materialities, the relations through which all space is produced remain stubbornly material (Massey, 2005). The ‘new’ material turn has placed matter at the heart of analysis of space, including online, reminding us, as Lehdonvirta (2010: 885) put it, that the online world is less an ‘open frontier’ than a ‘built-up’ area. This materiality is relevant not just because, as Fayard (2012) notes, virtual space involves ‘a lot of stuff’, hard and software, but because all users of such spaces are embodied and embedded in particular places (Hines, 2015) and because online spaces have socio-material consequences. In other words, space – online or otherwise – is constituted through, but also shapes, social relations (Lefebvre, 1991): it can keep people in (their) place or help them move. Not surprisingly, then, online space is often made sense of through material metaphors. This is true, too, of the sharing of emotion in online space, though the metaphors used in this context speak to the more ethereal aspects of offline space: atmosphere (Tucker and Goodings, 2017), intimacy (Michaelsen, 2017) or ambiance (Thompson, 2008). The relational emphasis on understanding space, noted in the above discussion of materialities, is often framed in an online context in terms of sharing. Despite being under-conceptualised (Kennedy, 2015), sharing is key to analysing online space because the Internet is both a space where sharing happens and is made up of spaces constituted through such sharing. Because the online realm is made up of relational spaces, it is also emotional. In addressing the sharing of emotion specifically, I am concerned with both the meanings identified above: that is, to give and receive emotion, and to do so jointly within the same space – but also with a third, less researched dimension, that is, to have the same understanding of an emotion (or why that emotion was expressed) as another. Research on the role of technological affordances, social relations and norms on sharing across different social media suggest a complex interplay between the three (Bucholtz, 2013). To understand this interplay involves going beyond what it means technically to be a Twitter user/follower (Bruns and Moe, 2014) to examine the meaning Twitter users give to the sharing of emotion and how this sharing constitutes, as much as is mediated by, space. Constituting safe(r) spaces ~~~~~~~~~~~~~~~~~~~~~~~~~~~ When people talk about safe spaces, notions of the relational and emotional are present but not always foregrounded. Safe spaces are ‘imaginary construction [s]’ (Stengel, 2010:524) that involve complex boundary work in relation to the imagined ‘unsafe’ (Rosenfeld and Noterman, 2014). The impossibility of wholly or indefinitely achieving exclusion or inclusion ensures that the boundaries and the spaces created are always porous, contestable and shifting and for this reason it makes sense to refer to safer rather than safe spaces. Discourses about ‘safety’ in relation to mental health, race, class or sexuality (Haber, 2016), for instance, are sometimes framed around the exclusion of others and, at other times, through inclusion – for instance, the idea of a safe space for all. The idea of being ‘safe to’ express oneself, without repercussion and perhaps in contestation of dominant discourses, is core to the creation of a space as safe but cannot always be separated from the idea of being ‘safe from’. This speaks to a long history of the public sphere as a space of surveillance and exclusion but also, especially for marginalised groups, as a space of radical potential, a ‘haven’ away from the oppressions of the private sphere (Haber, 2016). Online spaces, because of their relative anonymity and ease of access, have been seen as potentially offering increased access to safe(r) spaces, though these same features can also increase the risk of feeling unsafe. Safe(r) spaces depend then not just on who is included or excluded and what is expressed but on how those who are present listen and respond. Reflecting a wider neglect of listening in the social sciences (Brownlie, 2014), much research on digital sharing still tends to focus on expression (Crawford, 2009). Yet listening to/reading emotional expression in public online spaces, including Twitter, involves considerable emotional work. The therapeutic remains a dominant narrative shaping understandings of the ‘right’ amount of sharing online as offline (Illouz, 2008): too much information (‘tmi’) is problematised as leading to an over-exposed self. Conversely, undersharing raises concerns within social networks about those who ‘go silent’. In both cases, listening is implicit but critical. Some have suggested there are ‘rules for sharing’ experiences of emotional pain and suffering online (Sandaunet, 2008), and that these are space- and subject-specific (Hess, 2015). But the fluidity and complexity of sharing (through both expressing and listening) in moderated and unmoderated spaces (Skeggs and Yuill, 2015) suggests that ‘rules’ may be too rigid a description for what shapes such sharing. Instead, emotional reflexivity, the way in which emotions are drawn on to negotiate social life, might offer greater analytic purchase (Burkitt, 2014). Radar's introduction into Twitter was a moment of contestation around the meaning of an online safe space: positioned by some as a form of digital outreach that could constitute a safe space but understood by others as breaching a pre- existing safe space. In what follows, I explore these framings through a focus on bloggers' reflexive accounts of what it means to be emotionally ‘safe(r)’ online, but I begin by outlining how these accounts were identified and analysed.","The following analysis draws on a sample of blogs about Radar identified from a search of all tweets that contained #Samaritans Radar or keywords ‘Samaritans’ and ‘Radar’, identified via the Twitter API in the two weeks following the launch of the app.7 This search produced 6540 tweets and the blogs drawn on here are the ones mentioned in the top 100 (ranked by number of retweets and favourites). The resulting 34 blogs were produced by a range of authors including academics, journalists, mental health writers and activists. Typically, they were not opinion pieces written by detached commentators but were from engaged Twitter users. The project focused on blogs accessed through Twitter rather than communication on Twitter itself. While analysis of tweets might allow for exploration of reaction to Radar on Twitter, the question of how Twitter users’ account for their sharing of emotion on this platform, in other words, their emotional reflexivity, is better accessed through their expansive and discursive blog writing. In other words, blogs offered Twitter users (and researchers) a degree of distance from their Twitter practice. There are good analytical reasons, however, for accessing this blog sample through Twitter. First, the Twitter dataset, produced as part of a larger study on emotional distress and digital outreach, offered the possibility of a comprehensive sampling frame of active Twitter users writing on this topic. Though there are limitations to the Twitter API that mean neither the Twitter nor the blog dataset can be treated as complete, it would not have been possible to sample these blogs systematically via a general internet search engine. Second, sampling through Twitter allowed identification of blogs and opinions that were being actively engaged with – or at least publicised – by Twitter users following the Radar debate. These blogs might reasonably, then, be regarded as a ‘long form’ version of the Twitter debate about Radar. Finally, analysis of the Twitter dataset suggested, not surprisingly perhaps, that themes relating to safety and the sharing of emotion emerged in both the Radar-related blogs and the larger dataset of tweets from which the blogs were identified. This offers some reassurance about the prevalence of these themes beyond the blog sample – again, this would not have been possible if the blogs had been identified through Google. As noted, however, my analytical concern here is not with quantifying discourses but rather with how those who supported and resisted the app framed their understanding, hence the decision to focus solely on blogs. As with other forms of documentary analysis, blog analysis can involve structural, thematic and narrative dimensions (Elliot et al., 2016). In particular, blogs can be read as performances of narrative identity and as constituting particular communities (Schuurman, 2014). While a focus on such community formation could inform further analysis, the focus here is on mapping ways of writing and thinking about being safe(r). To this end, blogs were coded thematically and were then re-read with a focus on the use of metaphors. How metaphors shape the way we think has long been of academic interest and now includes a focus on the online (Lakoff and Johnson, 2011; Markham, 2013). Ethically, no assumptions were made about the public nature of blogs unless the blogger's page was explicitly linked to an organisation, such as a newspaper (AOIR, 2012). In all other cases bloggers were contacted via email to advise them of the research project and to seek consent to quote from, and link to, their blog. Only where consent was given were blog data used in these ways.8 The analysis draws on a particular understanding of emotion. Much research on the sharing of emotions through social media has drawn on affect (Hillis et al., 2015). Understandings and definitions of affect vary widely: for some, it involves a visceral, bodily or ‘gut’ reaction beyond intent, though increasingly there is a move towards recognising its relational and discursive elements (Veletsianos and Stewart, 2016). Many of these framings of affect share a focus on its productivity, on what it does (Kennedy, 2015) and have, therefore, been useful in helping to think through how social media both engages and produces emotion (Hillis et al., 2015) In work focused on media affordances, however, how emotions are constituted by and constituting of particular relationships, as well as spaces, can get lost. To this end, the analysis which follows is informed by an understanding of our capacity for emotional reflexivity and of emotions as constituted, and made sense of, through relational processes (Brownlie, 2014; Burkitt, 2014).","At first glance, analysis of the blogs suggested that the response to Radar could indeed be understood through the polarised framework of emphasising benefits (harm reduction through outreach) or risks (those associated with surveillance). The app was introduced on the basis that it could potentially save lives and, initially, was positively received. In the two week period after its launch, however, the metaphors appearing in the blogs shifted towards the discourse of surveillance. The announcement of the petition for the app to be withdrawn, for instance, described Radar as having been launched ‘behind people's backs' on ‘unsuspecting Twitter users’.9 One blogger who writes about data protection10, referred to the app through the metaphors of ‘profiling’ (http://bit.ly/2A9W9uc) while another, who writes on mental health and disability, referred to ‘Orwellian’ practices (http://bit.ly/2Bcoo8y). The peculiarly human nature of the work the app was being programmed to do – identify emotional distress – meant it was metaphorically positioned by some bloggers as cyborg-like, a cross-over between a ‘suicide bot and concerned friend’ (Hess, 2015). To a degree this points to the particular ambiguity of the Radar app: introduced by an organisation and hence akin to traditional centralised surveillance practices, yet dependent on Twitter users signing up for the app to be reminded of what their followers have been up to, hence closer to peer to peer surveillance. The app's ambiguous position was further accentuated by the commercial context in which it was developed (Mason, 2014). Drawing on metaphors of experimentation, one mental health blogger suggested Twitter users were being positioned as ‘subjects for your tech’ (http://bit.ly/2Bcoo8y), reflecting an unease with such developments even when, as in the case of Radar, they were packaged as ‘big data for good’. On first reading, then, the conventional (individualised) polarities associated with digital outreach, of reducing harm or increasing the risk of surveillance, are evident in discussion of the app. As the debate unfolded it became more polarised, a process accentuated by the adversarial nature of the platform. Yet it is also clear that both ‘sides’ acknowledged the significance of finding support online and both expressed anxieties about making Twitter a safe space. Re-reading the blog data through an understanding of safe(r) spaces as materially, relationally and emotionally achieved helps make sense of polarised views of how digital caring happens or is imagined to happen. Discourses, often metaphorically expressed, run through all these dimensions and, in what follows, I consider online safe(r) space discretely through each. This, though, is a heuristic separation as they are in practice co-constituting.","Twitter, like all space, is constantly under construction, produced through everyday relationships, practices, emotions, materials and the discourses that revolve around all of these. In launching Radar and in subsequent statements, Samaritans sought to introduce the app as offering Twitter users a ‘second chance'11 to see tweets from someone they knew who might be struggling to cope. In other words, they sought to make the technical case that the tweets that triggered app alerts were already appearing in the Twitter feeds of the app subscribers.12 Bloggers' response to this justification, however, reflected a by now well-established critique of the dualistic understandings of public/private developed through work on the nature of visibility on social network sites (Marwick and boyd, 2014). In particular, some bloggers described how they constituted different spaces within ‘public’ Twitter through the strategic placing in tweets of the @ sign or a period in order to restrict or expand their potential audience (for an explanation of how this works in practice, see http://bit.ly/1fkfKbI). In this extract, for instance, a blogger who writes on digital privacy explains that such practices allow him and others to use the Twitter space in ways that are less public than those deployed by celebrities such as Stephen Fry: Not all Stephen Fry. Not all tweets equally visible – can make more intimate and more public by using @ and also by using ‘.’. (http://bit.ly/2BqLrNR) Reflecting a tendency, noted earlier, to perceive digital space through physical metaphors (Haber, 2016), Samaritans in their naming of the app and through their core metaphor of Radar as a ‘safety net’ sought to make sense of digital space through material offline practices and objects. Similarly, offline place metaphors were important for how bloggers attempted to understand Twitter as a digital safe space. Place has long been associated with understandings of safeguarding: ‘the basic character of dwelling is safeguarding’ (Heidegger, 1978:352 cited in Martin, 2017) and this is also the case with online dwellings. In the following extracts, bloggers are trying, in particular, to think through the potential for privacy on Twitter. One blogger, an academic, turns to offline surveillance practices (CCTV) to do so, while another turns to offline places (the pub, the dinner table) and communicative practices (whispering): my office window looks out on a public street – whatever people do there is public. There would still, though, be privacy issues if I installed a video camera in my window to tape what people did outside. (http://bit.ly/2BxoUOF) You can have an intimate, private conversation in a public place – whispering to a friend in a pub, for example […]. Chatting around the dinner table when you don't know all the guests – where would that fit in?’ (http://bit.ly/2BqLrNR) Crucially, though, these place metaphors are also a means of imagining who is listening (Litt and Hargittai, 2016). Like those who have written about online ‘places’ as having ‘residents’ and ‘clusters of friends and colleagues’ (White and LeCornu, 2011 cited in Veletsianos and Stewart, 2016), bloggers constitute Twitter, not just as a material and discursive space but also as a relational one.","From the outset, Radar was conceived by Samaritans as a relational tool; a means for all of us to look out for each other online. This is consistent with a shift towards the democratisation of helping: the idea that one does not need to be trained to offer help to others (Lee, 2014). Those who initiated Radar had in mind a diffuse but benign audience: in other words, the ‘attitude of the whole community’ (Mead, 1934) was imagined to be supportive. In an update following the launch, Samaritans suggested that the app could be aimed at those ‘who are more likely to use Twitter to keep in touch with friends and people they know’.13 One blogger, who writes about technology, also suggested that the app might work better if it was restricted to reciprocal arrangements with people ‘who the at- risk person is following back – so theoretically only friends’ (http://bit.ly/2A9YdSY). Different platforms offer different understandings of safe relations and this focus on friends is closer to the generalised ‘attitude’ of the Facebook rather than Twitter community. Facebook involves known networks, and indeed most research on emotional support has been focused on the platform for this reason (Burke and Develin, 2016). Yet there are powerful norms that make it difficult to seek emotional support on Facebook (Buehler, 2017) and, for some, its very interconnectedness makes it a less than safe space for sharing emotional distress. Affordances of other platforms may mean that sharing is experienced as potentially less stigmatising with feelings expressed, for instance, through images and through second (Instagram) or ‘throwaway’ (Reddit) accounts (Andalibi, 2017) in a way that is not about maximising visibility. Twitter followers are not necessarily friends in the sense of intimates; nor are followers who are friends, and are intimates, necessarily able or willing to take on the responsibility of ‘looking out’ for others. Indeed, there are significant risks in assuming friends and/or followers are in a position to reply or, as a blog below from a mental health service user suggests, that their response would be useful or welcomed. Indeed, they might in fact have a ‘chilling’ effect (Hess, 2015): It's not just having the information, it's being able to *do* something useful with it. If it just flags up that you need to try and connect with and support a person, that's one thing, but what if people who know very little about mental health sign up and wake up to a worrying email? Would they have enough info to call police/ambulance? SHOULD they – would the person welcome this? (http://bit.ly/2Bcoo8y) One blogger who writes on disability and mental health expressed concern that app users might be potentially deluged by ‘ghost patterns’ (Crawford, 2014): ‘if you use [Radar], you're going to get a hell of a lot of spam, unless you only follow a few close friends' (http://bit.ly/2zEvAOz). Such anxieties, however, might also be about whether apps can capture the relational nuances of friendships and about the authenticity of those friendships that depend on the app: would a real friend need technological help to become aware of a known other's emotional upset? If individuals need to be poked to remember to check on someone, they very much might not be the type of person someone [ …. ] wants to talk to. (http://bit.ly/2AchWiH) There are ways of displaying friendship, online and offline, that shore up authenticity; and, conversely, friendship norms can be breached through inappropriate sharing or use of technology (Bazarova, 2012): Friends are people who although I have only met them via Twitter care enough about me to check in with me before they go to bed and when they wake up, who states openly, “I am really concerned about you, how can I help?” They don't need an app to record what I have been saying while they were asleep or at work, because they simply message me as soon as the[y] wake/return asking how I am and telling me they are thinking of me. (http://bit.ly/2Bcoo8y) The clash of norms arising here is from the use of an algorithm14 in relation to the emotional support and specifically about the automation of friendship: the speeding up, delegation and mechanisation of what it means to be a friend. In the context of online dating, this is what Peyser and Eler (2016) refer to as the ‘tinderization’ of emotion. Research is beginning to emerge on the limitations of ‘conversational agents’ such as Siri in responding to distress (Miner et al., 2016) and for Wachter-Boettcher (2016) this is a reflection of the political economy of social media, ‘an industry willing to invest endless resources in chasing “delight” but not addressing pain’. These criticisms are with the friend-like sensitivity of the app but the bloggers' concerns are about what using such an app in the context of friendship says about the nature of that relationship. Other bloggers were focused less on friendships than on types of relationships seen as ‘unsafe’ because they are not supportive. These include digital bystanders; those who choose not to respond to tweets about distress. Illustrating the complexity of ethics of care in digitally mediated relations, these may be people who do not want to help but they may also include those who tactfully ‘disattend’ to messages believing they are not intended for them (Brake, 2014: 45). Either way, lack of acknowledgment of emotional distress could leave those who have tweeted about their distress with the impression, as one mental health blogger described it, ‘that nobody cares enough to respond’ (http://bit.ly/2zEvAOz). The public space of Twitter can also be inhabited by those who are the antithesis of ‘friends’ or even ‘digital bystanders’: those whom some bloggers referred to as ‘stalkers’ and ‘abusers’. As with safe spaces, safe relations are defined by their opposite. Drawing on metaphors which speak to vulnerability and fear of violation, some bloggers suggested that in the ‘wrong hands’, the app acts like an advertisement for our ‘empty homes’ or ‘lost children’. This speaks to the porousness of safe(r) spaces: the inability to separate out a safe space constituted by well-intentioned use of the app from other spaces, on and offline, that are, as the mental health blogger below notes, potentially less benign: What if your stalker was a follower? How would you feel knowing your every 3am mental health crisis tweet was being flagged to people who really don't have your best interests at heart, to put it mildly? (http://bit.ly/2zrS8hv) Radar had been introduced as a way of listening and being empathic. But not all listeners are benign and, as a result, some bloggers feared that being listened in on, rather than listened to, might lead some to withdraw from the informal support that exists from simply being on Twitter. For these Twitter users, apps such as Radar, are disruptive of a pre- existing ‘positive’ community formed in a ‘natural, human way’. In describing this sense of community, the metaphors used are system related – Twitter as ‘an online ecosystem’. Positioned as ‘ad hoc’, this relational space is understood by some as being built up through individuals’ sharing distress and thus can be thought of as an ‘authentic space’. It is a community then in the sense that it involves a fusion of weak and strong ties (Gruzd et al., 2011) and, crucially, has an affective component: sharing emotions constitutes a sense of ‘we ness’ or ‘digital togetherness’ (Marino, 2015). While this imagined space is finely calibrated, constituted by the honesty of what is said, significantly, there is not necessarily an expectation of direct intervention from those listening. This speak to a particular understanding of emotional expression (and listening) in public spaces which is at odds with the understanding that informed the development of the app, outlined in this extract from Samaritans: People often tell the world how they feel on social media and we believe the true benefit of talking through your problems is only achieved when someone who cares is listening.15 Awareness of the app, however, meant that some bloggers imagined Twitter as less safe and the app as silencing rather than amplifying. This was a view expressed by some of those who were already known to be using Twitter for emotional support, such as the mental health blogger below. Those not publicly using Twitter in this way may, of course, have very different views but they are less likely to be known about: It will cause me harm, making me even more self aware about how I present in a public space, and make it difficult for me to engage in relatively safe conversations online for fear that a warning might be triggered to an unknown follower. (http://bit.ly/2AchWiH) Bloggers discussed the possibility of switching Twitter accounts to ‘private’ to avoid the algorithm but this recalibration of the imagined Twitter space and its relationships, would make it, for some, more akin to Facebook's relational space(s) with its attendant expectations of distress being shared in particular (undemanding) ways (Buehler, 2017). In the final section, understandings about emotional expression embedded in the above accounts of safe(r) spaces are analysed more closely.","As Pedersen and Lupton (2016) have noted there is still surprisingly little focus on the role of emotions in the sizeable body of literature on digital sharing and surveillance. In the specific context of sharing emotional distress, Radar was premised on the belief that there is such a thing as a digital cry for help and that it can be identified by algorithm as such. Indeed the app, as noted, was positioned as a second chance to hear such online cries. Radar was intended to create a sense of safety by increasing awareness of those in difficulty and indirectly a sense of obligation to act caringly towards them. Metaphorically, as described by a journalist below, the app was intended to catch people on the precipice: The point of Radar, however, is to catch people on the edge of the cliff face who would be truly grateful of a helping hand to pull them back up, but for whatever reason, didn't feel they could say this outright.16 The assumption here is that the meaning of expressed emotion is shared. This is the third understanding of sharing emotion I introduced earlier: that two or more people share an understanding of why an emotion is expressed. In the case of emotional distress, however, those who express emotion and those who then hear or read it are not necessarily sharing the same understanding. Expression of emotional distress can be a cry for help but it can also have and achieve other ends. A cathartic model of self-expression, for instance, suggests that emotional expression is a form of release (Solomon, 2008). Other perspectives focus on emotional expression as self- actualisation or as a means of making sense of our emotional experiences by sharing them. Bennett (2017) has suggested that there is also a symbolic model for understanding emotional expression: that emotions, like poetry, are an end in themselves - a way of speaking to whom one is (Murthy, 2012). What is missing from some of these psychological and philosophical frames, however, is a strong enough sense of the relationships and contexts within which the sharing of emotions takes place. In their work on academics who share personal information online, for instance, Veletsianos and Stewart (2016) note this group tend to do so with clear intent and a strong awareness of social context. Some of the Radar-dataset bloggers, such as the mental health blogger below, made it clear that they would not consider it safe for them to express their distress on Twitter in a way that could be interpreted as a cry for help: No matter how suicidal I've been, I would never, ever tweet something like ‘I can't go on’ or ‘I'm so depressed’. [ …] I also don't want to cause a fuss, so being aware of specific trigger words that may cause a fuss would make me consciously not use them. Furthermore, I am private about my mental health […] I don't want my health problems to be noticeable. (http://bit.ly/2AchWiH) This practice of public tweeting with the intent of not being noticeable or ‘making a fuss’ returns us full circle to the question of what constitutes a safe ‘public’ online setting, and why people choose to express and listen in such spaces. For some there is safety in being one of many when disclosing pain. This can be thought of as quiet public disclosure, constituting and constituted by diffuseness and a dispersed sense of emotional connection. This is as potentially public as those who ‘life stream’ but is driven less by a desire for visibility per se than, as the mental health blogger below suggests, ‘mutual witnessing and display’ (Warner, 2002: 13, cited in Haber, 2016: 393). A way, in other words, to ‘welcome care’ from ‘known’ networks (Veletsianos and Stewart, 2016) It's a place I can express myself without worrying that I will upset relatives, where I can say the unsayable because other people understand. (http://bit.ly/2Bcoo8y) This saying of the unsayable is a way of feeling differently in digital spaces which are often regulated towards optimism (Pedersen and Lupton, 2016) or what Michaelsen (2017) framed as ‘a better future of emotional sameness’. Some bloggers suggested that it was the ‘processing’ work of the algorithm which made their messages (and hence their emotions), public and potentially unsafe, not the act of tweeting or expressing emotion in itself. For these bloggers, the interplay between the material, relational and emotional positions the algorithm as part of an online interaction order (Mackenzie, 2016). These concerns tune into anxieties about emotional surveillance and about whether or not we can control information about ourselves, including about our emotions (Brandimarte et al., 2013). This brings us to a key distinction between outreach and surveillance: the notion of consent. Becoming aware that one is part of an infrastructure of ‘noticing’ online that one has not directly consented to can lead to an unease with, and a reaction against, the reading of such information, not least because it may not speak to who people believe they ‘really’ are (Ball et al., 2016). Though others, of course, may position information gained through emotional surveillance as authentic exactly because it happens outside of platform users' full control or even knowledge (Horning, 2016).","It is in the nature of safe(r) spaces that they are contestable, open to both boundary maintenance and change. Digital safe(r) spaces can be even more porous and hence boundary skirmishes in relation to them are potentially even more fraught. This article, through exploring the reaction to Radar, has been concerned with what makes an online space feel emotionally safe(r). Those concerned with safeguarding or digital outreach and those concerned with the guarding of safe spaces against surveillance both write about Twitter as a material, emotional and relational space, but they constitute it differently (Lien and Law, 2011). Working with metaphors helps us to understand what and why things matters to people and is revealing of how the Radar app processed and also produced, a great deal of emotion: strong emotions linked to notions of territoriality, belonging and safety. So while there is value in thinking about the online in terms of what we do (Jurgenson et al., 2018) rather than as a place or space, all space, as argued earlier, is constituted and for many users of online platforms, doing and space are interconnected. To respond to the sharing of emotions, including cries for help, in a way that is disembedded from and does not acknowledge their spatial context creates upset, and this can be read in the metaphors of intrusion and violation that permeate the blog discussions. Focusing on the material, emotional and relational context of safe spaces, however, also reveals points of connection particularly in relation to valuing peer relationships and the benefits of sharing of emotional distress. Those concerned with both surveillance and outreach, however, also risk positioning emotional distress as there for the seeing/hearing: depending on one's point of view, ready to be caught in a ‘safety net’ or turned into ‘sensitive data’ by the app. The analysis presented here suggests emotions are not so easily read or categorised. For some bloggers, the sharing itself is what keeps the distress at bay; for others, there may well be no obvious emotional content to their sharing and yet they are distressed. Even if distress is correctly identified, however, there are still questions about how online proximates and strangers should be cared for. The analysis points to concerns about the automation of support online and also to the claim that some users may leave online public spaces such as Twitter if they believe these spaces to be ‘monitored’. For digital outreach to be enabling rather than undermining of existing peer support, therefore, several issues need to be considered by organisations. These include a need to be explicit about organisational assumptions about what constitutes support and how consent to be supported is sought and given online; being aware of the potential for outreach to be both silencing and amplifying and of the sensitivity of the tipping point between keeping a digital eye out and emotional surveillance; and recognising the specific risks of using algorithms in relation to the most human of activities - responding to emotional pain. While only implicitly present in this article, trust runs through all these considerations and is core to what constitutes feeling ‘safe(r). Drawing on the meaning of safe(r) spaces repositions the sharing of emotions as being about more than the commercial interests of social media platforms or the norms relating to Twitter, and highlights that people share distress or emotions because of their sense of trust in relationships they believe they have or desire to have. Any organisation engaging with digital outreach needs to understand how their actions can either shore up or breach such trust, and hence vitally affect people's sense of being safe. A dualistic focus on outreach or surveillance in relation to emotional distress online reduces possible responses to either direct intervention or silence – digital Samaritan or bystander - missing other nuanced practices, including what I have referred to here as quiet public disclosure of emotional pain. Moving beyond dichotomous accounts of online sharing of emotional distress creates more ambiguity but should be an important first step for organisations in working out the fine line between digital caring and surveillance and what it might mean, in both senses, to look out for each other online."],["Social relationships and interactions contribute to daily emotional well-being. The emotional benefits that come from engaging with others are known to arise from real events, but do they also come from the imagination during daydreaming activity? Using experience sampling methodology with 101 participants, we obtained 371 reports of naturally occurring daydreams with social and non-social content and self-reported feelings before and after daydreaming. Social, but not non-social, daydreams were associated with increased happiness, love and connection and this effect was not solely attributable to the emotional content of the daydreams. These effects were only present when participants were lacking in these feelings before daydreaming and when the daydream involved imagining others with whom the daydreamer had a high quality relationship. Findings are consistent with the idea that social daydreams may function to regulate emotion: imagining close others may serve the current emotional needs of daydreamers by increasing positive feelings towards themselves and others. --------------------------------------------------------------------------------","Social interactions and relationships are vital for a healthy, happy and meaningful life (e.g. Baumeister, Vohs, Aaker, & Garbinsky, 2012; Diener & Seligman, 2002; Holt-Lunstad, Smith, & Layton, 2010) and contribute to daily emotional well-being. For example, people report feeling happiest when socializing (Kahneman, Krueger, Schkade, Schwarz, & Stone, 2004) and during interactions with friends (Csikszentmihalyi & Hunter, 2003), and feelings of social connectedness are predicted by social activities and supportive interactions (Reis, Sheldon, Gable, Roscoe, & Ryan, 2000). Despite the emotional benefits of social interaction, a substantial portion of each day is typically spent in the absence of social activity and/or separated from close significant others (e.g. at work). However, even in the absence of social interaction the mind will invariably drift to imagine others and mentally simulate past and possible future social scenarios. Estimates suggest that we spend an inordinate amount of time daydreaming (Klinger & Cox, 1987), which is often social in nature (Mar, Mason, & Litvack, 2012). What would the impact of imagining others during daydreaming activity be on momentary feelings: could the influence of others on emotional well-being emerge from the imagination as well as from real events? In the present research we use experience sampling to explore whether everyday social daydreams are associated with increased positive social emotion and whether this depends on who is being daydreamed about. Daydreaming and its social content ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Whilst reading a novel, walking to work, or during everyday activities, the mind has a proclivity to drift to unrelated thoughts, images and feelings. Such daydreaming activity can occur as mind-wandering when attention becomes decoupled from one’s current task (Smallwood & Schooler, 2006) but can also occur when there is no specific task at hand, such as during a commute to work or when relaxing on a beach (Klinger, 2009). Daydreaming can be defined as mental content experienced during a state of normal waking consciousness that is stimulus-independent and task-unrelated, because it is neither a direct reflection of the current sensory environment nor related to the thinker’s current mental or physical task (e.g. Stawarczyk, Majerus, Maj, Van der Linden, & D’Argembeau, 2011). Defined in these ways, daydreaming occupies a substantial proportion of waking thought – a figure estimated to be between 30% and 50% (Killingsworth & Gilbert, 2010; Klinger & Cox, 1987) – and is thought to represent a psychological baseline to which people return in the absence of external demands (Mason et al., 2007). Daydreams are proposed to reflect an individual’s current goal commitments and they occur when individuals’ encounter goal- relevant cues in situations that do not lend themselves to attaining those goals (Klinger, 1975; Klinger, 1996; Klinger, 2009; Klinger, 2013). For example, hearing a friend’s name in a song on the radio may act as a reminder that the friend has an upcoming birthday, which then triggers thoughts and images about what gift to give, what the birthday party might be like, who will be there, and what conversations might unfold. In this way, daydreams are goal-relevant and involve mentally pursuing or seemingly attaining goals when doing so in reality is not possible (Baird, Smallwood, & Schooler, 2011; Klinger, 2013). Building upon this, emerging evidence indicates that daydreaming may be predominately social in nature and centered on social goals and needs. This is unsurprising given that the need to feel close and connected with others is fundamental and drives behavior and thought content towards the formation and maintenance of close, positive social bonds (Baumeister & Leary, 1995; Ryan & Deci, 2000). Mar et al. (2012) demonstrated that 73% of a large sample (N = 17,556) reported that other people are ‘frequently’ or ‘always’ involved in their daydreams whilst less than 1% reported that others are ‘never’ involved. Likewise, Song and Wang (2012) found that other people featured in 71% of sampled task-unrelated thoughts, and Andrews-Hanna et al. (2013) provide evidence that the tendency to think about others represents a major dimension of thought content. Neuroimaging data also lends converging support for the social nature of daydreams; A meta-analysis of 12 neuroimaging studies reported substantial overlap between brain regions involved in daydreaming and those involved in social cognition, suggesting a predisposition to generate social thoughts during daydreaming activity (Schilbach, Eickhoff, Rotarska-Jagiela, Fink, & Vogeley, 2008).1 Daydreaming and emotion ~~~~~~~~~~~~~~~~~~~~~~~ Why would daydreams influence feelings? Daydreams are imaginary experiences that resemble their simulated target, generally via visual and auditory imagery (Andrews-Hanna et al., 2013; Klinger & Cox, 1987). Imagining events or experiences can evoke the feelings that would arise if the simulated event were occurring (Kosslyn, Ganis, & Thompson, 2001). Indeed, the capacity of imagination to evoke and change feelings associated with the imagined subject matter is well established. Asking participants to imagine emotional events is a widely used technique to induce desired mood states (Westermann & Spies, 1996) and guided imagery is often employed in therapeutic interventions to promote positive feelings and reduce negative feelings (e.g. Hutcherson, Seppala, & Gross, 2008; Lewis, O’Reilly, Khuu, & Pearson, 2013; Panagioti, Gooding, & Tarrier, 2012). Given the capacity of the imagination to evoke feeling states, it seems plausible that social daydreams would induce social feelings associated with the imagined experience and underlying social goals and needs. Social daydreams may therefore play a role in shaping people’s everyday feelings in relation to their social goals and needs, such as feelings of love and connection. Previous research regarding the link between daydreaming and emotional well- being has tended to focus on its relationship with negative affect. For example, there is evidence to suggest that daydreaming may be detrimental to well-being due to its associations with dysphoria (Smallwood, O’Connor, Sudbery, & Obonsawin, 2007), depression (Carriere, Cheyne, & Smilek, 2008; Giambra & Traynor, 1978), rumination and self-focused attention (Marchetti, Van de Putte, & Koster, 2014) and feeling less happy in daily life (Killingsworth & Gilbert, 2010). However, research increasingly acknowledges that daydreaming is unlikely to be a homogenous experience and has begun to explore the conditions under which daydreaming is associated with negative and positive emotion. For example, the relationship between daydreaming and emotion may depend on its phenomenological and emotional content (Andrews-Hanna et al., 2013; Poerio, Totterdell, & Miles, 2013), temporal focus (Ruby, Smallwood, Engen, & Singer, 2013), interest in thought content (Franklin et al., 2013), personal lay theories (Mason, Brown, Mar, & Smallwood, 2013) and current depressive symptomology (Marchetti, Koster, & De Raedt, 2012). As an extension to these factors we propose that the social content of daydreaming will also have an impact on how daydreaming relates to emotional well-being and, in particular, to positive social feelings rather than negative emotion more generally. Research supports the proposition that imagining interactions and relationships can have a positive impact on social feelings. For example, across two studies, Kumashiro and Sedikides (2005) found that participants instructed to visualize a close positive relationship expressed warmer and more positive other-directed feelings compared to participants who had visualized a close negative, or neutral, relationship. Additionally, a large body of research on imagined contact indicates that imagining positive and neutral interactions with out-group members can promote positive feelings towards others including positive affective attitudes, increased out-group trust, and reduced inter-group anxiety (Miles & Crisp, 2014). These findings indicate that deliberately imagining interactions and interpersonal relationships can evoke positive social feelings. Whether similar effects apply to naturally occurring social daydreams that derive from an individual’s personal goals, rather than laboratory-based directed mental simulations, is an open question. However, recent correlational evidence has associated the tendency to daydream about close others with greater socio-emotional well-being (Mar et al., 2012). Daydreams about non-close others did not show a positive association with well-being, which suggests that the quality of the relationship in the daydream may be important. Given that interactions within close relationships are most likely to elicit positive social feelings in daily life (e.g. Laurenceau, Barrett, & Rovine, 2005) daydreams about close others may be especially likely to elicit positive social feelings. Another factor that may influence social feelings is the thematic content of social daydreams. The social goals underlying and influencing daydreams may be approach-oriented, i.e., concerned with the attainment of positive end-states (e.g. affiliation) or avoidance-oriented, i.e. concerned with the prevention of negative end-states (e.g. social rejection). Daydreams involving the mental pursuit of social approach goals would be more likely than those involving social avoidance goals to be associated with positive social feelings because the former engages positive cognitions and the latter engages negative cognitions (Elliot, Sheldon, & Church, 1997; Tamir & Diener, 2008). Although individual social daydreams may be associated with increased or decreased positive social feelings, there is also reason to suspect that, as a general pattern, social daydreams will be associated with increased positive social feelings. We make this prediction from a study by Johannessen and Berntsen (2010) that assessed participants’ current goal commitments which often referred to social life categories including “love, intimacy and sexual matters” and “friends and acquaintances”. Importantly, participants reported their specific goals to be related to achievement rather than avoidance. This suggests that daydreams will be predominately associated with mentally pursuing desired social goals, which in turn, should increase the positive social feelings associated with their imagined pursuit or attainment. The present research ~~~~~~~~~~~~~~~~~~~~ Building on the ideas presented above, we sought to explore whether social daydreams would be associated with increased social feelings by choosing to focus on feelings of love and connection. We used experience sampling methodology (Bolger, Davis, & Rafaeli, 2003) to examine naturally occurring daydreams and associated feelings. In addition to reports of social daydreams and social feelings, we sampled non-social daydreams and non-social feelings to serve as points of comparison. On the assumption that social daydreams commonly represent attempts to mentally pursue social approach goals we predicted that social daydreams, but not non-social daydreams, would be associated with increased love and connection. We did not make specific predictions about whether social and non-social daydreams would differ in their association with changes in non-social feelings. Daydreams with and without social content could both relate to changes in non-social feelings. We also measured the emotional content of daydreams to rule out the possibility that the predicted increases in social feelings could be attributed to the emotional, rather than social, content of daydreaming. In addition to our main prediction, we explored whether the effect of social daydreams on social feelings might depend on who was involved in the daydream by measuring the relationship quality between the daydreamer and the most central other person involved. We predicted that relationship quality would moderate the effect of social daydreams on positive social feelings whereby daydreams involving higher quality relationships would be positively associated with increases in positive feelings.","One hundred and one volunteers (81 women, 20 men; Mage = 22.32 years, SD = 5.17) were recruited to the study. It was described as an investigation into the content and nature of daydreams and advertised via email, flyers at a public engagement event, personal contacts and referrals. Of the participants, 49 were undergraduate psychology students, 22 were postgraduate students, 20 were in full-time employment, and 10 were non-psychology undergraduate students. In exchange for their participation, undergraduate psychology students were given study credits; all other volunteers were entered into a prize draw to win shopping vouchers equivalent to $33, $50 and $83. Experience sampling protocol ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We used a signal-contingent experience sampling protocol (Wheeler & Reis, 1991) to sample daydreaming and associated feelings. Participants were signaled four times via text messages to their cell phones to answer online questionnaires about their two most recent social and two most recent non-social daydreams. The questionnaires were answered by following a survey link sent within the messages. Participants received the four messages on one day between 10 am and 10 pm at individually randomized times within four three-hour blocks (between 10:00–13:00, 13:00–16:00, 16:00–19:00, 19:00–22:00), with the constraint that consecutive signals were at least one hour apart. The order of the questionnaires (social, non-social) was also individually randomized for each participant. We randomized the time and order of questionnaires to prevent anticipation of signals, to sample daydreams and feelings across a range of times and daily activities, and to counteract potential order effects and demand characteristics. Measures were kept brief so as not to unduly interfere with participants’ daily routine. Social feelings Two items, taken from Crocker, Niiya, and Mischkowski (2008), measured the positive, other-directed feelings of love and connection. Participants indicated how loving (“How loving did you feel before/after your daydream?”) and connected with others (“How did you feel before/after your daydream?”) they felt before and after their daydream on 7-point scales from not at all to extremely. Non-social feelings Participants indicated how they felt before and after their daydream (“How did you feel before/after your daydream?”) on the following dimensions: sad-happy, anxious-calm, and excited-bored (reverse-scored). Responses were made on a 7-point scale (e.g. 1 = sad, 7 = happy). These items were chosen to measure the pleasure (valence) and arousal (activation) dimensions of core affect (Remington, Fabrigar, & Visser, 2000); specifically, pleasure (sad-happy), pleasant deactivation (anxious-calm) and pleasant activation (bored-excited). Emotional content A single item measured the emotional content of daydreams. Participants indicated the emotional content of their daydream (“The emotional content of this daydream was…”) on a 7-point scale (1 = very negative, 4 = neutral, 7 = very positive). Relationship quality Three items were used to provide a quality index of the relationship between participants and the most central person involved in their daydream. Participants rated their general feelings of closeness (“In general, how close do you feel to them?”), liking (“In general, how much do you like them?”), and trust (In general, how much do you trust them?”) towards the most central person in their daydream on 7-point scales from not at all to extremely. These items were chosen to reflect indicators of high-quality interpersonal connections (Niven, Holman, & Totterdell, 2012). We combined items to create an overall score for relationship quality; internal reliability was high, α = .92.","Participants attended an individual training session during which they were given a written and verbal description of daydreaming. A daydream was defined as a series of connected thoughts and/or images where that mental content is not about whatever mental or physical activity one is engaged in at the present moment. Participants were told that daydreams could be brief but should consist of more than a single thought or image. Social daydreams were defined as daydreams where another (real or imaginary) person or people are involved; non-social daydreams were defined as daydreams that did not involve another person or people. Examples of daydreams, including social and non-social ones, were provided. When participants indicated that they understood what counted as daydreaming, they were provided with written instructions for the study followed by a demonstration of the text message with online questionnaire link and verbal explanation of the meaning and response of each questionnaire item. Finally, participants nominated a date to complete the study and were free to choose whatever day they liked as long as it represented a typical day in their life. On the nominated day, participants followed the online questionnaire link sent via text and, after entering their unique participation number, indicated their social and non-social feelings before and after their last (social or non- social) daydream. All five items referring to feelings were asked twice (with reference to before and after the daydream) but the order of all 10 question items was individually randomized to minimize response bias. Participants then rated the emotional content of their daydream, and for social daydreams, completed items indexing relationship quality. Participants were asked to report on their last social or non-social daydream before each text message but were not asked about the time lapse between the daydream and reporting its content. Participants were also asked to provide a short description of the daydream. Several additional questions regarding the daydream were also asked but were not relevant to the hypotheses tested here. Response rate ~~~~~~~~~~~~~ Overall, 383 of a possible 404 daydreaming questionnaires were completed (192 social and 191 non-social daydreams) corresponding to a 95% response rate. Participants’ daydream descriptions were examined by the first author, in line with the definitions given to participants, to ensure that daydreams were accurately categorized as social or non- social. Twelve non-social daydreams were excluded from the dataset because they contained references to other people which suggested that they may have been instances of social, rather than non-social, daydreams. We chose not to reclassify these as social daydreams because doing so would have led to an unbalanced design. Therefore, the following analyses were based on 179 non-social and 192 social daydreams. Were social daydreams associated with increases in positive feelings? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To examine whether social, compared to non-social, daydreams were associated with increases in positive feelings, we conducted a series of 2 × 2 × 2 (Daydream Type [social, non-social] × Time [pre, post] × Questionnaire [questionnaire 1, questionnaire 2]) triply repeated-measures ANOVAs with each feeling state (i.e. happiness, calmness, excitement, loving and connected) as the dependent variable. The results of these analyses are presented in Table 1. Social feelings There was a significant interaction between daydreaming type and time on feeling loving (F(1, 73) = 14.70, p < .001, ηp2 = .17). Post-hoc repeated measures t-tests indicated that participants reported feeling significantly more loving after (M = 4.86, SD = 1.23) compared to before (M = 4.20, SD = 1.17) social daydreams (t(73) = −3.64, p = .001, d = −.39, 95%CI [−.75, −.21]) and significantly less loving after (M = 3.70, SD = 1.34) compared to before (M = 3.92, SD = 1.06) non-social daydreams (t(73) = 2.06, p = .043, d = .18, 95%CI [.01, .45]). Similarly, there was a significant interaction between daydreaming type and time on feeling connected with others (F(1, 73) = 14.28, p < .001, ηp2 = .16). Repeated measures t-tests indicated that participants reported feeling significantly more connected with others after (M = 4.67, SD = 1.23) compared to before (M = 4.22, SD = 1.26) social daydreams (t(73) = −3.40, p = .001, d = −.39, 95%CI [−.70, −.19]). Participants also reported feeling marginally less connected with others after (M = 3.61, SD = 1.28) compared to before (M = 3.80, SD = 1.18) non-social daydreams (t(73) = 1.89, p = .063, d = .15, 95%CI [−.01, .38]), although this result was statistically non-significant. Overall, social daydreams were associated with increased, whereas non-social daydreams were associated with decreased, feelings of love and connection (see Fig. 1). Non-social feelings There was a significant interaction between daydreaming type and time on happiness (F(1, 73) = 5.72, p = .019, ηp2 = .07). Post-hoc repeated measures t-tests indicated that participants reported feeling significantly happier after (M = 4.91, SD = 1.23) compared to before (M = 4.53, SD = 1.06) social daydreams (t(73) = −2.67, p = .009, d = −.33, 95%CI [−.64, −.10]), but not non-social daydreams (t(73) = .88, p = .383, d = .09, 95%CI [−.14, .35]). Social daydreams were associated with increased happiness but there was no change in happiness for non- social daydreams (see Fig. 1). The same interaction effect was not observed for feelings of calmness or excitement. However, there was a significant main effect of time for these feelings. Participants were significantly less calm after (M = 4.50, SE = .14) compared to before (M = 4.82, SE = .13) daydreaming (F(1, 73) = 7.74, p = .007, ηp2 = .10) and significantly more excited after (M = 3.31, SE = .10) compared to before (M = 4.20, SE = .11) daydreaming (F(1, 73) = 68.06, p < .001, ηp2 = .48). This suggests that daydreaming was associated with decreased calmness and increased excitement, which may be indicative of an overall increase in the arousal dimension of core affect. Post-hoc analysis: Were social daydreams regulating people’s feelings? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our results indicate that social, but not non-social, daydreams were associated with increased happiness, love and connection. We were interested in whether this effect might be regulatory; that is, whether social daydreams might be compensating for low levels of happiness, love and connection. If that was the case, then we would expect increases in happiness, love and connection to be observed for participants who scored low, but not for those who scored highly, on these feelings before daydreaming. To explore this, we created each participant’s average scores for feelings of happiness, love and connection before and after social daydreams over the two time samples. We then ran a series of repeated measures t-tests to examine differences between feelings of happiness, love and connection before and after social daydreams separately for those ‘high’ and ‘low’ in the associated feeling before daydreaming. We classified each participant as ‘low’ or ‘high’ using a median split of their average feeling state before social daydreams. The results were consistent across feeling dimensions: increases in happiness, love and connection were only observed for those participants scoring ‘low’ and not for those already ‘high’, on the associated feeling before social daydreaming. Participants low in happiness felt significantly happier after (M = 4.54, SD = 1.14) compared to before (M = 3.86, SD = .71), social daydreaming (t(56) = −4.41, p < .001, d = −.71, 95%CI [−.99, −.39]); participants low in feelings of loving felt significantly more loving after (M = 4.03, SD = 1.26) compared to before (M = 3.24, SD = .92) social daydreaming (t(45) = −3.90, p < .001, d = −.71, 95%CI [−1.20, −.43]); and participants low in feelings of connection felt significantly more connected with others after (M = 4.20, SD = 1.30) compared to before (M = 3.04, SD = .99) social daydreaming (t(44) = −5.68, p < .001, d = −.99, 95%CI [−1.56, −.77]). In contrast, there were no significant differences observed between feelings of happiness, love and connection, before and after social daydreams for participants that were ‘high’ on the associated affective state before daydreaming (all p’s > .1) Inspection of mean scores and 95% confidence intervals also suggests that results for ‘high’ scorers were not due to a ceiling effect in this group: mean ratings (on a 7 point-scale) and confidence intervals before daydreaming were 5.47 (95%CI [5.32–5.60]) for happiness, 5.11 (95%CI [4.98–5.31]) for love, and 5.19 (95%CI [5.05–5.34]) for connection. For comparison, we performed the same set of analyses with non-social daydreams. Levels of happiness from before and after non-social daydreams were not different for participants low (t(57) = −.96, p = .340) or high (t(40) = 1.41, p = .168) in happiness before non-social daydreaming. Participants low in feelings of love and connection before non-social daydreams did not report significant increases in these feelings after non-social daydreams (loving: t(59) = −.12, p = .906; connected: t(42) = −1.22, p = .230). However, participants high in feelings of loving before non-social daydreams felt marginally less loving after (M = 4.74, SD = 1.04) compared to before (M = 4.99, SD = .57) non-social daydreams (t(38) = 1.88, p < .068, d = .24, 95%CI [.01, .49]). Similarly, participants high in feelings of connection felt significantly less connected after (M = 4.36, SD = 1.15) compared to before (M = 4.77, SD = .78) non-social daydreams (t(55) = 3.16, p = .003, d = .40, 95%CI [.16, .67]). These results support a regulatory explanation of our findings: the positive emotional outcome of social daydreaming was only found for participants who would benefit from it the most (i.e. ‘low’ scorers), but not for participants already experiencing positive feelings (i.e. ‘high’ scorers). The fact that the opposite pattern of results was observed for non-social daydreams also suggests that these effects were not a result of regression to the mean. Did the emotional content of social daydreams account for changes in feelings? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We sought to establish whether increases in happiness, love and connection for social daydreams were attributable to the emotional content of those daydreams.2 Multi-level modeling (Hox, 2010) was used to examine this possibility because the measure of daydreaming emotional content was time-varying. We restructured the data so that time points (i.e. questionnaire responses) were nested within individuals and then ran a series of multi-level models. First, we confirmed that social daydreams were associated with increased happiness, love and connection: the fixed effect of time was significant in models predicting happiness (B = −.29 (.13), t(283) = −2.32, p = .021, ICC = −.15, 95%CI [−.54, −.04]), loving (B = −.46 (.13), t(281) = −3.66, p < .001, ICC = −.20, 95%CI [−.70, −.21]), and connection (B = −.48 (.13), t(281) = −4.07, p < .001, ICC = −.20, 95%CI [−.73, −.23]); the respective feelings were significantly greater after compared to before social daydreams. Next, we included the emotional content of social daydreams as a predictor in each model. The emotional content was a significant (p < .001) predictor in all models indicating that the emotional content of daydreams was positively associated with happiness, love and connection, before and after social daydreams. After controlling for the emotional content of daydreams, the fixed effect of time remained significant in models predicting happiness (B = −.29 (.12), t(282) = −2.54, p = .012, ICC = −.19, 95%CI [−.52, −.07]), loving (B = −.46 (.11), t(281) = −4.17, p < .001, ICC = −.27, 95%CI [−.67, −.24]) and connection (B = −.48 (.12), t(281) = −4.07, p < .001, ICC = −.25, 95%CI [−.71, −.25]), indicating that the association between social daydreams and increased happiness, love and connection could not be solely attributed to the emotional content of those daydreams. Did the effect of social daydreams on positive feelings depend on relationship quality? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The measure of relationship quality was significantly (p < .001) negatively skewed indicating that participants’ daydreams overwhelmingly involved significant others. Attempts to transform the variable to normalize the distribution were unsuccessful so a median split procedure was applied (MacCallum, Zhang, Preacher, & Rucker, 2002). We dichotomized the variable to represent ‘high’ and ‘low’ quality relationships for the sample: low (n = 186) = 1–5.67; high (n = 198) = 6.00–7.00. We then ran multi-level models examining whether feelings of happiness, love and connection, were significantly greater after, compared to before, social daydreams separately for daydreams that involved low, and high quality relationships. For low quality relationships, the fixed effect of time was non-significant for models predicting happiness (B = .12 (.15), t(113) = .80, p = .423, ICC = .05, 95%CI [−.17, .41]), loving (B = −.09 (.15), t(114) = −.53 p = .600, ICC = −.04, 95%CI [−.41, .24]) and connection (B = −.19 (.15), t(114) = −1.23, p = .223, ICC = −.08, 95%CI [−.51, .12]), indicating no significant change in feelings from before to after daydreaming. In contrast, for high quality relationships, the fixed effect of time was significant in models predicting happiness (B = −.68 (.17), t(118) = −3.94, p < .001, ICC = −.51, 95%CI [−1.02, −.34]), loving (B = −.82 (.14), t(119) = −5.82, p < .001, ICC = −.49, 95%CI [−1.08, −.53]) and connection (B = −.75 (.16), t(118) = −4.82, p < .001, ICC = −.38, 95%CI [−1.06, −.44]), indicating more positive feelings after compared to before daydreaming. Specifically, after daydreams involving high quality relationships, feelings of happiness (M = 5.41, SD = 1.42), love (M = 5.43, SD = 1.26) and connection (M = 5.31, SD = 1.28) were greater than feelings of happiness (M = 4.74, SD = 1.10), love (M = 4.63, SD = 1.27) and connection (M = 4.57, SD = 1.47) prior to daydreaming. These results indicate that social daydreams were associated with increases in feelings of happiness, love and connection, only when participants’ daydreams involved people with whom they had a high, but not low, quality relationship.","We investigated the impact of social daydreams on momentary feelings to explore whether the positive influence of others on emotional well-being could emerge from the imagination as well as from real events. We sampled naturally occurring social and non-social daydreams and associated social and non-social feelings before and after those daydreams using experience sampling methodology. Because daydream activity is commonly concerned with the mental pursuit of social goals, we proposed that social daydreams would induce positive social feelings associated with the imagined experience and underlying social goals and needs. Our results support this hypothesis; everyday daydreams with social, but not non-social, content were associated with increases in feelings of love and connection. Social daydreams were also associated with increases in happiness which may reflect the tendency for happiness to be a positive emotion linked with social interaction (e.g. Csikszentmihalyi & Hunter, 2003; Kahneman et al., 2004). Our results suggest that people’s everyday social feelings are shaped by their imaginary, as well as actual, social worlds, and that daydreams can be a source of positive other-directed feelings. Importantly, the possibility that increases in happiness, love and connection were due to the emotional content of daydreams was ruled out because increases in these feelings were still present when statistically controlling for emotional content. This suggests that imagining the pursuit of social goals is not simply a pleasant experience, but also one that is associated with a similar emotional outcome to that which would occur as a result of actually pursuing those same goals. Additionally, increased positive feelings were only observed when the relationship between the daydreamer and most central person in the daydream was classified as ‘high’, paralleling the pursuit of social goals in reality. In the same way that different forms of social interaction differentially contribute to social feelings (Reis et al., 2000), with interactions within close relationships more likely to elicit positive social feelings (e.g. Laurenceau et al., 2005), imagining close others appeared to elicit greater feelings of love and connection. This is also consistent with previous research linking daydreams about close others to socio-emotional well-being (Mar et al., 2012). Additional analyses support the idea that social daydreams may function to regulate emotion because increases in happiness, love and connection were present only when participants were low, but not high, in these feeling before daydreaming. This pattern of results suggests that social daydreams may have compensated for deficiencies in social feelings serving the emotional needs of the daydreamer at the time. In daily life, feelings of social disconnection may act as a cue or trigger to pursue or attain goals that would alleviate these feelings (e.g. physical contact with a close other). However, in situations that do not lend themselves to social contact, daydreams may provide an opportunity to simulate desired contact, which fosters feelings of love and connection. In this way, daydreams involving close others may function as an imaginary substitute when close others are not immediately available in reality (c.f. Maner, DeWall, Baumeister, & Schaller, 2007; Pickett, Gardner, & Knowles, 2004). A prediction from this is that, relative to non-social daydreams, daydreams about close others should increase during periods of social disconnection. Future research might explore this during transitional periods where socio-emotional needs are likely to be pertinent (e.g. homesickness, Watt & Badger, 2009), and investigate whether social daydreams can promote adaptation. Of course, social daydreaming may also become dysfunctional if it comes to unduly displace actual social interaction (Somer, 2002), but future studies will need to ascertain for whom social daydreaming is adaptive or maladaptive and under which circumstances. Future research might also explore the feasibility of encouraging certain types of social daydreaming as a means of enhancing well-being and personal relationships. Daydreaming and thinking about close others have previously been identified as emotional regulation strategies, although not in conjunction. Daydreaming has been reported as a strategy to relieve boredom (e.g. Fisher, 1987; Singer, 1966) and as an emotion-focused coping strategy (Endler & Parker, 1990), and recurrent daydreams have been seen as a source of comfort in times of distress (Greenwald & Harder, 1997). Similarly, thinking about close others is reported as an example of a cognitive emotion regulation strategy (Parkinson & Totterdell, 1999), and imagining close others can regulate negative emotions about an impending stressful event (McGowan, 2002) and alleviate the distress of recalling a distressing memory (Selcuk, Zayas, Günaydin, Hazan, & Kross, 2012). Our results support and extend these findings by providing evidence that daydreaming about close significant others may be an adaptive emotion regulation strategy to alleviate feelings of social disconnection in daily life. Previous research regarding the link between daydreaming and emotion well-being has associated daydreaming with negative mood and depressive symptomology (e.g. Carriere et al., 2008; Giambra & Traynor, 1978; Smallwood, O’Connor, Sudbery, & Obonsawin, 2007) but other research suggests that this relationship may depend on factors such as the thought content of daydreams and individual differences (e.g. Andrews-Hanna et al., 2013; Poerio et al., 2013; Ruby et al., 2013). The current research indicates that the association can sometimes be positive, depending on whom the daydreamer focuses, how the daydreamer feels at the time, and which emotions are examined. Although current findings are consistent with the idea that social daydreams function to up-regulate social feelings, our correlational design prevents causal interpretations from being made. We cannot confirm whether feelings of social disconnection caused participants to imagine close others which then caused them to feel happier, more loving and connected with others. If social daydreams were functional for up-regulating feelings, then low levels of happiness, love and connection should predict the occurrence of social, rather than non-social, daydreaming. However, because participants reported on either their last social or last non-social daydream rather than their last daydream of any type, we cannot shed light on this issue. A future study might address this shortcoming by experimentally inducing feelings of social disconnection and examining the social content of daydreaming activity during a subsequent task. Another limitation is the use of retrospective reports for daydreams, associated feelings before and after daydreaming and measures of emotional content and relationship quality. Participants reported on their most recent social or non-social daydream at four quasi-random intervals within four, three-hour blocks, but were not asked to estimate how long ago their daydream was experienced. The time between experience and recall may have influenced the validity of reports in ways that we cannot control for or explore (Bradburn, Rips, & Shevell, 1987). However, given that daydreaming is thought to occupy between 30% and 50% of waking thought (Killingsworth & Gilbert, 2010; Klinger & Cox, 1987), we suspect that the interval between experience and recall would have been relatively small (i.e. minutes rather than hours) and hence the potential effects on accuracy would also be small. Another possible consequence of the use of retrospective reports, particularly with reference to feelings before and after daydreaming, is that our results may reflect a demand characteristic or participants’ own lay theories concerning how they should have been feeling before and after social and non- social daydreams. Although we cannot rule out these possibilities, we took steps to minimize potential demand characteristics. Pre- and post- daydreaming feeling measures were individually randomized meaning that participants could have completed the questions concerning their feelings on each dimension (happy, calm, excited, loving, connected) referring to before and after their last daydream (i.e. 10 items) in any possible order. In addition, each question (e.g. “How loving were you feeling before your daydream?”) was completed on participants’ smartphone screens individually such that participants were unable to view their previous responses. If participants reported emotion change to fit with their possible views on the study, then they would have had to remember their responses for each individual measure to use as a reference point for reporting feeling change. Although we cannot be certain that lay theories about the influence of social and non-social daydreams on social feelings did not influence participant responding in the current study, this issue has been addressed in previous research (Poerio et al., 2013) which found no evidence to suggest that participants believe that daydreaming has either predominately positive or negative effects on mood, or that lay beliefs are consistent enough across participants to systematically bias results. Although this does not specifically shed light on lay theories concerning how social daydreams relate to social feelings, there is no reason at present to suspect that people associate social daydreams in particular with increases in positive feelings. To address concerns associated with retrospective sampling of daydreaming and affect, a more intensive time-sampling approach could be used in future research where participants report on current daydreaming activity and current affective states at separate time points. However, whether this methodological benefit would outweigh the additional participant burden would need careful consideration (Stone, Kessler, & Haythornthwaite, 1991). Overall, the present research represents an important first step in establishing and exploring how social daydreams shape and regulate social feelings in daily life. The idea that positive social interactions and interpersonal relationships are vital for well-being is well accepted (Baumeister et al., 2012; Diener & Seligman, 2002; Holt-Lunstad et al., 2010). However, the current findings offer a broader conception of the role that others play in socio-emotional well-being by demonstrating that the influence of others on positive emotion can emerge from the imagination as well as from real events. Our findings indicate that love can really be a triumph of the imagination, albeit not in the sense originally intended in the popular quotation.3 Although research on daydreaming has experienced a resurgence in recent years (Callard, Smallwood, Golchert, & Margulies, 2013) the field would benefit from examining social aspects of daydreaming to uncover the ways in which people’s imaginary, as well as actual, worlds contribute to their socio-emotional functioning.","The research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest."],["Background/purpose: Diasporic associations and hometown groups fuel transnational exchanges and circulations. Their role has mostly been understood in terms of broader calculative agendas related to ethnic and national cultural politics. In South Africa, classical Indian singers, dancers and instrumentalists are an important part of these transnational landscapes. This paper focuses on the individual actors giving shape to these flows, and explores how a range of subjectivities is entangled with the materialities and forces present in classical performance spaces. Methods and results: Drawing on fieldwork in Durban, South Africa, it explores how, and why organising actors assemble the matter of classical performance spaces. The paper also explores interconnections to Bollywood as another emergent diasporic site both in tension and accord with classical Indian performances. Conclusion: Drawing from a feminist social practice approach, this paper argues that diaspora associational life is assembled through agents negotiating different gaps and discrepancies arising from the material and affective inhabitation of diasporic worlds. --------------------------------------------------------------------------------","On the 9th of May 1993 after a gap of over four decades, diplomatic relations between India and South Africa were formally restored. Since then, the gradual transnational (re)materialisations of South African Indian1 life have arisen out of intersecting sets of commitments to develop closer ties with India. These include new desires of the Indian government to engage people of Indian origin in South Africa as part of its overseas diaspora strategy, the growing role of South African Indian diaspora entrepreneurship, and the increasing influence of Indian and South African Bollywood promoters and film producers (Dickinson and Bailey, 2007; Ebrahim, 2008; Munjal et al., 2013). Amidst these activities, various South African Indian conservative religious and cultural associations, formed during the apartheid-era system of separate development, are developing closer ties with India through transnational circulations of classical Indian performers, dancers and musicians. The term ‘South African Indian’ expresses a positionality that is constituted through a multiplicity of competing identifications, and expressed and negotiated in a diverse range of interconnected transnational spaces (Landy et al., 2004). Yet, as Hansen (2005) has argued, diaspora associations draw upon classical Indian music and dance to negotiate, order and control the boundaries and expressions of Indian diasporic identities, contributing to new contestations and dilemmas surrounding South African Indian ethnic, national and gender politics (see also Ganesh, 2010; Radhakrishnan, 2005; Ramsamy, 2007; Rastogi, 2008; Vahed and Desai, 2010). In addition to being locations for negotiations over categorical politics, South African Indian classical performances are also sites of affective and embodied experiences. The quote at the start of this paper is taken from a description of a performance by an Indian Carnatic singer at a 2004 Swami Thygaraja music festival sponsored by the Indian Academy of South Africa, an association involved in the promotion of closer ties between South Africa and India. The description alludes to deeply embodied and felt experiences of being enthralled by the music's sonic affects. Whilst there is important critical work that has been done in thinking through the implications of South African Indian diaspora associations' transnational connections for collective negotiations of Indianness (e.g. Ganesh, 2010; Radhakrishnan, 2005), this paper aims to critique, and open out, these discussions by exploring classical Indian performances as a specific site for the production of a more ambiguous and complex self. To do so, it shifts the focus of analysis from the types of diasporic identifications represented by classical performances towards the spaces themselves; and examines how individual organising actors relate to the human and non-human materials, forces and elements assembled therein. I use the case study of South African Indian associational life to call for scholars of diaspora and hometown associations to interrogate and engage with the wider turn to practice and performativity in the humanities and social sciences in order to enliven critical understandings. Diaspora actors are often understood as possessing broader calculative and strategic agendas related to political identity discourses, positions and practices in the homeland (Quinsaat, 2013). Whilst recent work on hometown associations has begun to produce more nuanced understandings of the range of subjectivities produced through the activities of diaspora and hometown associations (Mercer et al., 2008), this paper explores how an affective, relational, social practice understanding of migrant subjectivities can extend these theorisations. To do so, it is guided by feminist scholarship that explores the material and affective dimensions of the ‘body-subject’ (Grosz, 1994; Probyn, 2004). Feminist interrogations that view body- subjects as assemblages or collections of, in Barad's (2001) terms, ‘intra-acting’ non- human and human components, have reoriented understandings of identity away from pre- existing named and sorted categories towards the multiple forces and intensities that both constitute and circulate between bodies and “provide the backdrop to and are active in producing what comes to be understood as ‘a’ subject” (Colls, 2012: 439). In this understanding, subject-formation is a multiple process of becoming, rather than being, one that is not wholly predefined by neither categorisations, nor stable and finished. Such a reading is useful for exploring the multi-dimensional agencies that inhere in associational life; and second by providing openings for further understanding the contingencies, discrepancies and ambivalences involved in marshalling and cajoling transnational circulations and flows of bodies and materials through sites of associational life. In the next section of the article I will describe briefly previous approaches to diaspora associations, before outlining debates that have considered the relationship between affect and migration in the production of mobile subjectivities. Following an overview of research into South African Indian transnationalities, the discussion in the first half of the analysis draws attention to how organising actor's subjectivities emerge from and choreograph the materialities of classical Indian performance spaces. Of specific interest are the matters and forces that contribute towards the making of perceived bonds between organisers, audiences and performers. The paper then adds a further layer to questions of complex subjectivities by discussing the overlaps between classical performance sites and the broader array of transnational sites through which ideas of India are produced and negotiated, focussing specifically on Bollywood. Of interest are the points of tension and accord between Bollywood and classical Indian performances, and the subjectivities that emerge in their negotiations. In doing so, the paper hopes to raise some questions for future scholars of diaspora associations to engage with and mark new directions for study.","Diasporic associational life, such as festivals, music and dance performances, meetings, religious expression and worship, is a site for the production of transnational and translocal belongings (Mercer et al., 2008). It provides a platform for the staging and reproduction of ancestral cultural traditions that can act as mechanisms for the expression of the dominant device of diaspora: the recreation of a homeplace in new settings (Gal et al., 2010). These mechanisms are of course also deployed by individuals in the transnational social spaces of their everyday lives (Ley, 2004); a key difference is that diaspora associations and groups employ such spaces in order to perpetuate a sense of a collective identity sustained by reference to common memories or myths about an ancestral homeland's location, history and achievements (Tölölyan, 2010). Analyses of diaspora associations have most commonly focused on the production of long-distance national identifications as a discursive political process (Gal et al., 2010). A wide body of scholarship explores the agency of diaspora associations in the politics of their home countries, attending to the dynamics of home-making involved (Brinkerhoff, 2011; Lyons, 2006; Sheffer, 2003). Related work focuses on the ways in which associations, and the organising elites involved, can direct flows of remittances, resources and knowledge in order to reproduce their own economic and political agendas or positions (Mohan, 2006). Scholarship exploring Hindu associations in North America perhaps exemplifies this approach, drawing attention to the transnational practices that allow diaspora elites to influence a range of Indian political and economic concerns, such as remittance activities, homeland identity preservation, and sectarian nationalist politics (Biswas, 2004; Mathew and Prashad, 2000). This kind of work is valuable in understanding the contested nature of the practices of diaspora associations and the dilemmas over competing identities they call to light. However, recent scholars of African diasporas have pushed us to reconceptualise the role of associations in mobilising around the shared and contested predicaments that arise from ruptures with home and the forging of local socialities (Mercer et al., 2008). Studies focussing on the convivialities, obligations and the gendered subjectivities on which associations' practices turn (Faria, 2011; Kleist, 2010; Mercer and Page, 2010) have dislodged dominant scholarly assumptions that have in the past essentialised diasporic and hometown associations as divisive, or even dangerous. By drawing attention to the ruptures and inconsistencies involved, what ensues is an understanding of the subjectivities underpinning diaspora elites' mobilisations of homeland ties as differentiated and multiply located. Whilst there is still further critical work to be done in understanding the intersectionalities of identity that inhere in transnational socio-geographic flows and connections – and their politics – there is also room to consider an even more expansive notion of diasporic selfhood, one that is attuned to the affective forces of social life that are weaved from a distributed range of experiences and material exchanges that cannot be “cleanly or clearly cleaved into a set of named, known and represented identities” (Anderson and Harrison, 2010: 10). My use of the term ‘affective’ here is drawn from previous work that articulates the embodied, physiological states that enable people to recognise themselves as subjects in ways that are not wholly predefined (Massumi, 2002). Most studies of diaspora associations have focused on organising actor's agency in terms of broader calculable economic and political agendas (such as development, remittances or hometown politics), and the categorical identities that underpin them. But, as the above scholarship emphasises, agency must also be understood in terms of its affective dimensions, ones that are made through a range of desires and intimacies (Berlant, 2000). Understandings of diaspora associations need to also consider what other kinds of forces motivate the organising individuals that comprise diaspora associations, and the roles those forces play in transnational exchanges and circulations. Holding open understandings of the self rather than privileging pre-defined concepts or symbolic categorisations has become central to deepening understandings of migration. Specifically, renewed emphasis has been placed on the performances and practices of material circulations across transnational and translocal space (Coe, 2011; McKay, 2007; Tolia-Kelly, 2004), and on the complex register of emotional experiences and affects that arise and are transformed through transnational circuits and connectivities (Andrucki, 2013; Boehm and Swank, 2011; Christou, 2011; Ho, 2009). This work mirrors broader geographical scholarship on the range of localities through which multiple co- present relations are entangled with the production of self-in-society (Kraftl and Adey, 2008; Stewart, 2007; Thien, 2005). Conradson and Mckay (2007) affirm the importance of the affective states that accompany the dialectics of mobility and emplacement, but also the instabilities such constructions call forth. The corporeal intensities of the transnational and translocal circulations and encounters with an array of multiply-located bodies and materialities, they argue, are linked to the formation, and also breakdowns, of individual and community selfhoods (see also Faier, 2013). Two connected interventions in cultural studies complement this work. First, scholarship examining musical performances, rituals and festivals organised by diaspora cultural associations offers insights into the ways that migrants move, feel and experience a dynamic and plural conception of home through animated sites and embodied practices (Farah, 2005). This work mirrors the broader turn towards practice-orientated approaches to music and performance that escapes categorical identity fixings by asking what music does rather than what it represents (Revill, 2005; Wood et al., 2007). The social practices and sensory forms associated with diasporic musical performance, such as narrative themes, linguistic registers and material styles, as Gilbert and Lo (2010) argue, “performs and activates a wide range of links with homeland and hostland” (151) and act as sites of “polycultural exchange but also as a zone of heightened affect” (155) where imagined pasts and futures can be expressed. Whilst many have argued that diasporic music and song are consciously engineered towards wider strategic goals, this work highlights the sensory experiences and meanings that bring diasporic cultural forms into being and give them shape. This work does not disregard collective inscriptions of homeland imaginaries, but fleshes out the underlying embodied practices and affective capacities that bring those inscriptions into being and also work to destabilise them. Second, and relatedly, scholars of queer diaspora have drawn attention to the multiple performance practices that contest, negotiate, and heighten ideas of diasporic belonging (e.g. Manalansan, 2003). Gopinath's (2005) work on South Asian public cultures, for example, unsettles the connections between South Asian diasporic culture and the (re)production of normative nationalisms, arguing that a range of feelings, emotions, dreams and subject positions are realised through (dis)affinities with an array of socio-spatial identifications. By identifying the contexts in which diasporic nationalisms begin to break down, and the unstable assumptions on which they are based, this work escapes categorical fixings by retheorising diasporic subjectivities as a rich assemblage of sites, practices and materialities produced through multiple locations and temporalities.","Indians in South Africa have multiple historical origins, producing a complex of identifications that are far from homogenous. Migration from India to South Africa began around 500 B.C, intensifying under Natal's indentured sugar plantation labour system in the nineteenth-century (Kuper, 1960: 1). Approximately 152,000 linguistically (Telugu, Tamil, Hindi, Urdu) and religiously diverse (variants of Hinduism, Christianity, Islam) indentured labourers arrived from the Madras Presidency and Northern Indian districts between 1860 and 1911 (Lemon, 2009: 131). Approximately 30,000 ‘passenger’ traders from Gujarati villages followed, part of the long established Indian trader network that connected East Africa, East Asia, Mauritius and the West Indies (Bhana and Brain, 2000: 36). ‘Passengers’ were equally diverse linguistically (Urdu, Gujarati) and religiously (variants of Hinduism, Parsi, Christianity, Islam) (Kuper, 1960: 7–8). Their descendants now number some 1.3 million, approximately 2.6% of the South African population (Statistics South Africa, 2010). In colonial and tapartheid Natal, classical Indian song and dance came to play a particularly important role in sustaining the different Indian vernacular communities and promoting ties to India. According to Bhana and Bhoola (2011: 18), associations were based around promoting so-called ‘traditional’ Indian values: the vernacular languages, correct behaviours and rich cultural and historical traditions of their region of origin. Some examples cited by Bhana and Bhoola include the Arya Pratinidhi Sabha (formed in 1925) supporting Hindi teaching and various Gujarati Hindu associations promoting the traditions and vernaculars of Gujarat, such as the Surat Hindu Association (formed in 1907). By the mid to late twentieth century, the significance of music in maintaining Indian regional identities was further heightened by the decline of other traditional emblems of Indianness, such as language and caste consciousness (Ebr- Vally, 2001). The strictures of apartheid played an important role in this process as the vernaculars associated with these broader regional distinctions disappeared almost entirely with the enforced teaching and use of English in Indian schools (Landy et al., 2004). Furthermore, different caste or regionally-focussed cultural associations pooled resources to sustain Indian vernacular traditions, leading to a redefinition of Indianness within broad regional, linguistic and religious parameters (Bhana and Bhoola, 2011). Bhana and Bhoola describe how this was the case for the multiple Gujarati Hindu communities in 1943 with the formation of the umbrella organisation Kathiawad Hindu Seva Samaj (KHSS), and again in 1993 with the merger of the KHSS with the Surat Hindu Association and the Saptah Mandir into the Gujarati Sanskruti Kendra. Broadly defined, Indian vernacular traditions have remained the basis of a loose, shared sense of community, which continues to be promoted in the post-apartheid period by the associated cultural group through their transnational circulations of Indian performers (Hansen, 2005). Although India is materialised in the lives of many Indian South Africans through performing arts programmes, Bollywood, food, clothing, festivals and increasingly heritage tourism, most South African Indians have little interest in claims that they belong to an Indian diaspora (Landy et al., 2004). Whilst diaspora associations continue to promote closer ties to India through traditional performances, rituals and festivals, for the majority of South African Indians, identification with India as a place of deep-rooted ancestral connection became obsolete over hundred and fifty years of indenture, permanent settlement and apartheid (Rastogi, 2008). The Indian government placed international sanctions on trade and domestic travel between 1963 and 1993, severing South African Indian's material flows and connections, and their spatial association with particular Indian locales (Landy et al., 2004). Furthermore, during the later stages of apartheid there was an attempt by some associations to shed Indianness in favour of a broader black struggle/liberation consciousness (Desai, 1996). The descendants of the original migrants have long considered South Africa their permanent home: in 1960, 95 per cent of the South African Indian population were born in South Africa (Ginwala, 1985: 3). Scholarship of contemporary Indian South African diaspora associational life has primarily focused on the role of classical Indian song and dance in the processes of ethnic boundary (re)making, drawing attention to contestations surrounding the authenticity, symbolism and identity politics associated with particular cultural forms (Kaarsholm, 2011; Vahed and Desai, 2010). Recent work has also demonstrated the complex positioning and articulation of Indian transnational life within the disciplinary tenets of faithed, racialised and gendered South African nationalist discourses (Radhakrishnan, 2005; Vahed, 2005). However, to focus on diaspora associational practices solely in terms of broader identity politics would be to miss the ways that diaspora associational practices take shape in relation to the heterogeneous elements and materialities of bodies, technologies and places that comprise them. In most cases, South African Indian associational life has been studied as a source of controversy, with scholarship primarily concerned with the impacts of associations' activities rather than with the physical spaces themselves per se. One exception is Radhakrishnan's (2003) account of how movements, embodiments and affects of diasporic performances attempts to, but also undermines, broader goals of resolving appropriate performances of Indian authenticity with the tenets and demands of multicultural citizenry in the Rainbow Nation. Radhakrishnan's study is instructive because it makes a good case for why there ought to be an affective reconceptualisation of cultural spaces and performance. If we are to properly understand the role and significance of musical performances for the elites concerned, then there needs to be a move towards an approach that appreciates the multifaceted constitution of their diasporic selves.","The present paper emerges out of a wider ethnography examining contemporary South African Indian transnational connections conducted in two periods of fieldwork in KwaZulu Natal Province, South Africa in 2004 and 2005. Although 26 participants (including academics, local officials, community activists and South African Indian newspaper and radio-station staff) contributed to the overall study, the analysis in this paper is based on fourteen interviews with the organisers of seven Indian South African Hindu associations involved in promoting closer ties between India and South Africa.2 In order to capture those with a strong desire to propagate ancestral Hindu cultural traditions, interviewed representatives were those who participated regularly in the outreach and social events promoted by the Indian Government's KZN-based branch of the Indian Council for Cultural Relations (ICCR). Interviewees represented the heterogeneity of South African Indian Hindu communities (Tamil, Gujarati, Telugu), reflecting the background of associations with strong connections to India. I also conducted twelve interviews with local government officials, Indian high commission staff and attendees at events. The fieldwork also involved extensive observational work, attending, for example, local events, celebrations and public meetings organised by diaspora associations. As part of my research, I interviewed several of the organisers and attendees before, during and following an event to capture their impressions. Interview and observation work was supplemented by visual and textual analysis of news articles, promotional material, reports and websites. The research took a narrative approach to the analysis of interview, observational and textual material in order to explore the motivations, discourses and technologies of diaspora associational practices. A narrative analysis aims to capture an individual's process of story-telling as they attempt to imbue their recollections with significances (Riessman, 1993). Narration is a process of story-telling at particular biographical and historical moments, capturing a person's interpretations within contexts and conventions that include not only the broader political and economic milieu but also the social contexts of the interview itself (Plattner and Bruner, 1984). In constructing narratives, Sandelowski (1991) argues, “events are selected and then given cohesion, meaning and direction; they are made to flow and are given a sense of linearity and even inevitability” (163). Using such an approach, analysis of interview and textual materials involved examining the ways respondents portrayed themselves and others, how individuals constructed past and future life events, taking note of how respondents explained their thoughts and feelings, and taking into account the social context of the interview itself.","After a meal at a nearby restaurant, Vinod3 (male, late 50s), his wife, daughter and I get into a car to leave to attend a classical Kathak performance at Kendra Hall in Durban by Shri Abhimanyu Lal. In the car, Vinod, who has been closely involved in promoting Indian classical song and dance since 1993, chattered excitedly about the roots of Kathak in Hindu scripture, pausing every now and then to tell me about the difficulties of getting performers from Gujarat to Natal because of apartheid-era restrictions. Vinod felt that the lifting of diplomatic restrictions had opened up a sense of opportunity. He argued that it “was a new experience for us to have India at our disposal … we had to take advantage of this”. As we got out of the car and entered the main hall, decorated with a red carpet, twinkling lights and strings of chrysanthemums, Ramesh (male, late 60s) greets us, pointing out the design of the dome and the pillars. The hall, he explained, had been built over two years by a team of Indian artisans, or karigars, and felt, he said, “like our own bit of India in South Africa”. Once we arrived, we joined the buzz of the crowd, who were milling around the entrance styled impeccably in saris, sherwanis and kurtas, so that Vinod and his wife could catch up with the latest gossip with their acquaintances before the evening's events began. The emotions, materialities, and visceral experiences that saturate musical and artistic performances are the focus of this section. In the narrative that follows, I trace organising actor's recollections of participating in, and creating, Indian classical musical events, connecting individual and shared experiences of the materialities and embodiments of musical forms to the production of organising actor's subjectivities. Ramesh, Vinod, and other key actors like him in Durban's Hindu associations, have been involved in formulating closer ties between India and South Africa since the resumption of formal diplomatic ties in 1993. These opportunities have emerged from the Government of India's diaspora outreach strategies, which both promote and regulate cultural flows. 48% of the South African Indian population live in the Ethekwini municipality, which encompasses the former Indian townships of Chatsworth and Phoenix, and around 80% live within the broader province of KwaZulu-Natal (Statistics South Africa, 2010). Under apartheid's Group Areas Act, Indians were forced to live in circumscribed parts of Natal (Maharaj, 1997), and whilst the lifting of apartheid restrictions and processes of desegregation have seen out-migration to other provinces, Durban remains a hub for classical Indian performing arts. In 1996, the Durban branch of the Indian Ministry of External Affairs' Indian Council for Cultural Relations (ICCR) was established, and has been central to the organisation of an intensive programme of India–South Africa cultural exchanges. In addition to such cultural programming, often organised in collaboration with local diaspora groups, the ICCR have organised numerous delegations and bilateral arts agreements. In 2010, the year of the arrival of the first indentured sugarcane labourers from India, a mini Pravasi Bharatiya Divas (Overseas Indian Day) was held in Durban with the aim of marking and building upon the city's historic links with India (Singh, 2011). Due in part to these efforts, circulations of classical Indian performers between India and South Africa have dramatically accelerated since 1993. Amongst migrant populations, music and dance may be used to recreate the culture of the past, but it is bound up with maintaining group identity through evoking connections and memories of a common homeland or place, allowing, notionally, a return (Baily and Collyer, 2006). Like many similar classical performances (Landy et al., 2004), the Kathak performance at Kendra Hall described above emerged out of a desire to bolster Indian vernacular identities. Vinod explained that he felt compelled by a sense of obligation to help promote a wider interest in Gujarat for those whom he felt were “losing touch with their roots”. He explained, “if we lose the language, we lose our culture and our community”. Yet, the vignette above also illustrates that the reproduction of collective identities are about much more than the performances themselves and the wider body of meaning they signify; they are saturated with a multidimensional spatial and temporal and array of objects, memories, experiences, events, materials, peoples and relationships that together coalesced into a wider performance environment. Many organising actors recounted how the multi-sensorial presences of classical arts performances provided a route to an embodied sense of wellbeing. One evening, I sat with Venita (female, 50s) over tea and sweets discussing the Kathak performance at Kendra Hall. Although such performances are related to the bolstering of vernacular identities, such spaces are also imbued with an experiential value. Most organising actors I met expected to feel content and happy in classical performance spaces. When speaking of her experiences, Venita for instance recounted that “I can just relax and enjoy the music … being dressed up”. In addition to its experiential value, many respondents felt that performance spaces were an important part of their social lives. The Indian community as an over-arching socio-legal construct was introduced and institutionalised during apartheid's system of separate development, during which period dense, tight-knit social networks developed that were based around intersections of religious, regional and class identities (Kuper, 1960). Contemporary transnational landscapes connecting South African Indians to different regions of India reproduce these social forms, even if the underpinning vernacular languages have disappeared (Vahed and Desai, 2010). Festivals, performances and events were where many organising actors expected to meet old friends and acquaintances based on these existing social networks, and share stories about family life and achievements. Venita described the Kathak performance as “one of the few times I can spend with my husband, meet up with our friends and catch up on the family gossip … I would say that is the main reason I went”. The type of social network reproduced depended on the type of performance attended. More regionally specific cultural programming was where organising actors would meet friends and acquaintances that shared linguistic commonalities. An organised visit by dancers from a famous Tamil Bharatnatyam school would provide opportunities for socialising with friends and acquaintances from within the Tamil community. Since a broader South Indian diasporic identity developed under the Indian cultural institutions of apartheid (Ganesh, 2010), many Telugu Indians would also attend. The social relations enacted in these spaces often also encompassed adult children, who might have been schooled in regional Indian performing arts by their parents. For instance, I met Saroji in 2005, along with his daughter Pritty, at a classical South Indian Bharatynam dance recital he had organised. Pritty (early 40s) had trained in the dance as a teenager. Whilst she left KwaZulu Natal to train as a doctor in Cape-Town four years ago, she often returned to Durban to attend Bharatnatyam performances and festivals with her parents. Asking why she attended, she said: “I like to see the dancers, their clothes, the makeup, the ways they move their hands and how they place their feet, just like I was trained”. In contrast, bigger festivals such as ‘Melodies of India’ (2004), amongst others, were designed to showcase a range of classical Indian performances from different regions of India, but might also include more contemporary music such as Bhangra, a hybrid cultural form that emerged from Punjabi diasporas in the United Kingdom but whose popularity has spread elsewhere in the diaspora. Actors from across the spectrum of South African Indian associations would be invited, even if their associations weren't specifically involved. These spaces also produced considerable overlaps between Hindu and Muslim communities, because of the political dynamics of apartheid and intersecting trajectories of migration (Kaarsholm, 2011). Muslim Gujarati's might attend a performance by Gujarati Kathak dancers, even though Kathak is ostensibly based on Hindu scriptures, because this was a cultural form promoted under apartheid and something they grew up with. Conversely, Vahed (2005) has shown how the Islamic festival of Muharram in South Africa, known pejoratively by the white authorities as ‘Coolie Christmas’, incorporates Hindu rituals and scriptural elements through broader constructions of a unified Indian identity as a base for political mobilisations. The social aspects of classical performances allow people to attach a range of significances and values to the performances beyond representations of a community, or a place. This is also what motivates Krish (male, 50s). For him, the physical minutiae of the performances are what matters more than what the performance symbolises, because it helps entangle him in feelings of mutuality with others watching the performances. Discussing a visit by some Indian Kuchipudi dancers, Krish said “you feel close to people when you go to these performances because we all appreciate the effort that has gone into doing things right”. The specific content of the performances allows collective recognition of cultural forms and finesse, and in doing so generates ideals of shared enjoyment and common purpose. Krish and Pritty are examples of some of the ways that relationships to the different physical and material properties of diaspora performances allow the reproduction of South African Indian social life, often over several generations and across differences. Organising actors developed performances not only to maintain a range of socialities and relationships, but also to take advantage of the burgeoning opportunities to be captivated by seeing performances live. In the spring of 2005 I attended a performance by Shobana Rao, a ghazal singer and her troupe from New Delhi. After the performance, I caught up with Hitesh (male, late 50s), a South African Indian who had organised the performance and invited me along. After the show had finished, people streamed out of the hall, discussing animatedly the performance they had just experienced. The music's stylistic tropes generated a buzz between audience members as they discussed the more populist, rather than purist, rendering achieved by Shobana and her musicians through the lightness of the ghazals' composition. They also pushed at individually felt experiences and sensations. Hitesh had “felt happy just listening to the individual notes of the Indian ghazals … watching someone play the tabla with skill and precision”. For Hitesh, and others, this particular ghazal performance was not merely a symbolic representation of India, but enabled engagements with the diverse, sonic presences that animated this particular performance. Whilst the ghazals staged here conserve ideas of Indian traditions, this is not a place-bound essentialisation of India. It a much more distributed manifestation that comprises experiences and sensations arising from relations to the objects, people, movements and the sonic properties of its physical (re) production and practice in specific sites. Affective experiences and relationships were also central to the production of the self as a knowledgeable cultural authority. Although the cultural politics of South African Indian diaspora associations allow actors to reproduce their elite positionings (Hansen, 2012), these also have emotional dimensions. During my fieldwork, Krish had organised a performance by vocalist and instrumentalist M. Balamuralikrishna because “I wanted everyone to have the same awe inspiring experience of seeing one of the great Carnatic playback singers in the same room”. In describing Balamuralikrishna's performance after the event, Krish said “I could tell that everyone enjoyed feeling connected to our spiritual roots. It made me proud to be able to provide that for them”. Accommodating a perceived desire amongst South African Indians for traditional experiences, performances reflected elites' desires to demonstrate to others that they could source appropriate cultural forms, but also that they were best able to guide the affectual basis of relationships to India. Nitin (male, late 70s) for instance stated that: [my] task is to bring India to the vast majority who won't be able to go, and even if they do go, they won't to be able to source Indian culture in true form … I feel it is my job to provide the right experiences for others At the same time that building transnational exchanges were a way of controlling affects felt, they also gave elites like Hitesh and Nitin a feeling of accomplishment in transforming their material practices of transnational exchange into a tangible, durable good that helped South African Indians achieve some degree of affectual knowledge. The value of cultural performances for Hitesh and Nitin lay not only in their own experiences of classical Indian performances, but also in their interpretations of the affects achieved in other people.","As the above section demonstrates, elites frame classical Indian performances around a complex, and ambiguous range of subjectivities. But classical performances are just one of a set of interconnecting sites in which Indianness is materialised in South Africa (Landy et al., 2004). These also include online worlds, literature, comedy, clothing and food consumption practices, and leisure spaces such as Bollywood films and entertainment shows. Since the end of apartheid, contemporary Bollywood films and entertainment shows (where actors, musicians and singers appear together to recreate a film, scene or dialogue), particularly those that are styled to appeal to the North American Indian diaspora,4 have acquired increasing salience as a public space for transnational materialisations of India. With the growth of modern South African Indian diasporic dispositions that encompass sites in India, North America and afar, Bollywood, as Hansen (2005) has noted, also threatens the integrity of the South African Indian experience provided by diaspora associations. In this section I examine the porosities, circulations and overlaps between classical performance spaces and Bollywood as a means of exploring further the entanglements between organising actor's diasporic subjectivities and material practices. There are important overlaps between classical Indian song and dance and the contemporary Bollywood film industry. Playback singers have been integrated into and are essential to popular Indian cinema since the 1930s, with recordings based on classical Sanskrit drama, folk theatre (Morcom, 2007). With playback singers, such as M. Balamuralikrishna discussed above, achieving the same prominence as actors, their repertoires are a central part of musical culture throughout the Indian diaspora, lending their voices both to film compositions, classical recordings and live performances (Bhattacharjya, 2009). Many of the elites interviewed grew up in the 1960s and became familiar with key playback singers of the era, and often invited them to perform in South Africa both during and after the lifting of diplomatic restrictions. Anxieties about the materialisations of contemporary Bollywood in South Africa have centred on whether clothing, attitudes and consumption practices appropriately represent India (Radhakrishnan, 2005). Contemporary cinematic themes and narrative plotlines have essentially remained unchanged, since they reproduce the same clichés and tropes of the village dramas of early Indian cinema, which itself ultimately derives from classical Sanskrit drama and folk theatre (Morcom, 2007). Respondents found going to watch a Bollywood movie a pleasurable, and family friendly, leisure practice. But contemporary Bollywood encompasses a changed set of capitalist and consumer aesthetics, in part to appeal to younger diasporic audiences in North America and Europe (Therwath, 2010). As observed elsewhere in the Indian diaspora (Mohammad, 2007), these changes have reconfigured the relationships between the South African Indian diasporic self and place (Ebrahim, 2008). Some elites found this troubled traditional norms. For instance, Akesh (male, 60s) says that: Bollywood has really taken over from the appreciation of traditional culture. But Bollywood doesn't show real Indian culture. The dances and the music look Indian but it does not feel Indian because now it is too Westernised. The girls don't wear saris anymore, and often it is not even filmed in India anymore The minutiae of the modern Bollywood aesthetic discussed here are contrasted with perceived ‘real’ Indian culture, the kinds provided through classical performances. Perceptions of Bollywood are also mediated by generation and class, with Indian cinema most often associated with younger people and the Indian working classes, whilst classical performances are often the preserve of wealthy elites (Hansen, 2005). But these tensions are also underpinned by feelings of loss akin to Marilyn Ivy's (1995) view of nostalgic desire as encompassing a range of emotions including anxiety around the erosion of felt experiences (see also Boym, 2001). Most elites however, were ambivalent about Bollywood. Some respondents spoke of the ways that Bollywood had generated economic benefits for South Africa's film industry and supported closer relations between Indians and Africans (see also Ebrahim, 2008). For Pritty, growing interest from non-Indians and the mainstreaming of Indian cinema in broader South African cultural life through its incorporation into public television programming had allowed her to not only craft, but rewrite a new sense of self: 10 years ago I had to be apologetic for being Indian. There has been a fundamental change thanks to Bollywood … When in High School I felt ashamed of being Indian and I didn't want to talk about India. It's gotten to the stage where I'm proud to be Indian. It fills me with pride that's where I'm from, there's my roots. At one point I was so ashamed about my name, really embarrassed Often, those elites who criticised the perceived dominance of contemporary South African Indian life by Bollywood were at once intrigued and drawn to the affects that its styles, registers and fashions produced. One afternoon I travelled to Sally's (female, late 60s) house to interview her over tea and sweets. As we talked, she rummaged around in a drawer and pulls out a CD of playback recordings from the movie Lagaan, given to her by her grandson. Notes floated around us as she described the pleasure she felt in keeping up with new releases so that she could talk with her grandchildren about the latest heartthrobs, and a sense of enjoyment in hearing them pick up and incorporate Tamil words learnt from films into everyday speech. To varying degrees, elites were beginning to capitalise on the new potentialities provided by Bollywood. Both the ICCR and KwaZulu-Natal Ministry for Arts, Culture and Tourism encourage Indian diaspora associations to incorporate contemporary Bollywood styles and aesthetics into their repertoire of cultural programming. Adapting their classical programming by incorporating modern dances and aesthetic styles, often organised in conjunction with local dance schools and contemporary choreographers, were part of these plans. Young Indian South Africans from Durban dance schools are one source of performers. Forty years ago, dance choreography was modelled precisely on Sanskrit traditions. According to Radhakrishnan (2003), as dance schools have developed and expanded particularly in the post-apartheid era, choreographers have begun to incorporate more modern Indian aesthetics, developing new hybrid styles. It is also a relational process, with teachers innovating and adapting classical and modern Indian styles to the particularities of the South African setting, by for instance incorporating repertoires from traditional Zulu folk dances. In fascinating ethnographic detail Radhakrishnan describes how these new fusion choreographies pose a radical challenge to the boundedness of classical Indian cultural traditions. Whilst respondents acknowledged that this variability had increased interest from younger generations, nevertheless the allure of their correct embodiments remained. For example, Athul stated that new dance styles are “not really India … But I am happy that some of the younger generation are at least getting a feel for some of the Indian traditions”. Amongst some respondents, I heard approval of Bollywood since it allowed some hope that it could act as a precursor for a broader interest in India and an eventual recovery of Indian languages. For example, when I met him in 2004, Dev's organisation (male, late 50s) was campaigning to pressure Indian film distributors to screen Tamil language films in KwaZulu-Natal. He talked about the Hindi romances he had watched in the 1960s with affection, describing the grandeur of visiting the Shah Jehan theatre in Durban's Indian Quarter and the ritual and performative qualities of disparaging, gently, the films; clearly they were an important part of his youth. But he also felt that for the current and younger generations, Tamil films had future potential to increase interest in local festivals and events such as Tamil New Year, or even encourage some form of physical visit to India, because of they way films could deepen the knowledge of South Indian traditions and languages. Like Athul and Dev above, I heard other actors speak of their hopes that interest in Bollywood would provoke younger people to one-day visit India to discover their ancestral roots and connections. Some spoke of their hopes that younger South African Indians would eventually embrace more classical dance forms, or begin learning an Indian language, both of which had seen declining uptake. For instance, Krish said that it was only through “knowing the true India, the kind we bring through our work” that the younger generations could really recover their roots. Whilst Bollywood, for some, posed a threat to the long-term appreciation of classical Indian forms, for both Dev and Krish, generating spatial and temporal connections between different Indian diasporic sites and practices was a concrete means through which diaspora associations could anticipate and realise future South African Indian diasporic selfhood.","Scholars of diaspora associations more broadly, and South African Indian cultural organisations specifically, have mainly focused on the political and economic outcomes of their activities, on negotiations over underpinning categorical identity politics, and on the complex landscapes of home that are produced, resisted and contested. Less work examines the affective ‘doings’ of associational life and the forces and matter that comprise it. In this paper, I have explored organising actor's relationships and interactions with the constitutive elements of classical Indian performance spaces, as one way of exploring the diverse meanings and significances that can be drawn from them by the actors participating and assembling those spaces. The paper has shown the ways that a range of subjectivities emerges from and is negotiated through the material constitutions of classical Indian performances, as they form points of connection and overlap with a range of different sites through which India is experienced and negotiated. I conclude by showing how such an approach can enhance scholarly understanding of the work of diaspora associations as well as pointing to future directions for research. First, by examining classical Indian performances through the material practices and affective forces of individual actors, the paper has deepened understandings of the role of agency in constituting sites of diaspora associational life. Whereas most studies of diaspora associations have focused on their wider calculable economic and political agendas, the agency discussed here is relational and affectual, one that produces and reproduces a wide range of subject positions. A discussion of their different points of connection, discrepancies and tensions has shown how, in the process of negotiation, organising actors can anticipate, realise and redefine past, current and future forms of South African Indian diasporic subjectivity. Second, the paper has sought to illustrate an alternative way in which diaspora associations' role in producing social differences (and similarities) might be explored in scholarship. This study of diaspora associations began not from understandings of India as a pre-existing physical location, but as a spatial array of interconnected sites, practices and non-human and human forces. This is not to dismiss the categorical orderings, politics and social differences that also underpin their practices, but to suggest that heterogeneous materials, bodies, objects and multidimensional presences come to compose their saliences and longer-term durabilities. In doing so, this suggests that the social differences produced through associational practices are full of instabilities, tensions and gaps, leaving room for critical refusal and transformation."],["Trait emotional intelligence (EI) was measured and self-estimated in a UK sample of 128 managers (52.3% female), recruited at a professional services firm. Participants' measured scores were compared to standardization sample data and gender differences in measured and estimated scores, as well as in estimation bias and accuracy were examined. As hypothesized, managers' global trait EI scores were significantly higher than those of the normative sample of the measure used, although the scores of female participants were largely responsible for this difference. Gender-specific hypotheses were confirmed for measured scores (differences only hypothesized at the factor level) and estimation accuracy (males estimating their trait EI more accurately), but not for estimated scores (female participants had higher estimates, but the opposite was hypothesized). Further, female managers showed signs of estimation bias. © 2014 Elsevier Ltd. -------------------------------------------------------------------------------- MEASURED AND SELF-ESTIMATED TRAIT EMOTIONAL INTELLIGENCE IN A UK SAMPLE OF MANAGERS -------------------------------------------------------------------------------- Management of human capital has been portrayed as one of the major settings for the relevance and application of emotional intelligence (EI). In part, the importance which EI has been ascribed in the managerial world is linked to its marketing potential within this context; the construct is readily sellable in the form of assessments, training programs, and interventions. One the other hand, the occupational demands associated with various types of management draw on the specific characteristics subsumed by the prevailing EI models and measures (e.g., Bar-On, 1997; Petrides, 2009a). Emotion-related qualities seem to be fundamental to professional success and adjustment within this diverse capacity, suggesting that managers may constitute a high EI population. Although there has been a surge of studies on managerial samples or in managerial contexts, much of this research has treated “EI” as a general concept, rather than considering the two more specific constructs tapped by various measures. Since the construct’s inception and popularization (Goleman, 1995), the field has gradually diverged into two streams of research, focusing on two complementary dimensions termed ability EI and trait EI, respectively. Ability EI concerns emotional-related abilities measured through maximum-performance tasks, whereas trait EI refers to the emotion-related personality dimension assessed through typical- performance measures. It has been argued that any typical-performance measure of EI is most appropriately interpreted through the trait EI lens, independent of the underlying model (Petrides & Furnham, 2001). This assertion and the distinctiveness of the two constructs is supported by non-significant to modest correlations between typical- and maximum-performance EI measures and moderate to strong correlations between measures based on the same method (Van Rooy, Viswesvaran, & Pluta, 2005). The operationalization-based split into two relatively distinct constructs, which has implications for the interpretation of findings gathered with a given measure, needs to be considered in research with special-interest populations, such as managers. One cannot generalize from one construct (i.e., trait or ability EI) and its operational vehicles to the other, as divergent findings can be expected from the two (Petrides & Furnham, 2001). The focus of the present study is on managers’ trait EI, and a concise review of studies assessing trait EI in managerial samples is provided next.","We retrieved 10 studies in which managers’ EI was assessed with typical-performance measures and, thus, representative of trait EI.1 The samples used in these studies varied considerably in geographic locations and ethnicities (e.g., China, UK, Australia, and Israel), occupational sectors (e.g., CFOs, restaurant franchises, public services, retailers, construction industry) and managerial levels. Seven of the ten studies employed workplace-oriented EI scales (Angelidis & Ibrahim, 2012; Gardner & Stough, 2002; Sy, Tram, & O’Hara, 2006). Unfortunately, these types of EI measures are unlikely to reveal much about managers’ trait EI (relative to the general population), since they were standardized on samples comprising managers, leaders, or people in similar roles. Therefore, we restrict our focus on the results gathered with general-population scales. Different general EI scales were used in three studies. The Trait Meta-Mood Sale (TMMS; Salovey, Mayer, Goldman, Turvey, & Palfai, 1995) was administered to an Australian female- only sample of managers from various industries (Downey, Papageorgiou, & Stough, 2006). Sample scale means were 3.94 for Attention (SD = 0.57), 4.22 for Clarity (SD = 0.57), and 4.23 (SD = 0.58) for Repair. In comparison, a sample of undergraduate students had scale means of 4.10 (SD = 0.52) for Attention, 3.27 (SD = 0.70) for Clarity, and 3.59 (SD = 0.90) for Repair (Salovey, Stroud, Woolery, & Epel, 2002). The Bar-On (1997) Emotional Quotient Inventory was administered to a sample of 191 middle managers (line managers; 69% male) working for a major UK retailer (Slaski & Cartwright, 2002). The overall EI sample mean of 94.4 (SD = 12.5) was lower than the normative sample mean of 100. Moreover, Schutte et al. (1998) Assessing Emotions Scale was completed by a sample of 98 senior managers (89% male) employed as CFOs in local government authorities in Israel (Carmeli, 2003). The sample mean was 3.71 (SD = 0.37), which was above the normative sample means for women (M = 3.45, SD = 0.46) and very similar to that of men (M = 3.78, SD = .50). The number of relevant studies is too sparse and their findings insufficiently consistent to suggest that managers are particularly high in trait EI. Importantly, the samples used in these studies varied widely in occupational sectors and managerial levels, making it difficult to tease apart the effects of management and work-domain. Another limitation concerns the use of different measures varying in subscales, with one (the Trait Meta-Mood Scale) comprising three weakly interrelated factors. A benchmark measure of trait EI, the Trait Emotional Intelligence Questionnaire, was used in one managerial context (Mikolajczak, Balon, Ruosi, & Kotsou, 2012), but no sample means were reported in this study. Furthermore, studies on managerial samples have tended to neglect the role of gender, despite its importance in EI research (e.g., Siegling, Saklofske, Vesely, & Nordstokke, 2012). Another pertinent factor not previously considered is managers’ holistic self-evaluation of their emotional adjustment. Self-perceptions are important for several reasons and have been studied for some time, particularly in the context of IQ and performance. It is conceivable that they have a profound influence on the kind of tasks people engage in or avoid, and on the kind of careers pursued. Further, positive self- perceptions are linked to mental health, in contrast to negative self-evaluations, which are linked to negative affect and depression (Petrides & Furnham, 2000). Although previous research has examined EI self-perceptions in university students, with a particular focus on gender differences (Petrides & Furnham, 2000; Petrides, Furnham, & Martin, 2004), self- perceptions of managers may differ in myriad ways from university samples in terms of perception accuracy, bias, and gender differences. Present study This study examined the trait EI profiles of a general managerial sample comprising of managers from different levels and not tied to any specific type of service. Participants’ trait EI scores were examined for gender differences and compared to normative sample data. Departing from the bulk of management-related studies, in which trait EI was assessed with workplace-oriented scales, this study used the Trait Emotional Intelligence Questionnaire (TEIQue), a scale designed to measure the construct comprehensively in the general population. We also examined managers’ overall self-estimates of trait EI, focusing on gender differences, estimation bias, and estimation accuracy. These self-perceptions were referenced against the TEIQue model to facilitate direct comparison with the measured trait EI scores. The following hypotheses were tested: H1: Participants’ measured trait EI scores will be higher than those of the normative sample of the TEIQue. Although our review of the literature did not yield conclusive evidence, this hypothesis is based on the particular importance of emotional resilience and socioemotional functioning in the managerial world. The argument is that emotionally resilient people are more likely to be selected for, or to advance to managerial positions. H2a: There will be no gender difference in managers’ global trait EI scores. Although the normative sample mean is significantly higher for males (Petrides, 2009b), gender differences were not apparent in other samples (e.g., Siegling et al., 2012) and female managers may be particularly well adjusted compared to women in the general population. However, as has been quite reliably found, we also hypothesized, H2b: Male managers will score higher on the Self-Control factor than female managers, who will be higher on the Emotionality factor. H3: Male managers will have significantly higher estimated global trait EI scores than female managers when controlling for measured scores, consistent with previous findings from participants recruited at British universities (Petrides & Furnham, 2000). This hypothesis also reflects self-enhancing and self-derogatory biases in men and women, respectively, which have been demonstrated for self- evaluations more generally. H4: Male managers will have more accurate estimates than female managers, also based on previous findings in British university students (Petrides & Furnham, 2000).","We invited 339 managers from senior, middle, and junior levels at a large professional services firm to participate in this study. Of this group, 128 (37.8%) managers with a mean age of 38.0 years (SD = 7.5, age range: 26–59 years) participated (three participants [2 male, 1 female] did not indicate their age). The gender split amongst the participants was almost equal (52.3% female), but the representation of the three managerial levels was uneven; the majority came from middle management (n = 79, 50.6% female), whereas similar sample proportions were senior (n = 27, 40.7% female) and junior managers (n = 22, 72.7% female). The mean ages of male and female participants were 39.1 years (SD = 7.9) and 36.9 years (SD = 7.1), respectively. The average length of time worked at the firm was 6.2 years (SD = 6.0) for the overall sample, 6.7 years (SD = 6.9) for male managers, and 5.8 years (SD = 5.0) for female managers. The majority of respondents (78.1%) indicated their ethnic background as Caucasian, others as Black, Asian, and Indian/Pakistani. Educational backgrounds in terms of the highest level of education attained varied considerably: 2.5% GCSEs/O-levels, 15.6% A-levels or similar, 53.9% BA/BSc or similar, 21.1% MA/MSc or similar, and 2.3% MBA (six participants did not indicate their highest level of education). After providing demographic and background information, trait EI was assessed and self-estimated. The study was conducted anonymously online. Trait EI The short form of the TEIQue (Petrides, 2009a) was sufficient for the purpose of this study. It contains 30 items from the full form (two items represent each of 15 facets) and can be used to measure global trait EI and the four factors derived from the full form: Emotionality, Self-Control, Sociability, and Well-Being. Respondents complete the items on a 7-point Likert scale, ranging from 1 (completely disagree) to 7 (completely agree). The internal consistencies in the present study were acceptable and consistent with those reported for the standardization sample (Petrides, 2009a). Specifically, Cronbach’s alphas were .91 for global trait EI, .85 for Well-Being, .70 for Self-Control, .73 for Emotionality, and .78 for Sociability Estimated trait EI Participants gave overall self-estimates for global trait EI and each of the four TEIQue factors. Definitions of the four factors, as shown in Petrides (2009b), were presented to the participants. As a description of global trait EI, participants were shown a visual illustration of the trait EI model integrating its four factors and 15 facets (see Fig. 1). Upon referencing these descriptions, they were asked to give their estimate for each of the factors and global trait EI on a scale ranging from 1 (extremely low) to 21 (extremely high). This range was considered adequate to yield sufficient variability in responses and precision in estimating one’s trait EI. Prior to statistical analysis, these estimates were divided by three to make them directly comparable to the measured TEIQue scores","Relationships amongst age, gender, and managerial level, as well as correlations of age with measured and estimated trait EI were examined to identify potential confounds. Measured trait EI scores were compared to the normative data (both general and gender- specific), as reported in Petrides (2009a), using one-sample t tests. Participants’ trait EI profiles and the role of gender in trait EI scores were examined using ANCOVA, controlling for any identified confounds. When examining gender differences in estimated scores, measured scores as well as any confounding variables were controlled. ANCOVAs, again controlling for any confounds, were executed to examine any bias towards over- or under-estimation of trait EI scores. Specifically, measured and estimated trait EI scores were compared for the overall sample and for each gender. Correlations between estimated and measured scores, also computed separately for the overall sample and each gender, were examined as an indicator of estimation accuracy. Accuracy was then compared between female and male managers using Steiger’s Z statistic. Preliminary analyses ~~~~~~~~~~~~~~~~~~~~ With missing responses highlighted to the participants electronically (without forcing responses), there were no missing data points. The mean ages of male and female participants were similar, t(123) = 1.68, p = .10, whereas ages increased significantly across managerial levels, F(2, 122) = 15.00, p < .0001. Participant age also correlated with measured global trait EI, r(125) = .21, p = .02, and its Self-Control factor, r(125) = .23, p = .009, but with none of the self-estimates (p > .05). Thus, we aimed to control for age in the main analyses. A qui-square test examining the relationship between managerial level and gender did not reach significance, χ2(2, N = 128) = 5.21, p = .07, indicating that managerial level would be an unlikely confound. Trait EI profiles ~~~~~~~~~~~~~~~~~ Table 1 shows descriptive statistics for measured and estimated trait EI scores for the overall sample and each gender. One-sample t tests showed that the overall sample had significantly higher scores than the standardization sample on global trait EI, t(127) = 4.06, p < .0001, Self-Control, t(127) = 3.59, p < .001, and Well-Being, t(127) = 3.20, p < .01. Concerning gender-specific norms, there were no significant discrepancies between male managers’ measured trait EI scores and those of the standardization sample. However, female managers’ scores were significantly higher than the standardization sample scores on global trait EI, t(66) = 4.66, p < .0001, and three of the four factors: Well-Being, t(66) = 3.93, p < .001, Self-Control, t(66) = 4.30, p < .0001, and Emotionality, t(66) = 2.97, p = .004. A 5 (trait EI factor) × 2 (gender) mixed-design ANCOVA on measured trait EI scores, controlling for age revealed a significant interaction, F(2.91, 354.70) = 8.11, p < .0001, partial η2 = .06. Mauchly’s test for the within-subjects factor was significant in this analysis, χ2(9) = 384.86, p < .0001, and, hence, the degrees of freedom were adjusted using Greenhouse-Geisser estimates of sphericity (ε = 0.73). While the factor scores appear to be relatively uniform for the overall sample, F(2.91, 354.70) = 2.16, p = .09, the significant interaction is indicative of a gender difference in trait EI profiles. Therefore, follow-up comparisons of the factor scores both within and between the genders were conducted. One-way within-design ANCOVAs controlling for age revealed a main effect for female managers’ measured trait EI scores, F(2.91, 190.08) = 3.93, p < .01, partial η2 = .06. As Mauchly’s test of sphericity was also significant in this analysis, χ2(9) = 184.66, p < .0001, the degrees of freedom were again adjusted using Greenhouse-Geisser estimates (ε = 0.74). Pairwise comparisons adjusted using Sidak’s correction revealed significant differences between most of female managers’ factor scores (p < .05), except between Well-Being and Emotionality. In contrast, there were no significant differences amongst the measured trait EI scores of male managers. One-way between-design ANCOVAs controlling for age only revealed a significant gender difference on the Emotionality factor, F(1, 122) = 11.31, p = .001, partial η2 = .08, indicating that female managers scored significantly higher on this factor. Estimation bias ~~~~~~~~~~~~~~~ Two (score type: measured, estimated) × two (gender) mixed-design ANCOVAs controlling for age were executed to compare each measured score to its corresponding estimate and examine the role of gender in self-estimates. The analyses revealed no significant main effects for score type on any factor. However, there were significant interactions of score type and gender on global trait EI, F(1, 122) = 11.04, p < .01, partial η2 = .08, Well-Being, F(1, 122) = 5.06, p = .03, partial η2 = .04, Emotionality, F(1, 122) = 5.92, p = .02, partial η2 = .05, and Sociability, F(1, 122) = 10.41, p < .01, partial η2 = .08. After applying Bonferroni’s correction for multiple comparisons, the interactions on global trait EI and Sociability remained significant. The previous analyses revealed only a gender difference on the Emotionality for measured scores. Hence, to probe the interactions, we first examined if there were gender differences on any of the estimates for global trait EI or Sociability. One-way between-design ANCOVAs controlling for age and the corresponding measured trait EI score revealed significant differences on estimates of both global trait EI, F(1, 121) = 13.06, p < .001, partial η2 = .10, and Sociability, F(1, 121) = 10.01, p < .01, partial η2 = .08. Thus, female participants had significantly higher estimates on global trait EI and on the Sociability factor. To probe these interactions further, we examined if any of the estimated global trait EI or Sociability scores of each gender differed significantly from the corresponding measured scores. One- way within-design ANCOVAs, controlling for age, only revealed a significant difference between female managers’ estimated and measured global trait EI scores, indicative of an over-estimation bias, F(1, 64) = 4.37, p = .04, partial η2 = .06. However, this difference did not hold up following adjustment for multiple comparisons. Estimation accuracy ~~~~~~~~~~~~~~~~~~~ Table 1 also shows the correlations between participants’ measured and estimated trait EI scores. The correlations were consistently significant and within a moderate to strong range. However, the magnitude of correlations is indicative of systematic differences in estimation accuracy across factors (as each pair of scores is different with no score used in more than a single pair, it seems inappropriate to compare them statistically). The strength of estimation accuracies across factors (from strongest to weakest) were in the following order: Well-Being, Self-Control, Emotionality, and Sociability. A pattern that emerges from the gender-specific correlations is that men’s measured and estimated trait EI scores were consistently more highly associated than those of women, despite the slightly smaller number of male participants. However, the only significant gender difference in these correlations was on global trait EI, Z = 2.63, p < .01. The difference in associations on the Sociability factor was also significant, Z = 1.70, p < .05, but it did not hold up after applying Bonferroni’s correction. Thus, male managers estimated their global trait EI more accurately than female managers.","The results support H1, which was based on the notion that managers constitute a high trait EI population. Conceptually, trait EI is particularly relevant to the occupational demands shared by managers from all kinds of backgrounds (e.g., demonstrating composure in high-stress periods, dealing effectively with employee turmoil, being responsive to employees’ needs). The overall sample mean (for global trait EI and two factors) was above the normative average, but gender-focused analyses showed that only female managers were above the gender-specific normative data for either global trait EI or particular factors. While the differences for male managers may well be significant in larger samples, the fact that it they were more pronounced, and only significant for female managers may hold key implications, subject to consistent replication in further research. Consistent with H2a, male and female managers did not differ on global trait EI. Although the standardization sample mean is significantly higher for men, this finding supports our reasoning that women who are high in trait EI are more likely to advance to, and be considered for managerial positions. The fact that only female managers were above the gender-specific standardization-sample mean for global trait EI and three factors lends further support to this idea. H2b was partially supported, since female managers were higher on the Emotionality factor than male managers, as previously reported for university samples (Petrides, 2009a; Siegling et al., 2012). Previous research within the university population has also shown that men tend to score higher on the Self-Control factor, but this difference did not replicate in our managerial sample. Relative to previous findings, it appears that the Self-Control factor, in particular, contributed to the non-significant gender difference at the global construct level. Contrary to H3, which derived from previous findings on university students and more general gender differences in self-perceptions (see Petrides & Furnham, 2000, for a discussion), female managers had significantly higher overall estimates than male managers. Yet, the results are consistent with the findings from another study, in which measured scores were unadjusted (Petrides et al., 2004)—in comparing the results of these studies, including ours, it is also important to consider differences in the measurement of measured and self-estimated scores. Follow-up analyses revealed a trend towards overestimation on the part of female managers, whereas male managers’ self-estimates were aligned with their measured scores. Thus, where trait EI is concerned, our results are indicative of a self-enhancing bias in female managers and a lack of bias in male managers. As we had no hypotheses for any gender-related estimation biases, however, this particular result, which strictly speaking was not significant after adjusting for multiple comparisons, needs to be replicated in comparable samples. Our last hypothesis (H4), which concerned gender differences in estimation accuracy, was supported by the data. Male managers’ estimates of global trait EI were more accurate than those of their female counterparts, consistent with previous research in university students (Petrides & Furnham, 2000). Our results speak to the external validity of this gender difference, which was replicated in a special-interest population and by means of different measures of both measured and estimated trait EI than those used in Petrides and Furnham’s (2000) study. The higher correlations in the present study are presumably an effect of deriving estimates with reference to the same model as the one underlying the measured scores. The findings surrounding participants’ measured trait EI have potential key implications for understanding the role of this personality dimension in management. Most generally, they suggest that emotion-related personality traits play a central role in the selection and advancement of managers. They also suggest that, of those who pursue management-related careers, trait EI may be particularly instrumental for women, relative to the female standardization sample. That said, we are neither in a position nor willing to argue that trait EI is more important for women to function in managerial contexts, as is indicated by the non-significant gender difference in global trait EI. Rather, high trait EI women are more likely to end up in management than women with average trait EI levels, whereas this does not seem to be the case for men. Holistic self-perceptions are important in that they influence the kind of tasks people prefer, the career choices they make, and how far they are willing to push themselves. It is possible that the threshold of emotion-related self-evaluations necessary for seeking managerial positions is particularly high for women; only women who are above average in their emotional self-perceptions may pursue managerial careers. It has also been noted that positive self-perceptions are conducive to mental health, and negative self-evaluations to psychological distress (Petrides & Furnham, 2000). Inflated self-perceptions may be adaptive to the extent that they do not mislead people into tasks, projects, or even careers that exceed their actual capacities considerably. To that extent, high self-perceptions can help people perceive the demands of their occupation as less threatening or stressful. As female managers had higher self-estimates than male managers, despite having similar measured trait EI scores, they may approach various situations more confidently and be somewhat less vulnerable to job-induced psychological problems, such as burnout. It is important to acknowledge that other factors cannot be ruled out as explanations for the above-average trait EI scores found in this managerial sample. While the ethnic compositions were similar and both samples were based in the UK, the average age of the standardization sample is about eight years younger and it is not representative of the general workforce. Nevertheless, a different article in this issued showed that trait EI distinguished between leaders and non-leaders employed by the same company, even after controlling for age and other control variables (Siegling, Nielsen, & Petrides, in press). A second limitation to be addressed in future research is that both measured and estimated trait EI were based on self-report. Consequently, some of the variance in both variables was likely influenced by participants’ self-perceptions, suggesting that our results provide an overstatement of estimation accuracy. A way to circumvent this problem is to measure trait EI with one of the available 360 forms (i.e., through peer or close-other ratings)."],["This study aimed to understand more fully some of the factors that influence decisions as related to air defence in a naval vessel's operation room. The study considered the impact of decision criticality (DC) and task load (TL) on measures of accuracy, confidence, and within-subjects confidence-accuracy (W-S C-A; a measure of metacognition). Personality constructs, workload, and situational awareness were also assessed. Participants were allocated to either a high, moderate, or low TL condition. Each took part in a computer-generated simulated air defence scenario where they were required to make a range of decisions and provide a corresponding confidence rating for each decision taken. Results showed that low DC increased confidence in decisions and high DC increased decision accuracy. Thus, DC significantly impacts decision confidence and decision accuracy. In addition, those less tolerant of ambiguity were less accurate in their decision-making. Future studies should take account of these factors. --------------------------------------------------------------------------------","General Audience Summary Air defence decision-making is often conducted in a complex and uncertain environment. It is therefore important that the individuals faced with this task are able to make accurate and confident decisions under varying degrees of stress and criticality (i.e., the consequence associated with a decision). The purpose of this study was to examine external factors (e.g., task duration/stress) and internal factors (e.g., personality constructs) that may impact air defence operator's decision-making abilities. In this study a measure of within-subjects confidence-accuracy was used. This measure considers the relationship between decision confidence and decision accuracy by assessing individual awareness of the accuracy of decisions made. For the task, a realistic set of scenarios, which varied in task difficulty, were designed with subject matter experts.","were required to make a range of decisions which varied in criticality and then rate how confident they were that they had made the best decision given the situation. The results demonstrated that the criticality of the decision impacted both decision accuracy and confidence. Low decision criticality increased confidence in decisions and high decision criticality increased decision accuracy. The implications of this research include an increased understanding of the understanding of decision criticality on decision-making in critical environments. The introduction of a novel method which has potential application in terms of informing the selection and in the training of personnel who are required to make accurate and confident decisions under conditions of uncertainty and stress is also highlighted. It is important to note that these inferences are based on findings from a novice sample and that non-trained staff are unlikely to make decisions in critical environments. Participants ~~~~~~~~~~~~ Sixty participants were recruited through opportunity sampling from the University of Liverpool. The participants consisted of 30 females and 30 males with a mean age of 26 years (SD = 3.96). None of the participants had any prior experience in naval warfare operations as the study was initially interested in the OR role and novice capacity to the task. The sample size was decided upon by design, power, and previous studies using G*Power software with an effect size of 0.8 and significance level of .05 (Faul, Erdfelder, Lang, & Buchner, 2007). The study received approval from the University of Liverpool's Institute of Psychology Health and Society Ethics Committee, and a favourable opinion from the Ministry of Defence Research Ethics Committee. Design ~~~~~~ A mixed measures quasi-experimental design was employed. Independent variables (IVs) were task load and decision criticality. As such, a 3 (Task Load: low, moderate, high) × 3 (Decision Criticality: low, medium, high) mixed analysis of variance (ANOVA) was conducted, with repeated measures on the last factor. The dependent variables (DVs) were confidence, accuracy, W-S C-A, personality constructs (openness to experience, conscientiousness, extraversion, agreeableness, and neuroticism; NEO-PI-R; Costa & McCrae, 1992), decision tendencies (i.e., tolerance of ambiguity; Budner, 1961; and decision style; Roets & Van Hiel, 2007), subjective mental workload (NASA TLX; Hart & Staveland, 1988), and situational awareness (SART; Taylor, 1990). Decision logs To ensure as high a level of ecological validity as possible in a quasi- experimental design, an air defence scenario was created with the guidance and assistance of subject matter experts (SMEs). The use of SMEs to assist in the experimental design is highly beneficial as SMEs are able to provide unique insight into the appropriate and relevant situations that are likely to be met and applied in the study context. Four SMEs with extensive knowledge of naval warfare and many years of experience in both Air Warfare Officer (AWO) and Principal Warfare Officer (PWO) roles were used to acquire the domain specific knowledge needed to provide the optimum ecologically valid options for decision making. The scenario uses a realistic set of events within a peace enforcement (PE) operation. A series of events and associated event decision logs were also created and agreed upon by SMEs. The event decision logs specify three decision options of reasonable equivalence for each event presented to the operator. SMEs agreed upon one option per decision made as the optimal/best decision option given the current situation. Computer scenario The visual display used as the stimulus for the experiment was created using Virtual Avionics Prototyping Software (VAPS XT). The screen depicted a quasi- realistic radar screen which included an airlane, a no-fly zone (NFZ), a coastline, and a border. A textbox to display additional information to assist decision-making and a timer which counted down from 20 s at each decision event were also included (see Figure 1). The algorithms used to animate the visual display symbols were created using Matlab/Simulink. The symbology used is as specified by APP-6c (NATO, 2008). Microsoft Movie Maker was used to edit the video (e.g., to apply timers). The SMEs verified the display as sufficiently realistic. Questionnaires The Situation Awareness Rating Technique (SART; Taylor, 1990) was used to measure SA. To measure WL, the NASA-TLX (Hart & Staveland, 1988) was utilised. NASA-TLX is a subjective workload assessment tool. Personality was assessed by the NEO-PI-R (Costa & McCrae, 1992).","Participants were randomly allocated to a high, moderate, or low TL condition. Participants first completed participant demographic forms which collected data on age, gender, and occupation. Participants were also asked to complete paper-based questionnaires to gauge the relevance of a number of measures across groups (e.g., general personality constructs, thinking and reasoning) where they may be relevant to particular questions. Following this, participants were given the task booklet to read. The task booklet provided participants with information needed to assist them in the decision- making task, including air defence terminology and symbols. Once they had read the booklet, participants undertook a practice trial. The practice trial involved a series of decision events which allowed participants to familiarise themselves with the task and the procedure. The duration of the trial was kept limited so as not to fatigue the participants before the experimental task (Barnes-Yallowley; personal communication, 2015). The questionnaire booklet presented three separate decision options based on the events of the scenario. One choice was required to be selected by placing a tick by the option they believed to be the “best option given the current situation.” Participants were then required to rate how confident they were in the option chosen on a Likert scale, where 0 = not at all confident to 5 = extremely confident. After 20 s, the screen was blanked out to signal to the participants that the allocated decision time has ended. All participants then undertook the experimental air defence scenario, following the same procedure as described for the practice. Thirty decision events were presented during the experimental simulation. A decision event was defined as an occasion where a decision may need to be made by an operator. For example, an unknown data link track appears on the screen. Decision criticality was varied across the decision events presented (i.e., 10 high, 10 medium, and 10 low DC). Decision criticality related to the consequence of that decision. The decision held a higher criticality if there was a greater risk should the decision taken be incorrect in the high DC decisions (e.g., aircraft demonstrating hostile intent) compared to low DC (e.g., new track on radar screen) and the event occurrences varied depending on TL condition. The scenario stimulus ran for 20 min, 30 min, or 45 min for the high, moderate, and low conditions, respectively. The high condition involved an increased frequency of decisions, multiple decisions, and more aircraft tracks to monitor on the screen in comparison to the low condition which was characterised by one aircraft track to monitor at a time, decisions based only on the one track, reduced frequency of decision events, and longer periods of no action. Once the scenario had finished, participants completed situational awareness and workload questionnaires. Participants were fully debriefed to ensure each understood the nature of the study and given the opportunity to ask further questions.","To assess the differences in means a number of statistical analyses were performed on the data for accuracy, confidence, and W-S C-A using analysis of variance (ANOVA). A manipulation check was carried out to assess the differences in TL (see analysis of workload and situational awareness below). No significant differences were found in WL, there were differences in SA. The TL manipulation was not significant. An alpha level of .05 was used for all statistical tests. Accuracy ~~~~~~~~ The accuracy of the decisions was decided on by the SMEs. When designing the decision log and generating the decision options, one of the decision options was voted to be the best decision given the current situation. Participants were scored 1 for an accurate response or 0 for an incorrect response. The maximum total score was 30 and the maximum mean for each DC was 10. To examine the mean differences between TL and DC in accuracy, an ANOVA was conducted. Confidence ~~~~~~~~~~ Participants were asked to rate confidence in each decision, from 0 (not confident at all) to 5 (extremely confident). The maximum confidence score in total was 150 (50 for each DC). A 3 × 3 mixed ANOVA was carried out to assess the impact of TL and DC on decision confidence. As Mauchly's test of sphericity was found to be significant, the Greenhouse–Geisser estimate for df was used (see Table 2). Percentage Confidence in Correct and Incorrect Responses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To consider the variation in the data and examine the low correlations displayed for W-S C-A, percentage confidence in correct and incorrect responses was calculated. The W-S C-A correlation indicates the relationship between confidence and accuracy; however, a high correlation suggests both being highly confident in correct decisions as well as low confidence in incorrect decisions. Similarly, a negative W-S C-A correlation would suggest that individuals are highly confident in incorrect responses or not confident in correct responses. Thus, by examining the percentage confidence in incorrect or correct responses, the direction of the confidence (over/under confidence) can be investigated. Workload and Situational Awareness ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To assess the relationship between workload (WL) and situational awareness (SA), a series of Pearson's correlations were calculated. A significant negative relationship was found between SA and WL, r(58) = −.53, p < .001. Higher levels of reported WL were related to lower feelings of SA. A one-way ANOVA was conducted to assess the relationship between SA and TL. There was a significant effect of TL condition on SA, F(2, 57) = 6.44, p = .003. Participants in the low TL condition reported higher levels of subjective SA (M = 21.40, SD = 4.67) than participants in the high TL condition (M = 14.30, SD = 5.42) p = .002. No significant relationship was found between WL and TL, F(2, 57) = 3.00, p = .06. As a non- significant relationship was found, this would suggest that the manipulation check was not successful. As an exploratory analysis, the 6 dimensions of the NASA TLX (mental demand, physical demand, temporal demand, performance, effort, and frustration) were also examined to determine whether differences existed across the conditions. One-way ANOVAs were conducted with TL across the different dimensions of workload (see Table 4). An interesting finding was that the attentional supply was higher for moderate than high TL; this suggests that participants may have struggled more with the potential uncertainty that the moderate condition might have brought to bear. It appears participants were more able to deal with the ends of the spectrum where real differential could be identified (i.e., low and high conditions). Relationships Between WL, SA, Accuracy, Confidence, and W-S C-A ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To establish whether a relationship existed between WL, SA, accuracy, confidence, and W-S C-A a number of Pearson's correlations were carried out. Results revealed a significant negative relationship was found between overall WL and confidence, r(58) = −.42, p = .001. As subjective measures of workload increased, confidence in decisions decreased. In addition, a significant strong positive relationship was found between overall SA and confidence, r(58) = .63, p < .001. Higher scores in subjective SA were related to higher scores of confidence in decisions. However, no significant relationships were found between SA, WL, and W-S C-A, or between SA and accuracy or WL and accuracy in decisions; all comparisons, p > .05. Furthermore, no significant relationship was found for between- subjects confidence and accuracy, p > .05. The findings suggest that decision confidence influences both WL and SA. In this study accuracy was found to be unrelated to WL and SA. Personality Constructs ~~~~~~~~~~~~~~~~~~~~~~ This study was also interested in establishing whether accuracy, confidence, or W-S C-A were related to psychometric measures of personality. For this, Pearson's correlations were conducted. A significant negative relationship was found between tolerance to ambiguity and accuracy, r(58) = −.34, p = .008. Those who scored higher on the tolerance to ambiguity scale (i.e., less tolerant) were less accurate. In addition, a significant negative relationship was also found between decision style and accuracy r(58) = −.35, p = .005. High scorers on the decision style scale were less accurate. Decision style explicitly probes the need for quick and unambiguous answers. To investigate whether there were individual differences in participants experiences of WL and SA, correlations were conducted on each measure of the NEO-PI-R (openness to experience, conscientiousness, extraversion, agreeableness, and neuroticism). The results showed a significant relationship between openness to experience and WL, r(58) = −.28, p = .03. High scorers on the openness to experience scale reported lower levels of WL during the task. No other relationships were found to be significant, p > .05. These findings therefore suggest that some cognitive constructs are involved in decision accuracy and individual differences in participants’ feelings of WL.","A novel method of measuring metacognitive ability to assess the impact of decision criticality (DC) and task load (TL) on measures of confidence, accuracy, and W-S C-A was used with mock air defence operators. Personality constructs, workload (WL), and situational awareness (SA) were also assessed. DC impacted both decision confidence and decision accuracy. Low DC was found to increase confidence in decisions while high DC increased decision accuracy. Cognitive constructs were also found to be related to decision accuracy. The findings suggest that accuracy increases with DC. Participants made more accurate decisions in high DC than both low DC and medium DC conditions. This outcome supports previous findings that criticality influences performance (Hanson et al., 2014; Wheatcroft et al., 2017). Research has also shown that task performance increases when participants find the task more important (Kliegel, Martin, McDaniel, & Einstein, 2004), perhaps indicating that individuals believed high DC decisions to be important within the task context. This could relate to participants applying more attention and effort to decisions with greater consequences for an incorrect decision. Future research could examine decision processes and which mechanisms lead to increased accuracy. The study also demonstrates that DC influenced confidence. Individuals were significantly more confident in low DC decisions than medium DC decisions, lending support to previous literature that confidence decreases as difficulty increases (Chung & Monroe, 2000; Kebbell, Wagstaff, & Covey, 1996). No significant differences were found between high and low DC or medium and high DC. This could be due to increased uncertainty as condition criticality increased in the conditions relative to low DC, but high DC was perceptually transparent. Nevertheless, it is the corresponding confidence relative to an individual's awareness of the accuracy of decisions that is most important. W-S C-A remained unaffected, with no significant differences evident in W-S C-A across TL and DC. Some research has shown that training and experience improve calibration (Lichtenstein, Fischhoff, & Phillips, 1977). It would therefore be beneficial to conduct further studies using naval participants with appropriate experience. The high TL condition did not impact decision confidence, accuracy, or W-S C-A. However, the manipulation check was not significant; this could be one reason why no differences were found for some variables. The results support previous research which demonstrates that confidence is a relatively robust and general trait (Stankov & Lee, 2008). Confidence remained high, irrespective of accuracy. No relationship was found between decision confidence and accuracy. The accuracy scores were just below chance but individuals displayed elevated confidence—the means of both correct and incorrect scores were around 70%, suggesting that individuals are unaware of incorrect responses and overconfident in some decisions taken. This is consistent with previous literature that has demonstrated a general tendency for overconfidence (Lichtenstein et al., 1977). Importantly, none of the participants had any prior knowledge of air defence decision-making. Consequently, an additional explanation for elevated confidence levels can be explained by the Dunning–Kruger effect (Kruger & Dunning, 1999) where unskilled/novice individuals assess their ability to be too high. The study also examined how SA and WL relate to decision confidence, accuracy, and W-S C-A. It was found that SA was related to decision confidence. Individuals who reported higher levels of SA were also more confident in their decisions. However, these findings should be taken with caution. SA was measured subjectively and a confidence bias has previously been found in SA reporting (Sulistyawati & Chui, 2009). Thus, it might be that individuals are generally confident in their assessments of SA performance. Importantly, SA was not related to accuracy in decisions taken. Individuals may privately believe they had a better understanding of the situation than they demonstrated. Conversely, WL was found to be negatively related to overall decision confidence. Higher reported levels of WL resulted in lower levels of decision confidence. This is important for decision-making; reduced confidence in decisions taken could lead to increased WL. Individuals may seek out more information to support or contradict decision certainty. Further analysis of the WL dimensions showed significant differences in individuals’ feelings of temporal demand in the high TL condition. Participants felt more time pressured and reported the speed at which the task events occurred to be higher. Time pressure has been linked to individuals using different decision-making strategies (Maule, Hockey, & Bdzola, 2000); thus, speed of decision events may play an important role in critical environments. Further research with experts may demonstrate a mitigated effect. The investigations into broad personality constructs were found to be unrelated to confidence, accuracy, W-S-C-A, or SA. This suggests the constructs may not be related to these measures or sufficiently salient in decision-making processes. Relationships did exist with other measured constructs. Individuals less tolerant to ambiguity were less accurate in their decisions, and high scorers on the decision style scale were also less accurate. Budner (1961) argues that individuals who are less tolerant find ambiguous situations threatening. It is probable that a lack of tolerance hindered individuals’ ability to make accurate decisions. Tolerance to ambiguity is a likely desirable trait for accurate air defence decision- making. Further, although not replicated in this study, Wheatcroft et al. (2017) found an intolerance of ambiguity was negatively related to W-S C-A (i.e., a greater tolerance of ambiguous conditions was related to increased W-S C-A). Further research is warranted to investigate the relationship between decision-making and ambiguity tolerance in critical environments. Outcomes also showed WL to be negatively related to openness to experience, providing support for the findings that some aspects of personality impact the perception of WL (Chiorri, Garbarino, Bracco, & Magnavita, 2015). Although the study was initially interested in the OR role and novice capacity to the task, one limitation was the use of novice participants rather than experts. However, it may be beneficial to the NDM paradigm to understand how expertise is developed (Hoffman and Klein (2017); Sala & Gobet, 2016). For instance, Klein, Hintze, and Saab (2013) developed the shadow box technique which helps novices understand the decision-making processes of experts. Therefore, the use of novices in this study does allow for a baseline comparison. The use of novices may also help to understand the training needs of less experienced decision-makers. The authors will use experts to further validate the work in future research. Research has found that practice can degrade certain aspects of metacognitive performance (Jackson, Kleitman, & Aldman, 2015). However, this study minimised this effect by reducing potential fatigue. It has been argued that NDM research should use a mixture of measures to reduce the limitations of using a single methodology (Lipshitz et al., 2001). This paper aimed to introduce the W-S C-A measure to assess an element of metacognitive ability in air defence operators using a realistic decision-making scenario together with combined objective and subjective measures. The proposed method and outcomes will provide a wider view of metacognition in critical decision-making environments. The broader implications include the potential for the approach to be used to prioritise training and selection, with the aim of improving effective air-defence decision-making.","This study found that decision criticality has significant impact on both decision accuracy and confidence. Future work should consider the impact of decision criticality and tolerance of ambiguity on accurate air defence decision-making.","All named authors were involved in the conception and design of the study, critical revision of the article, and final approval of the published version. The first author was responsible for data collection and data analysis, interpretation, and in drafting the final article.","No conflicts of interest are declared."],["Scales measuring procrastination focus on different aspects of unnecessary and unwanted delay, delay in task implementation – an increased gap between intention and action – being a core characteristic. However, an inspection of existing procrastination scales reveals that the scales do not distinguish between two facets of implemental delay, onset delay, and delay related to sustained goal striving. We trace this failure to an imprecise understanding of “delay,” another core concept in procrastination. This paper discusses the relationship between onset and sustained delay in procrastination, and then describes a new scale attempting to measure these two facets of task implementation. In two studies (aggregated N = 465) we demonstrate, using exploratory and confirmatory factor analysis, that although onset and sustained action procrastination measures correlate, they are still separate facets of implemental procrastination. Problems with onset delay seem to be particularly important, increasingly so in high procrastinators. Implications, as well as suggestions for further research, are discussed. -------------------------------------------------------------------------------- MEASURING IMPLEMENTAL DELAY IN PROCRASTINATION: SEPARATING ONSET VS. SUSTAINED GOAL STRIVING -------------------------------------------------------------------------------- Motivated behavior extends over time. Following a decision, the individual must plan how to implement it, then follows goal striving or implementation, and finally goal attainment if successful (e.g., Achtziger & Gollwitzer, 2018). People may procrastinate – delay unnecessarily – in all these stages (e.g., Svartdal & Steel, 2017), and scales have been developed to measure procrastination in each. For example, the Decisional Procrastination Scale (DPS, Mann, 1982, unpublished; Mann, Burnett, Radford, & Ford, 1997) focuses on delay in decision-making and onset of implementation. General procrastination scales, such as the General Procrastination Scale (GPS; Lay, 1986) address various examples of implemental delay. Finally, McCown and Johnson's Adult Inventory of Procrastination Scale (AIP; McCown, Johnson, & Petzel, 1989) includes items related to promptness, meeting deadlines, and timeliness. A closer inspection of the procrastination literature reveals, however, that although implemental delay is a key feature of the procrastination problem (e.g., Klingsieck, 2013; Steel, 2010), this core concept has rarely been explicated in the procrastination literature. By definition, procrastination is a dysfunctional delay of intended behavior, with procrastinators tending to demonstrate larger intention-action gaps compared to non-procrastinators (e.g., Steel, Brothen, & Wambach, 2001). However, as reviewed by Sheeran and Webb (2016), the intention-action gap addresses three rather distinct phases of goal pursuit – initiation, maintenance, and close when the goal has been attained. Sheeran and Webb noted that different factors are involved when problems occur in these phases. For example, failure to get started may be rooted in factors such as forgetting to start, indecision about means, second thoughts, and failure to engage in preparatory behaviors. On the other hand, not keeping goal pursuit on track relates to factors such as failure to monitor goal progress, competing goals, bad habits, disruptive thoughts and feelings, and low willpower. In support, Steel and Weinhardt (2018) adopt a similar three-stage model for their Goal Phase System and note that the motivational forces operating during the beginning of a goal are not necessarily the same as those later, particularly during goal-striving or realization. For example, high expectancy or confidence assists the initial goal choice but can have a negative effect during goal- striving as those overconfident in their abilities may underinvest in allocation of time and resources. Similarly, impulsiveness can have a neutral or positive effect with goal choice, interacting with extrinsic rewards, but often hampers goal striving until just before deadlines.1 Unfortunately, the procrastination literature seems to have focused primarily on delay in goal pursuit initiation, often neglecting that procrastination manifests itself also in goal maintenance pursuit (cf. Gollwitzer, 2014). One reason for this situation may be that delay – a defining criterion for procrastination – is ambiguous. Often, “delay” is used as a temporal judgment of unnecessary delay (e.g., as in item 2 in the DPS, “Even after I make a decision I delay acting upon it”), and some seem to restrict procrastination to such cases (Tice, Bratslavsky, & Baumeister, 2001, p. 63). However, “delay” is also used in another meaning, referring to indirect delays as a result of impulsive diversions and other forms of wasting time during intention implementation (as in item 12, GPS, “In preparing for some deadline, I often waste time by doing other things”). We argue that both usages are legitimate, but that the first seems to be the default interpretation whereas the second has often been overlooked in the procrastination literature, and especially so in scales measuring procrastination. In the next paragraph, we expand on these two arguments. Then we present, in two studies, a new scale attempting to measure two facets of implemental delay, onset delay, and delay in sustained goal striving.","Procrastination is defined by two core characteristics, the first being the delay of some intended behavior, the second that this delay is chosen despite realizing the negative consequences of the delay (Klingsieck, 2013; Steel, 2007). The latter criterion implies that procrastination is “irrational” in the sense that the individual acts against better judgment, often referred to as akrasia (Andreou & White, 2010). This understanding of procrastination implies that internal norms and cognitive-affective evaluations of delay play an important role in identifying procrastination (Milgram & Naaman, 1996; van Eerde, 2000). Furthermore, procrastination must be distinguished from rational forms of delay, as many forms of delay of intended behavior may be adaptive, rational, and beneficial. Importantly, a definition of procrastination in terms of delay despite better judgment may leave the impression that all forms of procrastination are defined in terms of timing. Given a model of goal-directed action flow from deliberation → planning → action → evaluation (e.g., Achtziger & Gollwitzer, 2018), timing would primarily be relevant in the transitions between deliberation and decision (decisional procrastination), between intention formation and intention realization (the intention-action gap), and finally in goal attainment (e.g., timeliness). In these cases, unnecessary or irrational delay may be observed in accordance with the above definition. However, many instances of intended acts unfold over longer periods where dilatory behaviors may manifest themselves in ways that create indirect delays in goal striving, yet without involving delays according to temporal criteria. First, engaging in competing activities may delay goal-directed behavior in an indirect way, as less time – and less adequate time – is spent on the prioritized goal activity (cf. Lay, 1986, Study II). Accordingly, several scales contain items addressing preference for other activities, as for example the GPS (Lay, 1986) item 1 (“In preparation for some deadline, I often waste time by doing other things”) and the IPS (Steel, 2010) item 4 (“When I should be doing one thing, I will do another).” Note that for these items there is no mention of delay per se; delay in goal implementation results from engaging in competing activities. Accordingly, experience sampling of procrastination (e.g., Pychyl, Lee, Thibodeau, & Blunt, 2000, Table 1) focuses on activities that people actually do at the moment (e.g., watching TV) versus what they should have been doing (e.g., studying), defining “procrastination” as doing something other than what one should do. Second, a closely related variant of diversion during goal- striving occurs when goal-directed behavior is impulsively diverted to more tempting situational alternatives, indirectly creating delays in realizing goals (Schouwenburg, 1995; Steel, 2007; Steel, Svartdal, Thundiyil, Brothen, 2018). Such impulsive diversions may address distractions and temptations that are available to the individual, and susceptibility to them indicate present-bias preferences characteristic of procrastinators (Steel et al., 2018). Examples include giving in to more pleasurable alternatives compared to continued goal striving (e.g., watching TV instead of reading), dysfunctional forms of mood-regulation (Sirois & Pychyl, 2013; Tice & Bratslavsky, 2000), and escape/avoidance from something aversive (e.g., escaping from a stressful or boring task). Such forms of delay reduce the amount of time spent on focal tasks (e.g., Tice et al., 2001, Experiment 3). Finally, procrastination during goal-striving may result from other overlapping factors, such as poor self-monitoring, competing goals, dysfunctional habits, poorly formulated goal intentions, low willpower, concentration problems, tiredness, and task aversion (Gollwitzer, 2014; Schouwenburg, 1995; Sheeran & Webb, 2016; Steel et al., 2018). Also, note that a conception of procrastination as a breakdown in self-regulation (Steel, 2007) implies that much of the procrastination problem expresses itself during the goal- striving phase. Concluding from these examples, we suggest that “delay,” although being a core criterion for procrastination, has remained inadequately clarified in the procrastination literature. Whereas some forms of procrastination are easily identifiable according to strict temporal criteria, others are less so and satisfy the delay criterion only indirectly. Interestingly, Lay (1986), defined procrastination as “the tendency to postpone that which is necessary to reach some goal” (p. 457), implying that procrastination does not refer to a single act of delay, but to a tendency demonstrated over time to delay goal-relevant acts. In effect, is seems to be important to recognize that procrastination is a dynamic phenomenon unfolding over time, manifesting itself in tendencies to act in ways that prove suboptimal in attaining goals. Such delays may result from explicit delays as well as be a by-product of maladaptive strategies in sustained goal-directed behavior. In both cases, the delays and strategies must be “irrational” or akratic for them to be regarded as procrastination. How procrastination scales measure different goal phases ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The procrastination literature is unclear on the various stages of intended goal realization, and so are scales attempting to measure procrastination. They do not distinguish between delays in getting started (onset delay, defined in terms of a temporal gap between intention formation and relevant goal-striving behavior) and failures in keeping goal pursuit on track (delays created in indirect ways during goal-striving). We examined 18 procrastination scales to determine their coverage of these implementation facets. Some scales include items that address delayed onset, as the DPS (item 2, “Even if I make a decision I delay acting upon it”), the GPS (item 1, “I often find myself performing tasks that I had intended to do days before”), and the Unintentional Procrastination Scale (UPS,(item 1, “I rarely begin tasks as soon as I am given them, even if I intend to;” Fernie et al, 2017). In contrast, very few scales address issues during goal-striving that may create unnecessary delay in more indirect ways. One exception is the Academic Procrastination State Inventory (APSI; Schouwenburg, 1995). The APSI asks respondents to indicate how “frequently last week did you engage in the following behaviors or thoughts,” listing a number of examples related to disruption of goal striving (e.g., “Gave up studying because you did not feel well, ” Drifted off into daydreams while studying,” and “Experienced concentration problems when studying”). However, this scale also contains items not addressing sustained goal striving (e.g., “Forgot to prepare things for studying”). Items in this scale were intended to cover the three important facets of procrastination according to Schouwenburg, 1995, lack of immediacy in intentions and behavior, a discrepancy between intention and behavior, and a preference for competing activities, and were not intended to measure failure of goal striving per se (cf. also Patzelt & Opitz, 2014). Other scales may address procrastination during sustained goal striving but are ambiguous as to whether they also refer to onset delay. For example, the GPS item 12 “In preparation for some deadline, I often waste time by doing other things” may refer both to initiation and sustained goal striving. This ambiguity also applies to the two IPS items discussed (IPS 2, “If there is something I should do, I get to it before attending to lesser tasks” - R), and IPS 4, “When I should be doing one thing, I will do another”). Table 1 summarizes items from existing procrastination scales that address delays in onset and sustained goal striving2. As is seen from the table, we have found only one scale addressing delay during goal striving, and this scale is intended for use in academic settings. Fig. 1 illustrates the phases of intended action, separating the implementation phase in onset versus sustained goal striving (cf. Sheeran & Webb, 2016). As is indicated in Fig. 2, procrastination in these phases may take different forms, and measuring them independently may be important. For example, delayed onset implies that effort during sustained goal striving must be increased if a task is to be completed within a fixed timeframe (e.g., Steel et al., 2018, Fig. 1). In this case, even though the onset and sustained action procrastination measures may be positively correlated, slow onset in high procrastinators predicts lower sustained procrastination because slow starters must catch up to finish in time. Another issue of interest is which of the two facets of procrastination – slow onset and delays in sustained goal striving – best predict overall procrastination score. If, as suggested, delayed onset is compensated by increased effort during goal striving, one may suspect that onset delay is a main determinant of overall procrastination. Finally, as procrastination implies reduced time spent on important tasks (Lay, 1986; Tice et al., 2001), one may ask whether reduced time can be traced to late onset, to impulsive diversions during goal striving, or to both. Given the compensation hypothesis discussed, it is likely that onset delay may be the best predictor of reduced time spent on important tasks. Onset delay As seen in Table 1, the GPS (Lay, 1986) contains two items addressing implemental delay, the DPS (Mann, 1982) one. Both items are included in the Pure Procrastination Scale (PPS; Steel, 2010; 6 and 8, from the GPS; 2 from the DPS). Furthermore, the Aitken Procrastination Inventory (API; Aitken, 1982) contains seven items explicitly addressing onset delay (e.g., “Even when I know a job needs to be done, I never want to start right away”; “It often takes me a long time to get started on something”). Finally, the Volitional Components Inventory / Volitional Components Questionnaire (VCI/VCQ (Kuhl & Fuhrmann, 1998) contains four items addressing swift action when action possibility presents itself, thus being inconsistent with delayed onset (e.g., “When something needs to be done I start without hesitating,” “When a task needs to be done, I like to do it right away”). Most of these items address onset delay (or the opposite), but they generally fail to address two important features of procrastination, delay in the context of an explicit intention, and general delay versus procrastination. Hence, we slightly modified some items (marked in italics in Table 2) to reflect intentions and akratic delay. The authors discussed a pool of 18 possible items and selected, based on content analysis, six items for use (Table 2). The two first items are from the GPS (also used in the PPS) and may serve as a benchmark for the other onset items, as they have proven successful in numerous studies of implemental delays (e.g., Svartdal & Steel, 2017). Sustained goal pursuit delay As discussed, whereas onset delay primarily relates to a timing criterion (“I will start tomorrow, not today”) or prioritizing (“I will do X rather than Y, even if both are possible”), sustained goal striving relates to commitment and having the necessary time, focus, and energy to implement intentions over time. For aversive and boring tasks, other alternatives become tempting, and especially so when one gets tired, and exhaustion may result if the task is demanding, also increasing the likelihood of being tempted to do other and more attractive things (Steel et al. 2018). Hence, procrastination in this action phase may relate to, among other things, tiredness, exhaustion, focus (being distracted), temptations, nature of task (difficult, boring), tasks decided by others, taking a break to escape for a while, give in after minor setbacks, dislike hard work, typically not finish tasks, change one's mind, and others (e.g., Schouwenburg, 1995). Table 3 lists the items selected for the present studies. All were custom made for this study, but several were inspired by items in the APSI scale (Schouwenburg, 1995). Delay in reaching the intended goal Finally, to cover the end of goal pursuit, we included items to measure meeting deadlines and timeliness (Table 4). Appropriate items are found in existing scales, especially in the AIP; (McCown et al., 1989; Svartdal & Steel, 2017). We also added custom items. The present studies ~~~~~~~~~~~~~~~~~~~ Study 1 assessed items purporting to measure different facets of procrastination (Tables 2–4) using exploratory factor analysis (EFA), both for item functioning as well as for factor structure, the overall expectation being that items would organize into a three- factor solution - onset, sustained goal striving, and timeliness. Study 2 further examined these items using confirmatory factor analysis (CFA), testing the factor structure as suggested by Study 1. Having established measures of onset procrastination versus sustained goal striving procrastination, we compared these measures against a well- established procrastination scale, the Irrational Procrastination Scale (IPS; Steel, 2010). This scale correlates highly with other general procrastination scales (Svartdal & Steel, 2017, Table 7). We expected for these data that the three facets of procrastination would correlate moderately to highly with the IPS score, and – as discussed – that the onset subscale would demonstrate an especially close relationship to overall procrastination score. Next, as delayed onset narrows the timeframe available for task completion, we examined the relations between the onset and sustained procrastination measures. Although these measures should correlate rather highly, a “compensation” hypothesis predicts that individuals delaying onset must work even harder to complete tasks in time. Hence, high procrastinators should demonstrate high onset procrastination scores but relatively lower scores on the sustained goal striving procrastination measure, whereas the opposite pattern should be observed in non-procrastinators. In both studies, we also administered a simplified model of a typical task completion sequence based on the model presented in Fig. 1 and asked participants to indicate the perceived difficulty associated with each stage. Task difficulty is associated with procrastination (Steel, 2007), and this procedure therefore served as an independent measure of problems associated with the various phases of intended action. Here we expected that the onset phase would be especially prone to be perceived as difficult, and especially so with increasing overall procrastination score. Finally, we measured self-reported time spent on self-directed academic work. An overall expectation is that procrastination limits the time available on a given project (Lay, 1986), suggesting a negative correlation between self-directed academic work and procrastination score. Note that direct and indirect delays have a common effect, as both limit the time available for goal-directed work. Separating onset delay and sustained goal striving delay may give an additional perspective to this picture, and again we expected onset delay to be particularly predictive of time spent on academic work.","The sample comprised 170 students aged 18 to 48, mean age = 25.56 (SD=5.47). The majority of participants (85.7%) were females. Procedure and ethics Participants were recruited by circulating a survey on mailing lists and social media at several universities and high schools in Norway. All were informed that participation was voluntary and anonymous and that they could withdraw from the study at any time. Participants were given information about the purpose of the study together with a link to the online survey (www.qualtrics.com) and gave informed consent by pressing a “start the survey” button. The current project is a part of a larger project on procrastination, which has ethical approval from the Regional Ethical Board in Tromsø, Norway (REK nord 2014/2313).","The item sets shown in Tables 3–5 were presented sequentially and rated on a 5-point Likert scale with higher scores indicating increasing agreement. Then the survey presented a simplified model of a typical task completion based on the model presented in Fig. 1 and asked participants to indicate, on a 1-5 scale, the perceived difficulty associated with each stage. Finally, all responded to the six-item version of the IPS (Svartdal & Steel, 2017). The original IPS (Steel, 2010) features nine items, of which three are reversed. In the six-item version, the reversed items are deleted. Both versions demonstrate good internal consistency, α = .85–.93, and the full and reduced versions correlate highly. All scale items are rated on a 5-point Likert scale, with higher scores indicating more procrastination. Analysis Exploratory factor analysis (EFA) was employed using principal axis factoring. An eigenvalue > 1 and the scree plot test were used to determine underlying factors. An oblique rotation (oblimin) was applied since the extracted factors were expected to be correlated. Analyses were performed in IBM SPSS statistics 25.","The EFA was appropriate as indicated by a Kaiser-Meyer-Olkin (KMO) measure of sampling adequacy of above 0.906 and a highly significant Bartlett's Test of Sphericity (p<0.001). The initial EFA analysis produced three underlying factors onset, sustain, and timeliness, as indicated by eigenvalues above 1 (8.71, 1.89 and 1.50) and the scree plot, confirming our overall expectation of the factor structure. Six items were reexamined based on low factor loadings and/or double loadings, resulting in deletion: Onset items 4 (“Often I start so late that I miss the deadline”) and 6 (“When I am to start tasks I planned to do, I often end up doing something else instead”) demonstrated double loadings on the onset and timeliness subscales. Sustain item 7 (“I can work on other things (check mail, FB, and so on) while doing academic work”) was deleted due to low factor loading. The timeliness item 1 (“I get short of time”) produced a double loading on onset and timeliness subscales. Item 5, suggested to measure timeliness (“When I work on an assignment, I typically finish before others”), loaded mainly on onset. Finally, “I lag behind on study work” produced a double loading on onset and timeliness. We then repeated the EFA with the 13 items retained. Again, three factors with eigenvalues > 1 (5.86, 1.63, and 1.42) appeared, accounting for 68.5% of the variance. The oblimin rotated loadings for each of the retained items are reported in Table 5. In conclusion, these data support a three- factor model, separating the two implemental facets onset and sustained goal striving, as well as confirming that timeliness is a separate facet of procrastination (e.g., Svartdal & Steel, 2017). Table 6 presents means and correlations between the three subscales, the IPS, and difficulty ratings of the action phases. As expected, the onset subscale correlated highly (r = .80), with the IPS, whereas correlations were lower for the other subscales. To assess the relation between the onset and sustained procrastination scores to general procrastination, a regression with IPS score as the dependent variable and subscale scores as predictors indicated that only onset scores significantly predicted IPS score, β = .70 (p<.001), whereas the sustained and timeliness subscales did not, β = .12 and β = .10, respectively. Hence, onset procrastination seems to be the better predictor of general procrastination, consistent with it being the first factor extracted. To further assess the relation between the two facet measures and general procrastination, we performed an ANOVA with the procrastination facets as dependent variables and IPS levels (IPS levels 1-5 as defined by the individual's absolute IPS score) as a categorical predictor. As argued, the expectation for these data was that both onset and sustained scores should increase with increased levels of general procrastination, but that this change would be especially pronounced in the onset measure and less so in the sustained measure. The ANOVA indicated an overall significant effect, F(4, 164) = 53.70, p < .001, η2 = .57, reflecting that both facet scores increased with increasing IPS levels. Importantly, the interaction effect was significant, F(4, 164) = 7.40, p < .001, η2 = .15, indicating that onset problems escalated more than sustained goal striving problems over increasing general procrastination levels. These results are displayed in Fig. 3 (left panel). Note in this figure also that timeliness scores increase only modestly over increasing IPS levels, and whereas overall onset and sustained procrastination means were quite similar (2.74 and 2.83), the overall timeliness mean was significantly lower (1.63), F(2, 328) = 135.06, p < .001, η2 = .54. Fig. 3 (right panel) shows the mean problem ratings for the four task completion phases over different procrastination levels. Difficulty ratings increased with increasing procrastination levels, F(4, 163) = 32.38, p < .001, η2 = .44. In addition, the phase * procrastination level interaction was significant, F(12, 489) = 2.10, p = .015, η2 = .05, reflecting that onset scores were especially sensitive to overall procrastination levels. Table 6 displays the correlations between the mean onset, sustained, and procrastination scores and action phase difficulty ratings. Notably, both the onset and sustained procrastination scores correlated highly, r = .67, with their corresponding difficulty ratings in the onset and sustained action phases. Finally, the relation between time spent on self-directed academic work and procrastination score (IPS) was negative, rS = -.33. The onset score correlated somewhat higher with academic work, rS = -.38, the sustained lower, rS = -.22, indicating again that onset problems contribute more to reduced time available for academic work.","Study 2 was performed as replication of Study 1, allowing for the use of CFA on a sufficiently large independent sample.","The sample comprised 295 students between the ages of 19 to 55, mean age = 26.28 (SD=7.56) years. The majority of participants (82%) were females. Procedure and ethics Participants were recruited by circulating the survey by email and on social media using assistants at several universities throughout Norway. See Study 1 for procedure and ethics information.","The material was identical to that of Study 1, except that the timeliness item 2 was not included in the data collection3. Hence, only two items were specified as indicators of the timeliness factor. The onset and sustain factors had four and six indicators, respectively, corresponding to those used in Study 1. Analysis CFA was employed using two estimators, robust maximum likelihood estimation (MLR) due to deviation from normality, and robust weighted least square (WLSMv) that is appropriate for ordinal data and relatively smaller samples (Kline, 2016, p. 326). Model fit to data was examined using standard fit indices (Brown, 2015), specifically the comparative fit index (CFI), the Tucker-Lewis index (TLI), the root-mean-square error of approximation (RMSEA), and standardized root-mean-square residual (SRMR). An RMSEA less than 0.05, CFI and TLI values greater than 0.95, and SRMR less than 0.05 represent a well-fitting model. RMSEA as high as .08 indicates a reasonable fit, whereas RMSEA values ranging from .08 to .10 indicate mediocre fit and estimates above .10 indicate poor fit (Browne & Cudek, 1993; MacCallum et al., 1996). In evaluating the models, we conducted chi-square difference tests. When conducting chi-square difference tests using the MLR estimator, it is necessary to adjust the chi-square using the Satorra-Bentler scaling correction (for details, see Bryant & Satorra, 2012). When using the estimator for categorical data (i.e., WLSMV), the difference test procedure using chi-square is applied (for details, see Asparouhov & Muthén, 2010). Analyses were performed with Mplus version 8.3.","First, the fit of a three-factor model was compared to a single-factor model. As seen in Table 1 (Appendix), the single-factor model produced a poor fit across estimators. The three-factor model had a good fit with the WLSMv considering CFI>.95, TLI>.95, and SRMR<.05, and a mediocre fit considering RMSEA=0.087, whereas MLR produced reasonable fit (i.e., CFI>.90, TLI>.90, SRMR=.052) and mediocre fit (i.e., RMSEA=0.082). The chi-square difference tests indicated support for the three-factor model (WLSMv Chi-square difference = 209.986, df=3, p<0.001; MLR Chi-square difference 250.361, df=3, p<0.001). Second, based on modification indices, errors of sustain items 1 and 4 were correlated with sustain item 2 and 3, and sustain item 2 was correlated with sustain item 3. In both cases, these steps were reasonable, as these items demonstrate considerable overlap in meaning. Both the single-factor and three-factor models were modified, resulting in a well-fitting three- factor model and a poor-fitting single-factor model with both estimators. Again, a chi- square difference test was employed yielding support for the three-factor model (i.e., WLSMv: chi-square diff=192.973, df=3, p<0.001; MLR chi-square diff=197.566, df=3, p<0.001). Finally, an alternative model with sustain item 3 removed (because of high similarity to sustain item 2) produced an even better fit using both estimators (see Table 1, Appendix). Again, based on modification indices using both estimators, errors of sustain item 3 was correlated with sustain item 1 and 5. This modified model produced excellent fit with the WLSMv (χ2 = 53.77, df=39, p=0.058; CFI=0.997; TLI=0.995; RMSEA=0.025, and SRMR=0.025) and with the MLR (χ2 = 41.61, df=39, p=0.358; CFI=0.998; TLI=0.997; RMSEA=0.016 and SRMR=0.029). Thus, the CFA confirmed the results of Study 1 supporting a three-factor model. The result of the CFA (MLR) is presented in Fig. 4 (left panel). As noted, in Study 1 only onset score significantly predicted overall procrastination score (IPS). Fig. 4 (right panel) displays the corresponding analysis for the present study, using SEM to model how the three subscales predict IPS score. The model demonstrated a good fit, CFI=0.974; TIL=0.968; RMSEA=0.041(90%CI 0.033 – 0.058); SRMR=0.041. As is seen from the figure, onset significantly predicted IPS score, β = .68, whereas the two other subscales demonstrated substantially lower but still significant beta values, β = .18 (sustain) and .16 (timeliness). Table 7 presents means and correlations between the three subscales and the IPS. The results were similar to those of Study 1, the onset subscale correlating highly (r = .80) with IPS, and lower correlations for the other subscales. The relation between the facet measures and general procrastination demonstrated similar patterns as those observed in Study 1. First, the ANOVA indicated an overall significant effect of the onset and sustained procrastination facets over increasing levels of IPS, F(4, 223) = 69.99, p < .001, η2 = .56. Second, as in Study 1, the onset * sustained interaction was significant, F(4, 223) = 6.78, p < .001, η2 = .11, reflecting that onset problems escalated more than sustained goal striving problems over increasing general procrastination levels. These results are displayed in Fig. 5 (left panel).4 The timeliness scores increased over increasing IPS levels, but levels were lower compared to the onset and sustained procrastination scores. Overall onset and sustained procrastination means were quite similar (2.86 and 2.92), whereas the overall timeliness mean was significantly lower (2.43), F(2, 446) = 18.13, p < .001, η2 = .08. The problem ratings for the four task completion phases over different procrastination levels demonstrated similar results as those of Study 1, see Fig. 5, right panel. First, difficulty scores increased with increasing procrastination levels, F(3, 223) = 44.49, p < .001, η2 = .37. Second, the phase * procrastination level interaction was significant, F(9, 669) = 5.70, p < .001, η2 = .07, demonstrating that onset scores were especially prone to increase as a function of general procrastination (IPS) levels. The relation between time spent on self-directed academic work and IPS score was negative, rS = -.28. Again, the onset score correlated somewhat higher with academic work time, rS = -.40, the sustained lower, rS = -.18, indicating that onset problems contribute more to reduced time available for academic work.","The distinction between delayed onset and disruptions during goal-striving is well established (e.g., Sheeran & Webb, 2016). Whereas the first primarily relates to timing and manifests itself directly in terms of later onset, delay in sustained goal striving occurs in a variety of direct and indirect ways. Thus, while delayed onset may address the prototypical instance of procrastination, the Latin procrastinus literally referring to something “belonging to tomorrow,” failure to stay on track toward a goal may be rooted in a variety of disruptive factors, indirectly creating delay in goal striving. The latter facet of procrastination is rarely addressed in procrastination research, nor in scale construction. The lack of procrastination scales to address both facets of implemental procrastination is unfortunate, even more so because this circumstance reflects an insufficient understanding of the phenomenon in procrastination literature. The present paper attempted to theoretically differentiate the two facets and to present a scale to measure them. The scale proposed in the present paper explicitly addresses two phases of task implementation, onset and continued goal striving. Results of Study 1 suggested a three-factor structure corresponding to the suggested constructs onset delay, delay during goal striving, and timeliness. The onset delay measure correlated highly with a standard measure of general procrastination (IPS; Steel, 2010), whereas the two other measures demonstrated lower correlations with the IPS. Study 2 replicated the first study with a larger sample, allowing for a specific test using CFA of the factor structure. The results confirmed the three-factor solution identified in Study 1. Correlational analyses repeated the findings from Study 1, with a high correlation between the IPS and the onset delay measure, and lower correlations with the sustained goal striving and timeliness measures. Both studies indicate support for the specific role of onset delay in procrastination. Specifically, an analysis of the relations between the three subscales over different levels of procrastination (as measured by the IPS) demonstrated that the onset delay means were lower relative to the sustained means for lower levels of procrastination, whereas for higher levels of IPS the levels reversed. This shift indicates that onset delay becomes an increasingly important problem as overall procrastination tendencies increase. These results were supported by an independent measure of perceived difficulties associated with the typical phases of planned action, decision, getting started, keeping up sustained action, and finishing in time. Again, onset was identified as the most problematic phase, and especially so by high procrastinators. Finally, self-reported time spent on self-directed academic work was correlated with general procrastination score, but higher so with the onset procrastination measure. In essence, the German saying “Aller Anfang ist schwer” (“Every beginning is hard”) and the Norwegian expression “dørstokkmila”(“the doorstep mile”) get a deeper meaning according to these data, as starting something may be hard for everyone, and is even harder as overall procrastination tendency increases. As delayed onset narrows the timeframe available for task completion, but many tasks have a fixed time frame for completion, the present data indicate that late onset may be compensated by increased effort during the sustained goal striving phase. However, working harder under stricter time regimes fosters stress and suboptimal conditions for work, which, in turn, probably lead to lower academic achievement if, at all, the task is completed. Limitations and future research ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our results indicate that onset delay is particularly important in procrastination, but an alternative interpretation is that procrastination during goal-striving is more difficult to monitor and hence to report and measure. Specifically, onset delay is a construct addressing timing, given an intention to act. Even if there may be various specific reasons for delaying onset (e.g., indecision about means, second thoughts), the result as perceived by the actor is a failure to start as intended. In contrast, delays during task execution may be due to diverse and complex factors that work indirectly to create delay. For example, low willpower can make situational temptations appear attractive, causing impulsive diversions from planned action (e.g., Steel et al., 2018). Although such diversions are themselves detectable, delays in task execution because of them may not be equally obvious to the individual. Hence, as we are notoriously poor at monitoring ourselves attentively when we attempt to self-regulate (Baumeister & Heatherton, 1996), delays during task pursuit may be more difficult to self-report. Furthermore, the fact that delayed onset often makes it necessary to increase goal striving effort to catch up may indicate to procrastinators that they work as hard as, or even harder, compared to others, even to non-procrastinators. Future research should assess these issues, also in situations that are not compromised by differences in onset (e.g., in controlled situations examining disruptions in goal striving with no differences between participants in onset delay, cf. Tice et al., 2001, Experiment 3). Turning to the limitations of the study, the sample size used in Study 1 is quite small. However, the factor loadings are generally high (i.e., loadings ranging from .44 to .91). Studies have shown that moderate to high loadings was the major factor in determining reproducibility (Guadagnoli & Velicer, 1988; Tabachnick & Fidell, 2013). For instance, Guadagnoli and Velicer (1988) suggested that a sample size of N≥100 was sufficient if four or more variables per factor had loadings above .60. Further, the present research is based on two student samples, suggesting that future studies should assess factor structure as well as relations between onset delay and delay in sustained goal striving in samples from the general population as well. Such studies should also assess scale properties over gender, age, and even cultures (e.g., Kankaraš & Moors, 2010). Items used to assess sustained goal striving were inspired by the work of Schouwenburg, 1995, and would profit from scrutiny and possibly adding other items. In this respect, a limitation in Study 2 is the post-hoc modification to improve the CFA model fit indices. Several measurement errors were allowed to correlate among indicators of the sustain-subscale, which, as pointed out by Hermida (2015), moves the model testing from being confirmatory to becoming an exploratory analysis. Error correlations are likely due to sampling error, which may restrict cross-validation of the structure in future studies (Hermida, 2015; Grant, 1996). Further, the underlying structure may be masked if a relevant omitted variable is estimated through measurement error (Cortina, 2002; Landis, Edwards, & Cortins, 2009). However, the model fit of the non-modified three-factor model was generally in the range of an acceptable fitting model. In particular, the alternative three-factor model (Appendix, Table 1) produced a good fit to the data. This again suggests that items would profit from scrutiny and potentially adding other items, and future research should address cross-validation and replication. Another limitation of the present research is that we have utilized only one general procrastination scale, the IPS (Steel, 2010), as a reference. Other scales may demonstrate different relations to the subscales discussed in the present paper, even more so because different scales focus on different facets of procrastination. Hence, future studies should include other procrastination scales to verify the close relation observed between overall procrastination and onset procrastination, and the relatively weak relation between overall procrastination and procrastination in the goal-striving phase. Another important step in validating the differentiation between onset procrastination and procrastination in the goal-striving phase is to investigate whether the two facets relate differently to motivational and volitional variables. As noted, the motivational forces operating during the beginning of goal pursuit are not necessarily the same as those important during later goal striving (Steel & Weinhardt, 2018). Thus, motivational variables, such as expectancies and values, should be related strongly to onset procrastination, whereas volitional variables, such as the ability to shield distractions or willpower in general, should relate more strongly to procrastination in the goal- striving phase. In a future study, we will relate the scale of the two facets to instruments measuring the different forms of motivation (as in the Self-Determination Theory; Deci & Ryan, 2008), strategies of regulation of motivation (Grunschel, Schwinger, Steinmayr & Fries, 2016), volition (Kuhl, 1984), and energy (Steel et al., 2018)."],["The present paper provides a review of research on medical students' attitudes to people with intellectual disabilities. The attitudes of medical students warrant empirical attention because their future work may determine people with intellectual disabilities' access to healthcare and exposure to health inequalities. An electronic search of Embase, Ovid MEDLINE(R), PsycINFO, Scopus, and Web of Science was completed to identify papers published up to August 2013. Twenty-four studies were identified, most of which evaluated the effects of pedagogical interventions on students' attitudes. Results suggested that medical students' attitudes to people with intellectual disabilities were responsive to interventions. However, the evidence is restricted due to research limitations, including poor measurement, self-selection bias, and the absence of control groups when evaluating interventions. Thus, there is a dearth of high-quality research on this topic, and past findings should be interpreted with caution. Future research directions are provided. © 2014 The Authors. --------------------------------------------------------------------------------","The electronic databases Embase, Ovid MEDLINE(R), PsycINFO, Scopus, and Web of Science were used to search for manuscripts that examined medical students’ attitudes to people with ID. The search was conducted within the titles and abstracts of English language journal articles published before the end of August 2013. Search terms were: (attitud* or aware* or behave* or belief* or bias* or discriminat* or emotion* or experience* or feeling* or opinion* or perception* or perspective* or prejudice* or stereotyp* or stigma* or view*) and (down* syndrome or developmental* delay* or developal* disab* or intellect* challeng* or intellect* disab* or learning disab* or mental* deficien* or mental* handicap* or mental* retard*) and (medic* adj4 clerk* or medic* adj4 intern* or medic* adj4 school* or medic* adj4 student* or medic* adj4 undergrad* or medico or md student* or student doctor* or student physician*). Review process ~~~~~~~~~~~~~~ The authors discussed and established clear inclusion and exclusion criteria. They agreed to only include studies that investigated medical students’ attitudes towards people with ID and/or their healthcare. Given the limited amount of research on this topic, studies that used measures of attitudes to people with disabilities (i.e., studies that did not use ID-specific measures) to assess participants’ attitudes to people with ID were included, as were studies whose participants were a combination of medical students and professionals or other students. The authors agreed to exclude the following types of articles: examinations of medical students’ views on training in ID, which did not assess participants’ attitudes towards people with ID and/or their healthcare (e.g., Burge, Ouellette-Kuntz, Isaacs, & Lunsky, 2008; Burge, Ouellette-Kuntz, McCreary, Bradley, & Leichner, 2002); studies without a focus on ID (e.g., Beausoleil, Zalneraitis, Gregorio, & Healey, 1994; Wonkam, Njamnshi, & Angwafo, 2006); and research without medical students (e.g., Boyle et al., 2010; Parchomiuk, 2013). Then, the first author reviewed the literature. Nine hundred and thirty-six items were imported into Zotero and 377 duplicates were removed, leaving 559. After reading their titles and abstracts, 507 clearly irrelevant items were deleted. The remaining 52 articles were read in full, with 28 irrelevant articles removed after this examination. This process resulted in the retention of 24 studies that examined medical students’ attitudes towards people with ID. While the Critical Appraisal Skills Programme (CASP; 2013) checklist for evaluating qualitative work guided the review of Karl, McGuigan, Withiam-Leitch, Akl, and Symons (2013), the Cochrane Public Health Group's (n.d.) quality assessment tool informed the review of the twenty- three quantitative papers. The latter focused attention on the following topics: selection bias, allocation bias, confounders, blinding, data collection methods, withdrawals and dropouts, analysis, and intervention integrity. Overview of studies ~~~~~~~~~~~~~~~~~~~ Twenty-four articles published between 1968 and 2013 met the inclusion criteria, all of which reported on separate studies. Studies mostly were conducted in the UK (n = 9), followed by the USA (n = 8), Australia (n = 3), Ethiopia (n = 2), Canada (n = 1), and China (n = 1). Eighteen studies sampled medical students only (e.g., Hall & Hollins, 1996; Khandelwal & Workneh, 1987) and 6 used samples that included medical students and other groups (e.g., healthcare professionals; Handler, Bhardwaj, & Jackson, 1994). All studies used surveys (with closed and/or open-ended questions) to assess students’ attitudes; no focus groups or interviews were conducted. Twelve studies used a pre-test post-test design, 10 cross-sectionally analysed attitudes, 1 was experimental, and another was qualitative. Using the aforementioned critical appraisal tools, each study's strengths and limitations were determined. Strengths included low attrition rates and attention to inter-group contact theory (Pettigrew, 1998) to explain medical students’ attitudes. However, these strengths were offset by disadvantages. For example, most studies employed ad-hoc measures with questionable psychometric quality; no study blinded researchers to the intervention; and only Sinai et al. (2013) reported a power calculation. The studies are reviewed in the following sections and an overview is given in Table 1. Studies on attitude interventions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Findings suggested that interventions disparately affected attitudes; however, there were methodological concerns. Research suggesting minimal or no attitudinal change Sinai et al. (2013) investigated attitudes towards the community inclusion of persons with ID among fourth-year medical students in the UK. The students reported favourable attitudes and these remained unchanged after a 14-week neurosciences block that included ID teaching. However, results should be interpreted with caution. It is unclear if participants attended the teaching block, and self-selection bias may have influenced results as only 136 and 133 students completed the questionnaire beforehand and afterwards, respectively, despite 387 students invited to participate. An amended, shortened version of the Community Living Attitudes Scale–Mental Retardation (CLAS-MR; Henry, Keys, Jopp, & Balcazar, 1996) was used, whose psychometric properties have not been assessed. Also, mean imputation for missing data was employed, a strategy that should be avoided (Allison, 2001). Laking (1988) compared UK medical students who had, and had not, completed a course on ID psychiatry. A modified version of the Attitudes to Disabled Persons Scale (ATDP; Yuker, Block, & Campbell, 1960) was employed. Items were changed with “mentally handicapped” replacing “disabled,” which is poor psychometric practice because word substitution is unlikely to produce items that optimally measure the intended latent construct. Students were not randomly assigned to conditions (i.e., course completion or not) and there appears to have been a self-selection bias (i.e., most students who completed the course reported previous contact with this group, which may not be representative of medical students). Also, listwise deletion was used for cases that did not complete the ATDP, a suboptimal strategy for the management of missing data (Allison, 2001). The two groups reported comparable attitudes and Laking (1988) suggested that the ATDP might not be sensitive enough to detect changes in attitudes over time. May (1991) also studied ID teaching's impact on UK medical students’ attitudes. In general, most students supported the rights of this group; however, before teaching, only 42%, 33%, and 13% supported their rights to have children, leave home upon adulthood, and attend mainstream schools, respectively. Although students were more likely to support people with ID's right to attend mainstream schools after the intervention, results suggested that teaching typically did not improve attitudes. However, the “crude measuring instruments” (May, 1991, p. 241) might have been unable to capture attitudinal change. Research suggesting worsened attitudes Khandelwal and Workneh's (1987) study demonstrated that an intervention might deleteriously affect attitudes. They found that the attitudes of 100 Ethiopian medical students worsened after a six-week full-time course in psychiatry. The course covered various conditions including ID, with students completing a measure, designed by the authors, before and after. Participants’ responses suggested that, upon completion of the course, more students believed that people with ID were unable to work or marry. For example, beforehand, 35% of students believed it was impossible for someone with ID to get married; however, afterwards, this figure increased to 65%. The intervention's non-specificity to ID, and the assessment tool's narrow focus, may be limitations. Research suggesting improved attitudes: intellectual disabilities-specific measures Several studies reported that interventions led to self-reported improvements in attitudes among medical students (e.g., Fishler, Koch, Sands, & Bills, 1968; Hall & Hollins, 1996; May et al., 1994; Simeonsson, Kenney, & Walker, 1976; Thacker, Crabb, Perez, Raji, & Hollins, 2007). Using a sample of 12 American medical students (two did not complete post-test measures), Simeonsson et al. (1976) found that participants reported more positive attitudes towards people with ID after training on the topic. The authors also found more positive self-reported attitudes among participants that had better experiences of persons with ID. However, descriptive statistics only were given and psychometric support for their measure was not provided. Fishler et al. (1968) also researched American students (N = 36), finding that they were less likely to rate sterilisation and custodial as important areas in ID, and more likely to rate medical and psychological as important areas, after clinical experiences in the area. Despite these experiences, and contrary to Fishler et al.’s expectation, students’ advice regarding institutional versus home care for children with ID did not change. However, analyses may have lacked power due to the small sample. The effects of ID training on American medical students’ (N = 39) beliefs about people with ID's functionality also have been examined (Widrick et al., 1991). Scores on the Prognostication about Mental Retardation Scale (Wolraich & Siperstein, 1983) suggested that students were more optimistic about what people with ID can achieve after the intervention, with people with mild ID ascribed the greatest functional ability, followed by persons with moderate and severe ID, respectively. Students’ comments, which also were recorded, suggested that they believed the intervention and, in particular, meeting with this population, increased their expectations about people with ID. Boyd et al. (2008) examined the efficacy of an intervention that aimed to reduce 101 American students’ difficulty with working with people with developmental disabilities. Results suggested that the intervention, which involved training with a virtual patient, achieved a reduction in students’ perceived difficulty with providing care to this population. However, only four participants were medical residents, therefore limiting the relevance of this study to understanding medical students’ attitudes to people with ID. Hall and Hollins (1996) found that, among 28 medical students in the UK, attitudes towards people with Down's syndrome improved on 7 of 10 items after taking part in a workshop with actors with ID. For example, students were less likely to report that people with ID have little sense of humour and act like children most of the time. Thacker et al. (2007) used the same measure to examine a teaching intervention's effects on the attitudes of medical students in the UK towards people with ID. Again, the intervention involved actors with ID. Thacker et al. (2007) stated that, compared to 14 students who did not take part in the role- plays, the 26 students who did reported relatively positive attitudes. It was unclear whether the students were randomly allocated to attending or not, or if attendance was volitional. Further, neither Hall and Hollins (1996) nor Thacker et al. (2007) provided psychometric information about their measurement tool; thus, its reliability and validity are unknown, making the interpretation of results difficult. Research suggesting improved attitudes: generic measures Studies that used measures of attitudes towards persons with disabilities in general also suggested that ID teaching/training enhanced medical students’ attitudes (e.g., Tracy & Graves, 1996; Tracy & Iacono, 2008). However, such measurement is problematic as scales non-specific to ID may omit critical aspects of students’ attitudes towards this clinical group. Tracy and Graves (1996) examined whether an optional teaching unit on developmental disabilities influenced the attitudes of 25 Australian first-year medical students. At the beginning and end of the unit, students reported their thoughts and feelings towards people with disabilities and the patients’ families. Before teaching, 56% of participants expressed discomfort and lack of confidence working with people with disabilities, and 92% wanted to become more knowledgeable about the area. Afterwards, 92% reported that their attitudes had changed over the course of teaching, with qualitative comments typically suggesting attitudinal improvement and identifying inter-group contact as an important change mechanism. However, due to the measure's non-specificity to ID, it is possible that the students’ attitudes towards interacting with people with ID remained unchanged or worsened, whilst their comfort interacting with people with other disabilities increased. As measures’ psychological constructs should be specific to the research goals (DeVellis, 2003), the validity of such findings is questionable. Tracy and Iacono (2008) evaluated changes in 128 Australian fourth-year medical students’ attitudes towards interacting with people with disabilities after training on developmental disabilities and communication skills. The students completed the 20-item Interaction with Disabled Persons Scale (Gething, 1994), which measured discomfort interacting with persons with a disability, before and after the intervention. Results suggested that the students were more comfortable interacting with people with disabilities after the intervention, with 77% of students valuing the opportunity to meet people with disabilities during the intervention. However, as with Tracy and Graves (1996), these findings are difficult to interpret due to the measure's lack of specificity. Andrew, Siegel, Politch, and Coulter (1998) also used a generic measure of attitudes to those with disabilities in their evaluation of training, which included experiences with children with developmental disabilities. Little information was given about the chosen measurement tool and its psychometric properties are unknown; however, descriptive results suggested that students enjoyed and learned from the experience. Most students reported that their attitudes at least moderately changed, with 30% indicating unchanged attitudes. Attitude change was mostly attributed to a new awareness of family dynamics, and the most commonly reported behavioural intention arising from the intervention was a need for greater sensitivity when interacting with children with disabilities. Research suggesting improved attitudes: qualitative work Karl et al. (2013) qualitatively examined medical students’ written responses to an Internet survey on their reflections about a clinical experience, in which they met patients with developmental disabilities and worked with professionals in this area. A survey was used to avoid interviewer and response bias; however, the author did not describe consideration of the relationship between the researcher and participants as recommended by CASP (2013), and interviews or focus groups may have produced richer data. Results suggested that, after the intervention, students better understood the need to overcome communication barriers; were more comfortable caring for this population; and were more aware of diagnostic overshadowing and this group's right to equal healthcare standards. Cross-sectional attitudinal studies that did not evaluate interventions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ While cross-sectional research has provided snapshots of medical students’ attitudes towards this population, studies frequently lacked methodological rigour. Lennox and Chaplin (1995) used four attitudinal items to examine the attitudes of 128 psychiatric trainees and 27 medical officers in Australia. Despite 30% of participants reporting that they would personally prefer not to treat people with ID and a psychiatric disorder, the majority of participants endorsed the need to investigate psychiatric symptoms among persons with severe ID, and recognised the utility of psychotherapy for persons with ID and a psychiatric disorder. No information on item generation was provided, and a rationale for the inclusion of only four items was omitted. Li, Tsoi, and Wang (2012) found that 280 Chinese students of education or medicine reported comparably favourable attitudes towards the inclusion of persons with ID. Participants with more experience with this population, and females, reported more positive attitudes. However, the use of the Mental Retardation Attitude Inventory-Revised (Antonak & Harth, 1994) among Chinese people may be questioned because its factor structure was not replicated among a sample of Chinese people (Hampton & Xiao, 2008). Ouellette-Kuntz et al. (2012) found that 258 Canadian medical students with experience of people with ID were more likely than those without such experience to score higher on sheltering (e.g., the belief that this population should be protected). Further analysis revealed that 88.5% of those with experience typically reported meeting with five or fewer persons with ID. Thus, their experience and consequent understanding may have been limited (Ouellette-Kuntz et al., 2012). Supervision's salience to attitudes emerged, with those who reported positive supervision experiences scoring higher on the empowerment of people with ID, and lower on the need to protect them in the community (Ouellette-Kuntz et al., 2012), than students who reported negative experiences of supervision. Whilst interesting, this study may have been limited by the authors’ decision to use the CLAS-MR (Henry et al., 1996), as it only measures attitudes towards community inclusion and neglects a focus on medical students’ attitudes to providing healthcare to people with ID. Holt and Bouras (1988) used a short questionnaire based on McConkey and McCormack (1983) to examine 166 British medical students’ attitudes towards people with ID. Findings predominantly indicated that students held favourable attitudes towards this clinical group, with 10% saying that they wanted to work in services for people with ID and participants typically disagreeing that people with ID would always act like children. Although encouraging, results may be explained by students’ socially desirable responses and the measurement tool's psychometric qualities are unknown. Wishart and Johnston (1990) examined stereotypical beliefs about children with Down's syndrome among different groups of British people, including 10 medical students. The role of previous contact with this group also was studied. In general, participants with more experience were less likely to endorse stereotypes, and medical students reported less stereotypical beliefs than other groups, including mothers with children with Down's syndrome. However, the measurement tool's content validity is questionable, and no psychometric information was provided, reducing the interpretability of the findings. Prognostic beliefs among 136 medical students and 149 healthcare professionals in the USA also have received empirical attention (Handler et al., 1994), with students reporting lower expectations about people with ID than their qualified peers. Perhaps, counterintuitively, students’ beliefs were unrelated to having a family member with a disability or working with people with disabilities. Compared to medical students in earlier years, fourth-year medical students reported more optimistic beliefs about this group's potential. Students were most pessimistic about people with severe ID, followed by those with moderate ID, and lastly persons with mild ID. Khandelwal and Workneh (1986) used vignettes to assess 60 Ethiopian medical students’ attitudes to various conditions, including ID. Ninety-two per cent of students said the person with ID was ill; 62% regarded it as a very serious illness; and 20% said the prognosis would worsen. Only 7% reported that the person with ID had the same ability to marry as anybody else, while 82% and 92% said the person would have at least some difficulty living at home and working, respectively. Scott and Rutledge (1997) used an uncited ATDP to investigate the attitudes of 80 American first-year medical students to people with ID. The authors claimed the scale's reliability and validity when measuring attitudes towards those with disabilities; however, its specificity to ID and psychometric properties were not detailed. Scott and Rutledge suggested that scores on the ATDP indicated that most participants did not have negative attitudes towards people with ID. Most participants reported that they were willing to work with this population and believed that people with ID should live in the community. Experiment on attitudes ~~~~~~~~~~~~~~~~~~~~~~~ St. Claire (1993) examined the role of social identification among 7 doctors and 38 medical students in the UK. The author hypothesised that, compared to participants whose personal identities purportedly were activated; those with activated clinical identities would report more negative beliefs about people with ID and be more likely to attribute ID to children. Participants were randomly assigned to either condition and therefore received questionnaires titled, “Medical diagnosis and visual cues” or “Personality and person perception.” Participants in the clinical identity condition reported more negative beliefs than those in the personal identity condition, but people in both conditions were equally accurate distinguishing between children with and without ID. However, as a manipulation check suggested different social identities might not have been activated, this study's findings should be interpreted with caution.","This literature review identified 24 articles regarding medical students’ attitudes towards people with ID. The majority of the evidence reviewed consisted of evaluations of teaching/training interventions that sometimes resulted in improved self-reported attitudes. As these interventions often involved students interacting with people with ID (e.g., Hall & Hollins, 1996), findings are consistent with intergroup contact theory, which posits that contact between groups usually reduces prejudice (Pettigrew, 1998). Thus, opportunities for medical students to gain experience with this clinical group may be a key component of future attitudinal interventions. However, as recommended by Corrigan and Penn (1999), interventions to reduce stigma “should not be accepted on faith” (p. 765); instead, their theoretical underpinnings and empirical support warrant scrutiny. This point seems particularly salient, as ID stigma research has not used systematic approaches with conceptual models (Ditchman et al., 2013). To address this omission, future research may experimentally examine interventions characterised by intergroup contact under optimal conditions of equal status between groups, shared goals, cooperation between groups, and organisational support (Allport, 1954); high levels of intimacy between groups; and minimal differences between the persons with ID involved and their stereotype (Corrigan & Penn, 1999). As the number, frequency, and quality of contacts may be important (Morin, Rivard, Crocker, Boursier, & Caron, 2013), the roles of these variables should be assessed. Also, as students’ attitudes towards persons with ID may be associated with their supervision (Ouellette-Kuntz et al., 2012), future research may examine if quality of placement supervision moderates the effectiveness of interventions on students’ attitudes and future clinical behaviours. In line with other areas of ID research (Ditchman et al., 2013; Rose, Rose, & Kent, 2012; Werner, Corrigan, Ditchman, & Sokol, 2012), there is a need for scale development. Specifically, a measure of medical students’ attitudes to people with ID is needed if the efficacy of interventions is to be determined in a valid manner. As precise definitions of psychological constructs facilitate valid measurement (Eagly & Chaiken, 2007), the conceptualisation of medical students’ attitudes to persons with ID requires empirical attention. According to Eagly and Chaiken (2007), attitudes may be: (a) covert or overt; (b) cognitive (e.g., thoughts and beliefs), behavioural (e.g., intensions and overt actions), or affective (e.g., feelings and emotions); and (c) conscious or unconscious. Eagly and Chaiken (2007) described explicit and implicit attitudes, noting that the former represent evaluations reported by the person holding the attitude, and the latter represent spontaneous emotional reactions that the person may not be consciously aware of. As explicit and implicit attitudes may predict volitional and spontaneous behaviour, respectively, both warrant empirical attention (Eagly & Chaiken, 2007). Further, people may hold an explicit attitude and an implicit attitude towards the same entity, and each may be differentially affected by an intervention (Wilson, Lindsey, & Schooler, 2000). Thus, future research may wish to examine the effects of pedagogical interventions on explicit and implicit attitudes of medical students.","This review suggests that teaching and training may improve medical students’ attitudes, with interventions driven by intergroup contact theory (Pettigrew, 1998) holding promise. However, the review also identifies the need for more robust research to accurately understand (a) medical students’ attitudes towards people with ID and (b) the kinds of interventions that improve these attitudes. Attitude enhancement is the ultimate goal of research on ID stigma (Ditchman et al., 2013). Indeed, if tomorrow's doctors’ attitudes towards this population do not improve, efforts to reduce health inequalities experienced by people with ID (Emerson & Baines, 2010; Turner & Robinson, 2010) may well have limited success."],["With their Duplo task, Rubio-Fernández and Geurts (2013) challenged the assumption that children under 4 years of age cannot pass the standard false belief test. In an attempt to replicate this task on a sample of 73 children aged 32–51 months, we added a standard change of location false belief task as well as a Duplo true belief task. Performance on the latter is crucial for interpreting answers in the Duplo false belief task as to whether they reflect evidence for understanding or merely exhibit a difference in guessing rate. We found (a) a greater variability of response types in both Duplo tasks, (b) no evidence that responses in the Duplo tasks reveal earlier competence than those in the standard false belief test, and (c) a reassuring correlation between false belief tasks, suggesting that the Duplo task does pick up understanding of belief in light of the standard test. --------------------------------------------------------------------------------","The classical finding that the standard false belief test—also known as the “Maxi” task (Wimmer & Perner, 1983)—is passed by 4 years of age (Wellman, Cross, & Watson, 2001) has been challenged by the use of indirect indicators such as anticipatory looking (Clements & Perner, 1994; Southgate, Senju, & Csibra, 2007), looking time (Onishi & Baillargeon, 2005), and neural activity (Southgate & Vernetti, 2014; Kampis, Parise, Csibra, & Kovács, 2015). Although the replicability of these findings with very young infants from 6 months to 2 years has been questioned recently (Kulke & Rakoczy, 2018; see also special issue of Cognitive Development edited by Sabbagh & Paulus, 2018), the finding that correct anticipatory looking occurs just before children turn 3 years old has been replicated in seven of seven studies (Kulke & Rakoczy, 2018). This early evidence preceding verbal answers in the standard test has been attributed to an implicit understanding of belief (Perner & Clements, 2000). The justification for this claim resides in the finding that children show no knowledge of the agent’s belief in a direct test, but only in an indirect test, commonly seen as a criterion for implicit knowledge (Reingold & Merikle, 1993). When directly asked where a mistaken agent will go to get an object that was transferred without his or her knowledge, children answer with the object’s actual location when at the same time their eye gaze shows some awareness that the agent will go to where he or she thinks it is. Alternatively, the eye gaze might be evidence for explicit understanding that remains obscured by the test question due to processing limitations (Baillargeon, Scott, & He, 2010) or misleading pragmatics (Helming, Strickland, & Jacob, 2016). Findings from Rubio-Fernández and Geurts (2013, 2016) provide potentially important evidence for resolving this controversy. Their claim is that by lowering processing demands, children become able to respond correctly to a verbal command, which is deemed possible only with explicit knowledge. In Rubio-Fernández and Geurts's (2013) Duplo false belief (DFB) task, 3- and 4-year-olds are told about a girl who stores bananas in one of two boxes. While she is turning away, the bananas are transferred to the other box, the same as in the standard false belief (SFB) story except for the following four differences. First, the girl never went out of sight; she merely stepped to the side where she could not see the manipulation. Second, the experimenter transferred the bananas instead of introducing an additional story character. Third and fourth, when the girl returned, children were not told what she wanted and were not asked a question; they were simply requested to finish the story operating the girl doll themselves. This was to avoid any need to inhibit knowledge of the bananas’ real location, which is thought to be particular difficult in the SFB task (Baillargeon et al., 2010; Leslie, 1994; Setoh, Scott, & Baillargeon, 2016). Mention of the girl’s desire to get the bananas also creates a strong pull towards their real location because the girl wants to go where they are and not where she thinks they are (Perner, Rendl, & Garnham, 2007). Omitting the girl’s desire is, however, a risky undertaking. Without knowing what she wants, children may make her do anything, for example, go to the now empty box to put something else in there. They will make her go to where she believes the bananas are only if they assume that she wants to get the bananas. Because there is no indication of children making this assumption, it is essential that a Duplo true belief (DTB) control task is administered. Rubio-Fernández and Geurts’ (2013, 2016) findings were impressive. Nearly all children (mean age of around 3½ years) made the doll go to the full box in the DTB condition (91% in Experiment 1 of the 2013 study1 and 100% in Experiment 1 of the 2016 study). In contrast, 80% of children made the doll go to the empty box in the DFB condition, whereas only about 22% gave the correct answer to the SFB question (deceptive content task in Experiment 1 of the 2013 study). The data have important consequences for theories explaining early sensitivity to beliefs. First, the fact that children move the doll to the believed location of the bananas as early as they show anticipatory looking in the false belief task (Clements & Perner, 1994) speaks against an implicit knowledge of belief because intentional actions are typically carried out consciously. Second, the data question the explanation that the spontaneity of looking behavior accounts for earlier evidence over the elicited responses in the SFB task (Baillargeon et al., 2010) because children’s responses in the DFB task are also elicited. Third, the data speak against the pure processing load account by Baillargeon et al. (2010) that the SFB task requires children to understand the story plus the experimenter’s question. Although no specific question is asked, the experimenter nevertheless elicits a response with a range of more general questions. The data also support some proposed explanations. First, the fact that the belief-related responses decline drastically with the mention of the bananas’ location or the girl’s desire (Rubio-Fernández & Geurts, 2013) goes well with the claim that the lure of reality cannot be suppressed (Setoh et al., 2016). Second, these effects could also be due to children’s difficulty in switching from their own perspective, enforced by mentioning the bananas’ location or the girl’s desire, back to the girl’s perspective (Rubio-Fernández & Geurts, 2013; Helming et al., 2016), as opposed to children needing to inhibit their representation of reality. Given their theoretical importance, the Duplo results warrant a careful replication, especially given that the employed experimental design has some weaknesses to be ironed out. In particular, there was only one experiment in which children were randomly assigned to the DFB and DTB conditions (Rubio-Fernández & Geurts, 2016, Experiment 1). Unfortunately, no SFB task was administered to the same sample. The claim that 80% of belief-directed answers at 3 years 7 months of age could be made only by comparing them with performance on the SFB test from a different, slightly younger sample (3 years 5 months) in Rubio-Fernández and Geurts (2013). Arguing for a performance difference purely on age is risky given that the large meta-analysis of SFB tests by Wellman et al. (2001) suggests a non-negligible number of samples in the age range of 3 years 3 months to 4 years that showed a mean correct answer rate of 80% or higher (see Fig. 2 in Wellman et al., 2001). Existing replications are either conceptual (Białecka-Pikul, Kosno, Białek, & Szpak, 2019; Dörrenberg, Wenzel, Proft, Rakoczy, & Liszkowski, 2019; Rubio-Fernández & Geurts, 2016) or did not include a true belief control (Kammermeier & Paulus, 2018). Without a DTB task, Kammermeier and Paulus (2018) could not decide whether the approximately 20% more frequent correct DFB answers compared to SFB answers reflected earlier evidence from the Duplo task or a difference in guessing rate. This in turn led to an exchange of opinions (Paulus & Kammermeier, 2018; Rubio-Fernández, 2018) as to what the baseline performance of children who do not understand belief might be. We argue that this discussion can be settled only with an experimental inclusion of the DTB control because it ensures that children ascribe the intended goal—of getting the bananas—to the agent. The current study, therefore, is the first to combine a direct replication method with a true belief condition. On the group level, the DTB task allows us to see whether the majority of children intuitively do ascribe to this goal. However, if this is not as cogent as the data of Rubio-Fernández and Geurts (2013, 2016) suggest, we cannot assume that every empty box response in the DFB task reliably indicates false belief understanding. In this case, DFB responses should be assessed in relation to DTB responses on the individual level. Consequently, before investigating the interesting specific effects claimed to be at work, we decided to replicate the study in a within-participants design including the DFB, DTB, and SFB tasks. This has the advantage of getting a more precise estimate of understanding if children give the correct answers to the DFB task as well as the DTB task. Rather than differences from chance level, we took the DTB condition as the baseline for the interpretation of DFB performance. To get the DTB condition involved, we relied on the joint probability between correct responses on the DTB and DFB tasks. Hence, if the DFB task reliably detects false belief understanding earlier than the SFB task, we should find (a) a significant performance difference between DFB × DTB and SFB (at the group level, between participants) or (b) significantly better performance in the DFB task than in the SFB task in the subgroup of DTB passers (at the individual level, within participants). In addition, the design allowed us to see whether the DFB performance correlates with that of the SFB task.","Rubio-Fernández and Geurts (2013) tested whether the percentage of correct responses was above chance of 50%. They found a large effect size (Cohen’s h = 64) for their sample of n = 25. On this basis, we computed the necessary sample size for obtaining 80% power following the standard recommendation outlined by Cohen (1988) and implemented in the pwr package by Champely (2013). Although this calculation indicated a sample size of n = 15, we decided to test 75 children for the following reasons. First, it is inadvisable to use fewer participants for replications than in the original studies (Etz & Vandekerckhove, 2016). Second, to allow comparisons among the three tasks in case children needed to be excluded due to transfer effects of the within-participants design. A total of 77 children from six nursery schools and two recreational facilities (Toy Museum and Indoor- playground) in the city of Salzburg, Austria, volunteered for this study. Of this sample, 4 children needed to be excluded because 2 children were not cooperative, 1 child had comprehension problems, and 1 child had an inaccurate date of birth given to the experimenter (after correction, it turned out that the child was too old for the sample). The final sample consisted of 73 children (28 girls) aged 32–51 months with a mean age of 41.92 months (SD = 4.87). Parents gave written consent, and if they agreed the session was recorded. Children received a small toy for participating. Design ~~~~~~ Children participated in a DFB task, a DTB task, and a change-of-location SFB task. The order of tasks was counterbalanced with Latin square. To ensure a direct replication of both the DFB and DTB tasks (Rubio-Fernández & Geurts, 2013), the original banana story was always used as the first Duplo task,2 and to allow for a within-participants comparison, a parallel version (toy car story) was developed. Whereas color and position of the containers was kept constant, the direction of the objects transfer was counterbalanced.","Children were tested by a female experimenter in a separate room in the day-care center or recreational facility. They sat at a table with the children sitting at a right angle to the experimenter, and each session lasted about 10 min. A total of 30 children were accompanied by a parent (n = 18), an older sibling (n = 3), or an educator (n = 9). Caregivers were seated in a chair some distance behind the children and were carefully instructed not to interrupt the testing procedure. To avoid procedural differences between our tasks and those of Rubio-Fernández and Geurts (2013, 2016), Paula Rubio-Fernández kindly sent us her comments on a videotape of each condition.3 Her recommendations were implemented in our final versions; see detailed instructions in Appendix A and videotapes of the procedure in OSF (Open Science Framework), https://osf.io/feg6u/. Duplo tasks All props used in the two parallel versions of the tasks were Duplo toys. The main protagonist was a figure (boy or girl) of approximately 6 cm in size. For the containers, we used either two pink boxes (6.4 × 6 × 3 cm) with a door (blue and yellow) or two treasure chests (5.5 × 4 × 4 cm; green and brown). The transferred object was a bunch of bananas (4 cm) or a toy car (3.3 cm). At the beginning of the experiment, children were allowed to explore the toys used in the specific task. After a while, the experimenter suggested acting out a story with the toys. She then placed the containers approximately 20 cm apart from each other on the table, facing the child. The approximate distance between containers and the child was 30 - 40 cm. Lastly, the object was placed approximately 20 cm in front of the containers and 10–20 cm in front of the child. Introduction In the original version of the Duplo task, the experimenter told the child about how the protagonist loved eating bananas and, because she already had eaten a banana today, she stored them in one of two fridges. In the parallel version, the protagonist was a boy who loved playing with a toy car, which he then stored in one of two toy boxes. False belief condition After the object was put into the container, the experimenter explained that the protagonist now wanted to take a walk. The experimenter walked the Duplo figure to the opposite edge of the table, where it was then placed with its back to the containers. During the transfer of the object and when asking epistemic prompts [see (1) and (2) in the “Epstemic prompts” section below], the experimenter whispered and behaved conspiratorially; for example, she put her finger to her mouth and had a slight smirk on her face. True belief condition After the object was put into the container, the experimenter placed the Duplo figure facing both containers. Next, the transfer of the object and the epistemic prompt question (2) were acted out neutrally and without any conspiratorial cues. Announcing that the protagonist now wanted to take a walk, the experimenter walked the Duplo figure to the opposite edge of the table and back to the scene. Epistemic prompts To make the child attend to the epistemic states of the protagonist, the following questions were asked: (1) (before transfer) “Can she/he see me from there where she/he is standing?” and (2) (after transfer) “Did the girl/boy see what I did?” The DFB condition included both kinds of prompts, whereas the DTB condition included the after-transfer prompt only, and both questions were treated as prompts rather than as control questions. Thus, irrespective of what the child answered, the experimenter continued by saying “No, she/he cannot see me!” and “No, she/he did not see what I did!” in the false belief condition and “Yes, she/he saw what I did!” in the true belief condition. Prediction test After the experimenter had the Duplo figure come back from the walk, she placed it in front of the two containers and asked the child, “Do you want to play with the girl/boy now? What will happen next?” If the child did not respond spontaneously, the experimenter encouraged the child: “You can take the girl/boy now if you want to. What will she/he do next?” If the child still did not respond, the experimenter repeated the question: “What will she/he do next?”4 Traditional false belief task ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A PowerPoint animation of the standard change-of-location false belief story (Wimmer & Perner, 1983) was presented on a tablet and narrated by the experimenter. In the protagonist’s absence, a toy was transferred to a new location and children were asked to predict where the protagonist, who had returned to continue playing with the toy, will look for the toy (test question: “Where will Maxi look for the ball first?”). Seven control questions were asked to make sure that children had understood relevant story facts (“Where did Maxi put the ball?”, “Where is the ball now?”, “Who placed it there?”, and “Where did Maxi place the ball in the beginning?”) and could still remember them at the end of the story (“Where is the ball now?”, “Where did Maxi place the ball in the beginning?” and “Did Maxi see that his sister put it there?”). If children were not able to answer a question correctly, the story was repeated and the correct answer was given.","For each task some children needed to be excluded due to experimenter errors (DFB: 1; SFB: 1), caregiver errors (DFB: 2), lack of cooperation by the child (DTB: 3; DFB: 2; SFB: 2), ambiguous responses to the test (DFB: 2; DTB: 2), starting to enact the false belief scenario in the DTB task (5), and failing on 50% or more of the control questions5 in the SFB (6). The number of children excluded in each task can be seen in Table 1 together with children’s answers to the epistemic prompts (DTB and DFB) and their responses to the tests. Confirmatory analysis ~~~~~~~~~~~~~~~~~~~~~ In both Duplo tasks, children had problems in giving answers or giving correct answers to the epistemic questions. Following Rubio-Fernández and Geurts (2013, 2016), we treated these as prompts and not as questions to be answered correctly. In the test, a surprisingly large number of children did not give any response, and several said “don’t know” or did something other than moving the doll to one of the boxes (e.g., let the protagonist go for a walk). In line with Rubio-Fernández and Geurts (2013), children who did not move the doll to one of the two locations (full/empty) in the DFB and DTB tasks but showed other responses are treated as exclusions; hence, success rates are relative to the sum of definite box responses (Σ±). For the confirmatory analysis, we first assess children’s test reactions against chance level (two-choice binomial test, hypothetical probability of success = .50, two-tailed) and then compare performance in the DFB and SFB tasks between and within participants. In none of the tasks did performance differ from chance. Success rates in the DFB (n = 43), DTB (n = 37), and SFB (n = 64) tasks were 58% (p = .36), 62% (p = .188), and 41% (p = .169), respectively. For between-participants comparisons, we use children’s response to the first task administered (Table 2). A chi- square test with Yates correction revealed no significant difference in children’s performance on the DFB and the SFB tasks, χ2(1, N = 31) = 0.382, p = .458. A total of 39 children gave a valid (full or empty box) response in both the DFB and SFB tasks and, therefore, can be used for within-participants comparison (see left panel of Table 3). Neither on the individual level is there evidence for earlier understanding revealed by the Duplo test than by the standard test, exact McNemar, χ2(1) = 0.75, p = .387, odds ratio (OR) = 0.50, 95% confidence interval (CI) [0.11, 1.87]. Transmission effects To assess these results properly, we need to make sure that our within- participants design did not work against the Duplo tasks because tasks presented earlier might have had a detrimental effect on Duplo tasks later in the series. There is, however, no sign of such an effect. The percentage of correct answers did not differ from first, to second, to third positions for any of the tasks, all χ2(2) ≤ 1.08, p ≥ .583, ΦCramer ≤ .13. The same held true for Duplo story context (banana or toy car), χ2(1) ≤ 2.74, p ≥ .098, ΦCramer ≤ .27,6 and for direction of transfer in the Duplo tasks (left → right or right → left), χ2(1) ≤ 0.18, p ≥ .668, ΦCramer ≤ .07. In the original study (Rubio-Fernández & Geurts, 2013, Experiment 1), the DFB task always followed the SFB task. To control for a priming effect, we compare responses in the DFB task when it was administered immediately after the SFB task (n = 23) with when it was administered as the very first task (n = 20). There is no evidence for such an effect (DFB as first task: 7 empty, 4 full, and 9 other responses; DFB after SFB task: 11 empty, 8 full, and 4 other responses), χ2(2) = 3.96, p = .138, ΦCramer ≤ .30. Extended baseline Because there is no convincing argument for excluding from the analysis children who responded in an unexpected way (Kammermeier & Paulus, 2018), we also report the success rates relative to all responses (other responses also coded as incorrect) in Tables 1 and 2. Analysis of the DFB task in light of the DTB task As we argued, empty-box responses in the DFB task should be interpreted as evidence for understanding belief only if they correspond with full-box responses in the DTB task. Therefore, we compare the joint probability DFB × DTB with performance in the SFB task on both the group and individual levels. The within- participants design allows us to identify the DTB × DFB response pattern of individual children shown in the right column of Table 3. Unfortunately, only 24 children responded with a definite box in both tasks, and only 6 children showed the desired response pattern. On these grounds Table 3 provides no evidence of any understanding because as many children showed the opposite pattern. Correlations Despite the fact that Duplo task performance did not outstrip SFB performance, there was a reassuring correlation with age in months for the DFB task, rs(43) = .39, p = .009, 95% CI [.102, .618], and for the SFB task, rs(64) = .34, p = .005, 95% CI [.103, .541]. No significant correlation was found for the DTB task, rs(37) = − .16, p = .301, 95% CI [−.46, .173]. Furthermore, the DFB and SFB tasks are significantly correlated, rs(39) = .39, p = .014, 95% CI [.085, .628], which suggests that the Duplo task does pick up understanding of belief in light of the standard test. When children start to understand the concept of belief, as mirrored in their SFB performance, the rationale of the DFB task also becomes comprehensible. For SFB passers, the narrative of the DFB task becomes less ambiguous given that 20 of 26 children (77%) responded with a definite box, whereas only 19 of 33 SFB non-passers (58%) did so.","There were two main findings: (1) greater variability of children’s test responses than originally reported and (2) no evidence of earlier competence in the Duplo procedure than in the SFB test. We take the first of these findings to be a result of the open-ended nature of the Duplo task, leaving children with a wide field of interpretations, thereby making them susceptible to small environmental cues. In our sample, more than a third of children (35% in the DFB task and 41% in the DTB task) either said that they did not know what to do (5% and 13%, respectively), continued the story idiosyncratically (15% and 13%, respectively), or did not respond at all (15% and 16%, respectively). Response variability was also larger in Kammermeier and Paulus (2018) than in the original studies. Those authors argued that the problem is the open response format paired with a lack of control questions. This is likely given that we cannot be sure whether children understood the task as intended by the experimenter. For instance, in the Duplo task, children might assume that the Duplo girl wants to either look for her bananas or look for an empty box for something else. Dörrenberg et al. (2019) managed to make this apparently clearer with 94% and 85% correct answers for the DTB task.9 However, the more consistent performance in the DTB task did not produce earlier evidence for understanding belief over the SFB task in the DFB condition. Another potential reason for the high variability in children’s reactions is that responding to the Duplo task requires children to switch perspectives.10 This feature distinguishes the Duplo task from other early false belief paradigms in which children take the same third-person perspective throughout the task. When asked to continue the story, they switch from being a third-person observer of the toy agent to being a first-person controller of that agent. The interpretation of responses is based on the supposition that children incorporate what they have learned about the agent (as a third-person observer) into their play (as a first-person actor). It is not explicitly controlled whether or to what extent children actually do this. Idiosyncratic and “don’t know” responses underline this problem, and there is no assurance that this does not also occur in empty-box and full-box responses. Hence, the switch between perspectives is another potential source for uncontrolled variability. More important, assessing performance on the DFB, DTB, and SFB tasks on the same sample of children allows us to interpret children’s responses without needing to rely on the average age of different samples. Only children who gave correct responses in both the DFB and DTB tasks can be claimed to have some understanding of belief. However, there is no evidence on the individual level, nor is the joint probability of correct DTB and DFB answers significantly different from the proportion of correct SFB answers. Thus, there is no support for the claim that the Duplo task provides earlier evidence of understanding belief than the SFB task. In contrast, our joint probability is significantly lower than the corresponding values computed from the data reported by Rubio-Fernández and Geurts of .73 in 2013 and .80 in 2016. Both of them are clearly beyond the borders of the confidence interval (.24–.66). On these grounds, our data also fail to replicate the theoretically essential result of Rubio-Fernández and Geurts (2013, 2016); the data are significantly different from theirs and provide no evidence for their theoretical claim. The volatility of response rates highlights a potentially general problem of early false belief studies. In looking time and anticipatory looking paradigms, infants are not told a story, and thus—as in the Duplo task—it is not clearly communicated what the story agent wants. Children need to figure it out from the agent’s behavior in a few familiarization trials, which in many cases might not be sufficient. This might be one reason why the data from these techniques have been difficult to replicate."],["We collected short video clips of speakers and created five types of stimuli: (1) the original videos, (2) the audio tracks only, (3) single pictures only, (4) speech content, and (5) stick-figure animations displaying body motion. Participants rated these stimuli on a brief Big Five personality inventory. We then used ratings of the incomplete information conditions to predict ratings of the original video condition. Impressions in the audio track condition were strong predictors throughout all trait ratings. However, other cues were also non-negligible contributors to an overall impression. People even make sense of parsimonious cues, e.g., an animated stick-figure. Thus, presenters on a public stage are not only judged by what they say but also by how they move. --------------------------------------------------------------------------------","When people form impressions about others, words often seem to affect them less than the observed outward appearance and nonverbal behavior. Visual and auditory cues affect people’s judgments of their interaction partners; judgments that are made spontaneously, effortlessly, and without conscious processing (Ambady, Bernieri, & Richeson, 2000; Sunnafrank & Ramirez, 2004). On the one hand, first impressions can be misleading and a source of prejudices and stereotyping (Zebrowitz & Montepare, 2005). On the other hand, they can pick up relevant information about one’s social environment. After being exposed to brief extracts of nonverbal or verbal information, people are able to assess other people’s actual personality, their job performances, or a CEO’s abilities to generate company profits (Albright, Kenny, & Malloy, 1988; Ambady et al., 2000; Borkenau & Liebler, 1992a,b; Hecht & LaFrance, 1995; Rule & Ambady, 2008; Scherer, 1978). Irrespective of whether they are the key to someone’s actual personality traits or abilities, snap judgments can have a strong impact on impression formation and decision making. This is greatly important for those who enter the public arena. Politicians and leaders who vie for media attention and try to win the approval of an audience have to be aware that they are not only judged by the content they present. Nonverbal and salient cues are assumed to be processed efficiently and easily remembered and for this reason they can dominate over verbal information (Clark & Paivio, 1991) Their impact may even be more prevailing nowadays because news reports have been undergoing a shift from political content to image bites (see Bucy & Grabe, 2007; Stewart, 2010). Moreover, the flood of information that people are confronted with daily imposes an additional cognitive load, which increases the propensity to take mental shortcuts when making decisions (Olivola & Todorov, 2010). Thus, people may tend to choose their leaders not after careful deliberation but on the basis of superficialities. Empirical studies underscore that appearance cues and nonverbal behaviors influence judgments of politicians and other leaders. Research on the perception of charisma showed that potential leaders who display expressive non-verbal behaviors (i.e., more body gestures, more variations in intonation, more eye-contact, etc.) are seen as more charismatic than persons who showed less non-verbal behaviors (Awamleh & Gardner, 1999; Gardner, 2003; Holladay & Coombs, 1993, 1994). Moreover, experiments using manipulated voices revealed that vocal cues affect people’s attributions of leadership qualities and their voting behavior (Klofstad, Anderson, & Peters, 2012; Tigue, Borak, O’Connor, Schandl, & Feinberg, 2012). People even read leadership qualities, such as, competence, trustworthiness, or dominance, into photographs of political candidates. Interestingly, the consensus among such ratings is strong enough to make them reliable predictors of hypothetical voting decisions and actual election outcomes (Antonakis & Dalgas, 2009; Ballew & Todorov, 2007; Banducci, Karp, Thrasher, & Rallings, 2008; Little, Roberts, Jones, & DeBruine, 2012; Olivola & Todorov, 2010). All these empirical findings point to nonverbal cues as the prevailing influence on the perception of leaders and politicians. However, some researchers who compared the relative impact of visual, vocal, and verbal information on judgments of politicians came to different conclusions (Krauss, Apple, Morency, Wenzel, & Winton, 1981; Nagel, Maurer, & Reinemann, 2012). They found that speech content dominates over nonverbal information. This does not undermine the role of nonverbal behaviors in human communication but indicates that their influence varies with the situation in which behaviors are performed as well as with audience motivation and involvement (Allwood, 2002; O’Sullivan, Ekman, Friesen, & Scherer, 1985; Petty, Cacioppo, & Schumann, 1983). In summary, regardless of on which communication channel (i.e., vocal, visual, or verbal) people’s first impressions are based when making inferences of social relevance, research clearly shows that “thin slices” of behavior or appearance cues can have a strong impact on which traits and abilities people read into their social environment. Motion cues as social information ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Body motion is a form of nonverbal channel that comprises hand gestures, movements of the head, or position shifts of the whole body. Although the current study provided data on the interplay of different communication channels in impression formation we mainly focus on the relative role of body motion. Empirical studies have found that people are very adept at extracting information from motion cues. Even abstract stimuli such as circles or triangles flitting around on a screen, are often interpreted as animal or human behavior and elicit attributions of intentionality and personality (Heider & Simmel, 1944; Koppensteiner, 2011; Scholl & Tremoulet, 2000). Research on human body motion was strongly influenced by the point light approach, which was introduced by Gunnar Johansson (1973). He and his colleagues attached point lights and reflective markers to persons’ major joints and filmed them so that only a set of dots were visible in the resulting movies. Supporting the role of motion in human perception, these kinds of stimuli only became human-like when the dots were moving, whereas observers who saw single pictures of the movies perceived nothing but a random distribution of dots. Other researchers were inspired by Johansson’s methodical approach and demonstrated that such point light displays contain enough information to make quite accurate guesses of other people’s age and sex (Dittrich, Troscianko, Lea, & Morgan, 1996; Kozlowski & Cutting, 1978; Montepare & Zebrowitz-McArthur, 1988; Troje, 2002). Research using more elaborated versions of this technique or alternative methods of motion capture revealed that motion cues convey information of social relevance (Blake & Shiffrar, 2007; Chouchourelou, Matsuka, Harber, & Shiffrar, 2006). People appear to be able to perceive affect in arm movements (Pollick, Paterson, Bruderlin, & Sanford, 2001), emotions in movements of the whole body (Atkinson, Dittrich, Gemmell, & Young, 2004; Clarke, Bradshaw, Field, Hampson, & Rose, 2005), and personality in patterns of human gait (Thoresen, Vuong, & Atkinson, 2012). The frequency and duration of motion and other kinematic features play a role in mating behavior (Bente, Donaghy, & Suwelack, 1998; Grammer, Honda, Juette, & Schmitt, 1999) and affect the way females judge the attractiveness of male dancers (Neave et al., 2011). Moreover, self- ratings and observer-ratings of personality on scales measuring sensation seeking or the Big Five personality dimensions are related to the motion behavior of dancers (Bechinie & Grammer, 2003; Hugill, Fink, Neave, Besson, & Bunse, 2011; Luck, Saarikallio, Burger, Thompson, & Toiviainen, 2010). In the domain of politics, variations in body motion influence how people judge the personality and health of politicians (Kempter, 1998; Koppensteiner, 2013; Koppensteiner & Grammer, 2010; Kramer, Arend, & Ward, 2010). Such results clearly show that motion cues are an equally important nonverbal communication channel as appearance cues (e.g., clothes), vocal cues (e.g., voice pitch) and facial expressions. The speed, the duration, and the flow of a gesture or variations in the movements of the whole body seem to have a non-negligible impact on how people form first impressions of their social environment. The present study ~~~~~~~~~~~~~~~~~ Human communication works on different levels ranging from symbolic information mostly conveyed by verbal content to information that has no definite signal character and is expressed by certain qualities of motion. The studies of Mehrabian (1972) and those on the perception of charisma (see above) provide evidence that under some experimental conditions, visual information exerts a dominant influence in impression formation followed by voice quality and speech content on the last position. In addition, comparison of trait judgments based on full channel information (i.e., video with speech) with trait judgments based on incomplete information only (e.g., silent videos or voice only) indicates that cues from different communication channels convey redundant information (Borkenau & Liebler, 1992a; Friedman, Oltmanns, Gleason, & Turkheimer, 2006). However, there are variations. Some traits appear to be preferably ascribed to visual cues, while other traits are preferably ascribed to cues from other modalities (Friedman et al., 2006; Gifford, 1994; Naumann, Vazire, Rentfrow, & Gosling, 2009; Zebrowitz-McArthur & Montepare, 1989). Moreover, when asked to identify emotional expressions or statements of agreement and disagreement in political debates people are more accurate when multimodal information is available (Bänziger, Mortillaro, & Scherer, 2012; Mehu & van der Maaten, 2014). In this study we examined to what degree body motion affects social judgments relative to information from other verbal and nonverbal communication channels. To accomplish this, we broke down short video clips of politicians making a speech into five different versions of stimuli. Independent samples of participants then judged either the full channel version of the speeches, sound only, single pictures taken out of the speeches, the content of the speeches read by a computer voice, or stick-figure animations displaying the body movements of the speakers. Measures of the participants’ first impressions were obtained using a brief questionnaire that builds on the five-factor model of personality (i.e. Big Five), because the five-factor model had already been successfully applied in numerous “thin slices” studies (e.g., Borkenau & Liebler, 1992a,b; Friedman et al., 2006). These measures allowed us to estimate to what extent information from different communication channels influences snap judgments of the speakers and how different “portions” of nonverbal and verbal information are related to body motion. Our study extends previous research in several aspects. First, other studies on nonverbal communication often instruct actors to display specific behaviors. In contrast to that, the stimuli we used were real politicians that had given their speeches in the German parliament. Hence, ecological validity of the displayed behaviors is high. Second, although our experimental design was inspired by previous studies on the role of different communication channels (Borkenau & Liebler, 1992a,b; Friedman et al., 2006), these studies did not examine the behavior of speakers in a public arena. Finally, we applied other tools for stimulus preparation and stimulus presentation than other studies in the field. By translating the body movements of the speakers into animated stick figures, we diminished the influence of confounding variables and were able to determine for which traits body motion was a strong predictor and to what degree it interacts with information from other nonverbal and verbal sources. In previous work we already used stick-figure animations as stimuli and related data-driven descriptors of body motion (e.g., measures of amplitude height) to judgments of personality. This revealed that people form first impressions on the basis of simple cues embedded in the behavioral stream. However, it was unclear whether and to what extent body motion affects judgements of speakers when vocal and other cues are also available to observers (Koppensteiner & Grammer, 2010). The work presented here, which is based on the same method of motion capturing but on a new set of stimuli, is a first step to overcome such limitations. It is suggested that rapidly evaluating other people’s intentions, assessing one’s social environment on broad trait categories and making fast decisions on whether to approach or avoid an individual has an adaptive function (e.g., Buss & Greiling, 1999; Oosterhof & Todorov, 2008). Misjudgments can have negative consequences and for this reason people may base their social evaluations on cues from different communication channels because this makes evaluations more reliable. Previous work has already shown that first impressions of extraversion, agreeableness and emotional stability are related to motion cues and gesturing (Borkenau & Liebler, 1992b; Gifford, 1994; Koppensteiner, 2013; Koppensteiner & Grammer, 2010). Hence, we hypothesized that for these personality dimensions judgments of the speakers’ body movements (i.e., stick-figure animations) can serve as predictors of people’s judgments in the full channel condition (i.e., video plus sound). Moreover, we expected pronounced links between ratings of the voice only condition and ratings of body motion, because it is well established that gesturing accompanies speech (e.g., McNeill, 1985). Appearance cues and vocal cues have been found to be related to all dimensions of the Big-Five (Borkenau & Liebler, 1992b; Friedman et al., 2006; Naumann et al., 2009). Consequently, we expected to replicate such findings in our study. Research on nonverbal communication revealed speech content to play a minor role when people form first impressions (e.g., Awamleh & Gardner, 1999). However, there seems to be a link between conscientiousness and speech content (e.g., Friedman et al., 2006), which we also expected to reveal in our data. Stimulus preparation and experimental conditions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Sixty speeches taking place in the German parliament (30 male and 30 female speakers) were randomly selected out of three parliamentary sessions (November 29 & 30 and December 14, 2012). The selected speakers were non-prominent members of the parliament and were unknown to our participants. At random, we extracted video segments with a length of 15 s from these speeches. Afterwards, the video segments were used to create five different stimulus types for the experimental conditions described below. Full channel condition For the full channel condition we merely removed the lower portion of the video clip, which gave information about the speaker’s name and party affiliation. Consequently, in this experimental condition the participants of a rating experiment assessed the original video clips, which included all visual and all auditory information. Static visual condition To eliminate information on body movement we chose one picture out of the frame series of each video segment. Each of these pictures showed the speakers adopting similar body postures. For more detail, we chose pictures with the speakers standing upright having their arms positioned at their sides. In accordance with the length of all other stimulus types pictures were shown for 15 s in the rating experiment. Original voice condition For this experimental condition, we extracted the audio track from each original video segment. Thus, visual information was removed and the participants assessing such stimuli only listened to what was being said. Artificial voice condition We extracted what said during the selected video segments from the transcripts of each parliamentary session. These transcripts were converted into audio files using the document reader software Ghostreader. Because all transcripts were read aloud by the same computer voice during the rating experiment, they served as the control condition for speech content and wording. Stick figure condition The main focus of the study was on the analysis of body motion in relation to information presented in the other conditions. To this end we created stimuli, which display body motion without confounding information, e.g., appearance cues. To accomplish this, we used a program that allows running through the video clips step by step. In the first frame of each video clip, we positioned landmarks on the speaker’s forehead, the hollow of the throat between the collar bones, ears, shoulders, elbows, hands, a spot in the middle of the body near the navel, and the corners of the lectern. Motion behavior occurring between single frames of the video clips was traced by rearranging the landmarks with the computer-mouse and software routines that automatically tracks positions shifts (Koppensteiner, 2013). Based on this data, we created stick figure animations, which served as abstract representations of the speakers’ body movements. To reduce the workload during the encoding process, we only used every third frame and filled in missing frames by linear interpolation.","A total of 308 Caucasian participants (i.e., students of the Faculty of Life Sciences) was recruited at the University of Vienna for taking part in the rating experiments. Sixty- five persons (37 females and 28 males; M-age = 23.3 years, SD = 5.8) participated in the full channel condition, 60 people (35 females and 25 males; M-age = 22.8 years, SD = 3.4) in the static visual condition, 63 (33 females and 30 males; M-age = 23.8 years, SD = 4.4) in the original voice condition, 60 participants (35 females and 25 males; M-age = 23.4 years, SD = 4.5) in the artificial voice condition, and 60 persons (33 females and 27 males; M-age = 22.5 years, SD = 3.7) in the stick figure condition. Participants received a financial compensation of €5.","Participants were approached in person and asked to take part in a short rating experiment. After accompanying them to our laboratory, we informed the participants briefly about the experimental set-up and the task to be accomplished. Subsequently, participants performed the rating tasks on their own using a computer-controlled interface. They were instructed to rate the stimuli after they had watched or listened to them. Pictures in the static visual condition were shown for 15 s, but participants were told that they could start their ratings whenever they felt to be ready for them. In the full channel condition, the original voice condition, and the artificial voice condition, participants wore headphones (AKG K 272 HD) connected to the computer to optimize sound quality. Stimuli were presented on the left-hand side of the user interface, while rating scales were displayed on the right hand side. Participants completed their ratings by dragging a slider to the favored position between the right pole (i.e., named, strongly disagree) and the left pole (i.e., named, strongly agree) of the scale using a computer mouse. The slider position corresponded to a position on a scale divided into 200 subunits. All ratings started with the slider being in the neutral position (i.e., middle position on the slider bar). To examine personality ratings of the speakers, we used the German version (Muck, Hell, & Gosling, 2007) of Gosling, Rentfrow, & Swann’s (2003) Ten Item Personality Measure (TIPI). This questionnaire is based on the five-factor model of personality and covers personality dimensions openness, extraversion, conscientiousness, emotional stability, and agreeableness. In each experimental condition, each participant rated a subset of 20 stimuli that were randomly selected from the 60 stimuli available. Statistical analysis Trait ratings were averaged for each condition and each stimulus. This yielded a dataset containing 60 speakers that were rated on ten items in five experimental conditions. Corresponding items of the TIPI questionnaire were turned into the Big Five personality dimensions according to the instructions by Gosling et al. (2003). Because each personality dimension only comprised two items, internal consistency of the questionnaire was determined by calculating the Spearman–Brown coefficient (Eisinga, Te Grotenhuis, & Pelzer, 2013). To estimate the relative influence of voice, speech content, appearance, and body motion on judgments of personality, we correlated ratings collected in the experimental conditions presenting incomplete information with ratings of the full channel condition (i.e., original video clips). We also performed multiple regression analyses using ratings in the full channel condition as criterion, ratings of the incomplete information conditions as predictors and the speakers’ sex as control variable. We expected these predictors to be intercorrelated and distorted due to multicollinearity. For this reason we also calculated so-called relative weights. Relative weights or relative importance weights range between 0 and 1, are unaffected by multicollinearity, and therefore very helpful when interpreting each predictors’ relative contribution in the regression model (Johnson, 2000; Kraha, Turner, Nimon, Zientek, & Henson, 2012; Lorenzo- Seva, Ferrando, & Chico, 2010). Moreover, bivariate correlations between predictor variables provided insight into the interrelations between ratings of the incomplete information conditions. Previous studies comparing different communication channels found correlations of .35 or higher between nonverbal cues and personality ratings (Borkenau & Liebler, 1992b; Friedman et al., 2006; Koppensteiner, 2013). We used different modalities as predictors in a regression analyses and assumed their effects to add up. Consequently, we expected the multiple regressions we performed to explain at least 20 percent of overall variance. On the basis of such an effect size, an alpha level of .05, a power level of .8, and five predictors an a priori power analyses suggested an optimal sample size of 51 stimuli. All statistical analyses were carried out in the program R (R Core Team, 2013).","Descriptive Statistics are presented in Table 1 and internal consistencies between corresponding items of the questionnaire in Table 2. Coefficients ranged from .06 to .94 and thus were unacceptably low in some cases. In particular, the item pairs anxious–easily upset and calm–emotionally stable did not combine properly in the stick figure, in the original voice, nor in the full channel condition. This might have been due to a weakness of the questionnaire on this dimension or to the choice of stimuli. Politicians on a public stage may rarely show body movements or produce vocal cues in which people perceive qualities associated with the items anxious and easily upset. To sum up, internal consistencies we obtained for emotional stability were very low and for this reason we discuss results for this personality dimension on the basis of single items (Tables 3–5). Statistical analyses of the relationships between incomplete information conditions and the full channel conditions are presented as correlation coefficients, relative weights, and regression estimates (see Tables 3 and 4). Results for extraversion revealed a strong relationship of the ratings in the stick-figure condition and the original voice condition with the ratings in the full channel condition. In addition, there was also a noteworthy relationship between ratings in the static visual condition and the full channel condition. These relationships were reflected in all types of analyses we applied (i.e., see correlations in Table 3, and estimates and relative weights in Table 4). We therefore came to the conclusion that in our setting, judgments of extraversion were mainly affected by the speakers’ voice and by their body motion as well as—to minor degree—by the speakers’ appearances. The content of the speeches had no important effect on ratings of extraversion. Perceived conscientiousness in the full channel condition yielded high correlation coefficients with ratings of conscientiousness in the static visual condition, in the original voice condition, and the artificial voice condition (i.e., computer voice presenting speech content). Similar results were obtained for estimates of the multiple regression and the relative weights (Tables 3 and 4). Consequently, impressions of conscientiousness were mostly guided by vocal information. However, appearance cues and speech content also played a role. Body motion was revealed as the weakest predictor. Ratings of perceived openness in the full channel condition were strongly related to ratings in the original voice condition and the static visual condition (Tables 3 and 4). Correlation coefficients also provided a noteworthy effect size for speech content (Table 3). In conclusion, participants mainly tended to ascribe openness to the speaker’s appearances and their voices. Full channel ratings of agreeableness showed strong bivariate correlations with ratings in the original voice condition, ratings in the stick- figure condition, ratings in the artificial voice condition and a noteworthy one for ratings in the static visual condition (see Table 3). These patterns of relationships were not in accordance with the ß-weights of the multiple regression (Table 4). It only provided strong estimates for the static visual condition and the original voice condition. This was an indicator of multicollinearity, which distorts ß-weights and, therefore, undermines accurate interpretation of the data. Relative weights, which are unaffected by multicollinearity, replicated the patterns found by the bivariate correlations (Table 3). Thus, we concluded that ratings of agreeableness were guided by appearance cues, by body motion, by vocal cues and by the verbal content the speakers presented. Due to the low internal consistencies results for emotional stability were analyzed on the level of single items. The item pair calm–emotionally stable yielded notable correlation coefficients and relative weights for all conditions, whereas regression estimates labeled ratings of the original voice condition as strong predictors and ratings of body motion and appearance as predictors with a moderate influence (Tables 3 and 4). The item pair anxious–easily upset provided a strong correlation of full channel ratings with original voice ratings and noteworthy ones with ratings of static visual cues and with ratings of body motion (Table 3). A similar pattern was provided by the regression estimates and the relative importance weights (Table 4). Analyses of the interrelations between ratings in the incomplete information conditions revealed strong correlations between voice and the content the speakers presented for nearly all personality ratings (Table 5). This is not very surprising, because the voice only condition comprised nonverbal vocal cues as well as the content of the speeches. We also found strong relationships between body motion (i.e., ratings of stick-figure animations) and voice for the personality dimensions extraversion and agreeableness (Table 5). This complements results obtained in the regression analyses. It shows that body motion was not only a communicator of agreeableness and extraversion but also is accompanied by vocal cues that convey similar information. Thus, for judgments of some personality dimensions there appears to be a coupling between voice and body motion. Perceived extraversion and conscientiousness showed notable relationships between different communication channels (Table 5). For instance, we found a link between appearance cues and speech content for these personality dimensions. This hints that people are able to form expectations about what politicians will say on the basis of their appearance or form expectations about their appearance on the basis of what they say. Original voice condition ~~~~~~~~~~~~~~~~~~~~~~~~ The most prominent finding of our study was that all personality ratings of the original voice condition, in which the participants only listened to an audio track of the selected speeches, were strongly related to corresponding ratings of the full channel condition, in which participants watched the unaltered original video clips. This was mostly due to paralinguistic cues such as intonation and pitch because the strong impact of the speakers’ voice on personality judgments did not disappear when speech content was controlled. Other studies also revealed that vocal cues are an important influential factor in impression formation (Borkenau & Liebler, 1992a; Friedman et al., 2006). However, they did not play such a predominant role as in our study. Participants were, of course, aware that they were judging politicians. For this reason, they might have attended more to the speaker’s voices than to other information. The majority of studies that compared the impact of different communication channels did not use politicians for their experiments or instructed actors to display certain behaviors (e.g., Awamleh & Gardner, 1999; Borkenau & Liebler, 1992a; Friedman et al., 2006). This is different from our study design and might explain why vocal cues were a strong contributor. Actors may overdo their acting and display stereotypical behaviors thereby drawing attention to cues that are less salient when real politicians present themselves. On the other hand, judging politicians may raise different expectations than judging people drawn from an average population. Speech content ~~~~~~~~~~~~~~ Judgments in the original voice condition did not only convey speech content but also information about intonation, voice pitch, and other qualities of the voice. To separate such vocal cues from speech content, we also collected ratings of the speeches’ wording read by a computer voice and related these ratings to ratings of full channel condition. As expected full channel judgments of conscientiousness, but also judgments of agreeableness and one item pair of the personality dimension emotional stability (i.e., calm, emotionally stable) could be predicted by speech content to a certain degree. In other words, participants read some traits and qualities into what the speakers said and not only how they said it. Unlike some other studies we were unable to show that verbal content dominates over nonverbal cues in judgments of politicians (Krauss et al., 1981; Nagel et al., 2012). Nevertheless, the results obtained indicate that people integrate speech content when forming first impressions. Static visual information ~~~~~~~~~~~~~~~~~~~~~~~~~ Static visual information, which was presented as single pictures, yielded notable relationships with the full channel condition. Previous research has already shown that appearance cues alone—clothing styles, physiognomic features and facial expressions, just to give a few examples—affect people’s first impressions (e.g., Naumann et al., 2009). We have supported such findings by demonstrating that static visual cues contribute to overall impressions of a person. Consequently, people form an impression on the basis of static cues before any motion behavior occurs or a word is spoken. The documented effects were particularly strong for ratings of conscientiousness and openness. Participants had no information about party membership, but static visual cues may already reflect a conservative or a more progressive mindset and this may more strongly guide attributions of openness and conscientiousness than attributions of other personality dimensions. Future studies including facial expressions, physiognomic features, and clothing style as variables in the analyses could support such assumptions. Body motion ~~~~~~~~~~~ The study’s main focus was to estimate the relative role of body motion when people judge politicians giving a speech. To accomplish this, we translated the speakers’ body motion into animated stick-figures. Full channel ratings of extraversion, agreeableness, and an item of emotional stability were strongly related to corresponding ratings of these abstract stimuli. Although other communication channels, in part, showed stronger effects over a wider range of ratings, it is still surprising that a parsimonious cue, as simple as a stick-figure, can guide social perceptions. In previous research, where we applied the same method of motion capturing, we revealed that people associate personality traits with salient and simple nonverbal cues that are embedded in the behavioral stream (Koppensteiner, 2013; Koppensteiner & Grammer, 2010). In accordance with the results obtained here we found strong relationships of certain motion descriptors with ratings of agreeableness and extraversion. For instance, speakers who produced expansive vertical movements (i.e., mostly up and down movements of the hands) were judged as highly extraverted while speakers who showed less such vertical movements and a greater overall variation in motion amplitudes tended to be judged as agreeable. Strong effects were also found for emotional stability. Speakers who displayed jerky movements were perceived as less emotionally stable than speakers who displayed smoother movements. This might explain why one item of emotional stability also produced pronounced effects in this study. We did not examine body movements on the level of motion cues in the work presented here. Consequently, the next step in this line of research would be to use motion descriptors such as those presented in our previous work and test in what way they are related to cues from other communication channels (e.g., auditory cues). Extraversion is often attributed to conspicuous and expressive behaviors (Kenny, Horner, Kashy, & Chu, 1992). For this reason, it is a trait that may be easily read from simple behavioral cues, including body motion. Similarly, appearing friendly is crucial when approaching or avoiding a stranger and therefore cues of agreeableness may also be easy to detect and even visible in body motion. Politicians often attack their opponents, make points with great vigor, and underline their arguments with expressive gestures. Trait ratings belonging to emotional stability—e.g., calmness—might be influenced by these expressive displays; this might explain why the participants of our experiments can use body motion for such attributions. Perceived agreeableness and extraversion of the stick-figure animations showed a strong coupling to perceived agreeableness and extraversion in the voice only condition. This underlines that speech is accompanied by body motion and gesturing (McNeill, 1985) and indicates that cues people associate with extraversion and agreeableness convey essential social information. To sum up, like other research in the field, our findings indicate that people perceive social meaning in a great variety of visual, auditory, and verbal information. In addition to this, we demonstrate that, for some important traits, abstract animations displaying the speaker’s body movements predict people’s first impressions.","We randomly extracted short video clips from speeches that were given in the German parliament. These clips were then turned into five different types of stimuli, which were rated on the Big Five personality taxonomy. Stimuli were either (1) short extracts of the original videos, (2) audio sequences of the speeches, (3) speaker’s body movements turned into stick-figures, (4) single pictures of the speakers, or (5) the content of the speeches read by a computer voice. To estimate the influence of these different communication channels on impression formation we related ratings in the incomplete information conditions to ratings in the original full channel video clip condition. The role of different communication channels ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ People make sense of brief displays of nonverbal and verbal information and although such snap judgments can be a source of false beliefs and prejudices, they often convey socially relevant information (e.g., Ambady et al., 2000). Studies examining the influence of different communication channels on impression formation found that people read information into and from a variety of cues. Most research suggests that visual cues have a prevailing influence on social judgments and that speech content only plays a minor role (e.g., Awamleh & Gardner, 1999; Mehrabian, 1972). Other researchers were unable to replicate such findings, which supports the idea that the influence of nonverbal and verbal cues varies according to context and audience involvement (e.g., O’Sullivan et al., 1985; Petty et al., 1983). In this study, we found a prevailing effect of vocal cues on impression formation but the relative role of different communication channels may, of course, depend on different factors that can hardly be tested and controlled in a single study. Future studies should thus extend the current experimental set-up by manipulating the salience of cues (e.g., manipulating vocal parameters), refine analyses by describing behaviors and appearance cues in more detail (e.g., clothing style), and include different contextual information (e.g., party affiliation, status). Also, politicians acting in the public arena may trigger differential judgments than the stimuli some other researchers used to investigate the impact of different communication channels (e.g., Borkenau & Liebler, 1992a; Friedman et al., 2006). As in our study, Krauss et al. (1981) as well as Nagel et al. (2012) collected ratings of politicians and found that people mainly judge them by verbal content. This is partly in accordance with our findings because we also revealed a non-negligible relationship between speech content and some trait ratings. However, the outcomes of the regression analyses suggest that, in our study, prosody and voice qualities had a marked impact on the participants’ judgments. Overall, the results obtained show that people’s first impressions are guided by cues from different communication channels. Put simply, in real life encounters there is no single cue that creates a first impression. Our analyses supports this by showing that in most cases single modalities (e.g., body motion, voice only) explained less variance than combinations of these modalities. Consequently, when making their judgments people seem to rely on a great variety of nonverbal and verbal cues. Also, some modalities such as voice and body motion are intertwined and communicate redundant information to a certain extent, while other modalities do not show such a redundancy. For instance, conscientiousness was more easily attributed to speech content than extraversion. Irrespective of that our results support the idea that people’s first impressions are guided by information from different communication channels (Bänziger et al., 2012; Mehu & van der Maaten, 2014). The relative role of body motion ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ By translating the speakers’ nonverbal behavior into animated stick-figures, we created an abstract stimulus that served as representative of body motion. It has already been shown that parsimonious cues displaying motion behavior communicate socially relevant information (e.g., Blake & Shiffrar, 2007; Chouchourelou et al., 2006). In this study, we extended such research by comparing the influence of cues conveyed by a moving body with information from other communication channels. We were able to show that, for perceptions of extraversion and agreeableness, body motion is a major player. Given that our stick- figure animations reduce information to an abstract stimulus with an artificial appearance, the results obtained are impressive. They indicate people do not only ascribe intentions and personality traits to abstract representatives of body motion, but they also use motion cues to form impressions of their interaction partners. Rapidly categorizing another individual’s intentions, traits and behavioral tendencies may be an ability that has been formed during human evolution (Buss & Greiling, 1999; Fiske, Cuddy, & Glick, 2007). However, some traits are assumed to be of higher social relevance in first encounters because they inform the decision whether to approach or avoid an individual. For instance, Oosterhof and Todorov (2008) found that judgments of neutral faces are reducible to the independent dimensions of valence (i.e., represented by facial trustworthiness) and dominance. They further suggest that evaluations on these dimensions are overgeneralizations for inferring harmful intentions. Other research has shown that attributions of extraversion to a stick-figure animation show a strong relationship with perceived dominance, while perceived agreeableness is strongly related to perceived trustworthiness (Koppensteiner, Stephan, & Jäschke, 2015). This supports the conclusion that people have a higher sensitivity for cues of extraversion and agreeableness because these belong to basic categories of social evaluation with a direct link to survival. For this reason, overgeneralizations with regard to potential threats may not only occur when faces are assessed but also when people categorize body motion. Furthermore, in contrast to facial information, body motion is recognizable from greater distances. It is thus conceivable that motion cues are socially relevant because they allow decision making as another individuals approach, before these individuals come too close. Future directions ~~~~~~~~~~~~~~~~~ In real life situations, nonverbal and verbal information does not disintegrate into pieces. Our results, and other research, suggest that some cues are intertwined. A well- known link, for instance, is between speech and gesturing (e.g., McNeill, 1985). Gestures are sometimes used to emphasize what is being said or go together with voice intonation. Hence, an overlap in the communicative value of motion cues, speech content, and voice quality is well established and the results obtained in the present study provide further support for this. Future research could investigate such an overlap in more detail by extracting certain motion cues as we have done in previous studies (see Koppensteiner & Grammer, 2010) and relating them to vocal cues. Follow-up studies with a more in depth analyses on the interrelations between different cues would also build a bridge to research that investigates how observers react to violations of such linkages (Weisbuch, Ambady, Clarke, Achor, & Weele, 2010). Presenting contradictory nonverbal and verbal information could show that there is indeed redundancy and that different cues might form one single cue with a common communicative value. This might also provide further support for work showing that judgments on the basis of multimodal information are more accurate (Bänziger et al., 2012; Mehu & van der Maaten, 2014). A next step in this kind of research would be to investigate how different nonverbal information influences decision making on important issues. Researchers have already shown that facial photographs from political candidates can be used to predict hypothetical or actual election outcomes (e.g., Antonakis & Dalgas, 2009; Little et al., 2012; Olivola & Todorov, 2010). Moreover, other studies reveal a link between vocal cues and leadership qualities (Klofstad et al., 2012; Tigue et al., 2012). Taking these as a starting point future research could extend into the domain of leadership research to examine how the interplay of different cues or multimodal cues affects decision-making.","In summary, our findings indicate that people make use of verbal and different nonverbal cues when forming first impressions of politicians. Static visual cues, body motion, verbal content, and vocal cues were related to people’s impressions. Vocal cues were revealed as a strong predictor for all trait ratings in our study, yet other communication channels also proved important. The impact of body motion was particularly strong for ratings of agreeableness and extraversion were also linked to corresponding ratings based on the speakers’ voices. In other words, with regard to these personality dimensions body motion is coupled to information from the vocal channel. Cues communicating agreeableness and extraversion may reflect social abilities, which are helpful in social encounters and facilitate establishing interactions. Hence, people may be more sensitive to cues of extraversion and agreeableness; for this reason, they also read these traits into body motion."],["This study contrasted two forms of mother-infant mirroring: the mother's imitation of the infant's facial, gestural, or vocal behavior (i.e., \"direct mirroring\") and the mother's ostensive verbalization of the infant's internal state, marked as distinct from the infant's own experience (i.e., \"intention mirroring\"). Fifty mothers completed the Adult Attachment Interview (Dynamic Maturational Model) during the third trimester of pregnancy. Mothers returned with their infants 7 months postpartum and completed a modified still-face procedure. While direct mirroring did not distinguish between secure and insecure/dismissing mothers, secure mothers were observed to engage in intention mirroring more than twice as frequently as did insecure/dismissing mothers. Infants of the two mother groups also demonstrated differences, with infants of secure mothers directing their attention toward their mothers at a higher frequency than did infants of insecure/dismissing mothers. The findings underscore marked and ostensive verbalization as a distinguishing feature of secure mothers' well-attuned, affect-mirroring communication with their infants. © 2014 Elsevier Inc. --------------------------------------------------------------------------------","In many mammalian species, mothers and infants engage in a rich repertoire of species- specific, reciprocal, dyadic interactions. Non-human primate mother–infant pairs show capacity for mutual eye gaze, reciprocal lip smacking, and vocal and gestural mimicry (Bard et al., 2005; Ferrari, Paukner, Ionica, & Suomi, 2009; Mancini, Ferrari, & Palagi, 2013). Human mother–infant dyads participate in communicative exchanges that are far more complex and affectively enriched (Beebe et al., 2010; Brazelton, Koslowski, & Main, 1974; Carpenter, Nagell, & Tomasello, 1998; Feldman, 2007; Gergely & Watson, 1996; Lavelli & Fogel, 2013; Malatesta, Culver, Tesman, & Shepard, 1989; Sroufe, 1996; Tronick, 1989). The infant routinely directs a broad range of affectively nuanced expressions to the mother (Bennett, Bendersky, & Lewis, 2005; Colonnesi, Zijlstra, van der Zande, & Bogels, 2012; Messinger, 2002). The mother sequentially mirrors the infant's signals as she empathically delivers her finely tuned response (Jonsson & Clinton, 2006; Lavelli & Fogel, 2013; Papousek & Papousek, 1989; Stern, 1985). In turn, the infant attentively responds, organizing his1 behavior with respect to the mother's input (Beebe et al., 2010; Bigelow & Walden, 2009; Cohn & Tronick, 1987; Soussignan, Nadel, Canet, & Gerardin, 2006). A relatively synchronous flow of affective communication is one of the key indicators of secure mother–infant attachment (Beebe et al., 2012; Crandell, Fitzgerald, & Whipple, 1997; Feldman, Gordon, & Zagoory-Sharon, 2011; Isabella & Belsky, 1991; Lundy, 2003). Maternal mirroring, or emotionally attuned responsiveness, has received extensive attention in the study of mother–infant behavior (Bigelow & Walden, 2009; Fraiberg, Adelson, & Shapiro, 1975; Gergely & Watson, 1996; Jonsson & Clinton, 2006; Lavelli & Fogel, 2013; Lyons-Ruth, 2000; Stern, 1985; Winnicott, 1967). Mirroring is a construct closely tied to that of secure attachment. Maternal attachment security is a critical determinant of the mother's capacity to provide adequate mirroring for the infant (Main, Kaplan, & Cassidy, 1985; Pederson, Gleason, Moran, & Bento, 1998; Tarabulsy et al., 2005; van IJzendoorn, 1995; Whipple, Bernier, & Mageau, 2011). Well-attuned maternal mirroring, in turn, is a necessary antecedent to the development of secure attachment in the infant (Ainsworth, Blehar, Waters, & Wall, 1978; Belsky, Rovine, & Taylor, 1984; Bigelow et al., 2010; Bretherton, Biringen, Ridgeway, Maslin, & Sherman, 1989; De Wolff & van Ijzendoorn, 1997; Isabella, 1993; Jaffe, Beebe, Feldstein, Crown, & Jasnow, 2001). In the early literature that followed Ainsworth's pioneering work on infant attachment (Ainsworth et al., 1978; Ainsworth & Wittig, 1969), mirroring was often studied as an aspect of the broader construct of sensitive responsiveness, which encompasses heterogeneous sets of maternal behaviors (Belsky et al., 1984; De Wolff & van Ijzendoorn, 1997; Grossmann, Grossmann, Spangler, Suess, & Unzner, 1985; Isabella, 1993; Main, Tomasini, & Tolan, 1979). While theoretically important distinctions had been made between types of mirroring generated by the mother, mirroring was coarsely defined as a generic construct under the rubric of sensitivity, and the fine-grained distinctions were overlooked in the early studies. In his seminal volume on infant development, Stern (1985) drew a stark contrast between mirroring of the external behavior and mirroring of the internal state, which was echoed with some variation by later developmentalists. In imitation, the mother mirrors and replicates the infant's external cues—facial, gestural, or vocal. The mother need not tune into the infant's internal experiences in order to imitate his external behavior. In contrast, a more sophisticated form of mirroring necessitates that the mother “get inside” the mind of the infant and “read” the affective state that underlies his overt behavior (Stern, 1985, pp. 138–139). This form of mirroring moves beyond the mere matching of the infant's external signals. What the mother observes and mirrors here is not the infant's external behavior per se, but his subjective internal state. Whereas a close within-modal match is found between the mother and the infant in imitation, the mother's mirroring of the infant's internal state is often cross-modal. As Stern (1985) famously observed (p. 140), the mother may match the feeling state conveyed by the infant's vocalization (e.g., exuberant “aaah!”) with her body movement (e.g., performing a shimmy with her upper body for the duration of the “aaah!”), or match the feeling state captured in the infant's movement (e.g., hitting a toy) using her voice (e.g., saying “kaaaaa-bam” in rhythm with the hitting movement). Thereafter, important empirical advances were made in the literature by Fonagy (1991) and Meins (1997, 1999), who led converging lines of research underscoring the mother's mentalizing capacity. These were respectively termed parental reflective functioning (Fonagy, Gergely, Jurist, & Target, 2002; Fonagy, Steele, Steele, Moran, & Higgit, 1991; Fonagy & Target, 1997; Slade, 2005) and maternal mind–mindedness (Meins, Fernyhough, Fradley, & Tuckey, 2001; Meins et al., 2003), referencing a mother's capacity to adequately mirror her infant's subjective internal state (see Sharp & Fonagy, 2008 for a detailed review of relevant constructs). High levels of reflective functioning and maternal mind–mindedness have been reported in mothers who are securely attached (Arnott & Meins, 2007; Demers, Bernier, Tarabulsy, & Provost, 2010; Fonagy et al., 1991; Slade, Grienenberger, Bernbach, Levy, & Locker, 2005). Others have demonstrated that the secure mother's accurate perception and reflection of her infant's internal state are causally related to the key features of the infant's self-development, including self- awareness, self-regulation, and self-efficacy (Bigelow et al., 2010; Fonagy, Gergely, & Target, 2007; Lyons-Ruth, 2000; Mcquaid, Bibok, & Carpendale, 2009; Nadel, Prepin, & Okanda, 2005; Schore, 2005; Tronick & Beeghly, 2011). Far less consensus and empirical support, however, exist on what constitute the essential ingredients of such mirroring and what mechanisms mediate these developmental effects. Recent research has begun to address this gap. Gergely (2007) has undertaken a fine-grained analysis of the nature of maternal mirroring. He proposed that markedness and ostensiveness were essential ingredients of mirroring (Gergely, 2007; Gergely & Unoka, 2008a). The putative mechanisms mediating the developmental functions of the marked and ostensive mirroring were also articulated. At birth, infants are understood to be incapable of differentiating universal categories of emotions that they experience, such as anger, fear, or sadness (Camras, 2011; Gergely & Watson, 1996; Walle & Campos, 2012; Widen, 2013). To infants, their affective experience is one of undifferentiated visceral arousal with overarching positive or negative valence, rather than one characterized by well-defined, discrete emotions (Fonagy et al., 2002, 2007; Gergely & Watson, 1996, 1999). Central to Gergely's proposal is the hypothesized role of the mother's marked, ostensive mirroring in the infant's emerging capacity for subjective awareness of his discrete internal states. When provided consistently to the infant, the mother's marked, ostensive mirroring is proposed to serve as the essential foundation upon which the infant learns to organize and make sense of his internal experiences (Gergely & Unoka, 2008a, 2008b). The mother's marked affective communication (Fonagy et al., 2002, 2007) is one in which the mother demonstrates her understanding of the infant's internal state, while concurrently signaling that she is not experiencing the same state herself. The mother accomplishes this by displaying the infant's affect in a schematic and exaggerated manner. Consider the mother mirroring her infant's distress. The mother exaggerates her display of distress; she slows down her expression as she ensures that it is seen by the infant. Some aspects of the distressed affect are made salient in the mother's expression, while other peripheral aspects are ignored. The mother may also mix in other emotions in her expression (e.g., distress intermingled with concern). What is shown in the mother's mirroring response is the schematically modified display of the infant's distress, which is perceptually distinguishable from the mother's expression of her own distress. Trevarthen (1977), Fogel (1993), and Stern (1985) had previously noted the qualitatively distinct nature of the mother's mirroring from the infant's original affective display, which was captured in their descriptions of “echo,” “elaboration,” and “affect attunement,” respectively. Gergely elaborated on the functional significance of the mother's marked mirroring, particularly underscoring its role in developing the infant's capacities for organizing and regulating his internal states (Gergely, 2007; Gergely & Unoka, 2008a). In marked mirroring, the mother's exaggerated display, coupled with her soothing tone, serves to mitigate the potentially arousing effect of direct imitation (e.g., the mother crying when the infant cries), while simultaneously making salient to the infant central aspects of his internal experience. The mother's marked response is often accompanied by what Gergely calls ostensive communicative cues, which manifest the mother's intention in displaying the affect (Csibra, 2010; Csibra & Gergely, 2009; Egyed, Kiraly, & Gergely, 2013). The term “ostensive” is borrowed from the communication literature (Russell, 1940; Sperber & Wilson, 1995), which posits the inherently dual nature of intention (i.e., informative and communicative) in human communicative acts (Grice, 1989). In a communicative act, the communicator intends to convey the desired information (“informative intention”) by making this intention evident to the addressee (“communicative intention”). Ostensive cues are signals employed by the communicator to reveal that she has a communicative intention directed toward the addressee. The mother's gaze at her infant, the slight tilting of her head toward him, her direct eye contact, the “motherese” intonation, and the calling of the infant's name all constitute ostensive cues that prototypically accompany the mother's marked mirroring, and signal to the infant that her expression concerns him and what unfolds within him. These ostensive signals orient the infant toward his own face and body, setting the stage for him to learn that this display matches his own subjective internal state. Gergely proposes that such instances of marked, ostensive communication, to which the infant is hard-wired to attend (Colombo, Frick, Ryther, Coldren, & Mitchell, 1995; Farroni, Csibra, Simion, & Johnson, 2002; Parise & Csibra, 2013; Senju & Csibra, 2008), repeatedly orient him to his subjective internal states (Csibra, 2010; Gergely, 2007; Gergely & Jacob, 2012); through this process, the infant comes to develop awareness of, and later adequate control over, his internal experiences (Fonagy et al., 2007; Gergely & Unoka, 2008a, 2008b). In Gergely's model, the mother's marked and ostensive mirroring functions to prompt the infant to look to the mother as a way of learning about himself, which serves as an impetus for the infant's subsequent self-development. In the present study, we contrast two types of maternal mirroring. The first is imitation, or what we hereafter refer to as direct mirroring. We consider this to be a rudimentary form of mirroring, which allows the infant to see his facial, gestural, and vocal behavior directly replicated by the mother in the same modality (i.e., when the infant frowns, he sees the mother frowning; when the infant coos, he hears the mother cooing back). We see this imitation as akin to the mother holding up a physical mirror to the infant. The second is what is hereafter called intention mirroring. Intention mirroring is the type of mirroring that is characterized by marked and ostensive verbalization, and lies at the crux of sensitive mothering as Gergely has hypothesized (Gergely & Unoka, 2008a, 2008b). Rather than acting as a mere physical mirror, as in the former, here the mother holds up an intention mirror to the infant, allowing him to see, through her, his own intentions, feelings, and attitudes, many of which he may not have otherwise made sense of (e.g., as the infant frowns, he sees the concerned look on the mother's face; gazing at the infant, the mother states in a motherese voice, “Aww, you didn’t like that”). Here the mother uses verbalization as a vehicle for representing the infant's internal state originally conveyed in his facial, gestural, or vocal behavior. Utilizing a micro-analytic coding system devised to distinguish between the two types of mirroring, we examined, at 7 months postpartum, the use of mirroring in mothers who were prospectively assessed to be securely attached compared to those insecurely attached during a modified still-face procedure (MSFP; Koos & Gergely, 2001). The MSFP is a three-phase procedure, in which the mother interacts freely with the infant in the first and third phases, but is instructed to maintain a motionless and neutral ‘still face’ during the second phase, suddenly depriving the infant of maternal contingency and henceforth inducing stress in the infant (Koos & Gergely, 2001; Tronick, Als, Adamson, Wise, & Brazelton, 1978). The MSFP thereby offers an opportunity to observe moment-by-moment exchange between the mother and the infant in the presence of and during recovery from an interpersonal stressor. Developmentalists have pointed to 7 months as the juncture at which the external environment develops a particular importance in the infant's developing awareness of his subjective internal states. Stern (1985) observed that, starting at 7 months, domains of mother–infant relatedness expand significantly to include the dyad's subjective internal states. Gergely also noted that infants demonstrate rudimentary mentalizing abilities at around 7 months (Gergely, 2011; Kovacs, Teglas, & Endress, 2010). Furthermore, behavioral (Walker-Andrews, 1986) and electrophysiological (Grossmann, Striano, & Friederici, 2006) evidence suggests that infants’ abilities to recognize and process cross-modal correspondence in emotional stimuli are initially seen to emerge at around 7 months. Therefore, we conducted the MSFP at 7 months to capture early experiences of coordinated communication between mother and infant seen during this formative juncture. Our primary aim in the study was to investigate whether the infant- directed communication of securely attached mothers at 7 months could be reliably distinguished from that of insecurely attached mothers by the extent to which intention mirroring was used. Three hypotheses were addressed in the present study. First and foremost, we hypothesized that securely attached mothers would engage in more intention mirroring than insecurely attached mothers during the free-interaction phases (i.e., first and third phases) of the MSFP. We also predicted that secure mothers would show an increase in intention mirroring during the third phase relative to the first phase, demonstrating sensitivity to the infant's experience of stress during the still-face phase (Leerkes, 2011; McElwain & Booth-LaForce, 2006). We did not expect to find differences in the frequency of direct mirroring, either facial/gestural or vocal, between the two mother groups. Second, we hypothesized that infants of securely attached mothers would direct their attention more frequently to their mothers, compared with infants of insecurely attached mothers. We tested this hypothesis by comparing the two infant groups on the frequency of their gaze toward and away from the mother during the still-face phase. Whereas the infant's gaze during the free-interaction phases may be directly confounded by the mother's behavior, the still-face phase was considered apropos for this examination because the mother's behavior is controlled across participants. Third, we tested the possibility that the relationship between maternal attachment security and the infant's attention toward the mother would be mediated by the mother's use of intention mirroring. Through examining these hypotheses, we attempted to carry out an empirical substantiation of an aspect of Gergely's model concerning the role of the mother's marked, ostensive mirroring in shaping the infant's attention and readiness to learn from his primary social environment—his mother.","First-time mothers were recruited during the third trimester of pregnancy through local prenatal clinics and community advertisements. Of 116 participants initially recruited, 61 met eligibility criteria, and 50 completed the MSFP procedure 7 months postpartum. Enrolled women were between ages 19 and 41 (M = 27.9 ± 4.8), and were generally from middle to high socioeconomic backgrounds. None of the participants had a history of past or present alcohol or substance abuse, nicotine use during pregnancy, or were on psychotropic medications at the time of the study. Each participant provided written informed consent in accordance with the protocol approved by the local institutional review board.","We adopted a prospective design: mothers’ attachment was assessed prenatally during the third trimester of pregnancy; mothers returned with their infants 7 months postpartum and completed the MSFP. Maternal prenatal attachment Maternal attachment was assessed using a modified version of the Adult Attachment Interview (AAI; Crittenden & Landini, 2011; George, Kaplan, & Main, 1985). The AAI is a semi-structured 1–1.5-h interview comprising probes that elicit attachment- related autobiographical memories, usually those involving childhood experiences with parents. The coding is determined by the participant's style of discourse in describing attachment-related experiences and their impact on present functioning. The AAI yields three basic categories that parallel Ainsworth's classification of infant attachment (Ainsworth et al., 1978): ‘secure,’ ‘insecure/dismissing,’ and ‘insecure/preoccupied.’ Those who are classified as secure describe their experiences in a balanced manner, flexibly integrating cognition and affect as they recount their past history and process attachment-related information. Those with insecure/dismissing attachment tend to be cognitively organized; they describe events in a temporally ordered manner, while inhibiting or distorting any display of negative affect. In contrast, individuals with insecure/preoccupied attachment are organized around their feelings; they oscillate between intense affect and draw causal relations that are erroneous and contradictory (Crittenden & Landini, 2011). The AAIs were audio-recorded, transcribed, and blindly coded by reliable raters in accordance with Crittenden's Dynamic Maturational Model (DMM) of Attachment and Adaptation. Fifty percent of transcripts were double-coded to ensure inter-rater reliability; there was 77% agreement for the AAI classification, with kappa of .66 (p < .001). Discrepancies were resolved through conferencing between coders. Of the 50 mothers who completed both the AAI and the MSFP, 25 (50.0%) were classified as having secure attachment, 16 (32.0%) had insecure/dismissing attachment, five (10.0%) demonstrated insecure/preoccupied attachment, and the remaining four (8.0%) alternated between or showed a combination of insecure/dismissing and insecure/preoccupied attachment patterns. Due to the small size of the latter two groups, all analyses were conducted comparing the two predominant attachment groups—those with secure attachment versus those with insecure/dismissing attachment. Mother–infant behavior at 7 months postpartum The MSFP followed the standard still-face procedure (Tronick et al., 1978), except that the mother and infant were seated next to each other, separated by a divider and facing a one-way mirror (Fig. 1(a)). The infant was placed in a high chair in the observation room facing the one-way mirror. The mother sat directly adjacent to her infant facing the same mirror. The divider placed between the mother and infant precluded direct face-to-face communication, but they were able to see each other reflected in the mirror. On the opposing side of the one-way mirror were two cameras, generating a split-screen recording of the mother and infant. The dyads were videotaped as they interacted with each other during the three 2-min phases: (1) the baseline normal interaction phase, (2) the still-face phase, during which the mother assumed a neutral face, and (3) the recovery phase, in which the mother resumed free interaction with the infant (Fig. 1(b)). The start of each phase was signaled to the mother via an intercom. While each phase was recorded for 2 min, some variation in timing was noted due to infant behavior and parent compliance with procedure instructions. Trained raters, who were blind to the mother's attachment status, coded the videotaped interaction. Forty percent of the videotapes were double-coded to establish inter-rater reliability. Coded variables are as follows. Maternal mirroring variables Two forms of maternal mirroring were coded: (a) direct mirroring and (b) intention mirroring. Direct mirroring (akin to holding a physical mirror) was defined as the mother's non-verbal attunement behavior in which the mother directly imitated her infant's expressions. We coded direct mirroring in two different modalities, facial/gestural and vocal. Maternal behavior was coded as direct mirroring if it entailed a direct replication of the infant's preceding behavior in the same modality without the use of verbalization. Common examples noted were the display of a maternal smile shortly following an infant smile (facial/gestural), and a maternal vocalization “brrr” following an infant's “brrr” (vocal). The intention mirroring (i.e., holding an intention mirror) was defined as the mother's non-imitative, marked, ostensive, verbal attunement. This form of mirroring was distinguished by the following: (1) clear indication that the mother was reflecting the infant's experience rather than displaying her own internal state (markedness); (2) a specific signal which made manifest to the infant the mother's communicative intention to present some new relevant information for the infant by referencing and acknowledging his intentional, emotional, or attentional state using ostensive cues (e.g., her direct eye contact, “motherese” intonation, or calling of the infant's name); (3) accuracy and appropriateness from the perspective of an onlooker, which suggested that the mother's response was aligned with the infant's putative experience in terms of timing, content, and intensity (operationalized in terms of the subjective judgment of an independent coder who saw the interaction as reflecting reasonable congruence between the mother's expression and the infant's assumed experienced state). As this coding scheme aimed to capture the mother's interest in and understanding of the infant's internal state, only those attributions that explicitly acknowledged the infant's subjective state (e.g., “You are feeling hot.”) were coded, whereas maternal verbalizations that were ambiguous or perceived as commenting on the infant's physical state (e.g., “You are hot.”) were not coded. Likewise, whereas the mother's simple imperatives (e.g., “Don’t cry!”) did not qualify, similar statements that referenced the infant's state in a marked tone of expression (e.g., “Oh, you are so upset. This is so upsetting. Oh dear, oh dear. What's the matter?”) were coded as intention mirroring. The most commonly observed examples of intention mirroring during the MSFP occurred when the mother recognized that the abrupt loss of maternal responsivity in the still-face phase might be both puzzling and distressing for the infant. In such an instance, the mother might look into the infant's eyes and remark in a marked (exaggerated) tone of voice as she transitioned out of the still-face phase: “Wow, what was that? That was crazy, wasn’t it? What happened to mommy?” Good to excellent inter-rater reliability was demonstrated for the mirroring variables, with the intraclass correlations of .80 (direct, facial/gestural), .60 (direct, vocal) and .83 (intention). Infant attention variables Infant attention was quantified by the frequency of gaze fixations toward and away from the mother. Our primary interests were (a) fixations on the mother's image in the mirror and (b) fixations away from the mirror, although we also recorded fixations on the infant's own image in the mirror. Fixation was defined as an eye gaze that remained stationary for a minimum of 1 second. The intraclass correlation for the infant gaze fixations was .84 (p < .001). Additional mother and infant characteristics Several mother and infant characteristics were also examined as potential confounds. Mothers were screened for symptoms of depression and personality disorders using the Beck Depression Inventory-II (BDI-II; Beck, Steer, & Brown, 1996) and Personality Disorder Questionnaire 4+ (PDQ-4+; Hyler, Skodol, Oldham, Kellman, & Doidge, 1992), respectively. Maternal parenting stress was assessed using the Parenting Stress Index (PSI; Abidin, 1995), and maternal temperament was measured via the Adult Temperament Questionnaire-Short Form (ATQ; Rothbart, Ahadi, & Evans, 2000). We also collected information on the infant's daycare status (i.e., number of hours per week the infant was cared for by someone other than the mother). Infants were screened for developmental delays using the Bayley Scales of Infant and Toddler Development, Third Edition, Screening Test (Bayley-III; Bayley, 2005). Infant temperament was evaluated using the Infant Behavior Questionnaire- Revised (IBQ-R; Gartstein & Rothbart, 2003) For details on the psychometric properties of these instruments, see Shah, Fonagy, and Strathearn (2010). Data analysis Mother–infant dyads were classified into secure and insecure attachment groups based on the mother's AAI. Between-group comparisons were made on all measured sociodemographic and behavioral variables using chi-square statistics and t-tests for categorical and interval data, respectively. All analyses were performed using SPSS version 20 and STATA version 12.1. Hypothesis 1: maternal direct mirroring versus intention mirroring We analyzed the frequency of direct mirroring (facial/gestural and vocal) and intention mirroring, adjusting for the total length of time for which codable data were available in each respective phase. The mirroring variables were inspected for normality via quantile–quantile plots of residuals against fitted values. Square root transformations offered the closest approximation to normality. We probed for the main and interaction effects of maternal attachment status (secure vs. insecure) and phase (phase 1 vs. 3), using mixed-effects linear regression models that included a subject-level random intercept and a random coefficient for phase. The model was fitted by maximum likelihood estimation, and nested models were contrasted using likelihood-ratio chi-squares. Hypothesis 2: infant gaze toward versus away from the mother We analyzed the frequency of infants’ gazes toward and away from the mother, adjusting for the total number of fixations recorded for each infant during the still-face phase. The total number of fixations was quantified as the sum of the infant's fixations on the mother, on himself, and away from the mirror. While used in the calculation of total fixation frequency, the infant's fixations on himself were of less interest in the present study and were excluded from the remainder of the analysis to minimize multicollinearity. The proportions of fixations computed were arcsine transformed and submitted to the 2 × 2 mixed ANOVA, with maternal attachment status (secure vs. insecure) as a between-subjects factor and gaze direction (toward mother vs. away from mirror) as a within-subjects factor. Hypothesis 3: mediating role of maternal intention mirroring We used the non-parametric bootstrapping procedure (Preacher & Hayes, 2008) to test the model in which maternal intention mirroring was specified as a mediator between maternal attachment and infant gaze direction. This procedure uses ordinary least squares regression to estimate the total, direct, and indirect effects of a predictor on an outcome through a proposed mediator, and provides bias-corrected confidence intervals for the indirect effect. This approach makes no assumptions about the shape of the sampling distribution of the indirect effect, generates estimates based on empirically derived bootstrapped sampling distribution, and has been recommended over traditional approaches to mediation analysis (i.e., Sobel test or causal steps approach; Mackinnon, Lockwood, & Williams, 2004). A total of 5000 bootstrapping samples were utilized in the present analysis. Participant characteristics ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Mother and infant characteristics are shown in Tables 1 and 2 for the secure and insecure/dismissing groups. No significant differences were observed between the two groups for any of the measured sociodemographic or behavioral variables. Hypothesis 1: maternal direct mirroring versus intention mirroring Means and standard errors of the maternal mirroring variables are shown in Table 3 for the two attachment groups. Maternal direct and intention mirroring did not correlate with each other (rfacial/gestural direct & intention = .021, p = .85; r vocal direct & intention = .196, p = .08), while the two forms of direct mirroring were significantly correlated (rfacial/gestural direct & vocal direct = .29, p = .009). Maternal direct mirroring As hypothesized, maternal attachment status was not a significant predictor of direct mirroring, either alone (facial/gestural, βAAI = .001, 95% CI = -.05 to .06, z = 0.03, p = .97; vocal, βAAI = −.03, 95% CI = −.09 to .03, z = −0.89, p = .37) or in interaction with phase (facial/gestural, βAAI × phase = −.006, 95% CI = −.06 to .05, z = −0.20, p = .84; vocal, βAAI × phase = .005, 95% CI = −.06 to .06, z = 0.15, p = .88; Figure 2). Both secure and insecure/dismissing mothers engaged in a higher frequency of facial/gestural direct mirroring during the first phase of the MSFP compared to the third phase (βphase = −.04, 95% CI = −.08 to −.01, z = −2.49, p = .01). No difference was found in the frequency of vocal direct mirroring between the two phases (βphase = −.01, 95% CI = −.05 to .02, z = −0.68, p = .50). Maternal intention mirroring A significant main effect was found for maternal attachment status (βAAI = −.08, 95% CI = −.15 to −.02, z = −2.73, p = .006), with secure mothers displaying intention mirroring at a frequency greater than twice that of their insecure/dismissing counterparts (Figure 2). Both mother groups engaged in intention mirroring more frequently in the third phase as compared to the first phase (βphase = .04, 95% CI = .01 to .07, z = 2.79, p = .005). Maternal attachment and phase did not interact significantly in the prediction of intention mirroring (βAAI × phase = .01, 95% CI = −.04 to .06, z = 0.46, p = .64). Hypothesis 2: infant gaze toward versus away from the mother Means and standard errors of the gaze fixation variables are presented in Table 4 for the secure and insecure/dismissing attachment groups. The 2 (secure vs. insecure maternal attachment) × 2 (infant gaze toward mother vs. gaze away) mixed ANOVA yielded no significant main effects of maternal attachment status (F(1, 39) < 0.001, p = .98) or gaze direction (F(1, 39) = .283, p = .60). However, a significant interaction of maternal attachment and gaze direction was found (F(1, 39) = 6.393, p = .02). Consistent with our hypothesis, infants of secure mothers directed their gaze more frequently to their mothers compared to infants of insecure/dismissing mothers (t(39) = 2.38, p = .02). The reverse pattern was seen for gazes directed away, with infants of insecure/dismissing mothers looking away more frequently than infants of secure mothers (t(39) = −2.06, p = .046; Figure 3). Hypothesis 3: mediating role of maternal intention mirroring Results indicated a non-significant mediating effect of maternal intention mirroring. Regression analyses did not reveal a significant relationship between maternal intention mirroring and infant's gaze directed toward the mother (b = −.33, t = −.69, p = .49) or directed away from the mirror (b = .22, t = .39, p = .70).","We contrasted a rudimentary form of maternal mirroring (i.e., direct mirroring) with the mother's marked and ostensive mirroring (i.e., intention mirroring), the type of mirroring that has been theorized to serve as an impetus for the infant's subsequent psychosocial development (Fonagy et al., 2002, 2007; Gergely, 2007). As hypothesized, direct mirroring did not distinguish between mothers who were prospectively assessed to be secure and those assessed to be insecure/dismissing. However, the two groups of mothers showed a significant difference in their use of intention mirroring, with the frequency in secure mothers observed to be more than double that of insecure/dismissing mothers. A notable difference was also found in the frequency with which infants directed their attention to their mothers. Infants of secure mothers directed their gaze toward their respective mothers at a higher frequency than did infants of insecure/dismissing mothers. Although Gergely proposed markedness and ostensiveness as essential ingredients of the mother's affectively attuned communication (Gergely, 2007; Gergely & Unoka, 2008a, 2008b), this is the first study, to our knowledge, that directly examined these elements as part of the mother's affect mirroring communication. Two aspects of our intention mirroring variable should be noted in considering our results. First, as opposed to direct mirroring, which is primarily concerned with the mother's matching of her behavior to her infant's external behavior, intention mirroring was coded when the mother went beyond the behaviors and remarked on the infant's subjective internal experiences. As has been described in the previous literature (Fonagy et al., 2002; Meins et al., 2001; Sharp & Fonagy, 2008), the process of intention mirroring draws upon the mother's complex higher-order metacognitive capacities, such as parental reflective functions or maternal mind–mindedness, which enable her to make sense of the infant's unobservable internal states. In this respect, the construct of intention mirroring encompassed what previous studies have identified as critical elements of maternal sensitive responsiveness. Second, however, to be coded as intention mirroring, the mother's acknowledgment of her infant's subjective internal state had to be delivered in a manner that generates understanding in the infant that her mirroring display concerns his internal experiences (Fonagy et al., 2002; Slade, 2005). In other words, to use Gergely's terms, the mother's use of marked and ostensive cues constituted a critical aspect of intention mirroring, which distinguished our intention mirroring variable from extant empirical constructs. The distinct nature of our concept of intention mirroring emerges from Gergely's fine-grained analysis of the functional significance of markedness and ostenstiveness in maternal mirroring (Gergely, 2007; Gergely & Unoka, 2008a). Intention mirroring shares similarities with Meins's mind–mindedness (1999) and Stern's affect attunement (1985) in that it concerns the mother's recognition and reflection of the infant's internal state. However, Gergely diverged from the primary intersubjectivist view of Stern (1985) and other theorists (e.g., Meltzoff, Trevarthen), which hinges on the assumption that infants have an inherent capacity to access their subjective internal states and to perceive the ‘sharing’ of these states by their mothers (Meltzoff, 2002; Meltzoff & Gopnik, 1993; Trevarthen, 1993; Trevarthen & Aitken, 2001). Detailing criticisms of this view (Gergely, 2007; Gergely & Csibra, 2005; Gergely & Watson, 1999), Gergely contended that the infant's ability to recognize his discrete internal states, and the sharing thereof, is a developmental outcome made possible by a unique form of maternal mirroring that enables this capacity to be fostered in the infant (Fonagy et al., 2002, 2007). In proposing this view, Gergely underscored specific elements of maternal mirroring—markedness and ostensiveness—that achieve this end. Meins similarly but independently proposed and demonstrated the functional significance of the mother's tendency to comment on her infant's mental states; namely, it facilitates the infant's developing understanding of the mind (Meins, 1997; Meins et al., 2003). However, Meins's theory did not discuss the putative mechanisms that mediate this link, although in later writings Meins and colleagues also emphasized the importance of measuring “appropriateness” in addition to simple mental-state talk in achieving robust predictions from maternal mind mindedness to the development of the child (Osorio, Meins, Martins, Martins, & Soares, 2012). Gergely's distinct contribution lies in spelling out how the mother's appropriate affectively attuned communication, when using marked and ostensive cues, can be accessed via the infant's rudimentary abilities. The developmental significance of this is assumed to be in the understanding of emotional experience, which has an interpersonal aspect in Gergely's theory. A recent study with primary school children provided confirmatory evidence in that emotional validation by the mother predicted higher emotional awareness, whilst emotional invalidation reduced emotional awareness in the child (Lindberg, 2013). Our measure of intention mirroring did not correlate with that of direct mirroring, indicating that distinct processes may underpin the two forms of mirroring. Also of note is that intention mirroring was significantly associated with maternal attachment security, while direct mirroring showed no relationship. Whereas secure and insecure/dismissing mothers did not appear to differ in their ability to respond on a behavioral level, as assessed by direct mirroring, insecure/dismissing mothers were significantly less able than their secure counterparts to accurately extract meaning from their infants’ behavior and respond to their underlying internal states using marked and ostensive cues, as assessed by intention mirroring. We had also hypothesized an increase in secure mothers’ intention mirroring during the third phase of the MSFP, the phase in which infants undergo recovery from the stress of the still-face phase. The hypothesized increase was observed not only in the expected group of secure mothers but also in insecure/dismissing mothers, and was accompanied by a decrease in facial/gestural direct mirroring in both groups. While little attention has been directed toward mothers’ responses in the still-face literature, a decrease2 in the frequency of maternal direct mirroring has previously been reported during the third phase (Bigelow & Walden, 2009). Our documented increase in intention mirroring, coupled with a decrease in facial/gestural direct mirroring, suggests that mothers may be more inclined to go beyond simple facial/gestural imitation and attend to their infants’ underlying needs in the face of infant dysregulation. This tendency appears to be present in both secure and insecure/dismissing mothers, although the frequency of intention mirroring was observed to be consistently higher in secure mothers. In Gergely's model, the function of the mother's intention mirroring is to help the infant recognize that what he sees displayed externally by the mother congruently matches his internal experiences. In mother–infant dyads where changes in the infant's internal states repeatedly effect visible external changes in the mother, the infant is thought to routinely look to the mother, his “intention mirror,” for a perceptual representation of his emotional and intentional states. Our study provided partial support for this model. Infants of mothers who were more proficient in intention mirroring (i.e., secure mothers) looked to their mothers more than infants of mothers who were less proficient (i.e., insecure/dismissing mothers). However, despite the significant difference that the two attachment groups demonstrated in both maternal intention mirroring and infant gaze direction, infant gaze direction was not directly associated with intention mirroring in our laboratory situation. Maternal attachment has been robustly associated with the quality of the affective communication that the mother provides for her infant (Arnott & Meins, 2007; Slade et al., 2005; Tarabulsy et al., 2005; Whipple et al., 2011). The link we report herein between maternal attachment security and intention mirroring is in line with this research. Evidence also exists that, by 7 months of age, infants develop consistent expectations about their mothers’ patterns of responsiveness, which helps guide and regulate their end of the communicative exchange (Hains & Muir, 1996; Legerstee & Varghese, 2001; Mcquaid et al., 2009). This capacity has been demonstrated in the still- face or replay phases, where infants who have been exposed to high levels of maternal mirroring continued their attempts at engagement with their mothers (e.g., continued gaze or smile), even in the absence of their mothers’ typical level of attunement (Bigelow & Walden, 2009; Legerstee & Varghese, 2001; Mcquaid et al., 2009). In noteworthy contrast, infants in these studies who were accustomed to low levels of mirroring displayed relatively little effort to carry on their side of the communication. While these results were obtained on measures of generic mirroring, and mothers were distinguished on the basis of the amount of mirroring they provide, we have demonstrated here that the type of mirroring may matter. We have shown that the above pattern of results was replicated when mothers were distinguished on the basis of intention mirroring. No noteworthy finding emerged, however, with regard to direct mirroring alone. Contrary to our expectation, our hypothesis that intention mirroring may mediate the link between the mother's attachment and the infant's attention toward the mother was not confirmed in our laboratory. The lack of association seen in our data between intention mirroring and infant gaze raises the possibility that the relationship between maternal attachment security and infant gaze direction may be mediated by aspects of maternal attachment that are independent from intention mirroring. While it is difficult to rule out this possibility, it also seems plausible that the mediational link, which may have taken shape over a period of months while patterns of mother–infant interaction were developed, may not have been evident during a 6-min structured interaction in the lab. In line with the previous studies (e.g., Bigelow & Walden, 2009; Legerstee & Varghese, 2001; Mcquaid et al., 2009), the differences seen in our infants’ attention toward their mothers in the still-face phase, during which maternal behavior was held constant, may reflect differences in the infants’ interactive histories with their mothers and the expectations that the infants have subsequently come to form. Infants of secure mothers may have directed frequent attention to their mothers during the still-face phase, given their routine experience of their mothers’ intention mirroring. Indeed, despite the lack of association at the micro level in the lab, there may have been a general pattern of mediation at the macro level, linking the mother's attachment security, her history of intention mirroring, and the infant's pattern of attention toward the mother. The absence of association seen at the micro level may also serve to corroborate results from prior studies that the variation in the mother's behavior (e.g., mirroring) on a short-term basis does not alter the general expectations of the infant for responsive interactions. Several limitations of the study should be recognized. First, our sample consisted largely of middle- to upper-class mothers of average to above-average intelligence, and therefore may not have been representative of the general population. Second, we were not able to obtain a large enough sample of insecure/preoccupied mothers and their infants. Future research should examine direct and intention mirroring in this group. It would be of interest to evaluate whether intention mirroring could discriminate between different subtypes of insecure attachment. Third, the present study did not measure contingency in maternal mirroring, a construct that Gergely emphasized alongside markedness and ostensiveness (Gergely & Unoka, 2008a; Gergely & Watson, 1999). Fourth, we did not code the valence of the infant's signals to which maternal mirroring was directed. There is evidence from brain imaging and neuroendocrine research that disrupted maternal attunement may be strongly characterized by the mother's disengagement from and denial of negative infant cues (Kim, Fonagy, Allen, & Strathearn, 2014; Kim, Fonagy, Koos, Dorsett, & Strathearn, 2013). One may therefore postulate that the low levels of intention mirroring we documented in insecure/dismissing mothers may be more specific to the infant's negative internal states (e.g., distressed state). This would be a fruitful area for further investigation. The present study is the first attempt to examine markedness and ostensiveness as distinguishing features of the mother's well- attuned, affect-mirroring communication with her infant. We have found evidence for high levels of marked, ostensive mirroring in securely attached mothers, who were also frequently the focus of their infants’ attention. Mothers with insecure/dismissing attachment were low in this form of mirroring, and were also less frequently the target of their infant's attention. While its direct links to infant behavior should be explored further in future research, the marked, ostensive mirroring may be more accurate than extant mirroring constructs in capturing the essence of securely attached mothers’ affective attunement to their infants."],["Background: Type 2 diabetes is a major public health problem. Effective diabetes self-management involves people engaging in multiple health behaviours, including physical activity. Walking is an effective, accessible and inexpensive form of physical activity, yet many people with Type 2 diabetes do not meet recommended levels. The present study aimed to: 1) identify demographic, motivational and volitional factors predictive of walking in people with Type 2 diabetes mellitus, and 2) test whether accounting for the perceived impact of other goal pursuits (goal facilitation and goal conflict) improved the prediction of walking. Methods: A theory-based cross-sectional study using the Health Action Process Approach was conducted in adults with Type 2 diabetes across Scotland. Assuming a 50% response rate 1000 questionnaires were mailed to achieve the target sample size (N = 500). Demographic information was collected, and intentional (outcome expectations, social support, risk perceptions), motivational (intention, self-efficacy), volitional (action planning, action control) and multiple goal (goal conflict, goal facilitation) factors were assessed as predictors of physical activity in general and walking specifically. Results: The final sample comprised 411 respondents. The majority (60%) were non-adherent to physical activity recommendations. Of 411 respondents, 356 provided walking data. Body Mass Index and age were the only demographic and anthropometric factors predictive of walking (overall R2 = 0.04). When motivational factors were added, intention and self-efficacy added to the prediction (overall R2 = 0.07). When volitional factors were added, only action control was predictive of walking (overall R2 = 0.08). Finally, goal facilitation explained an additional 7% variance in walking when added to the model (final overall R2 = 0.15). Conclusion: There was low adherence with physical activity recommendations in general and walking in particular. When testing predictors of motivational, volitional and competing goal constructs together, action control and goal facilitation emerged as predictors of walking. Future research should consider how walking can be embedded synergistically alongside other goal pursuits and how action control may help to ensure that they are pursued. --------------------------------------------------------------------------------","Diabetes is a common non-communicable chronic disease. The global prevalence of 8.3% is expected to increase to 10.1% by 2030 (IDF, 2013). In Scotland, the prevalence of diabetes is 4.7%, slightly above the UK average (S.D.S.M. Group, 2012). Almost 90% of patients with diabetes have Type 2 diabetes (WHO, 2006) and their life expectancy is up to 10 years less than people without Type 2 diabetes (Diabetes UK, 2012a, 2012b). Diabetes is a chronic, metabolic disease characterized by increased levels of blood sugar. Diabetes occurs either when the pancreas produces no or insufficient insulin, or when the body cannot effectively use the insulin it produces. Type 2 diabetes results from the body’s insufficient production and/or ineffective use of insulin. Hyperglycaemia (an increased concentration of glucose in the blood) is a common effect of uncontrolled diabetes and over time leads to serious damage to the heart, blood vessels, eyes, kidneys, and nerves (WHO, 2015). Type 2 diabetes has non-modifiable (genetic) and modifiable (environmental and behavioural) risk factors (Alberti, Zimmet, & Shaw, 2007). Genetic predisposition is aggravated by behavioural factors including smoking, being overweight, abdominal obesity and lack of physical activity (Stumvoll, Goldstein, & van Haeften, 2005). Good management of these behavioural factors can prevent or delay onset of diabetes, and many of its complications (WHO & IDF, 2004). The recommended regimen for managing Type 2 diabetes includes eating healthily, being physically active (moderate intensity) for at least 30 min on most days, smoking cessation, and taking medication (e.g., oral hypoglycaemic drugs, insulin, antihypertensive and lipid lowering drugs) (D.P.P.R. Group, 2002; Hallal et al., 2012; WHO, 2012; Zimmet, Alberti, & Shaw, 2001). Evidence suggests that regular physical activity reduces the risk of coronary heart disease, stroke, diabetes, hypertension, colon cancer, breast cancer and depression and is the main factor in weight control (WHO, 2010). For example, trials have demonstrated the benefits of undertaking physical activity in preventing Type 2 diabetes, improving glycaemic control and aerobic fitness, as well as decreasing the risk of cardiovascular disease and overall mortality (Sigal, Kenny, Wasserman, Castaneda-Sceppa, & White, 2006). Physical inactivity is the fourth leading risk factor for mortality worldwide, accounting for 6% of deaths (WHO, 2010), and approximately 30% of the disease burden due to diabetes and ischemic heart disease (WHO, 2010). There is evidence to suggest that patients with Type 2 diabetes engage in even less physical activity than the general population (39% versus 58%) (Morrato, Hill, Wyatt, Ghushchyan, & Sullivan, 2007), and the level of physical activity in those who do participate is low (Badenhop, 2006). However there is wide inter-country variation, with recent studies showing that adherence to recommended physical activity in Type 2 diabetes ranges between 9% and 69% (Broadbent, Donkin, & Stroh, 2011; Morrato et al., 2007; Nelson, Reiber, & Boyko, 2002; Plotnikoff, Brez, & Hotz, 2000; Serour, Alqhenaei, Al-Saqabi, Mustafa, & Ben-Nakhi, 2007; Shultz, Sprague, Branen, & Lambeth, 2001; Thomas, Alder, & Leese, 2004). A World Health Organization (WHO) report showed that adherence with physical activity recommendations by people with Type 2 diabetes ranged from 7.7% to 55% across different countries (WHO, 2003). The high prevalence of Type 2 diabetes, along with low level of physical activity, highlights the need for new approaches to improve an individual’s adherence to physical activity recommendations. These new approaches need to be acceptable, accessible and inexpensive to increase the probability of adoption. Walking as a specific form of physical activity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Walking is the most common form of physical activity and is an important component of total physical activity in adult populations (Monteiro et al., 2003; Morris & Hardman, 1997). Walking is acceptable, accessible and inexpensive; it requires no specific facilities, can be integrated easily into a daily routine, and is generally safe (Monteiro et al., 2003; Morris & Hardman, 1997). The energy expenditure of walking at a moderate pace of 5 km/h (3 miles/hour) can meet the definition of moderate intensity physical activity (Ainsworth et al., 2000). However, data from the National Health and Nutrition Examination Survey on 2896 patients with Type 2 diabetes in the US showed that 46% of participants did not report any walking for exercise (Gregg, Gerzoff, Caspersen, Williamson, & Narayan, 2003). While the literature to date on behavioural determinants of physical activity focuses on more generic descriptions of physical activity, given the above-mentioned benefits of walking, our aim was to focus specifically on understanding the factors associated with walking in people with Type 2 diabetes. Behavioural determinants of physical activity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There is a large body of evidence on the biological, sociological, psychological, and environmental factors that influence physical activity (Bonner, 2010). Non-modifiable factors (e.g., age, gender) can help to identify sub-groups that are likely to be physically inactive, whereas modifiable factors (e.g., intention, self-efficacy) provide potential targets for increasing physical activity (Schwarzer, 2008; Wing et al., 2001). A number of theories summarise the relationship between modifiable factors and behaviour to generate testable hypotheses. The Health Action Process Approach (HAPA) (Schwarzer, 1992) is a comprehensive social cognition model which accounts for motivational factors including outcome expectation, social support, risk perceptions, intention, and self- efficacy, and as well as contemporary theoretical development in volitional (post- intentional) processes including action planning and action control (Schwarzer, 2008). The HAPA describes intention as a function of self-efficacy, outcome expectations and risk perceptions. Intentional processes are then related to action via volitional processes involving planning and action control, further supported by self-efficacy and impacted by available barriers and facilitators such as social support (Schwarzer et al., 2003). Self- efficacy is a main influential factor, referring to a person’s perceived capability of performing a desired behaviour (Schwarzer et al., 2003). Outcome expectations refer to perceived positive and negative outcomes of engaging in the health behaviour; the more the beneficial outcomes and the fewer the negative outcomes that are perceived, the more likely it is that an individual will intend to engage in the behaviour (Schwarzer et al., 2003). Risk perceptions refer to the minimum level of perceived risk, which must exist before an individual starts to consider the benefits of possible behaviour and their capability to undertake those behaviours. Strong intention is an often necessary but rarely sufficient precondition for action (Orbell & Sheeran, 1998). Post-intentional (volitional) processes such as action planning and action control can help to ensure intentions are translated into action. Action planning involves linking goal-directed action to environmental cues by specifying the when, where, whom, and how to enact a behaviour to help translate intention into action (Darker, French, Eves, & Sniehotta, 2010; Gollwitzer, 1999). In addition, more active self-regulatory efforts can further supplement the translation of intention into action. Action control, i.e., self-monitoring of behaviour, being aware of monitoring standards and expending effort in goal pursuit, is a self-regulatory process for ensuring intention enactment (Carver & Scheier, 1982; Sniehotta, Scholz, & Schwarzer, 2005). The HAPA has been applied to understand physical activity across numerous studies. Some studies focus on the entire HAPA model (Barg et al., 2012; Bonner, 2010; Caudroit, Stephan, & Le Scanff, 2011; Renner, Spivak, Kwon, & Schwarzer, 2007; Scholz, Schuz, Ziegelmann, Lippke, & Schwarzer, 2008; Scholz, Sniehotta, & Schwarzer, 2005, 2008; Schwarzer et al., 2007; Sniehotta, Scholz, et al., 2005; Sniehotta, Schwarzer, Scholz, & Schüz, 2005), whilst others focus on more specific components of the model (Barg et al., 2012; Lippke, Ziegelmann, & Schwarzer, 2005; Schwarzer et al., 2007; Sniehotta, Scholz, & Schwarzer, 2006; Sniehotta, Schwarzer, et al., 2005). Few studies have applied the HAPA to the behaviour of people with Type 2 diabetes. Bonner (Bonner, 2010) used the HAPA in Type 2 diabetes and showed that self- efficacy and outcome expectations were predictive of physical activity intention, and intention (but not self-efficacy or action planning) predicted physical activity levels. No study has yet used the HAPA model to understand physical activity in people with Type 2 diabetes focusing specifically on walking as an inexpensive and accessible form of physical activity (Lippke & Plotnikoff, 2014). Towards multiple behaviour approaches ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Most popular social cognition models of health behaviour focus on understanding a single health behaviour at a time. The ecological validity of such an approach has increasingly been questioned (Presseau, Tait, Johnston, Francis, & Sniehotta, 2013). In everyday life, individuals pursue multiple goals and perform multiple behaviours alongside the single health behaviour that is typically the focus of tests of behavioural theory. These goal pursuits compete for time and energy such that pursuit of some may help and/or hinder the pursuit of a particular health behaviour, such as physical activity in general or walking specifically. The extant literature has predominantly managed the concept of considering multiple goals by focusing on the impact of goal conflict on health behaviour. Goal conflict can be described as occurring when the pursuit of multiple personal goals leads to situations where they interfere with one another. For instance, working, childcare, relaxing and socialising may be common personal goals that have the potential to conflict with walking by taking available leisure time, energy or other resources that might otherwise be used go for a walk. The evidence on the link between goal conflict on physical activity-related behaviour is mixed. There is a lack of support for this relationship in between-subject predictive studies (Li & Chan, 2008; Presseau, Sniehotta, Francis, & Gebhardt, 2010; Riediger & Freund, 2004). However, a study investigating actual time spent pursuing goals that conflict with physical activity within-subjects was negatively predictive of objectively assessed physical activity (Presseau et al., 2013), and a study investigating goal conflict in more resource constrained contexts has also shown that goal conflict is negatively predictive of behaviour (Presseau, Francis, Campbell, & Sniehotta, 2011). As people with Type 2 diabetes engage in self-management regimens that inherently involve pursuing multiple behaviours and goals, it is plausible that goal conflict may be a useful additional construct in this population. By comparison, goal facilitation has received less research than goal conflict, yet is recurrently shown to be predictive of physical activity-related behaviours. Goal facilitation involves instances where the pursuit of other personal goals sets the stage or makes it more likely that physical activity will take place (e.g. socialising with friends that involves walking in the park), or inherently involves physical activity (e.g. commuting to work can be facilitative of physical activity when involving active travel). The presumption is that the more one’s other personal goals are aligned with physical activity, the greater the physical activity. Goal facilitation has been demonstrated to positively predict physical activity (Riediger & Freund, 2004), a relationship that is maintained even when controlling for intention and self-efficacy (Presseau et al., 2010). However, it is not clear whether these relationships persist when accounting for volitional (planning, action control) processes, which could in themselves involve managing competing goals. For instance, action planning may involve describing other goals that facilitate engaging in physical activity, whereas coping planning may involve identifying barriers that in themselves are actually competing goal pursuits (Presseau, Boyd, Francis, & Sniehotta, 2015). This conceptual overlap issue could be addressed empirically by investigating whether indicators of goal conflict or goal facilitation remain predictive of physical activity when controlling for volitional factors. Furthermore, it is not clear how either goal conflict or goal facilitation relate to walking behaviour specifically, which may have different levels of perceived conflict and facilitation than other forms of more intensive physical activity. The present study aimed to: 1) identify demographic, motivational and volitional factors predictive of walking in people with Type 2 diabetes, and 2) test whether accounting for the perceived impact of goal pursuits (goal facilitation and goal conflict) improved the prediction of walking.","This was a cross-sectional theory-informed postal questionnaire study undertaken with people with Type 2 diabetes from the Grampian and Tayside regions of Scotland. All English-speaking adults (>18 years) diagnosed with Type 2 diabetes were eligible to participate. Patients with serious end stage illness and patients with mental disability were excluded. Questionnaire development ~~~~~~~~~~~~~~~~~~~~~~~~~ A qualitative study was initially conducted using the Theoretical Domains Framework (TDF) (Michie et al., 2005) to identify which theoretical domains and constructs were relevant to understanding the adherence of people with Type 2 diabetes to physical activity recommendations in general and walking in particular. The results were used to identify relevant items that were included in a draft questionnaire. The questionnaire explored physical activity in general, and walking in particular. Pre-piloting of the questionnaire was undertaken with five people using the “think aloud” method (Jones, 1989; Lundgren- Laine & Salantera, 2010) where participants verbalised their thoughts. Three participants with Type 2 diabetes were recruited from the Scottish Diabetes Research Network (SDRN) (see later) and three were colleagues with Type 2 diabetes in the Centre of Academic Primary. Minor revisions were made prior to the pilot study. The questionnaire was piloted with 50 people with Type 2 diabetes, selected randomly from the SDRN list, replicating the distribution process planned for the main survey (pre-notification letter, questionnaire and covering letter, and a reminder letter and replacement questionnaire after two weeks). To assess test-retest reliability, respondents were sent a second copy of the questionnaire two weeks after returning their first questionnaire. Sample and recruitment ~~~~~~~~~~~~~~~~~~~~~~ The sample size for this study was influenced by two factors: 1) having acceptable precision for the estimation of adherence with physical activity (any precision within ±5% that would be clinically and statistically acceptable) and 2) the resources (time and money) available to undertake the research. To achieve a balance between these two items, a sample size of 500 patients was required. As previous research has shown compliance with physical activity to range from 19 to 30% (midpoint: 25%) (Kamiya et al., 1995; Kravitz et al., 1993), this allowed estimation of adherence with physical activity of 25% with precision within ±3.8% (95% CI 21.2%–28.8%). Previous research in community samples indicated a 50% response rate was likely, therefore 1000 questionnaires were mailed to achieve the target of 500 evaluable responses. Participants were recruited from the Scottish Diabetes Research Network (SDRN), a register of patients with diabetes in Scotland who have consented to be contacted about potential participation in research studies (SDRN, 2010). All SDRN registered patients in Grampian (n = 388) were identified and invited to participate, supplemented by a random sample of 612 of the 1279 patients registered in Tayside exclusive of those who had taken part in the pilot study. A pre- notification letter with a reply slip, that they could use if they did not want any further communication, was sent to these 1000 patients two weeks before the questionnaire and accompanying invitation letter were mailed. Two weeks after the first mailing, a reminder letter and another copy of the same questionnaire were sent to non-respondents. The questionnaire was piloted with 50 people with Type 2 diabetes, selected randomly from the SDRN list, replicating the distribution process planned for the main survey (pre- notification letter, questionnaire and covering letter, and a reminder letter and replacement questionnaire after two weeks). To assess test-retest reliability, respondents were sent a second copy of the questionnaire two weeks after returning their first questionnaire. Physical activity and walking The questionnaire included items assessing time spent being physically active in the last seven days based on the short version of International Physical Activity Questionnaire (IPAQ) (IPAQ, 2002). It measures physical activity over a short time frame. The IPAQ was developed by consensus in 1998–1999 with support from the WHO to enable the cross-national assessment of physical activity in adults aged 18–65 years (Craig et al., 2003; Macfarlane, Lee, Ho, Chan, & Chan, 2007; Papathanasiou et al., 2010). The short format of the IPAQ asks about three types of activity in the four domains. Walking, moderate-intensity activities and vigorous-intensity activities are the specific types of activity which are assessed by the IPAQ short form (IPAQ 2002). This version generates a total score by summation of the duration (in minutes per day) and frequency (days) of walking, moderate-intensity activities and vigorous-intensity activities. The IPAQ measures energy as Metabolic Equivalent of Task (MET). The IPAQ has been used in a number of international studies (Craig et al., 2003; Guthold, Ono, Strong, Chatterji, & Morabia, 2008) and acceptable reliability and validity has been reported (Craig et al., 2003; Hagstromer, Oja, & Sjostrom, 2006; Hallal et al., 2010; Macfarlane et al., 2007; Papathanasiou et al., 2010). An international reliability and validity test of the IPAQ was conducted in 14 centres in 12 countries and reported that it has acceptable reliability and validity at least equal to other established self- report tools for physical activity in diverse populations of 18–65 years (Craig et al., 2003). We focused specifically upon understanding predictors of walking as the primary outcome of interest given the wording of our predictors focused upon walking. Walking was assessed using the total time or energy (150 min or >600 MET minutes/week) spent on walking measured by the IPAQ and served as the dependent variable in all predictive analyses. However we also aimed to describe overall adherence to physical activity recommendations. Adherence to physical activity was assessed by comparison with two different recommendations. Firstly it was assessed by comparison with the Scottish Intercollegiate Guideline Network/WHO (SIGN, 2010; WHO, 2010) advice of at least 150 min of vigorous/moderate (no walking included) combined physical activity per week (equal to at least 600 MET1). Secondly it was assessed accordingly to the IPAQ criterion of 600 MET minutes/week of any combination of walking, moderate-intensity or vigorous-intensity physical activities (IPAQ, 2002). According to IPAQ, <600 MET minutes/week, 600–2999 MET minutes/week, and >3000 MET minutes/week are considered as low, moderate and vigorous physical activity, respectively (IPAQ, 2002). Predictors of walking The questionnaire assessed a number of potential demographic and theoretical predictors of walking: demographic variables, self-efficacy, outcome expectations, risk perceptions, intention, action planning and control, social support, goal facilitation and goal conflict. The demographic variables age, gender, education, and employment items were defined using the England household version of the 2001 Census questionnaire (OFNS, 2002). All theoretical items were worded according to the TACT principle (Target, Action, Context, and Time), specifying the behaviour of interest as: “To increase (my) own walking level by 20% during the normal daily routine in the forthcoming month” and described in detail below. Self-efficacy Self-efficacy was assessed using six items ranging from 1 (strongly disagree) to 5 (strongly agree) in relation to perceived capability to increase walking despite the presence of barriers (Schwarzer et al., 2003). The stem “I am confident that I can increase my walking by 20% in the next month even if ….” had response options such as: “the weather is bad”, “it is hard for me physically”, “I do not have much time”. Outcome expectations Two facets of outcome expectations were assessed (Schwarzer et al., 2003), with scores for each item ranging from 1 (not at all) to 4 (exactly true): there were six items to assess positive outcome expectations, and three items to assess negative outcome expectations. The stem “if I increase my walking by 20% in the next month ….” had response options such as: “I would feel better afterwards”, “it would take up a lot of time”. Risk perception Risk perception refers to the respondent’s belief about their vulnerability to health problems, or specifically in this patient group for their diabetes to worsen (Schwarzer et al., 2003). Absolute and relative vulnerability were assessed using six items with response options ranging from 1 (strongly disagree) – 7 (strongly agree). The items measuring absolute vulnerability had a stem “If I am not physically active … ” and response options such as: “ … I am concerned that my health in general will become worse”, “ … I am concerned that my diabetes in general will become worse”, “ … I will worry about getting a serious medical condition”. The items measuring relative vulnerability had a stem “If I am not physically active … ” comparing myself with an average person of my age and sex, then I will be at higher risk of … and response options such as: “ … my diabetes gets worse”, “… having a serious medical condition”. Intention Intention refers to a participant’s intention to increase walking (Schwarzer et al., 2003) and was assessed by four items with response options ranging from 1 (completely disagree) to 5 (totally agree). Intention was measured by items such as “I intend to walk more in the next month” and “I am motivated to walk more to improve my health in general”. Action planning Action planning consisted of items assessing the extent to which participants had a plan about when, where, and how to increase their walking (Schwarzer et al., 2003). Action planning was assessed using four items (Sniehotta, Scholz, et al., 2005; Sniehotta, Schwarzer, et al., 2005). All items had response options ranging from 1 (completely disagree) to 4 (totally agree). The stem “I have made a specific plan about ….” had response options such as: “… when to increase my walking in the next month”, “ … where to increase my walking in the next month”, “ … what to do if something interferes with my intention to increase my walking in the next month”. Action control Action control refers to perceived self-monitoring, awareness of standards and effort (Sniehotta, Scholz, et al., 2005; Sniehotta, Schwarzer, et al., 2005) to increase walking of participants. Action control was assessed using six items and all items had response options ranging from 1 (strongly disagree) to 4 (strongly agree). The stem “During the last week I ….” had response options such as:“… regularly thought about my intention to be regularly physically active”, “ … I have consistently checked to see whether I am physically active enough”. Social support Social support items assessed support from colleagues, friends and household members to increase walking using a modified version of the Molloy social support tool (Molloy, Dixon, Hamer, & Sniehotta, 2010). All items (17 items) had response options ranging from 1 (strongly disagree) to 7 (strongly agree). Social support (friends/colleague) was measured by items such as “I have a friend/colleague who thinks that I should increase my walking”, and “I have a friend/colleague who encourages me to increase my walking”. Social support (household) was measured by items such as “I have somebody to encourage me to increase my walking on the regular basis”, and “I have somebody to walk with me”. Goal conflict and goal facilitation Goal conflict (5 items) and goal facilitation (3 items) items focus on the extent that a participant’s personal goals conflicted with physical activity and were adapted from general goal conflict and facilitation scales (Riediger & Freund, 2004). All the items had response options ranging from 1 (never, not at all, or completely disagree) to 5 (very often, a great deal, or completely agree). The items measuring goal conflict consisted of a stem “How often does it happen that, because of the pursuit of another personal goal, you do not invest ….” and response options such as: “ … as much time in participating in regular physical activity as you would like to?”, “ … as much energy in participating in regular physical activity as you would like to?” Goal facilitation was measured by items such as “To what extent do other things you do in everyday life help you to participate in regular physical activity?”, and “How often does it happen that you do something in pursuit of a personal goal that is simultaneously beneficial for participating in regular physical activity?” Data management and analysis ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Data were entered into SPSS version 20 and 10% of all data were double entered and checked for quality assurance. Few errors (n = 11 or 0.1% of entered fields) were identified and corrected, with no evidence of systematic errors. The primary outcome measure was the IPAQ walking criterion (MET minutes/week). A sensitivity analysis was conducted using total MET minute/week. The extent of missing data varied across variables. The variables with the greatest and smallest amount of missing data were walking level (13.4%), and diabetes management method (2.1%). We used multiple imputation (Klebanoff & Cole, 2008) to account for missing data which addresses missing data issues in the most robust manner possible. All model testing was conducted on multiple imputed data and results presented as pooled estimates. Hierarchical multiple regression analyses were conducted to test the sequential contribution of demographic, motivational, volitional and multiple goal constructs as predictors of walking. Ethics approval ~~~~~~~~~~~~~~~ Ethics approval for the study was granted by North of Scotland Research Ethics Committee (NRES) (Ref 10/S0802/4). Response rate ~~~~~~~~~~~~~ Of 1000 people contacted, 35 withdrew at the pre-notification letter stage. Of the 965 questionnaires mailed, 426 were returned (compared to the target sample size of 500). Of these fifteen were excluded (five received after the agreed deadline (15/07/2012), seven with excessive (>90%) missing data, three because of participation in the pilot study). Most questionnaires (373/426; 87.6%) were returned by people who had responded to the pre- notification letter. No significant difference was found between participating and non- participating respondents in terms of gender and age suggesting the final sample was representative. The final evaluable sample comprised 411 respondents. Socio-demographic characteristics of respondents ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The mean age of respondents was 65.5 years (SD 9.7); 57.4% (n = 236) were men. Most were married (60.6%), did not live alone (63.3%) or were retired (62.3%). A quarter (26%) had no formal educational qualification. Most participants (92.7%) were either overweight (BMI 25.0–29.9) or obese (BMI ≥ 30.0). The mean average BMI was 34.0 (SD 5.9) and 31.4 (SD 5.1) for women and men, respectively. Descriptive statistics and bivariate correlations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As shown in Table 1, which presents findings across all 411 respondents, the mean total physical activity measured as Metabolic Equivalent of Task (METs) was 1732 min/week (Inter Quartile Range (IQR) 485, 4398; median 200). Based on SIGN and WHO guidelines, which exclude walking, almost 60% (n = 236) of patients did not adhere to physical activity recommendations (<600METs); however this proportion was reduced to 28% using the IPAQ (Metabolic Equivalent of Task (MET) minutes/week) measure which includes walking (Table 2). Men had higher median levels of physical activity than women. According to the IPAQ categories nearly 36% and 35% of participants reported moderate and vigorous levels of physical activity during the last week (Table 2), but the median time (hours/week) spent for both moderate and vigorous physical activity was zero (Table 1). The median duration of walking was 5.25 h per week. The proportion of total physical activity reported as walking was 65.6%. As shown in Table 3, which presents findings for the 356 respondents providing walking data, BMI, action planning, action control and goal facilitation were significantly associated with walking behaviour, and outcome expectations, social support, risk perceptions, self-efficacy, action planning, action control, and goal conflict were significantly associated with walking intention. The Cronbach’s alpha of different subscales of HAPA questionnaire are presented in Table 3 indicating that most subscales of the questionnaire had a good internal consistency. The negative outcome expectations scale was omitted from any analyses due to low observed internal consistency. Predicting walking ~~~~~~~~~~~~~~~~~~ The hierarchical multiple regression was conducted in four steps. First, demographic factors and predictors of intention from the HAPA were included. Next, motivational factors from HAPA were added, then volitional, and finally multiple goal constructs. At each step, we tested whether the added factors contributed to explaining additional variance in walking beyond factors in the model from the previous steps, and which specific constructs explained this additional variance. In Step 1 of the hierarchical multiple regression, walking was regressed against demographic factors (BMI, age, sex) and HAPA theory-based predictors of intention (outcome expectations, social support and risk perception). As shown in Table 4, only BMI and age predicted walking, explaining 3.7% of the variance in walking. In Step 2, HAPA motivational constructs (intention and self- efficacy) were added, with intention and self-efficacy adding to the prediction (ΔR2 = 0.03). In Step 3, the volitional constructs of action planning and action control were added, with only the latter adding significantly to the prediction (ΔR2 = 0.01) and intention and self-efficacy no longer significantly contributing to predicting behaviour. In Step 4, the multiple goal constructs of goal conflict and goal facilitation were added, with the latter significantly adding to the prediction of behaviour (ΔR2 = 0.07) whilst action control no longer significantly predicted behaviour.","The study showed that the majority (60%) of Type 2 diabetic patients were non-adherent to physical activity recommendations as defined by SIGN/WHO. Most of the physical activity undertaken by people with Type 2 diabetes was walking (65.6%). Action control and goal facilitation were predictive of walking. Goal facilitation explained a further 7% of the walking variance. Non-compliance of the majority of respondents with the SIGN recommendation (SIGN, 2010), for physical activity is consistent with the Scottish Health Survey (The Scottish Government, 2012) which showed that 61% of the general population aged 16 and over did not meet physical activity recommendations. Other evidence suggests that patients with Type 2 diabetes may be even less physically active than the general population (Morrato et al., 2007). This was also the finding of a study in USA of 23,283 adults, which showed that only 39% of individuals with Type 2 diabetes were physically active compared with 58% of those without diabetes (Morrato et al., 2007). The median duration of walking reported in the current study was 5.25 h per week (IQR 1.5, 12). The proportion of walking as a percentage of total physical activity was 65.6% suggesting that in some cases walking was the main type of physical activity undertaken by patients. This finding reflects the behaviour of the general adult population (Monteiro et al., 2003; Morris & Hardman, 1997); therefore developing and evaluating interventions to increase and maintain this behaviour are important. Walking is a common, accessible, inexpensive Type of physical activity. Walking provides diverse health benefits of physical activity with few adverse effects. There is a large body of evidence about the positive effect of walking to improve health in people with Type 2 diabetes. This suggests that focusing on walking as a form of physical activity to improve peoples’ adherence with physical activity recommendations is important and could be an effective way to improve physical activity. In terms of the existing literature one study conducted with cardiac rehabilitation patients, was found that measured action control as a predictor of physical activity (Sniehotta, Scholz, et al., 2005; Sniehotta, Schwarzer, et al., 2005). That study reported that each of the three factors of planning, self-efficacy and action control made unique contributions to translating intention into action (Sniehotta, Scholz, et al., 2005; Sniehotta, Schwarzer, et al., 2005). A study conducted in students confirmed associations specified by the HAPA at the intrapersonal level: outcome expectancies and self-efficacy, but not risk awareness, were positively associated with intentions for physical exercise. Physical activity was positively associated with intentions, self- efficacy, action control, but not with action planning (Scholz, Keller, & Perren, 2009). These findings are in accordance with the results of this current study. Another study conducted in Type 2 diabetic patients participating in a Diabetes Self-Management Education (DSME) (Bonner, 2010) showed that self-efficacy was the strongest predictor of behavioural intention, followed by positive outcome expectancy. The study (Bonner, 2010) revealed that behavioural intention, but not self-efficacy and action planning could significantly increase initiation of a minimum level of physical activity. The current study showed some degree of support for the tenets of the HAPA, whilst demonstrating the importance of considering multiple goal pursuit in people with Type 2 diabetes. The majority of respondents did not engage in physical activity at recommended levels. Action control and goal facilitation were shown to be predictors of physical activity when considered alongside other HAPA and demographic factors. Findings in relation to the HAPA with respect to intention (step 2 of the regression) and action control (step 3) were consistent with previous research (Sniehotta, Scholz, et al., 2005; Sniehotta, Schwarzer, et al., 2005) and extend these findings by demonstrating the role for multiple goals constructs on physical activity (in this case, goal facilitation). Conversely we did not show a predictive role for action planning and in step two, there is an unexpected negative predictive relationship between self-efficacy and walking behaviour, although this becomes insignificant when the additional predictors in steps three and four are added. Both findings are at odds with the HAPA model and most of the literature investigating these relationships (Sniehotta, Scholz, et al., 2005; Sniehotta, Schwarzer, et al., 2005). Self-efficacy showed no significant bivariate relationships with walking which may be due to the fact that the target behaviour was ‘increasing walking by 20%’ which equates to large absolute changes for more active respondents. Moreover, self- efficacy was significantly correlated with intention, so that the negative beta- coefficient in the second step of the hierarchical regression analysis may be reflective of an artefact, a statistical suppressor effect. Action planning showed a weak bivariate correlation with walking and was significantly correlated with action control so that when action control was simultaneously controlled for, there was not a unique predictive relationship between action planning and walking. In the final model, neither of these variables was significant. Findings regarding multiple goal constructs are also consistent with earlier research showing that perceived goal facilitation but not perceived goal conflict were predictive of physical activity (Presseau et al., 2010, 2013; Riediger & Freund, 2004). There is now growing evidence across a range of studies with diverse populations that particularly support the role of goal facilitation as a key factor in physical activity and, with the present study’s findings, walking specifically. Goal facilitation is an indicator of the extent to which a target behaviour (in this case, walking) “fits” synergistically alongside the other behaviours and goals that individuals pursue in daily life. Findings from this study continue to support the role of goal facilitation and also underscore its potential importance in understanding health behaviours; indeed, even when controlling for predominant theoretical constructs reported in the literature, the relationship between goal facilitation and walking robustly accounted for additional variability in walking. With increasing recognition of the importance of considering the wider context of multiple goal pursuit when understanding performance of a given health behaviour, the present study further contributes evidence suggesting that goal facilitation may be a key indicator in the move towards developing models that explicitly account for the impact of multiple goal pursuit. There is also mounting lack of support for the role of goal conflict in understanding physical activity. There may be a range of reasons for this. For instance, when considering the totality of an individual’s goal pursuits, individuals may be better able to perceive helpful goal relationships than conflicting ones. Individuals may not be aware of the extent that their competing goals interfere with their physical activity. When using diaries to assess actual time spent in pursuit of goals that conflict with physical activity over time, goal conflict has been shown to be predictive of objectively assessed physical activity (Presseau et al., 2013). This suggests that measures of perceived goal conflict may need to be supplemented with behavioural assessments. This also presents opportunities for feedback interventions by showing individuals which of their behaviours is most interfering with their physical activity. In addition, when focusing the goal pursuit context to a specific time and place rather than all of everyday life, both goal conflict and goal facilitation have been shown to predict behaviour (Presseau et al., 2011). The utility of the HAPA to explain and possibly predict adherence with physical activity in addition to the demonstrated added contribution of considering goal facilitation suggests clear opportunities for developing and evaluating novel, theory-based interventions for promoting walking in people with Type 2 diabetes. The present study extends the literature by demonstrating the role of multiple goal pursuit and goal facilitation in particular in a population sample of people with Type 2 diabetes. In addition, the findings extend the theoretical literature by demonstrating that goal facilitation predicts independent variability in health behaviour over and above all contemporary single-behaviour cognitions. This is important as it provides further evidence for moving beyond on of health behaviours in isolation. This study is the first to specifically consider the role of goal facilitation in relation to walking by people with Type 2 diabetes. The importance of goal facilitation as a key predictor of walking, points to possible interventions to increase walking behaviour. Indeed, Darker et al. (Darker et al., 2010) used a variation of action planning – facilitation planning – in their walking intervention, which was successful in increasing and maintaining the increased walking behaviour. Planning when, where and how to perform behaviours may facilitate action. To some extent, these may be preparatory behaviours, but goal facilitation encompasses the broader spectrum of valued goals pursued in everyday life and may not necessarily be preparatory in nature, whereas preparatory behaviours may not have any intrinsic value to the actor. Nevertheless, the functional similarities between preparatory behaviours and goal facilitation are noteworthy and future research should consider these two constructs in more detail. Strengths and limitations ~~~~~~~~~~~~~~~~~~~~~~~~~ The present study is strengthened by its large sample size, robust development and inclusion of theoretical factors as determinants of walking. Although the sample of 411 (356 for the main analysis) was slightly short of the target of 500, this did not impact substantially upon the precision of the estimates achieved: 40% with precision within ±4.7% (95% CI 35.3%–44.7%) of respondents being categorised as adherent with physical activity recommendations compared with the original estimate of 25% with precision within ±3.8% (95% CI 21.2%–28.8%). The study also had limitations. Firstly, the cross-sectional study design only allows association, and not causation, to be inferred. While there is no obvious suggestion of multicollinearity, the modest bivariate correlations between predictors in the model should be considered in interpreting the relative contribution of predictors in the model, particularly with respect to factors which were not zero-order correlations, and were not bivariately associated with walking but which were associated with walking when included in the multivariate analyses (i.e. age and self-efficacy). Future research should aim to replicate findings using a prospective design or by embedding such questionnaires in a theory-based process evaluation alongside a trial (Sedgwick, 2014). A further limitation is that the study may have overestimated levels of physical activity in people with Type 2 diabetes. People living in Grampian and Tayside have slightly better self-reported general health than the total population of Scotland (72% and 69.6% in Grampian and Tayside, respectively versus 67.9% in Scotland) (The Scottish Census, 2011). Therefore, their self-reported physical activity, used as the main outcome in this study, may also be higher than the general national population. A further cause of over estimation could be that due to the patient population in the current study i.e. patients with Type 2 diabetic registered with the SDRN may be more engaged with their disease management compared with patients not registered with the SDRN. Social desirability bias could also contribute to any over-estimation of self-reported physical activity. The IPAQ has in fact been shown to overestimate self-reported time spent in physical activity compared with accelerometer measured activity (Ekelund et al., 2006; Hallal et al., 2012). The assessment of physical activity in the population (Van Hees, 2012) is challenging. Some tools include any type of walking as a physical activity (e.g. the IPAQ) (IPAQ, 2002) whereas other scales (e. g. The Rapid Assessment of Physical Activity [RAPA]) (University of Washington, 2006) do not. The recommended level of physical activity is at least 150 min of vigorous/moderate combined physical activity, in both SIGN and WHO guidelines (WHO, 2010). If walking is considered a physical activity, (SIGN, 2010) 72% of participants were compliant with the guidance, but this reduces to 40% if walking is not included, as in the SIGN guideline. This demonstrates the variation which arises when different tools are used. The use of an internationally relevant and valid tool allows comparisons to be made across studies. Finally, items for some constructs (risk perceptions, action control, goal conflict, goal facilitation) were measured in reference to physical activity, whereas others (outcome expectancies, social support, action planning, intention and self-efficacy) referred specifically to walking. While no obvious pattern of association seemed to preference one or the other conceptualization and walking is inherent to physical activity, future research could ensure greater correspondence of all items with walking.","Low physical activity in people with Type 2 diabetes is an important factor in terms of disease management. The majority of respondents did not engage in physical activity or walking at recommended levels. When testing motivational, volitional and competing goal constructs together as predictors of walking, Action Control and Goal Facilitation were shown to predict walking and could form the basis for developing novel, theory-based interventions for promoting walking in people with Type 2 diabetes.","The authors declare that they have no competing interests.","Masoumeh Namadian (m.namadian@zums.ac.ir) Social Determinants of Health Research Centre, Zanjan University of Medical Sciences, Iran and Centre of Academic Primary Care, University of Aberdeen, UK; Margaret C. Watson (m.c.watson@abdn.ac.uk) and Christine M. Bond (c.m.bond@abdn.ac.uk) Centre of Academic Primary Care, University of Aberdeen, UK; Justin Presseau (jpresseau@ohri.ca) Centre for Practice-Changing Research, Ottawa Hospital Research Institute, Ottawa, Canada; Falko F. Sniehotta (falko.sniehotta@newcastle.ac.uk), Institute of Health and Society, Newcastle University, UK. This article is based on data reported in the first author’s doctoral dissertation."],["People read dominance, trustworthiness and competence into the faces of politicians but do they also perceive such social qualities in other nonverbal cues? We transferred the body movements of politicians giving a speech onto animated stick-figures and presented these stimuli to participants in a rating-experiment. Analyses revealed single body postures of maximal expansiveness as strong predictors of perceived dominance. Also, stick-figures producing expansive movements as well as a great number of movements throughout the encoded sequences were judged high on dominance and low on trustworthiness. In a second step we divided our sample into speakers from the opposition parties and speakers that were part of the government as well as into male and female speakers. Male speakers from the opposition were rated higher on dominance but lower on trustworthiness than speakers from all other groups. In conclusion, people use simple cues to make equally simple social categorizations. Moreover, the party status of male politicians seems to become visible in their body motion. --------------------------------------------------------------------------------","Nonverbal cues affect impression formation (Ambady, Bernieri, & Richeson, 2000; Borkenau, Mauer, Riemann, Spinath, & Angleitner, 2004) and decision making in the public arena (Rule & Ambady, 2011; Zebrowitz & Montepare, 2008). For instance, attributions of dominance, trustworthiness, competence and other personality traits to specific facial features of political candidates can be reliable predictors of hypothetical and actual election outcomes (Antonakis & Dalgas, 2009; Chen, Jing, & Lee, 2014; Little, Roberts, Jones, & DeBruine, 2012; Olivola & Todorov, 2010; Oosterhof & Todorov, 2008; Poutvaara, Jordahl, & Berggren, 2009). Apart from facial and other nonverbal cues people also read socially relevant information from and into body motion. They recognize emotions in arm movements (Pollick, Paterson, Bruderlin, & Sanford, 2001), in whole body gestures (Atkinson, Tunstall, & Dittrich, 2007), and in movements displayed during interpersonal dialog (Clarke, Bradshaw, Field, Hampson, & Rose, 2005). Moreover, they perceive personality traits in dance movements (Hugill, Fink, Neave, Besson, & Bunse, 2011) and in walking behaviors (Thoresen, Vuong, & Atkinson, 2012), and health related cues and personality in the body movements of politicians giving a speech (Koppensteiner, 2013; Koppensteiner & Grammer, 2010; Kramer, Arend, & Ward, 2010). Humans and animals appear to use expansive body postures, expressive body movements, and broad gestures to display power and dominance (Carney, Hall, & LeBeau, 2005; De Waal, 2007; Eisenberg & Reichline, 1939; Mehrabian, 1972; Tiedens & Fragale, 2003). Also, people adopting open and expansive postures (i.e., power posing) have an enhanced sense of power (Carney, Cuddy, & Yap, 2010; Huang, Galinsky, Gruenfeld, & Guillory, 2011; Park, Streamer, Huang, & Galinsky, 2013). All this implies that there is a link between dominance and expansiveness in body postures and body movements. Communicating dominance as well as building connections to followers are vital abilities for leaders and politicians. Skilled self-presenters communicate their dominance by reassuring followers in noncompetitive contexts while threatening those who would jeopardize group stability (Stewart, Salter, & Mehu, 2009; Stewart, Waller, & Schubert, 2009). Consequently, speakers may not only present themselves differently according to their personality and their rhetorical skills but also according to situational factors such as their role in parliament. In the present study we selected brief video clips of politicians and mapped the body movements of the speakers onto animated stick-figures to control for appearance features. Then we asked people to judge these stimuli on dominance as well as on two additional basic social categories, namely trustworthiness and competence (Fiske, Cuddy, & Glick, 2007; Oosterhof & Todorov, 2008). In line with previous research we examined to what degree expansiveness in body postures and in body motion is related to judgments of dominance. To track down the relative contributions of body motion and expansiveness, we also investigated the influence of the quantity of motion. Analyses of the relationship of perceived trustworthiness and perceived competence with nonverbal cues were exploratory, because there is no research upon which to derive clear hypotheses. Our measures were simple because first impressions of speakers' body movements seem to be guided by simple and salient cues (Koppensteiner, 2013). This study pursued several aims. In accordance with previous research we intended to show that people's impressions of a speaker's dominance are guided by expansive body postures and expansive body movements. In addition, we investigated whether the quantity of body motion also influences ratings of dominance. We also explored whether ratings of trustworthiness and competence, which have been shown to be important qualities in judgments of faces, are related to our “nonverbal measures”. Males tend to challenge someone else's status by dominance contests (Mazur & Booth, 1998). Moreover, male speakers appear to display motion behaviors that lead to higher ratings of extraversion than the motion behaviors displayed by female speakers (Koppensteiner & Grammer, 2011) and perceived extraversion of speakers' motion behaviors is positively related to perceived dominance (Koppensteiner, Stephan, & Jäschke, 2015). Thus, we expected male speakers to show more dominance displays than female speakers. Finally, we examined whether speakers (i.e., their stick-figure animations) from the opposition and speakers from the government are judged differently on dominance, trustworthiness and competence and whether such differences show an interaction with gender.","At locations throughout the University we recruited 60 participants (33 females and 27 males; age M = 24.2 years, SD = 4.1) to take part in our rating experiment (see also Koppensteiner et al., 2015). Participants received a financial compensation of €5. Stimulus preparation ~~~~~~~~~~~~~~~~~~~~ Using a random number generator we selected 60 speeches from parliamentary sessions (29.11.2012, 30.11.2012, 14.12.2012) of the German Parliament. Deviations from random selection were necessary to reach equal numbers of male and female speakers (i.e., 30 different males and females) and nearly equally sized groups representing the parties. From each of the selected speeches we extracted brief, randomly chosen video segments with an average length of 15 s. Thirty-two speakers belonged to opposition parties (i.e., SPD, Bündnis 90/Die Grünen, Die Linke) and 28 speakers belonged to the government (i.e., CDU, FDP). To encode behavior we used the program SpeechAnalyzer, which runs through a movie frame by frame. In the first frame of each video clip, so called landmarks were positioned on the speaker's forehead, the hollow of the throat between the collar bones, ears, shoulders, elbows, hands, a spot in the middle of the body near the navel, and at the corners of the lectern (see also Koppensteiner, 2013; Koppensteiner & Grammer, 2010). Shifts in the positions of these body regions were automatically tracked by software routines based on optical flow (e.g., landmark of left shoulder in frame one was moved to position of left shoulder in frame two). As these software routines are prone to error, landmark positions often had to be corrected doing drag and drop operations with the computer mouse. This procedure of behavior encoding on the basis of landmark shifts yielded a time series of two dimensional marker positions that were used to create stick- figure animations (Fig. 1) representing the speakers' body movements. We only used every third frame in the encoding process; linear interpolation was used to fill in missing frames.","Participants were brought to our laboratory and asked to rate stick-figure animations of speakers. They received instructions on how to use the rating program and performed the rating tasks on their own (i.e., no experimenter present) using a computer-controlled interface. Stick-figure video clips were presented on the left-hand side of the user interface; rating scales that were named dominant, trustworthy, and competent were displayed on the right hand side. Participants completed their ratings by dragging a track bar control to the right pole (i.e. named strongly disagree) or the left pole (i.e., named strongly agree) of the rating scales using a computer mouse. The scales were divided into 200 subunits, with − 100 being the minimum value and + 100 being the maximum value (i.e., similar to a visual analog scale). Time to complete the ratings was unrestricted. Each participant rated a subset of 20 randomly selected stick-figure animations (i.e., each participant rated her/his own set of stimuli, which were presented in randomized order). All video clips were presented without sound. Analysis ~~~~~~~~ Coordinate data obtained during the behavior encoding was used for analyses of body motion. Previous studies have shown that the horizontal and vertical components of body motion affect impression formation differently (Koppensteiner, 2013; Koppensteiner & Grammer, 2010). For this reason we decomposed the speakers' body movements into horizontal (x) and vertical (y) components by calculating distances between the coordinates of different landmarks. We determined the magnitude of vertical hand movements: the sum of lectern(y) – right hand(y) and lectern(y) – left hand(y), and vertical body movements: throat(y) – lectern(y). In a second step we determined horizontal hand movements: throat(x) – right hand(x) and left hand(x) – throat(x) and horizontal body movements: throat(x) – origin(x). This gave four time series of changing landmark distances from which we extracted the amplitudes between successive local maxima and local minima. The sum of all vertical amplitudes served as an estimate of the overall vertical distances a moving body produced (i.e., overall vertical expansiveness in motion) while the sum of all horizontal amplitudes served as an estimate of the overall horizontal distances a body produced (i.e., overall horizontal expansiveness in motion). The sum of the number of local minima and maxima, on the other hand, served as an estimate of the vertical quantity of motion and as an estimate of the horizontal quantity of motion without including the influence of motion amplitudes (see Figs. 1 and 2). In addition to this we also used single positions of the speakers' bodies in our analyses. The maximal horizontal distance between the speakers' left and right hand served as a measure of maximal horizontal expansiveness. Similarly, the sum of the maximal vertical distances between the speakers' hands and the lectern and between the speakers' throats and the lectern served as measures of vertical expansiveness. In summary, we created three types of measures: (1) a measure that captures both expansiveness and the quantity of motion, (2) a measure that only captures the quantity of motion, and (3) a measure that captures expansiveness without including the quantity of motion. Amplitudes and distance measures were corrected for variation in body height by dividing them by the maximum vertical distance between the lectern and the speakers' foreheads (i.e., stick-figure height). Each stimulus was rated by a range from 18 to 22 participants (M = 20). For the statistical analyses we averaged these ratings. Motion variables tend to be highly interdependent, which affects the interpretation of regression coefficients of multiple regressions. For this reason we used simple bivariate correlations (with bootstrapped confidence intervals) to examine the relationships between the variables. We also divided the sample of stimuli into four categories: male and female speakers belonging to the opposition and male and female speakers belonging to the government. The differences between these groups in trait ratings and motion measures were assessed by calculating means with bootstrapped confidence intervals. As we expected interactions between speaker gender and role in Parliament (opposition or government), we also calculated a series of two-factorial Type III ANOVAs. All statistical analyses were carried out in the program R (R Core Team, 2013).","Correlations between the indices of body motion and expansiveness (Table 1) revealed a wide range of interdependencies. For instance, horizontal and vertical measures of motion were strongly linked. Specifically, speakers that showed a great a deal of expansive movements along the vertical axis (up and down movements) also showed expansive movements (e.g., hands moving from the left to the right) along the horizontal axis. A similar relationship was found between horizontal and vertical quantity of motion. Measures of static expansiveness (i.e., maximal distance between hands and hands held up high) were strongly correlated to measures of expansiveness of motion (i.e., amplitudes). Consequently, speakers who produced expansive body postures also tended to produce expansive movements and speakers who produced expansive movements also showed a high quantity of motion. Because our nonverbal measures showed strong interdependencies (Table 1) we also created a composite measure of body motion by adding up vertical and horizontal amplitudes. Stick-figures rated high on dominance tended to be rated low on trustworthiness (r(58) = −.57), those rated high on trustworthiness were also rated high on competence (r(58) = .76), and no relationship (r(58) = −.11) was found between dominance and competence (see also Koppensteiner et al., 2015). Thus, in part dominance and trustworthiness were mutually exclusive categories. Indices of body motion, expansiveness and our composite measure of motion showed profound relationships with ratings of dominance and trustworthiness. Perceived competence did not yield a meaningful relationship with any of our measures (Table 2). Stick-figures adopting an expansive body posture, displaying expansiveness in overall motion behavior and a great deal of body movements were rated high on dominance. These findings indicate that people not only ascribe dominance to static cues of expansiveness but also to motion cues. Low to moderate interrelations between expansiveness of body postures and the quantity of motion – a motion measure independent of amplitude – further supported such an interpretation (see Table 1). The negative relationship between trait ratings of dominance and trustworthiness was also reflected in correlations between motion measures and trustworthiness. There were negative relationships throughout (Table 2). Overall, this can be condensed into a simple formula, namely that high ratings of trustworthiness were linked to low activity in body motion. In the second step we analyzed whether the stick-figures were judged differently on dominance, trustworthiness, and competence depending on the speakers' gender and their party status (i.e., opposition or government). ANOVAs yielded significant interactions between gender and the politicians' role in parliament (i.e., opposition or government) for perceived dominance and perceived trustworthiness (Table 3). Inspection of the mean ratings for dominance provided detailed insights into the differences between the four groups we investigated (Table 4 and Fig. 3). Perceived dominance was highest for males belonging to the opposition parties. In particular, when compared with the ratings of males and females belonging to the government, this result was impressive (confidence intervals show no overlap). Findings for trustworthiness were similar to findings of dominance. Again, the ratings of males belonging to the opposition stood out (Table 2). They received the lowest ratings on trustworthiness of all groups. The composite measure of overall motion provided no significant interaction between gender and party position (Table 3). Means and their confidence intervals showed that overall motion was highest for males from the opposition (Table 4). This was in line with this group's ratings on dominance. However, the differences in trait ratings were not convincingly reflected in our motion measure. Therefore, people's ratings must be guided by additional nonverbal cues.","To display dominance humans and animals make broad gestures or adopt expansive body postures (e.g., De Waal, 2007; Mehrabian, 1972). We investigated whether people perceive such dominance cues in the body motion of politicians giving speeches. For this purpose we translated short excerpts from different speeches into stick-figure animations and extracted the “maximum expansiveness” (i.e., expansiveness of a single posture), the overall “expansiveness of motion” and “the quantity of motion”. All three measures were strongly interrelated and good predictors of perceived stick-figure dominance. This indicates that speakers who adopt expansive postures also produce such expansive postures frequently. Moreover, expansive postures lead to impressions of dominance and perceiving such expansive postures frequently may even reinforce such impressions. In contrast to dominance, trustworthiness was associated with low expansiveness in motion and a low number of movements. Such a negative relationship between dominance and trustworthiness was also reflected in the trait ratings. Therefore, displays of high dominance have a negative impact on impressions of trustworthiness. Taking a closer look at the results revealed that our “nonverbal measures” were not as strongly related to trustworthiness than to dominance. In other words, people perceived dominance more easily than trustworthiness in the nonverbal cues we extracted. This is not surprising because dominance displays are not subtle in general and intended to attract attention. Stick- figures perceived as trustworthy were also perceived as competent, yet we found no relationship between our “nonverbal measures” and competence. Despite being one of the basic social categories in judgments of politicians (Fiske et al., 2007) competence was not conveyed by the cues we extracted. It is conceivable that competence is associated with motion patterns our simple measures failed to capture. However, it is also possible that people, although able to perceive competence in facial cues, are unable to perceive competence in body motion. First impressions may represent a cognitive adaptation that helps to decide whether to approach or to avoid someone (Oosterhof & Todorov, 2008). The results obtained may be interpreted in a similar way and the ratings may be classified as “positive versus negative”. However, previous studies using stick-figure stimuli show that the body movements of speakers allow more than such simple categorizations (e.g., Koppensteiner et al., 2015; Koppensteiner, 2013). In the present study we only extracted conspicuous dominance cues, which may have a predominant influence on first impressions, but more in depth social evaluations may be based on more complex motion cues. Also, it is not clear whether social categories such as those applied here or the Big Five personality dimensions sufficiently capture the information people perceive in motion cues (see also Thoresen et al., 2012). To clarify this follow-up studies are needed. Politicians change their rhetoric depending on the situation or their role in Parliament (Pancer, Hunsberger, Pratt, Boisvert, & Roth, 1992; Tetlock, 1981). This study is a first step toward investigating whether such flexibility in self-presentation also becomes apparent in body motion. Although the results obtained need to be backed by additional studies, they support such an assumption regarding male speakers in the opposition. Stick-figures representing male politicians in the opposition were perceived as more dominant but less trustworthy than stick-figures representing politicians of the government. However, to provide more clarity future work needs to elaborate on this and test whether former opposition members change their nonverbal performance when they become members of the government. Although a previous study using stick-figure stimuli revealed differences between male and female speakers for perceived extraversion and emotional stability (Koppensteiner & Grammer, 2011), in the present study we only found noteworthy gender differences when including party status. The results obtained suggest that dominance displays, that negatively affect perceptions of trustworthiness, are predominantly shown by male speakers expected to control and criticize the work of those who are in power. Male and female speakers of the government, on the other hand, may intend to communicate integrity and on the level of body motion this might look similar for both genders.","Snap judgments are not only guided by the outward appearance and facial expressions but also by salient and simple motion cues. We show that speakers displaying expansive movements and a great deal of body activity are rated high on dominance and low on trustworthiness. In addition, females and male members of the ruling parties appear to communicate their views in a different way than males from the opposition and such differences appear to be discernable in body motion. This hints that politicians not only express their positions with words but also with their bodies."],["Objectives: To: a) identify motivational profiles for exercise, using Self-Determination Theory as a theoretical framework, among a sample of parents of UK primary school children; b) explore the movement between motivational profiles over a five year period; and c) examine differences across these profiles in terms of gender, physical activity and BMI. Design: Data were from the B-Proact1v cohort. Methods: 2555 parents of British primary school children participated across three phases when the child was aged 5–6, 8–9, and 10–11. Parents completed a multidimensional measure of motivation for exercise and wore an ActiGraph GT3X + accelerometer for five days in each phase. Latent profile and transition analyses were conducted using a three-step approach in MPlus. Results: Six profiles were identified, comprising different combinations of motivation types. Between each timepoint, moving between profiles was more likely than remaining in the same one. People with a more autonomous profile at a previous timepoint were unlikely to move to more controlled or amotivated profiles. At all three timepoints, more autonomous profiles were associated with higher levels of MVPA and lower BMI. Conclusions: The results show that people's motivation for exercise can be described in coherent and consistent profiles which are made up of multiple and simultaneous types of motivation. More autonomous motivation profiles were more enduring over time, indicating that promoting more autonomous motivational profiles may be central to facilitating longer-term physical activity engagement. --------------------------------------------------------------------------------","The purpose of the present exploratory study was to adopt a person-centred approach to: a) identify motivational profiles for exercise amongst adults, using Self-Determination Theory (SDT) as a theoretical framework; b) explore the stability of and movement between motivational profiles over a five-year period; and c) examine differences across these profiles in terms of gender, accelerometer-estimated physical activity and BMI. Design and participants ~~~~~~~~~~~~~~~~~~~~~~~ This study uses data from the longitudinal B-Proact1v cohort study. A more detailed outline of the study can be found elsewhere (Jago et al., 2017, 2019; Jago, Sebire, et al., 2014; Jago, Thompson, et al., 2014). In brief, B-Proact1v aimed to examine physical activity and sedentary behaviour of primary school children and their parents. Data were collected on three occasions, between January 2012 and July 2013 when the child was in Year 1 (ages 5–6), between March 2015 and July 2016 when the same child was in Year 4 (ages 8–9), and between March 2017 and May 2018 when the same child was in Year 6 (ages 10–11). A total of 57 schools participated in the first data collection, and the same schools were invited to take part in subsequent phases, with 47 participating in the second phase and 50 in the third phase. Across the three timepoints, data were collected from 2555 parents/caregivers from 2132 families: 1195 were involved at time 1, 1140 at time 2, and 1233 at time 3. Prior to data collection, the study received ethical approval from the School for Policy Studies Research Ethics Committee at the University of Bristol and written consent was obtained from all participants at each phase of data collection. Exercise motivation Motivation to exercise was measured via the Behavioural Regulation in Exercise Questionnaire-2 (BREQ-2; Markland & Tobin, 2004). Grounded in SDT, the 19-item measure assesses five forms of behavioural regulation; intrinsic (4 items e.g. ‘I exercise because it’s fun’), identified (4 items e.g. ‘It’s important to me to exercise regularly’), introjected (3 items e.g. ‘I feel like a failure when I haven’t exercise in a while’), external (4 items e.g. ‘I exercise because other people say I should’), and amotivation (4 items e.g. ‘I don’t see the point in exercising’). Participants recorded their responses on a 5-point Likert scale ranging from 0 (not true for me) to 4 (very true for me). The subscales demonstrated internal consistency at each time point (Table 1). Moderate-to-vigorous physical activity Parents wore an ActiGraph wGT3X-BT accelerometer on their waist for five days, including two weekend days. Accelerometer data were processed using Kinesoft software (v3.3.75; Kinesoft, Saskatchewan, Canada) using 60-s epochs. A valid day was defined as at least 500 min of data after the exclusion of periods of non-wear time of over 60 min, whilst allowing up to 2 min of interruptions. Analysis was restricted to those parents who provided at least three days of valid data, to ensure reasonable estimates of typical daily activity whilst maximising sample size (Aadland & Ylvisaker, 2015; Tudor-Locke et al., 2005). The average number of MVPA minutes per day were used in the analysis, derived for each participant using population-specific cut points for adults (≥2020 counts per minute; Troiano et al., 2008). Participant characteristics Parents reported their date of birth and gender. BMI was calculated from self- reported height and weight as weight (kg)/height (m2).","First, confirmatory factor analysis using maximum likelihood estimation was conducted to obtain weighted factor-scores for each subscale of the BREQ-2 measure (intrinsic motivation, identified regulation, introjected regulation, extrinsic regulation, and amotivation). Doing so provides subscale estimates that consider the contribution of each item to the latent variable they are measuring. Model fit was assessed using multiple indices as follows; the Chi-square index, comparative fit index (CFI), standardised root mean square residual (SRMR), and root mean square of approximation (RMSEA). The thresholds for good fit used were >0.90 for the CFI, <0.08 for the SRMR, and <0.06 for the RMSEA (Hu & Bentler, 1999). Longitudinal invariance of the measurement model was sequentially tested via a series of increasingly constrained models. Invariance was indicated by a change in CFI of ≤0.01 (Cheung & Rensvold, 2002). The generated factor scores were used as input variables in subsequent analyses, in line with guidance (DiStefano, Zhu, & Mindrila, 2009). We used latent profile analysis (LPA) as the primary data analysis approach to explore and identify motivational profiles for exercise. In LPA, each participant is assumed to belong to one of a set of underlying profiles, and the analysis estimates the probability of membership to each profile for each person. As an extension of LPA, latent transition analysis was used to additionally estimate the probability of moving between profiles at different time points. We used a three-step approach (Asparouhov & Muthen, 2014; Nylund-Gibson, Grimm, Quirk, & Furlong, 2014) to conduct the LPA and transition analyses in MPlus (version 7, Muthen & Muthen). In a fourth step we explored the associations of profile membership with gender, BMI and MVPA. The syntax for all main analyses is available as supplementary material. Step 1 and 2- identification of latent profiles and obtaining classification errors Through step 1 we identified the motivational profiles for each timepoint. A sequence of models, with an increasing number of profiles from 2 to 7, were examined to ascertain whether more complex (i.e. more profiles) or parsimonious (i.e. fewer profiles) models provided the best description of the data. The models were estimated using data from all three timepoints, with each time point assumed to be independent of the others. Based on recommendations (Nylund, Asparoutiov, & Muthen, 2007), and in line with previous papers (Jago et al., 2018; Lindwall et al., 2017), several criteria were used to determine the most appropriate model. Statistically, the log-likelihood, the Bayesian information criterion (BIC) and the sample-adjusted Bayesian information criterion (SSA-BIC) were considered, with lower values indicating better model fit (Henson, Reise, & Kim, 2007; Yang, 2006). Relative entropy and the class membership probabilities were used to identify potentially problematic models in which some classes have small proportions of membership. We also considered the theoretical alignment and interpretation of the final profiles in terms of different levels of behavioural regulation (Vansteenkiste & Mouratidis, 2016) and, with this in mind, sought to choose the most meaningful model with the smallest number of profiles. In order to explore assumptions about variance, we compared the model with no constraints, correlated indicators and equal variances across timepoints (Morin, Meyer, Creusier, & Bietry, 2016). In the second step, to include the measurement error in assignment of individuals to profile, we conducted latent profile analysis separately for each set of latent profile indicators, fixing the measurement parameters so that the profiles were the same as in step 1. This allowed us to obtain profile variables and classification errors. This step was repeated for all three timepoints. Step 3- Transition Between Profiles Across timepoints In the third step, we estimated the movement between motivation profiles across the three timepoints, keeping the latent profiles at each timepoint fixed and accounting for measurement error in profile assignment. Step 4- associations of profile membership with gender, BMI and MVPA We examined the associations between profile membership and gender, BMI, and accelerometer-estimated MVPA, via the Wald test using the BCH method, which includes classification error and is robust to violations of assumptions (Yang, 2006). All analyses accounted for clustering of parents at the family and school levels. Between timepoints, attrition was largely attributed to school drop-out, accounting for the drop out of 244 parents at time 2 and 167 parents at time 3, or to families moving to schools not involved in the project, accounting for a total of 253 parents. Therefore, as most missing data was explained by school-level rather than individual-level factors, all model parameters were calculated using full information maximum likelihood, which uses available information from participants at all time points and handles missing data within the analysis model, under the assumption that data are missing at random. Models were estimated using multiple start values (500 starts and 100 sets) to check convergence. Preliminary analysis ~~~~~~~~~~~~~~~~~~~~ 1023 participants provided valid accelerometer measurements at time 1 (86% of those in the study at time 1), 925 at time 2 (81% of those in the study at time 2), and 891 at time 3 (72% of those in the study at time 3). Table 1 shows descriptive statistics and proportions of missing values for questionnaire measures in these participants. Correlations between variables are presented in supplementary material (Table S1). At each timepoint, most participants were female (72–76%) and mean BMI was between 25 and 26 kg/m2. At time 1, the mean age was 37.8 years increasing to 41.3 at time 2 and 43.0 at time 3. The average daily minutes of MVPA increased across timepoints, from 49.5 min per day at time 1 to 50.4 min per day at time 2 and 51.86 min per day at time 3. At all three timepoints, the sample had similar motivational distributions, with low levels of amotivation and high levels of both identified regulation and intrinsic motivation. Factors scores from BREQ-2 were derived via CFA.1 The model showed acceptable fit to the data at time 1 (χ2 = 574.38, df = 142, p < .0005; CFI = 0.93, RMSEA = 0.05 (90% CI = 0.05, 0.06), SRMR = 0.05), time 2 (χ2 = 496.66, df = 142, p < .0005; CFI = 0.95, RMSEA = 0.05 (90% CI = 0.05, 0.06), SRMR = 0.05), and time 3 (χ2 = 418.44, df = 142, p < .0005; CFI = 0.96, RMSEA = 0.04 (90% CI = 0.04, 0.05), SRMR = 0.04). Factor determinacy scores ranged from 0.86 to 0.97 and were deemed to provide a good estimate of the true factor score (Table S2). The results of the invariance testing provided evidence for the equivalence of the measurement model across timepoints (Table S3). Prior to running the main analyses, behavioural regulation variables were checked for univariate and multivariate outliers, and no meaningful outliers were detected (Tabachnick & Fidell, 2013). Step 1 and 2- identification of latent profiles and obtaining classification errors Table 2 shows the indicators of model fit for latent profile models containing 2–7 profiles. The log-likelihood decreased as the number of profiles increased and the 6-profile model had the lowest BIC indicating a better fit than the models containing fewer profiles. Whilst the 7-profile solution had a lower SSA-BIC, at each timepoint at least one profile had a very low probability of membership (<2%). The 6-profile solution had the next lowest SAA-BIC and all profiles had reasonable membership probabilities (over 4%). Compared to the 7-profile solution, the 6-profiles were more meaningful, with combinations of different types of behavioural regulation representing logical and theoretically-appropriate profiles. Additionally, we ran the model with alternative specifications (1000 starts and 200 sets) and the log-likelihood was replicated. We therefore chose to proceed with 6 profiles as the most appropriate model, based on a combination of log-likelihood, BIC, SSA-BIC and model interpretability. With regards to measurement invariance (Morin et al., 2016), models with different constraints produced similar results (Tables S5 and S6), and so, for parsimony and given that similar patterns of motivation exist across populations and contexts (Milyavskaya & Koestner, 2011), we proceeded with the model that assumed that the variances for each motivation variable differed between profiles but were equal across timepoints. Details of the six profiles are reported in Table 3, the profiles are represented graphically in Figure 1, and the probability of membership to each profile is presented in Figure 2. To ensure that interpretation is theoretically meaningful and appropriate, we have re-ordered the profiles to match the motivational continuum proposed in SDT (i.e. from the least to the most self- determined). The six profiles were labelled as: Strongly amotivated- Primarily amotivation with some external regulation Amotivated- Moderate levels of amotivation with low levels of all other types of regulation Controlled and amotivated- High levels of both introjected and external regulation accompanied by high levels of amotivation Low in motivation-low levels of all types of behavioural regulation Autonomously motivated and introjected- Predominantly introjected regulation alongside some intrinsic and identified regulation Autonomously motivated- Primarily intrinsic motivation accompanied with identified regulation. At all three timepoints, a greater proportion of participants were likely to be assigned to Profile 3 (low in motivation) or Profile 6 (autonomously motivated). Participants were consistently least likely to belong to Profile 1 (strongly amotivated; see Figure 2). Comparison of profiles at time 1, 2 and 3 Figure 2 illustrates the probabilities of profile membership at each timepoint. Proportions of participants in each profile were similar over time, with Profile 6 (autonomously motivated) and 4 (low in motivation) being the largest. At each time point, Profile 1 (strongly amotivated) had the lowest membership. The proportion of participants in Profile 5 (autonomously motivated with introjected) increased across timepoints, particularly across time 2 and 3. Step 3- Transition Between Profiles Across timepoints The likely patterns of movement between profiles across the three timepoints are shown in Figure 3. Between time 1 and 2, a large proportion of participants were likely to remain in the same profile (47%). Participants who transitioned between profiles (53%) were likely to move to more motivationally-positive profiles, with participants having the highest probability of moving to profiles that were more autonomous. The exception to this was those in Profile 4 (low in motivation) being equally likely to move to Profiles 5 and 6 (both characterised by strongly autonomous regulations) or Profile 3 (controlled and amotivation). If those in Profile 6 (autonomously motivated) at time 1 were to move (19%) they were most likely to move to Profile 4 (low in motivation). Between time 2 and 3, a similar proportion of participants were likely to remain in the same profile (45%). The probability of movement between profiles across time 2 and 3 was more varied, with a greater likelihood of moving to a wider range of profiles. In line with the stability found across time 1 and 2, those in Profiles 5 and 6 (both with high levels of autonomous motivation) and Profile 1 (highly amotivated) at time 2 were most likely to remain in the same profile at time 3. Those participants who were likely to move from Profiles 5 and 6 had the highest probability of moving to Profile 4 (low in motivation). Those who moved from profile 1 (high amotivation) were most likely to move to Profile 2 (amotivated), with very few moving to more self-determined profiles. For Profiles 3 (controlled and amotivated) and 4 (low in motivation) movement was more diverse, with a relatively balanced probability of moving to profiles characterised by more autonomous motivation or those characterised by more amotivation. Step 4- associations of profile membership with gender, BMI and MVPA Table 4 shows the associations of profile membership with co-variates at each timepoint. Across all timepoints, the proportion of female membership in a profile ranged from 66% to 87%. There were no consistent patterns in profile membership across each timepoint in terms of gender, but at both time 1 and time 2, Profiles 1 and 2 (strongly amotivated and amotivated respectively) had the highest proportions of female participants. At time 3, Profile 3 (controlled and amotivated) had the highest proportion of female participants. BMI ranged from 24 to 28 across profiles and timepoints. Those in Profile 6 (autonomously motivated) had the lowest mean BMI at each timepoint and, generally, those in Profiles 1 and 3 (strongly amotivated and controlled and amotivated) had the highest BMI. There was less consistency in the association between profile and MVPA across the three timepoints. However, at each timepoint, there was a 15–17 min variation in MVPA across profiles. Profile 6 (autonomously motivated) was consistently associated with higher MVPA, and Profile 1 (Strongly amotivated) consistently associated with the lowest MVPA. At each timepoint, participants in Profile 5 (autonomously motivated and introjected) engaged in less MVPA than participants in Profile 6. For all timepoints, profiles with higher MVPA also had a lower BMI indicating an association between these variables.","The evidence presented in this paper indicates that people have multiple simultaneous motivations for engaging in physical activity, providing further support for the complex multi-dimensional nature of physical activity behaviour. We identified six distinct motivational profiles that represented different combinations of motivation types spread across the continuum of motivation proposed within SDT. Further, whilst exploratory, this paper provides the first evidence for the movement of people between profiles over time, and the findings suggest that motivation is dynamic, with most participants moving between profiles across timepoints. Profiles consisting of strong endorsement of more autonomous forms of motivation were the most stable. The six profiles identified were qualitatively different and distinct. Two profiles were characterised predominantly by a lack of motivation to exercise (amotivation), with participants in Profiles 1 and 2 having high and moderate levels of amotivation respectively. Profile 3 (controlled and amotivated) consisted of high levels of external regulation alongside amotivation and introjected regulation, indicating that a lack of interest in exercise, pressure from others and guilt and shame about not engaging in exercise do occur simultaneously. Participants in Profiles 5 and 6 reported moderate to high levels of both intrinsic and identified regulation, characterised by enjoyment and personal value of exercise, but Profile 5 had additional high levels of introjection. The structure of the profiles followed a logical progression along the theoretical motivation continuum, with combinations of closely-related regulation types indicating that similar types of motivation are more strongly correlated than disparate regulation types (Ryan & Deci, 2000). However, one profile did not align with the theoretical propositions within SDT, and that is Profile 4 (low in motivation), characterised by no distinct regulation type. A similar profile has been seen in previous papers (Lindwall et al., 2017) despite the theoretical expectation that low levels of both controlled and autonomous regulation types would be complimented with high levels of amotivation. This profile is therefore difficult to explain, and may be the result of response category artefact, where this profile comprises individuals who provided mid- range responses on the Likert scale (Nadler, Weston, & Voyles, 2015). Alternatively, given that the BREQ-2 questions refer specifically to exercise, this profile may represent a group who did not find the questions relevant due to not engaging in exercise (known as ‘exercise aschematic’; Kendzierski, 1990). In this sample, this profile represents a substantial group of participants (up to 30%) and so qualitative interviews with individuals likely to be classified to this profile could help to provide clarity on the source of this profile. Several of the theoretically-meaningful profiles identified in this paper align with those found in previous profile analyses (Bechter et al., 2018; Lindwall et al., 2017). Specifically, Profile 6 (autonomously motivated), Profile 5 (autonomously motivated and introjection), Profile 1 (strongly amotivated), and Profile 4 (low in motivation) replicate profiles previously identified in both active and non-active adult populations (Lindwall et al., 2017). There have also been similar profiles identified in research with young people (Bechter et al., 2018). Collectively, this provides further confirmation of the validity of the profiles and indicates that similar combinations of behavioural regulations are observed across different samples, countries, and age-groups, providing further support for the universal nature of motivation as conceptualised in SDT. The patterns of MVPA for each profile provide support for the construct validity of the profiles and the continuum of motivation proposed within SDT (Ryan & Deci, 2017). At all three timepoints, individuals most likely to be in Profile 6 (autonomously motivated) engaged in more MVPA than those in any other profile, with those in Profile 1 (strongly amotivated) consistently engaging in the least. Further, the difference in MVPA between the motivation profiles provides further support for person- centred analysis as a method to provide additional insight into the role of each regulation type and the way in which they may combine to influence behaviour. In particular, the profiles identified through this analysis indicate that introjection may not occur in high levels in isolation, but rather in two qualitatively different profiles, combined either with external regulation (as in Profile 3) or with identified and intrinsic regulation (as in Profile 5). Most of the SDT literature has found little association between introjected regulation and physical activity (Duncan et al., 2010; Standage et al., 2008), but the different combinations of behavioural regulation may mean that traditional analysis methods have masked the differential influence that introjected regulation can have, depending on which other behavioural regulations it occurs alongside. Further, a similar profile characterised by high self-determined and introjected regulation was found in previous studies associating motivation profiles with self- reported exercise behaviour (Lindwall et al., 2017), but this evidence indicated that introjected regulation did not have a detrimental influence on physical activity when found in combination with autonomous motivation. In contrast to this, our findings suggest that the presence of introjected regulation alongside autonomous motivation is associated with lower levels of MVPA compared to when autonomous motivation is experienced alone. Collectively, the evidence suggests that the addition of introjected regulation to an otherwise autonomously motivated person will, at best, have no impact or, at worst, undermine behaviour. Further work is needed to clarify the role of introjected regulation in determining physical activity behaviour. The data presented here show that the autonomously motivated profile was associated with a lower BMI at each timepoint. This may be linked to the association between more autonomous motivation and higher levels of MVPA but may also represent a wider association between autonomous motivation for exercise and other weight control behaviours such as diet, sometimes referred to as ‘motivational spill-over’ (Mata et al., 2009). Additionally, at all three timepoints, participants in Profile 1 (strongly amotivated) had a high average BMI. This is consistent with previous research with adolescents and adults showing that individuals with a higher BMI are more likely to report high levels of amotivation and individuals who are have a healthy BMI are more likely to report intrinsic regulation for exercise (Ersoz, Altiparmak, & Asci, 2016; Hwang & Kim, 2013). The data therefore highlight a need for further examination of the associations between motivation, physical activity and body weight. In the current sample, across all timepoints individuals were more likely to move between profiles than to remain in the same one, indicating that motivation for exercise is relatively dynamic. More autonomous profiles were the most stable across timepoints and participants likely to belong to profiles characterised by strong controlled motivation or amotivation were most likely to move to other profiles. This movement was most commonly to more autonomous profiles, which is consistent with the principle of SDT that humans have an innate desire to seek out situations and environments that are psychologically fulfilling and, over time and given satisfaction of autonomy, competence and relatedness needs, will become more self-determined in their motivation (Deci & Ryan, 2000). The findings also suggest that individuals with low levels of motivation may be motivationally vulnerable, in that over time they have a similar probability of moving to more autonomous or more amotivated profiles. This presents a potentially important opportunity to intervene to attempt to inspire people towards more stable autonomous forms of motivation. From a public health perspective, these findings suggest that strategies to promote greater physical activity engagement should seek to foster more stable autonomous motivation by developing physical activity environments that support, rather than thwart, the basic psychology needs of autonomy, competence and relatedness (Ryan & Deci, 2017). The stability of the profiles identified in this paper indicate that doing so could have longer-term benefits for physical activity behaviour. Environments that provide choice about when and how to be active, opportunities for success, skill building, and an optimal level of challenge, a personally-relevant rationale for being active, and the opportunity to develop strong connections with others have been shown to promote greater physical activity engagement in the long-term (Samdal et al., 2017). There are several opportunities for future research with a clear need for more person-centred analyses of exercise motivation. As this is the first exploration of the transition between profiles, further research is needed to identify whether the stability of and movement between profiles is consistent across samples. Additionally, further work is needed to ascertain the reasons why individuals may move between motivational profiles and so the longitudinal assessment of wider theoretical constructs, such as need satisfaction and well-being, is needed. Qualitative work allow the origins of the low motivation profile to be explored, specifically focusing on whether this profile represents a genuine group of individuals, is the result of the measure used, or if it is an amalgamation of individuals with different motivational profiles. Strengths and limitations ~~~~~~~~~~~~~~~~~~~~~~~~~ The exploratory person-centred analysis, longitudinal data and objective estimates of physical activity are particular strengths of this study, allowing the investigation of the interplay between different types of motivation and associations with physical activity, as well as exploring the stability of motivation over time. However, it is important to highlight the limitations of this work. First, the sample was largely female and therefore the profiles may be more representative of those found amongst women rather than men. Additionally, the sample were all parents of primary-school aged children and were mostly mothers, therefore may not represent the wider adult population. However, given that previous studies have found no gender differences in behavioural regulations (Guerin et al., 2012), and given the universality assumption of SDT, we would not anticipate major differences in the profiles we identified in a more balanced sample. There may, however, be differences in the proportions of parents belonging to particular profiles when compared to the wider population (McIntyre & Rhodes, 2009; Solomon-Moore et al., 2016). Additionally, our sample were high in autonomous forms of motivation, low in amotivation and generally active, which limits the generalisability of the profiles to groups with lower autonomous motivation and activity levels. The nature of the study in which this analysis was nested is likely to have influenced this sample, attracting parents who enjoy and value being active themselves, and so future work should aim to recruit a more representative and varied sample. Further, whilst we assumed that data were missing at random, it is possible that parents who provided full data may have higher quality motivation than those with missing data. It is also to be noted that BMI was based on self-reported indicators and therefore may not be accurate. The consistent associations between BMI and motivation profiles across the three timepoints provide strong support for this relationship, but future research should adopt objective measurements of height and weight to provide additional clarity. Additionally, whilst we did not have sufficiently powered sample to do so, future work should seek to explore the associations between profile transition and MVPA and BMI and additional covariates, such as need satisfaction and need frustration.","This paper provides evidence for the experience of multiple simultaneous reasons for engaging in exercise and that more autonomous motivation profiles are associated with higher levels of accelerometer assessed MVPA and lower BMI. The latent transition analysis provides the first evidence that profiles characterised by autonomous forms of motivation are more stable over time than less self-determined profiles. This indicates that once individuals establish and personal value of enjoyment of exercise, this persists over time, and so promoting more autonomous motivational profiles may be central to facilitating long-term physical activity engagement."],["Musical hallucinations (MH) account for a significant proportion of auditory hallucinations, but there is a relative lack of research into their phenomenology. In contrast, much research has focused on other forms of internally generated musical experience, such as earworms (involuntary and repetitive inner music), showing that they can vary in perceived control, repetitiveness, and in their effect on mood. We conducted a large online survey (N = 270), including 44 participants with MH, asking participants to rate imagery, earworms, or MH on several variables. MH were reported as occurring less frequently, with less controllability, less lyrical content, and lower familiarity, than other forms of inner music. MH were also less likely to be reported by participants with higher levels of musical expertise. The findings are outlined in relation to other forms of hallucinatory experience and inner music, and their implications for psychological models of hallucinations discussed. --------------------------------------------------------------------------------","Auditory hallucinations (AH) are defined as the conscious experience of sounds that occur in the absence of any actual sensory input. Although the most frequently reported form of AH are auditory verbal hallucinations (AVH), phenomenological surveys have also shown that a substantial minority of people also report musical hallucinations (MH): that is, the perception of music when none is playing. For example, one survey of 100 people with psychosis and AVH found that 36% also described the occurrence of MH (Nayani & David, 1996). The most frequent reports were of hearing choral music, with orchestral music and pop music also evidenced, although specific frequencies were not provided. More recently, McCarthy-Jones et al. (2012) analyzed data from a semi-structured interview with 199 psychotic patients who reported AVH, finding that a smaller proportion (compared to Nayani and David) of approximately 15% also experienced MH. The two largest phenomenological surveys of AVH, then, suggest that MH occur in a substantial minority of people who hear voices; however, since both primarily focus on AVH, few details of MH are described beyond prevalence. Whilst questionnaire measures used to assess proneness to hallucinations (e.g., Launay-Slade Hallucination Scale; Morrison, Wells, & Nothard, 2000) in the general population do include items relating to non-verbal hallucinations, responses to individual items are rarely reported; thus, we know little about either the prevalence or phenomenology of MH in clinical or non-clinical samples. Indeed, non-verbal hallucinations have been somewhat neglected in the psychological literature, with only a small number of studies investigating risk factors and basic phenomenological features of MH. Surveys focusing exclusively on MH have suggested that they may occur in around 16% of individuals with a diagnosis of schizophrenia (Saba & Keshavan, 1997), and as many as 41% of individuals with obsessive compulsive disorder (Hermesh et al., 2004). Other risk factors include hearing impairments, old age, and social isolation, although these may not be independent factors (Evers & Ellger, 2004). Few surveys have specifically investigated the phenomenology of MH, further than reporting the most frequent styles of music. Saba and Keshavan did report on several details of MH in a small sample of individuals with a diagnosis of schizophrenia, showing that the majority included both instrumental and lyrical elements, which tended to be familiar to the individual. Patients tended to appraise the MH fairly positively, with the most frequent description of the experience being ‘soothing’ (62%). Many experiences of MH were described as perceived as emanating from the external environment, and approximately half were described as outside of volitional control, which Saba and Keshavan argue should be considered a key feature of MH. Whilst this study provided important information on the experience of MH in schizophrenia, the sample size (16 participants reporting MH) was low, and the questions on phenomenology relatively limited. Golden and Josephs (2015) recently reviewed medical records of individuals, including 393 cases of MH, grouping the data into five categories: MH associated with neurological disorder, psychiatric disorder, structural brain damage, drug toxicity, and those not otherwise classifiable. The study mainly reports on brain regions associated with MH, but does note that many individuals with psychiatric disorders found that the experiences were ‘mood-congruent’ (e.g., sad music when they were feeling depressed). Indeed, within psychiatric patients reporting MH, depression seems to be the most common diagnosis (69%), along with hearing loss or tinnitus (Golden & Josephs, 2015; Rocha et al., 2015; Teunisse & Olde-Rikkert, 2012). A case series presented by Warner and Aziz (2005) of patients referred to old-age psychiatric services, though, only found a rate of hearing loss of 33% in patients with MH – perhaps surprisingly low given a mean age of 78 years. They also note that many patients were not distressed by the MH, and so speculate that the prevalence of such phenomena may be higher than previously thought if individuals do not seek medical attention. Due to the nature of these studies, however, no participants from non-clinical populations were included. Other studies have also used stringent inclusion criteria: for example, Evers and Ellger, in a review of the etiology of MH, deliberately excluded musical ‘pseudohallucinations’ (those experienced as internal to the individual). The distinction between ‘true’ hallucinations and pseudohallucinations is no longer thought to be clinically significant (Copolov, Trauer, & Mackinnon, 2004), and research into AVHs typically includes both internally and externally located perceptions (Nayani & David, 1996). It is unclear to what extent MH are experienced as internal or external, but it is possible that previous research has omitted a significant number of cases by using overly strict inclusion criteria. Furthermore, it is unclear to what extent the attributes assessed in the small amount of previous research could also be applied to other forms of ‘inner music’1 (Fernyhough, 2016, p. 238). Musical imagery, for example, is the generation of music in one’s own head, not necessarily instigated by any external percept. It is frequently reported by many individuals in the general population (Bailes, 2007; Williamson et al., 2012), and often used by musicians to rehearse or aid reproduction of music, in the form of notational audiation (Brodsky, Henik, Rubinstein, & Zorman, 2003). Musical imagery can also occur involuntarily (INMI) with little or no volitional control. One form of INMI, ‘earworms’ (also referred to as ‘sticky tunes’ or ‘stuck songs’), are typically defined by their repetitiveness and persistence (although there is some debate in the literature regarding how to precisely define the experience – see below). Previous research has indicated that the frequency of earworms is affected by exposure to, and rehearsal of, music (Liikkanen, 2012), and as such is elevated in musically trained individuals (Beaty et al., 2013; Floridou, Williamson, Stewart, & Müllensiefen, 2015), with one experience sampling study in musicians finding musical imagery occurring in as many as 32% of randomly sampled episodes, with 58% of these samples noted as being due to having recently heard or rehearsed music (2007; Bailes, 2006). Earworms tend not to be associated with negative emotions, unless the reported duration is particularly lengthy (presumably due to unwanted persistence) (Floridou et al., 2015). To our knowledge, no research has directly compared self-reported experiences of musical imagery and earworms to MH, and, in fact, the boundary between earworms and MH is somewhat unclear in much of the literature. For example, Hemming (cited in Williams, 2015) defines MH as INMI that reaches a pathological level (presumably reflected in distress experienced by the individual), implying that MH are simply a more extreme, persistent, or distressing version of earworms. In support of this, the aforementioned study by Saba and Keshavan (1997) distinguished MH from musical imagery purely in terms of volitional control. In contrast, Williams argues that whilst both MH and earworms are involuntary, only MH are experienced as located in the external environment. However, as discussed above, other forms of auditory hallucination, for example AVH, are often experienced as internally located (Daalman et al., 2011; Nayani & David, 1996), yet are typically still classified as hallucinatory experiences. An open question, then, is the extent to which earworms and MH share phenomenological attributes (e.g., control, perceived location), and whether MH can be distinguished on other aspects of musical experience (e.g., type of music, frequency, duration, familiarity, level of acoustic detail). Research into musical imagery and earworms has also investigated their effect on mood and behavior, but again, these have not been directly investigated in comparison to MH. For example, Williamson, Liikkanen, Jakubowski, and Stewart (2014) showed that 74.6% of individuals reported humming or singing along in response to earworms, whilst only 10.9% reported attempting to suppress them.","in other studies have also reported that bodily movements in response to earworms (e.g., tapping a foot to the beat) are relatively common (Floridou et al., 2015). Finally, frequency of earworms was associated with self-reported obsessive-compulsive traits, perhaps similarly to the persistence of intrusive thoughts in obsessive-compulsive disorder (Beaman & Williams, 2010). In contrast, little is known about typical affective and behavioral responses to MH. There are, then, several large gaps in the MH literature, which we sought to address in the present study. Firstly, very little research has been conducted regarding the phenomenology of MH much further than asking about the broad style of music experienced. Secondly, there is some confusion within the literature over precise definitions as to what constitutes an MH, leading previous studies to use different inclusion criteria. As mentioned, Saba and Keshavan (1997) used ‘volitional control’ as a key indicator of MH; yet, this alone fails to distinguish the experience from that of earworms. On the other hand, some authors seem to have equated earworms and MH (Hermesh et al., 2004), whilst others have equated earworms and INMI as referring to the same phenomenon (Farrugia, Jakubowski, Cusack, & Stewart, 2015). Williams (2015) has argued that INMI should be used as a broader term, defining any type of musical imagery outside of conscious control, with the term ‘earworm’ being restricted to a type of INMI characterized by its repetitiveness. Within Williams' framework, MH would be categorized as a form of INMI defined by their perceived externality and pathology (e.g. hearing impairment and/or brain damage). Williams, then, offers perhaps the most rigorous categorization of forms of inner music, but, due to a lack of previous research, does not discuss other potential differences between MH and earworms. We sought to conduct an exploratory survey of the phenomenology of MH, to investigate potential similarities and differences between MH and other forms of inner music. Participants were asked to pick a category that they felt best described their inner musical experience (musical imagery, earworm, musical hallucination) based on basic definitions (see Supplementary Materials), or specify that they regularly experienced multiple different types of inner music (henceforth referred to as the ‘mixed experiences’ group). These categories were then compared on a number of phenomenological attributes, both based on those reported in previous research, and areas that have not previously been investigated. Based on previous literature, it was expected that MH would be more likely to be experienced as coming from the external environment, whereas earworms would be characterized by a lack of volition and repetitiveness, compared to musical imagery. As well as collecting demographic information, we also asked about the presence of psychiatric diagnoses and hearing impairments, as well as prior musical experience, given that previous literature suggests these may be key predictors of the presence of MH. Furthermore, we asked about a number of other features of the experience (frequency, duration, familiarity, feelings of anticipation, triggers), musical details (perception of lyrics, instruments, intensity, harmony, melody) and effects on behavior (effects on the body, effects on mood, effects on relationships with others). Participants were also given the chance to provide further information about their experiences in free text boxes on many questions. The aim was, therefore, to provide a more detailed and nuanced study of phenomenological aspects of MH as compared to other types of inner music, than has previously been conducted. Supplementary data associated with this article can be found, in the online version, at https://doi.org/10.1016/j.concog.2018.07.009. We sought to conduct an exploratory survey of the phenomenology of MH, to investigate potential similarities and differences between MH and other forms of inner music. Participants were asked to pick a category that they felt best described their inner musical experience (musical imagery, earworm, musical hallucination) based on basic definitions (see Supplementary Materials), or specify that they regularly experienced multiple different types of inner music (henceforth referred to as the ‘mixed experiences’ group). These categories were then compared on a number of phenomenological attributes, both based on those reported in previous research, and areas that have not previously been investigated. Based on previous literature, it was expected that MH would be more likely to be experienced as coming from the external environment, whereas earworms would be characterized by a lack of volition and repetitiveness, compared to musical imagery. As well as collecting demographic information, we also asked about the presence of psychiatric diagnoses and hearing impairments, as well as prior musical experience, given that previous literature suggests these may be key predictors of the presence of MH. Furthermore, we asked about a number of other features of the experience (frequency, duration, familiarity, feelings of anticipation, triggers), musical details (perception of lyrics, instruments, intensity, harmony, melody) and effects on behavior (effects on the body, effects on mood, effects on relationships with others). Participants were also given the chance to provide further information about their experiences in free text boxes on many questions. The aim was, therefore, to provide a more detailed and nuanced study of phenomenological aspects of MH as compared to other types of inner music, than has previously been conducted. Participants ~~~~~~~~~~~~ Participants were invited to take part in an on-line survey, advertised via social media and a research project website (http://hearingthevoice.org/2014/11/04/round-and-round-the- phenomenology-of-inner-music). Rather than aiming to recruit a sample representative of the national population, the aim was to recruit participants who reported hallucinatory experiences or regular inner music, to investigate phenomenological features of these experiences. There were 276 respondents to the questionnaire. From these, 7 participants were excluded who did not respond to a sufficient number of questions (<10% response rate) on the ‘Phenomenology of Inner Music Questionnaire’ (see below), whilst 14 that described an alternative form of inner music in the free text box (e.g., musical memory) were excluded. Thus, the sample analyzed consisted of 255 participants (105 male, 144 female, 6 other), with a mean age of 39.4 (SD = 13.3, range = 18–74). The majority of participants were from English-speaking countries (e.g., 42.4% USA, 34.1% UK) but there were also respondents from other European countries (e.g., Denmark, Germany, France, Finland) and non-European countries around the world (e.g., Israel, Mexico, South Korea). Some form of hearing loss was reported by 15.8% of participants, with most of these being described simply as hearing loss/hypacusis (80%) and/or tinnitus (37.5%). Mean reported length of hearing impairment in these individuals was 18.3 years (SD = 17.3). 42.7% of participants reported having previously received a psychiatric diagnosis, with the largest proportion of these being for depression (60.6%) or anxiety (30.3%), with smaller numbers of participants having diagnoses of bipolar disorder, obsessive compulsive disorder, attention deficit hyperactivity disorder, or schizophrenia/schizoaffective disorder/psychosis. 9.0% of participants reported some form of neurological disorder, whilst 26.4% reported being on some form of medication for a psychiatric/neurological disorder. See Table 1 for demographic information of the sample. Previous musical experience Preliminary questions asked about musical expertise, musical preference, practising and listening habits. See Appendix 1 for a full list of questions. Preliminary questions asked about musical expertise, musical preference, practising and listening habits. See Appendix 1 for a full list of questions. Phenomenology of Inner Music Questionnaire The Phenomenology of Inner Music Questionnaire was designed as a preliminary exploration of the experience of MH, musical imagery, and earworms, including items based on previous literature on MH, but also items that have only been used in relation to imagery and earworms. Firstly, based on short definitions, participants were asked to classify their inner music as either musical imagery, earworm, or musical hallucination. A free text box was also provided for participants to expand on this description. Based on previous research into inner music or MH, further questions asked about aspects of the experience encompassing frequency, duration, level of detail (melody, harmony, intensity, presence of instruments, presence of lyrics), familiarity, effects on mood and behavior, amount of perceived control, triggers, and likelihood of being mistaken for an external stimulus. All questions required the participant to respond on a Likert scale, although many questions also provided a free text box to elicit more detailed responses. (See Appendix 1 for full questionnaire.) The Phenomenology of Inner Music Questionnaire was designed as a preliminary exploration of the experience of MH, musical imagery, and earworms, including items based on previous literature on MH, but also items that have only been used in relation to imagery and earworms. Firstly, based on short definitions, participants were asked to classify their inner music as either musical imagery, earworm, or musical hallucination. A free text box was also provided for participants to expand on this description. Based on previous research into inner music or MH, further questions asked about aspects of the experience encompassing frequency, duration, level of detail (melody, harmony, intensity, presence of instruments, presence of lyrics), familiarity, effects on mood and behavior, amount of perceived control, triggers, and likelihood of being mistaken for an external stimulus. All questions required the participant to respond on a Likert scale, although many questions also provided a free text box to elicit more detailed responses. (See Appendix 1 for full questionnaire.) White Bear Suppression Inventory (WBSI) The WBSI is a 15-item scale designed to measure the tendency to suppress unwanted thoughts. Various studies have implied different factor structures underlying the WBSI; here, we used the subscales identified and used by Muris, Merckelbach, & Horselenberg (1996), assessing tendency to have intrusive thoughts (e.g., I have thoughts that I cannot stop) and thought suppression (e.g. I always try to put problems out of my mind). For each question, participants are required to rate their agreement on a Likert scale from 1 (Strongly Disagree) to 5 (Strongly Agree). These subscales have previously been shown to have acceptable internal reliability (Jones & Fernyhough, 2006). Revised Launay-Slade Hallucination Scale (LSHS-R, auditory items) Tendency to experience auditory hallucinations was assessed using the 5-item LSHS-R (McCarthy-Jones & Fernyhough, 2011; revised from Morrison et al., 2000) (e.g., I hear people call my name and find that nobody has done so). For each question, participants are required to indicate agreement with each question on a 4-point Likert scale, ranging from 1 (Never) to 4 (Almost Always). Total score can range from 5 to 20. It has previously shown high internal reliability (Cronbach’s α = .73) (McCarthy-Jones & Fernyhough, 2011). Varieties of Inner Speech Questionnaire (VISQ) The VISQ is an 18-item scale designed to assess phenomenological features of inner speech. It consists of four subscales: evaluative inner speech (e.g., I think in inner speech about what I have done, and whether it was right or not), dialogic inner speech (e.g., I talk back and forward to myself in my mind about things), other people in inner speech (e.g., I experience the voices of other people asking me questions in my head) and condensed inner speech (e.g., I think to myself in brief phrases and single words, rather than full sentences). Each item is scored on a Likert scale from 1 (Never) to 6 (All of the time). Each subscale has previously shown high internal reliability (Cronbach’s α > .8) and acceptable test-retest reliability (>.6) (McCarthy-Jones & Fernyhough, 2011).","Given that our main area of interest regarded the phenomenological differences between MH and other forms of inner music (musical imagery, earworms), participants were categorized by the main type of experience they reported. If participants indicated in the free text box that they frequently experienced more than one form of inner music (for example, reporting both frequent MH and earworms), but indicated in the free text box that one of these was much more prevalent than the other, they were categorized according to their most prevalent experience. If participants indicated that they experienced more than one type of music, but did not report relative frequencies, they were categorized in a separate ‘mixed experiences’ group. This categorization was performed separately by two of the authors (PM, BA, k = .77), and any disagreements (n = 9) were resolved by discussion between authors. Thus, the sample was split into four groups (musical imagery, earworms, MH, mixed experiences) for between-group analysis. We used chi-square analysis to investigate associations between type of musical experience and presence of a psychiatric/neurological diagnosis, hearing impairment, and level of musical expertise, and further explored significant results using odds ratios (OR) in conjunction with 95% confidence intervals. Due to mainly ordinal and non-normally distributed data, non- parametric ANOVAs (Kruskal-Wallis) were used to test for differences between the categories of inner music, for each phenomenological attribute. Where appropriate, Mann- Whitney U tests were used to investigate differences between MH and other individual categories; note that post-hoc tests were only conducted between MH and other categories (as opposed to between all different categories) to limit the number of tests performed. Based on our areas of interest, we split the analysis into four main sections: (1) demographic and etiological information; (2) basic characteristics and acoustic details of inner music; (3) location and controllability of inner music; (4) effect of inner music on behavior and mood. Qualitative examples given by participants are included throughout as illustrative examples of different aspects of their phenomenology. (All qualitative examples given are taken from participants in the MH group.) Bonferroni corrections for multiple comparisons were applied within each section (e.g., in Section 3.2, 10 tests are performed, so the alpha level is corrected to .05/10 = .005; in Section 3.3, 5 tests are performed so the alpha level is corrected to .05/5 = .01; in Section 3.4, 11 tests are performed, so the alpha level is corrected to .05/11 = .0045). Missed items in the VISQ, LSHS-R or WBSI were replaced with the mean from other items in the same (sub)scale. Demographics and categories of inner music ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The first question of the Inner Music Questionnaire asked people to choose the category that best described the music they heard. Of the 255 participants, 17.3% reported musical hallucinations, 40.4% reported earworms, 27.1% reported musical imagery, whilst 15.3% were categorized in the ‘mixed experiences’ group. To investigate phenomenological differences between MH and other musical experiences, ‘inner music category’ was used as a between- subject variable. Only 1.2% of participants reported that they could ‘Never’ accurately describe their inner music (in response to Q5) but, since they continued to provide responses to the questions, these participants were not excluded from the sample. Table 1 summarizes basic demographics of the sample, whilst Table 2 shows these basic demographics broken down by the category of inner music reported by the participant. For the main effect of inner music type in this section, a Bonferroni-corrected alpha level of .05/9 = .006 was used. A chi-square analysis indicated that there was not a significant association between presence of a psychiatric diagnosis (χ2(4) = 4.33, p = .228, φ = .131) or neurological diagnosis (χ2(3) = 1.36, p = .714, φ = .073) and the category of inner music. There was also no significant association between presence of a hearing impairment and category of inner music (χ2(3) = 0.89, p = .828, φ = .059). There was a significant association between level of musical expertise and type of inner music reported (χ2(15) = 36.93, p = .001, φ = .381), although in this analysis 25% of cells had an expected count of < 5, due to the low number of semi-professional (n = 23) or professional (n = 9) musicians (at least compared to other groups) in the sample. This violates a key assumption of chi-square analysis (Howell, 2010) and, as such, the sample was collapsed into two groups: non-musicians (non-musicians and music-loving non-musicians; n = 110) and musicians (amateur, serious amateur, semi-professional, and professional musicians; n = 145). Again, a chi-square analysis indicated an association between musical expertise and inner music type (χ2(3) = 20.99, p < .001, φ = .287). Further analysis suggested that musicians were less likely to report MH (OR = 0.27, 95% CI [0.14–0.54]), but more likely to report mixed experiences (OR = 3.23, 95% CI [1.42, 7.33]) compared to non-musicians. There was little difference in the proportion of musicians in either the imagery (OR = 1.37, 95% CI [0.78, 2.41]) or earworm (OR = 0.78, 95% CI [0.47, 1.28]) groups. Non- parametric ANOVAs (Kruskal-Wallis) with inner music category as the independent variable suggested a similar pattern of results for number of hours spent practicing musical instruments per week (χ2(3) = 17.82, p < .001), with MH being associated with less music practice than imagery (U = 1084.5, p = .011), although not significantly different to earworms (U = 2168.5, p = .806) or mixed experiences (U = 623.5, p = .040) at the corrected alpha level. However, this pattern of results seemed to be specific to actually practising music, and did not hold for time spent listening to music (χ2(3) = 2.23, p = .526). Basic characteristics of musical hallucinations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The questionnaire asked about a number of basic characteristics of inner music experiences, such as style of music (Q4), frequency (Q2), duration (Q3), familiarity (Q10), repetitiveness (Q22), and whether music was experienced in its entirety or was shortened (Q20 and 21). Participants were also asked whether their inner music included attributes such as melody, harmony, intensity, instruments, and lyrics (Q6). Table 3 presents descriptive statistics and the results of group contrasts, for the basic characteristics and acoustic details of the different categories of inner music. For main effects of inner music type, a Bonferroni-correct alpha level of .005 (.05/10) was used for all statistical tests in Section 3.2. Classical music was the most frequently reported style of inner music in participants with MH (47.7%), whereas rock music was the most frequently reported style in the musical imagery group (64.2%), and pop music in the earworms group (66.0%). The mixed experiences group reported that both pop music and classical music were the most frequent style of their inner music (both 62.5%). See Table 4 for a full list of reported music styles and their frequency. A non-parametric (Kruskal- Wallis) ANOVA with inner music frequency (Likert scale responses; see Appendix 1) as the dependent variable showed a significant main effect of inner music category (χ2(3) = 33.73, p < .001), with Mann-Whitney U tests (Bonferroni corrected alpha levels at .05/3 = .017) indicating that participants in the MH group reported the experience as occurring significantly less frequently than those in the imagery group (U = 651.5, p < .001), earworm group (U = 1153, p < .001), or the mixed experiences group (U = 468, p < .001). There was also a significant main effect of inner music type on familiarity (χ2(3) = 56.96, p < .001), with MH being rated as less familiar than imagery (U = 738, p < .001), earworms (U = 866, p < .001), or mixed experiences (U = 573.5, p < .001). For example, a typical response from a participant in the MH group in a free-text box was: “I can tell the style, but it’s not songs I’ve heard before. It’s new songs.” A non-parametric (Kruskal-Wallis) ANOVA with inner music frequency (Likert scale responses; see Appendix 1) as the dependent variable showed a significant main effect of inner music category (χ2(3) = 33.73, p < .001), with Mann-Whitney U tests (Bonferroni corrected alpha levels at .05/3 = .017) indicating that participants in the MH group reported the experience as occurring significantly less frequently than those in the imagery group (U = 651.5, p < .001), earworm group (U = 1153, p < .001), or the mixed experiences group (U = 468, p < .001). There was also a significant main effect of inner music type on familiarity (χ2(3) = 56.96, p < .001), with MH being rated as less familiar than imagery (U = 738, p < .001), earworms (U = 866, p < .001), or mixed experiences (U = 573.5, p < .001). For example, a typical response from a participant in the MH group in a free-text box was: “I can tell the style, but it’s not songs I’ve heard before. It’s new songs.” There was also a main effect on repetitiveness (χ2(3) = 21.22, p < .001), with MH being reported as less repetitive than earworms (U = 1341, p < .001) or imagery (U = 1117, p = .011), but not significantly different to the mixed experiences group (U = 618, p = .032), following corrections to the alpha level (.05/3 = .017). There was also no significant effect of inner music type on duration (χ2(3) = 11.60, p = .009), or the extent to which participants reported inner music being shortened (χ2(3) = 5.31, p = .151) or a complete song (χ2(3) = 5.77, p = .123). Almost all participants reported being able to perceive melody in inner music (97.3% in overall sample), with most also being able to perceive harmony (77.4%), intensity (66.5%), instruments (77.0%) and lyrics (75.9%). Chi-square analysis indicated no association between inner music category and whether harmony (χ2(3) = 3.02, p = .389, φ = .109), intensity (χ2(3) = 5.12, p = .163, φ = .142) or instruments (χ2(3) = 5.69, p = .128, φ = .149) were perceived, although there was an association between inner music type and whether lyrics were perceived (χ2(3) = 23.02, p = < .001, φ = .300). (Given that >97% of participants reported being able to perceive a melody, there was insufficient variation to test the association between inner music type and melody.) Further analysis showed that MH were less likely to include lyrics than other types of inner music (OR = 0.21, 95% CI [0.10, 0.41]), whereas earworms were more likely to include lyrics (OR = 2.13, 95% CI [1.14, 3.98]). Meanwhile, the imagery (OR = 1.37, 95% CI [0.70, 2.68]) and mixed experiences (OR = 1.29, 95% CI [0.56, 2.98] groups were no more or less likely to include lyrics. One participant who experienced MH commented: “...[they] tend not to involve voices but do involve many different instruments. I have occasionally heard other genres of music, including with voices, but I have not been able to understand the lyrics.” Perceived control, location, and anticipation of musical hallucinations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Further questions asked whether participants felt control over their inner music (Q23), whether they could anticipate the experience (Q24), whether they knew the trigger of the experience (Q25), whether they ever mistook it for an externally located stimuli (Q26), and whether they felt like the experience was of their own creation (Q11). Table 5 summarizes responses from these questions by category of inner music. For main effects of inner music type, a Bonferroni-corrected alpha level of .01 (.05/5) was used for all statistical tests in Section 3.3. There was a main effect of category of inner music on perceived control (χ2(3) = 11.52, p = .009), with participants reporting MH also reporting less control over the experience compared to participants in the imagery group (U = 1033, p = .003). However, there was no significant difference between perceived control of MH and earworms, or between MH and mixed experiences (ps > .129). A typical comment related to the perceived effortlessness of MH, with little or no control: “Imagination to me is a conscious effort – this isn’t.” “Like a radio station, I just have to wait for the next song.” Others, meanwhile, described techniques to stop the MH, such as distracting themselves with another activity: “I just think of something specific as a distraction (e.g., I’m thinking about typing this response correctly, so my internal background music is switched off for now).” When asked about the frequency with which their inner music was mistaken for coming from the external environment, there was a significant main effect of inner music category (χ2(3) = 95.10, p < .001), with MH being more likely to be mistaken for an external percept than imagery (U = 460, p < .001), earworms (U = 656, p < .001) and mixed experiences (U = 581, p = .008). For example: “There is usually a moment where I am not entirely sure if it is external or internal, but the quality of sound and the apparent feeling of “proximity” is strange when it is a hallucination. In other words, it might sound soft as if it should be coming from far away, and yet it does not sound as though it is coming through any barriers like walls... it is almost as if it is coming through earphones, closer to me than the outside world, but not exactly ‘in my head’” “Remembering music is like a faint shadow of real music...whereas hallucinating it really involves hearing it” There was a significant main effect of inner music type on the extent to which participants reported that their inner music was their own creation (χ2(3) = 39.09, p < .001), with participants in the MH group, counterintuitively, rating this attribute more highly than participants in the earworm group (U = 1255, p < .001), although there was no difference between the MH and imagery or mixed experiences groups at the corrected alpha level (ps > .30). There was no main effect of inner music category for reports of being able to anticipate the experience or knowing what triggered the experience (see Table 5 for statistics). Effects of musical hallucinations on mood and behavior ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Participants were asked a number of questions about the effect of inner music on their behavior and mood, including whether they were able to hum along with the inner music (Q8), whether their body moved with the music they experienced (Q9), whether the experience reflected their own feelings (Q18), or was associated with negative affect (Q15) (see Table 6). A Bonferroni-corrected alpha level of .05/11 = .0045 was used for all statistical tests in Section 3.4. The extent to which participants reported that the inner music reflected how they were feeling differed between categories (χ2(3) = 19.22, p < .001), with MH being rated lower on this attribute than imagery (U = 825.5, p < .001), earworms (U = 1539, p = .003) or the mixed experiences group (U = 522, p = .002). Although there was a significant main effect of inner music category on the extent to which the experience made the participant feel excited (χ2(3) = 15.30, p = .002), further tests did not reveal any significant differences at the corrected alpha level between MH and any other categories (all ps > .019). Notably, there was no significant effect of inner music category on the experience contributing to the participant feeling depressed (χ2(3) = 8.21, p = .042) or anxious (χ2(3) = 6.79, p = .079) after corrections to the alpha level, providing no strong evidence that MH were associated with negative affect more than other forms of inner music. For example: “It’s never anything emotional playing, usually just ‘there’. I’ve gotten used to it for the most part, but the most it makes me feel is annoyed.” Participants reported being less able to hum along with MH (χ2(3) = 24.72, p < .001), compared to imagery (U = 777.5, p < .001), earworms (U = 1204.5, p < .001), or mixed experiences (U = 481, p = .001), as well as being less able to hum the music after the experience (χ2(3) = 22.15, p < .001) compared to imagery (U = 860.5, p < .001), earworms (U = 1313, p < .001), or mixed experiences (U = 589, p = .012). Participants also reported moving their body less to MH (χ2(3) = 28.88, p < .001), compared to imagery (U = 736, p < .001), earworms (U = 1442, p < .001), or mixed experiences (U = 416, p < .001). One typical comment was: “Unfamiliar, no lyrics, no compulsion to sing or hum along” Finally, there was no association between inner music category and the extent to which it was said to affect the participant’s relationship with others (see Table 3 for statistics). Associations between musical hallucinations, inner speech phenomenology, and intrusive thoughts ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Participants also completed a small number of quantitative self-report measures relating to phenomenology of inner speech (VISQ), intrusive thoughts and thought suppression (WBSI), and auditory hallucination-proneness (LSHS). There was no main effect of inner music type on any measure of inner speech phenomenology, intrusive thoughts, or thought suppression (all ps > .092). Unsurprisingly, there was a main effect of inner music type on hallucination-proneness (χ2(3) = 20.03, p < .001), with participants in the MH category (M = 9.60, SD = 2.61) scoring higher than those in the imagery category (M = 8.18, SD = 2.50, U = 958, p = .005) and the earworm category (M = 8.19, SD = 2.31, U = 1445.5, p = .003), although not significantly differently to the mixed experiences category (M = 10.21, SD = 3.22, U = 723.5, p = .469).","The present study provided preliminary evidence for a number of phenomenological differences between MH and other forms of inner music (summarized in Table 7). Whilst some of these differences fit closely with the typical conceptualization of hallucinatory experience (e.g., experienced as externally located, uncontrollable), others were less expected and may open avenues for future research. For example, the data suggested that MH are less likely to be experienced by musicians, less repetitive, less likely to include lyrical content, and are described as less likely to be associated with one’s own feelings, compared to musical imagery or earworms. Somewhat counterintuitively, the data also suggested that MH were more likely to be experienced as one’s ‘own creation’ compared to earworms. Importantly, our findings suggest that there are key differences between, in particular, MH and earworms. Moreover, given the self-report methodology used in this study, the data provides valuable insight into how people make judgements about their own experiences (that is, what do people classify as hallucinatory?). Here, the findings are compared to previous literature on MH, but several new findings that could form the basis for future research are also interpreted. A previous systematic review by Cope and Baguley (2009), focusing on etiological factors underlying MH, indicated that hearing loss, female gender, old age, and social isolation were all risk factors for MH. The present data did not indicate that females were more likely to report MH; while the proportion of females in the MH category was higher than other groups, this was not statistically significant. The data did not show elevated levels of hearing impairment in individuals reporting MH as opposed to other forms of inner musical experience, although, given the relatively low rates of reported impairment in our sample, this may be an issue of statistical power. Numerous studies have previously linked AH or MH to hearing impairment (Cope & Baguley, 2009; Griffiths, 2000; Kumar, Sedley, Barnes, Teki, & Griffiths, 2014; Linszen, van Zanten, Teunisse, Brouwer, Scheltens, & Sommer, 2018); one interesting question, therefore, would be to examine the phenomenology of MH that seem to be linked to hearing loss, compared to those that are not. Rates of psychiatric or neurological diagnoses were also similar across different forms of reported inner music, and, furthermore, individuals with MH were not significantly more likely to say that the experience made them depressed or anxious, compared to individuals reporting other forms of inner music. Previous literature has provided mixed evidence regarding negative emotions associated with MH. The most prevalent emotional description in Saba and Keshavan (1997) study was ‘soothing’; in contrast, Evers and Ellger (2004) found that 41% of participants described the experience as ‘frightening’. Such differences may partially reflect the different populations from which data was collected, with most previous research investigating MH in psychiatric or neurological patients. Our data, meanwhile, is consistent with the view that MH, and auditory hallucinations more broadly, can occur without significant distress and outside of any need for care (Johns et al., 2014). Indeed, the only demographic factor significantly associated with MH in the present data was level of musical expertise, with individuals reporting MH being less likely to classify themselves as a musician, playing fewer musical instruments, and spending less time practising music. Previous research has suggested that musical training is associated with superior performance on tasks requiring auditory imagery generation, for example to complete a musical sequence (Aleman, Nieuwenstein, Böcker, & de Haan, 2000), or to evoke spontaneous experiences of musical imagery, and importantly, voluntarily modify the experiences (Goycoolea et al., 2007). Similarly, Pallesen et al. (2010) showed that individuals with musical training showed increased performance on an auditory working memory task, showing greater levels of cortical activation in areas of the brain typically associated with cognitive control, such as the lateral prefrontal cortex and anterior cingulate cortex. It is possible, therefore, that enhanced cognitive control among musicians may contribute to the decreased prevalence of MH, although further research is needed to investigate this issue. One of the few studies to report phenomenological details of MH (Saba & Keshavan, 1997) reports data from 100 participants with a diagnosis of schizophrenia, finding that 16 of these individuals experienced MH. These experiences were not compared directly to experiences of musical imagery or earworms; consistent with our findings, however, a substantial number were rated as being perceived as emanating from the external environment, and approximately half were rated as not under volitional control of the individual. Saba and Keshavan, however, argued that if the individual reported any volitional control of the experience, it should not be defined as an MH, suggesting that this lack of control is actually a key feature of the experience. This is in contrast to our data, which suggests that while MH were rated as much lower in controllability than musical imagery, earworms were rated intermediately between the two categories. One possibility consistent with Saba and Keshavan’s argument is that the experience of volitional control can vary along a continuum, with imagery becoming hallucinatory when it is extremely uncontrollable. Level of volitional control may also be an important factor in clinical distress, with recent studies showing that individuals that report regular AVH but have no clinical diagnosis report higher levels of control, compared to those with a clinical diagnosis (Alderson-Day et al., 2017; Daalman et al., 2011; Powers, Kelley, & Corlett, 2016). On this view, earworms would simply fall on a midpoint of this continuum between musical imagery and MH. However, our findings also suggest that MH differ in a number of other ways from both musical imagery and earworms; for example, MH are much more likely to be experienced as externally located, consistent with the argument made by Williams (2015). Indeed, given that a substantial proportion of individuals in the Saba and Keshavan study (37.5%) perceived their MH as externally located, it is unclear why only volitional control, rather than perceived location, was chosen as the main criteria by which MH were defined. Another clear difference between earworms and MH, in our sample, was the reported familiarity of the perceived music, with MH being rated as less familiar to the participant than other forms of inner music. It is possible that this feeling of unfamiliarity may add to a feeling of alienness typically associated with hallucinations, as opposed to musical imagery or earworms. In contrast, previous studies have reported that as many as 78% of cases of MH were experienced as familiar music (Evers & Ellger, 2004). There are two possible methodological differences that may account for this discrepancy. Firstly, as already mentioned, the two referenced studies focused mainly on MH occurring in psychiatric and neurological disorders, whereas rates of diagnoses were much lower in the current sample. As such, our data may reflect MH occurring across a broader population, mainly consisting of those without a need for care, and without comorbid psychiatric symptoms or brain damage. Secondly, previous studies have deliberately excluded participants reporting ‘pseudohallucinations’ (which Evers and Ellger dismiss as being linked to ‘memory representations’), although this distinction is no longer typically used in hallucinations research. As such, the present data presumably encompasses a wider variety of experiences regarded as hallucinatory; indeed, given the self-report nature of this data, a sufficient level of insight regarding the hallucinatory nature of their experience was necessary for participants to report on the features of their MH. Our data suggested that less than half of participants reported the presence of lyrics in their MH, compared to rates of approximately 80% in other forms of inner music. This is, again, in contrast to Saba and Keshavan, who noted that MH consisting only of instrumental music was rare. It should be noted that the present questionnaire only asked about lyrical content rather than concurrent MH and AVH. Nevertheless, the low frequency of lyrical content was an unexpected finding, and highlights a key difference between MH and AVH, the most frequent form of auditory hallucination. Given that cognitive neuroscientific models of AVH specify a key role for brain networks involved in speech production and perception (Allen, Larøi, McGuire, & Aleman, 2008; Moseley, Fernyhough, & Ellison, 2013), this is suggestive that non-lyrical MH may be associated with different cognitive mechanisms than AVH, rather than simply differing in content. An interesting area for future research would be to investigate differences between the phenomenology and cognitive mechanisms underlying MH with and without lyrics, as well as investigating the prevalence and phenomenology of mixed AVH and MH. A further unexpected, and rather counterintuitive, aspect in which MH differed from other forms of inner music was in rating the extent to which the music was one’s own creation. Although MH were rated as lower in controllability and familiarity, they were actually rated as significantly higher on this attribute, in comparison to earworms. A frequent argument is that earworms occur due to unintentional re-activation of memory representations of previously heard music (Kvavilashvili & Mandler, 2004; Liikkanen, 2012). In this sense, they may be viewed as not one’s own creation, since the individual realizes that they are elicited by an external stimulus (that is, the individual is aware that they were not the author of the music). In contrast, a less clear link with previously heard music may lead to recognition of creative ownership of MH, despite a lack of controllability or familiarity of the music. This argument assumes a fairly high level of insight regarding one’s MH; that said, participants in this sample presumably required sufficient levels of insight to self- classify their experiences as hallucinatory. It is particularly interesting that volitional control and authorship can be dissociated in this way, which potentially provides evidence for a dissociation between individual sense of agency and sense of ownership in relation to MH (Gallagher, 2000); that is, our data suggest that individuals may feel a lower sense of agency, but a retained sense of ownership, over MH. This study, then, has provided data on several aspects of the phenomenology of MH. In comparison to the most widely studied type of inner music – earworms – MH are less controllable, more likely to be experienced as coming from the external environment, less familiar in content, less repetitive, less likely to include lyrics, and more recognizable as one’s own creation. These differences highlight a need for greater clarity in the use of definitions when talking about different experiences of inner music. Our data is not consistent with previous claims that MH can simply be thought of as earworms that have reached pathological levels (e.g., Hemming, cited in Williams, 2015), for two reasons: firstly, MH appear to differ from earworms on a number of phenomenological attributes, rather than simply being more persistent or distressing; secondly, many individuals in our sample reported experiencing MH without having a hearing impairment, psychiatric or neurological disorder, or any resulting distress. As such, more nuanced definitions of MH and earworms are needed. Previous research tends to have studied all involuntary musical imagery as one type of experience (Floridou et al., 2015; Müllensiefen et al., 2014; Williamson et al., 2014), with some conflating the terms INMI, earworms, and MH. Williams (2015) argues that the key difference between MH and earworms is the perception of spatial location, and suggests that INMI should be used as an umbrella term which can be further separated into earworms and MH. Our data support the importance of spatial location, but also suggest that this is not the sole difference, with aspects such as repetitiveness and familiarity also appearing to be important. Other instances of musical experience not included in this survey should also be investigated, including musical pareidolia (hearing music in other sounds) and musical memories. Future research should further investigate some of the findings presented here. One limitation of the present study was that it required participants to classify their own experience as hallucinatory. Ideally, such experiences should be enquired about via a face-to-face clinical interview; however, online surveys can reveal experientially rich and sometimes unexpected aspects of hallucinatory phenomenology (see, for example, Woods, Jones, Bernini, Callard, Alderson- Day, & Badcock, 2014; Woods, Jones, Alderson-Day, Callard, & Fernyhough, 2015). Additionally, although this study showed a link between a lack of musical training and MH, it is impossible to establish cause and effect from our data. Although findings of improved cognitive control following musical training are suggestive that this may reduce the likelihood of MH, a randomized controlled trial would be needed to make a direct link. Equally, it is possible that individuals who experience MH are less likely to pursue musical training, if their inner musical experiences are persistent or appraised negatively. Another important line of research will be to investigate the cognitive and neural correlates of MH, in comparison to musical imagery or earworms, about which little is currently known. An interesting question is whether (at least some) MH can be explained within a similar framework to AVH; that is, if MH can be explained as musical imagery that has been misattributed to an external source (Fernyhough, 2016). Studies investigating the association between MH and performance on tasks requiring the monitoring of self-generated actions could be the first step in this direction. Given the numerous phenomenological differences between musical imagery and MH highlighted above, however, a self-monitoring model may be too simplistic. Although research into MH is in its infancy, this study has provided preliminary evidence regarding a number of phenomenological aspects of MH that have not previously been addressed; future research should aim to replicate and extend these findings, which can be informative of inner musical experience in its many different forms."],["Previous research has highlighted that deaf children acquiring spoken English have difficulties in narrative development relative to their hearing peers both in terms of macro-structure and with micro-structural devices. The majority of previous research focused on narrative tasks designed for hearing children that depend on good receptive language skills. The current study compared narratives of 6 to 11-year-old deaf children who use spoken English (N = 59) with matched for age and non-verbal intelligence hearing peers. To examine the role of general language abilities, single word vocabulary was also assessed. Narratives were elicited by the retelling of a story presented non-verbally in video format. Results showed that deaf and hearing children had equivalent macro-structure skills, but the deaf group showed poorer performance on micro-structural components. Furthermore, the deaf group gave less detailed responses to inferencing probe questions indicating poorer understanding of the story's underlying message. For deaf children, micro-level devices most strongly correlated with the vocabulary measure. These findings suggest that deaf children, despite spoken language delays, are able to convey the main elements of content and structure in narrative but have greater difficulty in using grammatical devices more dependent on finer linguistic and pragmatic skills. --------------------------------------------------------------------------------","This paper provides a description of the development of story-telling abilities of deaf and hearing children who use spoken English. In addition to assessing macro- (global) and micro- (local) level narrative skills, probe questions were used following the story presentation to assess comprehension abilities. A scale was devised to assess the micro- level skills of cohesion, grammatical morphemes, and narrative and evaluative devices. While previous studies assessing narrative development in deaf children have used language dependent stimuli designed for hearing children, the current study uses a non-verbal story presented in video format that does not depend on deaf children’s receptive language skills. In contrast to the findings of previous studies, deaf children showed equivalent performance to their hearing peers at the macro-level; however, performance on micro-level narrative skills was poorer, and less relevant and detailed answers were provided to the inferencing probe questions than hearing peers. This paper thus highlights the strengths and weaknesses of oral deaf children’s language abilities.","Narrative is a powerful tool that all cultures possess for organizing and interpreting experience (Bamberg, 1997; Labov & Waletzky, 1967). Children learn to tell stories by taking part in narrative practices that their parents and other adults model to them (Van Deusen-Phillips, Goldin-Meadow & Miller, 2001). Profoundly deaf children are increasingly communicating in spoken English, yet even with advances in cochlear implant technology, they continue to lack full auditory access to the spoken language that surrounds them, and so consequently persist with communication delays (Marschark & Spencer, 2015). While there is a good understanding of deaf children’s oral language development, their ability to narrate a story in spoken language has previously been addressed in only a small number of studies (Crosson & Geers, 2001). This paper focuses on narrative development in oral deaf children and addresses a broad range of narrative skills at both the macro- (global) and micro- (local) level. Narrative skill encompasses the ability to communicate a story containing sequential information usually about a past or future event (Gleason, 2002), and is considered a cornerstone of children’s language development. Children’s emerging narrative ability is crucial for developing social skills (Miller, 1994) and has been shown to predict later literacy skills (Griffin, Hemphill, Camp & Wolf, 2004; Roth, Speece & Cooper, 2002). Typically developing children’s language shows a large proportion of personal narratives (Beals & Snow, 2002; Liles et al., 1995), In everyday conversation, children as young as 2–3 years naturally retell stories or recount a sequence of events, and as they get older children increasingly become able to deal with the discourse- pragmatic requirements that underpin narrative. Several concurrently developing, higher- level language and cognitive skills are necessary to form cohesive, coherent and structured narratives (Bamberg & Damrad-Frye, 1991). These include the mastery of a variety of linguistic (lexical, syntactic and pragmatic) skills, the ability to remember and order in sequence a series of events, and to establish and maintain perspectives of a range of characters (Norbury, Gemmell & Paul, 2014). Assessing narrative development ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Narrative is assessed for typical and atypical language development (Botting, 2002; Cleave, Girolametto, Chen & Johnson, 2010) and is typically measured for two factors: the global organisation of content, known as macro-structure; and a local linguistic level which measures devices used within and across sentences, known as micro-structure (Liles, Duffy, Merritt & Purcell, 1995). The macro-structure level focuses on two aspects: the ability to construct a hierarchical representation of the story’s main elements, including the sequencing of events, introduction to the characters and setting of the scene, complicating actions, the story climax and resolution, and internal response felt by the characters and plot evaluations (Norbury & Bishop, 2003); and also a measure of information provided for specific content (e.g., Pankratz, Plante, Vance & Insalaco, 2007). Studies with typically developing children show that at around aged 4 years, children begin to use the macro components (Trabasso & Stein, 1994), and by seven years of age, children are more able to structure a story with multiple episodes. By nine-ten years of age children can tell complete stories with substantial detail (Crais & Lorch, 1994). Micro-structure elements are assessed at the word and sentence level and include devices for achieving cohesion, such as coordinating (and, but, so) and subordinating (because, when, that, if) conjunctions. These devices provide connections from one event to another and create a clearly understood sequence (Berman & Slobin, 1994). A second measure of cohesion is the unambiguous use of reference to specify and distinguish characters in the narrative, both at first mention, and through the use of anaphoric pronouns to refer back to the named character (he, she, his, her). Micro-structure becomes more sophisticated with age (Liles, 1993; Liles et al., 1995) and depends on the ability to integrate syntactic and pragmatic information (Hemphill, Picardi & Tager-Flusberg, 1991) as well as the growth of perspective taking (Tager-Flusberg & Sullivan, 1995). Narrative measures are also used to evaluate other local language aspects in children with language learning difficulties (e.g. specific language impairment: SLI), such as frequent grammatical errors of verb tense and pronoun use (Cleave et al., 2010). In addition, during the school-age years, typically developing children develop elements related to evaluative comments (Norbury & Bishop, 2003) and improve their use of literate, decontextualized language (Curenton & Justice, 2004). These features can help reduce ambiguity in a story by increasing the explicitness of character, object and event descriptions, for example through the use of adjectives, adverbs (e.g., to specify manner: carefully), or information about spoken dialogue (e.g., said, shouted; Greenhalgh & Strong, 2001). It has been suggested that such language use is dependent on vocabulary development, and an ability to mentally represent objects absent from the immediate context (McGillicuddy- DeLisi & Sigel, 1991). Narratives also reveal the links between social cognition and language development through the assessment of children’s growing story comprehension and inference-making abilities. There is little written about inference making abilities in deaf children’s narratives, but more attention has been given to atypically developing populations with cognitive differences, such as Autistic Spectrum Disorders (ASD) and SLI (Norbury et al., 2014). When a series of probe questions based on elements not explicitly mentioned in a previously heard story are used, children with autism spectrum disorders (ASD) (Tager-Flusberg & Sullivan, 1995) and children with SLI (Bishop, 1997) were more likely to be literal in their responses, showing they had not understood the story’s underlying message: a skill that was shown to be closely linked to “theory of mind” (i.e., understanding the intentions of others; Premack & Woodruff, 1978). Narrative development in deaf children who use spoken language ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ With over 90% of deaf children being born to hearing parents, the restricted access to verbal and/or signed information means that this group faces significant difficulties in their language skills, including the ability to produce a coherent narrative (Crosson & Geers, 2001). Typically-developing hearing children have frequent opportunities to engage in narrative discourse, both in interactions with others and indirectly overhearing others recount their experiences. Telling stories about themselves at school, home and in other social settings is an everyday occurrence (Crais & Lorch, 1994). Deafness itself is not a barrier to full language development, for example deaf children of deaf parents has been shown to follow the typical narrative developmental milestones in British Sign Language (Morgan, 2002). In contrast, deaf children who are not exposed to a natural sign language by parents/carers with native level of fluency have reduced opportunities for interaction and particularly incidental learning (Morgan et al., 2014). In many countries the majority of deaf children have hearing parents who themselves do not sign, and instead choose to use oral language with their children (Marschark & Spencer, 2015). Currently these children are most often educated in a mainstream setting using a spoken language. The impact of deafness on general spoken language skills has been widely documented. For example, Geers, Nicholas and Sedey (2003) investigated expressive grammar and found that deaf children with cochlear implants (CIs) showed poorer morphological and syntactic skills than their hearing peers. On average, deaf children with (or without) implants have smaller receptive vocabularies than hearing children of the same age (Eisenberg et al., 2004; Spencer, 2004), and this difference persists over time (Blamey et al., 2001; Kirk et al., 2002). With advances in neo-natal screening and hearing aid technologies, spoken language skills of deaf children are gradually improving but it is less clear what changes are occurring for pragmatic and higher levels of language use as required in narrative (e.g., Rinaldi, Baruffaldi, Burdo & Caselli, 2013). Previous studies that have specifically investigated the spoken narratives of deaf children have focused on those with CIs and have shown that in general, they lag behind their hearing peers (Boons et al., 2013a; Crosson & Geers, 2001; Guo, Spencer & Tomblin, 2013; Worsfold, Mahon, Yuen & Kennedy, 2010). Crosson and Geers (2001) videotaped 8–9 year old oral deaf children with CIs on a story telling task and found that the deaf children, in particular those with poorer ability to discriminate speech using the CI, scored poorly on narrative structure and cohesion (use of conjunctions and character references) relative to hearing peers. More recent studies have focused on using story retell with the support of picture prompts. At the micro-level, Worsfold et al. (2010) found that oral deaf children with CIs were poorer at producing high-frequency morphemes (e.g., past tense, -ed) and used fewer subordinate clauses than their hearing peers when retelling “the bus story” (Renfrew, 1997). Using the same story retell method, Boons et al. (2013a) reported no differences between deaf and hearing groups in referencing story protagonists, but hearing controls outperformed deaf children on the number of subordinate clauses used. The deaf group also had a higher percentage of utterances with morphological, syntactic or semantic errors. Finally, Guo et al. (2013) showed in a longitudinal study that children with CIs used fewer tense markers on verbs in story retelling than age-matched peers with normal hearing. At the macro-level, with the exception of a high-scoring subgroup of children who were implanted early (Boons et al., 2013a), oral deaf children with CIs were reported to achieve lower scores than their hearing counterparts. The deaf group’s bus stories were poorer in plot structure and comprised fewer essential elements in story content (Boons et al., 2013a; Worsfold et al., 2010). A limitation of previous research using story retell with deaf children is that the task depends on receptive language skills. The deaf participant must listen to and speech-read the experimenter telling the story, and must be able to divide their attention between picture prompts and the story narrator, before retelling. A further limitation noted by Worsfold et al. (2010) is that deaf children may convey some of their story content by using gestures. Without videotaping the child, it is not possible to capture this element of the narration. It is possible that deaf children with spoken language delays are still able to produce narrative with the aid of gestural substitutions. Relevant evidence comes from deaf children who spontaneously developed home signs (a form of systematic gestures) and were able to use these to create rudimentary narratives (Morford & Goldin-Meadow, 1997; Van Deusen-Phillips et al., 2001). Finally, the use of mental state vocabulary and other evaluative devices in the narratives of deaf children using spoken English has received little attention to date. This is important given the consistent finding that oral deaf children display difficulties in mental state reasoning as evidenced by a delay in passing the false belief task (e.g., Schick, De Villers, De Villiers & Hoffmeister, 2007). A recent longitudinal study found that although length of time since CI significantly improved deaf children’s narrative performance, deaf children still used fewer evaluative devices and less mental state vocabulary compared to hearing peers, which was linked to a reduced opportunity to overhear discussions about people’s intentions and emotions (Huttunen & Ryder, 2012). In summary, research to date suggests that deaf children have difficulty with both macro- and micro narrative skills, yet assessment has generally depended upon verbal story retell methods designed for hearing children. The focus in much of this previous research has been with deaf children who wear CIs, while many deaf children using spoken language are still using hearing aids. Furthermore, there is scope to provide a more comprehensive assessment by additionally including probe questions to gauge deaf children’s understanding of the characters’ intentions and mental states. Finally, some studies have concurrently investigated deaf children’s spoken English narratives and vocabulary ability (e.g. Boons et al., 2013b), but have not examined the relationship between these two abilities. The current study aimed to address each of these factors. Present study ~~~~~~~~~~~~~ We investigated the narrative abilities of deaf children who use spoken English. The children were recruited from across the UK and were representative of deaf children who used both hearing aids and cochlear implants. The deaf children were compared with a hearing control group who were carefully matched for age and non-verbal intellectual ability. To overcome the limitation of using a measure that is dependent on receptive language abilities, a video clip of a story acted out silently by two actors was employed to elicit a narrative (Herman et al., 2004). The advantage of this elicitation method is that it relies on the children’s visual rather than auditory memory. This reduces the processing demand of dividing attention between the story pictures and communicating with the experimenter, which may enable the deaf and hearing children to complete the task on more equal level. Children were assessed on their macro level skills (content and structure) and comprehension was evaluated by probe questions, which assessed understanding of the mental state and intentions of the story characters. The children’s story telling was videotaped, enabling representational gestures to be included in the scoring of narrative content and structure. In addition, a novel grammatical scale for English was devised to assess micro-level narrative skills. The children were also assessed on their one-word expressive vocabulary. As a secondary aim, the relationship between expressive vocabulary and narrative skills was then examined. It was predicted that the deaf children would show comparable performance to hearing children in terms of narrative content and structure, given that the task is not dependent on receptive language skills. However, given previous reported delays in finer linguistic, pragmatic skills, and mentalizing abilities, it was expected that deaf children would be poorer in their micro-level narrative skills and their ability to answer the comprehension questions, relative to hearing controls. As the language used in narratives tends to be more decontextualized and requires the use of more elaborate vocabulary, as well as more exact syntactic marking of temporal and causal nature of events (Curenton & Justice, 2004), it was expected that there would be a positive relationship between vocabulary and micro-level narrative skills for both deaf and hearing groups. In addition, it was expected that a relationship between micro-level narrative skills, vocabulary and the ability to infer the mental states of others as measured by the probe questions would be found, given that language ability has been shown to be a strong predictor of theory of mind skills in both hearing (Milligan, Astington & Dack, 2007) and deaf children (Schick et al., 2007). On the other hand, it was reasoned that macro-level narrative skills would depend less on the children’s general language abilities, particularly in light of the evidence that even deaf children with limited language abilities but typical non-verbal intelligence are able to construct stories through home signs.","Fifty-nine deaf children (30 boys) were recruited based upon the following inclusion criteria: (1) pre-lingual deafness (congenital or occurrence at age ≤ 1 year), (2) aged between 6 and 11 years, (3) spoken English as the preferred modality of communication, (4) no known learning disabilities or concomitant disorders such as attention deficit or autism. The deaf children’s ages ranged from 6;0 to 11;8 (M = 8;9, SD = 1;8). Their non- verbal ability was derived from scores on the Matrix Reasoning subset of the Wechsler Abbreviated Scale of Intelligence (WASI; Wechsler, 1999) and their T-scores (M = 50; SD = 10) ranged from 30 to 69 (within 2SDs above/below the mean). Table 1 summarises the background characteristics of the deaf participants in terms of cause of deafness, level of hearing loss in their better ear and type of hearing device used. All children received auditory amplification or cochlear implants (CIs) and used these devices during testing. The mean age of first implant for the CI group was 3;5 (SD = 2;0, range = 1;0 to 10;2). The majority of the deaf children’s parents were hearing, but twelve had a deaf parent: 7 of these parents specified BSL as their own preferred language, and the remainder spoke English as a first language. All deaf parents however reported that their deaf child’s preferred language was spoken English. To gain a broadly representative sample the deaf group were recruited from specialist deaf schools (5 from day schools and 2 from residential schools) but the majority from mainstream schools across the UK (24 from schools with a specialist support unit and 28 from schools without specific provision). Forty-three parents (73%) had some level of education after leaving school (university or further education college). The majority of the deaf children were White British or White European (N = 49; 83%), 4 were Asian British, 2 were Black British, and 4 were mixed race or other. Table 2 shows the participant demographic information (age, non-verbal ability, gender and whether parents had further education) for deaf and hearing children. A group of 67 hearing children (37 boys) were recruited as a typically developing control group. These children were from a range of primary schools in rural and urban settings, and when possible were from the same schools and year groups as the deaf children ensuring similar demographic backgrounds to control for social status and match on chronological age. Table 2 shows that deaf and hearing groups did not significantly differ in terms of age (M = 8;10, SD = 1;6; range = 6;0 to 11;11) and non-verbal ability. There were no significant differences between groups in terms of gender, whether the parents had further education (N = 51) (Table 2), or ethnicity (χ2 (3) = 3.54, p = 0.32).","The UCL Research Ethics Committee gave ethical approval for the study. Children were recruited either by contacting deaf schools and specialist support units directly, or by establishing contacts with parents via the National Deaf Children’s Society. Informed written consent was obtained from parents/guardians prior to testing. Children gave verbal consent at the start of the testing session and were informed they could opt out at any time. Language measures ~~~~~~~~~~~~~~~~~ All children were tested using measures of narrative ability and spoken English expressive vocabulary. Narrative ability Children were tested on the Narrative Production Test (originally the BSL Production Test; Herman et al., 2004) with an English grammar adaptation. First the child watches a short, silent story on a laptop. The two children in the video act out a series of events without the use of language (see Table 3 for a descriptions of each story episode). Participants are instructed to watch the story carefully and to remember it so they can retell it immediately after viewing. To encourage the child to tell the whole story, the experimenter leaves the room and returns once the video has finished. The child is able to watch the film a second time if they wish. When the experimenter returns, the child is asked to tell the story and the experimenter listens to the child’s response without prompting. After completion, they are asked two probe questions to assess story comprehension and inferencing skills: (1) Why did the boy throw the spider? (2) Why did the girl tease the boy? The children’s narratives and responses to the questions were video recorded and then transcribed for analysis. All transcripts were checked against the video recordings by a second examiner. Discrepancies were discussed and agreement between examiners was obtained for all transcripts. Scoring narratives Table 4 provides an overview of the method used to score the children’s narratives. At the macro-level, the narratives were evaluated for content and structure following the scoring guidelines of Herman et al. (2004). Narrative content (i.e., the level of detailed information in the narrative) was scored by awarding one point for each mention of 15 specific story episodes (Table 3), plus a further point for mentioning any “additional information” in the story (e.g., the spider was horrible) giving a maximum of 16 points. As the stimuli material contains only gestures and actions, this prompted some children (deaf and hearing) to use gesture in their story retellings. This was mainly co-speech gesture, but on a few occasions children used silent mime e.g., a gesture to represent holding a sandwich up to the mouth and pretending to eat it. These gestures/mime were included in the scoring of story content for both deaf and hearing children, therefore both the video and transcribed speech were referred to when scoring narrative content. Narrative structure, the global organisation of story content, was scored using a high-point analysis (Labov & Waletzky, 1967) based on six key elements: (1) orientation (2) two complicating actions, (3) climax and (4) resolution. Each section is awarded 1 or 2 points depending on the amount of detail given. A further point is awarded for: (5), evaluation (i.e., where the child presents their own perspective on the characters’ feelings or expresses their own views). Responses to questions were also included; and (6) narrative sequence (i.e., correct order of story episodes). A maximum of 12 points was thus awarded for narrative structure. After extensive piloting and comparison of English narrative norms from other research, a scoring scheme was created to assess micro-level narrative skills in English for the same stimuli: a score for grammatical markers and narrative devices was generated by considering narrative cohesion, grammatical morphemes, and narrative and evaluative devices (Maximum 29 points). Responses to both the spontaneous story and the probe questions were included in scoring. Narrative cohesion included the use of referents to specify a character, and the use of conjunctions. A referential cohesion score (maximum 4 points) was based upon the first introduction of the story character(s) and whether references were consistently clear throughout. A maximum of 2 points for first introduction was scored in the following way: 0 points for no first mention 1 point for unspecified pronoun (e.g., the girl) 2 points for non-presupposing introduction using indefinite article(s) and noun or number (e.g. a girl). Reference maintenance points (maximum 2) were assigned based on the following: 0 points for unclear referencing 1 point for some ambiguity in references 2 points for clear references throughout (i.e., uses pronouns and contrasts characters effectively). A conjunction score (maximum 6 points) comprised the use of basic coordinating conjunctions (e.g., and, but), the use of logical markers (e.g., because, if) and the inclusion of subordinate clauses (e.g., the girl picked up the spider that was crawling across the floor). A maximum of 2 points were awarded for each based on the following scale: 0 points for no inclusion. 1 point for 1–2 uses. 2 points for 3+ uses. Nine types of English grammatical morphemes were analysed: articles, prepositions, regular verb forms, irregular verb forms, agreement in grammatical gender, agreement in grammatical person, use of negatives and use of modal verbs (maximum 15 points): 1 point was awarded for inclusion and correct use of articles throughout the narrative A maximum of 2 points were awarded for inclusion and correct use of prepositions: 0 points for no prepositions or rare correct use 1 point for including 2–3 prepositions (at least 2 different examples e.g., on, in, at) correctly (accuracy <50%) 2 points for 4+ prepositions correctly used (accuracy >90%) A maximum of 2 points each was rewarded for regular verb inflections (e.g., she walked/walks/was walking), irregular verb forms (e.g., he bites/he bit/had bitten), agreement in grammatical gender (e.g., sheshookherhead) and agreement in grammatical person (e.g., theywerebrother and sister) using the following scoring method: 0 points when errors were made most of the time (>50%) 1 point when errors were made some of the time (10–50%) 2 points when errors were rarely made (<10%) Errors included both omissions (e.g. the girl walk__ in; the boy __ angry) and commissions (e.g. the boy throwed the spider). A maximum of 2 points each were awarded for the correct inclusion of negatives, e.g. the girl didnt/did not know (excluding “I dont know”) and modal verbs, e.g., theremighthave been, heshouldhave got) using the following scoring method: 0 points for no usage 1 point for 1–2 occurrences 2 points for 3+ occurrences A maximum of 4 points was awarded for the inclusion of narrative and evaluative devices. One point was awarded for the inclusion of one or more examples of each of the following: Direct (e.g. the girl said no) or indirect speech or thought (e.g., the girl thought to herself) Adjectives e.g., lazy, hungry, bored Adverbs describing manner e.g., slowly, cunningly, carefully Intensifiers e.g., very, really, so; or de- intensifiers e.g., quite, nearly, almost Finally, the story comprehension and inferencing questions were allocated a maximum of two points per question depending on whether responses were partially or fully correct. The questions tested whether the children had understood the content of the story, as well as the intentions of the story characters (maximum 4 points; see Appendix A for example correct responses). Reliability of the narrative production test As there is no previously published reliability data for the Narrative Production Test used for English, intra-rater reliability of the test was assessed by two independent coders. All narratives were scored by both coders for structure and content, and relevance of answers to the probe questions. High inter-rater reliability was found for each score on each sub- scale of the test (Content: r (128) =0.98, p <0.001; Structure: r (128) =0.95, p < 0.001; Questions: r (128) =0.92, p < 0.001). The second experimenter also scored 110 randomly selected narratives (86%) for grammatical markers and narrative devices, and inter-rater reliability was also excellent (r (110) = 0.96, p < 0.001). Thirteen of the narratives (10%) were randomly selected and scored a second time by the same coder. An overall total score was calculated and a strong correlation between scores at both time points was found (r (13) = 0.98, p < 0.001). Vocabulary The expressive one word picture vocabulary test (EOWPVT; Brownell, 2000) was used to assess single word vocabulary production. The EOWPVT was standardised on children with normal hearing, but has frequently been used with deaf children as a measure of English vocabulary (Geers, 1997; Kyle & Harris, 2006; Moeller, 2000). The full test was administered as per the instruction manual. The children are presented with single pictures that test knowledge of primarily simple nouns (e.g., train, pineapple, kayak), but also some verbs (e.g., eating, hurdling), and category labels (e.g., fruit, food). The EOWPVT was developed in the USA and so a few pictures (n = 3) were substituted with alternative pictures to make the test more culturally relevant for children in the UK (e.g., raccoon with badger). Statistical analyses Independent t-tests were used to compare group means on narrative skills using raw scores. Significance criteria were set at p < 0.05 and Bonferroni corrections were applied to all multiple comparisons. A series of correlations were carried out to explore the relationship between narrative ability and age, nonverbal ability, and vocabulary. A hierarchical multiple regression was conducted to explore the extent to which vocabulary contributed uniquely to performance on the grammatical markers and narrative devices (micro-level narrative skills). Analyses were performed using SPSS v22.0. Post hoc power analysis (G*Power 3.1 software) showed sufficient power for the total group (n = 126, effect size (d) =0.64, Power = 0.97). Preliminary analysis ~~~~~~~~~~~~~~~~~~~~ Overall, the hearing group children (M = 41.91, SD = 7.78) had a significantly higher total Narrative Production Test total score (maximum score = 61) than the deaf group children (M = 35.88, SD = 10.70; t (124) = −3.65, p <0.001, Cohen’s d =0.64). The hearing children (M = 108.86, SD = 11.04) also had significantly higher standardised EOWPVT scores than the deaf children (M = 91.95, SD = 18.87; t (124) = −6.08, p <0.001, Cohen’s d = 1.09). To account for the heterogeneity of the deaf children, within group differences on overall scores on the Narrative Production Test were investigated according to type of hearing amplification (CI vs. HA) and level of hearing loss, groups were matched on age and non-verbal ability (ps > 0.05). No significant difference in total Narrative Production Test scores were found between deaf children using CIs (N = 22; M = 34.5, SD = 10.14) and those deaf children wearing hearing aids (N = 37; M = 36.70, SD = 11.07; t (57) = −0.76 p = 0.45, Cohen’s d = 0.21). There was no relationship between severity of hearing loss in the better ear and total narrative scores (mild-moderate: N = 10; M = 35.1, SD = 13.52, severe: N = 25, M = 36.48, SD = 9.82 or profound: N = 22; M = 34.72, SD = 10.49; p all > 0.05). Main group comparisons ~~~~~~~~~~~~~~~~~~~~~~ Table 5 displays means, standard deviations, group comparisons and effect sizes for the children (deaf and hearing) on each of the narrative skills subscales: content, structure, grammatical/narrative devices, and inference questions. Macro-level narrative skills Narrative content. For total scores on story content, the t-test showed that there was no significant difference between deaf and hearing children, suggesting the level of information recall in the narrated stories was similar in the two groups of children. Narrative structure. Similarly, there was no significant difference between groups on overall scores for global narrative structure indicating that the deaf and hearing children were similar in their ability to organise story content following key elements (i.e., including detail on the orientation, complicating actions, climax, resolution, evaluation and story structure). Micro-level narrative skills: grammatical markers and narrative devices Overall, the deaf group children obtained significantly lower scores for grammatical markers and narrative devices (p < 0.001; Table 5). Cohesion. The deaf children’s scores on the referential cohesion scale was significantly poorer then the hearing children (p < 0.001; Table 5), suggesting that hearing children made better use of reference (e.g., the use of anaphoric pronouns was less ambiguous). The hearing group also scored significantly higher on the conjunction score (p < 0.001), showing that they were more sophisticated in their use of temporal conjunctions and subordinate clauses in order to express semantic relations across their stories. Grammatical morphemes. The deaf group’s score for grammatical morphemes was significantly lower than the hearing group (Table 5). This suggests that deaf children made more omissions and errors with words that carry grammatical information. An example from an 8-year-old deaf child illustrates incorrect regular and/or irregular verb inflections, either omissions (e.g., he pick_ it up) or commissions (e.g., he putted); the omission of articles (e.g., on _ floor); and the omission of prepositions (e.g., he putted it _ the sandwich): “Then he saw the spider on floor. Then he pick it up. Then he putted it the sandwich.” Narrative and evaluative devices. There was no significant difference between groups for the use of narrative and evaluative devices (Table 5), suggesting that the deaf and hearing children were equally able to use evaluative language such as adjectives (e.g., the spider was horrible) or spoken information about the dialogue (e.g., the boy said, “give me the sandwich”). Comprehension and inference questions Finally, the hearing group’s mean score on the story comprehension and inference questions was significantly higher than the deaf group children (p < 0.001; Table 5) and the effect size was large (Cohen’s d = 0.74). This suggests that on average the hearing children demonstrated greater understanding of the underlying messages and provided more detailed explanations based on inferencing of the reasons for the characters’ actions. Appendix B shows two example narrative transcripts of a deaf and hearing child to further illustrate the group differences found in narrative abilities. Predictors of performance ~~~~~~~~~~~~~~~~~~~~~~~~~ Age and non-verbal ability were first investigated as predictors of performance on the narrative skills. Deaf children’s age was found to correlate moderately with scores of story content, r (57) = 0.47, p < 0.001, and structure, r (57) = 0.47, p <0.001, but not for inference questions or grammatical markers and devices. For hearing children, age had a weak-moderate correlation with all of the narrative skills (Content: r (65) = 0.39, p <0.001; Structure: r (65) = 0.38, p = 0.002; Inference questions: r (65) = 0.30, p =0.01; Grammar: r (65) = 0.33, p = 0.006 ps≤.05), and non-verbal ability (WASI matrix) correlated weakly with grammatical markers and narrative devices, r (65) = 0.34, p =0.004. Table 6 shows partial correlations (controlling for age and non-verbal ability) between vocabulary (EWOPVT) and narrative skills for both groups. The vocabulary measure (EOWPVT) correlated strongly with deaf children’s use of grammatical markers and narrative devices scores (p < 0.001) and there were weaker correlations with scores on inference questions and narrative structure (p < 0.05). The scatterplot in Fig. 1 illustrates the strong positive correlation between the residual scores of grammatical markers and vocabulary for deaf children. Vocabulary (EOWPVT) correlated weakly with narrative structure (p<0.05), but did not correlate with any of the other hearing children’s narrative skills (all ps >0.05). The relationship between each subscale of the Narrative Production Test showed a moderate to strong correlation between each section for deaf children. For the hearing children, mean scores on narrative content and structure strongly correlated, but the correlations with grammatical markers, while significant, were weaker (Table 6). There were no correlations between inference questions and other narrative subscales for hearing children. As performance on the grammatical markers and devices narrative subscale was weaker for deaf children we wanted to explore the contribution of vocabulary as a measure of language ability to children’s performance on this subscale, over and above age, nonverbal ability and a diagnosis of deafness. A hierarchical multiple regression was carried out across all participants (Table 7). In the first stage of the analysis, non- verbal ability (WASI matrix) and age were entered as independent control variables (IV) at step 1. The resulting multiple regression equation was statistically significant, F (2, 123) = 7.91, p =0.001, adj. R2 = 0.10. At step 2, with the entry of EOWPVT scores into the equation, there was a statistically significant increment in the prediction of variability in the children’s grammatical markers and narrative devices score, F (change) = 60.73, p <0.001. The overall model remained significant, F (3116) = 26.90, p <0.001, R2 =0.40, accounting for an additional 30% of variance. At step 3, a dichotomous IV: deafness (1, deaf; 0, hearing) was additionally entered as a dummy variable. The model remained significant, (F (4, 115), = 22.96, p < 0.001) and group accounted for only a further 3% of the variance (R2 =0.43). The final beta weights indicated that EOWPVT scores, age, and deafness all significantly independently contributed to predicting performance on grammatical markers and narrative devices. Therefore, children’s vocabulary skills (EOWPVT scores) contributed significantly to predicting variability in performance on grammatical markers subscale even after controlling for age and diagnosis of deafness.","As deaf children are starting to communicate exclusively in spoken language, the main aim of the current study was to compare deaf and hearing children’s narrative ability in spoken English at both macro and micro levels. Narrative is an important skill for children to master for several social-emotional and educational functions. A different method of elicitation was employed from the conventional picture prompt and verbal story retell, by showing all children a non-verbal story in video format, in order to reduce the demands on deaf children’s auditory memory. As predicted, there were no differences at the macro level of narrative (content and structure) between deaf and hearing children. Additionally, both groups of children displayed the same pattern of improved performance for content and structure with age. However, there were clear differences in micro-level skills; in particular, the deaf children’s performance was significantly poorer in terms of grammatical morphemes and narrative cohesion. These micro-level findings are consistent with previous studies, but our other results contrast with other findings that show that deaf children also lag behind typically developing peers on global narrative skills (Boons et al., 2013a; Crosson & Geers, 2001; Worsfold et al., 2010). There was also a key difference in narrative understanding and inferencing as measured by the probe questions, suggesting that linguistic development is important for deeper understanding of narratives. Equivalent performance between oral deaf and hearing children in narrative structure and content indicates that if the task is designed so that assessing story retell ability is not dependent on receptive language skills, deaf children are able to tell a coherent story at the global level. The dissociation between deaf children’s narrative macro- and micro- structure in the present study suggests that the latter is more dependent on purely linguistic and pragmatic skills. In support of this suggestion, micro-level narrative skills correlated strongly with deaf children’s vocabulary, whereas in terms of macro-level narrative skills, there was only a weak correlation between vocabulary and narrative structure for both groups. While micro-level narrative skills depend on an elaborate vocabulary and syntactic cohesion to clearly mark the temporal and casual nature of events (Curenton & Justice, 2004), macro-skills may depend less on linguistic skill and more on general cognitive mechanisms. The videotaping of all children in the present study enabled the coding of gesture to capture some additional content in children’s narratives that would otherwise be overlooked. While the children predominantly used co-speech gestures in their story telling, both deaf and hearing children used a number of representational gestures in their narratives to convey particular sequences of events (e.g., gesturing holding a sandwich up to the mouth to represent the episode where the girl pretends to eat a sandwich). Even deaf children with very limited language, reliant on an invented gesture system, have previously been found to recount stories of the same type and structure as hearing children when non-linguistic gestures have been coded (Van Deusen-Phillips et al., 2001). The findings of the present study support the argument that despite language delays in vocabulary and micro-level devices, deaf children experience social interactions, which can trigger an interest in recounting and linking past events. It is possible that the story telling function is robust in spite of reduced linguistic capabilities (Morford & Goldin-Meadow, 1997; Van Deusen-Phillips et al., 2001). Strengthening this possibility, deaf and hearing children showed comparative performance for narrative and evaluative devices including the use of direct or indirect speech, intensifiers, adjectives and adverbs of manner. This suggests that deaf children are aware of the importance of these elements in story telling. Consistent with previous studies, the deaf and hearing children’s performance was markedly different for micro-level skills that are dependent on more efficient linguistic and pragmatic abilities (Boons et al., 2013a; Crosson & Geers, 2001; Guo et al., 2013; Worsfold et al., 2010). The use of grammatical morphemes was notably different between the two groups of children. Deaf children were more likely to over-generalise regular verb rules (e.g., the boy putted), and make errors in the omission of articles, prepositions and verb inflections. This finding is expected because previous studies have found that even a moderate hearing impairment can impact a deaf child’s ability to perceive these difficult to segment morphemes, which leads to less well instantiated representations (McGuckian & Henry, 2007; Moeller et al., 2010). The deaf children also used fewer conjunctions and subordinate clauses, which are important for linking semantic representations across a narrative (temporally and causally) to form a well-structured, cohesive story (Crosson & Geers, 2001). The deaf group also had a greater tendency to introduce characters with ambiguous references. For example, using a definite article (the), rather than indefinite article, (a) plus noun (boy). In addition, they were also more likely to refer to both characters (i.e., the girl and the boy) as “he” throughout the story, creating confusion. These referencing errors and lack of syntactic cohesion suggest some deaf children are unfamiliar with discourse and pragmatic conventions presumably linked to reduced exposure to direct and indirect narrative language, and/or lack the pragmatic skill that requires an awareness of the needs and perspective of the listener (Bruner, 1986; Morgan et al., 2014). Therefore, despite being able to convey the rudimentary elements of the content and structure of a story, these findings suggest that a disruption to language acquisition has a detrimental effect on narrative skills in oral deaf children. Linked to social-cognitive influences on narrative, the deaf group provided less relevant and/or detailed answers than the controls to probe questions that focused on understanding a characters’ intentions or feelings. While deaf children are able to use emotion and mental state terms in their narratives (e.g. the boy was angry), our results point to a difficulty in determining the psychological causes of these mental states. Studies investigating narrative skills in children with autism (Tager-Flusberg & Sullivan, 1995) and SLI (Norbury et al., 2014) have also found this distinction between emotion and mental states. The deaf children’s poorer performance in answering the probe questions in the present study was expected given that deaf children generally show difficulty with theory of mind (false-belief) tasks (Peterson & Slaughter, 2006). Language ability is strongly related to theory of mind understanding in typically developing (Milligan et al., 2007) and deaf children (Schick et al., 2007). For the deaf group in the current study, grammatical markers showed a moderate positive correlation with the probe questions, suggesting that a threshold of linguistic skills are necessary to make causal links about others’ mental states and actions. The relationship between vocabulary and probe questions, while significant, was weaker. The precise role of language ability remains uncertain, but it is thought that reduced exposure to conversational interactions caused by deaf children missing out on the conversations that surround them in hearing families and educational environments is likely to impact the ability to give emotional explanations and engage in causal discourse (Morgan, Hjelmqist, & Meristo, in press; Rieffe, Terwogt & Cowan, 2005). It is important to highlight that a number of previous studies have shown that groups of deaf children implanted with a CI at a very early age (Boons et al., 2013a) and those with an early diagnosis of deafness (Worsfold et al., 2010) perform at the same level as their hearing peers in micro- as well as macro- narrative skills. However, Boons et al. (2013a) acknowledged the variability in spoken language skills within the early implanted children. In the present study, there was no difference between deaf children with conventional hearing aids and those with CIs in narrative performance; neither was there a difference based on level of hearing loss. However, among the group of CI users in the current study there was large variation in the age at implantation and length of exposure to auditory input, which might explain the lack of consistent findings. In conclusion, the deaf children in the present study were able to construct a narrative at the macro level, but showed a weakness with micro-structural devices that are more dependent on finer linguistic and pragmatic skills. More research is needed to explore the factors that drive the development and possible dissociation of macro- and micro- narrative skills in deaf children. The narrative task and subsequent coding presented in this study also has the potential to be used with other groups of children and to therefore have a broader impact across the field. The study of deaf children compared with other groups with atypical narrative skills will be informative in delineating the particular influences of sensory and neuro-cognitive impairment on this crucial aspect of language development."],["Over 98% of our genome is non-coding and is now recognised to have a major role in orchestrating the tissue specific and stimulus inducible gene expression pattern which underpins our wellbeing and mental health. The non-coding genome responds functionally to our environment at all levels, encompassing the span from psychological to physiological challenge. The gene expression pattern, termed the transcriptome, ultimately gives us our neurochemistry. Therefore a major modulator of mental wellbeing is how our genes are regulated in response to life experiences. Superimposed on the aforementioned non-coding DNA framework is a vast body of genetic variation in the elements that control response to challenges. These differences, termed polymorphisms, allow for a differential response from a specific DNA element to the same challenge thus potentially allowing ‘individuality’ in the modulation of our transcriptome. This review will focus on a fundamental mechanism defining our psychological and psychiatric wellbeing, namely how genetic variation can be correlated with differential gene expression in response to specific challenges, thus resulting in altered neurochemistry which consequently may shape behaviour. --------------------------------------------------------------------------------","The human genome has evolved to include a combination of both highly conserved regions of regulatory non-coding DNA (ncDNA) found across many species and human-specific regulatory DNA elements which together act to regulate expression of mRNA. This combination of DNA elements allows determination of where, when, how much and for how long, genes are expressed in the human brain in response to normal developmental, psychological and physiological cues, Figure 1. Many of these elements exhibit genetic variation which is not only associated with risk for a specific condition, but has also been demonstrated to alter the regulatory properties of the gene. The functional interpretation and analysis of ncDNA variation can be initially addressed in silico by overlaying its position on databases containing characterised and predicted functional elements within the genome, Box 1. The most easily accessible free database is the Encyclopaedia of DNA Elements (ENCODE; https://www.encodeproject.org/) which is a collaboration of research groups funded by the National Human Genome Research Institute [1,2••], this can be used in combination with a plethora of other database browsers [3] such as the University of California Santa Cruz (UCSC) Genome Browser (http://genome.ucsc.edu/) [4]. This review will begin with an introduction to the most conserved regulatory regions in the genome and how these may be functionally modified by the simplest and most extensively studied class of genetic variation, single nucleotide polymorphisms (SNPs). The review will then focus on human regulatory elements that are associated with neuropsychiatric conditions which are larger blocks of DNA variation such as variable number tandem repeats (VNTRs) and non- long terminal repeat (non-LTR) retrotransposons, Box 2.","Evolutionary conserved regions (ECRs) in the genome can be easily found using the ECR browser (https://ecrbrowser.dcode.org/) [5]. ECRs in this browser are typically defined as regions of sequence within the human genome that retain 70% or more sequence identity over a window of 100 bases when compared to the corresponding region of sequence in other species, this will frequently include exons in coding DNA. However, Pennacchio et al., were amongst the first to demonstrate that ECRs in the non-coding DNA (ncECRs) could be important, particularly in directing gene expression in the CNS. They determined by use of a transgenic mouse model that whilst ncECRs could direct expression in a broad range of anatomical structures in the embryo, the majority of the ncECRs tested directed expression to various regions of the developing nervous system [6]. Subsequently, consistent with this, a third of paralogous ncECRs examined were predicted to have regulatory activity in the brain [7], for example, deletion of ncECRs in the neuronal transcription factor Arx resulted in substantial alterations of neuron populations and structural brain defects in a trangenic model [8]. Furthermore, the combinatorial complexity of gene expression was exquistely demonstrated in a transgenic model of craniofacial morphology in which the action of multiple ncECRs driving expression of many genes resulted in a vast array of facial differences [9••]. These studies demonstrated that ncECRs can have important transcripitonal regulatory properties, therefore the expectation is that polymorphism in such domains has the potential to modify interactions with transcription factors and thus affect regulatory function.","Early studies of genetic variation correlated with mental health focused on DNA variation in exons encoding proteins. Most of these studies addressed SNPs; thus a SNP that changed an amino acid (non-synonymous change) or resulted in a truncation of the protein could be mechanistically relevant as it could alter protein function. However, with technological advances the ability to address SNPs in genome wide association studies (GWAS) rather than solely exons, demonstrated that the vast bulk of SNP variation associated with behavioural and psychiatric conditions was in ncDNA [10,11]. GWAS has led to significant discoveries in defining some of the genes involved in neuropsychiatric disorders and demonstrated there is genetic overlap between many of the major psychiatric disorders [12,13]. In several examples the proteins identified can work together to alter a key pathway underpinning wellbeing and mental health, such as those modifying calcium signalling [14]. Understanding the mechanistic significance of SNPs in ncDNA for a specific condition has been a much more difficult task than for SNPs found in exons. A SNP in ncDNA could be tagging a regulatory domain 10K+ bases from itself (a tagging SNP is representative of a large section of DNA that is inherited as one, thus the SNP is not the causative agent but rather highlights a region of DNA). Analysis of SNP variation within ENCODE and associated data sets can determine if it is present in a genomic region defined as a regulatory domain. In this scenario, the SNP could affect the efficiency or specificity with which a transcription factor, proteins which modulate the process of transcription, would bind to this regulatory DNA sequence, Figure 2. We and others demonstrated that SNP polymorphisms in ncECRs which correlated with known behavioural problems could modify the regulatory properties of the ECR including those associated with depression located in BDNF, BICC1 and galanin genes [15–19]. Furthermore multiple ncECRs may be required for appropriate gene expression, for example eight conserved ncECRs were identified at the schizophrenia- associated MIR137/DPYD locus of these, six were shown to be positive transcriptional regulators, and two negative transcriptional regulators in a human cell line model [18,19]. Bioinformatic analysis of this locus using the Psychiatric Genomics Consortium GWAS dataset for schizophrenia highlighted five of the ncECRs had genome-wide significant SNPs in, or adjacent to their sequence [11]. Epigenetic marks which are indicative of active or inactive chromatin, are often found at regulatory DNA. Genetic variation such as GWAS risk SNPs, can effect such epigenetic parameters impacting on long term regulatory changes in response to challenge [20]. Both local (gene specific) and global (multigene) epigenetic changes have been implicated in neuropsychiatric disorders [21,22] and the NIH Roadmap Epigenomics Consortium (http://www.roadmapepigenomics.org/) data can be utilised to analyse such data. For example, local methylation variation at the glucocorticoid receptor gene has been associated with prenatal and postnatal depression [23•] and global differences in methylation in astrocytes have been associated with depression [24]. Simplistically, methylation of regulatory regions is considered a repressor of transcription as it interferes with transcription factor binding by limiting the accessibility of specific DNA recognition sequences. The ability to rapidly address genetic variation on a ‘road map’ of regulatory domains has allowed the development of a significantly better understanding of how the ncDNA GWAS SNPs can be mechanistically involved in mental health issues [25]. This can be further updated within the UCSC browser which permits new, novel data to be overlaid on the existing data from ENCODE.","SNP variation is not the only example of ncDNA variation that can affect the regulation of gene expression. Many of the best characterised genetic polymorphisms correlating with mental health issues are found in repetitive DNA, Figure 2. These include the VNTRs [26], examples of which have been identified in key behavioural and mental health-related genes. VNTRs have been demonstrated to be both biomarkers and transcriptional regulators in genes such as the serotonin transporter, the dopamine transporter and monoamine oxidase A [21,22,27–32]. In these three examples, the primary DNA sequences of the VNTRs are rapidly evolving such that humans have their own specific VNTR sequences. All three of these monoaminergic genes contain a minimum of two VNTRs that have been demonstrated to act both independently and synergistically as transcriptional regulators whose function is further modulated by the repeat copy number within the VNTR [21,29,31]. The copy number of the repeat itself is also a biomarker for good mental health and wellbeing thus correlating function with phenotype [27,28,31,33,34]. Perhaps not unexpectedly VNTRs and GWAS SNPs in the same promoter may act additively or synergistically to regulate gene expression. This is exemplified by one of the promoters of the schizophrenia candidate risk gene, MIR137 [35,36•], where experimentally in vitro, the VNTR in the promoter can support differential reporter gene expression based on the copy number of the repeat within the VNTR, and inclusion of the promoter region encompassing the GWAS SNP can further modulate expression depending on the allele of the SNP present. This illustrates a route to identifying the potential functional significance of non-coding variants in transcriptional or post transcriptional regulatory mechanisms in areas distinct from the region of the DNA in which the GWAS SNP is found. The rapid evolution of VNTRs has been noted more globally for contributing to primate evolution; analysis in humans and non-human great apes identified that genes with VNTRs have higher expression divergence than those without [37]. The association of VNTRs with gene expression is reflected in the finding that VNTRs are enriched in promoter regions and locations close to transcriptional start sites for mRNA expression [26,38,39]. Generally, VNTRs have not been analysed as extensively as SNPs which may be attributed to the requirement to perform PCR to genotype each VNTR target and the inability to accurately identify such regions in the initial short read whole genome sequencing protocols. Improved depth and coverage in whole genome sequence combined with the development of bioinformatic programmes such as ExpansionHunter may improve the association of VNTRs and other repeat variants with neuropsychiatric conditions [40].","Non-long terminal repeat (non-LTR) retrotransposons are mobile DNA elements that can copy and paste themselves into new genomic loci and are therefore polymorphic for their presence or absence at specific loci in the genome. These retrotransposable elements (RTEs), also known as ‘jumping genes’, can range in size from a few hundred to 6000 base pairs and have been shown to be major modulators of gene expression at several levels. Non-LTR retrotransposons comprise three classes, long interspersed nuclear element 1 (LINE-1), Alu and ‘SINE-VNTR-Alu’ (SVA), Figure 2. LINE-1 expression has been implicated in many neuropsychiatric conditions such as depression [41], addiction [42], schizophrenia [43–45] and autism [46]. The other two classes have also been implicated in CNS function, for example variation in the Alu sequence within an intron of the TOMM40 gene is associated with non-pathogenic cognitive decline [47] and X-Linked Dystonia-Parkinsonism is associated with the presence or absence of a SVA in the TAF1 gene [48••]. SVAs contain several distinct VNTR elements and have properties consistent with transcriptional regulation [49,50•,51]. There has been tremendous interest in RTEs, due to their ability to make a copy of themselves which then can reinsert at a different locus in the genome of that cell. Depending on when this occurs, it results in either novel heritable germline variation or somatic mutation that can alter cellular function in only the individual affected. The former generates a large reservoir of de novo genomic variation in the population which, to date, is poorly characterised as it is very seldom annotated properly in the DNA sequence databases. The mobilisation or ‘jumping’ of RTEs is proposed to increase both with age within the adult CNS and to comprise one of the key mechanisms underpinning age-related CNS problems [52••,53,54•]; several instances where this variation has been addressed, have associated it with disease [48••,55]. More recently bioinformatic analysis has improved to allow robust identification of RTEs using programmes such as TEBreak (https://github.com/adamewing/tebreak) and MELT [56•].","The basic transcriptional mechanisms outlined in this review will also operate during development. Conditions such as schizophrenia and autism are often referred to as having a neurodevelopmental origin [57,58]. Modulation of the transcriptome during development could have a significant effect on the wiring of the brain and therefore how information is processed in the future. Early in vivo work using a mouse transgenic model indicated how a human serotonin transporter VNTR could differentially affect gene expression in key areas of serotonergic lineage based on the copy number of the repeat unit [32,59]. The regulation of regulatory domains during development will be determined by the co- expression of transcription factors as exemplified by the schizophrenia associated gene CACNA1C whose expression in development mirrors that of the transcription factor EZH2 an important regulator of the CACNA1C gene promoter [14]. An argument can be made that alterations in the transcriptome at specific times in foetal development could result in a physical change in neuronal connections that would be more difficult to correct than transcriptome changes in the adult [60].","The identification of variation in the non-coding part of the genome which affects the regulation of gene expression in part explains the often episodic nature of mental health conditions. In addition, it offers the potential for resolution of these conditions by a variety of interventions ranging from pharmaceutical to cognitive behavioural therapy which modify the signalling pathways targeting specific gene regulatory domains, modulation of the ‘stress’ driving such pathways would alter the transcriptome and hence brain chemistry, Figure 1. A prior exposure to trauma or stress could leave a molecular scar of that event, represented by an epigenetic change which alters parameters of transcriptional or post transcriptional regulation in the medium to long term [61]. It is often considered that the environmental challenge needed to affect mental health should be severe, which is not necessarily correct. For example ‘normal’ child development could also have an effect on mental health and wellbeing [23•,27,33,62]. Similarly a more general approach to maintaining good mental health via diet and exercise could play a role as they could affect the cellular signalling pathways that affect mental health [63,64]. However these issues only illustrate the complexity of defining ‘life style/environment’ and its effect on our wellbeing given the complex nature of life-long experiences in defining our transcriptome, which in turn affects the neurochemistry that ultimately shapes CNS function. It’s often said that the genome is the roadmap through which ‘life style’ shapes the individual, however one can argue that the roadmap is unique for each one of us and we all have our own route to travel [65].","None of the authors have interests to declare, thus we are stating officially ‘Declarations of interest: none'."],["Having international social ties carries many potential advantages, including access to novel ideas and greater commercial opportunities. Yet little is known about who forms more international friendships. Here, we propose social class plays a key role in determining people's internationalism. We conducted two studies to test whether social class is related positively to internationalism (the building social class hypothesis) or negatively to internationalism (the restricting social class hypothesis). In Study 1, we found that among individuals in the United States, social class was negatively related to percentage of friends on Facebook that are outside the United States. In Study 2, we extended these findings to the global level by analyzing country-level data on Facebook friends formed in 2011 (nearly 50 billion friendships) across 187 countries. We found that people from higher social class countries (as indexed by GDP per capita) had lower levels of internationalism-that is, they made more friendships domestically than abroad. --------------------------------------------------------------------------------","On the one hand, studies suggest that people from high social class groups are action- oriented in connecting with others (Keltner, Gruenfeld, & Anderson, 2003). At the individual level of analysis, high-social class individuals tend to feel more positive emotion, send out more approach-related signals (such as smiles or friendly eye contact), and approach others to the extent that they can be useful to fulfilling their needs and/or goals (Gruenfeld, Inesi, Magee & Galinsky, 2008). Select studies find that, depending on situational demands, people from high-social class groups tend to take responsibility and be more inclined to assist low-social class members (Brewer, 1988; Keltner, Gruenfeld, Galinsky, & Kraus, 2010; Overbeck & Park, 2001). In light of these processes, one might expect upper class individuals to form more friendships across national boundaries (Maddux & Galinsky, 2009). Objective conditions of the lives of upper class individuals—where they work, travel to, and are educated—would make it reasonable to predict that they will have more international friends. These findings and reasoning converge on the building social class hypothesis: upper class individuals and people from high-social class countries build up their international social capital through the formation of more international friendships relative to their lower class counterparts. A competing hypothesis is found in recent analyses of status, wealth, and social class (Kraus et al., 2012; Vohs, Mead, & Goode, 2006). This line of reasoning holds that high social class individuals are endowed with greater resources, and therefore less dependent upon others. As a result, the wealthy tend to be less socially engaged with others, in particular those from different groups than their own. In keeping with this restriction of social capital perspective, studies have found that higher class individuals show higher patterns of nonverbal social disengagement (e.g. doodling) compared to individuals from lower class backgrounds (Kraus & Keltner, 2009), they prove to be less responsive to others' suffering, and they tend to share less with others (Piff, Kraus, Côté, Cheng, & Keltner, 2010; Piff, Stancato, Côté, Mendoza-Denton, & Keltner, 2012; Stellar, Manzo, Kraus, & Keltner, 2012). By contrast, lower class individuals have perhaps more to gain from diversifying their social connections and prove to be more oriented towards reaching out and connecting with others (Piff et al., 2010). These findings and theoretical analysis lend themselves to a competing hypothesis that we tested in this investigation, the restricting social class hypothesis: upper class individuals will form fewer international friendships than low- social class people.","In the present research, we tested these two competing hypotheses about who is likely to form friendships with people from different countries than one's own. We did so in two complementary studies, one at the individual level of analysis, and a second at the national level of analysis. In Study 1, we tested the relationship between personal social class (income and social class) among individuals in the United States and the percentage of their total Facebook friends (a proxy of social relationships) that came from outside the United States. In Study 2, we examined the relationship between every friendship made on Facebook in 2011 and GDP per capita (as a proxy of social class). Several aspects of the Facebook platform allow us to overcome certain classic challenges in social sciences. First, Facebook's user base is massive, spanning over 1.3 billion users; thus, in our second study, our findings provide insights based on data from every corner of the earth and most walks of life. With growing concerns about the robustness of findings based on Western, educated, student samples—indeed, even cognitive psychology findings vary massively across cultures—such a huge, culturally, ethnically, and class mixed sample provides a vital step forward to generalizing effects (Henrich, Heine, & Norenzayan, 2010). Second, Facebook friendship structure mirrors real-world friendships (Wilson, Gosling, & Graham, 2012). In fact, unlike other social networking sites, real-world friendships tend to be a precursor to becoming Facebook friends (Ross et al., 2009). These findings highlight the validity of using Facebook friendships as a proxy for a person's social contacts. However, even given these findings, we suggest that Facebook friendships should ultimately be treated as a proxy for real relationships as there are almost certainly social ties on Facebook that are with individuals that users have only met once or not at all. Facebook friendships provide a convenient way to approximate a person's social sphere, but there is some error in this metric. Similarly, income, social class, and GDP per capita are powerful proxies for individual and national social class given the central role that money plays in people's determinations of social class (Kraus, Piff, & Keltner, 2011). Third, using Facebook friendships allow us to quantify the percentage of a person's friends that are international without relying on self-report—thus, many classic biases are not threats to the interpretation of the findings. Indeed, it would be exceedingly difficult to quantify the number of international social ties a person has since—by virtue of them being international—the individual is unlikely to see them often and thus more likely to forget them when asked to make a list of friends. Our research makes two principal contributions to the intergroup and social class literatures. By demonstrating how social class underpins the creation of cross-national friendships, we shed light on how people reach across social divides and form connections to disparate others, a phenomenon we know to be important for cultural change, increased chances of innovation, and less hostile intergroup attitudes (Maddux & Galinsky, 2009; Pettigrew & Tropp, 2008). In addition to underscoring how social class is a major driver of these cross-national friendships, we also provide support for the idea that, despite status being beneficial in many aspects of everyday life, it affects the composition of social networks in a way that reduces international diversity. Thus, our research strengthens and enhances a budding line of evidence (Kraus et al., 2012; Piff et al., 2012) that social class carries certain risks as well as advantages.","We recruited 1069 individuals from Amazon's Mechanical Turk who lived in the United States to participate in the study in exchange for $1.00. At the beginning of the study, participants completed a consent form which detailed all parts of the study. No deception was used. Of these individuals, 857 participants consented to authorizing our Facebook app to gather some information from their profiles automatically. This information included their total number of friends and their friends' current location—we note that friend location data was not available for all friends, but it was available for the majority of friends. Facebook researchers were not directly involved in Study 1—the data collection was done independently through the Cambridge team's own app. We focused only on the participants who had at least one Facebook friend (sample N = 815). These participants had on average 353 friends, and all together had 287,739 friends. For each participant, we examined their friends' current location, calculating the percentage of friends who lived outside the United States, which served as our metric of internationalism. In examining the internationalism histogram, we found extreme positive skew: the vast majority of participants had a small percentage of international friends (on average, 4%), with a small minority having high levels of international friendships. Since such outliers can extremely skew the results of a regression model, we followed the recommendations of Tabachnick and Fidell (2006), z-scoring all internationalism values, excluding any values more than 3 SDs away from the mean, z-scoring the remaining values, excluding any values of the new z-scores that are 3 SDs away from the mean, and so forth until the z-scores revealed no scores more than 3 SDs out. This procedure left 671 individuals for analyses (mean age = 28.6, sd = 9.2; 54% female), with internationalism scores ranging from 0 to 14% (mean = 4%, SD = 3%). Since these scores were still skewed positively, we natural log transformed the data, which created an approximately normal distribution of results. Importantly, we note that we ran all models without excluding the outliers, instead simply natural log transforming the data to account for skew—and all results in that procedure were consistent with the reported findings below. We also ran all models using a Poisson regression without removing any outliers, as a further alternative to taking the natural log, and again found the same pattern of results highly significant. We therefore present the results yielded by the first approach since it offers the best interpretability and protection from bias due to outliers, but note the conclusions of our paper are not dependent on the chosen method of analysis and/or outlier removal. Self-report measures Participants responded to two items aimed to measure their social class: (a) income, “What is your total annual household income?” (1 = Under $10,000, 11 = Above $100,000) and (b) social class, “Where would you place yourself on the following spectrum for social class?” (1 = Working class, 5 = Upper class). Participants also indicated how long they have been members of Facebook (months, years) and how frequently they use it (1 = Never; 5 = Several times a day). While there was some skew towards lower values for both these metrics, the skew was below 1 for both. Furthermore, when conducting all analyses using the logged versions of the variables, we found the same pattern of results. Thus, we have used the non-logged versions of the variables for ease of interpretation.","We included several key control variables in all models that could potentially suppress or confound the association between social class and internationalism. First, we controlled for age, because we found that older individuals tend to have higher social class and use Facebook less in our data. Second, gender is related to both social engagement (women slightly more likely to have more friends) and social class (trending towards lower income), so we controlled for this variable as well. Third, Facebook users with more friends had more international friends in our data—likely because they have larger social networks. Thus, in all our models, we controlled for participant age, gender, natural log number of total Facebook friends, length of using Facebook, and self-reported frequency of using Facebook. We note that all effects held even without controlling for age or gender. However, when removing all Facebook control variables, we did not find significant effects. This is to be expected, however, as Facebook can only function as an effective proxy of social relationship for people who actually use Facebook; therefore, usage rates need to be accounted for. We set up three regression models to test the building and restricting social class hypotheses—which yielded contrasting predictions concerning whether social class is positively or negatively predictive of internationalism. In each model, we tested a different social class indicator to establish robustness of our effects—specifically, in Model 1, we used income as the predictor; in Model 2, we used self-reported social class as the predictor; and finally, in Model 3, we used a composite of income and social class as the predictor. We did this because the two indicators of social class often only moderately correlate with one another, and can be thought of as objective and subjective forms of social class (Kraus et al., 2012). We found that higher income, b = − .05, CI95(− .08, − .02), rpartial = − .13, t(477) = − 2.96, p < .01, self- reported social class, b = − .18, CI95(− .29, − .07), rpartial = − .15, t(476) = − 3.33, p < .01, and the composite of the two, b = − .09, CI95(− .15, − .04), rpartial = − .15, t(479) = − 3.37, p < .01, were each negatively related to internationalism. In examining internationalism percentages between low (− 1SD social class) and high-social class (+ 1SD social class) individuals, the models indicate that low-social class people have nearly 50% more international friends (2.9% internationalism) than high-social class people (2.0% internationalism). Thus, these results provide support for the restricting social class hypothesis. Fig. 1 shows two world maps: (a) percentage of internationalism on Facebook and (b) GDP per capita. As these maps suggest, and in keeping with the individual level results from Study 1, there is a negative correlation between GDP per capita (national social class) and percentage of Facebook friends that are foreign in a nation (internationalismlog), r = − .18, CI95(− .32, − .04), t(172) = − 2.44, p = .02, thus providing support on a global level for the restricting social class hypothesis. Examining low-social class (− 1SD GDP per capita) and high-social class (+ 1SD GDP per capita) countries, people from low-social class countries had on average 35% of their friendships be international while people from high social class countries on average had 28% of their friendship be international. We also considered potentially controlling for variables that could influence the relationship. However, the difficulty with these variables (i.e., percentage of population with internet, rate of travel, immigration) is that they tend to be heavily correlated with GDP per capita. For instance, GDP per capita and total friendships per capita are correlated at r = .73. Thus, including them both in a model would greatly change the meaning of both variables, and in turn change the meaning of any potential correlation. Given these issues—and also the convergence of the global correlation with the individual level correlations presented in Study 1 with a full set of controls—we felt that the most prudent approach is to keep the model as a simple correlation between GDP per capita and internationalism without adding controls. However, net migration proved an exception to this because (a) it had a relatively moderate correlation with GDP per capita (r = .49) and (b) it represented a potential mediation mechanism of our effects. We therefore ran a regression model controlling for net migration. We did not find evidence to support a relationship between net migration and internationalism, b = .04, CI95(− .03, .12), rpartial = .10, t(172) = 1.27, p = .21; in contrast, the effect of GDP per capita not only remained significant and negative, but in fact increased in magnitude, b = − .03, CI95(− .05, − .01), rpartial = − .25, t(172) = − 3.34, p < .01.","In Study 1, we showed that among individuals living in the United States, self-reported social class is negatively related to internationalism, thus providing support for the restricting social class hypothesis for personal social class. In Study 2, we examined how GDP per capita predicts the percentage of Facebook friendships that are international in each nation in the world. In transitioning to data at the national level, we therefore tested the two competing hypotheses across cultures that vary dramatically in terms of their social values, economic development, religion, political organization, and self- construals, thus allaying some concerns about potential biases of Western samples (Henrich et al., 2010). It's important to note that Study 2 does not intend to extend the insights of Study 1 at an individual level—instead, Study 2 focuses on macro-level, cultural effects, and thus it would be premature to take its results as equivalent to conducting individual level studies around the world to provide universal evidence for our Study 1 individual-level effect.","Facebook provided data on every friendship formed in 2011 in every country in the world at the national aggregate level. These data set included a total of 57,457,192,520 friendships. From these data, we knew how many friendships were made within each country (domestic friendships) and also how many friendships were made between every country pair (international friendships). We note that we did not receive data on individuals—but rather only the macro-level, national sums. While most countries had at least tens of millions of friendships formed, a small number had relatively few. Thus, we removed from our analyses nations with less than 1 million friendships. This left 204 countries in the sample. We quantified social class as GDP per capita in 2011 (World Bank, 2012). Full data across all variables was available for 175 countries which make up the majority of the global population (5,291,704,711) and Facebook's user base (952,818,100). For these countries, we examined every domestic (a total of 48,458,812,050 friendships) and international friendship (a total of 7,572,368,093 friendships) made on Facebook in 2011. For each country, we calculated an internationalism score by dividing total international friendships by total friendships (2 ∗ domestic + international)—we note that domestic friendships were multiplied by 2 because every international friendship is counted twice (once for each nation part of the friendship) and to keep the percentages consistent, domestic friendships need to be counted twice (once for each person who can claim the friendship in the nation). Since GDP per capita was highly positively skewed, we normalized it by taking the natural log. We also included net migration per capita (latest figures for each country from World Bank) as a control to account for the possibility of migration explaining our effects. There was substantial skew in net migration; we therefore logged the scores to make them more normally distributed.","Friendships can provide wide-ranging benefits and can be closely related to health, trust and well-being. The barriers to forming friendships with individuals from groups that differ from one's own are well-known, and include in group bias, intergroup anxiety, and prejudice (Pettigrew & Tropp, 2008). In the present investigation, we relied on individual and national level data to ask how social class, both objective (income) and subjective (perceived social class), personal (individual level) and national (nation level), influence the likelihood of forming friendships on Facebook with individuals from different countries. The literature on power and social class justified two competing predictions. Given the association between power and social approach (Gruenfeld, Inesi, Magee, & Galinsky, 2008; Magee & Galinsky, 2008), one might expect the wealthy to reach out and form more friendships with people from different groups—the building social class hypothesis. In contrast, given the tendency for the upper class to be socially disengaged with other individuals (Hall, Coats, & LeBeau, 2005; Kraus & Keltner, 2009), and the well documented tendency for lower class individuals to be more prosocially oriented to others (Piff et al., 2010), one might expect the wealthy and higher social class to privilege friendships within one's own group, and not have as many international friends—the restricting social class hypothesis. In line with the restricting social class hypothesis, our results across two studies at the micro-individual and macro-national levels showed that people and nations with greater objective and subjective social class had fewer international friends on Facebook. In Study 1, people who reported that they had higher income and were from upper class categories had fewer international friends than individuals who had lower incomes and self-identified with lower class categories. In Study 2, wealthier countries, as indexed in per capita GDP, had fewer international friends than poorer countries. Despite having fewer resources at their disposal, it is actually low-social class people and poorer nations who tend to have friendship networks populated with more international connections. These findings make a number of important contributions to the social class and social network literatures. First, the negative link between social class and international friendships may underscore a tendency for high- social class people to accrue relational resources such as social support in their local vicinity. According to the self-reinforcing account of social class (Magee & Galinsky, 2008), high-social class people tend to think and act in ways that reinforce their own social class. However, by forming fewer international friendships (relative to their low- social class counterparts), high-social class people may be at risk of strengthening their local social class at the expense of improving their broader social class. Second, theory and research highlight the value of developing weak ties to others in distant social circles because these ties offer access to resources not likely to be found in one's immediate social circle (Agrawal, Kapur, & McHale, 2008; Burt, 2008; Granovetter, 2005). An encouraging sign is that low-social class people tend to have greater access to these resources on account of having more international friendships. Thus, contrary to research describing the plight of low-social class individuals with poor access to social resources (Magee & Galinsky, 2008), our findings point to international friendships as a means by which low-social class individuals can benefit from the resources inhering in social ties. Third, our study showcases a new approach to studying internationalism at the global level—through the usage of nation-level social media. This approach opens the doors for new investigations of the micro- and macro-forces that promote or inhibit internationalism. One particularly fruitful avenue is to study how cultural differences—on dimensions such as individualism-collectivism—are related to internationalism. Several datasets are available at the nation level that have cross-cultural scores, including Hofstede's Cultural Dimensions (Hofstede & Bond, 1984), Schwartz Cultural Values (Schwartz & Bilsky, 1990), and the World Values Cultural Dimensions (Minkov & Hofstede, 2010). These scores can be used to probe the relationship between culture and internationalism in future work. We do note that there are several important limitations to our work which require important future extensions. First, both of our studies were correlational, and thus making strong causal claims at this stage is premature. While we believe that it is most likely the causal arrow that flows from social class to internationalism, explicit experimental work is needed to test this thesis. Second, our sample in Study 1 consisted of MTurk users which are not representative of the general US population. While MTurk users tend to be as good—if not better—than typical undergraduate samples (Buhrmester, Kwang, & Gosling, 2011), future work should endeavor to recruit a nationally representative sample to further corroborate the results. Third, while we believe that Facebook friendships offer a powerful proxy for studying social relationships—and previous work does support the notion that Facebook friendships tend to mirror real world relationships (Wilson et al., 2012)—further replication of the findings using traditional self-report metrics of friendships would be a welcomed further validation of our approach. Fourth, we do not know the place of birth of participants in Study 1; thus, it is possible that migration flows could be a potential mechanism behind this effect. Future work should explore this issue in greater detail. However, we do note that we found evidence against a migration explanation in Study 2. Fifth, the effect sizes were of a small-to-moderate nature in both Studies 1 and 2. This is not surprising as there are likely a large multiplicity of factors involved in internationalism; however, it does highlight that we should be careful in over-interpreting the effects and their suggested consequences. In a world with more global connections than ever, some individuals are creating more international ties than others. Despite the benefits of having international connections, and the fact that high-social class people should be better positioned to travel and meet people from different countries, our results provide one empirical demonstration as to how low-social class people may actually stand to benefit most from a highly international and globalized social world."],["Testing infants in the laboratory is expensive in time and money; consequently, many studies are underpowered, reducing their reproducibility. We investigated whether the online platform, Amazon Mechanical Turk (MTurk), could be used as a resource to more easily recruit and measure the behavior of infant populations. Using a looking time paradigm, with users’ webcams we recorded how long infants aged 5 to 8 months attended while viewing children's television programs. We found that infants (N = 57) were more reliably engaged by some movies than by others and that the most engaging movies could maintain attention for approximately 70% of a 10- to 13-min period. We then identified the cinematic features within the movies. Faces, singing-and-rhyming, and camera zooms were found to increase infant attention. Together, we established that MTurk can be used as a rapid tool for effectively recruiting and testing infants. --------------------------------------------------------------------------------","Infants are difficult to recruit and test. Recruiting their busy caregivers requires broad advertising, collaboration with day-care facilities or maternity hospitals, and labor- intensive relationship building. As a consequence, it often takes long periods of time to recruit a sufficient number of participants. Once recruited, the schedules of the caregivers, infants, research staff members, and testing facilities must be coordinated, and practicalities such as transport must be resolved. Given these complexities, infant studies are relatively slow and expensive and also require patience and perseverance. This puts pressure on investigators to minimize the number of participants recruited, and as a result studies are sometimes under-powered, reducing their reproducibility (Peterson, 2016), reflecting a broader issue in psychology (Open Science Collaboration, 2015). Thus, to make it easier to conduct high-quality infant research, it is imperative to find ways in which to reduce these pressures. One solution may lie online. During recent years, the crowdsourcing engine Amazon Mechanical Turk (MTurk) has become a central marketplace, bringing together hundreds of thousands of workers from more than 100 countries to complete “human intelligence tasks” (HITs) through a web browser for modest remuneration (see Crump, McDonnell, & Gureckis, 2013, for a review; Buhrmester, Kwang, & Gosling, 2011; Kittur, Chi, & Suh, 2008; Mason & Suri, 2011; Pontin, 2007). Often these tasks involve image annotation, rating surveys, and demographic questionnaires using templates delivered by MTurk (Mason & Suri, 2011; Paolacci, Chandler, & Ipeirotis, 2010). However, by employing external websites, requestors can generate more complex tasks that meet the demands of their experimental needs (Buhrmester et al., 2011; Goodman, Cryder, & Cheema, 2013; Mason & Suri, 2011). In combination with MTurk’s simple interface and flexibility, this lends itself well to the fast and cost-effective collection of data. Recently, experimental psychologists have used MTurk to obtain data from adults on simple tasks (Lewis, Sugarman, & Frank, 2014; Piff, Stancato, Côté, Mendoza-Denton, & Keltner, 2012; Starmans & Bloom, 2012; Sweeny, Andrews, Nelson, & Robbins, 2015). Parental report measures of child behavior have also been documented (Schneider, Yurovsky, & Frank, 2015); however, to our knowledge, direct testing of infants has never been attempted. Therefore, the aim of the current study was to examine whether MTurk could be used to recruit and test infant populations. To address this, we implemented a task that aimed to quantify looking time to a set of video stimuli in infants aged 5 to 8 months. Specifically, infants viewed children’s television programs, and attention was quantified by measuring when infants fixated on the screen using their webcams.","Ethical approval was obtained from Western University’s health sciences research ethics board. We recruited infants aged 5 to 8 months using MTurk (Amazon, Seattle, WA, USA). All workers of MTurk remain de-identified and are referred to only by a unique worker identity code provided by Amazon. To participate, infant caregivers agreed to the Amazon MTurk Participation Agreement (https://www.mturk.com/mturk/conditionsofuse), which included the declaration that they were at least 18 years old, provided informed consent, and were required to have a webcam, speakers, and Adobe Flash. The same experiment was administered in two independent batches differing in compensation rate. During the first batch, 63 participants were recruited over the course of a week, reimbursing each with $1.25 (U.S.). To motivate increased participation, remuneration in the second batch was raised to $5.00, leading to 84 participants being recruited within the following 6 days. Altogether, 147 participants were recruited. However, due to the quality control requirements of our study and the difficulty of infant testing in general, 90 were excluded. The causes were technical issues associated with internet connectivity bandwidth (no webcam video being obtained from the server (n = 3), the webcam video becoming desynchronized from the movie (n = 43), the quality of the recorded video (infant’s eyes not being visible; n = 22), blurry video (n = 5), and the location of the webcam being changed (n = 1)). Participants were also excluded if they were not infants (n = 11), self-reported an age outside our specifications (n = 3), or did not fully complete the experiment (n = 2). Of the remaining 57 participants included in the study (Mage = 6.49 months, SD = 0.93), 9 were 5 months old, 19 were 6 months old, 21 were 7 months old, and 8 were 8 months old. Stimuli and video recording ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Ten movie clips between 9 and 13 min in length were used as stimuli. They were taken from popular programs designed to appeal to infants and children: Baby Einstein, Blue’s Clues, Curious George, Despicable Me, Dora the Explorer, The Program With the Mouse, In the Night Garden, Teletubbies, Timmy Time, and Up. To present the movies and record from the webcam, Flash was used (Adobe, San Jose, CA, USA). Content was provided from and recorded to a computer in the Amazon cloud (Amazon) running Wowza Media Streaming Server (Wowza, Golden, CO, USA) using real-time streaming protocols. Our software is available on request.","The HIT was created and posted on MTurk under the title “Infant Television Viewing.” Participants viewed a webpage detailing compensation rate, the allotted time to complete the HIT, the expiration date of the HIT, a short description of the task, and the required qualifications. After accepting the HIT, participants were directed to a webpage that provided an information sheet and asked for their informed consent. Consent was obtained via online checkbox and button press. This was followed by an evaluation of the suitability of participants’ computers, software, webcams, speakers, and internet connectivity. To do this, we recorded a brief 5-s video of caregivers and their infants and asked them to move and make sounds. This video was then played back to caregivers, and they were asked to check a box to indicate whether or not they were able to see and hear themselves. If they indicated they could not, they were thanked and excluded from participation. Otherwise, they were directed to a new webpage instructing them to position their infants on their laps in the center of the screen and in a well-lit room. This was specified to ensure that infants’ eyes were visible while recording the webcam videos. Once they were in a comfortable position, caregivers were instructed to press a “start” button to commence the experiment. Then 1 of 10 pseudorandomly selected movies was presented. Afterward, participants completed a short demographic questionnaire from which the ages and language backgrounds of their infants were obtained. Video annotation ~~~~~~~~~~~~~~~~ The time course of overt attention was measured from the recorded webcam videos. Using the free Anvil tool (Kipp, 2001), the experimenter and a second observer annotated when in the movie each infant fixated on the screen. The two observers agreed in their assessment of looking time 99.81% of the time with a corrected kappa of .41. Are some of the movies more engaging overall? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We first quantified overall looking time by calculating the proportion of attention in a movie for each infant. This proportion was then arcsine transformed to increase the normality of its distribution. To assess whether some movies were more engaging than others, we used a two-way analysis of covariance (ANCOVA) with factors of movie (10 levels) and infant age (in months). Post hoc t tests were then used to compare pairs of movies. Fig. 1 and Table 1 show arcsine-transformed proportions of attention by movie. There was a main effect of movie on attention time, F(9, 25) = 3.56, p < .001, η2 = .562 (Fig. 1). Post hoc comparisons using Tukey’s HSD (honestly significant difference) revealed that Curious George (M = .33, SE = .09) engaged infants significantly less than In the Night Garden (M = .74, SE = .06, p < .05), Teletubbies (M = .70, SE = .09, p < .05), and Timmy Time (M = .68, SE = .04, p < .05). Age was not found to modulate looking time, F(3, 25) = 2.74, ns, η2 = .248. Are some parts of the movies more engaging to infants than others? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To determine whether the looking times were consistent within a movie, we conducted a correlation analysis that assessed the time course of attention. Our hypothesis was that if particular parts of a given movie were more engaging than others, greater similarity would be seen for the time courses of attention from infants viewing the same movie than from those viewing different movies. A “binning” technique was applied in which each time series was divided into 0.5-s intervals and assigned a 0 (not looking) or 1 (looking). Due to the movies having different run times, we restricted the analysis of the time-series data to the duration of the shortest movie (552.50 s or 1105 bins). To quantify the similarity, the Pearson correlation between the time course of each participant and each other participant was computed. A correlation of 1 would reflect identical time courses, and a correlation of 0 would represent no correspondence. We tested the hypothesis that infants watching the same movie should have a more similar time course than infants watching different movies. To do this, the mean of within-movie correlations was calculated. This was then tested against a null distribution calculated by bootstrapping, randomly shuffling the matrix of pairwise Pearson correlations of participants and taking the mean of positions in the matrix that previously held within-movie comparisons. The process was repeated 100,000 times to build the null distribution. The proportion of null values that were greater than the true value was taken as the p statistic. Fig. 2 shows the binned time courses of attention for each of the infants grouped by movie. By eye, there do appear to be some places of some movies where infants start or stop paying attention together. However, not surprisingly, there are also large individual differences given that infants may have varying stimulus preferences or be spontaneously thinking of different things. Therefore, we evaluated statistically whether infant attention was modulated by the content of the movies. The correlation in the time course of attention across participants watching the same movie was positive but weak (r = .07). However, this was highly significant when compared with the null distribution calculated through bootstrapping (Fig. 3; p < .001), confirming that the movie content modulated infant attention. What features make the movies more engaging? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ For the five movies for which cinematic annotations were available (Fig. 4), we used linear regression to investigate what drove infant attention. Because no effect of age was seen in the overall looking time, we collapsed across ages and grouped infants by the movie they viewed (n = 28). From the 10 annotated features (Table 2), it was found that singing-and-rhyming, t(27) = 2.60, p < .05, camera zooms, t(27) = 2.16, p < .05, and faces, t(27) = 2.98, p < .01, significantly increased movie engagement (Fig. 5).","Our exploratory study supports the principle that MTurk can be used as an efficient tool to recruit infants. In a looking time paradigm, we showed that infants aged 5 to 8 months were engaged by different child-directed movies more so than others, that some parts of the movies were more engaging than others, and that the cinematic features of faces, singing-and-rhyming, and camera zooms within the movie increased attention. No age differences in attention were found. Most important, we report for the first time an online experiment capable of capturing and quantifying infant behavior directly. Our finding that infants demonstrate preference for faces and face-like stimuli even in the midst of distracters and dynamic visual scenes is supported by a substantial existing literature (Di Giorgio, Turati, Altoè, & Simion, 2012; Farroni et al., 2005; Franchak, Heeger, Hasson, & Adolph, 2015; Frank, Vul, & Johnson, 2009). Similarly, our finding that singing-and-rhyming attracts attention concurs with previous reports indicating preferences for engaging melodies over dissonant sounds (Costa-Giomi & Ilari, 2014; Nakata & Trehub, 2004; Trainor, 1996; Trainor, Tsang, & Cheung, 2002). High engagement toward both of these cinematic elements has been associated with internal biases reflective of maternal behavior and infants’ keen interest to attend to stimuli rich in social information (Bushnell, Sai, & Mullin, 1989; Nakata & Trehub, 2004). In addition, camera zooms also recruited increased visual attention. This agrees with recent evidence that optic flow—structured patterns of motion across the visual field—is processed in the brains of children aged 4 to 8 years (Gilmore, Thomas, & Fesi, 2016). Furthermore, even neonates have mechanisms that can quantify the degree of optic flow (Jouen, Lepecq, Gapenne, & Bertenthal, 2000). Regarding the more specific effect of optic flow on attention, an existing report found that preschool-aged children and toddlers demonstrated unchanged or reduced visual attention when camera zooms were present compared with when they were absent (Levin & Anderson, 1976; Susman, 1978). These authors suggested that camera zooms deter attention because of their tendency to disrupt the visual flow of content by taking the viewer from a whole perspective to a part perspective. The difference between these results and ours may be due to the different ages of the infants tested or to the very different stimuli. For example, the camera zooms in our stimuli served an artistic or communicative vision determined by the movie directors and, thus, may have been more congruent with the overall content in comparison with Susman’s (1978) study. MTurk recruitment was found to be easy and enabled us to collect a large infant data set in a relatively short period. Unlike in the laboratory where local participants are tested one at a time, MTurk permits workers from across the world to carry out the same HIT in parallel (Mason & Suri, 2011; Paolacci et al., 2010). The service is entirely online, which allows caregivers to recruit their children in the convenience of their own homes without needing to worry about the demands that a novel environment imposes on their children and restrictive participation time slots. Reflecting this increased convenience for participants, payments to participants were much lower than in typical studies. Although recruitment was easy, there was a trade-off in data quality. Due to our limited control in screening participants and their equipment online, many workers (∼40% of our data) needed to be excluded from our analyses. Although we implemented screening measures to constrain who could complete and view the HIT, the majority of exclusions were due to issues regarding internet connectivity rather than task performance. Therefore, although the internet is inherently involved in using MTurk, in order to enhance data quality when bi-directional video streaming is required, internet speeds could be better prescreened in future experiments. Imposing stricter screening procedures will likely improve data quality; however, given the low cost of recruiting participants, a feasible solution could be to just accept a substantial rejection rate. To maximize the potential of MTurk, it will be important to develop ways in which to eliminate as many interfering factors as possible. In future studies, it would be beneficial to gather additional information on the display configuration by querying information from the browser (e.g., the window size and screen resolution) and by asking users for information (e.g., the screen size, distance, and screen model). It might also establish, for example, whether a proxy for viewing distance can be obtained from infant face size, as recorded with the webcam. Ways in which to quantify lighting conditions and sound levels would also be advantageous. And further demographics, such as the sex and socioeconomic status of the infants, could be informative. Because the study took place in participants’ homes, the experimental context for each participant differed. We directed caregivers on how to seat their children for the experiment; however, we did not specify details relating to the environment in which it should be carried out. We recognized in parts of the webcam video that there were occasional distracters present in the room (e.g., toys, other people, telephone, television) that potentially could have added noise to our measures. In addition, because caregivers were aware of the movie content, they may have given unintentional cues to their infants even though caregivers were asked to remain still. Furthermore, we did not have control of the screen size or specify children’s viewing distance from the computer, likely affecting the stimulus visual angle and possibly adding further noise. Furthermore, we obtained a moderate kappa value (Viera & Garrett, 2005). Although such a value could be attributed to both coders being relatively new to video annotating, the study’s unconstrained viewing procedure could have made quantifying looking behavior more subjective than studies completed in the laboratory. It will be useful for future validations to include comparative measurements in a laboratory setting. These could not be conducted currently in our laboratory because it was winding down prior to a shift in location. Better understanding of the effect of online and offline contexts on infant behavior will enable us to elucidate the extent to which virtual studies produce similar outcomes to laboratory conditions. Emerging tools such as Lookit (https://lookit.mit.edu/) will be valuable in conducting these studies. Ideally, appropriate recording conditions would involve a well-lit environment where shadowing of the face is limited and where participants’ eyes are visible. In our task, we explicitly asked for infants to be in a well-lit room; however, we did not convey that infants’ eyes needed to be visible. Failure to do so may have resulted in the 22 participants being omitted from analysis. Therefore, in the future it would be beneficial for researchers to provide details pertinent to the experiment so that tasks are more likely to be carried out properly. Furthermore, this could potentially reduce rejection rates. In addition, online looking time paradigms should aim to have each participant’s face in the center of the screen. Our experiment asked parents to position their infants in the center of the webcam video to discern whether or not the infants were looking at the stimulus. However, what was “center” for one participant could have been definitively different for another participant. Due to the differences in screen and webcam parameters, viewing of the stimulus was likely idiosyncratic. As a solution, incorporation of facial recognition software could be implemented into the research design to standardize testing procedures and improve data quality overall. Moreover, in optimizing recording conditions, researchers should also specify that video viewing take place in an enclosed room limited in distractions. The opportunity to record with a webcam made employing a looking time paradigm possible. In the laboratory, looking time paradigms have been widely used to study a variety of domains in infants such as intentionality (Hamlin, Wynn, & Bloom, 2007; Woodward, 1998), emotion (LaBarbera, Izard, Vietze, & Parisi, 1976; Montague & Walker-Andrews, 2001), and speech preference (Cooper & Aslin, 1990; Maye, Werker, & Gerken, 2002). Having shown in this study that behavior can be easily captured, reviewed, and quantified, this opens up the possibility of putting similar paradigms and laboratory-based studies that use comparable equipment on MTurk. This is not to say that all studies can be employed online; some require specific equipment not readily available and, hence, require local testing; however, we suggest that select tasks have the potential of being administered online.","Our study demonstrates that MTurk is a powerful new tool for recruiting infant populations. We designed an online study that was capable of capturing infant behavior directly and, in particular, that informed us of stimulus features in movies by which infants were most engaged. Online infant testing could reduce the high costs experienced in running experiments in the laboratory, and by removing barriers to larger samples, this could lead to increasing data reproducibility."],["A diverse body of research has demonstrated that people update their beliefs to a greater extent when receiving good news compared to bad news. Recently, a paper by Shah et al. claimed that this asymmetry does not exist. Here we carefully examine the experiments and simulations described in Shah et al. and follow their analytic approach on our data sets. After correcting for confounds we identify in the experiments of Shah et al., an optimistic update bias for positive life events is revealed. Contrary to claims made by Shah et al., we observe that participants update their beliefs in a more Bayesian manner after receiving good news than bad. Finally, we show that the parameters Shah et al. pre-selected for simulations are at odds with participants’ data, making these simulations irrelevant to the question asked. Together this report makes a strong case for a true optimistic asymmetry in belief updating. --------------------------------------------------------------------------------","Numerous studies spanning behavioural economics (Eil & Rao, 2011; Krieger, Murray, Roberts, & Green, 2016; Krieger et al., 2014; Möbius, Niehaus, Niederle, & Rosenblat, 2012), psychology (Garrett & Sharot, 2014; Kuzmanovic, Jefferson, & Vogeley, 2015; Moutsiana et al., 2013) and neuroscience (Garrett et al., 2014; Korn, Prehn, Park, Walter, & Heekeren, 2012; Kuzmanovic, Jefferson, & Vogeley, 2016; Lefebvre, Lebreton, Meyniel, Bourgeois-Gironde, & Palminteri, 2016; Ma et al., 2016; Moutsiana, Charpentier, Garrett, Cohen, & Sharot, 2015; Sharot, Guitart-Masip, Korn, Chowdhury, & Dolan, 2012; Sharot, Korn, & Dolan, 2011; Sharot et al., 2012) have shown that people alter their beliefs to a greater extent in response to good news than bad news. This asymmetry can lead to a positive bias in beliefs regarding oneself, referred to as the superiority illusion (Hoorens, 1993; Kruger & Dunning, 1999; Svenson, 1981), and in beliefs regarding one’s future, referred to as unrealistic optimism (Calderon, 1993; Radcliffe & Klein, 2002; Shepperd, Grace, Cole, & Klein, 2005; Weinstein, 1980, for review see Sharot & Garrett, 2016). The latter phenomenon, first described by Neil Weinstein (Weinstein, 1980), has since been supported by a large body of evidence (Armor & Taylor, 2002; Regan, Snyder, & Kassin, 1995; Shepperd, Helweg-Larsen, & Ortega, 2003; Shepperd, Klein, Waters, & Weinstein, 2013; Weinstein, 1987; for a review see: Sharot, 2011, 2012). A recent study by Shah, Harris, Hahn, Catmur, and Bird (2016) has revisited the phenomenon of optimistic asymmetry in belief updating. The authors claim that an optimistic bias is unlikely, apart from very specific cases including sports fans and smokers. Whilst Shah et al. present empirical evidence from a number of experiments, findings across their studies are inconsistent and methodological issues lie at the heart of their experimental design. Nonetheless, the research questions they attempt to address are important and each constitute interesting tests of the robustness of optimistic updating. In this report, we carefully examine the work and reveal statistical, methodological and conceptual problems, which when corrected supports the robustness of the asymmetry in belief updating. In particular, the authors of that paper: (i) claim that not taking into account subjects’ beliefs regarding base rates can lead to misclassification of trials, which causes the update bias; (ii) claim to show biased updating in Bayesian agents; (iii) claim to empirically show pessimistically biased updating for positive life events. Whilst the report may sound compelling at first read, an informed examination of the evidence casts significant doubts on these claims. Specifically, in this report we detail that: (i) previous studies have shown that after taking into account subjects’ base rates, exactly as Shah et al. lobby for, an optimism update bias is still observed (Garrett & Sharot, 2014; Kuzmanovic et al., 2015); (ii) comparing human data to Bayesian agents shows that subjects are less Bayesian when updating their beliefs in response to bad news than good news, thus the optimistic update bias cannot be explained away as Bayesian. Shah et al.’s simulations of Bayesian agents rest on specific parameters that are incompatible with human data; (iii) Shah et al.’s finding of a pessimistic update bias for positive events is due to a confound in the set of stimuli they used. When avoiding this confound no pessimistic update bias is observed for positive events. Misclassification argument is of no empirical consequence ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In their paper (Shah et al., 2016), Shah et al. outline an issue that has been described and resolved in two previous papers (Garrett & Sharot, 2014; Kuzmanovic et al., 2015). The basic claim (outlined previously in Garrett & Sharot, 2014) is that participants may hold one estimate regarding their own likelihood of experiencing an event and another regarding the likelihood for someone like them.","Participants then receive information regarding the likelihood of “someone like them” to experience a specific event. Trials are classified according to whether the information is better or worse than the participant’s estimate of their own likelihood of experiencing that event. One can imagine a situation where a participant believes their own likelihood of being robbed, for example, is 10% but believes the likelihood for someone like them is 40%. They then learn that the likelihood for someone like them is in fact 30%. That trial will then be classified as a “bad news” trial (because 30% is worse than 10%) when in fact it should be classified as a “good news” trial (because 30% is better than 40%). This concern has been empirically tested and resolved before (Garrett & Sharot, 2014; Kuzmanovic et al., 2015). In two previous studies, participants’ estimates of base rates were elicited in addition to their self- estimates. Trials were then classified according to whether provided base rates were lower or higher than participants’ estimated base rates.1 In both studies, an optimism update bias persists even after this classification method is applied (Fig. 1). This is because only a small number of trials are in fact misclassified (Garrett & Sharot, 2014). Optimistic update bias cannot be explained away as Bayesian ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Previous studies have found that subjects are less Bayesian in updating their beliefs when receiving bad news than good news, consistent with the idea that asymmetric updating is not purely Bayesian (Eil & Rao, 2011; Möbius et al., 2012). Shah et al. test this same question using the belief update task for future life events. Their findings, across four studies, were inconclusive. We thus employed the same analysis carried out by Shah et al. on our previously published set of data from 32 participants (Garrett & Sharot, 2014), also classifying trials as advocated by Shah et al. (see Section 1.1). Our Materials and Methods have been reported in detail elsewhere (Garrett & Sharot, 2014) and we summarize them in Fig. 2. The calculations advocated by Shah et al. (2016), were used and summarized in Box 1. We found that comparing participants updates to that of a perfectly rational, Bayesian agent revealed that participants were less likely to update their beliefs in a Bayesian manner in response to bad news than good news (Fig. 3 shows the new analysis of previously published data (from Garrett & Sharot, 2014): t(31) = 2.53, p < 0.05, paired sample t-test). Similar findings, using the belief update task, were recently reported by Kuzmanovic and Rigoux (2016). Contrary to the claim of Shah et al., these results, as well as past ones (Eil & Rao, 2011; Möbius et al., 2012), support a true valence-dependent asymmetry in how humans update beliefs. Shah et al. also conduct a simulation using hypothetical data that consists of one specific base rate (30%) and 52 hypothetical agents, 93% of which are fixed to have likelihood ratios (LHR) less than 1. Likelihood ratios express how diagnostic the evidence available to a participant is. Shah et al. (2016), claim to show an optimistic update bias in the agents used in their simulation. However, the authors themselves note that if the likelihood ratios are greater than 1, the opposite pattern ought to ensue - greater updating for bad news compared to good news. If likelihood ratios are close to 1, no bias will emerge in either direction in Bayesian agents. This raises the question: what are the likelihood ratios that human participants use in this task? Using the same set of data we reported above (published previously in Garrett & Sharot, 2014), we found that the mean likelihood ratio across participants was 1.12 (s.d.: 0.43), which was not significantly different than 1: t(32) = 1.63, p = 0.11, one sample t-test against test value of 1. A separate data set from an independent research group found that likelihood ratios were 1.55 on average (Kuzmanovic & Rigoux, 2016). Despite having the empirical data to derive likelihood ratios from their own experiments, Shah et al. select to fix the vast majority of likelihood ratios in their simulation to those that are less than 1 (mean likelihood ratio = 0.70, significantly less than 1: t(51) = 9.66, p < 0.001, one sample ttest against test value of 1). This assumption, however, is not supported by the data. These simulations are thus irrelevant to the question of whether the optimistic update bias observed in participants is Bayesian.2 It is worth noting that following the equation in Box 1, on each individual trial the value of the LHR (less than 1, equal to 1, larger than 1) indicates whether a subject believes their own risk is lower/equal/higher than someone like them. However, this logic does not hold for average estimates. That is to say, it is mathematically possible for the average risk estimates of a subject to be lower for themselves than others, yet for the average LHR to be equal or higher than 1. This is indeed the case in Garrett and Sharot (2014). Seemingly pessimistic update for positive events is misleading ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ An optimistic update bias for positive stimuli has been described previously (Krieger et al., 2014; Wiswall & Zafar, 2015). Yet Shah et al. (2016), report finding an optimistic asymmetry for negative life events but a pessimistic asymmetry for positive life events. Their studies, however, fall into two methodological pitfalls that need to be guarded against when studying belief updating. We outline these below and then proceed to present the methodology and results of a new study that corrects for these pitfalls. Our new study uses both positive and negative life events but takes care to avoid the pitfalls we outline. Hence it constitutes a valid test of belief updating using both types of life events. We find an optimistic update bias for both negative life events and positive life events. METHODOLOGICAL PITFALL I: OBTAINING UNRELIABLE, MEANINGLESS, STATISTICS FOR POSITIVE EVENTS -------------------------------------------------------------------------------- In the original belief update task (Garrett et al., 2014; Moutsiana et al., 2013, 2015; Sharot et al., 2011; Sharot, Guitart-Masip et al., 2012; Sharot, Kanai et al., 2012), participants are asked to estimate their likelihood of experiencing 80 aversive events in their lifetime (first estimates). They are then presented with the likelihood of these events in their population (information) and subsequently asked to estimate their likelihoods again (second estimate). Trials are then divided into ones where participants received good news (they learn that an aversive event is less likely than they thought) and trials where participants received bad news (they learn an aversive event is more likely than they thought). Update is calculated as the difference between the first and second estimate. When attempting to adapt this task to study positive life events, researchers face potential confounds which if ignored, will lead to invalid conclusions. In particular, whilst validated statistics regarding the likelihood of encountering negative events in one’s life-time are well documented (such as the likelihood of suffering different type of illness or being a victim of crime), statistics about the occurrence of positive life events are not readily available. This is problematic, as the belief update task requires the use of many trials and stimuli. Yet, it is practically impossible to find even 40 positive life events accompanied by validated statistics. Shah et al. attempt to circumvent this problem by making up statistics for positive events whilst using real, valid, statistics for negative events (see Appendix A of Shah et al., 2016). This systematic asymmetry in the validity of base rates for positive and negative base rates presents a possible confound. For example, for positive events, participants were asked the following questions (try and answer these yourself): How likely are you in your lifetime to: ‘Attend a friend’s birthday party?’ ‘Eat at your favorite restaurant?’ ‘Receive a present?’ ‘Have family visit you at Christmas?’ ‘Being told you are special?’ Most people will have all these events happen to them over a lifetime, thus they are likely to enter estimates close to 100%. This makes it impossible to measure update in a desirable direction (people cannot increase estimates beyond 100%). Participants are likely aware that actual statistics for these questions are less likely to exist than for negative events such as robbery or lung cancer. Note that not all the positive events in Shah et al. are extremely likely, which may explain why mean first estimates reported from their studies are below 50%. Here, to generate meaningful stimuli for both positive and negative life events we altered the belief update task as follows: we asked participants to estimate their likelihood of encountering everyday life events in the upcoming month. First, we obtained the frequencies of such events by asking over 200 participants to report whether different common positive and negative life events occurred to them at least once in the past month. We then used this data to construct a list of base rates for each event (i.e. the likelihood of each life event occurring at least once in a given month in the sample). Next we ran the belief update task on an alternate, but demographically well-matched, set of participants, asking them to estimate their likelihood of encountering these events in the next month. METHODOLOGICAL PITFALL II: SKEWING THE DISTRIBUTION OF BASE RATE ARTIFICIALLY PRODUCES A FLIP FOR POSITIVE EVENTS -------------------------------------------------------------------------------- When investigating biases in belief updating it is important to use a list of base rates that are normally distributed around a mean that sits mid-scale. For example, if a rating scale ranges from 5% to 95% the ideal mean would be 50% and the set of base rates should be normally distributed around this mean. Running simulations we show below that failing to ensure this creates spurious biases in updating (unless one controls for “estimation errors”, see below). Consider four lists of base rates for life events, the distributions of which are shown in Fig. 4a and b. The first two lists are the actual ones used by Shah et al. in two of their experiments (experiments 3A and 3B). The distribution of each set of these base rates is positively skewed (i.e. the majority of base rates are rare and fall below the midpoint of 50%) for positive and negative life events (Fig. 4a). The third and fourth lists are the actual base rates we used in this study. They are normally distributed around 50%, (Fig. 4b). For each simulation we will randomly generate a first estimate for each trial for each “participant”. This will be a random integer drawn from a uniform discrete distribution between 0 and 100 for the first two lists of base rates (this is the range of estimates participants are permitted to enter in Shah et al.) or between 5 and 95 for the third and fourth lists of base rates (this is the range of estimates permitted in our experiment). We will then “present” our simulator with the actual base rate for the event on that trial (information) and generate a second estimate - a random integer between the first estimate and the information (also drawn from a uniform distribution). For example, for the question “how likely are you to go out of town for leisure in the upcoming month” our simulator may randomly select 10% (first estimate), it will then observe a base rate of 36% (information) and adjust its answer to a random number between 10 and 36, let’s say 30% (second estimate). Thus, the amount of update on this trial would be 20 (update is calculated such that positive numbers always indicate a move towards the base rate). We run 1000 such simulations (i.e. “experiments”) for each set of base rates (Fig. 4a and b). If the data produced by a simulation results in biased updating, this will indicate that the bias is due to a statistical artifact (the result of the mathematical constraints of the task) and not to an asymmetry in human learning. If however the simulation produces no bias in updating but a bias is observed for human data, this would suggest that the bias is due to asymmetric learning not to a statistical artifact. Our simulation clearly shows that when the distribution of base rates is skewed, an artificial bias in belief updating is observed (Fig. 4c). This is not the case however when the base rates are normally distributed (Fig. 4d). Importantly, when an artificial bias is detected it is observed in opposite directions for positive and negative life events creating a distinctive “flip”. Specifically, when base rates are positively skewed (Fig. 4a), the simulation shows greater updating for bad news than good news for positive life events (significant difference in 100% of our simulations, Fig. 4c), whilst for negative life events updating for good news is greater than bad news (significant difference in 100% of our simulations, Fig. 4c). Note that this flip in asymmetry is exactly the pattern reported in Shah et al. (Fig. 4e). When the distribution of base rates is normally distributed however, the simulation does not reveal asymmetric updating (Fig. 4d, significant difference in only 5% of cases for positive and negative life events). Why is an artificial bias produced for skewed distributions? The answer is relatively simple – if the majority of base rates are low numbers then on average there is more room to alter estimates when the first estimate is higher than the base rate, than when it is lower. In other words, the difference between the first estimate and the information given (this is known as the estimation error) will be larger when participants receive bad news for positive events and when they receive good news for negative events. Thus, updates will be greater when information is “bad” for positive events and “good” for negative events, and vice versa when base rates are skewed towards high numbers. This statistical artifact, however, can be corrected by controlling for “estimation errors”. If in the simulations above we control for estimation errors in our analysis no bias is observed for any sets of base rates. This was tested by running an additional 20 simulations (10 for positive life events, 10 for negative life events, 20 participants per simulation) for each set of skewed base rates. For each set of simulated data we then conducted a repeated measures ANOVA with valence (good news/bad news) as a factor and entering the difference in estimation errors between good news and bad news as a covariate. We controlled for estimation errors on a condition level, rather than on every trial, so that we could directly compare our findings to that of Shah et al. When the distribution of base rates was positively skewed (Fig. 4a) a valence effect between good news and bad news was not significant in 19 (95%) of simulations. Thus, to avoid false conclusions researchers should either use normally distributed base rates with a mean at the midpoint of the scale, control for estimation errors, or do both. In previous studies, estimation errors have either been equal for good news and bad news trials (e.g., Garrett et al., 2014; Sharot et al., 2011, both of which used a scale of 3–77) and/or in instances where differences exist, these are carefully controlled for (e.g., Garrett & Sharot, 2014; Korn, Sharot, Walter, Heekeren, & Dolan, 2014). In contrast, whilst the differences between estimation errors in response to desirable and undesirable information can account for the patterns of update observed in Shah et al.’s experiments, estimation errors are not controlled for in their Experiments 1 and 2. In Experiment 3a, when estimation errors are finally controlled for, the critical interaction between valance of stimuli and desirability of information is no longer significant. In Experiment 3b and 4 of Shah et al., it is stated that the results hold when controlling for estimation errors but statistics to support this are not provided. It is worth noting that the experiments conducted by Shah et al. are not replications of the original belief update studies in design, nor analysis. A few of the differences may have significantly affected the results: (i) out of the 80 negative stimuli used previously, Shah et al. select 38 creating a skewed set. (ii) Shah et al. incentivize participants to remember the base rates provided and require them to input these using a keyboard. This deviation from the original task is important because the novel manipulation may encourage participants to use information in a different manner than they would otherwise. (iii) In contrast to all past studies, Shah et al. “…exclude trials which were more than 3 interquartile ranges from the mean value of the analyses. These exclusions reduced the total number of trials across the conditions by approximately 2.5%” (exclusion is reported for Experiments 1, 3a and 3b but not for Experiments 2 and 4). One may assume that this exclusion does not matter, but in at least one documented case the exclusion altered the results from non- significant in a previous version of the Shah et al. manuscript to significant with exclusion of outliers done post hoc. (iv) Different factors are controlled for in each one of the five experiments reported in Shah et al. The reason to control for some factors but not others in the different experiments is not given. We were unable to assess the consequence of this practice, because statistics of the tests after controlling for varied confounding factors were not provided in most instances. We now proceed to describe a new study, in which we examine how individuals integrate good and bad news into their beliefs about the likelihood of experiencing positive and negative life events, taking care to avoid the two pitfalls we have outlined above. Participants 300 participants located in the United States completed the survey on Mechanical Turk. As in past studies of the belief update task (Garrett & Sharot, 2014; Moutsiana et al., 2013, 2015), we excluded participants with a high Beck Depression Inventory (BDI) score indicating potential depression. 73 participants were excluded for having a BDI score greater than 11 (final sample = 227). Participants were all between the ages of 20 and 30 years of age (inclusive). Completion of the survey took approximately 25 min and participants were compensated for their time. Task The survey began by collecting basic demographic information from participants (age, level of education, marital status, employment status, monthly income) then 2 training examples were presented to familiarize participants with the task. Participants were then presented with 100 different commonly occurring life events for 3 s each. These were a mixture of positive events (for instance: “Discovered a new song you like”, “Laughed at a joke”) and negative events (for instance: “Had an argument with a family member”). Whilst the event was displayed on screen, participants were instructed to recall whether this event had happened to them in the past 4 weeks. They were then asked to indicate either (1) Yes: This event occurred to me at least once in the past 4 weeks; or (2) No: This event did not occur to me in the past 4 weeks. The order of these two options was counterbalanced. Participants had unlimited time to make a response (Fig. 5a). After completing the survey, participants rated each event on a 5 point likert scale (1 = Very Negative; 2 = Negative; 3 = Neutral; 4 = Positive; 5 = Very Positive) and then completed the BDI (Beck, Ward, Mendelson, Mock, & Erbaugh, 1961) and Life Orientation Test Revised (Scheier, Carver, & Bridges, 1994). The survey was constructed and presented using web based survey service Qualtrics. Analysis For each event, the percentage of participants who indicated the event had occurred to them in the past month (out of all participants who completed the study and were included in the final sample) was calculated. Event selection A subset of the events (n = 54) were selected for use as stimuli. We selected positive and negative life events such that the distribution of base rates for each type of event (positive and negative) was normally distributed around a mean of 50%. Note that we were unable to verify how accurately participants were able to recall whether events occurred to them or not in the past month. However, since a base rate for a specific event will represent good news to some participants and bad news to others (depending whether a participant overestimates or underestimates the base rate), noise from such inaccuracies ought to cancel out over good news and bad news events in the belief update task. Participants 200 participants located in the United States (age range 20–30) completed the survey on Mechanical Turk. 56 participants were subsequently excluded for having a BDI score above 11 indicating possible depression. A further 2 participants were excluded because the range of their responses was limited, resulting in zero trials in either the “good news” bin or “bad news” bin, making comparison impossible (final n = 142, mean age: 25.74; mean BDI score: 2.80). There were no differences in age, education, income, marital status or employment status between this set of participants and participants that had completed the base rate survey used to construct the base rate statistics (all P > 0.20). Completion of the survey took approximately 1 h and participants were compensated for their time. Task The survey began with an attention check designed to filter out participants that did not read the instructions prudently. Then, demographic information was collected (age, level of education marital status, employment status, monthly income) and 2 training examples provided to familiarize participants with the task. In the first session, on each trial (54 trials in total) participants were presented with 1 of 54 life events (see Supplementary Material, Table 1 for list of events used) and asked to imagine the event happening to them in the month ahead. They were then asked to estimate how likely that event was to happen to them in the next 4 weeks. Participants were instructed to type in an estimate between 5% and 95% using a computer keyboard. Trials with responses outside this range were excluded from analysis (mean(s.d.) number of responses outside this range: 1.80(3.43)). Participants had 8 s to provide a response. Participants were then shown the base rate statistic of the event happening in the next 4 weeks, which ranged from 15% to 85% (see Fig. 5b). They were told that the statistic was the average likelihood of this event happening at least once in the next four weeks to someone from the same socioeconomic environment as them. In a second session, which took place immediately after the first session, participants were asked to re-estimate how likely each event (54 trials in total) was to happen to them in the next 4 weeks. As in the first session, participants had 8 s to provide a response. After completion of the task, we tested participants’ memory for the information presented. Participants were asked to recall the information previously presented of each event. Subsequently, participants were then asked to rate all life events according to how positive or negative they found them on a five point Likert scale (1 = very negative, 2 = negative, 3 = neutral, 4 = positive, 5 = very positive). This range of scale was used so that events could be clearly categorized into events that were considered negative (assigned a rating of 1 or 2), neutral (assigned a rating of 3) or positive (assigned a rating of 4 or 5). Participants were also asked to rate past experience with each event (“Has this event happened to you before?” From 1 = never to 6 = very often), as done previously (Garrett & Sharot, 2014; Garrett et al., 2014; Sharot et al., 2011; Sharot, Guitart-Masip et al., 2012; Sharot, Kanai et al., 2012). Three quarters of participants (75%) also rated all events on: vividness (“How vividly could you imagine this event?” From 1 = not vivid to 6 = very vivid); familiarity (“Regardless if this event has happened to you before, how familiar do you feel it is to you from TV, friends, movies and so on?” From 1 = not at all familiar to 6 very familiar); and arousal (“When you imagine this event happening to you how emotionally arousing is the image in your mind?” From 1 = not arousing at all to 6 = very arousing). The scores of these are reported in Supplementary Material, Table 2. Participants then completed the BDI and the Life Orientation Test Revised. The survey was constructed and presented using web based survey service Qualtrics. Analysis Life events were categorized as negative or positive for each participant individually according to their own evaluation. Specifically, events were classified as positive if the participant rated the event as 4 (positive) or 5 (very positive) in the ratings section of the task, and negative if rated as a 1 (very negative) or 2 (negative). Events with a neutral rating of 3 were excluded from the analysis (mean(s.d.) number of events with neutral rating: 7.17(5.67)). For each type of event, participants could receive either “good news” or “bad news” depending on whether the participant initially overestimated or underestimated the probability of the event relative to the base rate (see Fig. 5c). Specifically, if their first estimate was lower than the base rate presented, the information would be categorized as “good news” if the life event was positive and “bad news” if the life event was negative (column 1, Fig. 5c). If their first estimate was higher than the base rate presented, the information would be categorized as “bad news” if the event was rated as a positive life event and “good news” if the event was rated as a negative life event (column 2, Fig. 5c). Trials in which the initial estimate was equal to the statistic presented were excluded from subsequent analyses as these could not be categorized into either condition (less than one negative life event trial and less than one positive life event trial on average per participant). Belief update was calculated for each trial and participant as the difference between first and second estimate. As done previously (Garrett & Sharot, 2014; Garrett et al., 2014; Moutsiana et al., 2013, 2015; Sharot, Kanai et al., 2012) update was calculated such that positive scores indicate a move towards the base rate, regardless of event type and valence categorization, and negative scores a move away from the base rate. Update scores were then entered into two general linear models using the IBM SPSS statistics software; one for negative life events and one for positive life events, with information valence as a fixed factor (good news/bad news), participant ID as a random factor and absolute memory errors and absolute estimation errors on each trial as covariates. Since covariates vary from trial to trial, controlling for them on a trial by trial level is more precise than on a condition level as done previously by us (Garrett & Sharot, 2014; Garrett et al., 2014; Moutsiana et al., 2013, 2015; Sharot, Guitart-Masip et al., 2012; Sharot, Kanai et al., 2012) and others (Shah et al., 2016). Nevertheless, to allow direct comparison with Shah et al., we also entered mean update scores for each participant into a 2 (good/bad news) by 2 (positive/negative life event) repeated measures ANOVA controlling for the following covariates on a condition level: (1) the difference in memory for good news trials and bad news trials, both for positive and negative stimuli, (2) the difference in number of good news trials and bad news trials, both for positive and negative stimuli (3) the difference in absolute estimation errors for good news trials and bad news trials, both for positive and negative stimuli (estimation error = |first estimate − base rate|).","We observed an asymmetry in updating, such that participants updated more in response to good news than bad news. This optimistic update bias was significant both for positive life events (t = −2.88, p < 0.01, mean good news update = 8.71, mean bad news update = 7.79) and for negative life events (t = −4.88, p < 0.001, mean good news update = 10.66, mean bad news update = 6.62). There was also a significant effect of estimation errors (negative events: t = 13.43, p < 0.001; positive: t = 15.81, p < 0.001) and memory errors (negative events: t = −6.13, p < 0.001; positive events: t = −2.92, p < 0.01). Comparing the bias for positive life events and negative life events by entering mean update scores for each participant into a 2 ∗ 2 repeated measures ANOVA with desirability of information (good/bad news) and life event type (positive/negative life event) as repeated factors (controlling for differences in memory, differences in number of trials and differences in estimation errors) revealed the expected main effect of desirability of information (F(1, 135) = 6.29, p < 0.02), no effect of event type (F(1, 135) = 0.08, p = 0.78) and no interaction (F(1, 135) = 0.31, p = 0.58). Fig. 4f.","The current set of results strongly support a valence dependent asymmetry in how participants update their beliefs, consistent with a large body of fast growing research (Eil & Rao, 2011; Garrett & Sharot, 2014; Garrett et al., 2014; Korn et al., 2012; Krieger et al., 2016; Kuzmanovic et al., 2015, 2016; Lefebvre et al., 2016; Ma et al., 2016; Möbius et al., 2012; Sharot, 2011; Sharot & Garrett, 2016; Sharot et al., 2011; Sharot, Guitart-Masip et al., 2012). By pitting this asymmetry against three robustness tests suggested by critics of optimism (Shah et al., 2016) we find that it survives each of these tests, suggesting it is a pervasive phenomenon. Specifically, we first summarize past data showing that an optimistic update bias exists under a different type of classification (Garrett & Sharot, 2014; Kuzmanovic et al., 2015). Second, we apply the Bayesian analysis outlined by Shah et al. to a previously collected data set (Garrett & Sharot, 2014) and reveal that participants’ updates are more Bayesian in response to good news than bad news. Third, we test Shah et al.’s curious report of a distinctive “flip” in the update asymmetry for positive life events, with updates being greater for bad news compared to good news (the opposite asymmetry to that observed for negative life events). We find that their studies fall into two methodological pitfalls, which can account for their results. Here, we avoid these pitfalls and reveal an optimistic update bias, for both positive life events and negative life events. Moreover, whilst an optimism bias has previously been shown to exist for both everyday and significant life events (Strunk, Lopez, & DeRubeis, 2006; Weinstein, 1980), biased belief updating had until now only been revealed for the latter (Garrett & Sharot, 2014; Sharot et al., 2011). This is the first demonstration of an optimistic update bias for everyday life events. The asymmetry in belief updating can result in overly optimistic beliefs. Whilst biased, these beliefs may be adaptive as positive expectations reduce stress and anxiety facilitating physical (Taylor, Kemeny, Reed, Bower, & Gruenewald, 2000) and mental health (Garrett et al., 2014; Korn et al., 2014; Strunk et al., 2006). Furthermore, optimistic expectations enhance motivation and exploration (Bandura, 1989; Puri & Robinson, 2007) increasing the likelihood of gaining resources (Johnson & Fowler, 2011). However, alongside these benefits to optimistic expectations, optimism has also been suggested to reduce necessary precautionary action leading to ill preparedness in the face of natural disasters (Paton, 2003) and failure to adopt preventative measures to safeguard ones health (Shepperd et al., 2013) (see Bortolotti & Antrobus, 2015 for a discussion of the benefits and costs associated with optimistic expectations). On balance it is possible that the positive consequence of unrealistic optimism outweigh the negative, leading humans to evolve this asymmetry in belief formation (McKay & Dennett, 2010)."],["This article theoretically discusses Arlie Hochschild's (1983, 1998) concept of the ‘real’ and ‘false’ self (1983: 194) and how this holds together her model about how it is we manage our emotions. Hochschild draws on ideas about surface acting, deep acting and authenticity to support her theory of emotion management. In this discussion I argue that these ideas undermine the clarity of the theoretical model Hochschild tries to develop to explain emotion management. The first aim here is to demonstrate that this concept of the real and false self acts as an unnecessary conceptual linchpin making Hochschild's ideas about emotion management opaque. The second aim in this article is to theoretically engage with Pierre Bourdieu's (1984, 1990) concept of habitus as a way of overcoming Hochschild's idea of the real and false self. --------------------------------------------------------------------------------","This article discusses Arlie Hochschild's model of emotion management (1983: 35) and identifies inherent problems with her use of the ‘real’ and ‘false’ self as a conceptual linchpin (Hochschild, 1983: 194–195). My intention is to explain these problems with this emotion management model and offer an alternative for the ‘self’ that Hochschild describes by drawing on Bourdieu's concept of habitus (1984, 1990). The real self is considered by Hochschild to be the very core, or essence, of who we are as a person, and in contrast, the false self is ‘a part of “me” that is not really “me”’ (Hochschild, 1983: 194). My contention here is that there is no such thing as the real or false self, nor is it important to make such a distinction. Hochschild (1983) wanted to explain how it is that we can act differently in certain social settings by managing our emotions. She suggests that by managing our emotions we are able to work on the self and present to the world a persona that is expected, and fits in. Her model of emotion management was ground-breaking because it helped to open up debate about the invisible and unrecognised work people do in order to fit in with social expectations (see Mann, 2004; Bolton, 2005). In this article I want to undo the dependency on the concepts of the ‘real’ and ‘false’ self that is complexly bound up in this model. In the first part of this article I deal with this by showing how Hochschild repeatedly draws on the real and false self as a conceptual linchpin in her research (1979, 1983, 1997, 1998) and how this makes her ideas inconsistent and opaque. In the second part of this article I engage with Pierre Bourdieu's (1984, 1990) concept of habitus as a way of overcoming Hochschild's idea of the real and false self.","Hochschild developed a model (1979, 1983, 1998) to explain how we manage emotion in certain social settings and around certain people. This arose out of her research into flight attendants working for Delta Airlines in the United States of America (1983). Her research looks at how employees become who they are expected to be at work. Hochschild revealed that these flight attendants were expected to act in a particular way at work to fit in with the organizational expectations of the ideal female employee. For these female employees this included being perceived as caring, mildly flirtatious, and impervious to rude customers, as well as dressing in a particularly feminised way that included a certain way of wearing make-up, uniform and hair (Hochschild, 1983: 101–103). This finding in itself was revealing of constraints on female employees in particular (1983: 127–128). However, what made Hochschild's work distinctive at this time was that she offered an insight into how it is these female employees were managing to do all these things and become the right kind of employee (105–106). Before scrutinising Hochschild's emotion management model, it is worth briefly signalling how Goffman has influenced her early work. Hochschild wanted to depart from Goffman's construction of an individual that she argued is made passive to rules governing interactions (1959; Hochschild, 1983: 225–227). She departs from Goffman's ideas about the self as a collection of many roles and performances because she is concerned with what she sees as a lack of continuity. She argues that Goffman's account of reality provides ‘no structural bridge between all situations’ (1983: 225). That is – an explanation of how a person is the ‘same’ from one moment to the next. She finds this problematic for two reasons: firstly, because this would suggest that a person is governed by social rules as a passive individual who has a lack of interiority. She notes how Goffman seems to ignore times when an ‘individual introspects or dwells on outer reality without a sense of watchers’ (1983: 226). Even though Goffman later explored to some extent a person's inner world and their social context in ‘Asylum’ (1961), by mainly focusing on the emotion of embarrassment, he does not discuss the internalised feeling rules or capacity for agency which Hochschild sees as being ‘“inside” the actor’ (1983: 226) and fundamental to the management of emotion (1983: 228). For Hochschild, then, it is this interiority and agency that is the ‘bridge between all situations’ (1983: 225), and this brings her to the notion of an inner essence – or real self. Secondly, she does not think that Goffman properly accounts for how people are able to use prior expectations to help navigate new situations. She criticises him saying that there is ‘no overarching pattern that would connect the “collections”’ 1983: 225). For her, ‘the idea of prior expectation implies the existence of a prior self that does the expecting,’ (Hochschild, 1983: 231). She provides this example: When we feel afraid, the fear signals danger. The realization of danger impinges on our sense of self that is there to be endangered, a self we expect to persist in a relatively continuous way. Without this prior expectation of a continuous self, information about danger would be signalled in fundamentally different ways (Hochschild, 1983: 231). Hochschild (1983) is uncomfortable with the idea that a person may be different depending on the stage setting and context. For her, there is continuity in terms of how a person acts and feels and that this is only possible because of an inner ‘real self’ (1983: 34). She writes, ‘To develop the idea of deep acting we need a prior notion of the self with a developed inner life. This, in Goffman's account, is generally missing’ (1983, 227). For Goffman, there is no such thing as real or false performances signalling a true self. According to Goffman, all of our performances are real in the sense that they simply take place – there is no unchanging core that is the ‘real’ self, only an ongoing and increasing personal portfolio of roles (see 1959: 252–253). However, Hochschild (1983) identifies that an explanation of continuity between moments is under developed in Goffman's work. Goffman did not write in detail about a reflexive or agentic self as such, but the need to explain continuity between situations (as Hochschild tries to do) is not, in my view, achieved through the ‘real self’ as a conceptual linchpin, which I will now discuss further.","Hochschild's ‘Managed Heart’ model of emotions (1983) quickly developed into a typology to explain how it is that emotions are performed or concealed in certain social settings. She identified two different types of emotion management: emotion work and emotional labour. Hochschild describes emotional labour as: ‘the management of feeling to create a publicly observable facial and bodily display; emotional labour is sold for a wage and therefore has exchange value’ (1983: 7). Emotion work is slightly different to emotional labour; as Hochschild states: ‘I use the synonymous terms emotion work or emotion management to refer to these same acts done in a private context where they have use value’ (italics in original, Hochschild, 1983: 7; see page 181 in book for further discussion). Hochschild suggests that we may undertake emotion work in our day-to-day lives in order to present feelings in a more agreeable way to friends, family and acquaintances, for example, by hiding anger or embarrassment to preserve social relations (1983: 19–20). Hochschild develops her model by outlining the mechanisms that make emotion work and emotional labour possible. She focuses on surface and deep acting (1983: 48–49). According to Hochschild (1979, 1983), surface acting is a practice in which an individual offers a performance that displays the expected feelings they sense are in keeping with the feeling rules structuring that particular social interaction, regardless of whether this is how they feel or not. This surface acting of expected feelings, Hochschild suggests, is an insincere performance that the individual hopes is convincing to others, nonetheless (1983: 49). For instance, the flight attendant smiles to show happiness; whether she actually feels happy or not does not matter (1983: 127–128). To put it another way, we portray or mimic what we think is expected of us and conceal undesirable feelings. In short, Hochschild suggests that what we are doing is acting out or mimicking the ‘shoulds’ accorded by feeling rules that structure interactions, but we are not obliged to internalise these feeling rules as our own (1983: 118). Surface acting then is about knowing how to act in a given situation (1983: 48). This means knowing the implicit feeling rules structuring workplace interactions. Knowing how to display emotions is essential to being able to fit in within the workplace. To get surface acting right requires some attention to the audience, usually a customer or co-worker, in order to discern whether the emotional performance has been convincing to them. This is very similar to how Goffman (1959) describes the dynamics of performing a role during social interactions. The employee interacts with the other person whilst trying to pick up clues that their performance may possibly be viewed as unconvincing. The crux of surface acting is to offer a performance that leaves the other person convinced that they had a meaningful interaction. This person tries to conceal from the other person that they were performing emotional displays that were simply expected of them. Another aspect of Hochschild's emotion management model relates to deep acting. This involves a person trying to sincerely embody an emotion so that displaying it for the other person is no longer a fake but convincing performance and becomes ‘real’ (1983: 194). Hochschild describes deep acting as deciding ‘what it is that we want to feel and on what we must do to induce the feeling’ (1983: 47). The person tries to make their emotional displays seem authentic to themselves as well as the other person. Hochschild goes further and describes the practice of deep acting as working hard trying to feel a particular emotion. This involves using emotional recall of memories of a situation where the individual really had felt happy: this memory is then re-visualised, invoked and attached to their present circumstances to shape the mind and bodily behaviour. Hochschild states that by, trying to feel what we sense we ought to feel or want to feel (Hochschild, 1983: 43) we must undertake deep acting, this activity of working on emotions at a ‘deep level’ so that they are felt as ‘real’ is accomplished via a process of imagining, that is, to think about a desired emotion we wish to feel and imagine it as if this were true (1983: 43). Hochschild suggests that by inducing an imagined emotion via deep acting, the self will come to accept it as authentic and part of the real self. Deep acting requires that the individual suspends their ‘usual reality testing’ and instead ‘allow a make-believe situation to seem real’ (Hochschild, 1983: 42), in the hope that it will take on the qualities of being real at a later stage. Deep acting is not only used to theorise the inducement of imaginary feelings in Hochschild's framework, but it is also used to refer to how an individual might prevent a real feeling from emerging from the depths of the real self and mis- fitting the situation. Hochschild uses her concept of deep acting to explain how the individual attempts to convince themselves that they really feel something else other than what they are feeling, or else ‘block or weaken a feeling we wish we did not have’ (Hochschild, 1983: 43). Deep acting then involves ‘bad faith’ (Hochschild, 1983: 47), which is a problematic concept that suggests that individuals can intentionally deceive themselves. A person would have to know that they were trying to forget something that they know. For Hochschild, deep acting is also about lying to ourselves, but the lie is suppressed in the hope that the lie will disintegrate and the deception we began with will take the form of something real. Whilst Hochschild describes the practice of deep acting she is not clear or convincing about the purpose of deep acting. Hochschild tries to convince us that the point of deep acting is to make a feeling that is imagined seem real, so that it becomes real.","Hochschild presents an array of interesting concepts as part of her model of emotion management (1998, 2003). She offers a way to theorise deep acting but is never quite clear about what it is, what the purpose of deep acting is, or how it is done. Instead, she offers numerous examples that she suggests demonstrate deep acting and often relies upon the reader to intuitively know what deep acting is, how it is done, and why. Deep acting is necessitated in Hochschild's study because she senses that there are times when people do not feel as though they fit in as they are in certain situations. People are motivated to do deep acting to bring mis-fitting feelings more in line with what is expected in a given social situation and transforming them. Deep acting is unnecessarily complicated, however, by Hochschild's discussion of real feelings and false feelings. Deep acting tries to explain how a person can knowingly hold two (or more) contradictory feelings in place – neither has to be real or false. This perception of real or false feelings, I would suggest, arises out of a calculation the person makes about some feelings ‘fitting in’ with the social space they are in, and other feelings being viewed as mis-fitting. These mis-fitting feelings are concealed, although not forgotten, because they do not fit in case they might incur a social sanction. Holding these contradictions in place can be painful for an individual. We do deep acting because we want to fit in within the dominating structures of feeling. The desire to fit in is what I would suggest motivates deep acting, despite it being a painful thing to do. Therefore, the idea of real or false feelings is rather a distinction a person makes about the kind of feelings that fit in and those that do not. Deep acting is painful precisely because it is work on the self that forces an adjustment to who we are told we ‘ought’ to be or how we ‘should’ feel, and recognition that we do not presently embody this already. It feels strange also to act in a way that we are not used to. Put another way, taking on someone else's rules to govern our own emotions and behaviour makes us feel odd, ill at ease, and it feels wrong when we try to convince ourselves that this is our normal, everyday behaviour and way of feeling. Deep acting can make us feel anxious because it is hard to do, and yet oddly enough it can also help to reassure us that at least we fit in better in certain social spaces by doing it. This process of ‘becoming’ who we sense we should be, by making painful adjustments, is rife with tension and anxiety. Deep acting is painful and invokes tensions because we harbour thoughts that we should already be that person we are trying to become. What is more, there is an on-going debasement directed at the self for having to undergo deep acting in the first place. A painful tension arises when we undertake deep acting because we think we shouldn't have to try so hard to feel a certain way, to be a certain kind of person, we want to be that person already to feel a certain way already. Having to labour at it is a clear sign that we are not the person we feel we ought to be, we do not know what to feel. The pain of doing deep acting may ease over time though as one way of feeling is replaced by another perhaps more socially desirable way, and is assimilated as if it were already our own way of thinking and feeling. Tonkens (2012) also makes a similar argument, suggesting that Hochschild does not sufficiently clarify what she means by particular concepts and what they are supposed to do. Tonkens too is critical of Hochschild's lack of connections between her concepts and how they are supposed to relate to each other in this emotion management model. In Hochschild's explanations of deep acting, for instance, she tends to jump from talking about deep acting as being about making false performances of feelings feel authentic (1983: 35–36), then to how deep acting is also about suppressing real feelings, and then she moves to a different theme altogether within deep acting in which the individual is also trying to preserve the real self (an inner jewel) (1983: 34) and manage the false self (1983: 195). Even after this extensive use of deep acting as a form of emotional labour/work, she theorises deep acting as also including the practice of conscious and continuous self-deception in which the real self is fooled by an illusion which they have deceived themselves into believing (1983: 40–42). It is difficult to untangle what Hochschild means exactly by deep acting and what it is supposed to do precisely, and the reason for this is because she tries to make the concept account for too much within her theoretical framework. Hochschild is trying to use deep acting as a conceptual device that explains how we try to make ‘who we are not’ (what she sees as the false performances) become ‘who we are’ (part of the real self). Deep acting is a mechanism that she uses to explain how we try to become a person with particular feelings because of a sense that this is who we ought to be and how we ought to feel. Deep acting then is a form of emotional labour/work that has the purpose of becoming someone else in order to fit in with already structured expectations and rules. It is necessary to her theoretical framework that she is able to justify a distinction between the real and false self because she argues that it is the real self that is being exploited by capitalist organizations (1983: 34; see also Fineman, 2008). I discuss this separation of the real and false self shortly and how it is particularly problematic in her work, and hence why I bring in Bourdieu's idea of the habitus (1984) to resolve it.","When Hochschild talks about a real self she is describing a self that has honest and true feelings that are not subject to pretence or acting. For her, the real self constitutes a continuity, an embodied way of being and doing that is predictable and recognisable to others. Possessing a ‘real self’ tells others what we are like as a thinking and feeling individual. Hochschild compares her idea of a real self to an inner jewel or essence that makes us who we are. To do this, she sets up an agentic, choosing individual with an internalised sense of continuity: a real self. The real self is a formed identity that is reliably the same, from day-to-day, and it is in our possession to control. In this sense then, I would argue that Hochschild mistakenly interprets the anxiety that arises out of trying to become someone we sense we ‘should’ be as evidence of a separation between a real and false self – for example, ‘I wasn't really being myself’ (Hochschild, 1983: 262). And yet, this is based on the person's perception of what is real and false, and a sense of a core self. She uses these dispositions to evidence a ‘real self’ and ‘an inner jewel that remains our unique possession no matter whose billboard is on our back […] we push this “real self” further inside, making it more inaccessible’ (1983: 34). What is missing then is a distinction and reflection made between how the individual sees themselves, and how the self is being socially constructed in these narratives through dispositional histories. Working on an emotion so that it fits in with the social space the individual is in can be painful. Hochschild argues the more we have to work on ourselves to become comfortable with representing an emotion we think we ought to feel, the more inauthentic the emotions we are trying hard to embody become, and therefore we find ourselves getting further away from our real selves. She writes that ‘subtracting credibility from the parts that are in commercial hands, we turn to what is left to find out who we “really” are’ (1983: 34). Hochschild further argues that backstage is where the ‘real self’ can supposedly relax and emerge, and that it is in this space that more authentic performances occur rather than during front stage performances, which tend to be around customers and clients (1983: 192). Erickson (2011:121) notes how there has been increased attention around the idea of authenticity and ‘the real thing’ in a post-industrial climate. Cain's research (2012) looks at authenticity and the emotional labour of workers in practice at a care hospice in North America and it challenges Hochschild's idea of a split real and false self. Cain argues that this was not the case in her research and instead shows how workers felt that how they act around their patients is just as authentic and real as their behaviour in the staff room. These workers felt that they presented a more formalised and different ‘hospice identity’, rather than a ‘false self’ which Hochschild would suggest. In Cain's research this worker identity was important to how the participants' felt they should manage their own emotions at work. Hochschild is engaging in a phenomenological debate about a person's state of being in the world. She is advocating the notion of a real person, with real feelings, that possesses an alter ego which she describes as the ‘false self’ (1983: 194), which creates illusions and make believe that fool the real self. This false self is described as ‘a disbelieved, unclaimed self, a part of “me” that is not “really me”’ (1983: 194). According to Hochschild, the difference between what is the real self and what is false depends on what aspects of it we claim, so for instance, we may say ‘I wasn't really being myself at that party’. For Hochschild, we each possess a real self, or inner essence that we know to be true (1983: 34): the critique I am making here is that Hochschild's concept of the real/false self requires more critical engagement with why a person might feel that way about their identity.","Hochschild needs the concept of a ‘real self’ with a ‘real life’ (1983: 47) in her theoretical framework to explain what she sees as the exploitation of a person's inner essence by capitalist organizations. Hochschild is suggesting that being made to become someone else who is ‘other’ to us at work, and being told how to feel according to structured feeling rules, is a suppression of an individual's core self and their agency. So, feeling pressured to change because of feeling rules makes us shape our sense of self in a way that we do not necessarily do out of choice in the workplace, and in ways we are not necessarily accustomed to or feel comfortable with. Hochschild says, ‘The airline passenger may choose not to smile, but the flight attendant is obliged not only to smile but to try to work up some warmth behind it’ (1983: 19). Hochschild is concerned with how the transmutation of emotion work, that is, a personal choice to work on the self to become someone we feel we ought to be for others in our private lives, is exploited by capitalist organization and given an exchange value. This exploitation amounts to a suppression of the real self. Hochschild therefore needs to set up this construct of the authentic, agentic individual to justify her argument that capitalist organizations are exploitative of emotional labour and emotion work. Many workers have little choice but to work on themselves it would seem, although this is not true of all workers who can exercise power to negotiate workplace feeling rules (1983: 19, 89). For the most part, it is concerning to Hochschild that workers potentially lose their capacity to freely display how they feel. This capacity to show feelings is relinquished by the employee as part of the expectations of their contract and instead they must internalise workplace feeling rules if they want to get paid. She constructs the exploitation of a person's emotional system by capitalism, that is, the transmutation of emotion work into emotional labour, as an immoral act that infringes on the nobility and sanctity of a ‘real self’ (1983: 34). Referring to a real self, Hochschild draws upon Rousseau's ‘Noble Savage’ (1983: 192) to describe a person who is not subjected to any feeling rules, who feels spontaneously and without calculation. This real self is to be valued, protected and preserved, according to Hochschild, from the onset of capitalism and the demands of emotional labour, lest we become a ‘faceless soul beneath the mask’ (1983: 194). Hochschild needs to explain this act of exploitation, specifically of emotion, via an idea of valuing a real self (1983:192) which she sees as increasingly constricted by rules, deeply entwined in the matrix of capitalism, with little choice but to exchange an inner self and emotions for a wage. Bolton (2005: 39) is critical of Hochschild's theory of emotion management, saying that it ‘restricts the possibility of individuals ever being active agents, who through negotiation are able to break the “chain” of power and “make their own histories”’. Furthermore, Bolton is opposed to the idea that employees are coerced to align themselves with organizational rules and lose a ‘sense of self in the process’ (2005: 39). Instead, Bolton argues for a model of a person who reflects the modern, reflexive individual (Bolton, 2005; Bolton and Boyd, 2003; Giddens, 1991, 1994; Urry, 2000; Beck, 1992, 1994; see Atkinson, 2010, 2012; and Bourdieu and Wacquant, 1992; Du Gay, 1996). That is, someone who is able to navigate, negotiate and overcome feeling rules that have the capacity to constrain employees. She is suggesting then that employees can choose how to feel at work to a greater extent than which Hochschild allows for within her theoretical framework. According to Bolton, Hochschild is suggesting that capitalist organizations have: Appropriated all of our feelings so that there is no longer any room for sentiments, moods or reactions that have not been shaped and commodified via the ‘commercialization of intimate life’ (Bolton, 2005: 2). Bolton suggests that Hochschild's argument – that intimate life is increasingly commercialized through emotional labour – inevitably positions employees as passive, or, as ‘crippled actors’ (Bolton, 2005: 48). Based on her own research into caring work, Bolton finds Hochschild's concept of emotional labour lacking depth because she says that it only describes one particular kind of emotion management relative to the service sector (i.e. that emotions are the product that is sold). For Bolton, some emotions are freely given as part of social relations during interactions with others. Bolton identifies here then that Hochschild's concept of emotion work is under-developed. We do not always work on emotions at work because it is exchanged for a wage – sometimes working on emotions can be useful as part of social relations. So for instance, some employees that Bolton studied in the care industry (Bolton, 2000, 2005; Bolton and Muzio, 2008; see also, Mann, 1999) use their emotions to facilitate their interactions with others, but their emotions are not what is being sold. Bolton sees Hochschild's concept of emotional labour as having too narrow a focus on capitalism and puts forth her own typology of emotion management: ‘pecuniary’ with ‘prescriptive’, which relate to instrumental performances of emotion in the workplace rooted in economic and status gain and are empty of feeling; as well as ‘presentational’ emotion management, which refers to the ‘basic socialized self’ (Bolton and Boyd, 2003: 297), and ‘philanthropic’, which is an intentional and freely chosen act of giving emotion to customers and co-workers that are not prescribed by workplace feeling rules. In this sense, according to Bolton, workers seek these ‘unmanaged spaces’ (2005: 102) in order to express their true and ‘authentic’ selves. I agree with Bolton when she says that there is more to emotion management at work than just exchanging emotions for a wage. But Hochschild acknowledges this, too – she just doesn't advance this area of her research. She does write briefly about emotion work as a theoretical concept, which addresses how people manage emotion in their personal social relations, and how this act has use value (Hochschild, 1983: 7; see page 181 in book for further discussion). It is true that this is particularly under-developed. I suspect that Hochschild would agree with Bolton that people work on their emotions at work for the purposes of maintaining social relations but that capitalist organizations are slowly starting to encroach on this kind of emotion management, threatening the ‘real self’, to advance their own strategies and agendas. Bolton's assessment of Hochschild criticises her emotion management model for constructing a ‘crippled actor’ who is overly constrained by feeling rules in the workplace. Bolton instead contends that her studies show the opposite, that employees actively choose how to follow workplace feeling rules (2003) and retain control over managing their emotions. However Bolton's criticism is incorrect – the actor is not ‘crippled’ yet; Hochschild is warning about this being a possibility if capitalist organizations are permitted to commodify emotion management. Hochschild is trying to preserve the idea of free will and agency, that is, the intentional and choosing individual, just as much as Bolton. What Bolton appears to be arguing is simply a matter of degree – the degree to which an individual's agency to negotiate structured feeling rules is (un)constrained at work. Hochschild is not describing someone who is without agency, only that this agency is tightly managed in the workplace and that this could get worse. Both Hochschild and Bolton's approaches to emotion management still rely on the idea of an authentic, choosing individual, in short, a true self. Just as Bolton has criticised Hochschild for creating an individual with no agency, I too would criticise Bolton for conceptualising an individual with too much agency, drawing on the words of Brook, for ‘claiming a form of supra-autonomy for emotion work’ (Brook, 2009: 540).","What is missing from both Hochschild's (1983) and Bolton's (2005) theoretical framework is a way of explaining how the individual interacts with structured feeling rules that does not rely on a true or real self. The main criticism that is generally levelled at theories that are based on the notion of a real self with an intangible core is that it constructs a wholly agentic and reflexive individual (Bolton, 2005; Bolton and Boyd, 2003; Giddens, 1991, 1994; Urry, 2000; Beck, 1992, 1994) who is able to stand outside of social structure (Bourdieu, 1984). What is needed to resolve the problems that arise out of centring theory on a real self is a way to overcome the antimony between the personal and the social, and structure and agency. To put it another way, a conceptual bridge between the mind and the body is needed to better explain how emotion management occurs. Tonkens (2012) also identifies this problem in Hochschild's work (1979, 1983, 1998, 2003). Tonkens argues that, ‘There is a theoretical lacuna in Hochschild's work on how relationships between individual emotions, social interactions, and large-scale processes like globalization and commercialization relate to one another’ (2012: 196). She suggests that Hochschild struggles to overcome the analytical gap between core micro concepts and macro level concepts.","I want to suggest that Bourdieu (1984) offers a more fruitful way of explaining continuity across moments (which Hochschild attempts to do by drawing on the notion of a real self) in his theory of the self as embodied history. Bourdieu describes that a person's embodied history is the accrual of memories and knowledge that are embedded as dispositions (1984). This conceptual device is described by Bourdieu as a person's habitus (1984). A person's habitus disposes a person to think, feel and act in ways that are the outcome of their ‘conditions of existence’ (Bourdieu, 1990: 52). The habitus then is formed out of ‘systems of durable, transposable dispositions, structured structures predisposed to function as structuring structures’ (Bourdieu, 990: 53). This means that a person will feel in a way that is shaped by ‘generative principles’ (1990: 53) that provide a structure and logic with ‘no active conscious intent’ (Addison, 2016). This means that a person is a product of history and their particular moment in time, and as such delineates relationality across structure and agency. Using Bourdieu's concept of the habitus it is possible to think about the self more as a unique set of embodied dispositions that we use to strategize our actions and our feelings. This means that there is no need to think of the self as an intangible core or a real self that can be criticised for a ‘supra-autonomy’ (Brook, 2009: 540) that exists outside of social structure (Barbalet, 2001, 2002). Moreover, the notion of a ‘real self’ that emerges in Hochschild's interviews with flight attendants (e.g. ‘I was not being myself’) is much better framed as a discussion about reflexivity, in two ways: firstly, on an epistemological level – this involves scholars ‘ being reflexive regarding how knowledge is generated, produced, represented and legitimated’ (Addison, 2016: 18; see also Bourdieu and Wacquant, 1992); and secondly, by critically engaging with how the individual is being reflexive about themselves in the world (Skeggs, 2002) and builds a narrative around this (Lawler, 2008). Whilst some individuals may feel they possess an inner self that is real and authentic, I argue that this is a perception that Hochschild does not sufficiently interrogate. Hochschild used the concept of the real and false self in her model of emotions. In contrast, I want to draw on Bourdieu to argue that we use dispositions as knowledge of how to act and feel in certain situations, and knowledge of how to express, and importantly manage, our emotions. Using this model of the self avoids the criticism directed at Hochschild that individuals are overly structured by feeling rules (Bolton, 2005), as well as the criticism levelled at Bolton (2005) by Brook (2009) that individuals are able to stand outside of structure and seek out ‘unmanaged spaces’ (Bolton, 2005: 102) as agentic and autonomous individuals. Moreover, Bourdieu suggests that we are all already born into social games that have started without us. We are immersed in the social world and acquire dispositions as we grow that orient us to the correct way to do things in certain social spaces and around different people. We grow accustomed to the different rules and principles that structure different spaces meaning that we ‘fit in’. By thinking of the formation of the self in this way, it also avoids the problem of the apriori self – where we are born with a unique essence, or as Hochschild puts it, an ‘inner jewel’ (1983: 34) that makes us who we are and enables us to be autonomous. For Bourdieu, this is unnecessary: he explains the formation of the self as being socialized into the ways of game playing from birth through our habitus and position in the field so that we develop a ‘feel for the game’ (Bourdieu, 1990: 67). Playing games involves fitting in with the ‘right ways of being and doing’ (Bourdieu, 1986, 1990: 511) in certain social spaces. This practice of fitting in convincingly, and having the right habitus, involves (amongst other things) managing our emotions. Having a feel for the game Bourdieu describes as ‘the sense of the imminent future of the game the sense of the direction (sens) of the history of the game that gives the game its sense’ (1990: 82). How well we play social games depends on our embodied history (habitus), the position we hold in the field, and the different resources (material and embodied capital) we have at our disposal to assist us (Bourdieu, 1984). That said, it is important that the habitus is not viewed as a concept which portrays the individual as a cultural dupe. Whilst the individual is immersed in the social world generally partaking in practices that are familiar to their habitus, there are significant moments when the individual is not at ease and it is these points that produce critical reflexivity for the individual. Not being familiar with the game can lead to feelings of being out of place – like a fish out of water (Bourdieu, 1990). Put another way, this can feel like we do not know what to do, how to think or feel, in a certain situation (see Addison, 2016) and it is in these critical moments that change and agency happen. However, it is this feeling of being uneasy with our surroundings, like we don't fit in, that Hochschild mis-identifies as a splitting of the self – a false self then where we put on a performance of what we think is expected of us (I was not being myself). However, this feeling of unease and conscious performance, I would argue, is connected to an awareness of a lack of knowledge of how to act and feel in a situation. This emotional dissonance then is not an argument to support the idea of a false self, but rather indicates a feeling of being out of place in a certain social space and around certain people.","It has been my intention here to show that feeling out of place, or not being ourselves in certain social situations, is not down to a dichotomy between a real and false self as Hochschild tries to unsuccessfully set up. Rather, I have argued here that this feeling of ‘not being myself’ arises because we do not possess the required knowledge dispositions or embodied practice, or have the right embodied history, in order to act comfortably in certain situations – therefore we are reflexive of our social position and feel uneasy as if we are not ourselves. Instead of the idea of a real self constrained by feeling rules, I have argued that individuals use their embodied histories as a way of understanding and making sense of the prevailing dominant symbolic structures. Acquisition of this knowledge of how things ought to be done is sedimented as dispositions over time, creating a personal history, which is then drawn upon to strategize future practice. This means then that there is no need to argue, as Hochschild does, about what version of the self is ‘real’ and what is ‘false’: fundamentally all aspects of one's self and performances are real. Some employees sense that their embodied histories don't fit well within social space. These people may find that they repeatedly have to adjust their practices, even when they feel uncomfortable, in order to fit in with a legitimated value system structuring how they ‘ought’ to be (see Bathmaker et al., 2013; Reay et al., 2009). Then again, those who find themselves feeling ‘out of place’ (Reay et al., 2009) may subvert structured feeling rules in the workplace, they may develop alternate ways of fitting in within or without the rules, they may collude with colleagues and share anger and humour as a strategy for dealing with exploitations of their emotional labour. The concept of embodied history works much better with these ideas of emotion management then than Hochschild's idea of the real self. It has explanatory power without theorising performances as ‘false’. I argue that it is more appropriate to critically consider why acting a certain way, and managing emotion, in a particular space can seem odd and make us feel ill at ease. I have suggested here that our embodied histories can feel out of place in the workplace because we may be used to a different way of doing things in our day-to- day lives (see Addison, 2012; 2016). The workplace can exploit our ability to shape ourselves into someone we are told we ought to be, and this can hold us in a state of being ill at ease. And so, entering a space that has dominating structured feelings rules which are different to our own embodied feeling rules obliges us to manage and display our feelings, perhaps in uncomfortable and unfamiliar ways. However, I would again reiterate that this is not evidence of a real and false self, but rather highlights the self- conscious feelings we may have in certain situations and around certain people."],["Objectives: Deficiencies in perceptual and cognitive functions have been linked with antisocial and aggressive behavior. To test whether these putative relationships generalize to sport – a context where such behavior is common – we determined the extent to which pain thresholds and cortical activity in response to painful electrical stimulation were associated with antisocial and aggressive behavior in sport; we also examined their link to moral disengagement. Design: A cross-sectional design was used. Method: Ninety-four participants completed questionnaires, had their pain threshold determined, and then had their central and frontal pain-related cortical activity recorded while they were electrically stimulated at supra-threshold intensity. Results: Subjective pain thresholds were positively related while pain induced frontal alpha power was negatively related to antisocial behavior and aggressiveness. Central pain evoked potential amplitudes were negatively related to aggressiveness and moral disengagement. Conclusions: Sensitivity to and cortical processing of noxious stimuli were reduced in individuals who more frequently behave antisocially and aggressively when playing sport and who are more likely to use psychosocial maneuvers to justify their harmful behavior. Our findings reveal that pain-related deficits are a feature of individuals who engage in more frequent antisocial and aggressive behavior in the context of sport. --------------------------------------------------------------------------------","Sport is a social context where moral issues are highly relevant (for reviews see Kavussanu, 2008, 2012). Research has shown that during competitive games, team sport players deliberately foul, physically intimidate, and try to injure their opponents (e.g., Kavussanu, Seal, & Phillips, 2006; Kavussanu, Stamp, Slade, & Ring, 2009). Thus, it is important to improve understanding the factors associated with antisocial and aggressive behavior, which encompasses acts intended to harm or disadvantage another individual (Kavussanu, 2012) and harm another individual (Anderson & Bushman, 2002), respectively. Although much research has examined antisocial and aggressive behavior in sport from a social psychological perspective, more recently researchers have begun to investigate this important topic from a cognitive neuroscience perspective (e.g., Kavussanu, Willoughby, & Ring, 2012; Micai, Kavussanu, & Ring, 2015). Research in non-sport contexts has revealed differences in how the brains of antisocial and aggressive individuals respond to sensory and cognitive demands compared to other individuals (for reviews see Blair, 2001; Volavka, 1990, 1999). For instance, these reviews discuss evidence that violent individuals are characterized by structural and functional abnormalities in their frontal and temporal lobes. We aimed to extend these findings to the sport context. In team sports that involve physical contact between players, such as association football, basketball, field hockey, and rugby, antisocial and aggressive behaviors are relatively common occurrences during games (Bredemeier & Shields, 1986; Kavussanu, 2012). Accordingly, the current study determined whether abnormal cortical processing and perception of pain is a feature of individuals who engage more frequently in antisocial and aggressive behavior when playing competitive team sport. Pain sensitivity ~~~~~~~~~~~~~~~~ Antisocial behavior and emotional detachment are the two key defining features of psychopathy (Blair, 2001). Early clinical observations noted that psychopaths often fail to avoid punishment (Cleckley, 1959; Hetherington & Klinger, 1964). Experimental research has since documented that psychopaths are characterized by impaired aversive conditioning (Flor, Birbaumer, Hermann, Ziegler, & Patrick, 2002; Hare & Quinn, 1971; Lykken, 1957), blunted conditioned anticipatory arousal prior to an impending noxious stimulus (Hare, 1965), reduced blink responses to noxious stimuli (Benning, Patrick, & Iacono, 2005; Patrick, Bradley, & Lang, 1993), and reduced pain sensitivity (Fedora & Reddon, 1993; Hare, 1968; Hare & Thorvaldson, 1970; Schalling, 1971; Schalling & Levander, 1964). Taken together, these data suggest that the increased frequency of antisocial behavior in psychopaths may be linked to their relative insensitivity to aversive stimuli. Further support for this proposal comes from studies showing that pain sensitivity is lower in aggressive and violent individuals (Niel, Hunnicut-Ferguson, Reidy, Martinez, & Zeichner, 2007; Reidy, Dimmick, MacDonald, & Zeichner, 2009; Seguin, Pihl, Boulerice, Tremblay, & Harden, 1996). Seguin et al. (1996) reported that boys with higher pain tolerance to pressure stimulation were characterized by increased history of physical aggression based on teacher reports. Niel et al. (2007) used the response choice aggression paradigm and found that males with higher pain tolerance to electrical stimulation administered higher intensity shocks and more maximal intensity shocks to their opponents. Similarly, Reidy et al. (2009) found that male (but not female) participants with higher pain tolerances scored higher on self-reported measures of verbal and physical aggression. Although the mechanism underlying this pain-aggression phenomenon has yet to be identified, a number of candidates have been mooted. It has been suggested that pain tolerant individuals may underestimate the degree of pain inflicted on their victims or may have been toughened up by frequent fights (Sequin et al., 1996). Based on this evidence, we tested the possibility that relative insensitivity to pain may be a feature of athletes who engage more frequently in antisocial and aggressive behavior when playing sport. In team contact sports, physical contact during competitive games can lead to unpleasant sensory and emotional experiences associated with tissue damage (i.e. pain). Antisocial behavior and aggression in team contact sports might be linked with pain sensitivity for various reasons: Pain tolerant athletes may be more likely to commit physical antisocial and aggressive acts because they cannot empathize with their victims (Stanger, Kavussanu, & Ring, 2012; Stanger, Kavussanu, Willoughby, & Ring, 2012) because of impaired cognitive perspective taking or emotional empathic concern and personal distress (cf. Sequin et al., 1996). Pain-related evoked potentials ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Researchers (e.g., Bromm & Lorenz, 1998) often supplement subjective reports of pain with its objective neurophysiological correlates to paint a more complete picture of the psychological and physiological processes implicated in the perception and processing of noxious stimuli. However, to our knowledge, no study has assessed cortical evoked potentials to painful stimuli to explore the central processes underlying the antisocial behavior–pain relationship. The electroencephalogram (EEG) represents a means of assessing cortical activity that involves the recording of electrical activity on the scalp to detect voltages generated inside the brain. Evoked potentials represent the cortical activity elicited in response to the presentation of an exteroceptive stimulus, such as a painful electrical stimulus. The most commonly studied pain-related evoked potentials are the N2 and P2 potentials, which refer to the second negative and positive peaks, respectively, of the cortical response to a noxious stimulus and represent the cortical activity that results from processing a painful stimulus (Treede, Kenshalo, Gracely, & Jones, 1999). These scalp potentials are measured at the vertex because they are reliably found to be largest in amplitude at this location. Pain-related evoked potentials reflect pain processing, that is increasingly painful stimuli elicit increasingly larger potentials (Bromm & Lorenz, 1998). It is possible that attenuated pain-related evoked potentials are associated with the tendency to commit antisocial and aggressive acts. Frontal cortical activity ~~~~~~~~~~~~~~~~~~~~~~~~~ There is evidence to suggest that frontal dysfunction, assessed using EEG, is a feature of aggressive individuals (Volavka, 1990). For instance, one study noted that violent behavior in psychiatric patients was negatively correlated with frontal alpha band EEG activity, particularly resting activity in the left hemisphere (Convit, Czobor, & Volavka, 1991). Similarly, brain imaging studies have implicated reduced prefrontal cortical activity (Raine, Buchsbaum, & LaCasse, 1997) and frontal lesions (Damasio, Grabowski, Frank, Galaburda, & Damasio, 1994; Grafman et al., 1996) with antisocial and aggressive behavior. These observations are compatible with the proposal that aggressive behavior is determined by a circuit in the brain comprising the orbitofrontal cortex, anterior cingulate, and amygdala (Davidson, Putnam, & Larson, 2000). In EEG studies, prefrontal cortical activity is typically indexed by the amount of activity in the alpha frequency band: A fast Fourier transform is applied to the raw EEG waveform to yield the spectral power of the EEG signal with a frequency of between 8 and 12 cycles per second. High alpha activity was originally interpreted as cortical idling (Pfurtscheller, Stancak, & Neuper, 1996), but more recently has been viewed as reflecting a sensory gating mechanism involving inhibition of task-irrelevant and activation of task-relevant areas (Jensen & Mazaheri, 2010; Schurmann & Basar, 2001). Although it is possible to assess frontal alpha brain activity under resting conditions, recent research has found better results using stimulus induced activity (e.g., Coan, Allen, & McKnight, 2006). Taken together, there is sufficient evidence to suggest that relatively attenuated pain-induced frontal brain activity may be associated with the tendency to behave antisocially in sport. Moral disengagement ~~~~~~~~~~~~~~~~~~~ Moral disengagement refers to the psychosocial mechanisms people use to minimize negative affect when they engage in transgressive behavior (Bandura, 1991; Boardley & Kavussanu, 2011). It allows individuals to engage in conduct that violates their personal standards without experiencing intense negative emotions that usually accompany such behavior. Moral disengagement operates by mentally reconstruing harmful behaviors into benign acts, minimizing personal accountability for harmful behavior, misrepresenting the injurious effects that result from such behavior, and blaming the nature or actions of the victim. Previous research has found that players who have the propensity to morally disengage are more likely to report engaging in antisocial behaviors toward other players (Boardley & Kavussanu, 2011). Given the link between blunted emotion and antisocial behavior in violent offenders and psychopaths (e.g., Blair, 2001; Cleckley, 1959), it is possible that moral disengagement may be associated with attenuated sensitivity and responses to painful stimulation. We could speculate on how moral disengagement might be linked with reduced pain. The distortion of consequences mechanism operates on the consequences of detrimental behavior and downplays the harm caused to victims: Individuals who minimize the harm they cause are more likely to repeat such actions (Bandura, 1999). Accordingly, athletes who feel little pain when they are hit, kicked or punched may also underestimate the seriousness of the injuries they cause (Boardley & Kavussanu, 2011), and, therefore are more likely to act aggressively (Niel et al., 2007; Reidy et al., 2009; Seguin et al., 1996). Some moral disengagement mechanisms operate on the agency of action by obscuring or minimizing one's role in the harm one causes (Bandura, 1999); these are displacement and diffusion of responsibility. It is possible that players who feel less pain are also those who obscure and minimize harm to others. With these possibilities in mind, the current study investigated the link between moral disengagement and pain. The present study ~~~~~~~~~~~~~~~~~ In sum, research has highlighted a relationship between pain and antisocial/aggressive behavior in non-athletes. However, to our knowledge, no study has examined the links between pain sensitivity or cortical processing of painful stimuli and antisocial behavior, aggressiveness, and moral disengagement in sport. We aimed to extend previous research by obtaining both self-reported and cortical measures of pain to examine whether pain is related to antisocial behavior and aggressiveness (i.e., the tendency to become aggressive, Maxwell & Moores, 2007) in sport. The first purpose of the study was to determine whether subjective pain thresholds, pain induced frontal alpha activity, and pain-related evoked potentials are associated with antisocial behavior and aggressiveness in sport. We expected that more frequent antisocial behavior and greater aggressiveness would be negatively associated with the subjective experience (i.e., higher pain thresholds) and cortical processing (i.e., less pain-induced frontal alpha activity, smaller pain-related evoked potential amplitudes) of pain. A second purpose was to determine whether pain (measured by subjective and objective methods) is related to moral disengagement in sport. We expected that moral disengagement would be negatively associated with self-reported pain and cortical measures of pain-related processing. Our hypotheses were tested using new analyses performed on an existing dataset (Kavussanu, Willoughby, & Ring, 2012).","Ninety-four team sport athletes (48 males, 46 females), with a mean age of 20.95 (SD = 2.72) years and 8.51 (SD = 4.31) years playing experience were paid £20 for participating. Their main team sport was association football (31%), field hockey (26%), basketball (22%) rugby (18%), and water polo (3%). All sports were contact sports. We recruited athletes from these sports because contact sports have inherent potential for injury and therefore higher likelihood for antisocial and aggressive behavior to occur and moral issues to arise (Bredemeier & Shields, 1986). Athletes from a variety of contact sports were recruited to increase the generalizability of our findings. All participants were free from neurologic and psychiatric disorders and medications. They were asked to refrain from alcohol, caffeine and smoking for at least 12 h prior to testing. Antisocial behavior Antisocial behavior in sport was measured using the 8-item antisocial behavior toward opponents scale of the Prosocial and Antisocial Behavior in Sport Scale (Kavussanu & Boardley, 2009; Kavussanu, Stanger, & Boardley, 2013). Participants were asked to rate how often they engaged in different behaviors when playing their team sport. An example item is “Tried to injure an opponent”. Each item was rated on a 5-point Likert scale, anchored by 1 (never) and 5 (very often). Kavussanu and Boardley (2009) reported very good internal consistency (α = .86) for this subscale. Aggressiveness Competitive aggressiveness was measured using the 6-item aggressiveness scale of the Competitive Aggressiveness and Anger scale (Maxwell & Moores, 2007). The stem “When playing your team sport how often have you behaved, felt or thought that … ” was followed by six items measuring aggressiveness. An example item is “Violent behavior directed toward an opponent is acceptable”. Each item was rated on a 5-point Likert scale, anchored by 1 (never) and 5 (very often). Maxwell and Moores (2007) have provided evidence for the factorial validity and reliability (α = .84) of this scale. Moral disengagement Moral disengagement in sport was measured using the 8-item Moral Disengagement in Sport Scale–short (Boardley & Kavussanu, 2008). Participants were asked to indicate their level of agreement with a range of statements concerning thoughts and feelings they may have in sport on a 7-point Likert scale, anchored by 1 (strongly disagree) and 7 (strongly agree). This scale includes an item for each of the eight mechanisms of moral disengagement. An example item is “It is okay to treat badly an opponent who behaves like an animal”. Boardley and Kavussanu (2008) reported very good internal consistency (α = .85) for the scale. Noxious stimulus ~~~~~~~~~~~~~~~~ A noxious electrical stimulus was delivered using a nociception-specific concentric electrode designed to selectively activate A-delta fibers (Katsarava et al., 2006; Kaube, Katsarava, Kaufer, Diener, & Ellrich, 2000). Each stimulus consisted of a double pulse. Each rectangular wave pulse lasted 500 μs separated by 100 μs; this delay is below the threshold required to discriminate the two stimuli, thus they were perceived as a single stimulus. The stimulating electrode was placed on the supraorbital nerve above the left eye (Cuzalina & Holmes, 2005), and a constant current stimulator (Model DS7A, Digitimer Ltd, UK) provided the stimulation. Pain threshold ~~~~~~~~~~~~~~ The subjective pain threshold was determined using a two-stage procedure. First, the participant rated the electrical stimulation on a 5-point scale adapted from Tursky and O'Connell (1972): 0 (feel no sensation), 1 (feel any sensation), 2 (uncomfortable sensation), 3 (painful sensation), 4 (don't want to go any higher). The stimulation was increased from .2 mA in steps of .2 mA until the stimulation became painful, that is the participant reported “3” on the scale. Using this as an initial estimate, the participant's pain threshold was determined using an up–down staircase procedure: Stimulus intensity was increased by .1 mA if the prior stimulation was rated as not painful by the participant, or decreased by .1 mA if it was rated as painful (Levitt, 1971). When the participant had reported that the stimulation was painful on three non-consecutive occasions, the pain threshold was calculated as the average of the last two peak stimulation values. These procedures have been used in previous research (cf., Wilkinson, McIntyre, & Edwards, 2013). Pain induced alpha power and pain related evoked potentials ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A noxious electrical stimulus, at an intensity corresponding to 125% of pain threshold (M = 1.68, SD = 1.08 mA), which was perceived as a single pinprick-like pain, was delivered while participants viewed a small black fixation cross on a white screen on six trials. On other trials (data not reported here), it was delivered while participants viewed a picture. The pictures were neutral (e.g., players standing or moving), pleasant (e.g., players celebrating, semi-naked players), and unpleasant (e.g., players being hurt or badly injured) in valence (for further details see Stanger, Kavussanu, Willoughby, et al., 2012). Electrical stimulation occurred every 20–25 s. The EEG was recorded using a 32-channel BioSemi ActiveTwo system (BioSemi, Netherlands), at 512 Hz, and was re- referenced to average earlobe electrodes offline. Electrophysiological data processing was performed using EEGLAB (Delorme & Makeig, 2004). The EEG was high-pass filtered using a finite impulse response windowed-sinc filter with a half-amplitude cut-off at 1 Hz and a .4 Hz transition band. We performed a fast Fourier transform (1 Hz bins) on the artifact- free epochs, and then computed power (dB) in the alpha (8–12 Hz) frequency band at left and right frontal (F3 and F4) sites in the seconds before and during the two second window following onset of painful stimulation. These values were then log-transformed and averaged across sites. The key components of the pain-related evoked potential (see Fig. 1) are the amplitudes (measured in microvolts) of the second negative (N2) and positive (P2) peaks in the event related potential following painful stimulation (e.g., Edwards, Inui, Ring, Wang, & Kakigi, 2008). Thus, we calculated the N2 (in the 100–200 ms post- stimulation window) and P2 (in the 200–300 ms post-stimulation window) peaks. We focused on the Cz electrode, since this is where the N2 and P2 potentials are maximal (e.g., Katsarava et al., 2006). This was confirmed in the present study by examination of the scalp maps (see Appendix 1). The amplitude of the pain evoked potentials at the vertex (i.e., Cz) were measured relative to a 100 ms pre-stimulus baseline and calculated as the average amplitude of the seven data points around the peak value during the 100–200 ms time window for the N2 potential and the 200–300 ms time window for the P2 potential (i.e., the peak value and the three data points either side, corresponding to a window of approximately 12 ms).","The study protocol was approved by the local research ethics committee and each volunteer gave informed consent to participate. At the start of the testing session, participants completed the self-reported measures of antisocial behavior, aggressiveness and moral disengagement. Following instrumentation, their pain threshold was determined. After sitting quietly for five minutes, they were instructed about the next task: They were told to always focus on the screen located in front of them and that their forehead would be stimulated when either a fixation cross or picture was on screen. In the task, the participant's pain related evoked potentials and pain induced frontal alpha activity were recorded as described above. Descriptive statistics and alpha coefficients ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The means and standard deviations of the scales that measured antisocial behavior, aggressiveness, and moral disengagement are presented in Table 1. This table also presents alpha coefficients of the variables used in this study; internal consistency of all scales was good. On average, players reported that, when playing sport, they behaved antisocially toward opponents rarely or sometimes and reported aggressiveness rarely. They also reported moderate levels of moral disengagement. These scores are in line with those reported in previous research (e.g., Boardley & Kavussanu, 2008; Kavussanu & Boardley, 2009; Stanger, Kavussanu, Boardley, & Ring, 2013; Stanger, Kavussanu, & Ring, 2012; Stanger, Kavussanu, Willoughby, et al., 2012). Pearson correlations showed that moral disengagement was positively related to antisocial behavior, r(92) = .48, p < .001, and aggressiveness, r(92) = .67, p < .001. The means and standard deviations of the subjective pain threshold, pain induced frontal alpha power, and pain evoked potential amplitudes are also shown in Table 1. The pain threshold for the noxious trigeminal stimulus is compatible with prior research (e.g., Katsarava et al., 2006). The scalp map for the pain related evoked potential confirmed that the N2 and P2 pain evoked potentials were maximal at the central electrode site Cz (see Appendix 1). We conducted a series of one-sample t-tests to determine whether the pain related evoked potentials and pain induced alpha activity were significantly different from zero. These tests confirmed significant N2, t(93) = 16.07, p < .001, and P2, t(93) = 19.43, p < .001, pain evoked potentials (see Fig. 1). A one-sample t-test also confirmed that the pain induced frontal alpha activity was greater than zero, t(93) = 5.74, p < .001. Frontal alpha was lower in the seconds after noxious stimulation compared to the seconds before stimulation, t(93) = 5.04, p < .001, with the decrease averaging −.96 (SD = 1.85) dB. Correlation analysis ~~~~~~~~~~~~~~~~~~~~ Pearson correlations were computed between the pain variables (pain threshold, pain- induced frontal alpha, pain-related evoked potentials) and antisocial behavior, aggressiveness, and moral disengagement (see Table 1). The pain threshold was positively related to both antisocial behavior and aggressiveness: Players who acted more antisocially when playing team sport and players who reported more aggressiveness in sport tended to be less sensitive to noxious trigeminal stimulation. The pain threshold was not significantly related to moral disengagement. Frontal alpha power associated with the processing of noxious electrical stimulation was negatively associated with antisocial behavior and aggressiveness. These findings indicate that the players who acted more antisocially when playing team sports and players who displayed more aggressiveness in sport were characterized by less frontal alpha activity when exposed to painful stimuli. The negative N2 potential was positively related to moral disengagement, whereas the positive P2 potential was negatively related to moral disengagement and aggressiveness. In brief, smaller pain related evoked potential amplitudes were a feature of players who reported higher moral disengagement and aggressiveness in sport. Gender as a moderator ~~~~~~~~~~~~~~~~~~~~~ Reidy et al. (2009) reported that the pain–aggression relationship was moderated by gender, with the effect evident for males but not females. To investigate this possibility in our study, we conducted moderation analysis using bootstrapping (Preacher & Hayes, 2008) and PROCESS for SPSS Release 2.13 (Hayes, 2013) to examine whether gender moderated the relationships between the pain variables (pain threshold, pain-induced frontal alpha, pain-related evoked potentials) and antisocial behavior, aggressiveness, and moral disengagement. Bootstrapping was set at 5000 samples with bias corrected 95% confidence intervals; an effect was significant when the Confidence Interval (CI) did not contain zero. Results of these analyses indicated that the associations were not moderated by gender, ts = .03–1.45, ps = .15–.97, with one exception. Gender moderated the relationship between moral disengagement and the N2 component of the pain-related evoked potential, b = 5.566, 95% CI = .237, 10.895; t = 2.08, p = .04. Moral disengagement was associated with reduced N2 potential in males, b = 4.113, 95% CI = .790, 7.436; t = 2.46, p = .02, but not females, b = −1.453, 95% CI = −5.619, 2.713; t = .69, p = .49.","Deficiencies in pain processing have been linked with antisocial and aggressive behavior in diverse populations in non-sport contexts. To test whether these findings generalize to sport, we examined the extent to which pain thresholds and cortical activity in response to painful electrical stimulation were associated with antisocial behavior, aggressiveness and moral disengagement in the context of sport. Athletes' subjective pain threshold was positively related while pain induced frontal alpha power was negatively related to both antisocial behavior and aggressiveness. Moreover, central pain evoked potential amplitudes were negatively related to aggressiveness and moral disengagement. Thus, sensitivity to and cortical processing of noxious stimuli were reduced in both male and female athletes who behaved more antisocially, displayed more aggressiveness, and were more prone to morally disengage when playing competitive sport. Pain sensitivity ~~~~~~~~~~~~~~~~ In support of our hypothesis, relative insensitivity to painful electrical stimulation was a feature of athletes who engaged more frequently in antisocial conduct and who were more accepting of and willing to be aggressive when playing sport. This finding is compatible with the available literature in other contexts documenting reduced pain sensitivity in aggressive and violent individuals (e.g., Niel et al., 2007; Reidy et al., 2009; Seguin et al., 1996). In contrast to Reidy et al. (2009) gender did not moderate the pain–aggressiveness relationship. To date, no mechanism has been identified to account for these findings. Reduced pain sensitivity could lead to misattributions related to pain inflicted on a victim, whereby perpetrators who are less sensitive to and more tolerant of pain may misperceive the degree of pain experienced by another and, therefore, may be more willing to use violence during interpersonal conflict (Niel et al., 2007). Thus, athletes could engage in antisocial acts because they do not believe they are hurting their opponents as much as they really are. Our findings resonate with Niel et al.'s (2007) conclusion that insensitivity to pain (high pain tolerance) may increase the likelihood of aggression during an interaction where competition and provocation occur. Provocation occurs in competitive team sport, and therefore, aggressive competitors may not appreciate the consequences of their actions for their opponents when playing sport. The current data provide preliminary evidence to support the proposal that relative insensitivity to pain is related to the frequency with which players engage in antisocial behavior when playing sport because regular perpetrators may underestimate the amount of pain inflicted on their victims. Pain-related evoked potentials ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We found that the N2 and P2 pain related evoked potentials were blunted in athletes reporting high levels of moral disengagement, while the P2 component of this potential was also blunted in athletes reporting high acceptance of aggression and willingness to be aggressive in sport. Moral disengagement mechanisms allow players to engage in transgressive conduct without experiencing strong negative emotions, such as guilt. We found that players who used more moral disengagement were more likely to report engaging in antisocial behaviors toward other players, in line with past studies (e.g., Kavussanu, Ring, & Kavanagh, 2015; Stanger et al., 2013; Stanger, Kavussanu, & Ring, 2012). Since the amplitudes of pain evoked potentials reflect pain processing, the current findings suggest that players who are more aggressive and use psychosocial maneuvers to justify their harmful behavior are more likely to have cortical deficits in how they respond to painful stimulation. Extending the previous behavioral research linking pain insensitivity to antisocial and aggressive behavior (e.g., Fedora & Reddon, 1993; Hare, 1968; Hare & Thorvaldson, 1970; Niel et al., 2007; Reidy et al., 2009; Schalling, 1971; Schalling & Levander, 1964; Seguin et al., 1996), the current study indicated that relatively attenuated pain evoked potentials are associated with the tendency to accept and excuse aggressive acts when playing sport. Pain-induced frontal cortical activity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In line with our prediction, lower pain induced frontal brain activity was linked with increased frequency of antisocial behavior and greater acceptance of and willingness to aggress in sport. This finding is the first to link changes in frontal alpha to morally- relevant behavior in sport and is compatible with previous evidence showing that frontal dysfunction, assessed using EEG, is a feature of aggressive individuals (for review, see Volavka, 1990). In line with previous research (see Peng, Babiloni, Yanhui, & Hu, 2015), frontal alpha was suppressed in response to acute noxious stimulation, presumably reflecting the effects of a pain-related gating mechanism on frontal areas (Jensen & Mazaheri, 2010). Thus, antisocial behavior and aggressiveness were related to relatively low alpha oscillatory activity in the context of suppressed activity in the frontal regions of both hemispheres (cf. Convit et al., 1991). These results are also broadly compatible with brain imaging studies that have found a link between antisocial and aggressive behavior and prefrontal cortical activity (Raine et al., 1997) and frontal lesions (Damasio et al., 1994; Grafman et al., 1996); they are also in line with the suggestion that aggressive behavior is determined by a circuit in the brain comprising the orbitofrontal cortex, anterior cingulate, and amygdala (Davidson et al., 2000). In sum, by assessing pain induced frontal activation, we showed that reduced frontal alpha activation induced by noxious electrical stimulation is associated with the tendency to engage in harmful conduct when playing sport. Study limitations and future research directions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our study yielded some novel findings but also has some limitations which should be noted when interpreting the findings. First, only one stimulation site and one mode of stimulation were used, leaving open the possibility that different effects may be a feature of painful stimulation at other sites and other forms of painful stimulation. Accordingly, the generalizability of the observed effects from trigeminal nociception to other aspects of nociceptive processing remains to be confirmed in future research. Second, we employed a cross-sectional design and therefore causal relationships cannot be determined. Research using longitudinal and intervention designs that incorporate direct observations of behavior during sport participation and other assessments sensorimotor processes are now required. Finally, the correlation coefficients displayed in Table 1 ranged from r = .20 to .28, which, based on Cohen's (1992) definitions, where .10 is small, .30 is medium, and .50 is large, can be considered medium-to-small effects. These effect sizes should be used when interpreting the extent of the link between pain and antisocial behavior/aggressiveness. Nonetheless, this study has several strengths, including the use of a validated nociception-specific concentric electrode designed to selectively activate nociceptive afferents, state of the art equipment for measuring electroencephalographic signals, and a large sample size.","In conclusion, our findings revealed a consistent pattern indicating that players who perform and justify transgressive acts are characterized by relative insensitivity to noxious stimulation. Building on previous research showing that increased pain tolerance is associated with increased aggression (e.g., Niel et al., 2007; Reidy et al., 2009; Seguin et al., 1996), the current study is the first to confirm the pain–behavior relationship in objective neurophysiological indices of pain processing as well as a subjective measure of pain sensitivity. It may be worth noting that antisocial behavior seemed to be more closely associated with subjective pain threshold and alpha power whereas moral disengagement was more closely linked with cortical processing of pain. The reason for these differential relationships between moral variables and indices of pain sensitivity cannot be established from the present study. Accordingly, future research is now needed to establish the mechanisms underlying the pain-behavior and pain–cognition relationships. Our findings suggest that a profile comprising insensitivity to pain and blunted cortical responses to noxious stimulation may serve as a biobehavioral risk marker for antisocial and aggressive athletes. In sum, relative pain-related deficits are more likely to be a feature of individuals who engage more frequently in and excuse transgressive conduct in sport."],["We report the results of two experiments which test the potential of arts engagement for promoting prosocial intentions. Experiment 1 (N = 216) tested the impact of a participatory arts intervention (vs. a control condition) on children's empathy and interpersonal prosocial intentions. Experiment 2 (N = 174) tested the impact of a participatory arts intervention (vs. a control condition) on children's prosocial intentions toward outgroup members under competitive and non-competitive conditions. Experiment 1 showed that the participatory arts intervention significantly increased children's interpersonal prosocial intentions, but not their empathy. Experiment 2 showed that, under competitive conditions, the participatory arts intervention significantly increased prosocial intentions toward outgroup members, an effect that persisted for six months beyond the intervention. Under non-competitive conditions, the participatory arts intervention consolidated improvements in prosocial intentions toward outgroup members. Overall, the results confirm the hypothesis that participatory arts engagement can promote prosocial intentions during middle childhood. --------------------------------------------------------------------------------","Prosocial responses can broadly be understood as those aimed at benefitting others. Examples of prosocial responses include helping, sharing with, and caring for others (Abrams, Van de Vyver, Pelletier, & Cameron, 2015). Research shows that humans frequently and intuitively engage in prosociality (Crockett, Kurth-Nelson, Siegel, Dayan, & Dolan, 2014). During early childhood, children develop the ability to accurately take the perspective of another (e.g., in terms of goals, wants, needs and desires), to understand another's negative emotional state (empathy), to recognise when a goal is unfulfilled, and to detect the source of a problem (Dunfield, 2014; Paulus, 2018). These antecedents of prosociality as well as prosocial responses themselves become integrated and well- established by middle childhood (Dunfield, 2014). Experiences of, and engagement in, prosociality are essential for personal and societal wellbeing. For example, engagement in prosociality improves personal wellbeing among children (Flouri & Sarmadi, 2016), adolescents (Layous, Nelson, Oberle, Schonert-Reichl, & Lyubomirsky, 2012), and adults (Nelson, Layous, Cole, & Lyubomirsky, 2016). Human prosociality and cooperation are also essential for tackling societal problems such as environmental degradation, humanitarian crises, and inequality. Therefore, it is essential to understand the conditions that can promote greater prosocial engagement. Intergroup prosociality ~~~~~~~~~~~~~~~~~~~~~~~ Importantly, children from the age of three develop an awareness of social categories, and by age five they are more likely to help ingroup compared to outgroup members (Abrams et al., 2015; Nesdale, 2004; Over, 2018; Sierksma, Thijs, & Verkuyten, 2015). This intergroup bias in prosociality is evident across the lifespan (Levine, Cassidy, Brazier, & Reicher, 2002; Stürmer, Snyder, & Omoto, 2005), and has detrimental effects on societies across the world. Intergroup bias has a clear developmental trajectory. Specifically, intergroup bias tends to increase gradually from as young as three years and to peak during middle childhood (Raabe & Beelmann, 2011). Moreover, the extent to which children display intergroup bias, and the age at which they express intergroup bias, varies depending on the complexity of the intergroup context (Rutland, Killen, & Abrams, 2010). For example, when deciding whom to include in a club, children are more likely to justify inclusion on the basis of group membership and stereotypes when there is only space for one child (and two peers from different groups want to join) compared to when there are sufficient spaces for all children (Killen, Pisacane, Lee-Kim, & Ardila-Rey, 2001). In other words, children show less intergroup bias in straightforward (vs. complex) intergroup contexts and this impact of situational complexity on intergroup bias appears from as young as four years of age (Killen et al., 2001). Similarly, Abrams et al. (2015) demonstrated that children are less likely to help outgroup members in a competitive (i.e., complex) context than in a non-competitive (i.e., straightforward) context. Given that these types of situational complexities are inherent across societies and given that prejudice against minority groups is often associated with perceptions that they are competing for resources (Cuddy, Fiske, & Glick, 2007; Van de Vyver, Leite, Abrams, & Palmer, 2018), it is important to understand whether and how we can promote outgroup prosociality under competitive as well as non-competitive contexts.","Artistic practices transcend geographic and historic boundaries, and it has been contended that artistic expression is part of an evolutionary mechanism for creating and maintaining social ties within humans (Pearce, Launay, & Dunbar, 2015). Any person in any part of the world can engage in the arts in one way or another and can hence establish shared meaning through the experience or creation of arts. Arts cover a broad and inclusive range of activities where creativity and self-expression are key (Broadwood, Bunting, Andrews, Abrams, & Van de Vyver, 2012). Arts offer opportunities to express and share viewpoints, feelings, ideas, stories, and values. Arts can build connections between artists and audiences, as well as within audiences and participants. When people engage with the arts they are creating meaning for themselves. Collaborative arts projects in particular enable people to make sense of the world together (cf. Broadwood et al., 2012). Some might argue that arts engagement activates mental simulation which “involves mentally transcending the ‘here-and-now’ to occupy psychologically a different time (past or future), a different place, a different person's subjective experience, or a hypothetical reality. In other words, simulation involves conjuring up the experience of something other than that which one is currently experiencing” (Waytz, Hershfield, & Tamir, 2015, p. 337). There are two relevant conceptual frameworks (Broadwood et al., 2012; Tay, Pawelski, & Keith, 2018) which aid our understanding of the socio-emotional impacts of the arts. Specifically, Tay et al.'s (2018) recent model proposes that the arts can promote wellbeing, broadly construed to include prosocial behavior. They suggest that arts engagement can produce four groups of outcomes which include (1) immediate neurological, physiological, and psychological outcomes, (2) enduring socio-cognitive and psychological competences (e.g., self-efficacy, creativity), (3) physical and psychological wellbeing, and (4) positive normative outcomes (e.g., values, morality, and civic engagement). Tay et al. (2018) identify four psychological processes through which arts engagement can affect these outcomes which include: immersion (“feeling carried away”), embeddedness (building socio- cognitive competencies), socialization (creating connections and identities), and reflectiveness (socio-moral reflection). Based on an extensive review of the literature (Broadwood et al., 2012), the Arts and Kindness model proposes that the arts have the potential to act as a social psychological catalyst for promoting human prosociality. The model proposes that there are four key routes through which arts engagement can promote prosociality: emotion (somewhat akin to immersion), learning (akin to embeddedness), values (akin to reflectiveness), and social connection (akin to socialization). Broadwood et al. (2012) and Tay et al. (2018) both propose that arts engagement has the potential to promote prosociality. Both models also emphasize that routes from specific or more general arts engagement can include short or longer term influences, be proximal or distal from particular events, and may be weighted differently depending on the particular art forms or context. Empathy ~~~~~~~ Empathy can be defined as an emotional reaction elicited by and congruent with another's emotional state or condition (Eisenberg & Fabes, 1998). Both Tay et al. (2018) and Broadwood et al. (2012) argue that the arts have a strong potential to promote empathy. Indeed, many arts activities will naturally align individuals into states of togetherness and will transport participants into the artists', the protagonists', or even fellow participants' lived or imagined experiences (Tay et al., 2018). Such joint states of togetherness and shared perspective facilitate the capacities needed for empathy (Rabinowitch, Cross, & Burnard, 2013). Empirical evidence in middle childhood indeed demonstrates that engagement in drama (Goldstein & Winner, 2012) and in music (Rabinowitch et al., 2013) promote empathy compared to control conditions. Moreover, it is well- established that empathy is crucial for building positive interpersonal and intergroup relationships (e.g., Abrams et al., 2015; Fabes, Eisenberg, & Eisenbud, 1993). Therefore, and in line with Broadwood et al. (2012), we hypothesize that empathy may be an important mechanism in explaining the relationship between arts engagement and prosocial intentions in middle childhood. Adults ~~~~~~ The relatively small body of psychological research that has examined the impact of arts engagement on prosocial outcomes is promising. Among adults different studies of specific art forms (e.g., singing, dancing, reading, acting) have provided evidence that engagement with that particular art form can promote empathy (Mar, Oatley, Hirsh, dela Paz, & Peterson, 2006) or prosocial responses (Greitemeyer, 2009; Johnson, Cushman, Borden, & McCune, 2013; Wiltermuth & Heath, 2009). Using a representative and longitudinal sample of over 30,000 adults in the UK, Van de Vyver & Abrams (2017) established a reliable and substantively meaningful longitudinal relationship between arts engagement and subsequent prosociality, even when accounting for socio-demographic variables, income, and personality differences. Children ~~~~~~~~ Among children, joint music making has been shown to promote within-group prosociality among 4-year olds (Kirschner & Tomasello, 2010) and interpersonal prosociality among 8–9 year olds (Schellenberg, Corrigall, Dys, & Malti, 2015). Tangentially, engagement in synchronous movement promotes interpersonal prosociality during early and middle childhood (Cirelli, Einarson, & Trainor, 2014; Rabinowitch & Meltzoff, 2017a; Rabinowitch & Meltzoff, 2017b; Tunçgenç & Cohen, 2018) and intergroup bonding in middle childhood (Tunçgenç & Cohen, 2016). Given that children seem to naturally and readily engage in artistic activities, and are often asked to do so at school, it is surprising that there is relatively little research examining its impacts on their socio-emotional development. Specifically, the potential social benefits of arts engagement have rarely been researched in developmental and social psychology (see Goldstein, Lerner, & Winner, 2017; Van de Vyver & Abrams, 2017), and therefore represents an important area of investigation for applied developmental research.","The current paper builds on these related strands of research. Specifically, we test a number of novel research questions. Separate studies have shown that engagement in specific art forms (e.g., singing) promotes empathy and interpersonal prosociality, but research has not tested whether engagement in a participatory arts intervention also promotes empathy and interpersonal prosociality. In Study 1 we test the hypothesis that engagement in a participatory arts intervention also promotes empathy and interpersonal prosocial intentions. Moreover, while recent research has shown that synchronous movement promotes bonding with outgroup members (Tunçgenç & Cohen, 2016), no research has examined whether engagement in the arts can promote outgroup-targeted prosociality. Study 2 extends past research by testing whether participatory arts engagement can promote outgroup- targeted prosocial intentions across competitive and non-competitive contexts. Notably, past research has examined the immediate impact of some types of arts engagement on prosociality, but it has not tested whether effects endure. Study 2 tests the hypothesis that the impact of arts engagement on children's prosocial intentions persists over a period of 6 months. In summary, across two field studies using experimental and longitudinal designs, we test the impact of participatory arts interventions on children's prosocial intentions. In Study 1 we employ a 2 (Condition: experimental vs. control) × 2 (Time: pre-intervention vs. post-intervention) mixed model design and measure effects on children's empathy and interpersonal prosocial intentions. In Study 2 we employ a 2 (Condition: experimental vs. control) × 3 (Age: 5–6 vs. 7–8 vs. 9–10 years) × 3 (Time: pre-intervention vs. one month post-intervention vs. six month post-intervention) mixed model design and measure effects on children's outgroup prosocial intentions in a competitive and in a non-competitive context. Interpersonal prosociality is well- established by middle childhood (Dunfield, 2014) and therefore we do not expect or explore age-related differences within the age range in Study 1. In contrast intergroup bias is known to increase and become more context sensitive during middle childhood (Raabe & Beelmann, 2011; Rutland et al., 2010), so it is possible that age may moderate the hypothesized impacts of arts engagement on outgroup-targeted prosocial intentions. This is tested in Study 2. Participants and design A-priori power analysis revealed that to detect a small to medium sized effect (η2 = .03) with 90% power for a 2 × 2 mixed model ANOVA design, we required a total sample size of 122 participants. Two hundred and sixteen children (102 male, 114 female) completed the study1. Children were aged between 7 years and 10 years (grades 2, 3, and 4; mean age = 8.20, SD = 0.86). The study used a 2 (Condition: experimental vs. control) between participants × 2 (Time: pre-intervention vs. post-intervention) within participants design. Participants were sampled from across three demographically and geographically matched elementary schools in the UK. Overall, 96% of participants were born in the UK and this was consistent across condition (96% in the experimental condition and 95% in the control condition). One of the schools experienced the intervention (N = 140). The other two served as control schools (N = 76). The study consisted of two testing times which were approximately twelve months apart, 60 weeks prior to and 1–2 weeks after the intervention.","Procedure In the experimental condition children in grades 2, 3, and 4 engaged in a participatory arts program which was administered at school and during the school day by an independent arts organisation. Participatory arts involve engagement in a range of art forms, that inherently include the audience in the creative process, allowing them to become co-authors, editors, and observers of the work. Arts Council England define participatory arts as follows: “Participation is a malleable dialogue that informs the work of the artist, builds and develops audiences, engages with communities, promotes learning and forges routes into active experience and artistic creation of many kinds. Participatory arts are now mainstream and are central to the core programme of many large arts organisations” (Arts Council England, 2010). In the experimental condition (or intervention school) local artists worked with every pupil and staff member over a period of one week to document and celebrate good news stories. Many schools already have programs in place to discuss and try to promote prosocial attitudes and social engagement (Education Commission of the States, 2016; UK Government, 2018). However, the present interventions were not prescriptive to convey a moral message. Instead, local artists explored the potential of using creativity and arts activities to help children to engage with and express stories of kindness. Example activities included: writing a song, producing art installations, producing story books, and making community artboards. The control schools only exposed children to routine curriculum-based discussions of kindness. Children in grades 2, 3, and 4 were invited to complete the pre and post questionnaires. Empathy We employed the 10-item measure of children's empathy used by Abrams et al. (2015). Example items are: “I get upset when I see someone get hurt” and “seeing someone who is crying makes me feel like crying”. Children responded from 1 (big frown) to 5 (big smile). Reliability analysis revealed low Cronbach's alphas for this scale (.57 at Time 1 and .44 at Time 2). We conducted a follow-up factor analysis with the 10 empathy items (using maximum likelihood to extract one factor). Following guidelines from Comrey and Lee (1992) we retained only items with a “fair” factor loading (.45 or higher). The final scale consisted of three items (all loadings were above .52 at Time 1 and above .56 at Time 2). The three retained items were: “I get upset when I see someone getting hurt”, “Seeing someone who is crying makes me feel like crying”, and “It makes me sad when I see someone who can't find anyone to play with”. Cronbach's alphas were: .64 (Time 1) and .69 (Time 2). The three items were mean scored within each time point. Interpersonal prosocial intentions We adapted and extended Abrams et al.'s (2015) measure of prosocial intentions. Specifically, children were told, “Imagine you are playing at the park and there are lots of children there”. Six scenarios were then introduced to assess children's willingness to help, share, and comfort (e.g., “While you are playing one of the other children comes over to you. The child has nothing to play with and asks if you will share some of your toys. Would you share your toys with the child?” ; “Some children are making fun of another child and the child is getting upset. Would you go over and comfort the child?”). Participants responded from 1 (definitely not) to 5 (definitely would). The six items were mean scored within each timepoint. Cronbach's alphas were .78 (Time 1) and .82 (Time 2). Descriptives Means, standard deviations, 95% confidence intervals, and bivariate correlations are presented in Table 1. Empathy A mixed model ANOVA with Condition (experimental vs. control) as a between participants factor and Time (pre-intervention vs post-intervention) as a within participants factor, revealed no significant effects of Time, F (1, 213) = 0.001, p = .975, η2 < .01, Condition, F (1, 213) = 2.54, p = .113, η2 = .01, or the Time x Condition interaction, F (1, 213) = 2.81, p = .095, η2 = .01. Prosocial intentions The ANOVA revealed a significant main effect of Time, F (1, 214) = 5.49, p = .020, η2 = .03, a significant main effect of Condition, F (1, 214) = 6.77, p = .010, η2 = .03, and a significant Time x Condition interaction, F (1, 214) = 3.95, p = .048, η2 = .022. Pairwise comparisons showed that, in the experimental condition, prosocial intentions significantly increased following the intervention (M T2 = 3.98, SE T2 = .07, 95% CI [3.85, 4.11]) compared to baseline levels (M T1 = 3.67, SE T1 = .08, 95%CI [3.52, 3.82]) (p T1 vs. T2 < .001). In contrast, in the control condition, prosocial intentions did not differ between Time 1 (M T1 = 4.06, SE T1 = .10, 95%CI [3.86, 4.26]) and Time 2 (M T2 = 4.08, SE T2 = .09, 95%CI [3.91, 4.26]) (p T1 vs. T2 = .826). Pairwise comparisons also showed that, at Time 1, prosocial intentions were significantly higher in the control condition than in the experimental condition (p = .002). However, as prosocial intentions increased in the experimental condition, there were no longer any differences between the control and experimental conditions at Time 2 (p = .348). Participants and design The study used a 2 (Condition: experimental vs. control) x 3 (Year Group: 5–6 vs. 7–8 vs. 9–10 years) between participants x 3 (Time: pre-intervention vs. one month post-intervention vs. six month post-intervention) within participants design. A-priori power analysis revealed that to detect a small to medium sized effect (η2 = .03) with 90% power for a 2 × 3 × 3 mixed model ANOVA design the study required a total sample size of 162 participants. 174 children (74 male, 100 female) completed the study3. Children were aged between 5 and 10 years (mean age = 7.32, SD = 1.62) and were in either kindergarten, Grade 2, or Grade 4 in elementary school. Participants were sampled from across five demographically and geographically matched elementary schools in the UK. Overall, 87% of participants were born in the UK and this was consistent across condition (90% in the experimental condition and 82% in the control condition). Three of the schools were intervention schools (N = 105). Two of the schools were control schools (N = 69). The study consisted of three testing times. The first (pre-intervention) took place just before the intervention started, the second took place approximately one month after the intervention ended, and the third took place approximately 6 months after the intervention. Procedure As in Study 1, all pupils in the experimental schools took part in a participatory arts program during school time. However, only children in kindergarten, Grade 2, or Grade 4 completed the questionnaires. In the experimental condition local artists worked with every pupil and staff member to document and celebrate acts of kindness. Over a period of seven months, local artists (different to those in Study 1) explored the potential of using creativity and arts activities to help children to engage with and express stories of kindness. Example activities included: painting, producing comic books, contributing to exhibitions, and producing a public performance. Children also went out into the community to interview people to collect stories of kindness. For example, one school visited a local fire department and brought them drinks and biscuits. As in Study 1, the control schools did not receive any form of intervention beyond their routine curriculum-based discussions of kindness. Competitive outgroup prosocial intentions We employed Abrams et al.'s (2015) measure of competitive outgroup prosocial intentions. Specifically, children were told about a sandcastle competition involving teams from their own and another fictitious nearby school. The team that built the biggest and best sandcastle would win a big prize and trophy. Children were then asked three questions to measure their intentions to share, help, and comfort the child (e.g., “As you are building your team's sandcastle, you see a child from the other team running to pick up a spade. He falls down and begins to cry. You could go over to comfort him, but your team needs you to keep building the sandcastle. Would you go over and comfort him?”). Participants responded from 1 (definitely not) to 5 (definitely would). The three items were mean scored within each timepoint. Cronbach's alphas were as follows: .62 (Time 1), .67 (Time 2), and .73 (Time 3). Non-competitive outgroup prosocial intentions We employed Abrams et al.'s (2015) measure of non-competitive outgroup prosocial intentions. Specifically, children were told, “a few weeks later, you are playing together in the park and there are lots of children there, including some from your school and [another local] school”. Three new items were then introduced to assess children's willingness to help, share, and comfort (e.g., “Some children are making fun of a boy from [other local school] and the child is getting upset. The children leave and he begins to cry. Would you go over and comfort the child?”). Participants responded from 1 (definitely not) to 5 (definitely would). The three items were mean scored within each timepoint. Cronbach's alphas were as follows: .62 (Time 1), .75 (Time 2), and .70 (Time 3). Descriptives Means, standard deviations, 95% confidence intervals, and bivariate correlations are presented in Table 2. Competitive outgroup prosocial intentions We conducted a mixed model ANOVA with Condition (experimental vs. control) and Year Group (5–6 vs. 7–8 vs. 9–10 years) as between participants factors and Time (pre-intervention vs. one month post-intervention vs. six month post-intervention) as a within participants factor. There were no significant main effects of Time, F (2, 336) = 1.69, p = .186, η2 = .01, Condition, F (1, 168) = 1.71, p = .193, η2 = .01, or Year Group, F (2, 168) = 2.00, p = .138, η2 = .02. The Time x Year Group interaction was also non-significant, F (4, 336) = 0.71, p = .589, η2 = .01. However, the significant interactions between Time x Condition F (2, 336) = 3.78, p = .024, η2 = .02, and Year Group x Condition, F (2, 168) = 3.07, p = .049, η2 = .04, were qualified by a significant three-way (Time x Condition x Year Group) interaction, F (4, 336) = 2.98, p = .019, η2 = .03 (see Fig. 1 for means and standard errors). Baseline comparison: Pairwise comparisons showed that there were no baseline differences in competitive outgroup prosociality between the experimental and control conditions among 5–6 year olds (p = .776), 7–8 year olds (p = .610), or 9–10 year olds (p = .708). 5–6 year olds: Pairwise comparisons revealed that, among 5–6 year olds, competitive outgroup prosocial intentions did not significantly change across time in the experimental (all p's > .433) or in the control (all p's > .176) conditions. 7–8 year olds: Pairwise comparisons revealed that, among 7–8 year olds, competitive outgroup prosocial intentions increased following the intervention compared to the baseline, and stayed at the higher level 6 months later (p T1 vs. T2 = .028; p T1 vs. T3 = .014). In contrast, in the control condition, competitive outgroup prosocial intentions did not significantly change over time (p T1 vs. T2 = .076; p T1 vs. T3 = .063). 9–10 year olds: Pairwise comparisons revealed that, among 8–9 year olds, competitive outgroup prosocial intentions increased following the intervention compared to the baseline, and stayed at the higher level 6 months later (p T1 vs. T2 = .002; p T1 vs. T3 = .038). In contrast, in the control condition, competitive outgroup prosocial intentions did not significantly change over time (p T1 vs. T2 = .499; p T1 vs. T3 = .833). Non-competitive outgroup prosocial intentions We conducted a mixed model ANOVA with Condition (experimental vs. control) and Year Group (5–6 vs. 7–8 vs. 9–10 years) as between participants factors and Time (pre-intervention vs. one month post-intervention vs. six month post-intervention) as a within participants factor. Results revealed no significant main effects of Time, F (2, 336) = 1.35, p = .260, η2 = .01, Condition, F (1, 168) = 3.15, p = .078, η2 = .02, or Year Group, F (2, 168) = 1.48, p = .231, η2 = .02. The Time x Condition interaction was significant, F (2, 336) = 6.53, p = .002, η2 = .04. All other two-way interactions were non-significant (all p's > .103). The three-way interaction was also non-significant, F (4, 336) = 1.13, p = .340, η2 = .01. Baseline comparison: Pairwise comparisons showed that there were no baseline differences in non-competitive outgroup prosociality between the experimental and control conditions among 5–6 year olds (p = .337), 7–8 year olds (p = .537), or 9–10 year olds (p = .809). Time x condition interaction on non-competitive outgroup prosociality: In order to probe the significant two-way interaction of Time x Condition we conducted pairwise comparisons (see Fig. 2 for means and standard errors). These comparisons revealed that, in the experimental condition, non-competitive outgroup prosocial intentions remained stable over time (p T1 vs. T2 = .090; p T1 vs. T3 = .277). In contrast, in the control condition, non- competitive outgroup prosocial intentions significantly reduced at Time 2 and Time 3 compared to the baseline (p T1 vs. T2 = .011; p T1 vs. T3 = .004). Pairwise comparisons also revealed that non-competitive outgroup prosocial intentions did not differ between the experimental condition and control condition at Time 1 (p = .297). However, at Time 2 and Time 3, prosocial intentions were significantly higher in the experimental than in the control condition (p T2 = .017; p T3 = .015).","The present field studies drew on the Arts and Kindness model (Broadwood et al., 2012) to test whether engagement in participatory arts promotes prosocial intentions among children. Study 1 showed that children who engaged in a participatory arts program showed increases in interpersonal prosocial intentions, while children in the control condition did not. There were no effects on children's empathy. Moreover, in Study 2, children who engaged in a participatory arts program showed increases in competitive outgroup-targeted prosocial intentions, while children in the control condition did not. These positive effects of condition on competitive outgroup prosocial intentions were long-lasting. Specifically, within the experimental condition, competitive outgroup prosocial intentions remained high six months post-intervention. For non-competitive prosocial intentions, Study 2 showed that while non-competitive prosociality reduced over time in the control condition, it remained stable in the experimental condition. It is possible that there may have been a time of school year effect in the control condition, where children showed greater prosociality at the start of the study (early in the school year) than the middle or end, especially in the non-competitive context. This might reflect that children become more stressed and less contented as the pressure of work and tiredness cumulates during the school year (Connor, 2003; Hall, Collins, Benjamin, Nind, & Sheehy, 2004). In both contexts, the trend in the intervention conditions was the opposite, suggesting consolidation or even strengthening of prosociality over time. One reason that this might happen is that after children have created and produced artistic outputs, these remain as significant and salient reminders over time, and as children (and perhaps teachers, carers, and peers) reflect on these they gain increased purchase on prosocial motivation. This would be consistent both with the Arts and Kindness model (Broadwood et al., 2012) and with the evidence from the longitudinal analyses conducted on population level data from adults by Van de Vyver and Abrams et al. (2017). Researchers have highlighted the lack of research exploring the social psychological and social-developmental outcomes of arts engagement (Goldstein et al., 2017; Sigel & Gitomer, 1992; Van de Vyver & Abrams, 2017). A small body of recent social developmental research shows that engagement in specific art forms such as music (e.g., Schellenberg et al., 2015) or more specific physical movements such as synchrony (Tunçgenç & Cohen, 2016) can promote prosocial responses during middle childhood. The current research provides a novel contribution to this area of research by testing impacts of a participatory arts intervention, testing impacts on outgroup-targeted prosocial intentions, and testing whether effects are long- lasting. Overall, our results are in line with the Arts and Kindness model (Broadwood et al., 2012) and the arts and human flourishing model (Tay et al., 2018), and suggest that arts engagement has the potential to act as a social psychological catalyst for promoting prosocial intentions during middle childhood. Intergroup bias in prosociality appears early in childhood, peaks in middle childhood, and is evident across the lifespan (Abrams et al., 2015; Over, 2018). Rigorous and applied social-developmental research is essential in order to test effective strategies for promoting positive and inclusive intergroup attitudes, intentions, and behaviors in childhood (cf. Rutland, Cameron, Bennett, & Ferrell, 2005). We hypothesized that arts engagement would be effective for promoting inclusive prosocial intentions across competitive and non-competitive contexts. Interestingly, and in line with previous research (cf. Abrams et al., 2015), children's outgroup prosocial intentions varied by the competitiveness of the intergroup context. Specifically, age interacted with condition and time to affect prosocial intentions in the competitive context but not in the non-competitive context. Among 7–8 year olds and among 9–10 year olds, competitive outgroup prosocial intentions increased significantly following the intervention, and effects persisted over the 6 month period. However, among 5–6 year olds there were no changes in competitive outgroup prosocial intentions. In contrast, non-competitive prosocial intentions remained stable in the experimental condition and across age. It seems likely that, to some extent, prosociality is a normative (socially desirable) response when there are no competing motivations, but that this ‘default’ response must become more deliberative under situations such as competition (cf. Rutland et al., 2010). The arts intervention seems likely to have provided a basis for a more prosocial orientation under such circumstances. Study 2 revealed interesting longitudinal effects of condition by age. Specifically, 5–6 year olds' competitive outgroup prosocial intentions remained stable across time in both the experimental and control condition. In contrast, 7–10 year olds' competitive outgroup prosocial intentions significantly increased following the arts interventions and these increases were sustained over time. However, in the control condition 7–8 year olds' competitive outgroup prosocial intentions marginally reduced over time and 9–10 year olds' competitive outgroup prosocial intentions remained stable over time. These findings suggest that 5–6 year olds may be less responsive, at least in the relatively short term, to holistic arts interventions. In contrast, among 7–10 year olds, there were significant effects of condition on children's competitive outgroup prosocial intentions. These age-based variations may be due to children's increasing sensitivity to contextual factors (e.g., competition) that impact intergroup dynamics (Abrams et al., 2017; Abrams, Rutland, Palmer, & Purewal, 2014). In summary, the present evidence demonstrated that engagement in participatory arts: (1) could promote intentions to act prosocially toward outgroup members even in the context of intergroup competition, and (2) that it consolidated prosocial intentions toward outgroup members rather than decaying over the duration of the school year. These findings suggest that facilitating engagement with the arts across childhood can be an effective way to maintain and promote outgroup prosocial intentions. Limitations and future directions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In this paper we tested the impact of participatory arts engagement on both empathy and prosociality but only found evidence for the latter. However, the Arts and Kindness model (Broadwood et al., 2012) proposes three additional routes through which arts engagement can affect prosociality, namely by promoting different values, establishing connections, and learning about others. Further work needs to be done to examine how these different routes may be involved in the connection between arts engagement and prosociality, and whether each has greater or lesser importance during different periods of child development and socialization. Indeed, further social-developmental research on the arts will enable us to understand how relationships between arts and socio-emotional outcomes may vary by periods of development. This proposal is in line with Eisenberg et al.'s (1999) work which shows that prosocial thoughts, emotions, and behaviors become more consolidated, and individual differences become more evident as children grow into adolescence and then adulthood. We hope that this paper will inspire researchers to become more interested in the potentially wide and pervasive developmental impacts of the arts on socio-emotional outcomes. Relatedly, it is possible that the absence of effects on empathy may have been due to a measurement issue and that we were simply unable to capture the effect of participatory arts on empathy. Indeed, across the developmental literature there is uncertainty about how to best to measure empathy in this age range. We used an adapted version of Bryant's empathy index which is widely used across this age range. However, reliability coefficients for empathy scales, including the Bryant empathy index, are often lower than desirable (De Wied et al., 2007). Moreover, it is now well-established that empathy entails at least two separate components: sympathy (concern for another based on the apprehension or comprehension of the other's emotional state) and personal distress (an aversive, self-focused emotional reaction to the apprehension or comprehension of another's emotional state) (Eisenberg et al., 1999). However, these two components typically cross-load when measured in middle childhood (De Wied et al., 2007; Garton & Gringart, 2005), perhaps because “children do not differentiate the two or perhaps […] did not capture any subtle distinction that Davis was trying to make” (Garton & Gringart, 2005). In the present research we measured prosocial intentions rather than prosocial behavior. Although we cannot assert that children in the experimental conditions would also behave more prosocially, there are good reasons for believing that this might be the case. Many intention-behavior relationships are large and stable because they become habitual (e.g., voting or health behaviors). The challenge for interventions is to change intentions in ways that can change behavior. A meta-analysis of the intention-behavior relationship in adults showed that a medium-to-large change in intention is required (d = .66) in order to induce a small-to-medium change in behavior (d = .36) (Webb & Sheeran, 2006). Across the present two studies, children's engagement in participatory arts led to small to medium sized changes in their prosocial intentions (effect sizes varied between d = .28 to d = .50 for the significant effects of the intervention). Consistent with this, another study from our lab (Ali, Abrams, & Van de Vyver, 2019) does indicate that, in middle childhood, outgroup prosocial intentions are significantly positively correlated with outgroup prosocial behavior (r = .23). Even if the effects of the intervention on intentions have small or subtle implications for immediate behavior, the gain in prosociality could be expected to yield positive cumulative effects, especially if it is reciprocated by others. It would be desirable for future research to test the impact of arts engagement on behavioral responses directly. A further research question is whether there are important moderators of the relationship between arts engagement and prosociality. For example, research shows that individual differences in personality (e.g., agreeableness and openness) are associated, albeit to different extents, with arts engagement or with prosociality (Van de Vyver & Abrams, 2017). It is possible that the impact of arts engagement on prosociality may vary depending on these individual differences, and relatedly how aspects of children's temperament may augment or militate against the impact of such interventions. Applied implications ~~~~~~~~~~~~~~~~~~~~ The Arts are universal and are woven into culture and history. Young children readily engage in the arts. Although the skills, performance, and consumption of the arts are the subject of substantial research, their role as a societal glue, and in the social lives of children, has been relatively neglected (Sigel & Gitomer, 1992). Publicly financed arts provision is often at the front line of cuts during periods of austerity, often depriving the poorest in society of access (e.g., free entry to museums, etc). In the UK, for example, the arts sector has faced substantial reductions in public and private spending and arts and creative subjects have faced heavy cuts from the school curriculum (NCA, 2017). However, the costs of neglecting the arts may be greater than expected. For example, developmental research shows that engagement in the arts can actually promote cognitive ability in childhood (e.g., Bilhartz, Bruhn, & Olson, 1999; Tõugu, Marcus, Haden, & Uttal, 2017). Furthermore, the present research demonstrates that engagement in the arts promotes socio-emotional development in childhood, perhaps improving the lives of children, their peers, and ultimately their wider communities. In order to maximise the positive impacts of arts engagement, and in order to ensure that arts are accessible to all, this evidence contributes to the case for embedding arts engagement across the curriculum.","Overall, new and longitudinal evidence across two studies provides support for the hypothesis that middle childhood is a period in which the arts can act as a social psychological catalyst that promotes interpersonal and intergroup prosocial intentions. The results suggest that if we wish to sustain cohesive and inclusive societies, increasing access to, and engagement in, the arts from an early age may be a valuable means of doing so.","Neither of the experiments reported in this article was formally preregistered. The measures are available in the supplementary materials of this paper. The data are available on OSF (https://osf.io/wjby5/?view_only=ee6074f6c3964947a83bb848db57df77)."],["Four-month-old infants perceive continuity of an object's trajectory through occlusion, even when the occluder is illusory, and several cues are apparently needed for young infants to perceive a veridical occlusion event. In this paper we investigated the effects of dislocating the spatial relation between the occlusion events and the visible edges of the occluder. In two experiments testing 60 participants, we demonstrated that 4-month-olds do not perceive continuity of an object's trajectory across an occlusion if the deletion and accretion events are spatially displaced relative to the occluder edges (Experiment 1) or if deletion and accretion occur along a linear boundary that is incorrectly oriented relative to the occluder's edges (Experiment 2). Thus congruence of these cues is apparently important for perception of veridical occlusion. These results are discussed in relation to an account of the development of perception of occlusion and object persistence. --------------------------------------------------------------------------------","In Experiment 1 we adopted the method used previously to investigate young infants’ perception of trajectory continuity (Bremner et al., 2005, 2007; Johnson, Bremner et al., 2003) to manipulate the position of the deletion and accretion event in the object's path of movement so that these did not coincide with the position of the edges of the occluder. Thus in one condition the deletion event occurred earlier and the accretion event occurred later than they should given the width of the occluder (Fig. 3, upper images), and in the other condition the deletion event occurred later and the accretion event occurred earlier than they should given the width of the occluder (Fig. 3, lower images). These displays were designed so that the maximum occluder width and maximum separation between deletion and accretion boundaries maintained time and distance out of sight well within the range in which 4-month-olds perceive trajectory continuity when the occluder edges and occlusion events are co-located (Johnson, Bremner et al., 2003).","Twenty-four 4-month-old infants (M = 123.2 days; range 112–141 days; 14 girls and 10 boys) took part in the experiment. Two other infants did not complete testing, one due to fussiness and the other because of equipment failure. Twelve infants were assigned to each of the two conditions in such a way as to ensure that the mean age and the gender balance were comparable across conditions. Throughout the series, infants took part in only one experiment. In all experiments, participants were recruited by personal contact with parents in the maternity unit when the baby was born, followed up by telephone contact near test age to those parents who volunteered to take part. Infants with reported health problems including visual and hearing deficits and those born two weeks or more before due date were omitted from the sample. The majority of participants (over 95%) across both experiments were from Caucasian, middle class families.","A Macintosh computer and a Samsung 100 cm color monitor were used to present stimuli and collect looking time data. An observer viewed the infant on a second monitor, and infants were video recorded for later independent coding of looking times by a second observer. Both observers were unaware of the hypothesis under investigation. Using HABIT software (Cohen, Atkinson, & Chaput, 2000) the computer presented displays, recorded looking time judgments, calculated the habituation criterion for each infant, and changed displays after the criterion was met. The first observer’s looking time judgments were input with a keypress on the computer keyboard. Fig. 3 indicates the displays used in this experiment. The habituation display was presented against a black background with a 20 × 20 grid of white dots measuring 48 × 48 cm (27° x 27°) serving as texture elements. A blue occluder with vertical extent 21.5 cm (12.3°) was placed centrally. A 6.7 cm (3.8°) green ball moved back and forth from one side of the display to the other, moving at 16.5 cm/s (9.4°/s). In the early deletion late accretion display the visible occluder was 7 cm wide, but the ball disappeared and reappeared at invisible contours separated by 12 cm. In the late deletion early accretion display the occluder was 12 cm wide, but the ball disappeared and reappeared at invisible contours separated by 7 cm. It took 2500 ms for the ball to traverse the width of the display. Time from complete visibility to invisibility or the reverse was 400 ms. Time totally out of sight was 366 ms. (early deletion late accretion display) or 67 ms (late deletion early accretion display). Time completely in sight to the left and right of the occluder was 1332 ms (early deletion late accretion display) or 1634 ms (late deletion early accretion display). The animation was run as a continuous loop for the duration of the trial. In test displays, the inducing elements were removed and the ball moved back and forth at the same speed as in the habituation display. In the continuous trajectory test display, the ball was always visible. In the discontinuous trajectory display, the ball went out of and back into view just as in the habituation event. Procedure Each infant was seated 100 cm from the display and tested individually in a darkened room. The habituation display was presented until looking time declined across four consecutive trials, from the second trial on, adding up to less than half the total looking time during the first four trials. Timing of each trial began when the infant fixated the screen after display onset. The observer pressed a key as long as the infant fixated the screen, and released when the infant looked away. A trial was terminated when the observer released the key for two seconds or 60 s had elapsed. Between trials, a beeping target was shown to attract attention back to the screen. Following habituation trials, infants were presented with the two test trials in alternation, three times each, for a total of six trials. On test trials, half the infants in each condition were presented with the continuous trajectory first, and the rest viewed the discontinuous trajectory first. The second observer coded looking times from videotape for purposes of assessing reliability of looking time judgments. Interobserver correlations were high across the two experiments in this report (M Pearson r = 0.99). In this experiment we did not run control conditions consisting of test trials alone to assess intrinsic preferences because we have obtained null preferences in control conditions with identical test trials in previous work (Bremner et al., 2005; Johnson, Bremner et al., 2003).","Analysis of data from habituation trials indicated no difference between conditions in trials to habituation (early deletion late accretion condition, M = 7.83; SD = 1.27: late deletion early accretion condition, M = 7.75; SD = 1.66; t (22) = 0.138, p = 0.89) or in total habituation time (early deletion late accretion condition, M = 201.2 s; SD = 90.86; late deletion early accretion condition, M = 217.1 s; SD = 77.52; t (22) = −0.46, p = 0.65). Fig. 4 shows looking times at the two test displays for each condition. Infants in both the early deletion late accretion and late deletion early accretion conditions looked longer at the continuous test trial, although this was more marked in the early deletion late accretion condition. We can have confidence in assuming that longer looking at one test display is indicative of a novelty preference rather than a familiarity preference for two reasons. Firstly, infants were habituated to a standard criterion, circumstances under which familiarity preferences rarely occur (e.g., Fiser & Aslin, 2002; Johnson et al., 2009; Moore & Johnson, 2011). Secondly, several papers (Bremner et al., 2005, 2007; Johnson, Bremner et al., 2003) have reported systematic age related data that would be very hard to interpret on the basis of familiarity preference. Because looking time data tend to be positively skewed, violating an assumption of ANOVA, data in this and the subsequent experiment were log transformed prior to analysis. A 2 (display: early vs. late occlusion) x 2 (test trial order) x 2 (test trial type: continuous vs. discontinuous) x 3 (test trial block) mixed ANOVA yielded a significant effect of test trial type, F (1,20) = 8.05, p = 0.01, ηp2 = 0.29. The interaction between test trial type and display condition was not significant, F (1,20) = 0.82, p = 0.37, ηp2 = 0.04, nor were any other main effects and interactions.","The rationale for this habituation-test approach is that the direction of the novelty preference indicates how infants have processed the habituation stimulus. Thus infants' longer looking at the continuous test trial suggests that they processed the object’s trajectory in the habituation display as discontinuous. This is the opposite result as is obtained for this age group with comparable times and distances out of sight when the occlusion event coincided with the visible occluding edges, and it certainly provides no evidence that infants perceived continuity in these habituation events. Thus we can conclude that the occlusion event must coincide spatially with the occluding edges for infants to perceive an occlusion event indicating continuity of the object.","Experiment 1 makes it clear that when the occlusion events are separated by 2.5 cm from the edges of the occluder, infants do not perceive the event as a normal occlusion in which the object persists. Our second question was whether the same result would be obtained if the edges of the moving ball at the boundaries where deletion and accretion took place were oblique, but had contact with the visible occluder edges. In this case, the separation between occlusion event and occluder edges is reduced, but the occlusion event takes a different form, occurring as if at oblique edges. Intuitively, this disjunction in the orientation of occluding contours and occluder edges seems subtler than the complete separation of event and edge in Experiment 1. However, it is possible that misorientation of occluding contour and occluder edge is as important as a simple displacement in the horizontal dimension. Thus in Experiment 2 we exposed infants to a display in which the occluder had vertical edges but the occlusion event occurred at oblique contours. We know that 4-month-olds perceive trajectory continuity when a horizontally moving object passes behind an occluder with oblique edges (Bremner, Slater, Mason, Spring, & Johnson, 2016), so a negative result should not be due to occlusion at an oblique contour alone. However, because we had previously only tested infants' perception of occlusion across a very narrow oblique occluder, in this experiment we included a ‘standard' baseline condition in which the occluder edges were oblique and the occlusion event was congruent with these edges, and also a control condition in which infants were only exposed to the test trials (to assess any intrinsic preference for either trial).","Thirty-six 4-month-old infants (M = 122.9 days; range 109-138 days; 17 girls and 19 boys) took part in the experiment. Eight other infants did not complete testing due to fussiness. Twelve infants were assigned to each of the two conditions in such a way as to ensure that the mean age and gender balance were comparable across conditions.","The same apparatus as in Experiment 1 was used for stimulus presentation, video recording of the infant, and recording looking times. Fig. 5 indicates the displays used in this experiment. In the experimental condition, the habituation display was presented against a black background with a 20 × 20 grid of white dots measuring 48 × 48 cm (27° x 27°) serving as texture elements. An occluder with vertical and horizontal dimensions 21.5 cm (12.3°) and 7 cm (4°) was placed centrally. A 6.7 cm (3.8°) green ball moved back and forth from one side of the display to the other, undergoing progressive deletion and accretion at oblique linear contours at 55 ° to the vertical, intersecting with the left and right edges of the occluder at the points where the bottom and top of the ball respectively contacted the occluder. In the baseline condition, the vertical occluder was replaced by an oblique occluder at 55 ° to the vertical with length 21.5 cm and width 9.5 cm such that its occluding edges were congruent to the apparent edges at with deletion and accretion took place in the experimental condition. In both cases, time from complete visibility to invisibility or the reverse was 500 ms. Time totally out of sight was 233 ms and time completely in sight to left and right of the occluder was 1265 ms. Test displays contained no visible edges or background occlusion and the ball moved back and forth at the same speed as in the habituation display. In the continuous trajectory test display, the ball was always visible. In the discontinuous trajectory display, the ball went out of and back into view just as in the habituation event. All other aspects of the form of the habituation and test trials were the same as in Experiment 1. We also included a control group who saw only the test trials. Procedure Infants were first habituated to the habituation display, and were then presented with the two test displays in alternation, three times each, for a total of six test trials. On test trials, half the infants in each condition were presented with the continuous trajectory first, and the rest viewed the discontinuous trajectory first. Habituation and test trials were carried out according to the same criteria and procedures as in Experiment 1.","Analysis of data from habituation trials indicated no difference between conditions in trials to habituation (experimental condition, M = 7.5, SD = 2.28; baseline condition, M = 7.75, SD = 2.22, t (22) = −0.27, p = 0.79) or in total time to habituation (experimental condition, M = 151.2, SD = 45.81; baseline condition, M = 238.8, SD = 150.4; t (22) = −0.193, p = 0.076). As Fig. 6 indicates, infants in the experimental condition showed a consistent preference the continuous test display, whereas infants in the baseline condition showed a consistent preference for the discontinuous display, and infants in the control condition did not show a consistent preference for either test display. A 3 (condition: experimental vs. baseline vs. control) x 2 (test trial order) x 2 (test trial type: continuous vs. discontinuous) x 3 (test trial block) mixed ANOVA yielded a marginally significant effect of test trial type, F (1,30) = 4.1, p = 0.051, ηp2 = 0.12, and a significant interaction between test trial type and condition, F (2,30) = 14.74, p < 0.001, ηp2 = 0.5. To interpret this interaction, separate analyses were carried out for each condition. In the case of the experimental condition, there was a significant effect of test trial type, F (1,10) = 5.49, p = 0.041, ηp2 = 0.35, with infants looking longer at the continuous test trial. There were no other significant main effects or interactions. In the case of the baseline condition, there was a significant effect of test trial type, F (1,10) = 20.51, p = 0.001, ηp2 = 0.67, with infants looking longer at the discontinuous test trial, and a significant interaction between test trial type and test trial block, F (2,9) = 4.42, p = 0.046, ηp2 = 0.5, due to an increase in the test trial type effect across trial blocks. In the case of the control condition, a 2 (test trial order) x 2 (test trial type: continuous vs. discontinuous) x 3 (test trial block) mixed ANOVA yielded only a significant effect of test trial block, F (2,9) = 9.29, p = 0.006, ηp2 = 0.67. This is explained by longer looking on the first test trial block, and is a common effect in control conditions in which test trials are the only displays presented. The important finding is that there is no underlying preference for one test trial over the other. In summary, infants in the experimental condition looked longer at the continuous test display, the opposite of the result shown by infants in the baseline condition. Apparently, unlike infants exposed to the baseline display, they perceived the habituation event as an object moving on a discontinuous trajectory. Thus it appears that an occlusion event that is wrongly oriented relative to the visible occluding edges does not provide information for object continuity.","In both experiments, infants appeared to perceive the object’s trajectory in the habituation display as discontinuous. This is despite the fact that in both experiments the display contained both visible occluding contours, and deletion and accretion events. These are circumstances under which infants would normally perceive trajectory continuity when occluding contours and deletion/accretion are congruent, providing the time and distance out of sight is sufficiently short, which they were in these experiments. Thus these studies contribute important additional information regarding young infants’ perception of occlusion events, namely that the occlusion event must be spatiotemporally congruent relative to a visible occluder if it is to specify normal occlusion in which the occluded object persists. Although it is clear that multiple cues are needed to specify a normal occlusion event (Bremner et al., 2012), the effect of these cues is not simply additive; they must be spatially congruent to be effective. We must acknowledge an alternative interpretation of the present results. Maybe the isolated deletion and accretion events are perceived as similar to the discontinuous test display, and so preference for the continuous display reflects a simple perceptual novelty preference. We believe that such an interpretation is unlikely, because infants showed a preference for the discontinuous test display in the case when the habituation display had a Kanizsa figure as a virtual occluder and thus contained particularly isolated deletion and accretion events. However, even if this is the appropriate interpretation of performance in the present experiments, our main conclusion remains unchanged, namely that deletion- accretion and occluder edges must be congruent if infants are to perceive an occlusion event in which the object persists while invisible. In one sense, this conclusion should come as no surprise because to the adult eye these habituation displays do not look like cases of veridical occlusion. However, the results of this study attest to another respect in which young infants appear to perceive the world in adult-like ways. Specifically, there has to be precise spatial coordination between information for occlusion and information about the occluder. This is important in relation to the finding that young infants perceive the Kanizsa figure as an occluding surface. In this case, the explicit information specifying the occluding surface is distant from the occlusion event. Presumably the main point, however, is that although this information is distant it specifies a complete illusory surface that extends to and is congruent with the deletion and accretion events. What we cannot say is whether the degree of congruence required in the case of the Kanizsa figure would be the same as is required in the case of a visible occluder. Our suspicion is that a high degree of spatial congruence would be required, whether or not the occluder is real or illusory. It seems likely that, even early in development, the visual system is tuned to detect veridical events, and spatial congruence of cues may be particularly important in this process. The fact that spatial alignment of inducing elements creates the illusion of a surface in the case of the Kanizsa figure probably points to the importance of alignment as a principle of perceptual organisation that exists early in development. If that conclusion is correct, then we can expect that spatial congruence of events is important for perception of veridical occlusion. By misaligning or misorienting the occlusion event relative to the occluder, our manipulations may thus have violated a key principle of perception of events in which one object passes behind another. The aim of this work was to investigate the importance of cue congruence in 4-month-old infant's perception of occlusion events, with the focus on 4 months chosen because this appears to be a pivotal age in development of perception of object continuity across occlusion. However, it would be of interest to investigate older infants' perception of these displays. On the one hand, we might expect that older infant would not perceive trajectory continuity either, because these displays present non- veridical occlusion events and that we would expect a developmental progression towards detection of veridical events. On the other hand, we know that adults perceive an occlusion event on the basis of deletion and accretion alone (the tunnel effect), and at so some point in development in infancy or later we might expect participants to rely on this single cue alone. Thus far, it appears we have evidence that 6-month-olds do not use deletion-accretion alone to specify continuity (Johnson, Bremner et al., 2003), and future work will establish the age at which this emerges. It remains possible, however, that even in adults the presence of incongruent luminance defined occluding edges would disrupt perception of the tunnel effect."],["Sensorimotor contingency is one of the main factors to warp time perception. Voluntary actions such as saccades and hand movements affect the subjective perception of temporal duration. Although the perceived timings of action and stimulus are affected by whether an action was automatic or controlled, its effect on the subjective perception of duration has not been studied except in the case of saccade (chronostasis), which has been shown to be unaffected by the context of action initiation. Here we investigate the effect of the context of action initiation on duration estimation in the case of finger movement. The reproduced intervals were shorter when actions were initiated by automatic manner, compared to self-timed or cognitively controlled actions. The results are compatible with an internal clock model employing variable latencies for switch closure after action. --------------------------------------------------------------------------------","The processing of temporal information is ubiquitous in cortical computation. It is important for both perception and action, serving as an essential element of how the brain constructs models of the environment. Recently, time perception ranging from sub-second to several seconds has been extensively studied. Studies have shown that the perceived duration is influenced by the properties of stimuli (Xuan, Zhang, He, & Chen, 2007). In a successive presentation of identical stimuli, the perceived duration of an oddball stimulus is longer (oddball effect, Tse, Intriligator, Rivest, & Cavanagh, 2004). A visual onset expands the subjective time (Kanai & Watanabe, 2006). When the same stimulus is presented successively for several times, the first one is perceived as longer than the others (debut effect, Pariyadath & Eagleman, 2007). The debut effect disappeared when the stimuli were random images. These results indicate that the predictability of the stimulus affects the oddball and debut effects. Multisensory interaction (van Wassenhove, Buonomano, Shimojo, & Shams, 2008) and emotion (Droit-Volet & Meck, 2007; Doi & Shinohara, 2009) also affect the subjective duration. Thus, time perception is a highly complex cognitive process affected by various elements related to the stimuli. One of the factors that potentially affect time perception is sensorimotor contingency. It has been suggested that the intentional state is one of the important factors affecting the sense of time (Haggard, Clark, & Kalogeras, 2002). The perceived timing of key pressing and the subsequent tone were shifted so that they were closer to each other, when the subject voluntarily pressed the key followed by a tone with some delay. This “temporal attraction” has been termed “intentional binding effect” (Ebert & Wegner, 2010; Engbert & Wohlschläger, 2007; Humphreys & Buehner, 2010; Moore & Haggard, 2008; Stetson, Cui, Montague, & Eagleman, 2006; Tsakiris & Haggard, 2003). In contrast, when the movement was induced by transcranial magnetic stimulation (TMS), the temporal shift occurred in the opposite direction. The mode of initiation of an action (intrinsic or extrinsic, i.e., self-initiated (intention-based) or externally-triggered (stimulus-based), Jahanshahi et al., 1995; Waszak et al., 2005)) is one of the important contextual factors in sensorimotor contingency. In an externally-triggered condition, a subject generates a movement as a response to a sensory stimulus. A self-initiated action, in contrast, is driven internally. It has been suggested that different mechanisms are engaged in the intrinsic and extrinsic movements (Obhi & Haggard, 2004). Welchman, Stanley, Schomers, Miall, and Bulthoff (2010) showed that the speed of action in a reactive movement was faster than in a self-timed action. Imaging studies have shown that different neural mechanisms are engaged in these movements (Herwig, Prinz, & Waszak, 2007; Jenkins, Jahanshahi, Jueptner, Passingham, & Brooks, 2000; Keller et al., 2006; Taniwaki et al., 2006). Activities in the basal ganglia (Cunnington, Windischberger, Deecke, & Moser, 2002) and dorsolateral prefrontal cortex (François-Brosseau et al., 2009; Jahanshahi et al., 1995; Wiese et al., 2004) increase in self-initiated movements compared to the externally- triggered movements. In a monkey study, the firing rate of neurons in the putamen increased faster in a self-initiated than in an externally-timed movement (Lee & Assad, 2003). A study on human subjects also showed that movement-related cortical potentials in a self-initiated movement were larger than in an externally-triggered movement (Jahanshahi et al., 1995). The onset of hemodynamic response of pre-SMA in a self-initiated movement was earlier than in an externally-triggered movement (Cunnington et al., 2002). Some studies have suggested that externally and internally initiated movements have different effects on time perception. Haggard, Aschersleben, Gehrke, and Prinz (2002), for example, reported temporal attraction effects between action and auditory tone both in self- initiated and externally-triggered conditions. In a self-initiated condition, a subject intentionally pressed a key at his or her own timing, causing a tone after 200 ms. In an externally-triggered condition, on the other hand, the subject pressed the key as quickly as possible upon hearing tone. The temporal orders of action and tone were reversed between the two conditions. The subject reported the perceived timings of key pressing and tone. The perceived timings of action and tone became closer to each other in both conditions, indicating that the directions of the shift of the perceived timing of tone and action were reversed between the two conditions. Studies on chronostasis have shown that actions and subjective durations are closely linked (Yarrow, Haggard, Heal, Brown, & Rothwell, 2001; Yarrow, Haggard, & Rothwell, 2004; Yarrow, Johnson, Haggard, & Rothwell, 2004). In this illusion, the visual stimulus presented immediately after the saccade was perceptually dilated (Yarrow et al., 2001). The magnitudes of the lengthening effect were similar between the controlled and automatic eye movements (Yarrow, Johnson et al., 2004), with a constant effect across various intervals, indicating that chronostasis was not affected by the magnitude of volition. Park, Schlag-Rey, and Schlag (2003) found that not only saccades but also voluntary movements such as key press and utterance caused an overestimation of the perceived duration of its sensory feedback, extending the chronostasis studies in context. In general, there appear to be interconnections between voluntary actions and subjective durations, a point that needs to be investigated further. Different neural mechanisms and range of intervals are involved for saccades, finger movements, and other kinds of voluntary actions (Kandel, Schwartz, & Jessell, 2000). Thus, the results obtained in the chronostasis studies do not necessarily apply to voluntary movements in general. Identifying a common mechanism for saccades and other voluntary movements will contribute to the understanding of the relation between subjective duration and voluntary movements in the general context. The effect of chronostasis was not affected by the context of saccade initiation (Yarrow, Johnson et al., 2004), and was constant across intervals (Yarrow, Haggard et al., 2004). It is still unknown whether such is the case for other kinds of voluntary movements (e.g. key pressing). It has been shown that the perceived timings of action and effect were affected by the sensorimotor context (Haggard, Aschersleben et al., 2002). It is possible that the subjective duration in voluntary movements such as key pressing is affected by the sensorimotor context, in contrast to chronostasis. It is interesting to investigate whether the context of action initiation would affect the estimated interval. Here we use a temporal reproduction paradigm to investigate the effect of sensorimotor context on the perception of duration. The subjects were presented with sensory stimuli of various intervals. They were instructed to estimate the intervals and reproduce them through key pressing. In order to investigate the effect of sensorimotor contexts (e.g. self-timed versus externally-timed), the timings of the key pressing were constrained in various ways. A voluntary movement can be classified either as an automatic or a controlled process (Wegner, 2002). An automatic movement is reflex-like and fast (∼500 ms), with its conscious perception occurring after the motor execution (Welchman et al., 2010). On the contrary, a controlled movement requires cognitive processes depending on attention and working memory, takes longer time (>500 ms), and is more flexible. An action initiation within 500 ms after the Go signal can be regarded as an automatically controlled movement, while an action initiated between 1 and 2 s after Go signal can be regarded as cognitively controlled movement, requiring the subject to inhibit motor output for some interval while carrying it out before the deadline. Such a movement could not be achieved without a top down control. Our experimental setup incorporated these temporal constraints. It has been suggested that different neural mechanisms are involved in the estimation of temporal duration below and above 2–3 s (Poppel, 1997; Ulbrich, Churan, Fink, & Wittmann, 2007). Temporal processing of up to 2–3 s can be regarded as “time perception”, while those lasting more than 3 s is thought to be “time estimation” (Fraisse, 1984). In order to investigate the relation between the estimated intervals and sensorimotor contingencies, intervals in the range of 1–5 s were presented in the experiments, covering the “time perception” as well as the “time estimation” domains. General descriptions ~~~~~~~~~~~~~~~~~~~~ We conducted three experiments. All subjects were naïve about the purpose of the present study, except that in experiments 1 and 2 subject TH was one of the authors of this paper. All subjects had normal or corrected-to-normal vision, and were right-handed by self- report, except for one subject in experiment 3. The experiments were conducted in accordance with the Declaration of Helsinki. The experimental procedures were submitted to and approved by the brain and cognitive sciences ethics committee of Sony Computer Science Laboratories. Informed consent was obtained from all participants. The experiments were conducted using the self-timed and externally-timed conditions. In the self-timed condition, the subject was instructed to start the reproduction at his or her own timing. In the externally-timed condition, the subject was required to start the reproduction within a predetermined interval. There were two subconditions for the externally-timed condition. In the automatic condition, the subject was instructed to press the key within 500 ms after the trigger signal. In the controlled condition, the subject was instructed to press the key within 1–2 s after the trigger signal. It was assumed that the automatic and controlled conditions would induce reflex-like (Welchman et al., 2010) and cognitively controlled movements, respectively, the latter involving an active suppression and higher volitional states. In Experiment 1 and experiment 2 (visual condition), the sessions were controlled by a desktop PC (EPSON Endeavor MT8800) and a 21-inch CRT monitor (Sony Trinitron CPD-G520). The subjects responded by pressing a key (SANWA SUPPLY NT-11UBK, with a keystroke of 2.2 ± 0.1 mm). In experiment 2 (auditory condition) and experiment 3, the sessions were conducted on Mac Book Pro 15 inch model. The tones were presented through a headphone (Sony MDR-XD100). Experiment 1 ~~~~~~~~~~~~ In experiment 1, we examined the effect of movement under self-timed and externally-timed (automatic or controlled) conditions on the perception of temporal duration. In experiment 1(a), ten subjects (seven males and three females, mean age = 29.4, sd = 2.5) participated in a combination of self-timed and automatic conditions. In experiment 1(b), eight subjects (five males and three females, mean age = 29.3, sd = 2.4) participated in a combination of self-timed and controlled conditions. The subjects of experiment 1(b) were a subset of the subjects of experiment 1(a). Before the experiment, the subjects practiced pressing the key according to the conditions. In experiment 1(a), the subjects practiced to press the key within 500 ms after the Go signal (automatic condition). The words “too late” were displayed on the screen at the passage of 500 ms after the Go signal. Typically, a practice of several times was sufficient. In experiment 1(b), the subject practiced to press the key between 1 and 2 s after the Go signal (controlled condition). After the key pressing, the subjects were given textual feedback; “too early” if the subject pressed the key before 1 s after the Go signal, “SUCCESS” if the subject pressed the key between 1 and 2 s after the Go signal, and “too late” if the subject pressed the key later than 2 s after the Go signal. When a subject successfully pressed the key for ten continuous trials, the practice session was over and the experiment started. All subjects could pass this criterion within trials of up to 30. At the beginning of each trial, a white circle (0.95°) appeared at the center of the screen, which remained until the end of each trial. The subject was instructed to keep a natural posture without crossing the legs and putting the elbow on the desk, and to fixate on the circle while keeping posture and attention. 1 s after the appearance of the circle, the reference auditory stimulus was presented to both ears through a headphone (Fig. 1). The subject was instructed to memorize its duration. The duration of the reference stimulus was either 1, 3 or 5 s, presented in random order. After the offset of the reference duration, the Go signal was presented. The intervals between the offset of the reference duration and the Go signal were randomly chosen from 1, 1.5 and 2 s, to prevent the subject from anticipating the timing of the Go signal. After the Go signal, the subject was instructed to press the designated key with their index finger of the right hand and keep pressing for a duration matching that of the reference tone. There was not an auditory feedback accompanying the key pressing. The subjects were instructed to refrain from silently counting or keeping a rhythm while they estimated the duration. The self-timed and externally-timed (automatic or controlled) conditions were assigned to respective blocks. Before the session started, the subjects were instructed which of the self-timed, or externally timed (automatic or controlled) conditions applied to each block. The subjects conducted two blocks for each condition. The order of the conditions was counterbalanced among the subjects. The subjects were allowed to take rest between the blocks. One block consisted of six presentations of the reference duration. Thus, the reproduction of duration was conducted twelve times for each interval in respective conditions, resulting in 72 trials overall. In the externally-timed conditions (automatic or controlled), when the subjects failed to press the key within the predetermined limits, an error feedback was presented, with the trial terminated to proceed to the next trial. The missed trials were stacked at the end of the block to be executed later. This manipulation was designed to prevent the subjects from adapting to the reference interval. Experiment 2 ~~~~~~~~~~~~ In experiment 2, we further investigated how the difference of action initiation would affect temporal processing. In exp. 1, the movement conditions were separated by blocks, where the subjects knew which action would be required at the moment the reference intervals were presented. If the subject’s internal states such as attention and concentration varied between the movement conditions, these changes in internal states throughout the block could affect the temporal processing at the encoding period. Therefore, there was the possibility that the differences in reproduced intervals were due to those in the encoding period, not in the reproduction period. In exp. 2, the subjects reproduced the reference intervals under externally-timed (automatic or controlled) conditions, without receiving a prior instruction of which action to take before the reference intervals were presented. Five subjects (three males and two females, with average age = 29.0, sd = 2.2) participated in the auditory condition. Five subjects (three males and two females, with average age = 29.6, sd = 2.1) participated in the visual condition. The subjects were a subset of the participants in experiment 1. In the auditory condition (Fig. 2a), a white square (1.72°) was presented for 1 s to indicate the beginning of the trial. 1 s after the disappearance of the square, a sound (1000 Hz, 60 dB) indicating the reference intervals was presented to the subject from the headphone. The subjects were instructed to memorize its interval. The duration of the reference intervals were 1, 3, and 5 s, presented in a random sequence. 1.5 s after the offset of the reference interval, the Go signal was presented. The Go signals were low (500 Hz) or high (2000 Hz) sound with a duration of 20 ms, one of which instructed the subject to reproduce within 500 ms after the Go signal (automatic condition), while the other instructed to start the reproduction between 1 s and 2 s after the Go signal (controlled condition). The assignments of the Go signal pitches for alternative actions were counterbalanced among the subjects. Before the experiment, the subjects learned the assignment of the Go signal pitches for alternative actions. This learning session consisted of thirty-four trials in a block. In the first four trials a red circle (1.53°) was presented at the center of the screen after the Go signal. It was presented for the predetermined interval within which the subjects had to press the key in the automatic or controlled condition (appearing twice each in an interleaved manner). The subjects attended to the stimulus and learned the timing of key pressing. In the remaining thirty trials, the subjects had to press the key as indicated by the red circle on the screen in the first four trials (the automatic and controlled conditions appearing fifteen times each in a random manner). The practice session was over when the subjects successfully pressed the key within the designated time frame for more than ten cumulative times for both conditions. In the case of failure, the subjects were required to conduct one more block. Text feedbacks were provided on the timing of subjects’ key pressing (“too fast”, “SUCCESS”, or “too late”). All the subjects cleared this criterion in less than two blocks. The experiment started after twelve warm up trials. One block of experiment consisted of eighteen trials, in which each reference interval for respective conditions was presented three times. The subjects conducted six blocks. In all, the subjects reproduced each interval 18 times in both conditions. After the Go signal, the subjects pressed the space key on the keyboard, pressing for the duration of the remembered interval and then releasing the key. No feedback was presented to the subjects during the key pressing. The next trial started 1.5 s after the subject keyed off. The subjects were instructed not to use the strategy of counting or keeping rhythm during the encoding and reproduction of the interval. In the visual condition, a white square (3.34°) was presented as the reference for 1, 3 or 5 s (Fig. 2b). After the reference stimulus, a color (red or green) circle (1.91°) was presented for 50 ms as the Go signal. In the automatic condition, the subjects were instructed to press the key within 500 ms after the Go signal. In the controlled condition, the timing of key pressing was between 1 and 2 s after the Go signal. The green or red circle instructed the subjects which action to conduct. The two colors were randomly presented. The associations between the color and action were counterbalanced among the subjects. Before the experiment, the subjects learned the association between colors and actions (automatic or controlled) in the practice session, where one of the alternative Go signals was presented randomly, and the subjects practiced to press the key within the predetermined limit after the Go signal. The subjects then practiced the experimental task, in which they learned to correctly estimate and reproduce the reference intervals. In one block, each reference interval for each condition was presented three times, resulting in a total of eighteen trials. The subjects conducted six blocks. They reproduced each interval 18 times in both conditions. As in previous experiments, the subjects reproduced the presented interval by continuously pressing the key. There was no feedback while they were pressing the key. Experiment 3 ~~~~~~~~~~~~ In experiment 3 (control), we investigated whether the delay between the presentation of the reference and reproduction affected the reproduced intervals. Seven subjects (three males and four females, with average age = 31.4, sd = 3.4) participated. Six out of the seven were the subjects in experiment 1. In the experiment, a fixation appeared for 1 s at the center of the monitor, followed by a blank screen of 1 s. A sine wave sound (1000 Hz, 60 dB) was presented through the headphone as the reference interval. Reference intervals of 1, 3 or 5 s duration were presented in random order. The subjects were instructed to memorize the durations of the reference sound. After the presentation of the reference interval, the Go signal (sine wave, 2000 Hz) was presented. The interval (fixed throughout the block) between the offset of reference and the Go signal was 1 or 2 s. The two intervals were switched alternatingly with the conditions. The order of the two conditions was counterbalanced among the subjects. The subjects were instructed to press the key to reproduce the reference interval at their own timing after the Go signal. One block consisted of 18 trials, in which each reference interval was presented six times. There were 74 trials in total, in which the subjects reproduced each reference interval 12 times for each condition. Before the experiment, the subjects conducted 10 practice sessions, in which the interval between the offset of reference interval and the Go signal was 1.5 s. General descriptions ~~~~~~~~~~~~~~~~~~~~ The measures used for investigating the effect of movement condition on temporal processing were the mean reproduction intervals, the ratio between reproduced and reference intervals, and the coefficient of variation (CV, standard deviation divided by the mean). The mean produced interval would indicate directly how the subject perceived and reproduced each reference interval. If the movement conditions affected temporal processing critically, the mean reproduced intervals would be significantly different between the conditions. The ratio between the reproduced and reference intervals would represent the degree of deviation of the reproduced interval from the actual stimuli. The CV values would indicate the variability of temporal reproduction within the subject. As noted in the discussion section, CV is one of the useful measures for investigating cognitive models of time perception. Experiment 1 ~~~~~~~~~~~~ Table 1 and Fig. 3(a) show the results for the self-timed and automatic conditions in experiment 1. The reproduced intervals were significantly longer than the actual intervals at 1 s (one sample t-test, t(9) = 8.75, p < .001) and 3 s (t(9) = 2.86, p = .018), but were not significantly different at 5 s (t(9) = 0.72, p = .48) in the self-timed condition. In the automatic condition, only the reproduced interval at 1 s (t(9) = 4.02, p < .01) was significantly different from the actual interval (3 s: t(9) = 1.94, p = .083, 5 s: t(9) = −0.19, p = .85). The mean reproduced intervals, ratio and CVs were submitted to a 2 × 3 repeated measures analysis of variance (ANOVA). The movement conditions (automatic versus self-timed) and the reference intervals (1, 3, 5 s) were a categorical variable and a covariate, respectively. The reproduced intervals were significantly different between the conditions (F(1, 9) = 8.17, p = .018). The effect of reference intervals was also significant (F(1, 9) = 96.79, p < .00001), indicating that the subjects were sensitive to the intervals and could discriminate the durations. Their interaction was not significant (F(1, 9) = 4.86, p = .054). For the ratio, the effect of reference interval was significant (F(1, 9) = 102.77, p < .00001). The effect of movement condition (F(1, 9) = 4.16, p = .072) and the interaction (F(1, 9) = 0.089, p = .77) were not significant. For CVs, ANOVA detected a significant effect of reference interval (F(1, 9) = 9.01, p = .014) but not for the movement condition (F(1, 9) = 1.47, p = .25) and their interaction (F(1, 9) < 0.0001, p = .99), indicating that the subjects attended equally to the two movement conditions. The timings of key pressing for reproduction initiation after the Go signal were 319 ± 25 ms and 1272 ± 391 ms (mean ± sd) in the automatic and the self-timed conditions, respectively, with a significant difference between them (paired t-test, t(9) = 7.43, p < .0001). The timing of the key pressing in the automatic condition was significantly smaller than the designated limit of 500 ms (paired t-test, one tail, t(9) = −23.86, p < .000001), confirming that 500 ms was a sufficient interval to press the key in response to the Go signal. Table 2 and Fig. 3(b) show the mean reproduced intervals, ratio and CVs for each reference interval for the self-timed and controlled conditions in experiment 1. The reproduced intervals were significantly longer than the actual interval at 1 s (one sample t-test, t(7) = 8.49, p < .001) and 3 s (t(7) = 4.13, p < .01), but not at 5 s (t(7) = −0.39, p = .70) in the self-timed condition. Similarly, in the controlled condition, the reproduced intervals at 1 s (t(7) = 11.82, p < .001) and 3 s (t(7) = 3.04, p = .018) were significantly longer than the actual duration, but not so at 5 s (t(7) = 0.20, p = .84). As in the automatic condition, the data were submitted to the 2 × 3 repeated measures ANOVA. The reference interval significantly affected the reproduced intervals (F(1, 7) = 89.85, p < .0001). In contrast to the automatic condition, the effect of the movement condition on reproduced interval was not significant (F(1, 7) = 0.49, p = .50). The interaction between the movement condition and the reference interval was also not significant (F(1, 7) = 1.04, p = .34). In the case of the ratio, the effect of reference interval was significant (F(1, 7) = 111.05, p < .0001). Neither the effect of movement condition (F(1, 7) = 0.30, p < .60) nor the interaction (F(1, 7) = 0.001, p = .98) was significant. The ANOVA showed that the CVs were significantly different between different reference intervals (F(1, 7) = 5.70, p = .048). The effect of the movement condition (F(1, 7) = 4.56, p = .069) and the interaction between the reference interval and the movement condition was not significant (F(1, 7) = 0.74, p = .41). Contrary to the automatic condition, the subjects started the temporal reproduction earlier after the Go signal in the self-timed condition (mean ± sd, 1266 ± 249 ms) than in the controlled condition (1530 ± 97 ms). The timing of starting reproduction was significant (paired t-test, t(7) = −0.59, p = .035), suggesting that the difference of reproduced intervals in experiment 1(a) was not simply the result of the delay between the Go signal and the timing of key pressing. If the latency of key pressing had a significant effect on reproduced intervals, there would have been significant differences of reproduced intervals here. Experiment 2 ~~~~~~~~~~~~ Results for the mean reproduced interval, ratio and CV in the auditory condition of experiment 2 are shown in Table 3 and Fig. 4(a). The reproduced intervals were significantly longer than the actual interval at 1 s in both the automatic and controlled conditions, and at 3 s in controlled condition. (one sample t-test, automatic condition: 1 s (t(4) = 4.93, p = .0078), 3 s (t(4) = 2.32, p = .080), 5 s (t(4) = 0.17, p = .86), controlled condition: 1 s (t(4) = 5.88, p = .0041), 3 s (t(4) = 2.95, p = .041), 5 s (t(4) = 0.72, p = .50). A 2 × 3 repeated measures ANOVA detected significant effects of movement conditions (F(1, 4) = 35.84, p = .0039) and reference interval (F(1, 4) = 37.54, p = .0035). The effect of the interaction was not significant (F(1, 4) = 0.053, p = .82). As for the ratio, both the effect of movement condition (F(1, 4) = 14.39, p = .019) and reference interval (F(1, 4) = 49.29, p = .0022) were significant, while the interaction of them was not significant (F(1, 4) = 4.92, p = .091). We submitted the CVs to a 2 × 3 repeated measures ANOVA. The effects of the reference interval (F(1, 4) = 7.85, p = .048) was significant. The effect of the movement conditions (F(1, 4) = 1.43, p = .29) and the interaction (F(1, 4) = 1.12, p = .34) were not significant. The latency from the Go signal to the key pressing was 357 ± 34 ms (mean ± sd) and 1409 ± 143 ms in the automatic and controlled conditions, respectively. The latency was confirmed to be different between conditions (paired t-test, one-tailed, t(4) = −15.10, p < .0001). Results for the reproduced interval, ratio and CV in the visual condition of experiment 2 are shown in Table 4 and Fig. 4(b). The reproduced intervals were significantly longer than the actual intervals at 1 s, but not at 3 s and 5 s in both the automatic and controlled conditions (one sample t-test, automatic condition: 1 s (t(4) = 4.06, p = .015), 3 s (t(4) = 0.55, p = .61), 5 s (t(4) = −1.11, p = .32), controlled condition: 1 s (t(4) = 4.20, p = .013), 3 s (t(4) = 1.45, p = .22), 5 s (t(4) = −0.34, p = .75). A 2 × 3 repeated measures ANOVA detected significant effects of movement conditions (F(1, 4) = 8.90, p = .040) and reference interval (F(1, 4) = 14.94, p = .018) but not their interaction (F(1, 4) = 0.28, p = .62). For the ratio, the significant effects of the movement condition (F(1, 4) = 10.56, p = .031), the reference interval (F(1, 4) = 17.32, p = .014) and their interaction (F(1, 4) = 13.56, p = .021) were detected. We submitted the CVs to a 2 × 3 repeated measures ANOVA. The effect of the reference interval (F(1, 4) = 3.38, p = .13), movement conditions (F(1, 4) = 0.55, p = .49) and the interaction of them (F(1, 4) = 0.028, p = .87) were not significant. The latencies from the Go signal to the key pressing were 380 ± 10 ms (mean ± sd) and 1350 ± 130 ms in the automatic and controlled conditions, respectively. The latency was different between the automatic and controlled conditions (paired t-test, one-tailed, t(4) = −15.97, p < .0001). Experiment 3 ~~~~~~~~~~~~ Table 5 and Fig. 5 show the mean reproduced intervals, ratio and CVs in experiment 3. When reproduced intervals were submitted to a 2 × 3 repeated measures ANOVA, a significant effect of reference interval (F(1, 6) = 135.89, p < .0001) was detected. Neither the effect of movement condition (F(1, 6) = 1.28, p = .30) nor the interaction (F(1, 6) = 0.52, p = .49) was significant. For the ratio, the effect of reference interval was significant (F(1, 6) = 12.99, p = .011). The effect of movement condition (F(1, 6) = 2.26, p = .18) and interaction (F(1, 6) = 1.93, p = .21) were not significant. We submitted the CVs to a 2 × 3 repeated measures ANOVA. The effect of the reference interval (F(1, 6) = 2.79, p = .14), movement conditions (F(1, 6) = 0.093, p = .77) and the interaction (F(1, 6) = 2.16, p = .19) were not significant.","In this study, the subjects reproduced the reference intervals by pressing the key. The purpose of this study was to investigate whether the subjective duration was affected by the nature of voluntary movement of key pressing in various contexts of action initiation. In addition, we examined how the effect on temporal reproduction varied as a function of the reference intervals. The results for experiment 1 suggest that when the subject pressed the key in a reflex-like manner (within 500 ms, automatic condition), the reproduced intervals became shorter compared to the timing of self-timed key pressing. Limiting the timing of key press between 1 and 2 s after the Go signal (controlled condition) presumably made the action highly volitional and cognitively controlled, as the subject was required to suppress the movement for some time and then initiate the movement before the deadline. This suppression would have necessitated cognitive modulations from top-down control circuits including the prefrontal brain area. These results suggest that the temporal estimation was not significantly different between self-timed and cognitively controlled movements. On the other hand, the reproduced interval in self-timed condition was significantly different from the automatic condition, suggesting that the automaticity of action initiation has a significant effect on interval estimation within the range of several seconds. The results of experiment 2 suggest that when the subjects started the reproduction in the automatic and reflex-like manner, their reproduced intervals were significantly shorter than when movements were initiated in a cognitively controlled manner. Since the subjects did not know which movement condition would occur when the reference intervals were encoded, the difference of reproduced interval could not be attributed to the encoding phase. Comparison with results of experiment 1 would suggest that the online estimation of interval was affected by the context of action initiation. The results of auditory and visual conditions also suggest that the effect was independent of the modalities. These results are consistent with a model involving a common timing mechanism for both visual and auditory modalities. The results of experiment 3 show that the delay of reproduction period did not affect the reproduced interval. It has been suggested that the interval between encoding and reproduction period affected the reproduced interval (Wearden, Goodson, & Foran, 2007). The delay used in Wearden et al. (2007) however, was larger than those employed in our present experiments. In addition, they showed that larger delays shortened memorized intervals, contrary to our results in experiments 1 and 2. Therefore, the interval between the offset of the reference interval and the timing of key press might not be a critical factor for the reproduced interval in our experiment. The differences of reproduced intervals are likely to have been caused by the contexts in which the movements were initiated. It has been suggested that different mechanisms are engaged in the processing of intervals under and above about 3 s (termed “time perception” and “time estimation”, respectively, Fraisse, 1984; Poppel, 1997; Ulbrich et al., 2007). In experiments 1 and 2, the interactions between the reference interval and the movement condition were not significant. Thus, the effect of the movement condition on duration reproduction was constant over these ranges, suggesting that a common mechanism might underlie “time perception” and “time estimation”. In a review of brain imaging studies, Lewis and Miall (2003, 2006) suggested that different brain networks were engaged in temporal processing, depending on whether the movements were automatic or cognitively controlled. Cognitively controlled timing task involves the right dorsolateral prefrontal cortex (rDLPFC) and the right posterior parietal cortex (rPPC). On the other hand, automatically controlled timing task involved the supplementary motor area (SMA), left sensorimotor cortex, and the right cerebellum. It is possible that the significant difference of reproduced intervals in different movement conditions resulted from the activation of different neural networks between the conditions, while the absence of the significant differences in reproduced durations between the self-timed condition and controlled condition in experiment 1 is due to the fact that similar networks were activated in the two conditions. For example, the dorsolateral prefrontal cortex might be activated and play a significant role in both self-timed and cognitively controlled movements. In the present study, some results were consistent with those in the previous studies of chronostasis, while others were contradictory. As is the case with the chronostasis (Yarrow, Haggard et al., 2004), the effect size was constant across the estimated intervals. Our results, on the other hand, showed that when the subject began the temporal reproduction in a reflex-like manner (automatic condition), the reproduced intervals were shorter than those in the self-timed or cognitively controlled actions (controlled condition). This result is in contradiction to the results obtained in the study of chronostasis in which the lengthened effects by saccades were not significantly different between the high volitional and reflex-like eye movements (Yarrow, Johnson et al., 2004). Yarrow, Johnson et al. (2004) concluded that the effect of chronostasis was triggered by an efference signal arising in the superior colliculus. The difference between our results and the chronostasis studies may be attributed to the difference in the neural mechanism of saccades and finger movements, and the range of estimated intervals involved. Our results suggest relative differences of interval estimation depending on the movement conditions. The automatic action might have dilated the subjective duration, while the self-timed and cognitively controlled action might have compressed it. It is also possible that both effects have co-existed. The pacemaker- accumulator internal clock model has been developed to account for interval estimation (Treisman, 1963). One of the applications of the model is the Scalar Expectancy Theory (SET, Gibbon, 1977; Gibbon, Church, & Meck, 1984). In SET, there are three stages (clock, memory and decision), each stage containing a few components. In the clock stage, it is assumed that there is an internal clock which emits pulses with a certain rate. Those pulses then pass a gate of accumulator. The gate is open by default. When the cognitive process of estimating intervals is initiated, the switch of the gate is closed and pulses are accumulated. The interval is then measured by the amount of pulses in the accumulator. The SET predicts the scalar property of variance: The distribution of estimated interval would scale in proportion to the mean of estimated interval, so that the coefficient of variation would be constant, as the estimated interval changes. One of the factors affecting the clock rate is arousal. It has been suggested that a higher arousal level correlates with a faster clock rate. A click train might increase the arousal level, resulting in an overestimation for the following intervals (Penton-Voak, Edwards, Percival, & Wearden, 1996). Some features of the present results can be explained by the SET model. Specifically, the results presented here show that the difference of estimated intervals between action contexts were constant across the range of estimated intervals. A change in the clock rate would predict a proportional difference as a function of the interval, while a change in the latency of closing switch would predict a constant difference regardless of the estimated intervals. Our results are compatible with an account in terms of the closing latency: When the subjects initiated action in a reflex- like manner (automatic condition), the latency of closing switch was shorter. Our results showed that the CV values decreased as a function of estimated durations. The decrease in CV values is consistent with the results of Ulbrich et al. (2007), where the subjects reproduced intervals between 1 and 5 s. On the other hand, the SET model predicts that CV would be constant for changes of the interval. The discrepancies indicate that the SET model might have to be modified. It is possible that different mechanisms are engaged in measuring a temporal interval, depending on whether the interval is shorter or longer than about 3 s (Fraisse, 1984; Poppel, 1997; Ulbrich et al., 2007). Higher cognitive processes would be involved in the processing of longer intervals. Measurement of shorter intervals would be carried out mainly by lower neural mechanisms, where temporal processing may be affected by various noises. Because of these noises, the measured duration would be variable for shorter intervals. When the target interval is longer, the higher cognitive processes would be engaged in measuring the duration. The higher cognitive processes may compare the interval being measured with the interval stored in memory and compensate for the deviation caused by the noises. As a result, the measured intervals for longer durations would have lower variability than for the shorter ones. Such a compensating mechanism might have to be incorporated to supplement the SET model. There are alternative explanations for our results. Wenke and Haggard (2009) investigated how a subject’s sense of time changed when a temporal attraction happened. The results indicated that the threshold for a successful temporal discrimination for tactile stimuli on the subject’s finger became larger when the stimuli were presented right after their action, while the subjects pressed the key voluntarily. Within the framework of internal clock model, the clock would have slowed down transiently by the initiation of voluntary movements. Although their study suggested that this effect occurred only when the action led to a sensory feedback such as an auditory tone, our results can be explained by this transient clock rate change assumption. A slower clock rate in the self-timed condition would lead to longer reproduced intervals than in the reactive movement condition (i.e., the automatic condition). In addition, the opposite effect, in which the rate of the internal clock became faster transiently, can also be considered. Welchman et al. (2010) showed that reactive action was conducted quicker than self initiated movements. This phenomenon suggests that the progress of the internal clock would get transiently faster when the subjects initiated a movement in a reflex-like manner. This consideration also fits the present results, as a faster speed of internal clock would lead to shorter reproduced interval. It is possible that either one or both effects led to the present results. In experiment 1, the automatic and controlled conditions with limits on the timing of key press were more demanding than the self-timed condition. Therefore, the subject’s attentional allocation might have been different between these conditions. Task difficulty is known to affect temporal processing (Zakay, Nitzan, & Glicksohn, 1983). The dual task paradigm studies have also suggested that loads of temporal and nontemporal concurrent tasks affected the attentional allocation and resulted in the temporal processing of other tasks (interference effect) (Brown & West, 1990; Brown, 1997; Dutke, 2005). A concurrent task such as visual search and arithmetic demands attentional resources and interfere with the subject’s temporal task, leading to a decrease in the pulses accumulated in the memory system. In view of these results, in our experiments, it is possible that pressing the key within the predetermined time limit demanded in the subjects some degree of attentional allocation, resulting in longer durations in the externally-timed conditions (automatic and controlled) than in the self-timed condition. Subjects who participated both in the automatic and controlled conditions in experiment 1 reported that the latter was more difficult to execute than the former. If the task load and attentional allocation affected the temporal reproduction period critically, the reproduced intervals in the automatic condition would be longer than those in the self-timed condition. In addition, the dilation of reproduced intervals in the controlled condition compared to those in the self-timed condition would be more enhanced than in the case of automatic condition. However, such were not the case. Therefore, we conclude that it is not likely that attention and task load led to the present results. In the controlled condition, before starting the reproduction of reference interval, the subjects had to estimate a duration of 1–2 s. There is thus a possibility that this additional temporal estimation process might have affected the following reproduction process of reference interval. However, when comparing the CV of the controlled condition with other conditions in experiment 1 (b) and experiment 2, the statistical analysis did not detect significant differences, indicating that the variability of the reproduced intervals were not statistically different whether there was an additional temporal estimation before the reproduction. In addition, in experiment 1 (b), the reproduced intervals were not significantly different between the self-timed and controlled conditions, indicating that the additional estimation of the interval of 1–2 s before the reproduction of reference interval did not affect the reproduced interval significantly. Based on these observations, we conclude that the estimation of 1–2 s in the controlled condition did not significantly affect the reproduction process. In sum, we investigated the effect of the context of action initiation on temporal reproduction. The subjects reproduced three kinds of reference intervals by pressing the key. When the subjects tried to reproduce the duration by reflex-like movements, the reproduced intervals were significantly shorter than those for self-timed and cognitively controlled movements. In the context of the internal clock model, these results would indicate that when the subjects initiated to move in a reflex- like manner, the latency of switch closing would be faster, or the clock speed would become faster. Such changes would affect the perception of duration, as a result of the interaction between our action and environment."],["We explored links between two perfectionism facets and alcohol-related problems. We predicted perfectionistic cognitions and nondisplay of imperfection would indirectly predict alcohol problems through negative affect, coping motives, and conformity motives, but would be unrelated to quantity of alcohol consumption. Participants included 263 young adult drinkers collected from two sites using self-report surveys with a 21-day, once-per-day measurement. Participants were mostly Caucasian (78.3%), female (79.5%), and young (M = 21.37, SD = 1.89). Data were analyzed using multilevel structural equation modeling. Nondisplay of imperfection (but not perfectionistic cognitions) had a serial indirect effect on alcohol-related problems through negative affect, followed by conformity motives. Other findings varied across analyses (fixed vs. random) and analysis level (between vs. within). Open Data/Methods: https://osf.io/gduy4. --------------------------------------------------------------------------------","Heavy and problematic drinking are major societal concerns. Excessive alcohol consumption and alcohol-related problems are especially common among young adults, with over 40% of Canadian undergraduates endorsing at least one alcohol-related problem (e.g., arguing with other while drunk, neglecting responsibilities) per year (Adlaf, Demers, & Gliksman, 2004). Links between negative forms of perfectionism (e.g., perfectionistic concerns) and alcohol problems are inconsistent, with some studies suggesting a positive link (Flett, Hewitt, Whelan, & Martin, 2007; Rice & Van Arsdale, 2010), and others finding null effects of perfectionism on binge drinking (Flett et al., 2008; Mackinnon et al., 2011). However, there is a clear theoretical link between negative perfectionism (i.e., perfectionistic thoughts and behaviors aimed at avoiding negative outcomes). and negative reinforcement (Slade & Owens, 1998), which suggests that negatively-reinforcing drinking motives (i.e., consuming alcohol to reduce negative emotions; Cooper, Kuntsche, Levitt, Barber, & Wolf, 2015) may be one mechanism by which perfectionism can confer risk for alcohol-related problems. Given the role of perfectionism as a transdiagnostic risk factor for a wide range of psychopathology (Sherry, Mackinnon, & Gautreau, 2015), it seems prudent to investigate whether facets of this personality trait also confer risk for alcohol problems. The present study examined the relationship between perfectionistic cognitions, nondisplay of imperfection, avoidance-based drinking motives, negative affect, alcohol consumption, and alcohol-related problems using a 21-day daily diary study of young adults. Defining perfectionism ~~~~~~~~~~~~~~~~~~~~~~ Although variably defined by scholars, most researchers agree perfectionism is a multidimensional personality trait comprising two higher-order dimensions, labeled perfectionistic concerns and perfectionistic strivings (Dunkley, Blankstein, Halsall, Williams, & Winkworth, 2000; Stoeber & Otto, 2006). Perfectionistic concerns include the perception that others impose unrealistic standards and expectations upon oneself (i.e., socially prescribed perfectionism; Hewitt & Flett, 1991), constant worry about making mistakes (concern over mistakes; Frost, Marten, Lahart, & Rosenblate, 1990), doubts about getting things “right” (doubts about actions; Frost et al., 1990), and a sense of falling short of one’s own standards (discrepancy; Slaney, Rice, Mobley, Trippi, & Ashby, 2001). Perfectionistic strivings include setting high personal standards (personal standards, Frost et al., 1990) and demanding perfection of oneself (i.e., self-oriented perfectionism; Hewitt & Flett, 1991). Perfectionistic concerns are consistently associated with undesirable outcomes, such as negative affect (Molnar, Reker, Culp, Sadava, & DeCourville, 2006; Smith, Saklofske, Yan, & Sherry, 2016) and decreased well-being (Mackinnon & Sherry, 2012). Findings on the link between perfectionistic strivings and positive psychological outcomes are less clear. After controlling for perfectionistic concerns, some evidence suggests strivings are positively associated with well-being (Hill, Huelsman, & Araujo, 2010; Stoeber & Otto, 2006) and low negative emotionality (Lizmore, Dunn, & Causgrove Dunn, 2017; Smith et al., 2016), while others found null or mixed relations between strivings and positive characteristics (Graham et al., 2010; Mackinnon & Sherry, 2012). Perfectionistic cognitions are a private expression of perfectionism and reflect the frequency of automatic ruminative thoughts about perfection (Flett, Hewitt, Blankstein, & Gray, 1998). Flett and Hewitt (2014) note that “At the conceptual level, we were focused from the outset of perfectionistic automatic thoughts pertaining to the self” (pg. 662). That is, they believed this construct was more closely related to self-oriented perfectionism, which perfectionism scholars broadly agree is part of the “perfectionistic strivings” dimension (Smith et al., 2016). However, based on factor analytic data, Stoeber, Kobori, and Brown (2014) argue that the perfectionism cognitions are multidimensional, with features of both perfectionistic strivings and perfectionistic concerns. Supporting this view, Flett, Hewitt et al. (2012) found strong associations between self-oriented and socially prescribed perfectionism and perfectionistic cognitions. A study by Sherry, Gralnick, Hewitt, Sherry, and Flett (2014) also found moderate associations between perfectionistic thoughts and all three of Hewitt and Flett (1991) trait dimensions of perfectionism (i.e., self-oriented, other-oriented, and socially prescribed perfectionism). Similar to findings on perfectionistic concerns, research links perfectionistic cognitions to rumination (Flett, Madorsky, Hewitt, & Heisel, 2002) and depressive symptoms (Flett et al., 1998, 2002, 2007; Flett, Hewitt, et al., 2012; Wimberley & Stasio, 2013). Given this debate, perfectionistic cognitions may not cleanly fit into the two higher-order traits of perfectionistic strivings and concerns, though it does appear to be relatively maladaptive. Perfectionistic self- presentation is a public expression of perfectionism and is thought to have three facets: perfectionistic self-promotion (i.e., self-promoting behaviors aimed at demonstrating perfection), nondisclosure of imperfection (i.e., avoiding verbal admissions of imperfection), and nondisplay of imperfection (i.e., concealing imperfect behaviors from others; Hewitt et al., 2003). The present study focuses on nondisplay of imperfection. Nondisplay of imperfection is closely linked to social anxiety (Flett, Coulter, & Hewitt, 2012; Hewitt et al., 2003; Mackinnon, Battista, Sherry, & Stewart, 2014), depression (Hewitt et al., 2003), and negative affect (Hewitt et al., 2003; Mackinnon & Sherry, 2012; Mushquash & Sherry, 2012). Though all three facets of perfectionistic self-presentation correlate with negative outcomes, nondisplay of imperfection has the largest relationship with negative affect (Hewitt et al., 2003); thus, we chose to focus on this facet in the present study. Nondisplay of imperfection is also correlated, yet not redundant with, trait perfectionism (Flett, Nepon, Hewitt, Molnar, & Zhao, 2016; Hewitt et al., 2003). The largest correlations are found between socially prescribed perfectionism and nondisplay of imperfection, supporting the notion that these individuals are preoccupied with the consequences of not appearing perfect because they perceive others will be critical of them. Due to its malignant nature and robust association with socially prescribed perfectionism, nondisplay of imperfection is considered a close relative of perfectionistic concerns (Hewitt et al., 2003). Perfectionism cognitions and perfectionistic self-presentation as trait-states ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Perfectionism variables are typically conceptualized as stable trait-like aspects of personality, which reflect dispositional levels, or inherent qualities, that make up one’s character (Dunkley et al., 2000; Flett et al., 1998; Hewitt & Flett, 1991; Hewitt et al., 2003).1 The viewpoint taken in the present paper departs slightly from the common assumption of trait-like stability. Despite a core of trait-like stability, short-term longitudinal research also suggests these measures have day-to-day fluctuation. In a 21-day longitudinal study, Mackinnon et al. (2014) found state-like variations in perfectionistic cognitions and perfectionistic self-presentation, with variation from day- to-day accounting for about 15–16% of the overall variance in the constructs. Other authors found within-person fluctuation in perfectionistic concerns and strivings from day-to-day (Boone et al., 2012) and week-to-week (Mackinnon, Kehayes, Leonard, Fraser, & Stewart, 2017). In all three cases, perfectionism tends to be predominantly stable across time, but also has small fluctuations across measurement occasions. More importantly, this state-like variation in perfectionism variables predicts substantive outcomes, such as social anxiety (Mackinnon et al., 2014) and well-being (Mackinnon et al., 2017). Thus, the present study conceptualizes perfectionistic cognitions and nondisplay of imperfection as “trait-states,” which reflects that they are primarily stable traits, but also exhibit a smaller amount of state-like variability from day-to-day. Perfectionism & alcohol outcomes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Given that research and theory link negative affect to alcohol use and alcohol-related problems (Castellanos-Ryan & Conrod, 2012), and that negative perfectionism exhibits strong links to negative affect (Dunkley, Zuroff, & Blankstein, 2006), it seems reasonable to expect that perfectionistic concerns (and facets related to concerns) would be associated with alcohol outcomes indirectly through negative affect. Perfectionistic strivings have been linked to a small decrease in alcohol consumption and binge drinking, while perfectionistic concerns remain mostly unrelated to binge drinking (Flett et al., 2008; Mackinnon et al., 2011; Pritchard, Wilson, & Yamnitz, 2007; Simons, Christopher, & Mclaury, 2004). However, perfectionistic concerns have been linked to increased alcohol- related problems (i.e., negative consequences associated with alcohol use; Rice & Van Arsdale, 2010). Moreover, some authors found that people who formerly had an alcohol use disorder had higher levels of perfectionistic cognitions, self-oriented perfectionism, and socially prescribed perfectionism than those without alcohol problems (Flett et al., 2007; Sherry et al., 2012). This evidence suggests people high in perfectionistic concerns are at an increased risk for alcohol-related problems, even if they do not consume large quantities of alcohol. Perfectionism, negative affect, & avoidance-based drinking motives ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The relationship between negative perfectionism and negative affect (i.e., emotional states such as anger, sadness, and fear) is well-established. Hewitt, Flett, Sherry, and Caelian (2006) theorize individuals high in socially prescribed perfectionism experience negative emotions from day-to-day, as well as enduring levels of trait-like negative affect across the lifespan. Additionally, longitudinal research found self-critical perfectionism (a variation of perfectionistic concerns that includes harsh self-criticism) predicted negative affect 6 months and 3-years later, suggesting these individuals have chronic negative emotionality (Dunkley, Mandel, & Ma, 2014; Prud'homme et al., 2017). In cross-sectional research, perfectionistic concerns, perfectionistic cognitions, and nondisplay of imperfection are all positively related to negative affect (Dunkley et al., 2006; Dunkley, Zuroff, & Blankstein, 2003; Flett et al., 1998; Hewitt et al., 2003). Slade and Owens (1998) contend that people high in negative perfectionism (i.e., another dimension closely related to perfectionistic concerns) are largely motivated by negative reinforcement. That is, this personality style is driven primarily by the desire to avoid or escape negative experiences (e.g., to avoid making mistakes and the associated negative emotions). Because this form of perfectionism is linked to negative reinforcement and negative affect, and drinking motives theory states that some individuals consume alcohol to diminish negative emotions (Cooper, 1994), it seems reasonable individuals high in perfectionistic cognitions and nondisplay of imperfection drink for avoidance-based motives (i.e., they drink to avoid negative outcomes). Drinking motives theory proposes two avoidance-based reasons for consuming alcohol.2 Coping motives are self-focused avoidance motives and refer to drinking to avoid or diminish negative emotions, or to deal with threats to one’s self-esteem. Conformity motives are social avoidance motives and include drinking to gain approval from others, or avoid social disapproval. A recent meta- analysis by Cooper et al. (2015) found a small positive relationship between coping motives and quantity of alcohol use, and a small negative relationship between conformity motives and quantity of alcohol use. Yet, in this meta-analysis, only coping motives had a robust relationship with alcohol-related problems, while conformity motives did not. Because individuals high in perfectionistic concerns often drink for avoidance-based reasons, drinking can lead to adverse outcomes. Rice and Van Arsdale (2010) found young adults high in perfectionistic concerns reported more coping motives and had more alcohol- related problems than non-perfectionists (i.e., individuals with low levels of perfectionistic strivings and concerns). Unfortunately, these authors did not assess quantity of alcohol use or conformity motives – the latter is an important omission, given that it is also an avoidance-based motivation.3 In general, the literature finds that coping motives do not predict alcohol use, but result in alcohol-related consequences (LaBrie, Ehret, Hummer, & Prenovost, 2012; Merrill & Read, 2010; Merrill, Wardell, & Read, 2014; Patrick, Lee, & Larimer, 2011). Rationale and hypotheses ~~~~~~~~~~~~~~~~~~~~~~~~ Perfectionism is a transdiagnostic risk factor for psychopathology, particularly disorders that involve heightened negative affect (Sherry et al., 2015). Research shows a link between perfectionism facets and hazardous alcohol use (Flett et al., 2007; Sherry et al., 2012), however the mechanism underlying this relationship remains unknown. Though one prior study suggested drinking motives as a mediator (Rice & Van Arsdale, 2010), like most studies on perfectionism, it was cross-sectional and failed to account for daily variation in perfectionism dimensions. Our study methodologically improves upon Rice and Van Arsdale (2010) by using a 21-day, once-per-day measurement scale. Moreover, to our knowledge, no prior studies have examined whether perfectionistic cognitions and perfectionistic self- presentation (i.e., nondisplay of imperfection) predict avoidance-based drinking motives. Using a 21-day daily diary design analyzed using multilevel structural equation modeling (Preacher, Zyphur, & Zhang, 2010), the current study examined the relationships between perfectionistic cognitions, nondisplay of imperfection, coping and conformity drinking motives, and alcohol outcomes. Competing avoidance-based motives may explain why those high in perfectionistic concerns experience alcohol-related problems, but do not drink large quantities of alcohol. If individuals high in perfectionistic concerns drink for both coping and conformity motives, these motivations may compete, resulting in a null bivariate relationship between perfectionistic dimensions and quantity of alcohol use due to the supressing effect of conformity motives. Specifically, we hypothesized: Perfectionistic cognitions and nondisplay of imperfection will be positively related to alcohol problems, but unrelated to quantity of alcohol use. Perfectionism facets will positively predict negative affect, which in turn will positively predict coping and conformity motives, which will in turn both predict alcohol problems. That is, perfectionism facets will have a serial indirect effect on alcohol problems through negative affect and drinking motives. Conformity motives will have a suppressing effect when predicting quantity of alcohol consumption. That is, an inconsistent mediation model (MacKinnon, Fairchild, & Fritz, 2007) will result in no total effect of perfectionism facets on drinking quantity because coping motives will be positively related to drinking quantity, and conformity motives will be negatively related to drinking quantity. Though hypotheses were the same for the between-subjects (i.e., averaged across all 20 days) and within-subjects (i.e., co-occurring changes within any given day) portions of the model, the within-subjects portion of the model was of greater interest given its assessment of change from day-to-day. Open data and materials ~~~~~~~~~~~~~~~~~~~~~~~ All study materials, raw data, and syntax are open-access and can be located at https://osf.io/gduy4/). This includes all copies of questionnaires used (including measures not analyzed in this paper), the raw data for all measures, and the materials necessary to reproduce all analyses in this paper.","Participants were recruited from Halifax Regional Municipality (HRM) and Dalhousie University in Halifax, as well as from Montreal and Concordia University in Montreal. Eligibility criteria included: (a) being between 18 and 25 years of age; (b) having consumed a least 12 alcoholic drinks in the past year; and (c) having Internet access at home. We targeted emerging adults as they tend to engage in heavy episodic drinking (Johnston, O'Malley, Bachman, & Schulenberg, 2005) and experience alcohol-related problems (e.g., LaBrie et al., 2012; Hingson, Heeren, Zakocs, Kopstein, & Wechsler, 2002). Flyers advertising the study were placed around downtown Halifax and on Dalhousie University campuses. In addition, advertisements were posted online on Dalhousie’s and Concordia’s undergraduate participant pools in the Department of Psychology and Neuroscience, as well as an online local classified ad (Kijiji). Mean age of participants (N = 263) was 21.37 (SD = 1.89), with females making up the majority (79.8%) of the sample. Self-reported ethnicities were Caucasian/White (78.3%), African Canadian/Black (2.3%), Asian (7.7%), Middle Eastern (1.1%), Hispanic (2.7%), First Nations (0.8%), and Other (6.5%). Participants’ primary residence was either Nova Scotia (60.5%) or Quebec (39.5%). Since both samples underwent the exact same procedure, samples from Halifax and Montreal were combined into a single dataset, with regional differences explored at the end of the results section.","The current study was approved by Dalhousie University’s Social Sciences and Humanities Research Ethics Board and Concordia University’s Human Research Ethics Committee. Questionnaires were administered online using a custom-created database. Our online survey platform provider was Interceptum (https://interceptum.com/p/en). After viewing the recruitment ads, interested participants contacted our research assistant via email, who sent them more information. If they agreed to participate, they were sent an email with a link to the consent form and baseline questionnaire. After consent was electronically obtained, participants completed the baseline questionnaire on day 1 (∼30 min). Participants were required to complete the baseline questionnaire before beginning the daily questionnaires. In the present study, only demographic measures were used from the baseline questionnaire; the remainder of the measures were measured daily. On days 2 through 21, participants were sent an email with a link to the daily questionnaires. Respondents were instructed to complete the daily questionnaires every day at end of each day (i.e., right before bed), as they were asked about events that occurred in the past 24 h. Specifically, the time frame for daily questionnaires went from 4 am one day to 4 am the next day (e.g., “From 4 am on November 23 to 4 am on November 24”). Answering daily questionnaires took approximately 15 min. Because each question referred to a specific time and date, if participants forgot to complete a batch of questionnaires on a day, they could make it up by filling out two daily questionnaires the following day. After 48 h, participants were no longer able to complete that day’s questionnaire and did not receive compensation for that day. After participants completed the study, they were emailed a debriefing form. To reduce missing data, participants were compensated for each questionnaire they completed (i.e., $2 worth of Amazon gift cards per day, for a maximum of $42 over 21 days). Participants could alternatively receive course credit in an eligible psychology class. They were offered 1 bonus point for every 5 days they participated in the study for a maximum of 3 bonus credits and a $12 gift card. That is, every $10 earned could be exchanged for a credit point. Perfectionistic cognitions Perfectionistic cognitions were measured using the Perfectionism Cognitions Short Form. This adapted three-item scale was derived from Flett et al. (1998)’s Perfectionistic Cognitions Inventory by Mackinnon et al. (2014). Participants were to indicate how frequently (0 = not at all to 4 = all of the time) they related to each statement (e.g., “I expect to be perfect”) in the past 24-hours. This measure has demonstrated acceptable between and within-subjects reliability (R1F = 0.92, RKF = 0.99, RC = 0.74) and items had large factor loadings (0.88–0.95) in a prior daily diary study (Mackinnon et al., 2014). Nondisplay of imperfection Nondisplay of imperfection was measured using three items from the Nondisplay of Imperfection subscale of Hewitt et al. (2003)’s Perfectionistic Self-Presentation Scale. Mackinnon et al. (2014) created this short-form by choosing the items with the highest factor loadings from the original measure (e.g., “I was concerned about making errors in public”). This measure displayed good between and within- subjects reliability (R1F = 0.91, RKF = 0.99, RC = 0.72) and each item had large factor loadings (0.74–0.92) in a prior daily diary study (Mackinnon et al., 2014). Positive and negative affect scale (PANAS-X) Negative affect was measured using three reduced subscales of the PANAS-X (Watson, Clark, & Tellegen, 1988) and one scale from Mackinnon et al. (2014). Each subscale is made up of 3 items relating to feelings of guilt (“angry at self”), fear (“afraid”), hostility (“scornful”), and depressed affect (“blue”). Participants were asked to rate (1 = very slightly or not at all to 5 = extremely) how strongly they felt each emotion over the past 24 h. Items that make up the guilt, fear, and hostility subscales were derived from the long-form PANAS-X to better fit our daily design in order to reduce participant burden by choosing the items with the highest factor loadings in prior work. Items from the depressed affect scale show good reliability and validity (RF1 = 0.77, RKF = 0.99, RC = 0.81; Mackinnon et al., 2014). Short form guilt, fear, and hostility subscales used in the current study have not been previously measured, however the original negative affect subscale from the PANAS-X has great alpha reliability for past daily diary studies (α = 0.87; Watson et al., 1988). Drinking motives revised – short form (DMQ-R-SF) Drinking motives were measured using the Drinking Motives Revised – Short Form (DMQ-R-SF; Kuntsche & Kuntsche, 2009). This 12-item scale assesses four drinking motives: social (“Because it helps me enjoy a party”), enhancement (“Because it’s fun”), coping (“To forget my problems”), and conformity (“To be liked”). The current paper used only data from coping and conformity motives. Participants indicated how often (1 = strongly disagree to 4 = strongly agree) they drank for each reason in the past 24 h. The DMQ-R-SF has good internal consistency (Cronbach’s alpha ranging from 0.70 to 0.83), and excellent construct and predictive validity (e.g., coping motives were positively linked to academic problems and risky sexual behavior). Also, the factor structure is similar to the original DMQ-R (Kuntsche & Kuntsche, 2009). Alcohol problems checklist The Alcohol Problems Checklist (Simons, Gaher, Oliver, Bush, & Palmer, 2005) contains 12 problematic events that could occur as a result of consuming alcohol (e.g., “Taken foolish risks”). Participants checked “yes” or “no” for instances that may have occurred in the past 24 h because of their alcohol use. Therefore, a participant’s total score on this questionnaire could range from 0 (checking all “no”) to 12 (checking all “yes”). This measure was correlated with alcohol consumption and negative affect, demonstrating predictive validity (Simons et al., 2005). Daily alcohol use Alcohol use was measured by asking participants how many drinks they consumed in the past 24 h. Participants responded using a sliding bar that allowed them to choose any whole number of drinks that applied best to them. “One drink” was defined as half an ounce of absolute alcohol (e.g., a 12-ounce can or glass or bottle of beer or cooler, a 5-ounce glass of wine, or a drink containing 1 shot of liquor or spirits; National Institute on Alcohol Abuse and Alcoholism, 2003). Data analytic strategy First, descriptive statistics were used to summarize the dataset. Means and standard deviations were reported for total scores of all variables. Internal consistency of each variable was assessed using a multilevel adaptation of Cronbach’s alpha developed by Geldhof, Preacher, and Zyphur (2014). We first reported descriptive statistics on protocol compliance (e.g., days completed) prior to analysis. Because daily drinking motives questionnaires could only be completed on drinking days, only data from drinking days were analyzed in subsequent analyses. Following these descriptive statistics, data were analyzed using multilevel structural equation modeling with fixed slopes (Preacher et al., 2010). This approach uses all available data by partitioning the available variance into orthogonal between-subjects components (i.e., trait-like variance that does not vary across drinking days) and within-subjects components (i.e., state-like variance that changes across drinking days), analyzing each partitioned component of variance separately. Thus, most reported statistics (e.g., factor loadings, latent correlations, path coefficients) are reported twice, once at the between-subjects level and again at the within-subjects level. To assess whether multilevel modeling was appropriate, we examined intraclass correlations (ICCs). ICCs can range from 0 to 1, and indicate the proportion of the variance available to be explained at the between-subjects level. Generally speaking, these values should be neither too close to 0 nor too close to 1 for multilevel analysis to be appropriate (Preacher et al., 2010). Prior to hypothesis testing, the measurement model was tested using multilevel confirmatory factor analysis in Mplus 8 using MLR estimation for standard errors and fit indices, which is robust to violations of the normality assumption. Individual questionnaire items (3 items each) were used as indicators for latent variables for perfectionistic cognitions, nondisplay of imperfection, coping motives, and conformity motives. Item parcels were used for negative affect and alcohol problems (Little, Cunningham, Shahar, & Widaman, 2002). The negative affect latent variable incorporated four item parcels, using the subscale totals for guilt, fear, hostility, and depression as indicators. The 12 alcohol problems items were assigned to one of three parcels semi-randomly, by evenly distributing items based on their variance (i.e., avoiding the situation where all the rarely-endorsed items end up in a single parcel). Thus, there were 4 items each in parcel 1 (items 1, 10, 7, & 12), parcel 2 (items 3, 4, 2, & 9), and parcel 3 (items 8, 6, 5, & 11). Number of alcoholic beverages consumed was left as a manifest variable, as it was measured with only a single item. Fit indices, standardized factor loadings, and latent correlations were reported for the measurement model. Following the measurement model, a structural model was fit to the data as outlined in Fig. 1 to test hypotheses. Fit indicies, standardized factor loadings, and unstandardized path coefficients were reported for this analysis. Tests of indirect effects were conducted using a Monte Carlo method with 20,000 iterations (Selig & Preacher, 2008).4 Simulation studies suggest this approach is preferable to bootstrapping for multilevel models (Preacher & Selig, 2012). We tested serial indirect effects with two mediators (X → M1 → M2 → Y) using the formula outlined in Hayes (2018).5 If the 95% confidence interval of the indirect effect did not contain zero, we concluded that mediation occurred. A well-fitting model was defined as a Confirmatory Fit Index (CFI) and Tucker-Lewis Index (TLI) greater than 0.95, Root Mean Square Approximation of Error (RMSEA) values less than 0.06, and Standardized Root Mean Square Residual (SRMR) value less than 0.08 (Kline, 2011). Pilot testing measures ~~~~~~~~~~~~~~~~~~~~~~ Many of the measures were adapted from their original form by asking participants to report on “The past 24 h.” Thus, a small pilot study was conducted to test the psychometric properties of these measures prior to the main study. Criterion validity was assessed by correlating measures that asked participants to report on “The past 24 h” to identical measures that asked participants to report on “The past 3 years.” Internal consistency was tested using Cronbach’s alpha. A sample of 137 undergraduate drinkers (84.7% female) were recruited using the same criteria as the diary study, and completed questionnaires online. Both criterion validity and internal consistency were adequate for perfectionism cognitions (r = 0.66, α = 0.89), nondisplay of imperfection (r = 0.62, α = 0.86), alcohol problems (r = 0.43, α = 0.79), negative affect (r = 0.64, α = 0.94), coping motives (r = 0.78, α = 0.88), and conformity motives (r = 0.73, α = 0.87). Correlations were significant at p < .001. These pilot data suggested our measures would be reliable and valid for use in the diary study. These pilot data are open-access, and are available at https://osf.io/gduy4/. Protocol compliance ~~~~~~~~~~~~~~~~~~~ Twenty-one participants (8.0%) did not drink during the study; thus, daily data is missing for these individuals. However, out of these twenty-one participants, one participant did not complete any daily questionnaires (only the baseline), and two participants completed only one day of the daily questionnaires. All participants were required to complete baseline measures before beginning the daily measures, thus there was no missing data for demographic variables. Out of 20 days, participants completed on average 16.16 days (SD = 4.68), with an average of 4.70 (SD = 3.80) drinking days. See Supplementary Table 2 for frequencies of the number of drinking days for participants. On average, participants completed 7.87 (SD = 4.93) make-up days (i.e., submissions that were 1 day late), which suggests that these make-up questionnaires were essential to achieving low rates of missing data. With our design, it was possible to have a maximum 5260 data points (263 participants * 20 daily measurements). Of these, 1236 were drinking days with data, 3015 were non-drinking days, and 1009 were missing data (i.e., no data for the entire day). Thus, 29% of reporting days were drinking days. This is consistent with prior daily diary research on undergraduates, which found that 24% of reporting days were drinking days (Grant, Stewart, & Mohr, 2009). Thus, the final analysis included 1236 drinking days nested within 242 participants. Overall, there was comparatively little missing data and participants were generally good at filling out the daily questionnaires as scheduled. Non-drinking days and missing days with no data were selected out of the analysis using listwise deletion. Item-level missing data (<1%) were handled using a full information maximum likelihood approach.","Means, standard deviations, and alpha reliabilities are located in Table 1. Generally speaking, measures had adequate reliability. However, internal consistencies were low at the within-subjects level for alcohol problems, which suggests greater measurement error for this variable. Means were slightly higher for perfectionistic cognitions (d = 0.51) and nondisplay of imperfection (d = 0.52), and almost identical for quantity of alcohol consumption (d = −0.01) when compared to prior studies sampled from this population (Grant et al., 2009; Mackinnon et al., 2014). Means for negative affect (d = −0.42), coping motives (d = −0.19), conformity motives (d = −0.32), and alcohol problems (d = −0.41) were slightly lower, but comparable to those obtained in our psychometric study. Measurement model ~~~~~~~~~~~~~~~~~ ICCs ranged from 0.24 to 0.72, suggesting that there was substantial variation from day- to-day, but also trait-like stability across all 20 days (Table 2). Generally speaking, ICCs were large for perfectionism variables (0.65–0.72), moderate for drinking motives and affect (0.43–0.59), and small for drinking outcomes (0.24–0.29). This suggests that perfectionism variables are more trait-like, alcohol outcomes are more state-like, with negative mood and drinking motives falling somewhere in between. The measurement model fit the data well for all fit indices except for the overall χ2, χ2(300) = 555.46, p < .001, CFI = 0.97, TLI = 0.96, RMSEA = 0.03, SRMRwithin = 0.03, SRMRbetween = 0.06. Standardized factor loadings were all large and substantial, with the lower bound of the 95% confidence intervals for the factor loadings ranging from 0.44 to 0.89 at the within-subjects level and from 0.59 to 0.98 at the between-subjects level (Table 2). Generally speaking, factor loadings were more substantial at the between-subjects level than at the within-subjects level. Overall, these analyses suggested that the measurement model was adequate. Latent correlations from the measurement model are reported in Table 3. At the between-subjects level, perfectionism variables, drinking motives, negative affect, and alcohol problems were all positively intercorrelated with medium to large effect sizes (0.24–0.66). However, alcohol consumption was correlated only with alcohol problems (0.29), and was uncorrelated with all other variables at the between-subjects level. The within-subjects correlations mostly mirrored the between-subjects correlations, but with smaller effect sizes. Nondisplay of imperfection, drinking motives, negative affect, and alcohol problems were all positively intercorrelated (0.13–0.41). However, a few key differences emerged (a) Perfectionism cognitions were generally uncorrelated with other variables, save for a positive correlation with nondisplay of imperfection (0.38) and a weak negative correlation with alcohol consumption (−0.07); (b) the relationship between alcohol consumption and alcohol problems was slightly larger than the between-subjects correlation (0.47 vs. 0.29); and (c) alcohol consumption was weakly positively correlated with coping (0.07) and conformity (0.19) motives. In general, variables were correlated with each other in the expected directions, and supported moving forward with the structural model in subsequent steps. Structural model ~~~~~~~~~~~~~~~~ The final model is depicted in Fig. 1 and tests of indirect effects are reported in Table 4. Fig. 1 shows 95% confidence intervals for parameters; however, the point estimates, standard errors, and p-values are in Supplementary Table 5. This model fit the data well except for the overall χ2, χ2(300) = 555.46, p < .001, CFI = 0.97, TLI = 0.96, RMSEA = 0.03, SRMRwithin = 0.03, SRMRbetween = 0.06. The between-subjects analysis examined the trait-like component of variance that did not vary across 20 days. H2 was partially supported for the between-subjects model; nondisplay of imperfection (but not perfectionistic cognitions) had a serial indirect effect on alcohol problems through negative affect, coping motives, and conformity motives. However, H3 was not supported; there were no significant serial indirect effects when predicting number of alcoholic drinks consumed. Moreover, effects predicted by H3 were not in the expected direction, as the effects of coping and conformity motives were both positive. Overall, variables in the model predicted a substantial amount of variance in all outcome variables (R2 from 0.26 to 0.48) except for number of alcoholic drinks consumed (R2 = 0.03). The within-subjects analysis examined state-like variations from day-to-day (i.e., when X changes, does Y change with it?). Results were broadly similar to the between-subjects results, albeit with much smaller effect sizes (R2 from 0.03 to 0.20). Though the statistical significance of a few paths changed relative to the between-subjects model (e.g., conformity motives were a significant predictor of drinking outcomes in the within-subjects model), the overall pattern of the indirect effects was the same. That is, nondisplay of imperfection had a serial indirect effect on alcohol problems through negative affect, coping motives, and conformity motives, while all other indirect effects were non-significant. Though not directly hypothesized, there remain some interesting direct effects worthy of discussion. Negative affect had a significant, positive direct effect on alcohol problems. Moreover, it had a non-significant negative relationship with number of drinks consumed. Both of these are important to understand the pattern of results observed. The relationship between coping motives and alcohol problems was small after controlling for negative affect; in fact, it became non-significant in the within-subjects model. Nonetheless, because the serial indirect effect also incorporates the direct effect of negative affect on alcohol problems, the overall serial indirect effect can remain significant even in the presence of a non-significant relationship between coping motives and alcohol problems. It does, however, suggest that negative affect may be more important than coping motives when predicting day-to-day variation in alcohol problems. Similarly, in Fig. 1, it may look as though nondisplay of imperfection should have an indirect effect on number of drinks consumed through negative affect and conformity motives. However, because the effect of negative affect on number of drinks is negative (despite being non-significant), it exerts a small suppressing effect on the serial indirect effect, resulting in non-significance. Thus, it appears that a suppressing effect of negative affect may contribute to the lack of a total effect of perfectionism facets on drinking quantity. It is also worth noting that (once controlling for all other variables in the model) perfectionistic cognitions had a small negative direct effect on drinking quantity in the within-subjects model, and a small negative direct effect on alcohol problems in the between-subjects model, contrary to the expected positive relationship. Given the magnitude of the relationship, it may be spurious, but is worth noting. Regional differences ~~~~~~~~~~~~~~~~~~~~ Given that data was sampled from two sites in Canada (Halifax, Nova Scotia and Montreal, Quebec), we explored potential regional differences by entering province of residence (0 = Halifax, 1 = Montreal) as a between-subjects covariate into our final structural model. This model fit the data about equally well to the model without the covariate, χ2(313) = 575.49, p < .001, CFI = 0.97, TLI = 0.96, RMSEA = 0.03, SRMRwithin = 0.03, SRMRbetween = 0.05. Moreover, no null hypothesis test conclusions in Fig. 1 changed as a result of this covariate. However, Halifax participants had higher rates of drinking, 95% CI (−1.60, −0.48) and lower rates of coping motives, 95% CI (0.02, 0.17) in this SEM model.6","We also explored a random slopes model, as there was substantial variation in the slopes. In this model, variances were estimated for the random slopes, but the covariances of random components were constrained to be zero for model parsimony (i.e., a variance components model). We also used a Bayesian estimator for more efficient convergence of this complex model. Indirect effects were calculated with the delta method using MODEL CONSTRAINT in Mplus. These results are summarized in Supplementary Fig. 1 and Supplementary Table 4. Results of null hypothesis tests were broadly similar to the fixed slopes model, with a few differences. The 3 direct effects of both perfectionism variables on alcohol outcomes that were originally statistically significant in the fixed model became non-significant. At the within-subjects level, the relationship between perfectionism cognitions and conformity motives became statistically significant and the relationship between nondisplay of imperfection and coping motives became non-significant. For indirect effects, the the Nondisplay → Negative Affect → Coping → Problems indirect effect at the within-subjects level became non-significant, and the Nondisplay → Negative Affect → Conformity → Drinks indirect effect became statistically significant at the between level. In general, the effect sizes using R2 also tended to be much larger in this model. Overall, effects that differ across fixed vs. random slopes models are unstable, and should be considered less reliable than effects that replicate across analyses.7","The current study examined the relationship between perfectionistic cognitions, nondisplay of imperfection, avoidance-based drinking motives, negative affect, and alcohol outcomes. We used multilevel structural equation modeling to assess whether these perfectionism facets predicted alcohol problems when averaged across 20 days (i.e., between-subjects) and whether changes in these facets predicted co-occurring changes in alcohol problems from day-to-day (i.e., within-subjects). There was significant variance to be explained at both the between- and within-subjects level, suggesting the variables under study have both trait-like and state-like qualities. Perfectionistic cognitions and nondisplay of imperfection were the most trait-like, which is in line with the notion that perfectionism is a relatively stable personality trait (e.g., Hewitt & Flett, 1991). However, there was also significant day-to-day fluctuations. This supports other short-term longitudinal studies suggesting perfectionism dimensions can be interpreted as more dynamic, state-like features of personality (Boone et al., 2012; Graham et al., 2010; Mackinnon et al., 2014). Negative affect, coping motives, and conformity motives fluctuated more than perfectionistic cognitions and nondisplay of imperfection on a day-to-day basis, with approximately an equal amount of variance to be explained at both the between- and within- subjects level. Lastly, alcohol-related problems had the largest amount of daily variation, suggesting they are dynamic and may be influenced by situational and environmental factors. Nondisplay of imperfection and alcohol problems ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The core finding was that nondisplay of imperfection indirectly predicted alcohol problems (but not quantity of alcohol consumption) through negative affect, coping motives, and conformity motives. These findings replicate previous empirical literature on the link between nondisplay of imperfection and negative affect (Hewitt et al., 2003; Mackinnon et al., 2014), and negative affect and avoidance-based drinking motives (Cooper et al., 2015; Todd et al., 2005). These results are also in line with evidence indicating coping motives are robustly associated with alcohol-related problems independent of heavy drinking (Cooper, Frone, Russell, & Mudar, 1995; Kuntsche & Cooper, 2010; Merrill & Read, 2010; Piasecki et al., 2014; Simons, Gaher, Correia, Hansen, & Christopher, 2005). Individuals high in nondisplay of imperfection tend to engage in this self-presentation strategy to compensate for perceived imperfections, yet often have a difficult time modifying how they present themselves to others (Hewitt et al., 2003). This often leads to feelings of inadequacy and decreased self-efficacy. Thus, it is reasonable that these individuals drink to cope with negative emotions, resulting in adverse alcohol-related consequences. However, it is worth noting that the Bayesian random effects model found that the indirect effect of nondisplay of imperfection on alcohol problems through negative affect and coping motives was non-significant at the within-subjects level, and that the indirect effect of nondisplay of imperfection on drinking quantity through negative affect and conformity motives became statistically significant at the between-subjects level. Thus, these relationships are comparatively more uncertain, pending future replication. Perfectionism cognitions and alcohol problems ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Contrary to H2, perfectionistic cognitions did not predict negative affect at either the between- or within-subjects level once controlling for nondisplay of imperfection. Thus, it also failed to indirectly predict coping motives, conformity motives, alcohol-related problems, and quantity of alcohol consumption. Perfectionistic cognitions did tend to be positively correlated with all other outcomes except for alcohol consumption at the between-subjects bivariate level (Table 3), consistent with prior cross-sectional work (Flett et al., 1998; Flett, Hewitt, et al., 2012). However, it did have a small negative relationship with alcohol outcomes once controlling for nondisplay of imperfection in some models, consistent with some theorists’ contention that certain perfectionism facets are adaptive (Stoeber & Otto, 2006). Nonetheless, it did not predict outcomes at the within- subjects level (i.e., co-occurring changes from day-to-day), nor did it predict negative affect beyond nondisplay of imperfection in the structural model. Thus, perfectionism cognitions were not related to alcohol outcomes in a substantive way. Though the Bayesian random slopes model suggested an indirect effect of perfectionistic cognitions on drinking quantity through negative affect and conformity motives at the between subjects level, this effect was primarily due to the positive relationships between (a) negative affect and conformity and (b) conformity and drinking quantity, since perfectionistic cognitions had weak, non-significant relationships with negative affect and conformity. Stoeber et al. (2014) posit that perfectionistic cognitions should be assessed as a multidimensional construct. These authors proposed that the Perfectionism Cognitions Inventory (PCI) measures three distinct factors (i.e., perfectionistic concerns, perfectionistic strivings, and perfectionistic demands), each relating to different psychological outcomes. Stoeber et al. (2014) found that thoughts about perfectionistic concerns predicted negative affect, while cognitions about perfectionistic strivings predicted positive affect. Two of the subscales used to measure perfectionistic cognitions in our study (“I expect to be perfect” and “My work should be flawless”) were shown to have high factor loadings on the perfectionistic strivings dimension. By using the PCI-Short Form and failing to partial out these two factors, mutual suppression could have occurred (i.e., cognitions about perfectionistic strivings could have supressed the effect of cognitions about perfectionistic concerns), leading to the lack of association between perfectionistic cognitions and negative affect in the current study. It also may be that individuals high in perfectionistic cognitions have coping strategies to deal with frequent perfectionistic thoughts. Indeed, research has found that individuals high in personal standards perfectionism (i.e., perfectionistic strivings) engage in problem- focused coping strategies (Dunkley et al., 2000; Prud'homme et al., 2017) and positive reinterpretation (Dunkley et al., 2014). These tactics could lead to fewer negative emotions and thus fail to predict avoidance-based reasons for drinking and alcohol-related problems. Perfectionism facets and quantity of alcohol consumption ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In H3, we proposed conformity motives would have a suppressing effect, leading to inconsistent mediation (MacKinnon et al., 2007) between nondisplay of imperfection, perfectionistic cognitions, and quantity of alcohol consumption. That is, we proposed a positive effect of coping motives and a negative effect of conformity motives, leading to a null total effect of perfectionism facets on alcohol consumption. This hypothesis was unsupported; in fact, in the within-subjects model (and in both the within- and between- subjects models in the random slopes model), conformity was positively related to alcohol consumption. Nonetheless, there was a weak, negative, and non-significant trend between negative affect and number of drinks consumed, after controlling for all other variables in the model. This negative trend had a small suppressing effect on one of the serial indirect effects (nondisplay → negative affect → conformity → drinking quantity), and is consistent with prior work that suggests people high in negative affect experience more alcohol-related problems, without necessarily drinking large quantities of alcohol (Cooper et al., 2015). The generally positive relationship between conformity motives and alcohol consumption is inconsistent with past literature suggesting conformity motives have a small negative association with alcohol consumption (Cooper et al., 2015; Cooper, 1994). However, prior literature failed to disentangle between and within-subjects effects using longitudinal, multilevel analysis (Preacher et al., 2010); this makes the present findings somewhat difficult to compare to prior work. Limitations and future directions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our study has important limitations. Our sample was primarily young, Caucasian and female, which limits generalizability. Given that women have higher psychological distress, and men have higher levels of alcohol abuse (Adlaf et al., 2004), future studies might consider moderation by gender. Regardless, results of this study may generalize to men more poorly given the large proportion of women in the present study. Moreover, results should not necessarily be generalized beyond the regions sampled – even across two regions in the same country, we found minor differences in drinking patterns. The present study also limited its focus to two dimensions of perfectionism and two drinking motivations, given the complexity of our analytic strategy and page limits of most journal outlets. Given that other dimensions of perfectionism (e.g., self-oriented, socially prescribed, and other-oriented perfectionism; Hewitt & Flett, 1991) and drinking motives (i.e., social and enhancement motives) were measured in the present study, future researchers may wish to utilize our open-access dataset to test supplementary hypotheses with these variables in a separate manuscript. Some participants had few drinking days (Supplementary Table 2); this can impede estimation of the within-subjects effects. Future research may wish to extend the daily diary period (e.g., to a period of 2 months) to increase cluster sizes. Statistical power may also be a concern, as there was not enough existing information in the literature to specify all the parameters for a Monte Carlo a-priori power simulation. However, the present study could inform a power analysis for a subsequent replication. Thus, as is true of all research, conclusions should be considered tentative pending replication. Another important limitation is our reliance on self-report measures. Self- report measures are well-known to suffer from social desirability biases, where respondents answer untruthfully in order to present themselves in a more positive light. This could be especially problematic for individuals high in perfectionistic self- presentation (e.g., nondisplay of imperfection). Future research should combine self- report measures with informant reports to obtain a more accurate and comprehensive view of perfectionistic dimensions and how they relate to negative emotions, avoidance-based drinking motives, and alcohol outcomes. Moreover, a 24 h time lag is arbitrary, and asking participants to aggregate across a 24 h period makes it difficult to tease apart directionality; future research might utilize event-contingent measurement (e.g., right after each drink). Though we were limited by the need for short-form measures to reduce participant burden, future research might assess perfectionistic cognitions as a multidimensional construct (e.g., Stoeber et al., 2014). Furthermore, it is important to acknowledge that our within-subjects model assesses only co-occurring changes within a given day. Such a model does not establish temporal precedence, which is a necessary prerequisite for inferring causality. Moreover, reporting on consumption once per day may miss more nuanced processes that occur on a shorter timeframe. Thus, these relationships should not be interpreted as causal. Lastly, we only studied one type of risky behavior (i.e., alcohol-related outcomes), yet theory and research suggest risky behaviors (e.g., unsafe sex, illicit substance use) tend to cluster together (Jessor, 1987; Mobley & Chun, 2013; Vazsonyi et al., 2008). Thus, exploring whether individuals high in perfectionistic cognitions and nondisplay of imperfection are at an increased risk for other unsafe behaviors, such as risky sexual behavior or substance abuse, would be beneficial to examine. It may be that perfectionists’ tendency to engage in self-presentation strategies and rumination make other types of risky behaviors less appealing. In contrast, those high in negative forms of perfectionism may use marijuana or other drugs to cope with negative emotionality. Psychotropic drugs used to regulate mood (e.g., SSRIs) could also lead to extreme intoxication due to alcohol sensitivity, offering an explanation as to why these individuals experience problems while drinking without consuming large amounts of alcohol. All of these possible explanations represent interesting avenues for future research.","Overall, our hypotheses about perfectionistic cognitions and drinking quantity were unsupported, while our hypotheses about nondisplay of imperfection and alcohol problems were supported. Perfectionistic cognitions do not appear to predispose individuals to negative emotions and avoidance-based drinking problems. In contrast, nondisplay of imperfection conferred risk for alcohol-related problems via negative affect and negatively-reinforcing drinking motivations. Hewitt et al. (2003) propose that individuals who engage in perfectionistic self-presentation do so in an attempt to gain approval and a sense of belonging from others. Unfortunately, these attempts are often in vain as others tend to perceive their inauthenticity, leading to further social alienation. Understandably, these factors may contribute to alcohol misuse and alcohol-related problems. Nondisplay of imperfection and problematic drinking among young adults is a major clinical and societal issue. Broadening our understanding of perfectionistic dimensions and the mechanisms through which these trait-states function may aid us in developing personality-tailored interventions, which will hopefully lead to fewer alcohol- related problems for people high in nondisplay of imperfection in the future."],["Basic personality traits are believed to be expressed in, and predictable from, smart phone data. We investigate the extent of this predictability using data (n = 636) from the Copenhagen Network Study, which to our knowledge is the most extensive study concerning smartphone usage and personality traits. Based on phone usage patterns, earlier studies have reported surprisingly high predictability of all Big Five personality traits. We predict personality trait tertiles (low, medum, high) from a set of behavioral variables extracted from the data, and find that only extraversion can be predicted significantly better (35.6%) than by a null model. Finally, we show that the higher predictabilities in the literature are likely due to overfitting on small datasets. --------------------------------------------------------------------------------","Over the last decades, new data collection methods have provided new opportunities for research on human behavior. Online social networks or personal mobile devices do not only provide real-time data for studies on human activity and interaction, but can also serve as an external validation of e.g. more classical questionnaire- or interview-based studies. For example, the predictability of basic personality traits from smart-phone usage is currently an active area of research Crandall et al. (2010), de Oliveira, Karatzoglou, Concejero Cerezo, Armenta Lopez de Vicuña, and Oliver (2011), LiKamWa, Liu, Lane, and Zhong (2011), Verkasalo, López-Nicolás, Molina-Castillo, and Bouwman (2010), Chittaranjan, Blom, and Gatica-Perez (2011a), Chittaranjan, Blom, and Gatica-Perez (2011b), Williams, Whitaker, and Allen (2012), de Montjoye, Quoidbach, Robic, and Pentland (2013), Sekara and Lehmann (2014), Mollgaard et al. (2016b). Based on data from the Copenhagen Network Study (CNS), Stopczynski et al. (2014), we use smartphone data to quanitfy the predictability of the Big Five personality traits Digman (1990), openness (O), conscientiousness (C), extraversion (E), agreeableness (A) and neuroticism (N), commonly called the five factor model and abbreviated as OCEAN. The CNS data is to the best of our knowledge the largest and most detailed study of its kind. Specifically, we use the Big Five Inventory John, Naumann, and Soto (2008), which consists of 44 items. For each item, participants in the CNS study have expressed, on a discrete scale from 1 to 5, how much they agree with a given statement. The personality traits are then computed from a pre-determined linear combination of the 44 answers. Previous research has suggested that smartphone data can be used to predict the Big Five with surprisingly high accuracy de Montjoye et al. (2013). In contrast, we show using a broad range of features extracted from the CNS data that only extraversion can be predicted with some certainty. In the Methods section below and in the appendices, we provide a description of the features (predictor variables) we extract from the smartphone data and further consider their cross-correlations. In the Results section, we use a support vector machine model for the prediction and quantify its relative improvement over a null model where personality scores are randomly assigned. Finally, we briefly compare the scoring system behind the Big Five Inventory against alternative dimensionality reduction techniques in terms of predictability.","We use questionnaire-based data on the personality traits together with phone based-data from 730 freshman students starting in the year 2013 at the Technical University of Denmark. The phone-based data has been collected over a period of 24 months by custom software installed on smartphones given to the participants of the study, Stopczynski et al. (2014). The data consists of telecommunication logs (phone calls, text messages), online social networks (Facebook connections and interactions), and networks based on physical proximity. The physical proximity is measured through the Bluetooth signal strength, and can be used to monitor face-to-face contacts Sekara and Lehmann (2014). From the GPS data, we obtain information on the geo-spatial mobility Mollgaard, Lehmann, and Mathiesen (2016a). Out of the 730 participants, we only include data from individuals, which, we believe, have used the phone as a primary device. This implies discarding data from users that have written less than 10 text messages, made 5 phone calls or have 100 GPS data points, as well as users with no Facebook friends. These criteria were chosen as a simple heuristic for removing participants who very quickly stopped using the phone, as the subjects remaining after this removal had vastly larger amounts of data. These requirements reduce the number of participants in our study to 636. For comparison purposes, we consider a list features similar to those in de Montjoye et al. (2013). Furthermore, we repeat our analysis on the part of the Friends and Family (FF) dataset, Aharony, Pan, Ip, Khayal, and Pentland (2011), which is publicly available.1 The FF dataset consists of data from 52 participants, 38 of whom have sufficient call and location data for our analysis according to the selection criteria described above. We finally compare our analysis on both datasets with the results in de Montjoye et al. (2013). Table 1 presents a list of all the features we consider. The feature extraction process is described in detail in the following. A number of quantities based on location data are also computed. We extract the median and standard deviation of the users’ daily distance travelled, their daily radius of gyration (here simplified to be the radius of the smallest circle enclosing all coordinates visited by the user on each day) and the entropy of the time spent in various locations by the user. We identify the locations visited by clustering the GPS points sampled when a user is not moving. A user is defined to not move, if the user’s mean speed does not exceed 0.5 m/s in a period between two consecutive GPS points. As the uncertainty on civilian GPS locations can be up to 100 m Zandbergen and Barbeau (2011), a user moving at a speed of 0.5 m/s would need at least 400 s to move a distance larger than two times the uncertainty. For that reason, we consider only GPS points taken even further apart, i.e. 500 s apart. The GPS data points are filtered according to the following procedure. For each user, we include the first recorded GPS data point, we then exclude data points in the subsequent time window of 500 s and then again include the first data point sampled outside this window. From this new data point we repeat the procedure of excluding points in a subsequent window of 500s and so forth. We identify clusters (locations) in the GPS points by use of the DBSCAN algorithm Ester, Kriegel, Sander, and Xu (1996) and we compute the entropy of visits to those clusters by again applying Eq. (1). Finally, we estimate the fraction of time a user spends at home, where home is assumed to be the place where a user spend most of their weeknights. We finally extract a range of features concerning a user’s social contacts. This includes their number of Facebook friends and the fraction of the time users spend in the proximity of other participants in the study. This is estimated from repeated automatic scans by the Bluetooth ports. The entropy of the proximity is also calculated similarly to Eq. (1), as well as the time series parameters as described in Eq. (2). Classification ~~~~~~~~~~~~~~ We divide the scores on each of the five personality traits into tertiles, i.e. we assign a label of 0, 1 or 2 specifying whether they score low, medium, or high on that trait, corresponding to them lying in the bottom, middle, or upper third, respectively, of all the user scores for that trait. We do this for two reasons - first, this has been done in existing research Chittaranjan et al. (2011a), de Montjoye et al. (2013) and hence allows comparison between our results and those in the literature. Second, although regression approaches have nice accuracy metrics like the mean squared error (MSE), which provides a number for how far from the true values the prediction of the regressor typically is, this measure is not particularly meaningful on ordinal values like personality traits, where e.g. higher extraversion scores mean a person is more extroverted, but there’s no precise interpretation for a difference in extraversion score of, say, 0.2. In the first approach, we perform a number of cross-validation runs. For each training set introduced during the cross-validation, we first choose the n features with the strongest correlations with the personality traits and perform an extensive grid search in the hyperparameter space. As a consequence, both the hyperparameter values and the features included in the classifier will vary between each cross validation run, potentially making it more difficult to interpret the results. At the same time, however, this ensures that training and test sets are completely separated, and thus that we do not observe overly optimistic results caused by overfitting.","In de Montjoye et al. (2013) a mean relative improvement of 0.42 is reported, which is significantly above what is reported in another study Chittaranjan, Blom, and Gatica-Perez (2013). We note that significant improvements over a baseline classifier for traits other than extraversion appears contingent on (a) having few data points, and (b) using correlations on the full dataset for feature selection, thus allowing the model to be fit to noise. Hence, it seems likely that earlier reports of high predictability of human personality traits from phone metrics have been greatly overestimated due to overfitting enabled by a combination of small sample sizes and a large number of variables. We note that only the extraversion trait appears to be truly predictable from phone-based data. This is in good agreement with common sense, as phones by their nature are devices for inter-human communication. Further, some of the features used in the classifier are expected to be related to extraversion such as the users’ number of Facebook friends and the number of new contacts made during the first months of the study. Based on the Big Five Inventory, the personality traits are computed by reducing the 44 answers to five scores. Any dimensionality reduction of this kind will inevitably lose information available from the full set of answers. We have therefore performed a series of alternative reduction methods on the 44 items to see if we could improve our predictions of the personality traits (see the appendices). Both supervised and unsupervised dimensionality reductions have been used. Among the unsupervised methods, we have tried principal component analysis, independent component analysis and factor analysis. We have applied the methods directly to the answers to the 44 items in order to extract five dimensional objects keeping the most relevant information about the original 44 items. In the unsupervised reduction no information about the features is used. For the supervised reduction method, we try reduce the target variables (the list of items) by finding those items that can be best predicted from the predictor variables (the features). While both the supervised and unsupervised methods improve significantly the quality of our predictions, the overall picture is the same that predominantly items related to extraversion can be predicted with some certainty.","Using data from the Copenhagen Network Study, which, to our knowledge is the largest dataset simultaneously containing information about the Big Five personality traits and extensive information about smartphone usage patterns, we have shown that the extraversion trait can be predicted significantly better than a null model based on random classification. In contrast, the other personality traits are poorly predicted by our data. Our findings contrast previous studies, which report significant predictabilities across all traits. Given that we have carried out the analysis on datasets of two sizes using two feature selection procedures, and since we obtained high predictabilities only when (a) using full-dataset correlations for variable selection and (b) analyzing a small dataset, the combination of the two appears a likely explanation for the results previously reported in the literature. Regarding the generalizability of our findings, we note that all participants in the study were students at the Technical University of Denmark, and that findings are not necessarily generalizable to the population in general.","Data are part of larger study “Social Fabric” involving researchers at the Technical University of Denmark and University of Copenhagen. Due to privacy consideration regarding subjects in our dataset, including European Union regulations and Danish Data Protection Agency rules, we cannot make all data used here publicly available. The data contains detailed information on mobility and daily habits at a high spatio-temporal resolution. We understand and appreciate the need for transparency in research and are ready to make the data available to researchers who meet the criteria for access to confidential data, sign a confidentiality agreement, and agree to work under our supervision in Copenhagen. The “Social Fabric” study was reviewed and approved by the appropriate Danish authority, the Danish Data Protection Agency (Reference number: 2012-41-0664). The Data Protection Agency guarantees that the project abides by Danish law and also considers potential ethical implications. All subjects in the study gave written informed consent.","The authors declare that they have no competing interests."],["Generalised models of positive change following adversity do not fully account for differences in adjustment among populations who experience posttraumatic growth (PTG). The contributions of event intentionality, frequency of the adversity types, age at serious event, spirituality/religiousness, active coping, PTSD symptoms and social support were explored as predictors of PTG across three samples of university students (N = 101; Study 1), survivors of violent crime recruited from support services (N = 71; Study 2) and those working with survivors of adversity (N = 96; Study 3). The results of Study 1 revealed that age at serious event, active coping, PTSD symptoms and social support positively predicted PTG. Within Study 2, spirituality/religiousness, active coping and social support were the significant positive predictors of PTG. Finally in Study 3, spirituality/religiousness, active coping and social support were the significant positive predictors of PTG. Across all studies, event intentionality and frequency of adversity types did not determine PTG. These results indicate that while participants within each of the populations have the ability to experience PTG, different factors predicted whether PTG was observed. The findings offer greater insight into the multifarious nature of adjustment following adversity. --------------------------------------------------------------------------------","The experience of adverse events often leads to stressful or traumatic reactions and changes in psychological functioning. Nevertheless, some people exposed to adversity report positive changes as a result of their experiences. Such changes are characterised by a greater appreciation for life, the perception of new opportunities, increased feelings of personal strength, improved relationships with people and an enhanced religiosity or spirituality (Joseph, 2012; Tedeschi & Calhoun, 2004). This positive transformation is known as posttraumatic growth (PTG; Tedeschi & Calhoun, 2004), and contrasts with earlier literature that has long-considered only the negative consequences associated with adversarial events. Transformational theory of PTG ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The processes that underlie the development of PTG are thought to emerge in the same way as do negative effects. These are largely represented in existing models of growth, most notably the transformational model (Tedeschi & Calhoun, 2004). Broadly, the model proposes event-related cognitions and individual differences such as coping responses and social support are thought to play a key role in post-trauma outcomes and PTG. Adversarial events are usually experienced as traumatic if they are seismic enough to shatter world assumptions and pre-existing schemas (Tedeschi & Calhoun, 2004). A period of rumination generally follows, where attempts to reconcile world views with new trauma-related information are made to accommodate it into existing knowledge. This does not imply that PTG occurs in the absence of negative effects, as people exposed to adversity typically report co-occurring negative symptoms including those of posttraumatic stress disorder (PTSD). These negative symptoms appear to be part of the emotional struggle in which growth can occur (Lancaster, Klein, Nadia, Szabo, & Mogerman, 2015). However, as a generalised account of PTG development, the transformational model does not fully account for individual differences in adjustment following adversity. Active, religious and spiritual coping styles ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Although primarily focused on cognitive factors, the transformational model (Tedeschi & Calhoun, 2004) attends to contextual factors that may predict PTG. Cognitive appraisals can shape coping strategies that are employed to mitigate the most distressing aspects of the adverse event. Separately, two meta-analyses of 84 and 103 PTG studies respectively (Helgeson, Reynolds, & Tomich, 2006; Prati & Pietrantoni, 2009) revealed that active coping strategies and the use of religious or spiritual coping were closely associated with PTG, as people sought to find comfort and attempted to frame their experiences in a positive light. It is thought these processes are driven by an intrinsic need to move towards growthful outcomes and make sense of the experience (Joseph, Murphy, & Regel, 2012). PTSD symptoms ~~~~~~~~~~~~~ The period of processing adverse events is generally reflected by increased posttraumatic stress symptoms, marked by intrusive thoughts and flashbacks of the event (Joseph et al., 2012). Findings suggest PTSD cognitions display positive linear and curvilinear relationships with PTG (Kleim & Ehlers, 2009; Powell, Rosner, Butollo, Tedeschi, & Calhoun, 2003; Shakespeare-Finch & Lurie-Beck, 2014). Specifically, lower PTSD symptoms may signify that the person is less affected by the adversarial event and therefore less PTG is experienced. Moderate levels of PTSD symptoms suggest that the person's world has been challenged in some way, yet they are able to engage in cognitive processing necessary for growth to occur. Higher levels of PTSD symptoms are thought to overwhelm a person's coping resources and they are more likely to succumb to negative aftereffects and therefore experience minimal PTG (Joseph et al., 2012). Social support ~~~~~~~~~~~~~~ In addition to coping methods, the wider social-environmental context is implicated in the transformational model (Tedeschi & Calhoun, 2004). In particular, social support has emerged as a robust predictor of growth across numerous PTG studies (Linley & Joseph, 2004). The dynamics of interpersonal relationships provide emotional support that can mediate adjustment outcomes (Ullman & Peter-Hagene, 2014) and offer new perspectives that are crucial for PTG development (Tedeschi & Calhoun, 2004). The additional outlooks provided by others aid deliberate rumination and the development of narratives that help people to draw upon the beneficial aspects of the event (Tedeschi, 1999). It is during this process of cognitive engagement that the foundations of PTG are laid, allowing the individual to thrive. Age at time of serious event ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As well as the aforementioned psychosocial characteristics, early life adversity is also thought to be a significant determinant of outcomes in adulthood. For many people, their first experience of adversity occurs in childhood and can often shape conceptions of self- identity (Sutherland & Bryant, 2005). Developmental adversity places the individual at greater vulnerability to more negative effects such as PTSD in later life (Hagenaars, Fisch, & van Minnen, 2011). However, people exposed to adversity in childhood could also change as a result of their experiences in a positive way. At present, there are mixed findings with regard to age at event experience and the degree of growth reported. Studies have reported negative, positive or no relationships between age and PTG (Meyerson, Grant, Carter, & Kilmer, 2011). While the transformational model (Tedeschi & Calhoun, 2004) does not explain the role of temporal factors, such discrepancies may be due to differences in PTG measurement, or the wide demographic range of populations sampled (Linley & Joseph, 2004; Powell et al., 2003; Shakespeare-Finch & Lurie-Beck, 2014). For example, perceived event severity may vary between younger people who experience a novel adverse event compared to older populations with more life experience (Sutherland & Bryant, 2005). Taken together, age at event experience and its influence on PTG development is not fully understood. Event intentionality and frequency of exposure to adversity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In addition to investigating the factors relating to growth, the type of event people have been exposed to may influence the amount of PTG reported. Research has observed PTG in approximately 30 to 100% of survivors of breast cancer, transport accidents, natural disasters and those traumatised through bereavement (Linley & Joseph, 2004). The majority of event types explored to date are largely representative of experiences that are not intentionally perpetrated against people. The trauma literature distinguishes between such acts of nature and intentional events where harm is deliberately inflicted upon another person (Santiago et al., 2013). Intentional events have been associated with more adverse outcomes and magnified PTSD symptoms compared to non-intentional acts of nature (Santiago et al., 2013). While it has been suggested that intentional events may have profound effects on PTG development in populations who experience them such as sexual abuse survivors (Tedeschi, 1999), the direction of the effect is unclear as this has not received sufficient empirical support to date. Alongside the type of adversity, the frequency of exposure is thought to determine subsequent psychological adjustment. Specifically, the experience of multiple adversity is thought to intensify PTSD reactions compared to isolated events (Green et al., 2000). However, no studies have explored the influence of frequency of exposure to adversity on PTG development. As objective characteristics of the event, both the type and frequency of adversarial exposure are not represented in the transformational model which places greater emphasis on subjective interpretations of the event (Tedeschi & Calhoun, 2004). At present, there are no empirical investigations of this assumption and so the way in which intentional and multiple events are related to PTG, if at all, is not clear. Therefore, research that explores PTG in samples of people exposed to a diverse range of multiple intentional and non-intentional adverse experiences is warranted. Students The literature has considered growth from adversity in a wide range of samples. This has included survivors of cancer, transport accidents and military combat (Barakat, Alderfer, & Kazak, 2006; Linley & Joseph, 2004). Such research tends to use homogenous samples of people exposed to a specific type of adversity. As a consequence, this confines the study of PTG to narrow samples of survivors and excludes the potential range of intentional and non-intentional adversity that people may experience in their lifetime. One sample where a range of adversarial events could be considered is university students. Students form samples in many existing PTG studies (e.g. DeRoma et al., 2003; O'Connor, Cobb, & O'Connor, 2003; Prati & Pietrantoni, 2009) and there are benefits of doing so. They are a generally accessible population who have been potentially exposed to a range of adverse events rather than one specific stressor. This enables the exploration of both intentional and non-intentional adversity types. Furthermore, it could be argued that university students represent high functioning individuals who, despite previous adversity, are able to lead lives relatively free of the impairments that adversity can generate (Taku et al., 2007). For example, they are able to study academically at a high level. These may reflect a proportion of the trauma population who exhibit resiliency traits prior to the event, or even growth after the event that buffers against pathology such as PTSD (Bensimon, 2012). As such, university students are a high functioning population who provide a representative sample of people potentially exposed to a range of intentional and non-intentional adverse events in order to explore predictors of PTG. Survivors of violent crime In contrast to student samples, survivors of violent crime may represent a population who experience more frequent adversity of a deliberate nature. Some survivors of serious criminal acts are subject to a disproportionate number of intentional events in comparison to the non-traumatised population (Kunst, Winkel, & Bogaerts, 2010; Tedeschi, 1999). In particular, survivors of intimate partner violence and sexual assault are likely to experience sequential acts of victimisation in the context of interpersonal relationships (Felson, Ackerman, & Gallagher, 2005). Collectively, exposure to intentional and repeat events place people at great vulnerability to substance dependency, depression and elevated PTSD symptoms (Ruback, Clark, & Warner, 2014; Scarpa, Haden, & Hurley, 2006), over and above the influence of natural occurrences (Santiago et al., 2013). These additional difficulties impair every day occupational and social functioning to a great degree in violent crime survivors, where chronic adversity can negatively influence perceptions of available support and thus magnify distress (Hanson, Sawyer, Begle, & Hubel, 2010). While literature has increasingly explored the impact of multiple and intentional types of adversity on the maintenance of PTSD symptoms (e.g. Graham-Kevan et al., 2015), less is known about their role in promoting growth. Furthermore, there are no PTG frameworks accounting for multiple exposures. Research has explored PTG among samples with physical assault as the index adverse event (e.g. Kleim & Ehlers, 2009); however, such studies have not taken into account the diverse range of intentional and non-intentional adversarial experiences that survivors of crime often face. This could lead to differences in the processing of adverse events and the factors that contribute towards crime survivor's experiences of PTG. Trauma workers The study of PTG also has particular relevance to those who work with or support people who are exposed to adversity in their occupation (hereafter termed ‘trauma workers’). Trauma workers represent another proportion of the population who routinely are exposed to an elevated degree of adverse events (Cohen & Collens, 2013). However, unlike survivors of violent crime, potential traumatisation and PTG can occur indirectly through interactions with people who are also exposed to serious adversity (Cohen & Collens, 2013). While there are currently no explanatory models of vicarious or secondary PTG, recent studies have increasingly drawn attention to PTG emerging in this manner (Brockhouse, Msetfi, Cohen, & Joseph, 2011; Samios, Rodzik, & Abel, 2012). However, there is a paucity of research in relation to trauma worker's own direct experiences of adversity. This is surprising, as altruistic tendencies observed in trauma workers and similar professions are thought to stem from the experience of adversity in their own personal lives (Staub & Vollhardt, 2008). According to the transformational model of PTG (Tedeschi & Calhoun, 2004), the emotional salience and proximity to personal adverse events can trigger cognitive processing necessary for PTG, more so than adversity experienced in occupational contexts. Both personal and work- related adversity has been found to predict PTG in firefighters (Armstrong, Shakespeare-Finch, & Shochet, 2014). Despite repeat exposure to a range of adverse events at work and their own personal history of adversity, trauma workers are relatively high functioning by sustaining employment within emotionally demanding professions (Cohen & Collens, 2013). This may reflect trait resiliency or the buffering nature of PTG which allows trauma workers to reinterpret multiple adversity in a less threatening way (Bensimon, 2012; Samios et al., 2012). Therefore, the current research will focus on personal adversity and predictors of PTG in a high functioning sample of trauma workers with repeat exposure to indirect adversity. The current research ~~~~~~~~~~~~~~~~~~~~ Based on the existing literature, this research explored the contributions of event intentionality, frequency of adversity types, age at which the most serious event occurred, spirituality/religiousness, active coping, PTSD symptomology and social support as potential predictors of PTG. These predictors would be explored in three samples who represent survivors exposed to different types or frequencies of adversity. Study 1 explored the role of event intentionality, frequency of adversity types, age at serious event, spirituality/religiousness, active coping, PTSD symptomology and social support variables in a student sample. The student sample represents individuals with experience of a broad range of adversity types yet are able to study academically at a high level. Study 2 applied the same predictors to a sample of survivors of violent criminal victimisation who experience frequent intentional adversity that may negatively impact upon psychological functioning. Finally, Study 3 extended the findings of studies 1 and 2 by exploring the predictors of PTG in a sample of trauma workers who experience not only their own personal adversity, but are frequently exposed to adverse events indirectly yet remain able to continue in demanding roles. Taken together, this approach would allow the identification of individual differences and similarities in the development of PTG across a diverse range of samples that would not otherwise be revealed in single study designs.","In Study 1, it was expected that spirituality/religiousness, active coping, PTSD symptomology and social support would positively predict growth based on existing PTG literature. Given relationships between objective characteristics and posttraumatic stress symptoms, it was also expected that event intentionality, frequency of adversity types and the age at which the serious event occurred would be related to PTG.","One hundred and one students with prior exposure to adversity took part in the study. Table 1 presents demographic information for the sample. Participants were recruited via university posters and online postings on message boards and student forums. Questionnaires were accessed through a link provided on the websites where the potential participants could access information about the study and their rights as participants. Upon providing informed consent, participants completed the questionnaires, were debriefed and provided details of support services. They had the option to enter a prize draw for a £50 shopping voucher as compensation for their time. The study was approved by the university ethics committee and adhered to British Psychological Society ethical guidelines.","Demographic information including age, gender, sexuality, ethnicity and religion was collected. Traumatic Experiences Questionnaire (TEQ; Foa, Cashman, Jaycox, & Perry, 1997). The TEQ is a self-report measure of adverse experiences and includes 12 types of event such as exposure to accidents, natural disasters, sexual assaults and serious illness. In this study, the scale was adapted from Foa et al.'s (1997) original version to include two further items of parental neglect and occupational secondary traumas which account for other potentially traumatic events (Cohen & Collens, 2013; Hagenaars et al., 2011). The participant records the frequency of each event to the best of their memory. Intentional events were considered to involve directly perpetrated physical or sexual violence. Two additional questions invite the participant to record the item of the event they perceived to be most severe and the age this first occurred. The measure has been validated in samples of individuals exposed to adversarial events and demonstrates reasonable internal consistency (Foa et al., 1997), which was replicated in this study (α = .68). Beliefs and Values Scale (BVS; King et al., 2006). The BVS is a measure of religious and spiritual beliefs, where respondents are asked to indicate their agreement to 20 statements using a scale from 0 (strongly disagree) to 4 (strongly agree). It has been validated as a reliable measure in large and diverse samples (King et al., 2006). Example items include, ‘Although I cannot always understand, I believe everything happens for a reason’ and ‘I believe in a personal God’. An overall score is produced, with higher scores indicative of greater religiosity and spirituality. In the current study, Cronbach's α = .96. Brief COPE (Carver, 1997). The Brief COPE is a 28-item questionnaire assessing 14 coping styles on a four point scale from 0 (I haven't been doing this at all) to 3 (I′ve been doing this a lot). Participants rate which coping styles they employ; example items include, ‘I've been taking action to make the situation better’ (active coping), with higher scores representing greater use of the specific coping style. The Brief COPE has demonstrated good internal reliability and can be used as a short measure for coping in specific situations of interest (Carver, 1997). As with previous studies (e.g. Thornton & Perez, 2006), the active coping scale was of particular interest due to links between such coping styles and new perspectives in PTG development (Tedeschi & Calhoun, 2004). These two subscales were used in subsequent analysis. Reliability scores for the active coping scale was .78. PTSD-8 (Hansen et al., 2010). The PTSD-8 is a measure of posttraumatic stress symptoms, where respondents rate their agreement with eight statements on a four point scale from ‘not at all’ to ‘most of the time’. Participants were asked to identify their most serious event and indicate the symptoms they have experienced in the past 2 weeks. There are three subscales of avoidance, intrusion and hyperarousal which are represented with items such as ‘Recurrent thoughts or memories of the event’ and ‘Avoiding activities that remind you of the event’. Participants with a score of three or above on each subscale may display PTSD traits. It has been validated in samples of rape survivors, whiplash patients and survivors of disasters (Hansen et al., 2010). In the study, the overall scale was used with Cronbach's α = .88. Two-Way Social Support Scale (2-Way SSS; Shakespeare-Finch & Obst, 2011). The 2-Way SSS is a 21-item measure of giving and receiving emotional and instrumental social support on a scale from 0 (not at all) to 5 (always). There are four subscales of receiving emotional support, giving emotional support, receiving instrumental support and giving instrumental support. Example items include, ‘There is someone in my life I can get emotional support from’ and ‘There is someone who will help me fulfil my responsibilities when I am unable’. Higher scores endorse greater support. The scale has been validated in two community samples (Shakespeare-Finch & Obst, 2011) and the overall score for the measure was used in this study, demonstrating excellent reliability (α = .93). Posttraumatic Growth Inventory — Short Form (PTGI-SF; Cann et al., 2010). The PTGI-SF is a measure of growth, on a six point scale from 0 (no change as a result of crises) to 5 (very great change). Participants are asked to rate what extent they have changed since their stressful life event with 10 items such as, ‘I changed my priorities about what is important in life’ and ‘I discovered that I′m stronger than I thought I was’. It has been validated for use in samples including survivors of domestic abuse, bereaved persons and those with complex health needs, demonstrating similar reliability to that of the original 21-item version of the PTGI, whilst having the advantage of brevity (Cann et al., 2010). A total score is obtained, with higher scores reflecting greater perceived change. The PTGI-SF demonstrated high internal consistency in the current study (α = .89).","The prevalence of exposure to adverse events for participants is presented in Table 1. Of the sample, 83.2% experienced more than one adverse event type, with 68.3% experiencing two to five event types and 15% experiencing six to ten separate event types. In addition, 30.7% reported bereavement as the most serious event experienced among the range of adversity types. Means and standard deviations for the psychosocial measures are presented in Table 2. Pearson correlations revealed that age at serious event (r = .37, p < .001), spirituality/religiousness (r = .40, p < .001), active coping (r = .46, p < .001), PTSD symptomology (r = .28, p = .005) and social support (r = .35, p < .001) were all positively associated with reported PTG. Event intentionality and frequency of event types were not related to PTG. Multiple regression analysis was conducted to assess the contributions of the seven predictors towards PTG in the student sample. Using the simultaneous method, a significant model emerged, F (7, 93) = 10.27, p < .001; adjusted R2 = .39 in which age at serious event, spirituality/religiousness, active coping, PTSD symptoms and social support emerged as significant predictors of PTG. Event intentionality and frequency of event types did not predict PTG. The results are presented in Table 3. Collinearity diagnostics revealed that no two variables were highly correlated (Tolerance for all variables > .66; VIF for all variables < 1.52).","Study 1 showed expected relationships between PTG and a number of psychosocial variables among students. Specifically, spirituality/religiousness, active coping, PTSD symptomology and social support were positively related to PTG development. In partial support of the hypothesis, Study 1 indicated that these variables as well as age at the time of the serious event were also positive predictors of PTG. In particular, active coping methods, spirituality/religiousness and social support demonstrated stronger relationships with growth compared to the other variables. Contrary to the hypothesis, both event intentionality and frequency of event types were neither associated nor predictive of growth. Taken together, the results suggest that psychosocial factors are more closely related to adjustment of adversity compared to objective characteristics of the serious event (Tedeschi & Calhoun, 2004). Study 1 explored predictors of PTG in a high functioning sample exposed to a broad range of adversity. However, this does not fully account for the experiences of people exposed to particularly frequent and intentional adverse events above the normative population. One example of a population with more extreme and intentional adversity is survivors of violent crime, whose experiences can lead to poor social and occupational functioning (Hanson et al., 2010; Ruback et al., 2014). Therefore, Study 2 assessed the efficacy of the predictor variables used in Study 1 in relation to a sample comprised of survivors of violent crime. The purpose was to ascertain the extent to which the degree and type of adversity experienced in this sample mediated the influence of the psychosocial predictors of PTG.","In Study 2, it was hypothesised that the age the most serious event occurred, spirituality/religiousness, active coping, PTSD symptomology and social support would contribute towards PTG, based on the findings from Study 1. Given that survivors of violent crime may experience significant adversity that may serve as a catalyst for growth, it was expected that event intentionality and frequency of event types would be associated with PTG.","Seventy-one survivors of crime volunteered to take part in this study. Table 1 presents demographic information for the sample. Participants were recruited using messages advertised on websites provided by three victim services, which support female and male survivors of domestic violence, child sexual abuse and sexual assault respectively. Two participants were also sampled from a concurrent study using survivors of violent crime (Graham-Kevan et al., 2015). Procedures used to collect data were the same as outlined in Study 1.","Participants self-reported demographic information including age, gender, sexuality and ethnicity and religious beliefs. All measures in this study were the same as those described in Study 1. Trauma history was explored using the TEQ (Foa et al., 1997). Internal reliability for this scale in this sample as measured by Cronbach's alpha was α = .73. The degree of spirituality/religiousness was assessed using the BVS measure (King et al., 2006) and in this study, Cronbach's α = .96. The Brief COPE (Carver, 1997) assessed active coping styles with a Cronbach's alpha of .61. PTSD symptomology was assessed using the PTSD-8 (Hansen et al., 2010) and the reliability of the overall scale in this study was excellent (α = .91). The 2-Way SSS measure (Shakespeare-Finch & Obst, 2011) was employed to establish perceptions of social support. As with Study 1, the overall score for the measure was used in this study and the internal consistency of the items was high (α = .95). Finally, the brief version of the PTGI measure (Cann et al., 2010) was employed to explore reported PTG. The PTGI-SF demonstrated excellent reliability in this study (α = .91).","The prevalence of exposure to adverse events for participants is presented in Table 1. Of the participants, 94.4% experienced more than one adverse event type, with 56.4% experiencing two to five event types and 32.4% experiencing six to ten event types. Notably, around three-quarters of participants experienced sexual abuse (73.2%) and serious physical attacks or threats (76.1%). Nearly a quarter (23.9%) of the sample indicated that sexual abuse was the most serious adverse event they had experienced. Means and standard deviations for the psychological measures are presented in Table 2. Pearson correlations revealed that age at serious event (r = .25, p = .033), spirituality/religiousness (r = .50, p < .001), active coping (r = .37, p = .002) and social support (r = .40, p = .001) were all positively associated with reported PTG. PTSD symptomology, event intentionality and frequency of event types were not related to PTG. Multiple regression analysis assessed the seven variables as potential predictors towards PTG in the sample. Using the simultaneous method, a significant model emerged, F (7, 63) = 7.12, p < .001; adjusted R2 = .38. Table 3 presents the results of the regression in which spirituality/religiousness, active coping and social support emerged as significant predictors. There was no evidence of colinearity among the variables (Tolerance for all variables > .68; VIF for all variables < 1.46).","In line with Study 1, the results suggest that spirituality/religiousness, active coping and social support were all positive predictors of growth among survivors of violent crime. As in Study 1, objective characteristics of event intentionality and frequency of event types were unrelated to PTG development in the sample. Contrary to the findings of Study 1 and the Study 2 hypothesis, the age at which the serious event occurred and PTSD symptoms did not predict PTG. This highlights that although the students and crime survivors share some similar predictors of PTG and are able to report positive changes despite previous adversity, the populations are not identical. While Study 2 considered predictors of PTG among people exposed to frequent and often intentional adversity, little is known about cumulative adversity in samples that appear to function at a higher level. Trauma workers are not only exposed to adverse events through engagement with clients in occupational settings (Brockhouse et al., 2011; Cohen & Collens, 2013), but themselves experience adversity in their personal lives. There are few studies of predictors of PTG in trauma workers in relation to their own personal adversity (Armstrong et al., 2014). It is possible that exposure to repeat indirect adversity may buffer against negative symptoms from their own adversity (Samios et al., 2012). Therefore, the purpose of Study 3 is to investigate predictors of PTG in a sample of trauma workers in the aftermath of personal adverse events.","As with the student sample in Study 1, it was predicted that the age at which the serious event occurred, spirituality/religiousness, active coping and social support would positively predict PTG in trauma workers. However, as their job role may encourage the development of coping techniques such as buffering from negative symptoms, it may be that event intentionality, frequency of event types and PTSD symptoms would be unrelated to PTG.","Ninety-six trauma workers volunteered to take part in this study. Participants were recruited using professional forums and snowball methods. The final sample consisted of 21 counsellors, 11 mental health nurses, 29 psychotherapists, 17 psychologists, three psychiatrists and 15 social workers or support workers. Table 1 presents demographic information for the sample. Procedures used to collect data were the same as outlined in Study 1.","As with the previous two studies, participants completed demographic information including age, gender, sexuality and ethnicity and religious beliefs. All measures in this study were the same as those described in Study 1. The TEQ (Foa et al., 1997) exploring trauma history demonstrated a Cronbach's alpha of .69. The BVS (King et al., 2006) measured perceptions of spirituality/religiousness and was found to have excellent internal consistency (α = .96). As in the previous studies, the Brief COPE (Carver, 1997) was employed to assess active coping (α = .77) which demonstrated acceptable reliability. The PTSD-8 (Hansen et al., 2010) captured PTSD symptomology and the reliability of the overall scale was excellent (α = .90). Social support was measured using the 2-Way SSS (Shakespeare-Finch & Obst, 2011) and the internal consistency of the items was excellent (α = .94). PTG was again explored using the PTGI-SF (Cann et al., 2010) and reliability for the scale in this study was high (α = .92).","The prevalence of exposure to adverse events is presented in Table 1. 86.5% of the sample experienced more than one adverse event type, with 64.7% experiencing two to five event types and 20.8% experiencing six to ten event types. Like Study 1, most of the participants (26.0%) rated bereavement of a family member or close friend as their most serious adverse experience. Means and standard deviations for the psychological measures are presented in Table 2. Pearson's correlations showed that spirituality/religiousness (r = .37, p < .001), active coping (r = .35, p = .001), PTSD symptomology (r = .27, p = .009) and social support (r = .32, p = .002) were positively associated with overall PTG. The age the serious event occurred, event intentionality and frequency of event types were unrelated to PTG. A multiple regression analysis using the simultaneous method assessed the seven variables as potential predictors towards PTG among participants and produced a significant model, F (7, 88) = 7.11, p < .001; adjusted R2 = .31. Spirituality/religiousness, active coping and social support emerged as the three significant predictors of PTG and are presented in Table 3. As with Study 1 and Study 2, colinearity was not identified in this sample (Tolerance for all variables > .70; VIF for all variables < 1.44).","This study was the first to identify predictors of PTG among a sample of trauma workers, building on similar work in firefighters (Armstrong et al., 2014). As with studies 1 and 2, the findings provided support for the hypothesis that spirituality/religiousness, active coping and social support were necessary for PTG to occur in trauma workers. Neither event intentionality nor frequency of event types was linked to growth, consistent with the hypothesis and the prior two studies on students and survivors of violent crime. As with the crime survivors in Study 2, age at serious event did not predict PTG, which was contrary to the students in Study 1 and the hypothesis. Collectively, the results support the robustness of spirituality/religiousness, active coping and social support factors as predictors of growth. Meanwhile, objective characteristics and PTSD symptoms appear to vary among different populations of people exposed to adversity and do not influence PTG development.","In the research presented, the predictive ability of event intentionality, frequency of event types, age at serious event, spirituality/religiosity, active coping, PTSD symptoms and social support on levels of PTG was explored. These factors were assessed in three populations of survivors exposed to different types or frequencies of adversity where factors salient for PTG development may vary. Collectively, the predictive factors explained a significant proportion (between 30 and 45%) of the variance in PTG scores across the three studies. An encouraging finding was that regardless of life trajectory, participants in all three studies reported similar levels of PTG. This is perhaps not surprising given that PTG has been observed across a broad range of adversarial exposures (Linley & Joseph, 2004). Notwithstanding the apparent universality of PTG, this series of studies for the first time revealed some notable differences and similarities among the predictive factors that were salient for growth to occur in three populations studied. Active, religious or spiritual coping and PTG ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Across all three populations, active coping and spiritual or religious coping strategies were the most robust predictors of PTG. Earlier reviews of the literature report large effect sizes for coping methods on PTG development overall (Prati & Pietrantoni, 2009). This is perhaps not surprising given that active and religious or spiritual coping methods may reflect attempts to understand significant challenges brought about by adverse events (Tedeschi & Calhoun, 2004). Importantly, the findings indicate that regardless of life trajectory, people exposed to different types of adversity who employ active coping strategies perceived more PTG. The presence and degree of spirituality/religiousness was consistently associated with PTG in the three samples. This suggests that the use of existential beliefs can be found in the three populations investigated. Literature on the benefits of spiritual and religious coping in PTG development is widely available (e.g. Helgeson et al., 2006; O'Connor et al., 2003; Prati & Pietrantoni, 2009). In a paradoxical fashion, adversarial events not only shatter assumptions but can lead to greater engagement with existential, philosophical or moral questions that represent growth (Tedeschi & Calhoun, 2004). It therefore appears that such strategies can enhance the sense of meaning in life or bring about a new engagement with religion and spirituality for many people. However, not all participants recorded a religious affiliation and so it would be advantageous for future studies to distinguish between types of religious and spiritual beliefs and their individual contributions towards PTG. PTSD symptoms and PTG ~~~~~~~~~~~~~~~~~~~~~ Relationships emerged between PTSD and PTG in Study 1 only. The mixed findings across the three populations may be partly explained by psychosocial resources that survivors may draw upon in order to mitigate negative effects. Students with less life experience of adversity may attribute greater significance to early or novel experiences (Sutherland & Bryant, 2005). This may exacerbate symptoms as processing of the event occurs (Tedeschi & Calhoun, 2004), but not so much as to overwhelm the survivor, allowing growth from the event. The lack of relationship between PTSD and PTG among the survivors of violent crime appears contrary to assertions that growth and distress co-exist (Lancaster et al., 2015). However, this may be explained by adaptive attempts to normalise or dissociate from such experiences to minimise distress (Hagenaars et al., 2011). It is this numbness to emotional experience that may account for the lack of PTSD symptoms among the crime survivors. In addition, PTSD symptoms may be of a severity as to overwhelm the crime survivors, thus inhibiting growth (Shakespeare-Finch & Lurie-Beck, 2014). Furthermore, trauma workers are in a unique position to experience cumulative stressors through their roles (Cohen & Collens, 2013). It is possible that this exposure may buffer against PTSD symptoms and allow growth to occur (Samios et al., 2012), as reflected by lower PTSD scores for this group. Collectively, the present findings suggest that PTSD symptoms are particularly susceptible to the wider environmental and psychological contexts in which the samples function. Social support and PTG ~~~~~~~~~~~~~~~~~~~~~~ Social support also emerged as one of the most robust predictors of growth in all three studies. This reinforces earlier findings on the benefits of social support as a potential buffer against stressful events (Linley & Joseph, 2004). It has been suggested that a recognition of one's own vulnerability as a result of exposure to adversity can lead to increased sensitivity towards other people and the revision of schemas (Tedeschi & Calhoun, 2004). In addition, enriched social networks can bring out opportunities for disclosure that in turn promote positive outcomes (Ullman & Peter-Hagene, 2014). Importantly, social support appears to permeate across all types of adversity and populations, which highlights the significant role that the accessibility and maintenance of supportive networks play in post-event adjustment. Age at serious event and PTG ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ An additional aspect of these series of studies was the inclusion of age at the time the serious event happened. Findings indicated that this factor was relevant to PTG in Study 1 only. While previous reviews have reported ambiguous relationships between age and PTG development (Helgeson et al., 2006; Meyerson et al., 2011), it has been suggested that the nature of the participants sampled may account for such discrepancies (Shakespeare-Finch & Lurie-Beck, 2014). The student sample was younger compared to the violent crime survivors and trauma workers. Younger samples are more likely to be confronted with novel adverse events in childhood and adolescence, which can represent significant changes in a person's life (Sutherland & Bryant, 2005). The age at which the event occurred may be less salient for older samples that are more able to process both the positive and negative aspects of the experience (Barakat et al., 2006). Event intentionality, frequency of event types and PTG ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This was the first PTG study to determine that event intentionality and frequency of historical event types were unrelated to PTG development. The findings confirm the view that objective characteristics of the event are unrelated to growth (Joseph et al., 2012; Tedeschi & Calhoun, 2004), which had previously received no empirical support. It had also been speculated that intentional and frequent acts may in some way influence growth compared to isolated events (Tedeschi, 1999), given that chronic adversity is associated with more severe pathology (Green et al., 2000; Hagenaars et al., 2011; Santiago et al., 2013). However, this suggestion is not supported by the current findings. There is some evidence to suggest that frequent exposure to adversity can buffer against perceptions of severity by allowing people to prepare for subsequent events, which may constitute growth in itself (Armstrong et al., 2014; Kunst et al., 2010; Samios et al., 2012). In sum, the findings provide new insight into the role of event type and frequency on PTG development, where growth can occur regardless of prior exposure to adversity. This is an encouraging development for psychological interventions that could target the psychosocial factors most closely associated with growth. Implications, limitations and future research ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This research has important theoretical implications for understanding PTG among different samples of the general population who are exposed to adversity. The present studies contribute to recent literature that calls for an exploration of individual differences and commonalties in predictors of PTG (Lancaster et al., 2015), which are not duly accounted for in existing PTG models. The findings suggest that generalised models of PTG do not reflect the nuances of positive adjustment after adversity. Furthermore, the research draws attention to the role of cumulative events, which did not appear to influence PTG development although may buffer against posttraumatic stress symptoms. Currently, the transformational model of PTG (Tedeschi & Calhoun, 2004) only considers growth and processing in the aftermath of single, isolated incidents, which do not represent people who are exposed to multiple adverse events across the lifespan. Future research is encouraged to adopt a more holistic view of adversarial experiences and investigate PTG development in survivors of multiple adversity. Encouragingly, the present findings provide the first evidence that prior adversarial history does not affect the ability of survivors to report positive changes (Joseph et al., 2012; Tedeschi & Calhoun, 2004). This differs from the posttraumatic stress literature where intentional and frequent events are often associated with exacerbated negative symptoms (e.g. Green et al., 2000; Santiago et al., 2013) and suggests that the mechanisms that underpin both PTG and posttraumatic stress operate differently. Indeed the relationship between PTG and posttraumatic stress remains ambiguous (Lancaster et al., 2015; Shakespeare-Finch & Lurie- Beck, 2014) and future research would be directed to explore these relationships further. In respect of practical implications, efforts could focus on enhancing resiliency factors that predict PTG across a variety of populations. The present findings implicate active coping, spirituality/religiousness and social support factors which may promote growth. In the case of spirituality/religiousness, these findings do not imply that belief systems should be imposed or altered by clinicians; rather, these beliefs appear to be beneficial for PTG development. When targeted in psychological interventions, coping and social support factors could promote a better quality of life as a result of improved social and occupational functioning (Hanson et al., 2010) and allow people to be in a better position to consider the positive as well as negative aspects of their adverse experiences. There are strengths and limitations to research of this kind which should be noted. The study included a diverse range of adversarial experiences within each sample ranging from common normative life stressors such as bereavement and illness, to more seismic life-changing events. However, the modest sample sizes prevented the exploration of additional factors that differ among the samples. Second, data relied on self-reports which are advantageous in that participants are likely to identify with the questions and are more motivated to consider their own personalities rather than those of others (Paulhus & Vazire, 2007). However, retrospective accounts of adversarial history and associated adjustment may have been influenced by tendencies to over or under-report information.","Overall, the results across the three populations broadly support the salience of subjective interpretations in adjustment from adversity, in contrast to objective characteristics of the event, such as type and frequency. The studies provide greater understanding of the dynamic nature of psychosocial factors. In particular, coping and social support variables remain robust predictors of PTG regardless of population or prior experiences of adversity. However, the age at serious event and PTSD symptoms appear to display more nuanced relationships with PTG which are not reflected within existing PTG frameworks. Therefore, coping and social support factors could be the focus of interventions to not only reduce PTSD symptoms, but promote opportunities for PTG development."],["Background: Deaf and hard of hearing (D/HH) children and young people are known to show group-level deficits in spoken language and reading abilities relative to their hearing peers. However, there is little evidence on the longitudinal predictive relationships between language and reading in this population. Aims: To determine the extent to which differences in spoken language ability in childhood predict reading ability in D/HH adolescents. Methods: and procedures: Participants were drawn from a population-based cohort study and comprised 53 D/HH teenagers, who used spoken language, and a comparison group of 38 normally hearing teenagers. All had completed standardised measures of spoken language (expression and comprehension) and reading (accuracy and comprehension) at 6–10 and 13–19 years of age. Outcomes: and results: Forced entry stepwise regression showed that, after taking reading ability at age 8 years into account, language scores at age 8 years did not add significantly to the prediction of Reading Accuracy z-scores at age 17 years (change in R2 = 0.01, p =.459) but did make a significant contribution to the prediction of Reading Comprehension z-scores at age 17 years (change in R2 = 0.17, p <.001). Conclusions: and implications: In D/HH individuals who are spoken language users, expressive and receptive language skills in middle childhood predict reading comprehension ability in adolescence. Continued intervention to support language development beyond primary school has the potential to benefit reading comprehension and hence educational access for D/HH adolescents. --------------------------------------------------------------------------------","The difficulties D/HH children have in acquiring reading skills have been repeatedly demonstrated but longitudinal studies of D/HH children identifying aspects of underlying language skills that contribute to variation in reading skills are rare. The present study examines the stability of reading skills from middle childhood to adolescence in D/HH individuals and shows moderate stability in Reading Comprehension and a high level of stability in Reading Accuracy scores. It also shows that variation in the language ability of D/HH children in middle childhood is predictive of their Reading Comprehension in late adolescence. This predictive relationship is over and above the continuity in Reading Comprehension between childhood and late adolescence. The same was not found for Reading Accuracy; language ability in middle childhood was not a significant predictor of adolescent word reading skills. Within the D/HH group, differences in the relationship between severity of hearing loss and reading ability in late adolescence were accounted for by differences in language ability. The results of this study contribute to the case for continued targeted intervention on language skills for D/HH individuals beyond primary school. Maximising their language ability, and consequently their ability to use reading to access learning, is likely to enhance their educational attainment and subsequent life chances","Despite recent technological improvements, such as digital hearing aids and cochlear implants, and earlier diagnosis and management of babies born deaf or hard of hearing (D/HH), reading ability in D/HH children and young people continues to lag behind that of their hearing peers (Wauters, van Bon, & Tellings, 2006; Moeller, Tomblin, Yoshinaga- Itano, Connor, & Jerger, 2007; Harris & Terlektsi, 2011; Qi & Mitchell, 2011; Pimperton et al., 2016; Harris, Terlektsi, & Kyle, 2017a). Moreover, with increasing age a widening gap in reading achievement between D/HH and hearing children has been observed (Blair, Peterson, & Viehwg, 1985; Marschark & Harris, 1996; Kyle & Harris, 2010, 2011). As children get older, reading takes on an increasingly important role in enabling them to access the curriculum; they move from ‘learning to read’ to ‘reading to learn’. Thus the reading deficits shown by the D/HH population are likely to have an increasingly significant impact on their educational attainment and subsequent employment opportunities. The continuing importance of reading ability and educational attainment for the later occupational status of D/HH individuals in adulthood has been shown by Walter and Dirmyer (2013). Despite group-level deficits in their reading ability, D/HH children and young people show substantial individual variation in reading skills, with some reading at an age-appropriate level (Kyle & Harris, 2006, 2010Kyle and Harris, 2010; Harris and Terlektsi, 2011; Pimperton et al., 2016). The question of what drives this variation is an important one because identifying these drivers may raise potential avenues for intervention to support reading development in this group. In hearing children, the causal contribution of underlying language abilities to reading development has been repeatedly demonstrated through good quality longitudinal and intervention studies (see Hulme & Snowling, 2014, for review). Phonological language skills (e.g. phonological awareness, letter-sound knowledge) appear to be most important for reading accuracy (i.e. decoding written words into their phonological form) (Muter, Hulme, Snowling, & Stevenson, 2004; National Institute for Literacy, 2008; Bowyer-Crane et al., 2008; Caravolas et al., 2012) whereas non-phonological broader oral language skills (e.g. vocabulary, grammatical knowledge, morphological skills) appear to be most important for reading comprehension (i.e. understanding the meaning of what is read) (Nation, Cocksey, Taylor, & Bishop, 2010; Clarke, Snowling, Truelove, & Hulme, 2010; Fricke, Bowyer-Crane, Haley, Hulme, & Snowling, 2013). The phonological and non-phonological language skills identified as playing a causal role in reading development in hearing children are all skills which D/HH children find difficult to acquire and in which they show deficits relative to the skills of hearing peers (Moeller et al., 2007; Musselman, 2000; Wake, Hughes, Poulakis, Collins, & Rickards, 2004). Taken together this suggests that variation in reading skills of D/HH children may be driven by variation in these underlying language skills. Consistent with this, factors related to audiological experience, such as severity of hearing loss, age at identification and age at cochlear implantation, that influence language development in D/HH children, have also been found to relate to their reading development (Moeller et al., 2007; Archbold et al., 2008; McCann et al., 2009). The findings outlined above imply a similar role for language skills in causally driving the development of reading skills in D/HH children, as in hearing children. However, there is a dearth of good quality longitudinal and intervention studies to test whether, and if so which, language skills play a causal role in the reading development of D/HH children. The most extensive set of studies of the role of language in reading for D/HH individuals is on phonological coding and awareness (PCA). Mayberry, del Giudice, and Lieberman (2011) identified 25 studies that had examined PCA abilities and their relationship with reading abilities in deaf children and adults. The assessment of reading proficiency was based on some studies that measured reading comprehension and others reading accuracy. They estimated that 11% of the variance in reading proficiency was explained by PCA, similar to the 12% of variance that was thus explained in a meta-analysis of studies with hearing participants (Bus and van IJzendoorn, 1999). Around half of studies included in the Mayberry et al. (2011) meta-analysis showed evidence of an effect of PCA on reading proficiency in the deaf participants, while the other half did not. It is likely that differences between studies in terms of the format of the PCA tasks used (e.g. auditory input-oral response vs. visual input-nonverbal response), aspects of reading assessed (accuracy vs. comprehension), and populations included (e.g. children vs. adults, oral language vs. sign language users) will have contributed to the lack of consistency in results. Vocabulary and grammatical knowledge are key broader language skills that have been causally associated with reading comprehension development in hearing readers and have also been shown to relate to reading skills in deaf children (Geers, 2003; Barajas, Gonzalez-Cuenca, & Carrero, 2016). In line with this, the Mayberry et al. (2011) meta- analysis also reported that a broader measure of general language ability (signed or spoken) explained 35% of the variance in reading proficiency in those studies that included such a language measure, and was the factor with the strongest relationship to reading. They concluded that “deaf readers, like hearing readers, are more likely to become successful readers when they bring a strong language foundation to the reading process” (Mayberry et al., 2011, p.181). In commenting on the Mayberry et al. (2011) meta- analysis, Kyle, Campbell, and MacSweeney (2016) point out that the majority of the studies on language and reading in deaf children are based on correlational studies assessing the concurrent associations between these two aspects of development. They argue that longitudinal studies would provide a more robust test of the nature of these possible causal relationships. Kyle and Harris (2010, 2011) have reported on two such studies of D/HH children in middle childhood. They found that vocabulary knowledge at mean age 7 years 10 months predicted reading comprehension outcomes one year later (Kyle & Harris, 2010). The same significant longitudinal relationship was found between vocabulary knowledge at 8 years 10 months and reading comprehension at 10 years 11 months. This finding supports an association between vocabulary knowledge in the development of reading comprehension. They also showed that word reading (i.e. reading accuracy) at this age was significantly predicted, albeit to a lesser extent, by vocabulary knowledge and that neither reading comprehension nor word reading was predicted by phonological awareness. Similar findings on the predictive role of vocabulary knowledge on reading accuracy were found in the second longitudinal study in a sample of D/HH children tested at mean age 5 years 8 months, 6 years 8 months and 7 years 11 months (Kyle & Harris, 2011). An important element of both these studies was the demonstration that vocabulary was associated with later reading performance even after adjusting for earlier reading performance. In other words, it controlled for the fact that the strongest predictor of reading will be the same ability measured at an earlier time point, known as the ‘auto-regressive effect’. To summarise, there is strong evidence in hearing children that phonological language skills predict word reading accuracy and non-phonological broader language skills predict reading comprehension (Hulme & Snowling, 2014). The evidence base is weaker for D/HH children, with the majority of evidence coming from cross-sectional correlation studies, but there is some longitudinal evidence for the role of vocabulary knowledge development of reading comprehension and reading accuracy for these children too (Kyle & Harris, 2010, 2011). More longitudinal studies are needed to contribute to this evidence base and to clarify the nature of the relationships between language skills and reading proficiency in D/HH children and young people. This has not previously been addressed in older D/HH children and adolescents, in whom relationships identified between reading and language in the early stages of literacy acquisition may no longer apply. This paper reports on an analysis of relationships between language and reading measured on the first occasion at approximately 8 years and on the second occasion at approximately 17 years of age in a population-based sample of D/HH children with bilateral moderate-profound permanent childhood hearing loss (PCHL) who used spoken English as their primary form of communication. This sample was drawn from a prospective cohort study that addressed the impact of Universal Newborn Hearing Screening (UNHS) on early confirmation of hearing loss (Kennedy et al., 1998; Kennedy, McCann, Campbell, Kimm & Thornton, 2005) and subsequent language and reading outcomes (Kennedy et al., 2006; McCann et al., 2009; Pimperton et al., 2016; Pimperton et al., 2017). The main question addressed in this paper is whether expressive language, receptive vocabulary and grammatical skills in middle childhood predict reading accuracy and reading comprehension abilities in adolescence, after adjusting for the effects of auto-regressive continuities in reading. The longitudinal design of the study also makes it uniquely well-placed to provide novel evidence on the stability of reading accuracy and reading comprehension scores between middle childhood and adolescence for D/HH individuals.","The D/HH and hearing participants were all drawn from a 1992-97 birth cohort of 157,000 children born in eight districts of southern England (see Kennedy et al., 2006). Of 168 children in the birth cohort with PCHL, 160 were contactable and 120 gave consent to be included in the study. These 120 children in the D/HH group had been diagnosed with bilateral PCHL > 40 dB hearing level (dB HL) in the better ear which was not known to be post-natally acquired. Severity of hearing loss was categorised as moderate (40–69 dB HL), severe (70–94 dB HL) or profound (≥95 dB HL) according to four-frequency averaging of the better ear pure-tone thresholds at 0.5, 1, 2 and 4 kHz. Maternal education was classified according to the 2001 UK census. The hearing comparison group (HCG; N = 63) was randomly selected from babies born on the same day and in the same district as participants in the D/HH group. Seventy-six of the 120 D/HH and 38 of the 63 HCG participated at Time 2 (T2). Seventeen of the 76 D/HH participants did not complete the spoken language assessments at T2. This was either because they used British Sign Language (BSL) as their preferred language, rendering these spoken English assessments inappropriate, or because they had severe additional disabilities that precluded the development of sufficient language to attempt the tests. The analyses presented in this paper are based on the 53 participants with PCHL and 38 participants in the HCG for whom data were available on tests of reading and language at both T1 (mean age 8.0 years) and also at T2 (mean age 17.3 years). The requirement for reading and language measures to be available at both time points was necessary to allow all the longitudinal analyses to be conducted on the same sample of children. If reading and language measures were available, participants were not excluded because of the presence of additional disabilities. The two such participants both had learning disabilities i.e. nonverbal IQ < 70 at T1 (see 3.4 for results of sensitivity analysis). The demographic characteristics of the samples are presented in Table 1. Sample attrition over the approximately eight years between these two assessment time points, coupled with inclusion only of those participants who were spoken language users, reduced the sample size for the longitudinal analysis presented here. The annual rate of attrition was 4% since their assessment at primary school. This degree of attrition is relatively low for follow-up studies of long-term paediatric conditions (Karlson & Rapoff, 2009). Attrition was principally due to the participants not responding to requests to participate in later phases of the study (for details see Pimperton et al., 2017, Fig. 1). The summary statistics presented in Table 1 show that both the PCHL and HCG groups in the present study were similar in their demographic characteristics to those of the initial samples.","The study was approved by the Southampton and South West Hamphsire Research Ethics Committee. Written informed consent for participation in the study was obtained from principal caregivers at T1and T2 and from the teenage participants at T2. At both T1 and T2, each participant’s reading, language and non-verbal ability (N-VA) was assessed in a quiet room at home or school by a trained researcher. At the same time we collected, from participants and their families, information on characteristics including maternal education level and languages used in the home. The most recently available audiological data were taken from audiology and cochlear implant centre records – for participants with hearing aids from the last annual review and for those with cochlear implants, unaided pure-tone thresholds obtained during the original implant assessment. Materials The group mean score and standard-deviation scores in the HCG were used to derive z scores for the D/HH participants. The z-score for a D/HH participant is equal to the number of standard deviations of the distribution of scores in the HCG participants by which the D/HH participant’s age-adjusted score differs from the mean score of the HCG participants. Reading Reading Accuracy and Reading Comprehension at T1 were measured using the Wechsler Objective Reading Dimensions (WORD) (Wechsler, 1993) and at T2 using the newly available York Assessment of Reading for Comprehension Secondary Edition (YARC), (Stothard, Hulme, Clarke, Barmby, & Snowling, 2010). The YARC was used at T2 because, with minor adaptations approved by the test designers (see Pimperton et al., 2016), we could use this measure with both spoken and sign language users in our sample as it did not involve reading aloud. The difficulty of reading comprehension tasks undertaken in the YARC was determined by word reading performance rather than chronological age, which was also felt to be more appropriate for our D/HH sample. Language comprehension At both T1 and T2 the Test for Reception of Grammar (TROG-2; Bishop, 2003) and The British Picture Vocabulary Scale (BPVS-3; Dunn, Dunn, & National Foundation for Education Research, 2009), were used to assess receptive skills for spoken English grammar and vocabulary respectively. TROG-2 contains test items that assess understanding of increasingly complex grammatical contrasts, including plurals, passives, negatives, and relative clauses. In both tests participants point to a picture from a choice of four alternatives that corresponds to a spoken stimulus. The z scores from the TROG-2 and the BPVS were highly correlated at both Time 1 (n = 53, r = 0.82) and Time 2 (n = 53, r = 0.67). Accordingly the two z-scores were averaged to produce Language Comprehension scores at Time 1 and at Time 2. Expressive language Scores from age appropriate narrative assessments (Renfrew Bus Story Test (Renfrew 1995) at T1, and Expression, Reception and the Recall of Narrative Instrument (ERRNI; Bishop, 2004) at T2) provided an Expressive Language score. The Renfrew Bus Story Test was developed for 3–8 year olds and involved children listening to a story told by the researcher supported by a series of pictures that correspond to the story. They then retold the story using the pictures as prompts and their retelling was audio-recorded and transcribed by the administering researcher. Two z scores from this measure (amount of information in the narrative and the average length of utterances i.e. number of words used) were averaged into an Expressive Language composite score (see Kennedy et al., 2006 for details). In a reliability exercise, data for 15 randomly chosen participants were independently transcribed by a second rater. No discrepancy of word content was found between transcriptions but the point at which a sentence should pause (e.g. commas, full stops) was a subjective decision. The Bland-Altman method (Bland & Altman, 1986) was used for assessing agreement between the two transcripts on average length of the longest 5 sentences and the information score. There was an average difference of 1.1 units between the two transcripts on these two scores (0.43 and 1.53 respectively) and variability of these differences was acceptable (95% confidence intervals −1.80 to 2.65 and −5.68 to 8.75 respectively). At T2 the ERRNI was selected, as it was similar in design to the Bus Story Test but suitable for use with adolescents. It requires test-takers to produce a narrative based on a series of cartoon pictures then reproduce that narrative, though this time without the support of pictures. As with the Bus Story Test, the ERRNI narratives were audio-recorded and transcribed by the researcher who had taken the recording. The ERRNI produced three scores: an Initial score for the quality of participants’ initial narratives, a Recall score for the quality of their recalled narrative, and a Mean Length of Utterance (MLU) score which gave the average length of utterances across both the initial and recalled narratives. The Initial and the Recall z scores were highly correlated (n = 53, r = 0.72) and therefore averaged to form an Expressive Language-Information score at T2. The MLU measure was less highly correlated with the other two T2 Expressive Language measures (n = 53, r = 0.20 and 0.32) and was therefore analysed separately. Following Whitehouse, Line, Watt, & Bishop (2009), an inter-rater reliability exercise was carried out to check the reliability of the ERRNI scoring. A random sample of 12 narratives (12% of the total) was transcribed and scored by a second rater. There was good agreement between the two ratings for all three scores (intraclass r: Initial = 0.82; Recall = 0.90; MLU = 0.95). Speech intelligibility Parents’ assessment of connected speech intelligibility was assessed at T1 and T2 using the Speech Intelligibility Rating Scale (Allen, Nikolopoulos, & O'Donoghue, 1998). This 5 point scale gives short descriptors of each level with 1 being ‘completely unintelligible’ and 5 being ‘intelligible to all listeners’ and was developed for use in cochlear implant assessment. Non-verbal ability At T1 we assessed Non-verbal Ability using the Raven’s Standard Progressive Matrices (Styles, Raven, & Raven, 1998). At T2 the 20 min timed version (Hamel & Schmittman, 2006) was used. Participants were given twenty minutes to work their way through a series of progressively more complex matrix reasoning puzzles. Raw scores reflecting the total number of correct items out of a possible 60 were calculated. Data analysis Regression analyses were conducted using SPSS 24 (IBM Corp., 2016). Following the recommendations by Kraemer and Blasey (2004), all the independent measures were centred before being included in the regression analyses. For this longitudinal analysis it was important to compare directly the same D/HH participants relative to the same NH control group at both time points (T1 and T2). The language z scores at T1, which had initially been derived relative to the scores for all participants in the HCG at T1, were therefore recalculated relative to the scores of only those HCG participants at T1 that also participated at T2. These recalculated z scores were used in the subsequent analyses. A preliminary test was conducted to determine whether there were possible covariates that might be confounding predictors of reading at T2. As recommended by Kraemer (2015) the number of putative covariates was kept to a minimum by requiring them to be significantly correlated with either of the T2 reading measures. Forced entry was used to enter the independent variables in blocks with the T1 reading score entering at Step 1, the covariates entering in Step 2 followed by the T1 Expressive Language and Language Comprehension measures at Step 3. The R2 change from step 2 was used to test for any significant effects of T1 language on T2 reading. Only the R2 values and their associated F tests are reported, as these are unaffected by collinearity. For each regression, the distribution of residuals was tested for skewness and kurtosis and the Kolmogorov-Smirnov test of normality was applied. For the prediction of T2 Reading Accuracy and for T2 Reading Comprehension these tests all had p values greater than .05 indicating that the distribution of the residuals did not deviate significantly from normal. Some previous studies on reading development in the D/HH have excluded children with low non-verbal cognitive abilities (Kyle & Harris, 2011). Accordingly, a sensitivity analysis was conducted to determine whether their inclusion had an impact on the results in the present study. Comparison of the PCHL and HCG mean scores ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The Reading Accuracy, Reading Comprehension, Expressive Language, Language Comprehension and Non-verbal Ability scores at T1 and T2 for the PCHL group and the HCG are compared in Table 2. The scores for the HCG (n = 38) on all measures are mean = 0.00 and SD = 1. Scores of participants with PCHL were lower than those of the HCG on all language and reading measures with the exception of T2 Expressive Language-Information score and T2 Expressive Language-MLU score. There was a wider range of achievement at T2 than T1 in the DH/H group and the spread of the D/HH group reading scores was larger than that in the HCG with approximately 25% achieving reading scores above the HCG mean at T2. The spread of D/HH group reading scores had also increased by age 17 years and was larger than that in the HCG. At T1 those with PCHL had significantly lower Non-verbal Ability scores than the HCG. There is a degree of normalisation over time of Non-verbal Ability in the DH/H group and consequently there is no significant difference from the HCG at T2. The increase in Non-verbal Ability in the D/HH group was large enough to be clinically important and was also statistically significant (SMD = −0.60, 95%CI −0.98 to −0.21, t = 4.41, df = 52, p < .001). At T2 all those in the D/HH group were intelligible to listeners who had at least some experience of the speech of children with PCHL (Appendix A). Testing for potentially confounding covariates ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Correlations between potential covariates and T2 Reading Accuracy and T2 Reading Comprehension for PCHL (n = 53) were calculated (see Table 3). It should be noted that the Reading Accuracy and Reading Comprehension scores were age-adjusted z-scores and consequently age would not be a confounder. Gender was not related to either T2 Reading Accuracy or T2 Reading Comprehension and therefore was not included as a covariate. All the other covariates were significantly correlated with at least one T2 reading score and were retained. These covariates were mother’s education, first language English, severity of hearing loss and T1 Non-verbal Ability. Predicting Time 2 scores ~~~~~~~~~~~~~~~~~~~~~~~~ The results of forced entry stepwise regression predicting Reading Accuracy at T2 are shown in Table 4. There was a high level of stability in the Reading Accuracy scores (Step 1 R2 = 0.63, p < .001). This indicates that 63% of the variance in T2 Reading Accuracy scores at mean age 17 years was predicted by Reading Accuracy scores measured some 9 years earlier. The covariates jointly added 0.07 to the R2 (p = .04). The T1 language scores did not add significantly to the prediction of the T2 Reading Accuracy scores (Step 3 change in R2 = 0.01, p = .459). A regression analysis adding T1 Expressive Language and Language Comprehension sequentially (i.e. step 3a then step 3b) showed that neither explained significant unique variance in T2 Reading Accuracy i.e. there was no significant change in R2 at step 3b whichever order they were entered. Forced entry stepwise regression predicting Reading Comprehension scores at T2 showed that there was a moderate degree of stability in the Reading Comprehension scores (Step 1 R2 = 0.43, p < .001) i.e. reading comprehension scores at T1 accounted for 43% of the variance of scores of the same skill at T2 (Table 5). The covariates did not add significantly to the prediction of the T2 Reading Comprehension (Step 2 change in R2 = 0.04). However, the two T1 language scores made a significant contribution to the prediction of T2 Reading Comprehension (Step 3 change in R2 = 0.17, p < .001). A regression analysis adding T1 Expressive Language and Language Comprehension sequentially (i.e. step 3a then step 3b) showed that although there was much shared variance, each of the language variables also made a significant unique contribution to the explained variance. Language Comprehension explained a significant additional 6% (p < .001) of the variance in Reading Comprehension beyond that explained by Expressive Language. Expressive Language accounted for an additional 3% (p < .05) of the variance in Reading Comprehension beyond that explained by Language Comprehension. Sensitivity test for other disabilities ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The analyses reported in 3.1–3.3 were repeated with the exclusion of the two participants with learning disabilities i.e. non-verbal IQ < 70. No substantive changes were shown in the results in terms of the parameter estimates and their associated p values.","This study examined relationships between language and reading, measured in middle childhood and adolescence, in a population-based sample of D/HH children with bilateral PCHL > 40 dB using spoken English as their chief mode of communication. Approximately half of the sample had a moderate hearing loss and the majority had language skills within 2 SDs of their HCG peers. The D/HH participants had significantly lower scores than the HCG on reading measures in both middle childhood and adolescence. Such group-level deficits in reading replicate findings previously reported in children with hearing loss of varying severities (Moeller et al., 2007; Geers & Hayes, 2011; Harris & Terlektsi, 2011). In the present study at mean age 17 years, the reading scores of the D/HH were approximately 1 SD below those of the HCG and this gap was slightly greater than it had been for the same participants at mean age 8 years. A similar pattern of relative decline in reading ability was reported in 11 year olds by Kyle & Harris (2010), where children who were severe- profoundly D/HH made only 0.3 of a grade improvement in reading age per year between the ages of 8 and 11 years. Early superiority in reading skills may have enabled the HCG to read more demanding material more frequently than their D/HH peers, increasing the skill gap and resulting in a rich-get-richer ‘Matthew effect’ (Stanovich, 1986) which is likely to impact on their educational outcomes and subsequent life chances. When examining the stability of reading scores in D/HH individuals in this study, we found, as expected, that the majority of variance of reading scores at T2 was accounted for by scores of the same abilities measured at T1. This extends the demonstration in typically developing young people with normal hearing that early literacy was highly predictive of reading ability at 17 years of age (Cunningham & Stanovich, 1997). In that study this effect was mediated, in part, by the impact of early reading on exposure to print. For some D/HH individuals such a virtuous circle of early reading enhancing later achievement may be more difficult to initiate, given the difficulties they can experience in acquiring reading skills. The main focus of this paper was to determine the extent to which language abilities contribute to later reading abilities in D/HH children. We found that receptive and expressive spoken language abilities in middle childhood accounted for significant variance in adolescent Reading Comprehension but not Reading Accuracy. A key aspect of this result is that the analyses adjusted for Reading Accuracy or Comprehension performance in middle childhood when predicting the same skill in adolescence thus demonstrating that these language variables explain variance in reading comprehension development from middle childhood to adolescence. In addition, we found that Expressive Language and Language Comprehension measured at age 8 years each explained small but significant unique portions of the variance in Reading Comprehension at age 17 years in addition to the variance they explained in common. The finding that broader oral language skills predicted reading comprehension outcomes in this sample of D/HH adolescents is consistent with findings in hearing and in younger deaf children regarding the salience of these skills for predicting reading comprehension development. (e.g. Nation et al., 2010; Kyle & Harris, 2010). This consistency with the findings of Kyle and Harris (2010) is observed despite the differing distribution of hearing loss severity within the two samples. Taken together these consistent findings suggest a common role for broader language skills in facilitating reading comprehension for hearing and for D/HH children across the severity spectrum. Conversely the lack of predictive relationship between these same language skills and reading accuracy is consistent with findings from hearing children that reading accuracy is best predicted by measures of phonological skills (e.g. Muter et al., 2004). However, a significant limitation of our study was the absence of a measure of phonological ability at age 8 years, which meant we were unable to assess the predictive relationships between phonological language skills and reading accuracy and comprehension outcomes in this sample. Kyle & Harris (2010, 2011) found that receptive vocabulary is an important predictor of word reading in younger D/HH children. T1 vocabulary might have predicted T2 reading accuracy in our study, but as receptive vocabulary and receptive grammar scores were combined to create a composite receptive language score (see Methods) and vocabulary was not examined separately we are not able to say. The fact that receptive language was not an important determinant of reading scores in the present study may be due to differences in the ages or the distribution of hearing loss severity and communication modes between the two study populations. The receptive language tests selected for this study have been widely used with deaf children in research, clinical and educational settings. Although standardised on hearing children they require a non-verbal response which avoided wrongly scoring mis-pronounced responses as if they were the consequence of limited syntax or vocabulary ability. Furthermore the tests had been standardised on a wide age range allowing use of same measures at both time points. For the expressive language tests, the Bus Story (T1) requires high levels of oral language comprehension (the participant listens to then retells the story). However, the ERRNI does not (participants generate the story themselves from pictures). We have previously highlighted this difference (Pimperton et al., 2017) as a potential contributor to the reduction in the expressive language deficit of the D/HH group relative to the hearing group at T2. Improvement in speech intelligibility of the D/HH group from T1 to T2 (see additional Appendix A) is another potential contributor to this effect. This likely involvement of language comprehension in the expressive language measure at T1 is consistent with the finding that there is much shared variance between T1 language comprehension and T1 expressive language when predicting T2 reading comprehension; though T1 expressive language still predicts a small but significant amount of unique variance beyond that explained by T1 language comprehension. An unexpected finding was improvement in non- verbal ability in those with PCHL between the ages of 8 and 17 years. This does not appear to have been reported in previous studies. The finding is given weight by being obtained in a longitudinally studied sample where the same children provided non-verbal ability scores at the two ages. However, the finding may be sample specific and would benefit from replication in other samples e.g. the sample studied by Wake et al. (2004). Alternatively, it may relate to having sufficient language to support reasoning processes but we know of no evidence of such a mechanism accounting for this relative normalisation of non-verbal ability. There are consistent reports that severity of hearing loss is related to reading development in D/HH individuals (e.g. Moeller, Tomblin, & the OCHL Collbaration, 2015). What has been less studied is what might mediate this effect. It is unlikely that factors linked to the aetiology of the hearing loss also acted to influence reading development, except those that affected non-verbal ability (which was included in our analysis) or were associated with progressive hearing loss. In the present study, severity of hearing loss was related to reading score but the severity of hearing loss per se did not predict reading ability once the effects of language ability (using our composite expressive + receptive language score) were controlled. This suggests that the association between severity of hearing loss and the difficulties in acquiring reading may be mediated by the effect of hearing loss on language ability. This study has a number of strengths including the longitudinal design and the population-based sampling. In such a population based longitudinal study over a 9-year period there is inevitably sample attrition. In the present study the evidence suggests that in most respects the retained participants were representative of the original sample. As the teenagers assessed in this study were born in 1992–1997, the question arises whether the reading skills of D/HH children born more recently might show different levels of reading, and different relationships between language and reading skills, as a result of advances in hearing device technology and in early identification and intervention since the mid 1990s. A recent report of reading skills of children with severe-profound PCHL aged 5–7 years when recruited in 2013–14 (i.e. born in 2006–09) described little improvement in phonological awareness or reading ability compared with a similar study sample born 10 years earlier, but did report improvements in vocabulary levels (Harris, Terlektsi & Kyle, 2017b). Predictive relationships between the language and reading variables were consistent across the two cohorts in some cases (e.g. vocabulary) but differed in others (e.g. phonological awareness). A replication of the longitudinal analyses reported in the current study in cohorts born more recently would be valuable to address whether the pattern of findings reported here generalises to these later-born cohorts. Our findings provide support for the value of intervention to foster language development in D/HH children (Rees et al., 2015; Gilliver, Cupples, Ching, Leigh, & Gunnourie, 2016; Richels et al., 2016) but importantly suggest that continued support for broader oral language skills into the secondary school years could continue to be an effective way to enhance reading comprehension ability in this population. Maximising the reading comprehension skills in D/HH individuals of secondary school age is vital because of the increasing demands on comprehension skills as text complexity increases in the secondary school years, and the vital role that reading comprehension plays in broader educational success at secondary school.","This study extends our understanding of reading development in D/HH children and adolescents by highlighting the predictive relationships between language skills in middle childhood and reading comprehension in adolescence, even when adjusting for earlier reading skills. These findings contribute to the case for continued targeted intervention on language skills for D/HH individuals beyond primary school. Maximising their language ability, and consequently their ability to use reading to access learning, is likely to enhance their educational attainment and subsequent life chances.","The authors declare that they have no conflicts of interest in presenting the results in this paper."],["We investigated whether childhood factors that are amenable to intervention (parenting stress, child psychological problems and pain) predicted participation in daily activities and social roles of adolescents with cerebral palsy (CP). We randomly selected 1174 children aged 8-12 years from eight population-based registers of children with CP in six European countries; 743 (63%) agreed to participate. One further region recruited 75 children from multiple sources. These 818 children were visited at home at age 8-12 years, 594 (73%) agreed to follow-up at age 13-17 years.We used the following measures: parent reported stress (Parenting Stress Index Short Form), their child's psychological difficulties (Strength and Difficulties Questionnaire) and frequency and severity of pain; either child or parent reported the child's participation (LIFE Habits questionnaire). We fitted a structural equation model to each of the participation domains, regressing participation in childhood and adolescence on parenting stress, child psychological problems and pain, and regressing adolescent factors on the corresponding childhood factors; models were adjusted for impairment, region, age and gender.Pain in childhood predicted restricted adolescent participation in all domains except Mealtimes and Communication (standardized total indirect effects β -0.05 to -0.18, 0.01. <. p<. 0.05 to p<. 0.001, depending on domain). Psychological problems in childhood predicted restricted adolescent participation in all domains of social roles, and in Personal Care and Communication 9β -0.07 to -0.17, 0.001. <. p<. 0.01 to p<. 0.001). Parenting stress in childhood predicted restricted adolescent participation in Health Hygiene, Mobility and Relationships 9β -0.07 to -0.18, 0.001. <. p<. 0.01 to p<. 0.001). These childhood factors predicted adolescent participation largely via their effects on childhood participation; though in some domains early psychological problems and parenting stress in childhood predicted adolescent participation largely through their persistence into adolescence.We conclude that participation of adolescents with CP was predicted by early modifiable factors related to the child and family. Interventions for reduction of pain, psychological difficulties and parenting stress in childhood are justified not only for their intrinsic value, but also for probable benefits to childhood and adolescent participation. --------------------------------------------------------------------------------","Children with cerebral palsy (CP) experience restricted participation in life situations ranging from leisure pursuits to education and social roles (Beckung & Hagberg, 2002). Most children with CP live to adulthood, where they remain at higher risk of social disadvantage than adults without CP in terms of independent living, employment and establishing a family (Michelsen, Uldall, Hansen, & Madsen, 2006). Adolescence may be particularly challenging for young people with physical impairments (King, Brown, & Smith, 2003a). Delayed puberty, the psychological consequences of perception of body image, and fewer opportunities to socialise out of school may make this period more difficult. Medical care may be jeopardised as responsibility transfers from parent to young person and from child to adult health services. Adolescents with CP have restricted participation in daily activities and social roles which depends on the severity of their impairments (Donkervoort, Roebroeck, Wiegerink, van der Heijden-Maessen, & Stam, 2007). Participation of children and adolescents with CP is associated with the modifiable factors: pain (Fauconnier et al., 2009), psychological problems (Ramstad, Jahnsen, Skjeldal, & Diseth, 2012) and parenting stress (Majnemer et al., 2008). However, evidence is scarce about the modifiable factors in childhood which predict participation in adolescence. A study with a longitudinal design (Holmbeck, Franks Bruno, & Jandasek, 2006) can help to distinguish participation patterns determined by factors operating in adolescence from patterns determined by factors already operating in childhood. The objective of this paper is to evaluate how participation of adolescents with CP is associated with modifiable childhood factors: pain, psychological problems, and parenting stress. We studied whether these associations were mediated by participation in childhood or by the level of the same predictors in adolescence. Setting and participants ~~~~~~~~~~~~~~~~~~~~~~~~ The present work is part of a larger project, SPARCLE, which studies the participation and quality of life of children and adolescents with CP in Europe. The overall design of the project, including sample size calculations, is described elsewhere (Colver & Dickinson, 2010; Colver, 2006) and is summarised below. Children born between 31/07/1991 and 01/04/1997 were randomly sampled from population-based registers of children with CP in eight European regions (Table 1) that share a standardised definition of CP (Surveillance of Cerebral Palsy in Europe (SCPE), 2000). 743/1174 (63%) target families identified from registers joined the study. One further region, northwest Germany, ascertained 75 cases from multiple sources, using the same diagnostic criteria. The 818 children who entered the study were interviewed initially in 2004/2005, aged 8–12 years (SPARCLE1), and followed up in 2009/10, aged 13–17 years (SPARCLE2), when 594 (73%) remained in the study. Predictors of drop-out have been reported (Dickinson et al., 2006, 2012). Researchers from the nine regions visited families in their homes to administer questionnaires to parents and their children. The researchers had attended common training in order to maximise homogeneity of survey methodology across regions.","We evaluated participation using the questionnaire of Life Habits (LIFE-H) (Noreau et al., 2004) which is based on a social model of disability similar to the theoretical framework of the World Health Organisation's International Classification of Functioning (World Health Organisation, 2007) and has been validated in children with disabilities (Noreau et al., 2007). Wherever possible the adolescent completed the questionnaire; otherwise a parent completed it. It consists of 62 items divided into six domains of daily life activities (Mealtimes, Health hygiene, Personal care, Communication, Home life, and Mobility) and five domains of social roles (Responsibilities, Relationships, Community life, School, and Recreation). It includes fifteen “non-discretionary” activities, such as transferring into or out of bed, which are essential for daily living; and forty-seven further “discretionary” activities, such as exercise to optimise health, which may or may not be achieved. We recorded the discretionary items using three levels: not achieved because too difficult; achieved with difficulty; achieved without difficulty. A discretionary item could also be considered non-applicable if it was irrelevant, for example if the child or adolescent had no interest in that activity; we treated non- applicable items as missing responses. We recorded non-discretionary items in childhood using two levels (achieved with difficulty, achieved without difficulty) but in adolescence we used three levels (achieved with much difficulty, achieved with some difficulty, achieved without difficulty). The Life-H asks, for each item, how much assistance the young person requires, but we ignored this information because we wanted to assess participation without incorporating the influence of environmental factors (Fauconnier et al., 2009). In order to assess pain, we asked parents about the frequency and severity of their child's pain over the previous week; we recorded responses on six levels, but grouped them into three categories for analysis. We captured the psychological problems of the child or adolescent using the Total Difficulties Score of the parent- reported Strength and Difficulties Questionnaire (SDQ) (Goodman, 1997). We captured parenting stress using the Total Stress Score of the Parenting Stress Index Short Form (PSI-SF) (Abidin, 1995). Parents provided information about their child's impairments (walking ability as captured by the gross motor function classification system (GMFCS) (Palisano et al., 1997), fine motor function (Beckung & Hagberg, 2002), seizures, feeding, communication, intellectual ability (White-Koning et al., 2005)), family structure and parents’ educational qualifications. Statistical methods ~~~~~~~~~~~~~~~~~~~ Full details of the statistical methods are reported in Appendix Statistics and summarised below. For each participation domain, we translated hypotheses about the variables that might influence participation into one structural equation model which comprised a ‘measurement part’ that defined the latent constructs that underlie sets of observed variables; and a ‘structural part’ that hypothesised the links between these constructs. The measurement part (illustrated in Fig. 1 for one domain, Home life) considered participation in each domain to be an unobserved or latent variable, manifested by responses to the items in the LIFE-H questionnaire. We likewise modelled impairment and pain using latent variables, manifested respectively by the levels of the six individual impairments as recorded in childhood and by the frequency and severity of pain over the previous week (see Fig. 1). The structural part specified the hypothesised links between the variables, both latent and observed (see Fig. 2). For each participation domain, we hypothesised that participation in both childhood and adolescence may be directly affected by the concurrent factors: pain (Doralp & Bartlett, 2010), psychological problems (Ramstad et al., 2012), and parenting stress (Blum, Resnick, Nelson, & St Germaine, 1991). Additionally, we hypothesised that each factor in childhood could indirectly affect participation in adolescence via its influence on four mediating variables: participation in childhood, and the three factors in adolescence. We adjusted for impairment, region, gender, and age. Psychological problems, parenting stress, region, gender, and age were treated as observed variables; region and gender were categorical while the others were continuous. Our analysis followed the steps outlined by Kline (2011) for a structural equation model. We assessed model goodness of fit by the root mean square error of approximation (RMSEA), which indicates a good fit when RMSEA <0.05. We identified the childhood factors that were significantly related to each domain of childhood participation through preliminary analysis that was restricted to childhood variables. We then developed the full model, retaining only these significant childhood factors and similarly identifying adolescent factors that were significantly related to adolescent participation. We report the estimated indirect effects (β) of childhood factors (pain, psychological problems, parenting stress) on adolescent participation. The total indirect effects were the sum of the partial indirect effects via all four possible pathways (see Figures and Appendix Statistics). These indirect effects were standardised in order to compare the contributions of different predictors to participation; a standardised effect of β of a predictor on a specific outcome means that an increase of one standard deviation in the predictor is associated with a change of β standard deviations in the outcome. As it was of interest to compare the effect of modifiable childhood factors with the effect of impairment, we also noted the standardised direct, indirect and total effects of impairment on adolescent participation. We determined statistical significance (p-values) from the estimated standard errors of the unstandardised effects. Finally we undertook sensitivity analyses around drop-out, including all 818 SPARCLE1 participants and imputing missing data (van Buuren, 2007). We analysed the data using Mplus software, version 6.12 (Muthén & Muthén, 1998). Ethics ~~~~~~ In each country, we obtained ethical approval or a statement that only registration was required, as appropriate. We obtained signed consent from all parents and from young people who could give meaningful consent.","The average age of the children was 10.4 years at the first visit and 15.1 years at the second visit; 249 (42%) were girls. The average time between visits was 4.7 years (3.6–5.8 years). Table 1 shows the distribution of putative predictors of participation. The rate of missing data was low: not more than 3% for any variable. Parents reported that approximately two thirds of their children experienced at least some pain during the previous week. The Total Difficulties Score of SDQ was abnormal (>16) in 21% of adolescents, a proportion twice that in the general population (Goodman, 1997). The Total Stress Score of PSI-SF was abnormal (>90) for 34% of their parents, twice the proportion in the general population (Abidin, 1995). Table 2 shows the distribution of participation items for each age group. Non-response remained less than 3% except for the item of extra classes in childhood. Among adolescents, discretionary items were considered non- applicable by between 0% (getting a good sleep) and 52% (religious activities). Both items of the Community life domain were considered non-applicable by 26%, with marked differences between regions, so we excluded that domain from analysis as in our prior study (Fauconnier et al., 2009). The proportion of adolescents achieving an item without difficulty varied widely, from 31% (moving on slippery or uneven surfaces) to 93% (maintaining a loving relationship with one's parents). In preliminary investigations, we fitted the measurement model for each domain and generated participation scores for each adolescent. These varied significantly between regions for all domains except Relationships; excluding this domain, the variation between regions was 3–11% of the total variation in participation, depending on domain; in contrast, the variation between levels of GMFCS was 22–60% of the total variation. Predictors of adolescent participation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 3 and Fig. 3 summarise the estimated effects of early predictors on adolescent participation. The goodness of fit of the models was good (0.037 ≤ RMSEA ≤ 0.048). The models explained between 61% and 90% of the variance of participation in adolescence, except in the Relationships domain where the model explained only 41%. As expected, impairment predicted a significant restriction of adolescent participation in all domains, with standardised total effects ranging from β = −0.37 (p < 0.001) for Relationships to β = −0.88 (p < 0.001) for Mealtimes. Pain in childhood predicted a significant restriction of adolescent participation in all domains except Mealtimes; the size of its standardised indirect effect was most marked in Health hygiene (β = −0.18, p < 0.001) and Relationships (β = −0.14, p < 0.001) (indicating that an increase of one standard deviation in childhood pain was associated with decreases of 0.18 and 0.14 standard deviations respectively in these domains of participation). Psychological problems in childhood predicted a significant restriction in adolescent participation in all domains of social roles, effects ranging from β = −0.11 in Relationships (0.001 < p < 0.01) to β = −0.17 in Responsibilities (p < 0.001), and in the daily life activities of Personal care and Communication (β = −0.07 and −0.11 respectively, p < 0.001). Parenting stress in childhood predicted restricted adolescent participation in Health hygiene (β = −0.11, p < 0.001), Mobility (β = −0.07, 0.001 < p < 0.01) and Relationships (β = −0.18, p < 0.001). Adolescent participation in Mealtimes was not significantly associated with any of the childhood predictors. Pathways to participation ~~~~~~~~~~~~~~~~~~~~~~~~~ Direct and indirect effects of impairment were of similar magnitude in most domains, the indirect effects being mediated largely by childhood participation. The associations between modifiable childhood predictors and adolescent participation were largely mediated by childhood participation, which was a strong predictor of adolescent participation in most domains: a change of one standard deviation in childhood participation predicted a change of between 0.28 and 0.68 standard deviations in adolescent participation, depending on domain (see Fig. 3). The influence of childhood pain on adolescent participation was essentially mediated via its direct effect on childhood participation (see Table 3 and Fig. 3); the partial indirect effects of childhood pain via adolescent factors were generally small or negligible (partial β ≤ 0.04). Psychological problems in childhood predicted adolescent participation in Communication, Responsibilities and Relationships mainly via child participation (partial β = −0.07, −0.12 and −0.11 respectively) but they predicted adolescent participation in the domains of Personal care, School and Recreation mainly via psychological problems in adolescence. Parenting stress in childhood was also significantly related to adolescent participation in Mobility and Relationships via childhood participation (partial β = −0.07 and −0.11 respectively) but via parenting stress during adolescence in the domains of Health hygiene and Relationships. Sensitivity analysis, which imputed missing data for those who dropped out between childhood and adolescence, yielded similar results (data not shown).","Childhood participation was the main predictor of adolescent participation (a change of one standard deviation in childhood participation predicted a change of between 0.28 and 0.68 standard deviations in adolescent participation, depending on domain). Three factors in childhood (pain, psychological problems and parenting stress) predicted, in varying degrees, restricted participation at adolescence in all domains except Mealtimes. However, these effects were small: a change of one standard deviation in any of the three childhood factors predicted a change of at most 0.18 standard deviations in adolescent participation. Furthermore, these three childhood factors predicted adolescent participation largely via their effects on childhood participation, although in some domains early psychological problems and parenting stress in childhood affected adolescent participation through their persistence into adolescence. Effects of impairment were much larger than the effects of these childhood factors: a difference of one standard deviation in impairment was associated with a difference of more than 0.6 standard deviations in adolescent participation in all domains except Relationships, the main pathway again being via childhood participation. Strengths and limitations ~~~~~~~~~~~~~~~~~~~~~~~~~ Because sampling of the children was multinational and from population registers of children with CP, conclusions may be generalised to the population of adolescents with CP living in Europe. As in any regression, statistical associations do not prove causation. Unmeasured, shared causes could explain part of the associations; for instance parenting style may influence both parenting stress and participation. Although we based our hypothesised directions of effects on prior research (Dang, 2012), alternative directions of effects should be considered (King et al., 2003b). For instance, increased participation may improve the psychological health of the child (Dahan-Oliel, Shikako- Thomas, & Majnemer, 2012). Non-response by families targeted for recruitment to SPARCLE1 was 37% (Dickinson et al., 2006), and drop-out between SPARCLE1 and SPARCLE2 was 27% (Dickinson et al., 2012). The parents who dropped out between SPARCLE1 and SPARCLE2 had a higher level of stress when their children were aged 8–12 than those retained in the study. Although such differential non-response is likely to result in biased estimates of population means, it may be less important in the present study which estimates associations (Korn & Graubard, 1999). Nevertheless, we tried to minimise the effects of differential non-response and drop-out in two ways. Firstly, we adjusted for region and walking ability, which were predictors of non-response (Dickinson et al., 2006; Korn & Graubard, 1999). Secondly, we performed a sensitivity analysis which included all children who participated in SPARCLE1, imputing missing data; this yielded similar results to the primary analysis. We considered that the young person knew most about their participation; but if a young person could not self-report due to intellectual impairment we relied on parent-report. Parent-reported pain may differ from self-reported pain (Parkinson, Gibson, Dickinson, & Colver, 2010), but we chose it in order to have a common metric across the sample. Prior research has tended to examine the impacts of specific impairments on activities (Beckung & Hagberg, 2002; Fauconnier et al., 2009), but in our study we controlled for impairment using a single latent variable because children with CP are affected by several correlated functional limitations reflecting the global severity of the cerebral disturbance. The model did not explicitly take account of environmental influences such as the availability of specialised schools or the accessibility of transportation; however, by controlling for region, we took account of the regional variation in environments. As in any structural equation approach, other models may fit the data equally well (Kline, 2011). The prevalence of the predictors of participation in our sample was comparable to that in other studies; for example, pain was present in 50–70% of children and adolescents with CP (Doralp & Bartlett, 2010; Engel, Petrina, Dudgeon, & McKearnan, 2005; Hadden & von Baeyer, 2002), psychological problems in 39% to 54% of children with CP (Brossard-Racine et al., 2012; Goodman & Graham, 1996), and symptoms of stress in about 30% of mothers of children with CP (Manuel, Naughton, Balkrishnan, Paterson Smith, & Koman, 2003; Sawyer et al., 2011). Comparison with other studies ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We could not find other longitudinal studies of early predictors of participation of adolescents with disabilities. However, King et al. (2009) examined predictors of change in intensity of participation in leisure and recreational activities over a three year period of children with physical disabilities aged 6–15. They found that participation intensity declined over the years in recreational, active physical and social activities but not in skill-based and self-improvement activities. Factors associated with these changes varied with type of activity and the child's age and sex. Their conclusions emphasised individual variability and proposed that interventions should be tailored to the individual child. It is difficult to compare their study with ours, because they measured the intensity (frequency) of participation in leisure activities and therefore could not address, as we did, difficulty in participation in essential daily activities such as feeding and toileting. Our findings are consistent with findings from cross- sectional studies that pain limits daily activities in children (Tervo, Symons, Stout, & Novacheck, 2006) and adolescents with CP (Doralp & Bartlett, 2010). Pain reduces children's school attendance (Houlihan, O’Donnell, Conaway, & Stevenson, 2004) and predicts altered school functioning via fatigue (Berrin et al., 2007). Psychological problems in childhood may contribute to friendlessness and restricted participation in social roles (Doll, 1996), and be associated with decreased participation in children with CP (Ramstad et al., 2012). Parenting stress is associated with more coercive parent–child interactions (Plant & Sanders, 2007), and this could restrict the freedom of the child or adolescent to experiment with activities. Other studies of children and adolescents with disabilities have highlighted the range and complex inter-relationships of child, family and community factors that predict participation (Colver et al., 2012; King et al., 2003b, 2006, 2009; Orlin et al., 2010; Yeung & Towers, 2014). In particular, King et al. (2006) used structural equation modelling to undertake a cross sectional analysis of children aged 6–14 years with physical disabilities, including CP, to examine child, family and environmental influences on leisure and recreational participation. Family participation in social and recreational activities influenced the child's participation, but the standardised β coefficient was only 0.18. Child preferences had a stronger β coefficient of 0.28 but this may reflect not only the personality and interests of the child but also environmental factors; for example, a child that has experienced unfriendliness in leisure settings will prefer not to attend such settings (Colver, 2010). King et al. (2006) found only a small indirect effect of an unsupportive environment; however this small effect may be partly because they used CHIEF to measure the physical, social and attitudinal environment; CHIEF is a subjective measure of the frequency and extent of perceived environmental barriers on participation rather than a direct measure of the environment; it may therefore reflect differing expectations of participation rather than actual environmental barriers. In our study of the cohort aged 8–12, we found that their physical, social, and attitudinal environment influenced their participation in everyday activities and social roles (Colver et al., 2011; Fauconnier et al., 2009; Michelsen et al., 2009); the variation in adolescent participation between regions suggests environment is also an important influence on adolescent participation in several domains. Implications for clinical practice ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Important modifiable predictors of participation of adolescents with CP are pain, psychological problems and parenting stress in childhood, which are highly prevalent in families with a child with CP. The consistency of the associations we have found in this longitudinal study and their correspondence to clinical expectation suggest that these factors have a causal role. These childhood factors influence adolescent participation largely via their influence on childhood participation, which strongly predicts adolescent participation, and to a lesser extent via their influence on the corresponding adolescent factors. These findings highlight the importance of improving childhood participation in order to improve adolescent participation. Clinicians will want to intervene early to reduce pain, parenting stress and psychological problems, not only because these factors are intrinsically distressing but because of their likely effect on child participation. Pain management and the psychological aspect of pain in children with CP may be addressed by working on coping strategies (Jensen, Engel, & Schwartz, 2006). Psychological problems experienced by the child can be addressed (Beale, 2006) and interventions which target the family as a whole may improve the emotional and psychological symptoms of children with chronic conditions (Barlow & Ellard, 2004). Parenting stress can also be addressed directly (Barakat & Linney, 1992; Frey, Greenberg, & Fewell, 1989; Hinojosa & Anderson, 1991). Ideally, multidisciplinary care should start early in childhood and continue into adolescence. Implications for research ~~~~~~~~~~~~~~~~~~~~~~~~~ The most reliable way to assess whether the identified associations represent causal mechanisms would be to undertake randomised controlled trials of interventions to reduce child pain, child psychological problems and parenting stress, with long-term follow-up and measurement of participation as a secondary outcome. Trials of interventions aiming directly to improve participation are also needed as. Additionally, further follow-up of the same cohort would enable assessment of the long-term effects of childhood and adolescent precursors on adult participation. Ideally, future observational studies should have a sufficiently large sample size (e.g. over 1000 participants) to allow the use of person-centred analytic methods that identify different patterns of participation and their predictors (Bartko & Eccles, 2003; Shanahan & Flaherty, 2001).","SPARCLE 1 (visits in childhood) was funded by the European Union Research Framework 5 Program – Grant number QLG5-CT-2002-00636, the German Ministry of Health GRR-58640-2/14 and the German Foundation for the Disabled Child. SPARCLE 2 (visits in adolescence) was funded by: Wellcome Trust WT 086315 A1A (UK & Ireland); Medical Faculty of the University of Lübeck E40-2009 and E26-2010 (Germany); CNSA, INSERM, MiRe – DREES, IRESP (France); Ludvig and Sara Elsass Foundation, The Spastics Society and Vanforefonden (Denmark); Cooperativa Sociale “Gli Anni in Tasca” and Fondazione Carivit, Viterbo (Italy); Goteborg University – Riksforbundet for Rorelsehindrade Barn och Ungdomar and the Folke Bernadotte Foundation (Sweden). None of the above funders had any say in the design and conduct of the study; collection, management, analysis, and interpretation of the data; and preparation, review, or approval of the manuscript.","All authors declare that they have no conflicts of interest, including financial, personal or other relationships that could be perceived to influence this paper."],["In a seminal study, Yoon, Johnson and Csibra [PNAS, 105, 36 (2008)] showed that nine-month-old infants retained qualitatively different information about novel objects in communicative and non-communicative contexts. In a communicative context, the infants encoded the identity of novel objects at the expense of encoding their location, which was preferentially retained in non-communicative contexts. This result had not yet been replicated. Here we attempted two replications, while also including a measure of eye-tracking to obtain more detail of infants’ attention allocation during stimulus presentation. Experiment 1 was designed following the methods described in the original paper. After discussion with one of the original authors, some key changes were made to the methodology in Experiment 2. Neither experiment replicated the results of the original study, with Bayes Factor Analysis suggesting moderate support for the null hypothesis. Both experiments found differential attention allocation in communicative and non-communicative contexts, with more looking to the face in communicative than non-communicative contexts, and more looking to the hand in non-communicative than communicative contexts. High and low level accounts of these attentional differences are discussed. --------------------------------------------------------------------------------","Humans are expert learners. We learn implicitly, through mechanisms like statistical learning (Fiser & Aslin, 2001; Kirkham, Slemmer, & Johnson, 2002; Newport & Aslin, 2004), and explicitly from others through social learning (Csibra & Gergely, 2009, 2011; Tomasello, Carpenter, Call, Behne, & Moll, 2005). Social learning can occur either through observation (Meltzoff, 1988a, 1988b; Meltzoff & Moore, 1989), or through pedagogy, or explicit teaching (Csibra & Gergely, 2009; Csibra, 2007; Tomasello et al., 2005). Although teaching usually involves language, knowledge transfer can also occur in its absence. Two types of communicative cues have been suggested to be key to information transmission through teaching: ostensive cues such as direct eye contact or infant directed speech (IDS) convey the intention of communication, and referential cues such as pointing or gaze shifts direct attention to the source of the information to be learned (Csibra & Gergely, 2009; Csibra, 2010). According to the Natural Pedagogy theory (Csibra & Gergely, 2009) ostensive communicative cues signal to infants when to learn culturally relevant kind- generalizable information about an object. In the presence of these cues, infants would be biased to encode surface features, which support learning about object kinds, over spatio- temporal information. One method to investigate how infants encode object properties is the violation of expectation paradigm (VoE). The VoE paradigm is based on the assumption that infants look longer at events that violate their expectations (Onishi & Baillargeon, 2005; Teglas et al., 2011; Woodward, 1998), including when features of an object change (Krøjgaard, 2009; Mareschal & Johnson, 2003). Yoon, Johnson, and Csibra (2008) used a VoE paradigm to test the hypothesis that being communicated to should bias infants to encode surface features. In their study, the authors presented the infants with videos of communicative and non-communicative scenarios. The communicative videos included IDS, direct eye contact, and pointing, whereas the non-communicative videos included adult- directed speech (ADS), no direct eye contact, and reaching. In communicative scenes, an actress said ‘Hey baby!’ in IDS, while engaging in direct eye contact, and then pointed towards a novel object out of reach on the left or right side of the scene. In non- communicative scenes, an actress said ‘What’s this?’ in ADS, while looking at the object, and then reached towards the object. Screens then occluded the object and actress. After a few seconds, the occluders opened to reveal the object again. At the point of reveal, either the identity or location of the object had changed, or no change occurred. Infants looked longer at the identity change in the communicative condition and at the location change in the non-communicative condition (both in terms of first look and total look length). The authors concluded that infants encoded the identity of the object after being communicated to, as this was relevant to kind-generalizable learning. In contrast, infants encoded the location of the object in the non-communicative condition, due to this being the default, or perhaps the attempted reach enhancing the perceived graspability of the object. This double dissociation in the encoding of identity and location information suggested that the communicative cues did not merely increase overall memory, but elicited a specific memory bias towards identity information. There have been few papers so far attempting to replicate or extend this finding. Okumura, Kobayashi, and Itakura (2016) found that in a live study, infants showed an identity bias in a direct gaze condition. However, the authors did not replicate the finding of a location bias for the condition with no direct gaze, instead finding encoding of both identity and location. The authors suggest that, due to the video deficit effect (Anderson & Pempek, 2005), infants may have performed better in their study than in the original study, therefore managing to encode both identity and location in the non-communicative condition. Their findings suggest that instead of identity being preferentially encoded by infants after viewing communicative scenes, it may be that infants are able to encode both spatiotemporal and recognition- relevant features, but ostensive signals disrupt location encoding. Two studies following up on this result in adults drew the same conclusions (Marno, Davelaar, & Csibra, 2014, 2016), finding encoding of both identity and location information in the non-communicative condition, and only encoding of identity in the communicative condition. As there are only two studies investigating communicatively induced memory biases in infants, we felt it was necessary to replicate the original finding before extending this research ourselves. After being sent example videos from one of the original authors we noticed that in the communicative videos, the actress pointed for longer than she reached (6.8 s compared to 4.5 s). This difference raises the possibility that a longer duration of having the hand on screen could be responsible for inducing an identity memory bias. Therefore, in our stimuli both types of actions were performed twice for the same duration. Also, in the original study, the actress had bars stopping her from being able to reach the object. We did not use bars in our stimuli, instead having the actress be too far away from the object to reach it, in order to conceptually replicate the idea that the actress was unable to reach the objects, but without obscuring her. Lastly, we used eye tracking to investigate infants’ distribution of attention while observing the communicative or non- communicative scenes. Although Yoon et al. (2008) compared overall looking to the action scenes and reported no difference between conditions, this does not speak to where the infants were looking while viewing these scenes. A difference in attention allocation in the two contexts could be responsible for the memory biases observed. For example, as preverbal infants tend to follow referential cues only when these are preceded by ostensive cues (Senju & Csibra, 2008; but see Gredebäck, Astor, & Fawcett, 2018; Szufnarowska, Rohlfing, Fawcett, & Gredebäck, 2014), we thought that perhaps in the communicative context infants would look more directly towards the objects, and that this could enhance the encoding of features. As this part of the study was exploratory, we did not pre-specify specific hypotheses relating to looking during the pre-occlusion section of the videos. We report two replication attempts. The first replication, Experiment 1 (bar the changes outlined above), followed the methodological details reported in Yoon et al. (2008). In the second replication, Experiment 2, we made slight changes to the method and exclusion criteria following guidelines provided by Csibra (personal communication, November 2017). Both replication attempts were pre-registered, and all data, code, supplementary results and materials are openly available on the Open Science Framework (OSF) (https://osf.io/77gpt/).","Experiment 1 was conducted following the methods section of Yoon et al. (2008), bar some changes outlined in the introduction. Participants were recruited from a database of families at an infancy lab of a UK university, and were given a book as a gift for participation and £10 travel reimbursement. Parents gave informed, written consent before participation, and were free to withdraw their consent. All data were kept confidential. Both experiments were approved by the university ethics committee and adhered to the British Psychological Society guidelines. Replication Forty-two normally developing 9-month-old infants took part in the experiment. Of these twenty-four were included in the replication analysis (mean age: 274 days; range: 258 days to 286 days; 9 female; 23 Caucasian; 20 monolingual English). Exclusion criteria matched those used in Yoon et al. (2008): Infants were excluded for ceiling looking time for all trials (n = 4), experimenter error (n = 4), fussiness (n = 7), and not looking during one or more occlusion events (n = 3). An occlusion event in this experiment was defined as the time between the first frame of the clip when the occluder starts closing, and the first frame when the occluder is fully closed. Scene analysis We were able to use less stringent exclusion criteria for this analysis as our exclusion criterion of watching the entire occlusion event was irrelevant when looking at infant looking before the occlusion. Of the forty-two infants who took part in the experiment, 40 were included in the scene analysis (mean age: 276 days; range: 258 days to 322 days; 17 female; 39 Caucasian; 20 monolingual English). Two infants were excluded due to experimenter error.","We created the video stimuli and digitally added objects (on the left or right side of screen) and occluders on each clip. Infants first saw two familiarization trials. These familiarization videos were 29 s long and consisted of the actress moving around slightly to upbeat music, while looking either at the infant with direct gaze and smiling (communicative condition), or at the object with intrigue (non-communicative condition) (6 s). After this, yellow screens occluded the object and actress (3 s), there was a short break where the occluders stayed closed (5 s), the object screens revealed the object again (with no change) (2 s), and the object remained on screen for a maximum of 15 s (less if the infant looked away for two seconds, in which case the next trial was advanced). We always used the same two objects, but counterbalanced across participants for actress, side and order. In the test videos (Fig. 1), there was first an introductory sentence produced by the actress (“Hey baby!” with direct gaze for the communicative videos, and “What’s that?” with gaze to the object for the non-communicative videos) (4 s). This was followed by the action being executed once towards an object to the front of the actress on the left or right side (a point or a reach towards the object as if to try and grasp it) (4 s). The actions were completed at an equal distance away from the object, and were of the same duration. After completing the action once, the actress returned to the resting position and said either “Wow” (with direct gaze) while waving at the infant (communicative) or “Hmm” (without direct gaze) with a hand on her chin (non-communicative) (4 s). After this, she produced the pointing or reaching action again (4 s). Following this, screens moved to occlude both the object and the actress (2 s). The occluders stayed on the screen (5 s), after which the object screens reopened (2 s) to reveal the object. There had either been no change to the object, or it had changed in either identity or in location. The objects stayed on screen for a maximum of 15 s, or until infants looked away for 2 s. Test videos were 37 s long. Occluders produced sounds when opening and closing to direct infant attention to these events. We obtained the objects from the Noun database (Horst & Hout, 2015), and chose 6 pairs with medium similarity ratings. We used the first of each of these pairs as the initial object, and the paired objects for the identity change condition (i.e., when the object changed identity, the second object in the pair was what it changed into). Every infant saw all 6 of the objects at test. Which object was shown for which condition combination was pseudo-randomised into 8 trial orders (with 3 infants viewing each order). Which actress played which role (communicative or non-communicative) was counterbalanced across infants. Object, side of action, condition and outcome were pseudo-randomised in the 8 possible trial orders. Additionally, infants never saw more than two trials similar on any factor (e.g. no more than two left actions) in a row. All videos are openly available on the OSF.","Infants sat on their caregiver’s lap during the experiment in front of a 23-inch screen (seated approximately 0.6 m away). An eye-tracker (Tobii X120) captured infant looking times and locations on screen for the action scene analysis. We used Tobii studio 3.3.1 to present stimuli and gather eye-tracking data. A camera placed above the screen fed input directly into Tobii studio, from which videos of infants were later exported for offline looking time coding. We performed a 9-point calibration for all infants before beginning the experiment. After this calibration, we instructed parents not to talk to or interact with their infant, and the experiment began. Infants saw 2 familiarization videos (1 communicative, 1 non-communicative) followed by 6 test videos (for each of the communicative or non-comunicative familiarization video there was one no change, one identity change, and one location change test video). Therefore, each infant contributed one trial per sub-condition to the replication analysis, and three trials per condition to the scene analysis. Replication We performed all analyses on raw data in R (R Core Team, 2017). All code is openly available on the OSF. Videos of infants were exported from Tobii studio and infant looking was blind coded in ELAN (2019)(Version 5.0.0). A second blind coder performed secondary coding on 20% of videos. We found a high degree of reliability between coders: the average measure ICC was 0.97 with a 95% confidence interval from 0.95 to 0.99. The length of the first look was defined as the time between the infant looking towards the screen and their first look away, beginning at the first frame of the occluders opening. The total looking time was defined as the cumulative length of time of all looks towards the screen, beginning at the first frame of the occluders opening, and ending at the first frame of the next attention getter. We replicated the analysis by Yoon et al. (2008). We carried out a 2 × 3 repeated ANOVA on the length of first look to screen after object reveal as a function of action (communicative vs. non-communicative) and outcome (identity change, location change, no change), with follow up one-way ANOVAs and parametric and non- parametric pairwise tests. The same analyses were conducted for total looking times. Scene analysis We performed all analyses on raw data in R. All code is openly available on the OSF. We exported raw data from Tobii Studio 3.3.1, and visualized and analyzed them in R using the eyetrackingR package (Dink & Ferguson, 2016). First, Areas of Interest (AOIs) were created for face (630 × 310 pixels), hand (385 × 340 pixels), and object (380 × 340 pixels) areas of the videos. The data were then visualized as a timecourse. We performed a bootstrapped cluster based permutation analysis (Maris & Oostenveld, 2007) to establish during which time-points conditions differed significantly. This involved running a test on each time bin (17 ms) that quantified a significant difference between actions (communicative vs. non-communicative). We grouped into clusters the adjacent bins that showed a significant difference. We then shuffled the data and performed this same test on one thousand iterations of the shuffled data. This produced a table of the probability of each cluster appearing under the null hypothesis. Clusters that had a probability of less than 5% of appearing under the null hypothesis (i.e., p < 0.05) were considered to be significant. This test accounts for both Type 1 and Type 2 errors, by controlling the false-alarm rate while sacrificing little sensitivity (Maris & Oostenveld, 2007). Replication: first look length We compared the length of first look in a 2 × 3 ANOVA with action (communicative vs. non-communicative) and outcome (location change, identity change, no change) as within-subject factors (Fig. 2). There was a significant main effect of outcome [F(2,46) = 4.15, p = 0.02, ηp2 = 0.15]. There was no significant main effect of action [F(1,23) = 0.2, p = 0.66] or significant interaction between action and outcome [F(2,46) = 1.32, p = 0.28]. Paired comparisons by parametric (Student’s t) and nonparametric (Wilcoxon rank-sum) tests were carried out to assess the effect of each outcome on the looking times. Infants looked longer to the identity change than the no change outcome, regardless of communicative context. Looking was significantly longer for the identity change than the no change outcome [t(23) = 2.79, p = 0.008, ηp2 = 0.26; Wilcoxon’s Z = -2.45, p = 0.01], whereas looking time for the identity change and location change [t(23) = 1.47, p = 0.15] and location change and no change [t(23) = 0.95, p = 0.35] outcomes did not differ significantly. A nonparametric McNemar test found that this difference could be attributed to some of the infants rather than the entire group. Six of the 24 infants showed the behaviour reported in Yoon et al.’s study (longer looking at identity change than no-change after communicative scenes and longer looking at location change than no-change after non-communicative scenes), 9 infants showed only the identity bias (longer looking at identity change than no-change after communicative scenes), 5 infants only showed the location bias (longer looking at location change than no change after non-communicative scenes), and 4 infants showed the opposite pattern in both contexts (longer looking at the identity change than no-change after non-communicative scenes and longer looking at location change than no change after communicative scenes). This distribution was not significantly different from chance (McNemar’s p = 0.42). Replication: total look length Total looking length was compared in a 2 × 3 ANOVA with action (communicative or non-communicative) and outcome (location change, identity change, no change) as within-subject factors (Fig. 3). There was no main effect of outcome [F(2,46) = 0.66, p = 0.52] or action [F(1,23) = 0.01, p = 0.93], showing that overall infants did not look longer at test in either the communicative or non- communicative condition. There was a significant interaction between action and outcome [F(2,46) = 3.5, p = 0.04, ηp2 = 0.13]. We carried out separate 3-level one-way repeated-measures ANOVAs followed by paired comparisons by parametric (Student’s t) and nonparametric (Wilcoxon rank-sum) tests to assess the effect of each outcome on the looking times. In communicative context trials, there was no difference in looking time between the three outcomes [F(2,46) = 1.06, p = 0.35]. In non-communicative context trials, there was a main effect of outcome [F(2,46) = 3.91, p = 0.03, ηp2 = 0.15]. Further analyses revealed significantly shorter looking for location change compared to both identity change [t(23) = 2.36 p = 0.03, ηp2 = 0.19; Wilcoxon’s Z = -2.34, p = 0.02] and no change [t(23) = 2.21 p = 0.04, ηp2 = 0.18; Wilcoxon’s Z = -1.76, p = 0.08] (note: when using a non- parametric test the difference between looking time to location change compared to no change was not significant). There was no difference between identity change and no change outcomes [t(23) = 0.61, p = 0.55]. Scene analysis Figs. 3, 4 and S5 (see Supplementary materials on the OSF) show proportion looking to the face, hand, and object AOIs, respectively, for communicative and non- communicative scenes. Overall, infants showed similar looking patterns when viewing both types of scenes (looking towards the face the most, especially when the actress was speaking, looking towards the hand when the point/reach was being performed, and very little looking towards the object at any time point). A bootstrapped cluster-based permutation analysis found that during both of the periods where the hand action (point/reach) was not being performed (0–3000 ms, 6000–9000 ms), looking towards the face was significantly higher in the communicative condition (p = 0.027 and p = 0.048 respectively), and conversely, during the first time the action (point/reach) was performed (3000–6000 ms), looking towards the hand was significantly higher in the non-communicative condition (p = 0.024) (Fig. 4). Looking towards the object was very low overall and did not differ significantly between conditions. Overall looking time towards the scene did not differ between conditions [t(39) = 0.03, p = 0.98]. Replication Our results do not replicate the finding by Yoon et al. (2008) that infants show a memory bias for identity information after viewing communicative scenes, and a memory bias for location information after viewing non-communicative scenes. Neither do we find that communicative scenes disrupt the encoding of location information, which is preserved when viewing non-communicative scenes (Marno, Davelaar, & Csibra, 2016; Marno et al., 2014; Okumura et al., 2016). Instead, our results suggest that infants show longer looking to identity changes regardless of communicative context (measured by length of first look). However, we found that a minority of infants drove this effect, and that we do not find the results in the same direction for total look, weakening our belief that this truly indicates an identity memory bias. If we do interpret this result as an identity memory bias, this is surprising, given that the default for preverbal infants seems to be to encode location over surface features (Carey & Xu, 2001; Haun, Call, Janzen, & Levinson, 2006; Mareschal & Johnson, 2003; Xu & Carey, 1996). After this first replication attempt, we contacted Csibra, who generously provided detailed comments that led to some key methodological changes in Experiment 2. Scene analysis Our results showed that there are differences in where infants allocate their attention when viewing communicative and non-communicative scenes. Infants looked more towards the face when direct gaze and infant-directed speech were displayed, and more towards the hand when a reach was performed than when a point was performed. This shows that even before infants are actively communicating themselves, they are responding differently when they are being communicated to, compared to when they are not. As both communicative and non-communicative scenes involved speech and the same sequence of actions, we can reasonably assume that differences are due to more specific features of the two types of scene. This experiment cannot clarify whether these differences are due to low-level perceptual differences or a higher-level understanding of being communicated to. We found that infants looked more towards the face when direct gaze and infant- directed speech were used. However, in these scenes, this was also when the actress was waving to the infant. We could hypothesize that merely the movement of waving the hand is more salient than the moving of the head to look at the object in the non-communicative scenes, purely because there is more motion involved. Alternatively, consistent with Natural Pedagogy account (Csibra & Gergely, 2009), infants might look more towards the face when ostensive signals are present because they are prepared to learn from the interlocutor. Again, our results cannot differentiate between these two interpretations. We also found that infants looked more towards the hand when it was reaching than when it was pointing. This difference in looking occurred at a different time-point to where we saw differences in looking towards the face, suggesting that this is not merely the other side of the coin (i.e. when infants are not looking at the face they are instead looking at the hand), and is in fact a different process at play. One low- level interpretation for this result could be that the hand occupies more space when it is a reaching hand than when it is a pointing hand, which could draw the infants’ attention. Also, the reaching hand moves around a little, to show that the actress is unsuccessfully trying to reach the object, whereas the pointing hand does not move. Like the waving in the communicative videos, enhanced hand looking in this case could simply be the product of motion drawing the infants’ attention. Alternatively, as the goal for the reach is to grasp the object, infants might fixate on the hand in order to see what happens next (i.e. whether the person manages to reach the object), whereas in the communicative condition, the goal of the point (to communicate) has been reached as soon as the infant perceives and understands it themselves. Again, these data do not speak to which of these interpretations is more likely. Our question was not the mechanisms behind any attentional differences, but instead whether attentional differences could be driving memory biases, and so further research should investigate the reasons for these differences. However, as we find an identity bias regardless of action condition, these differences in where infants allocate their attention cannot be responsible for any memory biases in our experiment. It is also worth noting that infants are not looking at the object in either context, suggesting that they are not fully understanding the goal, as in both cases the goal is either to share attention about an object, or to reach the object. We know that 12-month-olds anticipate the goal of reaching actions by looking at the object, but 6-month-olds do not (Falck-Ytter, Gredebäck, & von Hofsten, 2006), so perhaps the infants in our experiment are too young to fully comprehend these goal directed actions (but see Kanakogi & Itakura, 2011).","After communication with Gergely Csibra (personal communication, November 2017), we were made aware of some important methodological differences between our replication attempt and the original study. Most importantly, we found that our interpretation of ‘occlusion event’ was not the same as theirs. We had interpreted the ‘occlusion event’ as the time where the occluders were currently closing, whereas they had meant it to mean the time where the occluders were closing, plus the entire period where the object is currently occluded. As the original exclusion criteria specified that infants should be excluded if they had not watched all of the occlusion events without looking away, this meant that we had essentially used a different exclusion criterion. When we went back to our data to check which infants would still be included with this new, very strict criterion, we found that none of them would be. Further discussion with Csibra made it apparent that in Yoon et al. (2008) the occlusion time had been wrongly reported as five seconds, when in fact, in the original stimuli this was actually only three seconds. This longer occlusion time, and the fact that we did not include any music during the occlusion period, may have been responsible for none of our infants showing continuous looking at the screen during the whole 5 s occlusion event for all six trials. Therefore, in the second replication attempt, we shortened the occlusion time to 3 s, added music to the occlusion period, and changed our exclusion criterion to match that of the original paper. Additionally, this second replication could serve as a confirmation for our action scene findings, as this analysis was exploratory in the first attempt.","Methods remained overall the same as in Experiment 1, but with the changes to occlusion duration and exclusion criteria discussed with Csibra. Replication Seventy-nine typically developing 9-month-old infants took part in the experiment, and of these twenty-four were included in the replication analysis (mean age: 273 days; range: 261 days to 287 days; 13 female; 22 Caucasian; 23 monolingual English). Infants were excluded for falling asleep (n = 1), ceiling looking time for all trials (n = 7), fussiness (n = 14), experimenter error (n = 1), and not looking during one or more occlusion events (n = 32). An occlusion event in this experiment was defined as the time between the first frame of the occluder beginning to close and the first frame of the occluder being fully open (this was a crucial difference to Experiment 1). Scene analysis Of the seventy-nine infants who took part in the experiment, sixty were included in the scene analysis (mean age: 272 days; range: 260 days to 289 days; 31 female; 55 Caucasian; 57 monolingual English). Infants were excluded for falling asleep (n = 1), fussiness (n = 11), less than 40% good eye- tracking data (n = 3), and experimenter error (n = 1).","Stimuli were the same as in Experiment 1, except for three changes. First (outlined above), the objects were occluded for three seconds instead of five. Second, there was a larger gap (roughly 5 times wider) between occluders when objects were fully occluded. This was to ensure infants could see that the object had not moved from one side to the other (and thus, a location change would be genuinely surprising). In order to make such gap larger, the third change was that the objects were made slightly smaller, and moved further apart, which also made the object spacing more comparable to the original study.","The procedure was identical to that of Experiment 1. Replication The analyses were identical to those of Experiment 1. Secondary coding was performed on 20% of videos by a second blind coder. A high degree of reliability was found between coders. The average measure ICC was 0.95 with a 95% confidence interval from 0.90 to 0.98. Scene analysis The analysis was identical to that of Experiment 1. Replication: first look length The 2 × 3 ANOVA revealed no significant main effect of outcome [F(2,46) = 0.13, p = 0.88] or action [F(1,23) = 0.62, p = 0.44] or significant interaction between outcome and action [F(2,46) = 1.32, p = 0.28] (Fig. S4) (Fig. 5). No statistical inference can be derived from this non-significant result (Lakens, McLatchie, Isager, Scheel, & Dienes, preprint). To determine whether the current data provides evidence for the null hypothesis (H0) relative to the alternative hypothesis (H1), a Bayes factor was conducted. Bayes factors (BF01) provide a measure of how likely the data are assuming H0 is true relative to how likely the data are assuming H1 is true. For the current analyses, a default Bayes factor with a wide cauchy distribution (scale of effect = 0.707) was calculated using the BayesFactor R package (Morey & Rouder, 2015), and yielded BF01 = 4.12. Thus, we can conclude that the data constitutes moderate evidence for the null hypothesis. Replication: total look length The 2 × 3 ANOVA revealed a significant main effect of action [F(1,23) = 4.33, p = 0.05, ηp2 = 0.16], with infants looking longer to all outcome conditions after viewing the communicative videos. There was no main effect of outcome [F(2,46) = 0.09, p = 0.92], or interaction between action and outcome [F(2,46) = 0.01, p = 0.99] (Fig. S5 on OSF). Scene analysis Our findings from Experiment 1 for the scene analysis were exactly replicated (Fig. S6–S8 on OSF). During both of the periods where the hand action (point/reach) was not being performed, looking towards the face was significantly higher in the communicative condition (p = 0.002 and p = 0.025), and conversely, during the first time the action (point/reach) was performed, looking towards the hand was significantly higher in the non-communicative condition (p = 0.038). There were no significant differences in proportion looking to the object. Overall looking time towards the scene did not differ between conditions [t(59) = 0.86, p = 0.39].","Our results show no evidence for memory biases, instead showing overall increased attention to the screen at test after viewing communicative videos. Our scene analysis results completely replicate findings from Experiment 1, suggesting that the attention allocation differences found are reliable.","We have described two attempts to replicate the results of Yoon et al. (2008). In the original study, infants looked longer at an identity change following a communicative context, and longer at a location change following a non-communicative context. These results were interpreted as a preferential encoding of object identity in a communicative context. In our first replication attempt, we found longer looking at identity changes regardless of context. In our second attempt (which, following communication with Csibra (personal communication November 2017), was better matched to the original study in terms of stimuli and exclusion criteria), we found no memory biases at all, and instead just higher increased overall attention to the screen after viewing communicative videos compared to non-communicative videos. It is important to note the discrepancy in the results from Experiment 1 and Experiment 2. If we assume that the changes in stimuli and exclusion criteria in the two experiments did not have a meaningful impact on infant looking, these results may be due to random fluctuations due to small sample sizes, as when the two Experiments are combined, we get moderate to high support for the null hypothesis (see supplementary materials). Alternatively, if we assume that there were key differences in the stimuli for the two experiments, we should compare results from Experiment 2 to those found by Yoon et al. (2008), as these are better matched. These results replicate neither the original study showing a double dissociation of identity and location memory biases, nor do they replicate other findings of impaired location memory in communicative contexts (Marno et al., 2014, 2016; Okumura et al., 2016) (Table 1) or even any memory effects in general (Blaser & Káldy, 2010; Kibbe & Leslie, 2011). It may be that the method itself is insensitive to measuring infant object memory in this specific paradigm. Our results could differ from those found by Yoon et al. (2008) because of small differences in our stimuli such as the size of the faces, the use of superimposed objects as opposed to real ones, or the distance between the two objects (video examples from Yoon et al. (2008) are available on the OSF for comparison). It is also possible that some of these differences had an effect on the higher exclusion rate in our study compared to the original (70% vs. 57%) due to our videos being potentially less engaging. We did also purposefully make the methodological changes of the absence of the bars, and of the matching of the duration of actions in both conditions. However, we have no reason to assume that any of these factors would affect the postulated creation of memory biases through the presence or absence of communication. Despite this, it is still impossible to know whether these (or other) changes could be responsible for the difference in results. Regardless, these small changes are not accounted for in current theory, which would predict that with our setup we would find the same result as Yoon et al. (2008). The original study is a key piece of evidence for the claim of Natural Pedagogy that ostensive signals not only enhance attention, but also specifically induce an expectation to learn kind-generalizable information. If small differences to a paradigm can disrupt this, then this might question the generality of the claims made by this theory, as the changes that we made should not affect the hypothesized mechanism. It is possible that the current study is a Type 2 error. This seems unlikely, given that we didn’t replicate in either of the two attempts, and find moderate support for the null hypothesis using Bayes Factor Analysis (and moderate to high support for the null hypothesis if we combine the results from Experiments 1 and 2 – see supplementary materials on the OSF). However, it does remain possible that the true effect size is smaller than that observed by Yoon et al. (2008), and that a higher-powered study is needed in order to find an effect. Further studies on communicatively induced memory biases in infants could also investigate small methodological changes, in order to see under what specific scenarios the original finding holds and advance theory about what information infants encode in communicative and non- communicative contexts. However, due to the extremely high exclusion rate in the second experiment (70% of infants tested excluded from the replication analysis), this may be a very demanding and resource consuming challenge. As out of six experiments (Table 1) only one has shown a specific identity memory bias induced by communication (as opposed to the loss of location memory), we believe that within this paradigm, the evidence against the experimental hypothesis is stronger than evidence for it, and we must consider the possibility that the original result may be a Type 1 error. In order to further study this hypothesis in a way that doesn’t require more than 50% data loss, we suggest the development of a new paradigm. Our attention allocation results are robust, with Experiment 2 completely replicating the results from Experiment 1. We found that at certain time points infants looked more to the face in communicative contexts than in non- communicative contexts, and more to the hand when it was reaching than when it was pointing. These results could be due to low-level perceptual differences between the two types of scene (e.g. with infants allocating their attention to where there is more movement), or to a high-level mentalistic interpretation of why infants would pay more attention to these areas (awaiting communication from the face in the communicative context, and awaiting the outcome of the reach in the non-communicative context). We believe there should be caution in attributing rich interpretations to phenomena that could also be explained by lean, attention-based interpretations (Haith, 1998; Heyes, 2016; Newcombe, 2002). We originally wished to investigate infant attention allocation in order to relate it to the observed memory biases. As we do not observe any differential memory biases for different contexts, we cannot relate these attention allocation differences to memory. What we can say is that attention allocation differences do not have an impact on what information is encoded or retained in our studies. Nonetheless, we feel that, when possible, eye-tracking data should be used in looking time studies to rule out attention allocation differences driving effects."],["Background Longitudinal research into the development of prosociality contributes vitally to understanding of individual differences in psychosocial outcomes. Most of the research to date has been concerned with prosocial behaviour in typically developing young people; much less has been directed to the course of development in individuals with developmental disorders. Aims This study reports a longitudinal investigation of prosocial behaviour in young people with language impairment (LI), and compares trajectories of development to typically developing age-matched peers (AMPs). Methods and procedures Participants were followed from age 11 years to young adulthood (age 24 years). Outcomes and results Participants with LI perceived themselves as prosocial; their ratings – though lower than those for the AMPs – were well within the normal range and they remained consistently so from 11 to 24 years. Two different developmental trajectories were identified for the LI group, which were stable and differed only in level of prosociality. Approximately one third of participants with LI followed a moderate prosociality trajectory whilst the majority (71%) followed a prosocial trajectory. We found evidence of protective effects of prosociality for social outcomes in young adulthood. Conclusions and implications The findings indicate that prosociality is an area of relative strength in LI. What this paper adds? To our knowledge, this is the first study to examine developmental changes in levels of prosociality from early adolescence to young adulthood in a cohort of young people with LI. Approximately one third of participants with LI followed a moderate prosociality trajectory whilst the majority (71%) followed a prosocial trajectory. We argue that prosociality is different to other areas of functioning in LI. Prosociality appears to be an area of relative strength and can act as a protective factor in social functioning. Prosociality was associated with better community integration in young adulthood and was significantly protective against friendship difficulties for individuals with LI. This paper also raises the thought-provoking issue of potential distal effects of early identification and intensive support for LI. It is important to note that all of the participants with LI in this study had been identified as having language difficulties in childhood and had received intensive intervention for their difficulties in language units attached to mainstream schools across England. The early identification of language difficulties and the context of early, intensive language support received in educational contexts such as language units may have nurtured socialisation processes and the development of emphatic concern, which in turn influence the development of prosociality later in young adulthood. More individual differences in prosociality have been reported for other samples drawn from a variety of schools with different educational provision and levels of language support and younger age groups, such as primary school-aged children with LI. --------------------------------------------------------------------------------","Prosociality involves behaviours that are positively responsive to others’ needs and welfare. Examples include being helpful and sharing, showing kindness and consideration, cooperating with others and expressing empathy and sympathy. Why and how prosociality develops is not fully understood but theories and evidence point to a multifactorial process, involving guidance from socialisation agents (such as modelling and reinforcement by parents or teachers, learning social and moral norms), genetic heritability, and emotional and social-cognitive development (Eisenberg, Fabes, & Spinrad, 2006; Jensen, Vaish, & Schmidt, 2014). Most of the research to date has been concerned with prosocial behaviour in typically developing young people; much less has been directed to the course of development in individuals with developmental disorders. Young people with disorders are at greater risk of social exclusion and so the extent to which they do manifest prosocial behaviours is an important question, with implications for our theoretical accounts of what factors influence progress in this domain and our understanding of what influences wellbeing in those with disabilities. In the present paper, we report a longitudinal investigation of prosocial behaviour in young people with language impairment (LI), followed through adolescence into early adulthood. Prosociality: developmental change and individual differences ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Given that multiple factors bear on prosociality, it is to be expected that prosocial behaviour will be subject to both developmental changes and individual differences. Prosocial behaviours are evident from infancy (Liszkowski, Carpenter, & Tomasello, 2008; Warneken & Tomasello, 2007) but they become more elaborate − and more nuanced − with development and, at any age, some individuals exhibit them more than others (Eisenberg & Fabes, 1998). From the toddler years through early childhood, children tend to show an increase in the frequency of prosocial behaviours (Eisenberg & Fabes, 1998). Through middle childhood, the findings are more mixed, with some studies suggesting stability (Cote, Tremblay, Nagin, Zoccolillo, & Vitaro, 2002; Flynn, Ehrenreich, Beron, & Underwood, 2015) but others finding modest declines (Kokko, Tremblay, Lacourse, Nagin, & Vitaro, 2006). During adolescence, some evidence points to a gradual decline in prosocial behaviours but with a possible rebound in late adolescence/early adulthood (Carlo, Crockett, Randall, & Roesch, 2007; Kanacri, Pastorelli, Eisenberg, Zuffiano, & Caprara, 2013; Spinrad & Eisenberg, 2009). At all of these stages, the overall picture is qualified by considerations including the beneficiaries of the behaviour, normative and situational variables − and individual differences, with different groups of individuals manifesting different trajectories (Nantel-Vivier et al., 2009). Within individuals, research by Eisenberg and colleagues on developmental trajectories has revealed significant, albeit modest, rank-order consistency in prosocial behaviours over time and contexts from the preschool years to early adulthood (Eisenberg, Miller, Shell, McNalley, & Shea, 1991; Eisenberg et al., 2002). Longitudinal studies of development from adolescence to adulthood remain sparse. Three main trajectory groups have been identified: prosocial (and increasing from adolescence 16/17 years to young adulthood 22/23 years), moderate prosocial, and low prosocial; the latter two groups having stable trajectories from adolescence to early adulthood (Kanacri, Pastorelli, Zuffiano et al., 2014). In order to distinguish the three trajectories found, Kanacri et al. refer to the prosocial trajectory as “high” prosocial (in relation to what they refer to as moderate and low). However, it is important to note that the scores for the participants they refer to as “high” prosocial are close to the average of the 1–9 point scale they used. Analyses from the same research group working with a large cohort of Italian children have revealed more variability when trajectories are modelled from early adolescence (age 13 years) to young adulthood (Kanacri, Pastorelli, Eisenberg et al., 2014). Taken together, findings suggest that individuals may show some fluctuations in prosocial development from childhood to young adulthood though radical shifts (e.g., from being low prosocial to becoming prosocial) are not common. Gender differences in prosociality have been consistently observed. Generally, girls score more highly than boys on measures of prosociality (Kanacri et al., 2013) and boys are less likely to follow a high prosociality trajectory (Nantel-Vivier, Pihl, Cote, & Tremblay, 2014). Prosocial behaviours: positive and protective? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Prosocial behaviours are conducive to positive social relations. Prosocial children are more accepted and more popular among their peers (Asher & Coie, 1990; Zimmer-Gembeck, Geiger, & Crick, 2005). In adolescence, prosociality is associated with social bonding and favourable friendship qualities (Cillessen, Jiang, West, & Laszkowski, 2005; Markiewcz, Doyle, & Brendgen, 2001). Prosocial behaviour in young adulthood has been found to be associated with greater involvement in the community (Kanacri, Pastorelli, Zuffiano et al., 2014). As well as contributing to positive social relationships, there is accumulating evidence that prosocial attributes and experiences may mitigate the effects of some factors that place young people at risk of adverse outcomes. Prosocial adolescents have been reported to be less likely to manifest antisocial and delinquent behaviour (Carlo et al., 2014; Pursell, Laursen, Rubin, Booth-LaForce, & Rose-Krasnor, 2008). Participation in prosocial peer relationships appears to provide support for children who have negative experiences (such as victimisation), facilitating coping and psychosocial resilience (Griese & Buhs, 2014; Martin & Huebner, 2007). Prosociality and language abilities ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Many factors are involved in the development of prosociality, and some of these are discussed in a large research literature (Eisenberg & Fabes, 1998; Eisenberg, Cumberland, Guthrie, Murphy, & Shepard, 2005; Eisenberg, Fabes, Spinrad, & 2006). However, an ability that may contribute to initiating and managing prosocial behaviours has received scant attention: language. Relatively little research has addressed the extent to which language ability bears on prosociality in children and young people. Yet language is the primary medium through which human beings communicate. It is possible to offer help to others, to share material possessions or emotions, to show kindness and consideration, to express empathy and sympathy without using language − but the likelihood is that most of these, and other, prosocial activities will involve speaking and listening, as do most human interactions from childhood through adolescence and beyond. Within this context, individuals with language impairment (LI) are of particular interest. How do they fare in prosocial skills, if they have deficits in expressing themselves and comprehending the subtleties of others’ language? Language impairment affects approximately 7% of children at school entry (Tomblin et al., 1997). Children with LI have problems putting words together (expressive language) and/or understanding what others say to them (receptive language) in the absence of learning difficulties or sensory problems such as deafness. There has been and continues to be much debate about the diagnostic criteria and terminology to describe the difficulties experienced by children and young people with LI (Bishop, 2014; Reilly et al., 2014). There is consensus however, that although LI is characterised by language difficulties during childhood, the disorder often persist into adolescence and young adulthood. There is also consensus that LI is heterogeneous and can be associated with difficulties beyond language. For example, motor functioning (Finlay & McPhillips, 2013) and memory abilities (Lum, Conti-Ramsden, Page, & Ullman, 2012). The few studies involving prosociality in children with LI have been mainly cross-sectional in design, have involved relatively small numbers of participants, and the findings have been mixed. For example, it has been found that children with LI attending primary school are rated by their teachers as being less prosocial and more prone to withdrawal than their peers. Nonetheless, overall levels of prosociality are not in the abnormal range (Brinton, Fujiki, Montague, & Hanton, 2000; Fujiki, Brinton, Morgan, & Hart, 1999; Hart, Fujiki, Brinton, & Hart, 2004), and standard deviations suggest large individual differences (Bakopoulou & Dockrell, 2016). The one longitudinal study of prosocial behaviours in children with LI (Lindsay & Dockrell, 2012) followed 65 children from 8 to 16 years, and examined prosocial behaviours using teacher report. On average, children with LI scored within the normal range, but there were individual differences. The children in this study exhibited stable trajectories, with a rise in prosociality evident between the ages of 8–12 years. Thus, the picture emerging to date shows that individuals with LI can certainly participate prosocially though, overall, they may do so less skilfully and less successfully than children without LI. Lindsay and Dockrell’s (2012) findings indicate increases in prosocial behaviour in those with LI in late childhood, which could reflect general developmental progress and/or gradual improvements in language abilities. Nevertheless, the amount of evidence available is small and only one study has addressed longitudinal trajectories in this population. Research on the associations between level of prosociality and outcomes in individuals with LI in young adulthood, has been scant (Brinton & Fujiki, 1999; Durkin & Conti-Ramsden, 2007, 2010; Mok, Pickles, Durkin, & Conti-Ramsden, 2014). In particular, an important question remains unanswered: Does prosociality confer protection against other developmental risks in the face of LI? The present study: questions and hypotheses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In this investigation we examine longitudinal development of prosociality from early adolescence (age 11 years) to young adulthood (age 24 years) in young people with and without a history of LI. The study was motivated by three main questions: Do adolescents and young adults with LI differ in prosocial orientation to age-matched, typically developing peers (AMPs)? Do those with LI show similar developmental trajectories to typically developing youth? And is there any evidence that being prosocial provides a protective factor, associated with more positive outcomes on other measures of social and behavioural functioning? With respect to differences between the groups in overall prosociality, the limited evidence available from studies earlier in development led us to expect that, on average, the LI group’s prosocial scores should fall in the normal range but somewhat lower than those of the AMP group (Lindsay & Dockrell, 2012). This would reflect the facts that individuals with LI have greater difficulties in participating in social life, tend to be less likely to initiate interactions, and have a lower sense of independence than AMPs (Brinton & Fujiki, 1999; Brinton, Spackman, Fujiki, & Ricks, 2007; Conti-Ramsden & Durkin, 2008; Durkin & Conti-Ramsden, 2010). These handicaps present impediments (though not necessarily insuperable barriers) to positive social interactions and helpfulness. An alternative hypothesis which should be acknowledged is that, as it is possible to behave prosocially with relatively little language (as demonstrated by infants and toddlers), it could be that those with LI could have adapted to their impairments by finding other ways of demonstrating prosociality. Whether young people with LI show similar or different trajectory patterns to those of AMPs remains an empirical question. In terms of their language development, those with LI continue to develop their language skills into adolescence (Conti-Ramsden, St. Clair, Pickles, & Durkin, 2012) and they follow similar language trajectories to AMPs, but with a lag (Rice, 2004). As developmental language problems tend to impact on many other aspects of development, it could be that patterns of prosocial development in this group would be similar to those of AMPs, but with the timing of any accelerations or declines delayed. Alternatively, it is possible that being ‘out of synch’ with the communicative skills development of the majority of one’s peers puts an individual at risk of lower engagement in social activity and hence affords less opportunity to develop prosocial skills. Finally, we examined whether different trajectories of prosociality were more or less protective of behavioural and social difficulties. Specifically, we examined friendship difficulties, community integration, aggressive behaviour and rule breaking. We predicted that having higher prosocial skills should be associated with more favourable outcomes in early adulthood in both LI and AMP groups. Ethics ~~~~~~ The study reported here received ethical approval from The University of Manchester. Participants with LI Participants with LI (used throughout for ease) had a history of LI and were part of the [name of study and references removed for blind-review]. The initial cohort of 242 children, which consisted of 186 boys (77%) and 56 girls (23%), were recruited from 118 language units across England and represented a random sample of 50% of all 7-year olds attending language units for at least half of the school week. Language units are specialised classes for children who have been identified with primary language difficulties. Individuals were contacted again at ages 8 (n = 232), 11 (n = 200), 14 (n = 113), 16 (n = 139), 17 (n = 85), and 24 (n = 84). The attrition observed was partly due to funding constraints at follow-up stages of the study. The sample of participants with LI did not differ between baseline and each of the follow up stages in standard scores of: age 11(receptive language (t(240) = 0.42, p = .676), expressive language (t(229) = 1.79, p = .076), or nonverbal IQ (t(231) = −.01, p = .991)), age 16 (receptive language (t(240) = −0.865, p = .388), expressive language (t(229) = −.64, p = .521), or nonverbal IQ (t(231) = −.188, p = .851)), and age 24 (receptive language (t(240) = −1.13, p = .261), expressive language (t(229) = −.45, p = .634), or nonverbal IQ (t(231) = −.60, p = .545)). Prosociality was ascertained at ages 11, 16, and 24 years. Thus, for the current investigation, analyses were undertaken for three time points only. These are referred to as time 1 (T1), time 2 (T2), and time 3 (T3). Participants were included in the analyses if data were available at least 2 of the 3 time points. At T1 (mean age 10 years 11 months, SD 5 months) and T2 (mean age 15 years 10 months, SD 5 months), there were 130 participants (92 male and 38 female). At T3 (mean age 24 years 5 months, SD 9 months) there were 84 participants (56 male and 28 female). There were 73 LI participants who provided data at all three time points. Age-matched peers (AMP) The comparison sample consisted of 65 AMPs (38 male and 27 female) and provided data at both T2 (mean age 15 years 11 months, SD 5 months) and T3 (mean age 23 years 11 months, SD 10 months). The comparison group of peers was selected to be of similar age, similar geographical area, and similar socioeconomic background as the young people with LI. The comparison group of AMPs were of a similar age to the sample with LI at each time point (T2: M 16.4, SD 0.4 years, T3: M 24.1, SD 0.9 years). AMP participants at age 16 (T2) came from similar geographical locations as the sample with LI. AMPs came from the same schools as the participants with LI as well as additional targeted schools to ensure a similar urban versus rural geographical distribution in both groups. In addition, participants in the AMP comparison group were sampled from selected demographic areas in order to ensure comparison peers came from a broad range of socioeconomic backgrounds, similar to participants with a history of LI. The LI and the comparison groups did not differ on household income at age 16years, T2 (χ2(10, N = 145) = 9.32, p = .501) nor personal income at age 24 years, T3 (χ2(5, N = 131) = 7.38, p = .194). AMPs had no history of special educational needs or speech and language therapy provision. At T2, 124 AMPs (76 males and 48 females) were recruited. Of these, 65 AMPs continued to participate at T3. Those who continued to participate at T3 had higher receptive language abilities (t(122) = 3.91, p < .001 95% CI [4.32, 13.2]) and PIQ scores (t(122) = 3.09, p = .002 95%CI [3.04, 13.92]) than those who did not. There were, however, no differences in gender (χ2(1, N = 124) = 0.46, p = .497) or expressive language abilities (t(122) = 1.34, p = .183 95% CI [−1.71, 8.92]) between those who participated at T3 and those who did not. The psycholinguistic profiles of the participants are shown in Table 1. Language and nonverbal IQ The Recalling Sentences subtest of the Clinical Evaluation of Language Fundamentals was used to assess expressive language (CELF-R, Semel, Wiig, & Secord, 1987; CELF-IV, Semel, Wiig, & Secord, 2003). At T1, the Test for Reception of Grammar (Bishop, 1982) was used to assess receptive language. The Word Classes subtest of the CELF was used to assess receptive language at T2 (CELF-R) and T3 (CELF-IV). Nonverbal IQ was measured at T1 and T2 using the Wechsler Intelligence Scale for Children Third Edition (WISC-III UK, Wechsler, 1992) and at T3 using the Wechsler Abbreviated Scale of Intelligence (Wechsler, 1999). Prosocial behaviour The prosocial subscale of the Strengths and Difficulties Questionnaire (SDQ) (Goodman, 1997) was completed by the participants (self-report) at all three time points. The scale has good internal reliability (Goodman, Meltzer, & Bailey, 1998). The scale consists of 5 items each being coded as 0 = Not true, 1 = Somewhat true, and 2 = Certainly true. The items were: “I try to be nice to other people”, “I usually share with others”, “I am helpful if someone is hurt, upset or feeling ill”, “I am kind to younger children”, and “I often volunteer to help others”. Sum scores for the subscale range from 0 to 10 and for self-report are categorised as “Normal” (6–10), “Borderline” (5), and “Abnormal” (0–4). In a population sample of adolescents (Goodman, Lamping, & Ploubidis, 2010), the construct validity of the SDQ was shown to be at an acceptable level (factor loadings 0.56-0.76). Agreement between parent report and self-report was modest (0.34) and test-retest correlations were good (0.62) (Goodman, 2001). The internal reliability of prosocial subscale of the SDQ in the sample was good (Cronbach’s α = .71). This was comparable to the internal reliability of the subscale in population samples of young people (Cronbach’s α = 0.64–0.72, Giannakopoulos et al., 2009; Van Roy, Groholt, Heyerdahl, & Clench-Aas, 2006). The prosocial subscale is positively skewed in the general population of young people (M 8.0, SD 1.7, Meltzer, Gatward, Goodman, & Ford, 2000). This is in contrast to the other subscales of the SDQ, which measure difficulties, and so are negatively skewed (e.g. emotional difficulties M 2.8 SD 2.1). In addition, we examined stability of the prosocial subscale across time. An exploratory factor analysis was run for each of the three time points using the five items on the SDQ prosocial subscale. Inspecting the scree plots and the eigenvalues determined the number of factors. The five items loaded onto a single factor with high eigenvalues at each of the time points (T1 = 2.34, T2 = 2.04, T3 = 2.34), suggesting stability of the prosocial scores across time Friendship difficulties At T3, a Friendship Difficulty Index (FDI) was created based on the Social Emotional Functioning Interview (SEF-I, Mawhood, Howlin, & Rutter, 2000). Participants were asked questions about their perception of acquaintances (range 0–2), description of current friendships (range 0–3), and their concept of friendship (range 0–3). Scores from the 3 questions were summed to create a total score (range 0–8). Higher summed scores indicated more friendship difficulties. The reliability of FDI in the sample was very good (Cronbach’s α = .84). Community integration At T3, the Community Integration Measure (CIM, McColl, Davies, Carlson, Johnston, & Minnes, 2001) was used. The 10-item checklist (e.g., I feel like part of this community, like I belong here) were scored on a 5-point Likert scale: 1 “Always disagree”, 2 “Sometimes disagree”, 3 “Neutral”, 4 “Sometimes agree”, 5 “Always agree”. Higher summed scores represent a higher level of community integration. The reliability of the CIM in the sample was very good (Cronbach’s α = .83). Aggressive and rule breaking behaviour At T3, two subscales of the Achenbach Checklist (Achenbach, 1991) were used: Aggressive Behaviour (15 items) e.g. “I argue a lot” and Rule Breaking (14 items) e.g. “I don't feel guilty after doing something I shouldn't”. All items were scored as 0 “Not true”, 1 “Somewhat or sometimes true”, or 2 “Very true or very often”. Higher summed scores indicated more difficulties. For both the aggressive behaviour (Cronbach α = .86) and rule breaking (Cronbach α = .72) the reliability of the both subscales was good. Informed consent ~~~~~~~~~~~~~~~~ The study reported here received ethical approval from The University of Manchester Research Ethics Committee, UK. Informed consent was obtained from all individual participants included in the study. Parents or legal guardians provided informed consent for all participants up to the age of 16 years. Participants themselves were asked if they wished to take part (at all phases) and provided written informed consent at ages 16 and 24 years.","The participants were interviewed face-to-face at school or at their home on the measures described above as part of a wider battery. Interviews took place in a quiet room, wherever possible with only the participant and a trained researcher present. Standardised assessments of nonverbal and verbal skills were administered in the manner specified by the test manuals. During the interview, the items were read aloud to the participants. The items and response options were also presented visually to ensure comprehension. The authors complied with APA ethical standards in the treatment of the sample. Latent class analysis ~~~~~~~~~~~~~~~~~~~~~ All statistical analyses were conducted using Stata/SE 13.1 (StataCorp, 2013). The ‘gllamm’ (generalized linear latent and mixed models; www. gllamm.org; Rabe-Hesketh, Skrondal, & Pickles, 2004) procedure command was used to model the changes in self-report prosocial scores across time. Latent classes (or groups) of individuals with similar patterns over time (Nagin & Odgers, 2010; Pickles & Davies, 1985) were identified using ordinal logistic models. Although the scale ranged from 0 to 10, there were only a small number of individuals who scored 0 or 1 (n = 3). Therefore, a score of 0 or 1 was recoded as 2. In doing this, the scale ranged from 2 to 10. The data was treated as missing at random. The gllamm command, which was used to for the latent class analysis, makes use of Maximum Likelihood Estimation to estimate model parameters. Intercept only, linear, and quadratic models were run with an increasing number of groups. The model used for further analyses was selected using both statistical goodness-of-fit criteria and interpretability. The Akaike information criterion (AIC) and Bayesian information criterion (BIC), which penalises more complex models, were used to assess the model fit. The most parsimonious model was the one with the lowest criterion value (Pickles & Croudace, 2010). The chosen model was then used to calculate for each participant the empirical Bayes’ estimates for the posterior probability of belonging to each trajectory group, and each participant was assigned to the trajectory group with the highest posterior probability. In addition, given the developmental period examined in this study (from childhood to young adulthood) and our aim to investigate mean-level differences over time, it was deemed necessary to test for scalar invariance of the SDQ prosocial subscale. We thus re-ran the above analysis using the gllamm command, and included link option (ologit) for conditional densities. Multiple links were specified using the lv option (time). The model still yielded a 2 class solution as the best solution, which suggests scale invariance can be assumed in the interpretation of the findings. Level of prosocial functioning ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Both groups of participants reported prosocial behaviours within the normal range (clinical cut-off ≤ 4, Goodman, 1997). Mean prosocial scores for participants with LI were 8.0 (SD 2.2), 7.8 (SD 1.9) and 7.9 (SD 1.9) at T1, T2, and T3, respectively and for AMP mean scores were 8.8 (SD 1.3) and 8.6 (SD 1.5) at T2 and T3. In each group, only a minority of individuals (between 2 and 6%) reported levels of prosociality in the abnormal range at one time point. There were no individuals in either the LI or the AMP group who scored consistently low, in the abnormal range, during the timeframe studied. Prosocial scores were submitted to a 2 (Group: LI or AMP) x 2 (Time: T2 & T3) mixed ANOVA, with repeated measures on the latter factor. This analysis yielded a significant main effect of group, F(1,336) = 14.0, p <. 0001, η2 = .04, but there was no main effect of time, F(1, 336) = .02, p = .90, nor an interaction between the two, F(1, 336) = .26, p = .61. Given the main effect of group, we undertook latent class analysis for LI and AMP separately. Trajectories of prosociality from early adolescence to young adulthood ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Intercept only, linear, and quadratic models were run with increasing numbers of classes starting with 1 class. For individuals with LI, the most parsimonious model was the intercept only 2-class solution. For the AMPs, it was the intercept only 1-class solution. The model fit statistics are shown in Table 2 and the trajectories are presented in Fig. 1. To aid with the understanding of Fig. 1, mean prosocial scores are presented in Table 3, which demonstrate the stability of prosociality over time. The two distinct LI trajectory classes (for ease “trajectory” henceforth) had mean scores of 8.6 (1.4) and 6.0 (1.8) respectively. Given that the population mean for the SDQ prosocial subscale for 5–15 year olds is 8.0 (1.7) (Meltzer et al., 2000), we refer to these classes as prosocial and moderate prosociality respectively. Seventy one percent of LI participants (n = 93) were classified as following a prosocial trajectory and 29% of LI participants (n = 38) were classified as following a moderate prosociality trajectory. There was a significantly larger proportion of females in the prosocial trajectory (89.5% of females vs 63.4% of males, (χ2(1, N = 131) = 8.88, p = 003). Age-matched peers all followed a prosocial trajectory with mean scores of 8.7(SD 1.4). It is known that the number of trajectory classes identified can depend upon the number of measurement occasions available (Lindsay, Clogg, & Grego, 1991). To investigate this potential effect further, models were fitted combining the LI and AMP participants into a single sample. The results were very similar to the findings examining LI and AMP samples separately. The best fitting model was a two- group intercept only model (prosocial and moderate prosociality) with a comparable number of LI participants in both groups as found with the LI sample only models. The majority of AMP participants were classified as following a prosocial trajectory. There were only 4 AMP participants following a moderate prosociality trajectory. Outcomes at age 24 ~~~~~~~~~~~~~~~~~~ A number of one-way ANOVAs were run to investigate differences between the three prosociality groups (LI Moderate Prosociality, LI Prosocial, & AMP) for outcomes at age 24 years (see Table 4). Post hoc comparisons between the prosocial vs moderate prosociality LI groups revealed that being in the LI prosocial trajectory was significantly protective in the social domain, specifically friendship difficulties and community integration. No significant differences between the prosocial vs moderate prosociality LI groups were observed in the behavioural domains as measured by the Achenbach subscales on aggression and rule-breaking. Comparisons between LI groups with AMP revealed some significant differences in social and behavioural domains. The correlations between language, PIQ and outcomes at age 24 (T3) for study participants can be found in the Appendix A. Language and prosociality: young people with LI are prosocial ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Participants with LI perceived themselves as prosocial; their ratings were well-within the normal range and they remained consistently so from 11 to 24 years. Mean prosocial scores for the group with LI were lower than those of their AMPs but still in the positive range according to SDQ norms. A history of language difficulties does not therefore preclude prosociality. On the contrary, prosociality appears to be a distinctive feature within LI. Children with LI tend to have problems across a range of social and behavioural measures. For example, using the same instrument, the SDQ, St Clair and colleagues (St. Clair, Pickles, Durkin, & Conti-Ramsden, 2011) found longitudinal evidence of hyperactivity, conduct problems, emotional difficulties and problems with peer relations during childhood and in adolescence in young people with LI. Data from the present investigation, indicate that prosociality is, in contrast, an area of relative strength, at least from early adolescence to young adulthood (and see also Lindsay & Dockrell, 2012; for broadly compatible findings in middle adolescence). To our knowledge, this is the first study to use latent class analyses to examine age-related changes in levels of prosociality from early adolescence to young adulthood that includes a sample of young people with LI. Analyses revealed two different developmental trajectories for the LI group, which were stable and differed only in level of prosociality. Approximately one third of participants with LI in this study followed a moderate prosociality trajectory whilst the majority (71%) followed a prosocial trajectory. These findings corroborate previous longitudinal research. Kanacri and colleagues (e.g. Kanacri, Pastorelli, Eisenberg et al., 2014) found that the majority of the participants in their Italian sample were prosocial and their scores were close to the average for the scale used from age 13–21 years, albeit, this trajectory showing some quadratic variation across time. These investigators also found a low prosocial trajectory, which was not evident in this investigation. More variation in prosociality may be evident in studies like those of Kanacri and colleagues (Kanacri, Pastorelli, Eisenberg et al., 2014; Kanacri, Pastorelli, Zuffiano et al., 2014) which involved larger samples (over 500 participants). There are two important points to note. First, all of the participants with LI had been identified as having language difficulties in childhood severe enough to warrant attending a specialist educational environment and not a mainstream classroom. All participants had thus received intensive intervention for their difficulties in language units attached to mainstream schools across England. All of the participants had continued to develop their expressive and receptive language skills during early adolescence to young adulthood (Conti-Ramsden et al., 2012). The early identification of language difficulties and the context of early, intensive language support received in educational contexts such as language units may have nurtured socialisation processes and the development of emphatic concern, which in turn influence the development of prosociality. Although it is likely that language units would have varied in their educational practice for inclusion (and access to non-affected peers), language units themselves afford opportunities for fostering prosociality, for example, helping others and working together. Lindsay and Dockrell (2012), for example, found more individual differences in prosociality in their sample of children with LI drawn from a variety of schools with different educational provision in two geographical areas (one city, one rural) in the UK. They found a higher proportion of children scoring in the “abnormal” range at one of the time points they studied (between 18 and 28% of children at 10, 12 and 16 years). The primary school years also appear to be a more vulnerable developmental period for children with LI. These children tend to be rated by their teachers as being less prosocial than their peers (Brinton et al., 2000; Fujiki et al., 1999; Hart et al., 2004). Future research that spans the primary as well as secondary school years would throw light as to potential developmental changes in prosociality in children and young people with LI. Second, gender differences in prosociality confirmed previous research that prosocial behaviours are strongly associated with gender (Carlo et al., 2007; Kanacri et al., 2013; Nantel-Vivier et al., 2014). There were a significantly larger percentage of females in the prosocial trajectory as compared to the moderate prosociality trajectory. It is important to underline these findings, as it is not always the case that gender differences observed in the general population are also observed in individuals with developmental difficulties, such as LI. For example, with this same cohort Conti-Ramsden and Botting (2008) found that the usual gender difference in mental health in adolescence (where there is vulnerability for females) was not evident in young adulthood. Prosociality is different. Prosociality appears to be an area of relative strength in young people with LI and it follows the gender pattern observed in the general population. Prosociality: higher levels of prosociality are protective in young adulthood ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We found significant small to medium effects for social outcomes in young adulthood. Our data suggest that prosociality also acts as a protective factor in social functioning for young people with LI. The results indicated that a prosocial trajectory as compared to a moderate prosociality trajectory was associated with better community integration in young adulthood and was significantly protective against friendship difficulties for individuals with LI. Comparisons between individuals with LI in the prosocial trajectory and same-age peers also revealed significant differences in relation to friendship difficulties. However, these findings should be interpreted within the context that both individuals with LI in the prosocial trajectory and age-matched peers were close to floor on the measure of friendship difficulties (group friendship difficulties means ranging from 0.1for peers to 0.6 for LI on a 0–8 point scale). It should also be acknowledged that while we have identified an association between prosociality and better friendships and better community integration, the association analyses cannot determine causal relationships, nor the direction of causality. In fact, a case could be made that the causal direction is the reverse: that is, that these more favourable social circumstances nurture prosocial behaviour. Nevertheless, our data are very much in line with previous research with typical populations in studies which do point to protective effects (Carlo, Crockett, Wilkinson, & Beal, 2011; Cillessen et al., 2005; Markiewcz et al., 2001). Prosociality, nonetheless, does not provide protection for all areas of functioning. In LI, associations of prosociality with behavioural functioning were weaker and non- significant. Comparisons with same age peers revealed individuals with LI exhibited significantly more aggressive behaviours in young adulthood regardless of their level of prosociality. These data have important implications for fostering the strengths of young people with LI. Harnessing and further developing prosocial tendencies may lead to better social outcomes for young people with LI. We are not claiming that prosociality is the only factor impinging on friendships and community integration. The picture is complex and there are individual differences. For example, we know that a third of this same cohort experience problems with friendship in adolescence and young adulthood (Durkin & Conti- Ramsden, 2007; Mok et al., 2014; St. Clair et al., 2011). Nonetheless, a medium size effect size was observed between LI and AMP groups for friendships in this study, suggesting that in LI, being moderately prosocial may not be enough to confer protection, a higher “dosage” of prosociality is likely to be required. To our knowledge, there has not been a systematic effort to build on the prosocial tendencies of individuals with LI in intervention programmes. It is more common to target areas of deficits rather than strengths. A good example is intervention research in autism. There is an abundance of programmes that target improving the social skills and prosocial behaviours of children and young people with autism spectrum disorders, although the effectiveness of such interventions has been limited (Bellini, Peters, Benner, & Hopf, 2007; Greenway, 2000). In future work, it will be important to examine prosociality in young people with LI longitudinally from an earlier point in development and for research to include both intervention and observational designs. The inclusion of a broader array of measures of prosocial behaviours (e.g. experimental tasks and direct observations) is also needed. Although the SDQ prosocial scale has good reliability and has been used extensively in the literature, different measures are sensitive to different aspects of prosociality and their concurrent use may elucidate potential causal pathways to better outcomes for young people with LI.","The authors declare that they have no conflict of interest."],["We examine the social-psychological and personality bases of support for radical right parties (RRPs), using cognitive-motivational approaches of ideology. Our comprehensive model includes core ideological variables which mediate personality traits (Big Five) or how individuals engage in social relationships and accommodate novel stimuli. Structural equation models were tested in an Austrian population sample to examine support for a RRP, the FPÖ. Our results suggest that a \"perceived immigrant threat\" and, in part, social dominance orientation are directly related to RRP support, whereas right-wing authoritarianism has consistent indirect (mediated) impact. Associations with lower Openness to Experience, lower Agreeableness, and to some extent also with Conscientiousness are mediated by the ideological variables. The conclusion discusses how RRPs' success and communication strategies can be linked to basic psychological motivations. --------------------------------------------------------------------------------","In recent years, Europe has witnessed growing electoral success of radical right parties (RRPs). As a “party family” RRPs are commonly characterized by their authoritarian beliefs, the return to traditional values, anti-immigrant and xenophobic stances, i.e., preference for an ethnically homogeneous population, as well as in-group/out-group thinking that portrays the existence of threats (e.g. Rydgren, 2007). Hence, the focus on grievances concerning immigration is considered a core feature of the RRP profile (Ennser, 2012; Ivarsflaten, 2008). Meanwhile, empirical research has tried to explain why voters support RRPs (see Van der Brug & Fennema, 2007). In terms of the social structure, RRP support was found to be more prevalent among less educated, lower-income, and younger voters (e.g. Lubbers, Gijsberts, & Scheepers, 2002; Oesch, 2008; Rydgren, 2007). With regard to policies, preferences on “new” issues, such as anti-immigration policies or EU- skepticism, are known to attract many RRP voters (e.g. Aichholzer, Kritzinger, Wagner, & Zeglovits, 2014; Ivarsflaten, 2008; Van der Brug & Fennema, 2007). Yet, the evidence concerning the role of core ideological dimensions, such as “egalitarian” or “authoritarian” attitudes, is contradictory (see Cornelis & Van Hiel, 2015; Dunn, 2015; Zandonella & Zeglovits, 2012). In addition, despite a large body of literature on basic personality traits as a factor in partisan orientation, few studies have attempted to link psychological traits (e.g. Big Five) to preference for RRPs (Zandonella & Zeglovits, 2012), extreme right-wing parties (Schoen & Schumann, 2007) or populist parties more generally (Bakker, Rooduijn, & Schumacher, in press). Furthermore, a coherent theoretical framework that links social–psychological factors of ideology and personality to core RRP stances is largely missing in the literature. In the present study, we anticipate that voters gravitate toward RRPs when they: (a) exhibit authoritarian attitudes (right-wing authoritarianism, RWA), i.e., motivational goals to seek group security and stability in societal order (Altemeyer, 1981; Duckitt, 1989); (b) exhibit competitively driven motivations to maintain hierarchical or superior–inferior relations between social groups (social dominance orientation, SDO) (Pratto, Sidanius, Stallworth, & Malle, 1994); and (c) perceive social threats to identity and cohesion induced by immigration and, hence, exhibit motivations to reduce that uncertainty and threat (Jost, Glaser, Kruglanski, & Sulloway, 2003). Finally, we propose that (d) these attitudinal factors fully mediate basic personality traits (Big Five) that might predispose individuals to uphold stability in social relationships or make them less open to new social situations or stimuli (see DeYoung, Peterson, & Higgins, 2002). After specifying our hypotheses, we analyze our theoretical model by using unique representative survey data from Austria. Our outcome variable is preferences for the Freedom Party of Austria (Freiheitliche Partei Österreichs), FPÖ, one of the most successful RRPs in Europe. Radical right party support and its relation to RWA and SDO ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ It is well established that basic cognitive–motivational goals drive our ideological orientations, namely advocating vs. resisting social change and rejecting vs. accepting inequality or RWA and SDO, respectively (see Duckitt & Sibley, 2009; Jost, Federico, & Napier, 2009, for an overview). Consistent with this framework, RWA and SDO are also believed to play a role in satisfying epistemic and existential motivations, namely reducing uncertainty and threat (Jost et al., 2003). Following Altemeyer (1981), the main perceptional and behavioral consequences (or lower-level structure) of RWA are: (1) to accept and to adhere to authorities as well as to social norms (“submission”); (2) to approve of and demand the punishment of people who deny the legitimacy of these authorities or deviate from these norms (“aggression”); and (3) to be sensitive to threats to a given social order (“conventionalism”). We thus anticipate that motivational goals of RWA foster RRP support as these are vital characteristics of RRP stances. In turn, SDO expresses competitively driven motivations to maintain or establish group dominance and superiority, i.e., people high on SDO support intergroup hierarchies and tend to arrange social groups in a superior–inferior order. Thus, SDO should predict a person's acceptance or rejection of ideologies and policies relevant to group relations (Pratto et al., 1994). We therefore expect SDO to be positively related to RRP preference. Even though RWA and SDO can be moderately to strongly positively correlated (Roccato & Ricolfi, 2005), these factors represent distinct predictors of numerous sociopolitical and intergroup attitudes, especially political orientation and forms of prejudice (Duckitt & Sibley, 2009; Sibley & Duckitt, 2008). However, previous evidence suggests that SDO, rather than RWA, might be more important for party preferences or more directly related to them (Cornelis & Van Hiel, 2015). Radical right party support and perceived immigrant threat ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To explain RRP support, we further consider a “perceived immigrant threat” (PIT), i.e., individuals' perception that immigration threatens their personal or the majority's societal value system, culture, social cohesion, or alleged ethnic homogeneity. Indeed, other authors have referred to this tendency as “cultural conflict” dimension (Kriesi et al., 2008), a “normative threat” (Stenner, 2009), or “symbolic threat” (Stephan & Stephan, 2000). Previous research suggests that this type of threat seems to matter most for RRP support (Lucassen & Lubbers, 2012; Oesch, 2008), even more so than economic or “material threats” (on this distinction see Lucassen & Lubbers, 2012; Stephan & Stephan, 2000), or that these types of threat by immigrants are not empirically distinguishable (Lucassen & Lubbers, 2012). According to a core component of RRPs' discourse, their supporters seek to mitigate perceived threats linked to immigration (PIT). RWA, SDO, and social threat ~~~~~~~~~~~~~~~~~~~~~~~~~~~ In a nutshell, Duckitt and Sibley's (2009) model suggests that scoring high in RWA makes individuals more sensitive or responsive to types of social threat. Indeed, research has shown that authoritarians are more responsive to threatening messages (e.g. Lavine et al., 1999) or that (extreme) right-wing individuals show stronger psychological, but also physiological responses, to negative or threatening stimuli (e.g. Hibbing, Smith, & Alford, 2014). We thus anticipate that RWA is an important antecedent of PIT. Another main hypothesis in Duckitt and Sibley's (2009) theoretical approach is that the relationship between RWA and political behavior (e.g. party preference) is, at least partly if not fully, mediated by perceived threats (i.e., PIT). As a consequence, RWA would only have an indirect effect on RRP support. SDO, on the other hand, is expected to be connected less strongly, if at all, to societal-level threats or normative threats (Onraet, Van Hiel, Dhont, & Pattyn, 2013). Instead, it will be related to threats explicitly activating competitiveness over relative superiority and dominance of groups. We nevertheless test, but do not expect mediation of SDO on voting preference via PIT. Radical right party support and its linkage to personality ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The literature on partisan orientation and individuals' personality mainly builds on the Big Five model. Based on the extant literature, we anticipate that RRP support is mainly predicted by lower scores on Openness to Experience (Caprara, Barbaranelli, & Zimbardo, 1999; Chirumbolo & Leone, 2010; Vecchione et al., 2011), higher levels of Conscientiousness (Chirumbolo & Leone, 2010; Schoen & Schumann, 2007; though with mixed findings: Vecchione et al., 2011; Zandonella & Zeglovits, 2012), and lower scores on Agreeableness (i.e., lower trust, altruism or compliance) (Bakker et al., in press; Chirumbolo & Leone, 2010; Schoen & Schumann, 2007; Zandonella & Zeglovits, 2012). Furthermore, preliminary evidence suggests that people low in Emotional Stability might prefer RRPs over other parties (Schoen & Schumann, 2007; Zandonella & Zeglovits, 2012), whereas Extraversion seems to play a negligible role in voters' behavior (see Gerber, Huber, Doherty, & Dowling, 2011; Schoen & Schumann, 2007; Zandonella & Zeglovits, 2012). Mediation of personality traits by ideological attitudes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Relationships between personality and political preferences are likely not to be direct, but rather indirect or mediated by ideological variables. In particular, RWA is assumed to have a unique foundation in personality, including social conformity traits or a combination of Conscientiousness and lower Openness to Experience (e.g. Brown, 1965; Duckitt & Sibley, 2009). In turn, Emotional Stability should be negatively related to feelings of threat and insecurity and could thus diminish RWA. SDO seems to be primarily related to lower Agreeableness (or lower trust, altruism, or compliance), higher Conscientiousness, and lower Openness to Experience (Sibley & Duckitt, 2008). PIT or prejudice might be rooted in traits that make people less capable of adapting to new stimuli and social situations, traits that entail lower levels of altruism or tolerance in social relationships, or traits that make them more anxious (see Brown, 1965). We anticipate that Openness to Experience, Agreeableness, and Emotional Stability are negatively associated with PIT, while Conscientiousness is positively associated with PIT (e.g. Gallego & Pardos-Prado, 2014; Sibley & Duckitt, 2008). Table 1 and Fig. 1 provide a summary of our hypotheses and the underlying model for our empirical analyses. Case ~~~~ To test our hypotheses we use data from Austrian voters and support of the FPÖ. In terms of its policy profile, the FPÖ is considered comparable to other RRPs in Europe (Ennser, 2012) and voting patterns are considered similar to other RRPs (see Aichholzer et al., 2014; Lubbers et al., 2002). In the last Austrian federal election 2013 the FPÖ gained 20.5% of the valid votes.","The data used in this study were collected by the Austrian National Election Study (AUTNES) Pre- and Post-Election Survey (panel) in 2013 (Kritzinger et al., 2013, 2014). The pre-election interviews were conducted face-to-face (CAPI) and the post-election interviews by telephone (CATI). Respondents were sampled from the Austrian population eligible to vote (i.e., aged ≥ 16), using an address-based multistage stratified clustered sample and selection of a random respondent within each household (total n = 3266, pre- election response rate: 61.8%; post-election n = 1504 or 46.1% re-interviewed). The sample composition was: 49.1% male; age: M = 45.3 years, SD = 19.5 years; education: 22.1% compulsory schooling, 47.3% lower secondary/vocational training, 18.4% admission to tertiary education, 12.3% college/university degree. Data were weighted for post- stratification and sample design adjustment.","All measures except actual vote choice were administered in the pre-election wave (see Appendix A for the exact question wording). For our dependent variable, we rely on two different operationalizations: first, we use the respondents' “propensity to vote” (PTV) for the FPÖ (pre-election wave, see Appendix A for exact question wording), which is measured on a quasi-metric 11-point scale (0 = very unlikely, 10 = very likely, M = 2.75, SD = 3.17). This measure refers to the general affiliation with a party beyond the vote in a specific election. Secondly, we employ a binary measure of respondents' “actual vote” (post-election), which is coded 1 if respondents voted for the FPÖ and 0 for all other parties, no party/invalid-answers, and excluding voters of the FPÖ splinter group BZÖ (n = 28) to avoid any overlap (16.2% FPÖ). The Pearson correlation between PTV and actual vote was r = .52. Big Five personality traits were measured using the German BFI-10 (Rammstedt & John, 2007). Our analyses support the hypothesized five-factor structure of the items, although it is not regarded as perfectly “clean” (i.e., items have non-ignorable cross- loadings) (see Table A1 in the Appendix A). We measure PIT based on five items that tap into the respondents' attitudes and emotions regarding immigrant and cultural threat (vs. enrichment), since emotional reactions as well as factual attitudes explain perceptions about immigration (Brader, Valentino, & Suhay, 2008). These are indicated by questions on concerns about (Muslim) immigrants as well as on respondents' specific emotions towards immigration, namely anger and anxiety. We also adjust for common method variance in the two questions capturing emotions (see Table A1). We measure RWA with seven items, spanning its facets “submission”, “aggression”, and “conventionalism” (Altemeyer, 1981), and apply correlated uniquenesses (CUs) to capture the conceptual overlap of items in each sub- dimension (see Table A1). Finally, we measure SDO with two items that strongly resemble the wording of the original SDO scale by Pratto et al. (1994) (see Appendix A). All Likert-type items use a fully labeled 5-point rating scale (1 = agree completely to 5 = disagree completely). Analysis ~~~~~~~~ We use structural equation models (SEM) to take into account unreliability or unsystematic measurement error in survey measures when analyzing the structural relationships between variables. Secondly, we take into account systematic acquiescence bias in agree-disagree items, using a response style factor as a control (Aichholzer, 2014). Thirdly, we allow for a “complex” (unrestricted) item-factor structure in the BFI-10 scale, applying ESEM (Asparouhov & Muthén, 2009), which is then held constant across dependent variables (see Table A1 for the measurement models). Further, we intend to control for socio-demographic heterogeneity and add age, education (1 = admission to tertiary education), and gender (1 = female) as controls for all latent factors and RRP support (see Fig. 1, detailed results not presented). All analyses were conducted in Mplus (Version 7) (Muthén & Muthén, 1998-2012), using linear SEM with full information MLR or WLSMV estimation for missing data. In order to evaluate the models' fit we review the goodness-of-fit indices CFI and RMSEA [90% CI] and interpret them jointly. A combination of cutoff values CFI > .90 and RMSEA < .08 is often regarded as acceptable and CFI > .95 and RMSEA < .05 is considered to be an excellent fit.","Overall, each structural equation model displays at least an acceptable or a good fit to the data. Furthermore, all models exhibit an equally good fit when allowing for direct effects of the Big Five traits on RRP support. Thus, it is reasonable to assume full mediation of effects from these variables. In Table 2 we present direct and total (direct + indirect) effects with fully standardized coefficients (β) separately for the two dependent variables: PTV and actual vote for the FPÖ. Our first empirical test examines how basic social-psychological factors of ideology are related to RRP support. With regard to direct effects, PIT strongly (β = .51 and .42) and SDO moderately (β = .22) influences the PTV for the FPÖ (no mediation), whereas the impact of SDO disappears when the respondents' actual vote choice is considered (β = .02, n.s.). One reason for this pattern could be panel attrition effects, since preliminary analyses suggest that the probability to remain in the panel increases with higher SDO but decreases with higher RWA, whereas PTV and PIT seem to have no impact. On the other hand, RWA has virtually no significant direct impact on RRP preference when including SDO and PIT, consistent with previous evidence suggesting negligible direct effects in the presence of the other variables. Nevertheless, it is evident that the effect of RWA is mediated by a strong relationship with PIT (β = .63 and .67), resulting in a highly significant positive total effect with RRP support (total β = .32 and .50), which is higher for actual voting. Finally, 36% or 39% of the variance in our dependent variables (R2) can be explained by the explanatory variables in our model. Our second empirical test investigates the relationship of personality traits with RRP support. Since our results corroborate full mediation via the ideological variables, we provide indirect (= total) effect estimates of the Big Five on RRP support. We find that the most consistent and moderate (negative) relationships with RRP support can be found for Openness to Experience (total β = −.17 and −.18) and Agreeableness (total β = −.12 and −.14), whereas Conscientiousness displays weak and inconsistent positive relationships with RRP support (total β = .05 (n.s., p = 0.061) and .09). Extraversion did not reveal significant or consistent relationships with the ideological constructs (apart from SDO: β = .08 and −.14), neither with RRP support. In turn, (lower) Emotional Stability only plays a moderate role for explaining RWA (β = −.15 and −.14), but is irrelevant for FPÖ vote preference. Finally, our results confirm most of the existing evidence on Big Five associations with RWA, SDO, and PIT as well as on the relation between RWA and SDO (see Table 2 for detailed results). On the one hand, this lends support to the criterion validity of our measures. On the other hand, it helps us to clarify how personality traits are related to RRP support. To sum up, we find that scoring low in Openness to Experience is directly, and most clearly, related to RWA, to SDO, and indirectly to PIT, whereas Agreeableness negatively correlates with PIT and in part also with RWA, resulting in their indirect negative association with RRP support. Surprisingly, however, we do not find the hypothesized relationship between SDO and lower Agreeableness. In turn, Conscientiousness is substantially related to RWA, but not to PIT or SDO, thus being only partly related with RRP support. With regard to socio-demographics and FPÖ vote preference, education was significantly negatively, though indirectly, related to FPÖ vote preference (total β = −.21 and −.28), because of consistent negative effects on RWA, SDO, and PIT (total βs = −.21 to −.33). In turn, higher age displays a moderate negative direct effect (β = −.20 and −.18), positive relationships with RWA, SDO, and PIT (total βs = .18 to .30), but virtually no total effect. Lower preference for the FPÖ among women only plays a direct role for PTV (total β = −.07), but not for actual vote.","In this study, we provide a comprehensive picture of the interplay between psychological aspects in voters' preference for radical right parties (RRPs), including personality (Big Five), authoritarianism (RWA), social dominance orientation (SDO), and a “perceived immigrant threat” (PIT). Putting the scattered pieces of the puzzle in order, we were able to comprehensively address the individual differences and social–psychological mechanisms that drive RRP support. Here, we investigated the case of Austria with an established RRP in its party landscape, the FPÖ, using representative survey data. To summarize, most of our initial expectations on the associations between the Big Five and RRP support were in line with the empirical evidence. Most importantly, we established indirect relationships (mediation hypothesis) which help us to understand why and how personality explains RRP support. Among basic social-psychological factors of ideology, PIT most clearly increases the likelihood of RRP support, SDO has a moderate direct positive impact for PTV, whereas the association with RWA seems to be indirect. There is reason to believe, however, that different conceptualizations and measurements of “authoritarianism” explain diverging evidence on its role in RRP support. In theoretical terms, our findings largely support the notion of cognitive–motivational goals in individuals, such as managing uncertainty or threat as well as maintaining stability in societal order and intergroup hierarchies, which manifest themselves in political ideology and voting behavior (Jost et al., 2003, 2009). In other words, individuals seek to support parties or politicians that match these goals. Furthermore, our results could be interpreted in the light of basic functions of personality traits, such as engaging in social relationships through trust and compliance, on the one hand, as well as the (in)ability of enjoying and processing novel stimuli and situations, on the other (see DeYoung et al., 2002). The exact personality origins of RRP support beyond the Big Five or its relation to their more narrow facets nevertheless deserve to be studied further.","What have we learned about the nature and the success of RRPs? With our proposed model we cannot explain the electoral success of specific RRPs at specific elections. However, learning about deeply rooted patterns in voters' political preferences adds new insights to electoral research, particularly in the light of insufficient explanatory power of more classical theoretical models (see Aichholzer et al., 2014). Furthermore, this also contributes to our understanding of how appealing political elites' and parties' political communication is to voters. Our results suggest that a successful RRP offers policy stances that deliberately address basic psychological motivations and cognitions that are rooted in voters' personalities and core ideological attitudes. For instance, triggering social threat and negative emotions associated with immigrants among voters seem to be effective in RRPs' voter mobilization (see also Brader et al., 2008). Some limitations of this research must also be considered: first, our results ideally require replication. Austria may be special due to its particular cultural and historical context (e.g., WWII) when it comes to political ideology and RRPs. That said, the alignment of RWA and SDO in the ideological spectrum and with their antecedents can be contingent on the ideological contrast among parties in a country (see Roccato & Ricolfi, 2005). Second, a common limitation in large-scale representative surveys are measurement instruments. Specifically, short measures of complex traits are inherently restricted in their breadth and, hence, more extensive and cross-culturally invariant scales would be desirable. Third, future research may want to investigate factors that activate and mediate SDO, such as actual experiences of intergroup competition (e.g. labor market disadvantage) which, in turn, may foster RRP support. We also expect future research to further investigate the RWA lower-level structure, because its sub-dimensions may differ in how they relate to aspects of social threats and, hence, to vote choice. Further theoretical work and empirical tests are required to fully understand the role of psychological aspects in supporting RRPs."],["Does personality predict how people feel in different types of situations? The present research addressed this question using data from several thousand individuals who used a mood tracking smartphone application for several weeks. Results from our analyses indicated that people's momentary affect was linked to their location, and provided preliminary evidence that the relationship between state affect and location might be moderated by personality. The results highlight the importance of looking at person-situation relationships at both the trait- and state-levels and also demonstrate how smartphones can be used to collect person and situation information as people go about their everyday lives. --------------------------------------------------------------------------------","We know that personality is linked to behavior. Several studies have shown, for example, that personality is linked to preferences for and success in various occupations (Judge, Higgins, Thoresen, & Barrick, 1999; Lodi-Smith and Roberts, 2007), maintaining satisfying intimate relationships (Ozer & Benet-Martinez, 2006; Roberts, Kuncel, Shiner, Caspi, & Goldberg, 2007), and how people choose to spend their free time (Mehl, Gosling, & Pennebaker, 2006; Rentfrow, Goldberg, & Zilca, 2011). Our understanding of these links is informed by interactionist theories, which argue that individuals seek out and create environments that satisfy and reinforce their psychological needs. An implication of this argument is that individuals experience higher positive affect and lower negative affect when in their preferred environments. Drawing on past research on person-environment interactions and using experience sampling and mobile sensing technology, the present research investigated whether personality traits moderate the associations between state affect and locations, which are associated with different types of situations. Personality is linked not only to behavior, but also to affect. In one study, situational characteristics (social vs. non-social context) and personality traits (Extraversion and Neuroticism) both predicted state positive and negative affect (Pavot, Diener, & Fujita, 1990). People experienced greater positive affect both when they were high in Extraversion, and when they were in social situations, but there was no interaction: people with all levels of Extraversion experienced more positive affect in social situations. In a related study, researchers found that positive affect (but not negative affect) follows a diurnal rhythm, with people socializing, laughing and singing more and more for the first 8–10 h after waking, and then doing those activities less and less (Hasler, Mehl, Bootzin, & Vazire, 2008). Further, this study found preliminary evidence that this diurnal cycle of positive affect may be amplified for people high in Extraversion. These findings have implications for within-person variation in personality, given that affect is assumed to mediate the influence of the situation on personality states (Mischel & Shoda, 1995). Although there is agreement that features of the environment affect how people think, feel, and behave, there is less agreement on which aspects of the environment have psychological implications (Fleeson & Noftle, 2008; Rauthmann et al., 2014). The physical environment, including location, is one objective characteristic often associated with a situation (Rauthmann et al., 2014; Saucier, Bel- Bahar, & Fernandez, 2007). The locations people regularly visit may be consistently associated with a constellation of factors (e.g., affect, sociability, recreation, goal pursuit), and thus represent types of situations that have psychological implications, and clear links to personality For example, the recently developed DIAMONDS taxonomy identifies several characteristics of situations that might feasibly be linked to locations (Rauthmann et al., 2014); work might be high in Duty and Intellect, and social places, such as restaurants and bars, might be high in Sociality and pOsitivity. These situational factors are connected to personality-related behaviors (Rauthmann et al., 2014), suggesting that locations should be as well. A challenge in studying affect as it is experienced in the various types of situations that people encounter in daily life is the repeated collection of data on mood and situation type. Methodological advances such as experience sampling have made it possible to collect repeated self-reports (e.g., of affect) as people go about their daily lives. However, to date it has been difficult to simultaneously collect objective information about the type of situation in which people find themselves. One exception is a recent paper that used repeated experience sampling to examine the relationship between a person’s personality traits, the types of situations they encountered, and state expressions of personality, all of which were self-reported (Sherman, Rauthmann, Brown, Serfass, & Jones, 2015). In this study, state personality (as manifested in behavior and emotions) was independently predicted by both personality traits and situation characteristics. The advent of mobile sensing technology provides a potential solution to the challenge of collecting repeated information about both behaviors and situations: detect the type of situation using the sensors built-into today’s ubiquitous smartphones. These devices come equipped with location sensors, an accelerometer that can detect a user’s physical activity, a microphone that can detect ambient noise in the environment, and various other sensors. The potential is great, but little research to date has made use of sensed information to examine psychologically relevant questions. In the present research, we explored the relationship between state affect and location, which can be thought to represent a type of situation. Our objective was to lay a foundation of preliminary knowledge in the under-explored domain of associations between situation types and state affect.","Participants were members of the general public who downloaded the free app from the Google Play store and installed it on their Android phone. The analyses reported herein include all users who provided data on the measures of interest (described below) from February 2013, when the app was released, to July 2015, when we began the analyses. A total of 12,310 users provided relevant momentary self-reports of location (i.e., they reported being at home, at work, or in a social type of situation). Of the users who reported demographics (N = 10,889), 44% of people who reported their gender were female, 71% of people who reported their ethnicity reported being White, and the most common birth year ranges were 1980–1989 (38%), 1990–1999 (32%) and 1970–1979 (18%). Given that these users may or may not have provided trait or state self-reports of personality, and may or may not have provided location sensor data (see Section 2.2 for why this is the case), the analyses described in the results section include different subsets of these users. Ethical clearance has not been granted for sharing the data supporting this publication. The Emotion Sense application ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Emotion Sense is a smartphone application that was designed to study subjective well-being and behavior. The app collects self-report data through surveys presented on the phone via experience sampling. By default, the app sends two notifications at random moments of the day between 8 AM and 10 PM, at least 120 min apart from one another. Clicking on a notification launches a momentary assessment, which includes measures of current affect, and measures assessing a single aspect of current behavior or context (e.g., location, physical activity, social interactions). In addition to the notification-driven surveys, the app also collects self-initiated surveys. These included longer measures of affect, and measures assessing multiple aspects of behavior and context. As well as collecting self-report data, the app also uses open-sourced software libraries (Lathia, Rachuri, Mascolo & Roussos, 2013) to periodically collect behavioral and contextual data from sensors in the phone. The data collected through the app is stored on the device’s file system and then uploaded to a server when the phone is connected to a Wi-Fi hotspot. Emotion Sense was designed to be a tool to facilitate self-insight, providing feedback about how participants’ mood relates to context and activity. In an effort to maintain user engagement over a period of weeks, participants could receive additional feedback by “unlocking” stages, in the same way that players can unlock different levels of a game after achieving certain objectives. Each stage had a particular theme (e.g., location, physical activity) that determined which behavior and context questions (e.g., “Where are you right now?”, “Compared to most days, how physically active have you been today?”) were asked in the self-report surveys. The second stage, related to location, is the only stage reported in these results.1 Trait personality Users reported their personality on the Ten-Item Personality Inventory (Gosling, Rentfrow, & Swann, 2003). They rated the extent to which ten pairs of words (e.g., “Extraverted, Enthusiastic”) applied to them, on a scale from 1 = Disagree strongly to 7 = Agree strongly. Affect Emotion Sense allows users to track and quantify their psychological well-being in various ways. On each self-report survey, whether notification-driven or self- initiated, users indicate their current feelings by tapping on a two-dimensional affect grid (see Fig. 1), where the x-axis denotes valence, from negative to positive, and the y-axis denotes arousal, from sleepy to alert (Russell, Weiss, & Mendelsohn, 1989). Location One way Emotion Sense assesses current location is through place self-reports. Users respond to the question “Where are you right now?”, indicating whether they are at “Home,” “Work,” “Family/Friend’s House,” “Restaurant/Café/Pub,” “In transit,” or “Other” (see Fig. 2). We treated both “Family/Friend’s House” and “Restaurant/Café/Pub” as social types of situations, and did not examine “In Transit” or “Other” reports. These locations were expected to be ones where people spent much of their time, and visited over and over again. Location is also sensed via the phone’s location sensors. For 15 min before each survey notification and at various intervals throughout the day, Emotion Sense determines the location of the phone by collecting the latitude, longitude, accuracy, speed, and bearing of the device from its location sensors. This is a coarse measure of location, with accuracies ranging from the tens to hundreds of meters, but measuring location more accurately consumes far more battery power. A three-step process allowed us to match self-report data (e.g., state affect and state personality) to places (i.e., home, work, social). First, we mapped each geographical location to a place by connecting the location sensor data to the location self-reports. Next, we mapped each self-report (e.g., state affect and state personality) to a geographic location by connecting the self-report to the temporally closest location sensor data. Finally, we mapped these self-report geographic locations to places using the mappings created in the first step. (See Appendix 1, Section 2.4 for more details.)","We investigated whether a person’s mood is related to the location they are in, and how that relationship is moderated by trait personality. How is affect related to location? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We created two sets of dummy codes: one set with Home as the reference group (to compare mood at Home vs. at Work, and at Home vs. in a Social type of situation), and the other set with Work as the reference group (to compare mood at Work vs. in a Social type of situation). We ran hierarchical linear modelling, using the lme4 package in R (Bates, Maechler, Bolker, & Walker, 2014), predicting mood from the place dummy codes (level 1), grouped by user (level 2). This analysis was carried out for both the location self- reports, and the locations sensed by the phone sensors. We report results for grid valence, but see Appendix 1 for analyses examining grid arousal and alternate measures of high and low arousal positive affect and negative affect. Given that degrees of freedom are not provided by the lme4 package, we approximated them from the number of level 2 units (i.e., users) minus the number of predictors in the level 2 equation, not including the intercept, minus 1 (Raudenbush & Bryk, 2002). Self-reported location Of the 12,310 users who provided momentary self-reports of location, 6759 reported their mood in at least two of the three targeted locations and were used in these analyses. These users provided, on average, 23 reports of location (from 2 to 305 reports, median = 18). The results from our analyses indicated that participants experienced a more positive mood when they reported being in a social type of situation (M = 0.34, SD = 0.18) vs. at home (M = 0.20, SD = 0.31), β = 0.29, t(6,758) = 38.84, p < 0.001, and when they reported being at home vs. at work (M = 0.18, SD = 0.25), β = −0.08, t(6,758) = −14.18, p < 0.001. (NOTE: Means and standard deviations were computed for each type of location for each person, and then averaged across people. They are unstandardized.) For more nuanced results, showing differences between high arousal and low arousal positive and negative emotions, see Appendix 1. Sensed location Of the 12,310 users who provided momentary self-reports of location, 7582 also provided relevant momentary self-reports of mood that could be linked up to location readings from the location sensor with some degree of accuracy (see Appendix 1, Section 2.4 for details). Of these, 3646 users reported their mood in at least two of the three targeted locations and were used in these analyses. These users provided, on average, 111 reports of mood (from 2 to 1117 reports, median = 88). Consistent with the results for self-reported locations, users reported a more positive mood when the phone sensors detected that they were in a social type of situation (M = 0.34, SD = 0.17) vs. at home (M = 0.23, SD = 0.32), β = 0.20, t(3,645) = 37.22, p < 0.001, and when the phone sensors detected that they were at home vs. at work (M = 0.20, SD = 0.27), β = −0.11, t(3,645) = −33.97, p < 0.001. For more nuanced results, showing differences between high arousal and low arousal positive and negative emotions, see Appendix 1. Does trait personality moderate the relationship between affect and location? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Next we tested whether an individual’s personality traits moderated the relationship between their location and their state affect. For example, do people high in Extraversion experience more positive affect in social types of situations compared to people low in the trait? We created two datasets: (1) a dataset that had only reports from home and work, and (2) a dataset that had only reports from home and in social types of situations. Then we ran regressions predicting state affect from the interactions between location (dummy coded as before) and each of the big five personality traits (standardized across users), entered simultaneously. When we found significant interactions involving personality, we examined the simple slopes for low and high values of the personality trait at −1 and +1 SD from the mean. Self-reported location Of the 6759 users who provided momentary self-reports of location and reported their mood in at least two of the three targeted locations, 1434 also self- reported trait personality and were included in these analyses. The more positive mood associated with being at home vs. being at work was moderated by Openness, β = −0.04, t(1,428) = −2.70, p = 0.01, and Agreeableness, β = 0.04, t(1,428) = 3.08, p = 0.002 (see Appendix 1 for full model statistics). The mood benefits associated with being at home vs. being at work were larger for people who were high in Openness, β = −0.14, t(1,428) = −7.24, p < 0.001, than for people low in Openness, β = −0.06, t(1,428) = −2.92, p = 0.004. The mood benefits associated with being at home vs. being at work were larger for people who were low in Agreeableness, β = −0.14, t(1,428) = −7.16, p < 0.001, than for people who were high in Agreeableness, β = −0.06, t(1,428) = −3.32, p < 0.001. The more positive mood associated with being in a social type of situation vs. being at home was moderated by Neuroticism, β = 0.03, t(1,428) = 2.02, p = 0.04. The mood benefits associated with being in a social type of situation vs. being at home were larger for people who were high in Neuroticism, β = 0.35, t(1,428) = 15.42, p < 0.001 than for people who were low in Neuroticism, β = 0.28, t(1,428) = 11.37, p < 0.001. Sensed location Of the 3646 users who reported their mood in at least two of the three targeted locations and provided momentary self-reports of mood that could be linked to location readings from the location sensor, 651 also self-reported trait personality and were included in these analyses. As with the self-reported locations, the more positive mood associated with being at home vs. being at work was moderated by Openness, β = 0.05, t(645) = 3.64, p < 0.001. Openness showed the opposite pattern to the self-reported location data: the mood benefits associated with being at home vs. being at work were larger for people who were low in Openness, β = −0.18, t(645) = −10.15, p < 0.001, than for people high in Openness, β = −0.09, t(645) = −4.73, p < 0.001. The more positive mood associated with being in a social type of situation vs. being at home was moderated by Extraversion, β = −0.04, t(645) = −1.98, p = 0.05, and Agreeableness, β = −0.05, t(645) = −2.34, p = 0.02. The mood benefits associated with being in a social type of situation vs. being at home were larger for people who were low in Extraversion, β = 0.23, t(645) = 8.04, p < 0.001 than for people who were high in Extraversion, β = 0.15, t(645) = 5.60, p < 0.001. The mood benefits associated with being in a social type of situation vs. being at home were larger for people who were low in Agreeableness, β = 0.24, t(645) = 8.04, p < 0.001, than for people who were high in Agreeableness, β = 0.15, t(645) = 5.70, p < 0.001.","A person’s momentary mood fluctuated in relation to their location (i.e., at home, at work, and in social types of situations. We used both self-reported location, and location sensed via smartphone sensors and found remarkably consistent results (on grid valence, which was reported in the main text, but also on alternate measures of high and low arousal positive and negative affect, which were reported in Appendix 1). People reported more positive affect in social types of situations than at home or work. Further, high arousal positive and negative affect were reported more at work than at home, whereas low arousal positive and negative affect were reported more at home than at work. We found some evidence that the more positive mood associated with being in a social type of situation vs. being at home was moderated by personality. The difference in mood reported in social types of situations vs. at home was greater for people who were low in Agreeableness, low in Extraversion or high in Neuroticism. These results are consistent with the finding that Extraversion and Agreeableness are positively related to high quality social relationships, whereas Neuroticism is negatively related to high quality social relationships (Lopes, Salovey, & Straus, 2003; Lopes et al., 2004). Perhaps social types of situations are associated with greater mood gains for those who generally struggle more with their social relationships. We also found some evidence that the more positive mood associated with being at home vs. being at work was moderated by personality. The difference in mood reported at home vs. at work was greater for people who were low in Agreeableness. This might suggest that people who are low in Agreeableness struggle more in their forced interactions with colleagues at work (thus experiencing more positive affect at home vs. at work), whereas they thrive in their chosen interactions with friends in social types of situations (thus experiencing more positive affect in social types of situations vs. at home). These moderation results should be interpreted with caution for several reasons. First, though significant, these effects were small, |β|’s ⩽ 0.05. Second, the results were not always the same for the self-reported locations as they were for the sensed locations. For example, the difference in mood reported in social types of situations vs. at home was larger for people high in Openness when the location was self-reported, but larger for people low in Openness when the location was sensed. Finally, these results conflict with past research, which found additive effects of trait personality and situation characteristics on expressed behavior and emotion, but no interactions (Sherman et al., 2015). Further research is necessary to determine whether these results are reliable. The current results are consistent with an average layperson’s intuition: when people are asked to reflect on the factors that affect their expression of traits, they often mention the location (Saucier et al., 2007). Although the physical characteristics of a location may be directly related to the expression of traits (e.g., people may express more Neuroticism in unfamiliar environments, or less Extraversion in green spaces), it is likely that the psychological characteristics that become associated with a particular location are actually driving the relationship. For instance, being at home may be associated with spending time with family, and taking on a more communal role, whereas being at work may be associated with competition, and taking on a more agentic role. Rauthmann et al. (2014) proposed eight situational factors with psychological consequences: Duty, Intellect, Adversity, Mating, pOsitivity, Negativity, Deception, Sociality. We expect that the three locations we examined varied in many of these factors (e.g., Intellect may be more of a factor at work than at home or in social types of situations, whereas Sociality may be more of a factor in social types of situations than at work or at home); these assumptions could, of course, be tested in a future study. Future work is needed to determine which characteristics of a situation are active in the locations that we examined, but also in locations more broadly. This knowledge about what types of locations are likely to present situations high in each type of characteristic would allow future mobile sensing studies to sample a more complete set of locations. The usefulness of mobile sensing for testing questions about person-environment interactions will depend on its ability to sample from locations with a wide range of situational characteristics, and its ability to accurately label locations according to their characteristics with as little user intervention and training as possible. Smartphone sensors provided a powerful means of passively detecting the type of situation a user is encountering, and reducing the burden of self-reports in the current study. Their usefulness in future studies may depend on how capable they are of detecting other psychologically relevant situational factors. In this study, we focused on using location sensors to learn the semantics of places, so that we could examine relationships between place, affect, and personality. Collecting data from other sensors could augment this analysis, as well as provide other signals that may characterize the type of situation that a user is encountering. For example, collecting Wi-Fi scans can be used for finer- grained, indoor localisation (Gao et al., 2011), and bluetooth sensors can be used to detect co-located bluetooth devices of other study participants (Eagle & Pentland, 2006). Smartphones’ microphones can be used to measure ambient noise (Lathia, Rachuri, Mascolo, & Rentfrow, 2013); with sufficient training data, these can even be used to analyze participants’ speech (Rachuri et al., 2010). Recent work has shown that the ambiance of a place can also be derived from photographs taken there (Redi, Quercia, Graham, & Gosling, 2015). Work is needed to establish the validity of data from various sensors to assess psychologically meaningful situational characteristics. One way to establish the validity of sensor data is to compare it to self-reports, as we did in this study. However, various decisions need to be made in order to sample sufficient data from the sensors at the right times and at the right level of detail. In this study, the place label (home, work, or social) given to a particular geographic location (latitude, longitude pair) was inferred by matching self-reported location to location sensor data (see Appendix, Section 2.4). This process has a number of limitations. First, if a participant has never self-reported their location in a particular geographic area, we cannot make any inferences about that place. While there are methods that attempt to infer place semantics based on other signals (e.g., time of day/day of week), we limited our analysis to those places that were explicitly labelled. We matched location samples with place self-reports based on a fixed temporal threshold of 15 min; reducing this constant may improve the accuracy of location inferences, at the expense of reducing the number of labelled self-reports. Finally, we assumed that a user was in the same place if the two location samples were less than 1600 m from one another. In highly dense urban environments, this may not necessarily be the case. Even if phone sensors are shown to measure situational characteristics with some level of validity, work is needed to understand the capabilities and limitations of the sensors. We used location data that was passively captured from participant’s devices in order to infer where they were at the time of reporting their state personality or their mood. While location sensors are broadly reliable (indeed, they are used for mapping applications), they do not work when a user has disabled their device’s location services, and they do not work in all types of situations, such as underground or, in some cases, indoors. Obtaining a highly accurate and recent location sample is possible, but it is a time- and battery-consuming task, since it requires turning on the GPS sensor, waiting for it to obtain a satellite fix, and then collecting the data. Instead, we resorted to collecting coarse-grained data, which is easier and quicker to sample and more energy- efficient, but may be less accurate. The current results bolster the call for further research of within-person variation in affect and personality. We found evidence that people experience fluctuations in state affect when they are in different locations. As research continues to identify the factors of a situation that are psychologically meaningful, future work will be able to investigate the interactive effects of various situational factors on people’s feelings and behavior. The combination of experience sampling and mobile sensing allowed us to collect large amounts of within-person data in the current work, and could prove invaluable to further study, to the extent that phone sensors can objectively measure psychologically active situational variables. The current study describes how people feel and behave in different locations, but as a field we are just beginning to learn about the full extent to which feelings and behavior vary by type of situation."],["Social conformity is a class of social influence whereby exposure to the attitudes and beliefs of a group causes an individual to alter their own attitudes and beliefs towards those of the group. Compliance and acceptance are varieties of social influence distinguished on the basis of the attitude change brought about. Compliance involves public, but not private conformity, while acceptance occurs when group norms are internalised and conformity is demonstrated both in public and in private. Most contemporary paradigms measuring conformity conflate compliance and acceptance, while the few studies to have addressed this issue have done so using between-subjects designs, decreasing their sensitivity. Here we present a novel task which measures compliance and acceptance on a within-subjects basis. Data from a small sample reveal that compliance and acceptance can co-occur, that compliance is increased with an increasing majority, and demonstrate the usefulness of the task for future studies of conformity. --------------------------------------------------------------------------------","The multitude of ways in which the behaviour and attitudes of others impact our own has been studied since the very inception of psychology (Cialdini & Goldstein, 2004). A particular focus has been the study of how groups influence the behaviour of the individual, with studies such as those of Asch on social conformity (Asch, 1951, 1955, 1956) some of the most well-known in the field. In these studies participants were asked to complete a simple perceptual task (judging the length of lines) in a group setting where, unbeknownst to the participant, the other members of the group were confederates of the experimenter. During critical trials, despite the task having an obvious answer, the confederates all gave the incorrect answer. Only a quarter of participants remained completely independent of the group, with the rest showing various degrees of conformity towards the group’s responses. Subsequent work has identified various types of group influence, individuated by factors including the circumstances of the influence (e.g. whether the group pressure is explicit or implicit), and the nature of the change brought about in the individual. With respect to the latter, of interest to the current study is the distinction between compliance and acceptance (Kelman, 1958; Nail, Di Domenico, & MacDonald, 2013). Compliance and acceptance can be distinguished based on the type of attitude change brought about by the social influence. Compliance occurs when the individual publicly agrees with the group but does not change their own attitude or belief, whereas acceptance occurs when the social influence causes the individual to internalise the belief or attitude expressed by the group such that it becomes their own. Compliance and acceptance are thought to arise primarily from normative and informational influence, respectively (Abrams, Wetherell, Cochrane, Hogg, & Turner, 1990; Deutsch & Gerard, 1955). Normative influence occurs due to the desire of individuals to be accepted by the group, or at least not to be publicly in conflict with the group. Abrams et al. (1990, page 98) suggest that “compliance with the demands and expectations of other group members and overt agreement with their views occur because of their power to reward, punish, accept or reject individual members.” In contrast, informational influence is thought to result in acceptance because it occurs when individuals look to others for evidence as to the state of the world. As such, its effects are thought to be maximal when the state of the world is ambiguous, or when the individual is uncertain about a decision or judgement (Cialdini & Goldstein, 2004; Cialdini & Trost, 1998; Deutsch & Gerard, 1955). Thus, both informational and normative influence may reduce conflict between beliefs held by the self and those received from others. However, normative influence involves the reduction of public conflict with others in a group, whereas informational influence results in a reduction of conflict between incompatible beliefs within the individual. Although it is commonly accepted that both types of social influence are typical in everyday social situations, social conformity effects obtained using the Asch paradigm are usually attributed to normative influence (compliance) only (e.g. Allen, 1965, 1975; Bond & Smith, 1996; Turner, 1991). The fact that the perceptual decision task has an obviously correct answer (participants almost never make errors on trials where confederates give the correct answer) is usually taken as prima facie evidence that results occur due to compliance to the group decision through normative influence. This view is not universally accepted however (e.g. Abrams et al., 1990; Turner, 1985), and claims of an informational influence are supported by a handful of studies that have compared levels of conformity using this task between groups of individuals who must respond publicly, and those who have the opportunity to make their responses in private. The logic of these experiments is that, by comparing individuals who give their responses in public with those who respond in private, the relative contributions of normative and informational influence (and hence compliance versus acceptance) can be established (see Fig. 1). Individuals who respond in private should experience little to no normative influence due to the fact that the group members are unaware if they have conformed or not, and thus any group influence should be due to informational influence alone. Comparison of the degree of social conformity in private and public groups therefore allows the extent of normative and informational influence to be established using the standard Asch paradigm. Results of studies which have compared public and private conformity effects support the existence of an informational effect in the Asch paradigm as well as a normative effect. For example, Asch (1956; Experiment 4) found that rates of conformity (the percentage of critical trials across all participants in which errors in the direction of the confederates’ judgements were made) dropped from 43% in public conditions to 12.5% in private conditions – demonstrating a substantial normative influence effect. However, the 12.5% conformity rate in the private conformity condition was higher than the 1% error rate observed in control groups who were not subject to group pressure to give incorrect answers. This indicates the presence of an informational influence, albeit of smaller magnitude than the normative influence. Abrams et al. (1990) reached a similar conclusion, noting that participants conformed on an average of 58% of trials when asked to respond publicly, but only 33% when responding privately (see also Deutsch & Gerard, 1955 [comparison of Experiment 1: face- to-face condition and Experiment 2: anonymous, no commitment condition], where they reported a 16% drop in conformity using a privacy manipulation). While it is logically coherent to compare public and private responses to identify the relative contribution of normative and informational influence in the Asch paradigm, the current implementation of this comparison can be improved. Thus far, the manipulation of public and private responses has been between groups: the public group make their responses as per the original paradigm, whereas in the private group, although the confederates give their responses publicly, the participant responds in private. Groups of participants who respond publicly are then compared with those who respond in private. One issue with this approach is that comparisons between groups have the potential to be affected by sampling error and do not account for subject-level variance; as is typical therefore it is likely that within-subjects comparisons would be more sensitive than between-groups comparisons. The loss of sensitivity with between-subjects comparisons is especially important when using techniques with a low signal-to-noise ratio such as functional Magnetic Resonance Imaging (fMRI), which have been used several times to examine the neural correlates of conformity-related processes (Berns et al., 2005; Campbell-Meiklejohn et al., 2012a; Klucharev, Hytönen, Rijpkema, Smidts, & Fernández, 2009; Klucharev, Munneke, Smidts, & Fernández, 2011). Perhaps more important is the potential confounding effect of asking participants to respond in private while the confederates, who the participant believes to be other participants, respond publicly. Abrams and colleagues have argued that the distinction between the way in which the group of confederates respond (publicly), and the way in which the participant responds (privately), may result in the participant feeling like an out-group member. This may cause them to anti-conform to the confederates, reducing the observed magnitude of any conformity effect due to informational influence. Ideally then, the private / public response manipulation would be on a within-subjects level and the participants and confederates would respond in a similar manner. The Asch task itself has been criticised on a number of methodological grounds, several of which were noted soon after the Asch experiments were originally published (e.g. Crutchfield, 1955). Chief among these criticisms is the fact that the original Asch task is an insensitive measure; the choice of only three response options, two of which are very obviously wrong, presumably means that a great deal of pressure to conform must be experienced before an incorrect option is chosen. Small effects of social influence are therefore likely to go undetected. The use of only three response options also means the size of any conformity effect is impossible to measure; therefore, the degree of conformity in the task refers to the frequency of conformity across trials, rather than the size of the effect on any one trial. The use of a continuous response scale would alleviate these problems, although as far as we are aware continuous response scales have only been used with ‘off-line’ conformity paradigms where participants do not interact directly with group members in real life, but instead receive false feedback as to the responses of a group who had previously completed the task (e.g. Campbell-Meiklejohn et al., 2012a, 2012b; Klucharev et al., 2009, 2011; Zaki, Schirmer, & Mitchell, 2011). Meta- analytic work has demonstrated that off-line conformity paradigms result in reduced conformity effects compared to on-line paradigms in which participants interact with group members, and off-line paradigms are thought to induce conformity through different mechanisms to on-line paradigms (Bond & Smith, 1996; Deutsch & Gerard, 1955; Levy, 1960). This paper therefore reports data obtained from a relatively small sample of participants primarily to illustrate a novel on-line social conformity task based on the original Asch paradigm. The task was designed to measure the effects of normative and informational influence on a within-subjects level utilising public and private responses – thus identifying whether both compliance and acceptance may be induced in individuals by the group. Furthermore, participant and confederate responses were obtained using the same method, eliminating a factor which may have resulted in the participant feeling like an out-group member and anti-conforming from group responses in previous studies. Finally, participants were able to use a continuous response scale, meaning that small effects of social influence could be detected. Participants were asked to take part in a study on ‘the effect of motor preparation on perception.’ They were told that they would be tested in groups of four, for efficiency, and asked to judge the colour of a patch on a central screen. They were informed that they would need to type their responses on some trials, and speak their responses on other trials as the experiment was designed to compare the effect of preparing a manual versus a vocal response on colour perception. The order of responding was fixed such that the participant was the last to give a response. Thus, randomly across trials, it was either the case that the participant and confederates all typed their responses (baseline trials in which no social influence was present), the participant typed their response after the confederates had spoken their responses out loud (private conformity condition), the participant and confederates gave spoken responses (public conformity condition), or the participant and two of the three confederates gave a spoken response while one confederate typed their response (‘reduced majority’ trials in which the participant responded publicly but where the majority consisted of 2 rather than 3 confederates). Reduced majority trials were also included to ensure the response style for confederates and participants did not differ across trials (i.e. all participants experienced private trials). Given previous results (Abrams et al., 1990; Asch, 1956; Deutsch & Gerard, 1955), three main predictions can be made. First, it was predicted that conformity would be observed during both public and private conditions, as indexed by the difference in participants’ responses between trials in which the confederate responses were congruent (i.e. accurate) and incongruent (i.e. inaccurate) with respect to the correct response. Second, it was hypothesised that both acceptance and compliance can be induced in the same individual to bring about conformity. Acceptance and compliance can be identified, respectively, by an informational influence effect (indicated by private conformity) as well as a normative influence effect (indicated by greater public compared to private conformity). Third, it was predicted that the normative effect (compliance) would be greater in magnitude as the size of the majority increased from 2 to 3 confederates (Asch, 1955). The presence of reduced majority trials, where only 2 confederate responses are available (compared to 3 on a standard public trial) allow for this comparison to be made. Participants and design ~~~~~~~~~~~~~~~~~~~~~~~ Twenty-two healthy adult female individuals (mean age = 21.2 years, SD = 2.8) were recruited via the King’s College London research recruitment website. Thus, all participants were either staff or students at King’s College London, from a wide variety of disciplines. For recruitment to the study, it was a requirement that participants had not previously studied psychology at college or higher education level. Only female participants were recruited in order to remove the potential for out-group effects based on the sex of the participant compared to that of the (female) confederates. Participants were informed that this was a study investigating the effect of ‘motor preparation on perception’ and that the investigators were interested in better understanding how motor preparation for spoken versus typed responses affects visual perception. They were also told that to speed up data collection, and depending on the number of participants signed up to the timeslot, they would be tested in a group of up to 4 participants. In fact, each participant performed the task alongside 3 other individuals, all of whom were confederates of the experimenter. On arrival participants reported normal or corrected to normal vision, and they were asked to confirm their area of study, or department in which they work, at King’s College London. Following the experiment, participants were fully debriefed as to the true nature of the experiment. During a funnelled debrief – whereby participants were asked a series of questions which started broad and open-ended and funnelled to more specific questions to gauge their level of awareness of the true study aims – it was apparent that only one participant believed their fellow participants in the experiment to be confederates (and this one additional participant was excluded, leaving 21 participants’ data for analysis). Room and confederate setup The same 3 female confederates performed the task alongside each participant. They were undergraduate psychology students from King’s College London (mean age = 20.3 years), and their mean age did not differ significantly from the mean age of the participant sample (p > .05). In each instance, Participant 1 (confederate) arrived 15 min prior to the testing slot to ensure they were always first to arrive, taking a seat at Position 1 (see Fig. 2 for seating position labels). Following this, the real participant would be allowed time to arrive (most arriving a few minutes early for the testing session). The experimenter would instruct the participant to enter the testing room and take a seat at any laptop. To reduce uncontrolled interaction between the participant and confederates prior to the task, they were told to sit quietly whilst reading the instruction sheet and consent form. Due to the layout of the room, every participant without prompting chose to sit at Position 4. With the testing room door open to the half way point, Position 2 was blocked for the participant to take a seat in this position, and of the two positions remaining, seating themselves at Position 3 would block another participant’s access to Position 4. Thus, all participants followed the same pattern of seating themselves at Position 4 without direct instruction from the experimenter. Finally, in each instance, Participant 2 (confederate) would arrive a minute or two after the participant and take a seat at Position 2, followed by Participant 3 (confederate) who would arrive shortly after calling to ask for directions to the testing room and take the last remaining seat (Position 3). The participant number assigned to the confederates was counterbalanced across testing sessions. Finally, the 4 testing laptops (all 15.6 in., ASUS-Z550C, running Windows 10) were labelled with a sticker denoting the participant number prior to the testing session. See Fig. 2 for positions corresponding to the participant number assigned during the experiment. Visual perception task The visual perception task was programmed and run using Cogent and Cogent Graphics for Matlab. In the task, participants were required to make judgements about the colour of squares presented on a large central TV screen (Samsung-DM65D, 65 in. display) running Windows 10 software; see Fig. 2 for location in the testing room). The colour of the squares could lie anywhere on a scale from white to black. Following the presentation of the square, a colour bar was presented on the TV screen. Participants were required to give a numerical response, matching the colour of the square with the same colour on the colour bar and reporting the colour’s numerical value (see Fig. 3). On some trials participants were asked to type their response using the numeric keyboard and on others they were asked to say their response out loud for the experimenter to write down. These instructions were given to each participant on their own laptop screen for each trial. They were instructed to pay very close attention to whether they should say or type their response on each trial, as this would sometimes differ across participants. For example, on any one trial, all participants could be required to say their responses out loud, or they may all be required to type their response. Alternatively, on some trials, one of the participants may be instructed to type their response whilst all other participants were instructed to say their response. Ostensibly in order that the experimenter could record the spoken responses by hand, participants were instructed to respond in ascending participant number order. Single trial structure The experimenter verbally introduced each new trial and the trial number was displayed (e.g. ‘Trial 1’, ‘Trial 2’, ‘Trial 3’) on the main TV screen and each participant’s laptop screen. Each new trial (see Fig. 3 for a visual depiction of one trial) began with the trial number as well as a coloured square being presented on the main TV screen for 3 s. The trial number was also displayed on the participants’ laptop screens during these 3 s. The square was then replaced by a colour bar on the main TV screen (showing every shade from black to white; with labels in 10 step increments from 0 to 100), and simultaneously, each participant was instructed on their own laptop screen whether they should type or say their answer on that trial. Participant 1 then made their response, followed by the other participants in ascending participant order. On trials where any participant was instructed to type their response, their laptop would beep after two digits had been pressed using the numeric keyboard. This acted as an auditory cue for the next participant to respond. Participants were instructed they could respond using any number on the colour bar scale between 01 and 99. At the end of the trial, all participants and the experimenter pressed the space bar to move the experiment onto the next trial. The protocol began with 10 practice trials, followed by 153 trials in the main experiment, including a short break after each block of 51 trials. The task took approximately 45 min to complete. Trial type manipulations To allow for public and private performance to be assessed within the same task, different trial types were introduced. First, to gauge the participants’ baseline performance when judging the colour of the squares, there were 18 ‘silent’ trials during which all participants were instructed to type their response. Second, ‘private’ trials were introduced (27 in total) during which the participant was required to type their response (and therefore this response remained private) whilst all other participants were instructed to say their response. All 4 participants (i.e. the real participant plus each of the three confederates) experienced their own set of 27 private trials where they were required to type their response. These trials constitute the private condition when the participant responded privately, and reduced majority trials when each of the confederates responded privately. Finally, during ‘public’ trials, all participants were instructed to say their response (27 in total). Although these trials constitute the public condition for analysis purposes it should be noted that on reduced majority trials the participant was required to make their response publicly after two responses from confederates, therefore these trials could be considered public trials with only two confederate responses rather than three. Based on the seminal studies of Asch (1951, 1955), one may expect to also see a conformity effect on reduced majority trials, although of lesser magnitude than on the public trials with three confederate responses. Across the whole experiment, responses on 30% of trials for each participant were typed, and 70% spoken. Confederate congruency manipulation In line with Asch’s original conformity studies (Asch, 1951, 1955, 1956), within each of the private and public conditions there was a ratio of 1:2 congruent to incongruent trials. Here, congruency refers to the relationship between the correct response and the responses given by the confederates. On congruent trials participants gave a correct answer, while on incongruent trials they gave an incorrect answer. Confederates were instructed as to which response to give via their laptop screens. During each trial, according to the trial type, a number was calculated for each confederate as a response by taking the correct square colour (e.g. 50) and adding or subtracting values from this. On an incongruent trial, 15 was added to or subtracted from this value (e.g. 35 or 65), as well as a small amount of jitter between 0 and 3 (e.g. values between 32 and 38 on a trial where the confederates responded lower than the correct response of 50, or between 62 and 68 on a trial where the confederates responded higher than the correct response of 50). On a congruent trial, however, only the jitter of 0, 1, 2, or 3 was added to or subtracted from the correct response (e.g. confederate responses could vary between 47 and 53 on a trial where the correct response was 50). This procedure was adopted in order to induce some slight variation between confederate responses to mask their status as confederates.","For each trial a ‘response discrepancy’ was calculated as the difference between the participant’s response and the correct response. These values were calculated relative to the direction of the responses of the confederates. Thus, positive values indicate a discrepancy in the participant’s response in the direction of the confederates’ responses and negative values indicate a discrepancy in the opposite direction to that of the confederates. For example, for a trial where the correct response was 50 and confederates responded higher on the scale, a response of 55 by the participant would be considered a response discrepancy of +5, whereas a response of 45 would be considered a response discrepancy of −5. However, for a trial where the correct response was again 50 but confederates responded lower on the scale, a response of 55 would be a response discrepancy of −5, whilst a response of 45 would be a response discrepancy of +5. This allows for a calculation of the degree to which participants shift their responses towards or away from the confederates’ responses during private and public trials. The private and public conformity effects were calculated as the difference between response discrepancies during their respective congruent and incongruent trials, thus providing a measure of the informational influence effect (indexed by the size of the private effect) and the additional normative influence during public trials (indexed by the size of the public effect minus the private effect). Trials for which response discrepancies were ±2.5 standard deviations from each participant’s mean for each condition were discarded with an a priori threshold of 15% lost trials for participant inclusion. As it was not necessary to discard more than 15% of data for any participant, all participants’ data were retained for full data analysis. Fig. 4 displays mean response discrepancies during each trial type. As observed in the original Asch paradigm, baseline performance was excellent; the mean absolute response discrepancy from the correct response during silent trials was 1.42 (standard error of the mean [SEM] = 0.16). When signed rather than absolute scores were analysed, response discrepancies during silent trials did not significantly differ from 0 [t(20) = 0.40, p = .696]. Thus, participants could perform the task with a high degree of accuracy when they were not subject to responses of the confederates. Moreover, performance did not significantly differ from 0 during congruent trials for either private [t(20) = 1.69, p = .106] or public [t(20) = 0.76, p = .456] conditions. However, in line with our first prediction, conformity was observed towards the confederates’ responses in both public and private conditions, whereby response discrepancies were significantly larger on incongruent trials relative to congruent trials in both the private (private congruent response discrepancy = 0.57, SEM = 0.34, private incongruent = 2.10, SEM = 0.36; t(20) = 3.10, p = .006, d = 0.96) and public conditions (public congruent = −0.26, SEM = 0.34, public incongruent = 4.07, SEM = 0.41; t(20) = 7.79, p < .001, d = 2.51). Data were entered into a repeated-measures two-way ANOVA, with condition (private or public) and congruency (congruent or incongruent) as within-subjects factors. This revealed no significant overall difference between performance in private and public trials [F(1, 20) = 3.67, p = .070, ηp2 = .16], but a significant main effect of congruency [F(1, 20) = 69.76, p < .001, ηp2 = .78] and crucially, a significant condition*congruency interaction [F(1, 20) = 12.66, p = .002, ηp2 = .39], showing congruency (or conformity) effects to be significantly larger during public compared to private trials. The conformity effect during private trials (private incongruent – private congruent response discrepancies) can be considered to arise from informational influence (i.e. acceptance of the group response) and this effect is significantly different from 0 [t(21) = 3.10, p = .006]. The additional conformity effect observed on public trials (public minus private effect) can be accounted for by normative influence (i.e. compliance to the group), and is also significantly different from 0 [t(21) = 3.56, p = .002]. Thus, in line with our second prediction, we see the presence of both informational and normative influence. Moreover, it is possible to compare the relative magnitude of informational influence (indexed by the private conformity effect; mean = 1.53, SEM = 0.50) and normative influence (indexed by public conformity minus private conformity effects; mean = 2.80, SEM = 0.79). The normative influence is numerically, but not statistically, significantly larger than that of informational influence [t(20) = 1.06, p = .303]. A significant congruency effect was also observed on public trials when responses of only two confederates were available to the participant [i.e. on reduced majority trials; t(20) = 9.05, p < .001, d = 2.73], whereby response discrepancies were larger during incongruent (mean = 3.21, SEM = 0.32) than congruent trials (mean = −0.59, SEM = 0.19). To examine our third prediction concerning the impact of the size of the majority (number of confederate responses available to the participant) on performance during public trials, we compared performance between public trials where responses of 2 versus 3 confederates were available to the participant. Response discrepancies during incongruent public trials were significantly larger during trials where the majority consisted of 3 confederates (mean = 4.07, SEM = 0.41) than 2 confederates (mean = 3.21, SEM = 0.32; t(20) = 2.72, p = .013, d = 0.51). Moreover, there was a significantly increased public congruency effect under a majority of 3 (mean = 4.33, SEM = 0.56) compared to 2 (mean = 3.27, SEM = 0.36; t(20) = 2.29, p = .033, d = 0.49) confederates. Finally, the magnitude of informational and normative influence can be compared for trials in which only two confederates’ responses were available to the participant. Once again, the normative influence here (mean = 1.74, SEM = 0.63) is numerically, but not statistically, larger than that of informational influence (mean = 1.53, SEM = 0.50; t(20) = 0.19, p = .851). The normative influence arising from 3 confederate trials was significantly greater than the normative influence arising from 2 confederate trials [t(20) = 2.29, p = .033, d = 0.32].","Compliance and acceptance are types of social influence which can be discriminated on the basis of their effect on the beliefs or attitude of the individual. Compliance occurs when the individual publicly accepts the group’s position but privately adheres to their own belief, while acceptance occurs when the individual internalises the belief or attitude of the group such that it becomes their own. Compliance is thought to be the result of normative influence; where the individual complies with group norms due to the social power of the group. Acceptance is thought to be the result of informational influence; where individuals seek information from others in order to determine the true state of the world. The impact of normative and informational influence can be determined through on- line conformity paradigms by comparing responses made by individuals in private and in public. Here we report a novel paradigm, based on that of Asch (1951, 1955), in which both normative and informational influence can be measured on a within-subjects basis. In line with the first prediction, as well as the original body of work by Asch (1951, 1955, 1956), results indicated significant conformity effects both when participants’ responses were made in private and when made publicly in front of a group. With respect to the second prediction, it is thought that private conformity results solely from informational influence – acceptance of the group response without social pressure to conform – whereas public conformity results from both informational influence and normative influence – responding in line with the group in order to publicly comply with them. Crucially therefore, as has been reported in past comparisons of private and public conformity (Abrams et al., 1990; Asch, 1955; Deutsch & Gerard, 1955), a greater conformity effect was observed during public than private trials, demonstrating the presence of both informational influence in the private condition and the addition of normative influence during public conditions. Thus, we see demonstrations of both acceptance and compliance within the same task. Finally, results support the third prediction of increasing normative conformity with increasing size of the majority, replicating the finding by Asch (1955) whereby the size of the conformity effect increased as group size increased from 2 to 3. The data reported here suggest that the paradigm is a valid test of social conformity. Based on Asch’s (1951, 1955) original paradigm, it compared performance on a test of colour matching in a baseline condition when participants were unaware of the group’s judgements, when they were aware of the responses of three confederates but could give their own response in private, and when participants were aware of the responses of two or three confederates and had to give their responses in public. In common with Asch’s original findings, participants conformed to the judgements of others in the group; despite the task being sufficiently easy (participants’ responses were extremely accurate when not influenced by the group), responses were more inaccurate when the group gave incorrect responses. Also in common with the findings of Asch was the observation that the degree of conformity was greater when participants were exposed to three confederates than two. This was observed even though the ‘degree of conformity’ in the original Asch paradigm refers to the frequency of conformity, whereas here it refers to the magnitude of the conformity effect on a per trial basis. The paradigm reported here has the advantage that measures of normative and informational influence are obtained from the same individual at the same point in time, facilitating sensitive comparisons as to their magnitude and enabling future studies using techniques such as fMRI which are reliant upon within-subject comparisons. In addition, the participant is not excluded from the group due to their private responses, avoiding the possibility of anti-conformity to the group judgements which may have influenced prior studies (Abrams et al., 1990). Finally, although limiting the direct comparison of effect sizes between the current paradigm and the Asch paradigm, the use of a continuous response scale allowed extremely subtle effects to be detected on a trial-by-trial and within-subject basis. While the present results are encouraging, and support the validity of the task as a measure of conformity and its ability to measure informational and normative influence, it should be noted that the current study was primarily aimed at task development and validation. Therefore, variance due to key individual differences previously investigated with the Asch-style paradigm was deliberately reduced. For example, due to the literature on sex differences in social conformity (Cooper, 1979; Eagly, Wood, & Fishbaugh, 1981; Larsen, Triplett, Brant, & Langenberg, 1979), both concerning the sex of the participant and also whether the group is of the same or opposite sex to that of the participant, the current study included only female participants and a group of female confederates. Moreover, we restricted our sample to young adults (aged 19–30 years) to reduce age-related individual differences. The sample size was also relatively small, likely providing an inexact estimate of the population effect size. It can be seen then, that although these results support the use of the task to measure compliance and acceptance, types of social influence thought to arise from normative and informational influence respectively, further work is required to establish the replicability of these findings, and how they are moderated by factors such as age, sex, and culture.","Here we present data from a small sample validating a task designed to identify compliance and acceptance as a result of group influence, and therefore to identify normative and informational influences on decision-making. Although heavily-based on the conformity paradigm developed by Asch (1951, 1955), the task has several advantages over previous versions of the Asch paradigm. First, public and private responding is manipulated on a within-subjects basis rather than between-subjects, providing a more sensitive measure of any difference in social influence in the two conditions. This feature also allows participants to respond in the same manner as confederates, reducing the likelihood that they will classify themselves as an out-group member. Second, participants were able to respond on a continuous scale – this enabled a graded measure of conformity to be established such that the magnitude of any effect on a single trial could be measured and subtle effects of conformity detected. Results supported previous claims of both normative and informational influences in Asch-type paradigms, with normative influence increasing with increasing majority size in such paradigms."],["Previous research suggests that simple structure CFAs of Big Five personality measures fail to accurately reflect the scale's complex factorial structure, whereas EFAs generally perform better. Another strand of research suggests that acquiescence or uniform response bias masks the scale's \"true\" factorial structure. Random Intercept EFA (RI-EFA) captures acquiescence as well as the complex item-factor structure typical for personality measures. It is applied to the NEO-FFI and the BFI scale to test whether an accurate model-to-data fit can be achieved and whether the \"clarity\" of the factorial structure improves. The results lend confidence in the general effectiveness of RI-EFA whenever acquiescence bias is an issue. Example M. plus code is provided for replication. © 2014 The Authors. --------------------------------------------------------------------------------","The evaluation of the structure and validity of new and existing measures of psychological traits or attitudes is a central task in quantitative empirical research. One of the most important statistical techniques in this area is the common factor model and its variants, Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA). In particular, this study focuses on assessing the factorial structure of “Big Five” personality trait measures by means of factor analysis. First, a relatively large body of literature has shown that the unrestricted EFA model is often better suited to reflect the complex factorial structure of Big Five scales than the more restrictive “simple structure” CFA approach (e.g. Asparouhov & Muthén, 2009; Booth & Hughes, in press; Borkenau & Ostendorf, 1990; Hopwood & Donnellan, 2010; Marsh et al., 2010). This is mainly because CFA, or more precisely, its notion of independent clusters of items assumes that indicators perfectly measure only one target trait, though this condition is rarely met. Second, the literature shows that self-report measures are commonly plagued by different sorts of method bias (see, for an overview, Podsakoff, MacKenzie, & Podsakoff, 2012). Particularly, acquiescence response style, also known as directional or uniform response bias (hereafter ARS), has been identified as one of the most important nuisance factors in personality measurement (e.g. McCrae, Herbst, & Costa, 2001; Rammstedt & Farmer, 2013). It therefore has become standard to mitigate this problem by using semantically balanced scales in questionnaires, i.e., using positively (pro-trait) and negatively (con-trait) worded items. Most importantly, ARS can distort the factorial structure, i.e., convergent and discriminant validity of the questionnaire items as systematic measurement error is introduced (Podsakoff et al., 2012). As a consequence, the “true” factorial structure, especially in multidimensional instruments (e.g. Big Five scales), is misrepresented. Several approaches have therefore tried to remedy this problem by means of adjustment techniques that mitigate the impact of ARS, for instance, by using a CFA model of ARS (Billiet & McClendon, 2000; Maydeu-Olivares & Coffman, 2006) or by subtracting the within- person mean response from each item, also called ipsatization (Rammstedt & Farmer, 2013).","In this study I investigate the properties of Random Intercept EFA (RI-EFA) (Aichholzer, in preparation). RI-EFA treats ARS as an individual random intercept (RI) factor extending the standard EFA model. In particular, the RI-EFA approach is applicable whenever multidimensional and semantically balanced scales of psychological constructs or attitudes are examined. This research demonstrates the usefulness of RI-EFA using two standard Big Five personality scales, the German versions of the NEO-FFI (Borkenau & Ostendorf, 1993) and the BFI (Rammstedt & John, 2005). In addition, I test four hypotheses that have been put forward in previous research: H1: Measurement models with complex item-factor structure (EFA) commonly fit better to Big Five data than restricted simple structure models (CFA) (e.g. Hopwood & Donnellan, 2010; Marsh et al., 2010). H2: Measurement models with adjustment for ARS fit the data better than models without such adjustment (e.g. Billiet & McClendon, 2000; Maydeu-Olivares & Coffman, 2006). H3: ARS adjustment improves the clarity of factor loading patterns in ways that approach the ideal “perfect simple structure” form (e.g. McCrae et al., 2001; Rammstedt & Farmer, 2013). H4: ARS adjustment improves the clarity of factor loading patterns in particular in populations that are prone to ARS (e.g. Rammstedt & Farmer, 2013; Rammstedt, Goldberg, & Borg, 2010). H1 is tested by comparing common model fit criteria of standard CFA with EFA. H2 is tested by comparing RI-EFA with the other models. H3 and H4 are assessed by comparing empirical factor loading structures with a perfect simple structure matrix (i.e., perfect −1/0/+1 entries). The contribution of this study is thus twofold: First, a novel method in the field of factor analysis is tested with regard to its applicability to established scales. Second, these scales are re-evaluated using this new method.","Recently, Aichholzer (in preparation) has presented the RI-EFA model that allows researchers to combine standard EFA, i.e., an unrestricted or complex item-factor loading matrix, with a random intercept (RI) factor that represents ARS or uniform response bias more generally. So far, the latter option was only available for restricted CFA models (Billiet & McClendon, 2000; Maydeu-Olivares & Coffman, 2006). Further note that the factor α has a constant loading vector of 1 on all indicators, regardless of their keying (pro- trait or con-trait), while its variance is freely estimated and tested to be non-zero.1 The RI-EFA model is thus a hybrid model that combines an EFA part where item-factor loadings are freely estimated and a restricted CFA part where item-factor loadings on the RI/ARS factor α are restricted to follow a predefined pattern. Due to its specific factor loading structure α must not be confused with other factors of personality that load on all items (see Anusic, Schimmack, Pinkus, & Lockwood, 2009). In other words, αi represents a uniform shift of individual item responses independent from substantial factors. Hence, it can be considered an issue of differential item functioning (DIF) or violation of measurement invariance. Furthermore, the model can be estimated quite easily applying the Exploratory Structural Equation Modeling (ESEM) framework (Asparouhov & Muthén, 2009) and the Mplus software (Muthén & Muthén, 1998–2012) (see the Appendix A for example Mplus code). Instruments ~~~~~~~~~~~ This study investigates the German 60-item NEO-FFI (Borkenau & Ostendorf, 1993) and the German 44-item BFI (Rammstedt & John, 2005). Both the NEO-FFI and the BFI are based on the Big Five taxonomy of Extraversion (E), Agreeableness (A), Conscientiousness (C), Neuroticism (N), and Openness to experience (O). All items were measured on a 5-point scale with endpoints labeled as 1 – strongly disagree to 5 – strongly agree in the NEO-FFI and 1 – does not apply at all to 5 – applies completely for self-reports in the BFI. In the sample used here, the Cronbach’s Alpha estimates for the hypothesized Big Five dimensions were .75 (E), .72 (A), .84 (C), .82 (N), and .66 (O) for the NEO-FFI and .81 (E), .76 (A), .80 (C), .74 (N), and .82 (O) for the BFI. Note that previous research has already investigated the factorial structure of these instruments using different factor analytic strategies (Booth & Hughes, in press; Marsh et al., 2010; Rammstedt & Farmer, 2013).","The NEO-FFI and the BFI were administered in a random sample of the German population (aged 18 and above) as part of a larger study on personality and political behavior (Schumann, 2004).2 The total sample with valid demographic data comprised n = 2508 respondents. Age ranged between 18 and 92 (Mean = 49, SD = 17) and 52% of the respondents were female, 48% male. The NEO-FFI was administered face-to-face, while the BFI was administered using self-administration as a drop-off (with n = 1492 participants in total). Analysis ~~~~~~~~ Analyses were conducted with standard simple structure CFA, EFA (or simple ESEM), and RI- EFA using linear MLR (maximum-likelihood with robust standard errors) estimation as well as WLSMV (weighted least square mean- and variance-adjusted) estimation for ordered categorical measures in Mplus 7 (Muthén & Muthén, 1998–2012). Global goodness-of-fit indices were inspected for each model in addition to the χ2-test. The reason for this is that χ2-tests commonly result in a rejection of the model when applied to large samples. According to common fit criteria (see Marsh, Hau, & Wen, 2004), acceptable fit is achieved when CFI > .90, TLI > .90, and RMSEA < .08, whereas excellent fit is achieved when CFI > .95, TLI > .95, and RMSEA < .05. Further, better fit of a model is also supported by lower BIC values. For all empirical analyses the theoretical five-factor structure was imposed to investigate the hypothesized structure of the NEO-FFI and the BFI. Solutions for the rotated factor loading matrix in the EFA part were computed applying oblique Quartimin rotation (for other factor rotation criteria see Asparouhov & Muthén, 2009).","I start by comparing the fit measures for the different modeling strategies (Table 1). First, the results reconfirm previous evidence suggesting that unrestricted EFA (or simple ESEM) fits Big Five data considerably better than standard simple structure CFA (see Table 1), which supports H1. Third, as suggested by Hopwood and Donnellan (2010), I compared the congruence of the empirical factor loading pattern to a perfect simple structure matrix by assigning each item to its hypothesized factor, using Tucker’s congruence coefficient c (see Table 2, detailed factor loading results using MLR estimation are available as supplemental materials). After partialling out the additional RI/ARS factor by means of RI-EFA, there is convincing evidence in support of further alignment (Δ of congruence) of items towards the hypothesized five-factor personality structure, which confirms H3. In other words, the matrix congruence to perfect simple structure is always larger for the RI-EFA factor loading patterns. Further, the variance explained by the RI/ARS factor can be computed using squared standardized loadings, 2.8% (NEO-FFI) and 7.5% (BFI) with MLR estimation, which is largely similar to previous findings (3–4%) (e.g. Anusic et al., 2009; Billiet & McClendon, 2000). Fourth, I look at subpopulations with potentially different amounts or unequal variance in ARS. A common proxy is the respondent’s educational level (Rammstedt et al., 2010) which was operationalized by contrasting “low” education (lower secondary education or less) vs. “high” education (admission to tertiary education or completed university degree), whereas “medium” education was omitted. The results reconfirm that the factor loading congruence to a hypothesized five-factor simple structure is, in general, lower among less educated respondents. Applying RI-EFA improves the clarity of factor loading patterns among both groups, though it improves especially among less educated respondents who filled out the BFI, which partly supports H4. In sum, the results suggest that applying RI-EFA makes a greater difference for revealing the “true” structure in the case of the BFI, whereas the impact of response bias on substantial results seems to be smaller for the NEO-FFI. However, it should be mentioned that differences between the two instruments may also reflect the use of different survey modes (F2F vs. P&P).","This research shows that RI-EFA can be used as an effective factor-analytic method. It provides a useful tool for researchers when they would like to investigate the factorial structure of multidimensional scales based on semantically balanced or heterogeneous items. The results support the improvement of model-to-data fit and suggest that the hypothesized factor structure became “clearer” (i.e., it approximated simple structure). Moreover, given that the identified bias resembles ARS, it may be the preferred method whenever heterogeneous populations with potentially different response behavior are studied or when specific measurement conditions apply that evoke such response bias. The main advantage which makes RI-EFA superior to other approaches, such as Principal Component Analysis (PCA) of ipsatized data (Rammstedt & Farmer, 2013), is that RI-EFA fits into the larger family of common factor models and structural equation modeling or ESEM (Aichholzer, in preparation). Therefore RI-EFA can be extended to testing measurement invariance over subgroups or over time as well as to testing covariates of the RI/ARS factor and, hence, causes of such bias. As has been shown, RI-EFA can also be extended to the case of ordered categorical indicators but also to binary items using WLSMV estimation (Asparouhov & Muthén, 2009). Nevertheless, there are some limitations to RI-EFA, some of which apply to the ESEM framework more generally (e.g. Booth & Hughes, in press; Herrmann & Pfister, 2013; Marsh et al., 2010). Firstly, RI-EFA is neither a strictly exploratory nor a strictly confirmatory method, but it might be valuable as a “first-step” method in the sense that it is less rigorous compared to the CFA approach. Second, different rotation criteria (Quartimin, Geomin, Varimax, etc.) should be examined as factor loadings and between-trait correlations are dependent on the rotation criterion used. Third, it is still a matter of debate to what extent the factor scores and associations with covariates will differ between the different approaches, i.e., CFA, EFA/ESEM, or RI-EFA (also see Booth & Hughes, in press; Herrmann & Pfister, 2013). Apart from that, quite often the items in a scale are not fully but partially balanced (e.g. NEO-FFI and BFI), that is, some subscales use an even number of pro-trait and con-trait items, whereas others are primarily keyed in one direction. Indeed, the ultimate goal of the RI/ARS factor is that it conveys all information from balanced subsets of items to also control for ARS in other items and subscales that are partly or completely unbalanced. One could also let a selection of balanced item sets define the RI/ARS factor (for a set of items in the BFI see Rammstedt & Farmer, 2013), whereas loadings are set to +1, and freely estimate loadings of the remaining items on that factor. Such a model can then be compared to the basic RI-EFA model (constant loadings of +1). Still, as previously noted (Aichholzer, in preparation), the construct validity of the RI/ARS factor itself will be higher with an increasing number of balanced items that come from heterogeneous constructs. However, less is known about the amount of balancing in scales in order to fully identify and accurately address the type of bias which is modeled here. Future studies might investigate these issues in more depth."],["Our visual system can rapidly process stimuli relevant to our current behavioral goal within various irrelevant stimuli in natural scenes. This ability to detect and identify target stimuli during nontarget stimuli has been mainly studied in adults, so that the development of this high-level visual function has been unknown among infants, although it has been shown that 15-month-olds’ temporal thresholds of face visibility are close to those of adults. However, we demonstrate here that infants younger than 15 months can identify a target face among nontarget but meaningful scene images. In the current study, we investigated infants’ ability to detect and identify a face in a rapid serial visual presentation. Experiment 1 examined whether 5- to 8-month-olds could discriminate the difference in the presentation duration of visual streams (100 vs. 11 ms). Results showed that 7- and 8-month-olds successfully discriminated between the presentation durations. In Experiment 2, we examined whether 5- to 8-month-olds could detect the face presented for 100 ms and found that 7- and 8-month-olds could detect the face embedded in rapid serial visual streams. To further clarify the face processing at this age of infants, we tested whether infants could identify upright and inverted faces in rapid visual streams in Experiments 3a and 3b. The results showed that 7- and 8-month-olds identified upright faces, but not inverted faces, during the visual stream, which reflected face inversion effects. Overall, we suggest that the temporal speed of face processing at 7 and 8 months of age would be comparable to that of adults. --------------------------------------------------------------------------------","The human visual system is highly sophisticated in terms of spatial and temporal perception. Regarding spatial resolution, human adults are most sensitive to the spatial frequency at 4–5 cycles per degree, and it takes approximately 6 months for infants to develop this acuity (Pirchio, Spinelli, Fiorentini, & Maffei, 1978). In terms of temporal resolution, adults could detect a target picture embedded in a sequence of nontarget pictures at a rate of merely 113 ms (Potter, 1976) or even 50 ms for letters (Nieuwenstein, Potter, & Theeuwes, 2008). Currently, it remains unclear whether and when infants develop such temporal visual faculties. Thus, in the current study, we focused on the development of infants’ processing of visual information in terms of the time domain by using the detection and identification of objects, especially faces. Faces convey critical information for social interaction; thus, the human visual system is intrinsically biased to prioritize faces over other nonface objects (e.g., Morton & Johnson, 1991; Reid et al., 2017). Nonetheless, little is known about the temporal aspects of infants’ face recognition during visually demanding circumstances of brief presentation. Previous studies have shown that the temporal thresholds of face visibility are immature in younger infants compared with adults who can accurately detect a human face presented for less than 100 ms (Gelskov & Kouider, 2010). Specifically, Lasky and Spiro (1980) examined the visibility thresholds of faces with visual masking and found that 5-month-old infants could not discriminate faces from other nonface objects presented for 100 ms followed by a visual pattern mask. Similarly, Gelskov and Kouider (2010) found that 5- and 10-month-old infants could detect faces at a duration greater than 150 ms, whereas no such detection occurs at a duration less than 100 ms. These results suggest that infants aged 5 months require a duration of at least 150 ms to detect faces. Another recent approach for the development of infants’ vision in the temporal domain has been made from electrophysiological studies, revealing that the neural mechanism of face detection would emerge at 5 months. Specifically, de Heering and Rossion (2015) used a fast periodic visual stimulation consisting of sequential images of objects at 6 Hz with human faces inserted in every fifth image. The researchers found 1.2-Hz right-lateralized occipito-temporal categorization responses to faces in 4- to 6-month-old infants. Similarly, Barry-Anwar, Hadley, Conte, Keil, and Scott (2018) demonstrated that 6- and 9-month-old infants could discriminate individual faces of apes with a 6-Hz presentation. Kouider et al. (2013) found that even 5-month infants show a face-related event-related potential (ERP) component of N290 and P400 in response to briefly presented (100 ms) face images with visual masking. These results suggest that the neural mechanisms begin to function for detecting visually masked faces even in 5-month-old infants. However, it is unknown whether these infants could identify a briefly presented face because the studies described above were designed to determine the ability of infants to detect faces. In addition, these studies using visual masking procedures only revealed the thresholds of face visibility. Therefore, infants’ ability to process the target face selectively among the nontarget but meaningful images in a short duration has not been investigated. The purpose of the current study was twofold. First, we examined whether infants could detect individual faces under temporally demanding circumstances of masked rapid serial visual presentations. Second, we further examined whether infants could identify a face by using the novel preference and face inversion effect. The presence of ability to identify the target face during rapid serial visual stream in preverbal infants expands the possibility that the infants’ cognitive functions, such as nonspatial temporal attention and visual working memory in the time domain, that these has not been investigated enough, are working from early age comparable to adults. To achieve this goal in the current study, we conducted four experiments. First, we examined whether infants could discriminate between the 100- and 11-ms visual streams during the rapid serial visual presentation. The literature shows that adults can detect and identify visual images embedded in a rapid stream of nontarget objects at a rate of 100 ms per item (Potter, 1976) or even faster at 13 ms per item (Potter, Wyble, Hagmann, & McCourt, 2014). The first sets of experiments examined whether infants could also detect the difference in presentation times (Experiment 1) or in faces (Experiment 2) at a comparable rate. Subsequent experiments were devoted to examining infants’ capability of identifying individual faces (Experiments 3a and 3b). Given the behavioral evidence that 5- and 10-month-old infants could detect a face presented for 150 ms (Gelskov & Kouider, 2010) and the neural evidence that infants aged 5 months showed the face-related ERP component for faces presented for 100 ms (Kouider et al., 2013), we predicted that the behavioral and neural system of face recognition would change substantially from 5 to 8 months of age, as evidenced by the findings that configural processing would be mature at around 7 months (Schwarzer & Zauner, 2003). Although this development continues through around 6 years (Freire & Lee, 2001) and even adulthood (Meinhardt-Injac, Persike, & Meinhardt, 2014; Mondloch, Le Grand, & Maurer, 2002), the findings of 7 months suggested early stages of the development. Thus, the current study enabled us to track the developmental changes in face detection and identification of relatively younger infants of 5 to 8 months.","In the current experiment, we investigated whether 5- to 8-month-old infants could discriminate between 100- and 11-ms rapid serial visual streams. It has been shown that adults can detect a target image fairly accurately at this rate (Potter, 1976). In general, adult studies using the rapid serial visual presentation task have presented stimuli at a rate of 10 items per second (e.g., Potter, 1976). As such, we investigated whether this presentation time would be applicable to infants. If infants could discriminate between 100- and 11-ms rapid serial visual streams, similar to adults, then infants should prefer the 100-ms stream over the 11-ms stream. However, if infants are unable to discriminate between these two streams, then they should show no preference for either stream.","In total, 20 infants from each age group participated in the experiment (5- and 6-month-olds: 13 boys and 7 girls, mean age = 164.6 days, SD = 18.06; 7- and 8-month-olds: 10 boys and 10 girls, mean age = 230.35 days, SD = 17.29). All infants were full-term at birth and healthy at the time of the experiment. An additional 8 infants were tested but were excluded from the analysis owing to their fussiness and/or crying. The infants for the current project were recruited through local newspaper advertisements. The current experiment and following experiments were approved by the ethical committee of Chuo University, and written informed consent was obtained from the parents of the infants participating in the experiments prior to testing. The current experiments were conducted in accordance with the Declaration of Helsinki guidelines.","All stimuli were presented on a CRT monitor with a refresh rate of 85 Hz and a resolution of 1024 × 768 pixels. There were two loudspeakers, and one was placed on either side of the CRT monitor. Each infant sat on a parent’s lap in front of the CRT monitor located approximately 40 cm away. There was a small hole below the monitor screen for a charge-coupled device (CCD) camera to record the infant’s behavior digitally throughout the experiment. This let the experimenter observe the infant’s behavior via a monitor connected to the CCD camera. The infant and the parent were located inside an enclosure made of plastic poles and a black cloth.","Stimuli were 30 color photographs of various indoor and outdoor scenes. On every trial, 15 images, which were pseudorandomly selected from 30 scene images, were presented sequentially, and the order of items was the same across the two streams. All stimuli were resized to 440 × 330 pixels, subtending 24.3° in width and 18.6° in height. The two streams were presented to either the right or left from the center of the monitor at a distance of 1.4° on each side from the center of the monitor. Procedure and design We presented two visual streams consisting of 15 images side by side for 10 s, and the presentation rate was 105 and 11.7 ms per item for the 100- and 11-ms streams, respectively; neither had an interstimulus interval. The order of images was maintained for each stream during a trial. Both streams were repeated throughout a trial, resulting in 100 and 909 images for the 100- and 11-ms streams, respectively. The side at which 100 or 11 ms appeared was maintained with the infants but was counterbalanced across infants. Three trials were conducted for all infants. We measured the infants’ looking behavior during the two streams using a preferential looking paradigm. Before the experiment started, the parents were instructed to close their eyes and not to talk to their infants throughout the experiment to prevent the parents’ behavior from influencing the infants. At the beginning of every trial, a cartoon (10.0° × 14.3°) was shown in the center of the monitor accompanied by a brief sound to obtain the infants’ fixation to the center of the monitor. The experimenter initiated each trial as soon as the infant fixated on the cartoon.","A coder, who was blind to the stimuli, measured the infants’ looking time for each stream by pressing one of two keys corresponding to the right and left presentation fields while the infants were looking at the monitor based on the offline video movie. When the infants looked away from the monitor, no key responses were recorded. In addition, the video recordings of the looking behavior of 10 infants were analyzed by a second observer who was not informed of the objectives of our study. The intercoder reliability was r = .98 (p < .05). We calculated the individual preference score for the 100-ms stream by dividing the infants’ looking time to the 100-ms stream across three trials by the total looking time across the three trials. We regarded the infants’ preference for the 100-ms stream as the index of discrimination of the two visual streams.","We computed the preference scores for the 100-ms stream across each infant to assess whether the infants differentiated between 100- and 11-ms streams. The mean preference scores to the 100-ms stream for the two age groups are shown in Fig. 1. A two-tailed one- sample t test (vs. the chance level of .50) revealed that the 7- and 8-month-olds preferred the 100-ms stream over the 11-ms stream, t(19) = 4.38, p = .01, d = 0.98. This pattern of results was not obtained for the 5- and 6-month-olds; no preference was found for this age group, t(19) = −0.02, p = .98, d = 0.00. We found a developmental difference in the preference score between 5- and 6- month-olds and 7- and 8-month-olds, t(38) = 2.14, p = .04, d = 0.68. One may argue that the preference for the 100-ms stream over the 11-ms stream in 7- and 8-month-olds is due to an aversion to the faster rate stream; however, if this were true, there should have been a preference for the 100-ms stream in 5- and 6-month-olds as there was in 7- and 8-month-olds. These results suggest that the 7- and 8-month-olds were able to discriminate the two streams, possibly reflecting that the 100-ms presentation rate would be available in infants.","The results of Experiment 1 indicated that the 7- and 8-month-old infants preferred the 100-ms stream over the 11-ms stream. However, this does not necessarily mean that the infants precisely perceived and identified every object represented in the images in the 100-ms stream. Rather, the results can imply that they hardly saw clear pictures from both streams and merely picked a limited number of segmented portions of the images more frequently in the 100-ms stream than in the 11-ms stream. In the latter case, we cannot conclude that the 7- and 8-month-olds could perceive the images embedded in the 100-ms stream. To exclude this concern, we examined whether 5- to 8-month-old infants could detect a specified target in the 100-ms stream in Experiment 2. Two female faces were presented as a target. Infants viewed two side-by-side streams of rapidly presented visual sequences at a rate of 100 ms per image. One stream contained an upright face, and the other stream contained its inverted version. We tested whether infants exhibit a preference for the stream containing the upright face. We predicted that infants who could detect the upright face should prefer the stream including the upright face more frequently than the stream containing the inverted face (Farroni et al., 2005; Mondloch et al., 1999).","In total, 20 infants for each age group who were full-term at birth and healthy at the time of the experiment participated in the experiment (5- and 6-month-olds: 9 boys and 11 girls, mean age = 173.4 days, SD = 14.39; 7- and 8-month-olds: 9 boys and 11 girls, mean age = 230.15 days, SD = 15.31). An additional 6 infants were tested but were excluded from the analysis owing to their fussiness and/or crying. A side bias was defined as a tendency to look at only one side at more than 90% of the looking time. Stimuli The apparatus was identical to that used in Experiment 1. The 30 color photographs presented in Experiment 1 and two color photographs of Japanese female faces taken in frontal view showing a neutral expression with a gray background were used. The image size of the two female faces was the same as the size of the natural images (24.3° × 18.6°). In every trial, pseudorandomly selected 14 natural images and one of the two female faces were presented sequentially. The order of the sequence was maintained for the same infant. The side on which the upright face was presented was also maintained for the same infant. The locations of the two streams were the same as those in Experiment 1 (1.4° apart from the center of the monitor). Procedure and design The procedure and design were the same as those in Experiment 1 except with the following changes. We presented two streams side by side for 10 s repeatedly; the upright female face or the inverted female face was embedded in either one of the two streams (Fig. 2). The female faces were always presented as the seventh image pairs in the streams. In every trial, 6 images preceded the face images, followed by 14 images, except that 3 images followed the final pair of the face images at the end of a trial. The presentation rate was 100 ms per item with no interstimulus interval for both streams. The order of images in a trial was maintained for each stream. Both streams continued until 100 images were presented. We conducted a total of six trials, in which the order of two female faces and their position orientation were counterbalanced across infants. The natural images presented on the right and left were identical. Verbal instruction in target detection task is impossible for infants; thus, we used preferential looking and a familiarization/novelty preference procedure. We presented the target face embedded in a visual stream several times, following the method applied in previous studies (Gelskov & Kouider, 2010; Lasky & Spiro, 1980).","In this experiment and the following experiments, the recording of the infants’ looking time for the two side-by-side streams (or pictures) was conducted the same as in the procedure for Experiment 1. We calculated the individual preference score for the stream containing the upright face by dividing the infants’ looking time to the stream including the upright face across the six trials by the total looking time across the six trials. As in the data coding of Experiment 1, an additional independent observer analyzed the video recordings of the looking behavior of 10 infants. The intercoder reliability was r = .91 (p < .05). We regarded their preference for the stream including the upright face as the index of detection of the upright face embedded in the stream.","Fig. 3 shows the mean preference scores for the stream containing the upright face for the two age groups. A two-tailed one-sample t test (vs. the chance level of .50) revealed that the 7- and 8-month-old infants preferred the stream containing the upright face, t(19) = 3.68, p = .01, d = 0.82. This result pattern was not obtained for the 5- and 6-month-old; no preference was found for this age group, t(19) = −0.43, p = .67, d = 0.09. We found a developmental difference in the preference score between infants aged 5 and 6 months and those aged 7 and 8 months, t(38) = 2.72, p = .01, d = 0.86. These results suggest that the 7- and 8-month-olds could detect the upright face embedded in one of the two rapid serial visual streams at a rate of 100 ms per item.","The results of Experiment 2 demonstrated that the 7- and 8-month-old infants could detect the upright face in a rapid stream visual presentation at a rate of 100 ms per item. Nonetheless, it remained unclear whether those infants identified the target as a specific person’s face or merely categorized it as a human face. We examined whether the 7- and 8-month-olds could discriminate a target face from a nontarget face by using a familiarization/novelty preference paradigm in Experiment 3a. Specifically, we introduced the familiarization phase and tests. We familiarized the infants to learn a target face embedded in a stream and subsequently tested whether they could discriminate between the target and nontarget faces. We expected that infants could reveal a novelty preference in that they could discriminate a target face from a nontarget face in the test when they learned a target face in the stream during the familiarization phase.","In total, 20 7- and 8-month-old infants who were full-term at birth and healthy at the time of the experiment participated in the experiment (11 boys and 9 girls, mean age = 230.15 days, SD = 14.67). An additional 5 infants were tested but were excluded from the analysis owing to their fussiness and/or crying.","The apparatus and stimuli were identical to those used in Experiment 2. A familiarization/novelty preference procedure was introduced to test whether the infants could identify a target face in a rapid sequence of visual images. The procedure consisted of the familiarization phase and tests. During the familiarization phase, 14 pseudorandomly selected natural images and one of the two female faces, as a target face, were presented sequentially in the center of the monitor at a rate of 100 ms per item with no interstimulus interval. One of the female faces was presented as a seventh image in the stream. Both the images and the order of images in a trial were maintained. The experiment consisted of pre- and posttests. During these tests, a target face (i.e., the familiarized face for the posttest) and a nontarget face (i.e., a novel face for the posttest) were presented side by side for 10 s to compare the infants’ looking behavior before and after the familiarization phase. The pretest was conducted to confirm that there was no preference bias for either one of the two faces used in this experiment, whereas the posttest was conducted to examine whether infants could discriminate the familiarized target face from a nontarget face. No-stream presentations were used during these tests. The distance from the center of the monitor was 1.4° for each face. Design As mentioned above, the procedure consisted of the familiarization phase and pre- and posttests. After all the infants received the pretest, they were familiarized with a target face during the familiarization phase. Six familiarization trials were conducted for individual infants. Finally, infants received the posttest to measure whether they could discriminate the target face from a nontarget face. Both pre- and posttests consisted of two trials in which the position of familiar and novel faces was counterbalanced across infants in the first and second trials. A trial began with a cartoon accompanying a brief sound to obtain the infants’ fixation to the center of the monitor. As soon as the infants fixated on the cartoon, the stimuli were presented. The intercoder reliability was r = .90 (p < .05) for the video recordings of the looking behavior of 5 infants. Familiarization phase The looking time at a single stream during the familiarization phase was averaged across the first three and last three trials for each infant. The mean looking time across the first three trials was 8.64 s (SD = 1.27) and that across the last three trials was 7.95 s (SD = 1.29). A two-tailed t test was conducted to confirm whether infants familiarized with a target face during the stream. The looking time across the first three trials was longer than that across the last three trials, t(19) = 2.41, p = .05, d = 0.54. The looking behavior toward a target face in the stream presentation declined throughout the familiarization trials, which suggested that infants were familiarized with a target face. Pretest and posttest To investigate whether the infants could discriminate the familiarized face (target) from a novel face (nontarget), we calculated the preference scores to the novel face (nontarget) during the pre- and post-familiarization tests. This index represents the novelty preference to a nontarget face in the post-familiarization test. We conducted a pre-familiarization test to establish that no preference bias was present before the familiarization. Novelty preference scores were computed by dividing the infants’ looking time at the novel face across the two trials by the total looking time across the two trials in the tests. The mean novelty preference scores for the tests are shown in Fig. 4 (left). A two-tailed t test revealed that the infants looked at the novel face for a longer period of time during the post- familiarization test than during the pre-familiarization test, as expected, t(19) = 3.43, p = .01, d = 0.77. In addition, we conducted a two-tailed one-sample t test (vs. the chance level of .50) separately for pre- and posttests and found that the infants looked at a novel face longer during the post-familiarization test, t(19) = 5.01, p = .01, d = 1.12. There was no such preference during the pre-familiarization test, t(19) = 0.52, p = .60, d = 0.12. These results suggest that the 7- and 8-month-old infants could discriminate the novel face (nontarget) from the familiar face (target). Because the face images were counterbalanced across infants, we can conclude that infants around this age could detect and identify a human face among a rapid sequence of visually presented natural images at a rate of 100 ms per item.","The results of Experiment 3a demonstrated that the 7- and 8-month-old infants were able to discriminate a target face from a nontarget face presented in a rapid visual stream. To unanimously conclude that the infants could discriminate the target as a face rather than a low-level visual pattern, we designed Experiment 3b in which the face inversion effect was introduced. The face inversion effect was reported to occur in 4-month-old infants (Turati, Sangrigoli, Ruel, & de Schonen, 2004), and preference for upright faces over inverted faces was observed in newborns and 3-month-old infants (Cassia et al., 2006; Mondloch et al., 1999). This preference has also recently been found in the human fetus (Reid et al., 2017). If infants were unable to discriminate the inverted face due to the face inversion effect, we can infer that the results of Experiment 3a were based on the infants’ face discrimination. However, if infants discriminated the inverted face, the results of Experiment 3a could be attributed to the discrimination of low-level features in the face images.","In total, 20 7- and 8-month-old infants, who were full-term at birth and healthy at the time of the experiment, participated in the experiment (15 boys and 5 girls, mean age = 233.75 days, SD = 16.22). The infants were recruited through local newspaper advertisements. Procedure, design, and apparatus The procedure, design, and apparatus were the same as those used in Experiment 3a combined with inverted faces. The intercoder reliability was r = .97 (p < .05) for the video recordings of the looking behavior of 5 infants. Familiarization phase We analyzed the results in the same way as in Experiment 3a. Specifically, the looking time at the single stream during familiarization phases was averaged across the first three and last three trials for each infant. The mean looking time in the first three trials was 9.28 s (SD = 0.85) and that in the last three trials was 8.23 s (SD = 1.54). A two-tailed t test to test the effect of familiarization revealed that the looking time in the first three trials was longer than that in the last three trials, t(19) = 3.41, p = .01, d = 0.76. The looking time declined across the familiarization phases. These results suggest that the infants were familiarized with the inverted target face. Pretest and posttest We calculated the preference score to the novel face (nontarget) for the pre- and post-familiarization tests across infants. The mean novelty preference scores for the tests are shown in Fig. 4 (right). A two-tailed t test indicated no significant difference in looking time at the novel face between the pre- and post-familiarization tests, t(19) = 0.04, p = .97, d = 0.01. In addition, the preference score during the tests did not differ from the chance level of .50 [pre-familiarization test: t(19) = 0.18, p = .86, d = 0.04; post-familiarization test: t(19) = 0.12, p = .91, d = .03]. These results suggest that the 7- and 8-month-old infants were unable to discriminate the inverted novel face from the inverted familiar face due to the face inversion effect.","The primary goal of the current study was to examine the infants’ ability to detect and identify individual faces under temporally demanding circumstances of masked rapid serial visual presentations. In Experiment 1, we tested whether infants could discriminate between 100- and 11-ms rapid serial visual streams. The results indicated that the 7- and 8-month-olds were able to discriminate the two streams, suggesting that the 100-ms presentation rate would be available in infants. In Experiment 2, we presented two visual streams, each containing an upright or inverted face to test whether infants could detect an upright face in 100 ms. The results showed that 7- and 8-month-olds could detect the face. In Experiments 3a and 3b, we investigated whether 7- and 8-month-olds could discriminate a target face that infants learned from a nontarget face using an upright face and an inverted face. The results showed that 7- and 8-month-olds could identify the individual face when the face was presented upright. However, they were unable to identify it when the face was inverted due to the face inversion effect (e.g., Turati et al., 2004). In sum, the current study revealed that infants aged around 7 or 8 months can not only detect the face but also identify individual face images among the nonface images presented at a rate of 100 ms. It is notable that 7- and 8-month-old infants could detect and identify individual faces presented even at a rate of 100 ms among nonface objects. To our knowledge, this is the first study to show successful identification of faces while viewing rapid serial visual streams in 7- and 8-month-olds. Previous studies displayed faces for approximately 20–70 s, and 7- and 8-month-olds identified neutral faces (Fagan, 1976; Otsuka et al., 2013; Righi, Westerlund, Congdon, Troller-Renfree, & Nelson, 2014; Tyrrell, Anderson, Clubb, & Bradbury, 1987; for a review, see Otsuka, 2017). In contrast, we presented the faces at a much faster rate (i.e., 100 ms per image), and the total duration amounted to 4.2 s at most. This result suggests that 7- and 8-month-olds can identify a face at a comparable rate to adults, who can detect and identify visual images embedded in a rapid stream of nontarget objects at a rate of 100 ms per item (Potter, 1976). Certainly, we could not rule out the possibility that infants learned the target and nontarget faces incidentally during the pretest because they were simultaneously exposed to these for 20 s.1 Regardless of this possibility, the significance of the current finding was that infants’ processing of face information occurs over a sequence of rapidly changing events. What Gelskov and Kouider (2010) focused on was the detection of a face appearing on the left or right side of a computer screen. In their study, the face was preceded by a forward mask that occurred in the same location where a face could appear. Their procedure measured detectability of a face with spatial uncertainty because the infants had no clue whether the face would appear on the left or right side of the screen. There were no distractors other than a nonface stimulus at the opposite side from the target. Therefore, Gelskov and Kouider’s stimulus configuration was relatively less demanding in terms of attention. In contrast, we embedded a target face in a rapid sequence of nontargets for more than 2 s. Therefore, the key issue here was the measurement of detectability of a face with temporal uncertainty under an attentionally demanding circumstance. In this sense, the current study is not a mere extension of a past study but rather revealed that 7- and 8-month-olds could detect the face when temporal uncertainty exists with visually demanding circumstances. The current finding is consistent with the literature demonstrating that the ability of identifying individual faces in 7- and 8-month-old infants is supported by a sufficiently developed higher-order mechanism of face processing at this age. It has been shown that the capabilities of configural processing and view-invariant face recognition are acquired at around 7 or 8 months of age (Fagan, 1976; Kobayashi et al., 2012; Nakato et al., 2009; Otsuka, Hill, Kanazawa, Yamaguchi, & Spehar, 2010; Schwarzer & Zauner, 2003). In terms of configural processing, 7-month-olds could perceive faces configurally (Cohen & Cashon, 2001; Kobayashi et al., 2012; Otsuka et al., 2010; Schwarzer & Zauner, 2003; Schwarzer, Zauner, & Jovanovic, 2007). Regarding view-invariant face recognition, 7- and 8-month-olds could perceive the face regardless of viewing orientations (Fagan, 1976; Nakato et al., 2009). It is plausible that 7- and 8-month-olds who have already acquired these abilities of face processing could identify the individual face embedded in a rapid visual stream. We suggest that these neural mechanisms support the ability to detect and identify faces found in the current study in 7- and 8-month-olds. In addition, previous studies have reported that infants aged even 3 months are sensitive to certain kinds of higher-order mechanisms such as configural processing (Bhatt, Bertin, Hayden, & Reed, 2005; Cohen & Cashon, 2004; Galati, Hock, & Bhatt, 2016; Joseph, DiBartolo, & Bhatt, 2015; Quinn, Tanaka, Lee, Pascalis, & Slater, 2013; Zieber et al., 2013). In the current study, we examined the infants’ ability to detect and identify faces during rapid serial visual presentation using configural face processing. Then, we found that the stimulus duration of 100 ms would not be sufficient time for 5- and 6-month-olds, but would be sufficient for 7- and 8-month-olds, to detect and identify upright faces. The temporal speed that would be sufficient for younger infants to detect and identify the face during the viewing of a rapid serial visual stream needs to be explored in future studies that systematically manipulate presentation rates. Importantly, a critical aspect of the current study is that infants can detect and identify faces under temporally demanding circumstances in which the images and faces are presented briefly at a rate of 10 items per second. Our findings on face identification in 7- and 8-month-old infants were inconsistent with the findings of Gelskov and Kouider (2010), which showed that 10-month-old infants could not detect the face presented for 100 ms followed by a scrambled face mask. In terms of the difference in experimental procedure, we used the rapid serial visual presentation task, whereas they used visual sandwich masking. Apart from this difference, we speculate that the difference in visual mask stimuli would be the main factor to elicit the inconsistent results. Gelskov and Kouider used the scrambled face as the forward and backward visual mask, whereas we presented various scene images as visual masking in a rapid serial visual stream. According to a previous study (Maguire & Howe, 2016), the effect of a visual mask by natural scene images is weak compared with scrambled noise images. Thus, the 7- and 8-month-olds in our study were able to detect and identify the face through using the natural scene images as a mask. One may argue that the target image and its subsequent nontarget mask images used in the current study differed in terms of low-level features (e.g., contrast, luminance), and this difference might have contributed to these 7- and 8-month-olds being able to identify a rapidly presented face easily. However, we believe that this possibility is unlikely because visual masking by natural scene images would occur in the current stimuli. The results in 5- and 6-month-olds in Experiment 2 provide affirmative evidence. Specifically, 5-month-olds could identify the face presented for 100 ms without a visual mask (Lasky & Spiro, 1980), whereas the 5- and 6-month-olds in our study could not detect the face presented for 100 ms during visual stream. These findings suggest that natural scene images serve as visual masking. Of course, embedding target faces among facelike distractors so that each image should contain a single large high- contrast pink object with a dark high-contrast hairline would make the task more difficult, and thus infants would need to be older to show the preference. To investigate such effects of masking during a rapid serial visual stream in infants, further studies like the adult studies are needed. The current study revealed that there is a developmental change in face recognition between the groups of 5- and 6-month-old infants and 7- and 8-month-old infants. The former could not detect a face embedded in a rapid serial visual stream, whereas the latter could. The failure of detection in the younger group of infants is consistent with a previous behavioral study in which 5-month-old infants were unable to detect a face presented for 100 ms followed by the visual mask stimuli (Gelskov & Kouider, 2010). An electrophysiological study, however, demonstrated brain activities in response to briefly presented faces at 5 months of age. Specifically, de Heering and Rossion (2015) found a right-lateralized occipito-temporal categorization response to a face presented at 6 Hz in 4- to 6-month-olds. These two apparently inconsistent results regarding the critical period of the emerging infants’ responses to faces can be reconciled by assuming that the emergence of brain activities in response to faces precedes behavioral responses. In general, the neural system should be organized before the emergence of infants’ behavior response (Jessen & Grossmann, 2015, 2019; Kouider et al., 2013; Nava, Romano, Grassi, & Turati, 2016). Therefore, it is likely that the current 5- and 6-month-olds might have shown responses in sensitive brain activities if measured, although those responses did not accompany explicit behavioral responses in the current study. However, by the age of 7 or 8 months, infants developed fully to reveal behavioral and explicit responses in detecting and identifying a face embedded in a rapid serial visual stream at a rate of 100 ms per image. The dorsolateral prefrontal cortex has been identified as neural substrates for functions of working memory and selective attentional control. These functions can be acquired in 7- and 8-month-old infants comparable to adults (Diamond & Goldman-Rakic, 1989). At around 7 or 8 months of age, the neural circuitry of fusiform face area (Kanwisher, McDermott, & Chun, 1997) and occipital face area (Rossion, 2008) regions were activated when rapidly detecting and identifying the briefly presented face in older infants and adults (Barry-Anwar et al., 2018; de Heering & Rossion, 2015; Gentile & Rossion, 2014). These neuronal areas might be concerned with the rapid detection and identification of faces in 100 ms similar to the mechanisms in adults. The current findings reveal remarkable capabilities of infants at this age, expanding the possibility that infants’ attentional functions can be tested by using the rapid serial visual presentation procedures. This encourages exploration of relatively less examined infants’ cognitive functions, such as nonspatial temporal attention, in future studies."],["This temporal binding effect is partially retrospective, i.e., occurs upon outcome perception. Retrospective binding is thought to reflect post-hoc inference on agency based on sensory evidence of the action - outcome association. However, many previous binding paradigms cannot exclude the possibility that retrospective binding results from bottom-up interference of sensory outcome processing with action awareness and is functionally unrelated to the processing of the action - outcome association. Here, we keep bottom-up interference constant and use a contextual manipulation instead. We demonstrate a shift of subjective action time by its outcome in a context of variable outcome timing. Crucially, this shift is absent when there is no such variability. Thus, retrospective action binding reflects a context-dependent, model-based phenomenon. Such top-down re-construction of action awareness seems to bias agency attribution when outcome predictability is low. © 2014 The Authors. --------------------------------------------------------------------------------","When a motor action is followed by a sensory event within a few hundred milliseconds, the subjective time of the action can be shifted towards that sensory event (Haggard, Clark, & Kalogeras, 2002; for a review, see Moore & Obhi, 2012). This binding of subjective action time by a subsequent stimulus depends on learning the underlying contingency (Moore, Lagnado, Deal, & Haggard, 2009; Walsh & Haggard, 2013). It is considered to reflect intentional (Haggard et al., 2002) and/or causative (Buehner & Humphreys, 2009) aspects of the learnt action – outcome association that contribute to a pre-reflective sense of agency (Moore & Obhi, 2012). Temporal action binding is partially retrospective, in so far as it occurs at the time of outcome perception (Moore & Haggard, 2008). This retrospective component of temporal binding has frequently been regarded as evidence in favour of a post-hoc implicit inference on agency based on sensory evidence, as opposed to a prospective implicit agency ascription during action preparation and/or execution (Moore & Haggard, 2008). This distinction relates to the broader question (Baldwin et al., 2003) of how awareness of prior expectations of an action-outcome (e.g. during unsuccessful trying) on the one hand, and the observation of physical action consequences on the other, are weighted and integrated in the process of sensing intention and causation (Wegner & Wheatley, 1999; Wolpert & Flanagan, 2001). These interpretations of retrospective action binding assume an underlying inferential, top-down process. In principal, however, a retrospective change of action awareness could be driven by two physiologically distinct mechanisms. A top-down process could re-construct action awareness by integrating prior knowledge and sensory evidence of the action – outcome association. A similar model-based process has been proposed, for example, for an integration of motor predictions and sensory re-afference in comparator models of motor control (Wolpert & Flanagan, 2001). Alternatively, retrospective action binding could reflect bottom-up interference of sensory outcome processing with on-going cognitive processes involved in action awareness, or, at a lower level, with the integration of visual information in the Libet clock paradigm typically deployed in these experiments (Libet, Gleason, Wright, & Pearl, 1983). In this case, the retrospective shift of subjective action time would be solely stimulus- driven, similar to the effects of backward masking on conscious perception, for example (Werner, 1935), and would be independent of an internal model of the action – outcome association. Previous evidence of retrospective temporal binding comes predominantly from studies that compare subjective action time in the presence and absence of a subsequent stimulus, i.e., two conditions that differ substantially in bottom-up drive (Moore & Haggard, 2008). In these studies, a shift of subjective action time due to outcome presentation reflects retrospective action awareness. Therefore, these studies cannot distinguish between model-based re-construction of action awareness and bottom-up interference with action awareness. Evidence for the idea that a top-down re-constructive process is involved in retrospective binding comes from a study that manipulated the contingency of an outcome on an action, specifically the probability of an outcome given no preceding action (Moore et al., 2009). This study held bottom-up drive constant across trials of interest and used a contextual manipulation instead to demonstrate that retrospective binding depends on an internal representation of how necessary an action is for a sensory event. Similarly, we here used a contextual manipulation while keeping bottom-up influences constant to address the question whether action awareness depends on an internal model of outcome timing. Previous studies have shown that binding of an outcome to an action depends on temporal outcome predictability (Haggard et al., 2002). We asked whether retrospective and prospective components of action awareness also reflect temporal outcome predictability. Moore and Haggard (2008) have previously shown that outcome presentation shifts subjective action time under conditions of low outcome predictability, as achieved by lowering the probability of outcome occurrence (the “whether” of the outcome). Here, we tested whether a contextual manipulation of temporal outcome predictability leads to a similar shift of subjective action time. In brief, we compared retrospective binding in two conditions that differ in temporal variability of the outcome with respect to the action (the “when” of the outcome) while keeping the structure and timing of trials of interest constant, i.e., bottom-up drive. Any difference in retrospective binding between conditions of variable and fixed outcome timing would reflect an influence of context, rather than any bottom-up interference with action awareness. In addition, Moore and Haggard (2008) demonstrated a shift of action awareness towards an expected outcome even when the outcome was omitted. The authors interpreted this shift as a prospective component of action awareness. To examine effects of temporal outcome predictability on this prospective component, our study introduced an expectation of outcome occurrence by presenting outcomes on 75% of trials in both “variable” and “fixed” outcome timing blocks.","We recruited 16 healthy volunteers (age 23.5 ± 2.78 years (mean ± SD), 10 females; all right-handed) via an online database. All participants gave written informed consent prior to participation with the right to exit the study at any time. The study was approved by the local ethics committee (University College London, UK). Participants received £10 per hour as reimbursement. Task ~~~~ Presentation® software (Neurobehavioral Systems, www.neurobs.com) was used to generate all stimuli and to program the paradigm. The task was presented on an LCD screen (vertical refresh rate 60 Hz) on a white background. The experiment was run in a dark and sound- proof room. Participants were seated 75 cm in front of the screen and kept their chin on a chin rest. Sound was presented binaurally at 74 dB SPL via Sennheiser HD 380 pro headphones. The task (Fig. 1) was a variation of the temporal binding paradigms used by Haggard et al. (2002) and Moore and Haggard (2008). On each trial, participants pressed a button with their right index finger once, at a time of their choice, and verbally reported the time of this button press with respect to the rotating clock hand of a Libet clock (Libet et al., 1983) (period of 2.55 s) at the end of the trial. The experiment was based on a 2 × 2 factorial design, with one factor (“trial type”, two levels: “action + tone” vs. “action only”) varied on a trial-by-trial basis and the other (“outcome timing variability”: “variable outcome timing” vs. “fixed outcome timing”) varied block-wise. There was an additional, blocked condition that served as baseline for subjective time judgements (“baseline” blocks). 75% of trials in “variable outcome timing” and “fixed outcome timing” blocks were “action + tone” trials, in which participants’ button presses were followed by one of two sine wave tones (750 Hz or 800 Hz, 100 ms duration). Each tone frequency was presented pseudo-randomly in half of these trials. In the remaining 25% of trials, no tone was presented. These trial frequencies of “action + tone” and “action only” trials were identical to Moore and Haggard (2008), where they were used to establish outcome expectation and quantify prospective action binding in trials of unexpected outcome omission. Here, they served to control for potential differences in prospective binding due to “outcome timing variability”: Stronger binding in “action + tone” trials compared to “action only” trials can be attributed to a re-construction triggered by the tone over and above any prospective shift of subjective action time. Tones followed the button presses either after a fixed delay of 250 ms (in “fixed outcome timing” blocks) or after a variable delay of either 50 ms, 250 ms or 450 ms (with equal probabilities) (in “variable outcome timing” blocks). At the end of “action + tone” trials, participants verbally reported whether the tone they had heard was a high-pitch or low-pitch tone, in addition to reporting the time of their button press. This incidental pitch discrimination task was introduced to prevent participants from ignoring the tones. There were three “fixed outcome timing” blocks and three “variable outcome timing” blocks, presented consecutively, with the order of the two conditions fully counterbalanced across subjects. Each of these blocks consisted of 24 trials. Outcome timing with respect to the action (fixed or variable) was made explicit at the beginning of each block. In addition, there were two “baseline” blocks, consisting exclusively of “action only” trials. These blocks were presented before and after the main experimental blocks and consisted of 9 trials each. Before the first block in which tones were played, participants were familiarized with the two tone frequencies by listening to them in an alternating sequence for as long as they wanted to. At the beginning of each trial, the clock hand started to rotate at a random position. Participants were instructed not to press the button during the first rotation of the clock hand. Rotation of the clock hand continued over a total of 350–400 vertical refresh frames and then stopped, largely independently of when participants chose to press the button. Only if participants pressed throughout the last 60 frames before the planned end of rotation, this rotation was extended for 60 more frames after sound presentation (75 frames after the button press in “action only” trials), so that the clock hand never stopped directly or very shortly after the button press. If participants did not press the button at all before the clock hand stopped rotating, a red “x” (font size: 1°) was presented in the centre of the screen, indicating a miss. These trials were repeated at the end of the respective experimental block. During the inter-trial interval, the clock was presented without the clock hand for 1000–1500 ms. Participants were discouraged from anticipating the clock hand position when pressing the button and from pressing in a stereotyped way. They were asked to fixate a stationary black circle at the centre of the clock at all times, rather than to follow the rotating clock hand with their eyes. Participants were instructed to report the clock hand position only after the clock had stopped rotating and a question mark (font size: 1°) had replaced the clock hand. Lastly, they were asked to be as accurate as possible in their timing reports and to use all integer numbers from 0 to 59, rather than just the multiples of five that were explicitly labelled around the clock face. After verbally reporting their subjective time judgement, the experimenter typed the value into a keyboard. The image of the clock consisted a circular annulus (outer diameter: 3°, inner diameter: 2.65°), presented at the centre of the screen, twelve radial lines, which represented the five-minute markers (length 0.5°, line width .075°) and which were labelled with the multiples of five between zero and 55 (font size: 0.5°) at 2.15° eccentricity, a central, solid black circle (diameter: 0.3°) and one clock hand (line width: 0.15°, length: 1.4°). All visual elements of the clock were black, the inside of the annulus was white. The clock hand rotated 2.353° upon every vertical refresh (∼17 ms), resulting in one full rotation every 2550 ms. Exact timing of picture stimulus presentation was monitored and trials were repeated at the end of the respective experimental block if at least one picture was delayed by one frame or more.","Our main focus was on participants’ time judgement errors, i.e., the difference between the subjective, reported time of the action and its objective time with respect to the Libet clock. Following previous studies on binding of subjective action time (Haggard et al., 2002), we subtracted the mean time judgement error in “baseline” blocks from mean time judgement errors in “variable outcome timing” and “fixed outcome timing” blocks for each participant. Positive time judgement errors therefore indicate temporal binding of the action by its (expected) outcome. We focussed on subjective time in “action only” trials vs. “action + tone” trials in which event timing was identical in “fixed outcome timing” and “variable outcome timing” blocks, i.e., on “action + tone” trials with an action – outcome delay of 250 ms. In addition, we report results for “action + tone” trials in “variable outcome timing” blocks that were based on the two alternative action – outcome delays separately (50 ms and 450 ms). To compare attentional demand across conditions, we analysed the variability (standard deviation) of time judgement errors (cf. Haggard et al., 2002) and sensitivity (d′) in the incidental pitch discrimination task across participants. A two-way repeated-measures analysis of variances (ANOVA) was used in all cases, followed by pairwise dependent-samples t-tests.","We were primarily interested in changes in retrospective binding due to contextual information on outcome timing variability, not due to changes in bottom-up drive. Therefore, we first focused on “action + tone” trials with identical trial event timing in both “variable” and “fixed” outcome timing blocks, i.e., on trials with an action – outcome delay of 250 ms. We compared these “action + tone” trials to “action only” trials in the two types of blocks (Fig. 2). A two-way repeated-measures ANOVA (“outcome timing variability” × “trial type”) revealed a significant interaction effect (F1,15 = 6.84, p = .019) and a main effect of “trial type” (F1,15 = 8.5, p = .011) on mean time judgement errors across subjects. Post-hoc dependent-samples t-tests showed that the interaction was due to a significant shift of subjective action time forwards in “action + tone” trials vs. “action only” trials in “variable outcome timing” blocks (mean ± SD: 31.7 ± 27.4 ms (“action + tone”) vs. 16 ± 24.9 ms (“action only”); t15 = −3.88, p = .001), an effect that was absent in “fixed outcome timing” blocks (mean ± SD: 23.3 ± 21.8 ms (“action + tone”) vs. 23.3 ± 21.5 ms (“action only”); p = .99). No other pairwise t-test was significant, suggesting that the observed main effect of “trial type” was predominantly driven by the difference between “action only” trials and “action + tone” trials in the ”variable outcome timing” block. Our paradigm required attention to the pitch of the tone in order to prevent participants from ignoring the tone. In order to test for potential differences in attention towards the tone across conditions, we analysed the standard deviation of time judgement errors, which has previously been used as an indicator of varying attention or, more generally, quality of information processing in action binding tasks (Haggard et al., 2002). “Outcome timing variability” and “trial type” showed no significant main effects or interaction effects on the standard deviation of time judgement errors across subjects (mean ± SD: 65.5 ± 19.2 ms (“action only” trials in “fixed outcome timing” blocks), 66.8 ± 22.5 ms (“action + tone” trials in “fixed outcome timing” blocks), 70.1 ± 24.4 ms (“action only” trials in “variable outcome timing” blocks), 65.4 ± 26.5 ms (“action + tone” trials in “variable outcome timing” blocks). Furthermore, there was no significant effect of “outcome timing variability” on pitch discrimination sensitivity (mean d′: 2.52 (“fixed outcome timing”) and 2.34 (“variable outcome timing”), t14 = 1.22, p = .24). For completeness, we report the shifts observed for “action + tone” trials in “variable outcome timing” blocks in which the action – outcome delay (and, therefore, bottom-up drive) differed from “fixed outcome timing” blocks. Compared to “action only” trials in “variable outcome timing” blocks, subjective action time in “action + tone” trials for which the action – outcome delay was 50 ms or 450 ms did not show the shift towards the outcome observed for the action – outcome delay of 250 ms. When the tone followed the button press after 50 ms, there was even a significant shift away from the outcome when compared to “action only” trials (mean ± SD in “action + tone” trials: −3.5 ± 18.2 ms; t15 = 2.92, p = .011) while there was no significant change due to outcome presentation 450 ms after the button press (mean ± SD in “action + tone” trials: 12.1 ± 22.9 ms; t15 = 0.81, p = .43).","We demonstrate that an essential aspect of action awareness – the experienced time of an action – can be re-constructed by a context-dependent process at the time of the action’s outcome. Specifically, we observe this re-construction of action awareness when the action is performed in a context of uncertain outcome timing. This context-dependency suggests a model-based, top-down re-constructive process that shifts action awareness based on previously experienced action – outcome intervals. If the reconstructive process had been directly triggered by the perceptual trace of the tone, and was independent of the action – outcome delay experienced on other trials, then it could be interpreted as a bottom-up backwards influence. This would be similar to the effects of backwards masking on conscious perception (Werner, 1935), and would not track changes in the time relation of an action and a subsequent stimulus. Context-dependency makes this bottom-up account unlikely. The majority of previous studies on retrospective action binding have compared conditions that differed in bottom-up drive (“action only” vs. “action + tone” conditions) (Moore & Haggard, 2008; but see Moore et al., 2009). Therefore, these studies cannot exclude the possibility that a retrospective shift of subjective action time is due to bottom-up interference with action awareness or, at a lower level, with the integration of the visual information of the Libet clock. Despite the possibility of such bottom-up driven interference of perceptual codes with action awareness that is unrelated to a representation of the action – outcome association, retrospective action binding has been considered to reflect implicit post-hoc inference on agency (Moore & Haggard, 2008; Voss et al., 2010). To avoid the difference in bottom-up drive present in previous studies, and to test whether re-construction of action awareness indeed tracks the temporal relation between the action and its outcome, we compared “action only” and “action + tone” trials in two contexts. These contexts differed with respect to temporal outcome variability while the structure and timing of trial events (of interest) was identical across contexts. The interaction between “trial type” (outcome presented vs. omitted) and “outcome timing variability” we observed suggests that a retrospective shift of subjective action time in our study results from an integration of sensory and contextual information on an action – outcome association. This context-dependency speaks in favour of a top-down process underlying retrospective action binding, which integrates prior knowledge and sensory re-afferences. Compared to previous studies on retrospective binding, notably by Moore and Haggard (2008), our study design introduces an additional aspect of outcome uncertainty. Moore and Haggard exclusively manipulated the probability of outcome occurrence while we also varied outcome timing with respect to the action. Our current results can neither exclude a retrospective binding component in “action + tone” trials when the action – outcome delay is fixed, nor, indeed, in “action only” trials, as the omission of an expected outcome in these trials may well constitute a sensory event (Sutton, Tueting, Zubin, & John, 1967) and thus alter action awareness retrospectively. In addition, since the probability of outcome occurrence was high across both operant conditions, we cannot exclude a prospective shift of subjective action time from the baseline condition to the operant conditions. However, our key finding is an enhancement of action binding in “variable outcome timing” blocks from “action only” trials to “action + tone” trials, i.e., a clear retrospective process that follows outcome presentation. In addition, we find no change in action binding upon outcome omission as a function of temporal predictability. This may suggest that, unlike retrospective binding, prospective action binding is more closely associated with the expected occurrence than the expected timing of an outcome. However, this negative result should be interpreted with caution. Our finding points to a stronger re-constructive component when outcome occurrence and timing are uncertain. This suggests that sensory evidence of the action – outcome association is weighed flexibly, depending on the temporal predictability of the outcome. As outcome timing becomes less predictable, sensory evidence is given a relatively stronger weight than outcome prediction and is used for a retrospective adjustment of action awareness. An interesting question is whether temporal outcome uncertainty interacts with specific action – outcome delays during this re-constructive process. In theory, enhancement of retrospective binding by temporal uncertainty could depend on a central estimate of the distribution of previously experienced action – outcome delays. Alternatively, enhanced retrospective binding could reflect time constants inherent to sensorimotor processing or an absolute time range over which an action is a priori expected to cause a sensory event. The present study cannot distinguish between these alternatives. Here, we find enhanced retrospective binding for an action – outcome interval of 250 ms but not for trials with shorter or longer delays (50 ms and 450 ms). Trials with shorter and longer action – outcome delays were necessary to build a context of temporal uncertainty in the “variable outcome timing” condition. These trials had no counterparts in the “fixed outcome timing” condition which could have confirmed that subjective action time is, in principle, shifted under operant conditions for these delays. Indeed, a recent study found no action binding at a fixed action – outcome delay of 50 ms at all (Walsh & Haggard, 2013). We know of no previous study that has systematically examined action binding across a broad range of action-outcome delays (whereas outcome (tone) binding (Haggard et al., 2002) and compound measures, which do not distinguish between tone binding and action binding (e.g. Humphreys & Buehner, 2009), have been studied across a broad range of action-outcome intervals). Because of this, we are cautious to draw strong conclusions regarding action binding in trials with an action – outcome delay of 50 ms or 450 ms in the “variable outcome timing” condition. In principle, our study design confounds two aspects of a contingent action – outcome association, namely temporal predictability and temporal control (Hughes, Desantis, & Waszak, 2012). Recent evidence supports the view that temporal binding reflects (voluntary) temporal control rather than action-independent temporal predictability (Desantis, Hughes, & Waszak, 2012; Haggard et al., 2002). Temporal binding is specific for action – stimulus associations and does not occur when one stimulus predicts another. By design, we cannot unambiguously attribute our finding of retrospective action binding to either temporal control or temporal predictability. This question goes beyond the purpose of this study, which was to distinguish between bottom-up interference and model-based inference in retrospective binding, and requires attention in future studies. A contextual influence on retrospective action awareness has been demonstrated in one previous study (Moore et al., 2009). This study varied the probability of a sensory event in trials in which no motor action was performed, thereby modulating the action – outcome contingency in a contextual manner, as in our study. The authors reported retrospective binding under conditions of high, but not low action – outcome contingency, i.e., evidence for an interaction between sensory evidence and context that is compatible with our own finding. Interestingly, we find enhanced retrospective binding when the outcome is not fully determined by the preceding action due to temporal uncertainty. This finding is in agreement with studies that describe retrospective action binding as a post-hoc adjustment of the subjective time of an action upon unpredictable outcome occurrence. From a theoretical perspective, this adjustment seems unnecessary when outcome predictability is high, i.e., when internal models are accurate, as in our “fixed outcome timing” condition. Under conditions of low temporal predictability, however, an internal model of outcome timing can, in principle, be used to evaluate an experienced action-outcome association retrospectively in relation to prior knowledge about the distribution of action-outcome intervals. The probability and timing of a sensory event given an action may affect re-construction of action awareness in a way that is distinct from the effects of contingency as shown by Moore et al. (2009). In summary, we demonstrate that retrospective action awareness results from a model-based, top-down process, rather than from bottom-up interference with action awareness. This re- constructive process integrates contextual and sensory information to infer on the action – outcome association and may contribute to establishing action – outcome associations under conditions of temporal outcome uncertainty."],["Reading and spelling abilities are thought to be highly correlated during development, and orthographic knowledge is assumed to underpin both literacy skills. Interestingly, recent studies showed that reading and spelling skills can also dissociate. The current study investigated whether spelling skills (indicating orthographic knowledge) are associated with the application of orthographic strategies during reading. We examined eye movements of 137 third- and fourth-graders who were either good or poor readers with or without a spelling deficit: 43 children with typical reading and spelling skills, 28 with isolated spelling deficits, 28 with isolated reading deficits, and 38 with combined reading and spelling deficits. Although we expected to find reduced reliance on orthographic reading processes among poor spellers, this was evident for the group with combined deficits only. Both isolated deficit groups applied sublexical and lexical processes in a similar amount to typically developing children. Our findings suggest that reading rests on orthographic strategies even if lexical representations are poor as indicated by a deficit in spelling skills. Findings also show that dysfluent reading does not result only from overreliance on decoding. --------------------------------------------------------------------------------","The importance of word-specific orthographic knowledge for spelling is beyond question because the precise letter sequence of a word must be fully recalled. However, word- specific orthographic representations are crucial not only for spelling but also for accurate and fluent reading (Ehri, 1992, 2005; Share, 1995). Thus, word-specific orthographic knowledge contributes to both spelling and reading (Conrad, Harris, & Williams, 2013; Rothe, Cornell, Ise, & Schulte-Körne, 2015). Moreover, findings from cognitive behavioral studies (Angelelli, Marinelli, & Zoccolotti, 2010; Burt & Tate, 2002; Monsell, 1987) as well as from neuroimaging research (Purcell, Jiang, & Eden, 2017; Rapp & Dufor, 2011; Rapp & Lipka, 2011) favor the view of a single orthographic lexicon (e.g., Behrmann & Bub, 1992) over independent lexica (e.g., Weekes, 1996) for reading and spelling. Thus, reading and spelling are supposed to share the same orthographic representations rather than relying on different orthographic representations (for reviews, see Jones & Rawson, 2016; Purcell et al., 2017).1 In line with this view, reading and spelling skills were reported to be highly correlated (Swanson, Trainin, Necoechea, & Hammill, 2003) and reading deficits were found to be often accompanied by spelling deficits (Angelelli, Judica, Spinelli, Zoccolotti, & Luzzatti, 2004; Landerl & Moll, 2010; Moll, Kunze, Neuhoff, Bruder, & Schulte-Körne, 2014). The pivotal role of orthographic representations for reading is also stated by the lexical quality hypothesis (Perfetti, 2007; Perfetti & Hart, 2001). According to this theory, skilled reading builds on high- quality representations integrating knowledge about a word’s phonological, orthographic, and semantic characteristics. In support of this view, reading speed was found to be affected by the quality (i.e., accuracy and stability) of orthographic representations, as indicated by spelling performance (Martin-Chang, Ouellette, & Madden, 2014; Ouellette, Martin-Chang, & Rossi, 2017). However, reading is thought to be easier and to involve less processing than spelling because in most alphabetic orthographies grapheme–phoneme correspondences are more consistent than phoneme–grapheme correspondences (Bosman & Van Orden, 1997). Furthermore, spelling requires retrieving the complete letter array from mind, whereas for reading recognition of printed letter strings is sufficient (Perfetti, 1997). Hence, spelling may require more precise orthographic representations than reading (Perfetti, 1992). Accordingly, isolated deficits in spelling in spite of adequate reading skills are generally acknowledged (ICD-10 [International Classification of Diseases–10th Revision]; World Health Organization, 2000). Interestingly, however, several studies across different languages reported dissociations in both directions: spelling deficits despite adequate reading skills as well as reading fluency deficits despite adequate spelling skills (Bar-Kochva & Amiel, 2016; Fayol, Zorman & Lété, 2009; Manolitsis & Georgiou, 2015; Moll, Kunze, et al., 2014; Moll & Landerl, 2009; Torppa, Georgiou, Niemi, Lerkkanen, & Poikkeus, 2017; Wimmer & Mayringer, 2002). Furthermore, different cognitive constructs were found to underpin spelling and reading fluency; whereas phonological awareness (PA; the ability to segment and manipulate speech sounds; Vellutino, Fletcher, Snowling, & Scanlon, 2004), is more strongly related to spelling, rapid automatized naming (RAN; the ability to quickly name aloud visual material; Denckla & Rudel, 1976) is more strongly linked to reading fluency (Furnes & Samuelsson, 2011; Moll, Ramus, et al., 2014; Vaessen & Blomert, 2013). Assuming that reading and spelling rely on the very same word- specific orthographic representations, the existence of marked dissociations between reading and spelling skills is not obvious. Investigating such dissociations provides an excellent opportunity to gain important insights into the role of orthographic knowledge during reading. For isolated spelling deficits, two explanations have been suggested; both assume that individuals with isolated deficits in spelling have problems in building up well-specified orthographic representations, but they differ in their explanation as to how this deficit is compensated for in reading. Frith (1980) argued that underspecified representations are sufficient to recognize words during reading even if they are not exact enough for spelling. Based on that, Frith assumed that isolated poor spellers rely on a partial cue reading strategy, that is, reading words by sight based on degraded orthographic representations without considering the exact letter-by-letter structure. In line with this interpretation, isolated poor spellers showed good performance when reading high-frequency words but made substantially more errors than typical readers in nonword reading. A different explanation was presented by Moll and Landerl (2009) based on German- speaking children. In their study, isolated poor spellers showed age-adequate reading performance for words as well as nonwords. More important, they did not show the typical advantage of real words over corresponding pseudohomophones (i.e., unfamiliar letter strings that sound like real words, e.g., rane for rain in English), which is taken as an indication of lexical access of the real word spelling (see also Francuz & Borkowska, 2013). Interestingly, they nevertheless showed adequate reading latencies compared with age-matched typically developing children even for words they could not spell. Therefore, it was assumed that they relied on highly efficient decoding strategies2 to compensate for their deficient orthographic knowledge, which is possible in an orthography with high consistency of grapheme–phoneme correspondences such as German. The current study went one step further by analyzing eye movement patterns in order to investigate to what extent children with isolated spelling deficits apply lexical processes during reading. In contrast to isolated deficits in spelling, isolated reading fluency deficits are hardly described in the literature, especially for opaque languages (but see Lovett, 1987). The phonological deficit view of dyslexia assumes that deficits in reading fluency result from slow and effortful grapheme–phoneme decoding compensating for lack of orthographic representations (Vellutino et al., 2004). However, recent findings suggest that poor readers do use lexical strategies during reading and do not simply rely on decoding (e.g., Gangl et al., 2018; Paizi, De Luca, Zoccolotti, & Burani, 2013). An alternative suggestion was presented by Wimmer (1993), who argued that dysfluent reading may result from inefficient visual–verbal access that affects both sublexical and lexical reading strategies. According to this view, dysfluent readers may have intact orthographic representations available and may even access them during reading, but in an inefficient way. Thus, children with isolated reading deficits may apply orthographic strategies but still show dysfluent reading due to delayed lexical access. In line with this interpretation, Moll and Landerl (2009) reported marked reading advantages of words over pseudohomophones for German-speaking poor readers, indicating orthographic access for these words. Nevertheless, overall reading times were seriously prolonged even for correctly spelled words, which was interpreted as delayed access to verbal word forms from existing visuo-orthographic representations. This visual–verbal access deficit view was further supported by studies showing that children with isolated reading deficits have marked deficits in RAN (e.g., Moll & Landerl, 2009; Torppa et al., 2017; Wimmer & Mayringer, 2002), which also requires fast and efficient activation of verbal word forms from the presented visual information (e.g., pictured objects or digits). In summary, studies on isolated deficits suggest that the extent to which orthographic knowledge is accessed during reading is associated with an individual's spelling ability, which provides a direct indication of orthographic knowledge. Investigating marked dissociations between reading and spelling skills, as in children with isolated disorders of reading or spelling, is a promising approach because specific hypotheses can be derived and tested; isolated poor spellers’ lack of orthographic knowledge should be reflected in reading either as access of underspecified representations (Frith, 1980) or as reduced reliance on lexical strategies (Moll & Landerl, 2009). In contrast, isolated poor readers’ unimpaired spelling may indicate that available orthographic knowledge is (inefficiently) accessed during reading. The current study for the first time systematically examined reading strategies of good and poor readers with and without spelling deficits by means of eye movement recording. Eye tracking has the major advantage that reading processes can be investigated “online” while they take place (Rayner, 1998), enabling us to assess the use of orthographic knowledge during reading more closely. In particular, we addressed the question of whether differences in orthographic knowledge, as indicated by spelling (dis)abilities, are associated with differences in accessing orthographic information during reading. Furthermore, we were interested in whether the same association holds for both good and poor readers. Do good readers rely on orthographic reading strategies even if they have an obvious deficit in orthographic spelling knowledge? Do poor readers apply different reading strategies according to differences in their orthographic spelling knowledge? In addition to eye-tracking measures, we also assessed cognitive constructs that have been reported as risk factors for specific learning disorders—PA, RAN, and verbal memory (e.g., Landerl et al., 2013; Moll, Göbel, Gooch, Landerl, & Snowling, 2016). We aimed to expand on earlier evidence that PA may be more strongly related to spelling (deficits), whereas RAN is most strongly related to (deficits in) reading fluency (Moll & Landerl, 2009; Moll, Ramus, et al., 2014; Wimmer & Mayringer, 2002). Corresponding differences in cognitive profiles for children with isolated reading or spelling deficits would be informative with respect to possible causes underlying the observed dissociations. Eye movements were recorded while children read short and long words, pseudohomophones, and nonwords. First, we were interested in children’s eye movement patterns during word reading because words can be read either by a decoding strategy or by an orthographic strategy (at least in consistent orthographies such as German), depending on whether or not word-specific orthographic representations are available (Coltheart, Rastle, Perry, Langdon, & Ziegler, 2001). We assessed a wide range of eye movement variables, which are indicative of lexical or sublexical reading strategies and which were found to differentiate between good and poor readers (Hawelka, Gagl, & Wimmer, 2010). We expected marked group differences if children rely on different reading strategies due to differences in orthographic knowledge. Namely, a sublexical decoding-based reading strategy should be reflected in short forward saccades interspersed with high numbers of fixations (De Luca, Borrelli, Judica, Spinelli, & Zoccolotti, 2002; Hawelka et al., 2010; Krieber et al., 2016). Furthermore, it should be rather exceptional that a word receives only one fixation because a dominance of singly fixated words would rather indicate a lexical reading strategy of directly accessing orthographic representations (Hawelka et al., 2010). We also examined marker effects of sublexical and lexical reading strategies on (a) gaze duration (i.e., the sum of all fixation durations on an item prior to moving to the next item, which is called “first pass”), (b) number of fixations (i.e., the sum of all fixations on an item during first-pass reading), and (c) landing position (i.e., the location of the first fixation on an item). Important markers of sublexical strategies are effects of length, that is, a systematic increase in processing load with increasing item length. Sublexical processing is reflected in longer gaze duration and higher number of fixations for long items compared with short items (e.g., De Luca, Di Pace, Judica, Spinelli, & Zoccolotti, 1999; Hawelka et al., 2010; Hutzler & Wimmer, 2004; Rau, Moeller, & Landerl, 2014). Indeed, particularly marked effects of length are reported for nonwords (e.g., De Luca et al., 2002; Hutzler & Wimmer, 2004; Rau et al., 2014). Sublexical processing is further reflected in the landing position by the constant tendency to fixate word beginnings irrespective of length (e.g., Hawelka et al., 2010; MacKeben et al., 2004). As marker effects of lexical processing, we analyzed word superiority effects in terms of word-pseudohomophone and lexicality effects, that is, an advantage for words compared with pseudohomophones and nonwords, respectively. More efficient processing of real words indicates availability of and access to orthographic representations for these letter strings (e.g., Di Filippo, De Luca, Judica, Spinelli, & Zoccolotti, 2006; Francuz & Borkowska, 2013; Grainger, Spinelli, & Ferrand, 2000; Moll & Landerl, 2009). If reading is mostly based on sublexical decoding, no marked differences would be evident between different types of stimuli in gaze duration, number of fixations, and landing position. We assumed that readers adapt their reading strategies according to their orthographic knowledge, reflected by their spelling performance. Children with age-adequate spelling performance (i.e., typically developing and isolated poor readers) were assumed to use orthographic reading strategies, which should be reflected in small length effects but marked word superiority effects compared with pseudohomophones and nonwords on gaze duration and number of fixations. Furthermore, we expected good spellers to adapt the landing positions to the length of words. However, in line with previous findings, isolated poor readers were assumed to experience delayed visual–verbal access, inducing prolonged fixation durations compared with typically developing children despite an adequate number of fixations. For poor spellers (i.e., isolated poor spellers and children with combined reading and spelling deficits), we were particularly interested to see whether we would find evidence for reduced orthographic strategies and overreliance on decoding (independent of their reading performance). If poor spellers predominantly rely on decoding strategies, we expected them to show marked length effects but reduced word superiority effects on gaze duration and number of fixations. Furthermore, they should fixate the beginnings of letter strings irrespective of length and stimulus type. We also expected differences between the two groups of poor spellers. Whereas isolated poor spellers’ adequate reading might result from a highly efficient compensatory decoding strategy, children with combined deficits in reading and spelling may be less efficient decoders due to an additional impairment in their visual–verbal access. Thus, prolonged fixation durations and increased numbers of fixations for words were expected for children with combined reading and spelling deficits compared with typically developing children.","Good and poor readers with and without a spelling deficit were selected based on an extensive screening procedure with 4123 German-speaking children at the end of third grade or beginning of fourth grade. Data were collected at two collaborating sites, Graz (Austria) and Munich (Germany). Spelling and reading fluency skills were assessed by standardized classroom tests (Müller, 2004; Wimmer & Mayringer, 2014; all tests are described in the next section). A deficit in spelling and/or reading fluency was assigned when test performance was at or below the 20th percentile. Scores at or above 25th percentile were defined as age-adequate performance. Because none of the children with isolated deficits had reading or spelling scores above the 75th percentile in the nonaffected domain, children with scores higher than the 75th percentile in any of the tests were not considered for the typically developing control group. Reading fluency performance was further assessed by an individually administered 1-min word and nonword reading test (Moll & Landerl, 2010). A deficit in reading was confirmed by scores below the 20th percentile on at least one subtest and below the 25th percentile on the other subtest. All children had normal or corrected-to-normal vision, a nonverbal IQ of at least 85, no clinical diagnosis of attention-deficit/hyperactivity disorder (ADHD), and an unremarkable score (see below) on a parental questionnaire for attention deficits (Döpfner, Görtz-Dorten, Lehmkuhl, Breuer, & Goletz, 2008). Overall, 137 children who fulfilled the criteria and received parental consent participated in the study, which was approved by the ethics committees of the University of Graz and the University of Munich. The sample consisted of 43 typically developing children (TD group), 28 children with isolated spelling deficits (SD group), 28 children with isolated reading deficits (RD group), and 38 children with combined reading and spelling deficits (RSD group). Reading fluency The standardized classroom reading fluency test (SLS 2–9 [Salzburger Lese- Screening für die Schulstufen 2–9]; Wimmer & Mayringer, 2014; parallel test reliability > .86 according to manual) comprises simple sentences (e.g., “Trees can speak.”), which children needed to read silently and mark as semantically right or wrong by circling a checkmark or a cross after each sentence. The number of correctly marked sentences within 3 min was scored. Reading fluency was further assessed by an individually administered 1-min reading fluency test (SLRT-II [Lese- und Rechtschreibtest]; Moll & Landerl, 2010; parallel test reliability between .90 and .95 according to manual) containing a word reading list and a nonword reading list. Children were asked to read aloud each list as fast as possible without making errors within a time limit of 1 min. The number of correctly read items was determined. Spelling The standardized classroom spelling test (DRT-3 [Diagnostischer Rechtschreibtest für 3 Klassen]; Müller, 2004; split-half reliability = .95 according to manual) consists of 44 words, varying in length and difficulty, that need to be written into the gap of sentence frames. The experimenter first dictated each word, then the full sentence was read out, and finally the word was named again. The number of incorrect spellings was scored. Attention rating A standardized questionnaire (Döpfner et al., 2008) made up of 20 items with a 4-point rating scale regarding symptoms of inattention (9 items), hyperactivity (7 items) and impulsivity (4 items) was given to the parents. A score higher than 1.10 for girls, and a score higher than 1.55 for boys, is suggestive of an AD(H)D diagnosis and was applied as an exclusion criterion in the current study (see description of participants). Nonverbal IQ The first part of the German version of the Culture Fair Intelligence Test (CFT 20-R; Weiß, 2006; reliability = .92 as described in the manual), including the four subtests Series, Classification, Matrices, and Topology, was given as an estimate of nonverbal IQ. Verbal memory The subtest Digit Span from the German version of the Wechsler Intelligence Scale for Children (Petermann & Petermann, 2011) was used to assess the verbal short- term and working memory. Rapid automatized naming Children named a matrix made up of 40 monosyllabic digits (8, 3, 5, 2, 9) arranged in five columns and eight rows as quickly and accurately as possible. Item order was randomized, and each item was presented once in each row. The experimenter recorded the time children needed to name all 40 digits as well as any errors. The number of items named correctly per second was determined. Phonological awareness A computerized phoneme deletion task was programmed with Presentation 16.3 (Neurobehavioral Systems, Berkeley, CA, USA). The task comprised 4 practice trials and 25 test trials (20 monosyllabic nonwords and 5 disyllabic nonwords) that were presented via headphones. Children needed to repeat each nonword first and then to pronounce it without a specified phoneme (e.g., “/fɔlt/ without /t/”–/fɔl/). The experimenter marked the response correctness by mouse clicking. Any nonword that was not repeated correctly was replayed up to two times. If it was still not pronounced correctly, it was discarded from analysis (0.9% of all items). The ratio of correct responses to the total number of responses was scored. Cronbach’s alpha was .79.","The eye movements of the children’s dominant eye were recorded inside a dimly lit room with an EyeLink 1000 Tower Mount eye tracker in Graz and an EyeLink 1000 Plus Desktop Mount eye tracker (SR Research, Toronto, Ontario, Canada) in Munich. To minimize head movements, the forehead rest was used. The experiment was run with Experiment Builder software (SR Research, Version 1.10.1241) on a 20-inch monitor (120-Hz refresh rate, 1024 × 768 resolution) and a 15.6-inch monitor (120-Hz refresh rate, 1280 × 960 resolution) in Graz and Munich, respectively. At both collaborating sites, the stimuli were presented identically with an uppercase letter height of about 0.62° of visual angle at a viewing distance of 65 cm. A spatial resolution of less than 0.5° of visual angle was achieved by using a standard 9-point calibration. Stimuli and design The experiment comprised 240 stimuli: 80 high-frequency words (mean absolute frequency of 1537.80 for 9- and 10-year old children according to the corpus of the database childLex [Schroeder, Würzner, Heister, Geyken, & Kliegl, 2015]) that were either short (three to five letters) or long (six to nine letters), 80 pseudohomophones, and 80 nonwords. Note that German has highly consistent letter–sound correspondences, but sound–letter correspondences are often inconsistent. The words were selected so that they included sounds that can be transcribed with different letters/clusters. This allowed us to create a pseudohomophone for each word by exchanging one phonologically identical grapheme. The pronounceable nonwords were derived from the words by exchanging one consonant or vowel grapheme per syllable for another consonant or vowel grapheme. The stimuli were matched on number of letters and bigram and trigram frequency (Fs < 1.5, ps > .20) according to childLex (Schroeder et al., 2015). All items and item characteristics are displayed in Appendix A. Stimuli were assigned to three blocks of 80 items each. Two blocks comprised words and pseudohomophones in equal parts, and the third block comprised nonwords only. Four pseudorandomized orders were created with the restriction that at most two words or pseudohomophones appeared in immediate succession and corresponding words and pseudohomophones did not appear together in the same block. Procedure Prior to the eye-tracking experiment, children were asked to spell the 80 word stimuli of the eye-tracking paradigm. Words were dictated by the experimenter, and children wrote them into sentence frames that provided adequate semantic content to each word. The number of correct spellings was scored (errors of capitalization were not considered because they constitute a specific grammatical error category in German). Because we selected high-frequency words, we guaranteed that even poor spellers could write a reasonable amount of these stimuli correctly (showing that they had orthographic representations available). This allowed us to investigate to what extent poor spellers access their existing orthographic knowledge during reading. On a separate day, children took part in the eye-tracking experiment to avoid recognition effects. The four equivalent versions of the eye-tracking experiment were randomly assigned to the participants. Each one started with a word/pseudohomophone block, followed by the nonword block and the second word/pseudohomophone block. Short breaks were provided between the blocks and after the first half of the nonword block. Each block was introduced with practice items. Feedback was given only during practice. Items were displayed in Arial font and appeared arranged in single lines separated by single spaces in black on a white background in the center of the computer screen after children had fixated a left-sided yellow smiley, which was linked to a fixation trigger, for at least 250 ms. Recalibration was done if no fixation was detected within 5000 ms, and afterward the experiment continued from the interrupted point. Lines consisted of 10 items each (8 target items and 2 filler items at the beginning and the end), and the first item of each line was displayed at the location where the smiley was shown. Children needed to read each item aloud without making errors while their eye movements were recorded and reading was tape-recorded. Reading accuracy was noted by the experimenter. Hesitations and inadequate word stress were not considered as errors. After the children had read a whole line, they were asked to look immediately at a small cross in the lower right corner of the screen. As soon as the cross was fixated, the line disappeared and the next trial started with the smiley on the left side of the screen. Cronbach’s alpha for words, pseudohomophones, and nonwords was high for reading accuracy (> .71) and number of fixations (> .92). Data treatment Overall data loss due to problems with calibration accuracy or because a child did not read the whole line was 2.10%. Following the standard procedure of eye-tracking data treatment (e.g., Jones, Snowling, & Moll, 2016; Moscati, Zhan, & Zhou, 2017), fixations shorter than 80 ms were excluded and fixations deriving from microsaccades (within 0.5° of visual angle) were pooled together. Furthermore, for each eye movement measure, individual outliers (i.e., data deviating more than 2.5 standard deviations from the individual mean of each condition [Length × Stimulus Type]) were removed (2.10%; Liversedge et al., 2016). Each target item was defined as a region of interest, and only correctly read items were considered for analyses. Group characteristics ~~~~~~~~~~~~~~~~~~~~~ A summary of the descriptive and cognitive measures for each group is provided in Table 1. Group differences in the standardized reading and spelling measures reflected our selection criteria; the deficit groups performed comparably to typically developing children in their unimpaired skills and comparably to each other in their impaired skills. Although our cutoff criteria for poor performance were somewhat lenient, Table 1 indicates that all deficit groups were seriously impaired in their affected skills with mean scores around the 10th percentile. Table 1 further shows that the groups were comparable in age, exhibited equally low parental ratings of ADHD symptoms, and did not differ significantly in their nonverbal IQ and verbal memory scores. In line with earlier findings (e.g., Moll & Landerl, 2009; Wimmer & Mayringer, 2002), the two groups of poor readers (RD and RSD) showed similarly poor performance in RAN compared with the TD and SD groups, who did not differ significantly from each other. In the PA task, only the RSD group showed significantly lower scores than all other groups. Reading and spelling accuracy for words, pseudohomophones, and nonwords ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In line with the literature on transparent orthographies (e.g., Wimmer 1993; Wimmer & Hummer, 1990), reading accuracy across groups was generally high for all stimulus types (see Table 2). Nonparametric Kruskal–Wallis tests indicated that reading accuracy scores of the SD group were largely comparable to typically developing children, who performed close to ceiling. Even for the two groups of poor readers (RD and RSD), who performed very similarly (except for lower accuracy for long pseudohomophones in the RSD group compared with the RD group), overall reasonably high reading accuracy rates were observed, although significantly different from typically developing children. Table 2 also shows the percentage of correct spellings of the experimental word stimuli. In line with our group selection, both groups of poor spellers (SD and RSD) showed significantly lower percentages of correct spellings than the two groups of good spellers (TD and RD). Thus, poor spellers were less familiar with the experimental word stimuli and had less exact orthographic representations available than good spellers. However, it should be noted that performance on the experimental task was relatively high; even children of the RSD group spelled on average two thirds of these high-frequency stimuli correctly. This was expected because we selected high-frequency words to ensure that even poor spellers had some orthographic representations available in order to be able to investigate to what extent they accessed orthographic knowledge during reading. Eye movement characteristics during word reading ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In a first step, we investigated children’s eye movements for real words only to get a first impression of their natural reading behavior. Following the approach by Hawelka et al. (2010), we compared the four groups on a large number of eye-tracking parameters that might be informative with respect to the application of lexical versus sublexical reading strategies. Given that reading accuracy rates were high for both short and long words across groups, analyses were based on correctly read words only. The first interesting finding evident from Table 3 is that children in the SD group did not differ significantly from typically developing children in any of the investigated eye-tracking parameters. Thus, contrary to our expectations, we did not find any evidence that they would show a compensatory overreliance on decoding mechanisms. Furthermore, we had expected that the RSD group would show more marked decoding strategies, whereas the RD group was assumed to predominantly rely on orthographic strategies. However, both groups of poor readers (RD and RSD) did show evidence for a more piecemeal reading style; apart from prolonged gaze durations, reflecting their reading fluency deficits, they showed significantly higher numbers of fixations and also more between-word regressions going back to a word after it was left already. They also showed a clearly lower percentage of words that received a single fixation only, suggesting that they were mostly unable to process words directly at a glance. Children with RSD showed more marked impairments than children with RD, with prolonged gaze durations, mean fixation, and single fixation durations compared with all other groups. Furthermore, they tended to show shorter forward saccades compared with typically developing children (p = .052). Across all groups, only between 10% and 15% of all within-word fixations during first pass were regressive (going backward) and groups also did not differ significantly with respect to their first fixation positions (i.e., landing positions) on words. Data analyses Effects of length and word superiority were investigated in the following three eye-tracking parameters: gaze duration as an indicator of temporal processing and number of fixations and landing position as spatial measures. To examine specific effects of length and word superiority, and to avoid spurious interaction effects (i.e., overadditivity effects) due to general performance differences between the groups, individual z-scores were calculated (Faust, Balota, Spieler, & Ferraro, 1999). This seemed important because marked group differences were already observed for word reading. For each variable of interest, individual z-scores were obtained by taking each individual’s scores, subtracting the overall mean across all correctly read items, and dividing it by the overall standard deviation. Thus, z-scores reflect an individual’s performance on an item relative to the performance on all other items (Paizi et al., 2013). The z-scores for the three eye-tracking parameters were analyzed separately by means of linear mixed-effect models (LMMs; for a similar approach see Paizi et al., 2013). These robust analyses are particularly appropriate for the repeated-measures designs we aimed to apply to investigate the marker effects of lexical and sublexical processing; LMMs allow treating participants and items as random effects in a single analysis, and they maintain individual variance by considering the whole data without the need for prior averaging as is known in traditional approaches such as analyses of variance (ANOVAs; Baayen, Davidson, & Bates, 2008). LMMs were run in R (R Core Team, 2017) and were fitted by the mixed function of the afex package (Version 0.16-1; Singmann, Bolker, Westfall, & Aust, 2016), which is based on the lme4 package (Version 1.1-12; Bates, Maechler, Bolker, & Walker, 2015). Group (TD, SD, RD, or RSD), length (short or long), and word superiority (words, pseudohomophones, or nonwords) were entered as fixed effects including interaction terms. Participants and items were entered as crossed random effects, for which we specified random slopes for each within-participants and within-items factor, respectively (see Barr, Levy, Scheepers, & Tily, 2013). To achieve better convergence, we removed random correlations (Barr et al., 2013). Likelihood ratio tests were used to compute p values for all fixed effects as implemented in the mixed function. To decompose significant main and interaction effects, post hoc contrasts were performed using the lsmeans package (Version 2.25-5; Lenth, 2016) and the Benjamini–Hochberg method for adjusting p values (Benjamini & Hochberg, 1995). In Table 4, the LMM results are presented separately for each dependent measure; in addition, LMM estimates are displayed in Fig. 1. Note that due to z-scores, main effects of group are negligible and, therefore, are not mentioned further. Raw scores are reported in Appendix B. Gaze duration If poor spellers (SD and RSD groups) relied strongly on decoding strategies, they should show marked length effects and small word superiority effects. Good spellers (TD and RD groups) were expected to show the reverse pattern. The analysis on gaze duration revealed reliable effects of length and word superiority (see Table 4); prolonged gaze durations were found for long stimuli compared with short stimuli and for pseudohomophones and nonwords compared with words (ps < .001). The significant interaction between length and word superiority indicated more marked length effects for nonwords compared with words and pseudohomophones (ps < .001). The group by length interaction tended to be significant. Post hoc contrasts revealed marked length effects across all groups (ps < .001). However, both groups of poor readers (RD and RSD) were less affected by length than typically developing children (ps < .05). The significant three-way interaction indicated that the reduced length effects in poor readers were limited to nonwords (ps < .001), which was rather surprising. Children with RSD also showed reduced length effects for nonwords compared with children with SD (p < .05). The Group × Word Superiority and Group × Length × Word Superiority interactions were also significant. Post hoc contrasts revealed that the RSD group exhibited smaller word–nonword discrepancies than all other groups for long stimuli (ps < .05). In addition, the RSD group also showed a smaller word advantage relative to pseudohomophones than the SD group (p < .05). These reduced word superiority effects provide the first evidence for overreliance on decoding strategies among children with RSD. Note that the SD group was also found to show smaller word–nonword discrepancies compared with the RD group (p < .05), even though not compared with typically developing children (p > .20). Importantly, however, all four groups showed marked effects of word superiority (i.e., shorter gaze durations for words compared with pseudohomophones and nonwords, ps < .001), indicating orthographic processing. Number of fixations Once again, the main question with respect to poor spellers (SD and RSD groups) was whether we would find evidence of overreliance on decoding strategies as indicated by marked length effects and small word superiority effects. For good spellers (TD and RD groups), the reverse pattern was expected. The LMM analysis on z-scores for numbers of fixations revealed significant main effects of length and word superiority, with higher fixation counts for long stimuli than for short stimuli as well as for pseudohomophones and nonwords than for words (ps < .001). The interaction between length and word superiority was again significant, showing that the word–nonword discrepancy was smaller for short stimuli than for long stimuli (p < .05). Furthermore, effects of length were more marked for nonwords than for words (p < .05). The group by length interaction as well as the three-way interaction failed to be significant, indicating that all groups were similarly affected by length. The Group × Word Superiority interaction was significant. Post hoc contrasts showed that the discrepancy between words and nonwords was significantly smaller in the RSD group than in the TD group (p = .01) and even marginally smaller than in the SD and RD groups (ps < .09). These reduced lexicality effects suggest that children with RSD relied more strongly on decoding strategies than all other groups. However, significant effects of word superiority (i.e., less fixations for words compared with pseudohomophones and nonwords) were obtained once more in each group (ps < .01). This means that even children with RSD relied at least to some degree on orthographic processes during reading. Landing position Reduced lexical strategies among poor spellers should induce similar landing positions at the beginning of letter strings irrespective of item type. Good spellers were expected to adapt their landing position to the length of words (but not nonwords) and, thus, to show marked length effects for those. Again, the main effects of length and word superiority were significant. As expected, shorter landing positions were observed for short stimuli than for long stimuli as well as for nonwords than for words and pseudohomophones (ps < .001). The marginally significant length by word superiority interaction was due to the fact that landing positions on words and pseudohomophones were affected by length (ps < .05), whereas those on nonwords were not (p > .87). The group by length interaction was also significant. Post hoc comparisons revealed smaller length effects on the landing positions of the RSD group compared with all other groups (TD and SD: ps < .001; RD: p = .06). Interestingly, there was no difference in the landing positions between short and long stimuli for the RSD group (p > .87; see Fig. 1). Thus, children with RSD fixated the beginnings of items irrespective of length and stimulus type, whereas all other groups adapted their landing positions to the length of the stimuli. This is a further indication for reduced orthographic reading strategies in the RSD group. The group by word superiority interaction and the three-way interaction were not significant.","The lexical quality hypothesis (Perfetti, 2007; Perfetti & Hart, 2001) postulates that orthographic knowledge plays an important role in spelling as well as efficient reading. Whereas most other studies on orthographic processing have focused on reading only, the current study implemented spelling ability as an indicator of orthographic knowledge. The assumption was that words that can be spelled correctly are represented in detail in the orthographic lexicon. Note that even though German has highly consistent grapheme–phoneme correspondences, the consistency of phoneme–grapheme correspondences is much lower. Thus, to spell words correctly, children must have exact word-specific orthographic knowledge available. The current study investigated the role of orthographic knowledge for reading by exploiting the fact that reading and spelling skills frequently dissociate during development. The first interesting finding was that, contrary to our expectations, children of the SD group showed no indication of overreliance on sublexical decoding strategies during reading due to lack of orthographic knowledge. The only exception was the smaller word superiority effect compared with matched nonwords, in contrast to the RD group in the gaze duration measure. Overall, SD children’s eye movement patterns were highly similar to typically developing children, and SD children also showed comparable effects of length and word superiority. Thus, SD children’s balance between sublexical and lexical reading procedures seems to be age adequate. In particular, Moll and Landerl’s (2009) finding of a nonsignificant word superiority effect compared with matched pseudohomophones, even for words children with SD had spelled correctly, was not replicated in the current group of SD children and was perhaps due to lack of statistical power in their relatively small sample of 14 children. The findings of the current study rather support Frith’s (1980) perspective that children with SD rely on partial cue reading strategies, that is, orthographic reading processes based on low-quality orthographic representations that are sufficient for word recognition but do not allow full retrieval of the correct letter sequence for spelling. However, note that the current findings are also not fully in line with Frith’s interpretation, suggesting that children with SD rely on partial cue reading strategies accompanied by deficits in decoding: Replicating Moll and Landerl’s (2009) findings, children with SD showed age-adequate decoding of nonwords. In part, this may be explained by the fact that in more transparent orthographies, such as German, the highly consistent grapheme–phoneme correspondence rules facilitate the development of good decoding skills. For the RD group, we had predicted evidence for applying these children’s orthographic knowledge indicated by good spelling performance during reading. Indeed, children with RD showed length and word superiority effects that were largely comparable to typically developing children. This finding is interesting with respect to the lexical quality hypothesis (Perfetti, 2007) because it shows that lexical representations that are of sufficient quality for spelling do not necessarily guarantee fast and efficient access during reading. Our evidence also contradicts the view that deficits in reading fluency are simply a consequence of a slow decoding-based reading style due to lack of orthographic knowledge. It rather supports the assumption that fluency deficits arise from a delayed visual–verbal access (Moll & Landerl, 2009; Wimmer, 1993). However, in addition to the age-adequate marker effects and prolonged gaze durations for words, which simply reflected their dysfluent reading style, children with RD also showed more fixations, regressions, and a lower number of words that received only a single fixation than typically developing children. The combination of standard effects of length and word superiority with increased numbers of fixations (forward as well as backward) has not been reported before. Typically, higher fixation counts are assumed to reflect decoding strategies (see Hawelka et al., 2010). However, in the current context, they might just as well be a consequence of RD children’s problems to access their orthographic representations; during a seriously prolonged access time, readers perhaps look harder at the words to be read, which may induce longer fixation durations as well as increased numbers of fixations and regressions. Note that the mean difference in gaze duration between RD and TD children was 205 ms, which easily allows for an additional fixation. A surprising finding was that for nonwords children with reading deficits (both RD and RSD groups) showed smaller length effects on gaze duration than typically developing children. In general, marked effects of length are expected, especially for nonwords among poor readers, reflecting a particular problem with decoding inherent in dyslexia (Rack, Snowling, & Olson, 1992; Wimmer, 1996). However, findings are inconsistent, particularly in transparent orthographies. Whereas some studies reported more marked nonword length effects for dyslexic readers compared with typically developing readers (e.g., De Luca et al., 2002; Hutzler & Wimmer, 2004; Ziegler, Perry, Ma-Wyatt, Ladner, & Schulte-Körne, 2003; Zoccolotti, De Luca, Judica, & Spinelli, 2008), others did not (Di Filippo et al., 2006; Di Filippo & Zoccolotti, 2012). Nevertheless, smaller nonword length effects for poor readers than for good readers have not been reported yet. These reduced nonword length effects in RD and RSD children might be explained by more practice in decoding if this were the dominant reading strategy among poor readers. However, the current study did not find evidence for stronger reliance on decoding, at least not for children with RD. Thus, this explanation seems unlikely. Another possible explanation is that poor readers may have a specific advantage in decoding digraphs and trigraphs at the subword level. This is suggested by an unexpected finding of a Dutch study (Marinus & de Jong, 2010); dyslexic readers read words and nonwords with digraphs faster than matched items without digraphs, whereas proficient readers did not benefit from the presence of digraphs. A possible explanation of this finding is that the slow reading style of poor readers may enable detection of digraphs before activation of incorrect phonemes due to strict letter-by-letter decoding, which would delay the reading process as supposed in the dual route cascaded model (Coltheart et al., 2001). In the current study, the long stimuli comprised more digraphs and trigraphs than the short stimuli. Thus, it is possible that poor readers had a specific advantage in decoding long nonwords, inducing a smaller length effect compared with typically developing children. Finally, the results for the RSD group revealed an overall inadequate eye movement pattern. Similar to the RD group, children in the RSD group showed prolonged gaze durations, increased numbers of fixations, more regressions, and less singly fixated words compared with typically developing children. However, they were clearly more strongly affected in their fixation durations than all other groups (longer gaze durations, mean fixation, and single fixation durations) and tended to show shorter forward saccades than the TD group as well. Importantly, the RSD group was the only group that showed evidence for reduced reliance on orthographic reading processes indicated by reduced word–nonword discrepancies on gaze duration and number of fixations compared with the other groups. This is in line with findings on dyslexic individuals showing reduced effects of lexicality compared with typically developed readers (De Luca, Burani, Paizi, Spinelli, & Zoccolotti, 2010; Di Filippo & Zoccolotti, 2012). Furthermore, children with RSD were found to consistently fixate item beginnings irrespective of length and item type. The findings confirm our assumption that children with RSD rely more on decoding strategies due to insufficient orthographic representations as well as delayed visual–verbal access (see also Hawelka et al., 2010). Nevertheless, note that significant word superiority effects on gaze duration and number of fixations were found among children with RSD as well, that is, shorter gaze durations and less fixations for words compared with pseudohomophones and nonwords, respectively. Thus, even children with RSD showed some spared reliance on orthographic strategies, at least for the easiest items. With respect to associated cognitive deficits, we focused on PA and RAN abilities because previous studies suggest that these are main predictors of spelling and reading skills, respectively. In line with previous findings (e.g., Moll & Landerl, 2009; Torppa et al., 2017; Wimmer & Mayringer, 2002), the current study found marked deficits in RAN among poor readers (both RD and RSD groups). Given that RAN is supposed to measure efficient verbal access from a visual symbol, this finding corroborates the assumption that reading fluency deficits arise from delayed visual–verbal access (Wimmer, 1993). Findings are less clear for the SD group. Whereas some previous studies found PA deficits among isolated poor spellers (e.g., Fayol et al., 2009; Torppa et al., 2017; Wimmer & Mayringer, 2002), the current study and others did not (e.g., Manolitsis & Georgiou, 2015; Moll & Landerl, 2009). The assumption is that children with SD may be able to compensate for an early deficit in PA, at least in more transparent orthographies (Wimmer & Mayringer, 2002; Moll & Landerl, 2009), so that deficits in PA at the end of primary school are mainly evident in children with RSD. This was also observed in the current study and supports Wolf and Bowers’ (1999) double-deficit view of dyslexia. To conclude, our findings suggest that good readers rely on orthographic strategies during reading even if they have deficits in orthographic knowledge as indicated by poor spelling performance. Thus, spelling deficits are not necessarily associated with decoding strategies during reading. Whereas spelling rests on fully specified orthographic representations, reading does not (Perfetti, 1992). However, even if adequate reading builds on orthographic processes, the opposite deduction—that poor readers’ dysfluency results from decoding—seems not to be true. We found strong evidence for orthographic reading strategies among children with RD, and even among children with RSD some spared reliance on orthographic strategies was observed. Hence, dysfluency may be primarily caused by deficient visual–verbal access rather than overreliance on decoding (Wimmer, 1993). Thus, the current study shows that spelling deficits, like reading deficits, are not systematically associated with deficits in lexical reading. Nonetheless, the buildup of high-quality word-specific lexical representations is crucial for the development of appropriate reading and spelling skills."],["Unfamiliar face recognition follows a particularly protracted developmental trajectory and is more likely to be atypical in children with autism than those without autism. There is a paucity of research, however, examining the ability to recognize the same face across multiple naturally varying images. Here, we investigated within-person face recognition in children with and without autism. In Experiment 1, typically developing 6- and 7-year-olds, 8- and 9-year-olds, 10- and 11-year-olds, 12- to 14-year-olds, and adults were given 40 grayscale photographs of two distinct male identities (20 of each face taken at different ages, from different angles, and in different lighting conditions) and were asked to sort them by identity. Children mistook images of the same person as images of different people, subdividing each individual into many perceived identities. Younger children divided images into more perceived identities than adults and also made more misidentification errors (placing two different identities together in the same group) than older children and adults. In Experiment 2, we used the same procedure with 32 cognitively able children with autism. Autistic children reported a similar number of identities and made similar numbers of misidentification errors to a group of typical children of similar age and ability. Fine-grained analysis using matrices revealed marginal group differences in overall performance. We suggest that the immature performance in typical and autistic children could arise from problems extracting the perceptual commonalities from different images of the same person and building stable representations of facial identity. --------------------------------------------------------------------------------","Face identity recognition is a complex skill performed rapidly and seemingly effortlessly by mature adults. The ability to discriminate between faces is present very early in development (Field, Cohen, Garcia, & Greenberg, 1984) and improves markedly between early childhood and adolescence (Bruce et al., 2000; Carey, Diamond, & Woods, 1980; Ellis, 1992; Mondloch, Geldart, Maurer, & Le Grand, 2003). Yet the emergence of adult face expertise follows a protracted developmental trajectory, with performance on tests of unfamiliar face recognition not approaching maturity until well into adulthood (Germine, Duchaine, & Nakayama, 2011; Susilo, Germine, & Duchaine, 2013). Much research has focused on the mechanisms underlying this lengthy course of development, including holistic, configural, and norm-based coding abilities (Crookes & McKone, 2009; Jeffery, Read, & Rhodes, 2013; Mondloch, Le Grand, & Maurer, 2002; Mondloch et al., 2003; Pellicano & Rhodes, 2003; Pellicano, Rhodes, & Peters, 2006; Taylor, Batty, & Itier, 2004; Turati, Sangrigoli, Ruel, & de Schonen, 2004), which continue to be the subject of much debate (McKone, Crookes, Jeffery, & Dilks, 2012). This research, however, has focused on individuals’ abilities to tell different faces apart (referred to here as between-person face recognition) or to match an image that differs in one experimentally manipulated way, such as facial expression or viewpoint, to a target image of the same identity (Bruce et al., 2000; Ellis, 1992; Mondloch et al., 2003; Sporer, Trinkl, & Guberova, 2007). Consequently, existing research cannot explain how we recognize the same “real” face in different contexts (Burton, 2013). Only one identified study has investigated individuals’ ability to recognize the same identity across several naturally varying images of the same face (referred to here as within-person face recognition). Working on the assumption that muscular movement, lighting, resolution, and depth of contrast can lead to considerable variability in photographs of the same face, Jenkins, White, Van Montfort, and Burton (2011) gave adult participants 20 naturally varying images each of two unfamiliar Dutch celebrities sourced from the Internet and asked them to sort the photographs by identity. No participant arrived at the correct solution of two identities. In fact, the median solution was 7.5 identities, and solutions ranged considerably from 3 to 16. These findings suggest that adults find it extremely challenging to process the wide variability in “real” photographs of the same (unfamiliar) face. The mechanisms necessary for the difficult task of within-person face recognition are as yet unknown, but it is possible that an averaged representation of facial identity, central to several models of between- face identity recognition (Bruce & Young, 1986; Burton, Jenkins, Hancock, & White, 2005; Valentine, 1991), plays a crucial role. According to norm-based coding models (Leopold, O’Toole, Vetter, & Blantz, 2001; Rhodes & Jeffery, 2006), a viewer creates an internal representation of an average face based on all of the different faces to which that particular individual has been exposed. Accurate between-person face recognition occurs through the positioning of facial identities as vectors from this norm in a multidimensional face space (Valentine, 1991). Burton et al. (2005) further proposed that a “population mean” is calculated for each separate facial identity and updated after every encounter to improve reliability. Successful within-person face recognition, therefore, may require a particularly fine-grained norm-based coding system in which an averaged internal representation is created separately for each facial identity. New images of the same face may then be coded as deviations from that particular identity’s “norm.” If judged to be close enough in multidimensional face space, commonalities extracted from this image may then be incorporated into this normed representation. Bruce (1994) suggested that it is continuous exposure to faces from various viewpoints, lighting conditions, and angles that helps us to derive stable averaged representations of faces through variation. According to Burton et al. (2005), these internalized representations adopt only the elements of identity that are consistent across many viewings (e.g., the structural aspects of the face such as the spatial relationship between the eyes and nose or between the nose and mouth) and discard the more superficial and changeable aspects of a face (e.g., haircut, lighting, expression). Jenkins et al.’s (2011) task, therefore, may be a particularly challenging one because it may require an ability to create and continually update an internalized representation of a new facial identity, reliably distilling the stable elements of identity within a series of 40 encountered images of only two unfamiliar people. Studies investigating individuals’ ability to distinguish different identities suggest that a norm-based face space is present in young children, including those as young as 4 years (Jeffery et al., 2013; Nishimura, Maurer, Jeffery, Pellicano, & Rhodes, 2008; Pimperton, Pellicano, Jeffery, & Rhodes, 2009). In fact, different responses to average faces, as opposed to “distinct faces,” in infants (Rhodes, Geddes, Jeffery, Dziurawiec, & Clark, 2002) suggest that it could be present even earlier. Whereas Crookes and McKone (2009) argued that general cognitive improvements in the visual system or in working memory are likely candidates for the mechanisms underlying age- related improvements in face identity recognition, Jeffery et al. (2013) suggested that the quantity and quality of dimensions of face space may undergo fine-tuning during development, leading to a more efficient and precise face identification system. To date, all of this research has examined norm-based coding within the context of between-person face recognition. There is no study examining within-person recognition of unfamiliar faces in different contexts in children. As challenging as it is for adults, Jenkins et al.’s (2011) task may be especially difficult for children because it may require participants to create and update an average of two facial identities repeatedly in a short space of time, extract the commonalities between images, and build a stable representation “online.” Within-person face recognition might also be particularly challenging for children diagnosed with autism, a neurodevelopmental condition that affects the way an individual interacts with and experiences the world around him or her (American Psychiatric Association [APA], 2013). Atypicalities in face processing are well documented in individuals on the autism spectrum (Rutherford, Clements, & Sekuler, 2007; Wallace, Coleman, & Bailey, 2008; Wilson, Palermo, & Brock, 2012; Wolf et al., 2008), including greater difficulty in recognizing unfamiliar faces in autistic children compared with children of similar age and ability (Boucher & Lewis, 1992; Klin et al., 1999). There have also been swathes of studies that failed to find any such differences. Indeed, Weigelt, Koldewyn, and Kanwisher (2012) reviewed 90 studies, of which only half found that autistic people perform worse than typical individuals (n = 46), with the other half finding no difference (n = 44), although they concluded that autistic individuals had difficulties in tasks specifically involving face memory and in discriminating eye gaze. Researchers investigating the source of potential difficulties in face recognition in autism have implicated weaker norm-based coding of facial identity (Ewing, Pellicano, & Rhodes, 2013; Fiorentini, Gray, Rhodes, Jeffery, & Pellicano, 2012; Pellicano, Jeffery, Burr, & Rhodes, 2007; Rhodes, Ewing, Jeffery, Avard, & Taylor, 2014). These experimental findings are consistent with a recent theoretical account that situates autistic perception within a Bayesian framework. According to Pellicano and Burr (2012), autistic individuals are less likely to use prior information to interpret incoming sensory information and, therefore, have difficulties in discerning the salience of incoming sensory information. In face identity recognition, less reliance on prior knowledge (i.e., knowledge accrued with experience) may translate to difficulties in discarding information that is irrelevant to identity judgments (e.g., the superficial and changeable aspects of facial images) and identifying the commonalities over repeated viewings necessary for face identification according to Bruce (1994) and Burton et al.’s (2005) models. If autistic individuals have difficulty in creating or updating an abstract representation of a person’s face, they should have difficulty in recognizing several unfamiliar images of the same person as belonging to the same face. The current study ~~~~~~~~~~~~~~~~~ The aims of this study were twofold. First, we sought to extend research on within-person face recognition using multiple images of the same identity to children and investigate age-related changes in this ability. In Experiment 1, we modified Jenkins et al.’s (2011) procedure to ensure that it was engaging and developmentally appropriate for children between 6 and 14 years of age. Children were shown 40 naturally varying photographs of two different male faces and were asked to sort them by identity. We predicted that, in line with other findings on the development of between-person face identity recognition, younger participants would perform less well than older participants, perceiving a greater number of identities among the 40 images of the two men than older participants. Second, we investigated the within-person face recognition skills of children with and without autism. In Experiment 2, we compared autistic children’s task performance with a subgroup of children from Experiment 1 matched for age and intellectual ability. We predicted that children on the autism spectrum would have difficulties in developing stable internal representations of newly presented unfamiliar identities (cf. Pellicano & Burr, 2012) incommensurate with their age and ability.","Participants A total of 77 children participated in this experiment, including 16 6- and 7-year-olds (M = 6;10 [years/months], range = 6;0–7;11, 6 girls), 28 8- and 9-year-olds (M = 9;3, range = 8;2–9;11, 9 girls), 20 10- and 11-year-olds (M = 10;11, range = 10;1–11;11, 14 girls), and 13 12- to 14-year-olds (M = 13;2, range = 12;0–14;9, 11 girls). One additional child was tested but excluded from analysis for not understanding task instructions (see below). In addition, 15 adults between 26 and 37 years of age also participated (M = 31;1, 6 women). Participants were recruited through mainstream schools and community contacts in the Greater London area. Stimuli Participants were presented with 40 laminated grayscale photographs of two male identities (20 of “Rob” and 20 of “Dom”) measuring 85 × 65 mm each. The individuals in the photographs were unfamiliar to participants. Because this study sought to examine children’s responses to naturally occurring variability in images of faces, we did not use experimentally manipulated images. Rather, and following Jenkins et al. (2011), the images encompassed a diverse range of natural photographs taken of the identities at different ages, from different angles, and in different lighting conditions. Photographs showed faces in front view with no obstructions. The images used can be seen in Fig. 1.","Procedure Participants were seen individually in a quiet room at their school, their place of work, or the university. For all participants—children and adults—the task was presented within the context of a game, with a cover story about a detective who needed their help to identify criminal suspects. Next, participants were presented with the 40 grayscale images in a shuffled deck and were instructed to sort the photographs into piles by identity. Specifically, participants were instructed to place photographs of the same person together and to place photographs of different people into different piles, so that they ended up with a separate pile for each different person. Participants were asked to repeat the task instructions to the experimenter in their own words before proceeding so that the experimenter could be confident that participants fully understood them. Participants were given 10 min to complete the task. Pilot testing suggested that this time limit was adequate. Indeed, during the experiment proper, the majority of participants (including half of 6- and 7-year-olds) did not make use of all the time available to them, indicating to the experimenter that they had finished the task well before reaching the time limit. Following Jenkins et al. (2011), no further instructions were given and no feedback was provided regarding accuracy. Photographs were numbered on the reverse from 1 to 40, and the experimenter used these numbers to record which photographs had been placed in each pile. The number of identities perceived by each participant (maximum possible = 40) and the number of misidentification errors (where the two different identities were perceived to be the same identity) were the dependent variables of interest. This study was granted ethical approval by the Institute of Education’s research ethics committee. Written informed consent was obtained by participating adults and from parents prior to their children’s participation in this study. Children gave their verbal assent to take part. Number of perceived identities Descriptive statistics are reported in Table 1. Consistent with Jenkins et al. (2011), adults (n = 15) often mistook images of the same person as images of different people, subdividing Rob and Dom into many perceived identities (median = 5, range = 2–28). Children (n = 77) fractionated the identities even further (median = 14.5, range = 2–40), frequently judging images of the same identity to be too dissimilar to go together. Two adults and none of the children arrived at the correct solution. Performance on the task varied considerably within each age group, as can be seen in Fig. 2A. To investigate age-related differences in the number of perceived identities further, we carried out a mixed-design analysis of variance (ANOVA) with age group (6- and 7-year-olds, 8- and 9-year-olds, 10- and 11-year-olds, 12- to 14-year-olds, or adults) as the between-participants factor and identity (Rob or Dom) as the within-participants factor. There was a significant effect of age, F(1, 87) = 3.80, p = .01, ηp2 = .15. Planned comparisons revealed that adults (M = 10.60, SD = 9.28) divided images into significantly fewer groups than both 6- and 7-year-olds (M = 18.81, SD = 9.34), t(29) = 2.45, p = .02, d = 0.88, and 10- and 11-year-olds (M = 17.55, SD = 8.51), t(33) = 2.30, p = .03, d = 0.78. There were no other significant differences (all ps > .10). There was a significant main effect of identity, F(1, 87) = 5.33, p = .02, ηp2 = .06, with images of Rob divided into significantly fewer piles (M = 8.30, SD = 4.72) than Dom (M = 9.07, SD = 4.82). There was no significant identity × age group interaction, F(4, 87) = 5.33, p = .88, ηp2 = .01. Misidentification errors In Jenkins et al.’s (2011) study, poor performance was principally a failure to unify images of the same identity, with adults unlikely to make any misidentification errors (placing images of two different identities in the same pile). To assess whether the same was true in our developmental sample, each of an individual’s “perceived identities” (image piles) was assigned a “0” if it (correctly) featured images of only one identity (i.e., either Rob or Dom) and a “1” if it featured images of both identities (i.e., Rob and Dom). These scores were summed to create a total error score for each participant (maximum score = 20). Similar to Jenkins and colleagues, misidentification errors among adults were rare (median = 0, range = 0–1), but these errors were more common in children (median = 2, range = 0–9) (see Table 1). The numbers of misidentification errors varied considerably among the childhood age groups (see Fig. 2B). To investigate developmental effects, a one-way ANOVA was performed on misidentification errors. There was a significant effect of age group, F(4, 87) = 7.86, p < .001, ηp2 = .27. Planned comparisons suggested that, broadly, misidentification errors decreased with age. Although 6- and 7-year-olds (M = 3.63, SD = 2.78) did not make significantly more errors than 8- and 9-year-olds (M = 2.21, SD = 2.28), p = .08, d = 0.56, they did make significantly more errors than 10- and 11-year-olds (M = 0.95, SD = 1.20), t(34) = 3.89, p < .001, d = 1.25, 12-to 14-year-olds (M = 1.54, SD = 1.39), t(27) = 2.46, p = .02, d = 0.95, and adults (M = 0.20, SD = 0.41), t(29) = 4.72, p < .001, d = 1.73. In addition, 8- and 9-year-olds made significantly more errors than 10- and 11-year-olds, t(42.69) = 2.49, p = .02, d = 0.69, and adults, t(41) = 3.37, p = .002, d = 1.23. Furthermore, 10- and 11-year- olds and 12- to 14-year-olds made more errors than adults (p = .02, d = 0.84 and p = .005, d = 1.31, respectively). Matrix analysis Neither of the two outcome measures described so far fully characterized task performance. A participant could receive a perfect score for the number of perceived identities (two) even if both of the participant’s identity piles featured Rob and Dom. Similarly, a participant could receive a perfect score for misidentification errors (zero) even if the participant sorted the 40 images into 40 different identity piles. A series of matrices, therefore, was created to visualize task outcomes and to derive an integrated measure of task performance (see Fig. 3). All 40 images (1–20 of Rob and 21–40 of Dom) were placed along both the x and y axes, yielding 400 cells to record incidences where two images of different identities were placed together (20 Rob × 20 Dom) and 760 cells to record incidences where two images of the same identity were placed together (380 Rob × Rob and 380 Dom × Dom) for each participant separately. The 92 performance matrices (one for each participant) were then combined into five “age-binned” performance matrices to compare patterns of performance across groups. The performance matrices for 6- and 7-year-olds (n = 16) and adults (n = 15) are illustrated in Fig. 3A and B, respectively. All five age-binned matrices can be seen in the online supplementary material. Each performance matrix had four “quadrants.” The top-left and bottom-right quadrants show the number of times participants correctly matched images of Rob/Dom with other images of Rob/Dom. The top-right and bottom-left quadrants, which are symmetrical across the diagonal, show the number of times participants incorrectly placed images of Dom with images of Rob. For 6- and 7-year-olds (n = 16), therefore, perfect performance would result in a value of “16” in every cell in the top-left and bottom-right quadrants (same identity match) and a value of “0” in every cell in the top-right quadrant and its duplicate in the bottom-left quadrant. To compare groups on an integrated measure of task performance, we first calculated a “match score” by summing the values in the top-left and bottom-right quadrants for each participant separately and a “mismatch score” by summing the values in the top-right quadrant for each participant separately. Each participant’s match score was then divided by his or her mismatch score to yield a “matrix score,” a ratio of correct sorting to incorrect sorting for each participant. A constant of 1 was added to each participant’s mismatch score beforehand to ensure that all mismatch scores were non-zero. Because of the large range in scores, and violations to the assumption of homogeneity of variance, nonparametric analyses were used on the matrix score. A Kruskal–Wallis nonparametric test on matrix scores revealed a significant effect of age group, H(4) = 39.97, p < .001 (see Table 1 for scores; higher scores indicate better performance). Mann–Whitney U tests further revealed that 6- and 7-year-olds performed significantly worse than 8- and 9-year-olds (U = 308.0, z = 2.05, p = .04, r = .31), 10- and 11-year-olds (U = 267.0, z = 3.41, p < .001, r = .57), 12- to 14-year-olds (U = 167.0, z = 2.76, p = .005, r = .51), and adults (U = 236.0, z = 4.59, p < .001, r = .82). In addition, 8- and 9-year-olds performed significantly worse than 10- and 11-year-olds (U = 382.0, z = 2.13, p = .03, r = .31) and adults (U = 393.0, z = 4.66, p < .001, r = .71), whereas adults also performed better than 10- and 11-year-olds (U = 268.5, z = 3.95, p < .001, r = .67), and 12- to 14-year-olds (U = 184.0, z = 3.99, p < .001, r = .76). Overall, the results suggest that within-person face recognition is a particularly challenging task, especially for children. Analyses on the number of perceived identities showed significant improvements in performance between young children and adults. Misidentification errors also decreased with age. A more fine-grained analysis using performance matrices showed a general pattern of age-related improvements, with adults performing better than older children, who in turn performed better than younger children on a summary performance measure. These findings suggest that the ability to integrate successfully different images of the same face into a single representation of facial identity follows a lengthy developmental trajectory, just like between-person face recognition (Germine et al., 2011; Susilo et al., 2013). Participants A total of 32 6- to 14-year-old children diagnosed with autism (Mage = 11;1, SD = 2;7, 5 girls) were recruited through advertisements, the Autism Spectrum Database–UK (http://www.ASD-UK.com), mainstream and special schools, and parent support groups in the Greater London area. All children had an independent clinical diagnosis of an autism spectrum condition according to DSM-IV (Diagnostic and Statistical Manual of Mental Disorders, 4th edition) criteria (APA, 1994). Parents completed the Social Communication Questionnaire (SCQ; Rutter, Bailey, & Lord, 2003), and autistic children were administered the Autism Diagnostic Observation Schedule (ADOS-G or ADOS-2; Lord, Rutter, DiLavore, & Risi, 1999; Lord et al., 2012) using the revised algorithm (Gotham, Risi, Pickles, & Lord, 2007; Gotham et al., 2008). All children with autism scored above the threshold for an autism spectrum condition on one or both of these diagnostic measures. An additional 5 autistic children were assessed but excluded from the dataset either because there was uncertainty over whether they understood the instructions (n = 4) or because they had a full-scale IQ score below 70, as measured by the Wechsler Abbreviated Scales of Intelligence–Second Edition (WASI-II; Wechsler, 2011) (n = 1). A subgroup of 32 typically developing children (Mage = 10;7, SD = 2;2, 12 girls) who participated in Experiment 1 was matched individually with the autistic children for age and cognitive ability. There were no significant group differences in terms of chronological age, t(62) = 0.84, p = .40, verbal IQ, t(62) = 1.12, p = .27, performance IQ, t(59.25) = 0.34, p = .74, or full-scale IQ, t(62) = 0.87, p = .38, as measured by the WASI-II (see Table 2 for scores). Given that there were no differences between males and females’ performance in Experiment 1 in terms of the number of perceived identities (p = .24), misidentification errors (p = .18), or the matrix total scores (p = .32), the autism and typical groups were not matched on gender. Of the children whose parents completed the SCQ, no child scored above the cutoff of 15 for autism specified by Rutter and colleagues (2003), indicative of low levels of autistic symptomatology. Procedure The stimuli and procedure were identical to those in Experiment 1. Informed consent was obtained from the parents of all children prior to participation, and verbal assent was obtained from participating children. Number of perceived identities Descriptive statistics for autistic and typical children are reported in Table 3. Autistic children sorted images into a median of 16.5 piles, and typical children sorted images into a median of 14 piles. As can be seen in Fig. 4A, there was wide variability in performance within groups, with the number of perceived identities ranging from 2 to 39 in the autism group and from 3 to 35 in the typical group. A mixed-design ANOVA on the number of perceived identities with identity (Rob or Dom) as the within-participants factor and group (autism or typical) as the between-participants factor revealed, unexpectedly, no main effect of group, F(1, 62) = 1.45, p = .23, ηp2 = .02. There was also no main effect of identity, F(1, 62) = 2.84, p = .10, ηp2 = .04, and no significant identity × group interaction, F(1, 62) = 0.90, p = .35, ηp2 = .02. Misidentification errors An independent-sample t-test on children’s misidentification errors (the number of their image piles featuring more than one identity) revealed similar numbers of errors in autistic (M = 1.94, SD = 2.18) and typical children (M = 1.91, SD = 2.49), t(62) = 0.05, p = .96, d = 0.01. Again, there was a large amount of variance within each group, with scores ranging from 0 to 9 in children with and without autism (see Fig. 4B). Matrix analysis Similar to Experiment 1, performance matrices plotting the number of times each of the 40 images was placed in a pile with each of the other 39 images were created for each autistic child (n = 32) and each typical child (n = 32). The values in these matrices were pooled to create two “group-binned” matrices, shown in Fig. 5. For autistic children (Fig. 5A), the values in the top-left and bottom-right quadrants ranged from 0 to 19, meaning that the number of times two images of Rob/Dom were correctly placed together ranged from never (0%) to 19 (59%). For typically developing children (Fig. 5B), the values in these same quadrants were largely similar, ranging from never (0%) to 23 (72%). Next, values in the top- right quadrant (duplicated in the bottom-left quadrant), which represent the number of times two images of different identities were incorrectly perceived as being the same person, were compared across groups. The values in these quadrants ranged from 0 to 6 in both groups, meaning that at most 6 children (19%) in each group mistakenly placed a pair of mismatched images together. Finally, and similar to Experiment 1, a total matrix score was calculated for each participant by dividing each child’s match score (the sum of the values in the top-left and bottom-right quadrants) by the child’s mismatch score (the sum of the values in the top-right quadrant) As before, a constant of 1 was added to each participant’s mismatch score beforehand to ensure that all mismatch scores were non-zero. Descriptive statistics on these matrix scores are shown in Table 3. A Mann–Whitney U test on children’s matrix scores revealed a marginally significant group difference (U = 370.5, z = –1.90, p = .057, r = –.24). These results suggest that children with autism have similar difficulties as typically developing children in recognizing the same facial identity across several images, with a trend toward poorer performance on the task overall.","Developmental improvements in face identity recognition are well documented. Although research has focused on same–different judgments of experimentally manipulated images, so far it has not considered individuals’ abilities to recognize the same identity across multiple viewings within a developmental context. In this study, an experimental task designed to test how well individuals recognize differing and naturally varying images of the same identity was extended to children for the first time. Similar to Jenkins et al. (2011), we found that within-person face recognition is a difficult task for adults, with only 2 of 15 participants arriving at the correct solution. Our findings further showed, however, that it is an exceptionally difficult task for children. None of the 77 children between 6 and 14 years of age arrived at the correct solution. Furthermore, we observed that misidentification errors, a rarity among adults, were much more common in young children. This means that as well as failing to recognize the same face across different contexts, children were also more likely to perceive different identities as being the same identity. These findings add to the wealth of existing research suggesting that face identity recognition follows a protracted developmental course (Germine et al., 2011; Mondloch et al., 2003; Susilo et al., 2013). Age differences in the overall number of identities perceived, and the number of times images of the same identity were placed together, indicate that within-person face identity recognition does not improve significantly within middle childhood but does improve significantly between childhood and adulthood. The number of misidentification errors appeared to follow a slightly clearer developmental path. Younger children made significantly more errors than older children, and older children made more errors than adults, who made almost no misidentification errors at all. One way of interpreting these findings is to view a greater number of identities perceived as a failure of within-person face recognition (a failure to recognize that several images are of the same identity) and higher numbers of misidentification errors as a failure of between-person face recognition (a failure to recognize two different identities as being separate). Working on the assumption that generally there is a greater difference between two images of different identities than between two images of the same identity, an ability to discriminate between two unfamiliar identity images might emerge earlier in development, with errors decreasing as the basic structure of “multidimensional face space” (Valentine, 1991), present from a young age (Jeffery et al., 2013; Nishimura, Maurer, & Gao, 2009; Nishimura et al., 2008; Pimperton et al., 2009), undergoes adjustments and improvements (Jeffery et al., 2013). An arguably more sophisticated system, which can efficiently carry out between-person face recognition, may be a necessary requirement for within-person recognition as measured by the current task. This sophistication may involve the abilities to weight or integrate information from multiple dimensions, which children appear to be less skilled at or inclined to do (Nishimura et al., 2009). Within a norm-based coding model, success on the current task depends on a viewer positioning images much closer together in face space, as vectors from the “norm” of that particular individual’s identity, before deciding whether to integrate that image into the averaged representation for that face. Using the concept of “stability through variation” (Bruce, 1994), Burton et al. (2005) proposed that an internal representation, or “population mean,” of individual facial identities (e.g., Dom) is readjusted after each new encounter to improve reliability. If we follow this model, in Jenkins et al.’s (2011) task, participants must begin to build this internal representation from scratch and then update it repeatedly as they sort through the 40 images and decide which ones to encode as a representation of an existing identity and incorporate into that population mean and which ones to keep separate. In accordance with this norm-based coding model, this particular identity’s vector in multidimensional face space must be continually adjusted in comparison with that person’s concept of an average face. The larger the variance among images of the same face, the more sophisticated this system needs to be to extract the commonalities between them. This may explain why an ability to recognize the same identity across several different images remains so difficult even for adults. We also examined for the first time within-person face recognition in autistic children. Our findings showed no clear differences between children with and without autism, of similar age and ability, on the number of identities perceived or on the number of misidentification errors made (image piles featuring both identities). Although this is not the only study to report similar performance between autistic and typically developing individuals on a face processing task (Yi et al., 2015), difficulties in face identity recognition are well documented in autism (Dawson, Webb, & McPartland, 2005; Weigelt et al., 2012), and so these findings might appear to be somewhat surprising. One possible explanation for these null findings is the considerable variability in children’s performance within each group, which may have precluded the possibility of detecting group differences in within-person face recognition in Experiment 2 (and indeed of detecting such age-related differences in Experiment 1). Alternatively, the difficulty of the task may have also prevented us from detecting differences in performance by children with and without autism. Future research should examine within- person face recognition in children with and without autism using a simpler task—perhaps by using fewer images to sort—or by assessing such skills in adolescents or adults (rather than children) with and without autism, when such abilities are more mature and potentially vary less between individuals. One further possibility relates to the nature of the task itself. Successful overall performance on the task requires an ability to tell different faces apart as well as an ability to tell faces together. Our integrated measure of task performance reflecting these abilities did reveal a marginally significant difference between the groups. One possible interpretation of a trend toward reduced overall performance in autism is difficulties in building stable “averaged” representations of a “normed” facial identity (Leopold et al., 2001; Rhodes & Jeffery, 2006; Valentine, 1991). Evidence from the face identity aftereffect task suggests that there may be reduced updating of face norms in response to experience in children with autism (Ewing et al., 2013; Pellicano et al., 2007) and in relatives of those with autism (Fiorentini et al., 2012). This is in line with a theory of autistic perception (Pellicano & Burr, 2012), which suggests that autistic individuals are less likely to use prior information to interpret incoming sensory information. Such a bias could make it difficult to discard information that is somewhat irrelevant to identity judgments (e.g., superficial and changeable aspects of facial images such as hairstyle, head angle, and lighting), increasing the likelihood of two images of the same identity being judged as different but also of two images of different identities being judged as the same. Caution is warranted, however, when interpreting this marginal result. Further replication of the result reported here is necessary before strong conclusions can be drawn. One strength of this study is that it used naturally varying images—“real photos”—unlike the vast majority of experimental studies where images of faces are often manipulated extensively and, therefore, vastly different from images of faces seen outside of the laboratory (Burton, 2013). Yet, even the stimuli used here are devoid of context, which should help individuals to build representations and guide their decisions in everyday life. The use of natural images in an experiment also brings considerable challenges because the substantial variance in the images makes it difficult to isolate underlying mechanisms responsible for this seemingly complicated task. In addition, the task itself was an “open” task, which means that participants could employ a range of different strategies to solve it. Creating and updating an abstract representation of each individual’s facial identity, therefore, is just one possible explanation for successful performance on the task; individuals may have also matched images based on featural or configural information. The extension of eye-tracking research to studies using more naturally varying images has the potential to yield insights into which strategies participants use for successful within-person face recognition by identifying which aspects of the face are attended to in order to perform the task. Further research could also use nonsocial stimuli to see whether these findings are selective to faces, for example, by comparing performance on recognizing several images of the same face with how well participants recognize several images as being of the same object (e.g., cars). Such research will be important in determining whether the age-related improvements found here are task specific, perhaps attributable to the large number of images participants were required to sort and keep track of. In conclusion, this study has demonstrated the considerable challenge posed by within-person face recognition for typical children, adults, and autistic children alike, and it highlights the need for greater understanding of the mechanisms that underlie this largely overlooked aspect of face perception. This work also has practical implications for working with children, especially children with autism. Photographs of adults supporting autistic children are commonly used in social stories and other tools designed to prepare the children to meet new people or to engage in new activities. The findings here suggest that some children may well struggle to recognize the same people they meet in person from an individual photograph."],["Background: Improving the quality of social care through the implementation of setting-wide positive behaviour support (SWPBS) may reduce and prevent challenging behaviour. Method: Twenty-four supported accommodation settings were randomized to experimental or control conditions. Settings in both groups had access to individualized PBS either via the organisation's Behaviour Support Team or from external professionals. Additionally, within the experimental group, social care practice was reviewed and improvement programmes set going. Progress was supported through coaching managers and staff to enhance their performance and draw more effectively on existing resources, and through monthly monitoring over 8–11 months. Quality of support, quality of life and challenging behaviour were measured at baseline and after intervention with challenging behaviour being additionally measured at long-term follow-up 12–18 months later. Results: Following intervention there were significant changes to social care practice and quality of support in the experimental group. Ratings of challenging behaviour declined significantly more in the experimental group and the difference between groups was maintained at follow-up. There was no significant difference between the groups in measurement of quality of life. Staff, family members and professionals evaluated the intervention and its outcomes positively. Conclusions: Some challenging behaviour in social care settings may be prevented by SWPBS that improves the quality of support provided to individuals. --------------------------------------------------------------------------------","Challenging behaviour remains a significant problem in supported accommodation settings for people with intellectual disabilities (cf. Department of Health, 2007). Almost half of residential services use restrictive responses such as physical intervention (Deveau & McGill, 2009). Challenging behaviour is associated with placement breakdown (Phillips & Rose, 2010) and the costly removal of individuals to more restrictive, out-of-area settings (Goodman, Nix, & Ritchie, 2006). Furthermore, it is associated with high rates of injury to care staff (National Task Force on Violence against Social Care Staff, 2001). Generally, challenging behaviour is treated as an individual problem requiring intervention by psychologists, psychiatrists or other behaviour support professionals (Royal College of Psychiatrists, British Psychological Society, & Royal College of Speech and Language Therapists, 2007). But many such professionals now adopt positive behaviour support (PBS) (Carr et al., 2002), an approach inevitably leading to a focus on the context in which challenging behaviour occurs – “the central independent variable in PBS is systems change” (Carr, 2007, p.4). Such change is not easily obtained with regular reports of difficulties implementing the proposed treatments both in social care (Ager & O’May 2001) and educational settings (Bambara, Nonnemacher, & Kern, 2009). The difficult behaviour presented in schools has been recognised as requiring a broader approach, more focused on prevention (Sugai & Horner, 2002). The development of school wide positive behaviour support in the USA reflects this (Horner et al., 2009) but there has been little attention to the potential for a similar approach in social care. A setting wide approach is consistent with theoretical developments in our understanding of challenging behaviour. Once seen as an almost inevitable concomitant of intellectual disability, it is now regarded as arising from the complex interaction of biological, developmental and environmental factors (Langthorne, McGill, & O’Reilly, 2007). In particular, it has become clear that certain characteristics of the social environment (such as social deprivation and aversive stimulation) may underpin the motivation of challenging behaviour (McGill, 1999). Altering such “motivating operations” (Michael, 2007; Simó-Pinatella et al., 2013) then becomes a theoretically viable approach to preventing or reducing the occurrence of challenging behaviour in those at increased biological risk (cf. Emerson & Einfeld, 2011). Such an approach would need to focus on improving the quality of social care especially in those areas known (through the development of individualized PBS strategies) to be associated with challenging behaviour. These include, amongst others, opportunities for choice (e.g., Dyer, Dunlap, & Winterling, 1990), predictable environments (e.g., Flannery & Horner, 1994), positive social interactions (e.g., Magito-McLaughlin & Carr, 2005), more independent functioning (e.g., O’Reilly, Cannella, Sigafoos, & Lancioni, 2006) and personalised routines and activities (e.g., Brown, 1991). Such an approach has been endorsed by the NICE guidelines on challenging behaviour (Murphy, 2017; NICE Guidelines, 2015) in which the term “capable environment” is used to summarise the characteristics of support that may reduce the risk of challenging behaviour. There remains, however, very little evidence of the impact of such an environment. Most intervention trials have been of psychotropic medication with mixed results leading NICE to recommend that medication not be used as a first-line intervention for challenging behaviour. A small number of trials have shown that cognitive behaviour therapy (Vereenooghe & Langdon, 2013) and individualized PBS can be effective (Hassiotis et al., 2009). There is also evidence that training staff in PBS is associated with reductions in challenging behaviour (MacDonald & McGill, 2013). However, the impact of improving the quality of social care remains untested. The current study set out to develop and evaluate an approach to improving the quality of social care in supported accommodation settings, drawing on work on quality improvement (e.g., LaVigna, Willis, Shaull, Abedi, & Sweitzer, 1994) and approaches to changing staff practice in residential settings (e.g., Mansell and Beadle‐Brown, 2012). The primary hypothesis was that intervention would be associated with reductions in challenging behaviour. Secondary hypotheses were that intervention would lead to improved quality of support and a better quality of life. A parallel study, the results of which are reported separately, investigated the outcomes of the intervention for social care staff. Design ~~~~~~ The study was carried out as a pragmatic, cluster randomised, controlled trial (RCT) (Hotopf, 2002). Intervention was implemented by a small team consisting of the Principal Investigator (PI), one full-time researcher and two part-time researchers. Two researchers implemented the intervention in each setting with one taking the lead and one a support role. Allocation of researchers to settings was geographically driven – the part-time researcher based in the North of England worked with settings in that region and the part- time researcher based in the South of England worked with settings in that region. The full-time researcher was involved in the intervention in all settings, either as lead or support. The PI supervised the intervention process through regular meetings and telephone conferences attended by the three researchers. Ethical and governance approvals ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The study received ethical approval from the Social Care Research Ethics Committee (REC Reference 12/IEC08/0018) including for the participation of persons lacking capacity to consent. Governance applications were made to and agreed by 14 local authorities covering all the settings (control and experimental) that participated. Approval was also gained from the Association of Directors of Adult Social Services. Staff and intellectually disabled participants with capacity to consent received comprehensive, accessible information about the project and provided written consent. Intellectually disabled participants lacking capacity to consent participated (consistent with the Mental Capacity Act) through signed declarations from personal or nominated consultees. Settings and participants ~~~~~~~~~~~~~~~~~~~~~~~~~ The study ran from 2012 to 2016 in residential settings for 1–8 adults with intellectual disability. Social care in all settings was initially provided by Dimensions, a not-for- profit provider supporting 3500 people with intellectual disabilities and autism in England and Wales. Settings were geographically spread with two clusters in the South and North of England. Dimensions was asked to identify 25–30 settings with an average of 4 adults with intellectual disabilities of whom (on average) two had a recent history of frequent and/or serious challenging behaviour. Additional inclusion criteria were that there were no significant changes planned (such as change of residents/tenants) and that residents/tenants and staff were likely to consent to participate. Thirty settings were identified and all contacted to confirm their meeting inclusion criteria and to begin seeking consent. Over approximately 6 months all settings were visited by researchers and the project discussed with the manager responsible. This led to the final identification of 24 settings where all residents/tenants and the great majority of staff had consented. Non-consenting staff participated in assessment and intervention procedures as part of their employment but did not complete measures or provide data. Intervention ~~~~~~~~~~~~ Researchers, using positive behaviour support principles, sought to improve the quality of practice within experimental group settings using the following intervention process: Managers and assistant managers of each setting attended a 3 h briefing session including presentations by Dimensions’ Director of Specialist Development, the PI and the research staff involved in providing the intervention. Two researchers were allocated to each setting as described above. All researchers engaged in intervention had a Master’s degree in Applied Behaviour Analysis or Intellectual/Developmental Disabilities, and several years of experience in providing applied behaviour analysis/positive behaviour support. The PI was a registered clinical psychologist and board certified behaviour analyst with extensive experience. Following an agreed timetable, researchers, in pairs, spent a week in each setting. The first two days involved observing practice, talking to staff and service users, and reviewing documentation. Researchers organised their findings into eight areas of social care − Activities and Skill Development, Health, Service Staff, Management, Relationships with Family and Others, Communication and Social Interaction, Wider Organisation, and Physical Environment. The eight areas were chosen based on research identifying their relationship with challenging behaviour. For example, substantial research shows the relationship between communication and challenging behaviour. When individuals understand what is going on and have effective ways of communicating their needs to the people supporting them, they are much less likely to display challenging behaviour. Similarly, and beyond the immediate context, aspects of the organisation may also impact on the occurrence of challenging behaviour. Organisational policies and practices, for example, should be informed by an understanding of challenging behaviour and ensure that frontline staff receive the support and leadership they need to work effectively. The specific areas identified represent an adaptation of previous attempts at evidenced taxonomies of those characteristics of mediators (Allen et al., 2013) and of environments (McGill, Bradshaw, Smyth, Hurman, & Roy, 2014) associated with less challenging behaviour. In each of the eight areas researchers identified, in discussion with the PI, the current picture, particular strengths, and areas where there was scope to support change within the setting. A comprehensive review of practice was presented to managers (including the manager’s manager where possible). Having agreed an outline improvement programme and initial actions for all parties, researchers wrote a draft programme in which a small number of outcome standards (supported by several process and monitoring standards) were set out in each of the eight areas (cf. LaVigna et al., 1994). A number of outcome standards were similar or the same in most/all settings or in at least some settings. All programmes also included standards idiosyncratic to one or two settings. The topics of standards are shown in Table 1. Researchers returned to the setting as soon as possible after the initial week to present and discuss improvement plans. Over the ensuing 8–11 months researchers engaged in a combination of the following activities tailored to each setting: Monthly meetings with manager to review progress against the standards set. Progress was assessed using a traffic lights system with each “green” (standard fully achieved) being worth two points, “amber” (standard partly achieved) one point and “red” (standard not achieved) no points. Documentary evidence was required to score. Points were totalled so that a percentage figure could be calculated for each setting each month. The percentage represented the proportion of standards that had been achieved/partly achieved. Following each monthly meeting, total scores were graphed and sent to the manager so that everyone in the setting could see the progress being made. Coaching staff and manager. In many settings staff received support to interact more effectively with the adults living there. This might, for example, be in the context of supporting participation in activity in ways that enabled engagement in the activity without provoking challenging behaviour. Supporting the development of documentation. While all settings had extensive documentation it was not all functional. For example, there were limited activity schedules and staff frequently decided on the day what was going to happen. For some residents/tenants this was problematic since they could not predict or influence what was going to happen. Staff and manager training. Where relevant to identified standards, more formal training was arranged. For example, an introductory session on autism was organised for a number of staff groups who supported individual(s) with autism but had little understanding of its influence on their behaviour. Utilisation of existing Dimensions resources. Managers were encouraged to draw in support from other parts of the organisation. This included the organisation’s behaviour support team, a coaching resource that enabled managers to receive support with difficult supervision issues and a resource that provided staff training specifically related to active support (Mansell & Beadle‐Brown, 2012). Utilisation of local professional resources from outside Dimensions. Managers and staff were encouraged to seek input from local Community Intellectual Disability Teams (CIDTs) and other sources of potential support. For example, a referral was made to gain bereavement support following the death of one of the people who lived in a setting. Progress chasing. Researchers played a local leadership role in which they encouraged achievement of the standards using a variety of means. This included encouraging staff or manager to keep following something up or taking on specific tasks themselves. Towards the end of the intervention, researchers reduced their input and sought to transfer outstanding activities to the manager or other Dimensions staff. During the same period, control group settings received no input from researchers. They were, however, able to access additional input in the same way as before. This included making referrals within (e.g., to the behaviour support team or for coaching assistance) and outside Dimensions (e.g., to the local CIDT). Intervention implementation ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Within the experimental group monthly monitoring procedures allowed the collation of data on the standards set. The mean number of standards across settings was 145 (SD = 19.3, Range: 118–180). Table 2 shows the mean numbers of outcome and overall standards set and their achievement in each of the eight areas. Table 2 also shows the extent to which “shared” standards (those described above as having been set in some, most or all settings) were achieved. Standards were achieved at higher rates in certain areas of social care (Activities, Management, Staff, Health) and at lower rates in others (Physical environment, Relationships, Communication, Organisation). The correlation between overall standard and shared standard achievement was 0.85 suggesting that this pattern applied whether the standards were shared or more idiosyncratic. Shared standards did seem to be more likely to be achieved, an average of 83% vs the average of 75% for all standards. Fig. 1 shows the percentage of standards achieved in each setting during the intervention. All settings started at 0% as standards were only included when clearly not achieved at initial assessment. While there was considerable overlap in standards set in different settings, the overall list of standards was idiographic to each setting and the percentage achieved cannot necessarily be meaningfully compared between settings. Overall average percentage achieved (median percentage in last data collection in each setting) was 80.1% (range: 29.7–92.3%). Service 7 was an outlier (see Fig. 1). This service was re-provided during the intervention for reasons unconnected with the intervention or challenging behaviour. As a result, the three adults were divided between two other settings supported by different providers. It was not meaningful to collect further data on the achievement of standards that did not necessarily apply to the new settings. While there was considerable variation (59.9–92.3%) in the remaining 10 settings, all made substantial and relatively steady progress towards achieving the standards set. Outcomes and other measures ~~~~~~~~~~~~~~~~~~~~~~~~~~~ The primary hypothesised outcome (reduction in challenging behaviour) was measured through the Aberrant Behavior Checklist-Community (ABC), a reliable and valid measure of the severity of challenging behaviour (Rojahn, Aman, Matson, & Mayville, 2003). Baseline questionnaires were completed prior to group allocation by staff who knew each adult well, typically their key worker. Subsequently, 3–6 months after the end of intervention, and 12–18 months after that, the same measure was completed. The secondary hypothesised outcomes (quality of life, and quality of staff support) were evaluated using non- participant observations made as close as possible to the period 4–6 PM in line with previous research (cf. Felce et al., 2000). Engagement in meaningful activity (as a measure of quality of life) was recorded using momentary time sampling during the 2-h period with observations made, by rotation, of all residents/tenants present. Engagement was defined as in the EMAC-R (Mansell & Beadle-Brown, 2005). Observers at baseline were research workers, one of whom had extensive experience in using the EMAC-R and provided training to the other. Observers after intervention were current or previous research workers with extensive experience in the use of EMAC-R. At the end of the 2-h period the observer completed the Active Support Measure (ASM) (Mansell & Elliott, 1996, revised 2005) which provided ratings of quality of staff support for activity, choice-making and other aspects of social care. Baseline data were gathered prior to randomisation of settings. Three-six months after the end of intervention, the same measures were carried out by observers with no previous involvement in the study and blind to group allocation. Inter-observer agreement was checked by having a second observer present during seven of the baseline and two of the post-intervention observations. Overall agreement about engagement occurred on 88.7% of observations. Exact agreement on ASM category scores was 79% with a weighted kappa of 0.72. Observers provided qualitative comments about each setting including their view on the setting’s group membership, and their reasons for this conclusion. Information on the characteristics of intellectually disabled participants was gathered using a shortened version of the Individual Schedule (IS) (Emerson et al., 1999) incorporating the Short Adaptive Behavior Scale (SABS) (Hatton et al., 2001). Questionnaires were completed prior to group allocation by staff who knew each individual well. Following completion of the intervention, staff in experimental settings, family members of the people supported there and external professionals (e.g. CIDT members) with significant involvement completed one-page questionnaires on their experiences and overall evaluation of the intervention’s impact. Questionnaires shared common items, expressed differently for different groups e.g. “Over the last year my relative’s health has improved” (family version), “For the people we support the project has improved their health” (staff version), “The project has improved the health of the people supported” (external professional version). Additionally, staff questionnaires included four items relating to their own experiences of the project (e.g., “The project has improved the way I am supported by my manager”). All items were rated on a 5-point scale from “strongly disagree” that the intervention had a positive impact to “strongly agree”. Sample size, randomization and blinding arrangements ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In an RCT of the impact of behaviour support on challenging behaviour in individuals living in community settings, Hassiotis et al. (2009) estimated an effect size of 0.8 using the ABC. Effect size in the current study was therefore set at 0.8; with power set at 0.8 and p < 0.05 (one-tailed), 20 participants would be needed in each of the two groups. Given the nature of the intervention, it would not have been possible to include participants from the same setting in both experimental and control groups. Accordingly, participants were grouped in clusters (where the cluster was the residential setting). With such an arrangement it was necessary to consider the within-cluster correlation. The magnitude of this effect is given by the formula: Deff = 1 + (m-1) × ICC where Deff is the design effect, m is the average number of participants in each cluster and ICC is the intra-cluster correlation coefficient. The latter was estimated from unpublished data on challenging behaviour across 13 settings (clusters) of an adult social care provider collected by the PI as part of a service evaluation to be 0.096. Taking this as 0.1 and assuming an m of 2, the Deff is then 1.1 and 22 participants would be required in each group. On the basis of an average of 2 participants per placement this implied 11 settings per group. Settings were expected to also accommodate a similar number of individuals who did not display challenging behaviour. In total, therefore, the expected required sample was, in both experimental and control groups, 22 people who currently present challenging behaviour together with 22 people who did not currently present challenging behaviour. As noted below, the actual samples obtained approximated these requirements. The 24 identified residential settings were allocated by the PI to experimental or control group using the computer programme MINIM (Evans, Royston, & Day, Undated). This method of allocation (minimisation, e.g., Treasure & MacRae, 1998) is particularly suitable for cluster trials containing a relatively small number of clusters since it ensures maximum similarity between groups in respect of variables that might influence outcome (cf. Turk et al., 2010). MINIM was set up to minimize differences between groups in respect of the following: geography (north vs south of England), number of staff in setting (above or below median), challenging behaviour (above or below median ABC score), number of adults without significant challenging behaviour (none vs one or more) adaptive behaviour (above or below median SABS score) and number of individuals with autism (0 vs 1 or more). Given the nature of the trial, only limited blinding arrangements were possible. All baseline data were gathered prior to group allocation so both settings and researchers were blind at this stage. All settings were aware, however, at subsequent datapoints, whether they were in the experimental or control group. Gathering of data following intervention was facilitated by researchers aware of group allocation since they had been involved in delivering the intervention. The exception to this was in respect of the collection of non-participant observation and rating data following intervention. Observers had not previously been involved in the study and were blind to group allocation. Analysis ~~~~~~~~ All data analyses were conducted in IBM SPSS version 24. The period between baseline and post-intervention data collection was 12–18 months. During this time there were a number of changes to adult participation affecting data availability (see Fig. 2). Consequently, the experimental group was reduced from 11 to 9 settings, and the control group from 13 to 12 settings. Data are presented throughout, and analysis conducted, primarily at the setting level as that was the focus of the intervention and of randomisation. To investigate the degree to which setting level analysis reflected outcomes for individuals, a sensitivity analysis was conducted of changes in the primary outcome measure between baseline and post-intervention. This also allowed the impact of changes in the participants present at each datapoint to be investigated. Inferential analysis involved the comparison of group means in difference scores between baseline and post-intervention. Thus, a “per protocol” approach was used with no data imputation for those lost to follow- up. While an “intention to treat” approach would normally be regarded as the gold standard, the primary focus of this trial was to evaluate intervention efficacy. Analysis of data at long-term follow-up (only possible for ABC scores) took a similar approach but is presented separately as there was more missing data and the study was originally planned with only baseline and post-intervention data points. Settings and participants ~~~~~~~~~~~~~~~~~~~~~~~~~ Eleven settings were allocated to the experimental group and 13 to the control group (see Fig. 2). Group composition, in terms of the variables used in allocation, is shown in Table 3. Group characteristics were compared using independent sample t-tests and, for categorical data, chi-square. None of the differences between groups were significant at p < 0.05, two-tailed. Demographic information was gathered on all residents/tenants of the 24 residential settings (see Table 4). Outcomes after intervention ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Average ABC scores for each setting are shown in Table 5. Setting average scores reduced in nine of nine experimental settings with the group mean reducing from 39.2 (range: 18.5–61) to 12.5 (range: 4–21). The control group mean reduced from 42.3 (range: 15.7–70) to 34.9 (range: 14–51.7) with seven of twelve settings reducing. The difference across time between groups was significant (t = 2.24, df = 19, p = .04, 2 tailed, Cohen’s d = 1.00, 95% CIs −37.20 to −1.29). Mean percentage ASM scores are also shown in Table 5. Mean scores increased in seven of nine experimental settings (remained the same in one and baseline data not gathered in the other) with the group mean increasing from 48.0 (range: 35.6–61) to 67.6 (range: 35.6–93.3). The control group mean reduced from 47.7 (range: 30.4–67.4) to 45.5 (range: 17.8–65.5) with five of twelve settings showing an increase. The difference across time between groups was significant (t = 2.88, df = 18, p = .01, 2 tailed, Cohen’s d = 1.28, 95% CIs 5.92–37.88). Mean percentage engagement in meaningful activity increased in experimental settings from 49% to 68.2% and, in control group settings, from 52.5% to 58.9%. The difference across time between groups was not significant (t = 0.75, df = 18, p = .46, 2 tailed, Cohen’s d = 0.34, 95% CIs −16.37 to +34.44). Post-intervention observers identified group membership of 19 of the 21 settings, correctly allocating control group settings in all instances and experimental group settings in seven out of nine instances. Kappa was 0.80. In noting their conclusions, observers made no comments suggesting blinding had been breached. All allocation reasons related to their observations of the quality of social care and/or its outcomes. Sensitivity analysis of changes in individual ABC scores ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ABC scores reduced in 21 of 26 experimental participants with the group mean reducing from 34.5 (range: 3–92) to 11.8 (range: 0–26). The control group mean reduced from 43.2 (range: 0–135) to 36.8 (range: 0–102) with 17 of 38 individuals reducing. The difference across time between groups was non-significant (t = 1.85, df = 62, p = .07, 2 tailed, Cohen’s d = 0.49, 95% CIs −33.89 to +1.24). Post hoc analysis of individual scores suggested that higher ABC scores were recorded in smaller settings in the experimental group at baseline, resulting in higher setting average scores at baseline compared to individual scores since settings were weighted equally however many individuals were supported in each. There was no evidence that change in individual participants across time had affected results – post-intervention data was not collected on four participants, three in the control group (one death, one data collection error, one left setting) and one in the experimental group (one death) – a repeat of the analysis of differences across time produced almost identical results to those in Section 3.2. Outcomes at long-term follow-up ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Setting average ABC scores at follow-up are shown in Table 5. Mean scores were 18.4 (range: 7.5–32.5) in the experimental group and 39.8 (range: 15–100) in the control group. Averages across the three time points are shown in Fig. 3. The difference from baseline to follow-up between groups was not significant (t = 1.58, df = 15, p = .13, 2 tailed, Cohen’s d = 0.85, 95% CIs −15.67–0.53) though the difference between groups at follow-up remained significant (t = 2.13, df = 15, p = .05, 2 tailed, Cohen’s d = 1.13, 95% CIs −42.88–0.01). Social validity ~~~~~~~~~~~~~~~ Seventy-two staff (from all nine experimental settings) completed a questionnaire. Of the 787 individual ratings made, 68% recorded the intervention as having a positive impact, 25% as making no difference and 7% as not having a positive impact. Smaller numbers of family members (N = 8) and professionals (N = 12) completed questionnaires. 52% of family ratings were positive, 39% neutral and 9% negative. 77% of professional ratings were positive, 23% neutral and 0% negative. As an additional indication of social validity, four experimental group services received awards.","The recent NICE guidance on challenging behaviour (Murphy, 2017) recognised the limitations surrounding previous research on challenging behaviour. In particular, there has been a shortage of research into the effectiveness of intervention employing the most robust designs (such as RCTs). As a result, our knowledge concerning intervention remains somewhat limited. The current study constitutes one of a very few RCTs on non- pharmacological treatment of the challenging behaviour of people with an intellectual disability. Its unique approach to “treatment” emphasised system-wide rather than individually-based intervention. This approach proved a feasible way of working with most experimental group settings. There were two exceptions. In one setting, the intervention was interrupted by external “turbulence” with changed commissioning arrangements (independent of the research) requiring individuals to move to other settings supported by different care providers. In the other setting, the difficulties were related to the intervention which was, in a sense, a catalyst for change. Initial assessment identified a range of poor practice but its exposure alienated many staff working in the setting so that they withdrew their consent. It proved possible to continue delivering a form of the intervention, albeit under considerable difficulties and without adequate data collection. While implementation across other settings was variable, an average of 80% of the standards set were achieved and the percentage achieved steadily rose across the intervention period. Challenging behaviour (as measured by the primary outcome measure, ABC total scores) reduced by over 2/3rds in the experimental group. This reduction was significantly more than in the control group and constituted a large effect size, albeit sensitivity analysis suggested that the strength of the effect at individual participant level was rather less. Much of the reduction was maintained at follow-up. On the measure of quality of social care, experimental settings improved significantly more than control settings, again a large effect size. Quality of life for service users (as measured by observed engagement in meaningful activity) also improved more in experimental settings but the improvement was not significantly greater than in the control group. The intervention was greeted very positively by staff, families and professionals engaged with experimental settings. There were a number of limitations to study design and implementation. A pragmatic RCT with usual care control is prone to the criticism that the intervention group received more “attention”. With the design used all that can be concluded are the effects of the “attention” received rather than the intervention components that were instrumental in producing these effects (Woods & Russell, 2014). The earlier description of the intervention reveals the authors’ assumptions about its critical components but further research would be necessary to endorse this. Similarly, it remains possible that control group staff suffered from “resentful demoralisation” (Woods & Russell, 2014, p.8) at not being included in the experimental group. However, on some of the measures employed, the control group improved from baseline, tending to argue against this possibility. No information was collected on what happened in control group settings between baseline and post-intervention. Given their continued access to the organisation’s behaviour support team, some are likely to have received significant input with respect to the challenging behaviour of individuals. In retrospect, monitoring the external input they received would have been useful. Attention should also be drawn to the routinely high staff turnover found in social care and its implications for longitudinal research. Inevitably, measures were often completed by different members of staff at different time points, introducing an additional source of variability. Future research might focus more on the use of measures that do not depend directly upon specific staff for their completion such as observation and routinely produced data such as incident reports. However, such measures also have their problems. Observation is costly and may not provide representative data, while incident reports are often of dubious reliability. RCT designs remain relatively unusual in social care and the funding available for the current study did not allow the involvement of independent statisticians, methodologists and health economists that would be regarded as typical in healthcare research. This, inevitably, increases the possibility of bias and means that important variables (such as costs) were not measured. In particular, it should be noted that the same people were involved in both delivering the experimental condition and in the process of collecting and analysing data. As far as possible the risks of bias associated with this were reduced e.g. by employing observers blind to group membership and encouraging staff to return questionnaires in sealed envelopes. Nonetheless, it would clearly have been much better to have had sufficient independent research capacity to avoid such risks more completely. A number of limitations also arose in the implementation of the study and should receive attention in future research. It would be useful to work with more than one social care provider to help test the generality of findings and to identify the characteristics of provider organisations required to support successful intervention. In the current study, the PI had a long-established relationship with the provider including at senior management level. It seems likely that this kind of relationship and support will enhance outcome but there is as yet no comparison on which to base such a judgement. The intervention was developed by the research team and has not yet been “manualized” in detail. This is clearly an essential step if attempts at replication and component analysis are to take place. In conducting the study it became clear that some aspects of measurement could have been improved. Four issues were identified. Firstly, the data produced through monthly monitoring of standard achievement was of unknown reliability and validity. Secondly, while direct observations are potentially very useful, the amount of time allocated to these in the current study may have been insufficient given the degree of variability that might be expected across relatively short periods. This may explain the failure to find significant differences between experimental and control group settings in engagement in meaningful activity. Thirdly, measurement of resident quality of life was limited and it would be useful to expand the range of measures used, perhaps exploring the use of ASCOT (Netten et al., 2010). Fourthly, it would be useful to measure the impact of the intervention on use of other services e.g., whether people are less likely to require psychiatric, psychological or other health input. In summary, the findings of this study are promising and suggest that research on challenging behaviour should continue to investigate intervention in the system of supports surrounding individuals at risk of developing or continuing to display challenging behaviour. This implies that health and social care providers should consider the scope for directing some of their behaviour support resources at systemic, preventative intervention. In the case of Dimensions, they decided, having seen the results of the intervention, to implement a new model of support based on that used in experimental settings (see https://www.Dimensions- uk.org/initiative/activate/). The promising findings suggest that further studies should be carried out. Ideally, these studies would be larger, involve multiple providers, allow investigation of costs as well as outcomes and contribute to investigation of the factors potentially responsible for positive outcomes. It might be argued that the findings of this study support a “social model” approach to challenging behaviour. That it is possible to substantially reduce challenging behaviour through interventions that are focussed more on the environmental context than the individual suggests an analogy with the prevention of disability (despite impairment) through the making of reasonable accommodations in the environment and wider society. This is both theoretically important, in extending our understanding of challenging behaviour, and practically important, in indicating steps we can take to its prevention."],["In line with the claim that regret plays a role in decision making, O'Connor, McCormack, and Feeney (Child Development, 85 (2014) 1995-2010) found that children who reported feeling sadder on discovering they had made a non-optimal choice were more likely to make a different choice the next time around. We examined two issues of interpretation regarding this finding: whether the emotion measured was indeed regret and whether it was the experience of this emotion, rather than the ability to anticipate it, that affected decision making. To address the first issue, we varied the degree to which children aged 6 or 7 years were responsible for an outcome, assuming that responsibility is a necessary condition for regret. The second issue was addressed by examining whether children could accurately anticipate that they would feel worse on discovering they had made a non-optimal choice. Children were more likely to feel sad if they were responsible for the outcome; however, even if they were not responsible, children were more likely than chance to report feeling sadder. Moreover, across all conditions, feeling sadder was associated with making a better subsequent choice. In a separate task, we demonstrated that children of this age cannot accurately anticipate feeling sadder on discovering that they had not made the best choice. These findings suggest that although children may feel regret following a non-optimal choice, even if they were not responsible for an outcome, they may experience another negative emotion such as frustration. Experiencing either of these emotions seems to be sufficient to support better decision making. --------------------------------------------------------------------------------","Regret is an aversive emotion that we experience when we believe that we would have obtained a better outcome had we chosen differently (Epstude & Roese, 2008). Although regret is aversive, it appears to be a functional emotion in the sense that it assists in decision making (Epstude & Roese, 2008; Zeelenberg & Pieters, 2007). Experiencing regret following a poor choice is likely to lead us to change our behavior (Ku, 2008). Moreover, when deciding how to act, we often try to anticipate and thus minimize the regret that might arise from our decisions (Zeelenberg & Pieters, 2007). Recently, the developmental psychology of regret has received considerable attention (Burns, Riggs, & Beck, 2012; O’Connor, McCormack, & Feeney, 2012, 2014; Rafetseder & Perner, 2012; van Duijvenvoorde, Huizenga, & Jansen, 2014; Weisberg & Beck, 2010; Weisberg & Beck, 2012). In a typical study, O’Connor and colleagues (2012) presented children with a choice between two boxes. In regret trials the prize in the chosen box was always less attractive than that in the non-chosen box, whereas in the baseline trial the prizes were equally attractive. Children selected a box for a prize and then rated on a 5-point scale how they felt about choosing their prize. The alternative prize was revealed, and children indicated, using a three- pronged arrow, whether they now felt happier, sadder, or the same about choosing their prize. By 6 or 7 years of age, children indicated feeling sadder on regret trials but not on the baseline trial, which was interpreted as evidence that children experience regret from this age. If regret is a functional emotion, then its emergence should have implications for children’s decision making. O’Connor and colleagues (2014) gave children the box choice task (Day 1) and then returned the next day (Day 2) to present children with the same choices. They hypothesized that if the experience of regret affects decision making, children who experience regret on Day 1 should be more likely to make a different choice in regret trials on Day 2 than those who do not experience regret. To control for a general preference to switch choices, adaptive decision makers were defined as those who were willing to pay a small cost to switch in regret trials on Day 2 but were not willing to pay this cost to switch in baseline trials. O’Connor and colleagues found that participants who experienced regret on Day 1 were significantly more likely to switch on Day 2, and this result held when controlling for age and verbal ability. O’Connor and colleagues established that this association was not due to children who experienced regret having better memory for the contents of the boxes by assessing memory in a separate study. They argued that the experience of regret leads to better decision making in this task, possibly by increasing the likelihood that, when faced with the same choice again, children spontaneously bring to mind and evaluate choice options (see O’Connor et al., 2014, for discussion). There are at least two issues of interpretation with these findings. The first is whether children’s decisions on Day 2 in O’Connor and colleagues’ (2014) study were really a result of experiencing regret on Day 1 rather than a consequence of some other negative emotion such as frustration (Rafetseder & Perner, 2012). Weisberg and Beck (2012) argued that if the emotion measured in this type of study is indeed regret, it should be affected by the level of responsibility one has regarding the outcome because, at least among psychologists, responsibility for an outcome is usually considered to be a necessary condition for regret (Zeelenberg, van Dijk, & Manstead, 2000). In a study of regret using a similar box choice task, Weisberg and Beck (2012) manipulated the extent to which children perceived themselves to be responsible for the choice; the outcome was determined by the child’s choice or the roll of a die. In that study, 6- and 7-year-olds felt sadder when they had chosen the box compared with conditions involving a die, suggesting that the negative emotion that they reported was regret. However, unlike O’Connor and colleagues (2014), Weisberg and Beck (2012) did not examine the impact of children’s negative emotions on decision making, and it remains possible that, rather than regret having a distinctive effect on children’s choices, other negative emotions would show the same relation with decision making. A further difficulty with interpreting an association between the reported experience of regret and adaptive decision making hinges on the important distinction between experienced and anticipated regret (Zeelenberg & Pieters, 2007). O’Connor and colleagues (2014) argued that the experience of regret directly affected children’s subsequent decisions. However, most theorizing on regret and decision making has focused on the effects of anticipating regret (predicting future regret and trying to avoid it) rather than its experience. Indeed, some theorists hold that the primary way in which emotion affects behavior is through simulation of emotional responses to outcomes that are in prospect (Baumeister, Vohs, DeWall, & Zhang, 2007). Perhaps the association between reported regret on Day 1 and decision making on Day 2 was mediated by anticipated regret in that those children who experienced regret on Day 1 were also likely to be those children who could anticipate regret. When faced with the same choice on Day 2, this ability to anticipate regret may have led to children switching choices to avoid future regret. If this is correct, it would suggest that simply experiencing regret might not have the functional role in children’s decision making that O’Connor and colleagues (2014) ascribed to it. Rather, in line with theorizing about adult decision making, regret affects decision making primarily because in choosing a course of action we attempt to avoid it. The distinction between experienced regret and anticipated regret is particularly pertinent in a developmental context because they have different developmental profiles. Using a paradigm first described by Guttentag and Ferrell (2008), McCormack and Feeney (2015) examined the age at which children are capable of anticipating regret. Children saw three boxes and were told that inside each box was a good prize, a medium prize, or nothing (in reality all three boxes contained a medium prize). One box was removed from the game, and then children chose between the remaining two boxes. On seeing that they had won the medium prize, children indicated how they felt on a 5-point scale. Next, they used a three-pronged arrow to predict whether they would feel the same, happier, or sadder if they discovered that the large prize was in the unchosen box. There was a developmental lag of at least 1 year between the experience and anticipation of regret; it was not until 8 years of age that children were able to accurately predict that they would feel sadder, whereas in a separate task using O’Connor and colleagues’ (2012) original paradigm nearly all 6- and 7-year-olds experienced regret. The current study was designed to address these two issues of interpretation with O’Connor and colleagues (2014) results: (a) whether the association they found between reported emotion and decision making is best described as one between regret (rather than some other negative emotion) and children’s choices and (b) whether experienced emotion (rather than anticipated emotion) underlies this association. To address the first issue (a), we used Weisberg and Beck’s (2012) manipulation of outcome responsibility to see whether the relation between emotion and adaptive decision making previously reported could confidently be attributed to regret. To address the second issue (b), we measured both experienced and anticipated regret to see which was the best predictor of adaptive decision making.","in the current study completed an experienced regret task on Day 1, completed the choice switching task on Day 2, and then completed an anticipated regret task. Participants ~~~~~~~~~~~~ A total of 229 6- and 7-year-olds (122 girls and 107 boys, Mage = 81 months, range = 72–95) were recruited from schools. One of these schools served a predominantly upper middle-class population, whereas the remaining six schools served middle- and working- class populations. The vast majority of participants were of Caucasian origin. Children were randomly assigned to one of three conditions: Self (n = 76), Dice–Self (n = 78), or Dice–Other (n = 75). There was no significant difference between conditions in relation to age or gender of participants.","The apparatus for the experienced regret task included the following: for the baseline trial, two different-colored boxes, each containing a smaller box with 1 plastic token inside; for the regret trial, two different-colored boxes, each containing two smaller boxes, one with 1 token and the other with 10 tokens inside. A two-colored die was used in the Dice–Self and Dice–Other conditions. An additional three gold boxes were used for the anticipated regret task, with each gold box having a distinct image on the lid along with a high-valued prize (Lego set or friendship bracelet set), a medium-valued prize (coloring pencils, bouncy balls, or stretchy toy), and a low-valued prize (paper clip). A 5-point scale with pictures of faces ranging from very happy to very sad was used to elicit children’s emotion ratings. Children indicated how they felt by placing a three-pronged arrow at the appropriate face. This arrow had one prong pointing upward, one pointing left, and one pointing right to enable children to report a subsequent change in feelings. Procedure Children were tested individually over 2 consecutive days in their schools. On Day 1, children were invited to play a game where they could win tokens that they could swap for stickers. All children were first introduced to the 5-point scale and trained in its use (see O’Connor et al., 2012, for the full training procedure). The training involved two puppets receiving or losing gifts, and children used the three-pronged arrow to indicate whether the puppet felt happier (leftward prong), sadder (rightward prong), or the same (upward prong) over four different scenarios. The experimental trials did not commence until each child answered the four training questions correctly. The baseline trial was always introduced first because previous findings suggest that this increases the likelihood of children experiencing regret in the regret trial (O’Connor et al., 2012; van Duijvenvoorde et al., 2014). The procedure for both trials was identical. Children in the Self condition were asked to select one of the two boxes, and the chosen box was then opened and the actual prize was revealed (1 token in both trials). Children indicated their emotional response on the 5-point scale. Next, the non-chosen box was opened, and the alternative prize (1 token in the baseline trial and 10 tokens in the regret trial) was revealed. Children used the three-pronged arrow to indicate whether they now felt happier, sadder, or the same after seeing the alternative prize. Children then watched as the tokens were returned to the appropriate boxes and their name was written on their chosen box. The experimenter explained that the game would be played again the following day, emphasizing that the same prizes would be in the same boxes. After children completed both trials, they swapped their tokens for two stickers. The Dice–Self and Dice–Other conditions were identical to the Self condition except that the roll of a colored die determined box selection. In the Dice–Self condition children rolled the die, whereas in the Dice–Other condition children observed the experimenter rolling the die. Trials were presented in the same order on Day 2; children were shown their name on their previously selected box and were reminded that the same prizes were in the same boxes as before. On Day 2, all children chose the box from which they would like a prize. Before selecting a box in each trial, the experimenter gave children 1 new token and told them that it was theirs “for now” but explained that switching boxes cost 1 token and choosing the previously selected box was free (this was required to control for a general tendency to switch choices; see O’Connor et al., 2014). Children were asked whether they would like to pay 1 token to switch boxes or choose the same box for free. Children received 1 token in the baseline trial regardless of box choice. In the regret trial, children won 10 tokens if they switched boxes or 1 token if they did not. Next, in the anticipated regret task, children were shown the three gold boxes along with the high-valued, medium-valued, and low-valued prizes. Children were asked which prize was the best and which was the worst. The experimenter explained that each of the boxes contained one of the prizes, whereas in reality each box contained a medium-valued prize. Children were asked to remove one box and then to choose between the remaining two boxes. Children rated how they felt about winning the medium-valued prize using the 5-point scale. The experimenter then said, “You won the [description of prize] from your box; that’s your prize to keep.” Next, the experimenter pointed at the non-chosen box and said, “What if the best prize is in this box, the box you did not chose …,” and then pointed at the chosen box and said, “… then how would you feel about choosing your box?” Experienced regret ~~~~~~~~~~~~~~~~~~ Table 1 shows the number of participants who felt happier, sadder, or the same after seeing the alternative prize in each trial broken down by condition. Binomial tests compared responses with chance performance (33%); in each condition, in the baseline trial children were more likely than chance to report no change in happiness rating (all ps < .001) and in the regret trial they were more likely than chance to report feeling sadder (all ps < .05). Although there was no overall effect of condition on changes in happiness rating in either trial, there was an association between condition and whether or not children felt sadder in the regret trial only,1 with children in the Self condition being more likely to say that they felt sadder than those in the Dice–Other condition (χ2 = 4.14, df = 1, p < .05, ΦC = .18). There was no significant difference between the Dice–Self condition and either of the other two conditions. Choice switching ~~~~~~~~~~~~~~~~ Children who paid to switch boxes in the regret trial, but not in the baseline trial, were classified as adaptive switchers. In Table 2, we report switching behavior in each trial broken down by condition and by whether children felt sadder in the regret trial only (i.e., sadder in the regret trial and happier or the same in the baseline trial). There was no significant association between switching behavior and condition, with similar numbers adaptively switching in all groups. A significant association was found overall between whether children reported feeling sadder in the regret trial only on Day 1 and whether children were classified as adaptive switchers on Day 2 (χ2 = 14.59, df = 1, p < .001, ΦC = .26), and this association held for each condition separately (all χ2s > 5.44, ps < .05). Therefore, children who had reported feeling sadder in the regret trial only were more likely to switch adaptively when faced with the same choice again. Anticipated regret ~~~~~~~~~~~~~~~~~~ Of the total sample, 18 children did not appropriately rank the prizes used in the anticipated regret task and were excluded from subsequent analysis. Of the remaining 211 children, 75 indicated that they would feel sadder if the better prize was in the non- chosen box, 101 that they would feel happier, and 35 that they would feel the same. The distribution of responses across the three categories differed significantly from chance (goodness-of-fit χ2 = 31.43, df = 2, p < .001), with happier being the most common response. The tendency of children of this age to predict that they would feel happier under such circumstances was also reported across three experiments by McCormack and Feeney (2015), who interpreted this as evidence that children adopt a summative approach (see McCloy & Strange, 2009); that is, any “more” is better even if they do not receive the larger prize. Predicting choice switching ~~~~~~~~~~~~~~~~~~~~~~~~~~~ A binary logistic regression analysis with children’s adaptive switching as the predicted variable examined the fit of the model with the predictors of age, feeling sadder in the regret trial only (and the same or happier in the baseline trial), and answering “sadder” to the anticipated regret question. Fully 65% of cases were correctly classified by this model, χ2(3, N = 211) = 23.90, p < .001, Nagelkerke’s R2 = .14. Feeling sadder in the regret trial only was the only significant predictor of adaptive switching (β = 1.09, Wald = 12.88, p < .001, odds ratio = 2.99). This was also the case within each condition separately.","We replicated O’Connor and colleagues’ (2014) finding that feeling sadder in a regret task on Day 1 is associated with adaptive choice switching on Day 2. Our main purpose was to examine this association in more detail, looking at (a) whether children’s negative emotion in our task on Day 1 is best described as regret and (b) whether anticipated regret, rather than experienced regret, is associated with better decision making. To address the first question, we manipulated children’s likely perceptions of their own responsibility for causing the outcome. Although we did not find a significant main effect of our manipulation, the findings were similar to Weisberg and Beck’s (2012) findings insofar as children who made a choice (Self condition) were more likely to feel sadder on discovering that they could have won a better prize than those who had the outcome determined by the experimenter’s die throw (Dice–Other condition). However, the findings differed from Weisberg and Beck’s findings in that in the Dice–Other condition, although only a minority of children reported feeling sadder, they did so significantly more often than chance. How should we interpret this pattern of findings? If we assume, along with Zeelenberg and colleagues (2000), that responsibility is a necessary condition for regret, we cannot describe the emotion experienced by children in the Dice–Other condition as regret. (The Dice–Self condition is more difficult to interpret because children may erroneously believe that they are responsible for the outcome in that condition; Weisberg & Beck, 2012). One possibility is that children feel something akin to frustration that they have not received the best prize, but (as noted by Rafetseder & Perner, 2012) this frustration could be either a result of thinking about what might have been (i.e., still a counterfactual emotion but not regret) or simply a result of comparing unfavorably what they have (1 token) with what else there is (10 tokens). Our findings do not allow us to distinguish between these possibilities, but the distinction is important because only in the first case is counterfactual thinking involved. Nevertheless, the fact that children were more likely to feel sad in the Self condition than the Dice–Other condition demonstrates that, regardless of how we interpret children’s negative emotions in the latter condition, feeling responsible for an outcome affected children’s emotions. Thus, it is plausible that at least some of what is measured in the Self condition is regret. As Rafetseder and Perner (2012) argued, children’s emotions in this type of task could be a mixture of regret and frustration, and indeed this may be the best description of what underpins emotion reports in the Self condition. Regardless of children’s level of responsibility for the outcome, feeling sad in the regret trial on Day 1 was associated with adaptive decision making on Day 2, with the strength of this association being similar across conditions. This suggests that, at least in this sort of very simple decision-making task, there may be no special relation between experiencing regret and making a better choice when faced with the same decision again; other negative emotions such as frustration (that may or may not involve counterfactual thought) can play a similar role. The exact mechanism has yet to be characterized, but it is possible that experiencing the negative emotion strengthens the memory for the events or makes it more accessible. Furthermore, our findings strongly suggest that it is the experience of such emotions, rather than their anticipation, that underpins better decision making. Most children were unable to predict that they would feel sadder if an unchosen box contained a better prize despite feeling sadder when faced with such circumstances in reality in our main task, a dissociation that replicates McCormack and Feeney’s (2015) findings. Thus, our results are a counterweight to the argument that emotions affect behavior primarily as a result of us simulating and anticipating them in advance of action (see Baumeister et al., 2007). Before children can anticipate a negative emotion such as regret, they can experience the emotion, and that experience, even controlling for any ability to anticipate it, has positive consequences for decision making. Although anticipated regret is likely to increase in importance later on when it starts to emerge (at around 8 years of age; McCormack & Feeney, 2015), our findings suggest that experiencing negative emotions such as regret can directly affect children’s decisions."],["We examined performance on implicit (non-verbal) and explicit (verbal) uncertainty-monitoring tasks among neurotypical participants and participants with autism, while also testing mindreading abilities in both groups. We found that: (i) performance of autistic participants was unimpaired on the implicit uncertainty-monitoring task, while being significantly impaired on the explicit task; (ii) performance on the explicit task was correlated with performance on mindreading tasks in both groups, whereas performance on the implicit uncertainty-monitoring task was not; and (iii) performance on implicit and explicit uncertainty-monitoring tasks was not correlated. The results support the view that (a) explicit uncertainty-monitoring draws on the same cognitive faculty as mindreading whereas (b) implicit uncertainty-monitoring only test first-order decision making. These findings support the theory that metacognition and mindreading are underpinned by the same meta-representational faculty/resources, and that the implicit uncertainty-monitoring tasks that are frequently used with non-human animals fail to demonstrate the presence of metacognitive abilities. --------------------------------------------------------------------------------","Humans are self-aware. By this we mean that they are capable not only of forming mental representations of themselves as living entities existing in a world that is separate from themselves, but also of forming representations of their own mental representations (i.e., meta-representations; Pylyshyn, 1978). In short, humans are metacognitive creatures: they entertain thoughts about their own thoughts.1 One important question concerns the relation between metacognition (awareness of one’s own mental states) and mindreading or “theory of mind” (awareness of the mental states of others). This has been subject to intense debate among philosophers, and these debates have increasingly spilled into the scientific domain. There are broadly three positions one can take (and which have been taken) concerning the relation between metacognition and mindreading. The first has a long tradition in philosophy, dating back at least to Descartes (1637, 1641). It is that we are aware of our own minds through introspection, and that we combine this self-awareness together with other forms of knowledge, and other kinds of ability, in order to know of the mental states of others. To have knowledge of others’ mental states, for example, one might rely on generalizations about the links between mental states and behavior that have been learned from one’s own case. Or one might use one’s imaginative abilities to figure out what someone else is seeing from a different perspective. In its most recent (empirically-informed) incarnation, the view is that to engage in successful mindreading, we employ imaginative simulations of the perspectives of other people, attributing mental states to them by introspecting in ourselves the results of these simulations (Goldman, 2006). Inherent to this view is that we have a privileged, non-inferential access to our own mental states via introspection, whereas mindreading is based on use of metacognition via mental simulation. We will refer to this as the “self-awareness is prior” account. A second possibility is that we are aware of our own mental states by introspection, but we employ a distinct mindreading faculty when reasoning about the mental states of other people (Nichols & Stich, 2003). On this view, the two capacities might have distinct evolutionary origins and paths, are likely to emerge separately in child development, and will be uncorrelated in task performance. We will call this the “two systems” view. Then finally, it has been claimed that there is no special faculty of introspection. Rather, self-knowledge and other-knowledge depend on the operations of the same underlying faculty/process of meta-representation, albeit relying on partially distinct inputs. For example, when attributing mental states to the self, the meta-representational faculty has access to one’s own inner speech and visual imagery, whereas the faculty only has access to observable behavior when attributing mental states to others. According to this view, this meta-representational faculty/process evolved in the first instance for other- directed mindreading, and this deployment of it is also the first to emerge in child development (Carruthers, 2009, 2011; Gopnik, 1993; Wellman, Cross, & Watson, 2001). We will refer to this as the “one system” view. Many possible lines of evidence can be brought to bear on the debate among these three views. First, evidence could be sought from studies exploring the association between mindreading and metacognition. Both the self-awareness-is-prior and the one-system views predict that measures of metacognition and mindreading should be correlated, since each maintains that the two abilities share processing resources (introspective ability, on the former account; a meta-representation faculty, on the latter). The two-systems view, in contrast, predicts that such measures should not be correlated significantly. Second, evidence could be sought from studies of metacognition in populations that have mindreading impairments. Arguably the “classic” disorder of mindreading is autism spectrum disorder (ASD).2 ASD is a neurodevelopmental disorder diagnosed on the basis of behavioral impairments in social-communication and behavioral flexibility (repetitive and restricted behavior and interests) (American Psychiatric Association, 2013). At the cognitive level, it is well established that autistic people have difficulties with mindreading (e.g., Yirmiya, Erel, Shaked, & Solomonica-Levi, 1998) and that these difficulties contribute to at least some of the social-communication difficulties that are diagnostic of the disorder (see Brunsdon & Happé, 2014). Mindreading task performance is reliably diminished among autistic people relative to the performance of age- and IQ-matched non-autistic individuals (with a large associated effect size of d = 0.88, according to meta-analysis; see Yirmiya et al., 2001) and is associated with reduced activation of a well-defined mindreading brain network (including the temporo-parietal junction, medial prefrontal cortex, precuneus and inferior frontal gyrus; e.g., Castelli, Frith, Happé, & Frith, 2002; see Schurz, Radua, Aichhorn, Richlan, & Perner, 2014, for review and meta-analysis). Given that individuals with ASD have clear difficulties meta-representing others’ mental states, the study of metacognition in this disorder has the potential to inform theory, as well as our understanding of ASD itself. Only the one-system view predicts that autistic people should be impaired on metacognitive tasks relative to neurotypical controls. In contrast, both self-awareness-is-prior and two-systems theorists explicitly predict that metacognition should be unimpaired in ASD, and cite evidence from studies of ASD to support the idea of a dissociation between metacognition and mindreading in this disorder (see Goldman, 2006; Nichols & Stich, 2003). If such a dissociation between metacognition and mindreading could be established in the case of ASD, it would represent a major challenge to the one system account of the relation between (and origin of) these abilities. Therefore, one key aim of the current study was to test these predictions by investigating metacognition among autistic people using a classic, widely used metacognitive task (together with classic mindreading tasks). A third source of evidence that has been used to evaluate the competing views in this debate comes from studies of metacognition in non-human primates (mostly macaques; see Beran, Smith, Coutinho, Couchman, & Boomer, 2009; Couchman, Beran, Coutinho, Boomer, & Smith, 2012; Smith, Beran, Couchman, & Coutinho, 2008; Smith, Beran, Redford, & Washburn, 2006; Smith, Couchman, & Beran, 2014; Smith, Redford, Beran, & Washburn, 2010; Smith, Shields, & Washburn, 2003; Washburn, Gulledge, Beran, & Smith, 2010). Based on the performance of macaques on measures designed to provide a behavioral index of metacognition, it has been claimed that monkeys are aware of their own states of uncertainty. That is, it has been claimed that macaques are capable of meta-representing their own knowledge/belief states and making strategic decisions on the basis of this meta-representation.3 The task that has probably been used most frequently to test metacognitive monitoring in non-human primates is the “strategic opt-out” uncertainty- monitoring task. In this paradigm, subjects are presented with a cognitive (“object- level”) task that involves making a discrimination of some sort (e.g., between a densely versus sparsely pixelated stimulus, or requiring them to select the longest among lines of varying length). Successful discrimination results in a valued reward, while failure results in a penalty (loss of reward, including a “time out” before the next trial). The crucial feature of such paradigms, however, is that subjects have the opportunity to opt- out of any given trial, moving immediately to the next trial without penalty or reward.","who make adaptive use of the opt-out option, selecting it in cases where the discrimination is difficult (where they are otherwise likely to make a mistake), are said to be monitoring and responding to their own states of uncertainty. In other words, the individual is said to be meta-representing their own uncertainty about the correct response on difficult trials, with meta-representation driving adaptive behavioral responses. The basic finding in the comparative literature is that macaques, chimpanzees, and humans all make adaptive use of the opt-out option (showing very similar response profiles), while pigeons and capuchins generally do not (Smith et al., 2014). Such claims are relevant to the debate about the relation between mindreading and metacognition, because there is no evidence that macaques are capable of meta-representing the epistemic states of conspecifics (Marticorena, Ruiz, Mukerji, Goddu, & Santos, 2011). If, therefore, this species is capable of meta-representing their own epistemic states, but not those of other macaques, then this would show the same dissociation between metacognition and mindreading that some have claimed is demonstrated by findings among autistic humans. If true, this would clearly represent a major challenge for the one system view and support the self-awareness-is-prior account or the two systems account. Such findings would also have methodological implications for the study of metacognition in humans, because they would provide a behavioral measure of the ability that could be employed with children or adults who have limited verbal communication ability and who could not therefore be tested using traditional metacognitive tasks. The claim that such tasks really do reveal any sort of metacognitive awareness is by no means uncontroversial, however. Some critics have argued against a metacognitive interpretation of the uncertainty-monitoring paradigm by claiming that its findings can be explained in terms of mere associative learning, rather than metacognition (Le Pelley, 2012) (but see Smith et al., 2014, for a strong case against such associative learning models). Other critics of a metacognitive interpretation of the findings have claimed that the data can be explained in terms of first-order prospective, affect-involving, decision-making (Carruthers & Ritchie, 2012), of the sort that humans also engage in regularly (Gilbert & Wilson, 2007; Seligman, Railton, Baumeister, & Sripada, 2013). On this view, the animals in question are uncertain and this uncertainty drives behavioral responses. However, uncertainty is a first-order state of anxiety directed at the primary response options, resulting from an appraisal that the act of selecting either one of them is unlikely to succeed. As a result, those options seem bad (are negatively valenced), whereas the opt-out response is neutral or mildly positive. In order to select the latter, participants just need to pick the option that seems best in the circumstances; they don’t need to be aware of, or meta-represent, their own state of uncertainty as such. In short, it has been claimed that while success in these tasks demonstrates cognitive decision-making, it fails to demonstrate metacognition. In the current study, autistic and (closely-matched) non-autistic adults completed a strategic opt-out version of the uncertainty monitoring task that is frequently used in the comparative psychology literature, as well as a classic “judgement of confidence” task that is used widely among humans to measure metacognitive ability. In judgement of confidence tasks, participants make some form of perceptual or cognitive discrimination and then make a judgement about the likelihood that their discrimination was correct. The closer the correspondence between actual knowledge and judgements of knowledge, the better the individual’s metacognitive accuracy. Investigating the performance of autistic individuals on the implicit opt-out alongside the traditional explicit judgement of confidence task provided an opportunity to test whether ASD involves a deficit in metacognition (indexed by diminished performance on the explicit judgement of confidence task) and whether the implicit opt-out task really measures metacognition. Finally, participants completed two classic tasks used to measure mindreading ability in humans. Use of mindreading tasks in the current study allowed us to investigate the extent to which implicit or explicit task performance relates to mindreading ability. One aim of the study was to establish whether metacognition is really unimpaired in ASD, as some have claimed. There is a paucity of research into metacognitive monitoring in ASD and findings from studies of this ability have high relevance for clinical practice, as well as our understanding of the cognitive profile of strengths and difficulties in ASD more generally (as we discuss further in the General Discussion). A second aim was to explore the extent to which meta-representation of self (on the judgement of confidence task) was associated with meta-representation of others (on the mindreading tasks). A third aim was to test the claim that strategic opting-out of difficult trials on uncertainty-monitoring tasks necessarily requires some form of meta-representation. If implicit uncertainty-monitoring tasks tap some of the same meta-representational resources as explicit judgement-of- confidence tasks undoubtedly do, then performance on the two tasks should be related significantly. In contrast, if implicit uncertainty-“monitoring” is really a form of first-order affective decision-making, then performance in the two sorts of task should be uncorrelated, because only the judgement-of-confidence task is meta-representational. In addition to their possible bearing on contrasting views of the relation between metacognition and mindreading, there is a second reason for investigating whether implicit uncertainty-monitoring tasks are genuinely metacognitive in nature. For if they are, this could provide researchers with an important and hitherto untapped tool for investigating the development of metacognitive abilities in young (even pre-verbal) human children. Moreover, such tasks could also provide a potentially important way of testing metacognitive abilities in nonverbal or minimally-verbal populations, including a significant proportion of people with ASD. We therefore set out to employ with our participants one of the implicit uncertainty-monitoring tasks commonly used in the comparative literature. We thus seek to address two inter-related sets of questions. One concerns the relation between metacognition and mindreading; the other concerns the metacognitive status of implicit forms of uncertainty-monitoring. In the study described below, we set out to investigate both sets of questions in a single framework. We presented both implicit (non-verbal) and explicit (verbal) uncertainty-monitoring tasks to a group of adults with ASD and a group of age-, IQ-, and sex-matched neurotypical controls. We also measured mindreading abilities in both groups. We hoped that this design would enable us to discriminate between the three differing views of the relation between metacognition and mindreading (self-awareness-is-prior, two-systems, or one-system), while at the same time testing whether one of the implicit uncertainty-monitoring tasks that has commonly been employed with non-human animals is really a test of metacognitive ability. The predictions of the various views outlined above are laid out in Table 1. Our own over- arching hypotheses were that (a) strategic opting-out on implicit (non-verbal) uncertainty-monitoring tasks does not require meta-representation, and (b) metacognition (as indexed by the judgement-of-confidence task) relies on the resources of the same meta- representational faculty as does mindreading. We therefore made the following more- detailed predictions for the outcome of the experiment. ASD-performance on the strategic opt-out task will not be impaired relative to the neurotypical performance. This is because we expect strategic opting-out to require only first-order representations, rather than meta-representations, and there is no reason (that we are aware of) to suggest that first-order valenced appraisals of stimuli would be challenging for intellectually able autistic adults. ASD-performance on explicit tasks will be impaired relative to neurotypical performance. This is because these tasks require explicit metacognitive judgments, and because such judgments implicate the same resources as the mindreading faculty, which is widely believed to be compromised in ASD. Implicit task performance (i.e., degree to which opting-out is adaptive/strategic) should not be predicted by metacognitive accuracy on the explicit judgement-of-confidence task in either group. This is because the two tasks do not tap into the same meta-representational resources. We expect the implicit task to require only first-order valenced decision making, whereas the explicit task requires meta-representation of one’s own uncertainty. Mindreading ability will not be related to implicit task performance. This is because the latter task is really a first-order one. Mindreading ability will be related to explicit metacognitive performance. This is because success in the explicit tasks requires self-awareness of one’s own state of uncertainty, and because we expect self-awareness to implicate the same core meta-representational resources that are employed in mindreading. These five predictions are laid out in the first row of Table 1. Subsequent rows represent the predictions of other possible combinations of views. Note that the first three rows represent the predictions of the three accounts of the relationship between mindreading and metacognition on the assumption that implicit uncertainty-monitoring tasks are not metacognitive, whereas the bottom three rows represent those predictions on the assumption that uncertainty-monitoring is metacognitive. Note, too, that the predictions of the self- awareness-is-prior and two-system views are the same across both sets of assumptions, with the exception that only the former predicts that there should be a correlation between mindreading abilities and explicit metacognitive performance. If the self-awareness-is- prior view predicts that mindreading and explicit metacognition should be correlated, and it is accepted that mindreading is damaged in ASD, then one might wonder why only the one- system account should predict that explicit metacognitive performance will be impaired in ASD (as we lay out in the second column of Table 1). But the self-awareness-is-prior account suggests that mindreading comprizes two distinct abilities: introspection/metacognition and imagination/simulation. The contribution of the former gives rise to the prediction that mindreading and explicit metacognition should correlate, but only the latter is thought to be damaged in ASD (Goldman, 2006). So there is no reason for a self-awareness-is-prior account to predict any impairment in metacognition in ASD. Proponents of the account could, of course, postulate independent damage to introspective abilities as well as to imaginative ones (impairment in which is thought to explain the absence of pretend play in childhood in ASD). But this is not a prediction made by self- awareness-is-prior accounts as currently formulated, and appears to lack any independent motivation. Participants ~~~~~~~~~~~~ Twenty-one adults with ASD and 22 neurotypical comparison adults took part in the study. All participants completed the Wechsler Abbreviated Scale for Intelligence-II (Wechsler, 1999), which provides verbal, performance, and full-scale IQ scores. All participants also completed the Autism-spectrum Quotient (AQ; Baron-Cohen, Wheelwright, Skinner, Martin & Clubley, 2001. The AQ is a valid and reliable measure of ASD traits in people with a full diagnosis and in the general population. Participants read statements (e.g., “I find social situations easy”; “I find myself drawn more strongly to people than to things”) and decide the extent to which each statement applies to them, responding on a four-point Likert scale, ranging from “definitely agree” to “definitely disagree”. Scores range from zero to 50, with higher scores indicating more ASD traits and scores ≥26 representing clinically significant levels of ASD-like traits (Woodbury-Smith, Robinson, Wheelwright, & Baron-Cohen, 2005). In addition, participants with ASD completed the Autism Diagnostic Observation Schedule (ADOS, a detailed observational assessment of ASD features on which a score o ≥7 is consistent with a diagnosis of ASD (Lord et al., 2000). Participant characteristics and matching statistics are presented in Table 2, together with their scores on the mindreading tasks described in Section 2.2 below. Participants in the ASD group had received verified diagnoses, according to conventional criteria (American Psychiatric Association, 2000; World Health Organization, 1993). No participant in either group reported current use of psychotropic medication or illegal recreational drugs, and none reported any history of neurological or psychiatric illness other than ASD.","Mindreading task #1. The Reading the Mind in the Eyes Task (RMIE) (Baron-Cohen, Wheelwright, Hill, Raste, and Plumb, 2001) is a widely used measure of mindreading in clinical and non-clinical populations. Participants were presented with a series of 36 photographs of the eye region of the face. On each trial, participants were asked to pick one word from a selection of four to indicate what the person in the picture was thinking or feeling. If participants felt more than one of the words was applicable, they were instructed to select the word they thought was most suitable. Stimuli were presented on screen to participants in a random order, and no time limit was imposed. Scores on the RMIE task range from zero to 36, with higher scores indicating better performance on the task. The proportion of items correctly identified by ASD and comparison participants is shown in Table 2.4 Mindreading task #2. We also employed a version of the Animations Task (Abell, Happé, & Frith, 2000) as a second measure of mindreading. The task, which is based on Heider and Simmel (1944), required participants to describe interactions between a large red triangle and a small blue triangle, as portrayed in a series of silent video clips. Four such clips are apt to invoke an explanation of the triangles’ behavior in terms of epistemic mental states, such as belief, intention, and deception. These clips comprise the “mentalizing” condition of the task and were employed in this study. Each clip was presented to participants on a computer screen. After the clip was finished, participants described what had happened in the clip. An audio recording of participants’ responses was made for later transcription. Each transcription was scored on a scale of zero to two for accuracy (including reference to specific mental states), based on the criteria outlined in Abell et al. (2000). Twenty percent of transcripts were also scored by two independent raters. Inter-rater reliability was excellent according to Cichetti (1994) criteria (intra-class correlations > 0.82). Accuracy (proportion) among ASD and comparison participants is shown in Table 2. Note that the Animations Task is a measure of spontaneous mental-state attribution. Those who score highly on this measure are spontaneously interpreting the movements of the geometric figures in the videos in mental- state terms (both affective and cognitive). Those who score low on this measure, in contrast, will tend to provide literal (non-mentalistic) descriptions of the movements they have observed. In addition, to reduce the number of statistical comparisons made and thus reduce the chances of making a type I error, we created a composite mindreading score (shown in Table 2), which was created by averaging performance across the Animations and RMIE tasks, as the two tasks correlated significantly in the current sample (r = 0.32, p < .04). The implicit strategic opt-out/uncertainty monitoring task. The implicit task was a perceptual discrimination task modeled closely on those that have been employed with non- human animals. On each trial, participants were presented with either two lines of differing length (in one version) or two patches of dots comprising differing numbers of dots (in the other version) (see Fig. 1a). In both versions, a red arrow was also presented to the right of the two stimuli (lines/patches). Participants were told that on each trial, they must click on one of the three response options within three seconds, and they were given a set of instructions outlining the consequences of each choice. They were instructed that they would start with a balance of £2.00, and would win or lose money depending on their responses. If the participant clicked on the longest line (or patch with greatest number of dots) they would gain 10p, whereas if they clicked on the shortest line (or patch with least number of dots) they would lose 10p. If they clicked on the red arrow, their score would remain the same and they would proceed immediately to the next trial. If they did not click on any of the options within 3 s, they would lose 30p. Participants were asked whether they understood the rules, and were told that this was important because they would be asked to recall them at the end of the task. They were given a final chance to clarify the rules before beginning the task, but no further information about the nature of the task or the study aims was offered. Participants completed 10 practice trials on which their responses did not count toward their prize money, before completing 60 experimental trials on which their responses did count. At the end of the task, the participant’s memory for the response rules was tested; the experimenter presented the participant with each rule and asked them to complete the payment amounts for each response type. The order in which tasks (implicit/explicit) were completed was fixed, such that all participants completed the implicit task before they completed the explicit task. This was to minimize the chance that participants would adopt an explicit metacognitive strategy in solving the implicit tasks. The primary dependent variable for object-level performance was the proportion of trials on which participants made accurate discrimination judgements. The primary dependent variable for “meta”-level performance was the difference between (a) the average perceptual discrimination difficulty on trials that participants opted out of making a discrimination judgment and (b) the average perceptual discrimination difficulty on trials that participants opted into making a discrimination judgment (i.e., trials on which they “took the test”). Trial difficulty was quantified as the similarity (in proportional terms) between the component items of each stimulus pair, with the smaller the value the more difficult the trial. For example, a trial on which the component items (length of the two lines, or number of dots in the two patches) differed by only 1% would be more difficult than a trial on which items differed by 10%. Evidence of adaptive opting-out would be indicated by participants opting out of trials on which the proportional difference between item pairs was significantly smaller than on trials that they opted into (e.g., opting-out of trials where there was only an average of 1% difference between stimulus pairs, but opting-into trials where there was an average of 10% difference between stimulus pairs). For the purpose of correlation analyses, the average trial difficulty of trials that were opted out of was subtracted from the average trial difficulty of all trials that were opted- into. The larger and more positive the resulting value (an “adaptiveness score”), the more adaptive/strategic participants’ opting-out was. If a participant never opted-out, the difference between options (opt-in/opt-out) could not be calculated. This was viewed as a lack of adaptive behavior, as participants did not appear to distinguish difficult from easy trials. Therefore, these participants were given a difference score of zero to reflect their lack of adaptive responding. The explicit judgement of confidence metacognition task. In the explicit task, participants completed the same perceptual discrimination task as they did during the implicit task, but with stimuli counterbalanced across tasks (those judging line length in the implicit task judged dot number in the explicit task, and vice versa) (see Fig. 1b). In the explicit task, however, the red arrow was not present, and participants were forced to choose between one of the two stimuli presented on the screen. On each trial, after a discrimination judgement had been made, participants were instructed (on a new screen) to make a confidence judgment relating to their response on the perceptual discrimination task. Participants could select either “yes” or “no” to the question “Are you confident?” by clicking on the appropriate box within 3 s. Participants started with a balance of £2.00 and were instructed that they would either win or lose money depending on their responses. If the participant clicked on the longest line (or patch with greatest number of dots) and clicked “yes” (indicating high confidence), they gained 10p, whereas if they clicked on the longest line and clicked “no” (indicating low confidence) they would gain nothing. If participants clicked on the shortest line (or patch with the fewest dots) and clicked “yes” (indicating high confidence) they would lose 10p, whereas if they clicked on the shortest line and clicked “no” (indicating low confidence) they would lose nothing. If they failed to respond on any trial within 3 s, they would lose 30p. Participants were asked whether they understood the rules, and told that this was important because they would be asked to recall them at the end of the task. They were given a final chance to clarify the rules before beginning the task, but no further information about the nature of the task or about the study aims was offered. Participants completed 10 practice trials on which their responses did not contribute to their prize money, before completing 60 experimental trials on which their responses did contribute. At the end of the task, the participant’s memory for the response rules was tested; the experimenter presented the participant with each rule and asked them to complete the payment amounts for each response type. The dependent variable for object-level performance was the proportion of trials on which participants made the correct discrimination judgement. Gamma scores (Goodman & Kruskal, 1954) were calculated to provide an index of overall judgement of confidence meta-level accuracy. This analysis is recommended by Nelson (1984) and is commonly used to analyze judgement of confidence tasks (e.g., Kelemen, Frost, & Weaver, 2000; Nelson & Narens, 1990; Nelson, Narens, & Dunlosky, 2004; Wojcik, Moulin, & Souchay, 2013). Gamma scores are a measure of association (between meta-level judgements of performance and actual object-level performance) and were calculated by comparing the number of correct meta-judgements with the number of incorrect judgements made by each individual. To calculate gamma scores the formula G = (ad − bc)/(ad + bc) was used, with “a” representing the number of “Yes” (confident) judgements made following a correct perceptual discrimination, “b” the number of incorrect “Yes” (confident) judgements made following an incorrect perceptual discrimination, “c” the number of incorrect “No” (not confident) judgements following a correct perceptual discrimination, and “d” the number of correct “No” (not confident) judgements following an incorrect perceptual discrimination. Gamma scores range between +1 to −1, where a score of 0 indicates chance-level accuracy, positive values indicate above chance accuracy, and negative values indicate below chance accuracy. When calculating gamma scores, the score cannot be calculated when two or more of the prediction rates (a, b, c, or d) are equal to 0. As such, the raw data were adjusted by adding 0.5 onto each prediction frequency and dividing by the overall number of judgement of confidence judgements made (N) plus 1 (N + 1). This correction is recommended by Snodgrass and Corwin (1988) and is routinely used when calculating gamma scores on metamemory tasks (e.g., Bastin et al., 2012; Wojcik et al., 2013). Implicit task performance ~~~~~~~~~~~~~~~~~~~~~~~~~ With regard to object-level performance, there were no significant differences between participants with ASD (M = 0.70, SD = 0.08) and comparison participants (M = 0.72, SD = 0.07) in the proportion of trials where stimuli were correctly discriminated, t(41) = 1.20, p = .24, d = 0.04. Thus, the two groups were very similar with respect to cognitive- level ability. Moreover, there was no significant between-group difference in the number of payment rules recalled, t(41) = 0.63, p = .53, d = 0.19. Finally, the difference between ASD participants (M = 0.94, SD = 0.10) and comparison participants (M = 0.87, SD = 0.17) in the proportion of trials that were opted into was non-significant (albeit moderate in size), t(41) = 1.69, p = .10, d = 0.50. Explicit task performance ~~~~~~~~~~~~~~~~~~~~~~~~~ With regard to object-level performance, there was no significant difference between participants with ASD (M = 0.69, SD = 0.07) and comparison participants (M = 0.70, SD = 0.07) in the proportion of trials on which stimuli were correctly discriminated, t(41) = 0.38, p = .71, d = 0.14. Thus, the two groups were very similar with respect to cognitive- level ability. Moreover, there was no significant between-group difference in the number of payment rules recalled, t(41) = 1.44, p = .16, d = 0.43. Finally, the difference between ASD participants (M = 0.80, SD = 0.16) and comparison participants (M = 0.76, SD = 0.13) in the proportion of confident judgements made was non-significant, t(41) = 0.93, p = .36, d = 0.27. Relation between implicit and explicit “meta”-level task performance ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ An initial correlation analysis indicated that the degree of strategic opting-out on the implicit uncertainty monitoring task was non-significantly associated with gamma score on the judgement-of-confidence task, r = 0.12, p = .43. In order to explore the extent to which performance on the explicit task predicted performance on the implicit task, a further regression analysis was conducted. The implicit task adaptiveness score was the dependent variable and judgement-of-confidence gamma score was used as the predictor variable. A Group × Gamma interaction variable was also included to establish whether judgement of confidence was a predictor of implicit adaptiveness score among one group of participants only. If strategic opting-out depends on meta-representation of own uncertainty, then meta-level performance on the judgement-of-confidence task (which clearly requires meta-representation of one’s own mental states) should predict the degree to which opting out was performed strategically. The results are shown at the top of Table 3. In Block 1, judgement of confidence Gamma was a non-significant predictor of implicit adaptiveness score, accounting for only 2% of variance in the latter. Gamma remained a non-significant predictor in Block 2 and the new Group × Gamma interaction-variable was also non-significant. Thus, the extent to which participants opted out strategically on the implicit task was not predicted significantly by their metacognitive accuracy on the explicit task among either group of participants. Relations between mindreading, judgement-of-confidence gamma, and adaptiveness of opting out ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In order to explore the relation between adaptive performance on the implicit and explicit tasks, on the one hand, and mindreading ability, on the other, two further regression analyses were conducted (Table 4). In the first analysis, the implicit adaptiveness difference-score was the dependent variable and the composite mindreading score was the predictor variable. In addition, a Group × mindreading interaction-variable was included to establish whether the extent to which mindreading was a predictor of the dependent variable differed across groups. In Block 1, the mindreading composite score was a non- significant predictor of implicit adaptiveness difference-score, accounting for <1% of variance in the latter. The mindreading composite remained a non-significant predictor in Block 2 and the new Group × mindreading interaction variable was also non-significant. Thus, mindreading ability was a non-significant predictor of the extent to which participants made adaptive responses on the implicit task and there was no significant between-group difference in this extent. In the second analysis, Gamma on the explicit task was the dependent variable and the composite mindreading score was the predictor variable. Again, a Group × mindreading interaction variable was included. In Block 1, the mindreading composite score was a significant predictor of explicit Gamma, accounting for 14% of variance in the latter. The mindreading composite score remained a significant predictor in Block 2, but the new Group × mindreading interaction variable was non- significant. Thus, mindreading ability was a significant predictor of the extent to which participants made accurate metacognitive judgements of confidence and there was no significant between-group difference in this extent.6","We employed an implicit opt-out task that was modeled closely on measures of (alleged) metacognitive ability widely employed with animals, as well as an explicit judgement-of- confidence task that is widely used with humans to measure metacognitive accuracy. Findings were closely in accord with our predictions. First, they indicate that autistic individuals have impaired metacognitive monitoring ability, as reflected by their diminished meta-level accuracy on the judgement-of-confidence task (a classic task, used widely to measure metacognitive ability among humans). Thus, autistic participants had significant difficulties both with meta-representation of others (they were impaired on each of the measures of mindreading undertaken) and with meta-representation of self (diminished judgement-of-confidence accuracy). This is contrary to the suggestion made by some that metacognition is intact in ASD and that that the disorder shows a dissociation between (intact) metacognition and (impaired) mindreading. The between-group difference in judgement-of-confidence accuracy was borderline large after controlling for other task variables (d = 0.78), and is comparable to the size of mindreading impairment typically observed among autistic individuals (d = 0.88; Yirmiya et al., 2001) and observed among autistic participants in the current sample (d = 0.65 for RMIE and 0.93 for Animations). This is compelling evidence that meta-representation of self is impaired in ASD and that this impairment is of a magnitude that is comparable to the known impairment of mindreading in this disorder. The finding of impaired metacognition in ASD is predicted by the view that metacognition and mindreading share a set of meta-representational resources. However, it is problematic for all alternative views of the relation between mindreading and metacognition, unless one were to postulate that a dedicated metacognitive system is independently damaged in ASD, in addition to the mindreading system that is widely-acknowledged to be impaired among people with this disorder. Yet none of the theorists in this debate have predicted such a “double hit”. Indeed, theorists who reject a one-system view are generally clear that they predict people with ASD should have preserved metacognitive abilities (Goldman, 2006; Nichols & Stich, 2003). Given the importance in driving scientific progress of disproving predictions, this finding of impaired metacognition is important for theory building. Moreover, our finding that mindreading abilities were predictive of explicit metacognitive performance on the judgement-of-confidence task suggests that mindreading and metacognition are not underpinned by two distinct mechanisms, as Nichols and Stich (2003) claim. Taken on its own, however, this result is consistent with Goldman (2006) self-awareness-is-prior account, since on his view mindreading abilities depend on metacognitive ones (together with simulation ability). But in fact our data raise a problem for the latter account also. For if, as Goldman assumes, it is simulation abilities that are damaged in individuals with ASD, then it will be these weaker simulation abilities that are responsible for the poorer mindreading performance of this group. That means that a significant proportion of the mindreading deficit in autistic individuals will result from an ability that is plays no part in metacognitive performance; namely, simulation ability. And in that case we should expect mindreading abilities to be less predictive of metacognitive abilities in autistic than in neurotypical individuals. But we found no evidence that this is so. These findings (that metacognition is impaired in ASD and that metacognitive ability is predicted by mindreading ability) provide support for one-system views of the relation between metacognition and mindreading. A challenge for one-system views, however, has been to explain the findings that have been taken to indicate that some non-human primates are capable of metacognition, but not mindreading. Our results are relevant to this issue too. In the present study, autistic participants had clear impairments in meta-representing self (on the judgement-of-confidence task) and others (on the RMIE and Animations tasks), but yet still opted out of implicit-confidence trials in the same strategic way that age- and IQ-matched neurotypical comparison participants did. The between-group difference in the degree to which opting out was adaptive/strategic was statistically negligible/small and non-significant, which contrasts with the significant and borderline large impairments in both mindreading (on the RMIE and Animations tasks) and metacognitive accuracy (on the judgement-of-confidence task) observed among autistic participants. Moreover, adaptiveness of opting out was not predicted by either metacognitive ability or mindreading ability; all associations with opting out were negligible in size. These findings suggest that opting out of difficult trials doesn’t necessarily require meta-representation (nor does it actually involve meta-representation in humans who undertake these tasks), but can be achieved using risk-based affective appraisals of likely success. Of course, it does not prove beyond doubt that non-human primates do not meta-represent their own states when they opt-out of difficult trials on equivalent versions of uncertainty monitoring tasks.7 Rather, it shows that it need not be the case that they strategically opt out by meta-representing self. In the current study, we observed that autistic participants had statistically significant and large difficulties with meta-representing self and others, yet nonetheless opted out of difficult trials strategically. Combined with the findings that the degree to which participants opted out strategically was not predicted by either metacognitive ability or mindreading ability (whereas mindreading ability did predict explicit metacognitive ability), we suggest that a high degree of caution should be taken when interpreting any findings of strategic opting out as indicating meta-representation of self. These findings are relevant to theory-building, of course, but also have methodological implications in that they suggest implicit uncertainty-monitoring tasks cannot be used to reveal nascent self-awareness abilities in non-verbal or minimally-verbal human populations. Couchman et al. (2012) point out, however, that when humans are queried after undertaking an implicit uncertainty-monitoring task of the sort employed here, they report that they opted out when they did because they were uncertain of the correct response. Isn’t this evidence that people respond as they do in these tasks because they are aware of their own uncertainty? But it is one thing to say that humans (and macaques) opt out because they are uncertain (which we agree with), and quite another thing to say that they opt out because they are aware of being uncertain. Since humans are consummate mindreaders, they will of course correctly judge that they responded as they did because they were uncertain. As such, they retrospectively explain their own behavior by appealing to epistemic mental states, which requires mindreading. However, this doesn’t show that they were concurrently aware of their uncertainty while they were responding, nor that they responded as they did because of such awareness. We submit that if humans sometimes make the latter sort of causal claim, they go beyond any evidence that could have been available to them at the time. (Indeed, no one should think that causal relations among mental states are introspectable.) Moreover, humans are known to engage in prospective valence-based decision-making in conditions of uncertainty (Gilbert & Wilson, 2007; Seligman et al., 2013). Indeed, valence is increasingly seen as the “common currency” for deciding among varying values under conditions of risk (Levy & Glimcher, 2012; Ruff & Fehr, 2014). Note that these models of human decision-making are not metacognitive ones. If it is granted that implicit uncertainty-monitoring isn’t genuinely metacognitive, then when considering the respective predictions of one-system (“metacognition and mindreading are grounded in a common meta-representational faculty”) and multiple-system views of the relationship between mindreading and metacognition, the only relevant rows in Table 1 are the first three. So the relevant differences are just that (1) the one-system view predicts that explicit metacognitive performance should be weaker in participants with ASD (who are known to be weaker at mindreading), whereas both the two-system account and the self-awareness-is-prior account do not; (2) the one-system view predicts that mindreading ability should be predictive of explicit metacognitive performance, whereas the two-system account does not; and (3) the one-system view predicts there should be no differences between groups in the extent to which mindreading ability predicts metacognitive performance, whereas the self-awareness-is-prior account predicts an interaction. Our findings support the one-system view in each case. If the one-system view is correct, however, then why is it that mindreading ability explains only 14% of the variance in explicit metacognitive performance? If metacognition and mindreading share a common meta- representational faculty, then one might expect that variability in mindreading ability should account for a larger proportion of the variance in metacognition than this. In reply, it is important to note it would be implausible to claim that mindreading is a monolithic system. Rather, it is likely to comprise a core set of resources (or “core knowledge”; Spelke & Kinzler, 2006), combined with attribution procedures and inference rules specific to certain kinds of mental state and/or certain sorts of task or situation. It seems likely, for example, that the mindreading resources required to identify an emotional expression from someone’s eyes (as required to succeed on the Reading the Mind in the Eyes Task) only partially overlap with those accessed when spontaneously interpreting the movements of seemingly-animated geometrical shapes on a screen (as required for the Animations Task). And then even on a one-system account, according to which metacognition and mindreading rely on the same faculty, many of the principles and resources involved in explicit forms of metacognition are likely to be different again. So it is not surprising that the measures of mindreading we employed should only correlate with explicit metacognitive performance to a limited degree, especially since the tasks differ from one another in many other respects (for example, in their executive demands). It is also important to note that this finding is a replication of a recent one by Williams, Bergstrom, and Grainger (2016), who found that the accuracy of explicit metacognitive judgements of confidence was associated significantly with mindreading ability in a large sample of neurotypical individuals. Thus, although mindreading ability is predictive of metacognitive judgement accuracy to a relatively modest degree, the association would appear to be a reliable one. Taken together, our findings are in keeping with only one of the theories considered, which is striking given that several predictions overlap across theories, highlighting how challenging it is to distinguish them empirically. The findings suggest, on the one hand, that traditional implicit uncertainty- monitoring paradigms of the sort used to test “metacognition” in non-human animals do not, in fact, measure awareness of one’s own mental states (or at any rate, not in humans). On the other hand, the results suggest that accuracy of explicit metacognitive judgements about one’s own mental states depends to a significant extent on mindreading ability. As a result, people with ASD, who have established mindreading difficulties, also show significantly weaker performance on explicit metacognitive tasks (but not implicit uncertainty-monitoring ones). The fact that we predicted exactly this pattern of results may be surprising to some autism researchers in the field, given that, in domains other than metacognition, explicit task performance is sometimes less impaired among autistic individuals than performance on implicit tasks (see, for example, Brosnan & Ashwin, 2018). There are various explanations for this pattern of performance (i.e., explicit superior to implicit) when observed among participants with ASD, such as compensatory strategy use on explicit tasks or alternative reasoning styles. However, these explanations rely on an underlying assumption that the implicit task in question depends on the same conceptual resources as the explicit task. When planning the current study we assumed, quite the contrary, that the implicit opt-out task did not rely on the same metacognitive resources as the explicit judgement of confidence task. Thus, the prediction that opt-out task performance would be undiminished in ASD was not out of keeping with findings in other domains. Of course, the question of whether intellectually high-functioning autistic adults would employ compensatory strategies to perform well on the explicit task, despite limited underlying metacognitive competence, was an open one. In the current study, this was clearly not the case, but we would not rule out the possibility that other studies could observe such a finding. The main conclusion from the current study is that there was a clear differentiation between the implicit and explicit tasks both in terms of levels of performance observed among autistic participants and in terms of associations with other cognitive tasks. This is in keeping only with the set of predictions that we made prior to beginning the study. In addition, our findings have implications for theoretical topics beyond those that have been our main focus here. They suggest, for example, that one shouldn’t uncritically accept metacognitive interpretations of adaptive responding on implicit uncertainty-monitoring tasks, not only when participants are non-human animals, but also when they are pre-verbal human infants (e.g., Goupil, Romand-Monnier, & Kouider, 2016). This is not to suggest that preverbal infants are incapable of meta-representing their own states of confidence or uncertainty. (After all, they do appear to be capable of representing others’ states of mind; e.g, Baillargeon, Scott, & He, 2010.) Rather, we are suggesting that care should be taken when selecting an experimental technique to demonstrate self-awareness in infants. One needs to be satisfied that the task taps into more than just first-order cognitive processes; and on our view, implicit uncertainty- monitoring tasks do not fit the bill. From a clinical perspective, these results suggest that intervention efforts designed to remediate mindreading deficits in ASD should have an added positive influence on metacognitive monitoring abilities among people with this disorder. Equally, however, the finding that the overlap between metacognition and mindreading was not absolute in the current study suggests that individuals with ASD would also benefit from targeted training in metacognitive monitoring. Previous research has shown that metacognitive monitoring predicts educational outcomes independent of general cognitive ability (Hartwig & Dunlosky, 2012; Thiede, 1999; Veenman, Wilhelm, & Beishuizen, 2004), and cognitive deficits can be improved by employing metacognitive strategies (Dunlosky, Kubat-Silman, & Hertzog, 2003; Murphy, Schmitt, Caruso, & Sanders, 1987). This is of particular relevance for individuals with ASD, many of whom experience significantly lower academic achievement based on their level of intelligence than would be expected, which impacts their life chances negatively (Estes, Rivera, Bryan, Cali, & Dawson, 2011). Therefore, training in metacognitive monitoring has the potential to improve both educational attainment and quality of life among individuals with ASD."],["The Self-Memory System encompasses the working self, autobiographical memory and episodic memory. Specific autobiographical memories are patterns of activation over knowledge structures in autobiographical and episodic memory brought about by the activating effect of cues. The working self can elaborate cues based on the knowledge they initially activate and so control the construction of memories of the past and the future. It is proposed that such construction takes place in the remembering-imagining system - a window of highly accessible recent memories and simulations of near future events. How this malfunctions in various disorders is considered as are the implication of what we term the modern view of human memory for notions of memory accuracy. We show how all memories are to some degree false and that the main role of memories lies in generating personal meanings. --------------------------------------------------------------------------------","The main contention of this paper is that when people remember they imagine and when they imagine they use memory. Imagining involves ‘working with memory’ (Moscovitch, 1992). Because of the intrinsic relatedness of memory and imagination we refer to what we term the remembering imagining system, RIS, (Conway & Loveday, 2015). The RIS is described further below. The constructive nature of autobiographical memory and autobiographical remembering are considered first.","Autobiographical memory (AM) is a complex cognitive system mediated by neural networks distributed through large areas of the neo-cortex and limbic system (see Cabeza & St. Jacques, 2007). Indeed, recent neuroimaging studies have found few differences between remembering, imagining the future, and what is sometimes termed ‘the default network’, all of which appear to share the same extensive distribution of interlocking neural networks (see Schacter, Chamberlain, Gaesser, & Gerlach, 2012 for a recent review). Autobiographical memory contains autobiographical knowledge, e.g. personal factual knowledge and cultural knowledge, such as the history of our times. It also contains episodic memories, e.g. fragmentary knowledge derived from experience (Conway, 2009). As such it forms a major part of the self (Conway, 2005; Conway & Pleydell-Pearce, 2000; Conway, Singer, & Tagini, 2004). Fig. 1 depicts partonomic knowledge structures in AM in which episodic memories are part of general events which in turn are part of lifetime periods which may themselves be part of broader themes such as work or relationship themes and the life story (Bluck & Habermas, 2001). Different levels of the knowledge structures are accessed by cues and the lines in Fig. 1 connecting different levels depict the action of cues and should not be taken as some sort of direct connection. So for example, knowledge represented in a lifetime period such as ‘Working at university X’ could be used as a cue to access a the general event, ‘Departmental talks’, which in turn contains knowledge that can access specific episodic memories. Thus, the entire complex knowledge base has patterns of activations arising and dissipating continuously in it as experience is represented in the mind and has the effect of activating associated long-term knowledge. The idea is that the AM system is labile and intrinsically responsive to cues. Occasionally, on a daily basis according to Berntsen and colleagues (Berntsen, Staugaard, & Sørensen, 2013), an AM may spontaneously come to mind. Such involuntary memories often appear to be the result of a specific cue, illustrating this cue-sensitivity. On other occasions control process elaborate a cue as a specific memory is sought for (Conway, 2005). This process of generating a memory takes time as a cue iterates through repeated cycles of elaboration and activation, until the sought-for memory has been constructed. In all cases the effect of a cue or set of cues is to form a stable pattern of activation within AM knowledge structures and it is that pattern of activation that is, temporarily, a memory. A specific AM will also always include the activation of episodic memories. One interesting hypothesis here is that it is only when an episodic memory or set of episodic memories are activated that a constructed memory enters consciousness, i.e. the rememberer becomes consciously aware of the memory. A possibility that then arises is that activation in the autobiographical memory knowledge base can, and indeed does, occur non-consciously. Experiences of ‘involuntarily’ remembering may then simply reflect unawareness of memory processing that occurred prior to the ‘involuntary’ recall. Indeed, Schank (1982) catalogued many instances of what he term ‘remindings’ – memories unexpectedly coming to mind often because of an abstract relation to a current situation or to another memory. So, for example, the structure of an event might cue a memory, as in the case of a person who recalled that he could not get his hair cur as short as he wanted when in England cueing a memory of not being able to get a steak cooked as rare as he liked, cf. Schank (1982). Remindings suggest that there maybe some non-conscious process that monitors memory for autobiographical knowledge that could help current problem solving. The non- conscious activation of memories may also underlie feelings such as déjà states, e.g. déjà vu, having seen before, and déjà vecu – having lived this moment before. We have proposed that in patients who often experience déjà states, particularly déjà vecu, their memory feelings may arise from memories that do not enter into consciousness but nevertheless are strongly enough activated to trigger recollective experience (Moulin, Conway, Thompson, James, & Jones, 2004). One of the problems with a memory system that has highly labile patterns of activation continually arising in response to continuously changing cues is that memories could, potentially, swamp consciousness. Thus, control process must modulate what memories become instantiated in consciousness at the same time as monitoring the cue- driven changing patterns of activation in long-term memory for information relevant to current goals. Finally, it seems to us highly possible that such non-conscious activation of autobiographical memory knowledge structures might, potentially, play a significant role in mediating a wide variety of social interactions, e.g. in influencing who we like, who we do not like, and who we are uninterested in, although this aspect of autobiographical memory has yet to be investigated. Autobiographical memory is then a complex multilayered distributed knowledge base in which cues constantly cause patterns of activation, some of which may stabilize into memories. It seems that this AM network is never totally inactive although, of course, activation will vary in strength at different times and may be modulated by control processes as suggested above. For instance, neuroimaging studies have found a ‘default’ network that is active when attention is unfocused but which is powered down when attention is focused (Schacter et al., 2012). The default network takes in many of the networks that feature in autobiographical remembering and, importantly, many of the same networks are active when future experiences are imagined. Thus, during ‘day dreaming’ the autobiographical memory knowledge base is active. In addition during sleep it is now clear that a major component of the AM system, the medial temporal lobe memory system is also highly active (see Solms, 1997). The AM knowledge base is, then, never totally quiescent and regions of it are always active, even during sleep. In general, however, it may be the case that control or executive processes exercise considerable control over which patterns of activation enter consciousness and also how patterns of activation are ‘shaped’ by cues. The latter by elaborating cues on the basis of initial cue activation and so directing a search. The former by denying patterns of activation access to consciousness, where they would disrupt current goal processing by turning attention to the pattern of activation in the AM knowledge base and away from other attention focused activities. Despite this inhibitory control the effect of cues is always to cause activation in the AM knowledge base. As proposed earlier it seems possible that such activation, that does not gain conscious representation, may nonetheless influence goal processing non-consciously. Whether this actually occurs is not known, but the Self–Memory System (SMS) model of AM put forward by Conway and Pleydell- Pearce (2000) allows for the possibility.","Fig. 2 shows a subsequent development of the SMS model that distinguishes between the working self, autobiographical memory and episodic memory. The working self consists of the conceptual self (Conway, Singer, et al., 2004) and the goal system. Memories are shaped in their construction by the working self and the working self determines what knowledge derived from experience becomes encoded into the AM knowledge base. Note that, as all experience is internal, encoding includes sensory-perceptual information, affect, thoughts, imaginings – it is a sort of derivation of the cognitive milieu over a limited time, perhaps from an event beginning to event ending (see Williams, Conway, & Baddeley, 2008). In this version of the SMS, AM consists of autobiographical knowledge. Autobiographical knowledge includes personal knowledge such as lifetime periods and general events (see Fig. 2) and also more generic and cultural knowledge. Lastly, the notion of episodic memory is developed further in this model of the SMS (Conway, 2009). Simple episodic memories consist of fragments of knowledge derived from experience, often in the form of visual images, together with a ‘frame’, i.e. contextualizing conceptual knowledge. Simple episodic memories can be grouped together, on the basis of shared cues, to form complex episodic memories, as shown in Fig. 2. Patterns of activation over episodic and autobiographical memory constitute specific autobiographical memories. The whole SMS is self-constraining and the working self determines what memories can be accessed whereas autobiographical memory constrains what goals, plans, beliefs, etc. can be realistically held (see Singer & Conway, 2011, for a psychoanalytic extension of this model).","The RIS is a window of accessibility of memories of the recent past and the near future (Conway & Loveday, 2015). Recently formed episodic memories e.g. those formed today, are highly accessible, however, as the retention interval increases these become progressively less accessible. Similarly near future (imagined) events, e.g. those anticipated to occur tomorrow, are also highly accessible and become progressively less accessible as the expectation interval increases. Conway and Loveday (2015) found the pattern of accessibility of specific memories and specific future events approximated to the curve shown in Fig. 3. Fig. 3 is idealized data rather than actual data and in the study it is based on participants listed all the memories they could for each day going back five days and all the specific future events they could envisage plausibly occurring on each day for the five following days. The curve in Fig. 3 shows the (idealized) percentage of memories recalled/imagined for each day. It is in the window of the RIS that memories and imagined events are constructed and so become associated with the current representations of the recent past and near future. One way to envisage the RIS is as a fish-eye lens in which items in the center of the lens are sharply defined and clear whereas items that lie progressively further away from the center become less focused and less distinct. In recent work in our laboratory (in progress) we have found that for events that progressively lie further in the future the more abstract, stylized, and schematic they become. They in effect loose their specificity and become more generic and scripted. Similarly, much of the more remote past becomes generic and inferred, with occasional specific memories. The RIS then is a window that extends the present moment some relatively short distance into the recent past and into the near future. In so doing we suggest that it is critical in supporting goal processing.","The concept of ‘accuracy’ in AM is complex. Memories might be thought of as falling in a 2-dimensional space of accuracy with ‘correspondence’ as one dimension and ‘coherence’ the other2 (see Conway, 2005). In this model each dimension would run from low to high creating four quadrants of: high correspondence-high coherence (unusually accurate memories), high correspondence-low coherence (memories of trauma might characterize this quadrant), low correspondence-high coherence (perhaps most memories fall somewhere in this quadrant), and low correspondence-low coherence (delusional and confabulated memories). The term correspondence refers to the case where a memory representation corresponds in some maximum way to a previously experienced event, in other words it is true to the event. Coherence refers to a memory representation that is coherent with other memory representations and self-beliefs (the conceptual self), in other words it is true to the self. Memories never fully correspond to our experience, only to parts of it, although they may be coherent with the self. Consider the following two memories: When I was 14 I went to see the Beatles at the Buxton Pavilion Gardens. It was absolutely packed and me and my friend Jennifer were in the middle of the crowd. You couldn’t hear the music because there was so much screaming. I am only 5 feet tall and remember thinking I’m going to faint. Anyway I fainted and was carried to the stage. I came round just as I was being carried past the Beatles on stage. I was then thrown out of the back door, where there were many girls sitting crying or in hysterics. The door opened again and my friend was turfed out. As we stood there the Beatles started playing Twist and Shout and we ran back into the hall to see the end of the show. A middle-aged man recalled his father distracting him when he was a young boy (about 4 years old) by asking who was the first man on the moon. He had been intensely interested in the moon landing when he was a young boy and this incident occurred while his father was on the telephone to his mother who had just given birth to his younger brother. He had a vivid and fond memory of his father placating him in this way, he was highly agitated by the birth, and in his memory he could ‘see’ his father on the telephone and almost ‘hear’ his voice. It was only decades later that he realized that his brother had been born in 1968, one year before the first moon landing. The first memory illustrates a common feature of all memories: time compression. The memory contains some vivid details, some of which may correspond to the experience. Nonetheless, an event that took place over several hours is condensed into a short paragraph, in turn derived from a memory that took merely seconds to construct. Moreover, many details will have been added non-consciously and automatically, such as details of clothes worn, lighting, sound levels, the weather, etc. According to our model this would be a moderate correspondence-high coherence memory. The second vivid memory may also contain some details that are true of the event, but the memory itself is of a supportive and caring father so it is a memory construction that is principally true of the self with which it is coherent. It is a low correspondence-high coherence memory. In general, memories are depictions of the past constructed in the RIS from the complex underlying knowledge base. They fall somewhere in the correspondence-coherence space with all memories having aspects of both. One suggestion is that the mean of the distribution of memories in the correspondence-coherence 2D space is, for most people, closer to coherence than it is to correspondence. In this model it is possible for a memory to be simultaneously correct and incorrect, after all the owner of the 2nd memory above, did have a brother and a caring father, who most probably did phone the mother after the birth, etc. There are also other interesting properties of memories that emerge according to their degrees of correspondence and coherence. One of interest is the extent to which an individual’s memories although a mix of both correspondence and coherence are, overall, biased to one or other of the dimensions. Perhaps such biases contribute to determining different personality characteristics and types.","In concluding I will make a brief and highly selective overview of some of the ways in which memory can be impaired following brain damage, in psychological illness, and the effects of this on the SMS. Confabulation ~~~~~~~~~~~~~ There is an extensive literature on confabulation resulting from various types of brain injury. I will consider two representative cases here. Patient O.P. was a 70 year-old woman who suffered lesions to the frontal and temporal lobes following an RTA (Conway & Tacchi, 1996). O.P. made a complete physical recovery but, on return home, her family complained of her constant ‘lies’. Her neuropsychological profile indicated spared IQ, language, and conceptual abilities. With some memory impairment and severe disturbance on frontal tests. O.P. could recall some memories very accurately but confabulated other memories about events relating to her family and these were nearly all focused on her grandson. They featured memories in which as a child he made implausible journeys to visit her, had major academic achievements, and later career achievements, all of which greatly pleased her. In fact he had been unsuccessful at school and in his career and lately had been unemployed for some years. He had little contact with her at any point in his life. Nonetheless, O.P.’s confabulations, or as Moscovitch (1995) termed them honest lies, were enduring and appeared to be highly emotionally positive for her. A case of coherence memories serving the purpose of providing a positive self-image during a time of considerable personal difficulty. Indeed, in many cases of confabulation following frontal damage it seems that the drive to make sense of the world persists but on the basis of false memories. For example, hospitalized frontal patients variously believe themselves to be in a hotel, at the office, waiting to play a round of golf, etc. Such meanings are not always positive and can be mixed with other confabulations that are emotionally negative, often highly negative (Fotopoulou, Conway, Birchall, Griffiths, & Tyrer, 2007). In either case, however, they reflect attempts to make sense of a world in which (false) memories constructed in the RIS drawn from knowledge in AM are constructed in ways that serve a disrupted working self. A working self in which personal beliefs and meanings are no longer grounded in and constrained by specific autobiographical memories. Patient P.S. (Hodges & McCarthy, 1993) suffered a bilateral paramedian thalamic infarction resulting in a temporally graded retrograde amnesia (RA) and dense anterograde amnesia (AA). His RA stretched back to his early 20s and he believed himself to be a naval rating in the 2nd world war at home on shore leave – he was in fact a 60-year old businessman. Thus, the working self in trying to make meanings out of experience could only draw up on a RIS that had existed in some form when P.S. had been about 19 years of age. Current phenomena of daily life were interpreted through this RIS and, for example, PS was surprised by how much science fiction there was on the television and insisted in a nightly blackout. The case of P.S. illustrates very powerfully how the working self and indeed the RIS are constrained by the autobiographical knowledge they are able to access, that is by the cues control processes are able to use to create patterns of activation in the knowledge base, i.e. construct memories. Psychological Illness ~~~~~~~~~~~~~~~~~~~~~ Baddeley and Wilson (1986) describe interviews with a series of schizophrenic patients. One striking feature of such patients is that their delusions were virtually always contradicted by other memories or autobiographical knowledge that, although wholly inconsistent with their beliefs no longer had the constraining power typical of a healthy autobiographical memory. In the 2D space of correspondence and coherence their memories would fall clearly in the quadrant denoting low correspondence and low coherence. For example, one patient, a man in his early 20’s, believed himself to be a famous rock guitarist, and confabulated memories of various ‘gigs’, while at the same time knew he had never learned to play the guitar and had never been to many of the places his delusions suggested. Another patient believed himself to be a Russian chess grandmaster, even though he knew he had never been to Russia and could recall being beaten at chess by other patients on his ward. Posttraumatic stress disorder (PTSD) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Memories that derive from traumatic experiences often contain errors and distortion (Conway, Meares, & Standart, 2004). One patient, for example, had been sexually abused by his grandfather from ages of 4 to 12. He had a persistent and intrusive flashback in which he was naked in a bathroom being pushed against the radiator by his naked grandfather. In this intrusive fragment he saw himself from an observer perspective as he was now, a balding 35-year old, and saw in the memory his grandfather as a frail 70-year old. In fact it became clear during therapy that he had been 6 at the time and his grandfather was in his 40s. The false (coherence) memory served the function of obscuring the fact that he had been a helpless victim. Another PTSD patient who had witnessed the 9/11 attack on the World Trade Centre in New York had a vivid and deeply disturbing flashback in which he was flying in complete silence above the plane as it approached the tower. During therapy he eventually retrieved the true memory of standing in the crowd watching the plane strike the tower and then experienced the intense feelings of apprehension and fear that he had had at the time and also the later guilt that he had about the experience. Thus, distorted and false coherence memories can have a function of creating meaning that allow the working self to avoid negative self images and intense negative emotions.","In the modern view of human memory memories are mental constructions. It is important to note that they are not reconstructions. They are not like videos, photographs or other recording media, even though they frequently contain mental imagery. They are transient constructions and although they may to some degree accurately represent the past they are time-compressed and contain many details that are inferred, consciously and non- consciously, at the time of their construction. Thus, all memories are to some degree false in the sense that they do not represent past experience literally. They can, of course, be wholly false but nonetheless be experienced as memories by a rememberer who may be unaware that the source of a memory is not experience but imagination. One of the main functions of memories is to generate meanings, personal meanings, that allow us to make sense of the world and operate on it adaptively. Memories are, perhaps, most important in supporting a wide range social interactions where coherence is predominant and correspondence often less central. A strong implication of this view is that false memories far from being damaging to the individual can often be of considerable benefit, particularly in maintaining a coherent, confident, and positive self. Our time-compressed memories with their inferred and remembered details, their mixture of coherence and correspondence, serve this function perhaps for all of us. In closing, then, we might reflect on a hypothetical, if fairly extreme, example of this. Consider a person who has had an adverse childhood. Their confidence and sense of self-worth have been undermined by an emotionally, psychologically, and perhaps physically abusive parent. Their subsequent life has been a series of disasters marked by poverty, failed relationships, substance abuse, and disturbed children of their own. Through whatever means they gradually develop a set of memories of having been sexually abused by the parent. These memories are based in part on memories of actual occurrences of abuse that they experienced, although not on memories of sexual abuse. Details of the latter gradually become incorporated into some of their childhood memories by a process of imagination (see Garry, Manning, Loftus, & Sherman, 1996). They then come to experience these high coherence-low correspondence mental representations as memories and memories that in turn support a powerful personal meaning. Namely that the individual has been victimized (which in this hypothetical example indeed they have, but not sexually) and it is the effects of the victimization that have led to their disastrous life. Thus, the ‘memories’ provide the basis of an explanation for subsequent important parts of their life. This, of course, is merely an illustration (although perhaps not such an uncommon one, see Conway, 2013) of how memories (and recall that all memories are to some degree false) maybe of value in supporting the generation of explanatory beliefs about one’s life. We suggest it is something we all do."],["There is substantial evidence that research studies reported in the scientific literature do not provide adequate information so that readers know exactly what was done and what was found. This problem has been addressed by the development of reporting guidelines which tell authors what should be reported and how it should be described. Many reporting guidelines are now available for different types of research designs. There is no such guideline for one type of research design commonly used in the behavioral sciences, the single-case experimental design (SCED). The present study addressed this gap. This report describes the Single-Case Reporting guideline In BEhavioural interventions (SCRIBE)2016, which is a set of 26 items that authors need to address when writing about SCED research for publication in a scientific journal. Each item is described, a rationale for its inclusion is provided, and examples of adequate reporting taken from the literature are quoted. It is recommended that the SCRIBE 2016 is used by authors preparing manuscripts describing SCED research for publication, as well as journal reviewers and editors who are evaluating such manuscripts. --------------------------------------------------------------------------------","The first phase to develop the SCRIBE 2016 consisted of two rounds of an online Delphi survey completed by SCED authors and methodology experts, resulting in 44 items to be discussed at a consensus conference. At the meeting, held in Sydney, Australia, in December 2011, participants reworked the items, resulting in a final set of 26 items for the reporting guideline, which are described in this Explanation and Elaboration article. The SCRIBE 2016 Statement (Tate et al., 2016) provides a detailed description of the methodology of this process.","This article provides examples of adequate reporting from the literature for each of the 26 items, along with a rationale for inclusion of the item and, where available, evidence of bias resulting from incomplete report- ing (the SCRIBE checklist of items appears in Table 1).","Item 1—Title: Identify the research as a single-case experimen- tal design in the title. Graded exposure in vivo in the treatment of pain-related fear: a replicated single-case experimental design in four patients with chronic low back pain. (Vlaeyen, de Jong, Geilen, Heuts, & van Breukelen, 2001, p. 151) Explanation. Although journals may place word limits on the title, it is important to include as much information as possible in the title, such as the intervention, the target behavior and the population. In particular, the title should explicitly mention that the study is a SCED because this differentiates the study from a case description. The abstract is sometimes copyrighted by the journal, so the title may be the only searchable information. Identifying the study as a SCED in the title will ensure that the article is appropriately indexed for bibliographic databases (such as PsycINFO or Medline). Note that using “SCED” as a key word will not be sufficient for this purpose because author key words are different from database key words and may therefore not be searchable in electronic databases. Item 2—Abstract: Summarize the research question, popula- tion, design, methods including intervention/s (independent vari- able) and target behavior/s and any other outcome measures (dependent variable), results, and conclusions. This study tested the effectiveness of Imagery Rescripting (ImRs) for com- plicated war-related PTSD [posttraumatic stress disorder] in refugees. Ten adult patients in long-term supportive care with a primary diagnosis of war-related PTSD and Posttraumatic Symptom Scale (PSS) score >20 par- ticipated. A concurrent multiple baseline design was used with baseline varying from 6 to 10 weeks, with weekly supportive sessions. After baseline, a 5-week exploration phase followed with weekly sessions during which traumas were explored, without trauma-focused treatment. Then 10 weekly ImRs sessions were given followed by 5-week follow-up without treatment. Participants were randomly assigned to baseline length, and filled out the PSS and the BDI [Beck Depression Inventory] on a weekly basis. Data were analyzed with mixed regression. Results revealed significant linear trends during ImRs (reductions of PSS and BDI scores), but not during the other conditions. The scores during follow-up were stable and significantly lower compared to baseline, with very high effect sizes (Cohen's d = 2.87 (PSS) and 1.29 [BDI]). One patient did clearly not respond positively, and revealed that his actual problem was his sexual identity that he couldn’t accept. There were no dropouts. In conclusion, results indicate that ImRs is a highly acceptable and effective treatment for this difficult group of patients. (Arntz, Sofi, & van Breukelen, 2013) Explanation. The abstract needs to provide an accurate, infor- mative and unambiguous overview of the study. It is important that all relevant information is included because many readers may not have access to the full article, or may choose to limit their reading of the study to the abstract. The CONSORT guideline for abstracts for randomized trials (Hopewell et al., 2008) provides useful information about how to write an abstract which, although written for RCTs, has applicability to SCEDs. A structured abstract can be useful and make the abstract easy to follow. It is important that the abstract clearly describes relevant features of the participant/s (including clinical details where appropriate), defines the depen- dent and independent variables, along with the SCED design used to examine their relationship. The target behaviors and any addi- tional outcome measures (e.g., for generalization) used in the study should be specified, along with the way in which the target behavior is measured. An accurate summation of the outcomes of the study needs to be clearly detailed, along with disclosure of any harms or adverse events, and conclude with a brief, cautious appraisal of the significance of the research.","Item 3—Scientific background: Describe the scientific back- ground to identify issue/s under analysis, current scientific knowl- edge, and gaps in that knowledge base. Verb production problems are an extremely common and pervasive aphasic deficit following stroke. Past research into word retrieval and production has mostly focused on nouns. More recently there has been increased interest in verb retrieval and verb processing disturbances [ref]. Unfortunately. One potentially useful speech pathology treatment for word production deficits involves the use of arm and hand gestures. What remains to be developed is empirical evidence to support or refute the suggestion that gesture is a potent treatment for verb retrieval deficits. This paper presents evidence. (Rose & Sussmilch, 2008, pp. 692–693) Explanation. The introduction is normally a discursive text that overviews the relevant literature and identifies the gaps in knowledge that the current study aims to address. Ideally the text is succinct and targeted to the main issues that frame the context of the study. The introduction should commence with what is known about the problem area and its interventions, what is yet to be understood, and how this study can address this gap. Item 4 —Aims: State the purpose/aims of the study, research question/s, and, if applicable, hypotheses. The purpose of the present study was to examine the effects of a written cueing treatment programme on verbal naming ability in two adults with aphasia. Treatment involved using a written cueing hierarchy, which was modelled after CART [Copy and Recall Treatment] and included verbal and writing components. (Wright, Marshall, Wilson, & Page, 2008, p. 524) Explanation. The purpose and aims of the study need to be clearly described, normally at the conclusion of the Introduction. These should take the form of research questions that define the independent variable, the dependent variable, and report if a formal relationship was assessed. Statement of aims and, if applicable, hypotheses, provide the reader with explicit directions regarding the way in which the design, methods and results should be read, given that these should all follow from the aims/research questions being asked. SECTION 3: METHOD—DESIGN (INCLUDING BOTH DESIGN STRUCTURE, AS WELL AS BROADER ASPECTS OF INTERNAL AND EXTERNAL VALIDITY), PARTICIPANT/S, CONTEXT, APPROVALS, MEASURES AND MATERIALS, INTERVENTIONS, ANALYSIS (ITEMS 5–18) -------------------------------------------------------------------------------- Item 5—Design: Identify the design (e.g., withdrawal/reversal, multiple-baseline, alternating-treatments, changing-criterion, some combination thereof, or adaptive design) and describe the phases and phase sequence (whether determined a priori or data- driven) and, if applicable, criteria for phase change. Example 1: Withdrawal/reversal design. An A-B-A-B single-subject design evaluated a token economy for increasing exercise in children with CF [cystic fibrosis]. Two advantages of this design are that it provides two evaluations of the treatment compared to baseline and it ends on a treatment phase, which is important from a clinical standpoint. The exercise diary data were used to determine when study phase changes were made. The specific criteria were the following: (a) there were three or more stable data points, (b) there was a predictable pattern in the data, or (c) there was no pattern, but the data points were predictably random. (Bernard, Cohen, & Moffett, 2009, p. 354–357) Example 2: Multiple-baseline design. A randomized, concurrent, multiple-baseline single-case design was applied.13 Participants completed repeated measurements during a baseline phase (phase A), an intervention period (phase B, 12 weeks) and a postintervention period (phase A = ). Phase A acted as a control and was therefore compared with phases B and A=. (Hoogeboom et al., 2012, p. 2) Example 3: Alternating- treatments design. The study used a multielement design with no baseline [ref] and a final “best treatment” phase [ref] to compare the effects of three contingent consequent events (Treatment A, adapted toys and devices; Treatment B, cause-and-effect commercial software; and Treatment C, instructor cre- ated video programs) on the frequency of stimulus activations. Students received intervention in alternating treatments followed by the best treatment phase, in which only the most effective intervention was delivered. (Mechling, 2006, p. 98) Example 4: Changing-criterion design. [This study evaluated] a prompting and shaping intervention, in the form of a changing criterion design. [ to] assist this athlete in the technical development of his vault and corresponding height cleared. A photo- electric beam. was set across the runway at a height of 2.30 m (5 cm above the mean height obtained during baseline). Intervention at this height continued until the participant displayed stability in his performance at a 90% level. Stability in performance referred to the successful repetition of three or more 90% performances. Following successful completion, intervention was also administered at the following heights: 2.35. 2.52 m. (Scott, Scott, & Goldwater, 1997, pp. 573–574) Explanation. In a broad sense, the design of a SCED encompasses all components of the methodology. For the purpose of this item, it specifically refers to the basic structure of phase sequencing and phase onset used in the investigation. That is, the design defines how the independent variable is manipulated, what the baseline conditions are, and when/how the independent variable is introduced/changed. Manipu- lation of these parameters allows the investigator to systematically control the independent variable in order to demonstrate the experimental effect. Moreover, using multiple phase changes in the design provides opportu- nity for repeated demonstration of the experimental effect on the same participant. This constitutes direct intrasubject replication (see Item 7 for intersubject replication and systematic replication). The design of SCEDs is crucial for determining the adequacy of control of threats to internal validity and the experimental effect (Horner et al., 2005; Kratochwill et al., 2013). Design characteristics need to be clearly and specifically reported, otherwise it is difficult for the reader to determine if the study has sufficient experimental control to ade- quately establish a functional cause- effect relationship between the de- pendent and independent variables and (b) evaluate the reliability of the results. In a small random sample of 20 reports using a single participant archived on the PsycBITE database, 45% of reports did not provide any information on the type of design used and an additional 20% incorrectly described the design (Tate et al., 2013b). Specific reporting requirements will vary among the experimental designs. A major strength of SCEDs is their flexibility and adaptability. In addition to the “classic” designs described below, elements of these designs can be creatively combined in a multitude of ways depending on the research/clinical question being addressed and the intervention being considered (e.g., see e.g., see Hayes, Barlow, & Nelson-Gray, 1999). A fundamental requirement in reporting SCEDs is that the basic structure of the design is explicitly and accurately described. This consists of reporting the following eight invariant features that apply to all designs: Basic information required for all designs, including withdrawal/reversal (A-B-A-B) design. The type of design (e.g., withdrawal/reversal) The number of phases (including baseline, experimental, maintenance and follow-up phases) The duration (length) of each phase The order in which the phases are sequenced (e.g., random- ized, counterbalanced, data-driven) The number of sessions in each phase The number of trials within each session in a phase (i.e., occasions when the dependent variable is being measured) The duration of sessions The time interval between sessions (see also Item 20) Additional descriptive information needs to be provided for spe- cific design types. Additional information required for multiple- baseline designs. The number of different (i.e., multiple) baselines (also referred to in the literature as data series, tiers, levels or legs) that the design contains. A graph alone is insufficient, because a graph represents the results that were obtained, rather than the design as planned, which may change in response to the intervention (see Item 6). Whether the baselines are across participants, behaviors or settings The method for determining treatment onset (e.g., response guided, randomization), or, if necessary, that there was no specific rationale or empirical basis Sometimes multiple-baseline designs also incorporate either a follow-up phase after the initial intervention phase or an “em- bedded” design (e.g., alternating-treatments). When such a vari- ant is utilized, the complete sequence of phases should be clearly stated in the design description. Whether or not the onset, and subsequent continuance, of data collection in each of the baselines occurred concurrently (i.e., at the same points in time) or nonconcurrently. If nonconcurrent, provide a rationale for this choice. Additional information required for alternating-treatments designs. Whether interventions were administered on the same day/session (e.g., Intervention 1 in the morning, Intervention 2 in the afternoon) or different days/sessions The way in which the order of the interventions was deter- mined (e.g., randomized, counterbalanced, Latin square) The detailed phase sequence (e.g., inclusion of a baseline preceding the intervention; a final “best treatment” phase fol- lowing the intervention) Additional information required for changing-criterion designs. All criteria or decision rules used to determine when a phase change occurs Whether the criteria are set a priori or are response guided Additional information required for adaptive designs. In these designs, the structure of the investigation (e.g., phase sequence/dura- tion, interventions, variations in intervention) is not fixed a priori, but depends, on an ongoing basis, on characteristics of the data (or responses) from early (or preceding) phases, Authors should also clearly describe the following: Features or characteristics of the data (operationally defined) that are used to regulate the study design Which specific aspects of the study design, at each and all steps in the study, are determined by which specific features of the data Item 6 —Procedural changes: Describe any procedural changes that occurred during the course of the investigation after the start of the study. BST [behavioral skills training] was implemented with each child individually in two 15- to 20-min sessions. If the child did not obtain a score of 3 during an assessment after the initial BST sessions and two booster sessions, an additional assessment session was turned into a training session. For Jake, who did not exhibit the correct behavior following BST or in situ training, an incentive phase was added. (Miltenberger et al., 2004, pp. 514–516) Explanation. There are occasions when, for a variety of reasons, the proposed implementation of a study changes from the original plan. Changes may occur to any component of the study, including (a) methodological design; (b) setting in which the study takes place; (c) target behavior and any other outcome measures, with respect to content, method or frequency of administration; (d) equipment or materials to deliver the intervention; (e) use of practitioners and assessors; or (f) the intended intervention. Changes involving participant attrition or early termination of phases are addressed in Item 19. The researcher may actively initiate changes that are a departure from protocol or they may be thrust upon the researcher as a result of external factors. With respect to changes in the intervention, one of the reported strengths of single-case methodology is the flexibility of implementation of the intervention (Connell & Thompson, 1986; Gravetter & Forzano, 2009). If adverse events occur or the intervention is not working sufficiently, then it is acceptable for the researcher to make alterations without necessarily compromising experimental control. Authors need to report any changes or departure from the original plan or protocol, such as those listed above, along with reasons. They should also provide a statement about their impact on the interpreta- tion of the results. Places to describe procedural changes in the report will depend on the type of change/s, but either the Method or Dis- cussion sections are appropriate. Item 7—Replication: Describe any planned replication. Example 1: Systematic replication. We employed a multiple-baseline design to test the efficacy of a recently developed approach for reducing school refusal behavior. To maxi- mize external validity, the intervention was tested using a systematic replication strategy, whereby only the major conceptual elements of the intervention were retained from previous applications. (Chorpita, Albano, Heimberg, & Barlow, 1996, p. 281) Example 2: Direct intersubject replication. A single- subject ABAC design with replication across three participants was employed. (Leon et al., 2005, p. 96) Explanation. Replication is a key feature of single-case method- ology and is important because it has the capacity to inform whether and the extent to which an intervention is generalizable and hence is important for external validity. In spite of its central significance, replication is not commonly reported, with just over half of the studies (54%) in the neurorehabilitation survey of Tate et al. (2014) describing replication. In addition, although it may be evident that there is replication in a study, it is not always explicitly stated. Three types of replication are described in the literature (Barlow et al., 2009; Gast, 2010a; Horner et al., 2005; Kratochwill et al., 2010, 2013; Sidman, 1960). Direct intrasubject replication refers to replication of the experimental effect within the design and addresses issues of internal validity (see Item 5). The other two types of replication are relevant to this item: (a) systematic replication (i.e., repeating the experiment with the same intervention but systematically changing characteristics of the in- dividuals, setting, interventionists and/or behaviors) and (b) direct inter- subject replication (i.e., repeating the same study but with additional individuals). Using a series of replications, Horner et al. (2005) and Kratochwill et al. (2010, 2013) note that it is possible to provide a strong basis for causal inference. They have proposed criteria for the purpose of establishing evidence- based treatments: (a) a minimum of five methodologically strong research reports, (b) conducted by at least three different research teams at three different geographical locations, and (c) with the combined number of cases being at least 20. Authors should clearly and specifically state the number and type of replications in the Abstract and the Method section. If the study is using systematic replication to build an evidence-base, authors may consider using the term systematic replication in the title. Authors need to clearly indicate whether the replication refers to (a) the replication of a previously published study or (b) intersubject repli- cation within the current study. Item 8 —Randomization: State whether randomization was used, and if so, describe the randomization method and the ele- ments of the study that were randomized. Example 1: Randomized sequence. Two treatments, one imitative and one cognitive-linguistic, were employed and treatment order was determined randomly. (Leon et al., 2005, p. 96) Example 2: Randomized sequence (with restricted randomization). A single case randomised experimental design with 12 phases was used. Each phase lasted for one week. During six treatment phases, participants wore the equipment and received cues (contingency elec- trical stimulation). During six no treatment phases, participants wore the equipment, but no cues were received. The phases were administered in a random order but always starting with a treatment phase, and no more than two consecutive phases were the same. (Wenman et al., 2003, p. 449) Example 3: Randomized onset. The start of the treatment phase was determined randomly for each participant, given the restriction that the baseline phase should last for at least 6 weeks (42 days) and at most 12 weeks (84 days). This means that the treatment phase could start on any day between the 42nd and the 84th days, resulting in a total of 43 possible assignments. (ter Kuile et al., 2009, p. 151) Explanation. The concept of randomization in SCEDs differs from its application in between-groups designs. In between-groups designs, randomization exclusively refers to allocation of participants to intervention groups (i.e., experimental vs. control). By contrast, in SCEDs, it refers to (a) the random sequencing of baseline and inter- vention phases, (b) the random determination of the commencement time for each phase, and (c) the combined randomization of both phase order and phase starting point (Kratochwill & Levin, 2010). If more than one participant is being studied in withdrawal/reversal designs, individuals can also be randomly assigned to intervention conditions. In multiple-baseline designs, random allocation of partic- ipants, behaviors or settings to each baseline of the design can also be implemented. Ferron and Levin (2014) and Kratochwill and Levin (2010) provide descriptions of these options. The sequencing of baseline and intervention phases in SCEDs may be randomized using either simple (unrestricted) or blocked (restricted) randomization strategies (hence, randomized order de-signs). The commencement point in time of each phase may also be randomized using simple randomization (hence, randomized start point designs). Randomization of both phase order and phase start point can also be combined in any given design (Kratochwill & Levin, 2010). Random allocation in SCEDs provides control of potential confounders related to time (Edgington, 1996; Kratochwill & Levin, 2010; Onghena & Edgington, 2005) which addresses at least two potential sources of experimental bias in SCEDs: history and maturation. Researchers can also use randomization to assign specific stimulus items to different stimulus sets (e.g., treated and nontreated stimuli), although this type of randomization does not control for experimental bias related to history and maturation. When randomization is used, authors should provide a reason for why it was used. They should also report specific details of (a) the basic randomization strategy (i.e., simple or restricted) and (b) those aspects of the design that were randomized (e.g., phase order, phase commencement, allocation of participants to interventions, allocation of participants, behaviors or settings to baselines). Authors need to describe any restrictions or modifications to the randomization pro- cess, which may be necessary for clinical or ethical reasons. If randomization was not used, authors need provide a reason why it was not used. They also need to report any decision criteria used to determine phase sequencing (e.g., counterbalancing), time points for phase onset (e.g., data driven), allocation of participants to interventions (if applicable), and/or allocation of participants, behaviors or settings to tiers. If decision criteria are based on participant-related reasons (e.g., clinical considerations, severity of behavior, participant needs), these need to be reported. Item 9 —Blinding: State whether blinding/masking was used, and if so, describe who was blinded/masked. Example 1: Blind assessor. The PQS [Psychotherapy Process Q-Sort] raters were not involved in any other aspect of the study procedures, and had no prior information regarding intended treatment approaches, design, or hypotheses. The raters were blind to the session number, treatment, and phase. Ratings of study tapes were made as part of PQS ratings of sessions from a larger sample [ref], and therefore were not rated consecutively or in comparison to one another. (Satir et al., 2011, p. 406) No suitable examples of blinding of participant or practitioner in the SCED behavioral sciences literature were identified. Explanation. Blinding (or masking) “refers to keeping trial participants, investigators (usually healthcare providers), or assessors (those collecting outcome data) unaware of an assigned intervention, so that they are not influenced by that knowledge” (Schulz & Grimes, 2002, p. 696). Lack of blinding or inadequate blinding in RCTs can inflate estimates of intervention effects by up to 17%, especially if trial outcomes involve subjective measures (Wood et al., 2008). It is not unreasonable to expect that this also applies to SCEDs. Blinding in SCEDs is reported very rarely, particularly for participants and practitioners (Tate et al., 2013b). Blinding is difficult to achieve in nonpharmacological trials involving surgery, psychological interventions or rehabilitation (Bang, Ni, & Davis, 2004; Boutron, Tubach, Giraudeau, & Ravaud, 2004). Given that the majority of SCEDs involve nonpharmacological interven- tions, this is a pertinent issue. Although difficult, blinding of persons providing the intervention in nonpharmacological trials is not insurmountable (e.g., e.g., Edinger, Wohlgemuth, Radtke, Marsh, & Quillian, 2001). Boutron et al. (2007) provide a selection of strategies that can be used for blinding participants. By contrast, blinding/masking of assessors is usually feasible. In reporting SCEDs, if blinding was not implemented, authors should state the reasons. When blinding was implemented, authors should clearly report who was blinded, and how the blinding was achieved. Going beyond basic reporting standards, authors may wish to consider reporting on procedures to assess the effectiveness of the blinding procedures and the outcome. Item 10 —Selection criteria: State the inclusion and exclusion criteria, if applicable, and the method of recruitment. We advertised the project in community newsletters and notices sent to local hospitals and rehabilitation professionals who typically worked with individuals with brain injuries. In the advertisements, we stated the following inclusion criteria: parents must have a documented brain injury, the children should be under the age of 10, and they must be demonstrat- ing behavioral difficulties with the injured parent. Through subsequent observation of the parent and child, we determined whether the child was demonstrating a serious level of oppositionality with the parent (e.g., noncompliance to more than approximately 50% of parent-delivered requests). (Ducharme, Davidson, & Rushford, 2002, pp. 586–587) Explanation. Readers of SCEDs need to know as much as pos- sible about the participant(s), within the boundaries of anonymity, because, until generality has been demonstrated, results are only representative of the conditions under which the investigation was conducted and for the individual/s who participated. In situations where participants were actively recruited into the study, inclusion and exclusion criteria should be provided. This in- formation will assist with the identification of factors that may influ- ence a participant's response to the intervention. It will also give an indication of the extent of the replicability and generalizability of research findings (see also Item 11). The description of the selection process should provide detail regarding who was recruited, and also the way in which the participants were recruited (such as newspaper advertisements, online recruitment targeting specific users, snowball- ing methods via other study participants or relevant professionals, distribution of brochures and leaflets). Information should be provided about the way in which selection criteria were applied (e.g., by use of diagnostic instruments including questionnaires and interviews). Details of methods, instruments ad- ministered, and classification or assessment criteria need to be clearly defined. Readers need to be able to unambiguously identify how participant selection was accomplished for successful replication of research (Horner et al., 2005). Method of recruitment may not be applicable in SCEDs in those situations where an individual presents with an issue that needs to be addressed. Such issues may be clinical (e.g., increasing a child's food intake) or nonclinical (e.g., improving an athlete's perfor- mance). In these circumstances, authors should instead state the reasons that the individual presented to the service and the reason for the intervention. Item 11—Participant characteristics: For each participant, de- scribe the demographic characteristics and clinical (or other) features relevant to the research question, such that anonymity is ensured. Three individuals with aphasia participated in the study. All had acquired aphasia secondary to a left hemisphere stroke. See Table 1 for participant's demographic data. Based on test performance and clinical judgment, all participants had nonfluent aphasia with good audi- tory comprehension. All participants exhibited significant word retrieval difficulties. Participants’ performances on the BNT [Boston Naming Test] were reviewed. P1's word retrieval errors consisted of a mix of semantic (e.g., boat for canoe) and phonemic errors (e.g., fesmask for mask). (Rider, Wright, Marshall, & Page, 2008, pp. 162–163) Explanation. Inclusion of standard baseline participant characteristics, including demographic information and functional status, ensures that the reader understands the presentation of the participants and will be able to interpret the findings (Higginbotham & Bedrosian, 1995). It is also important for generalization (Barlow et al., 2009) and facilitates meta-analysis of multiple studies (Robey, Shultz, Crawford, & Sinner, 1999). In spite of their relevance and importance, the systematic review of Maggin, et al. (2011) found that participant characteristics were incompletely reported, even for very basic demographic features such as sex (not reported in 33% of reports) and age (not reported in 42% of reports). The description of the participants of a study should include basic demographic information such as age, sex, ethnicity, socioeconomic status, geographic location, as well as diagnoses where indicated, and functional or developmental abilities (Wolery & Ezell, 1993). Any diagnosis used should include instrumentation and scores. Lane, Wolery, Reichow, and Rogers (2007)also recommend that baseline or environmental factors which serve to influence or maintain the participant's behavior during the initial baseline should be evaluated and reported. These features will go beyond simple description of sociodemographic, medical and functional status variables. Authors need to ensure that information provided does not lead to identification of the participant. This is a particular risk in the treat- ment of rare conditions. It is also important to protect an individual's privacy if the study involves stigmatized conditions. The participant in a SCED typically, but not always, represents an individual person. The participant, however, can also be a group whose performance generates a single score per measurement period (such as the rate of a particular behavior performed by all students within a classroom during a set period of time). In that situation they are generally referred to as a “case” or “unit.” It is important that the authors opera- tionally define what constitutes the group and provide criteria for its selection (see Item 10). Authors need to specify the baseline character- istics (to the same level of detail) for each participant (Wolery & Ezell, 1993) to allow the reader to ascertain the extent to which the results are generalizable (see also Item 24). Item 12—Setting: Describe characteristics of the setting and location where the study was conducted. The children and teacher comprised the full membership of a third grade general education classroom in a rural postindustrial Northeast elemen- tary school. All observations were conducted in their homeroom class during the SSR [sustained silent reading] period. During SSR, students sat at their desks, which were arranged in two rows of desks facing toward the teacher in the front of the room. (Methe & Hintze, 2003, p. 618) Explanation. It is critical to report the setting and location of the study because these factors have implications for the generalizability and applicability of the findings (see Item 24). The context of the intervention will vary according to whether it is provided in primary, secondary or tertiary health care, classroom/educational facility or community settings. The location of the study will also vary according to whether the intervention is offered in urban, rural or remote locations, and this will have an impact on the service delivery model. The setting is of particular interest for the reporting of SCEDs for two reasons. First, it may be an inherent a priori feature of the design to introduce the independent variable across a range of settings in a con-trolled manner. Multiple-baseline designs across settings are a case in point. Second, the detailed description of the relevant location and setting is central to the replicability of the study (see also Item 7). An indepen- dent researcher may be interested in varying the location of the study as one step toward systematic replication. In this scenario, a sufficiently detailed description of the setting is requisite information. Authors need to provide detailed information on the location of the study, including the number and type of settings, as well the practitioners or providers involved. There should be sufficient detail regarding the location and setting of an intervention to enable others to evaluate how different this is from their own situation. Item 13—Ethics: State whether ethics approval was obtained and indicate if and how informed consent and/or assent were obtained. Approval for this research study was obtained from the human subjects research Internal Review Board at participating institutions. If a woman was interested, she was given a pamphlet with information about the intervention and proposed research project before leaving the hospital, or at a follow-up appointment. Women then contacted the primary inves- tigator (SMB) for more information and/or to schedule an initial assess- ment appointment. An informed consent form was read, discussed and signed before beginning the initial assessment and a copy of this consent form was given to the participant for their records. (Bennett, Ehrenreich-May, Litz, Boisseau, & Barlow, 2012, p. 166) Explanation. It is virtually a universal requirement that research involving human participants requires review by and approval from an Institutional Ethics Committee. If the SCED was implemented as part of clinical care, ethics approval might not be required, as is the case in N-of-1 trials (Punja, Eslick, Duan, Vohra, & the DEcIDE Methods Center N-of-1 Guidance Panel, 2014). Reporting whether ethics ap- proval has been obtained is not, however, a feature of all research reporting guidelines. For example, the CONSORT Statement (Moher et al., 2010a) does not include this as a checklist item. Following the CENT 2015 guidelines (Shamseer et al., 2015; Vohra et al., 2015), the SCRIBE 2016 also includes the reporting of ethical approval as a checklist item for reporting SCEDs. Written informed consent should always be secured (Mechling & Gast, 2010). There may be instances when the participant cannot provide informed consent (e.g., if the participant is a minor or otherwise legally unable to provide informed consent). In this situation, their assent to participate should be sought, and consent also obtained from legal guardians or parents. It is not sufficient to merely state informed consent/assent was obtained. Rather, the process by which consent/assent occurred needs to be described so that it is clear who provided informed consent/assent. This is particularly important with vulnerable populations or where lim- ited disclosure is necessary (National Health and Medical Research Council, 2009). Item 14 —Measures: Operationally define all target behav- iors and outcome measures, describe reliability and validity, state how they were selected, and how and when they were measured. Example 1: Operational definition of the target behaviors and how they were measured. Topographies of targeted behaviors for Bob included self- injury. Self-injury consisted of face slapping, defined as a forceful contact between an open palm and cheek. Bob also displayed spitting, defined as spittle landing within 1 foot of another person. The primary dependent variable was the percent of intervals in which maladaptive behaviors were observed during each 10-min session. Data were collected using paper and pencil during consecutive 10-s intervals cued by an audio tape throughout the entire session, resulting in a total of 60 consecutive intervals. Partial interval recording was used: during each interval observers recorded whether or not any of the target maladaptive behaviors had been observed for any portion of the interval. Observers underwent training in behavioral observation. and dem- onstrated mastery prior to participating in the study. (Treadwell & Page, 1996, pp. 65–66) Example 2: Reliability of dependent variables when using non- standardized measures. Three different categories of interobserver agreement were calculated, as follows. Parent-therapist agreement: This category involved agreement between compliance data coded by the parent and those coded live by the research therapist. In this category, interobserver agreement was obtained on 39% of sessions conducted by parents, randomly selected from each of the phases across all children. Overall agreement averaged 92% for baseline (range 82%–100%) and 98% for treatment, generalization, and follow-up sessions (range 94%–100%). (Ducharme et al., 2002, p. 588) Explanation. A single item in the SCRIBE 2016 covers all as- pects of the dependent variable/s: the what, how, and when of mea- suring the effect of the intervention. As with other research designs, the measurement process in SCEDs also needs to be valid and reliable. Validity is enhanced by selecting dependent variables that are (a) relevant to the behavior in question and that best match the interven- tion, as well as (b) accurate in their measurement, which is facilitated when the behavior that is targeted for intervention is operationally defined. Reliability is enhanced when the behavior is measured in a manner that yields consistent results. What is measured? SCEDs commonly use a variety of dependent variables that play specific roles in the experiment. The primary outcome variable in SCED methodology is referred to as the target behavior. Target behaviors have three defining features in order to enhance quality of the study and minimize bias: they are specific, observable and repli- cable (Barlow et al., 2009). Other dependent variables frequently used in SCEDs may be considered akin to secondary outcome variables: Gener- alization measures are increasingly recognized for their important role in contributing to the external validity of the study (see also Item 24). In addition, SCEDs have a strong tradition in promoting experiments that address socially relevant behaviors and interventions. Additional mea- sures are often incorporated into a SCED to specifically measure social and ecological validity. Authors need to provide operational definitions of the target be- havior, which should be objective, clear and complete (Kazdin, 2011), in order to convey what does and does not constitute an instance of the dependent variable. In studies where there is more than one target behavior, the report should clearly identify and describe each of the target behaviors in detail. Other dependent variables used in the study (e.g., generalization measures, social validity measures) should also be described with equal clarity and precision. How is it measured? The “how” of measurement covers measure- ment procedures, including who selected and measured the target behaviors and other outcome variables, along with their training in the assessment procedures. Authors also need to provide information on the way in which the dependent variables were measured, justification for the selection of those measures, and detail regarding what consti- tutes a correct or incorrect response. Because the target behaviors are highly specific to the presenting case in SCEDs, formal psychometric evaluation of the measures will generally not have been established. It is therefore recommended practice that evaluation of interobserver agreement on the target behavior is conducted and reported. When standardized instruments are used, it is essential that psychometric details regarding reliability, validity and responsiveness of the instruments are reported. Any equipment used to measure the dependent variable/s should also have established measurement properties, which, if available, should be provided in the report. When is it measured? A distinctive feature of SCEDs, in contrast to group methodology, is that the target behavior (primary outcome variable) is measured repeatedly and frequently throughout all phases, including the baseline and intervention phases. Smith (2012, p. 519) observes that “the baseline measurement represents one of the most crucial design elements of the SCED.” In spite of this, baseline data were not available for 22% of SCEDs in Smith's systematic review. Moreover, in other phases of the study, the number of data points could not be readily identified. Target behavior/s need to be selected so that they are suitable for repeated and frequent measurement. Previously, the recommended minimum number of data points per phase was three (Barlow & Hersen, 1984; Beeson & Robey, 2006), but more recently professional guidelines recommend a minimum of five data points per phase (Horner et al., 2005; Kratochwill et al., 2010, 2013). There should be a clear description of the number of sessions in which the target behavior is measured in each phase, as well as the number of times it is measured in each session (i.e., number of trials per session). Frequency of measurement of other out- come variables depends on their role. Recommended practice is that generalization measures are probed continuously throughout all phases (Schlosser & Braun, 1994). By contrast, evaluation of social validity can only logically occur after the intervention has taken place. As with target behaviors, authors should clearly state the frequency and regularity of measurement of all other outcome measures, including phases during which such measures were taken. Item 15—Equipment: Clearly describe any equipment and/or materials (e.g., technological aids, biofeedback, computer pro- grams, intervention manuals or other material resources) used to measure target behavior/s and other outcome/s or deliver the interventions. Training in use of email interface: “Participants 1–4 used a mouse connected via USB port to activate the e-mail program. Participant 5 used a trackball instead of a mouse to accommodate his motor impairment. The instructor used a number pad. to control presentation. The program was run on. a laptop. An altered interface was also developed to assess generalization to a slightly differ- ent platform. It included additional buttons. as well as rearrangement of existing buttons to novel positions.” (Ehlhardt, Sohlberg, Glang & Albin, 2005, p. 571; p. 576) Photo Cue Cards [ref] were used as instructional stimuli. Four target photographs and four control photographs were selected for each child. Target stimuli are presented by child and instructional condition in Table 2. A stopwatch was used to time the length of experimental sessions. (Holcombe, Wolery, & Snyder, 1994, p. 53) Explanation. Target behaviors may be measured using behav- ioral observations and standardized scales (see Item 14) or, alterna- tively, equipment. Similarly, many interventions used in the behav- ioral sciences are accompanied by materials and equipment. Complete and accurate reporting of equipment and materials is central to the issues of replicability (see Item 7) and generalizability (see Item 24). A detailed description of the independent variable (i.e., the intervention) will include not only a description of the elements of the intervention (see Item 16), but also any specific equipment used, along with the way in which the equipment operates. Such equipment will include training manuals, computer programs, bio- feedback techniques or any other materials required to implement the intervention. When equipment is used to measure the dependent variable, the way in which it operates needs to be described, as well as its calibra- tion. In addition, its measurement properties, if available, should be reported (see Item 14). Item 16 —Intervention: Describe the intervention and control condition in each phase, including how and when they were actually administered, with as much detail as possible to facilitate attempts at replication. Check-in sessions were scheduled once a week during baseline to collect paperwork, monitor participant functioning, and to serve as an active waitlist control condition. Following the baseline period, women entered the inter- vention phase, which consisted of eight weekly sessions lasting approxi- mately 60 min. The therapist was the lead investigator of this study (SMB) who was a senior graduate student at the time of data collection. Women were told they could opt to include their partners in treatment sessions if they wished. (Bennett et al., 2012, p. 166) This description is accompanied by a detailed table (p. 165) de- scribing the session, primary goal of each session, and brief description of session content and homework. Explanation. Evidence of inadequate description of the intervention abounds in the broader health-intervention research literature using group methodology (e.g., Boutron et al., 2008; Dijkers et al., 2002; Glasziou, Meats, Heneghan, & Shepperd, 2008). The importance of describing the independent variable itself in sufficient detail to allow replication has been emphasized in the general SCED literature. The independent variable needs to be operationally defined, similar to the descriptions of the target behavior/s and the outcome measures (see Item 14). In situations where more than one intervention is used (e.g., A-B-A-C-A-D, or alternating-treatments design) each intervention should be described. In addition to ancillary aspects of the intervention (i.e., materials, manuals, stimulus items, equipment, software programs or applications; see Item 15), the nature of the intervention per se needs to be clearly specified. This description includes information on who delivered the intervention and in which mode, whether it be individual, group, distance, carer- or educator-focused, or whether telehealth or other technologies were used. It is critical to report the exact number, duration and frequency of the intervention sessions (Baker, 2012; Warren, Fey, & Yoder, 2007). Specifically, intervention intensity or dosage should be described in terms of the following: (a) dose form (i.e., the typical task or activity being used), (b) the dose (i.e., the number of times an active ingredient or teaching episode occurs per session), (c) the dose frequency (i.e., the number of intervention sessions per unit of time), (d) the total intervention duration (i.e., the total period of time in which an intervention is provided), and (e) the cumulative intervention intensity (i.e., the product of dose by dose frequency by total intervention duration; Warren et al., 2007). SCEDs typically compare the effect of an independent variable with either a control condition or another independent variable (i.e., another intervention). The control condition, often called the baseline condition in SCEDs, represents a period of time when the participant's target behavior is recorded repeatedly before the intervention is introduced. The same degree of specificity that is used to describe the experimental intervention should also be used for the control condition. Item 17—Procedural fidelity: Describe how procedural fidelity was evaluated in each phase. Adherence to treatment procedures was accomplished by use of a manual that described the steps of the programme. In addition, an independent, trained observer watched videotapes of one training session from each step of the training programme (25% of sessions), and tallied the opportunities for the experimenter behaviours and the number of times the experimenter used the expected behaviour (e.g., prompts, models, reinforcement). Adherence to the protocol was calculated using the formula (EA × 100)/ET, where EA = the experimenter behaviours. Procedural reliability was 95% to 100% for each step. (Hickey, Bourgeois, & Olswang, 2004, p. 630) Explanation. Wolery (1994) describes the collection of proce- dural fidelity data as having three main functions: (a) to monitor the occurrence of relevant variables, (b) to provide documentation that the experimental conditions occurred as planned, and (c) to provide information to practitioners about the use of the interventions. In spite of the seriousness of unreliable implementation of the experimental conditions, fidelity checks in SCEDs in the behavioral sciences are infrequent (27% in the series of Didden et al., 2006). This result is comparable to the data on reporting quantitative measurement of the independent variable in three reviews of the contents of Journal of Applied Behavior Analysis (Hagermoser Sanetti & Kratochwill, 2007). In recognition of the critical importance of the fidelity of the imple- mentation of study protocols, an item addressing adherence to the pro- tocol was introduced for the CONSORT Extension to Nonpharmacologi- cal Treatments (Boutron et al., 2008). Application of a prepared checklist or steps of a protocol is the best way to document procedural fidelity (see Borrelli et al., 2005; Dane & Schneider, 1998; Gast, 2010b, pp. 99–101). Steps taken to evaluate procedural fidelity should be reported by authors, including how it was measured and the results of its assessment. If procedural fidelity is evaluated during the course of the study, and found to be suboptimal, authors may consider going beyond basic reporting standards to describe whether and what steps were taken to improve procedural fidelity (Hagermoser Sanetti & Kratochwill, 2014). Item 18 —Analysis: Describe and justify all methods used to analyze data. Example 1: Use of visual analysis. The split-middle technique [ref] was employed to detect changes in the number of successful shots within phases and resultant trend lines (Barlow & Hersen, 1984). White proposed that level, slope, and mean score of the celeration line (or trend) line be assessed as three descriptive analyses for conclusions. Given that a point on the celeration line does not actually explain the performance level, for brevity we have chosen to concentrate on the slope of the celeration line. (Mesagno, Marchant, & Morris, 2009, p. 136) Example 2: Use of statistical analysis. As serial dependence. can bias the visual inspection,17 we checked our data in each phase for serial dependence using the lag-1 method.12 If data were found to be significantly correlated, we transformed the data using a moving average transformation, in which the preceding and succeeding measurements were taken into account.12,16 In addi- tion, randomisation tests for multiple-baseline single-case designs were carried out. We expected phases B and A9 to be superior to phase A in terms of our health outcome assessment. Therefore, we tested the null hypothesis that there would be no differential effect for any of the measurement times using a randomisation test of the differences in the means between the preintervention phase and the intervention or postintervention phase.17 A p value of <0.05 was considered statisti- cally significant. For the premeasurements and postmeasurements, we considered change scores of 20% on validated questionnaires as clin- ically relevant.32 We used Stata/IC 10.1 for Windows for the descrip- tive and visual analysis of the data and R version 2.14.1 for the randomisation tests.31 (Hoogeboom et al., 2012) Explanation. Both visual and statistical techniques can be used to analyze SCED data (see Appendix). They are considered complementary rather than mutually exclusive (Maggin & Odom, 2014; Parker & Brossart, 2003; White, Rusch, Kazdin, & Hartmann, 1989) and should, arguably, be used in combination (Davis et al., 2013; Smith, 2012). Visual analysis relies upon visual inspection of the graphed data to draw conclusions regarding the reliability and consistency of intervention effects (Lane & Gast, 2014). In the past, experts have argued that visual analysis is the most sensitive and appropriate way to detect intervention effects in SCEDs (e.g., Parsonson & Baer, 1986). It has been the traditionally preferred and most frequently used approach (Busk & Marascuilo, 1992), but has significant limitations (for discussion, see Lane & Gast, 2014; Smith, 2012). Authorities have proposed guidelines for systematiz- ing visual analysis (e.g., Kratochwill et al., 2010; Lane & Gast, 2014), but there is not yet complete agreement about decision- making criteria to guide the process. Statistical analyses have advantages in that they (a) use an explicit set of operational rules and replicable methods, (b) provide a direct test of the null hypothesis, (c) utilize precisely defined criteria for significance, (d) are useful when there is instability in the baseline or treatment effects are not well understood, and (e) can help to control for extraneous factors (e.g., Kazdin, 1982a, 1982b). There is no “universal gold-standard,” however, and authors should choose the analytic method for SCEDs that is guided by a number of considerations: (a) design requirements and data assumptions, (b) the research question being posed, (c) features of the data being analyzed, and (d) the interpretability, ease of computation and proven validity of the technique for analysis of SCED data. For detailed discussion of these issues, see Manolov, Gast, Perdices, and Evans (2014) and Shadish (2014). If statistical analyses are used, authors need to report whether underlying assumptions and other requirements pertinent to the technique were evaluated for the data set being analyzed. This informa- tion will allow the reader to determine the suitability of the analytic methods used. It is critically important that authors fully and clearly describe the method/s of analysis used, regardless of whether they select visual, statistical or both techniques. Authors should state if the method/s of analysis was prespecified before commencement of data collection, and those changes (if any) that were subsequently made as a consequence of limiting/problematic features of the data. The rationale for selecting the analytic technique in terms of its appropriateness to the study design and the research question should be clearly stated, and the source reference describing the technique should be cited. In the case of visual analysis, it is important to clearly report the features of the data that were selected for analysis (along with reasons) and whether a systematic protocol was used and, if so, which one. SECTION 4: RESULTS—SEQUENCE COMPLETED, OUTCOMES AND ESTIMATION, ADVERSE EVENTS (ITEMS 19–21) -------------------------------------------------------------------------------- Item 19 —Sequence completed: For each participant, report the sequence actually completed, including the number of trials for each session for each case. For participant/s who did not complete, state when they stopped and the reasons. Example: Deviation from protocol: Interruption of treatment. Ms. O had 24 sessions of psychotherapy over a one hundred and 47 day period, which included two periods of treatment interruption; once in the middle of the BCT [Behavior Change Treatment] phase, and once at the end of the BCT phase. Ms. O was randomly assigned to receive AFT [Alliance Focused Treatment] for the first 4-week therapy phase, BCT for the second 4-week therapy phase, and AFT for the last 4-week therapy phase. The two treatment interruptions coincided with the Thanksgiving and Christmas hol- idays, and occurred during BCT only. (Satir et al., 2011, p. 406) Explanation. A major benefit of SCEDs lies in their flexibility. As described in Item 6, it is possible during the conduct of the investigation to change and fine-tune procedures and interventions that do not appear to be working. Any such deviation from the original plan of the study, however, needs to be clearly reported (see also Item 25). The present item pertains specifically to the report of procedural variations that reflect changes in any of phase sequence, phase order, and/or number of sessions actually com- pleted by each participant. Changes to phase sequence may occur in response to clinical or ethical concerns, such as subsequently deciding to commence the investigation with a treatment phase rather than the predetermined randomized sequence. For similar reasons, a phase might be pre- maturely terminated, interrupted, extended, or substituted for rea- sons of lack of efficacy, harms, periods of absence, and so forth. Post hoc changes to the randomization schedule or the intended duration or structure of phases can significantly weaken the inter- nal validity of the study and therefore need to be reported. Missing data due to attrition may bias the results, especially if this reflects intentional or systematic noncompliance (Smith, 2012). Moreover, if participants do not complete all phases as planned, there may also be insufficient phase repetitions to dem- onstrate adequate experimental control. Attrition of participants within units (such as classrooms) over time may confound inter- vention effects especially if attrition is nonrandom and associated with the intervention itself. In order to minimize such bias, Horner et al. (2005) argue that, regardless of attrition or missing data, results of any participant for whom there is data for both a baseline and an intervention phase should be reported. If techniques for dealing with missing data are used (e.g., retrospective data completion, such as completing dia- ries), this also needs to be reported, given that such techniques can introduce significant bias in the data due to incorrect recollection/recall (Bolger, Davis, & Rafaeli, 2003). Accordingly, any changes to the sequencing of phases or missing data should be reported, along with their reasons so that the reader can evaluate the integrity of the results and their interpretation. Item 20 —Outcomes and estimation: For each participant, re- port results, including raw data, for each target behavior and other outcome/s. Example 1: Provide raw data (Fig. 2). Example 2: Statistical analysis. The three measures of treatment gains are shown in Table 2. Four patients. demonstrated significant treatment gains as determined by all three measures (C statistic, effect size, modified CDC [conservative dual criteria]). One patient. did not demonstrate significant gains on any measure. The patients who improved made variable gains in other areas as indicated by a significant increase in their WAB [Western Aphasia Battery] AQs [Aphasia Quotient] (mean increase = 5.58, SD = 2.32, t = 4.81, df = 3, p < .05). (Crosson et al., 2009) Explanation. Traditionally, SCED results are reported in graphical form which is, arguably, the clearest and most unambig- uous way of depicting the major features of the raw data. At minimum, SCED data for each session should be reported in graphic form. This does not preclude the additional reporting of raw data for each session (or for each trial within a session) for each phase of the study in a tabulated format. Although tabular presentation is acceptable (and even desirable for verification in meta-analysis), relevant features (e.g., consistency across similar phases) may not be directly obvious in this format. Alternatively, authors may provide information about where the raw numerical data set can be accessed. Raw data should be reported in the results section for each measurement point/session in each phase of the study for each participant, setting and target behavior/s. Even though aggregation of data (e.g., averaging results over several sessions), may provide a clearer view of the apparent intervention effect, it may also mask or misrepresent various important features. Such features may include (a) stability of the initial baseline phase, variability and trends within a phase, (c) degree of consistency between similar phases (e.g., intervention phases), (d) the degree of overlap between baseline and intervention phases, (e) magnitude of effect latency following intervention phase onset. Readers need to be able to critically evaluate all these aspects of the data. They also need to be able to draw their own conclusions about how adequately the investigators have taken any anomalies into account when appraising the clinical value of the intervention (Barlow et al., 2009). The metric used on the horizontal axis of graphed data should be in units of real time (i.e., days, weeks, etc.) rather than session number. This information allows the reader to more clearly inter- pret and critically evaluate the results of the study. Providing an exact chronology of the time interval between sessions allows the reader to accurately evaluate patterns of consistency between sim- ilar phases and effect latency following intervention onset. Carr (2005) argues that using a real-time metric on the horizontal axis is particularly relevant in multiple- baseline designs. The reason is so that the reader can determine the order in which sessions across participants, settings or behaviors were conducted relative to each other. More importantly, it also allows the reader to see how many sessions occurred in the initial A-phase of the second baseline (or third, fourth, etc.) of the design after the intervention was intro- duced in the first baseline. Results of any statistical analyses for the target behavior/s and other relevant outcome variables also need to be reported. This report should be done for each participant and setting in the study. Irrespective of the analysis used, authors need to clearly state which phases and features of the data were compared. Any changes or deviations in a preplanned analysis strategy should also be reported, as well as the reasons that it was necessary to make such changes. For statistical analyses, the value of the calculated sta- tistic/s, standard errors, and any associated probability level should also be reported. Item 21—Adverse events: State whether or not any adverse events occurred for any participant and the phase in which they occurred. The current reporting of adverse events is rare in SCEDs where behavioral interventions have been applied. Only a single example (Hoogeboom et al., 2012) was identified where authors mentioned that adverse events were monitored, but no information was provided on the way in which adverse effects were measured, nor were results provided. Explanation. The reporting of adverse events or harms is rare in SCEDs of behavioral interventions. Their report is a more common feature in medical interventions in randomized and non-randomized trials and observational studies (Golder, Loke, & Bland, 2011), even though it is often inadequate (Papanikolaou, Christidi, & Ioannidis, 2006; Vandenbroucke, 2006). This problem is despite the reporting of harms being a specific checklist item in the CONSORT 2010 Statement (Schulz, Altman, & Moher, 2010). Although there are no clear definitions of harms, the CONSORT Extension for Better Reporting of Harms in RCTs (Ioannidis et al., 2004) has recommended harms as the preferred term to describe an adverse event (as opposed to describing the “safety” of a treat- ment). Particular recommendations include that (a) authors make specific mention of harm in the title or abstract, as well as the introduction where harms are a primary outcome measure; and (b) there is specification of the harms in reporting of results, with special attention paid to discontinuations and withdrawals due to adverse events. The guideline also recommends that authors should present the absolute risk of each adverse event (specifying type, grade, and seriousness per arm of the trial), (b) describe any subgroup analysis and exploratory analysis for harms, and (c) provide a balanced discussion of benefits and harms with emphasis on study limitations, generalizability and other sources of infor- mation on harms (Ioannidis et al., 2004). The guide can be adapted to accommodate SCED methodology. The SCRIBE 2016 recommends that authors make full and explicit disclosure of any harms or adverse events that occurred to any participant during the course of the SCED trial, including the absence of these events. Loke, Price, Herxheimer, & the Cochrane Adverse Effects Methods Group (2007) suggest a framework to enable a systematic, manageable and clinically useful way to define adverse effects. It includes a predefined classification of adverse effects as diagnosed by the clinician, by test results, or by participant- reported symptoms (e.g., pain). SECTION 5: DISCUSSION—INTERPRETATION, LIMITATIONS, APPLICA- BILITY (ITEMS 22–24) -------------------------------------------------------------------------------- Item 22—Interpretation: Summarize findings and interpret the results in the context of current evidence. Results showed that the repeated reading program combining several research-based components. improved fluency on second-grade trans- fer passages for the three participants lending support to the existing literature on repeated reading [ref]. With research on repeated reading spanning decades and numerous studies demonstrating successful out- comes. this practice holds great promise as a strategy for improving reading fluency. However, as suggested by [the metaanalysis of] Chard and al. (2009), the current research literature on repeated reading is not sufficient for it to be designated as an evidence-based practice. (Lo, Cooke, & Starling, 2011, pp. 133, 136) Explanation. An early section of the discussion needs to provide a clear and concise summary of the findings of the study, including the strength of the intervention effect and the clinical importance of the findings. The results should be interpreted in terms that are specific to the study, as well as more generally with reference to the current literature. Interpretation specific to the study will benefit from taking into account the aims of the study, along with the robustness of methodology and procedures. Item 23 on limitations of the study is a separate item in the SCRIBE 2016 Statement, but has obvious rele- vance to the item on interpretation. Vandenbroucke et al. (2007) note that overinterpretation is a common problem in observational studies. For the STROBE Statement, they advocate that caution is exercised in interpreting the findings of a study. A cautious approach to interpretation is particularly pertinent to SCEDs in instances where findings are unclear or adequate replication has not occurred (see Item 7). In terms of interpreting the study findings in the context of the current evidence from the literature, the CONSORT Statement (Schulz et al., 2010) suggests that results of clinical trials are inter- preted with respect to the knowledge base, as synthesized in system- atic reviews. In some research areas using SCEDs, meta-analyses are available and it is recommended that information from these be incorporated when available (as in the above example of Lo, Cooke, & Starling, 2011, referring to the meta-analysis of Chard, Ketterlin-Geller, Baker, Doabler, & Apichatabutra, 2009). Item 23—Limitations: Discuss limitations, addressing sources of potential bias and imprecision. The validity of comparisons between AFT [Alliance Focused Treatment] and BCT [Behavioral Change Treatment] was compromised by several factors, including the significant intervention interruptions during BCT, the administration of psychotropic medication during the study. Furthermore, there may have been significant carry-over effects from one phase to another, as skills learned during one phase could not be “un- learned”. Limitations to measurement include possible limitation to the accuracy of the self-report nature of the kilocalorie intake, which was not verified by other sources of data collection (e.g., independent observa- tion). (Satir et al., 2011, p. 417) Explanation. It is often tempting for authors to focus on the positive findings of their results and to glide over the flaws in their study. However, all studies have limitations that can either bias or confound results. Important limitations reflect threats to the validity of the study that introduce potential bias in the findings. There are many potential limitations in SCED studies that compro- mise the extent to which results are unbiased, reliable and likely to generalize (e.g., see Horner et al., 2005; Tate et al., 2013b; Wolery, Dunlap & Ledford, 2011). These limitations include the following: (a) poor matching of the design to the type of intervention (see Item 5); inadequate replication (see Items 7 and 24); (c) absence of blind- ing (see Item 9); (d) lack of randomization (see Item 8); (e) impreci- sion with respect to the description of the participant's functional abilities (and, where applicable, diagnosis), making it difficult for readers to generalize to other cases (see Items 10 and 11); (f) problems with the operational definition of the target behavior or assessment of its reliability, insufficient number of data points in some or all of the phases to meet minimum standards (see Item 14); (g) absence of information regarding procedural fidelity (see Item 17); and (h) reli- ance on visual analysis for ambiguous cases, insufficient number of data points required for specific statistical analyses (e.g., the C sta- tistic requires at least eight data points per phase), use of statistical procedures that do not deal adequately with extant features of the data (e.g., trend, variability or auto-correlation), or whose underlying as- sumptions are not met (see Item 18). In terms of adequate reporting, authors need to provide a systematic discussion of specific limitations associated with their findings (as de- scribed above) which is also contextualized within the relevant literature. In this way the reader is provided with a realistic, critical appraisal of the contribution that the study makes to the field and the extent to which the results reflect a strong finding that is likely to be repeated. Item 24 —Applicability: Discuss applicability and implications of the study findings. The results of the present study replicate the findings of [ref] with regard to the effect of using DRA [differential reinforcement of alternative behavior], nonremoval of the fork, and stimulus fading to increase variety of food intake. The study extends previous findings by showing that the intervention package was effective independent of who fed the child. The effect of our treatment on John's consumption of nonpreferred foods did not generalize across settings in the absence of intervention in the home, but multisetting training led to transfer across settings and caregivers. Our study also shows that the treatment package described by [ref] was effective during typically occurring mealtimes with regularly scheduled food types, and that the treatment was effective for increasing the number and variety of originally nonpreferred foods. (Valdimarsdóttir, Halldorsdottir, & SigurÐardóttir, 2010, p. 105) Explanation. The concept of applicability or generality is based on the assumption that inferences can be drawn from the condition in which an intervention effect was demonstrated, to other conditions based on known similarities and differences be- tween these conditions (Gast, 2010a). The reader should be provided with a discussion of implications of (a) conceptual/theoretical considerations, (b) clinical/practical consid- erations, and (c) methodological considerations, which impact on the generalizability of the findings. Conclusions could be drawn, for example, from information provided about replication (Item 7), inclu- sion and exclusion criteria (Item 10), participant characteristics (Item 11), and generalization measures (Item 14). The replication process (see also Item 7) involves an increasing number of variations to the different dimensions (most importantly participants, setting and practitioner) that can be changed with every replication (Barlow et al., 2009; Sidman, 1960). Each replication will add information regarding the generalizability of the findings. The stage of the replication process (direct or systematic; see, e.g., Gast, 2010a) should be made clear. Similarities and differences between the current and previous studies need to be explicated. Reference should be made to factors that determined baseline performance, such as described in Items 10 and 11, because these factors are hypothesized to influence relations between the indepen- dent variable and the dependent variable in a very specific way (Horner et al., 2005; Wolery et al., 2011). If outcome measures additional to the target behavior (Item 14) are used for the purpose of generalization, authors should discuss the evidence to support (or not) a functional relationship between the independent variable and the gener- alization variable in the context of a theoretical framework, if available. If responses to the intervention differed across participants (or between studies), reasons and proposed causal relationships of this finding should be discussed and new propositions made to explain the finding.","Item 25—Protocol: If available, state where a study protocol can be accessed. No published SCED studies were identified that contained information on how to access a protocol. A number of published reports indicated that the research protocol of a SCED was reviewed by an institutional research committee, but this pertains to Item 13. Several SCED protocols were identified in trial registries (Lloyd at www.anzctr.org.au, Trial ID ACTRN12611000812998; Pool at www.anzctr.org.au, Trial ID ACTRN12611000531910; Wambaugh, Mauszycki, Cameron, Wright, & Nessler, 2013; Wambaugh, Nessler, & Wright, 2013, at www.clinicaltrials.gov) but personal communication with the authors indicated that the studies were not yet published (Lloyd & Sherrington, 2011) or the published report did not make reference to the protocol (Pool, Blackmore, Bear, & Valentine, 2014; Wambaugh et al., 2013a, b). Explanation. Trial protocol availability was a new item introduced for the CONSORT 2010 Statement (Schulz et al., 2010) and is also included in the CENT 2015 Statement (Shamseer et al., 2015; Vohra et al., 2015). Moher et al. 2010a, p. 21 indicate that the protocol refers to the planned methods of the “complete trial (rather than a protocol of a specific procedure within a trial).” The rationale for including the protocol item in the CONSORT 2010 Statement was based on published evidence of the discrepancy between the trial protocol and the subsequent published report (e.g., selective reporting of results, post hoc change in the main outcome measure). Such discrepancies continue to be documented (Dwan et al., 2011). Having a protocol readily accessible makes authors accountable for any changes that are made to the planned research design and analysis. As noted, however, a special feature of SCED methodology is its flexibility in terms of modifying the design and/or intervention after the trial has commenced (e.g., if the participant does not respond to intervention or specific research problems arise). These modifications do not necessarily compromise the experiment (Connell & Thompson, 1986; Gravetter & Forzano, 2009). As a consequence, in a SCED the design and/or intervention actually received may depart from an a priori protocol. Such departure is considered acceptable within SCED methodology, as long as the authors declare such departure and provide justification for it (see Item 6). Nonetheless, the flexibility to modify the design and intervention in SCEDs does not imply that an a priori protocol is not relevant. The SPIRIT 2013 Statement (Standard Protocol Items: Recommen- dations for Interventional Trials; Chan et al., 2013) provides guidance for the report of protocols of clinical trials and can be used as a guide for preparing a protocol on a SCED study. As noted in the CONSORT 2010 Explanation and Elaboration document (Moher et al., 2010a), there are many ways in which a protocol can be made available, including trial registries, journals that publish protocols, website of the journal publishing the main results of the study, author's institu- tional website, contact with the author. Reference to the study proto- col can be made either in the text or as a footnote. Item 26 —Funding: Identify source/s of funding and other sup- port; describe the role of funders. The study was financed by the Sint Maartenskliniek Nijmegen and Woerden, the Netherlands. All authors declare: no support from any organisation for the submitted work, no financial relationships with any organisations that might have an interest in the submitted work in the previous 3 years and no other relationships or activities that could appear to have influenced the submitted work. (Hoogeboom et al., 2012, p. 8) Explanation. Journals often request that authors disclose funding sources. It is important that readers of the manuscript can make a judgment as to whether the funders have control over knowledge dissemination or whether they provided funds but the researchers worked independently and autonomously. A consistently and overwhelmingly strong association between industry support and proindustry findings in group clinical trials has been demonstrated (Sismondo, 2008a, 2008b). The likelihood of finding a positive result when backed by industry funding has been estimated in the range of 3.6 to 4.05 times greater than when not (Bekelman, Li, & Gross, 2003; Lexchin, Bero, Djulbegovic, & Clark, 2003). Although methodological quality is not necessarily compromised in industry-funded research (Sismondo, 2008a), bias may infiltrate in other ways such as the decision to use less active controls (Bekelman et al., 2003). Provision of a grant number can be helpful because it enables readers to retrieve the details of the grant. Otherwise, one needs to know the following: Did the funder provide equipment or money? Did the funder have control over where the study was published or what was published? Were the funders the authors, or did they edit the manu- script? Any conflicts of interest should be stated explicitly. This information can be provided in the body of the text or in a footnote.","We developed the SCRIBE 2016 to assist investigators in the behavioral sciences to report SCEDs with transparency, accuracy, clarity and completeness. This article provides rationale for and ex- planation of the 26 SCRIBE 2016 items. It also includes examples of adequate reporting of specific SCRIBE 2016 items, drawing on arti- cles in the published literature. Authors reporting on one specific type of single-case methodology, the medical N-of-1 trial with multiple cross-overs, will find it helpful to consult the CENT 2015 Statement (Shamseer et al., 2015; Vohra et al., 2015), which was developed for that particular methodology. We welcome feedback from users of the SCRIBE 2016, which can be made through the SCRIBE website (www.sydney.edu.au/medicine/research/scribe). Since the first CONSORT guideline appeared in 1996 (Begg et al., 1996), medical journals have continued advocating the use of pre- scriptive reporting guidelines in the CONSORT tradition. The EQUATOR network (www.equator-network.org) is a useful resource to keep up- to-date with new developments in this field. This develop- ment has not occurred in the behavioral sciences to the same degree, although many authors of intervention studies in the behavioral sciences consult relevant CONSORT Statements (e.g., CONSORT Extension to Nonpharmacological Interventions; Boutron et al., 2008). The influence of CONSORT is such that the peak medical journals require that authors address all of the criteria of the relevant guideline in their report. The benefit has resulted in improved reporting (Turner et al., 2012). It will be advantageous to the behavioral sciences if journals publishing single-case methodology also endorse use of the SCRIBE 2016 in this way."],["Does disruption of prefrontal cortical activity using transcranial magnetic stimulation (TMS) impair visual metacognition? An initial study supporting this idea (Rounis, Maniscalco, Rothwell, Passingham, & Lau, 2010) motivated an attempted replication and extension (Bor, Schwartzman, Barrett, & Seth, 2017). Bor et al. failed to replicate the initial study, concluding that there was not good evidence that TMS to dorsolateral prefrontal cortex impairs visual metacognition. This failed replication has recently been critiqued by some of the authors of the initial study (Ruby, Maniscalco, & Peters, 2018). Here we argue that these criticisms are misplaced. In our response, we encounter some more general issues concerning good practice in replication of cognitive neuroscience studies, and in setting criteria for excluding data when employing statistical analyses like signal detection theory. We look forward to further studies investigating the role of prefrontal cortex in metacognition, with increasingly refined methodologies, motivated by the discussions in this series of papers. --------------------------------------------------------------------------------","The role of prefrontal cortex (PFC) in visual awareness remains hotly debated (Boly et al., 2017; Bor & Seth, 2012; Odegaard, Knight, & Lau, 2017; Tsuchiya, Wilke, Frassle, & Lamme, 2015). One way of addressing this topic is to disrupt prefrontal activity using (theta-burst) TMS, while examining any effects on subjective aspects of visual perception. An influential 2010 study took this approach, combining theta-burst TMS (tbTMS) with an innovative signal-detection-theory (SDT) analysis in a simple perceptual decision task (Rounis, Maniscalco, Rothwell, Passingham, & Lau, 2010). A major contribution of this study was to introduce the SDT quantity - meta-d’ - which quantifies, in a principled and statistically unbiased way (assuming the model is true), how subjective reports (such as visibility or confidence ratings) discriminate the accuracy of first-order perceptual decisions (such as whether a stimulus is a square or a diamond). Rounis et al. found that tbTMS to dorsolateral PFC (DLPFC) selectively reduced visual metacognition, while first- order perceptual accuracy remained unchanged by design. This finding, if it stands up, provides strong evidence supporting a causal role for DLPFC in visual metacognition, which in turn support [at least on some theories, (Lau & Rosenthal, 2011)], a causal role in conscious perception. However, attempted replications by our laboratory failed to find any reliable effect of tbTMS to DLPFC on visual metacognition (Bor, Schwartzman, Barrett, & Seth, 2017), concluding that, across studies, there was no reliable evidence for such an effect. This study in turn has been critiqued by authors from the original paper (Ruby, Maniscalco, & Peters, 2018), who go so far as to claim, on the basis of computational modelling, that the Bor et al. study in fact supports the findings in the original Rounis et al. study. We disagree with this surprising claim. We maintain that the results of our study should be understood following our interpretation in Bor et al.: that there is not (yet) good evidence that tbTMS to DLPFC disrupts visual metacognition. In substantiating our arguments, we encounter more general issues concerning what constitutes good practice in replication of cognitive neuroscience studies, and in setting criteria for excluding data when employing statistical analyses like signal detection theory (SDT). Discussion of these issues is, we believe, worthwhile beyond the current set of studies. We therefore welcome the continuation of this productive debate. Finally, we recognise that whether or not PFC is causally implicated in perceptual metacognition and conscious perception remains an important open question, and that other sources of evidence including lesion and neuroimaging studies will play an important role in addressing this question (Boly et al., 2017; Lau & Rosenthal, 2011; Odegaard et al., 2017).","In this spirit of open and constructive debate, the remainder of this paper describes our opinion on Ruby et al. (2018). Let us first summarise our understanding of their argument, the heart of which is statistical and computational. They focus on the fact that, in Bor et al. (2017), we excluded a substantial number of subjects from our analyses, in comparison to the original analyses in Rounis et al. (2010). In their view, our exclusions were not necessary since, according to computational simulations they performed, they did not affect false positive (FP) rates. Given that in Bor et al. we found a result which they claim was in line with Rounis et al., when we did not exclude subjects, they argue that we unnecessarily lowered our chances of finding a true effect (of TMS on metacognition) when we did apply our rigorous exclusion criteria. Ruby et al. then present a Bayesian analysis purporting to show that the data from Bor et al., support, rather than conflict with, Rounis et al. We disagree with this claim, as we will explain. Ruby et al. also selectively criticize some of the methodological changes we made, raising important issues about what constitutes an adequate replication. We take the opportunity here to revisit the reasons for these design decisions in the context of this broader concern. Finally, Ruby et al. appeal to evidence from a diverse range of alternative studies when discussing the overall question of the role of PFC in metacognition and consciousness. We do not think this appeal is relevant to the specific discussion about methods and analyses. Instead, it leads to further confusion since we do not claim our analyses undermine any of these alternative studies. Subject exclusion, FP rates, and the value of ‘clean’ data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Ruby et al. motivate their simulations by saying that our main reason for subject exclusion was to avoid a high FP rate. Ruby et al. claim that excluding the subjects as we did in Bor et al. (2017) was ‘unnecessary’, on the basis of their computer simulations showing that FP rates are unaffected by such exclusions. Although we still maintain that including unreliable meta d’ results (“unreliable” meaning no longer measuring what it purports to, no longer a smooth, linear function, so that now a small change in underlying psychological components, such as hit rate, can lead to a large jump in the measure) may well increase FP rates (see below for reasons), this is a secondary issue, and is a misrepresentation of our reasoning, which is the critical point of disagreement between us. We in fact performed these exclusions based on well-established theoretical arguments, elaborated in Barrett, Dienes, and Seth (2013) and fully described in Bor et al. Specifically, we excluded subjects when they had false alarm (FAR) or hit rates (HR) either less than 0.05 or greater than 0.95, at either the Type I (objective decision) or Type II (metacognitive) levels,1 in other words values that were numerically “extreme” by being close to or at the 0 or 1 limits. In Barrett et al. (2013), we showed that it is principled to apply these criteria because, otherwise, as we stated in the Bor et al. (2017) paper, “including such extreme results in the analysis is very likely to introduce instabilities [here referred to as “unreliabilities”] in measures reliant on type I and II SDT quantities, including type II d’ and especially various implementations of meta-d’ … Specifically, since the z function (i.e. the inverse of the standard normal cumulative distribution) approaches plus or minus infinity as HR or FAR tends to 0 or 1, SDT measures such as meta-d’ can take on extreme and highly inaccurate values with such inputs. In practice, we demonstrated from our data that unstable [in this commentary referred to as “unreliable”] meta-d’ – d’ values are significantly different from stable [reliable] values” (p. 15). As this is such a critical point, let us rephrase it here: When hit rates or false alarm rates (HR/FARs) are not near 0 or 1, there is a roughly linear relationship between changes in HRs or FARs and meta d’. Such smooth, approximately linear relationships are good indicators of a reliable measure (i.e. one that accurately reflects the psychological process it is capturing, and where a small change in some psychological component, such as HR, will lead to a small change in the measure). However, when HRs or FARs become too big or too small (i.e. close to 0 or 1), a small change to them causes a big change in meta-d’. In other words, at these (numerically) extreme values, the relationship between changes in HRs or FARs changes from smooth and linear to non-smooth and non-linear, and the meta d’ measure is no longer reliable. This is clearly illustrated in Barrett et al. (2013) Fig. 5, reproduced here (see Fig. 1), where type II FARs approaching 0 or HRs approaching 1 cause meta d’ values to significantly deviate from their previously linear path. Note that the unreliability of the measure seems to be a particular issue for the SSE method, which was the main version of meta d’ used both in the Rounis et al. (2010) and our Bor et al. (2017) studies (in order to keep close to the Rounis et al. methods). Although the above figure was for simulated data, this is an empirical problem too. Fig. 2 below shows that, for all datapoints from experiment 1 of Bor et al. (2017) those that have more extreme (close to 0 or 1) FAR and HR values, which appear to the right of the graph, have a far larger spread of meta d’ - d’ values, as would be expected from the arguments above. In theory, meta d’ - d’ should have a maximum value of 0. In practice, this is violated when sample sizes aren’t large, and this problem becomes exaggerated when type I and II HRs and FARs are extreme. Note that the figure demonstrates that meta d’ prime values markedly above the theoretical maximum of type I d’ can occur with either extreme HR and FAR values at the type I or II level. To put it even more bluntly, such subjects with extreme SDT values will have a meta-d’ measure that will not reflect their metacognitive performance, since the measure is effectively breaking down due to the extreme HR or FAR. These extreme meta d’ values can have an influential effect at the population level, depending on what statistical tests are used. Ruby et al. take the view that all that matters is whether subject exclusion will affect FP (or FN) values, e.g. “we emphasize that in classical hypothesis testing there are two main types of error: Type I and Type II. Within this standard framework, any statistical decision or justification of a choice of procedure is meaningful only with respect to its impact on these two kinds of error” (p. 39). Our view instead is that the proper application of statistical techniques requires confidence that measured variables, on which inference is performed, reflect the property of interest. If it’s clear that some portion of data, due to mismeasurement (e.g. subject responses orthogonal to question of interest, or extreme HR/FAR values as in the present case) do not reflect the property of interest, then they should be excluded for this reason, regardless of statistical concerns (although usually a likely corollary is that FP or false negative (FN) rates may well be reduced). From this perspective, the main purpose of the Ruby et al. paper, namely to use simulations to demonstrate that FP rates are unaffected by subject exclusions, becomes irrelevant. It is worth noting that these principled reasons apparently motivated the authors of the original Rounis study to mitigate these problems in their original design, by instructing subjects to keep their type II responses as even as possible, thus avoiding the extreme HR/FAR values. Having taken such clear methodological steps to reduce extreme values, it would then seem reasonable for Rounis et al. to have excluded their remaining subjects that included extreme HR/FAR values, as we did ours. However, we also note that had they done this, as Ruby et al. report, they would have reduced their n from 20 to 7. In other words, their attempt to avoid the extreme HR/FAR values by this design failed. Issues with the Ruby et al. simulations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A second point of disagreement has to do with the relative value of empirical data compared to simulations. In both the Rounis et al. and Bor et al. studies, including unreliable meta d’ measures, based on extreme HR/FAR data leads to a significant result (though in both studies this was a parametric test on non-gaussian data and in our study this was the case when using only one of various metacognition measures, and only on one of two large experiments), while excluding these results leads to a null. This is entirely consistent with the notion that including unreliable meta d’ scores increases one’s chances of a false positive, which to our minds provides the most parsimonious explanation of both these results. The Ruby et al. simulations may even lend credence to this interpretation, given that they show (Table 1, supplementary materials) that power should have increased in our between subjects design (where there was a positive result on one measure without excluding subjects) after exclusion. Principled analysis of actual empirical data is preferable to claims based on computer simulation of empirical data that are then used to reinterpret that data. Simulations necessarily have a range of parameters, some of which are inevitably set with somewhat arbitrary values, as compared to rigorous theoretical and analytical approaches applied to empirical data. In addition, simulations often embed assumptions about the nature of the data generating process which are by construction in line with the analysis methods, meaning that they will tend to overestimate the sensitivity and validity of the corresponding analyses.2 Accordingly, the Ruby et al. simulations make the assumption that type I and II decisions are based on the same evidence, distributed according to a standard SDT model. Furthermore, the simulation parameters used by Ruby et al. to model our (2017) data are drawn from the original Rounis et al. study empirical data, not from the empirical data in Bor et al. As Ruby et al. themselves emphasise, the two studies are not identical (since we made some minor modifications to improve the design, as we describe in detail in our paper). In which case, it would seem more appropriate to use the parameters from our own empirical data, which are bound to differ somewhat from their dataset.3 The potential (and as yet unknown) impact of these choices about parameter values again underlines the potential dangers of relying on simulations as compared to principled theoretical analysis. In addition, the analysis of Barrett et al. (2013) suggests that including extreme data points should weaken statistical inference since, as we have already noted, small random fluctuations in hit and miss rates at extreme ends of the spectrum will translate into large fluctuations in measured values of meta-d’. Ruby et al.’s simulations of Bor et al.’s Experiment 1 concord with weaker inference in this sense, but for Experiment 2 they surprisingly find the opposite, marking an intriguing discrepancy deserving of further research. One possibility to resolve this discrepancy would be to create simulations based on distributions of empirical hit and false alarm rates in Rounis et al, and use these to generate surrogate data, rather than starting from the levels of meta-d’ as currently done in Ruby et al. This approach would avoid making assumptions about how responses are generated in response to stimuli, and would require only the assumption of trial-to-trial independence. Issues with the Bayesian assumptions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Ruby et al. claim that DLPFC tbTMS in the Bor et al. study actually caused a metacognitive deficit, on the basis of a Bayesian analysis, using a prior of 0.5, partially based on the Rounis et al. data. We are not compelled by this logic since we have argued that the results found in Rounis et al. are based on invalid data (i.e., data obtained without subject exclusion). It therefore seems arbitrary and unjustified to assign any Bayesian prior (including a neutral prior of 0.5) to the probability that the effect found in that study is true. Indeed, the primary contribution of our study was to examine what happens when only valid data are included in the analysis. If a prior is anyway to be selected, what should that prior be? Based on only valid data, both the Rounis and Bor et al. studies found a null result. In addition, the Ruby et al. paper cites a relevant recent TMS prefrontal cortex metacognition paper by Rahnev and colleagues (Rahnev, Nee, Riddle, Larson, & D'Esposito, 2016), where TMS to prefrontal cortex actually enhanced perceptual metacognition. Ruby and colleagues conclude that these results “suggest that different parts of DLPFC perform different [metacognitive] functions.” However, given that tbTMS to both an anterior and dorsolateral region of the prefrontal cortex enhanced perceptual metacognition compared to a control site, the results of this new study are evidently relevant to any putative Bayesian prior, making a metacognitive impairment due to DLPFC TMS even less likely. If two null results and an opposite effect are input into a Bayesian prior, we assume the value would be very different to 0.5 and consequently the conclusions of the Ruby et al. paper about the likelihood of our results actually reflecting a DLPFC- induced metacognitive impairment would be very different as well. Did Bor et al. unintentionally replicate Rounis et al.? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Ruby et al do not merely disagree with our conclusions in Bor et al. They in fact say, “We therefore find it striking that Bor et al.’s (2017) results appear to demonstrate a successful replication prior to subject exclusion, despite their interpreting otherwise.” Because of the contentious nature of this claim, we believe it is worth going into the details of what we did and did not find. It is true that, using one particular method of computing meta-d’ (sum-square-error, SSE), in Experiment 1, when not excluding the subjects with unreliable meta d’ scores due to extreme HR/FAR values, on a one-tailed t-test, without multiple comparisons on non-parametric data we found a significant difference between the DLPFC and control (vertex) group (p = 0.03). We included this invalid statistical test (a parametric test on non-parametric data) explicitly to make a point: “We should emphasise, however, that this significant result, aside from being uncorrected for multiple comparisons [we had 4 experimental groups and three metacognitive measures], is not to be trusted as it includes data that invalidates the (parametric) assumptions underlying the analysis. We merely include this analysis to demonstrate how the inclusion of unstable [unreliable] values could potentially generate spurious significant results.” (p. 12). It is unclear whether a one-tailed t-test was appropriate here, given any prior expectation to replicate the Rounis et al. result was potentially nullified by other more recent studies showing the opposite results, such as Rahnev et al. (2016). Had we used a two-tailed t-test, this single significant result would not have been significant. Similarly, had we corrected either for the number of experimental group or number of metacognitive measures, the result would not have been significant either. Our suspicion was that the main parametric result from Rounis et al. was based on non- normal data, as was our significant result above. Had we carried out the more appropriate non-parametric equivalent test, this result would not have approached significance (Wilcoxon Rank Sum Test W = 113, p = 0.1949). Ruby et al. confirm the non-normality of their original data in the last paragraph of their supplementary materials document, and claim that the critical non-parametric equivalent interaction test was still significant, though unfortunately without providing any details on p-values or related quantities. Moreover, in the Bor et al. case the effect (if any) seems to be driven by an unpredicted boost to metacognition in the control group, rather than a reduction in metacognitive performance in the experimental (DLPFC) group, as interpreted by Rounis et al. This solitary statistical effect therefore does not seem to replicate Rounis et al., contrary to Ruby et al.’s claim. In addition, when we used the main alternative version kindly supplied by Rounis and colleagues (maximum likelihood estimation, MLE), or a new Bayesian method by Stephen Fleming4 (HMeta-d), without exclusion, there was also no such significant effect (with unreliable measures: MLE p = 0.66, HMeta-d p = 0.65; excluding unreliable measures MLE p = 0.69, HMeta-d p = 0.54). We therefore disagree that our study constitutes an unintended replication of Rounis et al., especially when taken in combination with the rest of our analyses which rigorously failed to find evidence for any effect, including Bayes Factor robust findings for a null. Furthermore, in Experiment 2 within the Bor et al. paper, we were even less successful in attempting to replicate the Rounis et al. findings. We designed our Experiment 2 as an adaptation of a within-subject design, and was therefore in this way a closer replication of Rounis et al. Our double- within-subjects protocol (potentially having 4 sessions of TMS and testing, instead of 2) was in fact designed to be sensitive to the presence of even just a single subject (out of the original 27) who would follow the pattern of the Rounis et al. study. Had we found just one subject who replicated the Rounis et al. results over the 4 sessions of session 2, we may have been able to describe this experiment as providing a partial replication. This was indeed our intention, contrary to Ruby et al. in their abstract claiming that we “adopted an experimental design that reduced their chance of obtaining positive findings.” However, we did not find even a single subject who followed such a pattern. The first session involved tbTMS to DLPFC for all 27 subjects. Of those 27, 10 were excluded due to extreme SDT values. Out of these 17 subjects, 7 showed an effect (in either direction) equivalent to the Rounis et al. study. Our original wording from Bor et al. makes clear the nature of our non-replication: “Of the remaining 7 participants, 3 showed the expected impairment, while 4 showed a clear metacognitive enhancement following DLPFC cTBS. 6 of these 7 participants also showed a clear metacognitive change for the vertex control session, and thus were not asked to return for the 3rd session (2nd DLPFC). Only 1 participant that showed a clear DLPFC cTBS metacognitive change in the first session also showed no change for the 2nd vertex cTBS session, and thus was brought back for the 3rd session (2nd DLPFC). This session, unfortunately, included unstable SDT values [here meaning extreme HR/FAR values close to 0 or 1], and thus the participant was not asked to return for a 4th session. If these instabilities [unreliable measures] are ignored, though, the metacognitive change for the 3rd session was very similar to the 1st session. Both sessions, however, showed a robust enhancement of metacognition for this single subject following DLPFC cTBS, as opposed to the impairment found in the Rounis study.” (Bor et al., 2017, p. 13). A comprehensive appreciation of our Experiment 2 therefore reveals that not even a single participant replicated the results in Rounis et al. Even when just considering the first two sessions, which would make our Experiment 2 even more closely resemble the original Rounis et al. study with its two sessions, not a single participant (out of 27) showed both a DLPFC tbTMS metacognitive impairment, as well as no metacognitive effect when tbTMS is applied to the control site (vertex). In terms of statistical power, which Ruby et al. critique for our Experiment 2, we note that their reported power for our Experiment 2 after excluding subjects was 0.308, which was almost identical to the Rounis et al. study (without subject exclusions) of 0.311. (The greatest statistical power across both the Rounis et al. and Bor et al. experiments was in our Experiment 1 after excluding subjects: 0.409.) Thus, on the basis of power calculations, it remains difficult to explain why our experiments failed to find a positive result, while the Rounis et al. study did find such a result. However, we emphasise that modelling of power and false positive rates are a poor alternative to rigorous analysis of the empirical data itself, as our Experiment 2, in particular, demonstrates. Finally, even if discounting all the arguments we make above, relying on outliers for a positive result – as in the original Rounis et al. study – suggests, at best, only a small effect size. Design differences between the Bor et al. and Rounis et al. studies ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In their manuscript, Ruby et al. also criticise the fact that we modified the experimental design of Rounis et al. in various ways. For example, they say “[Bor et al.] made several changes to the original study design, some of which are known to undermine the chance of finding meaningful results from the outset”. Indeed, we did make small changes to the Rounis et al. design. As we made clear in Bor et al., in each case the changes were made to improve the design with respect to its suitability to address the underlying experimental question. Some of these changes should be entirely uncontroversial; for instance, in removing some minor software ‘bugs’ in the experimental scripts which Rounis and colleagues kindly shared with us, and in ensuring consistency in behavioural responses, and so on (see Bor et al., p. 3). One notable change was that we used a between-subjects design rather than (as in Rounis et al.) a within-subjects design for our Experiment 1. This was mandated by our original aim to replicate and extend Rounis et al. by exploring the potential causal contribution of other regions within the frontoparietal network to metacognition. Our Experiment 2, however, utilised a double-repeat within- subject design and still failed to find any effect of cTBS on metacognition. This double- repeat design exemplifies that the changes we made were employed to maximize the chance of finding a true effect should one exist. The other design changes in Bor et al. (2017) were also made in a principled way, and captured in full the spirit of the original Rounis et al. experiment. These changes are fully described in Bor et al. (see, for example, p. 15) and we will just reinforce two points here. First, there were small changes to task instructions: we used confidence rather than visibility ratings (to ascertain metacognition more directly and to avoid additional working memory demands). However, in attempting to avoid the issues around relative, rather than absolute, metacognitive judgements, of the original Rounis study, we recognise that asking a different metacognitive question could have led to different results. Whether such a subtle change is sufficient to turn a significantly positive finding into a null result is a question we address in the next section. We also used an active TMS control (vertex) rather than sham TMS, to provide a closer match between experimental and control conditions. Part of our justification for this change was to address demand characteristics, with sham TMS versus real TMS potentially causing subjects to demonstrate a stronger impairment for that condition (DLPFC) for which they perceive is the real condition. However, although we continue to believe that a real TMS site is preferable to sham as a control, we also acknowledge that the choice of vertex itself is not perfect. First, a vertex control would not have induced as many peripheral nerve issues (facial twitching and pain) as the DLPFC stimulation in a subset of subjects (though would be similar to the parietal TMS stimulation we used in experiment 1). Second, the vertex control meant that the same region was stimulated twice, unlike any of the experimental conditions. It is difficult to know how such a difference would affect participants both psychologically and neurophysiologically. Interestingly, at one point, Ruby et al. (p. 15) criticise that we used only two levels of confidence judgements, suggesting that using more than two would have allowed more sensitive estimation of meta-d’. However, although we agree that more than two levels would have been beneficial, we used two levels because Rounis et al. also used two levels (in their case of visibility ratings rather than confidence). Had we used more than two levels, this would have changed the design considerably, and any subsequent null result might have been attributed to such a clear departure from the original protocol. What constitutes a good replication ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ These points speak to a broader discussion of what counts as an adequate replication in cognitive neuroscience [see, for example, (Szucs and Loannidis, 2017)]. Our view is the following: Replications in cognitive neuroscience are not (yet) like replications in experimental physics – cognitive neuroscience (broadly construed) necessarily embeds a greater degree of flexibility in how experiments are designed and conducted, given the range of theoretical views, experimental methods, and analysis techniques. In addition, for any particular experimental result to be described as ‘strong’, it should be robust to small variations in paradigm, especially when such variations should be expected to improve the design according to principled criteria. A related point is that, as Ruby and colleagues admit, both the original Rounis et al. and the Bor et al. studies were relatively underpowered. We attempted to include a high n and sufficient power in our studies, but the extent of subject exclusions due to extreme HR/FAR values was substantial, which weakened power, even when the data were cleaner. Our view is that the best way forward is to collect more empirical data, rather than run simulations. We freely admit limitations in both the Rounis and Bor et al. studies, and that a new experiment with more robust power could address these limitations. A new design could also improve on both studies by having MRI-guided TMS, using a continuous confidence or visibility judgement (or at least including many more options than the binary options we both used), with a sufficiently large n, and multiple controls (e.g. both sham and real TMS at various non-DLPFC sites). To our minds, such a study could definitively answer the question of whether tbTMS to DLPFC can impair visual metacognition, and would be far beneficial to any simulation or to debates about the perceived differences between two experiments. In general, lack of power is a persistent problem in cognitive neuroscience studies (Szucs and Loannidis, 2017), and if this is combined with the bias towards positively significant results, creates an atmosphere where Szucs and Loannidis (2017) have concluded that “more than 50% of published findings deemed to be statistically significant are likely to be false.” Consequently, both the original study, and the attempted replication, should strive for high statistical power. Furthermore, given the issue of failed replications, we believe that there should be a greater onus on initial publications to demonstrate robust effects, both in terms of power and the extent of the robust statistical findings. Is the prefrontal cortex involved in metacognition? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Ruby et al. imply that we argue against any role for PFC in metacognition or conscious perception. For instance, in the introduction they write that following our null result, “they… suggest that DLPFC might not be ‘critical for generating conscious contents’.” This quoted phrase is a misrepresentation of our position, and we urge readers to read the entire discussion from Bor et al. to retrieve the full context. There we raised the possibility that our findings in Bor et al. are consistent with so-called ‘no report’ paradigms which have independently cast doubt on the role of the DLPFC in generating conscious contents (e.g. Brascamp, Blake, & Knapen, 2015), but also raised other possibilities, such as that the prefrontal parietal network might be especially plastic, and adapt within seconds to TMS administration, or that tbTMS to DLPFC at ethically safe levels simply isn’t sufficient to induce an impairment in a subtle paradigm. We nowhere deny that there is other evidence, from other studies (including, for example, lesion studies) that speak to a role for PFC in metacognition, or indeed in generating conscious contents more generally (Bor, 2016). Indeed, we completely agree with Ruby et al. when the say “Ultimately, whether theta-burst TMS to DLPFC can robustly impair visual awareness concerns the specific method and details. More important is the general question regarding the role of the prefrontal cortex in metacognition and conscious perception”. Absolutely. But the only points of debate between Ruby et al. and Bor et al. concern the specific details and methods. Simply citing evidence from other studies relating to PFC and consciousness does not contribute to the present discussion; instead it falsely suggests that one should believe Rounis et al. simply because it may possibly fit better with some other independent experiments. However, Ruby et al. agree with us that establishing the role of PFC in awareness is a very important, interesting, and open question (Bor and Seth, 2012).","In summary, we continue to think the evidence so far, across studies, does not (so far) demonstrate that theta-burst TMS to prefrontal cortex can disrupt visual metacognition. Moreover, the findings from each study depend on the fine details of methods and analyses. Further investigations would be helped by open access provision of data (as done for Bor et al., but not for Rounis et al.), and simulation code (used in Ruby et al.) – so that interested researchers can draw their own conclusions in addition to the opinions presented in this series of papers. We also look forward to further studies with increasingly refined methodologies, motivated by the discussions here."],["Rhythm and metrical regularities are fundamental properties of music and poetry - and all of those are used in the interaction between infants and their parents. Music and rhythm perception have been shown to support auditory and language skills. Here we compare newborn infants’ learning from a song, a nursery rhyme, and normal speech for the first time in the same study. Infants’ electrophysiological brain responses revealed that the nursery rhyme condition facilitated learning from auditory input, and thus led to successful detection of deviations. These findings suggest that coincidence of prosodic cue patterns and to-be-learned items is more important than the format of the input. Overall, the present results support the view that rhythm is likely to create a template for future events, which allows auditory system to predict prospective input and thus facilitates language development. --------------------------------------------------------------------------------","Across cultures, caregivers and their infants use music to interact: for example, singing play songs and lullabies, reciting nursery rhymes, and rocking infants in time with music (Fernald & Kuhl, 1987; Obermeier et al., 2013; Papousek, 1996; Shannon, 2006; Trehub & Schellenberg, 1995; Wallin et al., 2001). Previous research has shown that this behavior results in beneficial socioemotional outcomes (Cirelli, Einarson, & Trainor, 2014; Kivijärvi et al., 2001), and supports early language development. Many studies have shown the benefits of both formal and informal musical activities to auditory, linguistic and literacy skills in children (e.g. Bergman Nutley, Darki, & Klingberg, 2014; Chobert, François, Velay, & Besson, 2012; François & Schön, 2011; François, Chobert, Besson, & Schön, 2012; Kraus et al., 2014; Linnavalli, Putkinen, Lipsanen, Huotilainen, & Tervaniemi, 2018; Moritz, Yampolsky, Papadelis, Thomson, & Wolf, 2013; Overy, 2003; Tallal & Gaab, 2006; Torppa et al., 2014; for a review, see Virtala & Partanen, 2018). For example, music training allows children to detect small pitch changes (Besson, Schön, Moreno, Santos, & Magne, 2007), and children’s music experience has been found to promote verbal memory (Ho, Cheung, & Chan, 2003), reading skills (Anvari, Trainor, Woodside, & Levy, 2002; Corrigall & Trainor, 2011), vocabulary (Forgeard, Winner, Norton, & Schlaug, 2008; Linnavalli et al., 2018), phoneme awareness (Anvari et al., 2002), and detection of prosody (Magne, Schön, & Besson, 2006). Music exposure changes the auditory cortical processing (Partanen, Kujala, Tervaniemi, & Huotilainen, 2013; Trainor, Lee, & Bosnyak, 2011) and increases the number of preverbal communicative gestures in 12-month-old children (Gerry, Unrau, & Trainor, 2012). Music captures and sustains infant attention, and infants are more engaged when listening to maternal singing than to maternal speech (Nakata & Trehub, 2004). In addition, musical activities seem to modify the phase of attention fluctuations (Jones, 1976) to the periodicities in the music, thus improving rapid temporal auditory processing and enhancing neural encoding of speech (Kraus & Chandrasekaran, 2010; Schön & Tillmann, 2015; Zhao & Kuhl, 2016). There have been fewer studies that investigate solely rhythm without the music accompaniment, yet the effects of rhythm have been observed from infancy to adulthood. Metric regularity of rhythmic speech and movement might have beneficial effects on speech processing. It has been found to facilitate word segmentation in 9-month-old infants (Curtin, Mintz, & Christiansen, 2005) and in adults (Leong & Goswami, 2015; Rothermich, Schmidt-Kassow, Schwartze, & Kotz, 2010; Schmidt-Kassow & Kotz, 2009; Schön et al., 2008). Rhythm also helps lexico-semantic integration (Rothermich, Schmidt-Kassow, & Kotz, 2012) and increases aesthetic appreciation and emotional processing (Obermeier et al., 2013) in adults. Rhythm perception is associated with foreign language learning in adults (Bhatara, Yeung, & Nazzi, 2015), and grammar skills in 5–7 year old children (Gordon et al., 2015). Moreover, children with developmental dyslexia have been found to have poor rhythmic perception (2011b, Flaugnacco et al., 2014; Goswami, 2011a; Goswami et al., 2002; Huss, Verney, Fosker, Mead, & Goswami, 2011). Goswami (2011, 2018) has suggested that rhythmic auditory input causes phase synchronization in oscillations of auditory cortex, and thus results in a more efficient neural sampling of the auditory signal. Bimodal auditory and visual stimuli in synchrony are highly salient, and deploy infants’ attention efficiently benefiting their perceptual processing and memory (Bahrick & Lickliter, 2000; Lewkowicz & Lickliter, 2013; Lewkowicz, 2000; Reynolds, Bahrick, Lickliter, & Guy, 2014). Music (Lebedeva & Kuhl, 2010; Peterson & Thaut, 2007; Schön et al., 2008; Thaut, 2005) and rhythm with or without music (Purnell-Webb & Speelman, 2008) have been found to enhance learning and memory in adults. This is because the structure of music and rhythms are likely to create a template for future events. Such structure helps adaptive auditory system to pick up sound regularities in a predictive manner (Rothermich et al., 2010; Schmuckler & Boltz, 1994; Winkler, Háden, Ladinig, Sziller, & Honing, 2009). Infants can detect pitch and rhythm changes at birth (Sambeth, Ruohio, Alku, Fellman, & Huotilainen, 2008; Telkemeyer et al., 2011; Trehub & Hannon, 2006) and even before that as near-term fetuses (Graniere-Deferre, Ribeiro, Jacquet, & Bassereau, 2011; Lecanuet, Graniere- Deferre, Jacquet, & DeCasper, 2000; Partanen, Kujala, Näätänen et al., 2013). Furthermore, infants are sensitive to prosodic rhythm, and can discriminate between languages representing different rhythm classes – even when the phonemic structure of speech is degraded (Nazzi, Bertoncini, & Mehler, 1998). Thus, it seems that infants are already sensitive to prosodic information carried by slow-varying patterns of spectral and amplitude modulation. Infants have been found to be able to extract and predict auditory regularities from sound frequency, duration, timbre (Háden, Németh, Török, & Winkler, 2015; Ruusuvirta, Huotilainen, Fellman, & Näätänen, 2009; Trainor, 2012) and stream of syllables (Teinonen, Fellman, Näätänen, Alku, & Huotilainen, 2009). François et al. (2017) found that newborn infants were able to detect statistical violations in pseudo-words in melodically enriched, but not in flat speech streams. They also showed that neonatal brain responses to deviations in melodically enriched speech correlated with expressive vocabulary at 18 months. Furthermore, newborn infants can extract temporal predictions from rhythmical regularities in the auditory input (Winkler et al., 2009). However, little is known about how neonates encode continuous speech, how much they are able to learn from it, and whether music or rhythmic speech helps them to do so. Current study ~~~~~~~~~~~~~ The aim of this study was to explore whether music and rhythm can be observed to facilitate learning from auditory input in newborn infants. To this end, we presented natural continuous stimuli with similar semantic content in the form of a nursery rhyme, a song and prose to infants and recorded their auditory event-related potentials (ERPs). We then probed the learning effects by introducing infrequent changes of vowel, word, pitch, and intensity changes to speech excerpts (similar to the optimal paradigm; Näätänen, Pakarinen, Rinne, & Takegata, 2004) and by recording the neural responses to them. In line with predictive coding theory (Rao & Ballard, 1999, 2010, Emberson, Richards, & Aslin, 2015; Friston, 2005; Friston, Kilner, & Harrison, 2006; Kayhan, Meyer, O’Reilly, Hunnius, & Bekkering, 2019; Kouider et al., 2015; Pickering & Garrod, 2013), we expected that sensory stimulation results in attempts to generate an internal model of its regularities in the infant brain. When an incoming stimulus deviates from the internal model, a quick and robust brain response is elicited. This response, known as prediction error, can be thought as a feedback signal to the higher levels of neural network that adjusts the neural representations of the internal model to minimize future surprises (Friston et al., 2006; Winkler, 2007; see also Näätänen, Paavilainen, Rinne, & Alho, 2007, for the mismatch negativity (MMN) elicited in an oddball paradigm). In the case of infrequent changes in our stimuli, possible prediction error brain responses would reflect learning, because without any learning from the long continuous stimuli, the infants’ internal model could not predict incoming stimulation (for predictions in speech processing and recognition, see DeLong, Urbach, & Kutas, 2005; Gagnepain, Henson, & Davis, 2012; Ylinen et al., 2016; Ylinen, Bosseler, Junttila, & Huotilainen, 2017). We hypothesized that the rhythm present in both nursery rhyme and song, and the melody present in song would enhance the extraction of predictions from auditory input and thus facilitate learning. We expected this to result in the strengthening of the prediction error response, which in this case reflected infants’ learning as a result of exposure.","Participants (N = 21; 8 boys, 13 girls) were healthy newborn infants born into Finnish- speaking families. The mothers or both parents of the newborn gave written informed consent to participate in the study. Electroencephalography (EEG) measurements were conducted in Jorvi hospital in Espoo, Finland. The study was approved by the Ethics Committee for Paediatrics, Adolescent Medicine and Psychiatry, Hospital District of Helsinki and Uusimaa. Initially 24 newborns participated in the study, but data of three participants were excluded due to technical problems or not reaching the criterion of 40 epochs. More details about participants can be found in Table 1. Study paradigm and stimuli ~~~~~~~~~~~~~~~~~~~~~~~~~~ A Finnish version of a well-known nursery rhyme Simple Simon (Kunnas, 1997) was recorded by a female native speaker of Finnish using three different conditions: she read two verses as metrically-spoken nursery rhyme, sang two verses to a melody (see Fig. 1), and read two verses in ordinary speech. To elicit as naturalistic stimuli as possible, the speaker was not given specific instructions with respect to acoustics, and thus the conditions differ by their duration, intensity, pitch and rhythmic structure (see detail in Table 4). However, she was aware that the stimuli would be directed to infants, and was given written instructions to read the stimuli as if she was reading to a toddler. The ordinary speech was modified from the original nursery rhyme. The modifications of this prose version included converting the atypical word order of the nursery rhyme into normal one, and adding modifiers (genitive attributes) and clarifying words, the meanings of which were implicit in the original nursery rhyme but which had been omitted from the nursery rhyme because of the meter (see Table 2). The nursery rhyme used in the experiment had a regular rhythmic structure (trochee) that was characterized by regular alteration of stressed and unstressed syllables that were cued by pitch and intensity. The song had a regular beat, but as compared to the nursery rhyme that used both pitch and intensity to cue stress, there was smaller acoustic difference between stressed and unstressed syllables (i.e., stressed syllables were less prominent). Rather than to cue stress, pitch was used to convey melody. Ordinary speech was, in turn, characterized by less regular pattern of stressed and unstressed syllables, since stress placement was determined by focus in the sentence (rather than meter or beat). In order to be able to examine the acoustic structures of nursery rhyme, song and speech in more detail, we utilized a number of acoustic measures. We identified and labelled the locations of vowel-consonant and consonant-vowel boundaries to segment the stimuli into vocalic intervals (n = 169), located between the onset and the offset of a vowel, or a vowel cluster, and consonantal intervals (n = 170), located between the onset and the offset of a consonant, or a cluster of consonants (Ramus, Nespor, & Mehler, 1999). This was done in Praat (Boersma & Weenink, 2018). When boundaries could not be identified by visual inspection of speech waveforms, a boundary was placed at a point where the preceding phone could no longer be heard. The measures we used were (i) the standard deviation of vocalic interval duration (ΔV), (ii) the standard deviation of consonantal interval duration (ΔC), (iii) the proportion of vocalic utterance duration out of total duration multiplied by 100 (%V), (iv) the standard deviation of vocalic interval duration normalized for speech rate variability (VarcoV), (v) the standard deviation of consonantal interval duration normalized for speech rate variability (VarcoC), (vi) the mean durational differences between successive vocalic or consonantal intervals divided by the sum of the same intervals, the normalized pairwise variability index (nPVI), and (vii) normalized pairwise variability index for vocalic intervals (nPVI-V) (Dellwo & Wagner, 2003; Grabe & Low, 2002; Lee, Kitamura, Burnham, & Todd, 2014; Ramus & Mehler, 1999; Ramus et al., 1999; White & Mattys, 2007a; 2007b; Wiget et al., 2010). White and Mattys (2007b) found especially VarcoV useful for discriminating different rhythm topologies: higher VarcoV suggests more salient contrasts between stressed and unstressed syllables (see Fig. 2). In line with Lee et al. (2014) we also utilized the pitch extraction algorithm in Praat (with standard settings: 75 Hz pitch floor, 500 Hz pitch ceiling, 0.45 voicing threshold) to estimate the median F0 of all vocalic intervals, and then calculated mean pitch in vocalic intervals and the standard deviation of the median pitches of vocalic intervals. These measures and other acoustic details of the stimuli can be found in Table 3. We were also interested in those prosodic stress patterns in our stimuli which are not revealed by durational measures, such as syllables, words and phrases (for discussion on general problems with durational rhythm metrics, see Arvaniti, 2012). Leong (2012) introduced the idea of using nested amplitude modulations (AM) as a measure of rhythmic information in speech: modulation range 0.9–2.5 Hz represents the stressed syllable rate, and 2.5–12 Hz range the syllable rate. The amplitude envelope (AE) of speech varies relatively slowly in time, and the amplitude rises coincide with stressed syllable onset, which occurs approximately twice per second in infant-directed speech (IDS) and nursery rhymes (Goswami, 2018; Leong & Goswami, 2015; Leong, Kalashnikova, Burnham, & Goswami, 2014). For example, Leong et al. (2014) used this method to find that dominant rhythmic patterning of IDS occurs at prosodic stress feet whereas for adult-directed speech (ADS) it happens at syllable rate. Descriptive analyses of the stimuli ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We extracted the amplitude envelopes of our stimulus signals using Hilbert transform and 40 Hz low-pass filter, and compared the envelopes and pitch contours of the three stimulus types with 2-Hz sine wave by calculating generalized linear regression, where either normalized signal envelope or pitch was the dependent variable and analytical (Hilbert transformed) 2-Hz signal the independent variable (see e.g. Parkkonen, Andersson, Hämäläinen, & Hari, 2008). Analytical signal was used to accommodate the phase of the signal. A higher value means that there is more 2 Hz rhythm in the signal. As indicated by Table 4, generalized linear regression analysis suggested that the highest degree of 2 Hz rhythm was found in the nursery rhyme signal. That is, the regularity of the rhythmic structure as indicated by intensity and pitch was highest in the nursery rhyme and lowest in speech. In addition, we examined the rhythmic regularity of our stimuli signals by computing the autocorrelation functions (ACF) of the nursery rhyme, song and speech signal envelopes that had been filtered into stressed syllable (0.9–2.5 Hz) and syllable (2.5–14 Hz) modulation rates. A Fourier transform was then applied to the ACFs in order to get a periodic power spectrum for each of the nursery rhyme, song and speech ACFs at both modulation rates (Leong, 2012). Fig. 3 shows these spectra. Based on this analysis nursery rhyme condition contained higher periodic power than the other two conditions at stressed syllable modulation rate, but song condition had slightly higher periodicity at the syllable modulation tier. These six verses were presented as repetitive standard stimuli to the infants for 10.4 min (12 repetitions/verse) during a learning phase. A testing phase (10 repetitions/verse, 8.6 min), which followed immediately after the learning phase, introduced multiple deviations to the syllables with word stress. The possible deviations were the change of the whole word, vowel, sound intensity, or pitch. The deviations were inserted to all types of verses avoiding two simultaneous ones. The order of the verses was counterbalanced across participant groups (4 participants/group). There were 10 differing versions of each verse, and each altered verse included each deviation type twice (see Table 5). There was only one deviation type per altered syllable, however, and the same deviation did not repeat on each syllable more than once. Intensity change was implemented as six dB increase or decrease of sound pressure level in all conditions. Sound frequency deviations were carried out by increasing or decreasing base frequency by 20% in poetry and prose verses, whereas in song verses melody was shifted by two to five whole tones using notes that did not break the original chord progression (see Appendix A). Infants were asleep in a crib during the measurement, facing randomly either to the left or to the right, thus partly obscuring one ear. The infants can be measured when they are asleep, as the elicitation of prediction error responses does not require focused attention (Cheour et al., 2002). Infants’ hearing was normal according to the clinical routine screening done with otoacoustic emission test (EOAE, ILO88, Dpi, Otodynamics Ltd, Hatfield, UK). The measurement was conducted by a registered nurse, who observed the infant and their apparent sleep stages throughout the EEG recording. The stimuli were presented via two loudspeakers located close to the corners at the foot of the crib, the approximate sound level being 74 dB SPL. EEG recording and data analysis ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ EEG was recorded from 12 channels using Ag/Cl-electrodes placed on the infants’ head according to the international 10/20 system (NeuroScan, Synamps 2 amplifier). Electrodes were attached to F3, F4, C3, Cz, C4, P3, P4, T7, T8, the two mastoids, and below the right eye. The sampling frequency was 500 Hz and the left mastoid was used as a reference. The data were re-referenced to the average of the two mastoids and bandpass-filtered (high- pass 0.5 Hz, low-pass 30 Hz, slope 48 dB/octave). Possible eye movements were corrected, and data were divided into epochs of −100-500 ms with respect to the beginnings of the syllables of interest. Epochs with artefacts exceeding 100 μV were rejected and omitted from further analysis. Epochs were then averaged for each stimulus type with a 100 ms pre- stimulus baseline. As the paradigm included many kinds of deviants, there were on average 60.7 accepted epochs per deviant. Therefore, some of the types had to be grouped to make sure there was a sufficient number of epochs per infant for averaging. The grouping was based on whether the changes were phonological in the Finnish language or not: word or phoneme changes are phonological in Finnish but fundamental frequency and intensity changes are not. The phonological status was expected to be important because the infants born to Finnish-speaking families1 must have been exposed to the Finnish phonology before birth. Based on previous findings on fetal learning (Partanen et al., 2013b, Partanen, Kujala, Tervaniemi et al., 2013), this might favour the processing of native phonological features over non-phonological ones. Therefore responses to frequency and intensity changes were grouped into the category of acoustic feature changes and word or phoneme changes were grouped into the category of word changes (see also Table 6)2 . Difference waveforms were calculated by subtracting the standard ERP (i.e. responses to unaltered syllables in learning phase) from the deviant ERP (i.e., responses to deviations) for each deviant type and each infant in each condition (nursery rhyme, song, prose). Using difference waveforms allows assessing the responses to deviants with minimal contribution of exogenous ERP components related to sound properties. Two time-windows (40–100 ms and 130–190 ms) were selected for statistical analyses based on earlier literature (especially Ylinen et al., 2017) and peaks revealed by visual inspection of grand average difference waves.","The presence of prediction error responses to changes in words and acoustic features in the selected time windows was examined by averaging together the mean amplitudes from F3, F4, C3, C4, Cz, P3 and P4. The averages were submitted to one sample t-test in order to assess whether the responses differed significantly from zero. The effect sizes for the t-tests were calculated using Cohen’s d values. The normality of the data was inspected visually from histograms depicting standardized residuals. The differences between different conditions were examined by submitting the mean amplitude to a 3 × 3 x 2 repeated-measures ANOVA with condition (nursery rhyme, song, prose), frontality (electrode lines F, C ja P), and laterality (electrode lines 3 ja 4) as factors. ANOVAs were run separately for each time window. When sphericity could not be assumed, Greenhouse-Geisser correction was used. Statistically significant (p < .05) main influences were analyzed in more detail with Bonferroni-corrected pair-wise comparisons. Cohen’s d values were calculated for each pair-wise comparison.","Statistically significant prediction error responses to word changes were found in newborns when deviations were presented within the nursery rhyme [40–100 ms: t(20) = −3.32, p = .01; 130–190 ms: t(20) = −2.35, p = .03] (Table 7). Word changes to deviations within song or prose did not elicit responses that differed significantly from zero, and neither did acoustic changes in any stimulus type (Figs. 4 and 5).","The current study examined whether the facilitative effect of music and rhythm on learning from auditory input can be observed in newborn infants. Infants were presented with three recorded versions of a nursery rhyme: one spoken metrically as a nursery rhyme, one sung to a melody, and one read as ordinary speech. These naturalistic conditions differed by their duration, intensity, pitch and rhythmic structure. We introduced word deviations and acoustical feature deviations in each version, and expected the changes to elicit prediction error responses, should the infants be able to extract predictions from the recently learned unaltered rendition. Only the word changes in nursery rhyme condition evoked significant prediction error responses. The responses to deviations in the nursery rhyme condition were significantly more negative than those in the song condition and were marginally more negative than those in the speech condition. Acoustic feature changes were not found to elicit prediction error responses in any condition. These results suggest that the rhythmic structure of nursery rhymes may facilitate learning from auditory input in newborn infants and may thus help future language development. Rhythm likely acts as a framework that allows the brain to form predictions of the future input, and aids temporal sequencing and segmentation (Huss et al., 2011; Kotz, Schwartze, & Schmidt-Kassow, 2009; Schön & Tillmann, 2015). The finding is supported by everyday-life experience: nursery rhymes are recited to and with babies and toddlers in many cultures. Perhaps nursery rhymes have always acted as naturally optimized educational auditory input to infants: their rhythm matches their linguistic prosody, which has been found to enhance phonological processing (Schön & Tillmann, 2015), extraction of phonological structure (Leong & Goswami, 2015) and lyric comprehension (Gordon, Magne, & Large, 2011). In addition, the auditory-visual synchrony present when an adult is reciting nursery rhymes promotes heightened attention and memory in infants (Bahrick & Lickliter, 2000; Lewkowicz & Lickliter, 2013; Lewkowicz, 2000). Previous research with adults suggests that both music and rhythm may help learning by creating templates for future events. Surprisingly, no such effect was found for the song condition in the current study, despite numerous studies that have found music to be beneficial for language learning and word segmentation (e.g. François et al., 2017; Kraus & Chandrasekaran, 2010; Putkinen, Tervaniemi, & Huotilainen, 2013; Putkinen, Saarikivi, & Tervaniemi, 2013; Tallal & Gaab, 2006). There are several possible accounts for this finding. First of all, it has been argued that melody may facilitate learning by attracting, maintaining and enhancing attention. For example, Nakata and Trehub (2004) hypothesized that short fixations associated with maternal speech might be linked to efficient learning and information transfer, while long fixations associated with maternal singing might be optimized to enhance interpersonal ties. Since the participants in the current study were asleep and thus not attending to the stimuli, the facilitating effect of the melody may have been weaker. Secondly, the lack of learning might have been due to a ceiling effect described by Thiessen & Saffran, 2004: single salient prosody cue might have the same effect as multiple cues. This possibility was speculated by Sambeth et al. (2008), who did not find a difference in infants’ responses when comparing singing and continuous speech. However, this would not explain why the nursery rhyme and the song conditions did not have similar effect. Thirdly, learning both the segmental (phonetic) and supra-segmental (melodic and/or rhythmic) patterns simultaneously from multi-dimensional signal such as a lyrical song might be more difficult for the newborn infants than learning only the segmental patterns. A previous study by Lebedeva and Kuhl (2010) found that melody facilitated phonetic recognition in 10-month-olds, but preliminary tests in 8-month-olds showed no difference between melodies or spoken strings. Music does contain additional information compared to nursery rhyme, so it is feasible that infants would use some of their cognitive resources to learn the melodic patterns when mapping of linguistic and musical information is not consistent, as in our song condition. Finally, the pattern of findings may suggest that stronger rhythmic structure of the nursery rhyme condition helps infants with probabilistic inference of words, and thus facilitates learning. Younger infants appear to segment speech stimuli based on probabilities and start to prefer prosodic cues only around eight months of age; possibly because they start to lose sensitivity to non-native contrasts (Johnson & Jusczyk, 2001; Thiessen & Saffran, 2007; Yeung & Werker, 2009). Our results support the previous findings from infants suggesting that exaggerated or variable pitch information within signal can facilitate phonetic recognition (Bergeson & Trehub, 2007; Lebedeva & Kuhl, 2010; Thiessen, Hill, & Saffran, 2005). As Tables 3 and 4 and Figs. 2 and 3 show, the current nursery rhyme stimuli, which led to the successful detection of word changes, had more pitch variation than the song or speech stimuli. It seems that pitch modulations might facilitate word segmentation as François et al. (2017) suggested, but here they were more pronounced in the nursery rhyme than the song. The difference in our and François’ and colleagues’ results might be due to us using a simple melody, that did not match well with word-stress pattern, whereas they had associated each syllable to a unique pitch. In line with François et al., we suggest that an exaggerated prosodic cue, such as pitch change, increases the saliency of to-be-learned item. What the present results add to those by François and colleagues is that the regular rhythmic patterns of the input, as introduced by the current nursery rhyme, seem to synchronize the auditory processing and to facilitate the prediction of the following items. This would also be in line with Goswami (2018) oscillatory temporal sampling framework. Perhaps the stronger frequency modulation and periodic power in the stressed syllable modulation tier in nursery rhyme stimuli lead to more efficient phase entrainment of neuronal oscillations in the auditory cortex, and thus more accurate temporal sampling of stressed syllables and more efficient learning. The song condition had higher periodic power in syllable modulation tier, which also suggests that while there was rhythmic regularity in the song stimulus, the stressed syllables were not as salient in the song condition as in the nursery rhyme condition. The present study is one of a few in which infants were presented with continuous speech. Thus, the natural stimuli are highly complex and acoustically not very well controlled. As follows, the current findings should be considered with some caution. When interpreting the results, it is good to keep in mind that infants’ brain responses are characterized by remarkable variability, the reasons of which are not yet fully understood. All conditions elicited positive responses in some infants, and negative responses in some other infants. The same phenomenon has been previously observed in traditional oddball paradigms (reviewed in Näätänen, Sussman, Salisbury, & Shafer, 2014) and also when infants have been exposed to the statistical auditory streams (Bosseler, Teinonen, Tervaniemi, & Huotilainen, 2016; Teinonen et al., 2009). In some conditions, these different response patterns may have cancelled each other out in averaging. Such variability might account for not observing significant responses to acoustic changes that typically elicit quite robust responses. In addition, the lack of significant responses may be caused by the use of continuous stimuli, the learning of which is expected to be more challenging than that of traditional oddball streams with simple acoustic deviations. Given the promising results, it would be valuable to replicate this study with a larger participant sample and continuous but more controlled stimuli. In the future, it would also be interesting to complement this study by exploring which factors in music and rhythm promote learning across different developmental stages. It might be that the effects of music and rhythms interact with effective attention allocation, and prediction errors would be associated with redirecting attention (Ylinen et al., 2017). Alternatively, language experience and developmental changes might be required before an infant can differentially attend to phonetic and melody information.","Our findings suggest that rhythmic structure of nursery rhymes may facilitate newborn infants’ learning from auditory input and may thus be beneficial to language development. Surprisingly, no such effect was found for a song in the present stimuli. This might be due to melody not including enough pitch variation, or simultaneous learning of the melody and phonetic content being more challenging to newborn infants than that of the phonetic content alone. In line with previous behavioral IDS studies (Curtin, Campbell, & Hufnagle, 2012; Spinelli, Fasolo, & Mesman, 2017), it seems that simultaneous occurrence of exaggerated prosodic cues and to-be-learned items helps infants to process language input. This could be taken into consideration for example when planning activities and interactions at home, music playschools, or when making educational music: linguistic and musical rhythm should maintain synchrony.","S.Y. and M.H. designed research; S.Y. and E.S. performed research, E.S. analyzed data; and E.S., M.H., and S.Y. wrote the paper."],["Saving energy at work might be considered altruistic, because often no personal benefits accrue. However, we consider the possibility that it can be a form of impure-altruism in that the individual experiences some rewards. We develop a scale to measure motivations to save energy at work and test its predictive power for energy-saving intentions and sustainable choices. In two studies (N = 293 and N = 94) motivations towards helping their organization and the planet were rated as important motivations, as was warm-glow (feeling good), indicating that impure-altruism does exist in this context. Energy saving was predicted by environmental concern and the desire to help one's organization. Notably, the stronger the motivations to promote one's reputation were, the weaker was the intention to save energy. Promoting motivations, particularly those that focus on benefits to the organization, may be an effective addition to environmental messages typically used as motivations in campaigns. --------------------------------------------------------------------------------","To help prevent damage to the environment due to climate change, the UN set a target to keep the earth's temperature rise to well below 2° Celsius above preindustrial levels within the Paris Agreement (UNFCCC, 2015). Within this, the EU has proposed to reduce its emissions to 40% below 1990 levels by 2030 (European Commission, 2012). Given pressures of climate change, energy security and affordability, there is an increasing interest across sectors in how to change current energy use. One key area of behavior change in this context is people's energy use behavior in non-domestic buildings (Janda, 2011; Schelly, Cross, Franzen, Hall, & Reeve, 2011; Schipper, Bartlett, & Hawk, 1989). It has been suggested that around 33% of greenhouse gas emissions in the UK and 17% in the US are released from shared buildings within the business sector (non-industrial) (DECC, 2011; United States Department of State, 2010). Current advances on reducing energy use in workplaces has mostly focused on improving physical infrastructure, appliances, system efficiency, or appointing key personnel with energy responsibilities (e.g., facilities managers, eco-champions) (Aragón-Correa, Matías-Reche, & Senise-Barrio, 2004; Christmann, 2000; Cordano & Frieze, 2000). There has been little investigation of how to encourage normal, individual workers (with no energy responsibilities) to change their own energy use behavior to reduce emissions. This gap in the literature is the focus of this paper. Since energy use is not an element of most employees' job assignments, and is usually not taken into account in performance evaluations, it might be argued that people simply will not care about, or act to save energy. The extent to which employees will try to reduce their energy use might depend on a number of motivations including if they see it as a key aim of their job (Rioux & Penner, 2001) or if they are motivated by more proactive prosocial behavior among employees, such as organizational citizenship behavior (Nisiforou, Poullis, & Charalambides, 2012; Schelly et al., 2011). The aim of the present research is to investigate what motivates employees to reduce their energy use at work when their job specifications do not include it. Indeed, energy saving can be considered an “extra-role” behavior (Ramus & Killmer, 2007) or an example of organizational citizenship behavior, as for the individual it is not normally directly or explicitly rewarded, but collectively is positive for the organization (LePine, Erez, & Johnson, 2002; Organ, 1997). Promotion of energy saving among employees ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Existing research on environmental behavior in the workplace shows that employees can be encouraged to adopt energy saving behaviors (Lo, Peters, & Kok, 2012). Using the Theory of Planned Behavior (TPB, see Ajzen, 1991), Greaves et al. (Greaves, Zibarras, & Stride, 2013) found that, in general, employees intended to turn off their computers when they left their desk for 1 h or more, and particularly if they think switching things off is a good thing (notably for the environment), and if the social norms of the workplace fit this behavior (see also Zhang, Wang, & Zhou, 2014). Goal setting has also proven an effective intervention (McCalley & Midden, 2002), as well as the use of rewards, (Handgraaf, Van Lidth de Jeude, & Appelt, 2013), individual feedback (Murtagh et al., 2013), group discussions (Werner, Cook, Colby, & Lim, 2012), and group feedback and peer education (Carrico & Riemer, 2011). We highlight that, to date, the purpose of energy saving in the workplace, that is for what or for whom employees would save energy, has not been studied as a precursor of energy saving intentions. Importantly, research indicates that to change (environmental) behavior, via any intervention or a communication, the goal [of the behavior] promoted by the intervention must be activated (Unsworth, Dmitrieva, & Adriasola, 2013). This gap in the existing literature indicates that previous research and applied interventions may therefore have miscommunicated energy savings in ignoring the reasons that employees may have for saving energy. We adopt a functional approach as we are interested in identifying the goal(s) which energy saving behavior helps to fulfill (Snyder, 1993). We want to investigate whether saving energy in the workplace can have multiple functions. Indeed people could, for example, have the goal to help their organization, and saving energy could have the function to help attain this goal. Other goals could include feeling good about themselves (warm-glow), gain reputation as a good person or just because no-one else does (reluctant altruism; Ferguson, 2015). In this, saving energy in the workplace could have multiple functions for different people or even multiple functions for the same person. We want to look at various potential motivations to save energy and investigate their importance, as these could be drivers, to different degrees, to adopt energy saving behaviors in the workplace. Motivations to save energy in the workplace ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Many studies focusing on interventions to reduce energy in the workplace do not specify the reason or goal they used to encourage people to reduce their energy use (Carrico & Riemer, 2011; Handgraaf et al., 2013; Staats, Leeuwen, & Wit, 2000) and indeed a lack of motivation has been highlighted in some instances (Murtagh et al., 2013). Whilst motivations to save energy often differ between individuals/user groups, determining commonalities would help to highlight the most effective ways to frame energy saving campaigns in different contexts. For example, cost is often a key motivation for users to save energy in residential contexts (Brandon & Lewis, 1999; Spence, Demski, Butler, Parkhill, & Pidgeon, 2017), and is often used to encourage people to save energy in behavioral interventions (Abrahamse, Steg, Vlek, & Rothengatter, 2005, 2007; Midden, Meter, Weenig, & Zieverink, 1983). However, research suggests that cost savings are often so low at the individual level that consumers may not consider behavior change worthwhile (Spence, Leygue, Bedwell, & O'Malley, 2014). In the workplace, for most workers or employees, saving electricity does not mean saving costs for oneself, as is the case for domestic use. Hence referring to energy in terms of costs might have a weaker impact on motivations in the workplace and other aspects of energy use may have a broader impact. On the other hand, cost saving potential as a collective may be much greater and motivating. In fact this technique of aggregating energy savings at the group level has been used successfully before, where university staff and students were told how much energy and costs in total would be saved if all the classrooms' lights were turned off every day. Though, we note this was not compared to other methods of calculating savings, e.g. in terms of carbon or non-aggregated (Werner et al., 2012). There is currently little evidence about whether, and how, motivations to save energy in the workplace context do differ from a residential context. Given the lack of cost incentive in the workplace, current research and interventions aiming at reducing the energy use of employees has mostly focused on the benefits of this behavior for the environment (Scherbaum, Popovich, & Finlinson, 2008; Unsworth et al., 2013) which may not capture the whole spectrum of motivations involved. One reason for this is that sometimes energy saving behavior is studied as one of several environmental behaviors (Bamberg & Möser, 2007; Lo et al., 2012). However, we propose that reducing one's energy use in the workplace could serve other functions and fulfill different goals other than environmental concern. Most people acknowledge the problems associated with climate change (Spence, Venables, Pidgeon, Poortinga, & Demski, 2010), but only a smaller proportion tend to feel they must or can do something to reduce it (Spence et al., 2010; Whitmarsh, Seyfang, & O'Neill, 2011). Targeting goals other than concern for the environment may therefore be useful in engaging every employee with saving energy in the workplace. We identify a number of theoretically relevant domains of motivation below. Pure, impure, reluctant altruistic, and selfish motivations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ At the organizational level, “corporate greening” has already been conceptualized as a pro-social behavior (Ramus & Killmer, 2007). At the individual level, given that most workplaces do not currently recognize or reward their employees for adopting energy saving behaviors, motivations to save energy at work may be mostly considered, at least in part, other-oriented or altruistic (Ramus & Killmer, 2007). Altruism is defined as a desire to maximize the welfare of others (e.g., by reducing their suffering) at a personal cost, without personal benefit (Andreoni, 1990; Ferguson & Lawrence, 2016). Indeed, saving energy for environmental reasons (i.e., to reduce carbon emissions) can be considered as an altruistic act, as the benefits will mostly affect others (e.g., the planet, future generations), while it will be costly to the individual (time, effort) (Sober & Wilson, 1998). Saving energy in the workplace could also be considered an altruistic act towards one's company. Employees might want to help their company reduce its energy costs by reducing their own energy use (Werner et al., 2012). In addition, helping their company reduce their energy use and its impact on the environment might help it increase or obtain a positive public image. Indeed, a survey of 8000 consumers in the United States revealed that 80% of high education/high income people would change brand if a company was negatively portrayed by the media on their social responsibility, and sustainability is now an important factor within corporate social responsibility (ACCA, 2004; DEFRA, 2006). This of course, should depend on the extent to which employees feel positive towards their company and their job, so should be affected by the organization's culture, and employees' commitment and identification towards the organization (Allen & Meyer, 1990; Mael & Ashforth, 1992). However, if a company is benefited through reduced costs or improved public image, an employee could feasibly indirectly benefit through the improvement in company status, e.g. with increased job security, or a potential increase in opportunities, so motivations here may not be purely altruistic. Finally, by helping their company's (green) image the employee will also be able to indirectly enhance their self- image as one working for a ‘green’ company. So by improving the organization's image the employee enhances their own positive image. This is an aspect of impure altruism as the ‘helper’ helps the target and might indirectly benefit themselves as well (Andreoni, 1990; Griskevicius, Tybur, & Van den Bergh, 2010). Research in pro-social behavior more generally, has shown that pro-social actions can also often be personally beneficial (Clary et al., 1998; Ferguson, Farrell, & Lawrence, 2008). Indeed, behaviors often considered as purely altruistic, such as donating blood, also often result in people feeling good about themselves for doing the right thing (Ferguson, 2015; Ferguson, Atsma, De Kort, & Veldhuizen, 2012), an effect called warm-glow resulting in impure altruism as both the donor and recipient gain (Andreoni, 1990). As well as being experienced at the time of donation, it is possible that warm-glow can be an ‘anticipated’ motivation for action (Ferguson & Masser, 2017). Furthermore, in the specific context of the workplace, other self-oriented motivations can exist for ‘altruistic’ acts, again leading to impure altruistic motivations. For example, doing the right thing at work can be a way of obtaining approval from management and colleagues, as it could be viewed as a “taking charge” behavior (Ramus & Killmer, 2007). Indeed, impression management is a strong predictor of positive behaviors in the workplace that go beyond one's job description (Rioux & Penner, 2001). Finally, altruistic acts can be performed when people think the cause is worthy and think others are unwilling to act (i.e., others are free riders). This form of reluctant altruism was first identified and described by Ferguson and colleagues, in the context of blood donations for first-time donors in the face of free-riding (Ferguson et al., 2012). Lack of trust in others to help, moral outrage and negative view of humanity are all facets of reluctant altruism (Ferguson & Masser, 2017; Ferguson, 2015) however in the context of saving energy at work lack of trust was the key parameter we focussed on. People then feel the energy to act because of negative feelings and lack of trust towards others who do not make any effort. So reluctant altruism-lack of trust could be an important motivator for prosocial behavior when free riding is high, as is likely to be the case in energy conservation in the workplace. Current research ~~~~~~~~~~~~~~~~ We propose that saving energy in the workplace can have multiple functions or utilities, and so is a benevolent act, or an example of impure altruism, where both other-oriented and self-oriented motivations exist. Specifically, we want to investigate whether employees can be motivated to save energy by two forms of altruism or other-oriented motivations: environmental concern and helping their organization without benefit to the self (direct or indirect), and two forms of self-oriented motivations: warm glow and reputation building. We are also interested in consistency between contexts, and whether energy saving functions at work could translate to goals at home, and so if we could find similar motivations to save energy at home. If this is the case, motivations at work and at home should be related. Furthermore, we are interested in whether these motivations will have independent effects on energy saving behaviors. Also, we postulate that the extent to which people are motivated to save energy at work and intend to do so, should depend on their general attitude towards their workplace; the more they evaluate their job and their organization positively, the more they should be motivated to reduce their energy use to help their organization, and the more they should try to adopt sustainable behaviors. We propose that motivations identified may mediate the relationship between attitude towards the workplace and energy saving intentions. Finally, we expect that these motivations should affect how people make sustainable choices in the workplace. In two studies, we measured motivations to save energy at work within five different organizations (N = 298 for Study 1 and N = 94 for Study 2): one private company (N = 75) and one academic institution (N = 218, 5 participants left this question blank) in Study 1; and two private companies and one non-government organization (NGO) in Study 2. We developed a scale of Motivations to save Energy in the Workplace (MEW) based on existing scales of motivations to volunteer, to donate blood, and to engage in citizen organizational behavior as well as theoretical considerations of self- and other- orientated motivations. There are important similarities between each of these behaviors and energy saving behavior at work: there is no direct benefit to the self, the beneficiaries are not people that are close to the actor, there is no obligation for the actor to perform these behaviors (Ferguson, 2015), and these behaviors need to be prolonged or repeated to be most effective at a communal level. We tested the predictive validity of the scale by exploring the relations between motivations identified for saving energy, reported energy saving intentions in the workplace, attitude towards the workplace, and motivations to save energy at home. In a second study among a different sample, we explored the effects of these motivations on a more comprehensive measure of intentions to save energy, and direct sustainable choices. Scale development ~~~~~~~~~~~~~~~~~ Based on the literature on altruistic behaviors such as blood donation, volunteering and citizen behavior (see above), we created an initial set of 28 items for the scale1. We distributed this scale in two business sites to measure the importance of each motivation, and the factor structure and validity of the scale were examined, as well as basic psychometric properties. To explore the antecedents and consequences of these energy saving motivations in the workplace, we included measures of general attitudes towards the workplace, and a scale measuring the self-reported frequency of environmental behaviors at work. In addition, we were interested in knowing whether motivations to save energy in the workplace would help predict sustainable choices. To test this, we conducted a second study to look at energy saving motivations, and we examined whether they would predict how people would spend money from and for their organization. We hypothesized the more people would be motivated to save energy at work, the more money they would choose to spend on sustainable products for their company.","Participants To recruit participants, we contacted “gatekeepers” to ask for permission to advertise the study and use internal mailing lists to recruit volunteers to participate in a study on energy use at work in two sites: a large private company and a university in the United Kingdom. Both companies had organized campaigns on sustainability in the past (e.g., cycle to work) but no campaign on energy was taking place at the time of the study. The large private company was represented by one gatekeeper, at the university there were multiple gatekeepers as all academic and administrative departments were contacted. Exact response rate cannot be evaluated here as it is unclear how many gatekeepers agreed to advertise the study and how (e.g., send emails to a mailing list or advertise on the intranet “message of the day”). Both the private company and the university total together approximately 10,000 employees. Questionnaires were filled in online using Qualtrics Research Suite©2 (Qualtrics, 2015). The incentive to take part was entrance into a prize draw to win one of five £10 (Sterling) (15.26 USD) shopping vouchers. Two hundred and ninety eight participants (198 women and 100 men) took part in the study. Their ages ranged from 18 to 65, (M = 39.64, SD = 10.8). Thirty-five percent of this sample had managerial responsibilities at the time of the study. Just under half of the participants (47.7%) had a postgraduate qualification, 29.2% had a degree or equivalent, 15.5% a high school qualification. This sample consisted of more women and was more highly educated than the general population in the UK when compared to the most recently available data obtained from the Labor Force Survey in 2006. Participants had been working in their current organization for an average of 7.74 years. Motivations to save energy We constructed 28 items to measures people's motivations to save energy at work. Items were statements adapted from scales of motivations to donate blood (Evans & Ferguson, 2014; Ferguson et al., 2008), motivations to volunteer (Clary et al., 1998; Snyder, 1993), and motivations to adopt organizational citizenship behavior (Rioux & Penner, 2001). Participants had to rate each statement about why they save energy such as “to help my organization save money on energy costs” or “because I'd feel proud of myself” on a 7-point scale according to how important each would be in their decision to save energy in the workplace. We constructed 25 items to measure people's motivations to save energy at home following the same process as for motivations to save energy in the workplace, e.g., items included, “to save money on energy costs” and “because I feel worried about the environment”. Isomorphism between the items at work and at home was ensured by using the same items when possible or changing only one word (e.g., replacing “my colleagues” by “my friends”). Three items were specific to the workplace and the organization's image and were not adapted to the home context. Behavioral intentions at work Environmental behavior intentions at work were measured by 14 items adapted from Whitmarsh and O'Neill (2010), see Appendix 1. Participants had to rate each behavior on how likely they would consider performing them on a scale from 1-very unlikely to 4-very likely, with a non-applicable option also provided. Cronbach's α = 0.73 for environmental intentions in general, and Cronbach's α = 0.65 for energy saving intentions, after eliminating one item: “use an extra electric heater” which did not cohere with other behaviors included. Reliability indexes were lower than expected here (for example, Whitmarsh and O'Neill obtained a Cronbach's α of 0.92). This could be due to the fact that less items were used here (as the context of a workplace offers less examples of sustainable behaviors than the home). Also, people tend to have less control over their environmental behaviors and their energy use in the workplace compared to the home, and this could result in more variability between which behaviors can be adopted or not and how frequently. General attitude towards the workplace We measured participants' general attitude towards their workplace by asking them about their job satisfaction (Hackman & Oldham, 1975), their commitment to the organization (Allen & Meyer, 1990), and their organizational identification (Mael & Ashforth, 1992). Job satisfaction was measured by 3 items (e.g., Generally speaking, I am very satisfied with this job) (Cronbach's α = 0.88 if one item is removed), commitment to the organization by 5 items (e.g., this organization has a great deal of personal meaning for me) (Cronbach's α = 0.83), and organizational identification by 6 items (e.g., When I talk about this organization, I usually say “we” rather than “they”) (Cronbach's α = 0.91). Participants had to rate each item on a 7 point scale from 1-strongly disagree to 7-strongly agree. Motivations to save energy at work Given that the present study is the first investigation of multiple motivations to save energy at work, we conducted an exploratory factor analysis on participants’ responses to the scale using Mplus 6 statistical software (Muthen & Muthen, 1998–2010). We used an MLR estimation3 as the scale contained 7 points. Results show that the best model is a structure of 6 factors for motivations to save energy at work, χ2 = 410.99, df = 204, p = 0.0001, CFI = 0.941 and TLI = 0.899, RMSEA = 0.058. These values of CFI and RMSEA are close respectively to the values of 0.95 and 0.06 that are recommended by Hu and Bentler (Hu & Bentler, 1999). Comparative indices for a structure with 5 factors were not as good, χ2 = 539.53, df = 226, p = 0.0001, CFI = 0.911 and TLI = 0.862, RMSEA = 0.068 and the Satorra- Bentler scaled chi-square difference test reveals that the difference between the two models is significant, TRd = 127.04, df = 22, p < 0.001 with the 6 factor a better fit (Satorra & Bentler, 2001). A structure with 7 factors did not improve the indices, χ2 = 427.74, df = 183, p = 0.0001, CFI = 0.931 and TLI = 0.867, RMSEA = 0.067 though the Satorra-Bentler scaled chi-square difference test reveals that the difference between the two models is not significant, TRd = 29.57, df = 21, p = 0.10. Environmental concern, warm glow and reputation building at work were found to be distinct motivations to save energy in the workplace, see Table 1. In addition to this, two further factors were identified. A factor that represents reluctant altruism was obtained, and altruism towards the company separated into two distinct factors: helping one's organization's finances and helping one's organization's image. We note that the factor concerning helping one's organization's finances contains only two items and that this is often considered as problematic. However, in this case the items are homogeneous and have face validity. Therefore the issues usually encountered with a low number of indicators should not be encountered (Gardner, Cummings, Dunham, & Pierce, 1998; Little, Lindenberger, & Nesselroade, 1999; Wanous, Reichers, & Hudy, 1997). The average correlation across factors identified was 0.33. Mean levels of reported importance were calculated for each factor and a repeated measures ANOVA, F(5, 285) = 171.85, p = 0.0001, η2 = 0.37, and post-hoc tests (Bonferroni adjusted) indicated that environmental concern (M = 4.89, SD = 1.54), helping one's organization's finances (M = 4.74, SD = 1.63), and warm-glow (M = 4.55, SD = 1.47), were considered the most important motivations to save energy in the workplace followed by helping one's organization's image (M = 4.04, SD = 1.62) and reluctant altruism-lack of trust (M = 3.72, SD = 1.53). Reputation building at work was not considered an important motivation (M = 2.3, SD = 0.99). If we look at both sites separately, the pattern is approximately the same, with participants in the private company scoring higher on motivations in general than participants at the university, F(1, 288) = 6.57, p = 0.01, η2 = 0.02, for the main effect of the type of site. We observe an interaction effect between site and type of motivation, F(5, 284) = 2.84, p = 0.01, η2 = 0.01, and post-hoc tests (Bonferroni adjusted) reveal that helping one's organization's image was rated as a more important factor in saving energy in the private company (M = 4.11, SD = 1.69) than in the university (M = 3.5, SD = 1.54), t(293) = −2.97, p = 0.003; the other types of motivations did not differ significantly between sites, see Fig. 1. Motivations to save energy at home Responses to the motivations to save energy at home scale, were examined in the same way as in the workplace. Results show that the best model is a structure of 5 factors, χ2 = 373.8, df = 205, p = 0.0001, CFI = 0.95, TLI = 0.92, and RMSEA = 0.053. The indices for a structure with 4 factors were not as good, χ2 = 536.97, df = 227, p = 0.0001, CFI = 0.91, TLI = 0.87, and RMSEA = 0.069. The Satorra- Bentler scaled chi-square difference test reveals that the difference between the two models is significant, TRd = 185.41, df = 22, p < 0.001, with the 5 factor a better fit. A structure with 6 factors slightly improved one index, χ2 = 338.24, df = 184, p < 0.001, CFI = 0.95 and TLI = 0.92, RMSEA = 0.054 and the Satorra- Bentler scaled chi-square difference test reveals that the difference between the two models is significant, TRd = 37.25, df = 21, p = 0.02. However the 6th factor did not contain any clear and unique items, with rotated loadings all inferior to 0.40 or loading equally on 2 factors. Environmental concern, warm glow and reputation building at home emerged as distinct factors, as they did in the workplace context. In addition a factor that represents costs and a factor that represents feelings of reluctant altruism-lack of trust were also obtained, see Table 2. The average interscale correlation was 0.36. Mean scores calculated for the motivations obtained indicated that participants rated costs as the most important motivation for saving energy at home (M = 6.33, SD = 1.04), followed by environmental concern (M = 4.91, SD = 1.7), and warm glow (M = 4.42, SD = 1.61). Again participants disagreed that feelings of reluctant altruism (M = 3.73, SD = 1.63) and reputation building (M = 2.32, SD = 1.14), were important motivations. The factor structure and the psychometric properties of the 6 subscales were invariant across gender and work status. Motivations at work and at home were consistent, see Table 3. We observe high correlations between both contexts regarding environmental concern motivations, r = 0.91, warm-glow motivations, r = 0.80, reputation building, r = 0.76, and, to a lesser extent, reluctant altruism, r = 0.60. Relationship between motivations, intentions to save energy at work, and attitude towards the workplace To validate the utility of the motivations we identify, we examined the extent to which each motivation would predict intentions to adopt environmental behaviors in the workplace. To investigate this, we ran a regression analysis (OLS), using the 6 types of motivations to save energy at work as predictors of environmental intentions at work. Results show that environmental behavior intention is predicted by environmental concern (B = 0.10, p = 0.001), helping one's organization's image (B = 0.09, p = 0.001), and reputation building (B = −0.08, p = 0.006). The whole model predicts 22% of the variance in environmental intentions at work (adjusted R2 = 0.22), see Table 4. Interestingly, the more people rate environmental concern and their organization's image as important motivations to save energy, the more they intend to adopt environmental behavior at work. However, the more they indicate reputation building at work as an important motivation, the less they intend to adopt environmental behavior. We were also interested in looking at energy-related behavior intentions in particular, for example, at work, turn off lights you're not using. To do this, we conducted a regression analysis using the mean score for energy related behavior intentions as a dependent variable and the 6 types of motivations as independent variables. Results reveal the same pattern; the more environmental concern (B = 0.09, p = 0.001), and helping one's organization's image (B = 0.09, p = 0.001) are important motivations, the more people intend to adopt energy saving behaviors at work. Conversely, the more people find reputation building motivations important, the less they intend to adopt energy saving behaviors (B = −0.13, p = 0.001). The whole model explains 14% of variance in energy-related behavior intention, see Table 4. Bivariate correlation analyses showed that motivations related to helping one's organization's image and finances are correlated with the three indices of attitudes towards the workplace. Furthermore, environmental intentions and in particular energy saving intentions correlated with commitment to the organization and organizational identification, but not with job satisfaction, see Table 5. Relationship between attitude towards the workplace and energy saving intentions and the mediating role of motivations to help organizational image The mechanisms underlying the relationships between contextual variables at work and energy saving intentions are of particular interest as these could inform potential interventions. Hence, we conducted mediation analyses to investigate whether the motivations to help one's organization's image would mediate the effects of commitment to the organization and organizational identification on energy behavior intentions. To do this, we used the MEDIATE method of Hayes and Preacher (2014) and due to the strict assumption of normally distributed data within the product-of-coefficients approach to mediation, we used bootstrapping to resample the data 10,000 times in estimating the indirect effects.4 Results show that organizational identification, B = 0.32, t = 4.05, p = 0.001, and commitment to the organization, B = 0.37, t = 4.35, p = 0.001, have a positive effect on motivation to save energy to help your organization's image. Motivation to save energy to help your organization's image predicts energy saving intentions, B = 0.09, t = 3.59, p = 0.001, and the indirect effects of commitment (B = 0.035, LLCI = 0.01, ULCI = 0.06) and identification (B = 0.03, LLCI = 0.01, ULCI = 0.05) on energy saving intentions are significant, see Fig. 2. Direct effects of organizational identification and commitment for the organization on energy behavior intention were both non-significant (B = 0.02, t = 0.66, p = 0.51, and B = 0.04, t = 1.09, p = 0.27, respectively). This pattern of results indicates that the effects of commitment towards the organization and organizational identification indirectly impact intentions to save energy in the workplace through motivations to help your organization's image. The results of Study 1 reveal that intentions, both of energy saving and of further environmental behavior, are mostly predicted by altruistic and impure altruistic motives such as environmental concern and wanting to help one's organization. Interestingly, self- oriented motives such as wanting to improve one's image at work (reputation building) negatively predicted intentions to save energy. However, Study 1 focused on self-reports of behavior intentions, so the question of whether these motivations can predict actual behaviors remain. Furthermore, measuring sustainable behaviors directly would reduce the reverse causality possibility that actually measures of intentions predict motivations. So the aim of Study 2 was to look at the effects of motivations to save energy at work on actual sustainable behaviors, measured online through a choice task. Furthermore, we were interested in examining the extent to which the results of Study 1 would replicate when using a more comprehensive measure of energy behavior intentions at work, that is, a measure that would also include social behaviors, such as discussion of issues with people in charge, or confrontation of people who waste energy. Participants A total of 94 employees (66 men, 27 women, one participant preferred not to state their gender) took part in Study 2. An opportunity sample was drawn from three small to medium sized companies, with an average response rate of 37.12%. The three companies were two private companies, and an NGO. Age ranged from 18 to 62, with a mean of 31.57, and a SD of 10.66. Fifty percent of this sample had managerial responsibilities at the time of the study. Before filling in the questionnaire, participants were randomly assigned to one of 3 scenarios, where they had to imagine that they were working in a company that had decided to reduce its energy use. Scenarios were accompanied by 1 of 3 displays, one showing energy use in terms of CO2 emissions, one in terms of costs, and the last was a combined CO2 and costs display. This resulted in 32 seeing the CO2 meter, 30 seeing the cost meter and the remaining 31 seeing both cost and CO2 readings on the same meter. These were used for another study and did not affect the motivations to save energy scores, so will not be discussed in the present study. Participants filled in a scale to measure their motivations to reduce their energy use at work as in Study 1. In addition, they completed measures of intentions to adopt energy saving behaviours at work and a measure of sustainable choices5. Motivations to save energy The same scale of motivations to save energy at work was used as in Study 1, except for the items that had not been included in the final factors for the analyses (see Table 1). So the new scale comprised of 24 items that participants had to rate on a 7-point scale. Energy saving intentions Participants rated each of 15 energy saving behaviors (Cronbach's α = 0.89) on how likely they would consider performing them on a scale from 1-very unlikely to 6-very likely, with a non-applicable option also provided. These included both individual behaviors e.g., “turn off communal office equipment (e.g., printer, copy machine, lab equipment) after using them”, and social behaviors “remind a colleague to switch something off to save energy”, see Appendix 2. Sustainable choices The study comprised an indirect measure of sustainable behavior: participants were asked in a fictitious scenario to distribute 100,000 sterling pounds into five purchases for their company. To do this, they were given a list of 20 possibilities, of which two were choices that would help their company save energy, e.g., ensure building operates at zero-carbon emissions by using Microgeneration (E.g. Solar Panels). Motivations to save energy at work We computed six mean scores for each of the six motivation types as in Study 1, reputation building (Cronbach's α = 0.83), environmental concern (Cronbach's α = 0.82), helping one's organization's finances (Cronbach's α = 0.88), helping one's organization's image (Cronbach's α = 0.86), warm glow (Cronbach's α = 0.85), and reluctant altruism (Cronbach's α = 0.65). The average correlation across factors identified was 0.48. A repeated measures ANOVA, F(5, 88) = 51.58, p = 0.0001, and post-hoc tests (Bonferroni) indicated that warm-glow (M = 4.43, SD = 1.36), environmental concern (M = 4.33, SD = 1.45), helping one's organization's image (M = 4.25, SD = 1.53), and helping one's organization's finances (M = 4.16, SD = 1.64) were considered the important motivations to save energy in the workplace, whereas reluctant altruism (M = 3.46, SD = 1.30) and reputation building at work (M = 2.5, SD = 1.08) were not considered important. This pattern was not affected by the company the participant was from, F(2, 90) = 2.46, p = 0.09. Prediction of energy saving intentions To examine whether the results of Study 1 would replicate, we ran a regression analysis (OLS), using the 6 motivations to save energy at work as predictors of energy saving intentions at work. Results show that energy saving intention is predicted by environmental concern (B = 0.25, p = 0.001), helping one's organization's finances (B = 0.30, p = 0.001), and reputation building (B = −0.35, p = 0.001). The whole model predicts 49% of the variance in energy saving behavioral intentions at work (adjusted R2 = 0.49). The more people rate environmental concern and their organization's finances as important motivations to save energy, the more they intend to adopt environmental behavior at work. As in Study 1, the more they indicate reputation building at work as an important motivation, the less they intend to adopt environmental behavior. Sustainable choices We summed the total amount of money that each participant distributed to energy saving purchases for their company, then log-transformed the score. The regression analysis using this score as a dependent variable and the 6 motivations as predictors reveals the amount of money spent on sustainable options is predicted by environmental concern (B = 0.08, p = 0.02) and helping one's organization's finances (B = 0.10, p = 0.007). Reputation building only had a marginal effect, (B = −0.09, p = 0.05). Interestingly, again reputation building had a negative impact on sustainable choices.","Our results provide a first exploration of motivations to save energy in the workplace and indicate that motivations to save energy at work are different from those at home, necessitating measures and interventions designed specifically for the workplace. Indeed, in the home, where electricity has direct costs to the individual, saving costs was the most important motivation. At work, helping one's company becomes an important motivation. Our research introduces a new measure of motivations to save energy in the workplace consisting of six factors: environmental concern, helping one's organization's finances, warm-glow, helping one's organization's image, reluctant altruism, and reputation building at work. However, motivations to save energy at work (MEW) indicated by our measure also correlated with motivations to save energy at home, showing consistency in people's rationales to save energy in different contexts. In two studies, the MEW scale showed predictive validity in its relationships with energy behavior intentions in the workplace and sustainable choices. Importantly we highlight that a range of motivations for saving energy exist in the workplace, beyond environmental imperatives, and these should be considered and taken into account when designing interventions for saving energy to ensure that maximum engagement with employees is achieved. Saving energy as impure altruism ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This research is the first, to our knowledge, to explore motivations of employees to save energy in the workplace. The cross sectional study among two large companies (public and private) obtained comparable results in both sites. Furthermore, a second study in three smaller sites obtained similar outcomes. Our results are in line with the literature on other types of pro-social behaviors and show that several types of motivations to save energy at work exist. These can be classified as self-directed: reputation building and warm-glow, or altruistic: towards the planet (environmental motivations), towards the organization (motivations about the company's image and finances), and finally reluctant altruism-lack of trust (Evans & Ferguson, 2014; Ferguson, 2015). Our data indicates that people are mostly motivated to save energy because of environmental reasons. In addition, behavior was often motivated by helping their company, and in order to gain warm-glow feelings. Energy savings may, therefore, often have aspects of impure altruism where others directly benefit but the individuals themselves also benefit from their actions. This supports previous research on pro-social behaviors, such as for example volunteering and blood donation, that revealed both altruistic and self-oriented drivers for such behaviors (Clary et al., 1998; Ferguson et al., 2008; Snyder, 1993). Interestingly, within other prosocial behaviors, factor structures of motivations were found to be different to that found here. Indeed, our reputation building factor included aspects related to evaluation by management similar to the “career” dimension of volunteering together with aspects related to norms or the “social” dimension of volunteering (Snyder, 1993). However aspects related to evaluation by management and social norms also seemed to be represented by one single “self-management” dimension within explorations of motivations to donate blood (Ferguson et al., 2008). In addition, our motivation towards the company factor included aspects close to identification with the organization as were found in the motivations to adopt organizational citizenship behavior (Rioux & Penner, 2001), as well as aspects close to the kinship dimension of blood donation. It could be the case that the current measure mixes levels of antecedents or motivations of energy saving intentions, and further research should try to draw out the causal sequence in the development of these motivations. Study 2 did not allow for a confirmatory factor analysis (given the low sample size), however the reliability analyses seemed to indicate that the 6-factor structure fitted the data of study 2 as well. Future research should aim at testing this 6-factor structure in large samples. Motivations to save energy in the workplace ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The different motivational factors to save energy at work identified here correlated relatively highly with each other, indicating that promoting one or fulfilling a specific motivation should not weaken another one. Also, self-oriented motivations did correlate with more altruistic ones, e.g. reputation building motivations correlated with helping your organization's image. Indeed, it is logical that a positive image of your organization could reflect positively on yourself, especially if people identify highly with their organization. Warm-glow motivations also correlated highly with environmental concern motivations, both in the workplace and at home. This is in line with recent work which considers environmental concern as both self-interested and pro-social (Bamberg & Möser, 2007) and work on self-serving goals and biospheric goals in saving energy (Bolderdijk, Steg, Geller, Lehman, & Postmes, 2013). Bolderdjik and colleagues explain that people want to maintain a “positive self-concept” and prefer presenting themselves and/or seeing themselves as “green” rather than “greedy”, implying that environmental concern could also have a self-oriented function. In our present study, warm-glow also correlated highly with motivations to help one's organization. Predictive validity of motivations to save energy in the workplace ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Altruistic motivations were the most important in our sample and were the best predictor of intentions to save energy at work across Study 1 and 2, and more broadly to adopt environmental behaviors in the workplace in Study 1. Whilst self-serving motivations were also important, reputation building motivations specifically were considered less important and actually negatively predicted energy saving intentions in both studies, using various measures of behavior intentions and actual behaviors. This might mean that people might not consider energy saving as something that would be recognized as a positive workplace behavior at the present time, or that it would not encourage them to do so even if it was; perhaps because it detracts, or is perceived to distract, from the main focus of their work. Indeed, this tendency of considering that saving energy in the workplace would not be done for reputation building or impression management, has been previously found for other organizational citizenship behaviors (Bolino, 1999; Dalal, 2005). Alternatively, people might think that saving electricity at work would hurt their reputation, as “do-gooders” are sometimes evaluated negatively by others. Indeed, people are then reminded that they are not doing the right thing themselves, and fear reproach (Minson & Monin, 2012; Rothgerber, 2014). Future research should explore potential explanations and variations of reputation building motivations between contexts. Whilst warm-glow was rated as an important motivation for people to save energy at work, it did not demonstrate predictive power for behavior intentions. This might be due to high correlations with the two main predictors of intentions: environmental concern and helping one's organization's image. Interestingly, reputation building motivations and warm glow motivations are both self-orientated but differ in the notion that one could be experienced as intrinsic (warm-glow), the other extrinsic (reputation building), hence demand characteristics might affect these in a different way. A comparable pattern of results was found by Zhang and colleagues (Zhang et al., 2014) on the prediction of attitudes towards electricity saving among office workers in China. They explored attitudinal beliefs rather than motivations to save energy and obtained a similar pattern of environmental benefits, organizational benefits and enjoyment (due to energy saving) having a positive effect on attitude, with anticipated extrinsic benefits to the self having little effect. It would be interesting to explore whether the negative effects of reputation building promotion on behavioral intentions are also due to the extrinsic focus of this motivation. Both results show that neither the belief that you will get a reward, nor the goal to save energy to obtain a reward (both extrinsic focused), seem to encourage energy saving behaviors in the workplace, whereas intrinsic focused motivations (environmental aims and warm glow/enjoyment) appear to be more successful. Notably, motivations to save energy in the workplace measured by our scale predict sustainable choices. We found that the more people found warm-glow motivations important, the more they wanted to participate in a second study on sustainability. We also found some preliminary evidence that environmental concern and motivations to enhance the company's image predict sustainable choices in financial decisions in the workplace. Organizational factors relating to motivations for energy saving in the workplace ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our data indicated that the more people are committed to their company, and the more they identify with it, the more motivations to help their company's image are important for them and the more, in turn, they intend to save energy at work. Energy saving in the workplace can hence be considered as an organizational citizenship behavior, as it is not normally specifically required, demands some extra effort, and correlates with commitment and identification. Interestingly, commitment and identification here had a positive effect on energy saving intentions when looked at individually (see similar findings from Feather & Rauter, 2004), but this effect disappeared when tested together. Their effect on behavior intentions is likely to be due to their common variance that could stem from general positive affect in, or positive attitude towards, the workplace (Bissing-Olson, Iyer, Fielding, & Zacher, 2013). These results show that energy savings in organizations can be promoted and increased through schemes focused on commitment and identification in the workplace as well as those targeted at more general well-being in the workplace. We highlight that motivations to save energy to help one's organization were only partially predicted by identification and commitment, so other organizational aspects that can affect this motivation should be explored in further research. Notably and in contrast, environmental motivations were only partly related to perceived contextual factors in the workplace (low correlations with organizational identification and no significant relationship with measures of commitment and satisfaction), highlighting that this is a universal factor that is not dependent on workplace context. Limitations ~~~~~~~~~~~ Our study used self-reports of behavior preferences and intentions rather than actual energy use, and these are not always well aligned with actual behavior. In particular, energy use is difficult to monitor and to control for individuals, so intentions to reduce energy may diverge somewhat from actual energy use. In Study 2, sustainable choices were used to look at consistency, but this measure could also be considered as cooperative behavior towards the organization. However, Study 2 provided data on actual sustainable choices that replicated the model found in Study 1, indicating that the same model can be found when not using self-reports of behavior intentions. Self-reports of the importance of the motivations may also be subject to social desirability. Notably, reputation building motivations are likely to vary according to the context of the organization, notably the priorities of that organization. We also note that altruistic vs self-oriented goals and motivations can be particularly affected by impression management and self- deceptive enhancement, as people may be reluctant to admit to certain “selfish” orientations. An interesting further study would be to see how participants react to reputation building motivations in the real world, and to look at real reactions to communications or larger scale interventions rewarding energy saving at work. Finally, we note that whilst we used a large sample of participants, this sample was not nationally representative of the UK population. Our sample in Study 1 was drawn from a sample of university staff and one large private company, and in Study 2 from 3 smaller companies. Also, the two studies differed in their gender proportions, with study 1 including more women and study 2 more men. These might have had an impact on our results as men and women can differ on their thoughts regarding climate change (McCright, 2010; Whitmarsh, 2011). Furthermore, our two study samples may have distinct characteristics, e.g. more positive views of their workplace, which could have influenced their motivations. Finally, respondents may be more altruistic than others, given that they were willing to help with this research. Our two studies also differed in the type of organizations they were taking place in, representing different types (e.g., for profit vs not for profit, private vs public), sizes, cultures and possibly attitude and behavior towards sustainability. These characteristics should be investigated in future research using larger, more representative samples, or in case studies that would investigate systematically the motivations in a single site (Tudor, Barr, & Gilg, 2008). Practical implications ~~~~~~~~~~~~~~~~~~~~~~ Our study reveals that various motivations to save energy exist in the workplace, so communications and interventions to save energy in the workplace should take people's personal motivations into account as impacts of interventions are stronger if the type of communication matches the individual's motivations (Shavitt & Nelson, 2002; Unsworth et al., 2013). Indeed, there is a “congruency” effect of motivations and campaigns that Clary and colleagues found for volunteering (Clary et al., 1998), that was also observed within literature on attitude functions (Shavitt & Nelson, 2002). Their studies showed that campaigns to promote volunteering were evaluated more positively if they were congruent with participants' motivations to volunteer. Furthermore, communications may result in being inefficient or even counterproductive if they focus on the wrong motivations. Indeed, it seems that reputation building motivations, and hence direct reputation building rewards, to save energy at work, could make people less willing to try to save energy. Further research should test this idea with interventions that go beyond simple text communications, as interventions that combine approaches, e.g. text communications alongside policy or structural changes, are usually found to be more impactful (Michie, van Stralen, & West, 2011; Spence & Pidgeon, 2009). Campaigns to save energy could usefully focus on integrating various motivations and goals within company policies, in particular to reduce the negative findings that people are less likely to be sustainable if motivated by reputation building. More research is necessary to explore why this is the case. It could be that such extrinsic motives would impair motivations if they are already internalized (Deci, Koestner, & Ryan, 1999), or that saving energy could be perceived as actually bad for one's reputation. For example, interventions could help increase the reputation of people who save energy, for example by clarifying the importance of energy saving for the company, while not insisting on its moral dimension (Minson & Monin, 2012), and making energy saving action more visible (Griskevicius et al., 2010). Alternatively it may be possible to utilize feedback on how environmental actions can help an organizations image, in order to encourage energy saving behavior without creating specific goals that could be perceived to be in conflict with work performance goals (Unsworth et al., 2013). Other self-oriented motivations, such as warm-glow, have a positive effect on sustainable choices, and may have positive effects on energy saving and campaigns. Importantly in the workplace context, environmental goals and work performance goals are often presented as potentially competing goals, meaning that when one is activated, it will be at the expense of the other (Unsworth et al., 2013). In the workplace, important work performance goals are thought to be most efficiently fulfilled by not saving energy at work (e.g., taking the elevator, keeping one's computer on to save time). These goals, higher in the hierarchy of goals (focal goals) compared to environmental concern (background goal), in the context of the workplace, will encourage people not to save energy (Fishbach, Zhang, & Koo, 2009; Kruglanski et al., 2002), or increase rebound effects or “fads” of energy saving behaviors (Unsworth et al., 2013). Our results show that promotion of energy saving policies at work can avoid these issues, because saving energy at work may then operate to fulfill work-related goals. Indeed, helping one's organization image and finance are often higher in the hierarchy of work goals than environmental concern, and less in conflict with work performance, which can lower the potential for competing goals in the workplace. Altogether, these results imply that in addition to environmental concern, energy saving motivations exist which may actually complement work performance goals, and provide guidance on the design of energy saving interventions and campaigns. Interestingly, introducing environmental issues into organizations' operations (de Burgos Jiménez & Lorente, 2001) does not necessarily mean that a system of direct or indirect “social” reward should be introduced. We see here that it could be counterproductive if done superficially, however real and significant changes to social norms and culture in the workplace (which change current motivations and relationships with behavior) could have substantial benefits. In relation to organizational culture, promoting employee identification and commitment to the company may help to promote energy saving more broadly through increasing people's motivation to act to promote the organization's image. Making energy saving or environmental goals part of an organization's image is likely to be beneficial in highlighting this to employees as a clear organizational goal."],["Families of missing people are often understood as inhabiting a particular space of ambiguity, captured in the phrase ‘living in limbo’ (Holmes, 2008). To explore this uncertain ground, we interviewed 25 family members to consider how human absence is acted upon and not just felt within this space ‘in between’ grief and loss (Wayland, 2007). In the paper, we represent families as active agents in spatial stories of ‘living in limbo’, and we provide insights into the diverse strategies of search/ing (technical, physical and emotional) in which they engage to locate either their missing member or news of them. Responses to absence are shown to be intimately bound up with unstable spatial knowledges of the missing person and emotional actions that are subject to change over time. We suggest that practices of search are not just locative actions, but act as transformative processes providing insights into how families inhabit emotional dynamism and transition in response to the on-going ‘missing situation’ and ambiguous loss (Boss, 1999, 2013). --------------------------------------------------------------------------------","Most of us in well resourced, democratic societies live with taken-for-granted securities in ordinary life in which our living loved ones are almost always contactable or known to be somewhere. For some, however, this sense of security is threatened when a family member or friend or colleague is missing, something that happens with surprising frequency with approximately 306,000 annual incidents in the UK (NCA, 2014). This paper considers what emotive actions accumulate in the space of absence for the people left behind, drawing on a funded research project in which UK families1 were interviewed about their experience of living with the absence of, and search for, their missing person. As we discovered, search/ing2 for a missing person is an emotional process, one also marked by (often competing) geographical knowledges and complex relationships with police officers charged with the task of locating the missing (this process may be significantly different elsewhere in the world, and see Edkins, 2011; for examples). We start by situating the paper with reference to interdisciplinary research concerning ambiguous loss and grief. This literature suggests that humans cope with absence via ‘continuing bonds’ with those who are gone, but that they also may become fixed or frozen by the trauma of their loss, especially in the case of ambiguous absence (Boss, 1999, 2013). In the next sections we explore the implications of these arguments through empirical materials related to family experiences of search/ing for their missing person and their communications with police officers. We disrupt a straightforward story of the freezing capacity of loss in relation to missing people, identifying the many ways whereby families are active agents in responding to this particular kind of absence. In this way we are highlighting how people might manage ambiguity, and thus elaborate Boss's work (also explored further below) as she rejects ‘stages’ or ‘steps’ of recovery from ambiguous loss, while also pointing researchers towards ‘movement, paradoxical possibilities of change, and diverse paths to resiliency’ (Boss, 2007: 108. See also the work of Glassock, 2006, 2011 and for a review of literature on loss and hope in a similar context see Wayland et al. 2015). We thereby explore search/ing as a key mode through which such emotional management happens, rather than (just) as a sign of frozen incapacity. Search/ing here is understood not as a unified category, act or feeling, but instead constituted by a diverse geography of shifting modalities, materialities and meanings.3 In doing so, we move from accounts of family liaison with police in efforts to locate missing people in the external world, to more reflexive and interior accounts of long-term search wherein the missing person finds a new place in the imaginations of family members. The paper concludes by suggesting that families of long-term missing people find new ways to live with ambiguous loss, closely connected to changing search experiences and evolving emotional geographies of human absence (and see also Parr and Stevenson, 2014).","Every year in the UK over 306,000 missing incidents are recorded by the police (NCA, 2014) with around 35% of these being adult missing persons (the concern of our paper), some involving repeat missing events by the same people. While the majority of cases are resolved within three days, others continue for much longer. It is difficult to gain accurate data of medium and long-term missing incidence, but about 1000 cases are outstanding every week in the UK and although 97% are eventually recorded as closed cases with no harm to the individual, 1% of cases are unresolved after a year according to the UK Missing Persons Bureau (the remainder being recorded as fatal incidence) (NCA, 2014). Despite the increasing (although patchy) statistical data on numbers of missing incidents profiled by age, gender and location, little is known about how missing people's absence affects lives over prolonged periods from the perspective of those left behind (for exceptions see Boss, 1999, 2006, 2012; Edkins, 2011; Holmes, 2008; Henderson and Henderson, 1998; Parr and Stevenson, 2013, 2014; Wayland, 2007; Wayland et al., 2015). The UK charity, Missing People, receives around 17,000 annual calls from families wishing to reach out for support from their 24 hour help-lines and dedicated counselling provision, and so the scale of emotional need is clearly significant. What kind of loss does such human absence provoke? How might we best understand this from geographical and other disciplinary perspectives? What spaces do people dealing with the absence of a missing person inhabit? How are the missing re-presenced and through what kinds of emotional, social and spatial practices? Such questions are ones that chime with contemporary writings, including those stressing how presence and absence exist in unstable and sometimes unexpected relationships (Till, 2005; Wylie, 2007; McCormack, 2010; Maddrell, 2013). In such work we are often reminded that geographies of absence are not always about spectral remains and ruination: ‘less as a disputed articulation or representation of the past and more as part of contemporary everyday activities, bodily experiences and contestations’ (Meier et al., 2013: 424). This emphasis on experiential and embodied qualities of absence might also be understood further through analysing feelings and materialities of loss, such as those that accumulate in the wake of missing episodes and journeys. Meier thus provokes us to understand the experience of absence further, and as something or someone as present, rather than something or someone that is simply recalled. In work on the loss of people made absent via death (Maddrell, 2009; Maddrell and Sidaway, 2010; Maddrell, 2013; Neimeyer, 2001; Neimeyer et al., 2006), rich discussion focuses on the relations surrounding end-of-life, with Maddrell (2013, p1) writing extensively on the liminal spaces of grief and how these are infused by ‘dynamic negotiations of absence–presence’. She notes how, for those in grief, the absent deceased person is simultaneously ‘nowhere, but everywhere’ (ibid: 4). Drawing on recent bereavement studies and the notion that the bereaved experience ‘continuing bonds’ with the dead, Maddrell argues that the absent dead are ‘given presence through the experiential and relational tension between the physical absence (not being there) and emotional presence (a sense of still being there)’ (ibid: 5). Maddrell (2009, 2013) has particular interest in how specific places of memorialisation and remembrance can help form bridging relations for absence–presence and be sites of existential encounter with the deceased. For Maddrell, memorials are important material spaces that enable continued relationships between the living and the dead, although they are not the only ways for this continuance to happen. For families of missing people, such material memorial spaces do not necessarily exist4 or feel appropriate, and so they may be left with more diffuse traces of the missing that reverberate through their everyday lives, in a manner not dissimilar to the absence–presence of the grieved-for dead, but perhaps experienced with a particular inflection precisely because they do not know if their person is still alive and somewhere. Families of missing people, like those in grief, also work to (re)presence the absent but via particular practices and spaces, such as celebrating birthdays, sending nightly text messages, setting up websites or using media to ‘witness’ the person's character and interests or call for their return (and see other examples in Edkins, 2011 and Wayland, 2007). Most poignantly, the re-presencing of missing people is usually attempted by the continued search for them by their families, a practical activity that happens in parallel to, and not always in partnership with, official police search enquiries. While we have acknowledged some comparisons between the experience of bereavement and the experience of knowing a missing person, we also want to draw out some differences, or particularities, as a precursor to understanding the family search efforts discussed below. Families of missing persons are often described as ‘living in limbo’ (Holmes, 2008), with the resultant states in which people find themselves often referred to as ‘ambiguous loss’, as noted earlier. Ambiguous loss is a term coined by the family therapist Pauline Boss who has worked with families of missing people, among others (Boss, 1999, 2002a,b, 2006, 2007, 2012). According to Boss (1999), ambiguous loss should not be confused with ‘ordinary loss’, as missing people are physically absent but psychologically present for their families. In Boss's view, this loss is different to the experience of death where one can know it is impossible for the deceased actually to return, even though there may be a powerful ‘continuing bond’ (Maddrell, 2013). Boss's research and practice leads her to understand the ambiguous loss of missing people as particularly difficult and disabling. She represents ambiguity in this case through a fixing language, elaborating its ‘freezing’ properties, as ‘they (the families) cannot make decisions, cannot act, cannot let go’ (1999, p61), due to the unknown cause of the absence. Ambiguous loss may also be evoked as a time/space of uncertain waiting that cannot straightforwardly be ended unless there is a body located (Hogben, 2006), although Boss hints that versions of ‘flexible waiting’ may actually be productive for some: “Those who wait endlessly for news about a lost person do not do so in vain if they find hope and optimism in their struggle. Indeed, they are able to find meaning in the midst of ambiguity because of their ability to remain optimistic, creative and flexible” (1999: 132). So, acknowledging the contradictory spatialities suggested by ‘frozen’ and ‘flexible’, as attached to those waiting for missing people to return,5 we now explore what these dynamics might look like in practice, particularly in relation to the role of search/ing. Here we question how and whether family search/ing enables ways of managing the emotions constituting ambiguous loss and the difficult absent–presence of their missing people, or prompt a kind of repetitive paralysis where apparent ‘flexibility’ in constant search/ing may be masking frozen loss. The paper draws on interviews with 25 families of missing people in the UK who were recruited via two police forces in England and Scotland and the UK charity Missing People database.6 The semi-structured interviews were led by a thematic topic- guide designed to reflect on the search experiences of families (the lead up to the disappearance, reporting the person missing, police and family search, dealing with returns). Family experiences in this regard ranged from having a relative missing for a few hours to a few weeks to 20 years, and included those who have had a return, alongside families who are still search/ing for loved ones (Parr and Stevenson, 2013). Of the families interviewed, 11 had relatives who were still missing. Interviews ranged from 1.5 to 3 h and involved couples and single interviewees, focus group and telephone interviews with some follow-up liaison. In what follows, we have deliberately avoided profiling multiple and named emotional testimonies of pain, anguish and profound psychological disturbance that many experience when dealing with a missing relative. This is partly because the emotional impacts of missing loss are noted elsewhere (Holmes, 2008), and partly because in our research project we are interested in the role of family search practices. Our exploration of such search practices also reflects how such practices feel and here we note that “emotion … is not a static thing-in-itself, but relationally constituted, dynamic, and so subject to shifts in position and relative power” (Thien, 2008: 312). As search/ing may involve subtle emotional transformations, this subtlety is what we seek to disclose. In what follows, we subtitle all our empirical sections with the adjective ‘search/ing’, precisely to convey a sense of sustained (maybe even ‘relentless’) family efforts to locate their missing people.","There are many dimensions to the search for missing family in different times and places and with different technologies of practice (for examples, see Edkins, 2011; Parr and Fyfe, 2012). In discussing UK family experience, it is acknowledged that we are necessarily privileging a partial representation of the processes and complexities involved. From the moment a person is noticed as absent, there unfolds not only a search to locate the person, but also a search for meaningful answers to questions such as ‘why?’, ‘how?’, ‘when?’, where? (Landsman, 2002). We might imagine a series of scenes as an absence unfolds: an absence ‘becomes’ via a concerned glance at a clock, confused phone-calling of friends and family, checking local workplaces, pubs or pathways and finally a call for police help. The arrival at a decision formally to report an adult missing is often fraught with anxieties around an individual's right to be absent and a strong desire to know their whereabouts. Once an absence is reported to the police, it triggers official risk assessment and search procedures that can leave some families worried about whether their missing family member will appreciate such intervention. This is especially the case for individuals experiencing mental health problems, where such intervention may result in a medical and legal process, such as a Mental Health Section and detainment in hospital. Some families have clearly struggled in such circumstances about whether and when to report the absence, with some families recalling instances where they have mounted significant searches of their own before calling in police assistance, especially in cases of repeated disappearance. For other families, a lack of knowledge of when and how to report someone missing has seemingly compounded the distressing nature of the initial stages: The police actually said to us “why did you leave it so long to contact us?” and I'm thinking “I thought they had to be missing at least forty-eight hours” and he said “no its a misconception, you know, if somebody isn't very well or has some kind of problems you can get in touch with us in a couple of hours if you are concerned.” (Judy, mother of Andrew, missing for 2 years) For the police, the ‘golden time’ for success in search for absent people is in the first 24 h after a person has last been seen, although there are barriers to the optimal use of this time-frame, as indicated above. Deciding if a person may be missing is tricky and families do not always feel qualified to make this decision alone. Quite often the decision to report a family member as missing takes place in conjunction with, or is prompted by, conversations with agencies and other family members. Feeling as though one can act in the face of another's absence is clearly a confusing emotional burden, but one with implications for the practicalities of effective search. From the first moment of an absence ‘becoming’, there do, indeed, seem to be ‘freezing’ forces at work in terms of how people respond to human disappearance (Boss, 1999).","Once a person is reported as missing to the police, an official search may take place.7 The Association of Chief Police Officers (ACPO, 2006: 94) guidance states that ‘search is a routine element of investigating reports of missing persons. It involves making an assessment of what the initial enquiries suggest are the most likely circumstances of the person's disappearance, and then concentrating the search in accordance with those circumstances’. Several families reported feeling confident in the police as search experts, recognising that the police are probably the most effective means to locate their loved one, given the knowledge and resources assumed to be at their disposal to carry out effective searches, as summarised here by Alice: They were very, very quick at getting searches up and running so there was no need for us to do anything like that. The police are the specialists, they know what they're looking for. (Alice, step-daughter to Martha, missing for 5 years) Family decisions not to search may be related to their perception of police as search specialists, although the majority of families did engage in some form of search themselves alongside or in response to police actions. In understanding the drivers for family search/ing, we might recognise not only an emotional need to be ‘doing something’ in the face of human absence, but also consider family relationships with police officers officially tasked with locating the missing person. Based on contingent information, police search strategy clearly varies in type and extent (Parr and Fyfe, 2012; Fyfe et al., 2014; Glassock, 2011; Gibb and Woolnough, 2007). However, and unlike Alice above, many families talked about feeling the level and scope of the police search to be inadequate. Knowing that ‘no stone has been left unturned’ in the search was represented as critical for a family's psychological and emotional welfare. One of the key factors contributing towards a positive experience of police liaison lay in a clear understanding of police decision-making about the parameters of any search. Yet, many interviewees, particularly those who asked precise questions about the geography of police search, remained dissatisfied: One of the things that I requested was the copy of the [search] map, and they went “well nobody has asked for it before” and they've got the map here so I'm peering across the table upside down at the map and I'm saying “well I really can't see” so I'd be turning it round like this and they'd be turning it back and saying “it's our map”. So I said “I'd really like a copy” and they went and did a copy in black and white. I was just furious. I said “how dare you? Go and do a colour copy” so they did a colour copy but there was no key on it, there was no legend, so all their little crosses and colours didn't mean anything to me and the police officer who was explaining it couldn't interpret either. She said “well we've done all of this” and I thought well it doesn't help me. (Sasha, wife to Bill, missing for 2 years). The literal mapping of search could convey the extent of search effort by the police, but some interviewees recalled being met with reluctance by officers to share such technical geographical details. Families interpreted this as a lack of police engagement and that their particular missing person was not central to policing process (Edkins, 2013; Parr and Stevenson, 2013). Some discussed feeling confused and frustrated, then, not only over the lack of clarity about who was in charge and how the search was going, but also due to a mismatch of expectations in relation to analytical and spatial parameters of police search: I just felt at the time that their analysis was poor. I was expecting a more detailed analysis of what he was wearing and the circumstances and his situation more from them. (Sasha, wife to Bill). For others, the lack of demonstrable or varied technical search activity led to deep and resentful attitudes towards the police: An hour searching for a young women, an hour search!, no heat seeking!, there was no scuba sonar! We are not even one hundred percent sure whether there was even a lifeboat, they just combed the beach and said “I don't know, we'll just leave it at that then shall we?” (Raquelle, sister to Libby, missing for 3 years). Where specific details of search were well communicated, it directly related to family perceptions of its standard, and their enhanced emotional handling of the (potential) loss. Below Sasha answered a question about what forms of search information helped her emotionally: There was a Search and Rescue Officer … He talked about how difficult it is to find a body after a certain length of time, how the first twenty-four hours are crucial, the fact that fourteen days have gone past made the search much more difficult and would need a specialist dog by then and he talked, but not in a frightening way, he talked about other environmental factors that you would look at …. I know it sounds gruesome, but in a way that actually was tolerable to listen to and I was able to acknowledge it. (Sasha, wife to Bill). Here Sasha, expectant and demanding of the police officers involved in her case, was appreciative of the shared environmental knowledge of the possible decay of her husband's missing body. Her need to understand the technicalities, confront the material details of possible death and share in the detailed geographical analysis of search was critical in her rejection of policing relationships where this service was not offered. Sasha wanted to be involved with the logic and rationale of search based on her knowledge of her missing husband (and see Parr and Stevenson, 2014 on witnessing the missing), so that she could be emotionally assured that highly professional and expert work was taking place. Although many families understood that they cannot obstruct or disrupt police investigation, many were sure that they could have more productive working partnerships, and that improved communication, information flow, content and task allocation could reflect and produce better police-family liaison. Moreover, such participation could help manage distressing emotions. Indeed, for Sasha, as for others, the reluctance of the police to share information and logic for search parameters led to her conducting her own search enquiries as a response to her strong needs in this respect.","In light of the varied experiences reported above and the support of an active charity (see Parr and Stevenson, 2013), the families participating in our research reported being forced or inspired to take search/ing into their own hands (and see Olsen, 2008). This section explores what families have to say about such practices, and Table 1 shows the range of practical search activities that the interviewees discussed. The table differentiates different types of activity – physical, documentary and virtual, social networking, liaison with other agencies/professionals and other practices – and captures the enormous lengths to which families may go in order to try to locate their loved one or information about them. Sally explained why the sheer emotional trauma of a loved one's absence can galvanise people into initial search actions: It's like a massive shock, but then you kind of feel like, well for me, I kind of felt like I had to take action. If someone dies you can't do anything about it, whereas [for missing people] you kind of feel like you have to do something. (Sally, daughter to Ned, missing for 7 days, returned). Advertising the absence via posters, phone-calls, door knocking and route-tracing are some of the very first search practices tried by families. Media reports of human absence (particularly children) can prompt intensive and large-scale family and community search/ing (see photograph 1): Intensive search may take place with police co-ordination, as in the photograph above, or without, as these memories of the turmoil of early search reveal: I got into organisational mode. So I contacted Missing People, got some posters printed out, got some of my old friends from home to hand them out around the town and then we started looking at places that we felt he might have gone to. We went to London and we stayed there for two nights and basically went round all the big stations and stuff like that. As soon as you turn up in London you are like I don't even know where to start you know what I mean, you are just walking through parks and things going “I don't even know what I'm looking for.” So I spent two days in London and then we went to [East town] briefly and then we drove down to [the South Coast]. (Sally, daughter to Ned). Many families pound the streets, draw up maps and follow lines of lateral thinking about where their missing members might be. This initial physical search is characterised by an intensity that belies the weight of the loss for families and the unbearable nature of the absence. Raquelle described how she first intensively searched the beach where her sister was last seen for the smallest of traces of her, and still finds it difficult to stop doing so: Looking, looking. Going up to that beach every day, every day … A lot of walking, beach combing, looking not just for my sister but for belongings, her house keys. I had my husband climbing up rocks and looking in little crevices, just to see, and the woodlands that were around there. I don't think I'll ever stop. (Raquelle, sister to Libby). Several interviewees reported on their detailed search for personal belongings and effects, as they actively seek both a person and their physical traces: I put up about two hundred and fifty posters on the hill saying we are looking for a haversack, we are looking for a black coat, black walking boats. (Sasha, wife to Bill). Such search/ing practices draw on latent knowledges of the personal geographies of their missing member. In practical terms, this can mean remembering and retracing their usual routes and routines, but also more in-depth and emotive appraisals of ‘where matters’ to the missing person and why. This difficult task might take in childhood haunts, sites of romantic significance, death places and graves, well-appreciated landscapes and favourite views, or general areas, regions and preferred pathways. For Pauline, whose son has been reported as missing many times, she related how she manages the physical search, bound up with a detailed knowledge of her son and the local area: I go out in the car to where I think his normal haunts would be, little paths he would take … it might be midnight, but usually about ten, eleven, twelve, the streets around our area tend to clear and there's decent street lights. So I go … I drive round and round … all the streets [… ], because of a pattern that he's followed in past experience. (Pauline, mother of Paul, missing for 2–4 days repeatedly). In exercising geographical imaginations about where matters to the missing person, a process of re-considering the person, what is known of the missing life, and the cause of the disappearance often occurs. It is in this reflexive space where transitional and dynamic emotions become most apparent, and a shift in effect from external to more interval search/ing may also occur (from pounding the streets to re- search/ing memories and emotions).","In imagining the spaces where a missing person might be, different scenarios are often considered over time. For many people who live with the ambiguous loss of a missing person over prolonged periods, different emotions emerge at different times and relate to changing understandings of what might have happened. Gladys, whose husband has been missing for 20 years, charted such change: from her first recognition that he had gone missing, ‘Just complete shock. Shock and fear. Just horrified’, to more hopeful ‘middle years’ where ‘I'd been over to [European island] where he'd [possibly] been sighted’, to latter feelings of abandonment, ‘he's always somewhere. It might not even be his face sometimes [in my dreams], I just wake up feeling deceived’. In such search/ing journeys, which sometimes occur over decades, the absent person resonates across subtly transitional emotional lives, and alongside changing capacities for living life. In parallel, the mode of search may begin to change: It's a bit of moving on, but it's also realising that he's made his decision. He's made that decision to go, for whatever reason, we don't know what that is and we haven't got any control over that … So therefore we might as well get on with what we're doing. I think we've coped by being able to reassure ourselves we've done as much as we could. (Charles, father of Simon, missing 2 years). Physical search practices, ones located in external space, and aided by calling, maps, walking, driving, posters and other technologies, may therefore diminish over time as other types of search activity take over (such as Internet use). Here search/ing becomes part of a more muted routine in the newly established life that follows absence in long-term cases. Many discuss a lessening of general search activity, undertaking it every week or every month instead of daily. Some gradually realise or come to believe that they are no longer search/ing for a living missing person, but rather for a dead body, and this is reflected in emotions and practices which may change as a result. Indeed, the initial emotional intensity and effort of search/ing can be extremely difficult to sustain, as Raquelle elaborated: I can play private detective until it sends me mad, you know, so I can't. You also have to slow yourself down a little bit because you still have to go to work and you still have to be mum and you still have to function and you do have to tell yourself “just stop, just slow it down a bit” because otherwise you would be out there until it would make you ill I think. (Raquelle, sister to Libby). In these shifts, it is, arguably, not only the practical external modes of search that change, but also the emotional search for understanding and meaning, and the place of the missing person in the family imagination of itself. In support services, health and counselling practitioners emphasise the importance of encouraging emotional reflexivity over time: ‘Empowering families to work this timeline [a trauma time-line] into a story of emotion as well as practical content, gives depth and perspective to their experience without it being just an external recount of events.’ (Wayland, 2007, p16). Although many families reported the difficult emotional consequences of ‘living in limbo’ (Holmes, 2008), they do not seem straightforwardly to embody the ‘frozen states’ that Boss's (1999) early work outlines. Rather, interviewees related nuanced strategies, actions and changing forms of emotional hope (FFMPU, 2005) that morph alongside their continued search/ing: from ‘hope of a reunion, to hope of information, which finally became hope of resolution’ (cited in Wayland, 2007:12). We explore these themes further, below. For some families in long-term cases, then, active search/ing is replaced by other practices that exist alongside changing senses of, and hopes about, the missing person. Such intimacies are more difficult to describe, but relate to an everyday alertness and a latent awareness that the missing person might be present or appear in routine or random public environments. For Alice, search/ing was hence replaced by looking, a qualitatively different experience, one in which the missing person is anticipated as possibly present but in ways other than via ‘conscious search/ing’: You are constantly looking even though you aren't searching … So even though it's not a conscious search, even today we are still looking. (Alice, step-daughter to Martha). Sasha suggested that even ‘just looking’ – as an anticipatory gaze - can transform into something else again, perhaps a practice of remembering, partly through simply being in places that were significant in the relationship with the missing person: So it's much more [… ], rather than ‘a look’ [… ], it's a remembrance of how we used to like coming here. (Sasha, wife to Bill). Search/ing in family narratives is thus represented as a transformative and transforming process, moving from an intense physical search to more documentary and virtual forms, and to practices of looking and remembering. It is critical that such transformations are not understood as modelling ‘normal stages’ of loss (as in early conceptions of ‘grief stages’) and, importantly, that these can be premised on non-linear relationships with forms of new information, technological advances or renewed energies or optimism, amongst other factors. Sometimes this can be a dynamic process associated with, or disrupted by, the location of sudden or suspected sightings, which also become a focus for changing geographical imaginations over time. Sasha discussed the changing imaginings that were bound up with her own idealisation of where her missing husband was likely to have ended his life. She related how her initially romantic vision for where he may have chosen to die is now tempered with a more realistic assessment of the likely ‘where’ of her husband's body: We walked there, walked our dogs there and we would often go back to walk there, so it seemed the most natural place for me. In hindsight, I think that's my kind of romantic ideal, because I think when you are planning to go missing with an end result, when you are going to end your life, I am not sure you are choosing to go to the most beautiful place, I think you are choosing to go to the place that you won't be found. (Sasha, wife to Bill). Over time, Sasha changed her view to incorporate a new, painful imaginary – that of an unknown and hidden location for her husband's body – as she gradually accepted that, in ‘doing absence’, missing people like her husband may seek to access precisely “the place where you won't be found” (Sasha). Geographical imaginations of ‘where matters?’ thus relate to the pragmatic process of search and police liaison, but also the changing ways in which those who are left behind refigure the absent person and the disappearance in their thinking and emotional life (and see also Parr and Stevenson, 2014).","Loss that is never ending can be crippling. To be ‘left behind’, with little or no evidence of where a loved one has gone and whether they will return, is an intensely difficult experience. Although it is recognised that adults have a right to go absent, families often struggle to cope with the possibility that their missing person has left deliberately and without trace, especially when it seems out of character. Regardless of the time period concerned, families often long for some form of communication that would signal connection or resolution, and help them to transition away from feeling ambiguous loss. For some, as new search leads and actions reduce overtime, ambiguity may remain, accompanied with what Horacek (1995) calls ‘shadow grief’, referencing a sense of loss that is not acute but persists and is part of a continuing relationship with the missing person. Gladys explained her struggle to manage the ambiguity and her sense of simultaneous connection and disconnection with her missing husband following twenty years of search/ing: I would like … only to see that face. I don't want to barge in and destroy anybody's life, it's not what I'm about. I just want peace of mind for myself. That's all I want, just peace of mind and to stop this never ending frustration and sadness. (Gladys, wife to Samuel). Until such a time, people like Gladys develop strategies in an attempt to live with absence, such as concentrating on practical issues, keeping busy to try to move away from the pain, and seeking the support of others (see Holmes, 2008). Learning to live with the constant demands of absent–presence in missing situations is complex, and interviewees still found it emotionally hard to use time for leisure instead of search/ing, although they did eventually do so: If I go to the ballet at the weekend, there is a little bit of you that says “oh you could have been looking at the map”. (Sasha, wife to Bill). The need to remain alert and aware for long periods of time, readying oneself for the potential trace of the missing person, is experienced as a form of ‘hyper-vigilance’ and a well recognised psychological issue for the people affected (FFMPU, 2010). As a result of the long-term effects of hyper-vigilance, family members are sometimes reluctant to leave home, even for short periods, in case of a return, but some do eventually manage this, if they put contingencies in place: Late Friday night, stayed there Saturday and came home Sunday. So I think that's the longest I've been away since Andrew went missing. I [texted] him, that's where we were going. And I said the keys are in the usual place if you want to come home. (Judy, mother of Andrew). Until a loved one is located or returned, many express the impossibility of giving up on search/ing or ‘moving on’. What may be more common in long-term missing situations, however, is a transition gradually allowing more flexibility in living everyday life, as Raquelle, Sasha, Judy, Gladys and Alice all discuss in their different ways. Here families become more flexible in their search/ing practice and mode, although still perhaps bound and limited by its repetition. Understanding such repetition as examples of ‘frozen loss’ may risk underestimating how repeated search/ing efforts change and how the felt ambiguity of ‘not knowing’ can transform from raw trauma into an ‘everyday remembrance’ of absence, lived out through muted practices of looking, and a latent awareness that the missing person is possibly still present somewhere. From a ‘family resilience’ perspective (Becvar, 2013), the quietly transformative experiences hinted at above might be understood to constitute a kind of resilience, what Masten (2001) calls the ‘ordinary magic’ emerging from ordinary processes in everyday life. Boss (2013, p288) elaborates this claim in the context of resilience theory and ambiguous loss, noting that, in order to move on from the ‘frozen states’ that a missing situation may produce in those ‘left behind’, it is necessary to ‘revise one's attachment to a person who is missing’. She argues that ‘because ambiguous loss is a relational problem, relational interventions are most effective’ in achieving this revision. For Boss and Carnes (2012), relational interventions bound up with an array of therapeutic and ordinary strategies, different ways of finding meaning in an irresolvable situation, and via accepting new forms of dialectical thinking: for example, ‘I have a son and he is missing, he is present and absent’. As geographers, we might also argue that the changes in the spatialities of search/ing (from ‘external’ physical practices through to more ‘internal’ remembrances and reconceptualisations) are critical in such relational dialectics and meaning-making. Flexible geographies of search/ing (reducing intensities and changing modes, re-imagining places of disappearance, accepting spaces of absent–presence) are arguably important in themselves as processes in the difficult task of living with ambiguous loss. Search/ing thus might be understood as comprised of relational spatialities, not just signs and symptoms of frozen or flexible ways of living with loss, but as a means through which emotional transformations might be lived out.","“The absence of people that have been, of things that have been but are not anymore, can hurt deeply. The experience of such absence can exert so much force, so much gravity, that it feels as if it pulls one's heart out” (Frers, 2013; 431/2). I still text him every single night, and I suppose I do that because it's a form of contact with him, as if I'm talking to him … [but] the messages have changed. (Jane, mother of Paul). In this paper we recognise that there are different spatialities of search/ing which are related to the ambiguous loss that accompanies the search for a missing person. We have suggested that there are numerous spaces that families of missing people inhabit, both material and emotional. In journeying from police liaison to street-search/ing to virtual spaces of tracing work to living everyday life with barely conscious practices of ‘looking’ and remembering in multiple public spaces, family members demonstrate their changing response to their unstable status as ‘left behind’. These changing search practices and geographies – deliberately not represented here as a stage model - are also accompanied by changing feelings, with the voices above describing tentative ways of responding to the loss endured. We understand these narratives to do more than demonstrate frozen states of psychological trauma where ‘families cannot make decisions, cannot act, cannot let go’ (Boss, 1999: 61), but, equally, neither do they represent a straightforward moving on whereby the ambiguity of loss is resolved. In arguing thus, we suggest that the experience of a missing person's absence is not just endured as a deep and forceful hurt (referencing Frers, above) for those left behind, but is also experienced as a processual emotional geography (referencing Jane, above). Here, spatialities of search become important entities through which ambiguous loss can be newly articulated and understood but never be entirely fixed. In Boss's later collaborative work (Boss and Carnes, 2012: 457), the authors advocate a new kind of search in the face of ambiguous absence; as they say, ‘when loss has no certainty, the search for meaning is excruciatingly long and painful, but it is the only way to find resiliency and some measure of peace’. For Boss (2004), a key step for those left behind is an acceptance of their lack of control and mastery over the situation. Instead of relentlessly search/ing for a physical presence or news of the absent person, a different kind of search for meaning may incorporate new relationships with the missing person. In previous work (Parr and Stevenson, 2014), we have suggested (and following Wayland, 2007) that this might happen through small ritualistic celebrations of the story and person so far. Indeed, it may be that in generating ‘ideas about how the missing person can be celebrated’, instead of just missed, that helps produce meaningful lives lived for those ‘left behind’ (and see Carnes, 2008). This is not the same as remembering the dead and experiencing the continuing bonds in grief (as discussed in the introduction): for in this case, families still hold open the possibility that the absent missing may one day speak back to address their place in such narratives. Such an approach acknowledges, and does not deny, uncertainty, and, Boss (2008, 19) argues that, ‘it is because of the mystery that we honour rather than memorialise …. by having a tribute, the uncertainty is not denied’. Family members often realise they may never know ‘the where’ of their dead/alive missing person and so a new space between loss and mourning (what Wayland calls ‘the space between grief and trauma’) may have to be found and occupied in order to live with the ambiguity of present-absence. We hope this article is one extra resource for that difficult journey."],["Distorted negative self-images and impressions appear to play a key role in maintaining Social Anxiety Disorder (SAD). In previous research, McManus et al. (2009) found that video feedback can help people undergoing cognitive therapy for SAD (CT-SAD) to develop a more realistic impression of how they appear to others, and this was associated with significant improvement in their social anxiety. In this paper we first present new data from 47 patients that confirms the value of video feedback. Ninety-eighty percent of the patients indicated that they came across more favorably than they had predicted after viewing a video of their social interactions. Significant reductions in social anxiety were observed during the following week and these reductions were larger than those observed after control periods. Comparison with our earlier data (McManus et al., 2009) suggests we may have improved the effectiveness of video feedback by refining and developing our procedures over time. The second part of the paper outlines our current strategies for maximizing the impact of video feedback. The strategies have evolved in order to help patients with SAD overcome a range of processing biases that could otherwise make it difficult for them to spot discrepancies between their negative self-imagery and the way they appear on video. --------------------------------------------------------------------------------","Video feedback can be used in a variety of different ways during a course of CT-SAD. The first time that it is used is in Session 3. In the preceding session, patients engage in an experiential exercise in which they have a social interaction under two conditions: first, while focusing their attention on themselves, evaluating their performance, and engaging in habitual safety behaviors; second, while focusing externally, getting lost in the interaction (as opposed to evaluating it) and dropping their safety behaviors. Typically, they discover that they feel less anxious and think they come across better to other people in the second condition, a discovery that is built on in subsequent therapy. In Session 3, video feedback of both social interactions is used to help patients compare their predictions of how they think they came across with how they actually came across. McManus et al. (2009) found that video feedback in Session 3 helped 94% of patients to see that they came across to other people more positively than they had predicted. Since the McManus study, improved availability and functionality of camera and smartphone technology has increased the options for using video feedback. Technology now makes it easy to capture as a still shot key moments that disconfirm patients’ negative beliefs, and we have found this to be a very useful technique for consolidating learning. Further experience in using video feedback has also made us more aware of some of the processing biases (outlined above) that can make it difficult for patients to fully benefit from viewing themselves on video and clinical strategies for overcoming these biases have been developed. In order to assess the impact of our current way of using video feedback, we present video feedback data from our most recent RCT (Clark et al., Trial registration: ISRCTN95458747) and compare it with that reported in the McManus et al. (2009) study.","Participants were 47 patients who met Diagnostic and Statistical Manual for Mental Disorders (DSM-IV; American Psychiatric Association, 1994) criteria for SAD and were treated with individual cognitive therapy for SAD following the general protocol outlined in Clark et al. (2006). Other inclusion and exclusion criteria were identical to the Clark et al. (2006) trial, which provided the main data for the McManus et al. (2009) report. Participants’ mean age was 32 years (SD = 8.82). Forty-nine percent (23) were female.","Here we report the data from Session 3 of CT-SAD. In this session, participants viewed the video recordings of the two social interactions they had engaged in during Session 2. In most instances, the Session 2 interactions involved having a conversation with a stranger. However, if patients felt this would not elicit a significant degree of anxiety, an alternative individualized task was chosen, such as having a conversation with a small group of people. Prior to watching the videos of both interactions in Session 3, patients were asked to form a mental image of the way they thought they would appear and to make clear-cut predictions about how they would come across. The videos were then watched and discussed before patients rerated how they thought they appeared. Details of how predictions were generated with patients and how the video was viewed and discussed are given below in the clinical guidelines section. Self-Perception Before watching each video, patients were asked to rate: (a) How anxious do you think you will look, on a scale ranging from 0 (not anxious at all) to 100 (the most anxious you have ever felt)? ; (b) the extent to which idiosyncratic feared outcomes might be evident (e.g., How boring do you think you will look? [ 0–100]); and (c) How would you rate your performance overall (0–100)? These ratings were repeated after each video had been viewed and discussed. Social Anxiety Participants completed the self-report version of the Liebowitz Social Anxiety Scale (LSAS-SR; Baker, Heinrichs, Kim, & Hofmann, 2002) prior to each treatment session. The LSAS-SR has demonstrated good test–retest reliability and validity (Baker et al., 2002). Internal consistency in our sample was excellent (α = 0.92). Effect of Video Feedback on Self-Perception Table 1 shows participants’ self-perception ratings before and after viewing the videos. Data from the two videos are combined. After viewing the videos, participants rated themselves as looking less anxious than they had predicted, felt that their feared catastrophes occurred to a lesser extent, and rated their overall performance as better than they had anticipated. Table 1 also shows comparable data from McManus et al. (2009). Inspection of effect sizes shows that the beneficial effects of video feedback in the present study were between 38% and 75% larger, depending on the measure used. To estimate the consistency with which video feedback improved patients’ self perception, a composite score was calculated in line with McManus et al. (2009). The composite score was the mean of the ratings of anxious appearance, social fears and overall performance. Internal consistency of the composite score was acceptable (α = 0.75). Ratings of overall performance were reversed so that higher ratings indicated a more negative appearance/performance on all variables. For 98% of patients their ratings after video feedback were less negative and more favorable than before viewing the video. Table 1 shows the composite scores. Impact on Social Anxiety In order to assess the impact of video feedback on participants’ subsequent social anxiety, we compared participants’ scores on the LSAS at the start of the video feedback session with their scores 1 week later. To determine whether any change was more than one might expect from the passage of time alone, we compared change that occurred during this week with change that occurred in two other time periods. The first was the interval between the initial assessment interview and Session 1 (usually 2 weeks). This interval provides an estimate of what happens with no intervention. The second interval was that between Session 1 and Session 2. In Session 1 therapists develop a personal version of the Clark and Wells (1995) model and socialize patients to therapy without trying to change beliefs and behaviors. This interval therefore controls for the effects of therapist attention per se. Table 2 shows LSAS scores at the beginning of the assessment interview; Session 1 (developing an individual formulation of the SAD); Session 2 (self-focus and safety behaviors experiment); Session 3 (video feedback); and Session 4 (1 week after video feedback). A repeated measures ANOVA was used to analyze the data. Mauchly’s test indicated that the assumption of sphericity had been violated, χ2(5) = 28.45, p < .001, therefore degrees of freedom were adjusted using the Greenhouse–Geisser correction. The results show that the LSAS score differed between the sessions, F(2.84, 124.82) = 15.37, p < .001. Pairwise comparisons between sessions 3 and 4 showed that there was a significant reduction in LSAS scores (from 73.4 to 65.5) in the week following the video feedback session (p = .003). By contrast, there were no significant changes in LSAS in the weeks following either the initial assessment (p = 1.00) or developing the individualized cognitive model (Session 1; p = .7).","Our present findings confirm and extend those reported by McManus et al. (2009). Video feedback in Session 3 of a course of CT-SAD was a highly effective intervention for changing negative self-perceptions and was associated with a significant reduction in social anxiety in the following week. The reduction in social anxiety was greater than in two similar time intervals earlier in the course of therapy, suggesting that video feedback has a specific effect, over and above therapist attention, although this result should ideally be confirmed in a between-subjects randomized allocation design. McManus et al. (2009) found that video feedback had remarkably consistent effects with 94% of participants showing an improvement in their self-perception. This consistency was replicated with 98% of patients in the present study stating that they came across to others more favorably than they had predicted, once they had viewed the video. Inspection of effect sizes indicates that the magnitude of the changes observed with video feedback in the present study were even greater than in McManus et al. (2009). As patient selection criteria were similar in the two studies, this suggests that our refinements in setting up, viewing, and discussing video feedback may have further enhanced the potency of the technique. In the next section we summarize the procedures we have so far found helpful for setting up and implementing video feedback, both within the standard video feedback session that occurs early in cognitive therapy and for its more general use throughout the course of treatment.","Our results suggest that video feedback is an excellent technique for helping patients to correct negative self-images and to also gain insight into the way in which their safety behaviors appear to others. There are many opportunities to use it during a course of cognitive therapy, both in the therapy office and also outside of the office when conducting behavioral experiments in the real world. Typically, the first time it is used is in Session 3 when therapist and patient have the opportunity to observe the two social interactions that form part of the self-focused attention and safety behaviors experiment in Session 2. The initial messages that patients learn from this experiment are that self- focused attention and safety behaviors make the problem worse, not better. In particular, they tend to increase anxiety, make it more difficult to focus on the interaction, and give patients enhanced access to internal information (negative images and feelings) that lead them to think that they are coming across to other people more poorly than they really are. The additional messages that patients often get from viewing the videos include: (a) that they come across more favorably than they think in both conditions; (b) some of the aspects of their behavior that they do not like are the unintended, observable consequences of their safety behaviors rather than an intrinsic feature of themselves. When used at other times in therapy, video feedback has a similar function. It also is a very good way of helping patients discover that they are less the subject of other people’s critical attention than they think. When video feedback is used, it is necessary to pay attention to the way in which the video recording is set up, how the patient is prepared in advance of viewing the recording, and how the video is subsequently viewed and discussed. With suitable attention to each of these aspects, it is often possible to overcome the substantial processing biases that have prevented patients from overcoming their negative self-perceptions before they entered therapy. Using the Video Camera As patients with social anxiety can become quite self-conscious about being recorded, it is best to make video recording a routine aspect of therapy, rather than something that is just introduced on an occasional basis for video feedback. We routinely record all therapy sessions using a small domestic video camera that is unobtrusively placed on a bookshelf at right angles to the therapist and patient’s eye line, so it is not in the normal field of view. In the initial assessment interview we explain to patients that we find it helpful to view sessions afterwards in order to reflect on progress and plan future interventions. We request written permission to make the recordings for this purpose. We also encourage patients to take audio recordings of the sessions for them to review afterwards, as we find this is a very good way of maximizing learning. Once permission for the recordings has been obtained, the videos can subsequently also be used for video feedback when therapist and patient together think this might be useful. As well as making the video less intrusive, routinely recording all sessions allows one to capitalize on unplanned therapy events that can be immensely informative. For example, when talking about a topic in the session patients may spontaneously mention that they feel they blushed a lot, had a panic, stuttered, or talked nonsense. In each case they can be asked to specify how they think they looked or sounded, before comparing their prediction with what was captured on the video. A similar process can be applied to the therapist’s behavior. For example, the patient may feel that one must always be perfectly fluent in one’s speech in order to be accepted. Being aware of this belief, the therapist may pause for a while in mid-sentence before carrying on or may start one sentence and then move onto another without completing the first sentence. Chances are that the patient is unlikely to have noticed this and will be surprised to discover that it happened. Reviewing the video with the therapist afterwards helps the patient see that the dysfluency had no real significance, even though the patient would have felt it was a serious social mistake if she had done it herself. One of the main aims of video feedback is to allow patients to see their behavior in context. For this reason, when videotaping an interaction it is best to show both people in the interaction, rather than zooming in on the patient. The latter tends to make it difficult for the patient to avoid feeling self-conscious when subsequently watching the video, and also prevents an appreciation of the true significance of behaviors. For example, patients who are concerned about fidgeting might notice their hands or feet moving in a zoomed-in shot and think this indicates that they are very fidgety. However, in a zoomed-out shot they are likely to see that the other person is also moving to a similar extent. It is important to elicit patients’ main concerns before setting up the video as knowledge about these concerns may have implications for the way in which the video is set up. For example, with patients who are concerned about blushing, it is important to have a color chart or other objects in the field of view that show different shades of red. This is not usually explained to the patient in advance (as it would make them excessively self-conscious). However, if they feel that they do blush during the recording they can subsequently be asked to point to the shade of red that they think matched the blush. Invariably they point to a much darker shade, which is a wonderfully graphic way of helping them discover that their blush is less severe than they feel. Other Participants Taking Part in Social Tasks During a course of CT-SAD patients are likely to have multiple interactions with other people in therapy sessions (conversations with a stranger, a presentation to a small audience etc.). As well as viewing the video of such interactions with their therapist, patients can also benefit from written feedback from the other participants (“confederates”) in the interaction. In order for this feedback to meaningfully reflect everyday life, it is important that the confederates are not informed about the patient’s personal fears (e.g., “I’ll sound stupid”; “I’ll have nothing to say”; “My lip will tremble”) as this would not be information that other people would have in a routine conversation. The confederate is encouraged to treat the patient like anyone else they would meet in life outside of the therapy session rather than someone they are trying to scrutinize. Behavioral experiments conducted outside of the office but in a public space can easily be recorded on a smartphone, tablet, or other domestic camera so that video feedback can be used to enhance the value of these exercises. The principles for setting up and viewing such out-of-the-office videos are essentially the same as those that apply to in-office videos. Identifying Patients’ Predictions in Advance of Viewing Before viewing a video, it is important to identify patients’ predictions of what they think they will see and to get them to visualize what these things look like. This provides them with the maximum opportunity to see the difference between their self-perceptions and reality. Patients are asked to rate (on 0–100 scales) the extent to which they thought their feared catastrophes occurred (“How anxious do you think you looked?” “How boring?” “To what extent to you think you sweated?” etc.) and to indicate how they think that will look. They should be as specific as possible. For example, if somebody says “I will look in a state—just awful,” we would want to elicit a more specific description of how they think they will look on video so that this can be compared to the actual video image. For situations like shaking, blushing, underarm sweating, and dysfluent speech, it is helpful to ask the person to demonstrate on video how they think it looked—for example, by intentionally shaking one’s hand or lips, pointing to the relevant shade of red in a color chart, indicating the size of a sweat patch, or re-creating a pause—so that this can be compared with how it actually looked in the original video recording. Once patients’ predictions have been clearly articulated, it is often helpful to ask them to close their eyes and create their own internal video by visualizing how they think they will appear. It is sometimes also useful to ask people to write short notes on how they think they will appear. Preparing an Unbiased Mode of Viewing It is common for patients to reexperience some of their anxious feelings while watching the video. These feelings may influence their perception of the video. For example, if they feel shaky they may see shaking in the video that would not be seen by others. To overcome this problem, we explain to patients that how one looks and how one feels may not be the same but it is impossible to discover this unless the two are kept separate. In order to do this, patients are asked to view themselves in the video as though they are watching a stranger, only making inferences about how they appear by using what they see and hear on the video, ignoring their feelings. To help them do this, we may encourage them to imagine they are watching a television show. When discussing the video with the therapist, we may ask them to refer to themselves as “that person” or to give themselves a different name. Some patients reexperience feelings from socially traumatic memories (such as being laughed at, ridiculed, or bullied) while watching the video. These feelings may also distort their perception of the video. If the therapist and patient have identified this problem, patients can be asked to specifically look for things in the video that are inconsistent with the past social trauma to help them clearly distinguish between then and now (e.g., focusing on everything about the people they are currently interacting with, which is different from the people involved with the traumatic experience). Some people find that it is very difficult not to turn on their habitual self-critical commentary when watching themselves. After discussing how this can take them away from what actually happens on the video, it can be useful to ask them to watch the video from a more compassionate stance, perhaps as they would if they were watching a close friend or somebody they like and respect. They can be asked to recall the last time they had a conversation with this friend and to consider how they would view their friend: THERAPIST: How do you listen when your friend Alex is talking? Do you just go with the flow of what he says or do you zoom in on how he says every word and ask yourself, “How boring does Alex sound?” “How weird does he look?” [ use patient’s own beliefs] PATIENT: (laughs) No! I just go with the flow. THERAPIST: Ok, we would like you to watch the people on the video in a similar way. Some people find that when they hear the sound of their own voice this automatically triggers the similar sounding self-critical commentary that they normally have in social situations. If the person is highly self-critical, to prevent this mode being activated as soon as the video starts, it can be helpful to watch the first 30 seconds or so with the sound off. This will help them to see that they look as normal as anybody else on the video. When the sound is then turned on, they are in a more appropriate cognitive set. A similar maneuver can be used for people who are highly critical of their physical appearance and find it difficult to look at anything else in the video. For these people, the therapist may initially cover up the patient’s image and let them focus on how other people are responding to them. After a minute or so, they can also be revealed. There is a risk that people may selectively zoom in on themselves looking for any imperfection, rather than watching the interaction in context. To avoid this, the therapist may say something like: “Imagine you walk into a coffee shop and see a conversation happening, look at the whole group, not just one person.” Viewing and Discussing the Video ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Once the patient has clearly articulated their negative self-image and they have been carefully prepared for viewing the video, therapist and patient watch it together. Sometimes the whole interaction is initially viewed without pausing. For highly self- critical patients it can be helpful to pause early in the viewing to check the following: “Are you watching that person (therapist points to the patient in the video) as you would anyone else, or are you watching as your worst critic? Are you looking at other people and how they respond, or just focusing onthat person?” Rewinding the Video to Capture Key Moments Once the video has been watched all the way through, it can be very helpful to rewind the video to look at particular moments that have significance to the patient. For example, rewinding to the moment when the person thought they had a panic attack, when they thought they paused for a long time, looked particularly anxious, looked fidgety or felt that they sweated. They can then be asked to compare how they looked at that moment with their expectation. Other helpful questions may include: “Does the other person seem to have noticed?” “Are they reacting as if they have seen a big mistake?” “Do the people onscreen look markedly different from each other?” “Is the other person also moving their legs and fidgeting?” “If an alien was looking at this would they think one of these people looked really odd?” As the therapist will be aware of the patient’s own self-images and has also seen how they actually behaved, the exact choice of questions will be determined by the therapist’s goal of helping the patient see which aspects of their self-image are distorted. Gaining Insight Into the Impressions Safety Behaviors Convey to Other People Patients are often unaware of the way that their safety behaviors appear to other people. Video feedback provides an ideal opportunity for them to gain insight, which in turn can help motivate them to drop the safety behaviors. For example, a patient who was concerned that his colleagues might see his hand shaking while drinking a beer often turned his back to his colleagues before taking a sip. This made him feel less self-conscious and so seemed a good way of coping with the situation until he saw what it looked like on the video. He then realized it will have appeared odd and may have conveyed to his friends that he was not really interested in them, when the exact opposite was the truth. As turning his back was an intentional strategy, he was able to choose not to do it in the future. Similarly, a patient who was concerned that other people might think she was stupid tended to run through a preprepared list of topics during conversations and to be distracted from the conversation by mentally monitoring how she thought she was coming across. On viewing the video, she realized that she was conveying the impression that she was not interested in other people and was just giving them a lecture. This was the opposite of the impression that she wanted to convey, so she experimented with just saying what came into her head and responding spontaneously to what people said. When she watched this on the video she realized that dropping her safety behaviors allowed her to come across to others in the open and friendly manner that she wished. Comparing Ratings Before and After Viewing the Video A key aspect of video feedback involves comparing patients’ ratings of how they thought they would appear with how they actually appeared once the video has been viewed. This comparison usually involves looking at all the specific predictions that the patient made. The initial 0–100 ratings that patients made in advance of viewing the video are compared with their ratings of the same concerns (looking anxious, sounding boring, etc.) after they have watched and discussed the video. To facilitate comparison a two-column table is constructed. Once the discrepancies have been tabulated, the patient is asked: THERAPIST: What do you notice when we compare these two sets of ratings? PATIENT: I look so much better than I thought I was going to. I look OK, not really that anxious, even though I felt it. THERAPIST: What does that tell you about how visible your anxiety is? PATIENT: Maybe it isn’t that visible to others. Maybe I come across OK. THERAPIST: So if your feelings aren’t that visible, are they a reliable judge of how you come across? PATIENT: No, I guess not. THERAPIST: So next time you are in a social interaction and feel you are coming across badly, you may want to bring to mind the image of how you actually looked on the video. Eliciting Feedback From Other People It can be very helpful to supplement video viewing with feedback to the patient from other people who might have been involved in an interaction. We routinely do this for the self-focused attention and safety behaviors experiment. After they have participated in both conditions, confederates are asked to think back to each interaction in turn and to provide feedback. We find a two-sided form, handed to the confederate as they leave the therapy room, helpful. The first side is largely blank. Confederates are asked to write a few brief notes about their general impressions. What tends to happen is that confederates mention specific points that make it clear they were interested in the patient and noticed various things about them. But what they noticed was mainly what was talked about, not the specific fears that the patient had (such as “my lip was shaking”). The second side covers the patient’s specific predictions (such as “I’ll sound boring”) and the confederate is asked to rate on 0–100 scales the extent to which this was true. Usually the confederates’ ratings are similar to the patients’ ratings after they have viewed the video, but sometimes the confederate is even more positive. When this happens it can be useful to discuss with the patient why the confederate may have been more positive: “Is it possible that your ratings are still partly influenced by your private feelings? This may be information that nobody else could have.” Therapists should use their discretion in deciding whether it is helpful or necessary to supplement video feedback with feedback from others. Most often we present feedback from others after patients have had a chance to view and discuss their video and it is essentially used as a way of further confirming the conclusions they have already reached. However, if the feedback is very positive and the therapist thinks that the patient’s self-criticism may make it difficult for them to view the video objectively, it can be useful to show the other person feedback first. This helps to establish a different cognitive set for viewing the video. Our data (see above) indicate that almost all (98%) patients view themselves more positively after viewing the video if the experience is set up and discussed in the manner presented here. In rare instances where that does not happen, it can be useful to focus the patient’s attention on how others in the video are reacting to them. This can help them see that features that remain prominent in their own mind have less significance to others. Feedback from the confederate is also helpful in such instances. Freezing the Moment of Disconfirmation and Consolidating Learning The aim of video feedback is to help patients see that they come across to others much better than they think. There are some moments in a video that illustrate this point more clearly than others. As video is a moving image, these moments can pass in and out of consciousness quite quickly. An excellent way of overcoming this problem is to capture the moment of disconfirmation as a still image. This can be done either by taking a still from the video or by taking a separate photograph. The latter is particularly useful in out-of-the-office behavioral experiments. For example, a patient reported feeling self-conscious when walking in the street even when not interacting with other people. She felt that people were likely to be hostile and predicted that if she asked someone for directions they would respond in an irritated manner, at the very least. She agreed with her therapist that she would test this out by stopping passersby and asking them the directions for the nearby rail station. The therapist accompanied and discretely took photographs on her mobile phone. The image reproduced in Figure 1 captures the moment when her negative prediction was convincingly disconfirmed: the stranger smiles and is helpful when she asks for directions. She pinned the image to her bulletin board at home to remind herself how people really respond to her. Still images involving other people, like the photograph in Figure 1, can be a helpful way to demonstrate to the patient that their feared concerns (I was boring, I looked sweaty, I looked shaky, I had nothing to say) were not as noticeable to others as they thought, or that even if others did notice, they did not react as negatively as the patient feared. Capturing these moments of disconfirmation can be particularly powerful when the patient feels their feared concern happened naturally (e.g., they forgot what they were saying, they naturally blushed mid-sentence). It can also be helpful for decatastrophizing experiments, where patients purposefully perform a feared concern in order to discover whether others react in the disapproving way they expect (such as adding water to their underarms to create the appearance of sweating, or purposefully trembling their hand when talking to a stranger). This can help patients to realize that they are much less the subject of other people’s critical attention than they initially thought. Still images can also be a wonderful way to capture moments when patients realize they come across much better than their internal feelings and self-perceptions tell them they do (e.g., I look bright red, panicky, have wide scared looking eyes). For example, a patient reported feeling she blushed 80% red while giving a presentation to a small audience during her therapy. Prior to viewing the video, she was asked to select the shade of red she felt she blushed at this moment using a color chart. When a still image was captured from the video at the moment she felt she blushed 80%, the color she selected from the color chart was held next to this image, providing a clear disconfirmation of her belief. A second image was then created: this contained the still taken from the video at the worst moment for the patient (when she felt she blushed red) side-by-side with the block of color she predicted she blushed from the color chart. The image reproduced in Figure 2 illustrates the contrast between the patient’s feelings and the reality. She kept this photo on her mobile phone and looked at it over the week whenever she felt she was blushing, as a reminder that “My feelings are not as visible as I think.” The contrast between two images that depict how patients feel and how they actually looked on video (as demonstrated in Figure 2) may seem stark to an objective observer. However, some patients who reexperience strong anxious feelings and/or find it hard to switch off their habitual self-criticism when viewing images may find it difficult to perceive contrasts that would be apparent to other people who did not experience their feelings or levels of self-criticism. In these instances, we have sometimes found it helpful to edit the image by removing identifiable features of the patient and isolating only the part of the person that they were most concerned about (e.g., showing a portion of their cheek that they felt went bright red; their smile that they felt looked like a grimace; their underarm that they believed was dripping with sweat, etc.). This isolated feature in still image captured from the video can then be compared to a visual calibration obtained before viewing. Removing identifiable features of the patient can help prevent the projection of their feelings and self-criticism into the image, and help them perceive the contrast between their self-perception and reality. Capturing two still frames from the video at different time points can also be another way to illustrate that patients’ internal feelings were not as noticeable as they thought. For example, one patient became extremely self-focused during a conversation with a confederate in therapy. He felt highly anxious at that moment and worried that his face looked odd. He and his therapist were able to isolate the moment on video and to also capture a still image from another part of the conversation when he did not feel particularly anxious, and was predominantly externally focused (see Figure 3). The patient was amazed to discover that he couldn’t see any difference between the two images. This helped him realize that his feelings are largely private. Creating a Still Image Flashcard In order to abstract key principles fitting the cognitive conceptualization of the patient’s social anxiety, and to consolidate and generalize learning, the patient and therapist may add some informative text to still images captured from the video (e.g., “I felt 80% anxious but I don’t look it—my feelings aren’t visible”; “I worried other people would laugh at me but they were friendly, this shows I’m acceptable”). Adding some of the key learning points in the patient’s own words, either handwritten on a printed copy or typed onto an electronic image that can be saved on a smartphone or tablet, can act as a powerful flashcard that the patient can use as a reminder the next time they enter a stressful situation. Rehearsing the Way They Looked on Video Negative self-images are often habitual and well rehearsed. For some people it can be helpful for them to intentionally bring to mind the pictures of how they really appeared on video when they appear anxious so these can counteract their habitual negative self-images. They can also remind themselves of how they look before going into a stressful situation. A Demonstration Video Clip Using the link below, readers can access a 7-minute video clip (Video 1) that illustrates some of the procedures described above. The role-played clip is based on a real cognitive therapy session that took 90 minutes. The clip illustrates: (1) Identifying a patient’s predictions in advance of viewing a video; (2) Preparing the patient to view the video; (3) Discussing the video and rewinding to capture key moments; (4) Comparing ratings before and after viewing the video; and (5) Freezing the moment of disconfirmation and consolidating learning. Summary of the Clinical Guidelines for Video Feedback ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Video feedback is a helpful technique that can be used throughout a course of cognitive therapy to help patients correct their negative self-perceptions and gain insight into the effects of their safety behaviors. Unfortunately, social-anxiety-related processing biases can make it difficult for people with SAD to see the difference between their negative self-image and what is actually shown on the screen. To reliably overcome these biases, particular attention needs to be paid to setting up the video recording; preparing patients to view the video; and subsequently watching and discussing the footage. This clinical guideline has provided details of many of the procedures we have found helpful to maximize the beneficial effects of video feedback. In selecting the most appropriate technique, therapists need to draw on their knowledge of the discrepancy between the patients’ negative self-images and how the patients actually come across. This in turn requires an ability to explore patients’ self-perceptions sensitively and in detail."],["Two studies examined the specificity of effects of explanation on learning by prompting 3- to 6-year-old children to explain a mechanical toy and comparing what they learned about the toy's causal and non-causal properties with children who only observed the toy, both with and without accompanying verbalization. In Study 1, children were experimentally assigned to either explain or observe the mechanical toy. In Study 2, children were classified according to whether the content of their response to an undirected prompt involved explanation. Dependent measures included whether children understood the toy's functional-mechanical relationships, remembered perceptual features of the toy, effectively reconstructed the toy, and (for Study 2) generalized the function of the toy when constructing a new one. Results demonstrate that across age groups, explanation promotes causal learning and generalization but does not improve (and in younger children can even impair) memory for causally irrelevant perceptual details. © 2014 The Authors. --------------------------------------------------------------------------------","A growing literature suggests that young children’s explanations play a crucial role in learning (Bonawitz, van Schijndel, Friel, & Schulz, 2012; Legare, 2012, 2014; Legare & Gelman, 2014; Roy & Chi, 2005; Siegler, 2002; Singer, Golinkoff, & Hirsh-Pasek, 2006; Sobel & Sommerville, 2009; Wellman & Liu, 2007). For example, generating explanations can improve acquisition of new material and its extension to novel cases (Crowley & Siegler, 1999; Lombrozo, 2006; Wellman, 2011) and can even accelerate difficult conceptual transitions such as acquiring an understanding of false beliefs (Amsterlaw & Wellman, 2006) or number conservation (Siegler, 1995). Despite the acknowledged importance of explanation during early childhood, however, little is known about how effects of explanation differ—if at all—from mere verbalization or general attention and whether and how effects of explanation are selective to particular kinds of learning. Here we explored whether prompting 3- to 6-year-old children to explain a mechanical device fosters causal–mechanical understanding more effectively than does observation or verbalization and, if so, whether such understanding comes at the expense of other kinds of learning. Educational research comparing self-explanation—that is, explaining to oneself or another person—with other activities suggests that explaining can be more effective for learning than alternative activities such as thinking aloud and reading study materials twice, especially when it comes to generalizing from study material to new cases (see Fonseca & Chi, 2011, and Lombrozo, 2012, for reviews). Although most research on self-explanation has focused on older children and adults, the limited research with younger children suggests similar effects. For example, research on problem solving among elementary school children comparing the effectiveness of self-explanation with alternative activities (e.g., solving practice problems) found that self-explanation was associated with greater conceptual and procedural knowledge (McEldoon, Durkin, & Rittle-Johnson, 2012). Rittle- Johnson, Saylor, and Swygert (2008) also demonstrated that explanation prompts facilitate transfer in children as young as 5 years relative to repeating problem solutions in problem-solving tasks. Notably, however, children in these previous studies of self- explanation were asked to explain why a particular solution or strategy was correct; that is, they explained some task-relevant feedback (see also Amsterlaw & Wellman, 2006; Crowley & Siegler, 1999; Rittle-Johnson et al., 2008). It could be that effects of explanation in young children result from the interplay of feedback with their explanations. For example, explaining could draw attention to the feedback and encourage children to rephrase it in their own words, thereby facilitating belief revision. Thus, it is an open question whether simply explaining one’s observations—in the absence of feedback of this type—can similarly improve young children’s learning and, if so, whether its impact results from the general use of language or from explanation per se. This question is especially important in understanding the role of children’s spontaneous explanations on learning throughout development (Legare, 2014). Another open question concerns the selectivity of explanation’s effects, especially during early childhood. In particular, are effects of explanation restricted to some kinds of learning, or do they extend more broadly? And do the benefits of explanation have any associated costs? Evidence from older children and adults suggests that effects of explanation can indeed be selective, improving some kinds of learning over others. For example, explanation can foster analogical transfer at the expense of memory for previous problems (Needham & Begg, 1991), privilege causal mechanisms over consistency with previous data in justifying causal judgments (Berthold, Roder, Knorzer, Kessler, & Renkl, 2011; Kuhn & Katz, 2009), and encourage learning about patterns instead of individual examples (Williams, Lombrozo, & Rehder, 2013). When it comes to effects of explanation during early childhood, however, comparable studies have not been performed and two distinct stories are quite plausible. On the one hand, explaining could boost general engagement or attention (e.g., Siegler, 2002), which might lead to relatively widespread benefits across wide-ranging measures of learning. On the other hand, consistent with research on older children and adults, explaining could privilege some kinds of learning (e.g., causal learning) at the expense of others (e.g., perceptual learning), leading to more selective effects and potentially even to impairments (see also Walker, Lombrozo, Gopnik, & Legare, 2014). Identifying whether and how effects of explanation are selective, therefore, is of both practical value (for informing educational practice) and of theoretical value (for helping to isolate the mechanisms by which explanation influences learning during early childhood). Building on prior work, we propose that explanation generates selective effects and that it does so by encouraging young learners to consider particular kinds of hypotheses, namely those that support good explanations (Legare, 2012; Lombrozo, 2012; Williams & Lombrozo, 2013). If explanations are typically judged as better when they invoke causal mechanisms or broad generalizations (for reviews, see Keil, 2006; Lombrozo, 2006, 2012), then explaining could selectively focus learners’ attention on hypotheses that exhibit these features. Consistent with these ideas, adults seek information about mechanisms when asked to explain an event (Ahn, Kalish, Medin, & Gelman, 1995; Keil, 2006; see also Bullock, Gelman, & Baillargeon, 1982; Koslowski, 1996; Lombrozo, 2010), and prompting adults to explain can foster the discovery and generalization of broad patterns (Williams & Lombrozo, 2010, 2013). In fact, a prompt to explain can lead even 3-year-olds to posit unobserved causes (Buchanan & Sobel, 2011; Legare, Gelman, & Wellman, 2010; Legare, Wellman, & Gelman, 2009) and favor generalizations that account for more observations (Walker, Williams, Lombrozo, & Gopnik, 2012), suggesting that prompts to explain might direct even young learners toward broad, causal patterns. In the two experiments that follow, we provide the first empirical investigation of the selectivity of effects of explanation on young children’s causal learning. Our studies differ importantly from prior work in that we did not provide children with direct instruction or corrective feedback. Moreover, we compared explanation with other tasks that were matched for time and verbalization, and we examined both causal and non-causal learning—learning about functional–mechanical properties versus memory for perceptual features such as color. Our first objective was to differentiate effects of explanation from those due to observation (Study 1) or verbalization (Study 2). Our second objective was to examine whether and how prompts to explain selectively benefit young children’s causal learning and whether this benefit comes at the expense of non-causal learning (Studies 1 and 2). Our final objective was to examine potential age-related differences in the effects of explanation on children’s learning (Studies 1 and 2). In particular, we examined whether learning benefits are a consequence of children’s age or instead follow from the content and quality of their explanations regardless of age. In two studies, we presented preschool- aged children with a novel mechanical toy with visible interlocking gears and examined learning using measures that assessed their understanding of the toy’s causal function, memory for the toy’s non-causal properties such as color, and (in Study 2) generalization of the toy’s causal function in constructing a novel toy. In Study 1, children were prompted to observe or explain the toy, and learning was assessed both as a function of the experimental manipulation and (for children in the explanation condition) based on the content of their verbal response (i.e., explanation vs. non-explanation). In Study 2, children provided verbal responses that were categorized as either explanations or non- explanations, and learning was assessed as a function of these categories. In the context of these studies, we defined explanations as responses that included functional or mechanistic information about the toy (i.e., about proximate causal processes or the function of a particular part or process) (see Kelemen, 1999; Legare et al., 2010; Lombrozo, 2009; Lombrozo & Carey, 2006). Comparing effects of explanation with observation and non-explanatory verbalization can shed light on the mechanisms by which each process contributes to learning. If explanation’s effects derive from greater attention or engagement, we might expect global improvements in learning relative to observation. If explanation’s effects are a consequence of producing a verbal response, we might expect a benefit for explanation relative to observation in Study 1 and comparable effects for explanation and other verbalizations in Study 2. In contrast, we predicted that, relative to both observation and verbalization, explanation would improve learning on measures of functional–mechanical understanding (Studies 1 and 2) and generalization (Study 2) but would have no effect on—or even impair—learning about non-causal properties such as color. Finally, our two studies also had the potential to differentiate several plausible trajectories for age-related changes in the effects of explanation on learning. One possibility was that explanation would help younger children more than older children because younger children may require greater scaffolding to achieve cognitive benefits that older children can achieve without prompting (for evidence that this is sometimes the case, see Walker et al., 2014). Another possibility was that explanation would help older children more than younger children because learning by explaining presumably requires some baseline level of verbal fluency and cognitive sophistication (Hickling & Wellman, 2001). A third possibility was that older children would be more likely than younger children to engage in (high-quality) explanation (Legare et al., 2010) but that effects of explanation among those who do explain would be relatively uniform across ages. Finally, effects of explanation could vary across development in their selectivity; for example, it could be that explanation privileges causal learning at the expense of perceptual memory in younger children but promotes causal learning without this penalty in older children. By investigating effects of explanation on learning across the 3- to 6-year age range—which spans a wide range of verbal, conceptual, and cognitive development—our studies could also help to differentiate these alternative developmental trajectories.","Study 1 examined the selectivity of explanation’s effects on children’s learning about a causal system. We presented young children with a novel mechanical toy and compared learning across conditions in which they were or were not prompted to explain (explanation vs. observation). We measured two kinds of learning: learning about causal–mechanical relationships and learning about causally irrelevant properties such as color. We hypothesized that explanation would direct learners’ attention toward causal mechanisms, resulting in better learning on measures of causal knowledge but not necessarily on measures of memory for causally irrelevant properties.","A sample of 95 children—31 3-year-olds (M = 41.83 months, SD = 3.40), 32 4-year- olds (M = 53.56 months, SD = 3.14), and 32 5-year-olds (M = 65.34 months, SD = 3.34)—participated in Study 1. Children were recruited from preschools in a major metropolitan area in the American Southwest. The sample was approximately gender balanced and primarily Euro-American and middle class, with an approximately equal number of children from each age group assigned to each condition. Children were tested in a quiet room in the preschool or a research laboratory at a major southwestern university; each session took approximately 10 to 15 min. An additional 5 children participated but were dropped from the final sample due to either inability to engage with the task (n = 2) or experimenter error (n = 3). None of the participants from Study 1 participated in Study 2.","Materials included a novel machine with five interlocking gears, including a crank that made a fan turn and non-functional peripheral parts connected to selected gears (see Fig. 1A). Each child saw the same machine. An additional three parts were used in a training trial, and an additional 10 parts were used to assess learning (see “Procedure” section below). Training task Each child participated in a training task where an experimenter demonstrated how the machine parts fit together. The child was presented with a base part, a gear, and a peripheral part (called a “topper”). The parts were similar, but not identical, to those used for the machine in the experimental task. The experimenter showed the child each part, labeling it as a base piece, gear, or topper, and then modeled how to put them together. Then the parts were taken apart, and the child was given the opportunity to put them together. Experimental task Following the training task, the experimenter placed the previously hidden machine in front of the child, pointed out the crank and fan, and turned the crank to demonstrate that this made the fan turn. In the observation condition, the child was told, “Let’s look at this!”, and then had 40 s to observe the machine (which was no longer in motion). In the explanation condition, the child was asked, “Can you tell me how this works?”, and then had 40 s to produce a verbal response.1 Learning measures After 40 s, the machine was removed from the child’s view and placed under the table. Unbeknownst to the child, the small pink gear was removed from the machine and the machine was placed on the table again. The experimenter indicated that this was the same machine as before but that one of the parts was missing. The child then participated in two learning tasks, with order counterbalanced across participants. For the causal choice task, five parts were presented, none of which was identical to the missing part. The choices consisted of one part of the correct size and shape but different color, one part of the correct shape but incorrect size, one part of the correct size but incorrect shape, and also a distracter part and a peripheral part seen before but not the correct shape (see Fig. 1B). The child was asked, “Can you point to which one of these parts you think will make it work?” This task was intended to assess children’s understanding of the causal contributions of the gears in making the machine work—an aspect of functional–mechanical understanding. For the color choice task, the experimenter presented the child with another five parts. All parts were the correct size and shape, but only one part was the same color as the original (see Fig. 1C). The child was asked, “Can you point to the piece that will make my machine look like it did in the beginning?” This task was intended to assess children’s memory for the color of the original gear, a property irrelevant to functional–mechanical understanding. After the choice tasks, the machine was again removed from the child’s view under the table. The experimenter removed all of the parts from the base and took the peripheral parts off the medium-sized purple gear, the large orange gear, and the small pink gear. The crank gear and the fan gear were not taken apart. The experimenter put the base on the table in front of the child as before. Then the experimenter put the gears and peripheral parts in front of the child in a predetermined random order (Fig. 1D) and asked, “Can you put my machine back together the way it was before and make it work?” The child was given 10 min to reconstruct the machine. Experimental task For children in the explanation condition, verbal responses were coded for the presence of explanations. If the child provided a mechanistic or functional explanation (i.e., made a claim about proximate causal processes or the goal or function of a particular part or process; see Kelemen, 1999; Lombrozo, 2009; Lombrozo & Carey, 2006), the verbal response was coded as an explanation (e.g., “That [the crank] spins around and the others spin around,” “It spins together”). All other verbal responses were coded as non-explanations. Non- explanations included descriptions, inquiries, statements of uncertainty, and no responses. Verbal responses were coded by two independent coders with agreement of 94% (kappa = .83). Disagreements were resolved by discussion. Learning tasks For the causal choice and color choice tasks, selecting the correct part was scored as 1 and selecting an incorrect part was scored as 0. For the reconstruction task, success was analyzed along two dimensions: successfully reconstructing the machine’s function (causal reconstruction score) and successfully reconstructing causally irrelevant components (topper reconstruction score). The child received a causal reconstruction score of 1 for successfully reconstructing the causal–functional relationships that would allow the machine to work (i.e., replacing the fan and handle correctly, replacing the gears correctly, interconnecting the gears, and spinning the handle during reconstruction) or otherwise received a score of 0. The topper reconstruction score reflected accuracy in pairing each gear with its non- functional peripheral part (1 = all correct, 0 = any incorrect). Because the causal choice task and the causal reconstruction score both reflect causal learning, these scores were combined into a single causal learning score that could range from 0 (success on neither task) to 2 (success on both tasks). Similarly, because the color choice task and the topper construction score both reflect learning about non-causal aspects of the machine, they were combined into a single non-causal learning score that could range from 0 to 2. Learning as a function of experimental condition Table 1 reports the proportions of correct responses as a function of condition (observation vs. explanation) for the four learning measures (causal choice, color choice, causal reconstruction, and topper reconstruction). To analyze performance, we conducted a repeated-measures analysis of variance (ANOVA) with type of learning score (causal learning or non-causal learning) as a within-participants factor and condition (observation or explanation) and age group (3-, 4-, or 5-year-olds) as between-participants factors. This analysis revealed a significant interaction between type of learning and condition, F(1, 89) = 38.30, ηp2 = .30, p < .001, as well as an interaction among learning score, condition, and age group, F(2, 89) = 7.48, ηp2 = .14, p < .001. Therefore, we analyzed performance for each learning score separately. Analyzing causal learning scores revealed that participants in the explanation condition (M = 1.04, SD = .65) performed significantly better than those in the observation condition (M = .49, SD = .66), F(1, 89) = 18.74, ηp2 = .17, p < .001, consistent with our predictions (Fig. 2). In addition, there was a main effect of age, F(2, 89) = 4.99, ηp2 = .10, p = .009; post hoc tests revealed that 3-year-olds (M = .48, SD = .57) performed significantly worse than 4-year olds (M = .88, SD = .71), t(61) = −2.41, p = .019, and 5-year-olds (M = .94, SD = .76), t(61) = −2.68, p = .010, who did not differ from each other, p = .734. Analyzing non-causal learning scores revealed that participants in the explanation condition (M = .46, SD = .62) performed significantly worse than those in the observation condition (M = .89, SD = .73), F(1, 89) = 11.15, ηp2 = .11, p = .001 (Fig. 2). In addition, there was a significant interaction between condition and age, F(2, 89) = 6.50, ηp2 = .13, p = .002. A prompt to explain led to significant impairments on the non-causal learning measures for 3-year-olds, t(29) = 2.29, p = .030, and 4-year-olds, t(30) = 4.34, p < .001, but not for 5-year-olds, t(30) = −0.86, p = .397. Verbal responses for the explanation condition The coding of verbal responses from children in the explanation condition resulted in 38 children designated as explainers (10 3-year-olds, 14 4-year-olds, and 14 5-year-olds) and 10 children designated as non-explainers (6 3-year-olds, 2 4-year-olds, and 2 5-year-olds). Although there was a trend for older children to produce more responses coded as explanations than younger children, the distribution of designations did not differ significantly across age groups, p = .17. We analyzed the relationship between children’s verbal response designation (explainer or non-explainer) and their performance on the causal and non-causal learning tasks with a repeated-measures ANOVA2 using type of verbal response as a between-participants factor (explainer or non-explainer), age group as a between- participants factor (3-, 4-, or 5-year-olds), type of learning as a within- participants factor (causal learning or non-causal learning), and learning score as the dependent measure. This analysis revealed the predicted interaction between designation and type of learning, F(1, 46) = 19.90, ηp2 = .32, p < .001. On measures of causal learning, children who explained (M = 1.21, SD = .58) outperformed those who did not (M = .40, SD = .52), t(46) = −4.03, p < .001. But on measures of non-causal learning, children who explained (M = .42, SD = .55) performed no better than those who did not (M = .60, SD = .84), t(46) = 0.81, p = .42. There was no reliable main effect of age, F(2, 44) = 1.59, p = .21, and a marginal main effect of designation, F(1, 44) = 3.18, p = .08.","The primary objective in Study 1 was to investigate the selectivity of explanation’s effects on young children’s learning. The data suggest that the benefits of explanation are indeed selective; although children prompted to explain performed better on measures of causal learning, they performed significantly worse when it came to non-causal learning. This suggests that effects of explanations did not result from an indiscriminate increase in attention or engagement but rather were a consequence of the specific processes invoked through explanation. We also found age-related improvement in performance; the 4- and 5-year-olds outperformed the 3-year-olds, and whereas 3- and 4-year-olds’ explanation-prompted improvements in causal learning were accompanied by a penalty for non-causal learning, this was not observed among 5-year-olds. Notably, although the majority of responses in the explanation condition included explanations, a substantial minority of responses did not. We found that the content of children’s verbal responses was predictive of their performance on causal learning measures; children who provided an explanation had higher causal learning scores than children who did not provide an explanation. Thus, it is possible that the content of children’s responses to an explanation prompt—and in particular whether they succeed in producing an explanation—predicts children’s causal learning performance beyond mere verbalization or description. Study 2 sought to examine how explicit prompts to explain versus describe influence the extent to which children generate explanations and how their explanations relate to learning. Doing so also provided a natural control for effects of verbalization because all children were prompted to provide a verbal response.","In Study 1, the explanation condition required a verbal response, whereas the observation condition did not. Moreover, participants in the explanation condition were instructed to explain with a potentially leading prompt—to tell the experimenter how the machine works. Thus, it is unclear whether children’s explanations in the absence of a prompt about how the machine works would similarly direct learners to causal understanding. Study 2 addressed these concerns by requiring verbalization in response to a very undirected prompt in all conditions and by once more analyzing performance as a function of the content of children’s responses. In Study 2, children received a general prompt to either explain the machine or to describe the machine. By using these undirected prompts, we hoped to experimentally alter the proportion of children providing explanations while also generating enough variability in the content of children’s responses within each condition to permit a comparison of children who did and did not explain regardless of experimental prompt. Thus, the experiment involved two conditions—describe prompt and explain prompt—with all verbal responses also coded for the presence of explanations. We predicted that the content of children’s verbal responses (i.e., generating an explanation) would predict performance on the learning measures; children who provided an explanation would perform better on causal learning measures than children who provided alternative responses and potentially mirror the impairments for non-causal learning observed in Study 1. Study 2 also included an additional measure—generalization from one mechanical device to another. At the end of the task, participants were invited to create a new device. Given the close relationship between explanation and generalization (e.g., Lombrozo, 2012; Lombrozo & Carey, 2006), we anticipated that explanation would increase the extent to which children generalized the functional properties of the first machine to the second.","A sample of 87 children (23 3-year-olds, 19 4-year-olds, 24 5-year-olds, and 21 6-year-olds) participated. Children were recruited from preschools in a major metropolitan area in the American Southwest. The sample was approximately gender balanced and primarily Euro-American and middle class, with an equal number of children from each age group assigned to each condition. Although Study 1 revealed main effects of age on performance, the single interaction between age and effects of explanation suggested a break between the 3- and 4-year-olds, for whom explanation led to impairments on non-causal learning, and the 5-year-olds, for whom it did not. Given our otherwise small sample sizes, therefore, we grouped children into younger (3–4 years) and older (5–6 years) age groups for analysis in Study 2: 42 3- and 4-year-olds (M = 45.9 months, SD = 6.73) and 45 5- and 6-year- olds (M = 70.48 months, SD = 7.14). Children were tested in a quiet room in their preschool or a research laboratory at a major southwestern university; each session took 10 to 15 min. An additional 5 children participated but were dropped from the final sample due to either inability to engage with the task (n = 2) or experimenter error (n = 3).","In addition to the materials from Study 1, there were 18 parts used during a generalization task (detailed below). Training task The training procedure was identical to that in Study 1. Experimental task Following training, the experimenter placed the previously hidden machine in front of the child. In all conditions, the experimenter pointed out the crank and fan and turned the crank to demonstrate that this made the fan turn. Then the child observed the machine (which was no longer in motion) for 40 s in one of two conditions. In the explain prompt condition, the child was prompted with, “Explain the machine to me,” which was followed up with “Can you explain anything else?” if the child stopped before the 40-s period ended. In the describe prompt condition, the child was told, “Describe the machine to me,” with “Can you tell me anything else?” as a follow-up prompt. In both conditions, the child’s responses were recorded on video. Learning tasks The learning tasks were identical to those in Study 1. Generalization task In the generalization task, the child was told, “Okay! Now I have new things to show you. Now it is your turn to make a machine.” The child was provided with a new base made up of six interlocking base parts, six new gears, and six new peripheral parts. The base was pre-constructed, and the peripheral parts and gears were laid out next to it in the same order for each child (see Fig. 1E). The child was given 3 min to create his or her own machine. Experimental task Children’s verbal responses in the explain prompt and describe prompt conditions were coded for the presence of explanations as in Study 1. If children provided a mechanistic or functional explanation (i.e., made a claim about proximate causal processes or the goal or function of a particular part or process; see Kelemen, 1999; Lombrozo, 2009; Lombrozo & Carey, 2006), they were coded as explainers (e.g., “You just have to move it [the handle] around and then the fan turns around and the bottom things [points to middle gears] turn around too, and that top [peripheral part] spins too”). Children who did not make such statements were coded as non-explainers. Verbal protocols were coded by two independent coders with agreement of 91% (kappa = .86). Disagreements were resolved by discussion. Learning measures and reconstruction task The choice and reconstruction tasks were coded as in Study 1. Generalization task Children’s behavior during the generalization task was coded for whether gears were placed on the base, whether gears interlocked, whether the handle and the fan were placed on gears on the base, and whether the gears with the handle and fan interlocked. Meeting all of these criteria successfully recreated the functional–mechanical basis of the original machine and corresponded to the causal generalization score, with success coded as 1 and failure coded as 0.","Table 2 reports the proportions of correct responses as a function of children’s verbal responses for the five key dependent variables (causal choice, color choice, causal reconstruction, topper reconstruction, and causal generalization) for younger and older children. Table 3 reports the proportions of correct responses as a function of children’s verbal responses for the five key dependent variables (causal choice, color choice, causal reconstruction, topper reconstruction, and causal generalization) by age in years (3-, 4-, 5-, and 6-year-olds). Fig. 3 reports the average causal and non-causal learning scores, as detailed below. Verbal responses The verbal response coding resulted in 38 children designated as explainers (12 younger and 26 older) and 49 children designated as non-explainers (30 younger and 19 older). Children were equally likely to be designated as explainers versus non- explainers across the explain prompt and describe prompt conditions (17 of 41 vs. 21 of 46), χ2(1, N = 87) = 0.1546, p = .694; that is, the prompts were equally effective at generating explanations, potentially because children failed to differentiate between a request to explain and a request to describe. Moreover, experimental condition (describe prompt or explain prompt) was not related to causal learning score, non-causal learning score, or generalization. In subsequent analyses, therefore, we report performance as a function of designation (non- explainer or explainer) while also taking into account that older children were more likely to be designated as explainers than younger children, χ2(1, N = 87) = 7.53, p = .006. Learning measures as a function of coded designation A repeated-measures ANOVA with designation as a between-participants factor (explainer or non-explainer), age group as a between-participants factor (younger or older), and learning score as a within-participants factor (causal learning score or non-causal learning score) revealed the predicted interaction between designation and learning score, F(1, 83) = 26.627, ηp2 = .243, p < .001. On measures of causal learning, children who explained (M = 1.71, SD = .46) outperformed those who did not (M = .80, SD = .74), t(85) = 6.71, p < .001 (Fig. 3). But on measures of non-causal learning, children who explained (M = .79, SD = .66) performed no better than those who did not (M = .96, SD = .68), t(85) = −1.17, p = .25 (Fig. 3). There was also a main effect of learning measure, F(1, 83) = 15.762, ηp2 = .16, p < .001, with higher scores for causal learning (M = 1.20, SD = .78) than non-causal learning (M = .89, SD = .67); a main effect of designation, F(1, 83) = 6.020, ηp2 = .068, p = .016, with higher scores for explainers (M = 1.25, SD = .45) than non-explainers (M = .88, SD = .53); and a main effect of age group, F(1, 83) = 11.137, ηp2 = .118, p = .001, with older children (M = 1.24, SD = .46) outperforming younger children (M = .82, SD = .50). Because older children were more likely than younger children to be designated as explainers, the preceding analysis potentially confounds designation with age. Therefore, we repeated the analysis within each age group to ensure that age did not drive the critical interaction between designation and learning measure. For the younger children, a repeated-measures ANOVA with designation (explainer or non-explainer) as a between-participants factor and learning measure (causal learning or non-causal learning) as a within-participants factor again revealed the predicted interaction between designation and learning measure, F(1, 40) = 13.915, ηp2 = .258, p = .001, as did the equivalent analysis for older children, F(1, 43) = 12.686, ηp2 = .228, p = .001. For older children, there was also a significant main effect of learning measure, F(1, 43) = 19.358, ηp2 = .310, p < .001, with higher scores for causal learning (M = 1.56, SD = .66) than non-causal learning (M = .93, SD = .65), as well as a significant main effect of designation, F(1, 43) = 6.422, ηp2 = .130, p = .015, with higher scores for explainers (M = 1.38, SD = .36) than non-explainers (M = 1.05, SD = .52). Generalization as a function of designation When analyzing generalization performance as a function of designation, explainers performed significantly better (32 of 38 succeeding) than non-explainers (27 of 49 succeeding), χ2(1, N = 87) = 8.31, p = .004. To test for differences across groups while accounting for the uneven distribution of ages across designations, we performed a logistic regression on generalization performance with age group entered in a first step and designation entered in a second step. This analysis revealed a significant effect of designation, p = .023 (β = −1.25, SE = .55), as well as a marginal effect of age group, with older children (36 of 45) more likely to succeed than younger children (23 of 42), p = .072 (β = .917, SE = .51).","Our objective in Study 2 was to replicate and extend key findings from Study 1. First, we succeeded in finding reliable effects of explanation when comparing the content of children’s responses rather than the experimental prompt that they received. As in Study 1, children who explained outperformed non-explainers on measures of causal learning but not on measures of non-causal learning. This result is important in establishing that effects of explanation do not derive solely from the use of language given that all children produced verbal responses. That effects of explanation were not eliminated when compared with alternative kinds of verbalization is especially striking in the context of our two studies given that the non-causal properties that we tested (e.g., color) were, if anything, easier to express linguistically than the causal properties (e.g., gear shape). The verbal prompt conditions (explain prompt and describe prompt) were equally effective at prompting children to produce explanations. Thus, our analyses were based on the content of children’s responses, and the effects of explanation in Study 2 are correlational. Nonetheless, our data demonstrate that children’s explanatory responses are predictive of learning. The effects of the content of children’s responses in Study 2 nicely complement the effects of prompted explanation in Study 1, as children’s responses were elicited by general prompts to explain or describe (as opposed to the explicit questions about how the machine works used in Study 1). Across both studies, therefore, we have evidence for a causal relationship between explanation and learning (Study 1) that is mirrored in the content of children’s responses (Studies 1 and 2).","Our studies had three objectives: (a) to differentiate effects of explanation from those due to observation (Study 1) or verbalization (Study 2), (b) to examine whether and how prompts to explain selectively benefit young children’s causal learning and may impair non-causal learning (Studies 1 and 2), and (c) to examine potential age-related differences in the effects of explanation on children’s learning (Studies 1 and 2). In Study 1, children prompted to explain outperformed others on causal learning but not on memory for causally irrelevant details. In both Studies 1 and 2, children who provided an explanatory verbal response outperformed those who provided other verbal responses on measures of causal learning, but again this benefit did not extend to memory for causally irrelevant details. Notably, when it came to generalizing a function from one toy to another, explainers outperformed non-explainers. Thus, our data suggest that effects of explanation are distinct from those of observation or other kinds of verbalization and result in selective rather than general benefits for learning, even leading to impairments for younger children. In our task, explaining directed learners to functional–mechanical understanding but not to causally irrelevant details, a trend that was observed across age groups. Why might explanation exert these selective effects? One possibility is that explanation is simply a constructive activity (Chi, 2009) or goal-directed activity (Nelson, 1973). We propose instead that although explanation likely recruits many cognitive processes that are shared with other kinds of constructive and goal-directed activities, explanation is especially conducive to causal learning and generalization. In the Introduction, we presented the hypothesis that explaining encourages learners to focus on causal mechanisms and generalization (Legare, 2012; Lombrozo, 2012; Williams & Lombrozo, 2013) precisely because invoking mechanisms and broad generalizations is characteristic of good explanations. Whereas other constructive and goal-directed activities may sometimes share these characteristics incidentally, they can also direct learners in a variety of alternative ways—depending on the nature of the task—and have consequences that are more or less diffuse. In the Introduction, we also identified several ways in which effects of explanation could change during the course of early development. We found several age-related improvements in general performance, yet we observed only two developmental changes involving explanation. First, in Study 1, 3- and 4-year olds who were prompted to explain exhibited impairments when it came to non-causal learning, whereas 5-year-olds did not. This suggests that the beneficial effects of explanation do have an associated cost, but one that may be realized only on more difficult tasks or when cognitive resources are taxed. Second, in Study 2 (and to a lesser extent in Study 1), we found evidence that older children were more likely than younger children to generate verbal responses that were designated as explanations. Importantly, however, the beneficial effects of explanation did not interact with age. We suggest that older children are more likely to engage in explanation, both spontaneously and in response to prompts, but that effects of explanation are predominantly a function of the content and quality of the explanation rather than children’s age. Understanding the ways in which explanation does—and does not—improve learning speaks not only to questions about the development of causal knowledge but also to questions about how to most effectively harness explanation for use in educational interventions. For example, research with young children has shown that explanation benefits learning by increasing the efficiency of strategy use whether the children generated the explanation or learned the explanation from the experimenter (Crowley & Siegler, 1999). There is also evidence that generating both self-explanations and explanations to others (i.e., caretakers) improves problem- solving accuracy at posttest, with the greatest benefits for problem-solving transfer from explanations to caretakers (Rittle-Johnson et al., 2008). Our findings suggest that these effects may derive in part from the role of explanation in directing children toward causal regularities (in these cases involving reasoning and the application of problem- solving strategies) and that explanation can be less beneficial—and perhaps even harmful—when the target of learning is more perceptual or descriptive. In educational research with adolescents and college students, learners who explain meaning in text either spontaneously or when prompted to do so understand more from the text and construct better mental models of the content (e.g., Chi, Bassok, Lewis, Reimann, & Glaser, 1989; Chi, DeLeeuw, Chiu, & LaVancher, 1994; Magliano, Trabasso, & Graesser, 1999; Trabasso & Magliano, 1996). However, there are large individual differences in spontaneous engagement in self-explanation and in the quality of students’ self-explanations. (e.g., Chi et al., 1989, 1994). Whereas a high-quality self-explanation would indicate an understanding of the meaning of the text, a low-quality self-explanation may involve description or restatement of the text. Across studies, we find similar variation in young children’s verbal responses to both explicit and more general explanation prompts and demonstrate that the quality of verbal responses (i.e., presence of causal explanation) predicts learning outcomes. Importantly, there is evidence that students can be trained to more effectively self-explain text and that the use of high-quality self-explanation strategies predict improved problem-solving performance (Bielaczyc, Pirolli, & Brown, 1995). For example, McNamara, O’Reilly, Rowe, Boonthum, and Levinstein (2007) demonstrated that self- explanation training improves reading comprehension to a greater extent than control tasks (i.e., reading out loud), and there is also evidence that self-explanation training increases engagement in elaboration, prediction, comprehension monitoring, the use of logic, and overall reading comprehension, particularly for students with lower prior content knowledge and for more difficult texts (McNamara, 2004; Ozuru, Briner, Best, & McNamara, 2010). Taken together, these studies provide evidence that self-explanation can be used productively as a learning strategy and that educational interventions can improve the quality of explanations. Our research provides evidence that analogous training methods may improve the quality of self-explanations in preschool-aged children. Our studies also suggest that the precise content of the explanation prompt can have an important influence on the content and quality of responses; although most children who were prompted to explain how the machine works in Study 1 produced a response coded as explanatory (38 of 48 children), a smaller proportion did so in Study 2 in response to the undirected prompt to explain the machine (17 of 41 children). Finally, our findings that explanation can improve causal learning and foster generalization also bear on claims about the nature of development. A vast literature on development has challenged Piagetian claims that young children are limited to concrete appearances (Piaget, 1930), demonstrating instead that they recognize abstract relationships (e.g., Gopnik & Schulz, 2007) and unobserved properties (e.g., Gelman, 2003; Wellman & Gelman, 1992). Nonetheless, little is known about the mechanisms by which children succeed in going beyond appearances, as understanding causal function often requires (Sobel, Yoachim, Gopnik, Meltzoff, & Blumenthal, 2007). We propose that explanation may be an especially powerful tool for causal learning during early childhood precisely because its effects are relatively selective, orienting young learners toward causal mechanisms and promoting generalization."],["Mimicry, the spontaneous copying of others’ behaviors, plays an important role in social affiliation, with adults selectively mimicking in-group members over out-group members. Despite infants’ early documented sensitivity to cues to group membership, previous work suggests that it is not until 4 years of age that spontaneous mimicry is modulated by group status. Here we demonstrate that mimicry is sensitive to cues to group membership at a much earlier age if the cues presented are more relevant to infants. 11-month-old infants observed videos of facial actions (e.g., mouth opening, eyebrow raising) performed by models who either spoke the infants’ native language or an unfamiliar foreign language while we measured activation of the infants’ mouth and eyebrow muscle regions using electromyography to obtain an index of mimicry. We simultaneously used functional near-infrared spectroscopy to investigate the neural mechanisms underlying differential mimicry responses. We found that infants showed greater facial mimicry of the native speaker compared to the foreign speaker and that the left temporal parietal cortex was activated more strongly during the observation of facial actions performed by the native speaker compared to the foreign speaker. Although the exact mechanisms underlying this selective mimicry response will need to be investigated in future research, these findings provide the first demonstration of the modulation of facial mimicry by cues to group status in preverbal infants and suggest that the foundations for the role that mimicry plays in facilitating social bonds seem to be present during the first year of life. --------------------------------------------------------------------------------","It is a common feeling: While talking to a friend or colleague, you suddenly realize that you are copying her behavior or accent. This tendency to spontaneously and unconsciously copy or “mimic” others’ behaviors has been suggested to play an important role in social interactions. For example, it contributes to the development of liking and rapport between strangers and makes social interactions more smooth and enjoyable (van Baaren, Janssen, Chartrand, & Dijksterhuis, 2009). It has been suggested that mimicry can be used as a strategy for social affiliation (Wang & Hamilton, 2012), and indeed studies have demonstrated that adults increase mimicry toward people they like and in-group members, while mimicry of out-group members is inhibited (for reviews, see Chartrand & Lakin, 2013; Hess & Fischer, 2017; van Baaren et al., 2009). Despite the important social functions that mimicry is hypothesized to serve, surprisingly little is known about the development of this phenomenon. While reports of neonatal imitation of facial actions (e.g., Meltzoff & Moore, 1983) have been subject to much criticism and doubt (e.g., Jones, 2009; Oostenbroek et al., 2016), recent studies that used a more objective measure of mimicry (i.e., electromyography [EMG]) have demonstrated that infants exhibit mimicry of emotional and nonemotional facial actions from at least 4 months of age (de Klerk, Hamilton, & Southgate, 2018; Isomura & Nakano, 2016) and that this early mimicry is modulated by eye contact (de Klerk et al., 2018). However, it is unknown when other social factors, such as group membership, start to modulate infants’ spontaneous mimicry behavior. In the current study, therefore, we investigated whether mimicry is modulated by cues to group membership in 11-month-olds. Previous research suggests that infants are sensitive to signals related to group membership from an early age (Liberman, Woodward, & Kinzler, 2017). For example, 8-month-olds expect agents who look alike to act alike (Powell & Spelke, 2013), and 10- and 11-month-olds show a preference for native foreign speakers compared to foreign speakers (Kinzler, Dupoux, & Spelke, 2007) and for puppets who share their food preferences (Mahajan & Wynn, 2012). Furthermore, 10- to 14-month-olds are more likely to adopt the behaviors of native speakers, showing a greater propensity to try foods they endorse (Shutts, Kinzler, McKee, & Spelke, 2009) and to imitate their novel object- directed actions (Buttelmann, Zmyj, Daum, & Carpenter, 2013; Howard, Henderson, Carrazza, & Woodward, 2015). However, in these latter studies where infants selectively imitated object-directed actions of native speakers, it is difficult to determine whether infants’ imitative behaviors were predominantly driven by affiliation with the model or by the motivation to learn normative actions from members of their own linguistic group, and most likely both learning and social goals played a role (Over & Carpenter, 2012). Whereas the conscious imitation of object-directed actions is considered an important tool for social and cultural learning, the spontaneous unconscious mimicry of intransitive behaviors is often thought to serve a predominantly social function such as signaling similarity and enhancing affiliation (e.g., Lakin, Jefferis, Cheng, & Chartrand, 2003; but see also Kavanagh & Winkielman, 2016). The current study, therefore, aimed to investigate whether nonconscious mimicry is modulated by cues to group membership in infancy. Considering the importance of facial information in our day-to-day social interactions, facial mimicry may provide a particularly strong affiliative signal (Bourgeois & Hess, 2008); therefore, we specifically focused on mimicry of facial actions. Whereas the majority of previous studies on facial mimicry used emotional facial expressions as the stimuli (e.g., Isomura & Nakano, 2016; Kaiser, Crespo-Llado, Turati, & Geangu, 2017), in the current study we investigated infants’ tendency to mimic nonemotional facial actions, such as mouth opening, to ensure that we were measuring motor mimicry processes without confounding them with emotional contagion (Moody & McIntosh, 2011). Despite infants’ early sensitivity to cues to group membership, the only previous study investigating the effect of group membership on young children’s mimicry found that 4-year-olds, but not 3-year-olds, showed modulation of overt behavioral mimicry by group status (van Schaik & Hunnius, 2016). Given that our previous work has shown that facial mimicry is already flexibly modulated by gaze direction at 4 months of age (de Klerk et al., 2018), this apparent insensitivity of behavioral mimicry to group membership cues in preschoolers may seem perplexing. One explanation for this finding may be that in the study by van Schaik and Hunnius (2016), the authors used a minimal group paradigm in which the in- and out-groups were defined by an arbitrary marker (i.e., t-shirt color). First devised by Tajfel (Tajfel, Billig, Bundy, & Flament, 1971), this paradigm is based on adults’ tendencies to favor those who share a superficial likeness with themselves. However, as van Schaik and Hunnius (2016) pointed out, it is not obvious that such sensitivity to superficial attributes should be present during early childhood. Alternatively, it could be that coding of overt behavioral mimicry provides a less sensitive measure compared to facial mimicry as measured by EMG. Here, we used language as a signal of group membership instead because it has been suggested that this is a particularly potent cue to social structure (Liberman et al., 2017) and previous research has shown that infants’ behavioral preferences are modulated by this cue from at least 10 months of age (Kinzler et al., 2007). In the current study, 11-month-old infants observed videos of facial actions (e.g., mouth opening, eyebrow raising) performed by models who spoke to the infants in their native language (English; Native speaker condition) or an unfamiliar foreign language (Italian; Foreign speaker condition) while we measured activation of their mouth and eyebrow muscle regions using EMG to obtain an index of mimicry. EMG captures the subtle muscle changes that occur during automatic facial mimicry and likely provides a more objective and sensitive measure of mimicry compared to that measured by observational coding. This is important for studying facial mimicry in developmental populations, especially considering the controversy surrounding facial mimicry in newborn infants (e.g., Jones, 2009; Oostenbroek et al., 2016). We simultaneously used functional near-infrared spectroscopy (fNIRS) to investigate the neural mechanisms underlying any differential mimicry responses, allowing us to potentially shed more light on the underlying cognitive mechanisms. Given infants’ preference for linguistic in-group members (Kinzler et al., 2007; Liberman et al., 2017), previous work demonstrating that behavioral, vocal, and facial mimicry are modulated by group membership in adults (e.g., Bourgeois & Hess, 2008; Weisbuch & Ambady, 2008; Yabar, Johnston, Miles, & Peace, 2006), and previous work demonstrating that even in infancy facial mimicry can be flexibly deployed (de Klerk et al., 2018), we hypothesized that infants would show greater facial mimicry of native speakers compared to foreign speakers. In terms of neural activation, previous functional magnetic resonance imaging (fMRI) research with adult participants has highlighted several regions that may be involved in the selective mimicry of linguistic in-group members. First, the temporal parietal junction (TPJ) has been suggested to play an important role in self–other differentiation (Uddin, Molnar-Szakacs, Zaidel, & Iacoboni, 2006) and in the control of imitative responses (Hogeveen et al., 2015; Spengler, von Cramon, & Brass, 2009). Furthermore, previous work has demonstrated enhanced TPJ activation for interactions with in-group members (Rilling, Dagenais, Goldsmith, Glenn, & Pagnoni, 2008) and during mimicry in an affiliative context (Rauchbauer, Majdandžić, Hummer, Windischberger, & Lamm, 2015). It has been suggested that the TPJ may be of particular importance in situations where salient affiliative signals lead to a greater tendency to mimic, requiring a greater effort to disambiguate one’s own actions from those of others (Rauchbauer et al., 2015). Based on these findings, we hypothesized that if linguistic group status is indeed perceived as an affiliative signal leading to enhanced facial mimicry, infants would exhibit greater activation of temporoparietal regions during the observation of facial actions performed by the native speaker compared to the foreign speaker driven by the enhanced effort needed to differentiate their own actions from those of the model. The medial prefrontal cortex (mPFC) is another area that may play a role in the current study given that it has been implicated in reasoning about similar others (Mitchell, Banaji, & MacRae, 2005), shows greater activation when interacting with in-group members (Rilling et al., 2008), and has been shown to be involved in modulating mimicry (Wang & Hamilton, 2012; Wang, Ramsey, & Hamilton, 2011). Based on these findings, we hypothesized that infants would also exhibit greater activation over prefrontal areas during the observation of facial actions performed by native speakers compared to foreign speakers.","A total of 55 11-month-old infants observed the stimuli (for a description, see “Stimuli and procedure” section) while we simultaneously measured their facial muscle responses using EMG and their neural responses using fNIRS. The final sample consisted of 19 infants who provided sufficient data to be included in the EMG analyses (Mage = 343 days, SD = 15.50, range = 316–375; 4 girls) and 25 infants who provided sufficient data to be included in the fNIRS analyses (Mage = 342 days, SD = 14.99, range = 310–375; 9 girls). This study was part of a longitudinal project investigating the development of mimicry, and we tested all the infants who were able to come back for this second visit of the project. Power analyses using effect sizes based on the previous visit of the project, where we found evidence for social modulation of facial mimicry at 4 months of age (de Klerk et al., 2018), revealed that a total sample size of 12 participants would have provided enough power (.95 with an alpha level of .05) to identify similar effects. Because we tried to get the infants to wear both the fNIRS headgear and the EMG stickers on their face, the dropout rate for this study was relatively high, but still comparable to other neuroimaging studies with a similar age range that used only one method (e.g., de Klerk, Johnson, & Southgate, 2015; Stapel, Hunnius, van Elk, & Bekkering, 2010; van Elk, van Schie, Hunnius, Vesper, & Bekkering, 2008). A total of 36 infants were excluded from the EMG analyses due to technical error (n = 8), because they did not provide enough trials for analyses due to fussiness (n = 14) or inattentiveness (n = 8), or because they constantly vocalized or repeatedly put their fingers in their mouth (n = 4)—factors that were likely to have resulted in EMG activity unrelated to the stimulus presentation. A total of 30 infants were excluded from the fNIRS analyses due to a refusal to wear the fNIRS headgear (n = 6), too many bad channels (n = 2), or because they did not provide the minimum of 3 good trials per condition due to fussiness (n = 14) or inattentiveness (n = 6). One additional infant was excluded from both the fNIRS and EMG analyses because she was exposed to Italian at least once a week, and another infant was excluded because the older sibling had recently been diagnosed with autism spectrum disorder—a disorder with a strong genetic component (Ozonoff et al., 2011) in which spontaneous facial mimicry has been found to be atypical (McIntosh, Reichmann-Decker, Winkielman, & Wilbarger, 2006). Because we did not selectively include only monolingual infants in the original sample for the longitudinal project, 4 of the included infants were bilingual but heard English at least 60% of the time. One additional infant was bilingual Italian (heard 75% of the time) and French (heard 25% of the time), and for him the Italian speaker was coded as the native speaker.1 Although one might expect weaker in- and out-group effects in bilingual infants, previous research suggests that both monolingual and bilingual children prefer in-group members who use a familiar language (Souza, Byers-Heinlein, & Poulin-Dubois, 2013). All included infants were born full-term, healthy, and with normal birth weight. The study received approval from the institutional research ethics committee. Written informed consent was obtained from the infants’ caregivers. Stimuli and procedure The experiment took place in a dimly lit and sound-attenuated room, with the infants sitting on their parent’s lap at approximately 90 cm from a 117-cm plasma screen. Infants were presented with videos of two models who spoke either English (Native speaker) or Italian (Foreign speaker). Infants first observed 2 Familiarization trials during which the models labeled familiar objects in either English or Italian. Thereafter, Reminder trials, during which one of the models labeled a familiar object, and Facial Action trials, during which the same model performed facial actions such as mouth opening and eyebrow raising, alternated (see Fig. 1). Facial Action trials started with 1000 ms during which the model did not perform any actions, followed by her performing three repeats of the same facial action, each lasting 3000 ms. The Reminder and Facial Action trials were alternated with 8000-ms Baseline trials consisting of pictures of houses, landscapes, and landscapes with animals to allow the hemodynamic response to return to baseline levels. The order of trials during the Familiarization phase was randomized, and the order of trials during the Reminder and Facial Action phases was pseudorandomized to ensure that infants saw a roughly equal number of eyebrow and mouth actions. Videos were presented until the infants had seen 12 10-s Facial Action trials (6 Native and 6 Foreign) or until the infants’ attention could no longer be attracted to the screen. Both the EMG and fNIRS analyses focused only on the Facial Action trials. The role of the models (Native vs. Foreign speaker) was counterbalanced across infants. To validate our procedure as tapping infants’ preference for native speakers, at the end of the session infants were encouraged to choose between the two models by reaching for a picture of one of the two models presented on cardboard. The cardboard with the pictures was brought into the room by an experimenter who was unaware which of the models in the videos had been the native speaker. The experimenter held the cardboard in front of the infants without saying anything to avoid biasing them toward choosing the native speaker. Video coding and data exclusion ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Videos were coded offline to determine which trials could be included in the analyses. Note that because the EMG and fNIRS signals have very different temporal resolutions and are influenced by different types of noise, the exclusion criteria for the EMG and fNIRS trials were not identical. For example, whereas facial mimicry as measured by EMG can be recorded on a millisecond scale, the hemodynamic response takes several seconds to build up. Therefore, each 3000-ms period during which a facial action was performed by the model in the video was treated as a separate EMG trial, whereas for the fNIRS analyses we treated the 10-s videos including three repeats of the same facial action as one trial (see Fig. 1). EMG trials during which the infants did not see at least two thirds of the action, or trials during which the infants vocalized, smiled, cried, or had something in their mouth (e.g., their hand), were excluded from the analyses because EMG activity in these cases was most likely due to the infants’ own actions. EMG trials during which the infants pulled or moved the EMG wires were also excluded, as were trials during which there was lost signal over either of the electrodes. Only infants with at least 2 trials per trial type (Native_Mouth, Native_Eyebrow, Foreign_Mouth, or Foreign_Eyebrow) and at least 6 trials per condition (Native_FacialAction vs. Foreign_FacialAction) were included in the EMG analyses (previous infant EMG research used a similar or lower minimum number of included trials; e.g., Isomura & Nakano, 2016; Turati et al., 2013). On average, infants contributed 11 EMG trials per condition to the analyses: 6 trials for the Native_Mouth trial type (range = 3–9), 5 trials for the Native_Eyebrow trial type (range = 2–8), 6 trials for the Foreign_Mouth trial type (range = 2–9), and 5 trials for the Foreign_Eyebrow trial type (range = 2–9). The number of included EMG trials did not differ between the Native and Foreign conditions (p = .509). fNIRS trials during which the infants did not attend to at least two of the three facial actions, or trials during which the infants were crying, were excluded from analyses. We also excluded Baseline trials during which the infants were looking at their parents’ face or their own limbs. Only infants with at least 3 trials per experimental condition (Native_FacialAction vs. Foreign_FacialAction)2 were included in the fNIRS analyses (Lloyd-Fox, Blasi, Everdell, Elwell, & Johnson, 2011; Southgate, Begus, Lloyd-Fox, di Gangi, & Hamilton, 2014). On average, infants contributed 5 fNIRS trials per condition to the analyses: 5 trials in the Native_FacialAction condition (range = 3–10) and 5 trials in the Foreign_FacialAction condition (range = 3–8). The number of included fNIRS trials did not differ between the two conditions (p = .83). EMG recording and processing ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Bipolar EMG recordings were made using pediatric surface Ag/AgCl electrodes that were placed on the cheek and forehead following recommendations by Fridlund and Cacioppo (1986) with an inter-electrode spacing of approximately 1 cm to measure activation over the masseter and frontalis muscle areas, respectively. The electrodes on the forehead were embedded within the fNIRS headgear, with the inferior electrode affixed approximately 1 cm above the upper border of the middle of the brow and the second electrode placed 1 cm superior to the first one (Fridlund & Cacioppo, 1986) (see Fig. 2A). A stretchy silicone band was tightly affixed over the headgear to ensure that the fNIRS optodes and EMG electrodes made contact with the scalp and the skin, respectively. The electrodes were connected to Myon wireless transmitter boxes that amplified the electrical muscle activation, which was in turn recorded using ProEMG at a sampling rate of 2000 Hz. After recording, the EMG signal was filtered (high-pass: 30 Hz; low-pass: 500 Hz), smoothed (root mean square over 20-ms bins), and rectified. The EMG signal was segmented into 3000-ms epochs, and the average activity in each epoch was normalized (i.e., expressed as z-scores) within each participant and each muscle group (masseter and frontalis regions) before the epochs for each trial type were averaged together. Because facial mimicry can be defined as the presence of greater activation over corresponding muscles than over non- corresponding muscles during the observation of facial actions (e.g., McIntosh et al., 2006; Oberman, Winkielman, & Ramachandran, 2009), we calculated a mimicry score per trial by subtracting EMG activity over the non-corresponding muscle region from EMG activity over the corresponding muscle region (e.g., on an eyebrow trial, we subtracted activity over the masseter region from activity over the frontalis region so that a more positive score indicates more mimicry). fNIRS recording and processing ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ fNIRS data were recorded using the University College London (UCL)–NIRS topography system, which uses two continuous wavelengths of near-infrared light (770 and 850 nm) to detect changes in oxyhemoglobin (HbO2) and deoxyhemoglobin (HHb) concentrations in the brain and has a sampling rate of 10 Hz (Everdell et al., 2005). Infants wore custom-built headgear with a total of 26 channels, with a source–detector separation of 25 mm over the temporal areas and 4 channels with a source–detector separation of 30 mm over the frontal area. The headgear was placed so that the third optode was centered above the pre-auricular point (directly over T7 and T8 according to the 10–20 system) (see Fig. 2B). Based on the understanding of the transportation of near-infrared light through tissue, this source–detector separation was predicted to penetrate up to a depth of approximately 12.5 mm from the skin surface for the temporal areas and a depth of approximately 15 mm from the skin surface for the frontal areas, allowing measurement of both the gyri and parts of the sulci near the surface of the cortex (Lloyd-Fox, Blasi, & Elwell, 2010). Previous research using coregistration of fNIRS and MRI using the same array design for the temporal areas has demonstrated that this design permits measurement of brain responses in cortical regions corresponding to the inferior frontal gyrus (IFG), superior temporal sulcus (STS), and TPJ areas (see Lloyd-Fox et al., 2014). The frontal array was designed to measure brain responses in mPFC (Kida & Shinohara, 2013; Minagawa-Kawai et al., 2009). Data were preprocessed using a MATLAB software package called HOMER2 (MGH–Martinos Center for Biomedical Imaging, Boston, MA, USA) and analyzed using a combination of custom MATLAB scripts and the Statistical Parametric Mapping (SPM)–NIRS toolbox (Ye, Tak, Jang, Jung, & Jang, 2009). Data were converted to . nirs format, and channels were excluded if the magnitude of the signal was greater than 97% or less than 3% of the total range for longer than 5 s during the recording. Channels with raw intensities smaller than 0.001 or larger than 10 were excluded, and motion artifacts were corrected using wavelet analyses with 0.5 times the interquartile range. Hereafter, the data were band-pass filtered (high-pass: 0.01 Hz; low-pass: 0.80 Hz) to attenuate slow drifts and high-frequency noise. The data were then converted to relative concentrations of HbO2 and HHb using the modified Beer–Lambert law. We excluded from analysis any channels that did not yield clean data for at least 70% of the infants. This resulted in the exclusion of three channels associated with a source that was faulty for a subset of the assessments (Channels 21, 23, and 24). As in previous infant fNIRS studies (e.g., Southgate et al., 2014), infants for whom more than 30% of remaining channels were excluded due to weak or noisy signal were excluded from analysis (n = 2). Note that for these 2 excluded infants, the intensities were very weak over more than 30% of the channels because the fNIRS headgear was not fitted properly. Our data analysis approach was determined a priori and followed Southgate et al. (2014). For each infant, we constructed a design matrix with five regressors. The first regressor modeled the Native Reminder trials (duration = 6 s), the second regressor modeled the Foreign Reminder trials (duration = 6 s), the third regressor modeled the Native_FacialAction trials (duration = 10 s), the fourth regressor modeled the Foreign_FacialAction trials (duration = 10 s), and the fifth regressor modeled the Baseline trials (duration = 8 s). Excluded trial periods were set to zero, effectively removing them from the analyses. The regressors were convolved with the standard hemodynamic response function to make the design matrix (Friston, Ashburner, Kiebel, Nichols, & Penny, 2011). This design matrix was then fit to the data using the general linear model as implemented in the SPM–NIRS toolbox (Ye et al., 2009). Beta parameters were obtained for each infant for each of the regressors. The betas were then used to calculate a contrast between the different conditions of interest for each infant. Although in principle the HbO2 and HHb responses should be coupled, with HbO2 responses going up and HHb responses going down in response to stimulus presentation, studies with infant participants often do not find statistically significant HHb changes (Lloyd-Fox et al., 2010; Lloyd-Fox, Széplaki-Köllőd, Yin, & Csibra, 2015). Therefore, as in previous infant fNIRS studies, our analyses focused on changes in HbO2. To ensure statistical reliability, and given that previous research has shown that our cortical areas of interest (e.g., STS, TPJ, IFG) are unlikely to span just one channel (Lloyd-Fox et al., 2014), we considered that activation at a single channel would be reliable only if it was accompanied by significant activation at an adjacent channel (Lloyd-Fox et al., 2011; Southgate et al., 2014). To implement this, we created a Monte Carlo simulation of p values at every channel of our particular array and, on 10,000 cycles, tested whether p values for two or more adjacent channels fell below a specific channel threshold. We repeated the simulations for a range of channel thresholds and then selected the channel threshold that led to a whole-array threshold of p < .05 for finding two adjacent channels activated by chance. The appropriate channel threshold was 0.0407, so we considered only effects present at p < .0407 in two adjacent channels to be interpretable. EMG ~~~ A repeated-measures analysis on the mimicry scores (activation over the corresponding muscle region minus activation over the non-corresponding muscle region) with condition (Native vs. Foreign speaker) and action type (Eyebrow vs. Mouth) as within-participant factors demonstrated a significant main effect of condition, F(1, 18) = 6.07, p = .024, ηp2 = .252. There were no other significant main effects or interactions. As can be seen in Fig. 3, infants showed significantly greater mimicry of facial actions performed by the native speaker compared to the foreign speaker. The mouth and eyebrow mimicry scores in the Native condition were not significantly different from zero (ps > .149), but the average mimicry score in the Native condition was, t(18) = 2.407, p = .027. This demonstrates that, overall, infants were significantly more likely to show greater activation over the corresponding facial muscles compared to the non-corresponding facial muscles when they observed the facial actions performed by the native speaker. (See the online supplementary material for the same analyses performed on the individual muscle activation as well as a depiction of the EMG signal over the masseter and frontalis muscle regions time-locked to the onset of the facial actions.) Baseline correction Note that the analyses reported above were planned a priori and followed our previous work (de Klerk et al., 2018). There are several reasons why we did not use baseline-corrected EMG values in these analyses. First, if we define mimicry as a relative pattern of muscle activation in which corresponding facial muscles are activated to a greater extent than non-corresponding facial muscles, it does not seem necessary to perform a baseline correction. Instead, by transforming the EMG activity to z-scores and calculating mimicry scores, we can measure infants’ tendency to selectively activate the corresponding facial muscles to a greater degree than the non-corresponding facial muscles during the observation of facial actions. The second reason why we did not subtract activity during the baseline period from activity during the trials was that one of the main reasons for excluding trials at this age were vocalizations, and infants tended to vocalize a lot during the baseline stimuli. To maximize the number of trials that we could include, and hence the number of infants that could be included in the analyses, we decided not to also require there to be a valid baseline preceding each trial. However, on request of the reviewers, we have recoded the videos of the sessions to identify all valid baselines and reanalyzed the data using baseline-corrected values. In these analyses, the effect of condition was no longer significant (p = .193) and there was no evidence for mimicry. (See the supplementary material for more details as well as potential explanations for this discrepancy in the findings.) fNIRS ~~~~~ t Tests revealed that the left temporal parietal cortex was sensitive to the linguistic status of the models. This region showed a significantly greater hemodynamic response (based on HbO2) both when the Native_FacialAction condition was contrasted to the Baseline condition and when it was compared to the Foreign_FacialAction condition (see Table 1). This effect was present at p < .0407 over two adjacent channels (Channels 12 and 13; see Fig. 4). No channels showed a significantly greater response in the Foreign_FacialAction condition compared to the Native_FacialAction condition. Although we found one channel over the frontal cortex (mPFC area) that showed a significant hemodynamic response for the Foreign_FacialAction condition compared to the Baseline condition, there was no significant difference between the two conditions over this channel. (See Supplementary Fig. 2 for the beta values for all channels.) Relationship between EMG and fNIRS data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We initially planned to perform correlational analyses to investigate the relationship between the EMG and fNIRS results. These analyses are underpowered (N = 15)3 due to the relatively high dropout rate, however, we report the results here for completeness. We found a marginally significant negative correlation between the Native_FacialAction > Foreign_FacialAction contrast over the left temporal parietal cortex (Channels 12 and 13) and the differential mimicry response (Native_FacialAction − Foreign_FacialAction), r(13) = −.487, p = .066 (see Fig. 5). Suggesting that infants who showed a greater hemodynamic response over the left temporoparietal area when observing facial actions performed by the native speaker compared with the foreign speaker showed less mimicry of facial actions performed by the native speaker compared with the foreign speaker (see Fig. 5). Note that because this correlational analysis is underpowered, it needs to be interpreted with caution. Choice task ~~~~~~~~~~~ Infants were significantly more likely to choose a picture of the native speaker compared to the foreign speaker, χ2(1) = 6.54, p = .011. Of the 27 infants included in the EMG and/or fNIRS analyses, 17 chose the native speaker, 5 chose the foreign speaker, and 5 did not choose.","This study is the first to demonstrate that the linguistic status of an observed model modulates facial mimicry in infancy and that this modulation is accompanied by changes in neural activation over the left temporal parietal cortex. These results show that one of the hallmarks of mimicry—that it is modulated by cues to group membership—seems to be present from at least 11 months of age. What is the psychological mechanism underlying infants’ tendency to selectively mimic the native speaker? It is important to note from the outset that the modulation of mimicry by the language of the model does not imply that infants conceived of the native speaker as a member of their in-group (Liberman et al., 2017). Infants are clearly sensitive to cues that correlate with group membership, such as language familiarity, and this sensitivity modulates their attention and expectations (Begus, Gliga, & Southgate, 2016). For example, infants may have considered the native speaker to be a more useful source of information (Begus et al., 2016) or a more competent person (Brooker & Poulin-Dubois, 2013) because she correctly labeled the familiar objects, whereas they could not understand the foreign speaker. As a result of this, the native speaker may have captured the infants’ attention to a greater extent, leading to increased encoding of her facial actions and, consequently, greater activation of the associated motor representations and greater mimicry—a process called input modulation (Heyes, 2013). Although the fact that we did not find any significant differences in the number of included trials between the conditions seems to speak against this interpretation, looking is not necessarily equivalent to attending (Aslin, 2012); therefore, we cannot completely rule out the possibility that there may have been enhanced encoding of the facial actions performed by the native speaker. It is also possible that, like adults, infants may have had a greater motivation to affiliate with the native speaker, and mimicking her facial actions could have functioned as a means to communicate their similarity to the model (van Baaren et al., 2009). Indeed, infants’ preference for the native speaker in the choice test seems consistent with the idea that infants have an early emerging preference to interact with familiar others (Kinzler et al., 2007). Although the facial mimicry we observed was subtle and mainly detectable by EMG, previous studies suggest that our emotions and social perceptions can be influenced by facial stimuli that we cannot consciously perceive (Bornstein, Leone, & Galley, 1987; Dimberg, Thunberg, & Elmehed, 2000; Li, Zinbarg, Boehm, & Paller, 2008; Svetieva, 2014; Sweeny, Grabowecky, Suzuki, & Paller, 2009). Therefore, regardless of the underlying mechanism—attentional effects or affiliative motivations—the increased mimicry of in-group members during infancy is likely to have a positive influence on social affiliation (Kavanagh & Winkielman, 2016; van Baaren et al., 2009) and, therefore, can be expected to be reinforced over the course of development (Heyes, 2017). A question that remains is how the facial mimicry that we measured in the current study relates to the overt mimicry behaviours observed in the original adult studies on behavioral mimicry (e.g., Chartrand & Bargh, 1999). In the adult literature, facial mimicry is generally placed in the same context as the spontaneous mimicry of other nonverbal behaviors such as postures and gestures (e.g., Chartrand & Lakin, 2013; Stel & Vonk, 2010; van Baaren et al., 2009). Indeed, facial mimicry as measured by EMG and mimicry of postures, gestures, and mannerisms share many properties; they both seem to occur without conscious awareness (Chartrand & Bargh, 1999; Dimberg et al., 2000), are influenced by the same factors such as group membership (Bourgeois & Hess, 2008; Yabar et al., 2006), and are supported by similar neural mechanisms (Likowski et al., 2012; Wang & Hamilton, 2012). Furthermore, it has been suggested that the subtle mimicry that can be measured by EMG may be a building block for more overt and extended matching (Moody & McIntosh, 2006, 2011). Potentially, this constitutes a quantitative change rather than a qualitative change, where the mimicry becomes overtly visible whenever the activation of the motor representation in the mimicker reaches a certain threshold. Future research will need to investigate the relationship between overt behavioral mimicry and subthreshold mimicry measured by EMG in more detail to determine whether they are indeed two sides of the same coin or rather distinct processes. The fact that EMG can pick up on relatively subtle mimicry may also explain the discrepancy between the current study’s findings and those of van Schaik and Hunnius (2016). In that study, 4-year-olds, but not 3-year-olds, showed modulation of behavioral mimicry (e.g., yawning, rubbing the lips) depending on whether the model was wearing the same t-shirt color as them. One possibility is that children younger than 4 years are not sensitive to superficial attributes such as t-shirt color (van Schaik & Hunnius, 2016) or that they lack experience with being divided into teams based on such attributes. Alternatively, it could be that coding of overt behavioral mimicry is less sensitive than facial mimicry as measured by EMG. Future research should investigate whether a minimal group paradigm might elicit selective mimicry in children younger than 4 years when more sensitive measures, such as facial mimicry as measured by EMG, are used. Such an approach would also allow researchers to investigate whether the selective mimicry effects found in the current study would hold when language, familiarity, and competence factors are controlled for across the in- and out-group members. It should be noted that although we found evidence for mimicry overall, the mimicry scores for the mouth and eyebrow actions separately were not significantly different from zero. In addition, when we performed baseline-corrected analyses (see supplementary material), the effect of condition became nonsignificant and there was no evidence for mimicry, suggesting that the mimicry responses we found here might not be as robust as those reported in previous studies (e.g., Datyner, Henry, & Richmond, 2017; Geangu, Quadrelli, Conte, Croci, & Turati, 2016; Isomura & Nakano, 2016). There are several possible explanations for this. First of all, most of the previous developmental studies on facial mimicry used emotional facial expressions as the stimuli. One possibility is that muscle responses to these stimuli reflect emotional contagion, a process in which the observed stimuli induce a corresponding emotional state in the child, resulting in the activation of corresponding facial muscles, whereas mimicry of nonemotional facial expressions, such as those used in the current study, reflects subtler motor mimicry (Hess & Fischer, 2014; Moody & McIntosh, 2011). Future research will need to investigate whether there are differences in the mechanisms underlying the mimicry of emotional and nonemotional facial actions. A recent study suggests that the development of facial mimicry is supported by parental imitation (de Klerk, Lamy-Yang, & Southgate, 2018). Thus, a second possibility is that the current study included a mixture of infants who receive high and low levels of maternal imitation, resulting in a relatively high level of variability in the mimicry responses, with some infants potentially not having received a sufficient amount of correlated sensorimotor experience to support the mimicry of the observed facial actions. We found a greater hemodynamic response in channels overlying the left temporal parietal cortex during the observation of facial actions performed by the native speaker compared to the foreign speaker. The temporal parietal region, and the TPJ in particular, has been suggested to play a critical role in disambiguating signals arising from one’s own and others’ actions (Blakemore & Frith, 2003), and some evidence suggests that it may play a role in these processes from an early age (Filippetti, Lloyd-Fox, Longo, Farroni, & Johnson, 2015). Although the left lateralization of our results seems inconsistent with previous research mainly implicating the right TPJ in separating self- and other-generated signals (e.g., Spengler et al., 2009), a recent study suggests that the bilateral TPJ is involved in this process (Santiesteban, Banissy, Catmur, & Bird, 2015). One interpretation of our fNIRS findings is that the Native speaker condition may have posed higher demands on differentiating between self- and other-generated actions because it presented infants with a highly affiliative context where the urge to mimic was strong (for similar results with adult participants, see Rauchbauer et al., 2015). In other words, a possible consequence of infants’ enhanced tendency to mimic the native speaker may have been an increased self–other blurring, which led to greater compensatory activity over the temporal parietal cortex. In line with this interpretation, the marginally significant negative correlation between the betas for the contrast Native_FacialAction > Foreign_FacialAction over the left temporal parietal cortex and the mimicry difference score suggests that those infants who showed a greater amount of temporal parietal cortex activation may have maintained greater self–other differentiation and showed a less pronounced selective mimicry response. However, given that this correlational analysis was heavily underpowered, this finding needs to be replicated in a larger sample. In addition, considering the limited amount of research on the role of the TPJ in self–other differentiation during infancy, future research will need to further investigate the role of this area in inhibiting mimicry responses during infancy. Finally, the spatial resolution of fNIRS does not allow us to say with certainty that the significant channels did not, at least in part, overlie the posterior STS (pSTS). Therefore, another possible interpretation of our fNIRS findings is that greater activation over the temporal parietal cortex during the observation of facial actions performed by the native speaker reflects input modulation, that is, the Native speaker condition may have captured the infants’ attention to a greater extent, leading to enhanced encoding of her facial actions as indicated by greater activation over the pSTS. This interpretation would be consistent with a previous study in which we found greater activation over the pSTS in the condition associated with greater facial mimicry in 4-month-old infants (de Klerk et al., 2018). However, given that the significant channels in the current study are located in a more posterior position on the scalp, they are unlikely to reflect activation of the exact same area. Future fNIRS–fMRI coregistration work would be beneficial to help tease apart the role of adjacent cortical areas in the modulation of mimicry responses. One potential concern may be that the differences in facial mimicry between the two conditions led to subtle artifacts in the fNIRS data that created the differences in the hemodynamic response over the left temporal parietal area. We should note that it seems unlikely that facial muscle contractions that cannot be seen with the naked eye would cause artifacts large enough to result in a significant difference between the two conditions. In addition, even if this did happen, it seems highly unlikely that this would have specifically affected two adjacent channels over the left hemisphere over a cortical area that is the farthest removed from the facial muscles rather than the frontal channels that are directly on the forehead. Unlike in adults (Wang et al., 2011), we did not find involvement of the mPFC in the modulation of mimicry in infants. One possibility is that our frontal array was not optimal for measuring responses in the mPFC, although previous studies using similar array designs have reported differential responses over this area (Kida & Shinohara, 2013; Minagawa-Kawai et al., 2009). Another possibility is that at this relatively young age, the selective mimicry responses were mainly driven by bottom-up attentional processes, such as the tendency to pay more attention to familiar others, whereas more top-down mechanisms start to play a role in modulating mimicry behavior only once myelination of the relevant long-range connections with the mPFC is more established (Johnson, Grossmann, & Kadosh, 2009). Future research measuring functional connectivity during mimicry behaviors over the course of development will need to investigate this further. Taken together, our results demonstrate that facial mimicry is flexibly modulated by cues to group membership from at least 11 months of age. Although the exact mechanisms underlying this selective mimicry response will need to be investigated in future research, these findings suggest that the foundations for the role that mimicry plays in facilitating social bonds seem to be present during the first year of life."],["To make sense of the visual world, we need to move our eyes to focus regions of interest on the high-resolution fovea. Eye movements, therefore, give us a way to infer mechanisms of visual processing and attention allocation. Here, we examined age-related differences in visual processing by recording eye movements from 37 children (aged 6–14 years) and 10 adults while viewing three 5-min dynamic video clips taken from child-friendly movies. The data were analyzed in two complementary ways: (a) gaze based and (b) content based. First, similarity of scanpaths within and across age groups was examined using three different measures of variance (dispersion, clusters, and distance from center). Second, content-based models of fixation were compared to determine which of these provided the best account of our dynamic data. We found that the variance in eye movements decreased as a function of age, suggesting common attentional orienting. Comparison of the different models revealed that a model that relies on faces generally performed better than the other models tested, even for the youngest age group (<10 years). However, the best predictor of a given participant's eye movements was the average of all other participants’ eye movements both within the same age group and in different age groups. These findings have implications for understanding how children attend to visual information and highlight similarities in viewing strategies across development. --------------------------------------------------------------------------------","Eye tracking is increasingly being used to try to infer what people are doing (Hayhoe & Ballard, 2005) or thinking (Kardan, Berman, Yourganov, Schmidt, & Henderson, 2015) based solely on how they looked at a visual scene. One advantage of this method over, for example, self-report or psychophysical testing is that it can readily be applied to populations that are difficult to evaluate, from young babies (Jones, Kalwarowsky, Atkinson, Braddick, & Nardini, 2014) to clinical populations, including autistic people (e.g., see Papagiannopoulou, Chitty, Hermens, Hickie, & Lagopoulos, 2014, for a review) and patients with Alzheimer’s disease (Crutcher et al., 2009). Although there has been a great deal of work looking at modeling patterns of fixations in adults, particularly looking at top-down and bottom-up influences (Itti, Koch, & Niebur, 1998; Xu, Jiang, Wang, Kankanhalli, & Zhao, 2014), and there has been some important work examining how these influences develop during a child’s first months and years (e.g., Amso, Haas, & Markant, 2014; Franchak, Heeger, Hasson, & Adolph, 2016; Frank, Vul, & Johnson, 2009), there has as yet been no systematic examination of different models applied to viewing dynamic scenes in school-age children compared with adults. We addressed this gap in the literature in this study. Development of fixation behavior ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Certain viewing behaviors, such as looking at faces and stimuli with social relevance, develop so early in childhood that they appear to be largely innate. For example, it is well established that newborn babies preferentially track faces and face-like stimuli in simple displays (Farroni et al., 2005; Johnson, Dziurawiec, Ellis, & Morton, 1991). Using static (complex) images, several studies have shown that infants aged 6 months and older orient to faces in images that contain non-face distractors (Di Giorgio, Turati, Altoè, & Simion, 2012; Gliga, Elsabbagh, Andravizou, & Johnson, 2009; Gluckman & Johnson, 2013). Frank, Vul, and Saxe (2012) used videos of objects, faces, children playing with toys, and complex social scenes with young children aged 3–30 months. They showed further that facial and bodily features that have social relevance, such as eyes, mouths, and hands, capture infants’ and toddlers’ attention and that this capacity to direct their attention to the stimuli that are potentially the most socially informative increases with age. In another study, Frank, Amso, and Johnson (2014) reported an age-related increase in looking at faces in complex videos (clips from Peanuts [Charlie Brown] and Sesame Street) in 3- to 9-month-old infants, which correlated with increased attentional orienting using a visual search task, suggesting that the extent to which infants show social preferences may well be underpinned by their ability to detect socially relevant stimuli in otherwise complex dynamic scenes. Frank et al. (2014) finding is also consistent with a recent report showing that infants over 4 months of age tended to look first and longest at faces, whereas 4-month-olds tended to look at the most salient object in a display (Kwon, Setoodehnia, Baek, Luck, & Oakes, 2016). Less is known, however, about the developments in fixation behavior that take place beyond early childhood. Kirkorian, Anderson, and Keen (2012) showed children (aged 1–4 years) and adults 20-min clips of television shows and found that younger children fixated more regions over a larger area than did older children. This variability was greatest immediately following scene cuts, which these authors proposed is due to an inability to suppress attention to irrelevant features (from the previous scene). More recently, Helo, Pannasch, Sirri, and Rämä (2014) examined differences in scanning behavior in adults and children aged 2–10 years and reported that fixation durations decreased and saccade amplitudes increased with age, at least with static images of naturalistic scenes, which these authors attributed to gains in general cognitive development. Similar age-related trends have been found in a preferential looking task, with eye movement response times falling with age from 1 to 12 years (Kooiker, van der Steen, & Pel, 2016). Interestingly, the response time to fixate highly salient targets reduced more rapidly with age than did fixations to less salient targets. Modeling natural fixation ~~~~~~~~~~~~~~~~~~~~~~~~~ One key question is precisely what is driving development in fixation behavior during childhood. Biologically inspired models have been developed to account for patterns of fixation within complex scenes and fall into two broad categories. Saliency models are driven by low-level (pixel-based) saliency and predict that eye movements are drawn to regions of visual information that differ locally in some basic feature (e.g., orientation, color) (Itti et al., 1998). In this scenario, fixation is mainly driven by a bottom-up process that relies primarily on sensory (rather than cognitive) processing (Xu et al., 2014). For simplicity, we refer to this class of model as “saliency” driven. The second class of models, top-down models, suggests that our eye movements are largely driven by cognitive or contextual factors. In this scenario, we expect that factors that affect top-down processing (e.g., age) will influence performance. Although age may also be considered a factor in bottom-up processing (as low-level sensory processing becomes more developed), this is unlikely to be the case for older children given that low-level sensory processes underlying acuity are adult-like by 3 years of age (Brown & Lindsey, 2009). Saliency-based models Most low-level models of fixation are based on Koch and Ullman’s (1985) and Itti et al. (1998) saliency model of visual attention in static images. These authors proposed that eye movements are preferentially driven by points of high image saliency, where the local statistics of an image patch differ from its surround. This model, and extensions of it, has been very influential (e.g., Baddeley & Tatler, 2006; Itti & Koch, 2000) and has spurred new analysis methods (e.g., Barthelmé, Trukenbrod, Engbert, & Wichmann, 2013), but it has only recently begun to incorporate the importance of features of dynamic images such as object motion and transformation. This is a major shortcoming because a reliance on static scenes is likely to minimize the relevance of ongoing semantic context such as social relevance and knowledge of cause and effect. We return to this point below. Top-down models Nonvisual factors influence interobserver variability in fixation patterns (particularly in response to dynamic stimuli). For example, when a video clip of people conversing is accompanied by a soundtrack that matches the visual content, observers are more accurate at localizing the face of the speaker (Coutrot & Guyader, 2014). Here, the sound needs to be understood as a human voice in order to drive eye movements toward the inferred speaker. But note that it is possible in some cases that the synchrony of sounds and visual transients (e.g., mouth movements) by themselves may drive eye movements, in which case audio can act as a more bottom-up influence. Other high-level processes have been shown to contribute to eye movements. For example, object “importance” (’t Hart, Schmidt, Roth, & Einhauser, 2013), social cues (faces or gaze; e.g., Birmingham, Bischof, & Kingstone, 2009), task instructions (Ballard & Hayhoe, 2009; Koehler, Guo, Zhang, & Eckstein, 2014), a person’s prior expectation about a scene (Eckstein, Drescher, & Shimozaki, 2006), and a person’s memory of a visual task (Ballard & Hayhoe, 2009) all can influence where a person allocates his or her attention in a visual scene and might not always be the region of highest saliency. A recent analysis of eye movements by Xu et al. (2014) investigated the potential influence of both bottom-up and top-down factors during fixation. They categorized static images into different attribute qualities: “pixel attributes” (low-level features akin to saliency), “object attributes” (e.g., object size, eccentricity), and “semantic attributes” (e.g., whether an object is being looked at by an individual in the scene). They reported that semantic-level attributes that would reflect top-down processing—particularly objects being gazed at, faces, and text—influenced observer fixations more than lower-level saliency. Dynamic stimuli It is increasingly recognized that fixation during prolonged presentation of static visual scenes might not be representative of natural viewing behavior of complex dynamic scenes. Consequently, dynamic stimuli (e.g., clips, movies) are increasingly being employed to gauge visually guided behavior. Yet, this new approach brings with it a new set of complications. For example, Dorr, Martinez, Gegenfurtner, and Barth (2010) reported that the type of dynamic stimulus used can largely influence performance. They reported a higher degree of variability of eye movements between observers when watching natural movies as compared with commercial movies. This finding arises largely because commercial movies suffer more from “center of screen bias” (e.g., relevant objects are framed in the center of the shot) as well as an increase in temporal structure arising from frequent cuts, whereas Dorr and colleagues’ natural movies were approximately 20 s uncut and were generally shot from fixed camera positions, meaning that the ongoing action in the clip would not necessarily have occurred in the center of the screen. In another example, Mital, Smith, Hill, and Henderson (2010) reported that gaze clustering across observers viewing a dynamic scene is determined mainly by motion within the clip. Although some of these studies have also quantified variability of eye movements, and studies have compared face and saliency models with adults and infants/toddlers (<4 years of age), there has been no systematic evaluation of different models that might account for changes in eye movement behavior in school-aged children viewing complex dynamic images. Furthermore, it has been shown that saliency is most influential in guiding the first fixations in static images, before top-down influences come into play (Parkhurst, Law, & Niebur, 2002). One expectation from this finding is that dynamic information may exacerbate the differences between early and later fixations. Specifically, we should expect that immediately after a scene cut, saliency will be more influential (we liken this to the first fixations in a static image), but that as the scene plays out, saliency influences will diminish. We also expect that, in general over the course of the clip, younger children’s viewing behavior will be better captured by a bottom-up saliency model than will that of adults. The current study ~~~~~~~~~~~~~~~~~ Here, we examined these changes in visual attention using a rich range of stimuli in which we measure the age-dependent variability in eye movements in school-age children relative to adults (taking into account scene cuts, which can lead to a change in semantic content in a film). We used longer-duration clips (3–6 min) of three different child-friendly movies (Roadrunner [cartoon], Night at the Museum, and Elf) to better examine the role of top-down (semantic) influences on eye movements as observers follow the story, and we sampled older children to allow these processes to have a greater influential role. These methods enabled us to test two main accounts of visual attention across age groups: one that relies on low-level saliency models and one that relies on the presence of faces in the scene. Although much research has looked at eye movements in infants or young children (e.g., Franchak et al., 2016; Frank et al., 2009; Kirkorian et al., 2012), we focused here on older children to chart how the development of cognition and interest in social objects is reflected in eye movements. We hypothesized that (a) we would observe an increased consistency in fixations with age as reported in adults (e.g., Dorr et al., 2010), and consistent with reports in infants (e.g., Frank et al., 2009) and young children (e.g., Kirkorian et al., 2012), and that (b) a model predicting fixation behavior based on faces would outperform saliency models (Franchak et al., 2016; Frank et al., 2009).","A total of 37 children from a range of ethnic backgrounds (identified by their parents as 19 Caucasian, 5 South Asian, 4 Black or Black/Caribbean, 3 Middle Eastern, 3 mixed race, and the remaining 3 unspecified) took part in the experiment during 1 week of “Brain Detectives,” a science club run at the UCL Institute of Education, University College London. Because several of the analysis methods we used to quantify performance (interobserver dispersion, cluster number, and normalized scanpath salience) depended on comparing fixation patterns within groups of individuals, it was not possible to treat age as a continuous variable in analyses. Therefore, we divided the children into two groups of approximately equal numbers and age ranges: “younger” (<10 years; n = 20, 9 girls, mean age = 7.8 years, range = 5.9–9.4) and “older” (≥10 years; n = 17, 12 girls, mean age = 11.9 years, range = 10.2–13.9). In addition, 10 adults (5 women, mean age = 31.9 years, range = 23.0–45.1) were tested at the UCL Institute of Ophthalmology, University College London. All participants reported normal or corrected-to-normal vision. Written informed consent was obtained from the adults and from children’s parents prior to their or their children’s participation in the experiment. The experiment was approved by Institute of Education and Institute of Ophthalmology research ethics committees.","Two identical systems were used for stimulus presentation and eye tracking. Calibration sequences and movies were presented using MATLAB (MathWorks, Natick, MA, USA) and the PsychToolbox (Brainard, 1997) running on Windows 7 PCs.","were displayed on LG W2363D LCD monitors (1920 × 1080 pixels, refresh rate = 60 Hz). The displays were calibrated using a photometer and linearized using lookup tables in software. Eye-tracking data were collected on two EyeLink 1000 systems with remote cameras (SR Research, Mississauga, Ontario, Canada) at 250 or 500 Hz. The reported average accuracy of the eye tracker is 0.5°, the spatial resolution is 0.05°, and the remote camera allowed for head movements of up to 22, 18, and 20 cm (horizontal, vertical, and depth, respectively). Viewing distance was approximately 57 cm, so that 1 pixel subtended approximately 1.6 arcmin and movies (at a resolution of 1708 × 960 pixels) covered 43.4 × 25.1° of visual angle. Children watched the videos individually while seated in a dimly lit room on a chair whose height could be adjusted so that their eyes were roughly level with the center of the screen. An experimenter was in the cubicle with them, seated at a different table, controlling the eye tracker. The experimenter explained the procedure to the children and informed them that they would be asked questions at the end of each clip. Adults viewed the videos at the Institute of Ophthalmology in a dimly lit room. In all cases, no chin rest was used. Stimuli ~~~~~~~ All participants watched two 5-min video clips and one 3-min video clip with their corresponding soundtracks to enhance comprehension of, and engagement with, movie content. One video was a cartoon (Roadrunner) and two others were taken from popular live-action children’s movies (Night at the Museum and Elf). In terms of shot segmentation, a cut is an abrupt transition from one shot to another that greatly affects visual exploration (Garsoffky, Huff, & Schwan, 2007; Smith, Levin, & Cutting, 2012). In the following, analyses were performed on each individual shot (see Table 1). As in Coutrot, Guyader, Ionescu, and Caplier (2012), shots were automatically detected using a pixel-by-pixel correlation value between two adjacent video frames. We ensured that the shot cuts detected were visually correct. Eye-tracking procedure ~~~~~~~~~~~~~~~~~~~~~~ Calibration routines were run at the start and end of each of the three clips with a custom-made cartoon character (size = 1°) appearing at each of nine positions: the four corners, the four midpoints of the edges, and the center of the movie frame. Each data sequence was used to parse the x/y position signal into saccades and fixations/smooth pursuits with a custom algorithm based on Nyström and Holmqvist (2010) (see below). In adults, eye-tracking data were successfully recorded in 10 of 10 adults across all clips. For the children, no data could be collected for 2 of 37 children (both in the younger age group) because the eye tracker failed to locate and track their pupil. Intermittent loss of the eye-tracking signal (due to either body or head motion while watching the clips) led to data from a further 3 children (2 in the younger group and 1 in the older group) missing from two clips, and 6 children (1 from the younger group and 5 from the older group) had missing data from one clip. Data pre-processing ~~~~~~~~~~~~~~~~~~~ If the calibrations at the start and end of a clip were different (by visual inspection), data from the corresponding clip were not used. For Clip 1 (Elf), 6 younger and 6 older children were excluded; for Clip 2 (Night at the Museum), 11 younger and 7 older children were excluded; and for Clip 3 (Roadrunner), 10 younger and 9 older children were excluded. This left 10 adults, 14 younger children, and 11 older children with usable data for Clip 1; 10 adults, 9 younger children, and 10 older children for Clip 2; and 10 adults, 10 younger children, and 8 older children for Clip 3. Data where “start” and “end” calibrations matched but there was a loss of signal for part of the clip were used, with the corresponding signal-less sections cut out. Saccades were identified based on the Nyström and Holmqvist (2010) iterative procedure that uses eye velocity and acceleration to set thresholds. Any period when velocity or acceleration exceeds an upper threshold for a minimum amount of time (10 ms) was classified as a saccade. The start- and end-points of the saccade were identified by going backward or forward in time until both the velocity and acceleration fell below a lower threshold. Saccades were identified to distinguish periods of fixation or smooth pursuit. The eye position data are missing during a blink, but there is also a period immediately before and after the blink when the dynamic occlusion of the pupil by the eyelid generates large, spurious motion signals in the eye- tracking data. The start- and end-points of blinks were identified using the lower threshold, as above, and all data during a blink were removed from the analysis. Estimating eye-tracking data quality is critical because systematic differences in the quality of raw eye position data can create a false impression of differences in gaze behavior between groups. This is particularly true when comparing children and adults because children are more prone to postural change during testing than adults, leading to generally poorer or more variable eye-tracking data. For this reason, we assessed our eye- tracking data using the precision metric proposed by Wass, Forssman, and Leppänen (2014). This metric quantifies the degree to which eye positions are consistent between samples (the higher the metric, the less precise the eye data). For each participant, we averaged this metric across the three video clips for an overall measure of precision.","Data were analyzed in two ways. First, we examined the variability between eye movements across observers for age-related changes, that is, an analysis based entirely on eye- tracking data (a gaze-based analysis). Second, we compared the eye-tracking data with the movie content using standard low-level (salience-based) models as well as a model based entirely on faces (a content-based analysis). Variability between participants: Gaze-based analysis As mentioned in the stimuli description, we performed our analyses on each frame of each individual shot. Interobserver dispersion To estimate the variability of eye positions between observers, we used a dispersion metric. This metric is commonly used in eye-tracking studies (Marat et al., 2009). For a given frame watched by n observers ((xi, yi) i∈[1, 2, ... n] the eye position coordinates), the dispersion D is defined as follows: The dispersion is the mean Euclidian distance between the eye positions of different observers for a given frame. The smaller the dispersion, the less scatter in the eye positions. Cluster number To quantify the number of points of interest attracting observers’ gaze, we performed a cluster analysis. For each frame, we clustered observers’ eye positions with the mean shift algorithm (Fukunaga & Hostetler, 1975). This algorithm considers feature space as an empirical probability density function. Mean shift associates each eye position with the closest peak of the dataset’s probability density function. For each eye position, mean shift defines a circular window of width w around it and computes the mean of the data points. Then, it shifts the center of the window to the mean and repeats the algorithm until it converges. This process is applied to every observer. All eye positions associated with the same density peak belong to the same cluster. The main advantage of mean shift compared with other popular clustering algorithms such as k-means is that it does not make assumptions about the number or shape of clusters. The only parameter to be tuned is the window width. Here, we took w = 100 pixels. Other w values between 50 and 200 pixels did not significantly change the results. Faces-based model We compared variants of a saliency model against a faces-based model. For each clip, the faces were labeled using a custom MATLAB script. When a face first appeared on the screen, an ellipse was drawn around it. When the face moved during the scene (due to either person or camera movement), the position, orientation, aspect ratio, and size of the ellipse were updated in a small number of “keyframes” and interpolated in between times. This led to a relatively fast and accurate labeling of the whole clip. Points within these ellipses were set to 1, and outside they were set to 0, to produce a binary map that was then filtered with a three-dimensional spatiotemporal Gaussian filter (spatial SD = 1.06°, temporal SD = 26.25 ms) similar to that used by Dorr et al. (2010) to produce a “face map” akin to the saliency maps described below. Given the sparsity of visual information in the Roadrunner cartoon, for this clip only we also labeled objects in a similar manner, with the exception being that the outlines could be polygons as well as ellipses. Saliency models Three variants of the Itti et al. (1998) saliency model were applied to each movie: (a) Itti et al. (1998) and (b) Harel, Koch, and Perona (2006) graph- based visual saliency (GBVS) with motion/flicker channels and (c) Harel et al. (2006) GBVS without motion/flicker channels. All saliency models seek to extract regions that somehow differ from their surroundings. Itti et al. (1998) separated the image into three different channels based on luminance, color, and orientation, and within each channel they identified regions of the image that differ from their immediate surround. These individual maps were then linearly combined to produce an overall map highlighting unusual regions, with the implicit assumption that these regions are likely to draw our attention. The GBVS model similarly splits the image into separate channels (luminance, color, and orientation plus two dynamic channels, motion and flicker, based on differences between consecutive frames) and uses a graph- based approach to highlight unusual regions. Although the GBVS model is not as intuitive as the Itti et al. model, it has been shown to perform better for static images (Itti et al., 1998), and its incorporation of motion and flicker channels makes it appropriate for dealing with dynamic images. Comparison of model performance We used two methods of assessing the performance of each of the models. The first is the normalized scanpath salience (NSS), which has been extended to allow analysis of moving images (Dorr et al., 2010; Marat et al., 2009), and the second is the area under a receiver operating characteristic (ROC) curve (Green & Swets, 1966) as proposed for eye-tracking analysis by Tatler, Baddeley, and Gilchrist (2005). Normalized scanpath salience On a frame-by-frame basis, the salience (or face) maps were normalized by subtracting the mean and dividing by the standard deviation (face maps for frames in which no faces were present were set to 0; see Fig. 1). Regions where the model predicts a high probability of fixation will have positive values, and low-probability regions will be negative. The normalized saliency value for each fixation is then averaged across participants to give frame-by- frame model performance as well as across the duration of the movie to give an overall model performance value. Correlation between outputs of a faces-based model and saliency models To quantify any overlap between face maps and bottom-up saliency maps, we calculated the correlation between a faces-based model and saliency models. For each frame, we computed a pixelwise correlation between bottom-up saliency maps (Harel et al., 2006, GBVS and GBVS + motion/flicker; Itti et al., 1998) and face maps. Frames without a face have been discarded from the analysis. Area under the ROC curve The area under the curve (AUC) is often used as a measure of model performance. We threshold the saliency or face maps at different levels to find regions of predicted fixations for that particular threshold. By comparing these regions with where the actual fixations occurred, we can extract the proportion of “true positives” (proportion of fixations within the predicted region) and the proportion of “false positives” (pixels that the model highlighted but were not fixated). By varying the threshold between 0 and 1, we produce an ROC curve (see panel in Fig. 4 in Results). The area under this curve is used as a measure of model performance. A value close to 1 indicates that the model explains the data well, whereas chance performance is .50. We derived confidence intervals for the AUC values via bootstrapping. Although an AUC analysis has a well-defined upper bound of 1, it is more appropriate to extract an empirical upper bound from the eye-tracking data themselves. A low AUC score could be due to a poor model, high variability between the looking strategies of different people, or a combination of the two. A modification of Peters, Iyer, Itti, and Koch’s (2005) NSS can be used to overcome this problem (Dorr et al., 2010). The NSS is calculated by using a “leave one out” approach—taking the scanpaths of n − 1 observers, spatiotemporally blurring these, and summing and normalizing them to produce a map of where these n − 1 people fixated over time. This map is evaluated as a predictor of where the nth person fixates, using the area under an ROC curve as above. This process is repeated n times, once for each observer, and the NSS score is the average AUC value. We used a similar technique to compare both between and within groups, that is, using all of the younger children’s fixations to build a fixation map and using this map to predict either the older children’s or the adults’ fixations. Statistical analyses To compare our findings for different models, age groups, and movies, we performed one-way and two-way analyses of variance (ANOVAs) as appropriate. For any significant differences between conditions, we also performed pairwise comparisons using t tests with Bonferroni correction for multiple comparisons. Eye-tracking data variability ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Using Wass et al. (2014) precision metric, we found no significant effect of age on precision: one-way ANOVA, F(2, 45) = 1.52, p = .23 (younger children: precision = .92, SEM = .19; older children: precision = .92, SEM = .19; adults: precision = .46, SEM = .03). Interobserver dispersion The shape of interobserver dispersion curves depicted in Fig. 2 (left panels) is conventional (Coutrot & Guyader, 2014). During the first 200 ms after a cut, dispersion is stable. During this period, observers’ gaze stays at the same locations as before the cut (latency period). Following this, dispersion decreases until 400 ms and slightly reincreases up to 1 s. This leads to the last stage, where the dispersion plateaus around a mean stationary value until the next cut. We ran a two-way ANOVA on interobserver dispersion with age (adults, older children, or younger children) and clip (Elf, Night at the Museum, or Roadrunner) as factors. There was a main effect of age, F(2, 707) = 20.76, p < .001, but not of clip, F(2, 707) = 2.87, p = .06. Post hoc Bonferroni tests showed that the dispersion was lower in adults than in older children, t(470) = −8.14, p < .001, or younger children, t(470) = −8.43, p < .001. There was no significant difference between the latter groups, t(470) = −0.29, p = .96. The interaction was also significant, F(4, 707) = 9.89, p < .001. For the Roadrunner clip, dispersion values were lower in younger children than in adults, t(74) = −2.20, p = .03, or in older children, t(74) = −2.68, p = .009. There was no difference between older children and adults, t(74) = 0.37, p = .71. Distance to center The global shape of distance to center is similar to the interobserver dispersion (Fig. 2, right panels). The stronger center bias around 200 ms after a cut has been reported previously (Wang, Freeman, Merriam, Hasson, & Heeger, 2012). It is due to a conjunction of factors, including the fact that the center of a scene is the optimal location to begin exploration (Tatler, 2007; Tseng, Carmi, Cameron, Munoz, & Itti, 2009). We performed a two-way ANOVA on distance to center with age and movie as factors. There was a main effect of age, F(2, 707) = 5.33, p = .005, and movie, F(2, 707) = 10.69, p < .001. Post hoc Bonferroni tests showed that the distance to center was significantly lower in adults than in older children, t(470) = −3.39, p = .002, or in younger children, t(470) = −4.00, p < .001. There was no significant difference between the latter groups, t(470) = −0.61, p > .90. Number of clusters The number of clusters quantifies the number of points of interest attracting observers’ gaze. Because this number is likely to increase with the number of observers, we normalized it by the number of participants. The number of clusters is stable across time except for a brief increase at around 200 ms after the beginning of the shot (Fig. 3). This increase can be explained by looking at the right panel of Fig. 3, where the temporal evolution of the number of recorded observers is depicted. We clearly see a loss of approximately 20% of recorded participants between 200 and 300 ms. This might be due to blinks induced by sharp cuts, leading to a brief loss of eye-tracking signal. Because the number of clusters is normalized by the number of participants, a decrease in the latter logically causes an increase in the former. We ran a two-way ANOVA on the number of clusters normalized by the number of observers with age and clip as factors. There was a main effect of age, F(2, 707) = 90.87, p < .001, and clip, F(2, 707) = 91.59, p < .001. Post hoc Bonferroni tests showed that for Elf and Night at the Museum, the number of clusters was significantly lower in adults than in older children, t(394) = −5.26, p < .001, or in younger children, t(394) = −11.76, p < .001, and was significantly lower in older children compared with younger children, t(394) = −7.16, p < .001. For the Roadrunner clip, the number of clusters was still lower in adults than in older children, t(74) = −8.40, p < .001, or in younger children, t(74) = −4.83, p < .001, but surprisingly it was higher in older children compared with younger children, t(74) = 4.32, p < .001.","Fig. 4 shows the model performances and the within- and between-group scanpath similarities for the three clips, broken down by age group (younger children in black, older children in red, and adults in blue). We performed a three-way ANOVA on the AUC with age, clip, and model and their interactions as explanatory variables. Note that the models included in this analysis were the three salience models: Harel et al.’s (2006) GBVS model, either with (GBVS+M+F) or without (GBVS) motion and flicker components, and Itti et al.’s (1998) multichannel saliency model (I&K). The faces-based model and the within-group NSS model showed that there were significant differences in AUCs across age group, F(2, 44) = 158.77, p < .001, clip, F(2, 44) = 124.92, p < .001, and model, F(4, 44) = 520.43, p < .001, as well as their interactions [Age × Movie, F(8, 44) = 8.37, p < .001; Age × Model, F(8, 44) = 4.06, p = .008; Movie × Model, F(8, 44) = 67.22, p < .001]. Post hoc Bonferroni tests showed that there were significant differences between the adults’ data and both the younger and older children’s data [younger t(16) = 15.30 and older t(16) = 15.60, both ps < .001], but there was no significant difference between the two children’s groups, t(16) = 0.26, p > .90). Post hoc Bonferroni tests also showed that there were significant differences between the results for Elf and each of the other two clips [Elf vs. Night at the Museum, t(16) = 13.70, and Elf vs. Roadrunner, t(16) = 13.60, both ps < .001], but not between Night at the Museum and Roadrunner, t(16) = 0.12, p > .90. Similarly, there were significant differences between each of the models except for the two GBVS models (with and without motion and flicker) [GBVS vs. GBVS+M+F: t(16) = 3.00, p = .085; all other pairwise model comparisons: t(16) > 7 and p < .001]. It is clear that for all three clips, the models predict the adults’ data better than the children’s data. A three-way ANOVA comparing the within- and between-age group NSS data for the three clips showed that there were significant differences in age groups in terms of how well their data could be predicted by data from another group (including its own), F(2, 26) = 27.80, p < .001, with post hoc Bonferroni tests showing that this was due to adults’ data being better predicted than those of either group of children [adults vs. younger children: t(20) = 7.00, p < .001; adults vs. older children: t(20) = 5.70, p < .001; younger children vs. older children: t(20) = 1.40, p = .5610]. However, there were no significant differences in how well each age group could be used to predict another age group (including its own), F(2, 26) = 1.73, p = .203. Correlation coefficients between the faces-based and salience models ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We found that correlations between the faces-based model and the bottom-up salience models were lowest for Elf (M = .31, SD = .001), followed by Night at the Museum (M = .35, SD = .001) and Roadrunner (M = .50, SD = .002). By way of comparison, the mean correlation coefficients between the maps of two bottom-up saliency models (Harel et al., 2006 [GBVS]; Itti et al., 1998 [I&K]) were .88, .90, and .87 for the respective three movies. Fig. 5 shows the normalized saliencies in the 2 s after a cut (averaged across all cuts in each clip and all observers in each age group; shaded regions show the 95% confidence intervals). All models perform relatively poorly immediately after a cut and rise to a peak at about 500 ms. The difference between the faces-based model and the salience models is most pronounced for Elf.","We tracked the eye movements of children and adults viewing dynamic stimuli (movie and cartoon clips). We performed two types of analyses on the resulting data: comparing fixation within and between age groups (gaze-based analysis) and examining how well fixation could be predicted from low-level (salience) or top-down (faces) features of the movies (content-based analysis). We report the following results. First, variance between eye movements decreases with age, consistent with comparison of eye movements in infants and young children (Franchak et al., 2016; Kirkorian et al., 2012). Second, all models that we tested predicted adults’ performance better than children’s performance, again suggesting greater homogeneity in eye movements in adults. Third, a face-based model performs at least as well as the low-level saliency models and significantly outperforms them for one of our clips (Elf). This difference in performance among the three clips highlights the importance of using different types of stimuli in testing for visual attention and is consistent with Dorr et al. (2010), who reported that interparticipant variability in adult eye movements is greater using natural stimuli than using commercial clips. For gaze-based analyses, the results showed that adults’ data were less variable than children’s data. The interobserver dispersion and number of clusters per person were lower for adults than for younger and older children, suggesting that adults look at a smaller number of objects or areas of interest when fixating a dynamic scene. The NSS analysis shows that the adults’ data are better predicted by other people’s data [whether those people are adults (within age groups) or children (between age groups)] than the children’s data. However, the adults’ data do not predict the children’s data (between) any better than the children’s data predict themselves (within). These findings suggest that although children on average look at more areas of interest in a scene (i.e., larger dispersion and numbers of clusters), there is significant overlap in the regions children and adults find most interesting. Because young children (4 or 5 years) have been found to have difficulty in maintaining accurate fixation (Kowler & Martins, 1982), the fixation heat maps for different age groups would be expected to be flatter for children (although they may have peaks in the same areas if they were looking at the same things as adults). This may account for some of the variability seen with our younger age group (the youngest child was 6 years old, only slightly older than the 4- and 5-year-olds in Kowler and Martins’s [1982] study), but given that fixation accuracy is likely to be more adult-like for the older children (10–14 years), this would not explain most of our results. Two recent studies have extended eye-tracking analysis to dynamic stimuli in children. Kirkorian et al. (2012) showed 1-year-olds, 4-year-olds, and adults a 19.5-min clip of Sesame Street and performed a gaze-based analysis. They reported reduced variability with age that they linked to increased influence of top-down mechanisms. Our results are largely consistent with this finding when extended to children in their teens (older child group). Interestingly, we also found that roughly 200 ms after a shot, there is an increase in the number of clusters (in all age groups) associated with a corresponding decrease in the number of participants. The most likely explanation for this is that scene cuts lead to eye blinks (Nakano, Yamamoto, Kitajo, Takahashi, & Kitazawa, 2009). Franchak et al. (2016) showed adults and young infants (up to 24 months) a 60-s clip of Sesame Street and also looked at interparticipant variability in eye movements. They reported that younger infants’ eye movements were weakly correlated with those of adults but that this interparticipant correlation increased with age (24 months). Furthermore, correlations between adults and infant groups were no greater than correlations within the infant group, which is consistent with the data presented here. For content-based analyses, we found that extending static models of faces (Frank et al., 2012) to dynamic stimuli leads to performance as good as, or better than, any of the low-level salience models of eye movements that we implemented (Harel et al., 2006; Itti et al., 1998). Comparing performance of the different models, we found that the Itti et al. (1998) model generally performed poorest and the faces-based model performed much better than the saliency models for Elf and is approximately as good as the best-performing salience model for Night at the Museum and Roadrunner. We also found that when there was only a weak correlation between our face maps and the low-level salience maps (as in Elf), there was a significant performance difference between the face and saliency models, with faces being more predictive of gaze behavior. If the faces based model performed well simply as a result of faces being highly salient parts of the image, we would expect the faces-based model to perform best when the correlation between faces and salience is highest. This is not what we observed. Therefore, we do not believe that the impressive performance of the faces-based model is simply due to faces being more salient than other objects. Interestingly, the GBVS model, which explicitly builds in motion and flicker components (and so might be expected to perform better on dynamic movies), does not appear to outperform the GBVS model that omits them. Indeed, for the cartoon clip, the opposite pattern of results was observed. This may be due to the style of animation, whereby a sparse scene made up of large blocks of uniform color typically contains one or two foreground objects of interest that are static in the frame (but moving in the world) and a number of (irrelevant and physically static) background objects that move across the screen behind them. It is worth noting that, surprisingly, the low-level models perform relatively well on these complex dynamic stimuli (average AUC = 80% vs. average AUC for faces-based model = 85%). A possible explanation for the lack of superiority of the face- based model for two of the three clips is that movies often contain more than one person in a scene, and this increases the number of possible face targets (also suggested by Franchak et al., 2016). Surprisingly, we found no clear spike in performance for the salience models shortly after a cut, which may have been expected in light of findings for static images where salience performs particularly well for the first fixation after an image is shown (Carmi & Itti, 2006; Parkhurst et al., 2002), although other authors argue that the effect is mainly an artifact of a center bias. In our results, there is no clear evidence of a decline in saliency model performance over time, which often occurs for dynamic stimuli (Carmi & Itti, 2006; Marat et al., 2009). It is likely that in our dynamic stimuli, the constant appearance of new salient regions promotes bottom-up influences at the expense of top-down strategies, inducing a stable consistency between participants over time. Overall, our results are consistent with earlier reports using young infants. For example, Kwon et al. (2016) reported an increase in the eye movements to faces in static images, and Frank et al. (2012) found an increase in fixations to socially relevant information in brief videos. Recently, Franchak et al. (2016) compared consistency of eye movements with saliency and the presence of faces in 1-min clips from Sesame Street, and although they reported that on average fixations were to the top quartile of salient regions, they noted that a model that relies on both saliency and faces accounts for (only) 41% of the variance in infants’ eye movements. However, the different age ranges, stimuli, and analyses in our study and theirs mean that we cannot directly compare our results. It is noteworthy that our procedure relies exclusively on modeling/quantifying eye-tracking behavior and, as such, is agnostic as to what may underlie these age-related changes. We postulate that most changes will reflect attentional orienting rather than motor or visual immaturities because we are not testing young infants (Farber & Beteleva, 2005). In fact, the youngest children tested were 6 years old. Although saccade latencies decrease with age (Fukushima, Hatta, & Fukushima, 2000; Salman et al., 2006), their accuracy and peak velocity is adult-like by 8 years of age. Smooth pursuit is age dependent. However, for targets moving at roughly 15°/s or lower, children’s pursuit of targets is adult-like (age 5 years and over) (Ego, Orban de Xivry, Nassogne, Yüksel, & Lefèvre, 2013). Given that, other than brief sections of the Roadrunner cartoon, most clips did not contain rapidly moving objects, we believe that any smooth pursuit immaturities would have a minimal influence on our results. Finally, we note that the dependence of our outcomes on the particular movie sample highlights the importance of using stimuli that vary in their semantic/saliency content. This will be particularly important when applying such techniques to study pseudo-naturalistic fixation behavior in clinical populations. For example, the presentation of people with autism spectrum disorder differs widely, and it may be that patterns of abnormal fixation in some classes of movies can provide pointers to classification of autism spectrum disorder subtypes. In summary, we have examined age-related changes in children and adults’ eye movements to dynamic visual scenes, namely popular child-appropriate cartoons and movies. We extended previous work in infancy and early childhood and showed that (a) variance in eye movements decreased with age; (b) a faces-based model outperformed saliency models, even for the youngest children; and (c) when there was an increased differentiation between faces and salient regions in a scene, participants looked more to the faces as in Elf. These findings shed light on the nature of visual attention during development and act as an important reference for understanding how such attention may develop differently in individuals with neurodevelopmental conditions."],["Inner speech is a commonly experienced but poorly understood phenomenon. The Varieties of Inner Speech Questionnaire (VISQ; McCarthy-Jones & Fernyhough, 2011) assesses four characteristics of inner speech: dialogicality, evaluative/. motivational content, condensation, and the presence of other people. Prior findings have linked anxiety and proneness to auditory hallucinations (AH) to these types of inner speech. This study extends that work by examining how inner speech relates to self-esteem and dissociation, and their combined impact upon AH-proneness. 156 students completed the VISQ and measures of self-esteem, dissociation and AH-proneness. Correlational analyses indicated that evaluative inner speech and other people in inner speech were associated with lower self-esteem and greater frequency of dissociative experiences. Dissociation and VISQ scores, but not self-esteem, predicted AH-proneness. Structural equation modelling supported a mediating role for dissociation between specific components of inner speech (evaluative and other people) and AH-proneness. Implications for the development of \"hearing voices\" are discussed. © 2014 The Authors. --------------------------------------------------------------------------------","Inner speech – the internal monologue that appears to accompany our daily lives – is an experience that will be familiar to many. Often, inner speech is simply defined as a silent form of speech; for example, Levine, Calvanio, and Popovics (1982) call it the “subjective phenomenon of talking to oneself, of developing an auditory–articulatory image of speech without uttering a sound” (p. 391). But the everyday nature of inner speech hides a wealth of complexity (Hurlburt, Heavey, & Kelsey, 2013). Developmentally, inner speech has been proposed to reflect the internalisation of external dialogue and self- directed private speech, facilitating higher order cognitive skills (Vygotsky, 1987). In adulthood, inner speech supports executive functions (Miyake, Emerson, Padilla, & Ahn, 2004) and facilitates self-evaluation and reflection (Morin & Michaud, 2007). It has also been argued that the form of inner speech in adulthood is shaped by its developmental origins. For example, Fernyhough (2004) has argued that inner speech is inherently dialogic, reflecting the external interpersonal discourse from which it came. Inner speech may also change its form as it develops, becoming syntactically and semantically abbreviated or “condensed”. In this sense, inner speech is likely to be more than simply silent, covert speech. Rather, it is a complex and dynamic activity, and one that needs to be examined in more depth. Various methods of investigating inner speech exist, including questionnaires (Duncan & Cheyne, 1999; Morin, Uttl, & Hamper, 2011), experience sampling (Hurlburt et al., 2013) and task-based methods (Murray, 1967). However, few have included the various developmental characteristics of inner speech, such as dialogicality and condensation, as well as the functions of inner speech, such as action control (Luria, 1961), in a comprehensive measure of everyday experience. One exception is the Varieties of Inner Speech Questionnaire (VISQ; McCarthy-Jones & Fernyhough, 2011), a scale that measures self-reported inner speech along four distinct dimensions: (i) dialogicality, the communicative and conversational quality of inner speech; (ii) condensation, or the extent to which inner speech is syntactically and semantically abbreviated, (iii) evaluative/motivational inner speech, such as saying “I should do this” to oneself, and (iv) the presence of other people’s voices in inner speech. An initial validation of the VISQ (McCarthy-Jones & Fernyhough, 2011) demonstrated two advantages of this approach. First, the VISQ can provide a more nuanced picture of the range of ways in which inner speech is used and experienced; in data from a student sample, dialogic and evaluative characteristics of inner speech appeared to be very common, while a substantial minority reported experiences of condensation and other people’s voices in their inner speech (McCarthy-Jones & Fernyhough, 2011). Secondly, given that inner speech is proposed to be the raw material of some auditory verbal hallucinations (‘hearing voices’) (Bentall, 1990; Frith, 1992; Shergill et al., 2003), and may play a causal or maintenance role in anxiety and depression (Harrington & Blankenship, 2002), the VISQ can be used to examine how specific types of inner speech may be related to psychopathological processes. In their original study, McCarthy-Jones and Fernyhough found that self-reported evaluative inner speech and the presence of other people in inner speech were positively related to anxiety. Alongside this, proneness to auditory hallucinations (AH) had bivariate correlations with evaluative, other people and dialogic inner speech, although after controlling for potentially confounding factors, only dialogic characteristics remained a significant predictor of AH-proneness (McCarthy-Jones & Fernyhough, 2011). The present study sought to extend this work by examining the interrelations between inner speech and two other constructs that have been related to the development and experience of psychosis-like phenomena: self-esteem (Krabbendam et al., 2002) and dissociation (Altman, Collins, & Mundy, 1997). Self-esteem refers to the set of positive and negative beliefs and feelings that people have about themselves. It has been related to the severity and content of hallucinations in clinical studies (Romm et al., 2011; Smith et al., 2006) and has been proposed to shape the interactions people have with the voices they hear (Paulik, 2012). In work with the general population, self-esteem has been observed to predict later development of psychosis (Krabbendam et al., 2002), and has been associated with increased hallucination-proneness (Gaweda, Holas, & Kokoszka, 2012; Gracie et al., 2007; Pickering, Simpson, & Bentall, 2008). Dissociation refers to a “lack of normal integration of thoughts, feelings and experiences into the stream of consciousness and memory” (Bernstein & Putnam, 1986, p. 727), and is typified by feelings of depersonalization, derealisation and absorption. Diagnostic overlaps between dissociative and psychotic disorders have often been noted (Allen & Coyne, 1995), and dissociation is strongly related to rates of psychotic experiences, including hallucinations (Altman et al., 1997; Glicksohn & Barrett, 2003; Kilcommons & Morrison, 2005; Merckelbach, Rassin, & Muris, 2000). More recently, specific links between dissociation and voice-hearing have been proposed (Moskowitz & Corstens, 2008), with dissociative experiences potentially playing a predisposing role or acting as a preliminary stage in the development of AHs (Perona-Garcelán et al., 2011; Varese, Barkus, & Bentall, 2012). In addition to their documented association with hallucinations, there are prima facie reasons to think that both self-esteem and dissociation might overlap with specific aspects of inner speech. Given its proposed role in self-assessment and reflection, evaluative aspects of inner speech would appear relevant to an individual’s self-concept and social rank. The presence of other voices in inner speech, in contrast, has a phenomenological similarity with dissociative experiences of depersonalization and detachment; a separation of the self and other. This can be seen in example items from the VISQ, such as “I experience the voices of other people asking me questions in my head” (McCarthy-Jones & Fernyhough, 2011). As such, the first aim of this study was to establish what links, if any, there are between varieties of inner speech, self-esteem and dissociative traits. We hypothesised that evaluative inner speech scores in particular would be strongly associated with self-esteem, while other people in inner speech would be closely associated with dissociation scores. The second aim of this study was to investigate how any potential relations between the varieties of inner speech, dissociation and self-esteem might influence the previously documented association between inner speech and AH-proneness (McCarthy-Jones & Fernyhough, 2011). If self-esteem and dissociation were to be found to be related to both inner speech and AH-proneness, then they could represent potential confounding variables, or act as either mediators or moderators of the link between inner speech and AH-proneness. Dissociation in particular has been proposed to act as a mediator in the development of hallucinations. In patients with schizophrenia spectrum disorders, dissociative traits positively mediate the effects of childhood trauma (Varese et al., 2012), and attentional focus (Perona-Garcelán et al., 2011) on hallucination proneness. As it is currently unclear how inner speech may become transmuted into AHs, an examination of whether dissociation and/or self-esteem are potential mediators is of considerable potential value. This was accomplished by testing the fit of a number of potential models of the relation between self-esteem, dissociation and the varieties of inner speech using structural equation modelling.","A sample of 156 students (22 male) was recruited from a United Kingdom university. The mean age was 20.02 years (SD = 2.11, range 18–31). Participants were invited to take part via a web advert and received course credit for participation. Participants indicated their age and gender and provided an email address if they wished to be contacted for a follow-up study (which is not reported here). All procedures were approved by the local university ethics committee. Varieties of Inner Speech Questionnaire (VISQ; McCarthy-Jones & Fernyhough, 2011) The Varieties of Inner Speech Questionnaire is an 18-item self-report scale measuring four dimensions of inner speech. Participants endorse items about their experience of inner speech in relation to dialogicality (“I talk back and forward to myself in my mind about things”), evaluative and motivational characteristics (“I talk silently to myself telling myself not to do things”), condensation (“I think to myself in words using brief phrases and single words rather than full sentences”) and the presence of other voices (“I experience the voices of other people asking questions in my head”). Responses are made on a 6-point Likert scale ranging from “Certainly does not apply to me” (1) to “Certainly applies to me” (6). Each subscale of the VISQ has high internal reliability (Cronbach’s α > .80) and moderate to high test–retest reliability (>.60). Rosenberg Self-Esteem Scale (RSES; Rosenberg, 1965) The Rosenberg Self-Esteem Scale is a widely used 10-item measure of global self- esteem with high test–retest and internal reliability (α = .8–.9; Fleming & Courtney, 1984; Robins, Hendin, & Trzesniewski, 2001). Respondents use a four- point scale (strongly agree, agree, disagree and strongly disagree) to classify evaluative statements about themselves, such as “On the whole, I am satisfied with myself”. Dissociative Experiences Scale – Second Revision (DES-II; Carlson & Putnam, 1993) The Dissociative Experiences Scale measures frequency of dissociative experiences on a 28-item scale. Participants indicate what percentage of the time they experience dissociative states ranging from 0% to 100%, e.g., “Some people have the experience of finding themselves in a place and have no idea how they got there. What percentage of the time does this happen to you?”. The original DES has been shown to have good reliability and validity; test–retest reliabilities range from .7 to .9 (Carlson & Putnam, 1993) with a mean internal reliability of .93 (Van IJzendoorn & Schuengel, 1996). The DES-II differs only in the use of incremental response options rather than a visual analogue scale. Revised Launay–Slade Hallucination Scale: Auditory Subscale (LSHS-R; Morrison, Wells, & Nothard, 2000) The Launay–Slade Hallucination Scale is a commonly used measure of hallucination proneness in non-clinical samples. The revision used here (Morrison et al., 2000) was developed specifically to separate factors relating to auditory and visual hallucination proneness. Five items relating specifically to auditory hallucination experiences (e.g., “I have had the experience of hearing a person’s voice and then found that there was no one there”) were used here following their use in McCarthy-Jones and Fernyhough (2011), in which they were observed to have good internal reliability (Cronbach’s α = .73). Items were scored on four-point scale ranging from “Never” (1) to “Almost Always” (4).","Relationships between the VISQ and other variables were first analysed using Pearson’s product moment correlation co-efficients. Secondly, hierarchical regression was used to assess the ability of inner speech, self-esteem and dissociation to predict hallucination proneness, with LSHS-R scores acting as the dependent variable. Structural equation modelling (SEM) was then used to test competing models of the relationship between the variables. All analyses were conducted in SPSS 20 and AMOS 7.0. Unless otherwise noted, all Durbin–Watson statistics and parameters regarding collinearity and normality of residuals were within acceptable bounds. A number of arguments have been made as to the most appropriate way in which to correct alpha to control for the altered experiment-wise error rate resulting from undertaking multiple comparisons. For example, it has been argued that a different approach is needed for exploratory and confirmatory studies (Bender & Lange, 2001). For confirmatory studies, Bonferroni corrections are generally viewed as appropriate (Bender & Lange, 2001). Thus, for the part of this study which involved confirmatory tests, i.e., which repeated the analyses of McCarthy-Jones and Fernyhough (2011) by examining four bivariate correlations (between the four VISQ subscales and AH-proneness), a Bonferroni corrected alpha was employed of α′ = .05/4 = .0125. For all other analyses, which were exploratory, α was retained at .05. This is in accordance with Perneger, 1998 argument that “simply describing what was done and why, and discussing the possible interpretations of each result, should enable the reader to reach a reasonable conclusion without the help of Bonferroni adjustments” (p. 1237).","Table 1 displays descriptive statistics for the four dimensions of the VISQ, as well as the RSES, DES and LSHS-R. Mean scores were highest for evaluative inner speech and dialogic inner speech, followed by other people in inner speech and condensed inner speech. If the presence of specific types of inner speech was defined as positive endorsement of the majority of items on each subscale (i.e., a response of 4 or higher), 80.1% of participants reported dialogic inner speech, 87.2% reported evaluative inner speech, 22.4% reported other people in inner speech, and 28.2% reported condensed inner speech. Internal reliability scores (Cronbach’s α) were good for each subscale (all >.70). Both mean scores and reliability co-efficients were comparable to the original VISQ data reported by McCarthy-Jones and Fernyhough (2011). Reliability for self-esteem and dissociation scores was high (α > .90). Cronbach’s α was lower for hallucination-proneness (0.69) although still well within the acceptable range (i.e., >0.6). Correlations of inner speech with self-esteem, dissociation and auditory hallucination-proneness ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 2 shows bivariate correlations among VISQ, RSES, DES and LSHS-R scores. Evaluative inner speech (p = .011) and other people in inner speech (p = .022) scores were significantly negatively correlated with global self-esteem scores, such that higher self- esteem was associated with less evaluative inner speech and fewer instances of other people in inner speech. Both other people and evaluative inner speech were also positively related to frequency of dissociative experiences (see table 2). AH-proneness was positively associated with dialogic inner speech (p < .001), evaluative inner speech (p < .001) and other people in inner speech (p < .001). Predicting proneness to auditory hallucinations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A hierarchical multiple linear regression was performed to assess the unique contribution of inner speech (VISQ), self-esteem (RSES) and dissociation (DES) in predicting auditory hallucination proneness. Age and gender were included in the first block as control variables. This was then followed by the four VISQ subscales in the second block (to replicate the findings from McCarthy-Jones & Fernyhough, 2011), RSES in the third block, and DES in the fourth block. All residuals were normal and measures of multicollinearity were in an acceptable range. While the first block was non-significant (p > .100), each of the second, F (6, 149) = 6.56, p < .001, third, F (7, 148) = 5.59, p < .001, and fourth, F (5, 214) = 8.67, p < .001, blocks significantly predicted AH-proneness. In Block 2 (R2 = .21), other people in inner speech, β = .250, p = .001, and evaluative inner speech, β = .198, p = .023, were the only significant predictors. This was also the case in Block 3 (R2 = .21); the addition of self-esteem made no significant change to the model, ΔR2 = .00, p = .873, and self-esteem failed to predict AH-proneness, β = −.012, p = .873. In contrast, the addition of DES scores in Block 4 significantly improved the model (ΔR2 = .11, p < .001, R2 = .32). Dissociation scores significantly predicted AH-proneness, β = .367, p < .001, while VISQ predictors were only evident at the level of trends for other people in inner speech, β = .134, p = .078, and condensed inner speech, β = −.130, p = .066. Thus, the predictive value of VISQ subscales was almost completely removed with the inclusion of dissociation scores. Mediator effects Structural equation modelling (SEM) in AMOS was used to further examine the relationship between inner speech, dissociation and AH-proneness. As the regression analyses showed self-esteem not to predict AH-proneness, this was excluded from the SEM. To allow for non-normal data distributions, an asymptotically free distribution (ADF) method was used combined with a Yuan–Bentler statistic, which is appropriate for sample sizes less than 200 (Yuan & Bentler, 1999). Three models were compared: (a) dissociation and inner speech scores as independent predictors of AH-proneness, (b) dissociation as a full mediator of the relationship between inner speech and AH-proneness, and (c) dissociation as a partial mediator of this relationship (see Fig. 1). Each model also contained covariate links between the VISQ domains where these were suggested by bivariate correlations. The first model represented a situation in which dissociation and inner speech present unique pathways towards increased hallucination proneness. The latter two models tested whether dissociation plays a role in mediating the relationship between inner speech and hallucination proneness, either by fully accounting for their covariance (full mediation) or only accounting for some of it (partial mediation). If dissociation was observed to fully mediate this relationship, inner speech could only be said to have an indirect effect on AH-proneness. Of the three models, only Model C (partial mediation) provided an adequate fit to the data, X2ADF (2) = 1.51, p = .471, TF (2, 154) = 0.75, p = .475, RMSEA = .000, GFI = .997, CFI = 1.00. Both Model A, X2ADF (6) = 20.69, p = .002, TF (6, 150) = 3.34, p = .004, RMSEA = .126, GFI = .962, CFI = .737, and Model B, X2ADF (6) = 16.47, p = .011, TF (6, 150) = 2.67, p = .017, RMSEA = .106, GFI = .969, CFI = .813 were significantly different from the data, indicating an inadequate fit. Bootstrap statistics were used to assess the significance of direct and indirect effects on AH-proneness. Direct effects on AH- proneness were observed for dissociation, p = .004, other people in inner speech, p = .042, and condensed inner speech, p = .025. A marginal but non-significant direct effect was also evident for evaluative inner speech, p = .063. Indirect effects on LSHS, via dissociation, were observed for other people in inner speech, p = .004 and evaluative inner speech, p = .006, but not condensed inner speech, p = .615. No direct or indirect effects were observed for dialogic inner speech (all p > .200). To further test the specific direction of these relationships, a model, D, was constructed that matched Model C but included VISQ scores as mediating variables and DES as their predictor (see Fig. 2). In such a scenario, dissociation would be associated with different characteristics of inner speech, and these in turn would lead to greater hallucination-proneness. However, this did not produce an adequate fit to the data, X2ADF (6) = 24.81, p < .001, TF (6, 150) = 4.00, p < .001, RMSEA = .142, GFI = .934, CFI = 0.664. Moderator effects To examine whether dissociation might moderate rather than mediate the effect of inner speech scores on hallucination proneness, a multiple linear regression was performed using the same predictors as above, plus mean-centred interaction terms between DES scores and the four VISQ subscales. This, however, yielded no new significant predictors of auditory hallucination proneness (all interaction effects p > .70, β < .01). Predictive relationships with overlapping items removed ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Finally, while aimed at nominally separable constructs, measures of inner speech, hallucination-proneness and dissociation can sometimes be very similar in the kinds of experiences they describe. It is therefore important to rule out any trivial relationships between these constructs that might be caused by individual items that happen to overlap. To address this, we reran each of our main analyses with subsets of items removed from our measures of inner speech and dissociation. A subset of items was removed from the VISQ and the DES (the RSES had little overlap with other scales). In the VISQ, two items explicitly refer to hearing other people’s “actual voices”; item 12 “I hear other people’s actual voices in my head, saying things they have never said to me before” and item 16 “I hear other people’s actual voices in my head, saying things that they actually once said to me”. Both these items were removed from the “Other People” subscale to avoid trivial overlaps with the LSHS-R. In the DES, two items were removed; one clearly overlapped with hallucination proneness and inner speech (item 27: “Some people sometimes find that they hear voices inside their head that tell them to do things or comment on things that they are doing.”) while another was deemed to be similar to questions about inner speech (item 21: “Some people sometimes find that when they are alone they talk out loud to themselves”). Each of the above analyses (correlations, regression and SEM) was rerun using the adjusted scores for other people in inner speech and dissociation. The new scores for other people in inner speech significantly correlated with LSHS-R (r = .24). RSES (r = −.21) and the adjusted DES (r = .31; all p < .05), albeit to a lesser degree than beforehand. When the new DES and other people VISQ scores were entered into the hierarchical regression, the results for each model were almost identical to those observed originally, with DES the single significant predictor in the final model (β = .382, p < .001). The adjusted score for other people in inner speech was a significant predictor in blocks 2 and 3, but was no longer evident (even at trend level) in the final model (p = .406). Re-analysis using SEM also produced similar results. When the adjusted scores for DES and other people in inner speech were included, only Model C yielded an appropriate fit to the data, X2ADF (2) = 2.79, p = .247, TF (2, 154) = 1.39, p = .253, RMSEA = .051, GFI = .995, CFI = .985. Bootstrap analysis yielded similar direct effects for DES (p = .004) and condensed inner speech (p = .033) on hallucination proneness, alongside a new direct effect for evaluative inner speech (p = .028). Indirect effects were evident for evaluative inner speech (p = .007) and the adjusted other people in inner speech (p = .008). Thus, removal of overlapping items (i) made little difference to the role of dissociation, (ii) reduced the role of other people in inner speech, but also (iii) highlighted a greater contribution of evaluative inner speech.","The present study examined the relationships among inner speech, self-esteem and dissociation, and how these factors combine to predict proneness to auditory hallucinations. The main finding was that dissociation mediated the relationship between inner speech quality and the tendency to experience auditory hallucinations. Specifically, dissociation mediated the effect of other people in inner speech and evaluative/motivational inner speech on auditory hallucination proneness. There was also some evidence to suggest direct effects of evaluative inner speech and condensed inner speech on proneness to auditory hallucinations. Correlational analysis indicated a relationship between self-esteem, evaluative/motivational inner speech and other people in inner speech, but no relationship was evident between self-esteem and hallucination proneness. Consistent with prior findings (McCarthy-Jones & Fernyhough, 2011), the four VISQ subscales showed good reliability and specific relationships with proneness to auditory hallucinations; dialogic and evaluative characteristics of inner speech and the presence of other voices in inner speech positively correlated with AH-proneness. However, although McCarthy-Jones and Fernyhough found dialogic inner speech to account for unique variance in auditory hallucination proneness (after controlling for all VISQ subscales, age, gender, anxiety, depression and visual hallucination-proneness), in this dataset it did not account for any unique variance in AH-proneness, at any stage of the regression or SEM analysis. Possible explanations for this are the collinearity of this particular domain of inner speech with other aspects of inner speech, and the use of different control variables (anxiety and depression) in McCarthy-Jones and Fernyhough’s original study. Given that evaluative inner speech and other people in inner speech correlate significantly with anxiety (McCarthy-Jones & Fernyhough, 2011), it may be that dialogicality is only predictive of hallucination-proneness once factors relating to anxiety are partialled out. Replication data are needed to clearly delineate which aspects of inner speech are the best predictors of AH-proneness and what role dialogic characteristics may play, if any. Dissociative traits were most strongly associated with the presence of other people’s voices in inner speech, although they were also linked to evaluative inner speech. This suggests that a greater tendency to experience other voices in everyday thinking may in itself be a dissociative tendency or could be a precursor of more developed dissociative states such as depersonalisation and identity confusion (Steinberg, 1991). When dissociation scores were included in the model for predicting AH- proneness, they accounted for some, but not all, of the predictive power of inner speech. The SEM analysis suggested that the best fit to the data involved a model where DES partially mediated the link between different subscales of the VISQ and hallucinations. Importantly, this was true even when items with potentially overlapping content were removed from the DES and VISQ, and highlighted a range of direct and indirect effects of inner speech, with other people in inner speech being fully mediated, evaluative inner speech partially mediated, and condensed inner speech appearing to show a small but unmediated effect. Given that some of the direct effects were only evident in the SEM and not in the hierarchical regression (particularly concerning condensed inner speech), these relations should be interpreted with caution and require further replication. Nevertheless, it appears that different characteristics of inner speech may be picking out independent and alternative routes to AH-proneness. These data support the role of dissociation as a mediating mechanism in the development of hallucinatory experiences. Varese et al. (2012) observed DES scores to positively mediate the effect of childhood trauma on hallucination-proneness in hallucinating patients. This finding has been replicated in a non-clinical university sample using alternative measures of two specific dissociative experiences, depersonalisation and absorption (Perona-Garcelán et al., in press). While such findings point towards a specific role for dissociation in the link between trauma and hallucination, the present data may suggest a wider mediating role for dissociative traits in the development of unusual perceptual experiences, as seen for example in data on attentional focus and hallucinations (Perona-Garcelán et al., 2011). In such a scenario, characteristics of inner speech could develop into hallucination via a dissociative stage, or there may be an additive effect between inner speech and dissociative traits that gives rise to hallucination proneness. Given that we did not observe any significant moderator effects in this dataset, these results appear to support the former scenario over the latter. The converse could of course be possible; it may be that having a tendency towards dissociation could lead to the development of specific kinds of inner speech – such as talking in other voices – and this in turn leads to the development of hallucinations. When this scenario was modelled with the VISQ scores as mediators (Model D) it did not produce an adequate fit to the data, which would suggest that inner speech does not play such a role. However, only longitudinal data on inner speech, dissociation and hallucination proneness could convincingly split apart these possibilities. Yet to be examined also are the potential relationships between earlier childhood experiences (including trauma and other adversity) and characteristics of inner speech. A tentative hypothesis would be that the presence of other voices and evaluative characteristics of inner speech are likely to have strong developmental roots, both based on their theoretical origin (Fernyhough, 2004) and links between early adversity and later hallucination proneness (Bentall, Wickham, Shevlin, & Varese, 2012). This is the first study to demonstrate links between specific varieties of inner speech and self-esteem. As predicted, a greater tendency to engage in evaluative forms of inner speech (“I talk silently to myself, telling myself not to do things”) was associated with lower self- esteem. Though such an association cannot demonstrate causality, it seems plausible to suggest that inner speech with a strongly normative and ruminative component could have a negative impact on self-concept and social rank. Equally, the former could be an expression of the latter, if lower self-esteem was taken to lead to more frequent evaluation and rumination in inner speech. In either case, this contrasts with the suggestion that evaluative aspects of inner speech generally play a positive role in enabling self-awareness and self-reflection (Morin & Everett, 1990); an increase in such tendencies could equally be associated with over-analysis and self-criticism. It is important here to consider the valence of the kind of inner speech that is occurring; some evaluations will be more positive (e.g., “I can do this”) and others negative (“I really shouldn’t have done that…”) and this will vary with context. The items in the VISQ are designed to be balanced in their appraisals, but further study with a greater range of positive and negative items is really needed to clarify the relationship between self- concept and inner speech use. Self-esteem was also negatively correlated with reports of other people in inner speech, indicating that experiencing other people’s voices in everyday verbal thinking is linked to lower self-esteem and a more negative self-concept. The lack of any relationship between self-esteem scores and AH-proneness contrasts with prior reports of self-esteem effects (Pickering et al., 2008) but is consistent with prior null findings in a similar population (Jones & Fernyhough, 2007). There are a range of caveats to be acknowledged in interpreting these results. First, these data come from a non-clinical, university sample and only indicate relationships between general population traits associated with psychosis. The correlations observed are broadly consistent with those reported in clinical studies (as in, for example, the case of dissociation and hallucination; Kilcommons & Morrison, 2005) but any extrapolation to clinical populations must be made with caution. Second, the present results are purely correlational and do not provide evidence of causal relationships between the data. Time-based data via longitudinal studies or use of experience sampling methods (ESM; Csikszentmihalyi & Larson, 1987) are required to assess causal relationships between these variables. Third, these data are reliant on self-reports for the measurement of experiences that may be hard to report reliably and accurately. Previously used inner speech questionnaires (although not the VISQ) have been criticised for showing a lack of convergent validity (Uttl, Morin, & Hamper, 2011) and it has been suggested that self-reported inner speech is likely to be overestimated in general questionnaires (Hurlburt et al., 2013). Similarly, the validity and reliability of self-reports for relatively infrequent, unusual phenomena such as hearing a voice or having a dissociative experience is hard to gauge in non-clinical samples, and the possibility of over-reporting cannot be ruled out. The mean DES score observed here, although comparable with other student/college studies (e.g., Merckelbach, Muris, Rassin, and Horselenberg, 2000) was relatively high, suggesting that some overestimation may have occurred. The internal reliability of hallucination-proneness scores, though in the adequate range, was also lower than the other measures used (α = .69). An alternative approach would be to use more items to examine hallucination- proneness, although that may be at the risk of losing specificity, as existing longer scales often conflate auditory experiences with visual hallucinations and other unusual phenomena (Morrison et al., 2000) The use of ESM is equally important here to examine whether day-to-day, random sampling of such experiences actually corroborates participants’ general trait reports. Alternatively, use of questionnaires and ESM could be combined with electromyographical techniques to assess covert muscle activity associated with inner speech (see, for example, Rapin, Dohen, Polosan, Perrier, & Loevenbruck, 2013). Finally, while the VISQ is an explicitly multi-dimensional scale, the measures used here for self-esteem, dissociation and hallucination-proneness represent unidimensional indices of complex phenomena. As the first study to investigate this group of constructs in relation to inner speech, we selected measures partly for simplicity and to limit the number of potential comparisons examined. In the case of hallucinatory experiences, auditory hallucinations were focused on here rather than visual experiences or other modalities because of the putative association of AH with inner speech (Fernyhough, 2004) and their proposed links to dissociation (Moskowitz & Corstens, 2008). Nevertheless, an important focus for future studies would be an assessment of the relationships among VISQ dimensions and separable aspects of dissociation (such as depersonalisation and absorption), and the specificity of such relationships for understanding auditory hallucinations.","In conclusion, these results highlight the importance of asking about a variety of experiences in inner speech. Different characteristics of inner speech relate to self- esteem, dissociation and AH-proneness in specific ways that would otherwise be missed by a less fine-grained analysis of how people talk to themselves. Further examination of these characteristics will shed light on how ordinary characteristics of day-to-day inner experience can, in some cases, become the extraordinary."],["Online learning environments are well-suited for tailoring the learning experience of children individually and on a large scale. An environment such as Math Garden allows children to practice exercises adapted to their specific mathematical ability; this is thought to maximize their mathematical skills. In the current experiment, we investigated whether learning environments should also consider the differential impact of cognitive load on children's math performance depending on their individual verbal working memory (WM) and inhibitory control (IC) capacity. A total of 39 children (8–11 years old) performed a multiple-choice computerized arithmetic game. Participants were randomly assigned to two conditions where the visibility of time pressure, a key feature in most gamified learning environments, was manipulated. Results showed that verbal WM was positively associated with arithmetic performance in general but that higher IC predicted better performance only when the time pressure was not visible. This effect was mostly driven by the younger children. Exploratory analyses of eye-tracking data (N = 36) showed that when time pressure was visible, children attended more often to the question (e.g., 6 × 8). In addition, when time pressure was visible, children with lower IC, in particular younger children, attended more often to answer options representing operant confusion (e.g., 9 × 4 = 13) and visited more answer options before responding. These findings suggest that tailoring the visibility of time pressure, based on a child's individual cognitive profile, could improve arithmetic performance and may in turn improve learning in online learning environments. --------------------------------------------------------------------------------","Extensive individual differences in learning trajectories show that in education there is no such thing as a one-size-fits-all approach. Adaptive e-learning systems, where an online learning environment is continuously adapting to accommodate differences between learners and changes over time for each individual (Park & Lee, 2003), may help to address this challenge and enhance children’s success. The idea behind this approach is that if pedagogical procedures are geared to adhere to their individual needs, students will be able to achieve a higher performance more efficiently (for a review, see Akbulut & Cardak, 2012). One example of such an adaptive e-learning system is Math Garden, an educational tool that adapts the difficulty of the math problems presented to children aged 4 years and over. The aim of Math Garden is that children always practice math skills at an appropriate individual level; in the case of Math Garden, items are chosen such that the probability of answering correctly is about .75 (Jansen et al., 2013; Straatemeier, 2014). In principle, emerging e-learning platforms allow the tailoring of the learning environment to individual students on a large scale. In contrast to the conventional classroom setting where teachers have a good sense of the pupils’ individual needs, in an e-learning context explicit information is required to reliably tailor the individuals’ learning environment based on these differences. Current adaptive e-learning systems such as Math Garden are well equipped to adapt to the specific math ability level of the student (Klinkenberg, Straatemeier, & van der Maas, 2011). However, the environmental context in online game-based learning environments, with its interruptions and distractions, poses a risk for the user in terms of sustained attention, engagement, and concentration (Terras & Ramsay, 2012). To maximize the learning potential offered by adaptive e-learning platforms, we also need to consider individual differences in the capacities to attend to, process, learn, and remember information when designing these technologies (Ramsay & Terras, 2015). When solving math problems, the overall load on an individual’s cognitive system, also referred to as cognitive load, can limit and interfere with performance (Sweller, 1988). This relates particularly to attention and working memory (WM). Working memory is the ability to control, regulate, and actively maintain relevant information in mind to accomplish complex cognitive tasks such as mathematical processing (Miyake et al., 2000). Many recent studies propose that individual differences in WM capacity in various domains (verbal, numerical, and visuospatial) are important predictors of math achievement (Bull & Lee, 2014; Dumontheil & Klingberg, 2012; Friso-van den Bos, van der Ven, Kroesbergen, & van Luit, 2013; Peng, Namkung, & Barnes, 2016; Raghubar, Barnes, & Hecht, 2010). WM can influence math achievement by helping to keep track of relevant information during problem solving but is also involved in selecting and switching to the most efficient arithmetic strategy (Barrouillet & Lépine, 2005; Cragg & Gilmore, 2014; Siegler & Lemaire, 1997; Wu et al., 2008). In online game-based learning environments, there is a great risk of overloading a player’s WM due to the rich number of multimedia elements and gamified features, which may limit the capacity for problem solving (Huang, 2011; Kiili, 2005; Moreno & Mayer, 2003). A cognitive overload on WM capacity may constrain both the acquisition of reasoning skills and the acquisition of knowledge (Baddeley, 1992; Eylon & Linn, 1988). The cognitive load experienced by individuals depends in part on their ability to selectively attend to relevant stimuli and, therefore, inhibit their attention to irrelevant stimuli such as distractors. Inhibitory control (IC) is the ability to prevent a response that is not relevant to the current task or situation (i.e., distracting stimuli or thoughts) and to control one’s attention, focusing on what one chooses and resisting interference (Diamond, 2013). IC skills have been found to predict mathematical performance in typically developing children, particularly in preschool and primary school children (Bull, Johnston, & Roy, 1999; Bull & Scerif, 2001; Espy et al., 2004; St Clair-Thompson & Gathercole, 2006). In online game-based learning environments, task-irrelevant distracting stimuli, such as gamified sounds, flashing objects, and alternative answer options, can trigger typically made errors. Similar to the Simon effect (Simon, 1969), where studies have found that irrelevant sensory stimuli in a task directly influence response selection and increase reaction time, the presence of irrelevant information in an online learning environment could interfere with performance in terms of accuracy and reaction time depending on one’s level of IC. Furthermore, Bull et al. (1999) and Rourke (1993) suggested that a lack of IC is also reflected in the type of errors children tend to make, for example, the inability to switch away from addition when multiplication is required (i.e., operant-related error). Interference and cognitive overload in a learning environment do not always stem from external stimuli but can also be internal in the form of worries about individual performance or about perceived time pressure (Ashcraft & Kirk, 2001; Mendl, 1999). These stressors can either drive people to use more efficient strategies (i.e., the best speed–accuracy trade-off within the constraints of the new situation) or compete with the attention that is normally allocated to the execution of the task (Caviola, Carey, Mammarella, & Szucs, 2017; Starcke & Brand, 2012). The latter is also known as the adverse effect of “choking under pressure”, where individuals perform worse than if there were no pressure (Baumeister, 1984; Beilock & DeCaro, 2007; Lewis & Linder, 1997). Critically, studies have found that people with high WM capacity are more affected by this dual-task environment and suffer more under pressure than those with low WM capacity (Beilock & Carr, 2005; Sattizahn, Moser, & Beilock, 2016; Wang & Shah, 2014). In addition, Sattizahn et al. (2016) found that individuals’ variability in attentional control processes influenced the effect of pressure. Those with poor attentional processes suffered decreased performance under pressure, reflecting that some individuals are able to prevent the interfering effect of pressure on their performance, whereas others with poorer attentional control are not. So, although increased WM and IC are generally associated with better math performance and efficient strategy use, many studies have found that this depends on the stressors in the environment. The purpose of the current study was to investigate the impact of stressors in the relatively new context of an online learning environment. One particular stressor, typical to a lot of online game-based learning environments, is time pressure, which is usually presented in the form of a gamified visual stimulus. For example, in Math Garden there is visual time pressure in the form of coins counting down every second, which is also incorporated into the game’s scoring rule for math performance (i.e., “high speed, high stakes” rule; see Maris & van der Maas, 2012). The advantage of using time pressure is that it provides the opportunity to relate speed of processing to the ability of the child, which is valuable with easy problems (Klinkenberg et al., 2011; van der Maas & Wagenmakers, 2005). In addition, in the case of games (similar to sports), the challenge of acting within a time limit can make the activity more enjoyable (Freedman & Edwards, 1988). Because time pressure itself is invaluable for most game-based learning environments, the current study addresses a different question: Should the visibility of the time pressure (in the form of a countdown) be adapted for individuals depending on whether it negatively affects math performance? Following the interference and overload theory, time pressure in the form of animated visual stimuli could be a distracting component that negatively interferes with solving math problems depending on the child’s level of IC and WM. However, the alternative situation, with no visible reminder of time passing by, requires attention to be allocated to time perception, which could result in suboptimal strategies in speed–accuracy trade-off in the main task (Brown & Perreault, 2017; Grondin, 2010; Matthews & Meck, 2016; Zakay, 1993). The purpose of this study was threefold. First, we investigated the association of individual differences in verbal WM and IC with performance of simple addition and multiplication problems in blocks of single or mixed operations in a game-based environment for primary school children. We expected that both verbal WM and IC would be positively associated with math performance and that higher IC would be associated with a reduced cost of switching between multiplication and addition. Second, we explored whether a particular feature of cognitive load, the visibility of time pressure, would affect arithmetic performance in general and whether this impact was different for various children depending on the level of WM and IC. We did not have a hypothesis regarding whether visibility or invisibility of time pressure would be associated with worse math performance because both features create a dual-task condition. Any effect on math performance was expected to interact with individual differences in WM and/or IC. Finally, whether learners are attending to or actively inhibiting their attention to irrelevant/distracting stimuli can be studied by looking at eye movements and fixations (i.e., moments when the eyes are relatively stationary and fixed on an object) using eye-tracking technology (Duprez et al., 2016; Wijnen & Ridderinkhof, 2007). In a learning environment, eye tracking can be used to investigate how learners interact with the stimuli and how the order and duration of their attending affect their problem solving. Eye-tracking data can also be used to improve the learning environment based on knowledge of how learners process the materials through their eye movements (Asteriadis, Tzouveli, Karpouzis, & Kollias, 2009; Garcia-Barrios et al., 2004). Using eye tracking, we explored differences in the locus of attention during the arithmetic task depending on whether time pressure was visible or not and on children’s levels of WM and IC. This study included data from a single time point and, therefore, will not inform our understanding of how individual differences and task features affect learning over time. However, a better understanding of how performance in online math tasks may be affected by these factors could allow a tailoring of the environment to individual learners, making sure that the task challenges, and therefore trains, their arithmetic skills rather than loading on other aspects of their cognitive capacity.","A total of 42 primary school children aged 8 to 11 years were recruited through a local voluntary participant database and word of mouth. Three children were excluded from all analyses because testing sessions were interrupted due to distress or tiredness. The final sample consisted of 39 children (19 male; Mage = 9.60 years, SD = 1.02, range = 8.00–11.50). For 3 children, insufficient eye gaze data were collected, leaving 36 children (18 male; Mage = 9.67 years, SD = 1.00, range = 8.00–11.50) for the eye-tracking analyses. The study was approved by the departmental ethics committee at the university. Informed consent was given by caregivers, and verbal assent was given by participants.","All stimuli were presented in MATLAB (The MathWorks, Natick, MA, USA) using the Psychophysics Toolbox (Brainard, 1997). During the first task, participants performed a math task on a computer (see Fig. 1A) similar in design to Math Garden (Straatemeier, 2014). The study took place in a lab setting, and all measures were completed in a single session lasting about 30 min in total. Before data collection started, condition assignment was randomized for a list of 40 participants using MATLAB. An additional 2 participants were tested to compensate for incomplete or withdrawn participants. There were three randomly ordered blocks comprising 20 multiplication problems, 20 addition problems, and 22 mixed multiplication and addition problems, respectively. All problems involved single-digit numbers between 1 and 9. For each arithmetic problem, participants were asked to choose one of six answer options, which consisted of the correct answer and the five most frequent errors made by children of similar age on that arithmetic problem, based on Math Garden data previously collected from a large Dutch sample (Fig. 1A). Participants had a maximum of 8 s to click on one of the answers, after which the correct answer was highlighted. In a between-participants manipulation, 19 children were randomly assigned to the visible time pressure condition, where the time limit of 8 s was visible in the form of coins counting down at the bottom right of the screen, similarly to Math Garden (Fig. 1A). The other 20 children needed to respond within the same 8 s, but there were no coins on the screen (no visible time pressure condition). After every trial, direct feedback on performance was given; the correct answer was circled in green, and the incorrect answer was circled in red in the case of an incorrect response. The measure of math performance was calculated with a scoring rule following the equation sij = (2xij − 1)(d − tij) (adapted from Maris & van der Maas, 2012). This rule imposes a speed–accuracy trade-off, where fast and correct responses result in a high score and incorrect responses result in a negative score. Player j responds xij on trial i (xij = 1 in the case of a correct answer and xij = 0 in the case of an incorrect answer) in time tij (in seconds; range = 0 to 8) before the time limit d (set to 8 s in this study) and obtains the score sij (range −8 to 8). Participants’ verbal WM was then assessed with a backward digit span task, where children were asked to repeat, backward, lists of single-digit numbers pronounced by the experimenter. After practice with two numbers, the first level consisted of four lists of three numbers; the child moved one level up (with an additional number) when at least three of the four lists were repeated back successfully. A WM score was computed as the total number of correct answers. IC was assessed with a computerized spatial incompatibility Simon task (adapted from Duprez et al., 2016; see Fig. 1B). Children were asked to move their mouse to either the top left or top right box depending on the color of the target square while ignoring its location. When the target was blue children needed to move their mouse toward and click into the left box, and when it was orange they needed to move their mouse toward and click into the right box. In half of the trials the location of the target was congruent with the correct response, and in the other half it was incongruent with the correct response (Fig. 1B). Participants completed 40 trials in a randomized order, which resulted in between 1 and 5 trials of the same type (congruent/incongruent) being repeated in a row. The measure of IC, referred to as the IC interference effect, was computed as the difference between incongruent and congruent trials mean reaction time (RT) divided by congruent trials mean RT, using correct trials only. A high score reflects a slower RT on incongruent trials (i.e., difficulty in inhibiting attention to irrelevant information) than on congruent trials (i.e., baseline processing speed). Eye tracking During the math and Simon tasks, children were seated at a distance of 60 cm in front of an eye tracker. Eye movements were recorded using a Tobii TX300 (Tobii, Danderyd, Sweden) at a sampling rate of 120 Hz. The raw data were classified into fixations and saccades using the gazepath package in R (R Core Team, 2013; van Renswoude et al., 2018). Gazepath uses an algorithm to categorize the data into fixations and saccades while accounting for individual differences and data quality. Fixations in the math task were labeled as one of the following three areas of interest (AOIs): (a) the question box, (b) one of the six answer options, or (c) the coins (i.e., the visible countdown of time) (Fig. 1A). Statistical analyses Data management and statistical analysis were performed using R software (R Core Team, 2013). For all independent variables, z scores were generated to standardize the scores for further analyses. In a first set of analyses, math performance was averaged over the three blocks (addition, multiplication, and mixed addition and multiplication) and compared between the visible time pressure condition and the no visible time pressure condition, covarying for age and WM score or IC interference effect, using between-participants three-way analyses of covariance (ANCOVAs). With a sample size of 39, the study had 80%, 90%, and 95% power to detect large eta-square (η2) effect sizes of .18, .22, and .26, respectively, when comparing two groups. Eta-square effect sizes have been classified as follows: small η2 = .02, medium η2 = .13, and large η2 = .26 (Cohen, 1988). An additional analysis investigated associations between IC and the cost of needing to switch between operations. We subtracted the average performance of the mixed block trials from the average performance on the trials in the single operation blocks for multiplication and addition problems separately. These cost measures were entered into ANCOVAs including the IC interference effect, visibility of time pressure, and age for multiplication and addition separately. Assumptions of the ANCOVAs were met, with analyses showing homoscedasticity and normality of the residuals. Eye-tracking analyses (n = 17 in the visible time pressure condition and n = 19 in the no visible time pressure condition) focused on correct trials (excluding 12.7% of trials) and trials where there was at least more than one fixation to ensure high eye-tracking data quality (excluding a further 1.2% of trials). The average number of fixations and the proportional duration of fixation on each AOI were calculated for each participant. An additional metric was the average number of answer option AOIs the participant attended to on a trial. We explored, in three-way ANCOVAs, whether these eye-tracking metrics differed according to the visibility of time pressure and whether this interacted with WM score, IC interference effect, or math performance. The data were checked for outliers using a criterion of |z score| > 3 for both the dependent and independent variables. No outliers were identified. In the regression analyses, Cook’s distance suggested one to three influential points for some behavioral and eye- tracking results. Analyses were repeated excluding these data points, and the results were strengthened except in one case, which is discussed further below. In addition, Bayesian ANCOVAs were performed post hoc for the results with null effects or p values just under the threshold (p < .05) using JASP (JASP Team, 2019). To quantify uncertainty about effect size and to obtain evidence in favor of a null hypothesis (Wagenmakers et al., 2018), we distinguished between experimental insensitivity (Bayes factor [BF]10 and BF01 < 3) and robust support for the alternative hypothesis (BF10 > 3) or null hypothesis (BF01 > 3) (Dienes, 2014) RT and accuracy ~~~~~~~~~~~~~~~ We performed two one-sided equivalence tests with alpha = .05 and no assumption of equal variance and found statistical equivalence between the visible and no visible time pressure groups for age, percentage female, verbal WM, and IC (Table 1). t Tests were run to test whether visibility of time pressure was associated with math performance. We did not find any difference between the groups in mean RT, t(38) = 0.82, p = .42, proportion of correct responses, t(38) = 0.45, p = .66, proportion of no response within the time limit, t(38) = 0.15, p = .88, or mean math score, t(38) = 0.77, p = .45. These comparisons indicate that the visibility of time pressure did not have an effect on math performance. The average overall math performance was 3.34 (SD = 1.34), meaning that the average score was correct and answered within roughly half the time limit (see Method for scoring rule). This measure was used for further analyses. Mean accuracy in the Simon task was high (M = .99, SD = .05). As expected, RTs differed between congruent and incongruent trials, t(38) = 6.17, p < .001. Participants were on average 150 ms slower in incongruent trials (Fig. 2A). The individual average IC interference effect was used as a measure of IC for further analyses (Fig. 2B). Impact of time pressure on math performance depending on level of IC and verbal WM ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The first analysis included only age and visibility of time pressure as predictors of math performance. This showed a positive association between age and math performance, F(1, 35) = 32.05, p < .001, ηp2 = .53 but not time pressure (p = .892, ηp2 = .00), nor was there an interaction between age and time pressure (p = .679, ηp2 = .01). The second analysis included WM score as a covariate (Table 2). WM score was positively associated with math performance, F(1, 31) = 13.15, p = .001, ηp2 = .24 (Fig. 3A), but there was no interaction with age or time pressure (all ps > .50, ηp2 < .01). The Bayesian ANCOVA showed that a null model with merely main effects for WM score and age was 11.4 times more likely than including any of the above-mentioned interactions or the main effect of time pressure. The third analysis (Table 2) included the IC interference effect as covariate. There was no main effect of IC interference effect (p = .62, ηp2 = .01), but there was a significant two-way interaction between time pressure and IC interference effect, F(1, 31) = 6.59, p = .015, ηp2 = .18, and a three-way interaction between time pressure, age, and IC interference effect on math performance, F(1, 31) = 4.55, p = .041, ηp2 = .13. Significant evidence for both interaction effects was demonstrated through Bayesian analyses (Table 2). To examine the two- and three-way interactions, separate multiple regressions were performed in the visible and no visible time pressure groups. In the group with visible time pressure, the IC interference effect and Age × IC interference effect interaction terms did not significantly predict variance in math performance (BF01 = 0.57, i.e., no evidence for either hypothesis) (Fig. 3B). In contrast, the group with no visible time pressure showed a negative association between math performance and IC interference effect, β = −.42, t(16) = 2.77, p = .014 (BF10 = 6.47, i.e., substantial evidence for including this effect) (Fig. 3B), and an interaction between age and IC interference effect, β = .43, t(16) = 2.629, p = .018 (BF10 = 8.84). The interaction effect showed that the association between math performance and IC interference effect was mostly driven by younger children (Fig. 3C). Operation switch cost ~~~~~~~~~~~~~~~~~~~~~ To investigate whether switching between operations led to a cost in performance, we compared the mean math scores of single operation versus mixed operation blocks for multiplication and addition problems separately. Paired t tests showed that children’s performance on multiplication problems did not differ between the mixed (M = 2.83) and single (M = 2.76) operation multiplication blocks, t(38) = 0.41, p = .341. For addition, children performed less well on the trials in the mixed block (M = 3.67) than on those in the single operation block (M = 4.00), t(38) = 2.51, p = .008. Therefore, children showed a cost of needing to switch between multiplication and addition on addition problems only. Because the ability to switch between arithmetic operations has been associated with IC in previous studies (Bull et al., 1999; Rourke, 1993), additional analyses explored whether IC predicted the ability to switch between addition and multiplication operations in the mixed operation blocks compared with the single operation blocks (Table 2). For addition problems, the IC interference effect predicted the performance difference between the mixed and single operation blocks, F(1, 31) = 5.06, p = .031, ηp2 = .13. Bayesian ANCOVA showed that a model including IC was 2.68 times more likely than the null model; no interaction with age (p = .302, ηp2 = .03) or time pressure (p = .153, ηp2 = .06) was found. Eye fixations and patterns ~~~~~~~~~~~~~~~~~~~~~~~~~~ Exploratory analyses investigated whether eye movements during the math task could give some insight into the behavioral findings. Analyses were performed on the mean number of fixations and proportion of total fixation duration on specific AOIs. The latter did not show any significant effect. The first analyses looked at the fixations on the question box AOI (e.g., 6 × 8 in Fig. 1A) because other studies have found that looking back and forth at the question is positively associated with attentional and WM load (Droll & Hayhoe, 2007; Orquin & Mueller Loose, 2013). ANCOVAs were run to test for associations with the visibility of time pressure in interaction with individual differences in IC and WM separately while covarying for age and math performance (Table 2). A significant main effect for time pressure, F(1, 28) = 12.02, p = .003, ηp2 = .43 (BF10 = 30.88, i.e., very strong evidence), showed that there were more fixations on the question box when time pressure was visible (M = 2.69) than when there was no visible time pressure (M = 2.02) (Fig. 4A). Second, because operation-related errors have been found to be associated with the level of IC (Bull et al., 1999; Rourke, 1993), fixations on the operation-related error answer options were investigated separately for addition and multiplication. ANCOVAs were performed to test for associations with the visibility of time pressure and the IC interference effect while covarying for age and math performance. For the addition problems with multiplication-related errors as answer options, we found no significant predictors (ps > .20, ηp2 < .05) (Table 2). For multiplication problems with addition- related errors, a significant interaction between time pressure and IC interference effect, F(1, 28) = 5.34, p = .018, ηp2 = .44 (BF10 = 5.21, i.e., substantial evidence), showed that the mean number of fixations on the addition-related error increased with increasing IC interference effect (β = .56) only when time pressure was visible. Finally, analyses were performed to investigate the mean number of answer options participants looked at before giving their answer and whether this was related to WM, IC, and the visibility of time pressure. An ANCOVA was performed with the average number of answer options attended to as the dependent variable, visibility of time pressure and IC interference effect or WM as independent variables, and age and math performance as covariates. The analysis with WM as a predictor showed no main or interaction effects, only evidence that a null model was 11 times more likely than including any of the predictors. The analysis with IC as a predictor showed a significant interaction between visibility of time pressure and the IC interference effect, F(1, 31) = 4.60, p = .039, ηp2 = .13. However, Cook’s distance highlighted that there was one influential point that drove this interaction. Consistent with this, only anecdotal evidence (BF10 = 2.90) for including this interaction to the null model was found in the Bayesian regression (Fig. 4C and Table 2). Follow-up regression analysis showed a trend for a positive association for the IC interference effect when time pressure was visible, β = .54, t(13) = 2.03, p = .063, but little evidence (BF10 = 1.23) in the Bayesian regression. No association between the IC interference effect and number of answer options visited was found when time pressure was invisible, β = − .21, t(16) = 0.82, p = .423; Bayesian regression showed anecdotal evidence for the null hypothesis (BF01 = 2.00).","This study combined behavioral and eye-tracking measures to test whether individual differences in verbal WM and IC in primary school children could predict their ability to solve arithmetic problems in different online learning environments, where visibility of time pressure was varied. The behavioral results showed that verbal WM was a positive predictor of arithmetic performance in general, in line with previous studies (see Raghubar et al., 2010, for a review), and that this association was independent of the visibility of time pressure. In contrast, individual differences in IC predicted arithmetic performance only when the same time pressure was not visibly illustrated by an animation. In addition, we found that this association with IC was mostly driven by younger children, similar to previous studies (Bull & Scerif, 2001). Eye-tracking results also showed that children fixated on different parts of the stimuli during the math task depending on the visibility of time pressure, their IC level, and age. Overall, these findings point out that the visibility of time pressure may affect performance of certain individuals in online learning environments and that possible constraints of attentional control (i.e., the amount of interfering information compromising cognitive resources) should be considered. Learning environments with both visible and invisible time pressure can create dual-task environments, leading to less attention to the main task of solving math problems. When time pressure is visible, the user has a constant physical reminder of timing (in this study in the form of an animated visual stimulus). Adding more visual stimuli and time pressure was suggested by previous studies to contribute to loading WM capacity, leading to suboptimal strategies and attention (Barrouillet, Bernardin, Portrat, Vergauwe, & Camos, 2007; Caviola et al., 2017; Terras & Ramsay, 2012). This impact can also be influenced by other individual differences such as math anxiety (Ashcraft & Krause, 2007; Caviola et al., 2017; Kellogg, Hopko, & Ashcraft, 1999), engagement, and attitude to learning (Barkatsas, Kasimatis, & Gialamas, 2009; Kebritchi, Hirumi, & Bai, 2010). Although the visibility of time pressure did not interact with individual differences in verbal WM in terms of math performance, the notion of visible time pressure as an increasing demand on WM resources is reflected in our eye-tracking results. Children made more fixations on the question in the visible time pressure condition than in the invisible time pressure condition, suggesting that they may have found it more difficult to keep the question in mind (Orquin & Mueller Loose, 2013). Although previous studies suggested that the impact of extra stressors on math performance depends on the ability to resist distractions (i.e., IC; Sattizahn et al., 2016), we showed that the performance of children was not affected by their level of inhibition when time pressure was visible. The higher number of fixations on answer options and on operation-related errors did suggest that for children with lower IC the task was more demanding in terms of decision difficulty and/or attentional resources (Orquin & Mueller Loose, 2013), but this did not result in lower performance. Time perception is intensively studied (for an overview of recent reviews, see Block, Grondin, & Gibbon, 2014) and involves diverse perceptual, motor, cognitive, and brain processes (Block & Gruber, 2014). One line of investigation in time perception concerns its bidirectional interference with higher-level executive cognitive processes such as mental arithmetic but also with executive functions (Block, Hancock, & Zakay, 2010; Brown, Collier, & Night, 2013). This interference occurs in a dual-task condition where time perception competes for the same attentional resources as the other task, leading to cognitive load. Because the interference is bidirectional, studies have also shown that lower IC is associated with less accurate time perception (Brown & Perreault, 2017; Meaux & Chelonis, 2005). This closely aligns with our finding that low levels of IC were associated with low arithmetic performance when children also needed to estimate time without a reminder. This could be due to an impairment of time perception, such that these children have trouble in deciding on an optimal speed–accuracy trade-off strategy. Therefore, for children with low IC, visualizing time pressure could reduce cognitive load, whereas children with high IC seem to be able to estimate time in parallel with solving arithmetic problems. One of the limitations of this study was the small sample size. This was due to the use of an eye tracker, which necessitated a lab setting. The use of a participant volunteer database and testing in a lab setting also likely biased our recruitment toward children from higher socioeconomic backgrounds and with higher cognitive abilities. The next step would be to replicate our findings with a larger heterogeneous sample from online learning environments such as Math Garden to ensure that the behavioral findings are reliable. In addition, we chose a between- participants design that minimizes the effect of learning and testing time, but a within- participants design would have had more power to detect interactions between the time pressure manipulation and individual differences in WM and IC. Future work will investigate whether learning, rather than performance at a single time point, can be improved based on an adapted environment, informed by the results of the current study. Although the purpose of this study was to implement these findings in an online adaptive environment, the arithmetic problems used were standardized to ensure that we could compare arithmetic performance within this sample size. Due to our wide age range (8–11 years), certain arithmetic problems were inevitably less challenging for some children; therefore, all analyses were covaried for age. Note, however, that there are large individual differences within year groups on arithmetic tasks (Straatemeier, 2014), so a more homogeneous sample in terms of age may have still shown considerable variability in arithmetic performance. To further investigate whether the associations among IC, WM, time pressure, and arithmetical outcome change with age, the difficulty level of the arithmetic problems should be adapted to the ability of the child. Finally, whereas we considered the coin countdown to reflect time pressure, it also indicated the potential reward to be gained when correctly solving the problem. Although the reward obtained was shown to both groups of participants when a trial was completed, the group with no visible coin countdown did not have a constant reminder of the potential reward. This reward cue difference between the groups may have led to some of the differences observed between the conditions. In conclusion, we found that the (in)visibility of time pressure, a key feature that is adaptable in a lot of online game-based learning environments and psychological tasks in general, may create cognitive overload and affect the application of knowledge and skills. Specifically, we showed that this aspect of online game-based learning environments may differentially affect children’s arithmetic performance as a function of their cognitive abilities. Measuring the individual levels of cognitive functioning, in particular WM and IC, is essential to allow children to perform and practice tasks at their highest level. In addition, the use of an eye tracker in this context allowed an in-depth exploration of how learners interacted with the different elements in the environment above and beyond accuracy and RT. Future work should focus on developing a broader online adaptive framework for learning mathematical skills and knowledge that adapts not only to children’s mathematical skills but also to their more general cognitive strengths and weaknesses."],["The aim of this study was to investigate how irrelevant speech, temperature and ventilation rate together affect cognitive performance and environmental satisfaction in open-plan offices. In Condition A, neutral temperature (23.5 °C), low intelligibility of speech (high absorption and low masking sound level) and high fresh air supply rate (30 l/s per person) were applied. This was contrasted to Condition B with high room temperature (29.5 °C), highly intelligible speech (low absorption and high masking sound level) and a negligible fresh air supply rate (2 l/s per person). Sixty-five participants were tested. In Condition B, performance decrement was observed especially in working memory tasks. Based on subjective assessments, mental workload, cognitive fatigue and symptoms were higher and environmental satisfaction was lower in Condition B. It was concluded that special attention should be paid to the design of whole indoor environment in open-plan offices to increase subjective comfort and improve performance. --------------------------------------------------------------------------------","Scientific interest towards subjective satisfaction in open-plan offices has increased because open-plan office has become the most usual office solution, mostly because of its high space efficiency (De Croon, Sluiter, Kuijer, & Frings-Dresen, 2005). Moreover, open- plan offices are also assumed to improve organizational productivity due to the enhanced exchange of information and communication and increased teamwork (Allen & Gerstberger, 1973; Hundert & Greenfield, 1969). However, many studies have shown that there are disadvantages in open-plan offices if the design of the indoor environment (IE) is inadequate. Increased cognitive workload (De Croon et al., 2005), concentration problems and fatigue (Haapakangas, Helenius, Keskinen, & Hongisto, 2008; Pejtersen, Allermann, Kristensen, & Poulsen, 2006) and the lack of speech privacy (De Croon et al., 2005) have been reported. Open-plan offices have also been associated with decreased job satisfaction (De Croon et al., 2005). Decreased satisfaction with IE has been indicated to have a connection with decreased job satisfaction (Veitch, Charles, Farley, & Newsham, 2007). The amount of annual sick leave has also been shown to be greater in open-plan offices, as assessed by employees’ self-ratings (Bodin Danielsson, Chungkham, Wulff, & Westerlund, 2014; Pejtersen, Feveile, Christensen, & Burr, 2011). One of the most commonly mentioned causes for these problems is poor acoustic conditions, i.e., disturbance caused by colleagues’ speech and poor speech privacy (Danielsson, 2005; Frontczak et al., 2012; Haapakangas et al., 2008; Pejtersen et al., 2006). Improper thermal conditions and poor air quality have also been reported as producing discomfort in open-plan offices (Haapakangas et al., 2008; Pejtersen et al., 2006). On the other hand, overall improvement of the IE can significantly increase environmental satisfaction in open-plan offices (Hongisto, Haapakangas, Helenius, Keränen, & Oliva, 2012). That is, differences between open-plan offices can be significant regarding on the quality of IE. The effects of IE on work performance and various components of environmental satisfaction have been studied in several laboratory experiments. However, most of the previous laboratory studies have focused on the effects of a single factor of IE. In the present study, we simultaneously manipulated three IE factors in order to examine their joint effects on task performance and environmental satisfaction. Fig. 1 depicts how our study was designed as a follow-up study to three previous studies each separately examining the effects of a single factor of IE. We next summarize the evidence for the effects of each factor examined separately. Effects of office noise ~~~~~~~~~~~~~~~~~~~~~~~ Office noise, especially irrelevant speech having sufficiently high intelligibility, has been shown to decrease performance in serial recall (e.g., Haapakangas et al., 2011; Haka et al., 2009), information search (Jahncke, Hongisto, & Virjonen, 2013), proofreading (e.g., Venetjoki, Kaarlela-Tuomaala, Keskinen, & Hongisto, 2006) and counting tasks (e.g., Buchner, Steffens, Irmen, & Wender, 1998). Our experiment was preceded by an experiment conducted in the same laboratory, which showed that the room acoustic design, where the intelligibility of irrelevant speech could be reduced, improved work performance (Haapakangas, Hongisto, Hyönä, Kokko, & Keränen, 2014). Moreover, subjective assessments confirm the negative impact of highly intelligible irrelevant speech; speech and other office activity sounds negatively affect subjective well-being, acoustic satisfaction and self-estimated performance (Evans & Johnson, 2000; Haapakangas et al., 2014; Haapakangas et al., 2011; Haka et al., 2009). It is important to study how different room acoustic solutions usually applied in open-plan offices can be used to reduce the negative effects of irrelevant speech. The effects seem to mainly depend on speech intelligibility (Ellermeier & Hellbrück, 1998; Hongisto, 2005; Jahncke et al., 2013) and not on the loudness of speech (Colle, 1980). Performance is expected to decrease with increasing Speech Transmission Index, STI (Hongisto, 2005). Subjective speech intelligibility can be objectively evaluated by measuring STI which ranges from 0.00 to 1.00, with large values representing highly intelligible speech (ISO 3382-3). STI can be reduced in open-plan offices by simultaneous application of high room absorption, high screens between desks and the use of masking sound (Bradley, 2003; Keränen & Hongisto, 2013; Virjonen, Keränen, & Hongisto, 2009). By reducing the STI values below 0.30, it can be expected that the negative effects on performance can be significantly reduced (Haka et al., 2009; Jahncke et al., 2013; Keus van de Poll, Ljung, Odelius, & Sörqvist, 2014) compared to a situation where the STI is above 0.50, which is, unfortunately, very typical in open-plan offices (Keränen & Hongisto, 2013; Virjonen et al., 2009). Our study involved an acoustic manipulation where the two most important factors of acoustic design were considered simultaneously: sound masking and room absorption. Absorption was used to reduce the reflection of sound from room surfaces and to reduce the overall speech level. Sound masking was used to reduce the signal-to-noise ratio of speech. Successful application of masking sounds in the office has been reported by Hongisto (2008) and Hongisto et al. (2012). Effects of high room temperature ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Room temperature can affect cognitive performance (see e.g., reviews of Hancock, Ross, & Szalma, 2007; Pilcher, Nadler, & Busch, 2002). However, the results of these reviews cannot be directly applied to office environments because the examined thermal conditions differed from usual thermal conditions in offices. The desirable room temperature in offices is between 21 °C and 25 °C depending on outside temperature, clothing, activity level and cultural differences. However, much higher temperatures, up to 35 °C, can be found in offices having insufficient cooling capacity or no cooling at all. When neutral temperatures (21–25 °C) have been compared to higher ones (above 26 °C), cognitive performance has been observed to decline at higher temperatures in short-term free recall tasks (Hygge & Knez, 2001), addition and visual tasks (Lan, Wargocki, Wyon, & Lian, 2011) and working memory tasks (Häggblom, Hongisto, Haapakangas, & Koskela, 2011). Maula et al. (2015; Fig. 1) performed an experiment before our study in the same laboratory environment. They found that high temperature (29 °C) affected the performance in working memory demanded N-Back task. However, temperature did not affect psychomotor, attention or long-term memory tasks. These results are consistent with the suggestion of Hancock et al. (2007) that the performance effects of room temperature are task-sensitive. Subjective assessments yield a more uniform picture of the effects of room temperature. High temperature has been reported to negatively affect mood, energy, motivation, concentration and the assessment of air quality (Lan et al., 2011; Maula et al., 2015). High temperature has also been found to increase self-rated intensity of somatic symptoms compared with neutral temperature (Lan et al., 2011). Effects of air quality ~~~~~~~~~~~~~~~~~~~~~~ Air quality is affected by the ventilation rate and emissions from the building, furniture and occupants (Wargocki, Bakó-Biró, Clausen, & Fanger, 2002). In the majority of laboratory experiments investigating the effects of air quality on performance, the air quality has been reduced by artificial material emissions, such as by installing old and polluting carpets in the room. The combination of artificial material emission and small ventilation rate has marginally decreased performance in typing and negatively affected subjective assessments of air quality and well-being (Wargocki, Wyon, Sundell, Clausen, & Fanger, 2000). Similar results have been found for high material emissions with a constant ventilation rate (Wargocki, Wyon, Baik, Clausen, & Fanger, 1999). When office buildings are renovated, furniture and decoration are likely to be changed and old material emission sources are usually removed. The emissions from new furniture and surface materials can cause relatively high concentrations of volatile organic compounds (VOCs) for a couple of months. The combination of high material emissions from new materials and small ventilation rate has been found to decrease objectively measured performance in typing, addition and memorization tasks and to reduce the acceptability of perceived air quality (Park & Yoon, 2011). In comparison, a previous study (Koskela, Maula, Haapakangas, Moberg, & Hongisto, 2014, Fig. 1) carried out in the same laboratory as our study investigated the situation where the occupants themselves were the strongest pollution sources of the room. Comparison between high (28 l/s per person, 600 ppm) and low (2 l/s per person, 2200 ppm) ventilation rates with negligible emission from furniture and building did not reveal any remarkable differences in subjective environmental assessment or objective performance despite the fact that the tasks required cognitive processes relevant to many types of office work and the exposure time was reasonably long (3.5 h). In contrast, decision- making performance has been shown to decline by high carbon dioxide (CO2) concentrations (Satish et al., 2012). Due to methodological differences, more research is needed to confirm the effects of air quality on task performance and subjective assessment of the work environment. It is also worth noting that there is no fundamental theory of what mechanisms, such as fatigue or working memory capacity, would explain the effects of poor air quality on cognitive performance. Combined effects of IE ~~~~~~~~~~~~~~~~~~~~~~ Only a few studies have focused on the effects of several simultaneously modified IE factors on objective performance, despite the fact that occupants are exposed to several factors of IE in open-plan offices. In addition, the results are inconsistent. Some studies have reported expected performance effects (Balazova, Clausen, & Wyon, 2007; Hygge & Knez, 2001). Also interaction effects between IE factors have been reported (Witterseh, Wyon, & Clausen, 2004). However, even major changes in IE conditions have not always affected performance (Balazova, Clausen, Rindel, Poulsen, & Wyon, 2008; Clausen & Wyon, 2008). On the other hand, subjective ratings have indicated more uniformly how several IE factors together affect human comfort. Simultaneous negative changes in IE factors, such as temperature, irrelevant speech (office noise), traffic sounds and air quality, have increased dissatisfaction (Balazova et al., 2007, 2008) and reduced the perceived ability to concentrate on and perform job-relevant tasks (Balazova et al., 2007; Clausen & Wyon, 2008). A wide variety of environmental conditions, tasks and procedures have been used in the aforementioned studies. First, performance is affected differently by different noise types (e.g., Banbury & Berry, 1998; Szalma & Hancock, 2011). Studied noises have been originated from office noise (Balazova et al., 2008; Clausen & Wyon, 2008; Witterseh et al., 2004), road traffic (Balazova et al., 2007; Clausen & Wyon, 2008), or ventilation (Hygge & Knez, 2001). It has been shown that intelligible speech interferes with working memory and performance because it is unpredictable and information-rich, while constant noise does not cause interference (Jones, Madden, & Miles, 1992). Second, the exposure times have varied from 20 min (Balazova et al., 2007) to 6 h (Balazova et al., 2008). Third, successful task performance requires several cognitive processes (Sörqvist, 2014); very different tasks relying on varied cognitive processes have been used to measure task performance. Fourth, in some studies participants have been told about the IE conditions before the experiment (e.g., Balazova et al., 2007) or they have been able to personally select the IE conditions (Clausen & Wyon, 2008). Finally, both between-subjects (e.g., Hygge & Knez, 2001) and within-subjects (Balazova et al., 2007, 2008) designs have been employed. All these variations can be assumed to affect the results on performance and subjective assessment of the IE via different paths. In our study, the environmental conditions and experimental procedures were selected so that both practical questions related to the work environments and the highest possible scientific quality could be met. Speech was used as the noise source since it is the most often complained noise source (e.g., Haapakangas et al., 2008). Several cognitive tasks were applied to cover a variety of different kinds of office work. A moderately long exposure time was applied to mimic a typical working period at the office desk (2 h). The participants were blind with respect to the experimental manipulations. Finally, a within-participants design was used to reduce inter-individual variability in cognitive performance and environmental effects. The aim of the study ~~~~~~~~~~~~~~~~~~~~ Our study represents the final experiment of a larger research programme (Fig. 1). The experiment, which combines three factors of IE (intelligibility of irrelevant speech, temperature and air quality), was preceded by three laboratory experiments that focused on each single factor. Out of these three experiments, performance effects were found in the experiment concerning the intelligibility of irrelevant speech (Haapakangas et al., 2014) and room temperature (Maula et al., 2015). The effects of these factors were also evident on subjective responses. However, the impact of low ventilation rate (2200 ppm CO2 concentration; Koskela et al., 2014) on subjective ratings was small and no effect on task performance was observed. Theoretically, it is interesting to study how the effects of these three factors add up when the conditions are presented jointly. For example, it is possible that the detrimental effects of the individual conditions are exacerbated when combined. Because IE problems are often multifaceted in workplaces, the simultaneous evaluation of these factors is also important for its practical significance. As far as we know, no experimental studies exist that have investigated the simultaneous exposure to the intelligibility of irrelevant speech, room temperature and air quality (ventilation rate) and the related effects on cognitive performance and subjective experience in an open-plan office. A variety of tasks was used to tap into the job-related cognitive processes required in typical office work (non-communicative tasks). Because subjective assessment has proved to be sensitive to capture the effects of IE, an extensive battery of questionnaires was also employed. The aim of this experiment was to investigate the simultaneous exposure to highly intelligible irrelevant speech, high room temperature and low ventilation rate and the effects of exposure on cognitive performance and environmental satisfaction. The study was carried out in open-plan office which was built in a laboratory environment. Two experimental conditions were investigated. Condition A was a combination of neutral temperature (23.5 °C), high outdoor air supply rate (30 l/s per person, 580 ppm CO2) and low intelligibility of surrounding irrelevant speech (high absorption and adequate masking in the room). Condition B was a combination of high room temperature (29.5 °C), highly intelligible irrelevant speech (no absorption or masking in the room) and a very low outdoor air supply rate (2 l/s per person, 1470 ppm CO2). The hypothesis was that Condition B would be inferior to Condition A with respect to the various subjective and objective measures used.","Sixty-five students (49 women and 16 men) from six faculties of the University of Turku took part in the laboratory experiment. The participants were 19–29 years old (M = 23, SD = 2) and were all native Finnish speakers. Each participant took part in the two conditions on two successive weeks. None of the participants reported any hearing difficulties, dyslexia or an attention deficit disorder and all had normal or corrected- to-normal vision. Participants were recruited via university email lists and were paid 50 euros (minus taxes) for their participation. Prior to the experiment, the participants were told that the aim of the experiment was to investigate work performance in an open- plan office environment. The participants were not informed beforehand about the experimental conditions. Design ~~~~~~ The experiment was carried out using a within-participants design, i.e., all participants were tested in both experimental conditions (two sessions), thus acting as their own controls. The within-participants variable was an indoor environment (IE) which had two conditions. The order of exposure to these experimental conditions was counterbalanced across participants in altogether thirteen groups. Half of the participants were first exposed to Condition A (32 participants) and the other half to Condition B (33 participants). Test environment ~~~~~~~~~~~~~~~~ The experiment was carried out in an open-plan office specially built for the purpose (Fig. 2). The room was carefully furnished and finished to resemble a normal work environment. The room was free from measurement apparatus usually found in laboratories. There were no measurement devices or other artificial artefacts visible in the room. The furniture and materials visible in the room were modern and commercially available. The furniture consisted of 12 identical desks, 1.3 m high screens between the desks, chairs, computers and storage units in the middle of the room. Six desks were reserved for the participants. The corner desks were reserved for the loudspeakers from which the background speech was played. Two desks were empty during the experiment. The height of the suspended ceiling was 2.55 m and suspension depth was 0.30 m. The suspended ceiling was made of 600 × 600 mm metal grid where 210 ceiling boards, six ventilation inlets, a ventilation outlet and 16 lighting units were installed. Approximately 88% of the ceiling (75 m2) was reserved for ceiling boards where either sound-absorbing (Condition A) or the non-sound-absorbing (Condition B) boards could be installed. Both boards had the same visual appearance so that the participants could not detect any differences between them. The walls were made of double drywall to provide good sound insulation to the surrounding office premises. Approximately 20% of the total wall area (18 m2) was reserved for porous linen pictures behind which sound-absorbing boards could be installed in Condition A. Artificial masking sound was produced from 14 loudspeakers placed evenly above the suspended ceiling so that the participants could not see them and the masking sound was most probably experienced as ventilation noise. The temperature and ventilation of the room could be controlled by an independent air-conditioning system. The air leakages in the room and ventilation ducts were minimized. Natural daylight had no access to the room. Two artificial windows were installed on one wall behind which lighting units were installed to resemble daylight. The illumination level of this artificial daylight was negligible in the desk area. The experimental conditions ~~~~~~~~~~~~~~~~~~~~~~~~~~~ The two experimental conditions were designed on the basis of our previous laboratory experiments focussing on single IE factors (Haapakangas et al., 2014; Koskela et al., 2014; Maula et al., 2015). The experimental conditions are described in Table 1. In Condition A, the IE was designed according to the most stringent target values used in open-plan offices in Finland (Class S2 of LVI 05–10440 en, 2008). In Condition B, the IE was designed to violate even the least demanding target values (Class S3 of LVI 05–10440 en, 2008). A neutral temperature has been found to be 23.5 °C in a study (Maula et al., 2015) similar to our study. This is very typical in office workplaces throughout the year in buildings equipped with modern air conditioning systems. Elevated temperatures up to 30 °C are, however, found during the summer season in many workplaces situated in buildings where cooling has not been installed. The temperatures in Conditions A and B were 23.5 and 29.5 °C, respectively, as used by Maula et al. (2015). Air quality in office-type buildings, where the occupants are the main pollutant sources, is usually determined by measuring the CO2 concentration. Concentrations up to 2500 ppm have been reported in office buildings (Seppänen, Fisk, & Mendell, 1999). Concentrations exceeding 3000 ppm and even up to 5000 ppm have been found, for example, in schools and meeting rooms with an insufficient ventilation rate and high crowdedness (Bakó-Biró, Clements-Croome, Kochhar, Awbi, & Williams, 2012; Seppänen et al., 1999). In large modern office buildings, the concentrations are normally below 1000 ppm (Apte, Fisk, & Daisey, 2000). The recommended upper limit for the CO2 concentration in offices is 1200 ppm in the lowest quality class S3 (LVI 05–10440 en, 2008). In quality classes S1 and S2, the recommendations are 750 ppm and 900 ppm, respectively. Values below 650 ppm are usual in modern Finnish offices. Condition A (580 ppm) conformed to this recommendation. Ventilation rate (30 l/s per person) was selected to certainly meet recommendations but not cause noise or the risk of draught. For Condition B, a CO2 concentration of 1470 ppm was applied. It should be noted that in the desk areas of open-plan offices, values exceeding 1500 ppm are rare. Ventilation rate for Condition B (2 l/s per person) was determined by ASHRAE standard 62.1 where the minimum ventilation rate is 2.5 l/s per person based on adapted persons. Room acoustics in Condition A corresponded to the most stringent requirements (Bradley, 2003; Hardy, 1957; LVI 05–10440 en, 2008; Virjonen et al., 2009). Speech intelligibility is reasonably high close to a speaker to enable normal face-to-face conversation but it declines fast with increasing distance resulting in low intelligibility of distant speech. Condition B corresponded to a situation where the room acoustic design does not conform even to the least demanding target values. This is nevertheless very typical in workplaces even though the room acoustic design principles have been published 57 years ago (Hardy, 1957). Speech intelligibility is very high both at short and large distances from the speaker so speech is expected to be disturbing independent of the distance of the speaker. The sound levels of background speech (the supposedly disturbing sound) and masking sound (the supposedly non-disturbing sound) in the desks are depicted more closely in Fig. 3. The lighting level at the table was constant, approximately 500 lux, in both conditions. The air velocity was less than 0.1 m/s in all desks and in both conditions. Implementation of thermal control and ventilation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The room temperature was controlled by six commercially available chilled beams in the ceiling and six electric radiators (total thermal power 5.5 kW) hidden below the tables of the unoccupied desks. The intake air was filtered with level F7 filters. Radiators were on during Condition B and chilled beams were used in Condition A. The room temperature was measured continuously in the corner desks (not visible to the participants) at a height of 1.1 m. Prior to the experiment, the implementation of thermal conditions in the desks (temperatures and draught) was also tested by installing six dummy bodies (60 W thermal load) at the desks. Based on this, it was deemed sufficient to monitor the temperatures during the experiment only at the corner desks. During Condition A, the CO2 concentration was planned to be low by means of high outdoor air supply. During Condition B, the desired CO2 concentration was planned to be high. Therefore, the rate of outdoor air supply was minimized by using a circulation duct in the technical room. To implement the poor air quality, eight employees worked in the office with this reduced ventilation rate for two hours to increase the CO2 concentration up to 1470 ppm before the participants arrived. By doing so, the human-based pollution rate (both CO2 and other human-based pollutants) was at the intended level right at the beginning of the experiment and the participants maintained the same pollution rate thereafter. There were typically six participants and the experimenter in the room at a time but the number of persons varied from four to seven because of occasional cancellations. The thermal load was compensated accordingly and temperature variation between sessions and within a session was negligible. However, the ventilation rate was not changed according to the number of persons present in the room resulting in some variation of CO2 concentration between sessions (see Table 1). The participants were instructed to wear trousers, long-sleeve shirt, t-shirt, socks and angle-length shoes. The estimated clothing insulation including office chair was 0.83 clo (ISO 7730) in both conditions. The participants’ main activity was typing and the estimated activity level was 1.1 met. According to the PMV-model, the temperatures 23.5 °C and 29.5 °C are estimated to be neutral and slightly warm, respectively (Maula et al., 2015). Production of irrelevant speech and masking sounds ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Speech was used as the sound source in both Condition A and B. Four loudspeakers in the four corner desks were used to simulate a situation where four persons would be talking on the phone one after the other about different topics so that a different plot was ongoing in each corner of the room (Figs. 2 and 4). In most previous experiments investigating the effects of speech on task performance, the speech levels have been manipulated electronically. Our experiment did not contain electronic manipulation of speech levels to demonstrate the effects of two extremely different room acoustic conditions in such way which could also be implemented in real open-plan offices. The sound power level of speech emerging from the loudspeakers was constant in both conditions. The speech heard by the participants depended solely on the room acoustic treatment of the room, i.e., the changes in the amount of room absorption materials and the level of masking sound (see Chapter 2.7). A four-channel sound file (wav file, 44.1 kHz) was used to produce the speech to four loudspeakers (Fig. 4). Only one corner speaker was active at a time. The speech was obtained from four radio programmes bought from YLE (The Finnish Broadcasting Company). The speech for each loudspeaker originated from a different radio programme so that the topic in each corner was different. In each programme, four participants (politicians, celebrities or specialists) were discussing a topic of common interest. The speech of each participant in each programme was isolated from the programme and placed to one channel of the four-channel sound file. The speech material of the speaker in each channel was thereafter edited so that the speech consisted of separate 5-to-25-s-long tokens. The four-channel speech file was then expanded and arranged so that there was no overlap between the channels. The order of the speakers was pseudo-randomized but the total amount of speech from every speaker was equalized. A silent period of 1–8 s was inserted between speech tokens. The sound level of each speech token was adjusted to the same A-weighted level. The editing was done using audio editing software (Adobe Audition 3.0). In addition, the spectrum of each speaker was modified so that the octave band spectrum shape deviated from the speech spectrum of ISO 3382-3 by less than 3 dB. The lengths of the two four-channel recordings used in the experiment were approximately 180 min. Two versions of speech recordings were created and counterbalanced across conditions. The versions included different topics and speakers. The loudspeakers having mouth-like directivity were used (Genelec 6010). The height was 1.2 m from the floor. The speakers were directed to the centre of the room. The participants could see the loudspeakers when they entered the room but the speakers were not visible while working. The sound power level of the speech was constant and equal between every channel. This was checked out by full-time sound power level measurements of each channel in a reverberant room according to ISO 3741. The speech level (effort) was set between normal and casual. The A-weighted mean level over all directions was 53.0 dB at 1 m distance in a free field (normal effort is 57.4 dB). The linear sound power levels (LW re 1 pW) were 51.3, 55.3, 53.6, 44.6, 38.9, 32.2 and 29.4 in octave bands 125, 250, 500, 1000, 2000, 4000 and 8000 Hz, respectively. The overall level of masking sound was 9 dB higher in Condition A than Condition B. However, the spectrum shape of the masking was constant and close to Brownian noise in octave bands 125–4000 Hz (Fig. 5) both in Condition A and B. This spectrum has been found adequate both in workplaces and in laboratory conditions (Hongisto, 2008; Hongisto et al., 2012, 2014). The difference in masking levels between the desks was negligible (±1 dB) because of the smooth distribution of masking loudspeakers. Implementation of room acoustic conditions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In Condition A, sound-absorbing boards (EN 11654:1997; class A, 20 mm mineral wool) were installed to the whole ceiling (75 m2). Sound-absorbing absorbers (EN 11654:1997; class A) were installed behind the porous linen wall pictures (18 m2). The materials had the highest available M1 classification so that the emission levels of various compounds were very low. The overall room absorption area including the furniture was very large, altogether 123 m2. The artificial masking system was set on so that the mean level in the desks was 45 dB (LAeq). As a result of efficient sound masking and high room absorption area, the intelligibility of the speech was moderate when the speech was originating from the nearest desk (2 m away) and reasonably low when speech was originating from a desk further away (6 m, see Fig. 2 and Appendix). It must be noted that the speaker producing the speech was changed every 6–33 s. Therefore, the STI of speech varied with time and the variation is indicated in Table 1. In Condition B, sound reflecting ceiling boards were installed (EN 11654; unclassified, plasterboard). The sound-absorbing boards were removed behind the linen wall pictures (EN 11654; unclassified). The overall room absorption area was very small mainly caused by furniture, altogether 30 m2. Ventilation sound was planned to be the only masking sound. However, the artificial masking sound was played at a level of 33 dB (LAeq) to reach the same masking level in all desks since the ventilation level varied from 33 to 36 dB between desks. The final masking level was reasonably low (36 dB) and the resulting masking efficiency was negligible as desired. As a result of the inefficient sound masking level and negligible room absorption area, the STI value of speech was very high and almost independent on the distance to the speaker (see Appendix). Neither the screen height nor the screen absorption was modified between Conditions A and B. The screen height was 130 cm and the screens were not sound-absorbing (EN 11654; unclassified). The room acoustic conditions were measured according to ISO 3382-3 (See Appendix). Cognitive tasks ~~~~~~~~~~~~~~~ Six cognitive tasks were used: a serial recall task, an operation span task, an N-back task, an information search task, a typing task and a story-writing task. The tasks mimicked the cognitive processes required in many types of office work. All tasks rely on several cognitive processes, but only the main processes that are typically related to these tasks in the literature are mentioned in the description of the tasks. This does not exclude the possibility that cognitive processes that are not typically associated with the particular task type may also have some effect on the performance of these tasks (Sörqvist, 2014). The serial recall task, the operation span task and the information search task were programmed with Visual Basic 6 (Microsoft) and the N-back task with E-Prime 2.0 (Psychology Software Tools Inc.). The serial recall task is a classic short- term memory task where participants have to recall randomly presented digits from 1 to 9 in the correct presentation order. Numbers were presented on the screen one by one at the rate of 1 per second with an inter-digit interval of 1.5 s. During recall, the numbers 1 to 9 appeared in a 3 × 3 array on the screen and participants responded by clicking with the mouse the numbers in the recalled order. Participants were instructed to guess or press ‘empty’ in case they did not know the correct answer in a certain serial position. After each trial, participants had 15 s to respond, after which a new trial began. A total of 12 trials were presented but the first two were excluded from the analysis. The main dependent variable was the percentage of digits recalled in the correct serial position. The task took about 7 min to complete. The operation span task is another highly used working memory task. The version used in our study is based on the original operation span task developed by Turner and Engle (1989). The task consisted of mathematical equation- word pairs, such as 3 × 3 + 7 = 22 and ‘BOOK’. First, an equation appeared on the computer screen. Participants had 10 s to decide whether the presented equation was true or false by clicking the appropriate option on the computer screen. After each equation, a word appeared on the screen and participants had 2 s to memorize the word. Then the next equation and word were presented. After a predetermined number of equation-word pairs was presented (the number of presented equation-word pairs increased gradually from three to eight), participants typed in all the words they remembered. Small misspellings were allowed and the precise order of recalled words was not required. In this task, the function of the arithmetic task is to interfere with memorizing the words. To ensure that participants focused on both equations and words, participants were instructed to get at least 85% of responses to equations correct. Participants received feedback on this after typing each word list. The materials used for words and equations are described in detail in Haka et al. (2009). Two matched versions of the task were constructed and counterbalanced between the conditions and sessions. Within each version, equations and words were presented randomly. The main dependent variable was the percentage of correctly remembered words. The task took about 12 min to complete. The N-back task (Gray, Chabris, & Braver, 2003) requires working memory and the ability to sustain attention. In this task, sequences of letters were presented on the screen one by one, each for 500 ms with a 2500 ms inter-stimulus-interval. Three difficulty levels were used (0-back, 1-back and 2-back). In 0-back, the participant's task was to press YES every time the letter ‘X’ appeared on the screen and press NO for all other letters. This is the baseline condition not taxing working memory. In 1-back, participants were required to respond whether the presented letter was identical to the one immediately preceding it. In 2-back, participants were required to response whether the presented letter was identical to the one presented two trials back. Participants were instructed to respond quickly, but accurately. Answers were given by pressing a key on the keyboard labelled YES (leftwards arrow) or NO (downwards arrow). Each set (0-, 1- and 2-back) included 30 + n letters (n = 0, 1 or 2) in a pseudorandom order. One set included 9 letters requiring a ‘YES’ response (30%). Upper and lower case letters were varied requiring participants to use abstract letter codes and preventing them from relying only on visual letter feature. The whole task included three blocks containing one set of each difficulty level, i.e., each difficulty level appeared three times during the task. The materials and the order of the difficulty levels were counterbalanced across the test blocks and participants. Six matching versions of the task were constructed and counterbalanced across the conditions and sessions. The main dependent variables were response accuracy (%) and reaction time of the correct responses (in ms). Reaction times deviating over 2.5 SD from the participant's mean were excluded. The reliability of the reaction time measurement was based on Maula et al. (2015). The task took about 20 min to complete. The information search task (Jahncke & Halin, 2012) requires working memory, attention and strategic executive functions. The task bears similarity to a problem-solving task. In this task, a table of 20 rows and seven columns was presented on the screen. In each row, an object was presented (e.g., a country) and each column described one aspect of the object (population, multilinguality, area, highest point in metres, major religion, etc.). In the descriptions, both categorical and numerical values were given. Participants were required to search for the object that met the presented criterion (see below for examples). Participants had a maximum of 60 s to respond, after which the next question appeared. The questions had the same grammatical structure. Altogether, four tables were used, two for both sessions. The task included 20 questions, 10 questions concerning the first table and 10 questions concerning the second table. Half of the questions were simple and half of the questions were difficult. In the case of difficult questions, the object could be nominated by investigating three columns, one with categorical and two with numeric values (e.g., “Which country is multilingual, its highest point over sea level exceeds 1000 m, and has the highest population?”). In the case of simple questions, participants had to follow two columns, one with a categorical value and one with a numeric value (e.g., “Which employee is a salesperson and has the highest annual income?”). Two matching task versions were constructed and counterbalancing across the conditions and sessions was made in relation to the task difficulty and the features of the objects and columns including categorical and numeric values across the tables. The main dependent variables were the percentage of total number of correct responses and the percentage of correct responses in simple and difficult questions separately. The number of exceeded response time was also analysed. The task took about 20 min to complete. The typing task is a routine task that requires different sub-processes of perception, attention and motor functions (Rumelhart & Norman, 1982). In the task, participants had to copy a text presented on paper by typing it using the computer keyboard. Participants were required to type quickly but avoid making mistakes. Participants were instructed to correct any typing errors they made. The text was a story (Juvonen, 2012a, 2012b) with a rich vocabulary and plot. The story printed on paper was placed vertically on a slightly slanted rack for easy reading; the participants were allowed to freely place the rack on the desk. Participants had 10 min to copy the text. Two stories with similar writing style from the same writer were selected and counterbalanced between the conditions and sessions. The number characters in the final text, writing fluency, the number of corrected and uncorrected errors, and the total time of pauses exceeding 2 s were analysed. Corrected errors refer to typing errors that a participant made but corrected, whereas the errors that a participant did not correct were labelled as uncorrected errors. There were no spelling errors in the original text that was given to the participants. Writing fluency was operationalized as the total number of characters produced during the given time (i.e., the sum of the number characters of the final text and the total number of deleted characters; Sörqvist, Nöstl, & Halin, 2012). The task was carried out and the data were collected using the ScriptLog program (Strömqvist & Karlsson, 2002). The story-writing task requires psychomotor performance, creativity and the ability to produce text in writing (Sörqvist et al., 2012). In the task, a photograph was presented on the screen. Participants were required to write a story about the scene depicted in the photograph. Participants were instructed to write whatever came to mind from that picture. The presented photographs displayed two different nature scenes (a forest road surrounded by green trees and a snowy mountain view with a small cottage in the middle). Participants had 5 min to write the story and they were instructed to write as much as possible but correct their typing mistakes. The presented pictures were counterbalanced between the conditions and sessions. The ScriptLog program (Strömqvist & Karlsson, 2002) was used to collect and analyse the data. The stories were analysed both quantitatively and qualitatively. In the quantitative analysis, the same variables as in the typing task were analysed. In the qualitative analysis, the quality of the stories (atmosphere, the existence of negative turning point, the invocation of the photograph and concreteness) was analysed by two raters. These factors were found adequate for most stories written by the participants. Inter-rater reliability was assessed using the overall percentage of agreement. Agreement varied between 53 and 94%. Questionnaires measuring subjective experience ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Five different questionnaires were created. Participants responded to the questionnaires at the baseline phase (questionnaire A), at the beginning of the acclimatization phase (questionnaire B, the first session or questionnaire E, the second session), before a short break during the experimental phase (questionnaire C) and at the end of the session (questionnaire D) (Fig. 6). In questionnaire A, background information including gender, age, sleep during the preceding night, ability to function normally, noise sensitivity, bodily symptoms and experienced level of tiredness and motivation were measured. Noise sensitivity was measured with three items from Weinstein (1978) and with four items from Noise-Q (Schutte, Marks, Wenning, & Griefahn, 2007). The intensity of the six symptoms was measured (Table 3). Experienced cognitive fatigue were measured with a modified version of the Swedish Occupational Fatigue Inventory (SOFI; Åhsberg, Gamberale, & Gustafsson, 1998; Åhsberg, Gamberale, & Kjellberg, 1995) including three factors (tiredness, lack of energy and motivation). Every factor included three items (tiredness: sleepy, yawning, drowsy; lack of energy: worn out, exhausted, drained; lack of motivation: uninterested, indifferent, passive). Questionnaire B included questions about experienced tiredness and motivation, and symptoms, but also questions about overall thermal comfort, satisfaction with temperature, draught and emotional reactions. Overall thermal comfort was measured using a 7-point scale (1 = cold, 7 = hot). Emotional reactions were measured with the modified version of the Zuckerman Inventory of Personal Reactions and Feelings (ZIPERS; Zuckerman, 1977), including four factors (Fear/arousal, Positive affect, Anger/aggression and Attentiveness) and one item measuring sadness (I feel sad). Fear/arousal was measured with three items (My heart is beating fast; I feel fearful; I am breathing fast), Positive affect with four items (I feel carefree; I feel gentle; I feel happy; I feel like I could act friendly to someone), Anger/aggression with two items (I feel like hurting or telling off someone; I feel angry), and Attentiveness with two items (Cronbach's alphas 0.53–0.89). Attentiveness is not reported in the results, as it overlaps in content with one symptom item (Difficulties in concentration). Questionnaire C contained questions about overall thermal comfort, satisfaction with temperature, draught, emotional reactions, experienced tiredness and motivation, and symptoms but also included questions of subjective workload. Subjective workload was measured with four items (mental demand, frustration, effort and performance) which were modified from the NASA Task Load Index, NASA-TLX (Hart & Staveland, 1988; Moroney, Biers, & Eggemeier, 1995). Workload questions were included because subjective awareness of distraction may lead participants to invest more effort in performance (Schlittmeier, Hellbrück, Thaden, & Vorländer, 2008). Questionnaire D contained questions about overall thermal comfort, satisfaction with temperature, draught, emotional reactions, experienced tiredness and motivation, symptoms and subjective workload. In addition, this questionnaire included questions of local thermal comfort (ten items), work effectiveness in a similar temperature (one item, Table 3), acoustic satisfaction, acoustic and visual privacy (three items, Table 3), experienced disturbance due to different factors present in the work environment (17 items) and subjective experience of the work environment (eight items). Acoustic satisfaction was measured with five items (habituation, disturbance, pleasantness, attention capture, and work efficiency in a similar sound environment). Items were combined to form a sum score (henceforth called Acoustic satisfaction; Cronbach's alphas 0.82–0.90). The disturbance caused by lightning, screen brightness and ergonomic conditions were also rated to rule out possible effects due to these factors. Questionnaire E contained questions about overall thermal comfort, satisfaction with temperature, draught, emotional reactions, experienced tiredness and motivation, and symptoms. In addition, this questionnaire included background information questions: the amount of sleep during the preceding night and ability to function normally. All questionnaires were presented with an Internet-based software (QuestBack, Finland), except for the thermal comfort questionnaire, which was also presented on paper.","The experiment was conducted between February and April 2012. Participants took part in the experimental sessions during two mornings on consecutive weeks. On the first morning, participants were asked to arrive at the laboratory at 8.30 a.m. for the baseline phase. The acclimatization and the experimental phase was carried out between 10.30 a.m. and 12.30 p.m. on both days. At the beginning of the first session, participants were given written information about the experiment and they signed the informed consent form. Participants were informed of the progress and the content of the test sessions, payment, confidentiality and their right to interrupt their participation at any time. The procedure is described in Fig. 6. The first session started with a baseline phase in silence in a room near the experimental room. During the break after the baseline phase, sandwiches and soft drinks were served in the hallway to ensure that the participants had eaten breakfast prior to the rather long experimental session. During the acclimatization phase, participants were acclimatized to the experimental condition and they practised the cognitive tasks (30 min). Performance was not measured during the acclimatization phase. Participants were unaware of the purpose of the acclimatization phase. The speech sounds were off. When the experimental phase started, the experimenter switched on the speech. Participants were instructed to ignore the sounds and to concentrate on the tasks. Before every task, the experimenter shortly repeated the task instructions, after which the participants performed the task at their own pace. A short break was given in the middle of the experimental phase to give the participants an opportunity to visit the restroom and drink water. During the break the speech was switched off. The temperature in the hallway to the restroom and in the restroom was similar to that in the laboratory. Conversation was not allowed during the experiment or the break. Prior to the beginning of the second session in the next week, sandwiches and soft drinks were served in the hallway. The acclimatization phase included questionnaire E and the revised versions of each task excluding the typing task. The experimental phase was identical to the first session. Participants were informed in detail about the aim of the study at the end of the second session. Statistical methods ~~~~~~~~~~~~~~~~~~~ SPSS (IBM SPSS Statistics for Windows 20, IBM Corp, Armonk, NY) was used for the statistical analyses. The normality of data was tested with the Kolmogorov–Smirnov test. The serial recall, the operation span, the typing and the story-writing tasks were analysed using repeated measures ANOVA with experimental condition (later: condition) as the within-participants variable. The quality of the stories in the story-writing task was analysed using Pearson's chi-squared test for independence. The N-back and information search tasks were analysed with a 2 (condition) × 3 (difficulty level) repeated measures ANOVA. A repeated measures ANOVA was also used for the questionnaire items that were normally distributed or when distributions were similarly skewed. Subjective workload was analysed with a 2 (condition) × 2 (exposure time: after one hour vs. after two hours) repeated measures ANOVA. Experienced tiredness and motivation, and emotional reactions were analysed with a 2 (condition) x 3 (exposure time: beginning, after one hour, after two hours) repeated measures ANOVA. Whenever needed, the homogeneity of variance was estimated with Mauchly's test of sphericity. When Mauchly's test indicated a violation of sphericity, the Greenhouse-Geisser correction was applied and the corresponding p-values are reported. When an interaction was found, paired comparisons between conditions were performed using t-tests for the variables that were normally distributed or when distributions were similarly skewed; otherwise the Wilcoxon signed-rank test was used. An alpha level of 0.05 was used in all analyses. The Benjamini-Hochberg procedure (Benjamini & Hochberg, 1995) was used for alpha-error adjustment in paired comparisons. In the N-back task, the data were combined over the two blocks, because the main effects of the block remained non-significant. One participant was excluded from the N-back analyses because of a misunderstanding of the instructions. Unlike in typical studies of the operation span task, participants who failed to achieve an 85% level on equation accuracy were not excluded, because the independent variable (speech intelligibility) could also have affected arithmetic performance (Jachnke et al., 2013; Schlittmeier et al., 2008). Instead, multivariate outliers were checked using Mahalanobis distance for identifying possible changes in performance strategy. As a result, one participant was excluded from the analysis. In the story-writing task, one participant was excluded from the analysis because he did not follow the instructions. The decrement in performance, D [%], was defined as D = 100·(1 − PB/PA) where PB and PA are the performance levels in Condition B and A, respectively. The results for the four items (mental demand, frustration, effort and performance) measuring subjective workload yielded a main effect of condition for all items and there were no interactions between exposure time and condition. Thus, the items were combined to form a sum score (henceforth called Subjective workload; Cronbach's alphas 0.66–0.77). During the experiment, the visual appearance of the office was also investigated using a between-groups design with two groups. Thirty-one participants experienced Conditions A and B in visual environment 1 and thirty-four participants in visual environment 2. Because the visual appearance did not significantly affect the objective or subjective results concerning the effect of condition, the groups were combined. There was one exception: the self-rated impairment of performance due to thermal conditions and air quality (heat, odours and stuffiness of the indoor air) was affected by the visual appearance. Thus, only the other group was included in the analyses of these variables (n = 31). For the remaining analyses, the participants from both groups were included. Performance measures ~~~~~~~~~~~~~~~~~~~~ In the serial recall task, a significant main effect of condition was found for the percentage of digits recalled in the correct serial position (F1,64 = 5.86, p = .018, η2 = 0.08; Fig. 7a). Performance was significantly better in Condition A (D = 6.7%). In the operation span task, a significant main effect of condition was found for the percentage of correctly recalled words (F1,63 = 10.84, p = .002, η2 = 0.15; Fig. 7b), with better performance observed in Condition A (D = 4.0%). In the N-back task, a significant main effect of condition was found for response accuracy (F1,63 = 4.01, p = .049, η2 = 0.06; Fig. 7c). Significantly more correct responses were given in Condition A. However, the performance decrement was negligible (D = 0.5%). There was no significant interaction between difficulty level and condition. That is, the effect of condition on accuracy was not affected by the difficulty level. Reaction times were not affected by condition nor was there an interaction between difficulty level and condition. The performance in the information search task was not affected by condition, neither by difficulty level (p > .05). Moreover, there was no significant interaction between difficulty level and condition (p > .05). In the typing task, a significant main effect of condition was found for the total number of errors (F(1,64) = 4.03, p = .49, η2 = 0.06; Fig. 7d). Significantly fewer errors were made in Condition A (D = 5.2%). This resulted primarily from the number of uncorrected errors: there was a marginally significant difference in the number of uncorrected errors (p = .05), but no difference in the number of corrected errors (p > .05). None of the other variables measuring typing fluency were affected by condition. The performance in the story-writing task was not affected by condition for any measure (p > .05). The mean values with standard deviations for the task variables where significant differences were not found (N-back, information search, typing and story- writing) are presented in Table 2. Subjective responses ~~~~~~~~~~~~~~~~~~~~ The results for the subjective assessments of the working conditions are presented in Fig. 8. The ratings for the conditions for working as a whole were significantly more positive in Condition A than B (Z = −5.97, p < .001). Similarly, the ratings of the possibility to work effectively were significantly higher (Z = −4.66, p < .001) and ratings of the riveting of the tasks were significantly more positive (Z = −2.18, p = .029) in Condition A. There was a significant main effect of condition for the sum score of Subjective workload (F1,64 = 15.02, p < .001, η2 = 0.19; Fig. 9a). Subjective workload was rated to be significantly lower in Condition A. The effect of condition on Subjective workload was not affected by exposure time, as indexed by a non-significant interaction between exposure time and condition. Acoustic satisfaction was significantly lower in Condition B (Z = −4.82, p < .001; Fig. 9b). The subjective assessments of the distraction of performance due to different sounds are presented in Fig. 10. Speech distracted significantly more in Condition B. Speech was more distracting both from nearby desks (Z = −2.95, p = .003) and desks from further away (Z = −5.53, p < .001). Moreover, speech from nearby distances was more distracting than speech from further away both in Condition B (Z = −5.20, p < .001) and in Condition A (Z = −6.29, p < .001). Similarly, a significant effect of condition was found for the distraction due to the sounds of computer tapping (Z = −2.22, p = .026) and other sounds made by other participants (Z = −2.68, p = .007). The distraction of the hum of ventilation, i.e., masking sound, was significantly higher in Condition A (Z = −4.70, p < .001). However, it should be noted that speech either from nearby desks (Z = −5.42, p < .001) or desks from further away (p = .06) distracted significantly more than the hum of ventilation in Condition A. That is, speech was the main acoustic distractor in both experimental conditions. Overall thermal comfort differed between conditions: participants felt warmer at the end of the exposure time in Condition B (Z = −7.00, p < .001); the difference between conditions was observed during the whole exposure time (p < .001). Using the 7-point scale of thermal comfort, participants felt warm (M = 6.3, SD = 0.7) at 29.5 °C and neutral (M = 3.6, SD = 0.9) at 23.5 °C after the second exposure hour (Z = −7.00, p < .001). A similar result was found for both local and overall thermal comfort during the whole exposure time (p < .001). The impairment of performance due to thermal factors is presented in Fig. 11. Overall, working efficiency was assessed to be significantly higher in Condition A (Z = −6.5, p < .001; Table 3). Heat was rated to interfere with performance significantly more in Condition B (Z = −4.60, p < .001), whereas cold (Z = −2.20, p = .028) and draught (Z = −2.08, p = .037) interfered significantly more in Condition A. However, the mean ratings of cold and draught were very low in both conditions. Self-rated air quality differed between conditions as expected (Fig. 11). Stuffiness of the indoor air (Z = −3.71, p < .001) interfered with self-rated performance significantly more in Condition B. No impairment caused by odours was reported in either condition. Condition also affected the perception of acoustic and visual privacy (Table 3). Participants would have preferred more screens around the desk in Condition A in order to make working more pleasant (Z = −2.20, p = .028). The distraction caused by other people and closeness of nearby desks did not differ between experimental conditions (p > .05). There was a significant main effect of condition for the perceived tiredness (F1,64 = 10.40, p = .002, η2 = 0.14), the lack of energy (F1,64 = 20.18, p < .001, η2 = 0.24) and the lack of motivation (F1,64 = 20.55, p < .001, η2 = 0.24; Fig. 12). Participants were more tired and the lack of energy and motivation was higher in Condition B. There was also a significant interaction between the exposure time and condition in the lack of energy (F2,128 = 9.62, p < .001, η2 = 0.13) and motivation (F2,128 = 9.56, p < .001, η2 = 0.13). Paired comparisons revealed that the lack of energy was significantly higher in Condition B in the beginning phase (t(64) = 7.05, p < .001, two-tailed) and after the first (t(64) = −3.68, p < .001, two-tailed) and the second (t(64) = −4.98, p < .001, two-tailed) exposure hour. There was no significant difference in the lack of motivation between the conditions in the beginning phase but the lack of motivation was significantly higher in Condition B after the first (t(64) = −3.10, p = .004, two-tailed) and the second (t(64) = −4.90, p < .001, two-tailed) exposure hour. In Condition B, an increase in the lack of energy (t(64) = 2.65, p = .013, two-tailed) and motivation (t(64) = 2.28, p = .033, two-tailed) was observed already between the beginning phase and the first exposure hour. The lack of energy (t(64) = 6.39, p < .001, two-tailed) and motivation (t(64) = 5.57, p < .001, two-tailed) continued to increase, being significantly higher after two hours of exposure compared to one hour of exposure. In Condition A, no significant difference in the lack of energy and motivation was found between the beginning phase and the first exposure hour, but the lack of energy (t(64) = 4.50, p < .001, two-tailed) and motivation (t(64) = 3.60, p = .001, two-tailed) increased significantly between the first and the second exposure hour. Regarding emotional reactions (Table 3), there was a significant main effect of condition for Positive affect (F1,64 = 4.18, p = .045, η2 = 0.06; Fig. 13) with higher values in Condition A. There was also a significant interaction between condition and exposure time (F2,116 = 4.15, p = .022, η2 = 0.06). Paired comparisons revealed that there was no significant difference between the conditions in the beginning phase or after the first exposure hour but Positive affect was stronger in Condition A after two hours of exposure (t(64) = −3.80, p < .001, two-tailed). In Condition B, Positive affect decreased significantly during the whole exposure time (p < .001) and difference was found between the beginning phase and the first exposure hour (t(64) = 4.10, p < .001, two-tailed) and between the first and the second exposure hour (t(64) = 3.86, p < .001, two-tailed). In Condition A, Positive affect decreased significantly between the beginning phase and the first exposure hour (t(64) = 3.33, p = .002, two-tailed) but no significant difference was found later between the first and the second exposure hour (p > .05). The ratings of Sadness, Fear/arousal and Anger/aggression remained low during the sessions. However, there was a main effect of condition on sadness (F1,64 = 6.13, p = .016, η2 = 0.09) with higher ratings in Condition B. There was also a marginal main effect of condition on Fear/arousal (F1,64 = 3.98, p = .05, η2 = 0.06). Participants reported higher Fear/arousal ratings in Condition B. Anger/aggression was not affected by condition. There was no significant interaction between condition and exposure time in these ratings. Different physiological symptoms (Table 3) were enquired on a scale from 1 (Not at all) to 5 (Very much). On the whole, the intensity of symptoms was very mild. Thus, a rating of 2 (Slightly) or higher was interpreted to indicate the existence of a symptom. The self-reported prevalence of different symptoms at the end of the experiment is shown in Fig. 14. Participants reported more feelings of being unwell in Condition B than in Condition A after the first (Z = −2.74, p = .009) and second (Z = −3.83, p < .001) hour of exposure. Symptoms increased with exposure time in Condition B (p < .05) but not in Condition A. The assessments of headache and nasal, throat and eye symptoms remained low in both conditions (Fig. 14), but the symptoms increased toward the end. Participants reported more headaches at the end of the session in Condition B than in Condition A (Z = −2.43, p = .030). No significant differences were found for other exposure times (p > .05). In addition, more throat symptoms were reported in Condition B after the first (Z = −2.20, p = .033) and the second (Z = −2.29, p = .033) hour of exposure compared with Condition A. Headache and throat symptoms also increased with exposure time in Condition B (p < .05) but not in Condition A. Ratings of nasal symptoms decreased during the first exposure hour compared to the beginning phase in Condition A (p < .01), but not in Condition B. However, a difference was not observed between the conditions (p > .05). Ratings of eye symptoms increased during the second exposure hour compared to the beginning phase in Condition B (p < .05). No significant changes were perceived in Condition A. However, no difference was observed between the conditions (p > .05). Difficulties in concentration were higher in Condition B; the experimental conditions differed both after the first (Z = −2.10, p = .046) and second (Z = −3.03, p = .004) exposure hour.","The aim of our study was to investigate the simultaneous exposure to highly intelligible irrelevant speech, high temperature and low ventilation rate and the effects of exposure on cognitive performance and environmental satisfaction in an open-plan office laboratory. As hypothesized, Condition B had detrimental effects on both cognitive performance and subjective experience. Condition A was perceived to be a better condition for office work. Cognitive performance ~~~~~~~~~~~~~~~~~~~~~ The results established negative effects of Condition B on cognitive performance. An effect was observed in the percentage of correct answers and in typing errors. In the serial recall and the operation span tasks performance was worse in Condition B in comparison to Condition A. Participants also made more errors in typing when working in Condition B. Thus, Condition B proved to be poor IE for working. The results from the single IE factors of prior experiments (see Fig. 1) are conceivable to reflect with our study where simultaneous exposure was used. Highly intelligible irrelevant speech reduced performance and interfered with the operation of working memory (Haapakangas et al., 2014). Similarly, high temperature reduced working memory performance (Maula et al., 2015). Instead, low ventilation rate did not have consistently effects on cognitive performance (Koskela et al., 2014). Based on previous findings concerning single factors, high speech intelligibility and high temperature decreased performance the most. The results of our study expound that simultaneous exposure to high speech intelligibility, high temperature and low ventilation rate decrease performance. Despite of the findings of Koskela et al. (2014), the self-rated distraction of work performance caused by the stuffiness of the indoor air, heat (Fig. 11) and speech from further away and nearby desks (Fig. 10) received higher ratings in Condition B in our study. Taking our experimental design into account, where the effects of individual factors were not investigated, we can only conclude that high speech intelligibility, high temperature and low ventilation rate may have together affected the performance results through subjective responses in our study. As pointed out above, performance differences were clearly found in the working memory tasks employed in our study. Regarding the N-back task, it is noteworthy that accuracy decreased in Condition B in comparison to Condition A, but the reaction times did not. Thus, response speed was maintained at a cost of more erroneous responses. This result is inconsistent with the speed-accuracy trade-off hypothesis (Hockey, 1984), which suggests that in noisy environments responses are given faster but with less accuracy. The reason may be that Hockey's experiment was conducted with pseudorandom noise instead of speech. The results of the N-back task shall be compared with the prior studies conducted in the same laboratory. Similar background speech (Haapakangas et al., 2014) did not decrease the accuracy in the N-back task when three difficulty levels (0, 1 and 2) were applied as in our study. However, an effect on accuracy was observed when four difficulty levels (0, 1, 2 and 3) were used in high temperature (Maula et al., 2015). Taken together, it appears that simultaneous exposure to three IE factors in Condition B exacerbated somewhat the performance decrement in the N-back task compared with that was observed with similar single factor manipulation alone. In the typing task, fewer typing errors were made in the better IE. In previous studies (Park & Yoon, 2011; Wargocki et al., 1999, 2000) the number of typed characters has been found to decrease with high material emissions in a similar text typing task while no significant difference has been observed in typing errors. On the other hand, Koskela et al. (2014; Fig. 1) obtained no effect of air quality on typing performance in a study conducted in the same laboratory as our study. A possible reason for this discrepancy in results is the fact that in our study not only ventilation rate but also thermal and acoustic conditions were manipulated. In other words, a set of poor IE conditions may need to be combined together to have an effect. Yet, methodological differences in manipulating air quality (human vs. material emission) and differences in task duration (10 min vs. nearly one hour) may also have contributed to the discrepancy. In the story-writing task, participants had to produce a new text instead of copying an existing text. Story-writing was not affected by the experimental condition. Keus van de Poll et al. (2014) observed a detrimental effect of irrelevant speech on writing fluency and pauses in story-writing with STI values below 0.34. Processing of semantically meaningful speech was assumed to interfere with writing performance. It should be noted that in our study the STI values in both experimental conditions exceeded the value of 0.34 and the effect obtained by Keus van de Poll et al. could not be confirmed. The information search task was not affected by the experimental condition. This is consistent with the study of Jahncke et al. (2013) who also observed no effect of speech intelligibility level on information search. It must be noted that neither Jahncke et al. nor our study included a silent condition (or STI 0.00) as a reference to different speech conditions so it cannot be concluded that speech does not affect performance in this task. However, the selected set of IE conditions in our study did not reveal any effect. It would be useful to investigate in the future whether this task is affected by speech or other IE conditions, because the task represents normal office work better than many other tasks normally used in experiments like this. Working memory has a central role in the tasks employed in our study. However, because cognitive performance relies on several cognitive processes, it is difficult to identify all processes that may have been interfered by indoor environmental factors. In the area of noise-related performance effects, one suggestion is that the effects of noise on performance result from attentional capture rather than the impairment of other cognitive processes (Sörqvist, 2014). Attentional capture is caused by a sudden auditory change which draws attention from the task to the deviant event (Hughes, Vachon, & Jones, 2007). The effect of attentional capture is assumed to be more detrimental to performance when the task difficulty increases. In our study, the level of task difficulty varied between the tasks. Thus, one possible reason why the effects of conditions were not found in all tasks might be the variation in task difficulty. Because there were also other adverse factors than auditory ones, we suggest that the results cannot be explained purely by attentional capture hypothesis. Moreover, high temperature has been related with increased stress which activates attentional resources to cope with stress (Hancock et al., 2007). This has been shown to lead to a situation where the capacity to process task-relevant information is reduced (Hancock et al., 2007). Thus, all environmental factors should be observed together with the impairment of other cognitive processes when the results are interpreted. Practical limitations of room acoustic design should also be considered when interpreting the overall effect of acoustic design on cognitive performance. Haapakangas et al. (2014) demonstrated that room acoustic design affects task performance and acoustic satisfaction but acoustic design has an effect only on speakers located at least 3–5 m from the listener. Nearby speech is difficult to control by room acoustic means. Our study confirms this finding. In our study and that of Haapakangas et al. (2014), speech intelligibility was temporally variable, because the location of the active speaker in the room varied every 6–33 s, and the distance between the speaker and the participant varied respectively. Thus, speaker's direction and distance was not constant. In both experimental conditions, task performance might have momentarily deteriorated when speech was heard from nearby desks. However, as task performance was overall better in Condition A, it may be assumed that performance was not adversely affected by speech when the speaker was far away from the participant. On the other hand, in Condition B speech was intelligible regardless of the location of the speech source. This probably explains the observed differences in working memory performance between the two experimental conditions. In many previous studies silence and highly intelligible speech has been compared to each other (Hongisto, 2005). If the high speech intelligibility condition, Condition B, would have been compared to silence, differences in performance would probably have been more evident. However, our study involves a realistic open-plan office environment where silence is not expected to exist for very long periods of time. Therefore, it is argued that the methodologies and the findings of this experiment have better practical relevance than most previous studies except that of Haapakangas et al. (2014). Subjective responses ~~~~~~~~~~~~~~~~~~~~ The working conditions were assessed to be significantly better in Condition A than in Condition B. On the one hand, the subjective measures support the findings on performance measures discussed above. On the other hand, they demonstrate even more pervasive effects than performance measures. This is in line with previous studies demonstrating that subjective measures are more sensitive to changes in IE than objective measures (Haapakangas et al., 2011, 2014; Schlittmeier et al., 2008). In addition, despite whether or not the open-plan office conditions are objectively measured as adequate, subjective experiences have an effect on what kind of meaning the occupant will give to the positive and negative characteristics of IE (Cox & Ferguson, 1994; Lahtinen, Huuhtanen, Kähkönen, & Reijula, 2002). Room acoustics When the acoustic environment was properly designed (Condition A), acoustic satisfaction was higher and distraction from different sounds smaller. This result was found even though the overall noise level was higher in Condition A due to the masking sound. In both experimental conditions, the most distracting sound was speech from nearby desks. Distraction by masking sound was measured by an item called “hum of ventilation” since the masking sound resembles the hum of ventilation. This was justified because the concept of sound masking is unknown among the general population. Even though the hum of ventilation (i.e., the masking sound) was perceived as more distracting in Condition A, the masking sound was more beneficial than harmful, since the most distracting sound, i.e., speech, was assessed significantly less disturbing in Condition A. Moreover, speech coming even from a distant location was rated more distracting than the hum of ventilation. Thus, it is evident from these data that speech was the main acoustic distractor. The findings are in agreement with previous studies (Haapakangas et al., 2011, 2014; Haka et al., 2009). Thermal conditions and air quality Stuffiness of the indoor air was rated as more interfering in Condition B. In the previous study (see Fig. 1), Koskela et al. (2014) did not find any significant change in experienced stuffiness when ventilation rate was reduced from 28 l/s·person (600 ppm CO2) to 2 l/s·person (2200 ppm CO2) while the room temperature was kept constant (23.5 °C). Instead, a significant change in stuffiness was observed by Maula et al. (2015) when high room temperature was compared with neutral temperature. This supports the previous result that when an individual feels warm in a room, air quality is also assessed to be poor (Lan et al., 2011). Based on previous results, it can be suggested that room temperature had a significant role in increased stuffiness ratings in Condition B. However, simultaneous exposure to high room temperature and poor air quality might also have an effect on stuffiness ratings together. Tiredness, energy and motivation Perceived lack of energy and motivation increased during the two-hour period, the increase being stronger in Condition B. In addition, participants were more tired in Condition B. Previously, a quite similar background speech as in condition B did not decrease arousal (i.e., tiredness) when silence was compared with a speech condition (Haapakangas et al., 2011). Similarly, Maula et al. (2015) did not find an effect of high room temperature on tiredness, energy or motivation. However, low ventilation rate (high CO2 concentration) increased the lack of energy and motivation although ratings remained rather low while the quality assessments of IE overall remained unaffected (Koskela et al., 2014). Based on these previous findings, it seems probable that the simultaneous exposure to highly intelligible speech, high temperature and low ventilation rate intensified the lack of energy and motivation more than the exposure to only one factor at a time could cause. Subjective workload In Condition B, subjective workload was higher than in Condition A which may reflect higher effort to compensate for the anticipated performance decrement. Experience of stress may be one mechanism leading to higher effort and subsequent disruption of performance (Hancock & Warm, 1989). As the higher subjective workload coincided with the performance decrement in Condition B, it appears that compensatory efforts were needed but were not sufficient to compensate for the worsened working conditions. Moreover, the enhanced effort hypothesis is an often- mentioned explanation for the observed differences between performance and questionnaire measures (Haapakangas et al., 2011; Schlittmeier et al., 2008). Subjective experience of disturbance, such as acoustic distraction, might lead individuals to invest more effort in their performance to compensate for the effect of distraction. In our study, the maintenance of good performance with the help of higher effort might have reduced the differences in performance between conditions; yet, significant differences were nevertheless found. Similar effects of compensatory effort have been reported by other researchers (Szalma & Hancock, 2011). Although enhanced effort could have diminished performance decrement during the two-hour work period in our study, subjective workload might accumulate across longer time periods if the IE conditions causing stress persist. In working life, individuals are exposed to specific IE conditions several hours per day. It has been shown that coping strategies are in use to combat the disturbance of sounds in open-plan office (Kaarlela-Tuomaala, Helenius, Keskinen, & Hongisto, 2009): increased effort, longer breaks, slower working rhythm and longer working days have been reported. Physical symptoms Our results showed that perceived somatic symptoms were at a lower level in condition where the IE was well designed, which is in agreement with previous findings. The prevalence of somatic symptoms was higher in Condition B although the intensity of symptoms was relatively mild. It is notable that even with a few hours of exposure the prevalence of somatic symptoms was increased (see also e.g., Lan et al., 2011). Emotions In previous studies, effects of IE on experienced emotions have been investigated very little. In our study, positive emotional responses decreased with exposure time in both experimental conditions, but emotional responses were more positive in Condition A after the second exposure hour. This is in agreement with the study of Lan et al. (2011) who found high temperature (30 °C) to be associated with negative mood. This result further reinforces the idea that indoor environment can affect emotional comfort even in the absence of somatic symptoms. Limitations of the study ~~~~~~~~~~~~~~~~~~~~~~~~ Limitations are related to methodological issues. First, participants were exposed to the experimental conditions on two separate days on subsequent weeks since it was not possible to build two open-plan offices. This might have caused measuring error despite a repeated measures design. Retest reliability of task performance, indicating the consistency of a test across time, has been indicated to differ depending on task type (Ellermeier & Zimmer, 1997), and this might be one error source. Second, a two-hour exposure time was used. Overall, the effects of noise on task performance have been shown to diminish with exposure time (Szalma & Hancock, 2011). This may not be the case with speech sounds (Haapakangas et al., 2014; Szalma & Hancock, 2011). Instead, as argued by Haapakangas et al. (2014), it is possible that cognitive and subjective impacts of noise might even increase over time as a result of an emerging stress response or decreasing compensatory resources. Increased exposure time associated with the intensity of thermal condition has been reported to increase the negative impact of temperature on performance (Hancock et al., 2007). This is in line with the Maximal Adaptability Model (Hancock & Warm, 1989). Because activation level in our study was low, it is probable that increasing exposure time by one or two hours would not have affected the results. Instead, it would have produced practical problems, such as a need for a lunch break. Overall, it is assumed that a significant increase in the exposure time might reveal more robust effects of temperature on task performance. In addition, the duration of thermal exposure might have more impact on performance and subjective responses when work is more mobile and the activation level is higher. Third, e.g., Hancock et al. (2007) have considered that time of day might affect the relationship between temperature and performance. In our study, the sessions were conducted in the morning to create optimal condition to perform without tiredness and thus, minimize the possibility of intervening variables. Our results should be applied with reservation when considering performance at another time of day. Fourth, the IE conditions were planned to combine the experimental conditions of prior experiments conducted in the same laboratory (Fig. 1). However, it is possible that one IE factor dominated the effect on performance or subjective response over the other factors and had more impact on the final results. The performance results should be applied with care to real workplace environments. Our results might only be valid for individually performed tasks in visual modality and requiring intensive use of working memory. Our results may not apply to teamwork or communication tasks. In the future it is important to also include tasks in the auditory modality. For example, it would be interesting and important to study how irrelevant background speech might affect oral communication.","This study provides strong evidence that the combination of high intelligibility of irrelevant speech, high room temperature and low ventilation rate impairs the perceived working conditions and cognitive performance. It is possible to suggest that by designing room acoustic conditions, thermal conditions and ventilation rate adequately, satisfaction with work environment is increased, somatic symptoms are decreased and the possible impairments of work performance can be avoided. The experiment was the final study in a series of four experiments. The simultaneous exposure to these three adverse factors might intensify some effects of IE. This is significant because IE problems are often multifaceted in workplaces. In practice, the ventilation rate, room temperature and acoustic conditions can vary significantly between office workplaces so that our suggestions should not be generalized to all possible configurations of indoor environment. Our study supports the view that special care should be paid to the holistic design of indoor environment in open-plan offices."],["Ever since the advent of high-rise architecture, in the late nineteenth century, the modern city has been a distinct locus of vertiginous experience. Whilst the correlation between vertigo and tall buildings might at first appear to be an obvious one, it is in fact a variable function of ever-evolving techniques and materials, as well as depending on the psychosocial conditions that underlie the experience of space at a given place and time. A great deal of public interest has been aroused over the past decade by the proliferation of glass-floored viewing platforms, which have become increasingly popular features of observation decks around the world. These platforms, often branded as ‘dare to walk on air’ experiences, are designed to challenge the user's perception of spatial depth. Whilst older types of viewing galleries, such as open decks with low walls, could provoke stronger feelings of height vertigo than new glass floors (not least because of the imagined agency of throwing oneself off the high point), the latter are distinct insofar as they are designed to conjure the thrill of walking over the abyss in a seeming state of suspension. The rise of glass floors gathered momentum in the mid-noughties, at a time of rapid growth of vertical cities (King, 2004) that saw a new wave of ‘supertall’ and ‘megatall’ buildings emerge around the world. Social implications of contemporary high-rise construction have been investigated from various perspectives, with regard to the vertical dimension of cities (Graham and Hewitt, 2012; Graham, 2016); the role of the skyscraper in architectural culture (Nobel, 2015) and in tourism-led urban regeneration (Leiper and Park, 2010); and the psychological influence of tall buildings on their occupants (Gifford, 2007). Meanwhile, the vertical visualisation of space brought about by the combined use of satellite imagery and digital technologies, such as Google Earth (Di Palma, 2008), has affected the conditions of embodied seeing as well as the physical experience of vertigo, ushering in a new ‘age of aerial vision’ (Gilbert, 2010). The recent surge of thrill-seeking practices such as rooftopping, which has sparked a broad diffusion of ‘vertigo inducing’ images on the web, is further signal of a wider shift in urban experience and representation (Deriu, 2016). Seen together, these phenomena are symptomatic of a wider socio-cultural condition that appears to be pervaded by a dizzying spatiality, particularly acute in vertical cities. The term vertigo crops up in architectural and urban discourse rather frequently, albeit mainly in a figurative sense. A case in point is the eponymous Glasgow exhibition (1999), where ‘The Strange New World of the Contemporary City’ was illustrated through an assortment of architectural projects that ranged widely in function and scale – from Tate Modern in London to the Ontario Mills shopping mall in California. In the exhibition volume, Tate Modern's architect Jacques Herzog remarked: ‘The word “vertigo” does not have auspicious connotations. In fact, it would seem to address the sinister and even dangerous side of things: fear of heights and the attendant dizziness. Or even a double anxiety: the fear of falling passively through no fault of one's own, and the fear of responding quasi-actively to the magical attraction of the abyss and thereby succumbing to its vertiginous appeal. “Vertigo” could be said to express an inescapable ambivalence and indeterminacy.’ (Herzog, 1999: 6) The show did not have a specific agenda, nor did it claim to present a consistent design approach. Instead, the curatorial strategy aimed to capture the generic state of ‘dizziness’ and ‘disquiet’ provoked by the global architectural landscapes of the 1990s (Moore, 1999). The provocative, and somewhat prophetic, title echoed the spiralling tension that runs through Alfred Hitchcock's Vertigo: an enduring point of reference for cinematic representations of the city as a protean emotional landscape. Two decades on, the ‘strange new world’ portrayed on the eve of the Millennium appears all too familiar. Indeed, the ‘double anxiety’ evoked by Herzog has meanwhile taken up a new dimension. Today, vertigo aptly describes the physical sensation that is induced by architectural elements such as the fashionable glass-floors that line many high-level walkways. This vogue is epitomised by the Shanghai World Financial Centre (SWFC), which was featured in the Glasgow exhibition when only the foundations had been laid. Its construction was eventually completed in 2008 following design alterations that increased the overall height of the tower to 492 m. The SWFC tower contains an exemplar of the current trend for immersive viewing experiences: besides panoramic vistas of the surrounding cityscape, the 100th-floor observation gallery, situated at 474 m of height, offers vertical views through a series of transparent glass panes built into the floor (see Fig. 1). The SWFC has since boasted ‘the world's highest observatory’, a record that may not last for long as the global race to the sky carries on unabated. The construction of similar design features around the urbanised world suggests that architectural vertigo has become a sought-after phenomenon. This trend raises questions about the conditions in which space is designed, perceived and experienced in the contemporary city. What bodily experiences are implicated in the ‘states of suspension’ induced by these platforms? What are their material and spatial properties? And how do these spaces relate to the wider socio-economic context in which they are produced? A cross-disciplinary approach will inform the investigation of these issues, while a series of design projects will serve to illustrate various manifestations of the subject. The proposed interpretation draws on theories of transparency and sensory experience of space, supported by insights from psychology, and ultimately critiques high- level glass platforms as ‘tourist bubbles’ that crystallise, quite literally, a social imperative of the present moment.","The term vertigo is fraught with multiple and ambivalent meanings that cannot be exhausted in an article. Some elucidation may nonetheless inform a critique of current architectural trends. Whilst in medical discourse vertigo is usually treated as a symptom of balance system disorders, in popular culture the word is used more loosely to evoke various sensations of giddiness, dizziness, and disorientation that are associated with a perceived loss of equilibrium. Dictionary definitions range widely, from the illusion of physical movement (‘the act of whirling round and round’) to the bodily perception related to it (‘swimming in the head’), and extend to figurative meanings (‘a disordered state of mind, or of things, comparable to giddiness’) (OED). In a figurative sense, the term has also been adopted to describe the precarious conditions of life in contemporary societies. For Young (2007: 12), ‘Vertigo is the malaise of late modernity: a sense of insecurity of insubstantiality, and of uncertainty, a whiff of chaos and a fear of falling.’ Accordingly, a generalised feeling of giddiness defines our ‘liquid’ modernity (Bauman, 2000), an epoch in which values that previously had a solid foundation, such as social status and economic position, have become increasingly fluid and unstable. The notion of ‘groundlessness’ has gained currency in art, architectural, and urban discourses over the past decade (Dorrian, 2009; Steyerl, 2011; Graham, 2016). Dorrian (2009) in particular draws comparisons between the ‘dissolution’ of ground evoked by Hitchcock, as well as by authors like Nabokov and Sebald, and the disorienting experience induced by contemporary architectures such as London's City Hall. At the bottom level of this Foster-designed building, visitors can stand on a giant aerial photomap of London, taking symbolic possession of the city from a vantage point that traditionally signifies a position of power and control. Dorrian (2009: 86) sees this as ‘an attempt to architecturally stage […] democratic transparency’, in a similar mould as the glass dome of the new Reichstag in Berlin – designed by the same architect. At City Hall, the abstract miniaturization of London produced by the aerial view somehow jars with the act of walking on the photomap while looking down in search of familiar clues. If, on the one hand, this embodied experience provides the visitor with a sense of grounding, on the other hand the photomap triggers a ‘vertiginous multiplicity’ that causes an opposite ‘ungrounding’ effect: a disorientation amplified by the non-hierarchical code of representation that distinguishes satellite imagery from cartographic maps (Dorrian, 2009: 91). The ‘vertiginous ungrounding’ described by Dorrian calls to mind philosophical conceptions of modernity as an existential condition perturbed by unprecedented degrees of freedom: namely, the ‘dizziness of freedom’ expounded by Kierkegaard in his 1844 treatise The Concept of Anxiety (LeBlanc, 2011). However, while Dorrian plays down the role of heights in the etymology of vertigo and its aforementioned cultural representations, there is evidence to suggest that elevation is in fact increasingly bound up with the dizzying experience of contemporary urban space. As we shall see, the current popularity of elevated glass platforms vividly illustrates how height vertigo is elicited in visceral ways through particular spatial arrangements. Remarkably, the notion of vertigo is conspicuous for its absence from the main texts that have defined the critical discourse on the perception of the city from above. Barthes (1983/1964) omitted it from his ‘Eiffel Tower’ essay, wherein he coined the term ‘architectures of vision’ (architectures de la vue) to describe the 19th-century structures that turned the ‘fantasy of a panoramic vision’ into material reality. Accordingly, the Tower epitomised the rise of a modern form of perception that made it possible to embrace a bird's eye view of the metropolis and thereby to comprehend it in its structure. By positing the city as a text to be read and deciphered, Barthes's argument privileged the power of intellection over the corporeal experience of space. A similar tendency to reduce the observation tower to a mere vantage point characterises De Certeau's (1984/1980) later description of Manhattan from atop the World Trade Center, wherein the ‘voluptuous pleasure’ of seeing the city as a whole is described as a purely visual act. The observation deck is the platform from which a dieu voyeur observes the city ‘like a picture’ and lays claim over it. The ensuing ‘fiction of knowledge’, concludes De Certeau (1984/1980: 92), ‘is related to this lust to be a viewpoint and nothing more.’ While these now-classic accounts remain critical for our understanding of urban modernity, and of the observation deck as one of its enduring topoi, they leave unquestioned the embodied experience of space that was induced by such unprecedented heights. How does one feel when confronting the city from high vantage points? This is not to diminish the cognitive function of viewing galleries, nor to dispute their significance as social spaces implicated in procedures of spectacle and surveillance; but rather to address another dimension of high-rise architecture that has become all the more relevant in light of the current vogue of glass floors. The design of platforms that expand the field of vision to the vertical axis affects the practice of viewing cities from above in ways that have not been interrogated as yet. In order to comprehend this phenomenon, let us turn to another theory that puts a different spin on the notion of vertigo. In his seminal book, Man, Play and Games, Caillois (2001/1958) revived the ancient Greek word ilinx (‘whirlpool’) to define one of the four fundamental categories of human play. With this term Caillois identified those games ‘which are based on the pursuit of vertigo and which consist of an attempt to momentarily destroy the stability of perception and inflict a kind of voluptuous panic upon an otherwise lucid mind.’ (Caillois, 2001/1958: 23) He remarked that ilinx is an ancient form of play, as exemplified by the rituals performed by whirling dervishes and pole-flying voladores. In the course of history, fairgrounds became traditional spaces for games of vertigo, with their ‘machines for rotation, oscillation, suspension, and falling, constructed for the sole purpose of provoking visceral panic.’ (Caillois, 2001/1958: 133) Crucially, Caillois observed that ilinx was a pervasive aspect of modern life, typified by the exhilaration induced, for instance, by speed and drugs. Historically, the quest of ‘voluptuous panic’ described by Caillois found new expression in the 19th century through a range of mechanised thrill rides epitomised by the rollercoaster. While gravity plays had long been popular social entertainments in funfairs and circuses, the amusement park turned the modern city into a playground for thrilling pleasures (Kane, 2013). These ‘mechanised entertainments’ emerged at a time when a series of bodily practices aimed at defying the force of gravity – such as high diving, parachuting, and funambulism – became mass spectacles. In this respect, the 19th century has aptly been called ‘the gravity century’ (Soden, 2003). It is interesting to note that Caillois singled out tightrope walking as the activity that most closely corresponds to ilinx: by turning the ludic – and, therefore, unproductive – nature of play into a performance art, the wire-walker ‘[moves] through space as if the void were not fascinating, and as if no danger were involved.’ (Caillois, 2001/1958: 137). This practice, which burgeoned as a public spectacle in the 1850s, over the twentieth century moved from the abysses of nature to those of cities. In recent decades, skillful acrobats like Philippe Petit and Nik Wallenda have displayed their ability to master vertigo in the midst of breathtaking urban environments (Johnston, 2013). Caillois's anthropology of play has inspired urbanists to engage with the ludic aspects of contemporary cities. Stevens (2007: 43), for instance, draws extensively on the concept of ilinx in an attempt to unearth the creative potential of public spaces: ‘[v]ertigo negates instrumental benefit and embraces risk for its own sake and the affirmation of human bodily experience.’ From a different angle, Caillois's theory might also inform the critical analysis of high-level glass platforms as stages where games of ilinx are socially performed. For the challenge of ‘walking on air’ in some way simulates the acrobatics of highwire walkers who confront the vertiginous heights of cities. As shown in the next section, the vogue of glass floors makes the thrill of ‘skywalking’ accessible to the general public within safe and protected environments – usually for the price of a ticket.","The use of glass in the design of walking surfaces is not new, and has a particularly rich background in retail interiors. Notably, the dramatic glass staircases designed by Eva Jiřičná from the 1980s onwards paved the way for a trend that, over the past decade, has been spread globally by the Apple store chain. In the noughties, while translucent glass became a popular means of allowing sunlight through multi-storey spaces, the design of transparent glass floors also gained momentum. A typical use can be found in exhibition spaces such as archeological sites and museums, where visitors are drawn to contemplate objects on display underneath their feet. Although transparent glass panes are usually placed over shallow spaces, there are exceptions that engage a bodily interaction with deeper voids. At the Acropolis Museum in Athens, for instance, structural glass was used extensively not only to exhibit the outdoors archeological excavations but also to create elevated, see-through surfaces that allow users to walk over the museum's central space (Self, 2014). Different considerations preside over the design of glass-bottomed platforms in tall structures, where the main purpose is to open up downward-looking perspectives onto the yawning void underneath. The first high-level transparent floor was fitted in the observation deck of Toronto's CN Tower, where in 1994 a horizontal grid of tempered glass panes was mounted on the viewing gallery 342 m above the ground (see Fig. 2). After a few early emulations, this project set a worldwide trend only a decade later. The CN Tower boasted the highest glass-floored observation deck in the world until 2008, when it was overtaken by the Shanghai World Financial Center (SWFC): by that time, the global craze for skywalks had become pervasive. Whilst the CN Tower heralded the glass floor as a means of enhancing the tourist appeal of high-level viewing galleries, the SWFC Observatory embodies its evolution to a full-fledged spatial concept. Located in the midst of Shanghai's financial district, the skyscraper symbolizes the boom of high-rise construction that took place in China amidst major changes in the global economic geography around the turn of the 21st century. Glass floors furnish the walkways that span the 55 m-long gallery on the tower's 100th floor – a transparent bridge suspended nearly half a kilometer above the ground. The ‘Skywalk 100’ epitomises a new type of urban observatory: a free-standing gallery that envelops the visitor's body within a glassy capsule-like space from which they can watch the city from above while enjoying the thrill of ‘walking on air’. The dizzying quality of this space is further heightened by the reflecting ceilings and the outward inclination of the fully-glazed side walls. As the SWFC website proclaims: ‘Walking on the three transparent glass-walled walkways, visitors will experience the feeling that Shanghai lies at their feet.’ (SWFC) Although the possibility of gazing at the city is still very much part of the attraction, the skywalk signals a move from Barthes's architectures of vision towards what might be called ‘architectures of vertigo’. No longer a pure machine for seeing, the observation deck is reconfigured as a machine for thrilling: a spatial enclosure that combines the visual spectacle of the city with an altogether more visceral experience. Over the past decade, glass floors have become recurring features of viewing platforms around the world.1 Two notable examples are located in Chicago, the historical birthplace of both the skyscraper and the observation wheel. In 2009 the Willis Tower (formerly Sears Tower) had its 103rd- floor Skydeck retrofitted with a series of retractable balconies designed by SOM – the firm behind the original 1970s tower. This feature, dubbed ‘The Ledge’, broke new ground in the design of high-altitude viewing platforms owing to its four all-glass boxes that protrude over 1.3 m from the building façade at 412 m above ground (see Fig. 3). The challenge to step on a fully transparent balcony at such dizzying heights is advertised as a brave act: ‘Get out on the Ledge – if you dare’ (Skydeck Chicago). Those who dare shall be rewarded: ‘With glass on the ceiling, floor, and all sides, it is truly, an unforgettable experience.’ (ibid.) Behind the platitude of this marketing pledge lies a deeper economic strategy that pits observation towers vying for tourists' attention against each other – globally as well as locally. In 2014 a new attraction called ‘Tilt’ opened across town at the former Observatory of the John Hancock Center, in overt competition with Skydeck Chicago. The feature consists of a glass-and-steel platform that tilts thirty degrees outwards, allowing visitors to lean with their bodies over a 300 m-chasm (see Fig. 4). The conception of this immersive space, branded ‘Chicago's Highest Moving Experience’ (360 Chicago), signalled that even one of the highest viewing galleries in the world, capable of 360-degree views of the cityscape, needed a facelift to keep up with the competitors. The Tilt presents an interesting variation on the theme of high- level transparent features: an alternate tendency towards the design of moveable elements aimed at stimulating a dynamic, vertigo-inducing experience. The game of architectural ilinx here requires that you let yourself go, quite literally: transported over the void, the visitor-cum-player enjoys the adrenaline rush provoked by the momentary disruption of sensory stability. The dynamics of thrill-seeking have been taken to a new extreme at the U.S. Bank Tower in Los Angeles, where an entirely glass-cladded attraction called ‘Skyslide’ was built in 2016 as an extension of the ‘Skyspace’ observation deck. Hanging on the building's façade at circa 300 m of height, the Skyslide entices visitors to ‘experience […] unparalleled views in a whole new way as they glide from the 70th to the 69th floor of the U.S. Bank Tower.’ (OUE Skyspace LA) Local commentators have noted that the Skyspace project aimed to bring tourism revenue to an area of Downtown L.A. with rising vacancy rates in commercial skyscrapers (Hawthorne, 2016). The costly makeover of the tower's observatory has been regarded as an attempt to reclaim, and reenchant, the view from above in a sprawling city where the aerial gaze is mainly associated with the helicopter view (ibid.). The Skyslide marks a further step in the design of high-altitude glass boxes. This transparent device induces the ‘voluptuous panic’ of a free fall into space: a fast yet safe descent that derives its allure, yet again, from the attraction of the void. Meanwhile, a parallel development has taken place in Europe, where glass platforms have become familiar attributes of tall structures – albeit at a more modest scale. In the U.K. the craze for glass floors has pervaded new-build and historical architectures alike. Portsmouth's Spinnaker Tower, whose sail-like shape echoes the Burj Al Arab hotel in Dubai, was completed in 2005 to spearhead the regeneration of the harbour area: the glass skywalk that was built on one of the viewing decks (100 m) is deemed to have played a part in the tower's touristic success. More recently, some of the country's major attractions have also been retrofitted with similar features. Blackpool Tower's ‘Walk of Faith’, initially built in the late-1990s as a simple glass-floor panel, was later expanded into a whole skywalk (116 m) as part of the refurbisment unveiled in 2011 by the ‘global leisure’ giant Merlin Entertainments. Less dramatic in terms of height, yet more interesting in other respects, is the 2014 retrofitting of London's Tower Bridge with glass floors on the panoramic walkways that connect the towers. While the Grade I-listed structure had previously undergone functional renovations, the insertion of 11 m-long, see-through platforms responded to the intent of attracting more users. The revamped twin galleries double up as venues for hire: visitors attending corporate events, private parties, or yoga classes can now revel in – or recoil from – the act of skywalking over the Thames. The relatively limited height (42 m), if compared with other such walkways, allows for a closer visual connection with the passersby and the traffic unfolding on the road bridge down below. The incongruous effect of transparency triggers, in people either side of the walkways, not only sheer curiosity but also an awareness of the reciprocal distances and positions within the newly-opened visual field (see Fig. 5). A similar intervention took place, around the same time, in the modern ‘architecture of vision’ par excellence – the Eiffel Tower. In their restyling of the Tower's first floor (57 m) architects Moatti-Rivière designed a series of glass platforms that aim to achieve an ‘augmented architecture’ (Trétiack, 2014). Besides improving the existing public spaces, the refurbishment included open-air, inward-looking terraces that open up the central void through a combination of horizontal and vertical transparency: ‘The project offers an improved experience of the Tower and Paris, an entertaining sensory experience, a journey of the senses and knowledge.’ (Eiffel Tower, n. d.) In a Living Architectures documentary, Alain Moatti explains the idea of magnifying the feeling of void that, historically, has been the primary characteristic of the Tower: a peerless monument which, with reference to Barthes, the architect credits with ‘the invention of third dimension in the city.’ (Bêka and Lemoine, 2014) In a telling scene, a visitor likens the sense of lightness he felt on a glass floor to levitation. The impression of floating on air is further enhanced by the physical contiguity between the horizontal surfaces and the inclined glass parapets that delimit the floor space, which are wholly transparent (see Fig. 6). Whilst at Tower Bridge the glass walkways were the main raison d’être of the renovation, in the Eiffel Tower they are comparatively minor features of a wider project. The restyling of the Tower is nonetheless symptomatic of a widespread tendency to produce ever-more immersive experiences of space. High-level platforms that challenge users to ‘walk on air’ manifest a pursuit of architectural ilinx – a contemporary version of the ‘voluptuous panic’ described by Caillois. These design features have become common not only in urban settings but also in natural environments, as shown by several projects aimed at heightening the scenic effects of landscapes around the world. The prime example is the cantilevered glass skywalk at Grand Canyon West, Arizona, opened in 2007 and reportedly inspired by Toronto's CN Tower.2 Vertiginous structures of this kind play a significant part in the promotion of landscapes that are branded as tourist attractions. As Porter (2015: 207) notes, they signal ‘a tourist development which, like the quest for the tallest building, is a never- ending exercise in one-upmanship, leading further and further into exaggeration.’ The glass floors that have concurrently appeared in observation towers provide an urban analog of the experience of natural landscapes. Indeed, the fact that these platforms were first introduced in cities and only later in natural landscapes is an interesting sign of how the urban environment, with its built-up gorges and canyons, has become a major source of attractions for thrill tourists. In our urban age, in which verticality is an increasingly common dimension of city life (Graham, 2016), architecture has become instrumental to the production of spatialised games of ilinx.","The notion of transparency is key to understanding the spatiality of glass-floored urban observatories. As Forty (2000: 286) notes, transparency is ‘a wholly modernist term, unknown in architecture before the twentieth century.’ Its basic sense, ‘meaning pervious to light, allowing one to see into or through a building’ (ibid.) became popular in the 1910s-1920s, when advances in glass manufacturing and frame construction made it possible to build self-standing glass enclosures. The coupling of structural frame with glass panes, which had an ancestor in the stained-glass windows of Gothic cathedrals, was such a breakthrough in modern construction that has been described as ‘the most significant development in architecture in the last millennium.’ (Fierro, 2003: viii-ix). A manifest expression of this shift was the curtain wall, which gained broad diffusion after World War II amidst a steep increase in structural glass production. It was in reaction to that trend that Rowe and Slutzky (1963) elaborated the concept of phenomenal transparency, a perceptual quality defined as the ‘illusion of spatial depth’ that characterised modernist architecture as well as avant-garde painting. This theory signalled an attempt to transcend the literal transparency of the International Style: that is, the basic optical quality of the modernist ‘glass box’. Rowe and Slutzky's sophisticated argument could do little, however, to halt the proliferation of curtain-walls worldwide. The architectural uses of glass have vastly expanded over recent decades, as this versatile material proved suitable to an increasing number of functions while also providing a symbol of political power (Fierro, 2003). By the turn of the 21st century, with ever more tower blocks dotting the skylines of cities, and the curtain wall defining a global aesthetic (Elkadi, 2006), literal transparency had become so widespread as to conquer the horizontal dimension. And yet, despite their growing popularity, transparent floors are barely mentioned in the literature about glass in architecture (Cruz, 2013). This is all the more remarkable if we consider the novelty of glass platforms that transpose the modernist frame construction onto the horizontal plane. The alteration of the floor into a see-through surface, akin to a horizontal window, calls to mind the notion of ‘fifth façade’ formulated by Le Corbusier in the late 1920s – when he envisioned new functions for the flat roof terrace after flying over South American cities. It might be argued that, by overturning the vertical window-wall onto a horizontal surface, the introduction of the glass floor has ushered in the sixth façade. This epithet befits the lower side of elevated buildings and overhanging building elements, insofar as they fulfil the condition of externality that is implied by the physiognomic etymology of the word ‘façade’: that is, in the case of glass platforms, the possibility of looking through from without as well as from within. For the purpose of the present argument it should be useful to distinguish, with some degree of approximation, between ‘low-level’ and ‘high-level’ glass platforms. The former, such as those at Tower Bridge and the Eiffel Tower, are usually installed in, or added to, the underside of existing structures and effectively operate as horizontal floor-windows that enable a two-way visual contact between inside and outside. Although the main goal of these platforms is to allow for a top-down vision, a bottom-up gaze from below is also afforded, to varying degrees, by the relatively low elevation. Higher platforms hanging several hundred meters above the ground, such as those at the SWFC Observatory and Skydeck Chicago, may not be clearly visible from the street but can nonetheless be seen from within the same buildings as well as adjacent ones. Moreover, the floors' undersides gain wide exposure through media representations, such as brochures and websites, where lower viewing angles are favoured by photographers to visualise the skywalkers' bodies. The recognition of the field of visibility that is opened up by these transparent surfaces prompts further questions about the spaces they define and the actual uses they enable. What kinds of responses and interactions are elicited by the sixth façade? A complex reciprocal relationship is established between internal users and external viewers. The simple fact that these platforms must remain tightly-sealed in order to perform their function means that visual contact through the glass occurs in a state of spatial separation. The bodies of those who ‘dare to walk on air’ are usually visible from below through a peculiar honeypot effect, as skywalkers perform their balancing acts in front of curious spectactors on both sides of the glass stage. There are nuanced interconnections between how people feel in these spaces and how the latter affect their emotions, since height vertigo manifests itself through a wide range of psycho-physiological responses. Depending of where one sits in the spectrum of ‘height tolerance’ (Salassa and Zapala, 2009), varying degrees of thrill or anxiety can be elicited by states of suspension that appear to defy the laws of gravity.","While skywalks are built in highly controlled and safe environments, they are designed to challenge the user's fear of heights by exposing them to the view lying underneath their feet. The structure encasing the glass panes, usually made of multiple-layer tempered glass laid out in a grid, doubles up as a comfort zone where hesitant users can reach for a sense of safety. Those who initially skirt around a glass platform often approach it through small and tentative steps as they confront the fear of the void. This experience transcends the visual field insofar as the act of ‘skywalking’ engages the subject's proprioceptive system as well as the optokinetic one – two distinct perceptual apparatuses whose signals to the brain, if discrepant, may trigger a sensation of dizziness (Yardley, 1994). In fact, the sensory stimulation caused by glass platforms is not confined to the five senses but rouses a sixth sense traditionally known as kinaesthesia, or the ‘muscle sense’. As Çelik (2006) points out, this sense was discovered in the early 19th century by psychologists who realised that muscles are capable of receiving sensations from the spinal cord and should therefore be considered to be sentient: ‘Kinaesthesia, the sense of bodily movement […] [refers] to those unclassifiable sensations that could not be traced accurately to one of the five known sense organs, but seemed to originate from the undifferentiated mass of the viscera.’ (Çelik, 2006: 159). Although the development of neuroscience has greatly expanded the scientific knowledge of body-mind relations, the notion of kinaesthesia remains relevant to the present discussion. It resonates with the prevailing conception of architecture as a field of multi-sensory experience, which, especially since the 1990s, has been largely informed by phenomenology (Holl et al., 2006). Pallasmaa (2006, 2012) in particular sought to ‘re-sensualise’ architecture by discerning the emotional states that are involved in the embodied experience of space, in contrast to the hegemony of ‘retinal architecture’ in modern western culture. By championing the return to a sensory architecture, Pallasmaa referred to the muscular tensions through which we apprehend the built environment, and on which we project our movements through a process of ‘bodily identification’ with place. He stressed the role of proprioception in the experience of space and recognised the unconscious desire to defy gravity that architecture can elicit: ‘The sense of gravity is the essence of all architectonic structures and great architecture makes us conscious of gravity and earth. Architecture strengthens verticality of our experience of the world. At the same time that architecture makes us aware of the depth of earth, it makes us dream of levitation and flight.’ (Pallasmaa, 2006: 37). Accordingly, the heightened consciousness of gravity that is produced by awe-inspiring architectures is the source of ‘memorable experiences’. What remains unaccounted in this theory is the realm of sensory experiences that, in the presence of spatial depths, can engender emotions of fear and anxiety as well as comfort and pleasure. Further insights from psychology shed light on the ‘thrill of transparency’ that makes glass floors so appealing – although not for everyone. The movements registered by the sensorimotor nerves define our overall body schema and thereby underpin the ways we organise our actions in space. As we have seen, the design of glass floors is intended to amplify the spectacle of the view from above by engaging the user's sense of balance when confronted with the sight of heights. When standing in front of a glass platform, a moment of realisation occurs whereby the incongruity of the view provokes an intense kinaesthetic feeling: an instinctive response that calls to mind the early experiments with depth perception based on the ‘visual cliff’ (Gibson and Walk, 1960). The kinaesthetic faculty that makes us aware of our bodily position in space (i.e., the ‘eye-object distance’) triggers the subjective feeling of imbalance that may cause discomfort and dizziness in ‘height intolerant’ subjects. Elevated glass floors are among those ‘encounter spaces’ that cause a high perception of risk in acrophobic sufferers (Andrews, 2007). At the opposite end of the spectrum are ‘height tolerant’ individuals who cope well with spatial depth and, in some cases, find excitement and exhilaration in the experience of altitude (Salassa and Zapala, 2009). The thrill of transparency induced by glass floors is what drives height-seeking individuals to revel in the psychological sense of danger that keeps others at bay. For thrill seekers, the sixth façade therefore constitutes a playground for the experience of the sixth sense. While the present essay does not claim to provide a systematic analysis, these observations begin to delineate the spatial and social context in which the experience of elevated glass floors is situated. The links between physical sensations and the emotions related to the experience of heights introduce a further degree of complexity due to their fundamentally subjective nature. Psychoanalytic research suggests that feelings of anxiety and pleasure associated with vertigo are interwoven, as bodily perceptions reflect – and reveal – our dynamic mental states (Quinodoz, 1997). The ways in which inner drives are channelled through actions and behaviours is bound to determine varying levels of height tolerance through an individual's life span. To consider emotions of pleasure and anxiety as inextricably bound up might therefore help us to understand the irregular and mutable occurrences of vertigo that often punctuate people's lives. These embodied experiences are, however, invariably situated within a material and social context: hence the importance of considering the wider conditions in which glass floors and the related ‘thrills of transparency’ are produced.","The idea of memorable experience advocated by the proponents of an architecture of the senses is an unholy bedfellow with the coeval ‘experience economy’ theory. With this influential term, Pine and Gilmore (1998, 1999) named a step change in ‘the progression of economic values’ whereby businesses seek to gain a competitive edge by staging experiences that are purportedly memorable: ‘While prior economic offerings – commodities, goods, and services – are external to the buyer, experiences are inherently personal, existing only in the mind of an individual who has been engaged on an emotional, physical, intellectual, or even spiritual level.’ (Pine and Gilmore, 1998: n. p.) The production of commodified, and increasingly customised, experiences spread from the entertainment industry to other sectors such as travel and retail, and extended to architecture as well (Lonsway, 2009). The tourist industry was among the first to embrace this process. Already in the 1970s MacCannell (1999/1976: 21) noted in his semiotic analyisis of tourist attractions: ‘Increasingly, pure experience, which leaves no material trace, is manufactured and sold like a commodity.’ Subsequently, Bauman (1996) pointed out that tourists, unlike other social types of travellers, are mainly driven by ‘pull’ rather than ‘push’ factors. In other words, their mobility is governed by aims (‘in order to’) rather than causes (‘because of’): ‘the tourist is a conscious and systematic seeker of experience, of a new and different experience, of the experience of difference and novelty – as the joys of the familiar wear off and cease to allure. The tourists want to immerse themselves in a strange and bizarre element (a pleasant feeling, a tickling and rejuvenating feeling, like letting oneself be buffeted by sea waves) – on condition, though, that it will not stick to the skin and this can be shaken off whenever they wish.’ (Bauman, 1996: 29) These ideas, which to a large extent are still valid today, shed further light onto the rise of glass floors as ‘experience design’ products. To contemplate a cityscape through a horizontal window-floor is in itself a ‘strange and bizarre’ attraction that draws scores of visitors to immersive viewing galleries. Skywalks arguably constitute typical examples of ‘tourist bubbles’ (Judd, 1999): self-contained and highly regulated spaces in which moments of leisure can be enjoyed within secluded and safe environments. Marking a shift from the traditional mechanisms of panoramic vision, these spaces presuppose an expanded function of the tourist gaze involving a multi-sensuous, kinaesthetic experience: in other words, they are stages for the performance of ‘embodied actions’ (Urry and Larsen, 2012: 190). These actions are almost invariably recorded on camera, as the act of photographing or filming one's body suspended over the void is a popular means of validating the memorable experience. Undeterred by the Global Financial Crisis of the late noughties, or perhaps even spurred by it, the skyscraper business has continued to tap into the growing market of vertiginous experiences. The design of multi-storey viewing galleries on top of ‘supertall’ buildings such as London's Shard and, more recently, New York's One World Trade Center has led to considerations that ‘observation decks have become cash machines.’ (Brown, 2014) Vying to attract visitors, today's viewing galleries stage the embodied experience of urban heights as an increasingly immersive event. Whilst vision is still predominant, its mechanisms and practices are increasingly augmented through immersive spatial experiences that find a parallel in the surge of technologies such as Virtual Reality, 4D cinema, and the like. A technological variation on the theme of the glass floor is shown by the ‘sky portal’ built into the observatory's floor at One World Trade Center, which opened to the public in 2015. This attraction consists of a round 14 ft-wide platform where visitors can stand and watch a live video image of the scene unfolding at street level. The thrill of the glass floor is simulated by a high-definition livestream: in lieu of real transparency, the screen technology reproduces views from above to be watched in a vicarious state of suspension. By replacing the direct sight of height with an immersive cinematic spectacle the ‘sky portal’ takes the suspension of disbelief to a new level: further evidence of how deeply architectures of vertigo are embedded in the structures of a buoyant experience economy. The pursuit of memorable experiences leads to the production of spaces that, while largely homogenised, at the same time boast their own distinctive features – or, in business parlance, ‘unique selling points’.","This study suggests that observation decks are increasingly conceived as spaces of visceral thrills. The vogue of high-level glass floors could easily be dismissed as a passing fad driven by the ‘form follows fun’ principle; and yet, upon closer scrutiny this phenomenon reveals deeper socio-spatial implications. When considered together, the cases discussed above show a consistent shift from the realm of architectures of vision towards what might be called architectures of vertigo. Although it probably remains that, while on high, ‘one can feel oneself cut off from the world and yet the owner of a world’ (Barthes, 1983/1964: 250), the panoramic view alone appears to be no longer adequate to meet the demands of urban observatories. These places reflect, and actively produce, a collective desire for vertiginous experience akin to the ‘voluptuous panic’ described by Caillois as ilinx. The proliferation of skywalks and sundry glass platforms signals that, amidst a thriving experience economy, designers have been perfecting new ways of harnessing the user's sensory responses to spatial depth. It can be hypothesised that the viewing subject central to modern scopic regimes is being superseded by a sentient subject whose feelings are put through ever more intense psycho-physiological stimuli. This subject embodies a visual sensibility that is no longer predicated on processes of abstraction and cognition but involves an expanded, and increasingly immersive, field of sensory experience. Glass floors in particular offer a kinaesthetic experience of space that combines the thrill of altitude with the thrill of transparency. As we have seen, these platforms owe their appeal to a horizontal transparency that provokes the exhilarating feeling of hovering over urban space as if in a state of suspension. Elevated glass floors reproduce the groundlessness of the present socio-economic condition and effectively reaffirm it in spatial form. By so doing, they partake in the ‘dreamworlds of neoliberalism’ (Davis and Monk, 2007) that are shaping the emotional landscapes of our cities. Indeed, the spaces described in this paper comply with the predominant subjectivity in which risks and insecurities are governed through environmental stimuli that discipline the activities of the body. They are aligned with a broad strand of architecture that, since the 1990s, has been designed primarily to stimulate ‘affects’ (Spencer, 2016): a ‘post-critical’ design approach that, by disengaging the user from rational and cognitive processes, bolsters up the dominant logic of neoliberalism. In a similar vein, skywalks are designed to embolden a dynamic and enterprising subject to enjoy a seemingly boundless degree of freedom, albeit an artificially staged one. This avid consumer of novel and ever-more thrilling experiences embodies the zeitgeist of our hyper-hedonistic age; that is, what psychoanalysts have identified as a social imperative of the present moment – ‘you must enjoy!’ (Recalcati, 2012: 111). In staging a playground for thrilling experiences, skywalks join a plethora of other games of ilinx that defy the laws of gravity and thereby exorcise the existential fear of falling. Thrill-seeking has been associated with those forms of ‘deep play’ (Ackerman, 1999) that entice people who yearn to escape the restraints of civilized life, particularly in big cities, to attain ecstatic and enlightening moments through adventurous exploits. Skywalks are arguably symptomatic of this wider psychosocial condition. The drive to challenge one's sense of balance is an innate form of human play, and there is something fascinating, as well as frightening, about the possibility of looking down through an elevated glass floor. These platforms hold the potential to reveal hitherto invisible spaces, to provoke an awareness of the vertical expansion of cities, and possibly reflections about its hubristic nature. However, their main social function seems to be an altogether different one. Far from inviting a reflexive aesthetic experience, the viewing galleries that draw visitors to ‘walk on air’ operate as tourist bubbles in which the encounter with the abyss underneath our feet is reduced to a themed spectacle. As the sixth façade becomes the stage of memorable experiences, the hyper-secure conditions in which these experiences take place ensure that our perception of verticality is domesticated and normalised within highly- controlled spaces. This phenomenon offers insights into the increasingly vertiginous environments of contemporary cities and provokes reflections about the social and spatial conditions they reflect and in turn produce."],["Speed-accuracy trade-offs are often considered a confound in speeded choice tasks, but individual differences in strategy have been linked to personality and brain structure. We ask whether strategic adjustments in response caution are reliable, and whether they correlate across tasks and with impulsivity traits. In Study 1, participants performed Eriksen flanker and Stroop tasks in two sessions four weeks apart. We manipulated response caution by emphasising speed or accuracy. We fit the diffusion model for conflict tasks and correlated the change in boundary (accuracy – speed) across session and task. We observed moderate test-retest reliability, and medium to large correlations across tasks. We replicated this between-task correlation in Study 2 using flanker and perceptual decision tasks. We found no consistent correlations with impulsivity. Though moderate reliability poses a challenge for researchers interested in stable traits, consistent correlation between tasks indicates there are meaningful individual differences in the speed-accuracy trade-off. --------------------------------------------------------------------------------","Response control is one of the cornerstones of cognitive psychology, and a topic of interest for both experimental and correlational approaches. Individual differences in tasks such as the Stroop (Stroop, 1935) and Eriksen flanker (Eriksen & Eriksen, 1974), have been linked to executive functioning (Miyake et al., 2000), impulsive behaviour (Sharma, Markon, & Clark, 2014), and a variety of neuropsychological conditions (Chambers, Garavan, & Bellgrove, 2009; Gauggel, Rieger, & Feghoff, 2004; Lansbergen, Kenemans, & van Engeland, 2007; Moeller et al., 2002; Verdejo-Garcia, Perales, & Perez-Garcia, 2007). From an experimental perspective, response control paradigms feature prominently in modelling and neurophysiological studies, where the goal is to characterise the general mechanisms responsible for the control of action (Bompas, Hedge, & Sumner, 2017; Logan, Yamaguchi, Schall, & Palmeri, 2015; Munoz & Everling, 2004). Though the application of these tasks across different disciplines is promising for the development of a coherent understanding of response control, recent work has illustrated that there are challenges to interpreting individual differences because they can arise from different sources, including strategic processes (Boy & Sumner, 2014; Hedge, Powell, Bompas, Vivian-Griffiths, & Sumner, 2018; Miller & Ulrich, 2013). Here, we ask whether strategic processes, often considered to be a confound in cognitive studies, represent a reliable and general component of decision making. Multiple processes underlying individual differences in response control ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In conflict tasks, such as the Stroop, flanker or Simon tasks, we typically subtract reaction times or errors in a baseline condition (congruent or neutral) from a condition in which the stimulus provides conflicting information (incongruent). When used in an individual differences context, the resultant RT or error costs are treated as an index of the individual’s ability to resolve competition between conflicting response options (e.g. Friedman & Miyake, 2004). However, the processes underlying behaviour are multifaceted, and variability in the magnitude of an RT cost or error cost cannot easily be attributed to a single mechanism (Hedge, Powell, Bompas, et al., 2018; Miller & Ulrich, 2013). For example, it has long been theorised that an individual’s reaction time reflects not only their ability to process a stimulus, but also their strategic choice to favour speed or accuracy (Pachella, 1974; Wickelgren, 1977). Indeed, one of the reasons why we use within- subject designs when examining differences between conditions in average RTs is to account for this so called speed accuracy trade-off (SAT). However, individual differences in strategy still contribute to variability in the RT costs. Individuals who favour accuracy over speed produce larger RT costs, as well as smaller error costs (Hedge, Powell, Bompas, et al., 2018; Hedge, Powell, & Sumner, 2018; Wickelgren, 1977). In order to dissociate contributions of strategy and ability in a cognitive task, we require a framework that characterises the contributions of both to behaviour. One such framework is that of sequential sampling models (Brown & Heathcote, 2008; McKoon & Ratcliff, 2013; Ratcliff & McKoon, 2008; Ulrich, Schroter, Leuthold, & Birngruber, 2015). These models assume that choice RT behaviour can be captured by a process of accumulating evidence sampled from the environment, until a boundary or threshold is reached. The rate at which evidence is accumulated represents the efficiency of processing, and the height of the boundary reflects the amount of evidence that an individual waits for before deciding on the response (i.e. their level of response caution, or strategy). By dissociating these processes, and for their ability to simultaneously account for both the RT and accuracy of responses, sequential sampling models could provide a useful window into individual differences in response control (see e.g. Hedge, Powell, Bompas, et al., 2018; White, Curl, & Sloane, 2016). Response caution as a meaningful component of response control ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To many researchers conducting choice RT tasks, strategic processes are considered a confound to the mechanisms of interest. For example, composite measures of RT and accuracy have been proposed with the explicit intention of providing a better control for SATs than traditional subtractions in studies of (e.g.) executive functioning (Draheim, Hicks, & Engle, 2016; Liesefeld & Janczyk, 2018). However, there is evidence that strategic control itself may be a meaningful measure of individual differences, as captured by sequential sampling models. For example, more cautious response strategies are often observed in healthy older adults, relative to younger adults (e.g. Ratcliff, Thapar, & McKoon, 2006; Thapar, Ratcliff, & McKoon, 2003). Studies have also observed changes in response caution in individuals with autistic spectrum disorders, though the studies vary in the direction of the effect (Karalunas et al., 2018; Pirrone, Dickinson, Gomez, Stafford, & Milne, 2017; Powell et al., 2019). Finally, in multiple response control task datasets, we observed a correlation between tasks in model parameters representing response caution, in the absence of correlations in parameters reflecting conflict processing (Hedge, Powell, Bompas, & Sumner, 2019).","Participants in the aforementioned studies are typically instructed to be both fast and accurate, such that the levels of response caution observed are interpreted as the individual’s ‘default’ strategy when given no explicit instruction to favour speed or accuracy. However, individuals are also able to flexibly adjust their strategy if instructed. In SAT paradigms, participants are instructed to prioritise speed in some blocks and accuracy in others, which is (primarily) captured in sequential sampling models by an individual decreasing or increasing their boundary (for a review, see Heitz, 2014). The extent to which individuals are able or willing to adjust their level of caution has also been the subject of individual differences research. Larger decreases in caution under speed emphasis relative to accuracy emphasis in a perceptual decision task were correlated with increased BOLD activation in the striatum and pre-SMA (Forstmann et al., 2008), as well as increased structural connectivity between those regions (Forstmann et al., 2010; though see Boekel et al., 2015 for a non-replication of the connectivity). An association has also been observed between response caution under speed emphasis and self-reported “need for closure” (Evans, Rae, Bushmakin, Rubin, & Brown, 2017). Need for closure is a personality trait theorised to reflect an individual’s preference for certainty over ambiguity (Webster & Kruglanski, 1994), from which Evans et al. predicted that a greater need for closure would lead to more urgent decision making. In line with this prediction, when the data were fit with the linear ballistic accumulator model (Brown & Heathcote, 2008), individuals with a greater need for closure set a lower threshold (Evans et al., 2017). In sum, the research to date suggests that individual differences in response caution and its strategic adjustments have the potential to inform our understanding of cognitive functioning in both healthy individuals and neuropsychological conditions. However, this promise is tempered by several unknowns. First, the psychometric properties of response caution and its strategic adjustments are not well understood. Test-retest reliability is an important consideration for individual differences research, reflecting the degree to which individuals can be consistently ranked on the dimension of interest (i.e. more or less cautious). Recent work has suggested that traditional measures of response control have sub-optimal reliability, and these concerns may also extend to model-based analyses (Hedge, Powell, & Sumner, 2018; Paap & Sawi, 2016). Though a few studies have examined the test-retest reliability of model parameters representing response caution (Enkavi et al., 2019; Lerche & Voss, 2017; Schubert, Frischkorn, Hagemann, & Voss, 2016), to our knowledge none have examined the reliability of strategic adjustments of caution in a SAT paradigm. A second consideration is the extent to which individual differences in strategic control adjustments can be generalised from a single task. Several studies have observed correlations in response caution between tasks when neither speed nor accuracy are preferentially reinforced (Hedge et al., 2019; Lerche & Voss, 2017; Ratcliff, Thompson, & McKoon, 2015), though those that have examined the SAT have used a single perceptual decision task (Evans et al., 2017; Forstmann et al., 2008, 2010). Here, we address this gap in the literature with two experiments. In the first, we apply a model of response control (the diffusion model for conflict tasks; Ulrich et al., 2015) to test-retest data from the flanker and Stroop tasks under different SAT instructions. This allows us to examine whether adjustments in control are reliable over time within the same task, and whether they generalise across tasks within the same cognitive domain. In the second experiment, we examine generalisability more broadly by comparing a response control task (flanker) to a perceptual decision task (random dot motion) commonly used in the decision making literature. To examine potential relationships with related constructs, we also collected data on self-reported impulsivity in both studies, as well as compliance and personality in Study 2. Participants ~~~~~~~~~~~~ Participants were 57 (6 male) undergraduate and postgraduate psychology students. Participants took part either for payment or for course credit. All participants gave their informed written consent prior to participation in accordance with the revised Declarations of Helsinki (2013), and the experiments were approved by the local Ethics Committee. Design and procedure ~~~~~~~~~~~~~~~~~~~~ Participants completed both the Stroop and flanker task in two 90 min sessions taking place approximately 4 weeks apart. A schematic of these tasks, as well as the random dot motion task used in study 2, can be seen in Fig. 1. We administered the UPPS-P, a self- report measure with subscales for different types of impulsivity, (Lynam, Whiteside, Smith, & Cyders, 2006; Whiteside & Lynam, 2001), after participants complete the behavioural tasks. Participants completed the tasks in a dimly lit room from a viewing distance of approximately 60 cm. Stimuli were presented on a 36.5 cm by 27.5 cm display (60 hz, 1280 × 1024). Eriksen flanker task Participants responded to the direction of a centrally presented arrow (left or right) using the z and m keys. On each trial, the centrally presented arrow (1 cm × 1 cm) was flanked above and below by two other symbols separated by 0.75 cm, so that flankers were individually visible. Flanking stimuli were either arrows pointing in the same direction as the central arrow (congruent condition), straight lines (neutral condition), or arrows pointing in the opposite direction to the central arrow (incongruent condition). Trials were separated by an interval of 750 ms. Stroop task Participants responded to the colour of a centrally presented word (Arial, font size 70), which could either be red (z key), blue (x key), green (n key) or yellow (m key). The colours were not purposely matched for luminance. The presented word could be the same as the font colour (congruent condition), one of four non-colour words (lot, ship, cross, advice; neutral condition), or a colour word corresponding to one of the other response options (incongruent). Trials were separated by an interval of 750 ms. For each session and task, participants completed 12 blocks, consisting of 4 each for speed, standard and accuracy instructions. Each block consisted of 144 trials, with 48 each of congruent, neutral and incongruent stimuli (192 trials total per congruency and instruction condition). The order of blocks was randomised, as was the order of trials within blocks. At the beginning of speed-emphasis blocks, participants were instructed “Please try to respond as quickly as possible, without guessing the response”. For accuracy blocks, participants were told “Please ensure that your responses are accurate, without losing too much speed”. For standard instruction blocks, participants were instructed “Please try to be both fast and accurate in your responses”. In speed blocks, if participants responded slower than 500 ms in the flanker or 600 ms in the Stroop, the message “Too slow” appeared on screen for 500 ms. In the accuracy condition, the message “Incorrect” appeared if participants made an error. In all blocks, the message “Too fast” appeared if participants responded faster than 150 ms in the flanker and 200 ms in the Stroop task (typically <1% of trials). Participants received feedback about both their average RT and accuracy after each block in all instruction conditions. Stimuli were presented until response. In the Stroop task, stimuli were presented for a maximum duration of 1950 ms. Trials exceeding this were rare (0.3% and 0.2% of trials in session 1 and 2). Data processing ~~~~~~~~~~~~~~~ Two participants were removed because they did not return for the second session. We excluded participants if there average accuracy across all instruction blocks fell below 60%. This resulted in more participants being retained for the flanker task (N = 47) than the Stroop (N = 43). These participants were retained for the reliability analysis in the flanker task, but were excluded when calculating correlations across tasks. We removed RTs less than 100 ms, and greater than the individual’s median plus three times their median absolute deviation for each condition (Leys, Ley, Klein, Bernard, & Licata, 2013). We did not code trials as incorrect on the basis that they exceeded our deadline for feedback in speed blocks, as changing the relationship between RT and accuracy would confound our modelling. The data are available on the Open ScienceFramework (https://osf.io/zag7c/). For the reliability analysis, we calculated Intraclass Correlation Coefficients using the psych package in R (ICC2; Revelle, 2018; Team & R Development Core Team, 2016). This value is the ratio of between-subject variance in the measure to the total variance, comprising between-subject variance, between-session variance, and error variance. The form of the ICC corresponds to a two-way random effects model for absolute agreement (Shrout & Fleiss, 1979). While the ICC is interpreted as a correlation, ranging from zero to one, different criteria are used to interpret the degree of reliability compared to interclass correlation effect sizes (Pearson’s R and Spearman’s rho). ICCs above 0.8 are typically considered excellent, while 0.6 and 0.4 are categorised as good and moderate reliability (Cicchetti & Sparrow, 1981; Fleiss, 1981; Landis & Koch, 1977). In contrast, Pearson’s R values of 0.5, 0.3 and 0.1 are typically interpreted as large, medium and small effect sizes respectively (Cohen, 1988). The higher convention for the ICC primarily reflects the application rather than the calculation, as high levels of reliability are typically a pre-requisite to correlational work. When calculated on the same data, intra and interclass correlations usually produce similar values (see supplementary material A for different calculations). The diffusion model for conflict tasks ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The diffusion model for conflict tasks (DMC; Ulrich et al., 2015) is a mathematical model of two-choice reaction time behaviour in response conflict tasks. It assumes that the response options are represented by an upper and lower boundary, here corresponding to the correct and incorrect response respectively. The decision processes can be described by a process of accumulating evidence from the stimulus until one or the other boundary is reached (see Fig. 2A). The reaction time on a given trial is determined by the time it takes for a boundary to be reached, plus the duration of sensory and motor (non-decision) processes. For mathematical details, see Ulrich et al. (2015). Boundary separation is the critical parameter for our current goal of measuring individual differences in response caution. Individuals who are more averse to making errors and slow their responses to avoid them should have higher boundary separation values. When participants are instructed to emphasise speed, this is primarily captured by lowering their boundary in that block (for reviews, see Heitz, 2014; Ratcliff, Smith, Brown, & McKoon, 2016). Recent evidence has indicated that the changes under speed emphasis are also reflected in non-decision time to a degree in non-conflict tasks, and sometimes also by a change in drift rate (see e.g. Rae, Heathcote, Donkin, Averell, & Brown, 2014). However, the sensitivity of DMC parameters to the SAT manipulations have not been examined. Given the time-consuming nature of the fitting process for our datasets, and the relatively large number of possible variants, we make the simplifying assumption that only boundary separation varies across SAT instruction conditions here. The DMC assumes that the accumulation process on a trial is a combination of processing from controlled and automatic pathways (De Jong, Liang, & Lauber, 1994; Ridderinkhof, 2002). The controlled route is responsible for processing the task-relevant stimulus feature (e.g. the central arrow in the flanker task), and is represented by drift rate parameter that is constant across conditions. Automatic activation is implemented as a re-scaled gamma function, described by two free parameters (amplitude and time-to-peak) and one fixed parameter (shape). Initially, the automatic activation receives a strong input, reflecting the capture of a prepotent response by (e.g.) the flanking arrows. After it reaches a maximum value (amplitude) at a specified point in time (time-to-peak), the automatic activation decreases, reflecting decay or active suppression (Hommel, 1994; Ulrich et al., 2015). In addition to the aforementioned parameters, which are typically the focus of interest, the model has two parameters describing variability in the starting point of the accumulation processes and variability in the duration of non-decision time respectively. Model fitting ~~~~~~~~~~~~~ For each participant and task, we estimated nine parameters: boundary separation under speed emphasis, boundary separation under standard instructions, boundary separation under accuracy emphasis, the amplitude of automatic activation (A for congruent trials, 0 for neutral trials, -A for incongruent trials), the time to peak automatic activation, mean non-decision time, drift rate of the controlled process, the shape parameter of the starting point distribution, and variability in non-decision time. Variability in starting points and non-decision time are captured by a beta and normal distribution respectively. As with Ulrich et al. (2015), the diffusion constant/within-trial noise (σ) was fixed to 4, and between-trial variability in drift rates was fixed to 0. We fixed the shape parameter of the automatic activation function to 2 for all tasks, following Ulrich et al. (2015). We accuracy-coded our data, such that the upper and lower response boundaries corresponded to the correct and incorrect response options. This allowed us to collapse across different stimulus configurations (e.g. a congruent flanker stimulus irrespective of whether the arrow was pointing left or right), and also to fit the same model to the four-choice Stroop data (Voss, Nagler, & Lerche, 2013). Though this level of abstraction is not ideal, it relates RT and accuracy to capture the strategic processes that we are interested in, and there is currently no extension of the model for four choice tasks. After excluding outlier RTs as described above, correct and incorrect RTs from congruent, neutral and incongruent conditions in each instruction block were separately binned into quantiles. We fit the DMC to experimental data using the similar approach to that used by the Diffusion Model Analysis Toolbox (DMAT; Vandekerckhove & Tuerlinckx, 2008). Correct RTs were binned using five quantiles (0.1, 0.3, 0.5, 0.7, 0.9). Incorrect RTs were binned using five quantiles if the total number of errors in that condition >10, otherwise they were not used. The application of five quantiles produced six bins per RT distribution (corresponding to: 0–10%, 10–30%, 30–50%, 50–70%, 70–90%, 90–100%). Therefore, participants’ fits would be based on either 6 or 12 data points per instruction and congruency condition, resulting in between 54 and 108 data points in total. These quantiles are commonly used when fitting sequential sampling models (c.f. Ratcliff & Tuerlinckx, 2002). We calculated the deviance (-2 log-likelihood) between observed and simulated quantiles, which was minimised with a Nelder-Mead simplex (Nelder & Mead, 1965) implemented in the fminbnd function in Matlab. We constrained the search such that all free parameters were positive, and the shape of the starting point distribution was greater than one. Initially, we fit each participant’s data using 5000 parameter sets that were randomly generated from a uniform distribution (see supplementary material B for maximum and minimum values). This was done to explore plausible starting points for our fitting algorithm. We then took the 15 best parameter sets resulting from this initial search, and submitted each of those to the simplex algorithm, in which we simulated 10,000 trials per condition at each iteration. The simplex was re-initialised 3 times to avoid local minima. After the process was completed, we took the single best fitting parameter set for each individual. This process took approximately 6 days per dataset, and was performed on Cardiff University Brain Research Imaging Centre’s (CUBRIC) high performance computer cluster. Behavioural data Reaction times and error rates for both tasks are shown in Fig. 3. To verify that the average performance reflected the expected manipulations, we conducted separate 3(instruction) × 3(congruency) repeated-measures ANOVAs on RTs and error rates for each session and task. In all cases we observed significant main effects for both congruency and instruction (all p < .001; see Supplementary Material C for full ANOVA results). Error rates and RTs increased for incongruent relative to congruent stimuli. Further, error rates increased and RTs decreased when participants were instructed to prioritise speed over accuracy. Reaction times and error rates for both tasks are shown in Fig. 7. As in Experiment 1, we verified that the average performance reflected the expected manipulations by conducting separate repeated-measures ANOVAs on RTs and error rate in each task. In all cases we observed significant main effects for both congruency/coherence and instruction (all p < .001; see Supplementary Material C for full ANOVA results). Error rates and RTs increased for incongruent (flanker) and low-coherence (dot-motion) stimuli relative to congruent and high-coherence stimuli. Further, error rates increased and RTs decreased when participants were instructed to prioritise speed over accuracy. Model parameters Descriptive statistics for the best fitting parameters can be seen in Table 1. Graphical displays of the model fits can be seen in Supplementary Material D. The values are numerically similar to previous fits we have observed in a non-SAT context (Hedge et al., 2019), with the Stroop showing a relatively slower time-to- peak and a higher value for the shape of the starting distribution (corresponding to less variability in start points). This reflects that the manual Stroop task does not tend to show fast errors (see Supplementary Material D). The model was successful in capturing the relative speed and accuracy of participants, though the data show more fast errors under speed-emphasis than the model. In both tasks and sessions, boundary separation was decreased under speed relative to neutral and accuracy emphasis, indicating that the parameter captured the SAT manipulation in the expected way. Descriptive statistics and graphical displays of the fits for the best fitting parameters can be seen in Supplementary material F. In both tasks, as expected, boundary separation was decreased under speed relative to neutral and accuracy emphasis. Values for the flanker task are similar to those observed in Study 1. Values for the DDM parameters fit to the dot-motion task were within typically observed ranges (Donkin, Brown, Heathcote, & Wagenmakers, 2011; Matzke & Wagenmakers, 2009). Within-task reliability of strategic adjustments of response caution ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We quantified strategic adjustments in response caution by taking the difference in boundary separation under speed-emphasis relative to accuracy emphasis for each individual. Strategic adjustments showed moderate reliability across both tasks (flanker ICC = 0.5, Stroop ICC = 0.40; see Fig. 4). To put these model parameter correlations in context of the behaviour from which they’re derived, we also examined the reliability of adjustments to RT and accuracy rates in isolation (averaged across congruency conditions). This led to a similar range of values, with ICCs from 0.46 to 0.68 (see Supplementary Material A for a full report). In other words, the reliability of the model parameters were not systematically higher or lower than the behavioural measures. Note that boundary separation is theorised to reflect a balance between RT and accuracy, and so would not have the same interpretation as either behavioural measure in isolation. See Table 2 for the reliability of all the DMC parameters. We also draw attention to the 95% confidence intervals (CI) given in this table. While a CI cannot be interpreted as an indicator of the precision of an estimate (c.f. Morey, Hoekstra, Rouder, Lee, & Wagenmakers, 2016), under similar assumptions as those used to calculate a p-value, it can be interpreted to contain the values we cannot reject based on our statistical test (Morey, Hoekstra, Rouder, & Wagenmakers, 2016). In other words, just as we reject the null hypothesis (ICC = 0) based on the interval for adjustments in the flanker task (95% CI: 0.26–0.69), we also reject values that correspond to excellent or clinically required levels of reliability (ICC > 0.7). Between-task correlation of strategic adjustments of response caution ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our second key question is whether strategic adjustments in response caution correlate between tasks, which we assessed using Spearman’s rho. We observed moderate to large correlations between tasks in each session (session 1 rho = 0.56, p < .001; session 2 rho = 0.40, p = .038; Fig. 5). Again, these were similar to the correlations observed in the adjustments of RTs and accuracy in isolation, which ranged from 0.31 to 60 (see Supplementary material A). We present the correlations between strategic adjustments of response caution and UPPS-P subscales in supplementary material E. Briefly, we see no consistent correlation across our datasets. As in study 1, we observed a large correlation in strategic adjustments in response caution (rho = 0.50, p < .001; Fig. 8). Thus, behavioural variability captured by parameters representing response caution do share commonality across tasks from different cognitive domains. As in Study 1, this was numerically similar to the correlation observed in the adjustments in RT and accuracy in isolation (both rho = 0.40). For correlations between self-report measures and strategic adjustments in response caution, see Supplementary material E. Correlations with self- report were generally small and inconsistent across the tasks. Reliability of other model parameters ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As we are the first to examine the test re-test reliability of the DMC parameters (not just the strategic adjustments to boundary), we present these in Table 2, along with the between task correlations. The reliability of the three main non-conflict parameters (drift rates, boundary separation in each instruction condition and non-decision time) ranges from moderate to good, and are similar to those observed for the standard drift- diffusion model (Lerche & Voss, 2017). For conflict processing parameters, the amplitude of the automatic activation showed moderate reliability in both tasks, whereas the time- to-peak was relatively poor. The between-task correlations are generally similar to that which we observed with these tasks in our other work (that did not include a SAT manipulation; Hedge et al., 2019). Interim discussion ~~~~~~~~~~~~~~~~~~ We discuss the implications of these values in more detail in the general discussion. First, we follow up on the observation that strategic adjustments in response caution correlate between the flanker and Stroop tasks in both sessions. This result is promising, and suggests that we can generalise our interpretation of individual differences in strategic control beyond a single task. However, it raises the question of whether it generalises outside of response control tasks, or if strategic adjustments may differ depending on the broad cognitive domain. This is particularly relevant as previous papers that have examined individual differences in strategic adjustments have used a perceptual decision task, rather than conflict tasks (Evans et al., 2017; Forstmann et al., 2008, 2010). To assess whether individual differences in response caution also generalise across cognitive domains, we conducted a second study in which participants performed the flanker task along with a random dot motion discrimination task under a SAT manipulation. Participants ~~~~~~~~~~~~ Participants were 81 (6 male) undergraduate and postgraduate psychology students. Participants took part either for payment or for course credit. Six participants that participated in Study 1 also participated in Study 2. The studies took place a year apart. All participants gave their informed written consent prior to participation in accordance with the revised Declarations of Helsinki (2013), and the experiments were approved by the local Ethics Committee. Design and procedure ~~~~~~~~~~~~~~~~~~~~ Participants completed both the flanker task and a dot motion discrimination task based on (Pote et al., 2016). The participants also completed a number of questionnaires for the purpose of exploratory analyses: the UPPS-P impulsivity scale, the NEO-FFI personality inventory (McCrae & Costa, 2004), the Gudjonsson Compliance Scale (Gudjonsson, 1989), and a Situational Compliance Scale (Gudjonsson, Sigurdsson, Einarsson, & Einarsson, 2008). The flanker task appeared as described above. Participants performed 12 blocks of 144 trials in total. Twelve participants did not complete all blocks within the allotted time, so data were only available for 11 (11 participants) or 10 (1 participant) blocks. In the dot motion task, each frame consisted of 50 white dots (5x5 pixels in size) displayed within an oval patch (14.7 cm high × 23.7 cm wide) in the centre of a grey screen (60 hz, 1680 × 1050). On each frame, either 30% (high coherence) or 15% (low coherence) of the dots were chosen as signal dots, which moved in a consistent direction (left or right) by 29 pixels. The lifetime of the dots was 3 frames. Non-signal dots reappeared in a random position on each frame. The stimulus was displayed for a maximum of 2000 ms, with a 500 ms ISI. Participants were asked to determine the direction of the coherent motion. Each block consisted of 120 trials, 60 of each coherence level. Participants performed 12 blocks in total, except for 5 participants who completed 11 blocks, and 1 participant who completed 10. Feedback relating to speed, accuracy or neutral blocks was given as described in Study 1. For the dot motion task, participants were informed that their responses were too slow in speed blocks if their RT exceeded 700 ms. Participants were informed that they were too fast in all blocks if their responses were shorter than 250 ms. Data processing ~~~~~~~~~~~~~~~ The same inclusion criteria and RT cut-offs described in study 1 were applied. After exclusions, 73 participants were retained for the analysis of between-task correlations. The drift-diffusion model ~~~~~~~~~~~~~~~~~~~~~~~~~ As the dot-motion task is not a conflict task, and the DMC extends the standard drift- diffusion model with conflict-specific parameters, we opted to fit the dot motion data with the standard drift-diffusion model (DDM; Ratcliff, 1978; Ratcliff & Rouder, 1998). Though the DMC is an extension of the DDM, it is possible that they capture variance associated with response caution in slightly different ways due to different parameterisations. However, we are interested in the conclusions that researchers would draw if they had used the model that was most appropriate for the task they had used. Critically for our purposes, strategic adjustments in response caution are conceptually captured by a change in boundary separation in both models. The primary difference between the DMC and the DDM is that, whereas accumulation rates in the DMC reflect a composite of controlled and automatic processes, accumulation rates in the DDM are determined by a single drift rate parameter. This means that the underlying accumulation rate in a given trial is constant over time, albeit subject to noise as in the DMC. Conditions with varying difficulty are captured by differences in average drift rates (see Fig. 6). Model fitting ~~~~~~~~~~~~~ We fit the DMC to the flanker data using the same process described for study 1. For the DDM, we used the Diffusion Model Analysis Toolbox (DMAT; Vandekerckhove & Tuerlinckx, 2008). Similar to our approach with the DMC, observed RT quantiles from correct and incorrect are compared to data simulated from the model, and the deviance minimised using a Nelder-Mead simplex (Nelder & Mead, 1965). As with the flanker task, for simplicity we assumed that only boundary separation varied across instruction condition. For each participant and task, we estimated eight parameters: boundary separation under speed emphasis, boundary separation under standard instructions, boundary separation under accuracy emphasis, drift rate for high coherence trials, drift rate for low coherence trials, mean non-decision time, starting point variability and non-decision variability. Between-trial variability in drift rates was fixed to 0.1, starting point bias was fixed to boundary separation/2, and within-trial noise was fixed to 0.1. Note that DMAT assumes uniform distributions for starting point and non-decision variability, whereas our implementation of the DMC uses a beta and normal distribution respectively (following Ulrich et al., 2015).","The aim of the current work was to examine whether individual differences in strategic adjustments of response caution are a reliable and generalisable dimension of response control. The answer to both questions is yes, though this is caveated by the magnitude of the effects that we observe. In Experiment 1, we observed moderate test-retest reliability in the change in response caution in both the flanker and Stroop tasks, as represented by the change in boundary separation between accuracy-emphasis and speed emphasis instructions. It is not trivial that these strategic adjustments show non-zero reliability, though the magnitude is below the levels typically considered good or excellent for conducting individual differences research (Cicchetti & Sparrow, 1981; Fleiss, 1981; Landis & Koch, 1977). The implication of this is that researchers interested in examining relationships between strategic adjustments in response caution and personality or brain structure will likely require large sample sizes to detect relationships, if they exist. With regards to generalisability, we show medium to large correlations in response caution adjustments across tasks conducted in the same experimental session. We observed this between conflict tasks (study 1), and between a conflict and a perceptual decision task (study 2). We focus our discussion on the interpretation of strategic control adjustments, and practical recommendations for researchers interested in response caution. Meaningful individual differences in default caution and its strategic adjustment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There is increasing evidence that there are meaningful individual differences in response caution (Evans et al., 2017; Forstmann et al., 2010; Hedge et al., 2019; Karalunas et al., 2018; Pirrone et al., 2017; Powell et al., 2019; Ratcliff et al., 2006a). Recently, we applied the DMC to three response control datasets, comprising the flanker and Simon tasks, flanker and Stroop tasks, and two variants of the Simon task. Our aim was to examine whether the model could uncover hidden correlations between mechanisms of conflict processing that are obscured in traditional measures (Hedge, Powell, Bompas, et al., 2018). Though we observed no correlation in the conflict parameters (amplitude and time- to-peak), we consistently observed correlations in boundary (see also Lerche & Voss, 2017; Ratcliff et al., 2015). This finding is mirrored in our results here, with boundary separation consistently showing correlation between tasks. The novel contribution of this work is that we also see correlation in the strategic adjustment in response caution, captured by the change in boundary separation between different SAT instructions. We manipulated participant’s levels of response caution through verbal instruction, which is the same method used by the previous studies that have examined individual differences in response caution adjustments (Evans et al., 2017; Forstmann et al., 2008, 2010). There are numerous alternative methods for eliciting a SAT (for a review, see Heitz, 2014). These include the use of payoff structures, in which participants receive different rewards and penalisations based on accuracy and/or RT (e.g. Fitts, 1966; Swensson & Edwards, 1971); and the use of response deadlines, where participants are informed that they must respond within certain time limits (e.g. Pachella & Pew, 1968). Heitz notes that verbal instructions are popular because they are easily understood by participants, and produce large effects with relatively few trials. However, just as the interpretation of the common instruction in choice RT tasks to be both fast and accurate is subjective, so too is the instruction to favour speed. Our reliabilities and correlations suggest that participants interpret these instructions somewhat consistently, though we do not know what the basis is for the criteria they set. In part, this is what we seek to understand by examining correlations with personality constructs such as impulsivity. To our knowledge there has not been a systematic examination of the consequences of the choice of SAT manipulation for individual differences relationships (though some have been compared experimentally, e.g. Dambacher & Hübner, 2013). It would benefit future research in this area to elucidate whether the method makes a difference. Do strategic control adjustments go beyond boundary separation? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A wealth of literature exists for the speed-accuracy trade-off, spanning both behavioural and neurophysiological approaches (for a review, see Heitz, 2014). In the context of the sequential sampling models, faster RTs and lower accuracy under speed emphasis are primarily attributed to reduced boundary separation: a relative decrease in the amount of evidence required to initiate a response (Ratcliff et al., 2016). However, performance under speed emphasis has also been captured by additional reductions in non-decision time, as well as sometimes lower drift rates (Rae et al., 2014; Starns, Ratcliff, & McKoon, 2012; Zhang & Rowe, 2014). Further, it has been argued that strategic adjustments can be captured by time-varying decision processes, such as urgency signals or collapsing boundaries (Cisek, Puskas, & El-Murr, 2009; Ditterich, 2006a, 2006b; Drugowitsch, Moreno- Bote, Churchland, Shadlen, & Pouget, 2012; though see Hawkins, Forstmann, Wagenmakers, Ratcliff, & Brown, 2015). Here, we fit a relatively simple model that only allowed boundary separation to vary across instruction conditions. Therefore it is possible that our fits absorbed variance in behaviour that might be captured by other parameters in a more complex model. Note that in the introduction, we highlighted the difficulty in translating assumptions from within-subject contexts to the study of individual differences. The SAT paradigm is also an approach that has largely been developed in within-subject experimental contexts, and the average best fitting model may not be appropriate for every individual. For example, we could ask whether every individual shows a decrease in boundary, non-decision time, and/or information processing parameters (c.f. Haaf & Rouder, 2018). Our results here provide a starting point for further examination; that we observe some reliability and cross-task correlation in response caution here suggests that there is reliable variance in the behaviour to be captured. Previous literature on the reliability of response caution ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To our knowledge, we are the first to examine the reliability of parameters of the diffusion model for conflict tasks. Previous work has examined the test-retest reliability of the standard drift-diffusion model (Lerche & Voss, 2017; Schubert et al., 2016), including applications to conflict tasks (Enkavi et al., 2019), though not in a SAT paradigm. Nevertheless, we can contrast our estimates of the reliability of boundary separation under standard instructions with theirs. Lerche and Voss (2017) reported one week reliability for a lexical decision task, a recognition memory task, and an associative priming task. They observed correlations of approximately r = 0.8 for boundary separation in all tasks (see maximum likelihood estimates in their Fig. 2). Schubert et al. (2016) report eight month reliabilities for three tasks, including a two- and four- choice variant of a visual choice RT task, a Sternberg memory scanning task, and a Posner letter matching task. Correlations for boundary separation between sessions ranged from r = 0.2 to r = 0.6 (see their Table A2). Recently, Enkavi et al. (2019) applied the hierarchical drift diffusion model (Wiecki, Sofer, & Frank, 2013) to reliability data from 15 choice RT tasks, including a three choice Stroop task. The average time between sessions was approximately 16 weeks. The reliability of boundary separation in the Stroop task was 0.29, which was slightly below the median reliability for all the tasks (0.31; see their HDDM values in Fig. 5). Taking these previous studies together, our results fall within the range of reliabilities previously observed, but the range is broad. It would be premature to suggest that there are systematic differences between tasks in the consistency of response caution that they elicit, though we note that it was relatively low for the Stroop task in both our Study 1 and in Enkavi et al. (2019) data. Model choice and model complexity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To examine the reliability of strategic adjustments in response caution, we applied the drift-diffusion model (Ratcliff, 1978), and an extended diffusion model for conflict tasks (Ulrich et al., 2015). Though the drift diffusion model is widely applied in SAT studies (e.g. Mulder et al., 2010; Ratcliff, 1985; Zhang & Rowe, 2014), there are alternative models for both conflict (Hubner, Steinhauser, & Lehle, 2010; White, Ratcliff, & Starns, 2011) and non-conflict tasks (Brown & Heathcote, 2008; Usher & McClelland, 2001). Several empirical and theoretical reviews have considered the relationship between different models, and it has been noted that there is often a high degree of mimicry between them, such that researchers would often reach the same conclusion irrespective of the model chosen (Bogacz, Brown, Moehlis, Holmes, & Cohen, 2006; Donkin et al., 2011; Ratcliff & Smith, 2004; White, Servant, & Logan, 2017). Nevertheless, we briefly consider the potential impact of this choice. Both Forstmann et al. (2010) and Evans et al. (2017) examined individual differences in response caution adjustments using the Linear Ballistic Accumulator model (Brown & Heathcote, 2008). Whereas in the DDM a single drift process represents the difference in evidence between to alternatives, the LBA consists of separate accumulators for each response alternative and a single threshold. Starting points for the accumulators in the LBA are drawn from a uniform distribution, and response caution is captured by the difference between the edge of the start point distribution and the height of the threshold. Forstmann et al. (2010) noted that applying the drift- diffusion model to their data did not produce the correlation between white matter strength and caution adjustments seen with the LBA (see their Supplementary Online Material). They suggested that this may be because the diffusion model captured the SAT manipulation in both non-decision time and drift rates, in addition to boundary separation. An imperfect mapping between the response caution parameters has also been noted when fitting one model to data generated from the other (Donkin et al., 2011). Given this discrepancy, researchers may wish to check the robustness of conclusions drawn from one model to another. Though we made the simplistic assumption that the SAT manipulation was specifically captured by changes in boundary separation in our fits, we added complexity by including parameters representing inter-trial variability in non-decision time and the starting point of the accumulation process. Including these variability parameters often produces better fits to empirical data at the sample average level (Ratcliff & Tuerlinckx, 2002), but they may also lead to poorer recovery of individual differences in the main parameters of interest, particularly with fewer trial numbers (Lerche & Voss, 2016; van Ravenzwaaij, Donkin, & Vandekerckhove, 2017). We reran some of our analyses without including the variability parameters, and it produced almost identical estimates for the reliability of strategic adjustments (see Supplementary Material G). This may be in part because we had a large number of trials, and a model that was quite well constrained across multiple conditions. Where researchers have smaller trial numbers, they may wish to implement a simpler model, or check that their conclusions are not specific to a particular parameter choice. Strategic adjustments and personality traits ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Willingness or reluctance to commit errors while attempting speeded responses has been linked to the concept of impulsivity (Kagan, 1966), and recently correlated with need for closure (Evans et al., 2017) and brain structure (Forstmann et al., 2010). However, we see little evidence for a correlation with any self-report impulsivity dimension in our data (Supplementary Material E; see also Dickman & Meyer, 1988). We also tested correlations with self-report compliance and the big five personality traits (neuroticism, extraversion, openness, agreeableness, conscientiousness; Digman, 1990; McCrae & Costa, 2004; McCrae & John, 1992). There were no consistent relationships. The absence of a correlation with impulsivity measures is particularly notable here, given the conceptual overlap between impulsivity and a lowered boundary. For example, Metin et al. (2013) examined whether differences in RT and accuracy in children with attention- deficit/hyperactivity disorder relative to healthy controls were best captured by “inefficient” or “impulsive” information processing in the context of the drift diffusion model. These corresponded to drift rate and boundary separation respectively. Despite the common terminology, our findings mirror a trend in the impulsivity literature to observe little to no correlation between behavioural and self-report measures (Sharma et al., 2014). It remains a possibility that there are non-zero correlations that we did not have sufficient power to detect. As we discuss in the next section, given that the reliability of strategic adjustments is suboptimal, we should expect correlations with other variables to be small. Practical considerations for future research ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The consistent between-task correlation in strategic adjustment indicates that the extent to which an individual adjusts their behaviour is not entirely task or domain specific. A practical consideration for researchers interested in response caution and its strategic adjustments, and are not specifically interested in a particular cognitive domain (e.g. response conflict) is that fitting the DMC to our response conflict tasks was substantially more demanding on time and/or computational resources than fitting the DDM to a perceptual decision making task. Note that this is not specific to the DMC (White et al., 2017), but rather reflects more complex models that do not have analytical solutions that allow faster estimation. Until faster methods can be realised (e.g. Mestdagh, Verdonck, Meers, Loossens, & Tuerlinckx, 2018), it may be more tractable to use non- conflict tasks to which the DDM or linear ballistic accumulator (Brown & Heathcote, 2008) can be applied. The reliabilities of strategic adjustments of response caution that we observe fall in the range typically interpreted as “moderate” (Cicchetti & Sparrow, 1981; Fleiss, 1981; Landis & Koch, 1977). We have recently discussed how reliabilities in this region are potentially problematic for examining individual differences in cognitive tasks (Hedge, Powell, & Sumner, 2018). The ICC reflects the relative contribution of between- subject variance (individual differences) and measurement error to variance in the variable of interest. In order to examine whether adjustments in response caution are related to trait measures (e.g. personality), we desire variance in our behavioural measure to also reflect individual differences that are stable over time. When measures are noisy, correlations with external variables will be weaker and require larger samples to detect. To put the ICCs we observe in context, they are similar to or exceed those we observed for several commonly used measures of response control and processing (e.g. flanker RT cost: 0.50, stop-signal reaction time: 0.43, Navon global precedence: 0; Hedge, Powell, & Sumner, 2018). Notably in the case of those traditional measures, we did not observe correlation between tasks in our previous study, despite it being commonly assumed that they share common mechanisms (see also Rey-Mermet, Gade, & Oberauer, 2017). Here, with strategic adjustments of response caution, we do consistently observe a correlation between tasks. Nevertheless, poor reliability corresponds to a reduced ability to detect correlation using those measures that must be compensated for (e.g. by increasing statistical power). It is possible that future work would benefit from developments in model-based analyses, by integrating individual differences measures of interest in to the parameter estimation (Evans et al., 2017; Turner et al., 2013; Wiecki et al., 2013). Here, we fit the models to each task and individual independently. In contrast, hierarchical models describe both the sample and individual simultaneously, as well as allow for regressors to be used to inform parameter estimation. For example, Evans et al. (2017) compared three different models when examining the relationship between response caution and need for closure. They first fit a hierarchical linear ballistic accumulator model in which parameters were determined by the behavioural data alone. The second model did not allow for individual differences in response caution, assigning everyone the same value, though differing between speed- and accuracy-emphasis. In the third model, rather than estimating response caution from the behavioural data, it was determined by a function that linked the parameter values to participants’ questionnaire values. Unsurprisingly, the first (unconstrained) model provided the best fit to the data. However, the third model outperformed the second, suggesting that there are common individual differences in the personality questionnaire and behavioural responses. In a second experiment by Evans et al. this improvement when comparing models contrasted against non-significant correlations between parameter estimates (and RTs) fit independently and subsequently correlated with the questionnaire values. Such joint modelling techniques may provide more powerful tests where appropriate.","The extent to which an individual prioritises accuracy or speed in choice RT tasks is commonly discussed but has less often been the focus of interest than individual differences in cognitive abilities. Here, we provide evidence that questions about individual differences in caution and its strategic adjustment are at least somewhat viable. On a given occasion, individuals show consistency in the extent to which they strategically adjust their levels of response caution across different tasks. Across time points, individuals show non-zero, but sub-optimal, levels of reliability in strategic adjustments. Though these levels of reliability raise power concerns for future research, we believe that our results and previous literature are evidence that there is value in pursuing such questions."],["This study investigated language function associated with behavior problems, focusing on pragmatics. Scores on the Children's Communication Checklist Second Edition (CCC-2) in a group of 40 adolescents (12-15 years) identified with externalizing behavior problems (BP) in childhood was compared to the CCC-2 scores in a typically developing comparison group (n=37). Behavioral, emotional and language problems were assessed by the Strengths and Difficulties Questionnaire (SDQ) and 4 language items, when the children in the BP group were 7-9 years (T1). They were then assessed with the SDQ and the CCC-2 when they were 12-15 years (T2). The BP group obtained poorer scores on 9/10 subscales on the CCC-2, and 70% showed language impairments in the clinical range. Language, emotional and peer problems at T1 were strongly correlated with pragmatic language impairments in adolescence. The findings indicate that assessment of language, especially pragmatics, is vital for follow-up and treatment of behavioral problems in children and adolescents. © 2014 The Authors. --------------------------------------------------------------------------------","Language is an important tool for social interaction as well as a means to control one's own and other's emotions and behaviors. Children who are able to use language to regulate their emotions and behave in a socially appropriate way are more likely to develop good peer relations and form new friendships (Im-Bolter & Cohen, 2007). Three intersecting areas of language – form, content, and use – are all essential ingredients for communication, and impairments within any of these areas may cause problems. The form and content components characterize language structure, whereas the use component characterizes pragmatics (Bloom & Lahey, 1978; Spanoudis, Natsopoulos, & Panayiotou, 2007). A growing body of research points to an association between behavioral and language development, and several studies have reported a substantial degree of overlap between language impairments and behavioral problems (Cross, 2011; Hill & Coufal, 2005; Mackie & Law, 2010). Children with language impairments frequently experience behavioral problems, and conversely, many children with behavioral problems show language impairments (Gallagher, 1999; Hartas, 2012; Ketelaars, Cuperus, Jansonius, & Verhoeven, 2010). Although this relationship is well documented in the literature, it seems to be less recognized in practice, and there is good evidence that language impairments are substantially underreported in children with psychiatric diagnoses (Cohen, Farnia, & Im- Bolter, 2013; Im-Bolter & Cohen, 2007; Law & Garret, 2004). Hill and Coufal (2005) claim that although students with behavioral disorders experience language impairments, their problems in this domain may be left as an “invisible” or “marginal” handicap unless systematic assessment is carried out. Symptoms that may be caused by problems in understanding or producing language may be perceived by adults as non-compliance, social withdrawal, or inattentiveness (Cohen, 2001). Children with Attention- Deficit/Hyperactivity Disorder (ADHD) and Autism Spectrum Disorders (ASD) commonly present co-existing problems related to language. Previous research has shown that in a large population derived sample of 5672 children aged 7–9 years, almost 60% of the children identified with symptoms of ADHD (n = 290) also fulfilled the criteria for language impairments compared to 5.7% of the typically developing control group (Helland, Posserud, Helland, Heimann, & Lundervold, 2012). Furthermore, in a clinical sample of 6–15 year old children with Asperger syndrome and children with ADHD, 90.5% and 82.1%, respectively, presented with clinically significant language impairments (Helland, Biringer, Helland, Heiman, 2012). In their review of studies of language skills in children identified with emotional and behavioral disorders, Benner, Nelson, & Epstein (2002) found that 71% experienced clinically significant language impairments. In their study of children aged 7–14 years referred to psychiatric services, Cohen, Menna and colleagues (1998) reported that children identified with language impairments showed more immature abilities with respect to resolving interpersonal conflicts than children without language impairments. Furthermore, parents often perceived these children as problematic and hard to manage compared to typically developing peers (Law & Garret, 2004). Language impairments refer to a broad spectrum of difficulties including limited vocabulary, expressive deficits, phonological deficits, comprehension deficits, and pragmatic language deficits. All these problems have been reported in studies of children with behavioral disorders (Gallagher, 1999). According to Tannock and Schachar (1996), pragmatic difficulties are the most frequently reported language problem. Pragmatics refers to the appropriate use and interpretation of language in different social contexts (Bishop, 1997). Children with pragmatic language impairments may speak fluently and well-articulated, but they have problems adhering to the needs of the conversational partner; they may make incorrect inferences, give conversational responses that are socially inappropriate or tangential, and interpret language in an over literal manner (Fujiki & Brinton, 2009; Poletti, 2011). Pragmatic language deficits are clinically relevant because they may have detrimental effects on the development of successful peer relations and negatively impact the child's quality of life (Gibson, Adams, Lockton, & Green, 2013). Gilmour, Hill, Place, and Skuse (2004) found that two-thirds of their sample of children with conduct disorder had pragmatic language impairments. They also identified pragmatic language deficits in about two-thirds of a sample of children with antisocial behavior, and suggested that these deficits may underlie the antisocial behavior. In line with this, Donno, Parker, Gilmour, and Skuse (2010) argue that pragmatic language deficits should be considered a possible contributory factor to behavioral problems in primary school children. According to Leonard, Milich, and Lorch (2011), pragmatic skills provide a unique contribution in the estimate of the children's social skills above and beyond the contribution of both hyperactivity and inattention. Recently, Mackie and Law (2010) reported clinical significant language impairments (pragmatic-, structural- and word decoding difficulties) in 91% of referred children. These findings strongly indicate that language impairments of some kind very often accompany behavioral disorders. Several explanations have been offered to account for the relationship between language- and behavioral problems (see Hartas, 2012); (1) language difficulties may lead to frustration and anger resulting in increased problems with social behavior and fewer opportunities to interact with peers, (2) behavioral problems, like inattention and hyperactivity, may contribute to language and literacy problems, (3) both language and behavioral difficulties co-exist and reciprocally influence each other, (4) the two conditions share an underlying deficit that may explain the association between language and behavioral problems (Hartas, 2012). All these explanations refer to the strong correlation between the two domains of problems. This is supported by the tendency that a wide range of problems seems to cluster within the same individual, and the high rates of comorbidity in child psychiatry (Posserud & Lundervold, 2013). The Early Symptomatic Syndromes Eliciting Neurodevelopmental Clinical Examinations model (ESSENCE) has been put forward to describe the more overarching dysfunction generally encountered within child psychiatry (Gillberg, 2010), and genetic studies also support the existence of larger, less specific set-ups of genes that together form a heightened vulnerability to a wide range of problems from intellectual disability to anxiety and more subtle motor problems (Cross-Disorder Group of the Psychiatric Genomics Consortium, 2013; Lichtenstein, Carlström, Gillberg, & Anckarsäter, 2010). The ESSENCE model was conceptualized also because developmental problems seem to change over time, depending on external factors, where a child may present with language problems in early childhood and then develop more overt ADHD symptoms in early school age. Inspired by this model, the current study aim at studying language difficulties within a broader group defined as having behavioral problems. The majority of studies investigating language impairments have been based on pre-and primary school children. There is mounting evidence that many of these children have enduring language problems that may negatively impact their long-term psychosocial and academic development (Cohen et al., 2013; Conti-Ramsden & Botting, 2008; Yew & Kearney, 2013). As children reach adolescence, demands on language competence increase and language skills become even more crucial for establishing and maintaining social relationships. Inadequate communication may cause misunderstandings, increase conflicts and deteriorate the quality of friendships, leaving children and adolescents at risk of stress, loneliness, and mental health problems (Durkin & Conti- Ramsden, 2010; Leonard et al., 2011). Adolescents with language impairments may see themselves as less socially accepted than their typically developing peers, and may also be perceived as withdrawn and unsociable by their peers as well as by their teachers (Im- Bolter, Cohen, & Farnia, 2013). In a recent study, Cohen and colleagues (2013) reported that clinic-referred youths aged 12–18 years were significantly impaired relative to a comparison group on measures of structural as well as higher order language function. Furthermore, their language impairments were associated with parent ratings of severity of externalizing psychopathology. The present study aimed to investigate language function in a group of adolescents with behavioral problems (BP). Based on previous research we expected to find more language related problems in the BP group than in the general population in childhood (part A) and in an age matched control group in adolescence (part B). Due to the importance of pragmatic language ability in adolescence, we finally asked if a measure of this ability when the adolescent was 12–15 years old could be predicted from parent reports of behavioral- and language problems approximately five years earlier (part C).","A group of children with behavioral problems (BP) were recruited among participants in the third phase of the first wave of the Bergen Child Study (BCS). The BCS is a longitudinal total population study of child mental health that started in 2002 with a screening questionnaire for all children attending 2nd to 4th grade in any school in the Bergen area (n = 9430). The response rate was high, with 97% of the teachers and 70% of the parents completing the BCS screening questionnaire including the Strengths and Difficulties Questionnaire (SDQ, Goodman, 1999), and four questions related to language function. From the first wave of the BCS, children were selected to a second and third phase according to screen status. The third phase (n = 329) consisted of a clinical assessment with the Wechsler's Intelligence Test for Children – third edition (WISC-III) (Wechsler, 2003) and the K-SADS-PL (Kaufmann et al., 1997; see Lundervold, Posserud, Ullebo, Sorensen, & Gillberg, 2011 for more details). The present study included children identified in this third phase (T1) with high symptom levels of an externalizing disorder according to the K-SADS-PL (defined by one or more definite symptoms of ADHD, Oppositional Deficit Disorder (ODD), or Conduct Disorder (CD)). These children were invited to a follow-up study when they were 12–15 years old (T2). The follow-up assessment included the K-SADS-PL and parent reports on the Children's Communication Checklist Second Edition (CCC-2; Bishop, 2003; Norwegian adaptation: Helland & Møllerhaug, 2006) and the SDQ. One child was excluded from the study because she was younger than the other children (11.11 years at T2), two children were excluded because of intellectual disability and seven children because the CCC-2 did not pass the consistency check (invalid). Thus at T2, the BP group consisted of 40 children (32 males, 8 females) in the age range 12–15 years (M = 13.47, SD = 0.82), see Fig. 1. The study was approved by the Data Inspectorate and the western Regional Committee for Medical and Health Research Ethics. Comparison group at T2 At follow-up (T2), comparisons regarding language abilities were made between the BP group and a comparison group (CO) of typically developing children. Thirty- seven children (18 males, 19 females) with a mean age of 13.54 years (SD = 1.14) who had participated in the Norwegian standardization of the CCC-2 (Bishop, 2003; Norwegian adaptation: Helland & Møllerhaug, 2006) served as CO group. The reason for using this group as comparison rather than adolescents from the same background study as the BP group was that the CCC-2 was only administered in the BP substudy of the BCS. The CO group did not have any problems regarding language or communication as reported by their parents, nor did they have any known learning disabilities or special education needs. The CO group had a more equal distribution of males and females compared to the BP group in which the majority were males. However, no significant differences were found between males and females on the General Communication Composite of the CCC-2; t(35) = 1.60, p = .11, which supports the use of the CO group for comparison. Strengths and Difficulties Questionnaire (SDQ): BP group at T1 and T2 Parents completed the SDQ as part of the BCS-questionnaire, when the children were 7–9 years old (T1) and at follow-up (T2) when they were 12–15 years. The SDQ is a brief screening questionnaire for behavioral and emotional problems designed for children aged 4–16 years. The questionnaire has been extensively validated in various countries, and reported internal consistency values (Chronbach's alphas) for the various scales have a mean α = .70 (Muris, Meesters, & van den Berg, 2003). Twenty-five items divided into five subscales (five items in each) are measuring emotional problems, conduct problems, hyperactivity–inattention problems, peer problems and prosocial behavior. A total difficulties score is computed by combining the first four subscales scores. Each item is scored on a three-point scale (0 = not true, 1 = somewhat true, and 2 = certainly true). On the first four scales a high score indicates problems, while a low score indicates problems on the last subscale (prosocial). The subscale scores are ranging from 0 to 10 and the total difficulty score is ranging from 0 to 40. Separate SDQ versions are available for parents, teachers and children, and in the present study data from the parent version are presented. Language composite (LC): BP group at T1 A set of four items relating to different aspects of language (phonology, expressive language, receptive language, and pragmatics) was included in the BCS- questionnaire and was completed by the parents. The language items were as follows: (1) cannot pronounce certain words or sounds; (2) cannot elaborate, explain or express himself or herself; (3) has difficulties understanding things that are being said; and (4) has difficulties having a conversation with others. These items were scored on a three-point scale, and a language composite score with a possible range of 0–8 was included in the present study. Children's Communication Checklist Second Edition (CCC-2): BP group and CO group at T2 The CCC-2 (Bishop, 2003; Norwegian adaptation: Helland & Møllerhaug, 2006) is a checklist designed to distinguish children with communication impairments from typically developing children and to identify pragmatic as well as structural language impairments in children aged 4–16 years. The checklist is to be completed by an adult who has regular contact with the child, in the present study it was completed by parents. A total of 70 items are grouped into 10 subscales (see Table 2) with seven items in each. The separate subscales assess speech, syntax, semantics, inappropriate initiation, stereotyped language, use of context, non- verbal communication, social interaction, and interests. The questionnaire is scored on a 4-point scale, indicating the frequency of the communicative behavior described, with a high raw score indicating poorer performance. An automatic scoring program that comes with the CCC-2 (Bishop, 2003) converts raw scores into scaled scores with a mean of 10 and a SD of 3. The first four scales measures structural aspects of language, the next four scales measure pragmatic aspects and the two last scales measure non-linguistic behavior. By summing the scaled scores of the first eight subscales, an overall measure of language abilities, the General Communication Composite (GCC), is derived. This composite is effective at discriminating children with communication impairments from typically developing children. In accordance with previous findings using the CCC-2 in a Norwegian sample, cut-off at or below 64-scaled scores on the GCC was selected for identifying children with language impairments (Helland, Biringer, Helland, & Heimann, 2009). An additional composite, the Social Interaction Deviance Composite (SIDC), is also computed to identify children with pragmatic impairments disproportionate to their structural language abilities. A negative SIDC implies that the child experiences difficulties with social interaction that are disproportionate to his/her general communication abilities. However, according to the manual (Bishop, 2003), this composite should only be interpreted if the GCC is below cut-off; an exception is scores of −15 or less, as such an extreme result is of clinical significance even with GCC within normal limits. Although not included as part of the CCC-2, a general pragmatic composite (PC) has been calculated in several studies (Bignell & Cain, 2007; Geurts & Embrechts, 2008; Helland, Helland, & Heimann, 2012). This is done by summing the scaled scores of the scales measuring coherence, inappropriate initiation, stereotyped language, use of context and, nonverbal communication (scales D–H). The PC was computed and used for analyses in part B of the present study. In the British standardization sample Bishop (2003) reported internal consistency values between .66 and .80 and inter- rater reliability between parents and teachers ranging from .16 to .79 for the CCC-2. The Norwegian version also presents with good internal consistency with alpha ranging from .73 to .89 and inter-rater reliability ranging from .44 to .76 (Helland et al., 2009).","All statistical analyses were run using SPSS, version 21. Part A: One-sample t-tests were conducted on the SDQ and the LC to evaluate whether the means of the BP group were significantly different from the means of the total population-based sample in the BCS, from which the BP group was derived when they were 7–9 years (T1). Part B: Students independent-samples t-tests were used to analyze differences between the BP and CO groups at T2 on a general measure of language abilities (GCC) and the subscales of the CCC-2. Bonferroni corrections were conducted due to multiple comparisons (alpha level of .005), and Cohen's d was computed to evaluate effect sizes. According to general guidelines, d's of 0.20, 0.50 and 0.80 should be interpreted as small, medium, and large, respectively. Part C: Longitudinal predictions for the BP group from 7–9 years (T1) to 12–15 years (T2) were investigated by running correlation analyses (Pearson product moment correlation) between SDQ and LC scores at T1 and the PC at T2, and a backward multiple regression analysis to evaluate whether SDQ and LC scores (T1) predicted pragmatic language abilities as measured by the PC at follow-up four years later (T2). See Table 5 for the sequence of variables included in the analysis. Additionally, correlation analyses were conducted to compare SDQ subscale scores at T1 and T2, and LC scores at T1 and T2. Part A: language abilities at 7–9 years (T1) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The BP group differed significantly (was more impaired) from the total sample on the total difficulties score and all subscale scores of the SDQ as well as on the LC at T1 (Table 1). Part B: language abilities at 12–15 years (T2) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the BP group altogether 70% (28 out of 40 children) obtained a GCC score in the clinical range. For the comparison group the corresponding number was 10.8% (four out of 37 children). On the SIDC, 13 out of the 28 children in the BP group identified with communication impairments obtained a score indicating pragmatic impairments that were disproportionate to structural language abilities. Additionally, one child scored in the clinical range (below −15) although his GCC was in the normal range. In the CO group three out of the four children identified with communication impairments showed disproportionate pragmatic impairments. A comparison between the two groups on the GCC revealed that the scores for the children in the BP group were significantly lower than the scores for the CO group (t(75) = 6.46, p < .001). As shown in Table 2, significant differences were found on all but one subscale (i.e., scale A measuring speech). The effect sizes (Cohen's d) were moderate for speech (0.6) and high (ranging from 0.8 to 1.9) for all subsequent subscales. Part C: predictions from 7–9 years (T1) to 12–15 years (T2) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ At follow-up (T2), SDQ and LC data were available for 29 out of the 40 children in the BP group. Bivariate correlation analyses between SDQ data from T1 and T2 showed highly significant correlations, with Pearson r ranging from .56 (conduct) to .69 (hyperactivity) (Table 3). Bivariate correlation analyses between LC data from T1 and T2 were statistically significant, Pearson r = .49, p < .01. Bivariate correlation analyses showed that the SDQ scales, measuring emotional problems, peer problems, total difficulties, and the LC at T1 correlated significantly with the PC reported at T2, see Table 4. To examine whether SDQ and LC (T1) could predict pragmatic language abilities in adolescence (T2), a backward multiple regression analysis included the SDQ subscales (emotional problems, conduct problems, hyperactivity problems, peer problems, pro-social behavior) and the LC score as predictors and the PC as the criterion variable. A significant regression equation was found F(2, 37) = 10,56, p < 001 (model 5), with an R2 of .363. Thus, the predictors accounted for 36% of the variance in pragmatic abilities at T2. The SDQ scale measuring peer problems and the LC both had significant effects on the PC; with beta values of −0.34 and −0.39, respectively (Table 5). Tolerance tests and VIF (variable inflation factor) did not indicate multicollinearity. The analyses were repeated for the prediction of the overall GCC score, showing that only the LC score at T1 independently predicted GCC score at T2 (beta value of −0.49; p = .001).","The present study followed a group of children with behavioral problems, investigating their language function in adolescence, and whether language function and mental health at 7–9 years predicted later pragmatic language impairment. The cross-sectional analyses of part A and B, exploring whether the children with behavioral problems at age 7–9 and 12–15 years could be differentiated from a typically developing comparison group regarding their language profiles, showed that children with BP scored higher on parent reported language problems in the general population at 7–9 years of age, and also obtained poorer scores on 9 out of 10 subscales of the parent form of CCC-2 at 12–15 years. In the longitudinal analyses of part C, peer problems and language problems reported in the BP group at 7–9 years were shown to be significant predictors of pragmatic language abilities in adolescence, whereas the overall measure of language abilities (GCC) was only predicted by the parent report on the four questions regarding language problems. As predicted, the group with behavioral problems scored significantly poorer than the comparison group on the GCC. Language impairments were far more common, with the vast majority (70%) scoring in the clinical range. Our findings confirm results reported by Benner and colleagues (2002) in their review of language impairments in children with emotional and behavioral disorders, as well as a report of language impairments (structural-, pragmatic-, decoding problems) in the majority of a sample of children with behavior causing concern at school (Mackie & Law, 2010). In our sample, the distribution of children with BP primarily displaying problems related to pragmatics (35%) and those displaying mainly structural language problems (35%) were quite equal. These findings are comparable to those of Donno and colleagues (2010) who identified pragmatic language deficits in 42% of their sample of disruptive children. Furthermore, they are in line with those reported by Gilmour and colleagues (2004) and Mackie and Law (2010), who found that two-thirds of their samples showed significant pragmatic language deficits. Our findings were somewhat more modest than in the last studies, which may be due to the fact that our BP group was identified as part of a population-based study. Still, the BP group differed significantly (more impaired) from the CO group on all the pragmatic subscales of the CCC-2, emphasizing that pragmatics is an area of language that is highly vulnerable in children with behavioral problems. These findings are in line with the results from a former study where children diagnosed with ADHD were found to differ significantly from typically developing children on the pragmatic subscales (Helland, Helland, et al., 2012). The CCC-2 profile revealed that the group of children with behavioral problems was impaired relative to the comparison group regarding all aspects of language except on the scale measuring speech. Conflicting results have been reported regarding speech; Geurts and Embrechts (2008) and Helland, Helland, et al. (2012) reported unimpaired speech in studies of children with ADHD, whereas Helland, Biringer and colleagues (2012), in a study of children with Asperger syndrome (AS) and children with ADHD, found that these clinical groups showed impaired speech relative to controls. A possible explanation for the finding of unimpaired speech in the present study may be that when children reach adolescence they may have outgrown their speech problems, while difficulties related to semantics, pragmatics and social relations are more likely to persist. Alternatively, initial speech problems, although no longer present, may have contributed to pragmatic impairments becoming more pronounced in adolescence as social situations grow more demanding and complex. Although the children with behavioral problems were inseparable from the comparison group on the scale measuring speech, they demonstrated significant impairments on the other scales measuring language structure, indicating that their language difficulties were not restricted to pragmatics but did affect other aspects of language as well. The latter finding aligns well with the recent results reported for clinic – referred adolescents by Cohen and colleagues (2013) as well as with our former studies of language impairments in children with ADHD, AS and typically developing controls (Helland, Biringer, et al., 2012; Helland, Helland, et al., 2012). The observed differences between the two groups on the CCC-2 scales measuring interests and social relations may indicate that children with behavioral problems experience considerable problems as far as friendship and peer acceptance are concerned, thus putting them at risk for increased levels of behavioral problems and mental health problems more generally. Language problems at T1 were only reported by parents on four general language items targeting a wide and unspecific range of language problems that can affect children. Still, the composite score on these four items predicted language problems as assessed by the CCC-2 five years later, even after controlling for psychopathology in the group of children with BP. Such a relationship over such a long time-span is almost surprising, underscoring the need for taking parental concerns of their child's difficulties seriously and to follow up their worries with further assessment. The significant prediction of pragmatic language abilities in adolescence from peer problems and language problems reported by parents in childhood, underline the close association between communicative abilities and social functioning The problems reported in childhood appear not to be transient; rather they seem to persist into adolescence, negatively affecting the development of successful social relationships, which may again lead to escalating behavioral problems. Limitations ~~~~~~~~~~~ Some limitations should be considered when evaluating the findings of the present study. The diversity of diagnostic subgroups within the BP group, children with ADHD, and children with ODD/CD symptoms, might be considered a limitation as small sample sizes prevent us from reporting separately for the diagnostic groups. On the other hand, the strong association in this heterogenic group shows that language should be an area of great concern irrespective of the nature of the behavioral problems. As only children in the BP substudy of the BCS performed the CCC-2, the comparison group at T2 was chosen from the sample of the Norwegian CCC-2 normative sample, where the gender distribution was different from the BP group. This could potentially have overstated the differences between the BP and the CO group; however, there were no gender differences in the CCC-2 scores in the CO group. The fact that language evaluations were solely based on parental reports is another limitation, and firm conclusions about the predictive value of early language problems awaits future large-scaled longitudinal studies. Finally, although we have stated that peer problems in this study predict pragmatic language deficits, the reverse could also be true. Pragmatic language deficits most definitely cause peer problems, and so the relationship between social difficulties and pragmatic language deficits is likely to be bidirectional.","Bearing in mind that language is commonly not an area receiving great attention in children with behavioral problems, our findings have some important clinical implications. Firstly, language assessment should be an integral part of the assessment procedure when children and adolescents are referred to mental health services with behavioral problems. Secondly, as pragmatic language deficits contribute to difficulties resolving interpersonal conflicts with others, pragmatics abilities should be an area of special concern taken into consideration when interventions and therapy plans for adolescents with behavioral problems are developed. Furthermore, it is possible that the lack of overt speech problems in children with behavioral problems may mask severe communicative problems. As most therapies are strongly language-based, verbal input should be modified to match the language level of the adolescents, the use of non-literal language should be monitored, and the clinician should be aware that what may appear as non-compliance may in fact result from problems understanding."],["A difficulty for reports of subliminal priming is demonstrating that participants who actually perceived the prime are not driving the priming effects. There are two conventional methods for testing this. One is to test whether a direct measure of stimulus perception is not significantly above chance on a group level. The other is to use regression to test if an indirect measure of stimulus processing is significantly above zero when the direct measure is at chance. Here we simulated samples in which we assumed that only participants who perceived the primes were primed by it. Conventional analyses applied to these samples had a very large error rate of falsely supporting subliminal priming. Calculating a Bayes factor for the samples very seldom falsely supported subliminal priming. We conclude that conventional tests are not reliable diagnostics of subliminal priming. Instead, we recommend that experimenters calculate a Bayes factor when investigating subliminal priming. --------------------------------------------------------------------------------","Exposure to a perceivable stimulus may influence or “prime” a response to another stimulus, even when the priming stimulus is just noticeable. More controversial are claims of priming induced by imperceptible or “subliminal” (i.e., below the threshold of perception) stimuli. Although many studies claim to have demonstrated subliminal priming, the phenomenon is still debated (Newell & Shanks, 2014). The debate continues because it is difficult to prove the subliminal part of the claim–that the prime stimulus was not perceived, not even slightly, by any observer–and thereby to rule out an alternative explanation of the observed priming: that it is solely attributable to the responses of observers who just barely perceived the prime. The main strategy to find support for subliminal priming has been to try to demonstrate a dissociation between a direct measure of prime stimulus perception and an indirect measure of prime stimulus processing (Reingold & Merikle, 1988). Two statistical methods are conventionally used to find statistical support for dissociation: a double t-test and a regression method. Here we simulate these methods and a less often used method based on Bayesian statistics. The simulations suggest that the latter method, but not the former methods, is suitable for evaluating whether experimental data support subliminal priming. Experiments on subliminal priming ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In a typical experiment examining subliminal priming, a sequence of stimuli is shown in each trial. First, one of two priming stimuli is briefly flashed, followed by a masking stimulus (to allow the priming stimulus to be processed but not perceived), and then a target stimulus is shown. Observers have two tasks in the experiment, one direct task regarding the prime stimulus (the direct measure) and one indirect task regarding the target stimulus (the indirect measure). In a typical experiment, the tasks are performed in separate blocks, beginning with the indirect task. In the direct task, observers decide which of the two possible priming stimuli was presented. For example, in a study by Kiefer, Sim, and Wentura (2015), the priming stimulus was an emotionally positive or negative stimulus and the direct task was to decide whether the prime stimulus was positive or negative. In analyzing data from such tasks, one of the two stimuli may be arbitrarily designated the “signal” and the other the “non-signal,” and the four possible stimulus answer combinations may be classified as hits (responding “signal” to the “signal stimulus”), misses, false alarms (responding “signal” to “non-signal” stimulus), and correct rejections. According to signal detection theory, the observer’s sensitivity, d′, to differences between the two stimuli is d′ = Φ−1(ph) − Φ−1(pf), where ph is the proportion of hits, pf the proportion of false alarms, and Φ−1 is the inverse standard normal cumulative distribution function (Macmillan & Creelman, 2005). Response bias, that is, the tendency to choose one over the other stimulus, may be quantified as the response criterion, i.e., c = −½[Φ−1(ph) + Φ−1(pf)]. For an unbiased observer, c = 0, because the proportions of hits and correct rejections are equal. In typical experiments, the direct measure of prime stimulus perception is d′. The indirect task varies depending on the type of priming the experimenter is examining. In a typical experiment, the indirect measure of prime stimulus processing is a congruency effect on reaction time for responses to the indirect task. Kiefer et al. (2015), also used a positive or negative stimulus as the target stimulus and the direct task was to decide whether the target stimulus was positive or negative. When the priming stimulus was incongruent with the target stimulus, the reaction time in the task was slower than if the priming and target stimuli were congruent. This type of congruency effect is the indirect measure of prime stimulus processing in typical experiments. Statistical analysis of subliminal priming ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The most common analytical method, here called the double t-test method, is to apply two t-tests, one for each of the two measures, and decide for or against subliminal priming based on the pattern of the p-values (see Table 1), typically using α = 0.05. Specifically, the double t-test method declares subliminal priming if (a) mean performance in the direct measure does not differ statistically significantly from chance (p > α) and (b) the mean effect in the indirect measure does differ statistically significantly from zero (p < α). The double t-test method is open to criticism on statistical grounds, because deciding whether the prime stimulus is subliminal is based on an unjustified interpretation of how the obtained p-values relate to the tested null hypothesis, H0: “True mean d′ = 0.” Specifically, the p-value is the conditional probability of obtaining the data or more extreme data given H0 [P(D|H0)], and therefore say nothing about the probability of H0 [P(H0)] (see e.g., Dienes (2014) and Gallistel (2009) for a good discussion of null hypothesis testing in the context of non-significant results). The double t-test method has also been criticized as the method often lacks the power needed to support that observers were subliminal (e.g., Finkbeiner & Coltheart, 2014; Gallistel, 2009; Macmillan, 1986; Rouder, Morey, Speckman, & Pratte, 2007; Wiens, 2006). Proponents of the double t-test method are also of course aware of this, but would argue that the strategy still can, and indeed does, do a good job of identifying subliminal priming. Therefore, the double t-test remains the most popular analytical method in research into subliminal priming (e.g., González-García, Tudela, & Ruz, 2015; Huang, Tan, Soon, & Hsieh, 2014; Jusyte & Schönenberg, 2014; Kido & Makioka, 2015; Kiefer et al., 2015; Lin & Murray, 2015; Marcos Malmierca, 2015; Norman, Heywood, & Kentridge, 2015; Ocampo, 2015; Ocampo, Al-Janabi, & Finkbeiner, 2015; Schoeberl, Fuchs, Theeuwes, & Ansorge, 2015; Wildegger, Myers, Humphreys, & Nobre, 2015). The double t-test method tests for subliminal priming at a group level. Observers, however, differ in their thresholds (e.g., Albrecht & Mattler, 2012; Dagenbach, Carr, & Wilhelmsen, 1989; Greenwald, Klinger, & Schuh, 1995; Haase & Fisk, 2015; Sand, 2016). To take individual differences in thresholds (and thus perception, given a specific prime stimulus intensity) into account, another conventional analysis is regression analysis (e.g., Jusyte & Schönenberg, 2014; Ocampo, 2015; Schoeberl et al., 2015; Xiao & Yamauchi, 2014). In regression analysis, the direct measure is used as the regressor and the indirect measure is the outcome variable. Specifically, the regression method declares subliminal priming if the intercept is statistically significantly above zero. Because these two conventional methods (double t-test and regression) remain popular today, we tested their robustness through simulations in which we assumed no dissociation between the direct and indirect measures. One analytical strategy not yet in widespread use is to calculate a Bayes factor to test whether or not a prime stimulus is subliminal at a group level. Calculating a Bayes factor, B, is the Bayesian equivalent of a null hypothesis significance test. The improvement of Bayesian statistics over null-hypothesis testing here is that B can lend support to H0 (which the p-value cannot) and H0 is what experimenters in this field want to support (Dienes, 2015). To calculate B, an a priori model of H1 (“True mean d′ slightly above 0”) needs to be specified. (See Dienes, 2015, for a discussion of how H1 can be specified with regard to subliminal priming.) Once H1 is specified, B can be calculated and it is the ratio of the likelihood of the observed data based on the two hypotheses, that is, P(D|H1)/P(D|H0). B > 1 thus indicates that the data support H1 over H0, while B < 1 indicates that the data support H0 over H1, and a B of approximately 1 suggests that the experiment was not sensitive. Although B is continuous, it has been suggested that B > 3 can be considered “substantial” evidence for H1, whereas a Bayes factor of 1/3 can be considered “substantial” evidence for H0 (Jeffreys, 1998). This criteria of B > 3 roughly corresponds to an α level of 0.05. We follow these criteria here when evaluating B. By calculating B, then, experimenters can test whether the data support subliminal priming (see Table 1). Curiously, a mixed strategy has also been used in which the direct measure is tested using B, whereas a p-value is still used to test the indirect measure (e.g., Norman et al., 2015). This mixed method is probably used because the main advantage of B over the p-value is when H0 is the hypothesis of interest. Here we compare how these two newer methods, which we call the double B method and the mixed method, perform when applied to the same samples as used in the conventional tests of subliminal priming.","We used simulations to evaluate the various scenarios described below. A simulation approach is motivated by the mathematical complexity of the scenarios, involving discrete and asymmetric distributions of observed d′ values (cf Miller, 1996). Using the statistical software R (see Electronic Supplementary Material for all scripts), we simulated scenarios in which we assumed no dissociation between the direct and indirect measures. The purpose of the simulations was to evaluate the extent to which the different analytical methods would incorrectly provide support for subliminal priming although all priming was actually supraliminal, that is, only participants who truly perceived the prime stimulus were primed by it. Following Miller (2000), we simulated an experiment in which d′ was used as a measure of an observer’s ability to detect a prime stimulus (direct measure of perception, d′d) and of the priming effect of the same stimulus on the observer’s performance on a related task (indirect measure of processing, d′i). Simulation 4 differs in that we used a congruency effect on reaction time as the indirect measure. (For an example experiment using d′ as both the direct and indirect measures, see Greenwald et al., 1995.) In the simulations, true sensitivity will be measured for each observer using a number, nd and ni, of two alternative forced-choice trials. This means that observed sensitivity will be a result of true sensitivity and random error in the measurement. We will denote the observed sensitivity d′ to distinguish it from true sensitivity, d′. Assumptions ~~~~~~~~~~~ The simulations were based on the following assumptions: (a) Performance on both the direct and indirect measures follows from the assumptions of the Gaussian equal-variance model of signal detection theory (Macmillan & Creelman, 2005). (b) The observed d′-values are composed of a true component, d′ (true sensitivity), and an error component that is independent of the true sensitivity and has an expected value of zero. (c) An observer’s true sensitivity remained constant during all trials of the experiment.1 (d) Observers’ true response criteria in both measures were unbiased, c = 0.2 (e) In all simulations, except the last, d′d = d′i. This implies absence of subliminal priming, because priming (d′i > 0) only happened for observers who perceived the prime stimulus (d′d > 0). (f) All N observers provided one observed d′d, based on nd trials used to measure direct prime stimulus perception, and one observed d′i, based on ni trials used to measure indirect prime stimulus processing. Simulation procedure ~~~~~~~~~~~~~~~~~~~~ The simulation procedure was as follows. (1) Observers d′d and d′i were sampled randomly from a specified bivariate distribution of d′-values (d′d = d′i in all except Simulation 5). The observers’ response criteria in the two measures were set to zero (i.e., cd = ci = 0). (2) A true probability correct score was calculated for each observer from the true d′-values. For an unbiased observer (c = 0), the true proportion correct is Φ(d′/2) where Φ is the standard normal cumulative distribution function (Macmillan & Creelman, 2005). Thus, pd = Φ(d′d/2) and pi = Φ(d′i/2). (3) The observed proportion of hits in the direct measure (phd) was simulated for each observer by randomly drawing a number from the binomial distribution based on pd and nd/2 and dividing that number by nd/2. The observed proportion of false alarms in the direct measure (pfd) was similarly simulated based on 1 − pd and nd/2. We simulated proportion hits and false alarms in the indirect measure in the same way. Observed proportions were determined from true scores plus the random error inherent in the set of binomial trials. (4) We calculated d′d as Φ−1(phd) − Φ−1(pfd) and d′i as Φ−1(phi) − Φ−1(pfi). (5) We subjected the set of d′d and d′i to the analyses of interest and stored the exact p-value and corresponding B. (6) We repeated steps one to five 10,000 times and used the stored p-values and B to calculate the proportions of each analytical outcome. To calculate B, we needed to specify both H0 and H1. H0 is easy to specify in the present application, namely true d′d = 0 and true d′i = 0 for the direct and indirect tests, respectively. For H1, the a priori probability of different possible population values needs to be specified. We wanted to specify H1 for the effect “observers (at a mean level) were marginally able to perceive the prime stimulus but lower sensitivities are more likely than higher sensitivities”. Based on previous experimental conditions using very difficult to perceive but not subliminal stimuli (e.g., Atas, Vermeiren, & Cleeremans, 2013; Haase & Fisk, 2015; Sand, 2016; Schoeberl et al., 2015), we chose to represent this using a normal distribution with a mean of 0 and a standard deviation of 0.2 but with the probability of values below 0 set to 0 (i.e., a half-normal distribution). That is, H1 was “true mean d′ slightly above zero, with larger d′ having smaller probabilities and d′ > 0.4 being improbable”; i.e., using the notation from Dienes (2015) we used a BH(0,0.2) to specify H1.","In the first simulation, we assumed a discrete bivariate distribution of true sensitivities in which 60% of observers could not perceive the prime stimulus and were not primed by it (d′d = d′i = 0) and 40% of observers could marginally perceive the prime stimulus and were primed by it (d′d = d′i = 0.2). This unrealistic but simple distribution was chosen to illustrate how random error will influence observed scores. In this simulation, 40 observers were drawn from this distribution (i.e., 24 observers d′d = d′i = 0 and 16 observers d′d = d′i = 0.2). In this simulation, true mean d′s across observers was thus 0.08 (0.2 × 0.4). Observed d′d was based on nd = 100 trials and d′i was based on ni = 200 trials, as previous studies have typically used more trials in the indirect than the direct task (e.g., Jusyte & Schönenberg, 2014; Kiefer et al., 2015; Lin & Murray, 2015; Ocampo, 2015; Schoeberl et al., 2015). Fig. 1 illustrates one sample drawn from this distribution. Although d′d = d′i assumed only two values (0 or 0.2), random error arising from the limited number of trials introduced variation in observed scores (Fig. 1A and B). In this sample, mean d′d was 0.04 (SE = 0.04) and mean d′i was 0.07 (SE = 0.03). We applied the double t-test method, mixed method and double B method to these data to compare the three methods. The double t-test resulted in p-values of 0.36 for d′d and 0.02 for d′i (illustrated in Fig. 1C as 95% confidence intervals). The double t-test method would therefore falsely indicate that subliminal priming occurred in this sample. Calculating B for the direct measure (with BH(0,0.2)) resulted in a B of 0.51. This B does not suggest that the sample was subliminal but rather that the data are inconclusive (encouraging experimenters to collect more data; a valid procedure using Bayesian statistics, Dienes, 2011). Thus, neither the mixed method nor double B method would support subliminal priming for this data-set. This simulation example (Fig. 1) followed the expected pattern of the results of 10,000 simulations (Table 2). This result suggests that the conventional double t-test method is unreliable but that calculating B for the direct measure is more reliable. Because the double B method reaches much the same conclusion as does the mixed method, we will omit reporting the results of the mixed method in the following simulations.","The discrete bivariate distribution of true sensitivities used in the previous simulation is of course not realistic. A more realistic distribution would include variation in true sensitivities between observers. True d′-values below zero are implausible in most applications, however, because a negative true d′ would imply that the non-target stimulus on average evoked stronger evidence for the target stimulus than did the target stimulus itself, or vice versa. Observed d′-values slightly below zero may of course occur due to random error. Therefore, in Simulation 2, a half-normal distribution with parameters μ = 0 and σ = 0.1 (corresponding to the standard deviation of the full normal)3 was used for d′d = d′i. In this simulation, true mean d′s across observers were 0.08. Otherwise, this simulation was identical to Simulation 1. We applied the double t-test method and double B method again in this more realistic simulation. We also tested a follow-up analysis (subgroup analysis) sometimes used to strengthen the support for subliminal priming and the regression method mentioned above. Double t-test and double B method. ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The result of applying the double t-test and double B method is summarized in Table 3. Again, the double t-test method would often encourage experimenters to wrongly conclude that subliminal priming had occurred. In comparison, the double B method would often encourage experimenters to collect more data or to conclude that the priming was supraliminal. As Bayesian statistics is relatively insensitive to stopping rules, simply collecting more data when B is inconclusive (1/3 > B > 3) is a valid procedure (Dienes, 2011). To test the robustness of this approach we assumed that in the samples B was inconclusive with regards to d′d (7.81%), researchers would test more observers, reanalyze their data after each observer and stop when a decisive B was reached. Using such a stopping rule would increase false support of H0 (B < 1/3) to 8.29%. As such, even when using such an “extreme” stopping rule, the double B approach would outperform the double t-test approach. Subgroup analysis ~~~~~~~~~~~~~~~~~ As experimenters are aware of the problems associated with claiming that a prime stimulus is subliminal based on a non-significant p-value, experimenters sometimes perform an additional analysis to support their conclusion. We call this analysis subgroup analysis and it entails dividing participants based on their observed d′d and basing the t-tests only on those who perform close to chance (e.g., Huang, Lu, & Dosher, 2012; Lin & Murray, 2015; Marcos Malmierca, 2015; Palmer & Mattler, 2013; Züst et al., 2015). The idea behind this is straight forward: If the indirect effect is apparent in the subgroup that performs worse, then the effect should be subliminal. How “performance close to chance” is operationalized varies between studies (e.g., based on median split of d′d or on the binomial distribution). To test this analysis in this simulation, for each sample, we excluded participants performing above the 97.5th percentile of expected d′d given d′d and 100 trials (subsampling based on other cut-offs led to similar results). We then applied the double t-test and double B only on the remaining subsample. In different samples between 0 and 25% of observers were excluded in this manner (mean 6%, standard deviation 4%-points). The result is not encouraging for the subgroup analysis. The double t-test method suggested subliminal priming in 54% of the subsamples and the double B method suggested subliminal priming in 17% of the subsamples. Both methods therefore perform worse when applied to a subgroup of participants than to the full sample.4 The reason for this is regression to the mean. Although d′d and d′i are perfectly correlated in this simulation, d′d and d′i are not (see Fig. 1). As such, an observer with a d′d of 0.2 may, due to random error, produce a d′d close to chance. However, the same observer is unlikely to be as unlucky in d′i, and is therefore more likely to perform close to his or her true sensitivity (d′i = 0.2). As such, subgroup analysis based on d′d trivially leads to a non- significant direct measure (or B < 1/3) without guaranteeing only subliminal effects in the indirect measure. For a Bayesian approach to sorting participants as subliminal or not, see Morey, Rouder, and Speckman (2009). Regression method ~~~~~~~~~~~~~~~~~ Another method of testing subliminal priming is via regression analysis. This method is sometimes used as the main analysis or as an additional analysis after the double t-test method (e.g., Jusyte & Schönenberg, 2014; Ocampo, 2015; Schoeberl et al., 2015; Xiao & Yamauchi, 2014). In this analysis, d′d is used as the regressor and d′i as the outcome variable. The idea is that if priming is independent of awareness then (a) the intercept should be above zero and/or (b) there should be no correlation between d′d and d′i. We are sympathetic to this method as the idea is to take individual variability in prime perception (d′d) into account. However, using simulations, Dosher (1998) and Miller (2000) have already demonstrated that regression analysis is not robust in this application. In short, random error in the predictor variable leads to underestimation of the correlation and, in turn, an overestimation of the intercept. One criticism of Miller’s simulations was that they were based on very large sample sizes irrelevant to most actual experiments (Klauer & Greenwald, 2000). Therefore, we decided to replicate Miller’s results by applying regression analysis in this simulation, where N = 40. Recall that in this simulation we assumed a perfect correlation between d′d and d′i and an intercept of zero. Regression analysis using null-hypothesis significance testing applied on the observed scores, however, resulted in 66% statistically significant intercepts and 62% non- significant correlations. We therefore verify that Miller’s conclusions are also valid for sample sizes more common in experiments. Having previously seen that the Bayesian approach performed better than the double t-test, we also decided to test the intercept by calculating B (with BH(0,0.2)). Here, however, when H1 was the hypothesis of interest, the Bayesian approach did not outperform conventional tests but resulted in 61% of intercepts supporting subliminal priming (B > 3). The reason that the regression approach is not reliable is error in the regressor. Therefore, to take error in both variables into account, we also tested the application of an orthogonal (Deming) regression (Ripley & Thompson, 1987). We then tested the resulting intercept using a p-value and B. Although the orthogonal regression controlled the error rate somewhat, 46% (p < α) and 42% (B > 3) of intercepts still supported subliminal priming. In conclusion, we agree with Dosher and Miller that regression analysis is not a robust way of testing subliminal priming.","A strength, but also a difficulty, of simulation is that it forces the researchers to explicitly specify all the assumptions on which a simulation is based. Of importance here is the underlying distribution of d′d and d′i. This distribution varies between experiments and, in any case, the true underlying distributions are unknown. Furthermore, experiments vary in their sample size (N) and number of trials (n) used. Therefore, we decided to vary these parameters based on previous studies (e.g., González-García et al., 2015; Huang et al., 2014; Jusyte & Schönenberg, 2014; Kido & Makioka, 2015; Kiefer et al., 2015; Lin & Murray, 2015; Marcos Malmierca, 2015; Norman et al., 2015; Ocampo, 2015; Ocampo et al., 2015; Schoeberl et al., 2015; Wildegger et al., 2015) systematically in two simulations. Simulation 3A ~~~~~~~~~~~~~ In Simulation 3A, we varied the underlying distribution of d′ and the sample size (N). We varied the underlying distribution by using various σ’s of the half-normal distributions, as shown in Fig. 2A. We then sampled from these distributions using various values of N. In all other respects, this simulation was identical to Simulation 2. For each combination of distribution and sample size, we ran 10,000 simulations. For each simulation, we applied the double t-test and double B method and noted the probability of false support for subliminal priming (p > 0.05 for d′d and p < 0.05 for d′i or B < 1/3 for d′d and B > 3 for d′i). Fig. 2B and C summarizes the results of this simulation. The y-axis shows the probability of false support for subliminal priming. For the double t-test method, there is a curvilinear relationship between the two parameters and false support: As the σ or N increases, the power of the indirect measures increases and thus also the proportion of support for subliminal priming. As the parameters increase beyond a certain point, however, the power of the direct measure also increases and the proportion of support for subliminal priming decreases. As for the double B method (Fig. 2C), this method is more robust under all parameter settings. Simulation 3B ~~~~~~~~~~~~~ In the simulations so far, more trials were used for the indirect than the direct measure, as this is a common practice in experiments (e.g., Jusyte & Schönenberg, 2014; Kiefer et al., 2015; Lin & Murray, 2015; Schoeberl et al., 2015). The difference in the number of trials results in smaller error (and thus more power) in the indirect than the direct measure. To test the impact of this, in Simulation 3B we therefore systematically varied nd and ni. In all other respects, however, the simulation was identical to Simulation 2. For each combination of nd and ni we applied the double t-test and double B method and noted the probability of false support for subliminal priming. Fig. 3A and B summarizes the results of this simulation. The y-axis shows the probability of false support for subliminal priming. For the double t-test method, although a greater number of nd and ni reinforces the probability of false support, an equal number of trials does not completely eliminate the issue (24% and 18% false support when nd = ni = 100 and 200, respectively). With this particular distribution and N, approximately 400 nd were necessary to reduce the proportion of false support for subliminal priming to 5%. As for the double B method, this method is again more robust under all parameter settings. The use of too few trials will usually result in a B encouraging experimenters to collect more data.","In the previous simulations d′ has been used for both the direct and indirect measures. This simplifies the presentation because the outcomes of both measures are on the same scale. In actual experiments a more common indirect measure is a congruency effect (between the prime and target stimulus) on reaction time (RT). We wanted to demonstrate that the results of the previous simulations generalize to experiments using a congruency effect as the indirect measure. Simulation 4 was identical to Simulation 1 except that the indirect measure was a congruency effect between median RT in incongruent – congruent trials. To simplify the simulation, we used the discrete distribution of d′d used in Simulation 1: that is, two types of observers were simulated, perceivers (40%, d′d = 0.2) and non-perceivers (60%, d′d = 0). We assumed no dissociation between the direct and indirect measures. In short, perceivers had a true congruency effect and non-perceivers did not. An ex-Gaussian distribution, i.e., a mixture of a normal and an exponential distribution, was used to model the positive skewness of RT data. The parameter values of the ex-Gaussian distribution are the μ and σ of the normal distribution and the mean of the exponential distribution. Here we obtained parameter values from Miller (1988). For all participants, RTs in congruent trials were drawn from an ex-Gaussian distribution with μ = 300 and σ = 50 and a mean of the exponential distribution of 300. As non-perceivers could not differentiate congruent and incongruent trials, their RT in incongruent trials was drawn from the same distribution as for congruent trials. Thus, non-perceivers’ true congruency effect was 0 ms. As perceivers could perceive the prime, their true RT was modeled as slower in incongruent than congruent trials and was drawn from an ex-Gaussian distribution with μ = 450 and σ = 50 and a mean of the exponential distribution of 150 ms. That is, perceivers’ true median difference (congruency effect) was 48 ms. In this simulation, true mean d′d across observers was thus 0.08 (0.2 × 0.4) and true mean across observers RT effect was 19.2 ms (48 × 0.4). This is representative of priming effect sizes previously reported in studies of subliminal priming (Dehaene & Changeux, 2011; Kouider & Dehaene, 2007; Van den Bussche, Van den Noortgate, & Reynvoet, 2009). For all observers, 100 trials were used to sample RT from the relevant distribution in each condition. Table 4 summarizes the results of applying the double t-test and double B method in the two simulations. To calculate B for the indirect effect, H1 was specified as a normal distribution with a mean of 30 ms and a standard deviation of 10 ms (i.e., BN(30,10); for d′d BH(0,0.2)). This model of H1 was based on priming effect sizes previously reported in studies of subliminal priming (Dehaene & Changeux, 2011; Kouider & Dehaene, 2007; Van den Bussche et al., 2009). As we can see, the results of the previous simulations generalized to a simulation using a congruency effect on RT as the indirect measure.","The previous simulations suggest that, due to random error in the measurements, the double t-test method leads to a high rate of false support for subliminal priming. Perhaps random error can also hide true subliminal priming? We tested this with a final simulation identical to Simulation 2 except that d′d was set to 0 for all observers (that is, we assumed that d′d and d′i were dissociated). Table 5 summarizes the results of applying the double t-test and double B method in this simulation. As can be seen, only 4% (α = 5%) of the direct measures were significant and very few simulations led to B > 3 for the direct measure. We conclude that random error lead to very few missed instances of subliminal priming for either analytical method. The slightly greater conservatism in classifying samples as subliminal when calculating B stems from the random error inherent in using only 100 trials to measure perception. Erring on the side of caution is not necessarily a fault, however. An inconclusive B will encourage researchers to collect more data; a valid procedure using Bayesian statistics, Dienes, 2011). Measurement error ~~~~~~~~~~~~~~~~~ In the simulations, assumption (c) (see Section 2.1 above) was that the observers’ true sensitivity remained constant during all trials of the experiment. This means that the measurement error simulated here is the random error inherent in the binomial nature of the two alternative forced-choice methods. As the simulations illustrate, such random error affects the conclusions that can be made. Another type of measurement error, not simulated here, stems from the fact that an observer’s true sensitivity may not remain constant throughout an experiment. Because of lapses of attention, motivation, or fatigue, a participant may perform sub-optimally. If such attention-driven measurement error affects both the direct and indirect measures similarly, this type of measurement error would not constitute a serious problem for experimenters investigating subliminal priming. However, because the direct task is often more difficult and is often performed at the end of the experiment, the direct measure may be more negatively affected by such measurement error (Newell & Shanks, 2014; Pratte & Rouder, 2009). Other factors that also have been suggested to negatively affect performance in the direct measure is task instructions (Lin & Murray, 2014) and prime-target stimuli similarity (Vermeiren & Cleeremans, 2012). As this would lead to underestimation of d′d, this would increase the risk of false support for subliminal priming beyond what the present simulations suggest. Effect size in the indirect measure ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In simulations 1–3, d′ was used for both the direct and indirect measures. This ensured that the true effect size was the same in both measures (because d′d = d′i). As actual experiments often use different outcome variables for the direct and indirect measures (e.g., d′d and a congruency effect on RT), the effect size and thus power do not need to be similar for the two measures. Perhaps marginally or infrequently perceiving a prime stimulus (d′d slightly above 0) could produce relatively strong congruency effects. This would lead to samples in which the mean d′d may approach chance but median priming effects are subtantial. As this would lead to more power in the indirect measure, the previous simulations may underestimate the extent of false support for subliminal priming in the literature.","The main result of the present simulations is that conventional tests of subliminal priming are not robust. A non-significant direct measure combined with a significant indirect measure does not constitute a reliable diagnostic of subliminal priming, as this result is common even when the only participants who are primed also perceive the prime. Furthermore, finding priming among a subgroup of observers performing close to chance is not a reliable diagnostic of subliminal priming because of random error in the direct measure. If anything, grouping observers based on performance on the direct measure leads to more error than does analyzing the full sample. Finally, as already suggested by Miller (2000), regression analysis will often result in a statistically significant intercept even though there is no dissociation between the direct and indirect measures. In summary, the tests of subliminal priming commonly used in the literature do not support the conclusion experimenters want to make. These simulations do not suggest that subliminal priming cannot occur or that it has not occurred in various studies. However, they do indicate that, with the experimental parameters used in many studies (e.g., González- García et al., 2015; Huang et al., 2014; Jusyte & Schönenberg, 2014; Kido & Makioka, 2015; Kiefer et al., 2015; Lin & Murray, 2015; Marcos Malmierca, 2015; Norman et al., 2015; Ocampo, 2015; Ocampo et al., 2015; Schoeberl et al., 2015; Wildegger et al., 2015), the error rate of falsely declaring priming to be subliminal can be as high as 30–50% (Figs. 2 and 3). This casts doubt on the strength of the evidence for some types of subliminal priming reported in the literature. In all simulations, testing subliminal priming based on Bayesian statistics outperformed the conventional analyses and was much less likely to falsely declare subliminal priming. This is because the Bayes factor distinguishes between data that support H0 and data that are inconclusive, which the p-value cannot do. Our recommendation for experimenters interested in subliminal priming is therefore to use Bayesian statistics to determine whether the examined prime stimulus is subliminal. Dienes (2015) presents a good discussion of how H1 can be specified in this context. Note, however, that calculating a Bayes factor tests only whether or not mean performance is at chance levels. A Bayes factor that supports H0 (i.e., B < 1/3) does not imply that the prime stimulus was subliminal for all participants. As such, even B < 1/3 does not exclude the possibility that priming was in fact was driven by a few observers who perceived the prime. The main difficulty in supporting a claim that priming is subliminal is that observers differ in their absolute thresholds (e.g., Albrecht & Mattler, 2012; Dagenbach et al., 1989; Greenwald et al., 1995; Haase & Fisk, 2015; Sand, 2016). This means that a certain prime stimulus intensity will likely be subliminal for some observers and supraliminal for others. Estimating a threshold for each observer on the other hand, may also be problematic (Rouder et al., 2007). It is therefore unfortunate that the regression approach based on taking individual variability in perception into account has a large error rate. However, part of the reason why regression often fails in this context is a lack of range in d′d. This occurs because experimenters try to use a prime stimulus setting that is subliminal for all observers. For another approach that use a range of d′d for which regression analysis is more robust, see Sand (2016) or Haase and Fisk (2015). There is a growing concern about false positive psychology (Simmons, Nelson, & Simonsohn, 2011) and the replicability of psychological findings (Open Science Collaboration, 2015). We submit that false positive findings may arise from faulty interpretation of non- significant results. Problems associated with interpreting p > α (failure to reject H0) and how this situation could be improved by Bayesian statistics, have been discussed at length elsewhere (Dienes, 2014; Gallistel, 2009). Our simulations also suggest that false positives may well be replicable if several experimenters use similar paradigms and questionable analytical strategies. We therefore hope that the field of subliminal priming will start using other experimental designs and statistical approaches to test subliminal priming."],["The current study tests the hypothesis that shy children's reduced word learning is partly due to an effect of shyness on attention during object labeling. A sample of 20- and 26-month-old children (N = 32) took part in a looking-while-listening task in which they saw sets of familiar and novel objects while hearing familiar or novel labels. Overall, children increased attention to familiar objects when hearing their labels, and they divided their attention equally between the target and competitors when hearing novel labels. Critically, shyness reduced attention to the target object regardless of whether the heard label was novel or familiar. When children's retention of the novel word–object mappings was tested after a delay, it was found that children who showed increased attention to novel objects during labeling showed better retention. Taken together, these findings suggest that shyer children perform less well than their less shy peers on measures of word learning because their attention to the target object is dampened. Thus, this work presents evidence that shyness modulates the low-level processes of visual attention that unfold during word learning. --------------------------------------------------------------------------------","Reports on the development of language usually begin by celebrating children’s remarkable ability to learn words. For example, it is often stated that by the time of their second birthday, children have typically acquired a vocabulary of more than 300 words (Fenson et al., 1994). Although these summaries are an impressive illustration of the speed of language acquisition, they often ignore an equally fascinating aspect of early language acquisition—its variability. For example, on closer examination, the data also show that 10% of 2-year-old children are able to produce more than 528 different words, whereas the 10% at the opposite end of the scale produce fewer than 66 different words. Interestingly, the majority of children in the bottom 10% will show no later difficulties with language development (Kelly, 1998). Research into variability in early language acquisition typically focuses on the role of environmental factors. We know that wide-ranging extrinsic factors such as socioeconomic status (SES), birth order, and differences in daycare quality exert an effect on children's language development. For example, it has been consistently shown that children from more affluent backgrounds acquire language more quickly than their less well-off peers (Hoff-Ginsberg, 1991). Similarly, children with older siblings show an earlier grasp of pronoun use than first-born children (Oshima- Takane, Goodz, & Derevensky, 1996) and, unsurprisingly, children in higher-quality daycare show advanced language and communication skills (Burchinal, Roberts, Nabors, & Bryant, 1996). The mechanisms underlying extrinsic effects on language development are often also explained in terms of the environment. Taking SES as an example, many argue that children raised in better-off families acquire language more quickly because their language input is easier to learn from (Hirsh-Pasek et al., 2015), contains more words overall (Hart & Risley, 2003), contains a greater variety of word types (Hoff, 2003), and is more likely to be in a child-directed register (Rowe, 2008). Thus, these accounts argue that the effect of SES on language development can be attributed to the effect of SES on children’s language input. Yet, variability during development is not only present in the environment but also exists within children. Even from birth, children show marked differences in their reaction to the environment. Some babies are consistently quick to settle after stressful events (e.g., sudden loud noises or inoculations), whereas others are highly reactive to such events, as demonstrated by a display of intense motor reaction and distress (e.g., Kagan & Snidman, 1991; Worobey & Lewis, 1989). Throughout development, further individual differences in response to the environment emerge. Some infants become easily distracted, whereas others show no difficulty in focusing attention on toys or events for long periods of time (Rothbart, 1981). Thus, such myriad individual differences suggest that, even given identical input, two children taken at random could process this input very differently. We know that individual differences in children's behavioral style, better known as temperament (Rothbart, 1981), can explain some variability in early language development. Differences in children's inhibition and discomfort in novel social situations (i.e., shyness; Putnam, Gartstein, & Rothbart, 2006), have been consistently shown to affect children’s language development. Shyness is negatively correlated with vocabulary size when measured via parent-report checklists of children’s vocabulary (Paul & Kellogg, 1997; Slomkowski, Nelson, Dunn, & Plomin, 1992), and these differences in productive vocabulary have also been confirmed experimentally, with shyer children being found to speak less than their less shy peers in both familiar and unfamiliar settings (Asendorpf & Meier, 1993; Crozier & Badawood, 2009; Evans, 1987). Various explanations have been put forward to explain the relation between shyness and measures of language development. Some have argued that shyness affects the propensity to respond (e.g., Smith Watts et al., 2014). According to this account, shyness does not affect language development per se; instead, shy children are less likely to demonstrate their language skills, leading to their reduced scores on language measures. Others argue that shyness exerts an effect on language development indirectly by modulating the environment; shyer children are likely to reduce their interactions in novel social settings and, in this way, restrict their exposure to language in comparison with less shy peers (e.g., Evans, 1993). This suggestion has some support, for example from Evans (1996), who found that children who are consistently reticent to talk in social settings, rather than those who are just slow to warm up, show reduced scores on tests of language ability. Generalization across previous reports of the relation between shyness and language is problematic due to differences in the precise operationalization and measurement of shyness. Whereas some have measured individual temperamental differences by examining children’s responses in a face-to-face task (e.g., Slomkowski et al., 1992), the difficulty with such an approach is that it does not allow any certainty that the measure reflects stable individual differences. It is possible that “shy-type” behaviors are exhibited during a single session for more transient reasons (e.g., tiredness). Other studies have dealt with this issue by capitalizing on caregivers’ experience of their children’s enduring behavioral style, using parent-report measures of shyness (e.g., Asendorpf & Meier, 1993). The Early Childhood Behavior Questionnaire (ECBQ; Putnam et al., 2006) is one such parent-report standardized measure of children’s temperament subdivided into different scales, one of which measures shyness by asking parents to indicate how often their child exhibits shy- type behaviors such as turning away from strangers and showing hesitation in approaching unfamiliar children. A recent account based on Putnam et al. (2006) approach argued that shyness can affect language development by modulating the low-level processes by which language is acquired (Hilton & Westermann, 2017). Shyer children demonstrate an avoidance of eye contact during social interaction (Putnam et al., 2006), indicating that shyness modulates attentional processing during these interactional episodes, and there is evidence that shyness may also modulate children’s allocation of attentional resources in the absence of social interaction: Pérez-Edgar and Fox (2005) demonstrated in a screen- based cueing task that whereas less shy children attended preferentially to cued locations associated with a reward, shyer children preferentially attended to locations associated with a penalty. Given that the earliest stages of language learning are governed at least in part by the cognitive systems that underlie attentional focus abilities (e.g., Dixon & Hull Smith, 2008), it is likely that any effects of shyness on attentional processing could also affect language learning. Much work on the critical role of attentional processes during language learning has examined how externally directing children’s attention affects their formation and retention of novel word–object mappings. Typically, children are presented with referent selection trials on which one novel object is presented alongside familiar competitors and experimenters record whether children select the novel object when asked for a novel label (e.g., “Where’s the blicket?” ; Horst & Samuelson, 2008). Then, after a 5-min break, children’s retention of this newly formed mapping is tested. Critically, by manipulating key elements of this task, such as the number and nature of the familiar competitors, researchers have uncovered evidence to suggest that children’s attention during the presentation of objects and their corresponding labels affects their learning of these word–object mappings. For example, Axelsson, Churchley, and Horst (2012) drew children’s attention to the novel object during labeling by flashing a light underneath it and partially covering competitors and demonstrated that children retained novel word–object mappings better under these conditions than when the novel object was only pointed to during labeling or when the light and covering were presented individually. Such work indicates that children who focus attention on targets during labeling are better able to learn the word–object mappings, and related work demonstrates that attention to the target can be modulated by extrinsic cues such as novelty (Horst, Scott, & Pollard, 2010; Kucker & Samuelson, 2012). Given the critical role of attention in successful word learning, it is possible that shyness affects language learning by modulating attentional processes during labeling episodes. However, there has been little examination of the intrinsic factors that may drive these differences in attention. Hilton and Westermann (2017) were the first to test whether attentional processes during word learning are affected by shyness. In their study, 24-month-old children were presented with sets of three objects, one of which was novel and two of which were familiar. When asked for a novel object using a novel word, shyer children selected objects at chance levels, whereas less shy children reliably chose the novel object. Furthermore, after a 5-min break, shyer children did not retain any of the (few) novel word–object mappings that they had formed during referent selection, whereas less shy children showed evidence of retaining these mappings. The authors argued that these findings could be explained in terms of differences in attention during labeling. Specifically, shyer children’s aversion to novelty (Kagan, Reznick, & Snidman, 1988) may have reduced their attention to the target (novel) object during labeling, disrupting encoding and, thus, formation of the word–object mapping. Of course, given that Hilton and Westermann’s (2017) study only measured children’s manual object selection, it was impossible to directly demonstrate whether shyness affected children’s word learning via attention during referent selection. Thus, the current study aimed to examine this possibility using a novel adaptation of the looking-while-listening paradigm (Fernald, Zangl, Portillo, & Marchman, 2008; Swingley, 2011). Children were presented with images of one novel object and two familiar objects on a screen, and their eye gaze across the objects was recorded while a familiar or novel label was heard. As well as providing a highly detailed picture of children’s online processing, this approach has the advantage of measuring implicit behavior. An explicit response, such as pointing at an object as in Hilton and Westermann’s (2017) study, has a social dimension, and it is possible that the random responses exhibited by shy children in that study related to social expectation. Recording eye movements in a task without an experimenter present addresses this possibility. In this study, we explored whether shyness as measured by the ECBQ modulates children’s attention to target objects during labeling. Based on the assumption that children’s mapping of a novel label to a novel object involves ruling out potential competitors (Halberda, 2006; Horst et al., 2010), children’s looking to competitor objects should be a critical step in the formation of a novel word–object mapping. However, in line with previous research (Axelsson et al., 2012; Horst et al., 2010), we expected that children who were better able to focus their attention on the novel object following this initial disambiguation would be better able to retain the mapping. Therefore, we also tested whether looking during labeling could predict retention of word–object mappings and whether this relation changes over development, as demonstrated in previous work (Bion, Borovsky, & Fernald, 2013).","A total of 32 typically developing 20- to 26-month-old children took part in the study. All children were monolingual English speakers. There were 16 children in the 20-month age group (M = 20 months 13 days, range = 19 months 11 days to 21 months 5 days; 8 girls). An additional 3 20-month-old children were excluded due to failure to complete the task (n = 2) or experimenter error (n = 1). There were also 16 children in the 26-month age group (M = 26 months 13 days, range = 25 months 19 days to 27 months 0 days; 8 girls). An additional 2 26-month-old children were excluded due to equipment failure (n = 1) or failure to complete the task (n = 1). These sample sizes are in line with previous research using a similar methodology (Twomey, Ma, & Westermann, 2018) and provide more than 80% power at the p < .05 level to detect an effect size of 0.56 in mutual exclusivity (referent selection) tasks (Bergmann et al., 2018; Lewis & Frank, 2018). Families were recruited by contacting parents who had previously indicated interest in participating in child development research. Parents’ travel expenses were reimbursed, and children were offered a storybook to thank them for participating. Written informed consent was provided by the participants’ parents. Prior to the study, parents completed the Oxford CDI (Hamilton, Plunkett, & Schafer, 2000), a British English adaption of the widely used MacArthur–Bates Communicative Development Inventory (Fenson et al., 1994). Vocabulary scores for 3 20-month-old children were not available because their parents failed to return the questionnaire. These missing values were replaced by the mean score. The 20-month-old group produced a mean of 127 words (range = 17–413) and understood a mean of 252 words (range = 104–413). The 26-month-old group produced a mean of 179 words (range = 158–386) and understood a mean of 347 words (range = 217–416). All vocabulary scores were within the normal range (Floccia, 2017; Frank, Braginsky, Yurovsky, & Marchman, 2017). As expected, the 26-month-old children had larger receptive and productive vocabularies than the 20-month-old children [receptive: t(30) = 3.81, p < .001; productive: t(30) = 4.61, p < .001)]. Stimuli and design ~~~~~~~~~~~~~~~~~~ During their visit to the lab, children took part in referent selection trials, which were presented on a computer screen, and retention trials based on their selection of the actual three-dimensional (3D) objects. Visual stimuli for referent selection trials consisted of digital photographs of known and novel objects. For each participant, the 12 referent selection pictures were randomly grouped into sets of 3, with each set comprising one novel object and two familiar objects, so that the objects comprising each set were consistent for each participant but varied across participants. Novel objects were a plastic tripod, a wooden roller, a wooden dumbbell. and a plastic tea strainer, the same novel objects used with a similar age group in a previous study (Hilton & Westermann, 2017). Known objects were toy versions of the following vehicles, animals, and household items: car, motorbike, elephant, fish, pig, ball, fork, and spoon. Each picture was of similar size on the screen (∼700 × 800 mm). Each novel object was assigned one of four pseudowords (cheem, koba, sprock, or tannin) chosen to be plausible pseudowords in English and having been used in previous word learning studies (Halberda, 2006; Horst & Samuelson, 2008; Markson & Bloom, 1997; Samuelson & Horst, 2007). For each trial, a short video was created. Each video began by showing the set of pictures bouncing onto the screen, accompanied by a short bouncing sound effect, in order to focus children’s attention on the screen at the beginning of the trial. Once the objects had finished bouncing (2000 ms after the start of the video), the audio stimulus automatically began playing. Audio stimuli consisted of three labeling phrases spoken by a male native British English speaker in infant-directed speech. The target object label appeared in the final position of each labeling phrase (“Look, there’s a _____! Where’s the _____? Wow, it’s a _____!”). The three labeling phrases were in the same order on all trials. Label onsets were 2500, 5200, and 9100 ms after stimulus onset, and the final label offset was at 10,000 ms. Once the auditory stimulus had finished playing, the objects stayed on the screen for a further 2000 ms before disappearing and leaving the screen clear for the next trial. All participants took part in 12 referent selection trials. Each set of pictures was presented three times; on 2 trials children heard the novel label (novel trials), and on 1 trial children heard one of the two familiar object labels (randomly selected; familiar trials). The order in which the sets were presented was randomized. The location of each object (on the left, in the center, or on the right) was pseudorandomized across trials, ensuring that the target did not appear in the same location on more than 2 successive trials, and the order of trial types was pseudorandomized, ensuring that a trial type (i.e., novel or familiar label) was not repeated more than three times in a row. Seven 3D objects were used for retention trials. Three randomly selected familiar objects (e.g., toy duck, car, and fork) acted as stimuli for the initial warm-up trials, and the four novel objects of which children saw images during referent selection were stimuli on test trials. Shyness questionnaire During their visit, parents completed the shyness scale of the ECBQ (Putnam et al., 2006). The ECBQ is a standardized parent-report measure of 18- to 36-month- old children’s emerging temperament, and the shyness scale asks parents to rate from 1 to 7 (1 = never, 7 = always) how often over the previous 2 weeks their children had demonstrated shy-type behaviors (e.g., “When playing with unfamiliar children, how often did your child seem uncomfortable?”). The average score across the 12 questions (Cronbach’s α = .86) for each child was calculated, resulting in a score between 1 (not at all shy) and 7 (extremely shy). To avoid demand biases in parents’ responses, three other questions taken from one other unrelated subdimension (perceptual sensitivity) were included in the questionnaire. Referent selection trials During referent selection, children were seated centrally 50 to 70 cm in front of a computer screen on their parent’s lap. A Tobii X120 eye tracker located beneath the screen recorded children’s eye gaze, and a video camera above the screen recorded parents and children throughout the procedure. Parents were instructed not to look at the screen or speak to their children as the videos were playing so as to avoid influencing their children’s behavior, and the experimenter monitored the testing session via a camera to ensure that this instruction was followed. Prior to the experiment starting, children’s eye gaze was calibrated using a 5-point calibration procedure. A child-friendly animation (a wobbling duck) was displayed in the four corners and center of a 3 × 3 grid accompanied by a jingling sound, and calibration accuracy was checked and calibration was repeated if necessary. Once the calibration procedure was completed, the referent selection trials were presented. Data coding and cleaning The sampling rate of the eye tracker was 60 Hz, and the raw data files were exported from Tobii Studio (Version 3.4) and analyzed in R (R Core Team, 2017). Three rectangular areas of interest (AOIs) were defined as the three areas of the screen where the stationary stimuli were displayed. All AOIs measured 385 by 408 pixels. AOI 1 covered the object on the left-hand side of the screen, AOI 2 covered the object in the middle, and AOI 3 covered the object on the right-hand side. There was a margin of 107 pixels between AOIs. Continuous gaze within an AOI was counted as a fixation. If continuous gaze within an AOI was interrupted for less than 60 ms, this interruption was recoded as a continuation of that fixation because the interruption was most likely due to eye blinks rather than children rapidly reorienting their attention (see Yu & Smith, 2011). All analyses were conducted on children’s looking times from 233 ms after the first label onset to allow for saccade preparation (Canfield & Haith, 1991; Haith, Hazan, & Goodman, 1988) until 12,000 ms when the objects disappeared from the screen. Retention trials After the 12 referent selection trials, children took a 5-min break by playing in an adjacent room, in line with Horst and Samuelson (2008). Children then took part in retention trials during which they sat on their parent’s lap at a table opposite the experimenter. Retention trials consisted of a three-alternative forced-choice task in which participants were required to select a target object from an array in response to a target word requested by the experimenter, a typical task used to measure children’s mapping of words to objects (e.g., Waxman and Booth, 2001; Dysart, Mather, & Riggs, 2016; Samuelson & Horst, 2007). Prior to the retention test trials, a series of warm-up trials was presented to ensure that children understood the task and were responding appropriately to the experimenter’s requests. On each warm-up trial, children were presented with three randomly selected familiar objects side by side on a tray divided into three sections. These familiar objects had previously acted as competitor objects during referent selection. After allowing children to look at the objects for 3 s, the experimenter requested one of the objects (e.g., “Where’s the car?”) before sliding the tray forward and allowing children to make their choice by pointing at or retrieving the object. On the next trial, the objects were rearranged out of sight of the children and another object was requested. These trials continued until children had correctly selected the target object on three consecutive trials. Across the initial three trials, each object was requested once and the target object appeared in each location on the tray once. On any further trials, the position of the target was randomly determined. To encourage participation, children’s correct responses were praised and corrections were offered when children did not at first select the correct object. Following the warm-up trials, the retention trials began. The procedure for retention trials was identical to that for warm-up trials except that no praise or corrections were offered so as to avoid biasing children’s responses. Instead, once children had selected an object, the experimenter replied “thank you” in a neutral tone, replaced the object, and then began the next trial. Across the 4 retention trials each novel object acted as the target once, and on each trial two other randomly selected novel objects acted as competitors. Pictures of these same novel objects had been presented during referent selection, and the experimenter requested the target object using the novel word with which it had appeared during referent selection. The order of trials and the position of the target on the tray were randomized. Looking during referent selection ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To establish whether shyness and trial type affected overall looking time, we submitted raw looking times to a linear mixed-effects model (LMEM) with main effects of shyness score (mean centered), competitor or target fixation (hit type; for all models, dummy coded: target = 1, Competitor 1 = 0, Competitor 2 = 0), and trial type (for all models, dummy coded: familiar = 1, novel = 0) and their interactions, by-participant random intercepts and slopes, and by-target random intercepts and slopes. Results are reported in Table 1 with significant predictors highlighted in bold. Initial analyses revealed no main effect of, or interactions with, age group, so we collapsed across this factor. All LMEMs were conducted using the lme4 package Version 1.1–12 in R Studio Version 3.2.4 (R Core Team, 2017) and had maximal random effects structures simplified until convergence (Barr, Levy, Scheepers, & Tily, 2013). The p values for fixed effects were obtained using sequential likelihood ratio tests. Overall, children looked more at the competitors than at the target (negative main effect of hit type), but children’s attention to targets and competitors differed according to trial type (significant hit type × trial type interaction). Follow-up post hoc tests were performed on the hit type × trial type interaction using the multcomp package Version 1.4–7 (Hothorn, Bretz, & Westfall, 2008), revealing that on familiar label trials children looked at the target more than at competitors (z = 4.87, p < .001), but this difference was not found on novel label trials (z = − 0.79, p > .99), as shown in Fig. 1. Table 1 also shows an interaction between shyness score and hit type. As can be seen in Fig. 2, this interaction could likely be explained by a negative relation between shyness and target looking, whereas no relation is visible between shyness and competitor looking. To examine whether this interpretation was supported by the data, we analyzed looking to the competitors and looking to the target in separate LMEMs. The results revealed that shyness was indeed related to target object looking [beta = 0.27, SE = 0.12, t = − 2.34, χ2(1) = 4.95, p = .026] but not to competitor object looking [beta = 0.12, SE = 0.11, t = 1.08, χ2(1) = 1.18, p = .28]. This finding demonstrates that shyer children looked less at the target object during labeling in comparison with less shy children. However, shyness was not associated with children’s attention to the competitor objects. Retention trials ~~~~~~~~~~~~~~~~ We next wanted to examine whether shyness and target looking during referent selection predicted retention scores. To do this, raw looking time to each novel target during referent selection was calculated. Retention trials were scored as 1 if children correctly retained the novel word–object mapping and as 0 if they did not. Trials on which no response was offered (n = 5) were excluded from the analyses. Given the previously observed differences in retention of word–object mappings between our two age groups (Bion et al., 2013), age group was reintroduced as a predictor in these models (dummy coded: 20-month group = 0, 26-month group = 1). Therefore, we submitted retention scores to a binomial LMEM with fixed effects of age, shyness score, and novel target looking during referent selection and by-participant and by-target random intercepts. Results are presented in Table 2 with significant predictors highlighted in bold. The positive main effect of target looking is evidence that looking to the target during labeling supports retention of the label–object association. However, children overall failed to retain any word–object mappings at levels better than chance, t(30) = 1.25, p = .22. The absence of a main effect of shyness suggests that retention and shyness were not directly related. However, because the results from the referent selection trials indicate that shyness modulates attention during labeling, this result leaves open the possibility that shyness could affect word learning. We return to this point in the Discussion.","The current work set out to examine whether shyness affects the processes modulating attention during labeling and children’s subsequent word learning. In line with previous research, we found that familiar labels encouraged target looking, whereas novel labels led to equal competitor and target looking. Critically, however, we found that shyness was linked to a reduction in children's attention to the target object regardless of whether it was familiar or novel. Furthermore, although we found no evidence of retention at the group level, we found that individual infants’ increased attention to the target during referent selection was positively related to retention of the novel word–object mappings. As expected, we found differences in looking depending on whether a familiar or novel label was heard. On familiar label trials, children were processing an already-learned word–object association and so showed heightened attention to the target object in comparison with competitors. Such behavior is in line with previous research showing that children focus attention on known objects that are being labeled (Fernald, Pinto, Swingley, Weinbergy, & McRoberts, 1998). In contrast, we also found that on hearing a novel label, children spread their attention equally across the competitors and the target, likely because the novel label trials entail disambiguation, requiring children to first rule out competitors as potential referents before mapping the novel label to the novel object (Halberda, 2006). This finding offers converging evidence that initial attention to competitors is a crucial aspect of novel word disambiguation (Axelsson et al., 2012; Halberda, 2006; Horst et al., 2010). Importantly, shyness reduced children's attention to the target object, regardless of whether the target was novel or familiar. This finding suggests that shyness modulates children's attention to labeled objects generally, and this is in line with Hilton and Westermann's (2017) finding that shyness affected not only children’s formation of novel word–object associations but also their selection of familiar referents. Interestingly, shyness did not affect overall looking time to the object array, suggesting that shyness did not reduce attention to the task but, critically, modulated attentional patterns across the potential referents. Thus, the current work supports the theory that shy children’s reduced performance during referent selection is related to their reduced attention to the target object rather than to a failure to engage with the task (Hilton & Westermann, 2017). Given the complex interplay of attention during labeling, there are several potential mechanisms to explain the effect of shyness on target looking. Shy children’s aversion to novelty (Kagan et al., 1988) may have caused children to look away from the novel object in favor of the familiar objects, but this explanation would also predict that shyness has no effect on target looking on familiar label trials. This was not what we found: Shy children looked less at targets on both trial types. Conversely, shyer children might have shown heightened attention to the novel object because we know that children attend more to stimuli that they find aversive (Field, 2006). However, this explanation does not support the finding that shyer children show reduced target looking on novel label trials. Because we know that shyness is linked to difficulty in engaging with novelty (Kagan et al., 1988), we argue that it is most likely that for shy children the novelty inherent in such a labeling episode (i.e., presence of a novel object and label, presentation in a novel voice) disrupts the ongoing attentional processes required to form the word–object association. This disruption is also pervasive enough to reduce shy children’s ability to focus attention on the referent of a familiar label. A critical contribution of the current work is that we found an effect of shyness on responses during referent selection even when no social interaction was required. Given shy children’s aversion to novel social stimuli (Putnam et al., 2006), it is unsurprising that shy children show difficulty in responding to experimental tasks that are presented face to face (Crozier & Hostettler, 2003) and that shy children show differences in processing stimuli that are strongly associated with social interaction such as faces (Matsuda, Okanoya, & Myowa-Yamakoshi, 2013). The current work, however, finds a relation between shyness and attention during referent selection even when no explicit response is required and when no unknown adults are present. Thus, the relation between shyness and attention seems robust across contexts—even those contexts that do not require or prompt social interaction. This finding suggests that we must be vigilant to possible effects of shyness when designing and interpreting results of developmental research. Unlike the referent selection trials, retention trials presented children with 3D objects. Nonetheless, a relationship between attention during referent selection and retention was found across this change in task: Although children overall looked equally across competitors and target on novel label trials, attention to the novel target was positively related to retention. Put differently, children who looked more to novel targets during labeling were more likely to retain the novel label–object mappings. Thus, results from the retention trials highlight the critical role of attention in predicting retention of newly formed word–object mappings, convergent with recent behavioral work (Kucker, McMurray, & Samuelson, 2018). The current study, therefore, suggests that although attention to competitors has been shown to affect children’s ability to retain label–object mappings (Horst et al., 2010; Zosh, Brinster, & Halberda, 2013), it is those children who are better able to focus their attention on the target that are more likely to retain the label’s meaning. This finding supports recent accounts of successful word learning as the product of two separate but related sets of processes: those processes that support disambiguation by driving elimination of competitors as potential referents, followed by those processes that support sustained attention to the target object, allowing for a robust association between the object and its label (see also Kucker, McMurray, & Samuelson, 2015). Interestingly, given the additional social component in the retention trials, we did not find a direct relation between shyness and retention. It is possible that this object selection measure, with a single data point per trial, was not sensitive enough to demonstrate a direct relation to shyness. Equally, although at the individual level retention was related to attention during referent selection, it is possible that the retention task resulted in a floor effect at the group level because it required children to transfer their learning from two-dimensional (2D) pictures of the objects to the actual 3D objects, something that we know is difficult for children overall (e.g., Krcmar, Grela, & Lin, 2007). These findings highlight the importance of exploring the processes underlying the relations among shyness, attention, and language development at the individual level as well as at the group level. Although we found an effect of shyness on attention during referent selection, we found no effect of age, which is noteworthy given that existing literature suggests that referent selection behaviors undergo a shift sometime toward the end of the second year of life. For example, Bion et al. (2013) found that 18-month-old children showed no heightened attention to a novel object on hearing a novel label, whereas 24-month-old children did. However, in our study, at both 20 and 26 months of age children showed no heightened looking to the novel object on novel label trials. Whereas Bion et al. presented participants with one novel target alongside a familiar competitor, the current study presented two familiar competitors alongside each novel target. Thus, it is possible that the additional competitor in the current study increased the cognitive load of the task so that children needed more time to rule out competitors as potential referents, reducing the time spent looking at the novel object to chance levels. This discrepancy between the findings of the current study and those of Bion et al. serves as an important illustration of the role played by competitors during novel word disambiguation. Overall, then, the current study demonstrates that shyness modulates the low-level attentional processes that unfold during word learning. Although future work is required to establish a direct link between shyness and retention, this finding may offer an explanation for shyer children's slower vocabulary development in comparison with less shy children (Spere, Schmidt, Theall-Honey, & Martin-Chang, 2004). Shyer children’s reduced attention to objects during labeling could mean that they require more exposure to the object–label pairing so as to allow them greater opportunity to encode the association. Over and above the predominant explanations for the negative relation between children’s shyness and vocabulary (i.e., reticence to demonstrate their language ability; Smith Watts et al., 2014), we identified a novel potential contributing factor: We showed that shyness also affects the early implicit processes that support one of the earliest stages of language learning. Early word learning is the product of a complex and dynamic set of cognitive processes, and we have come far in understanding the environmental factors that modulate these processes (e.g., Axelsson et al., 2012; Twomey et al., 2018). In the current study, we showed that shyness is an intrinsic factor that influences the attentional processes supporting word learning. Therefore, we must now consider the role of this and other individual differences in word learning so as to better understand why early language development is so variable."],["Tax compliance represents a social dilemma in which the short-term self-interest to minimize tax payments is at odds with the collective long-term interest to provide sufficient tax funds for public goods. According to the Slippery Slope Framework, the social dilemma can be solved and tax compliance can be guaranteed by power of tax authorities and trust in tax authorities. The framework, however, remains silent on the dynamics between power and trust. The aim of the present theoretical paper is to conceptualize the dynamics between power and trust by differentiating coercive and legitimate power and reason-based and implicit trust. Insights into this dynamic are derived from an integration of a wide range of literature such as on organizational behavior and social influence. Conclusions on the effect of the dynamics between power and trust on the interaction climate between authorities and individuals and subsequent individual motivation of cooperation in social dilemmas such as tax contributions are drawn. Practically, the assumptions on the dynamics can be utilized by authorities to increase cooperation and to change the interaction climate from an antagonistic climate to a service and confidence climate. --------------------------------------------------------------------------------","Citizens appreciate public goods such as schools or hospitals. Funding the public goods through taxpaying, however, represents a social dilemma in which the individual short-term interest to minimize paying taxes is at odds with the long-term collective interest to ensure sufficient tax payments for financing the public goods (Balliet & Van Lange, 2013). To overcome the social dilemma and to insure high tax compliance among citizens, tax authorities rely on two measures. Power measures such as audits and fines and trust related measures such as fair procedures (e.g., Allingham & Sandmo, 1972; Feld & Frey, 2007; Srinivasan, 1973). In research, the positive impact of both measures on tax compliance received empirical support (e.g., Muehlbacher & Kirchler, 2010; Wahl, Kastlunger, & Kirchler, 2010). Surface validity might suggest that power and trust are incompatible and the opposites of each other. In contrast, we assume that power and trust are related in a specific dynamic in which they mutually destroy or mutually foster each other and in turn influence tax compliance. However, distinct theoretical assumptions about the dynamics between power and trust are missing. The purpose of the present theoretical paper is to conceptualize these dynamics and to elaborate on how they might influence tax compliance. This conceptualization serves as the theoretical basis for empirical research and conclusions how to increase tax compliance in particular and cooperation in social dilemmas in general. There is little doubt that audits and fines are necessary to levy taxes, however, they are not the only determinants to ensure contributions. Experiments on tax behavior in the laboratory have consistently supported the positive impact of audits and fines on compliance (Blackwell, 2007). Nonetheless, the effects are rather weak. Field studies and surveys have yielded effects that are lower than, and sometimes the opposite of the predicted effects (e. g., Andreoni, Erard, & Feinstein, 1998). Additionally, Feld and Frey (2007) question whether audits and fines may destroy trust, as they crowd out the intrinsic motivation to cooperate among committed and cooperative citizens. Thus, besides “economic” determinants such as audits and fines, “psychological” determinants such as the motivation to comply, the attitudes of taxpayers towards the state, the government and taxation, transparency and understanding of tax laws, personal and social norms, and fairness perceptions were shown to impact tax compliance (Braithwaite, 2003; Kirchler, 2007; Torgler, 2003). Kirchler (2007) and Kirchler, Hoelzl, and Wahl (2008) endeavored to integrate the economic and psychological factors into a comprehensive two-dimensional framework, the Slippery Slope Framework (SSF). The dimension power of authorities aggregates economic determinants and is defined by taxpayers' perception of authorities' capacity to detect and punish tax evaders. The dimension trust in authorities covers psychological bases of tax compliance and results from taxpayers' general opinion that the tax law and regulations are clear and easy to follow, and that the tax authorities operate fairly and benevolently in the interest of the community. The SSF asserts that both the power of authorities and the trust in authorities can solve the social dilemma of tax compliance. On the individual taxpayer level, the framework differentiates between two motivations to comply with tax law, enforced compliance and voluntary cooperation. Enforced compliance results from the power of tax authorities, whereas voluntary cooperation is driven by the taxpayers' trust in tax authorities. On the aggregate level, the SSF postulates that power and trust define different interaction climates between tax authorities and taxpayers: while the exertion of strong power by the authorities fosters an antagonistic climate, high trust is the prerequisite of a synergistic climate (Kirchler, 2007; Kirchler et al. 2008). Fig. 1 depicts power and trust as independent dimensions, positively related to enforced compliance and voluntary cooperation, respectively, and to an antagonistic and synergistic climate, respectively. Empirical evidence generally supports the relevance of power and trust as determinants of compliance (Kogler et al., 2013; Muehlbacher & Kirchler, 2010; Muehlbacher, Kirchler, & Schwarzenberger, 2011; Wahl, Endres, Kirchler, & Böck, 2011; Wahl et al. 2010). For instance, in a representative sample of self-employed taxpayers, trust and power co-varied with tax compliance (Muehlbacher & Kirchler, 2010). Kogler et al. (2013) and Wahl et al. (2010) found that compliance is highest if both power and trust are perceived as high. This result suggests an additive effect of power and trust. Moreover, a dynamic relationship between power and trust can be assumed. In the conceptualization of the SSF, Kirchler et al. (2008) speculate about a dynamic relationship but they offer no elaboration of the possible interaction effects between power and trust. In contrast to surface validity, which might suggest that power and trust are incompatible, they assume that power and trust might not only weaken but also strengthen each other. So far, empirical studies in the tax behavior context suggest that power and trust are influencing each other positively (Kogler et al. 2013; Muehlbacher et al. 2011; Wahl et al. 2010). Nevertheless, in various research fields the theoretical conceptualization and the empirical evidence for the mutual effects of power and trust are inconsistent, which suggests that there is both a fostering as well as an eroding influence of power on trust (Adler, 2001; Bijlsma-Frankema & Costa, 2005; Das & Teng, 1998; Kumlin & Rothstein, 2005; Möllering, 2005). This inconsistency may originate from different conceptualizations of power and trust and from diverse operationalizations in empirical investigations. Therefore, we propose to distinguish between the independent qualities of coercive power and legitimate power. We further differentiate between reason-based trust and implicit trust. These distinctions will provide an explanation of the dynamics between power and trust. The aim of the present paper is to shed light on the effects of the mutual interaction of coercive and legitimate power on the one hand, and reason-based and implicit trust on the other hand, as well as to formulate assumptions regarding the consequences on the interaction climate between tax authorities and taxpayers as well as on tax compliance. Consequently, we extend the SSF by distinguishing between three types of interaction climates resembling for instance Alm and Torgler's (2011) interaction styles and respective qualities of cooperation comparable to Kelman's (2006) psychological processes of social influence. Hence, we not only integrate the dynamics between power and trust in well-established existing theories but more importantly show how these dynamics can be used to transform a hostile interaction into an interaction in which voluntary and committed cooperation prevails. We extend the SSF borrowing from the literature on social dilemmas, social influence, organizational behavior, and leadership; hence, we make predictions beyond tax compliance on general interaction climates, motivations to cooperate, and eventually, contributions to public goods which are regulated by authorities such as insurance funds, public transportations, or business organizations. All these cases represent social dilemmas, similar to the social dilemma of tax compliance, in which the short-term self-interest is at odds with longer-term collective interest (Van Lange, Joireman, Parks, & Van Dijk, 2013). Each single individual would be better off by not contributing to the public good but nonetheless taking advantage of the public good provision (Dawes, 1980; Ostrom, 2000). However, if all individuals would chose this strategy no public good would be provided and eventually, all would end up worse off than if all had cooperated (Dawes, 1980). Authorities, however, as intermediates are one possibility overcoming this tragedy of the commons by actively regulating and monitoring the individual contributions to the public good (Van Vugt, 2009; Van Vugt & De Cremer, 1999). Hence, although our predictions on the dynamics between power and trust are focused on tax authorities interacting with taxpayers, we propose, that these predictions apply to all authorities regulating individuals' contributions to public goods. In the remainder of this paper we first introduce the concepts of coercive power and legitimate power, and reason-based trust and implicit trust. Second, we speculate on the dynamics between the different qualities of power and trust and how these impact tax compliance. Third, we discuss the consequences of different qualities of power and trust for interaction climates and the respective motivations to comply. Fourth, the paper concludes with observations on the transformation from one type of interaction climate to another.","Power has received much attention in various scientific disciplines. Besides specific perspectives taken by different disciplines, there is considerable agreement on a general definition of power. Power is consistently defined as the potential and perceived ability of a party to influence another party's behavior (e.g., Freiberg, 2010; French & Raven, 1959; Molm, 1994). In research on the regulation mechanisms of citizens' behavior, two competing theories of power are widely recognized, the conceptualizations of coercive and legitimate power. The perspective on coercive power is based on Becker's (1968) economic approach which argues for strict control and punishment to influence individuals' utility functions and in turn, their behavior. The second and more recently developed approach by Tyler (2006) argues that legitimate power, i.e., the power of accepted authorities, is more appropriate and effective in shaping individuals' behavior than severe controls and punishment. We seek to integrate both perspectives of power in the SSF and refer to the social-psychological theory of the ”bases of social power” developed by French and Raven (1959), and Raven (1965). The bases of social power were initially conceptualized to explain relations between supervisor and employee, i.e. individuals. It can, however, be assumed that people's behavior in organizations, public institutions, and the state is shaped by the same perceptions and judgments of the dominant party as in bilateral relationships or small group settings (Tyler, 2006). French and Raven's approach distinguishes between coercive power, reward power, legitimate power, expert power, referent power, and information power. The different bases of power are seen as independent implying that authorities cannot only hold one of the bases of power but several bases of power at the same time. Moreover, the different bases of power can be integrated into a two-dimensional structure (Raven, Schwarzwald, & Koslowsky, 1998): the six bases of power fall into the two independent categories of harsh and soft forms of power. To be consistent with the terminology in the context of the regulation of citizens' behavior (Turner, 2005), we use the term coercive power for harsh power and legitimate power for soft power. In the following, the terms coercive power and legitimate power refer to our conceptualization and not to French and Raven's (1959; Raven, 1965) terminology. Perceived coercive power originates from the pressure applied through either punishment or remuneration. Our concept incorporates the two harsh forms of social power bases, i.e., coercive power and reward power. Whereas coercive power is based on the expectations of the influenced party that non-cooperative behavior will be punished (e.g., through monetary penalties or imprisonment), reward power operates through the expectations of the influenced party that obeying the rules of the powerful party will be rewarded (e.g., through awards or gratuities). Our concept of coercive power is consequently based on incentivizing and compulsion. Individuals who do not obey the rules of the authorities will face monetary, physical, social, or psychological costs (e.g., being fined or not receiving a reward, being excluded from future transactions). Perceived legitimate power originates from legitimization, knowledge, skills, access to information, and identification with the powerful party, and comprises French and Raven's (1959) soft forms of power, namely legitimate power, expert power, information power, and referent power. Legitimate power operates through the accepted right to influence others by means of, for instance, agreed election rules, the norm of reciprocity (Gouldner, 1960), social responsibility, and equity norms (Berkowitz & Daniels, 1963). Expert power operates through the attribution of knowledge and skills that leads to the perception that the expert has a high capacity to lead (Raven, 1992, 1993). Information power is based on sharing of valued information (Raven, 1965, 1992, 1993). Referent power results from the dependent party's identification with the influencing party (Raven, 1992, 1993). Our concept of legitimate power is based on the fact that the legitimate authorities use information, charisma, legitimization, and expertise to convince taxpayers that it is the right course of action to cooperate. Hence, in contrast to the conceptualization of Becker (1968) and Tyler (2006), in the current conceptualization the qualities of power are multifaceted. Coercive power includes not just deterrence but also audits, punishment and rewards, and legitimate power comprises not just acceptance of the authorities but also the legal position, distribution of information, identification with the authorities, and their expertise. Importantly, in our conceptualization of power, coercive power and legitimate power are not seen as opposing entities but as independent factors (Hofmann, Gangl, Kirchler, & Stark, 2014). Authorities can wield coercive power without legitimate power, legitimate power without coercive power as well as they can wield both qualities of power at the same time (Hofmann et al. 2014). Accordingly, the tax authorities can or cannot be perceived as having the means to punish and reward taxpayers and can or cannot be perceived as having procedural measures to make it acceptable and easy for taxpayers to contribute.","The importance of trust in social systems is broadly recognized. Despite notable differences in approaching the phenomenon of trust, there is wide agreement on defining trust as the willingness of a party to take a risk (Lewis & Weigert, 1985a) and “to be vulnerable to the actions of another party based on the expectation that the other will perform a particular action important to the trustor, irrespective of the ability to monitor or control that other party” (Mayer, Davis, & Schnorrman, 1995, p. 712). Two independent qualities of trust are distinguished: trust based on cognitive-rational processes and trust based on automatic-affective processes (Castelfranchi & Falcone, 2010; Lewis & Weigert, 1985a; McAllister, 1995; Nooteboom, 2002; Tyler, 2003). We draw on Castelfranchi and Falcone's (2010) conceptualization of trust and differentiate between reason-based and implicit trust. Reason-based trust corresponds to concepts of calculative trust (Coleman, 1994; Fehr, 2009), rational trust (Ripperger, 1998), and knowledge-based trust (Lewicki & Bunker, 1996). Implicit trust corresponds to concepts of identification- based trust (Tyler, 2001), habitus trust (Misztal, 1996), social trust (Welch et al. 2005), and affective trust (Jones, 1996). Reason-based trust results from a deliberate (rational) decision grounded on four criteria: goal achievement, dependency, internal factors, and external factors (Castelfranchi & Falcone, 2010). First, the trustor evaluates whether the other party is pursuing a goal that is important to the trustor. Second, it is evaluated whether the trustor depends on the other party. Third, a positive evaluation of internal factors of the other party, i.e., competence, willingness, and harmlessness, is required. Fourth, the external factors in decision-making include the perception of opportunities and dangers. In this sense, reason-based trust corresponds to trust developed by a rational agent who trusts that there are good reasons to expect the other will forgo opportunistic goals (Coleman, 1994; Fehr, 2009; Mayer et al. 1995). Implicit trust is defined as an automatic, unintentional, and unconscious reaction to stimuli (Castelfranchi & Falcone, 2010). The automatic reaction originates from associative and conditioned learning processes and memory and is expected to emerge in situations in which shared social identities are activated (Castelfranchi & Falcone, 2010; Coulter & Coulter, 2002). Social categories or groups serve as stimuli which provoke the perception that certain social practices and norms can be relied on and that every person, organization, or authority that falls into this category can be trusted (Castelfranchi & Falcone, 2010; Lewis & Weigert, 1985b; Messick & Kramer, 2001). It can be expected that an authority perceived as belonging to the same category like the taxpayer will be evaluated positively and implicit trust should be higher when compared to trust in authorities perceived as belonging to another category (Tanis & Postmes, 2005). In addition to social identities also the cue that the tax authorities are an official institution might serve for some as a sign which activates automatic trust. Having repeatedly successfully interacted with public institutions leads to automaticity in the interaction (Verplanken, 2006) and a situation in which implicit and habitual trust in the institution prevails (Misztal, 1996). Other cues which activate implicit trust might be signs of warmth in contrast to hostility or cooperation in contrast of competition communicated through tax authorities' communication (websites, brochures, buildings; Williams & Bargh, 2008). To conclude, implicit trust occurs without the conscious recognition of reasons to trust and thus, without considering competence or intention of the official institution. The different conceptions of trust correspond to the two-process theories of cognition in which it is distinguished between system 1 and system 2 (Kahneman, 2003; Sherman, Gawronski, & Torpe, 2014). System 1 is working fast, effortless, associative and often is emotionally charged, governed by habit and difficult to control and modify. System 2 is based on slow, effortful, serial and deliberately controlled cognition, relatively flexible and potentially rule-governed (Evans, 2008; Kahneman, 2003). Whereas system 1 describes the functionality of implicit trust, system 2 explains the mechanism of reason- based trust. However, the two-process theories also assume that reason-based trust and implicit trust are related (Evans, 2008). Depending on the circumstances, the two qualities of trust might have parallel as well as sequential relationships (Evans, 2008). For instance, reason-based trust and implicit trust are operating together if taxpayers might implicitly trust the tax authorities because of cues such as a friendly voice on the tax line and at the same time might gain reasons to trust as the same person on the tax line also offers a competent advice. On the other hand, reason-based trust and implicit trust might operate independently, if taxpayers are cognitively too lazy to consider whether the tax authorities give reasons to trust, and rather just implicitly trust without questioning the tax authorities as official institution. Taxpayers also might in principle mistrust official institutions and hence, only trust the tax authorities, if they have proven evidence that the tax authorities act benevolently and competently. For a sequential relationship, research suggests that after taxpayers gained relevant experience based on deliberatively considering tax authorities' trustworthiness, reason-based trust enhance or even changes its quality to fast, and implicit trust (Evans, 2008; Sun, Slusarz, & Terry, 2005). Thus, in the long-run implicit trust develops with increasing reason-based trust that in the end becomes implicit trust.","Depending on the quality of power and the way power is exerted and perceived, trust in the powerful party can either be strengthened or weakened (e.g., Castelfranchi & Falcone, 2010; Choudhury, 2008; Korczynski, 2000; Kumlin & Rothstein, 2005). Also the quality of trust can affect the perception of authorities' power. In the SSF, Kirchler et al. (2008) conclude that tax authorities which enforce compliance through hostile and coercive measures run the risk of losing trust, whereas tax authorities perceived as legitimate may gain trust and the voluntarily cooperation of trustors. Also, the SSF proposes that if the authorities gain trust, they also enhance their legitimate power (Kirchler et al. 2008). In this vein, we assume two strong mechanisms which in general regulate the dynamics between power and trust: coercive power and implicit trust mutually decrease each other and that legitimate power and reason-based trust mutually increase each other. Additionally, we propose two second order relationships such as that coercive power and reason-based trust are related to each other via legitimate power and that legitimate power and implicit trust are related through reason-based trust. In the following, these assumptions are presented in detail (Fig. 2). Coercive power and implicit trust are mutually decreasing each other. If coercive power manifests by strict controls and fines, particularly if addressed at the individual, it provokes deliberate reasoning regarding possible gains and losses and the risk of non-compliance and therefore interjects and destroys implicit trust (Kirchler, 2007). Additionally, coercion may weaken affective and social bonds and interrupt habitual cooperation (Balliet, Mulder, & Van Lange, 2011; Castelfranchi & Falcone, 2010; Kramer, 1999; Nooteboom, 2002; Tenbrunsel & Messick, 1999). Coercive power damages implicit trust and social bonds because asymmetrically established control mechanisms indirectly convey the message that the person to which coercive power is addressed, is not trusted (Das & Teng, 1998; Nooteboom, 2002). As a reaction, implicit and automatic trust cannot emerge and instead coercive power is assumed to lead to reactance and deliberate and strategic reasoning (Balliet et al. 2011; Kirchler, 1999; Kirchler et al. 2008). However, implicit trust also reduces coercive power. People who trust implicitly base their automatic trust on shared norms, signaled values, and habits. Accordingly, audits and fines, which are expressions of coercive power, are not perceived as necessary (Cummings & Bromiley, 1996; Dekker, 2004; Yamagishi, 1988). Implicit trust activates social control mechanisms and relational governance (Dekker, 2004), and it fosters spontaneous, unreflected cooperation (Castelfranchi & Falcone, 2010). Willingness to spontaneously cooperate with another party reduces the complexity of the social world (Luhmann, 2000), because control is not necessary (Das & Teng, 1998; Inkpen & Currall, 2004). Hence, those who implicitly trust might not demand tax authorities to increase their coercion. Legitimate power and reason-based trust are mutually amplifying each other. Legitimate power and reason-based trust are strongly entwined and can be seen as the two sides of the same coin. Hence, their mutual fostering influence is not only theoretically but also empirically well established (Bijlsma-Frankema & Van de Bunt, 2002; Das & Teng, 1998; Malhotra & Murninghan, 2002; Mulder, Van Dijk, De Cremer, & Wilke, 2006). There are several reasons for taxpayers to trust; parties with legitimate power are perceived as competent to provide assistance and support (Bijlsma-Frankema & Van de Bunt, 2002; Castelfranchi & Falcone, 2010). Additionally, legitimate processes can provide a “track record” of the behavior of the parties involved and thereby build up a positive reputation (Das & Teng, 1998). Hence, legitimate power provides reasons to trust the tax authorities. At the same time, reason-based trust also increases legitimate power because reason-based trust both emerges from and leads to the recognition of the legitimacy of the authorities and the acceptance of the authorities (Castelfranchi & Falcone, 2010). If the perception of shared goals prevails, it is likely that the party is also accepted as the rightful authority (Inkpen & Currall, 2004). Reason-based trust permits and requires the powerful authorities to influence taxpayers' behavior (Cullen, Johnson, & Sakano, 1995; Das & Teng, 1998). Coercive power and reason-based trust are related through legitimate power. There are reasons why coercive power is destroying trust, for instance, if it is applied without competence and there are reasons why coercive power can strengthen trust, for instance, if coercive power is perceived as being used competently only against tax evaders. Hence, coercive power in combination with legitimate power has a relationship with reason-based trust. It is assumed that authorities which are perceived wielding coercive and legitimate power strengthen reason-based trust, whereas authorities perceived as wielding coercive power and not legitimate power reduce reason-based trust. Coercive power which is perceived to be used in a legitimate way, for instance, in a procedural fair manner to effectively inhibit rule-breaking behavior (Bachmann, Knights, & Sydow, 2001; Hofmann et al. 2014; Van Prooijen, Gallucci, & Toeset, 2008), gives good reasons to trust. Hence, coercive power combined with legitimate power is assumed to be perceived as targeted to evaders and as a safeguard of honest taxpayers which fosters reason-based trust. In contrast, high coercive power combined with low legitimate power destroys reason-based trust because the authorities are perceived to wield audits and fines in an incompetent and random way and even might be seen to prosecute honest taxpayers. If coercive power is low and legitimate power is high, reason-based trust also is high, as the wielding authorities are perceived to lead through legitimacy alone. If both, coercive power and legitimate power are perceived to be low also reason-based trust is low. The authorities are not legitimated and weak; hence, they are not seen as capable to guarantee a fair tax system. As the impact of coercive power on reason-based trust depends on legitimate power, coercive power which is wielded without highlighting its legitimacy likely reduces reason-based trust in the tax authorities. In general, coercive power and reason-based trust are related over the perception of legitimate power. Legitimate power and implicit trust are related through reason-based trust. Legitimate authorities increase first, reason-based trust and second, through establishing a stable system of functioning cooperation also implicit trust. Trust initially based on rational consideration transforms into implicit trust through routine. The more routine taxpayers can develop interacting with authorities perceived as competent the more reason-based trust gradually changes into implicit trust. The different qualities of power and trust are assumed to be perceived by taxpayers in a specific way. Legitimate power and reason-based trust are assumed to be positively related, whereas coercive power and implicit trust are negatively related. These dynamics are assumed to hold not only for vertical relationships in which tax authorities are seen to affect a specific taxpayer's behavior but also for horizontal relationships in which a taxpayer's behavior is affected by the observation on how the tax authorities treat other taxpayers not the taxpayer in question. The relationship of qualities of power and trust holds not only, if power is perceived to be wielded on a taxpayer, but also if it is wielded on other taxpayers. Coercive power perceived to be directed at the individual reduces implicit trust, however, also if it is perceived to be directed to other taxpayers it may fuel the impression that a considerable number of taxpayers tries to evade taxes, hence, that the social norm of tax honesty is low. Accordingly, coercive power indirectly conveys the message that the other taxpayers cannot be trusted to pay their fair share of taxes and therefore need to be enforced (Mulder et al. 2006). Also legitimate power addressed to other taxpayers has the same effect compared to when it is addressed on the individual. Addressing the individual with legitimate power already conveys a message about the other taxpayers. Legitimate power is applied competently which means it is addressed in general to all taxpayers which fosters reason- based trust. Coercive power and legitimate power addressed to other taxpayers will increase reason-based trust because it leads to the perception that the right people, those who try to evade, are controlled whereas the honest taxpayers are protected. To sum up, two main mechanisms are assumed to determine the relationship between power and trust, a negative relationship between coercive power and implicit trust and a positive relationship between legitimate power and reason-based trust. Additionally, it is assumed that coercive power is independent from reason-based trust and that legitimate power is positively related to implicit trust. Whereas coercive power impacts reason-based trust only due to its perceived legitimacy, legitimacy is assumed to steadily increase implicit trust by increasing reason-based trust. Based on these assumptions and the empirical evidence provided by Kogler et al. (2013) and Wahl et al. (2010) we propose that the manipulation of high power and high trust was in fact a manipulation of coercive power and legitimate power. Whereas the manipulation of high power and low trust was likely perceived as coercive power applied with low competence to guarantee a fair system, low power and high trust were likely perceived as legitimate power without the means to enforce compliance with the law. Hence, only high power and high trust, thus, high coercive power and high legitimate power were perceived as a competent safeguard of cooperation which induced the overall highest compliance rates. THE IMPACT OF THE DYNAMIC OF POWER AND TRUST ON TAX CLIMATE AND TAX COOPERATION ------------------------------------------------------------------------------- Based on the presented assumptions regarding the dynamics between power and trust we extend the SSF and distinguish on the aggregated level between three interaction climates (cf. Alm & Torgler, 2011): an antagonistic climate, a service climate and a confidence climate. This differentiation is based on conceptualizations from organizational research distinguishing between market or price mechanisms regulating social interactions, authorities, hierarchy-based or bureaucratic mechanisms of regulation and finally, community or trust mechanisms of managing social interactions (Adler, 2001; Bradach & Eccles, 1989; Haslam & Fiske, 1999; Ouchi, 1979). We also hypothesize that the different interaction climates lead, in the long run, on the individual level to corresponding forms of cooperation by taxpayers: enforced tax compliance, voluntary tax cooperation, and committed tax cooperation. Again, current assumptions are in line with earlier research on social influence for instance by Kelman (Kelman, 1961, 2006), who concluded that three psychological processes determine individual reactions to influence: compliance, identification, and internalization. Compliance with rules is based on incentives, identification with roles is grounded in reciprocity and modeling and internalization of values is based on value congruence or perceived continuity of the own self-concept (Kelman, 2006; Kelman & Hamilton, 1989). Hence, in the extended SSF, we conclude that the dynamics between power and trust are the preconditions of three cooperative climates, the antagonistic, the service, and the confidence climate, with corresponding qualities of motivations to cooperate, enforced compliance, voluntary cooperation and committed cooperation. The conceptualization of the interaction between authorities and individuals builds on the classical psychological insight, that authorities' actions create a specific social atmosphere, hence a cooperative climate which in turn provokes on the individual level specific corresponding habitual reactions of cooperation (Lewin, Lippitt, & White, 1939; Schneider, 2013). Fig. 3 summarizes our assumptions and shows that coercive power favors an antagonistic climate and enforced compliance, whereas legitimate power and reason-based trust are the antecedents of a service climate and voluntary cooperation. Implicit trust is the base of a confidence climate and committed cooperation. In the antagonistic climate coercive power prevails and a “cops and robbers” attitude is predominant with taxpayers and tax authorities working against each other (Kirchler et al. 2008). Tax authorities perceive taxpayers as “robbers” who try to evade and escape the tax authorities. In turn, taxpayers may feel prosecuted and harassed by the tax authorities (“cops”) and may feel the necessity to “hide”. The antagonistic climate is characterized by mistrust and resentment and leads to a vicious circle in which coercive power and mistrust mutually reinforce each other. Thus, compliance in such a climate needs to be enforced. Enforced compliance is characterized, for instance, by the feeling that tax authorities are interested in catching taxpayers evading, independent of whether the wrongdoing is intended or not (Kirchler et al. 2008). These assumptions received empirical support through experiments showing that high in contrast to low coercive power leads to a perceived antagonistic climate and enforced compliance of taxpayers (Hofmann et al. 2014). The thoughts underlying an antagonistic climate that taxpayers can only be forced to comply with the tax law by dint of controls and fines match the standard economic paradigm of tax behavior (Allingham & Sandmo, 1972). The disadvantage of an antagonistic climate is — besides costly audits — that taxpayers are likely to develop motives of opposition and reactance (Braithwaite, 2009; Kirchler, 2007) which cause instability in tax behavior and tax collection: when the tax authorities lose power, taxpayers lacking the intrinsic motivation to comply are expected to engage in evasion. The service climate bases on legitimate power and reason-based trust. It is characterized by a “service and client” attitude which means that taxpayers and tax authorities collaborate on the basis of well- defined rules and standards. Tax authorities perceive taxpayers as clients who expect and deserve professional, fair, and supportive services. Taxpayers reciprocate this attitude by contributing their tax share. Taxpayers who perceive the authorities as being supportive and competent are likely to cooperate voluntarily. Voluntary tax cooperation reflects the view of taxpayers that paying taxes is an accepted obligation as well as a necessity if the state is meant to provide public goods (Kirchler & Wahl, 2010; Wahl et al. 2010). These assumptions also received empirical support through experiments conducted with taxpayers showing that high in contrast to low legitimate power leads to a perceived service climate and voluntary cooperation (Hofmann et al. 2014). The advantage of the service climate lies in its stability — a single event of inappropriate services provided by the tax authorities will not lead to reduced taxpayers' cooperation, because the taxpayers themselves want the tax system to work smoothly. A disadvantage of a service climate may be the bureaucracy entailed in producing elaborate written rules as well as complex procedures to treat taxpayers fairly, which results in substantial administrative overheads (Ouchi, 1979). In a confidence climate implicit trust prevails. Taxpayers automatically trust the tax authorities and cooperate without thinking about it. Taxpayers pay their taxes because they perceive the tax authorities to work on the basis of shared norms and values or simply cooperate out of a habit. Tax authorities on the other hand, reinforce implicit trust by showing respect to the honest taxpayers (Feld & Frey, 2002). Tax authorities perceive themselves as working in the name of the taxpayers; they show empathy and feel obliged to offer support. Taxpayers perceive the tax authorities as working for the good of the community and reciprocate by contributing their share because they feel intrinsic motivation as members of the same community. For taxpayers, tax compliance is a personal and societally shared norm that is binding. Shared perceptions and values prevail and taxpayers are personally committed to the tax system. Committed cooperation is characterized by taxpayers' feelings that paying taxes is a customary thing to do and a moral obligation also followed by fellow citizens. Taxpayers feel committed to the tax system as a whole and actively engage to make the system work. The main advantage of a confidence climate is that taxpayers do not follow the letter of the law, but comply with the spirit of the law. Specific and complicated tax legislation is not needed because taxpayers follow moral standards instead of specific tax rules. According to Sloterdijk (2010), the main benefit of a confidence climate which induces that taxpayers feel committed to contribute their share, is that taxpayers are in a position of self- determination and generosity where they actively participate in a vital democracy and take responsibility for their society. Undoubtedly, a disadvantage of a confidence climate is its vulnerability to free-riders if tax authorities are perceived to avoid controls and punishment of tax evaders (Ouchi, 1979). Adler (2001) adds for the organizational context that such a confidence climate should be reflective and grounded in open dialogue among the interacting parties to avoid blind and traditionalistic loyalty.","As a main contribution of the present theoretical elaboration on the dynamic between power and trust, it can be derived how authorities such as the tax authorities can change and enhance cooperative climates and cooperation in social dilemmas. A negative dynamic between coercive power and implicit trust and a positive dynamic between legitimate power and reason-based trust explain how tax authorities can solve the social dilemma of taxpaying by either creating an antagonistic climate with enforced compliance, a service climate with voluntary cooperation, or a confidence climate with committed cooperation. In the following, the present paper concludes how a change from one climate to another climate can be accomplished, how the present assumptions can fuel empirical research, and why the dynamic between power and trust explains tax compliance and cooperation in social dilemmas in general. As a practical implication of the dynamic between power and trust, it is possible to demonstrate how a change from one climate to another emerges and why some interaction styles are more “slippery” than others. Countries differ in their interaction styles with taxpayers and their tax honesty (Alm & Torgler, 2006; Kogler et al. 2013). In some developing countries authorities lack power and trust, hence, a state of instability and in extreme cases a state of anarchy prevails with low levels of compliance, whereas on the other extreme, stable countries such as northern European ones exist displaying high levels of power and trust in the authorities that guarantee high tax compliance (Kogler et al. 2013). Also, on the individual level differences prevail between taxpayers (Braithwaite, 2003). Individuals within a country differ in their motivation to be honest and can be distinguished into taxpayers who perceive low power and hold low trust and hence, intentionally evade, perceive low legitimate power and fail to comply due to, for instance, complex tax laws and complex procedures or perceive high levels of trust and cooperate out of commitment to the community. We are convinced that tax authorities possess the measures to transform, on the aggregate level, a tax climate of distrust gradually into a climate of confidence. On the individual level, taxpayers' enforced compliance can be transformed into voluntary and committed cooperation. In line with the full range leadership model (Avolio & Bass, 1991; Kirkebride, 2006), in which the laissez- fair leadership develops over various stages into transactional leadership based on incentivizing behavior and further into transformational leadership based on a vision and shared values (Antonakis, Avolio, & Sivasubramaniam, 2003; Kirkebride, 2006), we also propose a possible transformation from one interaction climate into another. Under circumstances of low power and of low trust, tax compliance will be at a minimum, no matter whether it is on a country or individual level. In such a situation, coercive power can be a starting point for tax compliance. Laboratory experiments indicate that in a social dilemma situation with low levels of trust, authorities using coercive power are efficiently increasing compliance (Van Vugt & De Cremer, 1999). However, the psychological effectiveness of coercive power lies in its potential to scare, deter and enforce taxpayers through efficient and strict audits and severe fines. Accordingly, coercive power precludes the emergence of implicit trust. Instead a vicious circle of mistrust and consequently stronger coercive power is likely to develop between the tax authorities and the taxpayers resulting in an antagonistic climate and enforced motivation of tax compliance. Additionally, the antagonistic climate is instable or slippery as it depends on permanent exertion of coercive power. To change an antagonistic climate into a service climate, the measures of coercive power, such as controls and punishments, have to be combined with accepted, legitimate power. Once legitimate power is established, reason- based trust is likely to increase and, as a result, a service climate is established with voluntary tax cooperation. Tax authorities can improve their legitimacy by improving their services such as establishing professional and comprehensible tax procedures or web and telephone services in order to be perceived as motivated, competent and benevolent (Alm & Torgler, 2011). Moreover, the service climate is not depending on permanent enforcement, thus, it is relatively stable compared to the antagonistic climate. A service climate is theoretically based on legitimate power and reason-based trust. It can be assumed that a service climate changes into a confidence climate and that voluntary cooperation changes into committed cooperation over the course of time and due to stable positive experiences (Verplanken, 2006). Cooperation founded on reason-based trust which is initially based on careful consideration of one's own risks and other's intents becomes automatic with routine and repeated positive experiences (Castelfranchi & Falcone, 2010; Dekker, 2004; Misztal, 1996; Nooteboom, 2002). Repeatedly positive experiences lead to the implicit expectation that the other party respects agreed norms and practices. Accordingly, reason- based trust decreases in the longer run, while correspondingly, implicit trust increases over time (Castelfranchi & Falcone, 2010). To promote a confidence climate, tax authorities could, for instance, establish contracts of fair play and long-term relationships with committed taxpayers (Adler, 2001; Alm & Torgler, 2011; Ouchi, 1979). Establishing fair play with enterprises and the guarantee of mutual collaboration on the basis of mutual trust are existing examples (e.g., Schepers, 2010; see also http://www.nltaxinternational.com/index.php/taxadvice/10; retrieved April 10, 2012). It can be argued that relying solely on trust as in a confidence climate is far too optimistic in a social dilemma context since there always will be citizens tempted to engage in egoistic profit maximizing activities and hence, trusting authorities are not perceived as being able to enforce compliance. On the other hand, a climate of confidence is easily destabilized by the emergence of suspicion caused by power mechanisms (Kramer, 1999; Nooteboom, 2002). As a consequence, the confidence climate might be as instable and slippery as the antagonistic climate. Whereas the emergence of egoistic free-riders not threatened by coercive power might bring the confidence climate to a collapse, perceived power measures, depending on their quality and severity, can easily change the confidence climate into a service climate or an antagonistic climate if they trigger rational consideration of authorities' intentions or are perceived as hostile prosecution. The attempt to describe the prerequisites of voluntary and committed cooperation is a worthwhile approach solving social dilemmas in assisting the transformation from an antagonistic climate to a climate of suitable services and confidence. Thus, long lasting and reoccurring experiences with legitimated authorities might convince recalcitrant taxpayers whose compliance is enforced to become responsible and self-determined citizens committed to the moral obligation of voluntary contribution to the common good. Although the proposed model allows for several theoretical predictions and practical implications, it has some boundaries. It could be argued that legitimate power and reason-based trust represent similar concepts that are highly related. We see them as the two sides of the same coin. However, legitimate power is the perception of influence whereas reason-based trust is the decision to be vulnerable based on an evaluation of the influencing entity and its environment; therefore they are similar, but not identical. On the other hand, these well-established and related definitions of power and trust also highlight how, in fact, deeply connected power and trust are. It is expected that the different interaction climates in a real life setting will never appear as sharply delineated as in theory. Therefore, future research should investigate the prevalence and overlaps of the different interaction climates, and how these affect tax compliance. Additionally, future research should not only consider trust in the authorities, but also trust in fellow citizens and believes about their motivation to cooperate (Eek & Rothstein, 2005). For instance, taxpayers may hold the opinion that they cooperate voluntarily, whereas other taxpayers only cooperate because they are enforced, and still other taxpayers cooperate out of commitment. Empirical studies should investigate whether such believes about others influence the dynamic between perceived power and trust. The present paper provides a theoretical frame for future research. The presented conceptualizations of the dynamics between power and trust as well as subsequent tax climates, motivations to comply, and tax compliance serve as a theoretical basis for empirical studies. As an example, in countries differing in their overall tax climate, surveys among taxpayers could be conducted relating power perceptions and trust in authorities with compliance intentions and perceived tax climates, and motivations to comply. Experiments manipulating perceptions of cooperative climate and different qualities of power and trust with scenarios could be run to test the proposed relationships between the different qualities of power, trust, tax climates, motivations to comply, and tax compliance (Hofmann et al. 2014). To conclude, by integrating of a wide range of psychological theories on cognition (Evans, 2008; Kahneman, 2003), leadership (Avolio & Bass, 1991; French & Raven, 1959), social and organizational climate (Adler, 2001; Lewin et al. 1939), compliance motivation (Kelman, 2006), situational influences (Alm & Torgler, 2006; Antonakis et al. 2003) and personal differences (Braithwaite, 2003) the SSF is extended. Starting off with the dynamic between power and trust, the interaction between tax authorities and taxpayers is explained which results in specific cooperative climates and corresponding forms of individual motivations to cooperate. Therfore, we are confident that the assumptions about the dynamics between power and trust are not only useful to understand and regulate tax behavior but can be transferred into other contexts related to cooperation in social dilemmas, managed by an authority. The present paper demonstrates that understanding the dynamic between power and trust can fuel research on cooperation in social dilemmas and can be utilized by authorities to foster voluntary and committed cooperation to guarantee the provision of public goods."],["Personal values are considered as guiding principles in one's life. Much of previous research on values has consequently focused on its relations with variables that are considered positive, including subjective well-being, personality traits, or behavior (e.g. health-related). However, in this study (N = 366) the negative 'dark' side of values is examined. Specifically, the study investigated the relations between Schwartz' (1992) ten value types and four different clinical variables - anxiety, depression, stress, and schizotypy with its subdimensions, unusual experience, cognitive disorganization, introverted anhedonia, and impulsive nonconformity. Positive relations between achievement and depression and stress, and negative relations between anxiety and hedonism and stimulation were predicted and found. Multiple regressions revealed that the ten value types explained the most variance in impulsive nonconformity and the least variance in unusual experience. Overall, values were better in predicting more cognitive clinical variables (e.g., cognitive disorganization) whereas clinical constructs were better in predicted more affective values (e.g., hedonism). Implications of the findings for value research are discussed. --------------------------------------------------------------------------------","Personal values are usually considered as cognitive concepts or beliefs that transcend specific situations and guide behavior and its evaluation (Maio, 2010; Schwartz, 1992). There have been different attempts to conceptualize personal values on an individual level. Based on the value approach by Rokeach (1973), Schwartz and Bilsky (Schwartz, 1992; Schwartz & Bilsky, 1987) have developed a motivational circumplex model with 56 personal values which can be grouped into ten value types (Fig. 1): universalism (e.g. equality, protection for the welfare of all people and for nature), benevolence (e.g. helpfulness, preservation of the welfare of people), tradition (e.g. respect, humility), conformity (e.g. obedience, honoring parents), security (e.g. safety and social order), power (e.g. authority, dominance), achievement (e.g. personal success, ambition), hedonism (e.g. pleasure, enjoying life), stimulation (e.g. exciting life, varied life), and self- direction (e.g. independent thought, creativity). One important feature of Schwartz' circumplex model is its motivational continuum. Two adjacent value types are motivationally similar, that is positively correlated, orthogonal value types are unrelated according to the assumptions of Schwartz (1992), and opposing value types are negatively correlated. Crucially, this prediction holds true for relations between values and variables such as personality traits (e.g., Parks-Leduc, Feldman, & Bardi, 2014). Thus, when the value types are plotted according to their proposed order along the x-axis and the correlation coefficients on the y-axis, the correlational pattern resembles a sine wave. Numerous studies have focused on the relations between values and constructs of positive affectivity such as subjective well-being and satisfaction with life. For example, Haslam, Whelan, and Bastian (2009) found that many value types were closely connected to positive affect, but not with negative affect. This is in line with other research that found no relations between personal values and neuroticism (Parks-Leduc, Feldman, & Bardi, 2014; Roccas, Sagiv, Schwartz, & Knafo, 2002). However, other studies have found relations between values and negative affectivity, although the results were inconsistent. Jarden (2010) for instance found negative relations of self-direction, stimulation, and hedonism value types with depressed mood. In a Chinese sample, the burnout dimension exhaustion was found to be related to conformity, but another dimension, losing interest, was not (Jia, Rowlinson, Kvan, Lingard, & Yip, 2009). A study among native American adolescents revealed negative relations between depression and tradition/benevolence. However, power/materialism and security/hedonism values did not display any associations with depression (Mousseau, Scott, & Estes, 2013). In a Brazilian sample, relations between values and psychopathy were observed (Monteiro, 2014): pleasure, success, and power were found to be positively related to the psychopathological subdimensions boldness and disinhibition. Kajonius, Persson, and Jonason (2015) found positive relations between hedonism, achievement, and power with the dark triad, which consists of machiavellianism, narcissism, and psychopathy. Silfver et al. (2008) found positive correlations of universalism and benevolence with guilt proneness. The present study aims to extend the small empirical support that personal values are also related to clinical constructs, in our case anxiety, depression, and schizotypy. Because of the little empirical research on this topic, we first turn to a set of constructs that has a similar theoretical and empirical base as personal values: personality traits. Numerous studies have investigated the relations between many clinical variables and personality traits. Personal values and personality traits ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Both personal values and personality traits have their common roots in the language, which are encoded as linguistic descriptors for individual traits and behavior (lexical hypothesis). Therefore, several studies considered values and traits as different components of personality (e.g., Saroglou & Muñoz-García, 2008). Another approach emphasizes the biological–motivational basis of values and traits (McCrae & Costa, 2008). This is because both constructs motivate individual behavior. Indeed, a behavioral genetics study showed that personal values as well as personality traits share common genetic factors (Schermer, Vernon, Maio, & Jang, 2011). Further support for a close link between personal values and personality traits of the five-factor model is provided by a meta-analysis that has found consistent correlational patterns (Parks-Leduc et al., 2014). For example, stimulation, self-direction, and universalism correlated positively with openness to experience, whereas security, conformity, and tradition correlated negatively. Agreeableness was highly correlated with benevolence and negatively with power. However, the relations between values and other personality traits are not as consistent as they are for openness and agreeableness. Neuroticism or emotional stability, for instance, showed no substantial correlation to any of the ten value types. Personality traits and clinical constructs ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Numerous studies have investigated the relations between the Big-Five and clinical variables, including anxiety and schizotypy. For example, neuroticism correlated positively and extraversion and conscientiousness correlated negatively with anxiety. Agreeableness and openness were mainly unrelated to anxiety disorders, as a meta-analysis revealed (Kotov, Gamez, Schmidt, & Watson, 2010). Schizotypal traits can be considered as both mild personality features and as a predisposition toward schizophrenia. Schizotypal traits reflected aspects of positive symptoms (unusual experiences), negative symptoms (introverted anhedonia), and impulsive nonconformity and cognitive disorganization (Lenzenweger, 2015). Mason, Claridge, and Jackson (1995) have proposed a multidimensional model of schizotypy with aspects of positive- schizotypy (reflecting the positive symptomatology of schizophrenia), asocial-schizotypy (reflecting antisocial, impulsive and tough-minded behavior), disorganized-schizotypy (reflecting a difficulty with attention and social anxiety), and negative-schizotypy (reflecting the negative symptomatology of schizophrenia). Meta-analyses on the relations between personality traits (Big-Five) and schizotypal traits showed that positive symptoms are positively related to openness to new experience (Samuel & Widiger, 2008; Saulsman & Page, 2004). The present study ~~~~~~~~~~~~~~~~~ The aim of the present study is to examine the relations between value priorities and the four clinical constructs anxiety, depression, stress, and schizotypy with its 4 facets. Based on previous findings described above and the common base of values and personality traits, the following hypotheses were derived. Valuing achievement is defined by Schwartz (1992) as demonstrating competence. This can include (time) pressure and a lot of demanding work, which in turn can lead to stress and even depressive symptoms. Therefore, it was hypothesized that stress and depressive symptoms are positively related to achievement. On the other hand, given that volunteer work and well-being are positively associated (Thoits & Hewitt, 2001), benevolence and universalism should be negatively associated with stress and depressive symptoms. This is also in line with the motivational continuum of the quasi-circumplex model (Schwartz, 1992), which predicts opposing pattern of results for opposing value types (cf. Fig. 1). Anxiety was expected to be negatively related to hedonism and stimulation, as for both of these value types, Schwartz stated that courage and outgoingness are needed to at least some degree. Furthermore, a negative relation to self-direction was not expected because those values are more cognitive than stimulation and hedonism. Previous studies have found relations between conservatism and different types of anxiety (e.g., death anxiety and fear of threat and loss (Jost, Glaser, Kruglanski, & Sulloway, 2003). Therefore, a positive relation between anxiety and security, tradition, and conformity was expected. Positive relations between the schizotypal subdimension impulsive nonconformity with stimulation and hedonism, as well as negative ones with security, tradition, and conformity were expected, because impulsive nonconformity can be considered as an extreme form of openness. This prediction is in line with the above discussed finding that the personality trait openness is linked to impulsive nonconformity (Samuel & Widiger, 2008; Saulsman & Page, 2004). Introverted anhedonia is predicted to be negatively related to stimulation and hedonism. However, we do not expect positive relations between introverted anhedonia with tradition and conformity because we consider them as conceptually different constructs. As valuing security implies harmony and stability (Schwartz, 1992), a negative relation with the cognitive disorganization subdimension of schizotypy was hypothesized. Unusual experience was expected to be unrelated to all value types. Furthermore, in an exploratory step, we investigated in a series of multiple regressions, which clinical construct can be best predicted by all value types and, reversing the dependent and independent variables, which value type can be best predicted by all clinical constructs. This can give us greater insight into the nature of personal values as it can help to reveal which psychological states (i.e., clinical variables) are completely unrelated and provide evidence of the discriminant validity of values. We only expected, based on the rational given above that rather affective value types such as hedonism and stimulation will more strongly predict the clinical variables than more cognitive value types such as self-direction and universalism (cf. Schwartz, 1992), because the clinical variables used are mainly affective.","Participants were 366 students of various disciplines from an East German university (Mage = 21.72, SD = 3.38, range = 18–39, 236 females). Participants volunteered to participate and were not compensated.","The full Portrait Value Questionnaire (PVQ-40) was used (Schwartz et al., 2001) in its German translation (Schmidt, Bamberg, Davidov, Herrmann, & Schwartz, 2007) to assess the 10 value types of Schwartz' (1992) value model. Participants were given a short description of a person (i.e., ‘portrait’) and were asked to rate how similar they are to this person on a 6-point Likert scale ranging from 1 very similar to 6 very dissimilar. To facilitate interpretation, all value items have been recoded prior to the analyses reported below. The reliabilities (Table 1) are similar to the ones reported in the original validation paper of the PVQ (Schwartz et al., 2001). Following the suggestions of Schwartz (1992) and Schwartz et al. (2001), the 40 items of the PVQ were centered in order to control for individual scale use tendencies. The short form of the Depression, Anxiety and Stress Scale (DASS-21; Henry & Crawford, 2005) was used to measure depression, anxiety, and stress. Example items are “I couldn't seem to experience any positive feeling at all” for depression, “I felt scared without any good reason” for anxiety, and “I found it difficult to relax” for stress. Each of the three clinical personality constructs was measured with seven items on a 4-point Likert scale ranging from 0 never to 3 almost always. Finally schizotypy was measured with the Oxford–Liverpool Inventory of Feelings and Experiences (O-Life; Mason, Linney, & Claridge, 2005), a short scale consisting of 43 items consisting of the four factors: unusual experience, cognitive disorganization, introverted anhedonia, and impulsive nonconformity. Answers were given on a yes–no response scale (0 and 1). The scales described were part of a larger survey, unrelated to the present study. Procedure The questionnaire was completed within one large group session.","The data file is available at osf.io/32ja6. Descriptive statistics and correlations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ First, the overall zero-order-correlations were calculated. As can be seen in Table 2 most hypotheses were supported. As predicted, stress and depression were positively related to achievement values, but also negatively with hedonism. Anxiety was negatively related to stimulation and hedonism values. Contrary to our expectation, we did not find significant relations between anxiety and security, tradition, and conformity, although all three correlation coefficients were in the predicted direction. Combining all the conservation values (security, tradition, conformity) into one variable did not result in a significant correlation either r(364) = .08, p = .07. Impulsive nonconformity correlated positively with stimulation and hedonism values, but also, somewhat surprisingly, with achievement and power values. As predicted, we found negative correlations between introverted anhedonia and stimulation and hedonism, but also, unexpectedly, positive correlations with power, security, and tradition. The correlation between cognitive disorganization and security was negative, as predicted. Unusual experience did not correlate with any value type, except for universalism. The pattern of correlations remained the same after controlling for gender and age. Multiple regressions ~~~~~~~~~~~~~~~~~~~~ In the next step, we explored which clinical variables are best predicted by personal values. Impulsive nonconformity was best predicted by all ten value types combined while unusual experience was the least predicted clinical variable. None of the 10 value types reached statistical significance in any of the seven multiple regressions after controlling for the other nine value types. Finally, a series of regression was conducted to test which value types were better predicted by all seven clinical variables. As can be seen in Table 3, hedonism and tradition were best predicted by the seven clinical variables, whereas universalism was very weakly predicted. The subdimensions of schizotypy explained in a hierarchical regression additional variance, if entered in a second step (columns 5 and 6 in Table 3). For most value types, schizotypy explained substantial more variance than anxiety, depression, and stress. For example, the four schizotypy dimensions explained 14% out of the total 16% variance of stimulation. Overall, no multicollinearity was observed (all VIFs < 2.1).","The aim of the present study was to investigate the relations between personal values and clinical aspects of personality. It focused on anxiety, depression, stress, and four subdimensions of schizotypy. First, all variables with the exception of the schizotypal subdimension, unusual experience, were related in a theoretical meaningful way with personal values. Stimulation and hedonism were negatively related to anxiety. The findings are in line with the results of Monteiro (2014), who has found positive relations between boldness and openness values such as stimulation and hedonism. Because openness to experience as a personality trait was unrelated to anxiety and depression (Kotov et al., 2010), this indicates that openness values and traits, despite being related (Parks-Leduc et al., 2014), differ in predicting clinical constructs. Second, stimulation and hedonism correlated positively with the schizotypal subdimension impulsive nonconformity. This is interesting, because impulsive nonconformity may be considered as an undesirable construct, as it refers to impulsive, aggressive, and asocial aspects of psychosis based on the Eysenck dimension psychoticism (Eysenck & Eysenck, 1975) and the hypomania construct (elevation of mood, feeling of grandiosity, risk taking etc.). In other words, stimulation and hedonism are both positively and negatively related to undesirable constructs (impulsive nonconformity and the DAS-scales, respectively). On the other hand, it was argued that impulsive nonconformity can also be considered as a beneficial trait, because of its relations to creativity (Acar & Sen, 2013; Cohen, Mohr, Ettinger, Chan, & Park, 2015). Further, we assume that benevolence is negatively related to impulsive nonconformity, because altruistic values such as loyalty and honesty require some reliability on the person, which may be incompatible with impulsivity. Introverted anhedonia was as predicted negatively related to stimulation and hedonism, likely because anhedonia is a key symptom of major depression (Pizzagalli, 2014) and those two value types are negatively related to depression (see above). However, somewhat surprisingly, introverted anhedonia was positively related to power, security and tradition, and negatively to benevolence. The latter is consistent with previous research reporting a negative relation between negative schizotypy and interest in social contact (Kwapil, Brown, Silvia, Myin-Germeys, & Barrantes-Vidal, 2012). The finding that introverted anhedonia is positively related to power contradicts previous studies that suggest that power correlates with extraversion (Parks-Leduc et al., 2014). Further research is needed to resolve this contradiction. Disentangling both constructs, introverted anhedonia and extraversion, may very well be a promising approach. As assumed, cognitive disorganization was negatively related to security, but also with power and self-direction. This can indicate that at least some structure is required for power and self-direction. Overall, the correlational pattern of openness values such as hedonism and stimulation with schizotypy was similar to the one found between the personality trait openness and schizotypy (Samuel & Widiger, 2008; Saulsman & Page, 2004). It is of theoretical interest that our findings, although predicted and meaningful, did in general not follow the expected sinusoidal pattern (Schwartz, 1992). That is, if a clinical variable is positively related to one value type, it should also be positively related to adjacent value types, unrelated to orthogonal value types, and negatively related to opposing value types. For example, albeit stress was negatively related to self-direction, stimulation, and hedonism, it was positively related to achievement and again (non-significant) negative to power, violating the assumption of a motivational continuum. A similar violation can be found for the other clinical variables, with the exception of impulsive nonconformity, which follows the proposed sinusoidal pattern well. This indicates that variables can be related to Schwartz' values without following the proposed sinusoidal pattern (cf. Schwartz, 1992). Finally, a series of multiple regression analyses revealed that the subdimensions of schizotypy with the exception of unusual experience were better predicted by all of Schwartz' (1992) ten value types than anxiety, depression, and stress. This finding is interesting from a theoretical point of view because it indicates that personal values are more strongly associated with cognitive variables than affective ones, which is in line with predominant definitions of values as cognitive constructs (e.g., Maio, 2010). The schizotypy subdimensions (Mason et al., 2005) represent the cognitive aspects of experiences in comparison to the dimension of the DASS, which reflect more negative affect (Henry & Crawford, 2005). This assumption is further supported by the fact that anxiety, depression, and stress are stronger related to affective value types such as hedonism and stimulation compared to cognitive value types such as self-direction or universalism. Our findings show that personal values can be positively related to negative constructs. In other words, essential principles that are personally important can be both negatively as well as positively related to behavior, feelings, and affect that are generally considered as negative and unwanted. Given that achievement, hedonism, and stimulation can be considered as value types with a strong personal focus (Schwartz et al., 2012), which are promoted in individualistic countries such as Germany (Hofstede, Hofstede, & Minkov, 2010), the findings also reveal a potential ‘dark side’ of individualism. Individualistic culture emphasizes the individual autonomy more and a low power distance with the consequences of more norm transgressions. On the other hand, more conservative/collectivistic values such as tradition and conformity are mostly unrelated to the clinical variables used in the present study. This is somewhat contradictory to previous findings, stating that individualism is in general positively associated with well-being and negatively with social anxiety (Diener, Diener, & Diener, 1995). Therefore it would be interesting to investigate whether the same pattern of relations can be found in collectivistic societies. Our study has also some limitations. Just as previous studies investigating the relations between personal values and clinical variables (e.g., Jarden, 2010; Jia et al., 2009; Mousseau et al., 2013), a non-clinical sample was used in the present study. Clinical samples would reveal further interesting insights about the structure and priorities of personal values with regard to the claimed universality of both (Schwartz, 1992; Schwartz & Bardi, 2001). Moreover, future studies could investigate whether the relations between personal values and clinical variables are mediated by the Big-Five traits, as are the relations between values and well-being (Haslam, Whelan, & Bastian, 2009). In conclusion, the present study shows interesting relations between personal values and clinical variables. Values were better in predicting cognitive clinical variables (e.g., cognitive disorganization) and more affective values (e.g., hedonism) were better predicted by them. In a broader framework, personality traits, personal values, goals, and needs should be integrated in a general theory of the structure of motivation to understand the underlying processes among the similar constructs (Schwartz, 2011). McAdams (1995) has already shown that traits and values can be hierarchically ordered on different personality levels. In a recent study by McGabe and Fleeson (2016) the role of traits for motivational processes, goal attaining, was examined. They found that person differed from each other in traits because they pursued different goals. We would like to add, this finding may have occurred because they have different personal values."],["Status updates are one of the most popular features of Facebook, but few studies have examined the traits and motives that influence the topics that people choose to update about. In this study, 555 Facebook users completed measures of the Big Five, self-esteem, narcissism, motives for using Facebook, and frequency of updating about a range of topics. Results revealed that extraverts more frequently updated about their social activities and everyday life, which was motivated by their use of Facebook to communicate and connect with others. People high in openness were more likely to update about intellectual topics, consistent with their use of Facebook for sharing information. Participants who were low in self-esteem were more likely to update about romantic partners, whereas those who were high in conscientiousness were more likely to update about their children. Narcissists' use of Facebook for attention-seeking and validation explained their greater likelihood of updating about their accomplishments and their diet and exercise routine. Furthermore, narcissists' tendency to update about their accomplishments explained the greater number of likes and comments that they reported receiving to their updates. --------------------------------------------------------------------------------","Why do some people write Facebook status updates that describe amusing personal anecdotes, whereas others write updates that declare love to a significant other, express political opinions, or recount the details of last night’s dinner? Since the inception of Facebook in 2004, status updates have been one of its most preferred features (Ryan & Xenos, 2011). Status updates allow users to share their thoughts, feelings, and activities with friends, who have the opportunity to “like” and comment in return. In spite of the central role of status updates in Facebook use, few studies have examined the predictors of the topics that people choose to write about in their updates. The current study took a step in this direction by examining the personality traits associated with the frequency of updating about five broad topics identified through a factor analytic approach: social activities and everyday life, intellectual pursuits, accomplishments, diet/exercise, and significant relationships. We also examined whether these associations were mediated by some of the motives for using Facebook identified in the literature (e.g., Bazarova & Choi, 2014; Seidman, 2013): need for validation (i.e., seeking attention and acceptance), self- expression (i.e., disclosing personal opinions, stories, and complaints), communication (i.e., corresponding and connecting), and sharing impersonal information (e.g., current events). A secondary purpose of this study was to examine whether people who update more frequently about certain topics receive greater numbers of “likes” and comments to their updates. Those who do may experience the benefits of social inclusion, whereas those who do not might experience a lower sense of belonging, self-esteem, and meaningful existence (Tobin, Vanman, Verreynne, & Saeri, 2015). Our results may therefore shed light on the status update topics that put Facebook users at risk of online ostracism. Below we review literature on personality traits and motives that are often linked with Facebook use. The Big Five ~~~~~~~~~~~~ According to the “Big Five” model of personality, individuals vary in terms of extraversion, neuroticism, openness to experience, agreeableness, and conscientiousness (Costa & McCrae, 1992). People who are extraverted are gregarious, talkative, and cheerful. They tend to use Facebook as a tool to communicate and socialize (Seidman, 2013), as reflected in their more frequent use of Facebook (Gosling, Augustine, Vazire, Holtzmann, & Gaddis, 2011), greater number of Facebook friends (Amichai-Hamburger & Vinitzky, 2010), and preference for features of Facebook that allow for active social contribution, such as status updates (Ryan & Xenos, 2011). We therefore predicted that extraversion would be positively associated with updating about social activities, and that this association would be mediated by extraverts’ use of Facebook for communication (Hypothesis 1). Neuroticism is characterized by anxiety and sensitivity to threat. Neurotic individuals may use Facebook to seek the attention and social support that may be missing from their lives offline (Ross et al., 2009). Accordingly, neuroticism is positively associated with frequency of social media use (Correa, Hinsley, & de Zuniga, 2010), the use of Facebook for social purposes (Hughes, Rowe, Batey, & Lee, 2012), and engaging in emotional disclosure on Facebook, such as venting about personal dramas (Seidman, 2013). Their willingness to disclose about personal topics led us to predict that neuroticism would be positively associated with updating about close relationships (romantic partners and/or children), and that the selection of these topics would be motivated by their use of Facebook for validation and self-expression (Hypothesis 2). People who are high in openness tend to be creative, intellectual, and curious. Openness is positively associated with frequency of social media use (Correa et al., 2010), and with using Facebook for finding and disseminating information, but not for socializing (Hughes et al., 2012). We therefore predicted that openness would be positively associated with updating about intellectual topics, and that this association would be mediated by the use of Facebook for sharing information (Hypothesis 3). People who are high in agreeableness tend to be cooperative, helpful, and interpersonally successful. Agreeableness is positively associated with posting on Facebook to communicate and connect with others and negatively associated with posting to seek attention (Seidman, 2013) or to badmouth others (Stoughton, Thompson, & Meade, 2013). The interpersonal focus of agreeable people and their use of Facebook for communication may inspire more frequent updates about their social activities and significant relationships (Hypothesis 4). Conscientiousness describes people who are organized, responsible, and hard-working. They tend to use Facebook less frequently than people who are lower in conscientiousness (Gosling et al., 2011), but when they do use it, conscientious individuals are diligent and discreet: they have more Facebook friends (Amichai-Hamburger & Vinitzky, 2010), they avoid badmouthing people (Stoughton et al., 2013), and they are less likely to post on Facebook to seek attention or acceptance (Seidman, 2013). Thus, we predicted that conscientiousness would be positively associated with updating about inoffensive, “safe” topics (i.e., social activities and everyday life), which would be mediated by the lower tendency of using Facebook for validation (Hypothesis 5). Self-esteem ~~~~~~~~~~~ People with low self-esteem are more likely to see the advantages of self-disclosing on Facebook rather than in person, but because their status updates tend to express more negative and less positive affect, they tend to be perceived as less likeable (Forest & Wood, 2012). Furthermore, anxiously-attached individuals – who tend to have low self- esteem (Campbell & Marshall, 2011) – post more often about their romantic relationship to boost their self-worth and to refute others’ impressions that their relationship is poor (Emery, Muise, Dix, & Le, 2014). We therefore hypothesized that self-esteem would be negatively associated with updating about a romantic partner, and that this association would be mediated by the use of Facebook for validation (Hypothesis 6). Narcissism ~~~~~~~~~~ Narcissistic individuals tend to be self-aggrandizing, vain, and exhibitionistic (Raskin & Terry, 1988). They seek attention and admiration by boasting about their accomplishments (Buss & Chiodo, 1991) and take particular care of their physical appearance (Vazire, Naumann, Rentfrow, & Gosling, 2008). This suggests that their status updates will more frequently reference their achievements and their diet and exercise routine (Hypothesis 7). Moreover, the choice of these topics may be motivated by the use of status updates to gain validation for inflated self-views, consistent with the positive association of narcissism with the frequency of updating one’s status (Carpenter, 2012), posting more self-promoting content (Mehdizadeh, 2010), and seeking to attract admiring friends to one’s Facebook profile (Davenport, Bergman, Bergman, & Fearrington, 2014). Response to status updates ~~~~~~~~~~~~~~~~~~~~~~~~~~ We examined whether people receive differential numbers of likes and comments to their updates depending on their personality traits and frequency of writing about various topics. People with lower self-esteem tend to receive fewer likes and comments because their status updates express more negative affect (Forest & Wood, 2012). We tested the possibility that they may also receive fewer likes and comments because they are more likely to update about their romantic partner (Hypothesis 8); indeed, people who write updates that are high in relationship disclosure are perceived as less likeable (Emery, Muise, Alpert, & Le, 2015). The associations of the Big Five traits, narcissism, and the other status update topics with the number of likes and comments received were examined on an exploratory basis to shed light on who may be at risk of receiving less social reward on Facebook, and whether it is because they express unpopular topics in their updates.","Data was collected from 555 Facebook users currently residing in the United States (59% female; Mage = 30.90, SDage = 9.19). Sixty-five percent of participants were currently involved in a romantic relationship, and 34% had at least one child. Fifty-seven percent checked Facebook on a daily basis, and spent an average of 107.95 min per day actively using it (SD = 121.41). Ninety percent of participants were recruited through Amazon’s Mechanical Turk and paid $1.00 in compensation; the rest were recruited through web forums for online psychology studies, and received no compensation.","Participants completed an online survey consisting of demographic questions and the following measures. Cronbach’s alpha coefficients are reported in Table 1. Big Five personality traits The 35-item Berkeley Personality Profile (Harary & Donahue, 1994) measures extroversion, neuroticism, openness, agreeableness, and conscientiousness with 7 items each (1 = Strongly disagree, 5 = Strongly agree). Self-esteem The 10-item Rosenberg Self-Esteem Scale (Rosenberg, 1965) measures self-esteem with items such as “I feel that I have a number of good qualities” (1 = Strongly disagree, 5 = Strongly agree). Narcissism The 13-item version of the Narcissistic Personality Inventory (NPI-13; Gentile et al., 2013) is derived from the original NPI-40 (Raskin & Terry, 1988) and measures three components of trait narcissism: need for leadership/authority, grandiose exhibitionism, and entitlement/exploitativeness. Items are rated on a forced- choice basis, such that one choice represents greater narcissism and the other less. Higher scores indicate greater narcissism. Facebook use Participants reported their number of Facebook friends, how many days of the week they check Facebook (0–7 days), how much time they spend actively using it on days they check it, and how frequently they update their Facebook status (1 = Never, 9 = 7–10 times a day). Topics of status updates Participants indicated how frequently they write about 20 topics in their Facebook status updates (i.e., verbal descriptions of their status excluding photos, videos, or emoticons). These topics were generated by the authors through laboratory group discussions. Responses were rated on a 5-point Likert scale ranging from 1 (Never) to 5 (Very often). To extract common themes across topics, we conducted principal axis factoring with promax rotation. This yielded four factors with eigenvalues greater than 1 that together accounted for 57% of the total variance. Five topics loaded on the first factor, which reflected social activities and everyday life (my social activities, something funny that happened to me, my everyday activities, my pets, sporting events). Four topics loaded on the second factor, which reflected intellectual themes (my views on politics, current events, research/science, my own creative output – e.g., art, writing, research). Three topics loaded on the third factor, which reflected achievement orientation (achieving my goals, my accomplishments, work or school). Two topics loaded on the fourth factor, which reflected diet/exercise (my exercise routine, my diet). Several topics did not meet Tabachnik and Fidell’s (2007) criteria that items must have a minimal loading of .32 on a single factor: three items (my children, my religious beliefs, and quotations or song lyrics) were below this threshold, and two items cross-loaded (my travels, my views on TV show, movies, or music). A final topic (my relationship with my current romantic partner) was not included in the factor analysis because it was only completed by participants currently involved in a relationship. Of the topics that did not load onto one of the four factors, we only further analyzed the frequency of updating about children and romantic partners as single variables because of our hypotheses regarding the associations of personality traits with updating about significant relationships. We also asked participants who they shared each status update topic with (no one, the public, friends only, close friends only), but because there was little variation across topics in these privacy settings, we did not examine this variable further. Motives for using Facebook We measured four motives for using Facebook by adapting items from a variety of sources (e.g., Hughes et al., 2012; Seidman, 2013) so that each began with “I use Facebook to…”. Use of Facebook for validation was measured with seven items that tapped attention-seeking (e.g., “I use Facebook to show off”) and need to feel accepted and included (e.g., “I use Facebook to feel loved”). Five items measured use of Facebook for self-expression (e.g., “I use Facebook to express my identity/opinions”). Three items measured use of Facebook to communicate (e.g., “I use Facebook to communicate with people I often see”), and eight items assessed use of Facebook to find and disseminate information (e.g., “I use Facebook to stay informed”). Participants indicated their agreement with these statements using a 1–7 Likert scale anchored with Strongly disagree (1) and Strongly agree (7). Likes and comments Participants indicated how many likes and comments, on average, they tend to receive when they post a typical Facebook status update.","Table 1 reports the descriptive statistics and Pearson’s correlations. Table 2 reports the results of regression analyses that examined the predictors of updating about each of the six topics (criterion variables), the four motives for using Facebook (mediating variables), and the number of likes and comments received to a typical update (criterion variable). Predictors included several control variables (frequency of updating one’s status, number of Facebook friends, sex, age) and the traits of interest (Big Five traits, self-esteem, narcissism). We conducted bootstrap tests of multiple mediation using Preacher and Hayes’s (2008) SPSS script to assess whether the motives for using Facebook mediated the associations of the personality traits with updating about certain topics. In these tests, the control variables and other personality traits were entered as covariates, and the four motives for using Facebook were entered as multiple mediators. Predictors of status update topics and motives for using Facebook ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 2 reveals support for Hypothesis 1: extraversion was positively associated with updating more frequently about social activities and everyday life, and with using Facebook to communicate. A further regression analysis showed that the use of Facebook to communicate predicted the frequency of updating about social activities and everyday life over and above the control variables and other personality traits (b = .25, p < .0001). Examination of the 95% bias-corrected confidence intervals (CI) from 1000 bootstrap samples revealed that the positive association of extraversion with updating about social activities and everyday life was mediated by the use of Facebook to communicate (b = .03, p = .05 (CI: .003–.05)). These results further confirm that extraverts use Facebook, and specifically status updates, as a tool for social engagement (Ryan & Xenos, 2011; Seidman, 2013). Hypothesis 2 was only partially supported: neuroticism was not associated with updating about any of the six topics or with using Facebook for self-expression, but it was associated with using Facebook for validation. Indeed, neurotic individuals may use Facebook to seek the attention and support that they lack offline (Ross et al., 2009). Consistent with Hypothesis 3, openness was positively associated with updating about intellectual topics, and with using Facebook for information. A further regression analysis showed that the use of Facebook for information and for self-expression predicted the frequency of updating about intellectual topics over and above the control variables and traits (b = .34, p < .0001 and b = .22, p < .001, respectively). The bootstrap test revealed that the positive association of openness with updating about intellectual topics was indeed mediated by the use of Facebook for information (b = .03, p < .01 (CI: .007–.05)). People high in openness, then, may write updates about current events, research, or their political views for the purpose of sharing impersonal information rather than for socializing, consistent with the findings of Hughes et al. (2012). There was no support for Hypothesis 4 – agreeableness was not associated with updating more frequently about social activities, significant relationships, or with using Facebook to communicate. Contrary to Hypothesis 5, conscientiousness was not associated with updating about “safe” topics such as social activities and everyday life; rather, it was associated with writing more frequent updates about one’s children. Furthermore, conscientiousness was not negatively associated with using Facebook for validation, but it was positively associated with using Facebook to share information and to communicate. The latter use predicted the frequency of updating about one’s children over and above the control variables and personality traits (b = .38, p = .01), but it did not significantly mediate the association of conscientiousness with updating about children. Thus, conscientious individuals may update about their children for purposes other than communicating with their friends. Perhaps such updates reflect an indirect form of competitive parenting. Consistent with Hypothesis 6, people who were lower in self-esteem more frequently updated about their current romantic partner, but they were more likely to use Facebook for self- expression rather than for validation. That the frequency of updating about one’s romantic partner was predicted not by the use of Facebook for self-expression but rather by communication (b = .24, p = .01) suggests that people with low self-esteem may have other motives for posting updates about their romantic partner. Considering that people with low self-esteem tend to be more chronically fearful of losing their romantic partner (Murray, Gomillian, Holmes, & Harris, 2015), and that people are more likely to post relationship- relevant information on Facebook on days when they feel insecure (Emery et al., 2014), it is reasonable to surmise that people with low self-esteem update about their partner as a way of laying claim to their relationship when it feels threatened. In line with Hypothesis 7, narcissism was positively associated with updating about achievements and with using Facebook for validation. Moreover, the use of Facebook for validation and for communication predicted the frequency of updating about achievements over and above the control variables and traits (b = .14, p = .02 and b = .13, p = .04, respectively). The association of narcissism with updating about achievements was significantly mediated by the use of Facebook for validation (b = .04, p = .05 (CI: .006–.07)), consistent with narcissists’ tendency to boast in order to gain attention (Buss & Chiodo, 1991). Also consistent with Hypothesis 7, narcissism was positively associated with updating about diet/exercise, but the use of Facebook for self-expression rather than validation was positively associated with updating about diet/exercise over and above the control variables and traits (b = .24, p < .01). Self-expression mediated the association of narcissism with updating about diet/exercise (b = .03, p = .03 (CI: .003–.04)), suggesting that narcissists may broadcast their diet and exercise routine to express the personal importance they place on physical appearance (Vazire et al., 2008). Predictors of likes and comments received ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As seen in Table 2, there was no support for Hypothesis 8: narcissism rather than self- esteem was associated with receiving a greater number of likes and comments to one’s updates. We then assessed whether the four topics common to the entire sample – social activities and everyday life, intellectual pursuits, achievements, and diet/exercise – predicted the number of likes and comments typically received to an update over and above the control variables and traits. Updating about social activities and everyday life was positively associated with the number of likes and comments received (b = .13, p = .05), as was achievements (b = .16, p = .01), whereas updating about intellectual topics was negatively associated (b = −.13, p = .04). Two additional regression models added the frequency of updating about one’s romantic partner or one’s children as predictors for participants who had a relationship partner or children. Only the frequency of updating about one’s children significantly predicted likes/comments (b = .23, p = .02). Bootstrap mediation revealed that the tendency for narcissists to report receiving more likes and comments was mediated by their higher frequency of updating about their achievements (b = .06, p < .01 (CI: .01–.18)). Thus, narcissists’ publicizing of their achievements appeared to be positively reinforced by the attention and validation they crave. Limitations and future directions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The main limitation of this study is that it was based on participants’ self-reported Facebook behavior. Narcissists, in particular, may not accurately report the number of likes and comments they receive to updates. More objective and precise estimates can be obtained in future research by coding participants’ actual status updates for topic themes and recording the number of likes and comments received to each topic. Another avenue for future research is to obtain direct evaluations of particular status update topics and of the likeability of people who update about these topics. That updating about social activities, achievements, and children was positively associated with Facebook attention, and updating about intellectual topics negatively associated, suggests that the former topics might be evaluated more positively than the latter. Yet these associations are at best a proxy for the likeability of these topics and of the individuals who write them. Considering that objective raters can accurately discern whether a person is narcissistic by looking at their Facebook page (Buffardi & Campbell, 2008), people may be correctly perceived as narcissistic if they more frequently update about their achievements, diet, and exercise. Furthermore, people may like and comment on a friend’s achievement-related updates to show support, but may secretly dislike such displays of hubris. The closeness of the friendship is therefore likely to influence responses to updates: close friends may “like” a friend’s update, even if they do not actually like it, whereas acquaintances might not only ignore such updates, but eventually unfriend the perpetrator of unlikeable status updates.","Taken together, these results help to explain why some Facebook friends write status updates about the party they went to on the weekend whereas others write about a book they just read or about their job promotion. It is important to understand why people write about certain topics on Facebook insofar as the response they receive may be socially rewarding or exclusionary. Greater awareness of how one’s status updates might be perceived by friends could help people to avoid topics that annoy more than they entertain."],["Using data from the Berlin diary study (N = 1223), we examined associations between the General Factor of Personality (GFP) and daily social experiences, self-esteem, and mood (positive and negative affect). As predicted, high-(vs. low) GFP individuals reported fewer daily interpersonal conflicts, better relationship quality, and better impressions on others. Also, relationship quality and daily impressions both mediated the relation between the GFP and mood and self-esteem. Multilevel analyses showed that, compared to low-GFP participants, high-GFP participants seemed less disturbed when experiencing conflict. In sum, the results were in line with the notion of the GFP as social effectiveness, with important consequences for people's daily social life and well-being. --------------------------------------------------------------------------------","In the personality literature, several studies suggest the existence of a General Factor of Personality or GFP (Figueredo, Vásquez, Brumbach & Schneider, 2004) which emerges due to the intercorrelations among more specific personality dimensions, such as the well- known Big Five. The GFP constitutes the socially desirable ends of those dimensions and has now been extensively replicated (e.g., Musek, 2007; Van der Linden, Te Nijenhuis & Bakker, 2010a). In terms of the Big Five, high-GFP individuals can be described as relatively open-minded, diligent, sociable, friendly and emotionally stable. Moreover, the GFP has shown criterion-related validity and is associated with various important life outcomes, such as job performance and leadership (Van der Linden et al., 2017). Despite such consistent findings, however, diverging scientific views exist on the interpretation of the GFP. One view is that the GFP represents social effectiveness (see Van der Linden, Dunkel & Petrides, 2016 for a review), implying the knowledge, abilities, and the motivation to generally behave in socially desirable ways. In contrast are views that the GFP merely represents a methodological artefact, due to, for example, socially desirable response bias (e.g., Schermer & Holden, 2019), common method variance (e.g., Chang, Connelly & Geeza, 2012), or other statistical artefacts (Ashton, Lee, Goldberg & de Vries, 2009; Revelle & Wilt, 2013). The different arguments for the substantive versus artefact views of the GFP have been discussed extensively in several review articles (Irwing, 2013; Revelle & Wilt, 2013; Van der Linden et al., 2016), and will therefore not be repeated here. The main point, however, is that there appears to be evidence for each of the different views. This is not surprising, given that the different explanations of the GFP need not necessarily be mutually exclusive (Davies, Connelly, Ones & Birkland, 2015; Dunkel, Van der Linden, Brown & Mathes, 2016). Here, using a comprehensive diary study, we aim to contribute to the literature by further testing the nature of the GFP. First, we test the GFP as a social effectiveness factor through its relations with daily social experiences. Second, we test to what extent the well-documented relation between the GFP and well-being and mood (Erdle & Rushton, 2011; Musek, 2007) is mediated by the presumed effective daily social experiences. And third, we test how the GFP moderates the relation between daily social experiences and daily well-being and mood. The GFP and daily social experiences ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Based on social effectiveness perspective, it can be expected that high-GFP individuals should, on average, also display higher effectiveness in their daily social interactions. Although the GFP has indeed been linked to various positive social outcomes such as peer- rated likeability and popularity (Van der Linden, Scholte, Cillessen, Te Nijenhuis & Segers, 2010b), to our knowledge there are no studies that have tested the relation between the GFP and social interactions using a diary study. The social effectiveness hypothesis implies that high-GFP individuals should be able to navigate more easily through social encounters and would consequently report higher levels of relationship quality, lower levels of interpersonal conflict, and would leave better impressions on others. Our first hypothesis thus states: H1. The GFP is negatively associated with daily (a) interpersonal conflict, and positively associated with daily (b) relationship quality, and (c) the impressions on others. Mediation of the GFP–well-being/mood relation by daily social experiences ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Socially effective behavior on a daily basis may partly explain why, high-GFP individuals often report enhanced self-esteem and mood (e.g., Musek, 2007). Generally, on days when people feel socially included, they also tend to experience higher levels of well-being than on days when they feel more socially isolated. This is known as the Sociometer theory (Leary, Tambor, Terdal & Downs, 1995), which can link the GFP to both self-esteem and indicators of social inclusion (e.g., relationship quality): higher GFP levels may be associated with higher levels of social inclusion, which in turn should result in higher levels of self-esteem and mood (e.g., Diener, 1984; Gable, Reis & Elliot, 2000). H2. The positive relations between the GFP, and self-esteem/positive affect and the negative relation with negative affect are, at least partially, mediated by (a) less daily interpersonal conflict, (b) better daily relationship quality, and (c) the enhanced daily impressions on others. Daily social experiences, well-being and mood: GFP moderation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Previous research has shown that personality traits can influence one's reactivity to daily social events in terms well-being and mood (e.g., Bolger & Schilling, 1991; Denissen & Penke, 2008). A similar role for the GFP can be expected: Because high-GFP individuals are more adapted to their social environment and have higher self-esteem (Musek, 2007), they can be expected to show less fluctuations in mood caused by social events. Specifically, even though higher GFP levels may imply less negative interpersonal events, obviously, sometimes negative social events, such as conflicts, will occur. Yet, when they do, part of the presumed social effectiveness may consist of the ability to adequately react to such negative events (Hengartner, Van der Linden, Bohleber & Wyl, 2017). For example, a higher GFP level may allow one to choose a more appropriate reaction to a conflict, thereby resolving it or preventing escalation. This notion fits with the meta- analytic finding that the GFP highly overlaps with emotional intelligence (Van der Linden et al., 2017). Therefore, the following hypothesis can be formulated: H3. The relations between (a) daily interpersonal conflict (b) daily levels of relationship quality, (c) daily impressions on others, and daily levels of self-esteem, positive affect and negative affect, are moderated by the GFP such that the relations are stronger for those with lower (compared to higher) GFP scores. The present study: using diary data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The majority of previous GFP studies used cross-sectional designs. Although informative, such designs, however, are limited because they provide a snapshot of ongoing psychological states and processes. In addition, they rely on people's imperfect ability to correctly recollect events or behaviors, which can lead to biases and inaccuracies. Accordingly, scholars have argued for the use of diary methods (Bolger, Davis & Rafaeli, 2003) assessing events and processes as they are naturally occurring, thereby increasing the ecological validity. Moreover, diary methods are assumed to be less susceptible to socially desirable response bias than cross-sectional designs (Barta, Tennen & Litt, 2013). This is especially relevant in light of the interpretation of the GFP as purely artefactual. Intuitively, it may be equally possible to over-report desirable events or traits on a daily basis as in a single measurement. However, daily reports are often found to be more accurate than single, one-time measurements (e.g., Presser & Stinson, 1998). Considering the above, our hypotheses are best tested with daily level data. To this end, we use data from the Berlin Diary Study by Denissen and colleagues (2005–2008), one of the largest diary studies in the world.","Data files, analysis scripts, and supplemental analyses can be accessed at https://osf.io/kywdf/.","The Berlin Diary Study (2005–2008) consisted of multiple phases, starting with a general questionnaire, including personality. Participants listed the friend and family member with whom they had most contact with, and their partner (if present). Then, for 30 days, participants filled out a daily questionnaire including randomly presented questions on daily well-being and daily interactions with the two or three identified others in the previous phase. For additional information on the study design we refer to Denissen and Penke (2008) and Denissen, Penke, Schmitt and Van Aken (2008), who previously used (parts of) these data. We decided to include participants who had completed at least 7 diary entries in order to minimize the influence of idiosyncratic days and assure participants’ commitment (Bolger et al., 2003). The final sample therefore included 1223 German participants (1055 women, 86%), with an average number of 19.28 (SD =6.81) daily reports. The average age was 29.47 (SD =10.49). Most people were either single (39%) or in a steady relationship (40%), without children (79% of the total sample). About 50% of the sample was relatively highly educated. Personality/GFP The Big Five Inventory (BFI; John & Srivastava, 1999) was used to measure Openness (O), Conscientiousness (C), Extraversion (E), Agreeableness (A), and Neuroticism (N), and to extract a GFP. Sample coefficient alpha's ranged from 0.72 to 0.90 (see Table 1). The BFI uses a 5-point Likert-scale format. Principal axis factoring was used to extract the GFP from the Big Five. The first unrotated factor explained 26.10% of the Big Five variance. The GFP factor loadings of O, C, E, A, and N were 0.36, 0.42, 0.66, 0.47, and −0.58 respectively (see Supplementary Materials for the convergence across different extraction methods). Daily social experiences Relationship quality. On a five-point scale, participants rated their feelings of enjoyment, interest, intimacy, power, important, calm, safe, wanted, and, respected, in the interactions with the identified persons. An overall index of relationship quality was created by averaging over all indicators across the identified others. Interpersonal conflict. Participants were asked whether they experienced (0 =not present, 1 =present) a conflict with the identified others on financial resources, communication, activities, life plans, encouragement, opinions, third persons, and “other topics”. Scores were summed over each day and the identified others: a zero indicated no conflict on that day. Because of the variable's skewedness, a dichotomized version with 0 indicating no conflict and 1 indicating any conflict was also created. Impressions on others. A subsample of the participants (N = 970) indicated (on a 7-point scale) the impressions they made on others during that day on eight different dimensions (competence, civility, ethical, artistic, sympathetic, orderly, psychical attractiveness, and tolerant). A total (mean) impression on others score was calculated.","To test H1 and H2, we aggregated the daily social and well-being reports. In the mediation analyses, due to the large sample size, we focus on the ratio (i.e., the effect size) of the standardized indirect effect to the total effect, rather than on significance levels. To test H3, we used multilevel regression analyses, as the data follow a hierarchical structure with days (Level 1) nested in individuals (Level 2). Multilevel analysis or hierarchical linear modeling (HLM) provides more accurate parameter estimates and significance tests than comparable ordinary least squares regression techniques by accounting for variance at each analysis level. In the present study, the intraclass correlations ranged between 0.32 and 0.54, indicating significant amounts of variance at both levels to justify multilevel analyses. In each multilevel model, we included GFP main effects, daily experiences, and their cross-level interaction. The daily predictors were person-mean centered (Nezlek, 2001), therefore, a participant's coefficient reflects daily fluctuations from his/her average level. Models were fitted using the nlme package in R (R Core Team, 2016; Pinheiro, Bates, DebRoy, Sarkar & R Core Team, 2016). Details on the multilevel procedure are in the Supplementary Materials.","The variables descriptives are presented in Table 1. Participants reported relatively few conflicts, on average 0.37 conflict per day. Conflicts and relationship quality, although related (r =−0.33), appeared to assess different aspects of interpersonal relationships. Relations between the GFP and daily social experiences ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ GFP scores positively related to daily relationship quality and impressions on others, and negatively to the number of daily conflicts (Table 1). These results support H1a-H1c. Interestingly, the relations between the GFP and the daily indicators of social effectiveness were roughly equal in size or larger than those involving the Big Five and these outcomes. Mediation Analyses ~~~~~~~~~~~~~~~~~~ There were sizeable correlations between the GFP and daily averaged self-esteem and mood (see also, Musek, 2007). As the GFP related to the mediators and outcomes, mediation analyses (Table 2) were permissible. Focusing on the direct/total effect ratio, the most important mediators were relationship quality and daily impressions. Daily impressions were the most important mediators of the GFP on the one hand, and self-esteem and PA on the other hand. For PA, about half of the total GFP effect was mediated by the daily impressions. Relationship quality was the most important mediator for NA. These results support the predictions in H2b and H2c. The relation between the GFP, and self-esteem and mood did not appear to be substantively mediated by the number of conflicts. Thus, only limited support was found for H2a. Moderation Analyses ~~~~~~~~~~~~~~~~~~~ The HLM-results are presented in Table 3. The hypothesized effects (H3a-H3c) were largely found for self-esteem and NA. Specifically, the cross-level interactions between the GFP and the various daily measures resulted in non-trivial decreases of random slope variance (between 0.10% and 4.37%, interpretable as R2-values). Taking the moderating effect of the GFP and NA as an example; the within-person SD of impressions was 0.60. Thus, having a bad day compared to an average day in terms of impressions (−1SD) results in a daily NA increase of about 0.13 and 0.19 for low-GFP and high-GFP individuals, respectively (Fig. 1A). These effects appear to be small, but should be considered against the within-person SDs of the outcomes (0.65, 61, and 0.55 for self-esteem, PA and NA, respectively). At first sight, Fig. 1A may suggest a floor effect because, compared to low-GFP individuals, high-GFP individuals scored lower on NA and thus have less opportunity to move down the scale. However, this explanation is at odds with GFP moderation of the relation between conflict and both NA and self-esteem. For example, for high-GFP individuals, there is enough leeway for interpersonal conflicts to negatively affect one's daily self-esteem (Fig. 1B). Yet, as expected, the negative effect of interpersonal conflicts on daily self- esteem is stronger (i.e., steeper) for those with lower GFP scores. None of the hypothesized moderations were found with positive affect. Interestingly, opposite to the hypothesis, the interaction between the GFP and daily impressions on PA was positive: a day with comparatively bad impressions resulted in a larger decrease of PA for high-GFP individuals. In conclusion, for daily self-esteem and NA, H3a through H3c were largely supported, while no support was found for daily PA (H3b).","The present study showed that (1) GFP scores were related to daily social experiences and well-being/mood, (2) daily social experiences partly mediated the relation between the GFP and well-being/mood, and (3) the GFP related to how individuals react to daily social events. To the best of our knowledge, this is the first time that the social effectiveness hypothesis of the GFP is studied using daily reports (with an N of ≈ 1220). This study may contribute to the literature in four ways. First, we found the GFP to be related to higher daily relationship quality and less conflicts. These outcomes can be viewed as indicators of social effectiveness (Denissen et al., 2008) and fit with previous findings showing that the GFP is related to positive social outcomes such as popularity (Van der Linden et al., 2010b), job performance and obtaining leadership positions (Pelt, van der Linden, Dunkel &, Born, 2017). Second, we showed how GFP scores were associated with leaving better daily impressions on others. Recent studies have confirmed that impression management is best seen as a stable, substantive trait related to self-control in social contexts (Uziel, 2010). This definition is similar to the substantive GFP interpretation. Accordingly, it can be argued that (successful) impression management may be inseparable from personality (cf. Danay & Ziegler, 2011). A third contribution is our provision of a potentially relevant mechanism for the strong relationship between the GFP and subjective well-being (e.g., Musek, 2007). Because social relationships have been proposed to be “the greatest single cause” of well-being (Argyle, 2001), it may not come as a surprise that any social skills associated with the GFP would allow for maintaining better social relationships, in turn, resulting in higher levels of well-being. Fourth, high-GFP (vs. low-GFP) individuals’ daily mood was less strongly influenced by daily fluctuations in social interactions and events. This is in line with the GFP as an adaptive trait that not only reflects social aptness, but that also cushions the impact of adversities (see also, Hengartner et al., 2017). Counter to our hypotheses, no moderating effects of the GFP were found on the relations between daily conflict/relationship quality and PA. One possible explanation is that the participants’ PA-levels resided around the midpoint of the scale, and daily social experiences may not be salient enough to warrant a reaction at such levels. In contrast, average self-esteem scores were relatively high and NA scores relatively low; a deviation from such higher levels will perhaps trigger a more direct reaction (thus allowing for GFP moderation). At this point, however, this explanation is rather speculative and should be tested in the future. Salience of a daily experience may also be responsible for the unexpected finding that fluctuations in daily impressions on others had stronger effects on the daily PA of high (vs. low) GFP individuals. By far the largest mediation effect was found for the GFP – daily impressions – PA link. Thus, as leaving a good impression on others may be especially important for PA, daily successes or failures in achieving this might to be more pleasing or disturbing, respectively, at higher GFP levels. Limitations ~~~~~~~~~~~ The main limitation was the exclusive use of self-reports, introducing possible common method bias. However, diary data can be assumed to partly reduce the biases associated with self-reports (e.g., recall bias and social desirability). In addition, by using daily within-person fluctuations, the influence of common method variance is reduced (Beal, 2015). Furthermore, it is unlikely to find cross-level interactions when large amounts of common method variance are present (Lai, Li & Leung, 2013). Still, lower GFP scores may be associated with quicker interpretation of a given social situation as a conflict, or with selecting oneself into conflicts (e.g., Bolger & Schilling, 1991). Future studies should include other-reports to remedy such drawbacks. Further, although a large community sample was used, the participants were relatively young, childless, and mostly women. Therefore, testing our results’ generalizability in more heterogeneous samples would be desirable.","This study revealed how the GFP, as a presumed social effectiveness factor, translates to day-to-day social experiences. Using an extensive diary design, high-GFP individuals were found to experience fewer interpersonal conflicts, and were less negatively influenced by potentially disruptive social events. It is not difficult to imagine how the effects of being socially adaptive and knowing how to respond in social situations on a daily basis will accumulate and eventually would affect broader life outcomes such as job performance and better social relations."],["A Single-Use Carrier Bag Charge (SUCBC) requires bags to be sold for a small fee, instead of free of charge. SUCBCs may produce 'spillover' effects, where other pro-environmental attitudes and behaviours could increase or decrease. We investigate the 2011 Welsh SUCBC, and whether spillover occurs in other behaviours and attitudes. Using the Understanding Society Survey (n = 17,636), results show that use of own shopping bags increased in Wales, compared to England and Scotland. Increased use of own bags was linked to increases in six other sustainable behaviours, although changes were significantly smaller in Wales for three of these behaviours. Increased own bag use was linked to stronger environmental views, but effects were weaker in Wales for two out of three measures. We conclude that the Welsh SUCBC effectively encouraged bag re-use, but with minimal changes in other environmental attitudes and behaviours, due to the external motivation to change behaviour. --------------------------------------------------------------------------------","In 2010, UK supermarkets provided 7.57 billion single-use plastic bags to shoppers, accounting for approximately 65,000 tonnes of plastic polymer (WRAP, 2013). The environmental impacts of plastic bags are can be seen from littering, damage to land and marine wildlife, oil and energy consumption, and non-biodegradable plastic bag wastage presents a long-term problem (DEFRA, 2013). Alongside voluntary agreements with producers and retailers, government policies enforcing a small charge on plastic bags have become more popular over the last decade, with various policies enacted in countries and sub- regions across the globe (Miller, 2012). Results of plastic bag charge policies have seen varied results, often with large variation in the size of the charge levied, the length of time the charge remained in place, and whether the customer or the retailer paid the charge (Ritch, Brennan, & MacLeod, 2009). One successful example is the Irish plastic bag levy introduced in 2002, which not only reduced plastic bag use by approximately 94%, but also has proved popular among the general public (Convery, McDonnell, & Ferreira, 2007). Plastic bag use in Ireland has increased since the levy was introduced however, and an increase in the levy in 2007 (from €0.15 to €0.22) appeared to further reduce bag use, suggesting that charges for bags may need to be revised for effectiveness (Clarke, 2014). Wales became the first country in the UK to introduce a minimum charge for single-use carrier bags in October 2011. The Welsh Single-Use Carrier Bag Charge (SUCBC) requires all businesses to charge shoppers £0.05 (approx. €0.07 or US$0.08) for each single-use carrier bag used.1 An alternative for shoppers is to purchase stronger carrier bags, which are designed to be re-used and brought along by shoppers to the shops. These re-usable bags are often marketed as “Bags for life”, which can be replaced for free once they have worn out. The Welsh SUCBC has so far proved very effective (Poortinga, Whitmarsh, & Suffolk, 2013), and the number of single-use plastic bags distributed since 2010 has fallen by 81%, with an associated decrease in plastic bags used per capita per month from 9.7 plastic bags in 2010, to 1.8 bags in 2012 (WRAP, 2013). Behavioural spillover ~~~~~~~~~~~~~~~~~~~~~ While the main effects of SUCBCs have focused on reducing the volume of single-use bags, the potential for a charge to produce additional spillover effects have been noted (Poortinga et al., 2013). Spillover is a phenomenon where an intervention targeted at increasing one behaviour may lead to an increase or decrease in other, untargeted behaviours (Thøgersen, 1999; Truelove, Carrico, Weber, Raimi, & Vandenbergh, 2014). Positive pro-environmental behavioural spillover effects have been reported in a number of contexts. Studies have found links between an increase in recycling and using less resources (Thøgersen, 1999); purchasing of organic goods, sustainable transport and recycling (Thøgersen & Ölander, 2003); greater purchasing of sustainable items and increases in several other sustainable behaviours (Lanzini & Thøgersen, 2014); and between fuel efficient driving and eating less meat (Van der Werff, Steg, & Keizer, 2013). The potential for spillover to enhance pro-environmental interventions has seen great interest for government policymakers as a cost-effective and non-intrusive way to change multiple behaviours (Thøgersen & Crompton, 2009), with the UK government highlighting the potential for “catalytic” behaviours to strengthen environmental lifestyles (DEFRA, 2008, p. 22). Several processes have been suggested to explain positive behavioural spillover effects. One approach utilises cognitive dissonance theory (Festinger, 1962), where a person acting pro-environmentally in one area, whilst neglecting another area, can develop an uncomfortable sensation of inconsistency which may spur other pro-environmental behaviours: especially if the person is already environmentally motivated (Thøgersen, 2004). An alternative explanation is based on self-perception theory (Bem, 1972) which states that we reflect on our actions to determine, in part, our identity. After performing a behaviour (e.g., recycling) we may be more favourable to other behaviours, based on the fact that we are now a ‘recycler’ (Holland, Verplanken, & Van Knippenberg, 2002). It has also been theorised that enacting a pro-environmental behaviour may generate more favourable views to wider environmental issues, which may influence the person in the future (Cornelissen, Pandelaere, Warlop, & Dewitte, 2008), and also that people may develop more skills and greater self-efficacy to enact other behaviours after successfully performing an initial action (Thøgersen, 2012). The evidence for behavioural spillover is however mixed, with a range of positive and negative results, and a lack of detailed investigation into the mechanisms behind spillover effects (Austin, Cox, Barnett, & Thomas, 2011; Truelove et al., 2014). The discussion of negative spillover also presents a problem to understanding the phenomenon, when the targeted increase in one sustainable behaviour is matched by a reduction in a non-targeted behaviour (Thøgersen & Crompton, 2009). People may be inclined to undertake small and easy pro-environmental behaviours, which are often over-emphasised for their environmental effectiveness (Pieters, Bijmolt, Van Raaij, & de Kruijk, 1998). Undertaking these small behaviours may reduce the likelihood of people engaging in other environmentally beneficial actions, justified by already “playing one’s part” to help the environment, without the need for further action (Thøgersen & Crompton, 2009; Thøgersen, 1999). Justifications based on current pro- environmental behaviour may extend further and allow people a “moral licence” to initiate un-sustainable behaviours that would be permitted after performing some other sustainable behaviours (Truelove et al., 2014). In a longitudinal experiment where residents of apartment buildings were encouraged to reduce water consumption, although water use decreased, there was an increase in electricity use, which may demonstrate a moral licensing effect (Tiefenbeck, Staake, Roth, & Sachs, 2013). The Welsh SUCBC and behavioural spillover ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A difficulty with evaluating behavioural spillover is the varied research on the topic. Evidence for behavioural spillover largely exists by correlating the frequency of two behaviours, and inferring the potential for spillover between them (Austin et al., 2011; Lanzini & Thøgersen, 2014; Truelove et al., 2014; Whitmarsh & O’Neill, 2010). Reviews of behavioural spillover have highlighted the scarcity of research evaluating people’s changes over time (Austin et al., 2011; Truelove et al., 2014), and thus far there only have been a handful of experimental (e.g., Lanzini & Thøgersen, 2014; Tiefenbeck et al., 2013) and longitudinal (e.g., Thøgersen & Ölander, 2003; Van der Werff, Steg, & Keizer, 2014) studies examining behavioural spillover. Additional research into behavioural spillover is warranted not only to improve our understanding of the phenomenon, but given the UK government’s interest in creating ‘catalyst’ behaviours to encourage other sustainable behaviours (DEFRA, 2008), there is a need to provide evidence for policy makers to implement effective and suitable measures (Truelove et al., 2014). Of interest to the current study is the work by Poortinga et al. (2013), who surveyed independent samples of respondents in England and Wales 2 weeks prior, and 6 months after the Welsh SUCBC implementation. Poortinga et al. (2013) found that in Wales the practice of bringing own shopping bags increased, and that support for the SUCBC increased, but were unable to detect any positive behavioural spillover effects. But by analysing independent samples before and after the SUCBC was introduced, Poortinga et al. (2013) evaluated results at a national level, and were unable to determine changes at the individual level where behavioural spillover effects would be observed. Longitudinal analysis of the Welsh SUCBC is therefore required for more detailed investigation. But there is also scepticism that a SUCBC could encourage positive behavioural spillover effects, with concerns over the influence of internal and external motives on behaviour (Austin et al., 2011). Models explaining behavioural spillover effects often use personal identity and threats to consistency as drivers of behavioural spillover (Truelove et al., 2014). One threat to the identity and consistency explanations for behavioural spillover is that a SUCBC represents an external motive to change behaviour in order to avoid a charge (Poortinga et al., 2013). As stated in self-determination theory (Ryan & Deci, 2000), motivation exists on a continuum from intrinsic motivation, where behaviour is enacted out of personal enjoyment and satisfaction, through to extrinsic motivation, where behaviours stem from compliance with external rewards and punishments. It has been argued that policies that use external charges to change behaviour could weaken potential positive spillover effects by weakening intrinsic motivation or reducing threats to personal identity and consistency, which may then lead to negative spillover effects (Truelove et al., 2014). The implementation of the Welsh SUCBC, and lack of such policies (at the time) in England or Scotland offers an opportunity to evaluate whether external regulation may be linked to positive or negative spillover effects. Therefore, we investigated possible behavioural spillover effects linked to the Welsh SUCBC, and whether this external pressure to change behaviour is linked to positive or negative changes in other pro-environmental behaviours, compared to UK countries with no external impetus to change (i.e., no SUCBC). To evaluate the topic, we analysed data from the Understanding Society Survey (USS). The USS is the largest longitudinal panel survey in the world, collecting data from approximately 40,000 households across the UK, and contains several measures of pro-environmental behaviours attitudes for analysis (University of Essex, 2015).","The paper has three aims. First, we will examine how bag re-use behaviour has changed in Wales, England, and Scotland. We hypothesise that the SUCBC will have increased bag re-use behaviour in Wales, and that the increase will be greater than in England or Scotland (Hypothesis 1). Second, we investigate whether increased frequency of taking own shopping bags is linked to increases in other pro-environmental behaviours. We hypothesise that because the SUCBC is an external motivation, increases in taking own shopping bags wil not lead to positive changes in other pro-environmental behaviours in Wales, but England and Scotland will demonstrate a positive link between an increase in own bag use and other behaviours (Hypothesis 2). Third, we will identify whether increased use of own shopping bags is associated with an increase in strength of pro-environmental views. Again as an external influence, we hypothesise that countries where bag re-use could be attributable to internal motivations (i.e., England and Scotland) will show positive links between increased bag re-use and stronger pro-environmental views, in contrast to Wales where external motivation from the SUCBC may have a weaker effect on changing personal views (Hypothesis 3). The Understanding Society Survey ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The USS collects data in ‘waves’ over 24 months; Wave 1 ran during 2009/10, and Wave 4 during 2012/13 with data collected using face-to-face interviews. Contents of each Wave vary, and Waves 1 and 4 included questions on environmental attitudes and behaviours, including a question on the frequency of bringing one’s own bag when shopping. This allows a longitudinal measure of changing attitudes and behaviour between 2009/10 and 2012/3, ideal for evaluating effects of the Welsh SUCBC implemented in 2011. In addition, the USS is designed to be representative of the UK population and allows comparison between Wales, England and Scotland. The USS is available online, and the 6th edition dataset was downloaded on 19th January 2015 (University of Essex, 2015). Environmental behaviours Waves 1 and 4 measured the frequency of 11 pro-environmental behaviours: leaving a TV on standby (reverse coded), switching off unused lights, running a tap while brushing teeth (reverse coded), putting on clothes instead of turning up home heating, not buying items because of wasteful packaging, buying recycled paper products, taking a bag when shopping, using public transport, walk/cycle short journeys, car sharing, and taking fewer flights. Behaviours were measured on a 5-point frequency Likert scale from “Always” to “Never”. Descriptive statistics of the frequency of these behaviours at Wave 1 and Wave 4, are shown in Table 1, with higher scores indicating greater sustainability. Views on environmental lifestyles Three items measured perceptions of pro-environmental lifestyles using varying Likert scales. First, the item “Which of these best describes how you feel about your current lifestyle and the environment?” was measured using a 3-point scale from “I’m happy with what I do at the moment” to “I’d like to do a lot more to help the environment”, where higher scores indicate greater desire to increase one’s environmental lifestyle. Second, the item “Which of these would you say best describes your current lifestyle?” was measured using a 5-point scale from “I don’t really do anything that is environmentally-friendly” to “I’m environmentally-friendly in everything I do”, where higher values indicate having a stronger environmental lifestyle. Third, the item “Do you agree or disagree that being green is an alternative lifestyle, it’s not for the majority” was measured using a 4-point scale from “Agree strongly” to “Disagree strongly”, where higher scores indicate greater disagreement that a green lifestyle is unsuitable for the majority. The three items had very low internal reliabilities (Cronbach’s α = 0.24 and 0.25 for Wave 1 and 4 respectively), and therefore are treated as individual measures. Demographics To control for variation in demographics between UK countries, and noting the link between demographic factors and environmental concern (Hawcroft & Milfont, 2010), we sought to control for possible confounding influences. The USS measured respondent’s age and gender, and included a calculated monthly income for each household based on all reported sources of income. Data preparation & final sample ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ All analyses used SPSS v.20. Data for Wave 1 and 4 were merged, selecting only participants who completed both Waves. We excluded people living in Northern Ireland (n = 1,242). The Northern Irish SUCBC was introduced in April 2013, and although outside the timespan of Wave 4, the charge had been discussed in the media at least 12 months beforehand (BBC, 2012), which may influence behaviours and attitudes in Wave 4 responses. The USS provides predetermined weights that allows generalisation of results to the UK population, and the weight d_indscus_lw was applied (Knies, 2014, p. 49). The final weighted dataset contained 17,636 respondents living in Wales, England, and Scotland, and all figures shown have weighting applied. A description of the final dataset is shown in Table 2. One-way ANOVA indicated that age was significantly different among the three countries, (F (2, 17,632) = 5.91, p = 0.003), with post-hoc tests (Tukey HSD) suggesting that the Wales sample was significantly older than the England sample, though of very small effect size (p = 0.022, Hedge’s g = 0.09), and no significant difference between England and Scotland (p = 0.05) or Scotland and Wales (p = 0.81). Pearson’ Chi-Squared test indicated there was no significant variation in gender between the three countries, X2 (2) = 4.88, p = 0.11. One-way ANOVA also indicated that monthly household income was significantly different among countries, F (2,17,632) = 17.11 p < 0.001, with post-hoc tests (Tukey HSD) indicating Wales had a lower monthly income than England (p < 0.001, Hedge’s g = 0.19) and Scotland (p = 0.004, Hedge’s g = 0.14) samples, with no significant difference between Scotland and England (p = 0.09). Changes in frequency of bringing own shopping bags ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ First we explore how respondents in Wales changed their bag re-use behaviour between 2009/10 (Wave 1) and 2012/13 (Wave 4) as compared to respondents in England and Scotland. Using a repeated measures ANCOVA, mean frequency of taking own shopping bags at Wave 1 and Wave 4 was compared for the three UK countries, controlling for age, gender and monthly household income. Results indicated that, in general, frequency of taking own shopping bags changed between Waves 1 & 4, F (1, 18,386) = 17.67, p < 0.001, partial η2 = 0.001, an extremely small effect where taking own shopping bags generally reduced over time. Covariate of age was not significantly related to changes in bag reuse (F (1, 18,386) = 0.91, p = 0.34), while gender showed an extremely small effect where women reduced how often they took their own shopping bags over time (F (1, 18,386) = 14.17, p = <0.001), and monthly income had an extremely marginal effect, F (1, 18,386) = 4.78, p = 0.029, partial η2 < 0.001. These effects were overshadowed by a stronger Time × Country interaction, F (2, 18,386) = 132.30, p < 0.001, partial η2 = 0.014, indicating that frequency of taking own bags varied over time between the countries. Using the frequency of taking own shopping bags, from ‘Never’ (1) to ‘Always’ (5), the significant interaction is displayed in Fig. 1. Using repeated measures effect size calculations (Lakens, 2013), simple pre/post comparison of mean bag use frequency scores for England and Scotland both show very small decreases, Hedge’s grm = 0.05. For Wales, the change in frequency score was Hedge’s grm = 0.48, a conventionally “medium” sized increase in the frequency of bringing own bags to shops (Cohen, 1988). Proportion ‘always’ taking their own shopping bag The proportion of people who indicated on the Likert-scale that they “always” took their own shopping bag when they went shopping was compared for countries at Wave 1 and Wave 4, with results highlighted in Fig. 2: At Wave 1, Chi-squared tests indicated no significant difference in the proportion of people ‘always’ taking their own shopping bag, X2 (2,17,102) = 1.11, p = 0.58, with around 47% of respondents in each country “always” doing so. At Wave 4, chi-squared test found a significant difference between GB countries, X2 (2,17,228) = 313.81, p < 0.001, Cramer’s V = 0.14, indicating while that 44% of English and 41% of Scottish respondents ‘always’ took a bag when shopping, 74% of respondent in Wales ‘always’ took their own shopping bag. With significant changes in the frequency of taking own shopping bags in Wales, and little to no change in behaviour for England and Scotland, Hypothesis 1 is supported. Change in bringing own bags and behavioural spillover ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Next, we investigated whether changes in bag re-use behaviour were associated with positive or negative changes in ten other pro-environmental behaviours measured within the USS. As no major differences in behaviour or demographics were observed between England and Scotland, these groups were merged into a general sample to simplify interpretation of results when comparing against Wales. To investigate positive or negative behavioural spillover and changes in lifestyle views, we emulated the method by Lanzini and Thøgersen (2014) where the dependent variable is the change of the target measure (i.e. Wave 1 of measure X is subtracted from Wave 4 score of measure X). Then, the baseline measure of the target variable (i.e. Wave 1 score of measure X) is added as a covariate in a regression model to control for regression to the mean effects, and then the change between Wave 1 and 4 of the frequency of taking own shopping bags (“Bag Change”) is included as an independent variable. This approach controls for baseline changes in the target variable, allowing any change in the target variable to be linked to changes in the catalyst behaviour. In addition, we control for respondent’s age, gender (dummy coded where female = 1 and male = 0), and monthly household income. To evaluate differences between countries, we included a dummy coded variable for Wales against England and Scotland samples (“Country”, where England & Scotland = 0 and Wales = 1), and an interaction between change in frequency of taking own bags and country; change in bag use was mean- centred prior to specifying the interaction term (Aiken, West, & Reno, 1991). To conserve space, a summary of coefficients for change in bag re-use behaviour predicting changes in other environmental behaviours, and the Country × Bag Change interaction (controlling for baseline and covariates), are shown in Table 3. Full details of model coefficients are available upon request. Results indicate that a general increase in taking own shopping bags is associated with an increase in six other pro-environmental behaviours. These effects are generally small, with standardised regression coefficients ranging around 0.11, below the conventional ‘small’ effect size of 0.20 (Cohen, 1988). In addition to these general effects, three interaction terms were significant, indicating that the Wales sample was significantly different than England and Scotland samples in their association between changes in taking own shopping bags and changes in other sustainable behaviours. For illustration, interactions were plotted using unstandardised regression coefficients using Jeremy Dawson’s excel macros2 and shown in Figs. 3–5. Simple slopes analysis indicates that for changes in frequency of turning off the tap when brushing teeth, a positive regression slope for England and Scotland is significant, B = 0.17, t (16,611) = 4.63, p = <0.001, and a positive but weaker slope for Wales is also significant, B = 0.03, t (16,611) = 3.85, p = <0.001. Simple slopes analysis show that for changes in frequency of wearing warm clothes instead of turning on the heating, the positive regression slope for England and Scotland is significant, B = 0.09, t (16,666) = 3.12, p = 0.002, and the positive but weaker regression slope for Wales is significant, B = 0.04, t (16,666) = 5.18, p < 0.001. Simple slopes analysis indicates that for changes in frequency of using public transport, the positive regression slope for England and Scotland is significant, B = 0.11, t (15,267) = 3.77, p < 0.001, and the positive but weaker regression slope for Wales is significant, B = 0.04, t (15,267) = 5.78, p < 0.001. Interaction results show that while England and Scotland samples show that an increase in bringing own shopping bags predicts an increase in three unrelated pro-environmental behaviours, suggesting positive spillover effects, the Wales sample showed much weaker and only marginal increases between changes in bringing own shopping bags and these three other pro- environmental behaviours. The results therefore partially support Hypothesis 2, indicating that where a general increase in bringing own shopping bags predicted an increase in six other behaviours, three of these links were significantly weaker for the Wales sample. Changes in bringing own shopping bags and changes in lifestyle views ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Next we investigated whether the observed changes in behaviour were associated with positive or negative changes in environmental views. Analysis used the same method as before, by including baseline measures of the target variable as a covariate, and an interaction term to compare results between Wales and the England and Scotland samples for possible links to the SUCBC. The USS included three questions that captured individual aspects of environmental lifestyle: desire to increase environmental lifestyle, strength of environmental lifestyle, and the perception of being ‘green’ as an alternative lifestyle. Changes in each measure, and their potential links with changes in frequency of taking own shopping bags, are explored below. Desire to increase environmental lifestyle Regression analysis predicting a change in people’s desire to increase their environmental lifestyle was run, where a higher score indicates a greater desire to increase one’s environmental lifestyle, as predicted by changes in frequency of taking own shopping bags, is shown in Table 4. Results indicate that an overall increase in taking own shopping bags is linked to a small decrease in the desire to do more environmental actions. However, with an interaction effect between change in bag re-use and country, regression slopes between the two samples were significantly different. This interaction is shown in Fig. 6. Simple slopes analysis indicates that the negative regression slope for England and Scotland was significant, B = −0.04, t (16,684) = 2.71, p = 0.007, and the negative regression slope for Wales was significant but weaker, B = −0.01, t (16,684) = 2.84, p = 0.005. Results show that in England and Scotland, an increase in taking your own bag when shopping was linked to a decreased desire to increase current environmental lifestyle (i.e., increased satisfaction with current level of behaviour), whereas in Wales, any changes in own shopping bag behaviour is linked to only a fractional reduction in the desire to do more environmental actions. Strength of current environmental lifestyle Analysis predicting changes in perceived strength of environmental lifestyle was run, where a higher score indicates a stronger environmental lifestyle, predicted by changes in taking own shopping bags, is summarised in Table 5. Results in Table 5 indicate that overall, an increase in bag re-use behaviour was associated with a very small increase in people’s perceived strength of their environmental lifestyle. No significant interaction effect was found, indicating that the Wales sample did not significantly vary from the English and Scottish sample. Perception of ‘green’ lifestyle for the majority Analysis predicting change in the perception of a ‘green’ lifestyle as acceptable for the majority, where a higher score indicates greater acceptance of a ‘green’ lifestyle, by changes in taking own shopping bags is summarised in Table 6. Results in Table 6 suggest a small, but overall positive association between increased use of own shopping bags and increased perception that a ‘green’ lifestyle is suitable for the majority. Notably, the interaction effect is significant, and illustrated in Fig. 7. Simple slopes analysis show that the positive regression slope for England and Scotland is significant, B = 0.05, t (16,386) = 3.01, p = 0.003, and the positive regression slope for Wales is significant, though weaker, B = 0.01, t (16,386) = 3.89, p < 0.001. Results indicate that a general increase in bag re-use behaviour is linked to a greater perception that a ‘green’ lifestyle is suitable for the majority. The Wales sample however, shows only a marginal increase in acceptance of green lifestyles linked to an increase in changes in bag re-use behaviour, compared to the stronger relationship seen in the England and Scotland sample. The results show that an increased frequency of bringing own shopping bags is generally linked to small changes in views of living an environmental lifestyle. However, for two of these links between own bag use and lifestyle views, a significant interaction indicates that the Welsh sample show weaker relationships than the England and Scotland sample; therefore these results partially support Hypothesis 3.","This paper describes an analysis of the Understand Society Survey (USS) dataset, a longitudinal and nationally representative survey of the UK population measuring pro- environmental views and behaviour in 2009/10 and 2012/13. With the introduction of the Single-Use Carrier Bag Charge (SUCBC) in Wales from October 2011, and with no such policies implemented in England or Scotland, the USS dataset is an ideal method to evaluate effects of the SUCBC on bag re-use behaviour. Additionally, the USS was analysed to establish if the SUCBC may have led to ‘spillover’ effects that occur when starting one pro-environmental behaviour leads to an increase, or a decrease, in other pro- environmental behaviours and views. We find that bag re-use behaviour increased significantly in Wales. The mean reported frequency of respondents in Wales bringing their own bags when shopping saw a significant and conventionally ‘medium’-sized increase between 2009/10 and 2012/13, whereas respondents in England and Scotland (with no SUCBC introduced) showed a small decrease in bringing own shopping bags. The reported frequency of people in England and Scotland “always” taking their own shopping bag showed no change between 2009/10, whereas in Wales, 74% of respondents reported “always” taking their own bag, an increase of 25%. Clear differences in behaviour changes in Wales, and a similar lack of changes in England and Scotland demonstrate the role of the Welsh SUCBC increasing bag re-use behaviour, and supports Hypothesis 1. The results also confirm previous research indicating that the Welsh SUCBC increased the proportion of Welsh people bringing their own bags to shopping trips (Poortinga et al., 2013). Alongside self-reported behaviour, our findings concur with objective reports of the number of single-use plastic bags used in Wales, which fell dramatically from before and after the implementation of the SUCBC (WRAP, 2013). Overall, it appears that the Welsh SUCBC had a clear and positive effect of reducing use of single-use plastic bags and increasing the frequency of people bringing their own bags when shopping. We also examined whether a change in the frequency of bringing own shopping bags was linked to changes in other behaviours and views. Results show that across the UK, a general increase of taking one’s own bag when shopping was linked to increases in six other sustainable behaviours: turning off the tap when brushing teeth, wearing warmer clothes indoors instead of turning up heating, buying recycled paper products, using public transport, walking/cycling short trips, and carsharing. Linked increases between taking own shopping bags and other behaviours may suggest that positive behavioural spillover was occurring. However, we do not believe that the increase in taking a bag when going shopping is a causal predictor of an increase in these other sustainable behaviours. In their review of behavioural spillover literature, Austin et al. (2011) conclude that the nebulous influences of external and internal influences should be considered before inferring spillover, and that causal inferences between two linked behaviours is often not possible. The sustainable behaviours that increased would also defy expectations that a small change in shopping behaviour (i.e., bringing a bag to the shops) could lead to an increase in more difficult behaviours in other domains (e.g., using public transport). Positive behavioural spillover effects are more likely among behaviours perceived to be similar (Thøgersen, 2004), and more likely to spillover to other low cost/energy behaviours (Lanzini & Thøgersen, 2014; Thøgersen & Crompton, 2009). What the results do show is that the frequency of various pro-environmental behaviours increases over the same time frame; as people increased one behaviour, they may have increased other, unrelated behaviours. Correlations between sets of pro-environmental behaviours has been examined before as evidence for potential behavioural spillover effects (Austin et al., 2011; Truelove et al., 2014; Whitmarsh & O’Neill, 2010). This analysis offers new insight by demonstrating that in a nationally representative sample, an increase in the frequency of one pro-environmental behaviour is associated with increases in other pro-environmental behaviours. The joint increases in frequency of various sustainable actions thus offer additional support to the idea that positive spillover may occur between different behaviours. Of note is that the observed effects in this study are small, with standardised regression coefficients (β) around 0.11 in size, below the conventional ‘small’ effect size of 0.20 (Cohen, 1988). It appears that the current results are not an isolated case however. In their review of pro-environmental spillover effects, Austin et al. (2011) reported that the strength of spillover effects were generally “weak” (p. 90). Using standardised regression coefficients, effect sizes of other longitudinal spillover results appear of similar strengths; Van der Werff et al. (2014) predicted intentions to eat less meat from eco-driving (β = 0.14), and Thøgersen and Ölander (2003) reported an interplay between recycling, organic food purchases, and sustainable transport choices that ranged from β = 0.06 to β = 0.13. Even within a tightly controlled experimental design, Lanzini and Thøgersen (2014) found that an increase in purchasing ‘green’ items led to increases in six other pro-environmental actions that ranged between β = 0.14 to β = 0.23. The current results are therefore comparable with the existing literature, but highlight the small effect sizes linked with behavioural spillover. Comparing Wales against England and Scotland, we find that of the six significant links between an increase in taking a pre-owned shopping bag and increases in other behaviours, three showed that changes in the other behaviours were significantly weaker in Wales; turning off the tap when brushing teeth, wearing warmer clothes indoors, and using public transport. In order to explain this result, we also investigated whether changes in taking own bags was linked to changes in attitudes. Results found that although a general increase in taking own shopping bags was linked to small increases in three measures of positive views of environmental lifestyles, two of these items showed significantly weaker effects in Wales when compared to England and Scotland. Recent investigations suggest that positive and negative spillover effects may also occur from sustainable behaviour to support for environmental policies (Lacasse, 2015; Truelove, Yeung, Carrico, Gillis, & Raimi, in press), and views on environmental lifestyle may also be affected. It therefore appears that although Wales saw a significant increase in frequency of taking own shopping bags (likely due to the SUCBC), this increase in behaviour is linked to lower rates of increases in some other behaviours, and lower increases in positive views on living environmental an lifestyle. We argue that the Welsh SUCBC, although effective at increasing the use of own shopping bags, was not as effective at encouraging wider changes because of the external pressure to change behaviour. In their review of behavioural spillover, Truelove et al. (2014) suggest that where behaviour is influenced by external pressure, be it an incentive or disincentive, positive spillover is less likely to occur as external pressures removes intrinsic motivation which could be a key motivator for positive spillover effects. In accordance with self-determination theory (Ryan & Deci, 2000), removing intrinsic motivation for bringing own shopping bags challenges many theories of how positive behavioural spillover could occur. Cognitive dissonance theory (Festinger, 1962) suggests that discomfort from inconsistency across behaviours would encourage other sustainable actions, but cognitive dissonance is not present when the inconsistency can be attributed to external pressures (Thøgersen, 2004). Self-perception theory (Bem, 1972) states that identity is modelled partly on actions, so that a ‘green’ identity is strengthened by performing sustainable behaviours (which then encourages other similar actions). Yet self-perception theory also highlights that a person will consider their actions and strengthen their identity only “if that behaviour appears to be free from the control of explicit reinforcement contingencies” (Bem, 1972, p. 6). Lastly, positive behavioural spillover pathways from increased knowledge and self- efficacy (Thøgersen, 2012) would also be challenged by the external disincentive of a SUCBC, as perceived competence does not enhance intrinsic motivation unless also matched by a perception of autonomy in the action (Ryan & Deci, 2000), which an external charge would likely remove. The external motive to change behaviour in Wales, and presumably internal motivations to change in England and Scotland (with no charge in place), would therefore explain the weaker changes in other behaviours and views on sustainable lifestyles seen in Wales. One positive indication from this analysis is that the Welsh SUCBC, although linked to weaker increases in other behaviours and attitudes, did not show signs of negative spillover effects; where other behaviours could actually decrease in frequency. Truelove et al. (2014) suggested that extrinsic motivations to act sustainably, such as price-based policies, may lead to negative spillover if not correctly enacted. With no evidence of negative spillover after the Welsh SUCBC, fears of negative spillover effects may be reduced, but as SUCBCs may need to be increased over time to maintain their effectiveness (Clarke, 2014), continued monitoring of long-term behaviour change is required. This analysis is the first longitudinal analysis to evaluate changes in individual behaviour and views after the implementation of a SUCBC. Previous work by Poortinga et al. (2013) evaluated the Welsh SUCBC using independent samples representative of Wales and England before and after the SUCBC. But without observing individual changes over time, their analysis could not evaluate the personal effects of the SUCBC. With longitudinal data from a nationally-representative sample of 17,636 respondents, the current analysis is a substantially higher-powered evaluation of spillover effects. Nonetheless, our analysis has some caveats to consider. Although longitudinal, we are unable to show causality between changes. For example, an increase in taking own bag use in Wales was likely due to the SUCBC, but we cannot predict the cause of the changes in taking own bags when shopping in England and Scotland. It may be that an increase in another, unrelated behaviour spilled over to taking own shopping bags in England or Scotland, but this cannot be determined. The length of time between measurements is also a factor: the Wave 4 survey was taken on average 36 months after the Wave 1 survey. Although useful for a large-scale evaluation, there may be several different influences on behaviour and attitude that occurred during these time points which cannot be conclusively ruled out. This analysis offers a very large sample for evaluation, but additional work on the effects of SUCBCs should also be considered, perhaps using more frequent measurements around the time of implementation of a charge.","We find that the frequency of taking own shopping bags significantly increased in Wales, compared to England and Scotland, which is most likely an effect of the SUCBC. We also evaluated the potential for the SUCBC to promote positive or negative ‘spillover’ effects, where the change in the targeted behaviour may have led to an increase, or even a decrease, in other behaviours. Results show that in general, an increase in taking own shopping bags is linked to a small increase in six other behaviours, but that this effect is significantly weaker in Wales than in England or Scotland. Additionally, increased use of taking own shopping bags had a small, positive link to stronger perceptions of living a sustainable lifestyle, but again two of these links were significantly weaker in Wales. We discuss the results in terms of extrinsic versus intrinsic motivation, and argue that the external motivation to change behaviour in Wales did not lead to changes in other behaviours or views, compared to the intrinsically-motivated changes in England and Scotland. Although the SUCBC appears effective at changing behaviour, we do not believe that positive behavioural spillover effects are likely to be encouraged by SUCBC policies."],["To examine the relationship between the Big Five and cognitive ability, we investigated whether we could replicate in a heterogeneous population sample the positive association between cognitive ability and Openness and Emotional Stability and its negative association with Conscientiousness. Besides analyzing the pure associations, we shed further light on sources of these associations by investigating potential moderating effects of education and labor force participation. Our results clearly replicate the previously found positive association between cognitive ability and Emotional Stability and Openness and the negative relationship between Conscientiousness and cognitive ability. The correlation between cognitive ability and Openness was found to be moderated by educational attainment, the negative association between Conscientiousness and cognitive ability was moderated by labor force participation. --------------------------------------------------------------------------------","The degree to which personality and cognitive ability are related is a question that has generated intensive research and intense debate. Some authors have concluded that intelligence test performance may be influenced by some non-ability traits but that intelligence and personality are two independent constructs (Zeidner & Matthews, 2000). Others (e.g., Ackerman, 1996) have argued that personality traits play a significant role in the development of intellectual skills. With regard to the most well-established model of personality, the Big Five, numerous studies and meta-analyses have indicated a substantial but comparatively modest association between personality and intelligence. The proportion of variance in cognitive ability explained by personality typically ranges between five and ten percent (Furnham, Dissou, Sloan, & Chamorro-Premuzic, 2007). Studies have consistently found a positive link between cognitive ability and Openness (see, for example, meta-analytical results by Ackerman & Heggestad, 1997 and Von Stumm & Ackerman, 2013) and Emotional Stability (e.g., Ackerman & Heggestad, 1997; Moutafi, Furnham, & Crump, 2003; Zeidner & Matthews, 2000), and a negative association between cognitive ability and Conscientiousness (e.g., DeYoung, 2011; Furnham et al., 2007; Moutafi et al., 2003; Soubelet & Salthouse, 2011). In particular, the negative association between Conscientiousness and cognitive ability appears contradictory at first sight, given that both intelligence and Conscientiousness are positively associated with work-related outcomes (e.g., Barrick & Mount, 1991; Gottfredson, 1997; Schmidt & Hunter, 1998). As one possible explanation for this “mysterious” effect (see Furnham et al., 2007), it has been suggested that the consistently found negative association between Conscientiousness and intelligence might be a methodological artifact caused by a sampling bias (e.g., Murray, Johnson, McGue, & Iacono, 2014; Soubelet & Salthouse, 2011). Almost all of these studies investigated only college student populations (see Furnham et al., 2007), and thus samples that are comparatively homogeneous with regard to education, age, labor market experience, and intelligence itself. Hence, the negative association might have been artificially created because individuals with low cognitive ability and low Conscientiousness were missing from the samples (see Major, Johnson, & Deary, 2014). Whether the negative association between Conscientiousness and cognitive ability can be replicated in a more heterogeneous adult population and can therefore be regarded as a general effect is a hitherto unanswered question. Besides offering methodological explanations, some researchers have also tried to explain the negative association between cognitive competencies and Conscientiousness substantively. Based on their finding that the Conscientiousness facet Orderliness, in particular, is negatively correlated with intelligence, Moutafi et al. (2003) argued that people with lower intelligence use planning and organization to compensate for their disadvantage on intellectual tasks (for a rebuttal, see Murray et al., 2014). Measuring cognitive ability ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Competence tests, such as those used in the Organisation for Economic Co-operation and Development’s (OECD) Programme for International Student Assessment (PISA) and the Trends in International Mathematics and Science Study (TIMSS), are highly correlated with intelligence. For TIMSS, Lynn and Mikk (2007), for example, report correlations with a general intelligence factor ranging between 0.92 and 1. Similar results can be found for PISA (e.g., Rindermann, 2006) and for the National Adult Literacy Survey (NALS; Gottfredson, 1997). It has even been debated whether the literacy, mathematics (or numeracy), and science competence tests measure general intelligence from a conceptual and empirical perspective (Rindermann, 2006). Other researchers have shown, however, that – in addition to general intelligence – such competence tests assess domain-specific competencies (e.g., Baumert, Brunner, Lüdtke, & Trautwein, 2007; Gottfredson, 1997). This debate notwithstanding, it has been clearly shown that – to a large extent – these competence tests measure general intelligence. Thus, these competence measures can be regarded as appropriate indicators of cognitive ability. For example, Hunt and Wittmann (2008) used the PISA 2003 competence measures as a proxy for intelligence to replicate results on country differences in IQ initially reported by Lynn and Vanhanen (2002). And Gottfredson (1997) used NALS data to show that general intelligence (g) is associated with cumulative life outcomes such as labor force participation or living in poverty. Besides being good indicators of cognitive ability, competence measures used in studies such as PISA, TIMSS, and the OECD-initiated Programme for the International Assessment of Adult Competencies (PIAAC; aka “PISA for adults”) have a further advantage compared to the commonly used IQ data – namely, that they are based on probability samples that are representative of the respective target populations (PISA: 15-year-olds; PIAAC: adults between 16 and 65 years of age) in the participating countries. Aim of the present study ~~~~~~~~~~~~~~~~~~~~~~~~ The present study examines the relationship between personality – in particular, the Big Five personality domains – and cognitive ability. As indicators of cognitive ability, we used the competence measures of the literacy and numeracy domains assessed in PIAAC. We investigated whether it was possible to replicate the previously found positive association between cognitive ability and Openness and Emotional Stability and its negative association with Conscientiousness. As it has been suggested that this negative correlation between Conscientiousness and cognitive ability might be caused by biased – primarily college student – samples, we investigated whether this effect was in fact less pronounced in a more heterogeneous adult population. Besides analyzing the pure associations, we aimed to shed further – or new – light on causes of these links by examining potential moderating effects of education and labor force participation. Sampling method and participants ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Data for the present study were collected in part within PIAAC. This programme compares cognitive skills such as literacy and numeracy across a large number of (mainly OECD) countries. For the present research, we analyzed the German PIAAC data.1 The target population were adults (aged between 16 and 65 years) randomly selected from local population registers in randomly selected municipalities throughout Germany. Participation in PIAAC was voluntary; an incentive of 50 euros was offered upon participation in the survey, which comprised a personal interview (average duration: 45 min) and a cognitive assessment lasting approximately 60 min. No time limit was imposed on the cognitive test. A detailed description of the sampling procedure and the technical implementation is given in Zabal et al. (2014). In addition to the PIAAC study conducted in 2012, 3758 of the original 5465 participants in Germany were re-interviewed in 2014 as part of the PIAAC Longitudinal Study (PIAAC-L).2 Data from the 2012 German PIAAC wave and the 2014 follow-up survey were combined and used for the present analyses.","The following variables from PIAAC 2012 were investigated: Verbal and numerical cognitive ability: Literacy and numeracy skills assessed in PIAAC were used as measures of verbal and numerical cognitive ability. Both competencies were assessed using a multistage adaptive testing design comprising a total of 58 items for literacy and 56 items for numeracy. Using a largely randomized procedure, respondents were allocated to the competence domains. Detailed information on the nature of the test and a selection of sample items are provided in the reader’s companion for the survey (OECD, 2013a, pp. 17).3 For each participant, 10 plausible values were estimated for each competency domain. (For details of the design and the IRT scaling process in PIAAC, see OECD, 2013b). Analyses of the cognitive data were run separately for each of the ten plausible values per domain. Results were then averaged within each domain. Education: Each respondent’s highest level of educational attainment was assessed with two separate questions (highest general education and highest vocational education qualification in the categories of the German education system), which were then mapped to the 1997 International Standard Classification of Education (ISCED 1997; [PIAAC 2012 variable: B_Q01a]).4 Forty-one respondents who reported that they had a foreign educational qualification for which they were unable to state the German equivalent were excluded from the analyses. In addition variables from the 2014 PIAAC-L follow-up were used for the present analyses: Labor force participation: All respondents were asked to report whether they were currently employed and, if not, what their current status was (PIAAC-L 2014 variable: perw_14). Based on this, a dichotomous variable was generated and used in the analyses (1 = in full-time employment, 0 = not in full-time employment).5 Respondents who indicated that they were still undergoing education were excluded from the analyses. The rationale behind dividing respondents into two groups – full-time employed and not full-time employed – was (1) to separate respondents for whom working is the most substantial part of their everyday lives from other respondents, (2) to create categories of similar size, and (3) to simplify the interpretation of the interaction term by creating a binary variable. In addition, a short version of the Big Five Inventory (BFI; John, Donahue, & Kentle, 1991) comprising three items per dimension was administered to respondents in the 2014 follow-up survey to assess their personality. This 15-item questionnaire – originally developed for use in the German Socio-Economic Panel (SOEP; Schupp & Gerlitz, 2014) – contains short statements, which are rated on a seven-point Likert scale ranging from 1 = “does not apply at all” to 7 = “applies completely.” Studies investigating the reliability and validity of this BFI-S have concluded that its psychometric properties were acceptable (e.g., Hahn, Gottschling, & Spinath, 2012). In the present sample, Cronbach’s alpha for the BFI-S scales ranged between 0.41 for Agreeableness and 0.69 for Extraversion.6 As our study design – though longitudinal – includes only one assessment of cognitive abilities and personality, respectively, we subjected the data to cross-sectional analysis.","To what degree is personality related to a person’s cognitive ability? To investigate this question, we correlated the scale scores for the Big Five dimensions with the estimates for verbal and numerical ability. Besides means and standard deviations for the five personality domains, Table 1 shows the resulting correlations between the Big Five and verbal and numerical ability. For both competency domains, the strongest correlation found was with Emotional Stability (0.11 and 0.14, respectively), indicating that emotionally stable persons have, on average, higher verbal and numerical abilities. In addition, both abilities were found to be significantly negatively related to Conscientiousness (−0.09 and −0.08, respectively). Smaller, but still significant, correlations were found between both cognitive abilities and Openness and Extraversion, indicating that introverts (−0.05 and −0.06, respectively) and open persons (0.05 for both domains) have, on average, higher cognitive abilities. In a second step, we analyzed the degree to which the Big Five personality domains incrementally predict cognitive ability. Studies have shown that intelligence is highly related to education and to work outcomes (e.g., Gottfredson, 1997; Schmidt & Hunter, 1998). We therefore investigated the extent to which the Big Five contribute to predicting intelligence over and above education and labor force participation. In a first step, we conducted regression analyses including only the Big Five personality domains. As a recent study (Major et al., 2014) demonstrated that there are also quadratic associations between personality and cognitive ability, we included both linear and quadratic relations in our analyses. In a second and third step, we then included in the regression analyses (a) the highest level of educational attainment and (b) labor force participation. As the highest level of educational attainment is a valid predictor only for those participants who have completed their initial formal education, we excluded all participants from further analyses who reported that they were still undergoing education (N = 544). We conducted the regression analyses separately for each of the ten plausible values and then averaged the regression coefficients across the ten analyses. The regression results for all three analyses – (1) Big Five only, (2) Big Five and highest educational qualification, and (3) Big Five, highest educational qualification, and labor force participation – are displayed in Table 2. When only the Big Five domains were included in the model, Emotional Stability was the strongest predictor of intelligence, followed by Extraversion, Openness, and Conscientiousness; both Extraversion and Conscientiousness were negatively correlated with cognitive ability. In addition to these linear effects, a small quadratic association of Conscientiousness with both cognitive abilities was also detected, which indicates that very high Conscientiousness scores, in particular, are associated with lower cognitive ability. Overall, the model explained four and six percent of the variance in the two domains, respectively. In a second step, we investigated the degree to which the Big Five explained additional variance over and above the highest educational qualification, as the primary predictor of cognitive ability. In addition, we analyzed whether the personality domains interacted with the educational qualification in predicting cognitive ability – in other words, whether the ability of persons with higher or lower education was more sensitive to personality effects. We therefore included the highest educational qualification in the regression. Results reveal that – for both domains – the effects of personality on cognitive ability decreased after controlling for education. Only Emotional Stability (0.06 and 0.10, respectively) and (low) Conscientiousness (−0.11/−0.07 and −0.09/−0.07, respectively) were found to have substantial associations with verbal and numerical ability over and above the educational qualification. In addition to these main effects for both skill domains, Openness significantly interacted with education level in predicting cognitive ability (−0.09 and −0.08, respectively), indicating that persons with a lower level of educational attainment benefit from high Openness with regard to their cognitive abilities. In a third step, we analyzed (a) whether the Big Five were still predictive of cognitive ability when both educational attainment and labor force participation (in full-time employment vs. not in full-time employment) were taken into account, and (b) whether this predictiveness varied across full-time employed and not full-time employed respondents. Results regarding the remaining main effects for personality after controlling for both educational attainment and labor force participation7 are slightly different for the two cognitive abilities: In the case of verbal ability, none of the Big Five domains proved to have a significant direct relationship with the ability level, whereas in the case of numerical ability, Emotional Stability (0.06) followed by Extraversion (−0.06) predicted a significant share of ability after controlling for education and labor force participation. However, the interactions of personality and labor force participation in predicting cognitive ability were absolutely parallel across both domains. Conscientiousness significantly interacted with labor force participation in predicting intelligence (−0.09 for both domains), which indicates that among persons in full-time employment, Conscientiousness is negatively associated with cognitive ability; no such relationship was found among non-employed persons or persons in part-time employment.8","The present study investigated the relationship between personality and cognitive ability. As previous studies have been criticized for investigating student populations only, we examined whether the previously found associations between cognitive ability and Openness and Emotional Stability and its negative association with Conscientiousness could be generalized to a heterogeneous adult population. By analyzing data from the German PIAAC and PIAAC-L surveys – using the competence estimates for literacy and numeracy as indicators of verbal and numerical cognitive ability – we also investigated the degree to which these associations replicated in a linguistic and cultural setting other than those featured in earlier studies that focused on US or British samples. Based on this heterogeneous sample, the magnitude of the variance explained by the Big Five personality domains is, overall, highly comparable with that reported by earlier studies in this field that were based on selective samples (see Furnham et al., 2007). In addition, our results clearly replicated the positive association typically found between cognitive ability and Emotional Stability; this replication was completely consistent across verbal and numerical ability. In addition, we were also able to confirm the positive relationship between cognitive ability and Openness. However, the latter effect was found to interact with the person’s level of educational attainment: High Openness was a predictor of cognitive ability only for persons with low educational qualifications. For highly educated persons, by contrast, no such relationship could be identified. Our results thus suggest that being open-minded and intellectually interested can be beneficial to the intellectual development of persons socialized in intellectually less stimulating surroundings – that is, persons who leave the education system early. Alternatively, it could be that persons with comparatively higher cognitive ability leaving the educational system early become more open-minded and curious, e.g. to retain intellectual stimulating surroundings. Theorists have debated possible explanations for the regularly found negative association between Conscientiousness and cognitive ability. One hypothesis that has been proposed is that this correlation is a methodological artifact caused by the sampling bias of previous studies. If this hypothesis that the negative correlation between intelligence and Conscientiousness applies only to highly educated college student populations were correct, an interaction between education and Conscientiousness should be found in predicting ability in a heterogeneous sample. We therefore investigated (a) the extent to which this correlation was replicated in data representing the full adult population and (b) whether we could identify this hypothesized interaction of education and Conscientiousness in this comprehensive data set. On the basis of these data, we were able to negate this hypothesis and to show that this negative association is not in fact caused by a sampling bias. Rather, in a heterogeneous population sample, too, there is a negative association between verbal and numerical ability and Conscientiousness. As suggested by Major et al. (2014), we also investigated quadratic effects and found a negative quadratic association between Conscientiousness and ability, which indicates that very highly conscientious respondents, in particular, show lower cognitive ability. In addition, our analyses revealed no interaction between Conscientiousness and education in predicting cognitive ability. Our results therefore support the assumption that there is a negative relationship between Conscientiousness and cognitive ability and contribute to further understanding this association. We could show that the relationship between Conscientiousness and cognitive ability is moderated by labor force participation and that the negative association between Conscientiousness and intelligence applies only to persons in full-time employment. Conscientiousness and intelligence are both highly relevant criteria for job success (e.g., Barrick & Mount, 1991; Gottfredson, 1997; Schmidt & Hunter, 1998). Our results can thus be interpreted as supporting the intelligence compensation hypothesis (Moutafi et al., 2003), which assumes that, on the labor market, people can compensate comparatively low cognitive ability with high Conscientiousness. By contrast, more cognitively talented persons fulfill their job requirements more easily and do not therefore need to be as conscientious. Alternatively, however, the interaction effect of labor force participation and Conscientiousness on cognitive ability found here might reflect personality differences among occupations. It could be the case, for example, that lower-skilled workers are more conscientious than high-skilled persons. Our study thus showed that the negative association between Conscientiousness and intelligence is not restricted to college student populations. Rather, our findings provide preliminary evidence that this association can indeed be found in the total population, albeit not in a uniform way: It is more pronounced among persons in full-time employment. Hence, our results support the assumption that low cognitive abilities can be compensated with high Conscientiousness. However, as the present study investigated the Big Five personality domains using a very brief instrument, further studies are needed that replicate the effects found here using longer Big Five instruments that also allow differential effects of the domain’s facets to be examined. Another limitation of the present study – or at least one difference between it and earlier studies – might be the fact that ability was assessed in a low-stakes setting, which may have affected the individual’s test motivation. As shown in a recent study (Duckworth, Quinn, Lynam, Loeber, & Stouthamer- Loeber, 2011), test motivation can have an impact not only on the test scores themselves but also on associations of ability with life outcomes, for example, with personality characteristics. In sum, our findings clearly replicate the simple positive association between cognitive ability and Emotional Stability and Openness and its negative association with Conscientiousness. In addition, we were able to show that the association with Openness is moderated by education insofar as only persons with a low level of education benefit intellectually from high Openness. Labor force participation moderates the negative association between cognitive ability and Conscientiousness, indicating that Conscientiousness is negatively linked to cognitive ability only among persons in full- time employment. Hence, our results contribute to understanding the associations between personality and cognitive ability."],["Background: The behavioral inhibition system (BIS) and behavioral activation system (BAS) are two neuropsychological systems hypothesized to underlie response to cues signaling potential reward and punishment, respectively, also in patient responses to chronic pain. Objectives: The aim of this study was to test these hypotheses by evaluating the relative contributions of BIS and BAS to the prediction of function in sample individuals with chronic musculoskeletal pain. Methods: 253 participants were administered a battery of questionnaires. Two linear regression analyses were performed to evaluate the contributions of BIS and BAS to the prediction of impairment and psychological function, and to determine if either or both moderated the effects of pain intensity on function. Results: After controlling for demographic factors, pain diagnosis, and characteristic pain intensity, BIS contributed significantly and independently to the prediction of pain-related physical impairment and psychological function. BAS activity had a significant and direct effect on psychological function only. No moderating effects of BIS or BAS on the association between pain intensity and function were identified. Discussion: The findings are generally consistent with a BIS-BAS 2-factor model of chronic pain, suggesting BIS and BAS activity as potential targets for chronic pain treatment. --------------------------------------------------------------------------------","Chronic pain is a major biopsychosocial problem worldwide. It has a negative impact on people's ability to exercise, engage in valued social and family activities, and maintain an independent lifestyle (Breivik, Collett, Ventafridda, Cohen, & Gallacher, 2006). Chronic pain also has a negative impact on psychological function domains, such as depression, anxiety, and perceived stress (Stubbs et al., 2016). However, pain does not have the same impact on everyone. The negative effects of pain are known to be influenced by a number of psychological factors, such as an individual's tendency to catastrophize about their pain (Craner, Sperry, Koball, Morrison, & Gilliam, 2017) and their trait anxiety sensitivity (Esteve, Ramírez-Maestre, & López-Martínez, 2012). Additional factors that have the potential to influence adjustment to chronic pain are the relative activation of two neurophysiological systems that have been hypothesized to facilitate approach and avoidance behaviors: the behavioral inhibition system (BIS) and behavioral activation system (BAS) (Jensen, Ehde, & Day, 2016). Gray's Reinforcement Sensitivity Theory (Gray, 1987; Gray & McNauhton, 2000) describes the BIS and BAS as neuropsychological systems that are activated in an automatic way in the presence of environmental or internal cues. Specifically this theory hypothesizes that BIS is activated in the presence of cues indicating the potential for punishment (e.g., pain). This system underlies and facilitates avoidance-related behaviors (e.g., withdrawal), emotions (e.g., anxiety), and cognitions (e.g., catastrophizing). On the other hand, BAS is activated in the presence of cues indicating the potential for reinforcement or the disappearance/omission of an expected negative stimulus. BAS activation facilitates approach-related behaviors (e.g., more activity, impulsivity), emotions (e.g., excitement, joy), and cognitions (e.g., self-efficacy; Bjørnebekk, 2007). Pain is associated with actual or potential tissue damage and its protective role often elicits attention and action, which occur by virtue of the withdrawal reflex it activates, the intrinsic unpleasantness of the pain experience, and the emotional anguish it can elicit (Woolf, 2010). A person's trait tendency for BIS or BAS to be activated in response to pain may therefore explain, at least in part, the variability observed in people's adjustment to pain, as reflected by measures of activity and psychological function (Renee & Cano, 2009). The BIS-BAS model of chronic pain (Jensen et al., 2016) proposes that pain is interpreted as an aversive or punishment-related stimulus by most people. This model therefore hypothesizes that more pain intensity would tend to result in activation of the BIS and subsequent negative psychological responses and physical impairment. In addition, and in support of this idea, significant associations between pain intensity and both impairment and distress are often found. For example, Saavedra-Hernánndez et al. (2012) showed that neck pain intensity is significant predictor of disability. Similarly, Moore et al. (2010) found that moderate and substantial pain intensity reduction resulted in improvements in many outcomes (sleep disturbance, depression, anxiety, and quality of life) such that they approached levels found in the normal (i.e., otherwise healthy) population. Thus, more pain intensity is hypothesized to result in (1) more BIS activation (2) less BAS activation behavioral activation and subsequent positive emotions (BAS inhibition). Moreover, because pain is an aversive or punishment-related stimulus, the association between BIS and BIS-related responses (as sensitivity to punishment system) and pain is hypothesized to be stronger than the associations between BAS and BAS-related responses (as sensitivity to reward system) and pain. In support of this idea, it has been found that cues that signal the occurrence of pain are more likely to increase the focus of attention on that cue, relative to “safety cues,” which result in a decreased chance that the person will experience pain (Van Damme et al., 2004) and that pain will interrupt behavior (Eccleston & Crombez, 1999). With respect to the relationship between BIS and BAS, a “separable subsystems” model (Corr, 2002; Gray & McNauhton, 2000) hypothesizes that the BIS and BAS work mostly independently. That is, individuals with greater BIS activity, compared with those with a less BIS activity, should be most sensitive to signals of punishment, regardless of their level of BAS activation; and individuals with greater BAS activity, relative to a less activity, should be most sensitive to signals of reward, regardless of their level of BIS activation. Thus, pain is thought to be a cue that directly activates the BIS and pain's impact on patient dysfunction (e.g., negative emotions and disability) is hypothesized to be mediated by BIS, at least in part, regardless of the level of BAS activity (Jensen et al., 2016). If pain influences BAS, then any of pain's negative effects on positive function (e.g., positive emotions and life engagement) would be expected to be mediated by BAS activity, separately and distinctly from any effects on BIS. On the other hand, a more recent “joint subsystems hypothesis” (Corr, 2002) postulates that BIS and BAS have the potential to influence each other's effects on both reward-mediated and punishment-mediated behavior. That is, these systems may work synergistically, such that the impact of one on function is influenced by the relative activation of the other. With this model, dysfunction is hypothesized to be greatest in people with both high BIS activation and lower BAS activation and vice versa (Corr, 2002). In support of this model, Corr (2002) found a significant BIS (Anxiety) x BAS (Impulsivity) interaction in reactions to experimental manipulations of punishment in a sample of volunteers recruited from a university population. However, to our knowledge, the potential moderating effects of BIS and BAS activation on their effects on patient function have not yet been examined in the context of chronic pain. The BIS-BAS model of chronic pain (Jensen, Ehde, & Day, 2016) hypothesizes that the two systems are distinct but not completely independent; thus, this model would hypothesize that significant BIS X BAS interactions predicting function might be found in some contexts but not others. Even though pain is hypothesized activate primarily BIS, it may also influence BAS to some degree, via two mechanisms. First, because BIS activation is hypothesized to inhibit BAS to some degree (but not completely), and vice versa, an increase in pain would be expected to inhibit BAS indirectly, via its effects on BIS. Second, because in some situations, pain may activate aggressive responses (a BAS “approach” response), an increase in pain has the potential to result in an increase in BAS activity in some settings and with some individuals (i.e., Muris, Meesters, de Kanter, & Timmerman, 2005). The combination of these two contradictory effects may act to result in an overall weaker association between pain and BAS activation. Thus, the BIS-BAS model of pain hypothesizes that experience of pain would result in (1) more behavioral inhibition and subsequent negative psychological function and (2) less behavioral activation and subsequent positive emotions. A greater tendency for engaging in approach behaviors, feeling of excitement and joy, and believing that one is capable of controlling pain is hypothesized to inhibit (although not necessarily completely eliminate) a tendency to avoid activities, experience fear, or have thoughts of helplessness. With respect to a possible BIS X BAS interaction effect, the BIS-BAS model of chronic pain hypothesizes that such interaction is possible in some contexts, but unlikely to emerge across all contexts. Existing research provides preliminary support for a BIS-BAS model of chronic pain (Jensen et al., 2016). For example, Jensen et al. (2017) found that patients with chronic pain scoring high in a tendency for BIS activation report more depressive symptoms. BIS has also been shown to moderate the associations between pain-related cognitions and psychological function. Specifically, individuals with chronic pain who endorse more BIS responding evidence stronger associations between kinesiophobia and depressive symptoms than those who endorse less BIS responding (Jensen et al., 2017). Moreover, a trait tendency towards BIS activation has been shown to be associated positively with pain catastrophizing (Muris et al., 2007) which is known to be associated with negative affect and disability in individuals with chronic pain (Quartana, Campbell, & Edwards, 2009). Also in support of the BIS-BAS model of chronic pain, Jensen, Tan, Chua, and BSoc (2015) showed that a higher frequency of severe headaches was associated with higher trait BIS and lower trait BAS scale scores in a sample of undergraduate students, with the association between BIS and pain stronger than that between BAS and pain. Consistent with this idea, Becerra-Garcia and Robles (2014) found that BAS was lower in patients with fibromyalgia, relative to a healthy control group. In addition, it has demonstrated that people with chronic pain have a reduced hedonic response to rewards, and this reduction is associated with smaller nucleus accumbens volume that is responsible of reward processing (Elvemo, Landrø, Borchgrevink, & Haberg, 2015). In part because of the fact that the BIS-BAS model of chronic pain is relatively new, research testing the model to determine its utility remains preliminary; more research is needed to evaluate the explanatory power of the model, and adapt it as needed based on empirical findings. Given these considerations, the aim of current study was to increase our understanding of the role that BIS and BAS responding may play in the physical and psychological function of individuals with chronic musculoskeletal pain. Based on the BIS-BAS model, we hypothesized that BIS activation and BAS activation would make significant and direct contributions to the prediction of physical impairment and psychological function (positive association with BIS and negative association with BAS), when controlling for demographic factors, pain diagnosis, and characteristic pain intensity. In addition, we hypothesized that BIS and BAS would moderate the association between pain intensity and the study criterion variables, such that those with more BIS and less BAS would evidence stronger associations between pain intensity and function. Finally, we examined the possible interaction between BIS and BAS as a predictor of function. A significant interaction would support the joint subsystems model (i.e., greater influence of BIS and BAS on the effects of each on function) with respect to chronic pain. On the other hand, if a significant BIS X BAS interaction did not emerge, this would support the separable subsystems model (i.e., less influence of BIS and BAS on the effects of each on function) in this context. Fig. 1 presents a graphic representation of the study hypotheses.","The study participants were recruited from two hospital pain units (the Hospital Costa del Sol Pain Unit and the Hospital Virgen de la Victoria Pain Unit, in Spain) and from a fibromyalgia association (“Asociación de Fibromialgia y Síndrome de Fatiga Crónica de Málaga AFIBROMA”, Spain). For the participants who were recruited from the hospital pain units, physicians in the units reviewed the clinical history of each potential participant, and invited them to participate if they met the study inclusion criteria. Interested participants were contacted by telephone to schedule an assessment. To recruit participants from the fibromyalgia associations, we contacted by phone with the chairpersons of associations and described the study to them. The chairperson then informed the organizations' members about the study via email, and interested members were invited to attend a meeting with research staff to hear more about the study. Those who remained interested following the meeting were enrolled in the study and scheduled for an interview for data collection. A total of 169 individuals were recruited from the pain units, and 84 individuals were recruited from the associations. Study inclusion criteria were: (1) being from 18 to 65 years old, (2) having a musculoskeletal pain problem for at least 3 months, (3) not having any other physical condition or illness in addition to the pain problem, and (4) not having a severe psychiatric disorder that would interfere with participation. After written informed consent was obtained, a psychologist met with the participants to obtain demographic information, pain and pain history information, and to administer the study questionnaires (described in the Measures section). The study procedures complied with the Declaration of Helsinki and received institutional review board approval by the University of Málaga Ethics Committee. Demographic variables ~~~~~~~~~~~~~~~~~~~~~ Participants provided basic information about their demographics including age, sex, marital status, highest level of education achieved, and employment status. Characteristic pain intensity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Characteristic pain intensity was assessed by asking participants to rate their current pain and worst, least, and average pain in the past two weeks on 0–10 numerical rating scales, with 0 = “No pain” and 10 = “Worst pain possible.” These ratings were then averaged into a single score representing characteristic pain intensity (Jensen, Turner, Romano, & Fischer, 1999). Trait BIS and BAS activity ~~~~~~~~~~~~~~~~~~~~~~~~~~ Trait BIS and BAS activity were assessed using the 20-item Sensitivity to Punishment and Sensitivity to Reward Questionnaire (SPSRQ-20; Aluja & Blanch, 2011). The SPSRQ-20 measures individual differences in trait tendency for BIS and BAS activation. Item are answered with a dichotomous “Yes” or “No” response, and are then summed to score into BIS and BAS scales (10 items each). A sample BIS item is, “Are you often worried by things that you said or did?” A sample BAS item is, “Do you like being the center of attention at a party or a social meeting?” The BAS and BIS scales demonstrated good (BAS) and excellent (BIS) internal consistency in the current sample (Cronbach's alphas = 0.81 and 0.91, respectively). Pain-related impairment ~~~~~~~~~~~~~~~~~~~~~~~ Pain-related impairment was assessed using the 30-item Impairment and Functioning Inventory for Patients with Chronic Pain (IFI-R; Ramírez-Maestre & Esteve, 2015). With the IFI-R, respondents are asked, first, if they performed a number of daily activities (e.g., sweeping the house, driving the car or dressing by themselves, or visiting friends) in the previous week. For each activity they did not perform, they were asked to indicate, yes or no, if they did not do the activity because of pain. A pain-related impairment score is then computed by summing the activities not engaged in due to pain; a higher score indicates more pain-related impairment. In this sample, the reliability of the impairment scale was good (Cronbach's alpha = 0.81). Psychological function ~~~~~~~~~~~~~~~~~~~~~~ Psychological function was assessed using the 5-item World Health Organization Well-Being Index (WHO-5; Bech, 1999). With the WHO-5, respondents indicate how they have been feeling over the last two weeks on a 0 (“At no time”) to 5 (“All of the time”) scale. Sample items include, “I have felt calm and relaxed” and “I have felt cheerful and in good spirits.” The internal consistency of the measure was excellent in the current sample (Cronbach's alpha = 0.90).","We first computed descriptive statistics to describe the sample. We then calculated Pearson correlations coefficients between the study variables to understand their univariate associations. Next, we examined the variables and their distributions for normality, homoscedasticity and multicollinearity to ensure that they met the assumptions for the planned regression analyses study (Tabachnick & Fidell, 2007). Finally, to test the study hypotheses we performed two multiple regression analyses (Cohen, Cohen, West, & Aiken, 2003), one for each criterion variable (i.e., pain-related impairment and psychological function). Given research that has shown that socio-demographic factors and pain diagnosis can influence important pain-related outcomes (e.g., Ando et al., 2013; Goldenberg, 2009; May, 2008), we planned to control for these factors in the analyses. In line with it, in each analyses, we first entered demographic (age, sex) and diagnostic group (fibromyalgia, low back pain, and limb [arm, hand, leg, or foot] pain, or other, dummy coded, being “other” the reference category) as control variables. We then entered characteristic pain intensity in step 2 and the BIS and BAS scale scores in in step 3. Finally, in step 4, we entered the BIS × Pain Intensity, BAS × Pain Intensity, and BIS × BAS interaction terms. The predictor variables (characteristic pain intensity, BIS score, BAS score) were centered prior to entry to avoid the biasing effects associated with multicollinearity that can occur when examining interaction terms. All analyses were conducted using the Statistical Package for Social Sciences (SPSS; Windows version 22.0, SPSS Inc., Chicago, IL). Sample characteristics ~~~~~~~~~~~~~~~~~~~~~~ Two hundred and fifty-three individuals participated in the study. They had a mean age of 52.51 years (SD = 9.85), and 206 (81%) were woman. Eighty-four (33%) reported a diagnosis of fibromyalgia, 75 (30%) of low back pain, 67 (26%) limb pain, and 27 (11%) other musculoskeletal pain problem. The mean pain duration was 10.06 years (SD = 12.23). Table 1 shows more details about the participants' characteristics. Descriptive analyses and correlations between variables ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Mean, standard deviations, and correlations among the study variables are presented in Table 2. The sample reported a characteristic pain intensity level that was moderate to severe, with a mean values of 6.32 (SD = 1.35; possible range, 0–10). The strength of the zero order associations between the predictor and criterion variables ranged from small (e.g., BAS with impairment, r = 0.16, p < 0.01; BAS with psychological function, r = 0.12, p < 0.05) to strong (e.g., BIS with impairment, r = 0.51, p < 0.01; BIS with psychological function 0.55, p < 0.01). With respect to assumptions testing, the skewness (range from −0.07 to 1.26) and kurtosis (range from −0.02 to −0.59) values did not exceed the standard cutoff of 3 (Tabachnick & Fidell, 2007) indicating adequately normal distributions for the study variables for the planned regression analyses. The lack of multicollinearity among the predictor variables was confirmed by variance inflation factors, as their values (range from 1.04 to 2.28 in both regression analyses) were substantially below the standard cutoff of 10 (Hair, Anderson, Tatham, & Black, 1995). Pain intensity and BIS and BAS activity as predictors of pain-related impairment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 3 presents the result of multiple regression analysis predicting pain-related impairment. As can be seen, and after controlling demographic variables (age and sex) and the diagnoses of the participants, we found that pain intensity contributed significantly to the prediction of pain-related impairment (R2 change = 0.05; p < 0.001). When pain intensity was controlled, BIS activity (β = 0.44, p < 0.001), but not BAS (β = 0.03, p = 0.671), made an additional significant contribution to the prediction of this criterion variable. However, none of the interactions made a significant contribution to the prediction of the criterion variable. Pain intensity and BIS and BAS activity as predictors of psychological function ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Both BIS activity (β = −0.53, p < 0.001) and BAS activity (β = 0.17, p < 0.001) made statistically significant and independent contributions to the prediction of psychological function, once demographic variables, pain diagnosis, and pain intensity were controlled (see Table 4). However, none of the interaction terms contributed significantly to the prediction of psychological function.","The primary purpose of this study was to evaluate the role that BIS and BAS may play in the physical and psychological function of individuals with chronic musculoskeletal pain as a test of the BIS-BAS model of chronic pain. The findings showed that, even after controlling for demographic factors, pain diagnosis, and characteristic pain intensity, BIS was independently and significantly associated with both pain-related impairment and psychological function. BAS was significantly and independently associated only with psychological function. Inconsistent with the study hypothesis, neither BIS nor BAS evidenced a moderating effect on the association between pain intensity and the function variables studied. These findings have important implications for understanding the potential role of BIS and BAS in adjustment to chronic pain. In line with the BIS-BAS model of chronic pain (Jensen et al., 2016), as well as previous research (Jensen et al., 2017; Muris et al., 2007), the results indicate that the BIS has a more predominant role in the prediction of function in individuals with chronic pain than the BAS. This is reflected both by the facts that (1) the BIS scale made significant and independent contributions to the prediction of both function criterion variables and (2) the association between BIS and both criterion variables was stronger than between BAS and the criterion variables. Also, Gray's theory (Gray, 1987; Gray & McNauhton, 2000) posits that BIS facilitates avoidance behaviors, and avoidance behaviors have been associated with chronic pain (Crombez, Eccleston, Van Damme, Vlaeyen, & Karoly, 2012). To the extent that future research identifies a causal role for BIS activation as influencing both physical and psychological function, these findings suggest that BIS activity may be a viable treatment target in chronic pain populations. Treatments which might decrease BIS activation (i.e., reduce avoidance behavior, maladaptive pain-related beliefs, reduce negative affect) may have benefits – at least in terms of individuals function – in people with chronic pain. As already noted, the study findings indicated that the BAS appears to be less important as a predictor of participants function than BIS, at least with respect to predicting impairment and psychological function. However, BAS did contribute unique variance to the prediction of psychological function in the study sample, over and above that accounted for by BIS. This role for BAS (reduced but still potentially important for some function domains) is consistent with the BIS-BAS model of chronic pain (Jensen et al., 2016) as well as the findings from other research. For example, Elvemo et al. (2015) showed that individuals with chronic pain had significantly reduced scores on reward responsiveness, but not reward drive (both as measured by Behavioral Inhibition/Behavioral Activation Scale; Carver & White, 1994), suggesting that having chronic pain may result in a reduction in hedonic responses to rewards. Moreover, research has shown that people with chronic pain have reduced nucleus accumbens volume (Elvemo et al., 2015); this area of the brain is implicated in the processing of reward, pleasure or positive reinforcement (Malenka, Nestler, & Hyman, 2009). If the current findings are replicated, it possible that, in individuals with chronic pain, BAS plays a greater role in emotional function and responding than behavioral responding. Thus, treatments that target BAS activity such as “positive psychology” interventions (Müller et al., 2016) would be expected to impact psychological function more than physical function, and so may be particularly important for individuals who endorse high levels of psychological dysfunction in response to pain. Research is needed to evaluate this hypothesis. The results did not support an interaction effect of BIS and BAS as predictors of function in our sample of individuals with chronic musculoskeletal pain. This findings are in line with the “separable subsystems” model (Corr, 2002; Gray & McNauhton, 2000), and inconsistent with previous human experimental research in undergraduate students (Corr, 2002). However, Corr (2002) notes that the “separable subsystems” model may be more appropriate in some contexts than others. For example, in the presence of strong appetitive/aversive stimuli, or in samples of individuals with “extreme” personality traits. The chronic pain context could potentially influence both of these characteristics. For example, chronic pain – especially when severe – can be viewed as a strong aversive stimuli. In addition, individuals with chronic pain may have “extreme personality” traits as a result of suffering for a long period of time (the mean pain duration of chronic pain in the sample of individuals who participated in this study was 10 years approximately). Thus, it remains possible that BIS X BAS interactions may emerge in samples of individuals with more mild pain, or who have experienced chronic pain for a shorter duration, consistent with the idea that BIS and BAS may work synergistically in some contexts and with some populations, but not others. Given that both BIS and BAS made significant and independent contributions to the prediction of psychological function, it is possible that overall treatment efficacy – at least on psychological function outcome domains – could be enhanced by targeting both an increase in BAS and a reduction in BIS activity as underlying mechanisms (instead of just one or the other). Research to evaluate the relative effects of existing (and new) treatments on each component of BIS and BAS could identify the potential “best combination” of treatments which maximally influence (reduce) behavioral avoidance, negative/maladaptive pain beliefs, and negative affect, and also influence (increase) approach behaviors, adaptive pain beliefs, and positive affect; such treatment combinations could potentially be more effective than treatments that target only BIS- or BAS-related domains. We had hypothesized that BIS or BAS levels could potentially moderate the association between pain and the criterion variables studied here. However, this hypothesis was not supported by the findings; BIS and BAS appeared to have direct effects on function that did not vary as a function of pain severity. However, it remains possible that BIS might increase the vulnerability of people to the consequences of pain, and/or BAS might provide individuals with more resources to help them when faced with the challenges associated with pain, even if these effects are similar across all levels of characteristic pain intensity levels. This possibility provides further support for the need to evaluate the potential benefits of treatments which effectively target and reduce BIS activity and increase BAS activity in individuals with chronic pain. A number of limitations should be considered when interpreting the current findings. First, we only used self-report measures in the study. Thus, it is possible that shared method variance may have influenced the findings, resulting in stronger associations between the predictors and criterion variables than would have occurred had different sources been used as sources for the study variables. Research that examines the associations between self-report measures of BIS and BAS and objective measures of patient function (e.g., actigraph measures of activity, significant other observations of patient behaviors) would be useful. A second limitation is that the study design was cross-sectional. As a result, it is not possible to draw causal conclusions from the associations found. Future research is needed determine the effects of changes in BIS or BAS (e.g., as might occur with treatments that target BIS and BAS activity) and subsequent patient function. Third, the sample included a larger number of women than men. Although the ratio of women is greater than of men in this health services, a sample with more men as well as with other type of chronic pain diagnoses is needed to evaluate the generalizability of the current findings. In addition, the most recent version of Gray's reward sensitivity theory includes a third system – a fight- flight-freeze system (FFFS) – that we did not evaluate here. We had a number of reasons for not including an examination of the FFFS in the current study. First, the goal of the current study was to evaluate the BIS-BAS model of chronic pain (Jensen et al., 2016), which does not take into account the FFFS, because the FFFS system is rarely stimulated in most situations; fight or flight responses do not usually occur on a daily basis. Thus, excluding this system allowed the model to keep more focused on those factors that predict day-to-day responses. In addition, like our 2-factor model (Jensen et al., 2016), none of the many other 2-factor models which incorporate the BIS and BAS or systems very much like them (Elliot, 1997; Gray & McNauhton, 2000; Harmon-Jones, 2004; Watson, Wiese, Vaidya, & Tellegan, 1999), also do not incorporate the FFFS as a part of their model. Moreover, scientists, including McNaughton and Corr (2008), note that the association between BIS and FFFS is very close. FFFS activation is thought to be preceded by BIS activation and they can therefore be combined into a single “punishment sensitivity” factor of personality (Corr, 2009). Thus, the distinction between the FFFS and BIS is thought be less than that between the BIS and the BAS. Also, to our knowledge, no one has yet developed a measure of FFFS activation that is comparable to the commonly used BIS/BAS measures, including the one used in the present study. Future research is needed to evaluate if, and how, the FFFS and other systems may interact with the BIS and BAS to impact adjustment to chronic pain. Despite the study's limitations, the findings provide new information regarding the role that BIS and BAS have as predictors of function in in individuals with chronic pain. The results are generally consistent with a model that argues that both BIS and BAS may explain differential responses to pain, and that BIS may play a larger role than BAS (Jensen et al., 2016). The findings also suggest that BAS may be only meaningfully important with respect to psychological function, while BIS may play roles in both impairment and psychological function. Additional research is needed to evaluate the generalizability of these findings in other chronic pain populations, as well as to study the potential causal role that BIS and BAS may play in adjustment to chronic pain. In addition, based on these findings, further research could analyze in detail how, and through what mechanisms, BIS and BAS are related to psychological function and emotional regulation in patients with chronic pain. In the same way, they could evaluate how the systems interact in the activity patterns of this type of patients (excessive avoidance or excessive persistent). Also, we recommend that future researchers incorporate the evaluation of additional subsystems when possible (e.g., as measures of these are developed) for understanding, and treating, chronic pain and its negative impact.","This work was supported by the Spanish Ministry of Economy and Competitiveness [PSI2013-42512-P]; and the Spanish Ministry of Education, Culture, and Sports [grant number FPU13/04928]."],["This study identifies the incidence and development of disabled children's problem behaviors (i.e., conduct, peer, hyperactivity, and emotional problems) during the early years. Using the Millennium Cohort Study, a nationally representative UK study, and a measure of disability anchored in the UK legal definition, we estimate growth curve models tracking behavior problems from ages 3 to 7. We examine whether disabled girls’ and boys’ behavior differs from their non-disabled peers, and whether it converges with or diverges from them over time. We investigate whether parenting and the home environment moderate associations between disability and behavior. We show that disabled children exhibit more behavior problems than non-disabled children at age 3, and their trajectories from ages 3 to 7 do not converge. Rather, disabled children, particularly boys, show increasing gaps in peer problems, hyperactivity, and emotional problems over time. We find little evidence that parenting moderates these associations. --------------------------------------------------------------------------------","The emergence of problem behavior during the early years may set children upon unfavorable developmental trajectories. This is particularly true in the case of early externalizing behavior problems (i.e., hyperactivity, aggression), which may lead to continued problems and poor academic achievement (see e.g., Campbell, Shaw, & Gilliom, 2000; Hinshaw, 1992). Boys and girls tend to exhibit problem behavior differently, with higher rates of externalizing problems documented for boys and, to some extent, more internalizing problems (withdrawal, depression) for girls (see, e.g., Baillargeon et al., 2007; Campbell, 1995; Keenan & Shaw, 1997; Midouhas, Kuang, & Flouri, 2014). Past research has shown that disabled children are more likely than their non-disabled peers to present behavior problems, including social and peer problems, conduct problems and oppositional behaviors, attention difficulties and hyperactivity, and internalizing problems, and that their problems are more likely to be within the clinical range relative to their peers (Alloway, Gathercole, Kirkwood, & Elliott, 2009; Baker et al., 2003; Eisenhower, Baker, & Blacher, 2005; Emerson & Einfeld, 2010; Landa, Gross, Stuart, & Faherty, 2013). Yet, we know little about the extent to which associations between disability and behavior are linked to children's developmental stage and whether they attenuate or intensify around the time of school entry. We know from decades of research the critical nature of the early years, in which both genes and the environment—and the interplay between the two—set into motion the development of brain structures that affect children for the rest of their lives (Shonkoff & Phillips, 2000). More proximally, children's development up to age 3 provides the building blocks for the increasingly complex social behaviors, emotional maturity, problem solving ability, and early literacy and numeracy skills that are critical leading up to school entry. For some children, early behavioral problems are temporary, resolved over the normal course of development, while for others they persist or even intensify in the early school years. School entry represents an expansion in children's developmental ecology from the primacy of parents and the home environment to incorporate the school context and peers. Whether disabled children's behavioral development tracks that of their non-disabled peers over the first few years following this transition to school is an important empirical endeavor, a better understanding of which will help to inform the timing of interventions for disabled children. A description of disabled children's early behavioral trajectories across four important domains of behavioral development is the first contribution of this paper. Our current understanding of the association between disability and behavior is limited by the focus on particular impairments or conditions and reliance on small-scale, localized studies, both of which hamper generalizability. Many common proxies for disability in UK-based studies, such as identification with special educational needs (SEN), may confound the measurement of disability with the measurement of problem behaviors (Keil, Miller, & Cobb, 2006; Keslair & McNally, 2009; Powell, 2003). Here, instead, we exploit an overarching measure of disability anchored in the UK legal definition, itself informed by the social model of disability, which distinguishes the impairments themselves from the societal conditions under which they become disabling (Oliver, 1990). Our measure, which takes account of the contextualized nature of limitations or impairments, was developed from the data in consultation with the leading UK child disability experts and validated against known correlates of disability. This measure defines disability as both longstanding and limiting daily activities (longstanding limiting illness; LSLI), in line with guidance on the legal definition. It incorporates long-term health conditions, mental health problems, and sensory impairments, among others, enabling us to capture a wide range of disabling conditions experienced by a nationally representative sample of young children in England. Use of this measure improves our understanding of the associations between disability, rather than specific impairments or conditions that may or may not be limiting, and behavior, the paper's second contribution. Given the importance of the family and home environment for young children's behavioral development, supportive and enriching experiences in the home could help mitigate the development of behavior problems for young disabled children. On the other hand, given increased levels of parenting stress associated with parenting a young disabled child (Baker et al., 2003; Hastings, 2002; Neece, Green, & Baker, 2012), it may be that less favorable family climates exacerbate differences in behavior problems between disabled and non-disabled children. To our knowledge, despite the wealth of research attesting to the importance of home environment on children's development, and the ways in which it can mitigate socio-economic disadvantage (see e.g., Siraj-Blatchford, 2010), research has not examined whether family environments promote greater convergence or divergence of behavioral trajectories between disabled and non-disabled children over time. The paper's third contribution is to investigate the moderating role of parental warmth and harshness and the home learning environment on disabled children's behavioral trajectories; notably to better understand which aspects of parenting and the home environment attenuate or exacerbate which problem behaviors. Using longitudinal data from the UK Millennium Cohort Study (MCS), a large, nationally representative sample of children born in 2000–2001, we examine four problem behaviors: conduct problems, hyperactivity, peer problems, and emotional symptoms. These distinct types of problem behavior have been shown to be important for children's development, and they may present differently over time for disabled and non-disabled children. We address the question of whether young disabled children growing up in England experience more behavioral problems, and in which domains, than their non-disabled counterparts at age 3, and if any initial gap in behavior widens between the ages of 3 and 7. Finally, we examine whether differences in behavioral trajectories are contingent on three aspects of parenting and the home environment. Behavioral problems and disabled children ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A large body of research attests to specific trajectories associated with the four types of childhood behavior problems (i.e., conduct problems, hyperactivity, peer problems, emotional symptoms), with the preschool and initial school years considered the time when most children learn to control early problematic behavior, particularly externalizing behaviors (Bongers, Koot, van der Ende, & Verhulst, 2003; Broidy et al., 2003; Campbell et al., 2000; Fanti & Henrich, 2010; Tremblay et al., 2004). While conduct, hyperactivity, and peer problems typically decline over this time (Flouri, Midouhas, & Joshi, 2014; Midouhas et al., 2014), emotional symptoms tend to be stable or increase (Bongers et al., 2003; Leve, Kim, & Pears, 2005; Midouhas et al., 2014). The exception to this general pattern is a small subset of children, comprising more boys than girls, who display high levels of physical aggression that persist (Broidy et al., 2003; Campbell et al., 2000; Tremblay et al., 2004). These studies do not, however, distinguish between disabled and non-disabled children. Studies that have explored the relationship between disability and behavior in the early years have shown that, relative to non-disabled children, disabled children experience more total behavioral problems, more serious and clinically significant problems, and more persistent problem behavior (Alloway et al., 2009; Baker et al., 2003; Eisenhower et al., 2005; Emerson & Einfeld, 2010; Guralnick, Hammond, Connor, & Neville, 2006; Landa et al., 2013; Midouhas, Yogaratnam, Flouri, & Charman, 2013), suggesting that general declines reported for conduct, hyperactivity, and peer problems in the early school years may occur later or not at all for disabled children. Further, a recent study found that disabled children were particularly susceptible to increases in internalizing symptoms (Hauser-Cram & Woodman, 2016). These findings are largely based on small, non-representative cross-sectional samples and tended to focus on one particular type of impairment and more global problem behavior (rather than specific types). While researchers have used the MCS, the data source used here, to explore links between disability and children's behavior (see e.g., Emerson & Einfeld, 2010; Midouhas et al., 2013), it has not previously been used to classify young children according to criteria aligned with the UK legal definition of disability, nor have disabled children's early behavioral trajectories been examined, focusing on the time leading up to and following school entry. Exploring four behavioral trajectories across a representative sample of children from England allows us to assess how disabled and non-disabled children may differentially respond to school entry. A potentially important element in understanding the behavioral trajectories of young disabled children is the role of parenting and the home environment. A large body of research has demonstrated that parenting characterized by high levels of warmth, cognitive stimulation and clear limit-setting is associated with favorable emotional and behavioral outcomes for children, with the opposite findings for parenting characterized by harsh, arbitrary discipline or emotional detachment (Baumrind, 1966; Belsky, 1999; Berlin & Cassidy, 2000; McLoyd, 1998). Parents can also provide materials and experiences within the home environment, such as reading and other learning activities that promote children's early behavioral development (de la Rochebrochard, 2012; Hall et al., 2013; Kelly, Sacker, Del Bono, Francesconi, & Marmot, 2011; Kiernan & Huerta, 2008). Yet, parenting a disabled child may yield less than optimal parenting behaviors. Parents of disabled children exhibit higher levels of stress, more coping difficulties, and more conflict than other parents, which may lead to increased child behavior problems over time (Baker et al., 2003; Eisenhower et al., 2005; Herring et al., 2006; Neece et al., 2012; Totsika, Hastings, Vagenas, & Emerson, 2014), although these studies did not differentiate between type of problem behavior. Parents' ability to parent positively depends, in part, on whether they can recognize and interpret their children's behavior and emotional states, which may be difficult with disabled children (Howe, 2006). Some parents successfully adapt to having a disabled child and are able accommodate their special needs, while others face continued challenges to their competence and confidence as parents, becoming stuck in negative interaction patterns (Bailey et al., 2006; Sanders, Mazzucchelli, & Studman, 2004). Unfavorable parenting behaviors, such as unresponsiveness, harsh discipline and negative control exacerbate both externalizing and internalizing behavior problems for disabled children (Campbell et al., 2000; Gilliom & Shaw, 2004), while positive parenting behaviors may buffer them from the development of future problems (Ellingsen, Baker, Blacher, & Crnic, 2014; Hauser-Cram & Woodman, 2016). One UK study found that parent-child relationship quality was a stronger predictor of young disabled children's global behavior problems at age 5 than was discipline or assessments of the family environment (Totsika et al., 2014). The present study aims to expand on the extant research to examine whether different aspects of parenting may have distinct influences on particular behavior problems from ages 3 to 7. A better understanding of these nuances could help inform the timing and content of interventions to support families with disabled children (Bailey et al., 2006). The current study ~~~~~~~~~~~~~~~~~ The present study explores the development of disabled and non-disabled children's internalizing and externalizing behavioral problems over the early years and entry into school, a time of rapid growth and development when children's developmental ecologies expand well beyond their home environments. Using data from a large-scale, nationally representative sample of children living in England, we are able to include a range of relevant child and family background characteristics. The study capitalizes on the longitudinal nature of the dataset, which is critical for understanding whether any early differences in behavioral problems between disabled and non-disabled children are stable, decrease or increase over time. The paper addresses the following questions: (a) Are there differences in rates of behavior problems, specifically conduct, hyperactivity, peer, and emotional problems, between disabled and non-disabled children at age 3? (b) Are observed patterns of development of behavioral problems moderated by child sex? (c) Do gaps in behavior between disabled and non-disabled boys and girls converge (decrease), diverge (increase), or stay constant from age 3 to age 7? (d) Are the observed patterns of behavioral development robust to the inclusion of family characteristics and parenting behaviors? and (e) Does growing up in positive and stimulating early home environments moderate any divergence in trajectories between disabled and non-disabled boys and girls? Our measure of disability, longstanding limiting illness (LSLI), aligns most closely with UK disability legislation namely the Disability Discrimination Act, 1995, which was subsequently incorporated in the Equalities Act, 2010. While it does not precisely reflect the terminology of the legislation or the guidance on interpretation of “longstanding,” it provides an approximation that matches the key elements of the law. By contrast with the medical model, which has dominated most extant research, our definition has its roots in the social model of disability (Oliver, 1990), which regards disability as the ways in which societal organization limits those with an impairment, rather than viewing the impairment itself as inherently limiting. Adhering to the social model enables us to perceive behavioral “problems” as manifestations of how social norms limit disabled children, thus linking LSLI to behavioral problems and their development over time. From the existing literature, we develop the following hypotheses. First, we expect that disabled children will exhibit higher initial levels of conduct problems, hyperactivity, peer problems, and emotional symptoms at age 3 than their non-disabled peers. While we expect that externalizing problems (conduct problems, hyperactivity) will decrease over time for all children, we expect that differences between disabled and non-disabled children, particularly boys, will become more pronounced from around the time of school entry (around 4.5 years in England). Given the ways in which children respond to difference and the fact that schools may enhance the potentially disabling environment for children (Baker & Donelly, 2001; Chatzitheochari, Parsons, & Platt, 2016; Connors & Stalker, 2006), we expect disabled boys and girls to exhibit increased peer problems over time relative to their non-disabled peers. Our hypotheses concerning emotional symptoms are more tentative, but in line with previous research, we expect stability or small increases in emotional symptoms over the early years, and that they may increase most for disabled girls. Given the importance of family environment for disabled children (Baker & Donelly, 2001) and the stresses for parents in families of disabled children (Dowling & Dolan, 2001), we expect that warm parenting and enriching home environments will lead to more convergence over time in disabled children's behavioral trajectories, particularly for conduct problems, hyperactivity, and emotional symptoms. Harsh parenting will likely only moderate the association between disability and externalizing symptoms (i.e., conduct problems, hyperactivity).","We use data from the longitudinal Millennium Cohort Study (MCS). This large-scale, multidisciplinary, nationally representative study follows approximately 19,000 babies born to families living in the UK between September 2000 and January 2002 (Plewis, 2007). The sample population was drawn from all live births in the UK over this period, which were registered for universal child benefit. Participants were selected from a random sample of electoral wards, disproportionately stratified to ensure adequate representation of all four UK countries, deprived areas, and areas with high concentrations of Black and Asian families. Probability weights available for both whole UK and separate country analysis ensure that oversampled groups are represented according to population proportions. Families have been surveyed when children were aged 9 months (wave 1), and 3, 5, 7, 11, and 14 years (wave 2–6). At each survey, the child's main caregiver (primarily mothers) and their partner (primarily fathers) were interviewed and carried out self- completion questionnaires. Physical measurements and cognitive assessments of children have taken place since age 3. We use data from the main caregivers' interviews at the first four waves of data collection, and children's cognitive assessments at age 3 (Centre for Longitudinal Studies, Institute of Education, University College London, 2012a, 2012b, 2012c, 2012d). Sample attrition occurred over the course of the study: 72% of the original sample was surveyed at age 7. We employed the relevant survey weights for analysis of separate UK countries (see below) at the fourth survey (age 7). These weights incorporated adjustment for initial non-response at wave 1 and for differential non-response over time. While weights do not fully resolve the potential bias introduced by differential attrition, comparison of the initial wave characteristics of the analytic sample with those of the original respondents revealed that, while the analytic sample tended to be more advantaged, the differences were not sufficient to imply substantial bias in estimates of relationships between variables in multivariate analysis controlling for these characteristics (see Appendix A; Wooldridge, 2007). Moreover, our approach is consistent with extant research on behavioral development using the same study (e.g., Fitzsimons, Goodman, Kelly, & Smith, 2017; Flouri et al., 2014; Midouhas et al., 2014), and thus facilitates direct comparison. Analytic sample and exclusions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We restricted our sample to the approximately 60% of MCS families living in England, since education and public health systems vary across the countries of the UK. We further restricted our sample to those families who both took part in the first four waves of data collection, including completion of the main caregiver interview and self-completion questionnaire. At wave 1 when children were 9 months, 11,533 families lived in England. Of these, 7387 (63%) took part in the first four waves of data collection, and around 6300 completed main caregiver interviews and had non-missing data on other key variables. Due to small variations in missing data on behavior, we have sample sizes ranging from 6277 to 6313 for the behavioral outcomes of interest. In terms of inclusion, we found that families with disabled children were as likely to have been continuously involved in MCS as families with non-disabled children (analysis available on request). Dependent variables ~~~~~~~~~~~~~~~~~~~ Our dependent variables are the four “problem” subsets of the parent-reported Strengths and Difficulties Questionnaire (SDQ). The SDQ is a brief behavioral screening tool for 3- to 17-year-olds that has been widely validated cross-nationally and cross-culturally for use in non-clinical settings (Goodman, 1997; Goodman, Meltzer, & Bailey, 1998). The full scale is provided in Appendix B. Parents indicated how true each of 25 attributes (both positive and negative) were of their child over the 6 months preceding the age 3, 5, and 7 interviews (waves 2–4), ranging from “not true” (0) to “certainly true” (2). Four problem scales (each comprising five items)—conduct problems (e.g., “often has temper trantrums or hot tempers”; Cronbach's α = 0.56–0.68), peer relationship problems (e.g., “picked on or bullied by other children”; Cronbach's α = 0.47–0.58), hyperactivity/inattention (e.g., “constantly fidgeting or squirming”; Cronbach's α = 0.71–0.78), and emotional symptoms (e.g., “has many worries, often seems worried”; Cronbach's α = 0.54–0.66)—were created by summing item scores (positive attributes were reverse coded), with a higher score representing more problems. While some of the alphas appear low for problem scales at certain ages, the subscales are well validated and extensively implemented, including in studies using the same data (Midouhas et al., 2013). More recently, and Flouri and colleagues (Flouri, Midouhas, & Narayanan, 2016) investigated the scales using structural equation modeling and concluded that the individual items loaded well on their latent constructs. Following standard practice (see e.g., Midouhas et al., 2014), we model the scores as continuous outcomes. Disability Disability was measured based on children's exposure to a longstanding limiting illness (LSLI) at 3, 5, or 7 years determined by two successive questions asking parents if: (a) the child had a longstanding illness, and (b) whether that illness limited daily activities. We conducted detailed exploratory analysis, including known correlates of disability, such as parental education, income, and employment status and developed our measure in discussion with the Council for Disabled Children, who provided insight into the meaning of changes in LSLI status across waves. We also conducted sensitivity analyses using special educational needs (SEN) and developmental delay (at 9 months) as alternative measures of disability. On this basis, we developed an indicator variable identifying children as disabled if they had an LSLI at one or more occasions between ages 3 and 7. LSLI included long-term health conditions, such as type 1 diabetes or asthma; mental health problems; and impairments, such as partial sight. Using this definition, 10% of the sample was disabled. Among disabled children, asthma was the most common condition (35%), followed by ear disorders (13%) and dermatitis or eczema (12%). Note that in line with the social model of disability, it is not the condition that defines whether or not the child is disabled, but whether it is experienced as limiting their activities. In additional robustness analysis, we re-estimated the models excluding 101 children with specific conditions that might overlap with our outcome measures (ICD10 codes: F80-F89 = disorders of psychological development; F90-F98 = behavioral problems). Our findings were robust to this narrower specification (results available on request), so we retained the analytic sample previously described. We additionally estimated our models using a time varying measure of LSLI at ages 3, 5, and 7 as a sensitivity analysis. Results were consistent with the findings reported here (available on request). A range of child, family and parent-child relationship variables that have been found to be significantly associated with child behavior and/or disability in previous research were included in all analytic models. Child characteristics Child age was measured in fractions of years centered at age 3. Centering enabled us to establish initial differences in behavior problems between disabled and non- disabled children. We also computed a quadratic age term to measure non-linearity in the development of behavior problems over time. Child's sex was included in all models. Through estimating an interaction term with LSLI we aimed to isolate any differences in behavior problem trajectories between (disabled) boys and girls, and to investigate the extent to which child sex moderated the relationship between behavioral difficulties and disability in these early school years. Where the interaction between disability and child sex was not statistically significant (i.e., for conduct problems), we did not include it in the final specifications. The British Ability Scale Naming Vocabulary scale (Elliott, 1996), a widely used assessment of young children's expressive verbal ability, administered at age 3 (wave 2), and therefore prior to any school influences on cognitive development, was used as a control for children's cognitive ability. The child is shown a series of pictures (e.g., shoe, chair, scissors) and asked to identify the objects. Children are shown up to 36 pictures, depending on their performance. Ability scores created using item response theory ranged from 10 to 141 (Connelly, 2013; Rasch, 1960). Family background characteristics Low income (poverty) was measured as family household income < 60% of adjusted median household income, in line with the UK definition of relative poverty. As well as a time varying measure when children were 3, 5, and 7 years of age (waves 2–4), low income status at wave 1 was also controlled to capture the different circumstances disabled children are born into. Maternal work status was captured as a binary time varying variable (1 = in work), to capture the role of work independently of family income, and to allow for the fact that mother's work status might respond to child disability over time. Mothers' initial work status at wave 1 was also controlled. Family structure and cohabiting father's (mother's partner's) work status was captured in a single time varying variable with three values: father not present (single parent family), father present and not in work, and father present and in work. As well as the time varying measure, we controlled for wave 1 father's work and family structure to capture antecedent influences. Parental education was based on the highest qualification held by a parent living in the household at wave 1. Qualifications were grouped according to the national qualification framework levels (https://www.gov.uk/what-different-qualification- levels-mean/overview), and were rated on a 5-point scale, ranging from no qualifications (0) to NVQ4 or 5 (4), which equates to a Bachelor's degree or higher. To control for maternal mental health, we used a reduced form of the Malaise Inventory (Rutter, Tizard, & Whitmore, 1970). At wave 1, mothers considered nine indicators of depression/anxiety (e.g., Are you easily upset or irritated? ; Do you feel tired most of the time?), and indicated for each whether they “generally” felt these symptoms. Items were summed, with higher scores indicating increased probability of depression or anxiety (range = 0–9; Cronbach's α = 0.73). While this was our preferred measure of maternal mental health and, since it was measured at wave 1 when children were 9 months, captured antecedent influences on child behavior and its evolution, in a sensitivity analysis we estimated an alternative measure of time varying (ages 3, 5, and 7) responses to the Kessler scale (Kessler et al., 2003). Since we did not identify any substantive differences to our results using this alternative measure, and rates of non-response to the Kessler scale were higher than for other measures, we retained the Malaise Inventory at wave 1 in our final analysis. Research has documented both favorable (Hall et al., 2013) and unfavorable (Stein, Malmberg, Leach, Barnes, & Sylva, 2013) associations between early child care usage and young children's behavior problems. Child care in formal settings or by non-family members may reduce the direct influence of home context or mother's work status. We control for use of center-based child care (nursery) and, for comparison, non- kin family-based child care (childminder) at age 9 months, relative to using neither of these external child care settings. Parenting When children were 3 years old (wave 2), parents reported on how frequently they engaged their child in six educational activities: going to the library (“not at all” to “once a week”), and reading, painting and drawing, being taught letters, being taught numbers, and singing, reading poems, or rhyming (“not at all” to “everyday”). Items were summed to create a home learning environment scale (M = 25.8, SD = 7.39, range = 0 to 42). This scale has been widely used (e.g., Chatzitheochari et al., 2016; Kiernan & Huerta, 2008; Parsons, Schoon, & Vignoles, 2014) and has shown strong links to children's cognitive and behavioral outcomes (de la Rochebrochard, 2012; Hall et al., 2013; Siraj-Blatchford, 2010). To capture parent-child closeness, parents' self-evaluation of how close they were to their child (“not at all” to “extremely”) at age 5 (wave 3) was used. As 69% of parents reported being “extremely” close to their children, we constructed a binary variable contrasting “extremely close” with all other responses. Harsh discipline in the home was captured at age 5 (wave 3), using seven items from Murray Straus's Conflict Tactics Scale (Straus & Hamby, 1997). The scale sums the number of discipline measures used by the parent (e.g., ignore, smack, shout at, send to bedroom/naughty chair, take away treats, bribe) together with how frequently they are used (1 = “never” to 5 = “daily”). The total score ranged from 7 to 34 (M = 17.78, SD = 4.01; Cronbach's α = 0.71; Johnson, Atkinson, & Rosenberg, 2015). Descriptive statistics for all measures, broken down by whether or not children had an LSLI, are given in Table 1.","We estimated linear mixed models of children's behavior (Rabe-Hesketh & Skrondal, 2012; Singer & Willett, 2003). This analytic technique capitalizes on the repeated measures of behavioral outcomes measured at three time points, when children were approximately 3, 5, and 7 years. We examined whether disabled and non-disabled boys and girls start with similar or different behavior scores at age 3, and whether disability status is associated with converging or diverging trajectories over the early years, while controlling for potentially confounding family and child characteristics. We estimate their associations at baseline (age 3). We also explored whether parenting and the home learning environment moderated associations between disability and children's behavior problems at baseline and over the early years. Level 1 represents within-child change in behavior problems from 3 to 7 years, and Level 2 the between-child variation in the expected mean of children's behavior problems at age 3 (random intercept, β00) and linear change from 3 to 7 years (random slope, β10). We included a fixed quadratic on age to account for the curved shape of children's average trajectories (β20). We examined whether average age 3 behavior problems (β01) and change over time in behavior problems (β11) varied according to disability status, as well as whether these relationships were moderated by child sex (β03, β13). The components in the first set of parentheses represent the fixed effects, and the components in the second set represent the random intercept and linear slope for each child, reflecting between-child variation in problem behaviors (u0i), their development over time (u1i), and the error term (eij). The quadratic slope was fixed in all models. We estimated the growth curve models separately for each of the four problems in a series of nested models. In model 1, we estimated an unconditional model with age, age squared and the random intercept and slope. In model 2, we estimated a model with disability and sex as predictors, as well as two- and three-way interactions between age, sex, and disability, retaining only statistically significant interactions for the final model 2 specification. In model 3, we incorporated the full set of family and parenting characteristics (i.e., time varying family poverty, maternal work status, and family structure; wave 1 family poverty, maternal work status, family structure, child care usage, and maternal mental health; wave 2 child cognitive ability and home learning environment; and wave 3 harsh discipline and parental closeness) as covariates in addition to the final model 2 specification. Inclusion of the time varying covariates enabled us to examine the average difference in change over time in behavior problems according to families' poverty status, maternal work status, and family structure, respectively. The time invariant covariates enable us to examine the influence of these factors at baseline (age 3). Finally, in model 4, we included two- and three-way interactions between each of the key parenting variables (i.e., home learning environment, harsh discipline, and parental closeness), disability, and age to identify any moderation effect of parenting. The series of models is illustrated schematically in Table 2, alongside the related research questions. The models were estimated using the mixed procedure in Stata 13.1 (Rabe-Hesketh & Skrondal, 2012). We present the results from the unconditional model 1 in Table 3; and in Table 4 we present the initial disability model (model 2) and the full model (model 3) for each behavioral outcome. As none of the interactions between disability and parenting were statistically significant for any of the outcomes, model 3 is the final model. To illustrate the key results and demonstrate the magnitude of the differences, we plot the four behavioral outcomes by disability and sex, using model 3 estimates (Figs. 1–4). Overall development of behavioral problems ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 3 presents model 1. It shows that, in line with previous results, conduct, hyperactive, and peer problem behaviors tended to decrease over time from around age 3, with a slight increase from around age 6, as illustrated by the positive value for age squared. Emotional problems increased over time and at greater rate as the child aged (inflection point at age 3.5). The random effects parameters reveal that there was substantial idiosyncratic variation in behavior problems between children at age 3 and over time. Unconditional behavior trajectories for disabled girls and boys ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 4 illustrates the role of disability in shaping behavioral trajectories. It presents the results from the models including age, disability, and child sex, and their interactions (model 2), and the final specification, which includes all the controls and parenting measures (model 3). We see from model 2 that disability tended to be positively associated with problem behavior at age 3: a substantial difference amounting to between a third and four-fifths of a point for boys on the behavioral outcome scale (typically double or more than double the gap between girls and boys). Prior to school entry, disabled children demonstrated more challenging behavior than their non-disabled peers. Child sex was a significant predictor of behavior problems at age 3, as well as their trajectories from age 3 to age 7. When we look at changes in problem behavior over time by disability, we see that there was an increasing gap over time between disabled and non- disabled children for peer problems, hyperactivity (p < 0.10), and emotional symptoms. For conduct problems, the gap was constant. Conditional behavior trajectories for disabled girls and boys ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Model 3 in Table 4 shows that the relationships between disability and behavioral problems were largely robust to the inclusion of the full set of family and maternal background characteristics. To clarify the pattern of the trajectories for disabled and non-disabled children, and the scale of the gap between disabled and non-disabled girls and boys, Figs. 1–4 illustrate trajectories derived from the full model estimates, with family and maternal background characteristics set to their mean values. Fig. 1 illustrates the relatively steep decline in conduct problems over time with a slight upswing one and a half to two years after their entry into school: an inflection point at age 6.4. This pattern was tracked by disabled children but at a higher level. Girls faced lower conduct problems than boys across the early years, but the disability gap for both boys and girls was constant. The gap in conduct problems between girls and boys was smaller than that between disabled and non-disabled children, as the figure makes clear. For hyperactivity, there was a slight decline over time for non-disabled boys that leveled off somewhat before they reached age 6 (the inflection point was at age 5.7). As Fig. 2 and Table 4 show, non-disabled girls started from a somewhat lower level of problems and faced a steeper decline, resulting in an increasing gap relative to boys. Disabled boys started with a larger gap compared to disabled boys, than did disabled girls relative to non- disabled girls. Disabled girls tracked the steeper decline exhibited by non-disabled girls, but both disabled girls and boys exhibited a slightly growing gap relative to their non-disabled peers over time (see Fig. 2 and the interaction effects in Table 4). The result is that by age 7, disabled boys had substantially higher rates of hyperactivity than either disabled girls or non-disabled boys, even if not as high in absolute terms as they were prior to school entry. Non-disabled children's peer problems largely declined over time, though there was a slight upswing as they reached age 6 (the inflection point was 5.9 years; see Table 4 and Fig. 3). By contrast, peer problems for disabled children started increasing by the time of school entry, such that the gaps between disabled and non-disabled children were at their greatest by age 7. The gap between disabled and non- disabled boys was substantially greater than that between disabled and non-disabled girls from the outset, and increased at a faster rate. This left disabled boys experiencing exceptionally high rates of peer problems by age 7. Emotional problems increased for all children over the school years. Indeed, they had already started increasing prior to school entry (inflection point at 4.3 years). However, they not only started higher, but also increased faster for disabled children (Fig. 4), particularly for disabled boys. Although disabled girls were more at risk of emotional problems than boys at age 3, and experienced a sharper increase over time than non-disabled girls, disabled boys experienced even more of an increase in emotional symptoms so that by age 7 they had the highest levels. This illustrates a pattern of divergence between disabled and non-disabled children that was particularly marked for boys. Associations between covariates and behavioral problems at baseline (age 3) and over time ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Before we turn to consider the role of parenting practices in shaping behavioral trajectories, we briefly discuss associations between the individual and family background covariates and behavioral problems at age 3. Overall, the inclusion of the covariates in the full model accounted for only part of the differences in problem behavior between disabled and non-disabled children, even though many were associated with behavior. As expected, maternal poor mental health had a strong positive association with child behavioral difficulties. Socioeconomic characteristics, namely initial poverty status and lone parenthood were associated with greater levels of problems across the domains. Time varying maternal employment was also associated with fewer problem behaviors, net of time varying poverty status, which was itself not significantly associated with problems. Early experience of external child care was associated with fewer behavioral problems at age 3. Finally, children who were more cognitively able at age 3 were less likely to exhibit behavioral problems. Parenting and home environment and behavioral problems ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We expected that parenting and home learning environment would both be associated with behavioral problems, and that they would also moderate the negative association of disability and specific problem behaviors. We found that harsh discipline was positively associated with greater levels of all four types of behavioral problems at baseline. Similarly, parental warmth as expressed in their closeness to their child was associated with fewer behavioral problems. Home learning environment was negatively associated with conduct problems and hyperactivity, but not peer or emotional problems. We found no evidence for moderation at age 3 or over time: Estimated interactions between parenting, LSLI, and age were small and did not approach statistical significance.","The early development of problem behaviors can have consequences for children's later outcomes. While most children “grow out” of the problem behaviors that are common in early childhood, others do not, and may show elevated levels over time. Early problem behaviors that do not attenuate give way to later problems, including mental health problems, substance use, and even crime (Caspi, Moffitt, Newman, & Silva, 1996; Fergusson, John Horwood, & Ridder, 2005; Roza, Hofstra, van der Ende, & Verhulst, 2003). Given that disabled children are at risk of disadvantage in adulthood across a range of domains (Berthoud, 2008; Janus, 2009; Lindstrom, 2011; Loprest & Maag, 2003), it is relevant to ascertain whether these inequities start to develop in the early years and whether their behavioral trajectories in the early school years are the same or different to those of non-disabled children. If disabled children are experiencing higher levels and different trajectories of behavior problems than their peers during this time, it could indicate a critical point for intervention. Our findings provide clear and consistent evidence that in their early preschool years disabled children suffer from more challenging expressions of behavior than their non-disabled peers, and that in the early school years their trajectories diverge rather than converge. Two points stand out from our findings. First, disabled children (girls and boys) face sharper increases or slower declines in problem behaviors across three out of the four domains compared to non-disabled children: hyperactivity, peer problems, and emotional symptoms. Disabled children demonstrate greater levels of hyperactivity in the early school years, while their peer relationships and emotional health also suffer, perhaps directly linked to their externalizing behaviors. Second, disabled boys are particularly vulnerable to the development of problems over these critical early years, with steadily widening gaps between themselves and their non-disabled peers, as well as with disabled girls, again with the exception of conduct problems. Even in conduct problems where all children exhibited sharp declines over time, disabled children still exhibited higher levels of these potentially serious behavioral problems at the beginning of their school career. While the higher rates of problems and their increase (or stability) were in line with our expectations for externalizing problems (conduct problems, hyperactivity), our hypotheses were more speculative for internalizing problems (peer problems, emotional symptoms), though we did anticipate that peer relationships would present more of a problem for disabled children as they adjusted to the social environment of school. The size of the effects, the clear escalation for several of the outcomes, and the ways in which disabled boys, in particular, were affected was more surprising. The challenges associated with the expansion of children's developmental ecologies at the time they enter school, where they are in contact with new and potentially demanding or unsympathetic social environments may be exacerbated for disabled children, notably, disabled boys. Disabled boys exhibited concerning rises in both peer problems and emotional symptoms over time. We anticipated that emotional symptoms would stabilize or increase over time, and perhaps increase most for disabled girls. Less anticipated, however, was the substantial increase for disabled compared with non-disabled children, and for disabled boys even more than disabled girls, perhaps related to the social difficulties they face. Research has found that some disabled children display less competence in social situations than their peers; they interact with peers less frequently and are less well accepted within social circles (Carter & Hughes, 2007). Children often learn social skills through observing and participating in social interactions. If disabled children are not afforded opportunities for social learning, however, they miss out on important developmental inputs that may affect their social competence for years to come. This may be particularly true for boys, since by the time of school entry, boys and girls exhibit different peer group experiences. Boys display denser and more hierarchical social networks with less prosocial behavior and more verbal and physical victimization than girls (Rose & Rudolph, 2006). Just as all boys face greater challenges in adjusting to peers, our findings show how disabled boys are particularly at risk, with disability status contributing to widening the sex gap. These findings clearly indicate that the early years provide a relevant period for intervention to prevent the escalation and entrenchment of disabled children's behavioral problems that may be consequential for their subsequent development. Disabled boys who are already vulnerable in early socializing environments and developmental settings may benefit, in particular, from practices and contexts that foster easier transitions into school, such as enriched preschool programs that more explicitly focus on developing children's socio-emotional school readiness skills (Bierman et al., 2008; Odom & Wolery, 2003). Social integration activities, such as the formation of structured play groups within inclusive classrooms, and small group prosocial skills training may both facilitate more positive peer interactions for disabled children to ensure they benefit from the natural learning that occurs in day-to-day peer group interactions (Carter & Hughes, 2007; Odom et al., 1999; Ofsted, 2011). Early intervention in supporting young disabled children's emotional resilience and curtailing maladaptive behaviors may be important in preventing the development of more serious mental health problems, and in preventing high levels of school alienation that have been demonstrated for disabled children in their teenage years (McDougall, DeWit, King, Miller, & Killip, 2004). Our findings point to the ways in which social contexts can “disable” children (Connors & Stalker, 2006), and imply that children's educational plans should incorporate a socio- emotional component. Successfully transitioning to school requires children to follow the rules and demands of the classroom, consistently navigate increasingly complex peer interactions, and appropriately self-regulate their emotions—actions that have been argued to pose difficulties for disabled children (McIntyre, Blacher, & Baker, 2006). Strong partnerships between families and school personnel could help to ensure behavior management strategies and decisions regarding children's care are consistent (Newman, McEwen, Mackin, & Slowley, 2009; Odom & Wolery, 2003). We expected that the home environment and parenting might moderate the development of behavioral problems among disabled children. Our findings showed a significant main effect between the home learning environment and children's hyperactive and conduct problems, though it did not moderate the impact of disability. More structured home environments support fewer externalizing problem behaviors, then, but do not seem to protect against peer or emotional problems. While closeness and harsh discipline were implicated in, respectively, lesser and greater levels of all four problem behaviors, we also failed to find any evidence that they moderated the effect of child disability. Nevertheless, given the higher levels of behavior problems among disabled children, the creation of stimulating outlets within the home and the provision of parenting support through early intervention could potentially have payoffs for them. Similarly, since disabled children are more likely to be growing up in poverty, in lone parent families, and with greater levels of maternal ill-health, all of which impact on behavioral problems (even if these do not account for the divergence in behavioral problems between disabled and non-disabled children), economic and maternal mental health support may go some way to preventing disabled children starting out on their differential behavioral pathways in the early years. Overall, this study demonstrates how disabled children face ongoing difficulties with social relations and ordered social contexts that cannot adequately adapt to or accommodate their impairments (Barkley et al., 2002). The early years during which children learn to regulate their behaviors and adjust to social expectations appear to be differently experienced by disabled children, and particularly by disabled boys. Instead of modifying their behavioral problems, the demands of responding to greater social influences and institutional settings seem to bring emotional costs, and difficulties in behavior and social relations. Home environments may initially shelter children from these external influences, but the early environment alone does not protect against these wider social challenges for disabled children. Our findings on the divergence over the early school years in children's peer and emotional problems, in particular, reflect the ways in which “difference” can be enhanced on primary school entry. One of the routes to these poorer outcomes is the daily separation of disabled children from their peers, illustrating one of the ways in which certain forms of compensatory provision can inadvertently stigmatize disabled children, impacting not only their educational, but also their behavioral development (Webster & Blatchford, 2015). Targeted programs that promote peer inclusion may offer more promising routes for intervention (Holt, 2007). This implies a better understanding of school effects and the mechanisms driving those effects could be crucial to developing appropriate, “non-disabling” environments. That is, the ability to identify institutional factors, school cultures, and teacher behaviors that are more or less supportive for disabled children and more or less conducive to stabilizing or reducing problem behaviors has the potential reduce the escalation of these problems. While our data cannot identify positive school cultures nor such institutional practices, future quantitative survey research would benefit from incorporating relevant school-level factors into analyses linking disability to behavioral problems to provide evidence of potential mechanisms and identify contexts for further interrogation. Such further research would facilitate crafting appropriate recommendations for English school systems. Our study is not without its limitations, including our dependence on parental (mother's) report of family context, child behavior, child disability, and her own parenting. In particular, relying on mother's report for behavioral problems means we cannot directly link it to their behavior at school. For our measure of closeness we rely on a single item. There was some variation in the reliability of our measures, and internal consistency for peer problems at age 3 was below 0.5, which raises some concerns for the reliability of our results at that age. Our longitudinal analysis was limited to three data collection points across our four-year period of interest, restricting the analytical purchase. Further, our home environment and parenting variables were each available at one wave only, preventing us from exploring whether children's experiences in the home and with their families at particular points in their development affected their behavioral trajectories. Nevertheless, our study represents a contribution to our understanding of young disabled children's behavioral problems. Using a nationally representative sample of children living in England, it provides clear and consistent evidence that disabled children experience greater behavioral problems in their early years and that these do not dissipate—and, in some cases, increase—over time. Child behavioral difficulties can have far reaching consequences and hence, without appropriate support or intervention, young disabled children may face an accumulation of adverse consequences that serve to compromise their well-being in adolescence and adulthood."],["There is conflicting evidence regarding the development of expert face recognition, as indexed by the face-inversion effect (FIE; de Heering, Rossion, & Maurer, 2011; Young and Bion, 1981) potentially due to the nature of the stimuli used in previous research. The developmental trajectory of the FIE was assessed in participants aged between 5- and 18-years using age-matched and adult stimuli. Four experiments demonstrated that upright face recognition abilities improved linearly with age (presumably due to improved memory storage capacities) and this was larger than for inverted faces. The FIE followed a stepped function, with no FIE for participants younger than 9-years of age. These results indicate maturation of expert face processing mechanisms that occur at the age of 10-years, similar to expertise in other domains. --------------------------------------------------------------------------------","While adult face recognition is one of the most impressive human visual skills given the ability to differentiate and recognise many thousands of faces (Ellis, 1986), the face recognition abilities of children are poorer (Adams-Price, 1992; Blaney & Winograd, 1978). Adult face processing is assumed to be based on some form of expert processing mechanism (Farah, Wilson, Drain, & Tanaka, 1998) that may well be specific to the processing of faces (Kanwisher, Tong, & Nakayama, 1998). Poorer face-recognition performance in children could be due to generally poorer cognitive, attentional, and perceptual systems (see e.g., Crookes & McKone, 2009) or a specific deficit in this expert face processing (e.g., Carey & Diamond, 1977). Expert processing is typically referred to as configural processing and is made up of three components (Maurer, Le Grand, & Mondloch, 2002): processing of the first order relations (i.e., two eyes level, side-by-side and above the nose); the processing of second-order relational information (i.e., idiosyncratic deviations to the basic template; Carey & Diamond, 1994); and holistic processing, which is processing the face as a gestalt whole (Rossion, 2008), integrating the multiple sources of information (Farah et al., 1998; Searcy & Bartlett, 1996). While researchers may not entirely understand what drives expertise in face recognition, there is consensus that faces are processed differently to objects and this is likely due to some form of configural processing (Piepers & Robbins, 2012). There are many sources of evidence to suggest that expert processing is not based on second-order relational information (Burton, Schweinberger, Jenkins, & Kaufmann, 2015) but is based on this final form of configural processing, known as holistic processing (Hole, George, Eaves, & Rasek, 2002; Mondloch & Desjarlais, 2010). Configural processing is usually contrasted with the featural coding, which is not indicative of expertise. Featural coding is typically defined as the processing of individual features in isolation (see Cabeza & Kato, 2000; Tanaka & Sengco, 1997). One method typically employed to assess configural coding is that of inversion (e.g., Freire, Lee, & Symons, 2000). Indeed, Sergent (1984) suggests that configural encoding is what is disrupted by inversion, whereas featural encoding is far less disrupted by inversion (see also Lewis & Glenister, 2003). Therefore, the face-inversion effect (FIE) is a reliable index of expert face processing (Edmonds & Lewis, 2007; Gauthier et al., 2000; Yin, 1969). While there is no doubt that face recognition is expert in adults, there is a debate about when this expertise develops. One theory suggests that there is an early development of expert face processing mechanisms complete by the age of approximately 5-years (Crookes & McKone, 2009; Gilchrist & McKone, 2003; Want, Pascalis, Coleman, & Blades, 2003). While face recognition improves with age, this view suggests that age-related improvements in face recognition are explained by general improvements in the ability to attend and focus on the demands of the task (Crookes & McKone, 2009). These general improvements increase with age and continue to develop throughout childhood and adolescence (Betts, McKay, Maruff, & Anderson, 2006; Pastò & Burack, 1997; Skoczenski & Norcia, 2002). An alternative view is that the expert processing mechanisms do not develop until around 10 years of age (Carey & Diamond, 1977, 1994), consistent with the notion that many forms of perceptual expertise take approximately 10 years of practice and development (Akhtar & Enns, 1989; Brodeur & Enns, 1997; Enns & Brodeur, 1989; Ericsson, Krampe, & Tesch-Römer, 1993; Pearson & Lane, 1991). Recently, a view was put forward that there might be differential effects for the development of face perception and face memory, with face memory developing late and face perception developing early (Weigelt et al., 2013). Wiegelt et al. have presented evidence highlighting that the mechanisms that control expert face processing are not necessarily the same as expert face memory. Memory for faces, apparently, develops later than the perceptual expertise for faces. Memory for faces can be revealed through an increase in hit rate and response bias (as hit rate represents more faces being stored in memory and more efficient encoding) without affecting false alarm rate (which better reflects poorer encoding and poorer access to memory: Hills, 2012). General memory (Chi, 1977; Dempster, 1981; Kail, 1992) and memory for faces (Flin, 1980), does improve with increased age. Consistent with the view that the FIE is a measure of expert face perception, then there should be sufficient evidence to establish whether face perception develops early or late. If face perception expertise develops late, then one would expect that children would show a smaller FIE than adults. The evidence for this is mixed. Most authors agree that face recognition abilities improve approximately linearly with age, reaching an asymptote at the age of 12 (Feinman & Entwisle, 1976), 17 years (Ellis, Shepherd, & Bruce, 1973; Golarai et al., 2007; Lawrence et al., 2008; O’Hearn, Schroer, Minshew, & Luna, 2010), or well into adulthood (e.g., Germine, Duchaine, & Nakayama, 2011; Susilo, Germine, & Duchaine, 2013) depending on the stimuli set used.1 However, Flin (1980, 1985) has reported a face recognition performance “dip” at age 11 years2 (see also, Carey, 1978, 1981; Carey, Diamond, & Woods, 1980). The improvement in recognition for inverted faces may also be linear, but at a slower rate. Using a novel (for this field) statistical procedure, de Heering, Rossion, and Maurer (2012) found that performance on the Benton Face Recognition Test (Benton, Sivan, Hamsher, Vareny, & Spreen, 1983) correlated with age, between the ages of 6-years and 12-years. This correlation was stronger for upright than inverted faces indicating that the magnitude of the FIE also correlated with age. Such an improvement for upright faces over inverted faces potentially reflects that the expert face processing system has developed and there is a general improvement in task performance or face memory. Alternatively, this improvement may reflect a protracted development of expert face processing skills. Data from matching tasks reveal that children younger than 10 years of age are more likely to be affected by paraphernalia and pose changes than children older than 10 years and adults (Diamond & Carey, 1977; Ellis, 1992a, 1992b; Freire & Lee, 2001; Saltz & Sigel, 1967). These results indicate that children are not coding faces in the most effective configural manner. Indeed, six- and eight-year-old children do not show the FIE when tested in matching paradigms (Carey & Diamond, 1977; Hay & Cox, 2000; Joseph et al., 2006; Schwarzer, 2000) or recognition paradigms (Goldstein, 1975) indicating a greater reliance on featural processing (Schwarzer, 2000). In these studies, the FIE was found by some ten- year-old participants indicating some individual difference in the development of expert face processing which may sometimes mask effects when development is tested cross- sectionally. These results indicate a qualitative shift in the way children code faces at age 10 from an inexpert to expert mechanism (Baudouin, Gallay, Durand, & Robichon, 2010; Mondloch, Leis, & Maurer, 2006). However, other authors have reported that the FIE is apparent in three- (Carey, 1981), five- (Fagan, 1972; Flin, 1983), or seven-year-old children (Young and Bion, 1981, 1982) leading to parallel improvements in recognition skills (Itier & Taylor, 2004). Proponents of the view that the FIE does not increase with age highlight that the studies that fail to show an FIE in younger participants suffer from floor effects (Young and Bion, 1981). Nevertheless, an age-by-orientation interaction is often found in studies that claim there is an FIE in younger participants,3 indicating that the magnitude of the FIE increases with age (Brace et al., 2001; Carey, 1981; Carey & Diamond, 1994; Flin, 1983; Goldstein & Chance, 1964). Any effect of age on the magnitude of the FIE would indicate that children rely more on featural rather than configural coding (e.g., Hay & Cox, 2000).4 There are a number of methodological and statistical issues with the studies on children's face recognition. Firstly, most of the studies conducted on face perception are cross-sectional. There are significant individual differences in face recognition ability (Li et al., 2010), in the amount of holistic processing participants engage in (Wang, Li, Fang, Tian, & Liu, 2012), and in terms of how faces are encoded (Bobak, Parris, Gregory, Bennetts, & Bate, 2017; Mehoudar, Arizpe, Baker, & Yovel, 2014). This means that, potentially, effects reported in the literature are due to cohort effects which may be unduly influenced by individual differences in studies with relatively small sample sizes. For example, the recognition blip observed by Flin (1985) and the change in FIE observed by Schwarzer (2000) may reflect cohort effects. A longitudinal study of face perception exploring the development of expert processing mechanisms has yet to be conducted. A longitudinal study would rule out such cohort effects. Secondly, many of the studies that explore the FIE in children are underpowered. When testing multiple age-groups, the necessary increase in error degrees of freedom mean that it becomes much more difficult to detect significant differences in the FIE due to age. A within-subjects (and thereby longitudinal design) would improve the statistical power of such studies. Only when these issues are addressed can studies adequately address the mechanisms of face processing employed by children. Thirdly, there are several statistical issues with existing work on the development of face recognition. Certain tasks do not adequately control for floor and ceiling effects. Floor effects cause a task to be too difficult for younger children to complete. This makes distinguishing any effect of inversion very difficult (a similar argument has been made by Weigelt et al., 2013). One reason for floor and ceiling effects potentially is the use of age-inappropriate stimuli. Given that the own-age bias exists in face perception, in which participants show a larger FIE for own-age than other-age faces (Anastasi & Rhodes, 2005; Harrison & Hole, 2009; Hills & Lewis, 2011; Kuefner, Macchi Cassia, Picozzi, & Bricolo, 2008),5 in order to avoid floor effects, faces should be age-matched to the participants. In adults, the processing of other-group faces has been theoretically linked to not using the most expert configural processing system (Hugenberg & Corneille, 2009; Michel, Caldara, & Rossion, 2006; Michel, Rossion, Han, Chung, & Caldara, 2006), which lowers performance in such tasks. This problem means that standardised tests of childrenʼs face recognition performance, such as the Cambridge Face Memory Test – Children (CFMT-C; Croydon, Pimperton, Ewing, Duchaine, & Pellicano, 2014) might underestimate performance. Finally, even when tasks are sensitive enough to detect differences, there is a further statistical issue: overall performance in younger children is lower than that of older children. This means that any effects of inversion may be harder to detect in younger children. In order to address this, a relative measure of performance needs to be considered (Goldstein, 1965). A relative measure takes into account the fact that childrenʼs overall performance will be lower than that of adults and therefore allows for smaller differences in performance in children to be equated to larger differences observed in adulthood. Even with a relative measure of the FIE, there is a potential issue with a younger children showing a more limited range of performance than older participants. This can be assess by ensuring that the variances in performance are equivalent for all groups of participants (which is an assumption of parametric data in any case). An alternative method to control for this is to match performance of upright faces in all children by presenting different numbers of stimuli. While matching performance addresses the issues of poorer performance in children, it creates a confound: face recognition is made up of face perception and face memory, therefore manipulating the number of stimuli prevents an analysis of face memory. It is for this reason, a relative measure of the FIE is the more appropriate technique for measuring face recognition performance in children. This paper presents a solution to these problems in order to establish whether children show the FIE to a similar level as adults and, by extrapolation, utilise expert face processing. Here, the FIE, as a measure of expertise, was assessed using a standard old/new recognition paradigm in children (from 5- to 15-years-old) and adults. Three possible developmental trends are possible: The magnitude of the FIE may increase with age as a product of experience (developmental induction); The magnitude of the FIE may be constant throughout development if the effect is not based upon experience (due to early maturation); Finally, developmentally-late maturation may occur in which the FIE appears at a particular critical age. These trajectories lead to the increasing inversion effect, the constant inversion effect, and the “all or none” hypotheses. The first two hypotheses can be explained by an early maturation of expert face processing mechanisms and the final hypothesis is derived from a late maturation of expert face processing mechanisms. In this study, a relative measure of the FIE was used in order to control for poorer general cognition in younger children controlling for lower absolute performance of the younger children and therefore avoids floor effects. Absolute floor and ceiling effects were avoided by choosing a manageable number of stimuli for the youngest children tested that produced sufficient errors in the adult participants (Ellis, 1992b). While there remains the possibility that younger children's performance was more restricted than that of older participants which might obscure results, our data indicate that there was roughly equivalent variance in performance for the younger children and the older children suggesting that we did not have floor effects in our study. Floor effects were also avoided by testing own-age faces for all participants in Experiments 1 and 2. In order to show that development is occurring within participants (avoiding cohort effects) we also conducted a longitudinal study of face recognition by testing the same children at different ages in Experiments 2 and 4.","Experiment 1 was a cross-sectional experiment executed in a similar manner to de Heering et al. (2012).","ranged from 5- to 15-years of age and an adult sample. The primary purpose of Experiment 1 was to understand the developmental trajectory of face recognition and expert face recognition as measured by the FIE. In order to do this, correlations and curve fitting was conducted for upright and inverted face recognition and the FIE separately. This will distinguish the developmental trend of the face processing. Participants Participants were 440 children (198 male) aged from 5,7 years to 15,4 years and 40 adults (aged 18–23 years; 12 male). See Table 1 for participant details. The age groups chosen were linked to the school age they were studying in rather than strictly their age (see Hills & Lewis, 2011). Adult participants were recruited through a university whereas the children were recruited through local mainstream schools. All participants had normal or corrected vision and were ethnically White as indicated by parent-report (or self-report in the adult group). All of the children were considered typically developing by their schools. No participants were familiar with any of the faces used in this experiment.","Frontal-view photographs of White childrenʼs (aged from 5- to 18-years) faces were collected by a research assistant. Two images of each child were collected: one presented during the learning phase and one presented during the test phase. The images were quite similar, taken a few moments apart, but the facial expression was slightly different. This was done to reduce pictorial recognition (Bruce, 1982). Photographs of 32 (16 female) children were collected in each age group. All parents provided consent for these photographs to be used in this research project. The faces all had similar hairstyles, positioned in a frontal view in neutral or mildly happy expressions (this was randomised). All extraneous paraphernalia and the background were masked using Adobe™ Photoshop™. The stimuli were collected from local schools that were not used for the experimental testing. Participants only saw faces of their own age. To confirm that faces of one age category were not more dissimilar to those in another category, a pixelwise similarity comparison was made between the faces in each age group (Haushofer, Livingstone, & Kanwisher, 2008). This was not significant (p > .56). Similarly, attractiveness ratings did not differ across stimuli types (p > .39). While this cannot rule out stimulus differences across ages, it indicates that differences are not substantial. The faces were presented 100 mm by 110 mm dimensions in 72 dpi resolution in the learning phase and 150 mm by 165 mm in the test phase. An inverted version of each face was created using the rotate function in Adobe™ Photoshop™. These were presented using Superlab Pro 2™ Research Software using a Toshiba Tecra M4™ Tablet PC. Design Twelve participant groups were tested as determined by their school year group. Participants viewed both upright and inverted faces. The faces were counterbalanced between participants such that each face was a target as often as it was a distracter. The faces were counterbalanced such that they appeared upright as frequently as they appeared inverted. Faces were presented in a random order (i.e., there was no blocking of face type). The dependent variables were recognition accuracy, measured in terms of the Signal Detection Theory (SDT e.g., Swets, 1966) measure d' and response bias, measured in terms of the SDT measure, C. Reaction time could not accurately be measured due to the experimenter keying the responses (see procedure). The relative FIE was established using the formula:FIE = (d'U − d'I)/(d'U + d'I)where d'U is the recognition accuracy of upright faces and d'I is the recognition accuracy of inverted faces. This measure produces a relative measure of the FIE and controls for differences in overall accuracy between participant groups. Procedure Participants were tested individually in a quiet brightly-lit room in their school (or in a quiet laboratory for adults). Participants sat 50 cm from the computer screen. This screen was positioned away from the Experimenter such that he could not see the contents of the screen. Participants responded verbally and the experimenter entered the responses into a standard computer keyboard. Thus, the Experimenter could not influence the participants’ responses because the stimuli were presented in a random order, and he was unaware of what was on the screen, preventing demand characteristics. Additionally, this ensured that the Experimenter could ensure the participants followed the instructions appropriately. A standard old/new recognition paradigm was employed involving three consecutive phases: learning, distraction, and test. In the learning phase, participants were shown half of the set of faces of their own-age (N = 16) of which half were upright and half were inverted. Faces were presented centrally in a sequential random order. Participants were instructed to rate each face for how attractive they thought the face was using a 1–9 Likert-type scale, with the anchor points “ugly” and “beautiful” (Light, Hollander, & Kayra-Stuart, 1981). If a participant did not understand the scale, it was explained to them using alternative synonyms. The presentation of each face was response terminated: There were no differences in presentation duration across participant group (see Table 3, F(11, 468) = 0.24, MSE = 1464959, p = .995, ηp2 = 0.01). There was a 150 ms blank screen inter-stimulus interval between each face. Immediately after this presentation, participants were given some control questions. These were: “What is your first name?” “What is your surname?” “What is your gender?” “How old are you?” \"When is your birthday?\" “What school do you go to?” “What school year group are you in?” “Where were you born?”. If the participant did not understand the question it was explained using simpler synonyms. Only participant age, birthday, and gender were recorded. These questions took no longer than 60 s to administer. Following this, the participants were given the test phase. In this, the participants saw all 16 target faces and 16 distractor faces sequentially in a random order and made an old/new recognition judgement to each face. Each presentation was in the centre of the screen and was response terminated. Half the faces were upright and half the faces were inverted. Orientation of the target faces was matched from learning to test. The participants were told to be as quick and as accurate as possible. Between each face there was a blank screen for 150 ms. Once this phase was completed, the participants were thanked and debriefed. Total testing time was no more than four minutes per participant.","Old/new responses from participants were converted to hit and false alarm rates and these were used to calculate d' and C using the MacMillan and Creelman (2005) method. In addition to reporting the traditional null hypothesis significance tests, we also report Bayesian statistics throughout this work. Bayesian analysis has the advantage that it is not based on the evaluation of significance levels that can be interpreted incorrectly (especially regarding non significant results, see Rouder, Speckman, Sun, Morey, & Iverson, 2009). The Bayes Factor (B10) provides the likelihood ratio of the experimental hypothesis being true compared to the null hypothesis being true (Dienes, 2011). A Bayes Factor of between 1/3 and 3 only provides anecdotal evidence; a value between 3 and 10 provides substantial evidence for the hypothesis; a value above 10 provides strong evidence for the hypothesis; a value between 0.1 and 1/3 provides substantial evidence for the null hypothesis and; a value less that 0.1 provides strong evidence for the null hypothesis (Jeffreys, 1961). We used the Bayes calculator provided by Dienes (2015). Recognition accuracy (d’) Zero counts in misses and false alarms were replaced with 0.01 (there were no zero counts in hits or correct rejections). With the number of stimuli in this experiment, d' ranged from 0 (chance recognition) to 4.48 (perfect performance). Fig. 1 presents the mean d' for upright and inverted faces grouped according to age and highlights, crucially, that there were no floor or ceiling effects in any condition. These data were first subjected to a 12 × 2 mixed ANOVA with the factors: participant age and orientation. Critically, there was strong evidence for a significant interaction between orientation and participant age, F(11, 468) = 13.19, MSE = 0.53, p < .001, ηp2 = .24, B10 = 3.18 × 1038. To decompose this interaction, two univariate ANOVAs were run: one for upright faces and one for inverted faces. The effect of age was larger for upright faces, F(11, 468) = 23.60, MSE = 0.75, p < .001, ηp2 = .36, B10 = 6.53 × 1045, than for inverted faces, F(11, 468) = 2.36, MSE = 0.27, p = .008, ηp2 = .05, B10 = 401679. Curve estimations were conducted for recognition accuracy, summarised in Table 2. The relationship between recognition accuracy and age was equally well represented by a linear, quadratic, and cubic function. Given that all three were significant, and that they all represent the data well, it is clear that this pattern indicates an overall linear improvement in recognition accuracy, but with a step change mid way through development, indicated by the quadratic and cubic functions. The correlations for upright faces were compared against the inverted faces using Fisher's r-to-z transformation. These revealed that the correlations were significantly stronger for upright faces than inverted faces (see Table 2). An analysis was conducted on the mean-relative FIE presented in the bottom panel of Fig. 1, and revealed strong evidence that the FIE depended on age, F(11, 468) = 6.84, MSE = 0.16, p < .001, ηp2 = .14, B10 = 4445395. Pairwise comparisons revealed that there was no significant difference in the FIE form participants aged 5- to 10-years nor between 9-years and adults (all ps > .18). Five- to 8-year old participants showed a significantly smaller FIE than participants age 11 and older (all ps < .030). Eleven-year old and older participants showed a significantly larger FIE than children under the age of 8-years of age. The FIE was compared at each age using a series of one-sample t-tests, with the significance levels reported in Fig. 1. Curve estimations show while the FIE increased with age according to an overall linear trend, the significant quadratic and cubic functions indicate a step change in the FIE during development. Response bias, C A final analysis was conducted on the response bias data (shown in Table 3), calculated using the Macmillan and Creelman (2010) method. This measures participants’ tendency to respond with an “old” response. The score typically ranges from −1 to +1, where 0 is a neutral bias. A positive number indicates participants are less likely to response with an “old” respond when they are unsure and a negative number indicates a tendency for participants to respond with an “old” response when they are unsure. There was strong evidence for a significant interaction between orientation and participant age, F(11, 468) = 2.60, MSE = 0.19, p = .003, ηp2 = .06, B10 = 1.69 × 1015. To decompose this interaction, two univariate ANOVAs were run: one for upright faces and one for inverted faces. For upright faces, the effect of age was larger with a lesser tendency to respond with a \"new\" response for older participants than younger ones, F(11, 468) = 2.78, MSE = 0.28, p = .002, ηp2 = .06, B10 = 232.64, than for inverted faces, F(11, 468) = 2.10, MSE = 0.07, p = .019, ηp2 = .05, B10 = 82.48. Curve estimations were conducted for response bias, summarised in Table 2.","The results from Experiment 1 indicate that face recognition improves with age. This is entirely consistent with previous research (e.g., Blaney & Winograd, 1978; Carey et al., 1980; Flin, 1985). Curve fitting revealed that upright face recognition improved more reliably with age than inverted face recognition. All measures of face recognition improved with age, suggesting that participants memory for faces improved with age (as indicated by hit rate) and the coding of faces becomes more accurate (as indicated by a reduction in false alarm rate). These results indicate that face memory and face perception are developing. The pattern of data observed here also indicates that there is a step-change in the FIE, whereby participants under the age of 9-years do not show the FIE, whereas participants 10-years and older do show the FIE. In this analysis, we have fitted curves to the data. These curves show a linear, cubic, and quadratic trajectory for the development of the FIE suggesting a qualitative shift in the use of configural coding. If the development of configural processing was due to developmental induction, we would only expect to see a linear relationship. Only a hypothesis predicting that there would be no FIE until a certain age, followed by a full-strength FIE is compatible with the data. This is as predicted by the late maturation of configural processing account.","In order to fully explore this step-change in the use of configural coding, a second experiment was conducted. This was a longitudinal study in which the same 9-year-old participants were followed up for two years following the original testing. We chose to follow these participants due to the observation in Experiment 1 that indicated a step change in the FIE between the ages of 9 and 11 years. This acts as an internal replication of the initial findings and critically offers a longitudinal approach to understand how face recognition develops. The longitudinal design also offers a window to explore development of a particular group of participants. Many of the drawbacks of cross- sectional designs are avoided in this approach as it allows researchers to see change in behaviour due to age. This is also one of a very small handful of longitudinal studies applied to face recognition.","The same nine-year-old participants in Experiment 1 were tested two further times at approximately (within one month) yearly intervals. Testing took place in an identical manner as in Experiment 1 with the same stimuli used for the ten- and 11-year-old children in Experiment 1 for the second and third testing times respectively. Therefore, the participants were viewing own-age faces at age 9-, 10-, and 11-years. This, therefore, led to a 2 × 3 within-subjects design with the factors: orientation of the face and participant age (9-, 10-, and 11-years of age).","The data were treated in the same way as in Experiment 1 and are presented in Fig. 2 and Table 4. The recognition accuracy (d’) data were subjected to a 2 × 3 within-subjects ANOVA with the factors orientation and age. Critically, these effects interacted, F(2, 78) = 4.20, MSE = 0.66, p = .018, ηp2 = .10, B10 = 6.68 × 108. This interaction was decomposed by conducting two one-way ANOVAs, one for upright faces and one for inverted faces. The improvement in recognition accuracy for upright faces was significant, F(2, 78) = 5.58, MSE = 0.94, p = .005, ηp2 = .13, B10 = 37.85, but it was not for inverted faces, F(2, 78) = 0.36, MSE = 0.32, p = .699, ηp2 = .01, B10 = 0.09. A trend analysis was conducted and this revealed that the improvement in face recognition accuracy for upright faces was linear, F = 11.32, MSE = 0.92, p = .002, ηp2 = .23, with no other significant developmental trajectories (all ps > .769). A one-way ANOVA was run on the relative measure of the FIE, revealing a significant effect of age, F(2, 78) = 4.05, MSE = 0.15, p = .021, ηp2 = .09, B10 = 18.77. One-sample t-tests showed that there was strong evidence the FIE was not present in the 9-year old participants, t(39) = 1.93, p = .061 effect size r = .30, B10 = 0.24, but was present in 10-year old, t(39) = 4.27, p < .001, effect size r = .56, B10 = 1606.76, and 11-year old participants, t(39) = 8.22, p < .001, effect size r = .80, B10 = 5.00 × 1013. A linear trend was observed, F = 9.19, MSE = 0.12, p = .004, ηp2 = .19. No other contrast pattern was observed (all ps > .44). Finally, a parallel ANOVA was run on the response bias (C) data, revealing a significant interaction, F(2, 78) = 7.09, MSE = 0.19, p = .001, ηp2 = .15, B10 = 1.27 × 1011. The change in response bias with age (with a lesser tendency to respond with a \"new\" response for older participants than younger ones) was significant for upright faces, F(2, 78) = 9.80, MSE = 0.28, p < .001, ηp2 = .20, B10 = 1.10 × 1090, but not for inverted faces, F(2, 78) = 0.43, MSE = 0.09, p = .649, ηp2 = .01, B10 = 0.02. This change in response bias followed a linear pattern, F = 11.92, MSE = 0.27, p = .001, ηp2 = .23, and a quadratic pattern, F = 7.87, MSE = 0.30, p = .008, ηp2 = .17.","We have demonstrated a significant improvement in the recognition of upright age-matched faces with development. This development is much greater for upright faces than for inverted faces. The effect of this is that the FIE is not observed when our participants were 9-years of age, but was present when they were older. This developmental study indicates that a step change in the emergence of the FIE. This important advancement in our understanding of the development of face recognition.","In order to be able to fully compare these results to previously published data, it is important to assess whether the developmental trends we have found exist when recognising adult faces. This is especially important given the use of the CFMT-C (Croydon et al., 2014) to assess deficits in children's face recognition (Bennetts, Murray, Boyce, & Bate, 2017) which use adult faces. If the results we found in Experiments 1 and 2 generalise to adult faces, then there is no concern with using such adult tests in children. To this end, we used a similar method to that use in Experiments 1 and 2, with different participants and different faces: specifically, adult faces. Participants, design, and procedure Participants were 242 children (115 male) aged from 5,7 years to 16,2 years and 22 adults (aged 18–22 years; 6 male). A participant summary is shown in Table 5. All other participant details were the same as in Experiment 1. The design and procedure of this Experiment was identical to that in Experiment 1.","Two images of 32 (16 female) adult faces from the Minear and Park (2004) database were used in this study. These were of adults in their late 20 s (and therefore did not match the age of any of our participants). One image displayed a happy expression and the other displayed a neutral expression. One of these was presented during learning and the other during test (this was randomised). Using two images of the same face reduces pictorial recognition (Bruce, 1982). The face images in this database were of males and females, all with similar hairstyles, positioned in a frontal view. All extraneous paraphernalia and the background were masked using Adobe™ Photoshop™. The faces were presented in the same way as in Experiment 1 and in the same dimensions.","The results for this Experiment were analysed in the same way as Experiment 1. Recognition accuracy (d') Recognition accuracy data are summarised in Fig. 3 and were subjected to a 12 × 2 mixed ANOVA, with the factors: participant age and orientation. This revealed strong evidence for a significant interaction between orientation and participant age, F(11, 252) = 8.66, MSE = 0.44, p < .001, ηp2 = .27, B10 = 1.45 × 1018. To decompose this interaction, two univariate ANOVAs were run: one for upright faces and one for inverted faces. For upright faces, the effect of age was larger, F(11, 252) = 13.52, MSE = 0.76, p < .001, ηp2 = .37, B10 = 1.88 × 1014, than for inverted faces, F(11, 252) = 2.30, MSE = 0.54, p = .011, ηp2 = .09, B10 = 70.47. Curve estimations were conducted for recognition accuracy, summarised in Table 6. The relationship between recognition accuracy and age was equally well represented by a linear, quadratic, and cubic function. The correlations for upright faces were compared against the inverted faces using Fisher's r-to-z transformation. These revealed that the correlations were significantly stronger for upright faces than inverted faces (see Table 6). An analysis was conducted on the mean-relative FIE presented in Fig. 3, and revealed that the FIE depended on age, F(11, 252) = 3.11, MSE = 0.17, p = .001, ηp2 = .12, B10 = 8.79. However, no pairwise comparisons between participant ages were significant following Bonferroni-Šidák correction. This highlights the critical point about using age-inappropriate stimuli and a lack of experimental power may hide real effects of development in the FIE. One-sample t-tests were used to compare the magnitude of the FIE at each age, shown in Fig. 3. Curve estimations show that changes in the FIE follow a linear, quadratic, and cubic function. Response bias, C A final analysis was conducted on the response bias data (see Table 7). There was a significant interaction between orientation and participant age, F(11, 252) = 0.70, MSE = 0.15, p = .737, ηp2 = .03, B10 = 45.72. To decompose this interaction, two univariate ANOVAs were run: one for upright faces and one for inverted faces. For upright faces, the effect of age was smaller and with anecdotal evidence, F(11, 252) = 1.88, MSE = 0.20, p = .042, ηp2 = .08, B10 = 2.15, compared to the strong evidence for inverted faces, F(11, 252) = 3.33, MSE = 0.16, p < .001, ηp2 = .13, B10 = 14.73. Older participants demonstrated a lesser tendency to respond with an \"old\" response than younger ones. Curve estimations were conducted for response bias, summarised in Table 6.","The results from this Experiment are consistent with Experiment 1: face recognition improves with age and more reliably so for upright faces than inverted faces. The step- change in the FIE, whereby participants under the age of 9-years do not show the FIE, whereas participants 11-years and older do show the FIE, was also found. Curve-fitting showed a linear, cubic, and quadratic trajectory for the development of the FIE suggesting a qualitative shift in the use of configural coding suggesting late maturation of configural processing.","Similar to Experiment 2, we retested the 9-year old participants in Experiment 3 one and two years later in order to show a developmental improvement in the recognition of faces from a longitudinal study.","The same procedure described in Experiment 2, except it was conducted on the 9-year old participants we tested in Experiment 3. The one change to the method was that we used two additional sets of faces from the Minear and Park (2004) database to ensure that the faces were unfamiliar to our participants. These were of the same age-range as those described in Experiment 3. This was a 2 × 3 within-subjects design with the factors: orientation of the face and participant age (9-, 10-, and 11-years of age).","The data was treated in the same way as in Experiment 1 and are presented in Table 8. The recognition accuracy (d’) data were subjected to a 2 × 3 within-subjects ANOVA with the factors orientation and age. Critically, these effects interacted, F(2, 42) = 8.90, MSE = 0.36, p = .001, ηp2 = .30, B10 = 3.22 × 1011. This interaction was decomposed by conducting two one-way ANOVAs, one for upright faces and one for inverted faces. The improvement in recognition accuracy for upright faces was significant, F(2, 42) = 6.07, MSE = 0.44, p = .005, ηp2 = .22, B10 = 14.67, but smaller for inverted faces, F(2, 42) = 3.24, MSE = 0.32, p = .049, ηp2 = .13, B10 = 2.32. A trend analysis was conducted and this revealed that the improvement in face recognition accuracy for upright faces was linear, F = 9.24, MSE = 0.57, p = .006, ηp2 = .31, with no other significant developmental trajectories (all ps > .689). A one-way ANOVA was run on the relative measure of the FIE, revealing a significant effect of age, F(2, 42) = 3.32, MSE = 0.13, p = .046, ηp2 = .14, B10 = 3.27. One-sample t-tests showed that the FIE was not present in the 9-year old participants, t(21) = 1.02, p = .545, effect size r = .22, B10 = 0.29, but was present in 10-year old, t(21) = 3.18, p = .005, effect size r = .57, B10 = 16.65, and 11-year old participants, t(21) = 3.74, p = .001, effect size r = .63, B10 = 251.31. A linear trend was observed, F = 4.83, MSE = 0.16, p = .039, ηp2 = .19. No other contrast pattern was observed (all ps > .35) (Fig. 4). Finally, a parallel ANOVA was run on the response bias (C) data, revealing no effect of age, F(2, 42) = 0.33, MSE = 0.13, p = .718, ηp2 = .02, B10 = 0.37. There was also no effect of orientation, F(1, 21) = 0.16, MSE = 0.11, p = .695, ηp2 = .01, B10 = 0.21. The interaction was not significant, F(2, 42) = 0.56, MSE = 0.13, p = .578, ηp2 = .03, B10 = 0.25.","In this Experiment, we have demonstrated a significant improvement in the recognition of adult faces with development. This development is much greater for upright faces than for inverted faces. The FIE was not observed when our participants were 9-years of age, but was present when they were older. This longitudinal study replicates the step change in the emergence of the FIE around the age of 10-years.","While there are many limitations regarding the comparison across different studies, especially considering subtle methodological differences across the studies,6 we compared the face recognition performance across Experiments 1 and 3 in a 12 × 2 × 2 Mixed ANOVA with the factors: age of participant, orientation of stimuli, and type of stimuli (age- matched, Experiment 1; adult faces, Experiment 3). Prior to this analysis, we compared the average stimulus pixelwise complexities for the children's faces and the adult faces and found that they were not significantly different (p > .32). Similarly, attractiveness ratings did not differ significantly across the adult and the children's faces (p > .16). Figs. 1 and 3 appear to show similar patterns that indicate the recognition accuracy of upright adult and child faces increases faster than that of inverted faces. In this comparative analysis, there are two main effects of interest: the main effect of stimuli type and the interaction between stimuli type and orientation of the stimuli. We ran this analysis on the d' recognition accuracy measure first. This revealed a main effect of stimuli type, F(1, 720) = 6.07, MSE = 0.62, p = .014, ηp2 = .01, B10 = 3.01. On average, performance with age-matched stimuli (M = 1.48, SE = 0.03) was approximately 10% greater than adult faces (M = 1.38, SE = 0.03) replicating the own-age bias (Anastasi & Rhodes, 2006). This effect interacted with orientation of the face, F(1, 720) = 6.63, MSE = 0.50, p = .010, ηp2 = .01, B10 = 5497, replicating Hills (2012). This interaction was revealed through a larger face-inversion effect on own-age faces (mean difference = 0.84, SE = .04), F(1, 479) = 252.57, MSE = 0.67, p < .001, ηp2 = .35, B10 = 1.45 × 10100, than other- age faces (mean difference = 0.65, SE = 0.06), F(1, 263) = 95.14, MSE = 0.58, p < .001, ηp2 = .27, B10 = 1.15 × 1024. All these results indicate that the use of other-age faces when studying face recognition in children may well be underestimating performance. Finally, the same interaction between stimuli type and orientation was observed in the response bias data, F(1, 720) = 16.56, MSE = 0.17, p < .001, ηp2 = .02, B10 = 2.79 × 1017. The effect of stimuli type was larger for inverted (mean difference = .12, SE = .03), F(1, 742) = 22.49, MSE = 0.11, p < .001, ηp2 = .03, B10 = 10425.45, than upright faces (mean difference = .07, SE = .04), F(1, 742) = 2.76, MSE = 0.26, p = .097, ηp2 < .01, B10 = 0.75.","Through four Experiments we have shown that there was a significant improvement in the recognition of upright faces with development (both longitudinally and cross sectionally) consistent with previous research (Blaney & Winograd, 1978; Carey & Diamond, 1977; Carey et al., 1980; Diamond & Carey, 1977; Goldstein, 1965; Goldstein & Chance, 1964; Saltz & Sigel, 1967). The longitudinal data showed numerically consistent results with the cross- sectional data for participants within three years. This indicates that recognition improvements over three years are relatively small but detectable with a suitable experiment. Here we shall first summarise the findings across the experiments before relating the findings to theoretical discussion presented in the introduction. Across all measures, the improvement in the recognition of upright faces was approximately linear (Carey, 1978; Flin, 1980, 1985), though had significant quadratic and cubic functions. This pattern is indicative of a slow gradual overall improvement in face recognition, but with a step change causing the cubic and quadratic functions. The statistics highlight that that there is less improvement in face recognition during mid-adolescence. This is consistent with findings of other cognitive abilities (e.g., Flin, 1983). Furthermore, there is quick rise in face recognition abilities between the age of 9- and 12-years. To highlight this step change in abilities between these ages, we have shown that the improvement in face recognition at these ages is exclusively linear (in the longitudinal data, Experiment 2). This indicates that this is the age driving the cubic function. These findings are interesting as they present an often ignored aspect within developmental work: while the age-related developmental trajectory may produce an overall linear pattern, other functions are present indicating that factors exist that may enhance or inhibit such development. These factors maybe individual difference variables, relating to hormonal, maturational, or environmental changes surrounding the individual. This linear improvement in the recognition accuracy of upright faces was significantly greater than the approximately linear improvement in the recognition of inverted faces, consistent with de Heering et al. (2012). In other words, the improvements to the recognition of upright faces were greater than the improvements in the recognition of inverted faces. This is entirely consistent with the notion of expertise developing for the recognition of faces that are encountered in the most common orientation. Relating to the general models of perceptual and skill development presented in the introduction, this is akin to developmental induction due to experience with upright faces outweighing experience with inverted faces and therefore leading to enhanced learning how to process these stimuli. The FIE, measured in relative terms, showed a different pattern of development. This showed no increase in the FIE until the age of 9-years, followed by an increase in the FIE over two years before plateauing. A close inspection of de Heering et al.ʼs (2012) data reveals a similar pattern. These data are consistent with an “all-or-none”’ late maturational model of the FIE (Carey et al., 1980; Itier & Taylor, 2004) with a slower development of configural coding than featural coding (Mondloch, Le Grand, & Maurer, 2002). Here, there is statistical evidence for this assertion: the change in the FIE with development followed primarily a cubic function. We present this assertion against a background that the measure of FIE we used should have enhanced the likelihood of us finding it in the youngest age groups. Since we did not, it clearly shows that configural coding (as measured by the FIE) is not employed for faces until about the age of 10-years. The present data also indicate that there is a general increase in memory for faces with development (see e.g., Brown, 1975; Chi, 1977; Dempster, 1981; Flavell, 1977; Kail, 1992) given that hit rate increased with age. However, memory for inverted faces did not improve with age. This is more consistent with Weigelt et al.’s (2013) data indicating a domain- specific increase in face memory with age rather than a general improvement in cognitive functioning (Crookes & McKone, 2009; McKone & Boyer, 2006). In the introduction, a number of potential explanations were presented for the FIE. Our data indicate a sharp increase in the FIE between the ages of 9 and 10 years. This indicates an increase in expertise in face recognition during this period. This expertise is likely to be the form of configural processing known as holistic processing (Maurer et al., 2002a). One theory of holistic processing is that inversion disrupts the perceptual field such that the entire upright face cannot be sampled from a single central fixation (Rossion, 2008, 2009). These effects mirror those observed in other areas of perceptual expertise that require many years of experience to achieve (Charness, Reingold, Pomplun, & Stampe, 2001; Ericsson, Krampe, & Tesch-Römer, 1993; Kundel & Nodine, 1975). An inverted face, on the other hand, cannot be processing from a central fixation: this has been found in eye-tracking evidence (Barton, Radcliffe, Cherkasova, Edelman, & Intriligator, 2006; Hills, Sullivan, & Pake, 2012; Xu & Tanaka, 2013). This suggests that with development, the perceptual field widens when viewing faces until, around the age of 10-years, it is sufficient to encode an entire face with one fixation. This hypothesis can be easily tested by exploring the eye-movements of children. There is evidence that the features used by children to process faces changes with age. Campbell, Walker, and Baron-Cohen (1995) has shown an external feature advantage at age 7-years that shifts to an internal feature advantage by 9-years of age. Further eye-tracking evidence supports a late development of this internal feature advantage (Kelly et al., 2011; Meaux et al., 2014; Senju, Vernetti, Kikuchi, Akechi, & Hasegawa, 2013). In addition, younger children should show more fixations with wider distribution over the features than older children and adults (Hills, Willis, & Pake, 2013). Indeed, Ge et al. (2008) have shown that older children tend to be able to recognise faces based on the most diagnostic features (the eyes for White faces and the nose for East Asian faces; Ge et al., 2008; Kelly et al., 2011; Liu et al., 2013). Younger children, on the other hand, rely on less diagnostic features. The refinement of eye-movements appears to occur at around 10 years of age. The development of eye-movements has been interpreted as evidence of the development of face expertise (Ge et al., 2008; Tanaka et al., 2014). It is thought that the progression to more sustained fixation on the diagnostic features may help the development of face processing expertise such that it becomes a more rapid and automatic process (Kelly et al., 2011) consistent with notion that face expertise involves a change in processing style and cognitive encoding (Diamond & Carey, 1977), from local to holistic (Hole, 1994; Tanaka & Farah, 1993) and configural processing (Leder & Bruce, 2000). The expert perceptual field theory has a great deal of support for it from other areas of congitive, attentional, and perceptual development. Expertise in many domains is associated with greater ability to \"chunk\" information and create more stable schemas (Braune & Foshay, 1983; Goldstein, 1975) with increased systematisation of knowledge (Karmiloff-Smith & Inhelder, 1974). Such systematised knowledge leads to improved memory when consistent with internal schemas (Simon & Barenfeld, 1969; Simon & Gilmartin, 1973 see also Chase & Simon, 1973a, 1973b; DeGroot, 1965, 1966; DeGroot & Gobet, 1996; Gobet & Simon, 1996a, 1996b). Wider perceptual fields are observed in expert footballers (Williams & Davids, 1997) and quicker processing from fewer fixations and greater chunking is observed in expert chess players (Charness et al., 2001; Reingold, Charness, Pomplun, & Stampe, 2001; Reingold, Charness, Schiltetus, & Stampe, 2001) and radiologists (Kundel & Nodine, 1975). Indeed, the automatisation of perceptual processing (Vurpillot, 1968) and attentional shifting requires less effort than inexpert processing as it is based on an unconscious and automatic process. This takes circa ten years to develop (e.g., Akhtar & Enns, 1989; Brodeur & Enns, 1997; Enns & Brodeur, 1989; Pearson & Lane, 1991) or occurs at the age of 10 years. The data fit with neuroscientific evidence that indicates the pattern of face-specific recruitment of cortical regions is not observed prior to the age of ten years (Aylward & Meltzoff, 2005 but see Golarai, Liberman, & Grill-Spector, 2015) or that the size of face-specific regions is smaller in children than adolescents or adults (Golarai et al., 2007) and that the magnitude of ERPs associated with face perception are different for children than adolescence and adults (Kuefner, De Heering, Jacques, Palmero- Soler, & Rossion, 2010). The key finding here is that the FIE appears to follow a maturational development (as highlighted by the cubic function), whereas the development of recognising upright faces appears to follow an induction pattern. Of course, the notion that there is ten years of development that is required for face recognition expertise is an overly simplistic argument. It may be that this trend appears at the age of 10 years due to biological and/or social constraints that are vital for expertise to develop rather than experience. The cause of expertise development at the age of 10 years could be due to hormonal changes altering the functioning of the attentional spotlight and eye-movements. Puberty is associated with a number of hormonal changes known to affect cognitive functioning especially those associated with frontal lobe functioning (Blakemore & Choudhury, 2006) due to the significant synaptic pruning that occurs in this region during puberty (Woo, Pucak, Kye, Matus, & Lewis, 1997; Zecevic & Rakic, 2001). Indeed, development during this period involves the development of attentional shifting, abstract reasoning, and inhibition (Yurgelun-Todd, 2007). These factors indicate that at this age, the perceptual system can utilise a more robust schema that inhibits irrelevant dimensions and focuses on the most diagnostic information for recognising faces. While there is hormonal and biological maturation occurring at this time, there are significant environmental changes that occur at the same age. School environmental changes at age 10–11 from small classes in smaller schools in primary education to larger classes at larger secondary schools (college). This results in exposure to more faces than previously encountered and typically a large number of new faces. This sudden increase in the amount of exemplar's entering the face-space may result in its refinement. Whatever the cause, this mechanism clearly results in a potentially unique and special processing for faces at this age. There is, of course, a limitation of the present work. By the age of 5-years, children will have been exposed to a significant number of faces right from birth. Indeed, there is evidence that the early visual system is set up to process faces (de Heering et al., 2008). Newborn visual acuity makes faces the single most important visual stimulus leading to early engagement with faces (Coulon, Guellai, & Streri, 2011). This explains why there are implicit measures of face processing in infants, measured using ERPs (de Haan, Johnson, & Halit, 2003) that appear to show differentiation between faces and objects (Peykarjou, Pauen, & Hoehl, 2016; Peykarjou, Wissner, & Pauen, 2017) and even between upright and inverted faces (Halit, De Haan, & Johnson, 2003; Peykarjou & Hoehl, 2013). While these measures do not show the ability to individuate faces, such data indicates significant learning about faces and the development of mechanisms to process faces prior to the age of participants tested here. The influence of this on subsequent face recognition has not been considered in the present study. Infants will not be exposed to as many faces as adults, and are less likely to encounter own-age faces than other-age faces given that infants spend more time with their parents and family members than at schools or nurseries. The method of discriminating between a small number of faces encountered is likely to be based on simple mechanisms, especially if these faces are more varied in terms of age (we are considering an environment where an infant has frequent contact with both parents, older siblings, and grandparents). In these situations, simpler age cues may be sufficient to identify and discriminate faces. However, when older, the number and types of faces that are encountered are likely to be more homogenous and therefore the methods used to discriminate these necessarily need to be more sophisticated. We used inverted faces in the present study as way of assessing development of inexpert processing. There is no reason to believe that the developmental trajectory for the recognition of objects would be different to that of inverted faces given that inverted faces are processed using the inexpert featural processing manner that is afforded to objects. We have also shown the effects of inversion to be greater when testing faces of one's own age relative to other-age faces. However, since we did not compare the recognition of own-age and other-age faces in the same children, we cannot rule out potential participant and stimuli effects for this difference. One caveat with the explanations presented here is that these results are specific to the FIE during recognition. A general theory that explains all effects of inversion would require a great deal more specification than is currently available. Furthermore, there is no reason to believe that inversion affects all aspects of face processing and configural coding in the same way. Another limitation with the present study is that it cannot address what the maturational mechanism actually is, or indeed if it is enhancement, facilitation, or maintainence (Coren, Ward, & Enns, 1999). Previously proposed mechanisms are the qualitative shift from featural to holistic processing (Carey & Diamond, 1994). This is entirely possible, but the cause of this shift is not clear. Similarly, it could be that there is a shift from using certain (presumably featural) dimensions to other (more configural) dimensions of face-space (Valentine, 1991). Indeed, the data presented by Hills, Holland, and Lewis (2010) is consistent with this presumption: children under 10-years of age can be adapted to facial distortions (a single eye moved) that adults cannot because they are coding faces according to each feature individually rather than the features together. Alternatively, it is plausible that it requires ten years for schema (in this case, prototype face) to become more fixed or robust. All of these interpretations lack the precise description of the mechanism and cause for this change. In conclusion, the expert processing mechanism (potentially configural or holistic processing) develops. However, there is a maturational component to this development, suggesting a stage where it is not largely utilised (prior to 10-years of age) to a stage where it is largely employed. This coding switch idea has been hypothesised before (Carey & Diamond, 1977), but adequate tests of it have not been forthcoming until now. This is not to suggest that configural processing cannot be done prior to the age of 10, but that its reliable and constant use is not likely before that age. Indeed, the effects of inversion may be unrelated to other measures of \"configural\" coding (such as the parts and wholes test; Konar, Bennett, & Sekuler, 2010; Wilhelm et al., 2010). Nevertheless, the results from this study are reliable given the internal replication and the robust statistical procedures employed."],["Although a number of studies have examined the developmental emergence of counterfactual emotions of regret and relief, none of these has used tasks that resemble those used with adolescents and adults, which typically involve risky decision making. We examined the development of the counterfactual emotions of regret and relief in two experiments using a task in which children chose between one of two gambles that varied in risk. In regret trials they always received the best prize from that gamble but were then shown that they would have obtained a better prize had they chosen the alternative gamble, whereas in relief trials the other prize was worse. We compared two methods of measuring regret and relief based on children's reported emotion on discovering the outcome of the alternative gamble: one in which children judged whether they now felt the same, happier, or sadder on seeing the other prize and one in which children made emotion ratings on a 7-point scale after the other prize was revealed. On both of these methods, we found that 6- and 7-year-olds' and 8- and 9-year-olds' emotions varied appropriately depending on whether the alternative outcome was better or worse than the prize they had actually obtained, although the former method was more sensitive. Our findings indicate that by at least 6 or 7 years children experience the same sorts of counterfactual emotions as adults in risky decision-making tasks, and they also suggest that such emotions are best measured by asking children to make comparative emotion judgments. --------------------------------------------------------------------------------","There has been a recent surge of research interest in the development of counterfactual thinking (for reviews, see Beck & Riggs, 2014; Rafetseder & Perner, 2014). Some of this research has focused on the development of emotions thought to require counterfactual thinking abilities, specifically regret and relief (Burns, Riggs, & Beck, 2012; McCormack & Feeney, 2015; O’Connor, McCormack, & Feeney, 2012; O’Connor, McCormack, & Feeney, 2014; Rafetseder & Perner, 2012; van Duijvenvoorde, Huizenga, & Jansen, 2014; Weisberg & Beck, 2010; Weisberg & Beck, 2012). Researchers studying early and middle childhood have primarily focused on attempting to pinpoint the age at which children first experience these emotions (O’Connor et al., 2012; Rafetseder & Perner, 2012; van Duijvenvoorde et al., 2014; Weisberg & Beck, 2010, 2012). Although this research has proved fruitful, there is still considerable disagreement over when regret can first be observed developmentally (Rafetseder & Perner, 2012; Weisberg & Beck, 2012). It is possible to identify three methodological issues that make it difficult both to resolve this disagreement and to integrate developmental findings with the larger body of research on regret conducted with adults. The current study directly addresses these issues. All of these developmental studies have used a simple paradigm requiring children to choose between two boxes to win a prize of stickers or candies. Children see the prize from their chosen option and then rate how they feel about that prize on an emotion rating scale ranging from very happy to very sad (the exact nature of this scale differs across studies). When giving these initial ratings, children have received only what is termed partial feedback; they have seen the prize resulting from their choice, but they have not yet seen the prize they would have won if they had made a different choice. Children are given complete feedback when they see both what they have won and what they would have won if they had made a different choice. Emotion ratings are subsequently made after complete feedback. If the prize from the unchosen option is better than that from the chosen option, children reporting that they now feel sadder are assumed to be experiencing regret about their choice; if it is worse, children reporting that they now feel happier are assumed to be experiencing relief. Although all of the developmental studies employed this basic procedure, they differ from each other in the exact way that counterfactual emotions are assessed; moreover, the task they use also differs in important ways from the type of task typically used to examine counterfactual emotions in adolescents and adults (Burnett, Bault, Coricelli, & Blakemore, 2010; Camille et al., 2004; Coricelli et al., 2005). We focus on three methodological issues stemming from these differences: the nature of the choices children need to make, the extent to which their emotional responses can be based on a single comparison between the prize received and the best or worst prize available, and the way in which emotion ratings are used to measure regret/relief. Each of these issues is described in turn. Choice and risky decision making ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Experiments examining regret and relief in adolescents and adults have used paradigms in which the choice that participants make is more complex than simply choosing between two boxes. Such studies typically align with research in the broader decision-making literature insofar as the tasks involve choosing between alternatives that vary in associated risk. In most tasks, the choice is between two gambles, with gambles giving participants the opportunity to win or lose points (e.g., Bault, Coricelli, & Rustichini, 2008; Coricelli et al., 2005; Mellers, Schwartz, & Ritov, 1999). For example, Burnett et al. (2010) asked participants to choose between pairs of gambles that differed in risk such as a choice between a gamble with a 50% chance of winning 50 points along with a 50% chance of losing 50 points (+50/−50) and a gamble with an 80% chance of winning 200 points but a 20% chance of losing 200 points (+200/−200). On a regret version of such a trial, if participants chose the +50/−50 option, they were then shown that they would have won 200 points if they had chosen the other gamble; on a relief version, if participants chose that option, they were then shown that they would have lost 200 points if they had chosen the alternative gamble. There is a good reason why studies of counterfactual emotions in adults have used tasks in which participants need to choose between options varying in risk: Researchers interested in the processes underlying risky choice have been examining whether counterfactual emotions are an important element of decision making (Coricelli, Dolan, & Sirigu, 2007; Mellers et al., 1999). Such a hypothesis stems from an influential tradition of formal models of the role of regret in economic choice (Bell, 1982; Loomes & Sugden, 1982). Moreover, neuropsychological studies using these tasks have attempted to identify the brain systems that underpin regret (Sommer, Peters, Gläscher, & Büchel, 2009). Studies using such tasks with brain-damaged patients who do not experience regret have helped to establish that this emotion may indeed play a role in risky decision making (Camille et al., 2004; Coricelli et al., 2007; but see also Levens et al., 2014). Although there are good methodological reasons for using simpler tasks with young children, the fact that the developmental studies on regret use a quite different task that does not involve risky choice means that it is difficult to integrate the findings of such studies with those from the research with adults. In particular, we do not know whether young children experience counterfactual emotions in the same sorts of decision-making situations as adolescents and adults or whether such emotions may play a role in explaining developmental changes in risky decision making. The current study used a task that closely resembled in structure the tasks used with adults insofar as it involved choosing between two gambles that varied in risk. However, the task was simplified. In the studies with older participants, not only does the magnitude of possible prizes vary across gambles, but the odds of winning prizes can vary as well (in the example given above, the odds in one gamble are 50/50 and those in the other are 80/20). By contrast, in our task only the magnitude of the possible prizes varied, with each gamble always involving a 50% chance of winning one of two prizes. For example, children needed to choose between a gamble in which there was a 50% chance of winning 10 points and a 50% chance of winning 7 points and a more risky gamble in which prizes of 16 points and 1 point were equiprobable. Although this simplified the task, it nevertheless allowed us to examine the developmental profile of counterfactual emotions in the context of risky choice. Single reference point versus multiple reference points ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Using a gambling task allowed us for the first time to examine the early development of counterfactual emotions in the context of risky choice. However, it also allowed us to address a further issue regarding the flexibility with which children make spontaneous counterfactual comparisons. One claim in the literature (see Rafetseder & Perner, 2012; Rafetseder & Perner, 2014) is that younger children’s emotions in a box-choosing task may be underpinned by just the simple thought that “I do not have the best prize” (or conversely “I do not have the worst prize”) and thus are best described as frustration rather than regret. This explanation implies that (a) younger children’s emotions are not the result of counterfactual thinking and (ii) the comparisons children make between what they obtained and what else was available are inflexible. That is, if children’s emotions are a result of such thoughts, we need only to assume that, unlike adults, children are making a single comparison between what they actually won and the other available prize. Adults’ emotional ratings on the gambling task reflect an ability to flexibly make different types of comparisons between the reward obtained and (at least) two other reference points. By flexibly, we mean that adults can shift their point of reference appropriately in response to the different types of information that they receive. Adults’ emotional responses indicate that they make a comparison following partial feedback between what they did win on the gamble they selected and what they could have won on that particular gamble. Using the example of a choice between two gambles of +50/−50 and +200/−200, when participants choose the +50/−50 gamble, unsurprisingly they report a negative emotion if they lose 50 points; negative emotions in such circumstances are interpreted as reflecting disappointment (Mellers et al., 1999). Regret and relief are assumed to be a result of a further comparison when complete feedback is provided (i.e., on being shown the result of the other unchosen gamble) between what they won and what they would have won had they chosen the other gamble. Using our example, a participant who lost 50 points might nevertheless report the positive emotion of relief on seeing that he or she would have lost 200 points had he or she chosen the other gamble. By using our simplified gambling task, we were able to examine whether children can make similar types of comparisons. Consider a situation in which children choose between a gamble to win either 16 points or 1 point (16/1) or either 10 or 7 points (10/7). Imagine that a child chooses the safer gamble, 10/7, and wins 10 points. That child already knows that he or she has not won the best available prize of 16 points. Nevertheless, the child may report feeling relatively happy because he or she did receive the best prize from the chosen gamble. If the child does report feeling relatively happy, we can be confident that the child’s response is not solely based on a single comparison between the actual outcome and the best (or worst) prize available. Now assume that complete feedback is provided, revealing that the alternative prize from the unchosen outcome is 16 points, and the child reports feeling sad. We can be confident under such circumstances that the child has flexibly used two different points of reference: what the child could have won in the chosen gamble and what the child would have won if he or she had selected the alternative gamble. Each of these different points of reference would yield different (and, crucially from the point of view of observing them, contrasting) emotional responses. Although we know that older children, like adults, can indeed use both of these points of reference (Habib et al., 2012), the design of previous studies with younger children has meant that it is not clear at what age such an ability is present. Within-trial versus between-trial comparisons of emotion rating ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The third issue that this study addresses concerns an important difference between existing studies in how emotion ratings are actually used and interpreted. In what we term within-trial procedures, participants make ratings following partial feedback, are subsequently given complete feedback, and then make a second rating. This is the method that has been used in the majority of developmental studies. By contrast, in what we term between-trial procedures, participants make only a single rating on any given trial either following partial feedback or following complete feedback (e.g., Burnett et al., 2010; Rafetseder & Perner, 2012). These ratings can then be compared across trials as necessary. The majority (although not all) studies with adults have used procedures in which a single emotional rating was collected following either partial or complete feedback (e.g., Burnett et al., 2010; Chua et al., 2009; Coricelli et al., 2005; Mellers et al., 1999; Zalla et al., 2014). As we will see, one possible explanation for the inconsistencies observed in developmental studies lies in whether within-trial or between-trial comparisons of emotion ratings have been used. Studies examining when regret first emerges developmentally that have used within-trial procedures have all found evidence of regret or relief in children from at least around 6 or 7 years of age (Burns et al., 2012; McCormack & Feeney, 2015; O’Connor et al., 2012, 2014; van Duijvenvoorde et al., 2014; Weisberg & Beck, 2010, 2012); indeed, using such a procedure, Weisberg and Beck (2012) argued that regret may even be observable in children as young as 4 years. Rafetseder and Perner’s (2012) study is the only one to examine the developmental emergence of regret using a between-trial comparison of emotion ratings. Notably, they found that in children under 9 years of age, there was no difference in ratings following partial versus complete feedback. On the basis of their findings, they concluded that regret was not apparent in children until around 9 years. Interestingly, it seems to be the case that whether further age differences are reported in levels of counterfactual emotions later in development also depends on which method is used. In a study comparing young adolescents, mid- adolescents, and adults using the gamble choice paradigm and a between-trial method, Burnett et al. (2010) found that there was only limited evidence of developmental changes. There was no evidence of developmental increases in the intensity of negative emotion in regret trials, and although they did find increases in the intensity of positive emotions in relief trials between young and mid-adolescents, there were no further changes after this age. By contrast, using a similar paradigm but with a within-trial procedure, Habib et al. (2012) found evidence of developmental increases in regret, with both 11-year-olds and an adolescent group showing less regret than an adult group. These studies differed in that Burnett et al. (2010) index of regret and relief was simply the intensity of reported emotion on complete feedback trials, whereas Habib et al. (2012) calculated the difference between emotion ratings involving partial and complete feedback taken within a single trial. Habib et al. speculated that the results of the two studies are not consistent because of this difference in the way that regret and relief were measured. Looking across the findings from developmental studies, it is clear that whether it is better to use a within-trial or between-trial procedure is an important methodological issue. Researchers working with younger children disagree as to which is the most appropriate method (McCormack & Feeney, 2015; Rafetseder & Perner, 2012). Rafetseder and Perner (2012) argued that it is not appropriate to use a within-trial method because children may feel compelled to change their response if asked to give a second emotion rating in a single trial; children’s reports of feeling sadder after complete feedback may simply reflect a tendency to give a different rating after double questioning. Rafetseder and Perner (2012) found some evidence to support this suggestion in a study where they directly compared a within-trial method with a between-trial method. Younger children were more likely to report feeling sadder after seeing the alternative outcome in the within-trial condition, where they made two emotion ratings, than in the between-trial condition, where they made only one emotion rating. However, O’Connor et al. (2012, 2014; see also McCormack & Feeney, 2015; van Duijvenvoorde et al., 2014) included an extra trial type to control for the effects of double questioning in which the prizes in both the chosen and unchosen boxes were identical. They found that very few children were likely to report feeling sadder in these control trials, suggesting that children’s tendency to report feeling sadder in regret trials was not an artifact of double questioning. Nevertheless, given the burgeoning number of studies on the development of counterfactual emotions, it would be useful to reexamine this methodological issue directly. The current study aimed to do so by using within-trial and between-trial comparisons of emotion ratings in the same study to directly examine whether these methods do indeed yield different patterns of developmental findings. The current study ~~~~~~~~~~~~~~~~~ The current study used our simplified version of the gambling task with children aged 6 or 7 years and an older group of 8- and 9-year-olds. Our youngest age group was selected because the majority of researchers have concluded that most 6- and 7-year-olds are able to experience regret, as measured by the box-choosing task. However, it is not clear whether this age group will show evidence of counterfactual emotions in the context of risky decision making. We included an older group because we were also interested in whether there would be developmental increases across this age range in the likelihood that participants experienced counterfactual emotions. Van Duijvenvoorde et al. (2014) demonstrated that the likelihood that children experienced counterfactual emotions even in a simple box-choosing task increased over this period of childhood, and this study allowed us to examine whether this was also the case in a risky decision-making task.","selected between gambles varying in risk (see Table 1), with a fixed 50% chance of each outcome. Children received many fewer trials than in studies with adolescents and adults, but we constructed these trials carefully to allow an examination of whether children were able to use two points of reference in their emotion judgments. There were two general types of trials: regret trials and relief trials. In regret trials, children selected between two gambles such as 10/7 and 16/1, where 10/7 is the safer option compared with the riskier 16/1 option. In Experiment 1, if children chose the safer option in regret trials, they always won the best prize from that option (10 points). However, they were also always shown that they could have won a better prize (16 points) if they had chosen the riskier option. In relief trials, children also needed to make a choice between a safer option and a riskier option such as between 9/6 and 14/1. If they chose the safer gamble, they always won the worse prize (6 points) but were shown that they would have won fewer points if they had chosen the risky gamble (1 point). Experiment 2 used a similar set of trials but, as will be described, also included additional regret and relief trials in which the choice of gambles and magnitude of the received prize were matched between these trial types. In Experiment 2, we also varied whether children initially won the best or worst prize from their chosen gamble. In Experiment 1, in addition to varying whether trials were relief or regret trials, we also varied whether participants made two emotion ratings per trial (one after partial feedback and a second one after complete feedback) or just one emotion rating (only after complete feedback). This allowed us to directly compare regret and relief as measured either using the within-trial method or the between- trial method. Of interest was whether the developmental patterns differed depending on the measure used and whether either measure was more sensitive. Participants Participants were 37 6- and 7-year-olds (18 girls; M = 84.4 months, range = 77–91) and 31 8- and 9-year-olds (13 girls; M = 110.8 months, range = 103–118). Children were recruited from schools local to the university of the first author, were predominately from lower- to middle-class backgrounds, and were of Caucasian origin.","The task used 12 cardboard boxes (16 × 26 × 15 cm). The 4 boxes used in the practice trials were painted blue and red, whereas the 8 boxes used in the experimental trials were painted black and white. The boxes were vertically divided into two colored sections, and each section contained a card representing the number of points won in two ways: as a printed symbolic number and in concrete form as a picture of a stack of discs, with the number of discs corresponding to the number of points won (see Figs. 1 and 2). Each box had a hinged lid that could be opened to reveal just one side of the box. Along the division on the outside of each box were two yellow stars indicating the number of points in the two sections displayed as symbolic numbers. There was no indication of which colored side contained which amount of points. Separately, there were also concrete representations of the possible points to be won in the form of cylindrical stacks of metal discs, with each disc representing 1 point. These stacks were placed to the side of each box (see Figs. 1 and 2). Two dice were used; the practice die was painted half blue and half red, and the experimental die was painted half black and half white. During the task, a horizontal number line from 1 to 100 with a small arrow was used to indicate the number of points accumulated during the game. A 7-point affective response scale was used, with cartoon faces varying in emotional expression from very happy on the left-hand side to very sad on the right-hand side of the scale. Children indicated their emotional responses using this scale and a three-pronged arrow. The three-pronged arrow had leftward-, rightward-, and upward-pointing prongs indicating feeling happier, sadder, and the same, respectively. Seven small pictures representing different emotion-provoking scenarios (e.g., a lost dog poster, a trophy) and two puppets with toys (a camera, a mobile phone, and a watch) were used in the pretraining session where children learned how to use the emotion scale. Prizes were small “goodie” bags containing age-appropriate toys. Procedure Children were invited to play a game and were told that if they won enough points by the end of the game, they would receive a goodie bag. The 7-point scale was first introduced along with the seven cartoon pictures. The cartoon faces were described from left to right as feeling “really really happy” through to “really really sad.” Children were then asked to indicate which face represented one of the aforementioned feelings. Next, children were shown the seven cartoon pictures and were told that different things can make us feel happy or sad. The pictures and the scenarios they represented were discussed before children were asked to rate how they would feel in each scenario (e.g., “Imagine you won a trophy; how would you feel?”). Each picture represented a different feeling on the scale. Answers were discussed and corrected if necessary. This technique was used so that children were comfortable with the idea of degrees of happiness or sadness and thus willing to use the full range of the scale. Children were then introduced to the two puppets. One puppet received a small toy and was described as feeling “really happy,” and the other puppet received two toys and was described as feeling “really really happy,” with the arrow pointing to the appropriate face on the scale. Each puppet received an additional toy, and children were asked to indicate, using the three-pronged arrow, whether that puppet now felt happier than before (leftward prong), sadder than before (rightward prong), or the same as before (upward prong). This demonstration was repeated but with both puppets losing their toys. Children who gave incorrect answers were corrected, and the trials were repeated. Full details of this training procedure are provided in O’Connor et al. (2012, Experiment 2). Next, a pair of blue and red practice boxes was introduced. Each of the two boxes was split into two sides, one red and one blue, and the boxes were opened to show children the two sections. The experimenter showed the two cards depicting the prizes for each box. In the first practice trial, the two cards for one of the boxes depicted 9 and 4 points, respectively, and the two cards for the other box depicted 8 and 5 points, respectively. The experimenter then placed one card in each colored section of each box, keeping the number on each card hidden. The experimenter pointed to the two stars on the outside of each box that depicted its two prizes (e.g., 9 and 4 points) along with the concrete representations of these in the form of the cylindrical stacks of metal discs that were placed in the front of each box. The experimenter made it clear that there was only one card per section and that children could not be sure which card was in which section. Children were then asked a series of comprehension questions to ensure that they understood how the points were allocated. Children who gave incorrect answers were corrected, and the process was repeated until they showed a full understanding of how the points were distributed. Children were told that to play the game they first needed to roll the colored die because this would determine whether they would be using the blue or red sections of the boxes.1 After the die was rolled, children were asked to choose just one box for a prize of points. The experimenter opened the predetermined colored section (blue or red) of the chosen box and revealed the printed card (e.g., 9 points). The corresponding section of the unchosen box was then opened, and the card was revealed (e.g., 8 points). Children completed one more practice trial with the two remaining blue and red boxes; in this second trial, one box contained 7 and 3 points and the other box contained 6 and 4 points. The presentation order of the practice trials was counterbalanced. Children then moved on to experimental trials that involved four sets of black and white boxes. Each experimental trial included a safer box and a riskier box; see Table 1 (Trials 1–4) for a breakdown of the number of points displayed on the outside of each box. Unknown to children, trial outcomes were manipulated in advance of the game by planting appropriate cards in the sections of the boxes; actual outcomes are depicted in Table 1 for each trial type, assuming that children made the safer choice. Children who chose the riskier box always won just 1 point because it was hoped that this would deter children from choosing the risky boxes. This decision was made because the critical trials for analysis were trials in which children chose the safer box.2 The experimental trials were presented in blocks of four, with each trial number (Trials 1–4 in Table 1) occurring once per block. For each child, the order of the trials remained constant across blocks but was varied between children. We did this by assigning children randomly to one of 12 different possible trial orders, with 6 of these orders having a regret trial as the first one in the block and 6 having a relief trial as the first one in the block. Because we were only interested in children’s performance on trials in which children made the safer choice (see also Burnett et al., 2010; Habib et al., 2012), the number of blocks and trials that children received varied depending on the number of safer choices they made; we detail below how this was managed. Two different methods were used to measure children’s feelings about their choices during the game. The within-trial method involved trials in which children made two emotional ratings, which had the following structure depicted stage by stage in Fig. 1. In Stage (i), children were shown the pair of black and white boxes with the possible points depicted on them and the corresponding representations of the number of points in stacked discs. In Stage (ii), the black and white die was thrown, determining whether children would get the black half or the white half of the boxes. In Stage (iii), children were asked to select one box from the pair, and depending on the die throw the experimenter opened either the black or white section of the chosen box and children were shown the number of points that they had obtained (in the example in Fig. 1, the child chooses the safer box, and the appropriate side is opened to reveal a card depicting 10 points; the stack of 10 discs is then placed beside the box, and the stack of 7 discs is removed to make it clear that the child has won 10 points). In Stage (iv), children were asked to rate how they felt about the outcome using the 7-point scale (emotion rating after partial feedback). In Stage (v), once this rating had been made, the alternative outcome was then revealed from the appropriate section of the unchosen box (in the example in Fig. 1, the child finds out that he or she would have won 16 points had he or she chosen the other box). In Stage (vi), children were asked to indicate, using the three-pronged pointer, whether they now felt happier, sadder, or the same (updated emotion rating after complete feedback). Thus, for these trials, children made two emotion ratings: one following partial feedback and one following complete feedback; these are labeled Two Rating trials. For the between-trial method, the measure of interest was whether children’s emotion ratings differed depending on whether they had received complete or partial feedback. This measure necessarily involved a comparison of emotion ratings from two separate trials of the same type; a rating after partial feedback was compared with a rating given on a separate but otherwise identical trial in which only complete feedback was provided. Ratings following partial feedback were taken from the Two Rating trials [i.e., at Stage (iv) listed above and as depicted in Fig. 1]. These were compared with ratings taken from a separate block of otherwise identical trials that had the same Stages (i) to (iii) as above but involved children making an emotion rating only once they had complete feedback. Fig. 2 shows a sample One Rating trial. As can be seen from the figure, these differed from Two Rating trials in that children saw the contents of both the chosen and unchosen boxes before making a single emotion rating. Children first completed two blocks of 4 trials, with each block consisting either of 4 One Rating trials or 4 Two Rating trials. Each block involved the 4 trials depicted in Table 1. Once these blocks had been administered, because only data from trials where the safer option had been chosen could be included in the analysis, we administered a single repetition of trials in which children had made the riskier choice. Testing proceeded after the administration of the first two blocks of trials as follows. If children had chosen the safer option in both the One Rating and Two Rating blocks for any specific trial number (from Trials 1–4 in Table 1), they were no longer given any more trials of that number because this provided the full data set needed for an assessment of performance on that trial number using both the within-trial and between-trial methods. This meant that if children had chosen only safer options on all 8 trials in the first two blocks (which was rare), testing was terminated. However, if children had made a riskier choice for a specific trial number in either the One Rating or Two Rating block, that specific trial number was repeated once as necessary in a subsequent block in order to obtain the data needed for that trial. For example, with regard to Trial 1 (10/7 vs. 16/1), if children chose the safer gamble (10/7) when given both One Rating and Two Rating versions of that trial, they did not repeat Trial 1. However, if they chose the riskier option (16/1) in the One Rating version of the trial, the One Rating version was repeated in a subsequent One Rating block (the same was true if they chose the riskier option in the Two Rating version of that trial). Thus, repetitions of the One Rating and Two Rating blocks included only trials for which children had not chosen the safer option. If children failed to choose the safer option again on a given trial number, we did not administer it a further time, and we set an upper limit of four blocks of trials per child to prevent fatigue with the task. This meant that children received variable quantities of trials, with a minimum of 8 and a maximum of 16 trials (M = 12.06, SD = 1.97). Children were shown a running total of the number of points they had won. The experimenter used the 1 to 100 number line to do this, recording the number of points children won during the game by updating the position of the arrow on each trial. All children received a goodie bag at the end of the game regardless of the actual number of points won.","The younger children completed on average 1 more trial than the older children (M = 12.51, SD = 2.01 vs. M = 11.52, SD = 1.81); this was due to the younger children making significantly more risky choices, t(66) = 2.22, p < .05. All subsequent analyses reported here are only of trials in which children chose the safer box. We began by comparing initial emotional ratings following partial feedback only from blocks of Two Rating trials to examine whether children’s responses to their actual prize varied depending on whether they had received the best possible prize from their chosen box. Ascribing a score of 1 = really really sad to 7 = really really happy, the average first emotion rating on the 7-point scale given in regret trials was 6.47 (SD = 0.92) and in relief trials was 5.42 (SD = 1.08); children were initially significantly happier after seeing the actual outcome in the regret trial (where they received the best outcome from their chosen gamble) than they were in the relief trial (where they received the worst outcome from their chosen gamble), t(57) = 6.51, p < .001. Within-trial analysis of counterfactual emotions This analysis focused on the second emotional ratings given after complete feedback in Two Rating trials using the three-pronged arrow. These ratings are always relative to the first ratings following only partial feedback and are categorical (i.e., happier than, sadder than, or the same as before the unchosen box was opened, depending on which prong of the arrow was chosen). Not all children chose the safer option for at least one Two Rating regret trial and one Two Rating relief trial, meaning that there were a small number of children who could not be included in the analyses; the analysis includes only the 31 6- and 7-year-olds and 27 8- and 9-year-olds who chose the safer option on at least one regret trial and one relief trial in a Two Rating block. The proportions of categorical responses (happier, sadder, or the same) given by each age group in regret and relief trials are shown in Fig. 3. A two-way analysis of variance (ANOVA) with a between-participants factor of age and a within-participants factor of trial type (regret or relief) was conducted on the proportion of sadder responses. There was a significant effect of trial type, F(1, 56) = 86.61, p < .001, ηp2 = .61, with more sadder responses in the regret trials, and a significant interaction between age and trial type, F(1, 56) = 8.94, p < .01, ηp2 = .14. Further analyses (making a Bonferroni correction for 4 tests, α = .0125) showed that the effect of trial type was significant for each of the two age groups: 6- and 7-year-olds, t(30) = 4.28, p < .01; 8- and 9-year-olds, t(26) = 9.37, p < .01. However, the proportion of sadder responses in regret trials increased with age, t(56) = −2.96, p < .01. A further two-way ANOVA with a between-participants factor of age and a within-participants factor of trial type was conducted on the proportion of happier responses. The main effect of trial type was significant, F(1, 56) = 98.84, p < .001, ηp2 = .64, with more happier responses in relief trials than in regret trials, but there were no other significant effects. These analyses demonstrate that children’s emotional responses varied appropriately between regret and relief trials and that there was an increase with age in the likelihood that children reported feeling sadder in regret trials. We also examined whether children were more likely than chance to report feeling sadder in regret trials and happier in relief trials. Because there were three possible choices (sadder, happier, or the same), we compared whether proportions of sadder responses in regret trials and happier responses in relief trials differed significantly from 0.33 using a one-sample t-test. These analyses showed that the older group gave significantly more sadder responses in regret trials and happier responses in relief trials than would be expected by chance, t(26) = 5.09, p < .001 and t(26) = 6.49, p < .001, respectively; however, although the 6- and 7-year-olds gave more happier responses in relief trials than would be expected by chance, t(30) = 4.06, p < .001, the proportion of sadder responses in regret trials did not differ significantly from chance, t(30) = 0.98, p = .34. Note, however, that this group did not respond randomly in regret trials (see Fig. 3); the younger children produced very few happier responses (unlike in relief trials), with most of their responses being either sadder or the same. Taken together, these findings suggest unambiguously that the older children experienced both regret and relief. However, whereas the younger group experienced relief, the younger children did not consistently experience regret. Between-trial analysis of counterfactual emotions In this analysis, children’s initial emotional ratings following partial feedback for a specific trial type [taken from Stage (iii) in Two Rating trials] were compared with their ratings from an otherwise identical trial type in which they made a rating only following complete feedback (taken from One Rating trials). To generate categorical data analogous to those used in the within-trial analysis, we examined the proportion of trials in which the latter ratings were the same as, sadder than, or happier than the former ratings (see Fig. 4). To make the required comparisons for analysis, it was necessary to have data from two trials of the same type in which children made the safer choice (one where children made an emotional rating following partial feedback and one where they made a rating only after full feedback) for at least one regret and one relief trial. There were 19 children from each age group who made a sufficient number of safer choices to generate these data. A two-way ANOVA with a between-participants factor of age and a within-participants factor of trial type was conducted on the proportion of sadder responses. The main effect of trial type was significant, F(1, 36) = 36.67, p < .001, ηp2 = .51, as was the interaction between trial type and age group, F(1, 36) = 7.74, p < .01, ηp2 = .18. Post hoc t-tests (making a Bonferroni correction for 4 tests, α = .0125) showed that the effect of trial type was marginally significant for the 6- and 7-year-olds, t(18) = 2.04, p = .06, and significant for the 8- and 9-year-olds, t(18) = 7.39, p < .01. The effect of age was significant for the proportion of regret trials in which children made a sadder rating, t(36) = −2.77, p < .01, with older children being more likely to report feeling sadder on these trials. A two-way ANOVA with a between-participants factor of age and a within-participants factor of trial type was conducted on the proportion of happier responses. The main effect of trial type was significant, F(1, 36) = 17.95, p < .001, ηp2 = .33, as was the main effect of age, F(1, 36) = 5.69, p < .03, ηp2 = .14, with younger children being more likely to report feeling happier than older children across both trial types. However, the interaction between age and trial type was not significant, F < 1. In summary, although these analyses are based on an entirely different set of emotional ratings, the findings are very similar to those reported for the within-trial analysis in suggesting that both age groups experienced regret and relief and an increased likelihood of children experiencing regret with age. Additional analyses To check whether the within-trial and between-trial measures differed from each other in terms of sensitivity, for each age group we used paired-sample t-tests to compare the proportion of sadder responses in regret trials yielded by each measure as well as the proportion of happier responses in relief trials yielded by each measure. These analyses showed that none of these proportions differed from each other (all ps > .30) except for the difference between the proportions of happier responses in relief trials for the older children, for whom the within- trial measure yielded a significantly higher proportion of happier responses, t(18) = 2.36, p < .05. The experimental procedure involved children completing different numbers of trials, meaning that children had quite variable exposure to different sets of outcomes, which could have affected their responses as they moved through the task. Therefore, we conducted a final set of analyses in which we focused just on the very first Two Rating regret trial and Two Rating relief trial in which children made a safer choice. In the first regret trial on which children chose the safer box, 19 of 33 (57.6%) 6- and 7-year-olds and 23 of 31 (74.2%) 8- and 9-year-olds reported feeling sadder following complete feedback. In both of these groups, the probability of children reporting that they now felt sadder was significantly greater than chance on a binomial test, both ps < .01, assuming that one third of responses will be sadder by chance. In the first relief trial where children chose the safer box, 20 of 32 (62.5%) 6- and 7-year-olds and 23 of 27 (85.2%) 8- and 9-year-olds reported feeling happier following complete feedback, in both cases a proportion greater than chance, binomial test, both ps < .01.","Children’s initial emotion ratings after partial feedback demonstrated that they felt sadder if they had received the worse prize from their chosen box than if they had received the better prize. However, children of both ages were able to flexibly update this initial emotional response once they had seen the prize in the unchosen box. Analysis of the within-trial data indicated that children were then likely to subsequently report feeling sadder if the prize in the unchosen box was better (regret trials) but happier if the prize in the unchosen box was worse (relief trials). However, although the younger children did not respond randomly in regret trials (they very rarely gave happier responses, unlike in relief trials), they did not report feeling sadder significantly more often than chance (many of the children reported feeling the same). Thus, although this group did vary their emotional responses across trial types, they did not all consistently experience regret. The likelihood that children experienced regret increased developmentally, with older children more reliably reporting feeling sadder in regret trials. The between-trial comparisons of emotional ratings given in separate trials under partial or complete feedback also revealed that children’s emotional responses in both regret and relief trials was sensitive to whether or not children had received full feedback. However, for this measurement the proportion of sadder responses in regret trials was only marginally significantly different from that in relief trials for the younger group, and the likelihood that children would report feeling sadder in regret trials also increased with age. Thus, the pattern of findings is generally consistent across both ways of measuring regret and relief. We note, however, that the within-trial method may be somewhat more sensitive, with the proportion of older children reporting feeling happier in relief trials being larger under this measure than under the between- trial measure. The within-trial method also has the advantage of requiring administration of half the number of trials as the between-trial method, which is very advantageous when testing young children.","The findings of Experiment 1 were primarily based on a comparison between the proportions of sadder/happier responses on regret trials and those on relief trials. However, these trials differed not just in terms of whether the prize in the unchosen box was better or worse than the one in the chosen box; the choices that children needed to make on regret trials were not completely identical to those that they needed to make on relief trials (cf. Trials 1 and 2 with Trials 3 and 4). Moreover, the actual prize that children received from their chosen gamble on regret trials (10 or 11 points) was always better than the actual prize that children received on relief trials (6 or 4 points). This latter aspect of the design was deliberate to provide us with a test of whether children could update their emotion appropriately on seeing the prize in the unchosen box. However, it might be argued that these differences between regret and relief trials mean that it is difficult to straightforwardly compare patterns of responses across these two trial types. In our second experiment, we used a similar task as in Experiment 1 but made three key changes. First, we added an additional 4 trials to ensure that regret and relief trials were matched both in terms of the choices children needed to make in each trial type and in terms of the size of the actual prize obtained. These are listed in Table 1 as Trials 1a to 4a. Second, because the results of Experiment 1 had indicated that there was no benefit in also including a between-trial measurement of regret/relief, we used only Two Rating trials, yielding only a within-trial measurement of regret/relief. Third, all children received the same number of trials to ensure that they were all administered an identical task.","Participants were 38 6- and 7-year-olds (19 girls; M = 79.6 months, range = 72–94) and 36 8- and 9-year-olds (17 girls; M = 107.8 months, range = 96–119 months). All children were recruited from the same population as in Experiment 1.","The stimuli used were the same as those used in Experiment 1. Procedure The procedure was very similar to that of Experiment 1, with the addition of Trials 1a to 4a in Table 1. The experimental trials were presented in blocks of 4 Two-Rating trials, with each block consisting of 2 regret trials and 2 relief trials. Block 1 consisted of Trials 1, 3, 2a, and 4a, whereas Block 2 consisted of Trials 1a, 3a, 2, and 4. All children received the same number of trials in this task (the two blocks of 4 trials shown in Table 1). Children were randomly assigned to one of 24 possible trial orders, with half receiving a regret trial first and half receiving a relief trial first.","The 6- and 7-year-olds tended to choose the riskier box (M = 4.82, SD = 1.27) significantly more often than the 8- and 9-year-olds (M = 3.31, SD = 1.33), t(72) = 5.00, p < .001. As in Experiment 1, only data from trials in which children chose the safer box were analyzed. We first examined whether children’s initial emotion ratings following partial feedback differed depending on whether they had received the best prize (Trials 1, 1a, 2, and 2a) or the worst prize (Trials 3, 3a, 4, and 4a) in their chosen box. The average first emotion rating given in trials where children won the best possible prize from their chosen box was 6.9 (SD = 0.27) compared with 5.7 (SD = 1.01) in trials where children won the worst possible prize; children were initially significantly happier in the former trials, t(68) = 9.13, p < .001. Subsequent analyses focused on the proportion of times children gave sadder, happier, or the same responses once they had been given complete feedback. To be included in these analyses, children needed to make the safer choice in at least one regret trial and one relief trial; here, 29 of the 6- and 7-year- olds and 34 of the 8- and 9-year-olds were included. Fig. 5 shows the proportion of sadder, happier, and same responses for each trial type. A two-way ANOVA with a between- participants factor of age and a within-participants factor of trial type was conducted on the proportion of times children reported feeling sadder after complete feedback. There was a main effect of trial type, with more sadder responses given in regret trials than in relief trials, F(1, 61) = 231.02, p < .001, ηp2 = .79. No other effects were significant. Because the results of Experiment 1 indicated that the proportion of sadder responses given by older children was significantly greater than that given by younger children, planned comparisons examined this age effect. However, although the proportion of sadder responses in regret trials was larger for the older group, the age effect was not significant, t(61) = −1.12, p = .27. A further two-way ANOVA with a between-participants factor of age and a within-participants factor of trial type was conducted on the proportion of times children reported feeling happier after complete feedback. There was a main effect of trial type, with more happier responses being given in relief trials than in regret trials, F(1, 61) = 398.06, p < .001, ηp2 = .87, and no other effects were significant. Although all children completed the same number of trials in this experiment, the number of trials used in the analyses to yield the data depicted in Fig. 5 differed between children because children differed in terms of the number of times they chose the riskier box (M = 4.08, SD = 1.50, range = 0–8). This might be seen as problematic because children with a tendency to make riskier choices will produce fewer data points but also potentially may be less likely to experience regret or relief. To check this, we examined initially whether there were correlations between the number of times children made a riskier choice and the proportion of sadder responses in regret trials and between the number of times children made a riskier choice and the proportion of happier responses in relief trials. Neither of these correlations approached significance (both ps > .17). We also reran the ANOVAs described above including number of riskier choices as a covariate and found a qualitatively identical pattern of findings, with only the main effect of trial type being significant in both analyses. These analyses indicate that whether or not children experienced regret or relief was not confounded by the number of riskier choices they made. We conducted a final set of analyses on the very first regret trial and relief trial in which children chose the safe option. In their first completed regret trial, 22 of 35 (62.9%) 6- and 7-year-olds and 29 of 35 (82.9%) 8- and 9-year-olds reported feeling sadder. In both of these groups, the probability of children reporting that they now felt sadder was significantly greater than chance on a binomial test, both ps < .01. In their first completed relief trial, 25 of 31 (80.6%) 6- and 7-year-olds and 32 of 35 (91.4%) 8- and 9-year-olds reported feeling happier, in both cases a higher proportion than would be expected by chance on a binomial test, both ps < .01. These findings are very similar to those yielded by the analysis of the full set of trials. In summary, as in Experiment 1, children were initially sadder (following only partial feedback) if they had received the worst prize from their chosen box. This suggests that children felt disappointed if their hope to win the best prize from that box was not realized. In this experiment, we found that children of both ages showed regret on seeing that the prize they would have won in the unchosen box was better and showed relief if the alternative prize was worse. In Experiment 1, we had found somewhat ambiguous evidence that the younger group experienced regret; although the youngest children were more likely to report feeling sadder on regret trials than on relief trials, this difference was only marginally significant using the between-trial measurement, and using the within-trial measurement these children did not differ from chance in the proportion of times they reported feeling sadder in the former trials. The best interpretation of those findings was that only some of the younger children experienced regret. However, in Experiment 2 even the younger group reported feeling sadder on regret trials 70% of the time. Older children were more likely to feel regret than younger children (in 81% vs. 70% of regret trials), although unlike in Experiment 1 this difference did not reach significance. Why might the younger group be more likely to show regret in Experiment 2 than in Experiment 1? This is a different group of children, and it is possible that individual differences in ability might explain the contrasting patterns of results (although there are mixed findings over whether children’s ability, as assessed using standardized measures, predicts reported regret/relief; see Burns et al., 2012; McCormack & Feeney, 2015; O’Connor et al., 2014). There are also notable methodological differences between the experiments that may explain the differing findings. Experiment 2 included a different set of trials than Experiment 1, and in the new trials the magnitude of the difference between the obtained reward and the counterfactual alternative was larger than in the original set of trials (see Table 1; compare Trials 1 and 2 with Trials 3a and 4a). However, this does not seem to explain the difference. We compared the percentage of times the younger children reported feeling sadder in Trials 1 and 2 with the percentage of times they reported feeling sadder in Trials 3a and 4a and found no difference; the younger children were sadder 70% of the time in the former trials compared with 72% in the latter trials. Our best guess is that performance was somewhat better in the second experiment in the younger group because the experimental procedure was less complex. Children made the same pair of emotion ratings on every trial, whereas in Experiment 1 they completed both One Rating and Two Rating trials. Moreover, children always completed the same 8 trials in Experiment 2, but the vast majority of children completed more than this number of trials in Experiment 1 because of the need to repeat some trials. We note that our analysis of the first completed regret trial in Experiment 1 did show that as a group the younger children reported feeling sadder more often than would be expected by chance, whereas this was not the case for the full trial set. Thus, we suspect that our younger children may have found the larger set of trials in Experiment 1 to be onerous.","This study was the first to examine regret and relief in younger children using the type of risky decision-making task used to study these counterfactual emotions in adolescents and adults. Taken together, the findings of both experiments suggest that children as young as 6 or 7 years do experience both regret and relief in this sort of task, although the evidence for regret in this age group was stronger in the second experiment. These findings are consistent with developmental findings from other studies in which children make a simpler choice between two colored boxes. The majority of these previous studies have also found that regret and relief are present from at least 6 or 7 years of age (Burns et al., 2012; McCormack & Feeney, 2015; O’Connor et al., 2012, 2014; Weisberg & Beck, 2010, 2012). However, they are not consistent with the findings of Rafetseder and Perner’s (2012) study, which failed to find evidence of regret in children until 9 years of age. We discuss our findings in relation to the three methodological issues raised in the Introduction. Choice and risky decision making ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This study is the first to use a version of the gambling task with children as young as 6 or 7 years to examine counterfactual emotions. Its findings indicate that, like adolescents and adults, children of this age do experience both regret and relief in circumstances that involve risky choice. The findings pave the way for using the gambling task with a broader age range of participants to examine whether there are developmental changes between childhood and adulthood in the likelihood of experiencing these emotions. One of the reasons why there has been intense interest in counterfactual emotions in the gambling task is that researchers have been trying to establish the roles of these emotions in decision making. Thus, it has been argued that regret and relief may have an impact on how people make risky decisions and that decision making will be different in the absence of such emotions (e.g., due to brain damage) (Camille et al., 2004; Coricelli et al., 2007). This claim is consistent with long-standing suggestions in both economics (Bell, 1982; Loomes & Sugden, 1982) and psychology (Zeelenberg, 1999; Zeelenberg, Beattie, van der Pligt, & de Vries, 1996; Zeelenberg & Pieters, 2007) that individuals attempt to minimize future regret when making decisions. Thus, one important reason for examining the development of regret within the context of risky decision making would be to explore whether developmental changes in regret are accompanied by or indeed explain changes in risky decision making. Single reference point versus multiple reference points ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our findings suggest that at least by the time children are 6 or 7 years old, their emotions are sensitive to two different reference points. When given partial feedback, children were less happy if they had won the worst of two prizes in their chosen gamble than if they had won the best of two prizes. This can be interpreted as disappointment (Mellers et al., 1999; Zeelenberg, van Dijk, & Manstead, 2000) and suggests that children hoped to win the better prize. As we have discussed, children’s emotions were also sensitive to the outcome of the unchosen gamble, as indicated by their emotion ratings following complete feedback that strongly contrasted with their ratings following partial feedback. This suggests that any explanation of children’s emotion ratings in this task just in terms of a simple comparison such as “I do [do not] have the best prize” is inadequate, whereas such an explanation may be sufficient to explain such ratings following complete feedback in previous studies with young children. Following a safer choice, children are already aware, even before partial feedback, that they do not have the best or worst prize; their subsequent emotion ratings following partial feedback seem to be sensitive to what they could have won given their choice and, following complete feedback, sensitive to what they would have won had they chosen differently. Does this mean that we can rule out an explanation of children’s emotions in terms of simple frustration rather than regret or relief (see Rafetseder & Perner, 2012; Rafetseder & Perner, 2014)? We are assuming that this question remains open even under circumstances in which participants have full responsibility for the choice that led to the outcome and full feedback has been given (see O’Connor, McCormack, Beck, & Feeney, 2015). Certainly, any explanation of children’s emotions in terms of frustration needs to be more complex than assuming a simple comparison between the prize obtained and the best or worst prize available; it would need to assume that this frustration can mutate within a single trial following complete feedback (from frustration to happiness in relief trials or from happiness to frustration in regret trials). One might argue that all that needs to happen is for children to systematically change the reference point by which they make their judgment regarding whether they have the best (or worst) prize and that this need not involve thinking counterfactually about their choices. We note, however, that children’s judgments following complete feedback do seem to be sensitive not just to whether or not a better (or worse) prize was available but also to whether children would have won that prize had they chosen differently. That is, they made the comparison not between their actual prize and the possible contents of the unchosen gamble but rather between their actual prize and the specific outcome they would have obtained had they chosen the other gamble. Although we cannot definitely rule out the possibility that children make these comparisons without thinking counterfactually, we note that neuropsychological evidence suggests that the regions of the brain that seem to be important for being sensitive to complete feedback in this sort of gambling task also seem to be important for counterfactual thinking (Barbey, Krueger, & Grafman, 2009; but see Van Hoeck et al., 2013). Within-trial versus between-trial comparisons of emotion ratings ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As described in the Introduction, Rafetseder and Perner (2012), Rafetseder and Perner (2014) argued that a between-trial method is the most appropriate one to measure regret, whereas all of the other studies with young children have used only a within-trial method. We directly compared both of these methods in Experiment 1 in the current study to examine whether they showed different developmental patterns. For both the within-trial and between-trial measures, the developmental patterns were very similar, and under either method there was some evidence of a developmental increase in the likelihood that participants experienced regret but no evidence of developmental changes in relief. Our findings suggest that Rafetseder and Perner’s pattern of developmental findings cannot be explained simply in terms of their use of a between-trial method for assessing regret. Of course, although the basic comparison we made in our between-trial method was in essence the same as that made in Rafetseder and Perner’s between-trial method (i.e., a comparison of emotion ratings following either partial or complete feedback on separate trials), our task differed from that of Rafetseder and Perner considerably, with the latter involving a simple box-choosing task. These differences included not just the types of choice children needed to make but also the number of trials, the nature of the emotion scale used, the rewards involved, and so on. However, it is hard to see why these differences would make it more likely, rather than less likely, that we would observe counterfactual emotions in younger children; moreover, we note that in their own study, Rafetseder and Perner found no consistent evidence of regret in children under 9 years of age even using a within- trial method. The main reason for comparing the between-trial and within-trial methods was to inform future studies of the development of counterfactual emotions. Although both of these methods yielded broadly similar results, the within-trial method requires administration of fewer trials and also seems to yield a somewhat clearer pattern of findings. In Experiment 1, we found that the difference in the proportion of sadder responses given in regret trials versus relief trials was only marginally significant for the younger children in the between-trial method and that evidence for relief in the older group was significantly stronger using the within-trial method. One possible explanation of the clearer pattern of findings using the within-trial method may lie in the fact that in Two Rating trials we explicitly asked participants to make a comparative rating, that is, to report whether they felt the same as, happier than, or sadder than before they had received complete feedback. This may have encouraged participants to proactively compare the actual outcome with the counterfactual outcome. Moreover, these comparisons did not hinge on the sensitivity of the scale used to report emotions because we used the three- pronged arrow method of describing relative emotions (see also O’Connor et al., 2012; O’Connor et al., 2014; Weisberg & Beck, 2012). By contrast, between-trial comparisons involved comparing separate ratings on the same 7-point scale, and whether such comparisons differ will depend on the sensitivity of this scale. Our findings suggest that future developmental studies should use the within-trial method, at least for this type of task.","This study has demonstrated that children as young as 6 or 7 years do experience counterfactual emotions in the context of risky decision making; as such, it may pave the way for future studies that examine whether these emotions have an impact on children’s decision making. Such studies are important because existing accounts of developmental changes in children’s decision making (for reviews, see Dhami, Schlottmann, & Waldmann, 2013; Jacobs & Klaczynski, 2005) have typically not considered how counterfactual thinking and its emotional consequences have an impact on children’s decisions, whereas it is widely argued that counterfactual cognition and its associated emotions play an important role in adult decision making (Connolly & Zeelenberg, 2002; Coricelli et al., 2007; Roese, 1999)."],["We examined how the strength of the size–weight illusion develops with age in typically developing children. To this end, we recruited children aged 5–12 years and quantified the degree to which they experienced the illusion. We hypothesized that the strength of the illusion would increase with age. The results supported this hypothesis. We also measured abilities in manual dexterity, receptive language, and abstract reasoning to determine whether changes in illusion strength were associated with these factors. Manual dexterity and receptive language did not correlate with illusion strength. Conversely, illusion strength and abstract reasoning were tightly coupled with each other. Multiple regression further revealed that age, manual dexterity, and receptive language did not contribute more to the variance in illusion strength beyond children's abilities in abstract reasoning. Taken together, the effects of age on the size–weight illusion appear to be explained by the development of nonverbal cognition. These findings not only inform the literature on child development but also have implications for theoretical explanations on the size–weight illusion. We suggest that the illusion has a strong acquired component to it and that it is strengthened by children's reasoning skills and perhaps an understanding of the world that develops with age. --------------------------------------------------------------------------------","The size–weight illusion refers to the perceptual experience of object weight that occurs when a person lifts equally weighted objects that differ in size (Charpentier, 1886, 1891). Namely, smaller objects typically feel heavier than larger objects of the same mass. Although the illusion has been studied for more than 100 years, its precise mechanisms are not completely understood and have yet to be explained satisfactorily. Over the years, researchers have proposed a number of different theoretical explanations for the illusion (Buckingham, 2014; Dijker, 2014; Saccone & Chouinard, 2019). According to sensorimotor explanations, the illusion is driven by the misapplication of fingertip forces during lifting (Dijker, 2014). When lifting two objects that weigh the same, people apply more force for the larger object than they do for the smaller one, which causes too much lift for the former and too little for the latter. According to these explanations, too much lift will make the object feel lighter, whereas too little lift will make the object feel heavier. However, Flanagan and Beltzner (2000) demonstrated how the motor system quickly learns to apply the correct amount of force for each object after only a few trials, whereas the perception of differences in weight remains the same in magnitude across many trials. This dissociation has been replicated several times (Buckingham & Goodale, 2010a, 2010b; Chouinard, Large, Chang, & Goodale, 2009; Grandy & Westwood, 2006), which undermines the necessity of the underlying mechanisms proposed by sensorimotor theories. Alternatively, other theories offer more top-down explanations, whereby the illusion arises from the brain comparing prior experiences and current sensory input, which in turn influences weight perception. Bayesian explanations, which have grown in popularity during recent years to comprehensively explain many forms of perception, posit that perceptual experiences are the result of an active process of formulating and testing hypotheses about the world. It then follows that experiences, or priors, are important in shaping perception (Geisler & Kersten, 2002; Gregory, 1980; Helmholtz, 1867). If one applies these principles to weight perception, then the perceived heaviness of an object should be influenced by any associations that we have developed over time between weight and other physical properties of objects. For example, small objects are usually lighter in the real world. According to Bayesian explanations, one should expect the smaller object in the size–weight illusion to weigh less, which consequently should make that object feel lighter during lifting. However, the reverse is experienced in the size–weight illusion, which is why the illusion is sometimes called an anti-Bayesian illusion (Brayanov & Smith, 2010). Nevertheless, as Peters, Ma, and Shams (2016) demonstrated through mathematical modeling, and discussed further by Saccone and Chouinard (2019), the size–weight illusion fits perfectly well within a Bayesian framework if one considers that the perceived heaviness of objects might be driven by expected density as opposed to size (Chouinard et al., 2009; Harshfield & DeHardt, 1970; Peters et al., 2016; Ross & Gregory, 1970). Manipulable man-made objects tend to be denser as they get smaller, and people consequently tend to estimate their weight according to this relationship (Peters, Balzer, & Shams, 2015). Thus, in the context of the size–weight illusion, the smaller and denser object affords more weight, which causes that object to feel heavier. If priors are indeed important for the size–weight illusion, then one might expect to find increases in the strength of the illusion during child development. It is conceivable that an adult with years of experience in lifting objects would have a more expansive and deep-seated repertoire of priors about their affordances and would use these priors for the purposes of perception with greater efficiency and influence than a young child whose motor skills are still developing and who has lifted fewer objects. Binet (1895) and Piaget (1969, 1999) proposed that illusions, such as the size–weight illusion, offered opportunities to understand typical development and tease apart perceptual mechanisms that might be innate from those that are acquired. Their arguments stem from their own research demonstrating how the strength of illusions can either decrease or increase with age depending on the illusion (Binet, 1895; Piaget, 1969, 1999). Namely, they argued that innate illusions decrease in strength as children age, whereas the strength of acquired illusions increases with age (Binet, 1895; Piaget, 1969, 1999). The Müller–Lyer illusion is an example of an innate illusion that decreases in strength with age. This was first demonstrated by Binet (1895) and has since been replicated several times (Brosvic, Dihoff, & Fama, 2002; Frederickson & Geurin, 1973; Hanley & Zerbolio, 1965; Pollack, 1970; Porac & Coren, 1981), but not always (Rival, Olivier, Ceyte, & Ferrel, 2003). Binet (1895) posited that it can be maladaptive to falsely perceive something differently than what it truly is and that children learn to suppress these forms of misperception, but only as cognitive faculties and an understanding of the world develop. Another possibility is that the perceptual system in younger children might exaggerate illusions as a way to compensate for not being able to account for sensory noise as effectively as more developed systems in older children (Duffy, Huttenlocher, & Crawford, 2006). More recently, Gandhi, Kalia, Ganesh, and Sinha (2015) demonstrated how children who gain sight for the first time after the surgical removal of congenital cataracts can see the Müller–Lyer illusion immediately after surgery. These reports on the Müller–Lyer illusion are difficult to explain within a Bayesian framework and call into question the necessity of priors. If experience were truly essential for perceiving the Müller–Lyer illusion, then the illusion would not be strongest during early childhood and the children in Gandhi et al.’s study would not experience the illusion immediately after gaining sight for the first time. Regarding the size–weight illusion, there is little research on the cognitive and sensorimotor explanations of the illusion in typically developing children. Table 1 provides a summary of all articles written in the English and French languages to date. Both the methods and results of these investigations are mixed. Consequently, the developmental profile for the illusion remains unresolved. Some studies demonstrate that the illusion is weaker in younger children and strengthens as children grow older (Philippe & Clavière, 1895; Rey, 1930), whereas others show the reverse findings whereby the strength of the illusion decreases with age (Robinson, 1964). The studies range in procedures from quickly administered perceptual ranking of stimuli (Dresslar, 1894; Flournoy, 1894; Philippe & Clavière, 1895) to lengthy testing sessions using the methods of constant stimuli (Pick & Pick, 1967; Rey, 1930; Robinson, 1964). Many of the earlier studies did not perform statistical analyses and failed to consider other factors that may have influenced their results such as manual dexterity and cognitive ability. Given the variability in procedures and the number of extraneous variables not considered, it is perhaps not surprising that this literature is contradictory and, therefore, warrants further investigation using modern-day methods and standards, which are far more rigorous. Currently, little is known about how the size–weight illusion might develop with age when other variables, such as motor and cognitive skills, are considered. Although there is evidence suggesting that the illusion might have a strong innate bottom-up component to it (Saccone & Chouinard, 2019), conceptual knowledge can nonetheless also influence weight illusions in a top-down manner (Buckingham, 2014; Saccone & Chouinard, 2019), as evidenced by the material–weight illusion in which objects that appear to be metallic feel lighter than objects that appear to be made of Styrofoam of the same size and mass (Buckingham, Cant, & Goodale, 2009; Seashore, 1898). Logically, conceptual knowledge can be obtained only when cognitive faculties are sufficiently developed to understand new experiences. Conceivably, a certain amount of manual dexterity is also required when forming an association between an object’s weight and its features because manual dexterity is necessary for gauging weight (Jones, 1986). Manual skills emerge during early infancy and continue to develop into adolescence (Mathiowetz, Federman, & Wiemer, 1985). For preschool children, manual dexterity improves as they learn to cut with scissors and to trace and copy lines and shapes. In primary school, children continue to improve their manual skills with handwriting and the use of computers. More experiences are acquired as these skills develop. Children’s repertoire of priors consequently becomes richer and more fine-tuned. Thus, in theory, previous experiences with objects can only begin to influence how children perceive their weight when cognitive and motor skills reach certain levels of proficiency. These are not new ideas. Piaget (1969, 1999) reasoned that children must first understand size and weight, be able to integrate the two collectively, and know that larger objects typically weigh more than smaller ones before children can experience the size–weight illusion. In line with this thinking, we hypothesized that illusion strength increases as children develop their manual and cognitive skills. Our findings demonstrate that the size–weight illusion increases in strength as children grow older and that the development of nonverbal cognition, but not motor or language skills, explains variability in the development of the size–weight illusion. Overview ~~~~~~~~ Children from primary schools from a regional center (Bendigo) and a major city (Melbourne) in Victoria, Australia, completed tasks that assessed susceptibility to the size–weight illusion, manual skills, receptive language, and abstract reasoning. Task order was counterbalanced across participants to reduce practice or carryover effects. Testing occurred over two or three sessions, lasting no more than 30 min each. All procedures were approved by the La Trobe University human ethics committee, the Department of Education and Training of Victoria, Catholic Education Melbourne, and the local schools. Legal guardians of all participants provided informed written consent and confirmed that their children were never diagnosed with a psychological, psychiatric, neurological, or neurodevelopmental disorder by a questionnaire prior to testing. Testing was administered at the children’s school in cooperation with classroom teachers to minimize disruption.","A total of 85 typically developing children participated in the study. Of these participants, 2 boys and 5 girls were excluded from the analyses based on having a standard score lower than 70 on either the Peabody Picture Vocabulary Test or the Raven’s Progressive Matrices (see below for more information about these tests), which is suggestive of an intellectual disability. An additional 3 boys and 3 girls were excluded from the analyses based on having an illusion strength index (see below on how this was calculated) exceeding ±2 standard deviations. Removing these participants in this manner helped to systematically and objectively remove noise from the data that would otherwise reflect various aspects of misunderstanding, noncompliance, or atypical development. This resulted in a final sample size of 72 participants (39 boys; 64 right-handers; mean age = 8.9 years, range = 5.6–12.5). Age was determined from the date of birth provided by parents on a form filled out prior to the children’s participation, which was confirmed by the children during testing. Size–weight illusion The task objects consisted of four plastic spheres created with a three- dimensional printer. We adjusted the weight of each plastic sphere by placing a ballast of lead pellets inside it. The ballast was held in place with compact foam to ensure that the center of mass corresponded to the sphere’s center. There were two pairs of objects (Fig. 1). The first pair consisted of two spheres weighing 120 g. One sphere had a diameter of 6 cm (volume = 113.10 cm3, density = 1.06 g/cm3) and was painted yellow (luminance = 62.2 cd/m2), whereas the other one had a diameter of 9 cm (volume = 381.69 cm3, density = 0.31 g/cm3) and was painted red (luminance = 8.2 cd/m2). The second pair consisted of two spheres weighing 405 g. One sphere had a diameter of 6 cm (volume = 113.10 cm3, density = 3.58 g/cm3) and was painted blue (luminance = 7.6 cd/m2), whereas the other one had a diameter of 9 cm (volume = 381.69 cm3, density = 1.06 g/cm3) and was painted green (luminance = 7.3 cd/m2). We also constructed a sliding measurement apparatus to allow participants to give a nonverbal magnitude estimate of the weight for each sphere (Fig. 1). This apparatus consisted of four sliders that were 30 cm in length. Each one was painted a different color corresponding to one of the spheres. Seated opposite the participant at a table, the experimenter gave the following instructions: “I’m going to hand you these colored balls. I want you to tell me how heavy they are using the sliding rings in front of you. If this side [indicating the participant’s left] means light and this side [indicating the participant’s right] means heavy, slide the ring and leave it wherever you think it should go.” These instructions were elaborated if needed, and the experimenter ensured that the participant understood the instructions before proceeding. The participant was then asked to extend his or her hands facing upward with the elbows elevated above the table. The experimenter then placed one object from one of the two pairs in each hand for the participant to weigh (e.g., the yellow sphere in the left hand and the red sphere in the right hand). The participant was given as much time as needed to assess the objects’ weights. Afterward, the participant was asked to use the color-coded sliders to indicate the perceived weight of each sphere, and the experimenter recorded the final measurements in millimeters. This was repeated for the next pair of objects (e.g., the blue sphere in the left hand and the green sphere in the right hand). These procedures were repeated with the objects placed in the opposite hands. The average position of the slider in millimeters for the two presentations of each object was taken as a measurement of its apparent weight. The order of presentation was counterbalanced across participants. Purdue pegboard test We also administered the Purdue Pegboard Test to assess manual dexterity (Tiffin & Asher, 1948). The test consisted of the participant manually inserting pegs into columns of small holes one at a time for 30 s with the participant’s preferred hand. The total number of pegs placed inside a hole was taken as the score. Hand preference was determined by the child’s self-report when asked whether he or she was right- or left-handed. Peabody Picture Vocabulary Test We measured receptive language with the Peabody Picture Vocabulary Test–Fourth Edition (PPVT; Dunn & Dunn, 2007). The child was presented with a series of pages containing four pictures and was asked to indicate which picture he or she thought best described the item word spoken by the administrator. The complete test consists of 228 trials. However, the number of trials administered to each child was determined by basal and ceiling rules in accordance with instructions from the test manual (Dunn & Dunn, 2007). Raw scores reported in our study reflect the total number of correct trials plus credit for all trials not administered below the basal start point. Standard scores were also calculated based on normative data obtained from the test manual to characterize verbal intelligence in our overall sample and across our age groups. Raven’s progressive matrices The Raven’s Progressive Matrices (RPM; Raven, Raven, & Court, 2003) is a nonverbal measure of general cognitive ability. The child was provided with a booklet of different patterns with a piece missing in each pattern. For each item, the child was required to select which piece from an array of different options best matched the missing piece. We administered two versions of the RPM, each designed for a different age group. The colored version was administered to children aged 5–9 years and consisted of 36 trials, whereas the standard version was used for the older participants and consisted of 60 trials. Raw scores reflected the number of trials that the participant got correct. For the purposes of data analysis and reporting, all raw scores on the colored form were converted to the scale of the standard form using the conversion table provided in the RPM manual (Raven et al., 2003). Standard scores were also calculated based on normative data obtained from the test manual to characterize nonverbal intelligence in our overall sample and across our age groups.","We carried out statistical analyses using GraphPad Prism–Version 7 (GraphPad, La Jolla, CA, USA), JASP software–Version 0.8 (University of Amsterdam, Amsterdam, Netherlands), and SPSS–Version 23 (IBM, Armonk, NY, USA). Before proceeding to any statistics, we first calculated a score of illusion strength from the magnitude estimates in the following manner: [(perceived weight of the small object − perceived weight of the large object)/(perceived weight of the small object + perceived weight of the large object)]. An overall score of illusion strength for the two pairs of objects was calculated by taking their average.1 This method of normalizing is used in many illusion studies (Chouinard, Noulty, Sperandio, & Landry, 2013; Chouinard, Peel, & Landry, 2017; Chouinard, Royals, Sperandio, & Landry, 2018; Chouinard, Unwin, Landry, & Sperandio, 2016; Schwarzkopf, Song, & Rees, 2011; Sherman & Chouinard, 2016) and allows for meaningful comparisons across studies. We used two approaches to analyze the data. The first consisted of comparing means between different age groups. To this end, we first divided our participants into quartile age groups. The quartile split ensured that a sufficient number of participants (n = 18) were evenly distributed in each group. Age Group 1 ranged from 5.6 to 6.9 years, Age Group 2 ranged from 7.0 to 9.3 years, Age Group 3 ranged from 9.3 to 10.8 years, and Age Group 4 ranged from 10.8 to 12.5 years. We then performed an analysis of variance (ANOVA) with age as a between-subject factor on illusion strength and the raw scores on the Purdue Pegboard Test, PPVT, and RPM. Raw scores were chosen for the three latter tests so that we could chart how these skills develop with age. Post hoc pairwise comparisons using Tukey’s honest significance difference (HSD) tests (Tukey, 1949), which corrected for multiple comparisons, were performed to test for differences between the various age groups when a main effect of age was obtained. We also performed one-sample t tests to determine whether or not the illusion strength index in each age group differed from zero, which provides an indication as to when the illusion might emerge during development. To account for multiple comparisons against zero, we applied a Bonferroni correction to the reported p values (i.e., pcorr = puncorr × number of tests comparing differences against zero). The second approach consisted of performing bivariate correlations and a forward selection multiple regression. Specifically, a correlation matrix of Pearson r coefficients was produced to assess for associations among age, illusion strength, the Purdue Pegboard Test scores, the raw PPVT scores, and the raw RPM scores. Again, raw scores were chosen on the three latter tests so that we could chart how these skills develop with age. To account for multiple correlations, we applied a Bonferroni correction to the reported p values (i.e.. pcorr = puncorr × number of bivariate correlations performed). The model for the forward selection multiple regression began with an empty equation. Age, the Purdue Pegboard Test scores, the raw PPVT scores, and the raw RPM scores were added to the model one at a time beginning with the one with the highest correlation with illusion strength until the model could no longer be improved. Given that we had no prior predictions on how to model the regression, this type of regression was favored over others for its exploratory and unbiased nature for determining which predictors should be entered into the model and are most important for explaining illusion strength. Participants with missing values were excluded from the analysis. The resulting standardized beta coefficients (β) and corrected p values arising from the multiple regression analysis are reported. All reported p values were corrected for multiple comparisons based on an alpha level of .05 unless specified otherwise. Comparison of age groups on illusion strength as well as on motor and cognitive abilities ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In summary, ANOVA demonstrated increases in illusion strength, manual dexterity, receptive language, and abstract reasoning with increasing age. Bivariate correlations and multiple regression ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 4 provides a correlation matrix of Pearson r coefficients among chronological age, manual dexterity, receptive language, abstract reasoning, and illusion strength. Age, manual dexterity, receptive language, and abstract reasoning were highly intercorrelated with each other (all rs ≥ .44, p ≤ .001). Illusion strength increased as a function of age, r(70) = .34, p = .034 (Fig. 3A) and abilities in abstract reasoning, r(69) = .33, p = .046 (Fig. 3B). In contrast, illusion strength was not correlated with either manual dexterity, r(70) = .13, p = 1.00, or abilities in receptive language, r(70) = .27, p = .217. To establish the importance of these different variables in predicting illusion strength, a forward selection approach was used in a multiple regression. In this analysis, age, manual dexterity, receptive language, and abstract reasoning were selected as predictors for entry. The first step, which included abstract reasoning as the predictor, was significant, F(1, 68) = 7.72, p = .007, and explained 10.2% of the variance in illusion strength. The standardized beta coefficient for abstract reasoning was significant (β = 0.32, p = .007). The forward selection analysis did not add age, manual dexterity, or receptive language as additional predictors, indicating that none of them significantly explained more variance in illusion strength beyond what was already shared with abstract reasoning.","We sought to characterize the development of the size–weight illusion in typically developing children and determine the contribution of manual dexterity, receptive language, and abstract reasoning underlying these changes. As hypothesized, the strength of the illusion increased with age. Manual dexterity and receptive language did not correlate with illusion strength. Conversely, illusion strength and abstract reasoning were tightly coupled. The multiple regression revealed that age, manual dexterity, and receptive language did not contribute significantly more than the 10.2% of variance in illusion strength already explained by abstract reasoning. Taken together, the effects of age on the size–weight illusion appear to be explained by nonverbal cognition. In the ensuing discussion, we outline how this study contributes to our understanding of the size–weight illusion and how our findings should be interpreted within the context of previous research that has characterized the development of the size–weight illusion in children. Factors contributing to the size–weight illusion ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Illusion strength increased with abstract reasoning skills as measured by the raw scores on the RPM. Conversely, illusion strength did not change with manual dexterity as measured with the Purdue Pegboard Test, nor did it change with language skills as measured with the PPVT. These findings provide clues about the mechanisms underlying the size–weight illusion. In particular, our findings underscore the importance of cognitive processing and are in line with Piaget’s ideas that certain cognitive faculties need to be developed before a child can experience the size–weight illusion to its fullest (Piaget, 1969, 1999). Namely, the illusion increases in strength as reasoning skills also increase. Reasoning skills are conceivably important for conceptually understanding size, weight, and how the two are distinct and typically associated with each other (Piaget, 1969, 1999). Future research can verify this by testing children’s understanding of size and weight and correlating this with illusion strength. Previous research demonstrates that children begin to conceptually understand size at 3 years of age (Smith, 1984) and weight at 5 years of age (Cheeseman, McDonough, & Clarke, 2011). If Piaget’s (1969, 1999) theory is correct, then forming associations between size and weight must proceed these stages. Only then can associations be reinforced to exert an influence on the illusion. In line with this thinking, the ordinary rectangle is perceived as an illusion in adults (Ganel & Goodale, 2003) but not in children aged 4 and 5 years (Hadad, 2018). In adults, the apparent width of the ordinary rectangle is contingent on its length. Longer rectangles are perceived as more narrow than shorter rectangles with the same width. Hadad (2018) examined the developmental profile of this illusion in children aged 4–8 years. The 4- and 5-year-olds could not see the illusion, whereas the 7- and 8-year-olds could. From these results, Hadad concluded that children can begin to perceive the ordinary rectangle as an illusion only after they gain the ability to process and combine width and length information. The lack of correlation between the PPVT and illusion strength is also revealing for two reasons. First, the PPVT measures receptive vocabulary, which is the ability to comprehend language. Hence, comprehension of task instructions cannot explain illusion strength in the overall sample. Second, the lack of a correlation suggests that the mechanisms underlying the size–weight illusion do not depend on language processing. This is not a particularly contentious finding. We are unaware of any size–weight illusion explanation that is centered on language processing. The lack of correlation between manual dexterity and illusion strength is also informative because it sheds light on the merits of sensorimotor explanations for the size–weight illusion (Dijker, 2014). It is conceivable that age-related improvements in manual dexterity are accompanied by more accurate and precise somatosensory information regarding the size and weight of objects. Yet our results reveal that more refined motor skills, and the possibility of more veridical size and weight information obtained by somatosensory channels, did not translate into a stronger illusion in the overall sample, nor did it explain age-related changes in illusion strength. These findings add to the growing evidence demonstrating a dissociation between how one handles objects motorically and their perceived weight (Buckingham & Goodale, 2010a, 2010b; Buckingham, Ranger, & Goodale, 2012; Chouinard et al., 2009; Flanagan & Beltzner, 2000; Grandy & Westwood, 2006). Nonetheless, sensorimotor explanations should not be discarded entirely (Saccone & Chouinard, 2019). The development of rudimentary motor abilities is likely to be an important precursor to the development of the size–weight illusion. Only by manually handling objects can one reinforce associations between size and weight, which can then strengthen the illusion. Further investigation in younger children on the size–weight illusion would be needed to test this explanation. Earlier research on the development of the size–weight illusion ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The examination of the size–weight illusion in children emerged during the 1890s (Dresslar, 1894; Flournoy, 1894; Gilbert, 1894; Philippe & Clavière, 1895). This pioneering work seemed to be geared toward determining whether the size–weight illusion first described by Charpentier (1886, 1891) was innate or acquired and develops with age, and also whether the task could be used as an index of intelligence that could differentiate typically developing children from delayed or disabled children. These earlier studies used either a matching paradigm (i.e., the children needed to indicate which stimulus from an array of comparison stimuli weighed the same as the standard stimulus) or a rank-ordering paradigm (i.e., the children were given an array of stimuli and needed to order them according to their apparent weight) (Table 1). The results from these early studies converge well with our findings in that they also demonstrate that the size–weight illusion is weaker in younger children. Philippe and Clavière (1895) tested children aged 3–7 years and found that the majority of them did not experience the illusion. The other studies tested older children from 6 years of age and found that the illusion was present in either all or the vast majority of the participants (Dresslar, 1894; Flournoy, 1894; Gilbert, 1894). Using a method of constant stimuli paradigm, Rey (1930) later confirmed these trends, with children aged 5 and 6 years being less susceptible to the illusion than children aged 7–14 years. Two other studies later emerged during the 1960s (Pick & Pick, 1967; Robinson, 1964). Neither converges with the earlier findings and with the findings obtained in our investigation. The first of these was by Robinson (1964). The study was influenced by behaviorism (Skinner, 1953), which featured prominently in psychological research at the time. Being concerned that the younger participants might not understand the concept of weight, Robinson (1964) introduced an intensive reinforcement training phase before testing them on the illusion. Namely, children as young as 2 years were trained by reinforcement to indicate which of two objects differing in mass was heaviest. The participants received a food reward whenever they got the answer correct. The introduction of this kind of reinforcement training likely influenced the outcome of the testing phase in which the author examined the magnitude of the size–weight illusion using a method of constant stimuli paradigm. The youngest children required more reinforcement training to reach the learning criterion than the older children, which could have magnified their subjective reports during the testing phase to please the experimenter. Perhaps for this reason, contrary to all earlier work, as well as the current investigation, Robinson demonstrated that the strength of the illusion decreased with age. The second study was by Pick and Pick (1967). The authors performed a series of experiments to characterize how the strength of the illusion increased with age in children aged 6–16 years when haptic, vision, or both types of cues specifying object size were available to the participants. The participants hefted the objects wearing a blindfold in the haptic-only condition, lifted the objects using strings in the visual-only condition, and hefted objects without a blindfold in the haptic and visual condition. The study yielded some interesting dissociations. The authors demonstrated that illusion strength (a) was the same for all ages when both haptic information and visual information were provided, (b) increased with age when only haptic information was provided, and (c) decreased with age when only visual information was provided. We view these results with some skepticism. It is unclear as to why the direction of change with age would depend on the sensory modality of available cues. Pick and Pick did not offer any explanation. In addition, this dissociation has not been described further since it was first reported by the authors more than 50 years ago. There is the possibility that the effect of age in the visual-only condition differed in its direction because of differences in the manner in which the participants lifted the objects as opposed to the manner in which object size was presented to them. We know of only one other study of the size–weight illusion in children that was performed since the 1960s. This study was performed by Kloos and Amazeen (2002) in preschool children aged 3–5 years. The study’s paradigm was simple (Table 1) and arguably more conducive to testing very young children than the paradigm used in our study. In short, the children were shown a picture of a mouse holding a block of cheese at the bottom of a steep hill while they held a task object representing the cheese in one hand. With the other hand, the children pointed to a position on the hill to indicate where the mouse might take a break if it needed to walk up the hill, which served as an index of the children’s perceived weight of the object they were holding. Using this paradigm, the authors demonstrated that preschoolers perceive the size–weight illusion. The effects of age were not investigated given the small age range tested. Kloos and Amazeen could have perhaps also demonstrated increases in illusion strength had they included older children in their sample. Nonetheless, their results are important. They suggest that the size–weight illusion might not be completely acquired with experience but that the illusion is present from early development. Our study further reveals that the illusion is reinforced by cognitive development, which we speculate is required for the acquisition of priors and their influence on perception. Methodological considerations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The study demonstrates that the size–weight illusion has a strong acquired component to it given that the age-related changes were linked to cognitive development. However, one should not consider the absence of an illusion in the youngest age group as evidence that the size–weight illusion is completely acquired. It may still have an innate element to it. After all, Kloos and Amazeen (2002) demonstrated that the illusion is present in children younger than those we tested. A simpler and perhaps more engaging task, like the one used by Kloos and Amazeen, could have increased this study’s sensitivity for detecting size–weight illusion effects in younger children. Another consideration is that the task objects had different colors and did not match in luminance (Fig. 1). We painted the spheres different colors to make them more interesting for the children and to help facilitate the matching of each one to a slider. Previous research has demonstrated that the luminance of objects can have a small effect on their perceived weight. Specifically, lighter objects feel heavier (De Camp, 1917; Walker, Francis, & Walker, 2010). This effect could have contributed to a stronger weight illusion in the red–yellow pair, which had a considerable luminance difference of 53.8 cd/m2, than in the blue–green pair, which had a negligible luminance difference of 0.3 cd/m2. Color can also influence the perceived size (Tedford, Bergquist, & Flynn, 1977; Walker et al., 2010) and weight (De Camp, 1917; Walker et al., 2010) of objects. However, these effects disappear when luminance is matched between objects (Walker et al., 2010), which is perhaps why color does not always influence the perceived weight of objects (Buckingham, Goodale, White, & Westwood, 2016; De Camp, 1917). Although we cannot discard the possibility that the luminance differences between the red and yellow spheres contributed to the weight illusion in this pair of objects, the effects of color (driven by luminance differences) on perceived weight is generally regarded as much weaker than those exerted by size (Buckingham, 2014; Saccone & Chouinard, 2019; Walker et al., 2010). In addition, one should consider that the Purdue Pegboard Test measures only fine finger dexterity and does not measure other aspects of manual performance such as applied forces and gross movements. The Purdue Pegboard Test was chosen over other tests because it is quick to administer and continues to be the most widely used clinical test for assessing manual dexterity after 70 years of existence (Yancosek & Howell, 2009). Its reliability and validity are excellent and are understood better than any other dexterity assessment (Yancosek & Howell, 2009). Future work could examine the application of fingertip forces, as is the case in many studies of the size–weight illusion in adults (Buckingham & Goodale, 2010a; Chouinard et al., 2009; Flanagan & Beltzner, 2000; Grandy & Westwood, 2006). To our knowledge, the recording of fingertip forces has never been performed in children experiencing the size–weight illusion. Closing remarks ~~~~~~~~~~~~~~~ The current study provides a novel investigation into the development of the size–weight illusion in primary school-aged children by considering a number of factors not considered in previous studies. We conclude that the development of the illusion coincides with age- related changes in nonverbal cognition rather than motor or language skills. We argue that the findings add to the growing amount of evidence supporting expectancy-based theories of the illusion, which depend on being able to form and apply conceptual knowledge."],["Young children experience wayfinding difficulties. A better understanding of the development of wayfinding abilities may inform strategies that can be used to improve these skills in children. The ability to learn and remember a route was assessed in 220 6-, 8-, and 10-year old children and adults. Participants were shown a route in a virtual environment, before they were asked to retrace this route until they had achieved two consecutive trials without error. The virtual environment contained (i) no landmarks (ii) landmarks or (iii) landmarks that were verbally labelled. Adults, 10-year-olds and most 8-year-olds learnt the route when landmarks were present, but not all the 6-year-olds were successful. All age groups of children improved when the landmarks were labelled. Children were much poorer when there were no landmarks. This is the first study to distinguish between route learning dependent on landmarks, and route learning without landmarks (i.e. dependent on directions). --------------------------------------------------------------------------------","The development of spatial abilities have been of interest to psychologists for decades (Acredolo, 1977, 1978; Bullens, Iglói, Berthoz, Postma, & Rondi-Reig, 2010; Hermer & Spelke, 1994; Nardini, Thomas, Knowland, & Braddick, 2009). Spatial tasks such as the reorientation task have been used to explore the developmental trajectory of spatial abilities. Wayfinding is a spatial ability that is used by most people every day and refers to the ability to learn and remember a route through an environment. Wayfinding involves learning routes successfully and there are various strategies for encoding and retracing a route (Kitchin & Blades, 2001). The two most important ones are a landmark based strategy in which an individual learns that a particular landmark indicates a turn (e.g., turn left at the sweet shop), and a directional strategy, when an individual learns a route as a sequence of junctions (e.g., turn left, then left again and then turn right). Such strategies are not mutually exclusive but can be used together for effective route learning. However, previous research has suggested that young children's route learning may be particularly dependent on recalling landmarks (Cohen & Schuepfer, 1980; Heth, Cornell, & Alberts, 1997). As discussed below there is much evidence that children rely on landmarks, but previous research has not been able to distinguish performance based on learning landmarks and performance based on learning directions, and we do not know whether, in the absence of landmarks, children's wayfinding will be negatively affected. This has often been an assumption (Kitchin & Blades, 2001), but for the reasons explained below this is an assumption that has rarely been tested. Landmarks are considered to be important for successful performance on spatial tasks such as the reorientation task (Lee, Shusterman, & Spelke, 2006; Nardini et al., 2009) and the ability to use landmarks has often been considered to be an essential foundation for children's successful wayfinding (Cornell, Heth, & Alberts, 1994; Courbois, Blades, Farran, & Sockeel, 2012). In one of the first theories of wayfinding, Siegel and White (1975) emphasised the importance of landmarks by arguing that landmarks were the first elements that children learnt visually when encoding a route. Only after learning the landmarks along a route did children relate the landmarks to particular turns. Finally, children could demonstrate an understanding of the relationship between different points of the route in a survey like representation of the environment (Siegel & White, 1975). Siegel and White (1975) emphasised how a landmark would cue a turn, so that when a child saw a specific landmark along a route he/she would know which turn to make next. Previous research has demonstrated the importance of landmarks in the development of children's wayfinding (see Kitchin & Blades, 2001). For instance, Cohen and Schuepfer (1980) showed children and adults a route displayed on six consecutive slides, and once the participants had ‘navigated’ their way through the slides, they were presented with identical slides with no landmarks and asked to recall the appropriate landmarks that they had previously seen. Younger children (6- to 8-year olds) recalled fewer landmarks than older children (11-year olds) and adults. Other research has demonstrated that children are dependent on landmarks remaining exactly the same as when they were encoded, because changes in the appearance of a landmark disrupts younger children's ability to use them when retracing a route (Cornell et al., 1994; Heth et al., 1997). Young children are more likely to rely on landmarks that are closely associated with a turn. Cornell, Heth, and Broda (1989) asked 6 and 12-year-olds to retrace a route through a University campus and told some children just to pay attention, some children to note landmarks that were near junctions, and other children were advised to remember distant landmarks that were not on the route itself but were visible from different parts of the route. The older children benefitted from advice about noting any landmark, but the performance of the younger ones only improved when their attention was drawn to landmarks very close to junctions. Young children's focus on landmarks near to choice points has been found when children have to name what they think are the best landmarks along a route (Allen, 1981), or when they are learning a route over a number of trials (Golledge, Smith, Pellegrino, Doherty, & Marshall, 1985). Taken together the past studies of children's wayfinding have demonstrated that young children are particularly dependent on landmarks closely associated with junctions/choice points. Despite the evidence above, previous studies of children's wayfinding have all been limited because of the impossibility of distinguishing between wayfinding that is based on purely landmark strategies (‘turn at the tree’) and wayfinding based on directions (‘first left, then right’). The routes used in previous experiments have been presented as slides, as films, or have been actual routes in real environments like towns, campuses or buildings (Kitchin & Blades, 2001), and in all such environments children's performance could well be based on recalling specific landmarks along the route, but it could also be based on recalling the sequence of junctions. Alternatively, children's performance could be a combination of both strategies – for example remembering that the first turn was a right turn, the second was by the shop, then the next was a left turn, the next was by a tree, and so on. Nothing in previous research excludes the possibility that even young children may rely on the sequence of turns (left, right) as well as on the presence of landmarks. One of the aims of the present study was to find out when children could learn a route in an environment with no landmarks, in other words in an environment where they were dependent on encoding just the turns. The route that children were asked to learn was through a maze in a VE. The maze was made of uniform and indistinguishable paths and brick walls so that the walls provided no wayfinding cues. In one condition we included salient landmarks at the junctions of the maze, and in another condition we removed all of the landmarks entirely. In the landmark condition children could learn the route through the maze by attending to the landmarks and/or by noting the sequence of the turns. In the condition without landmarks children could only navigate by remembering the sequence of turns. Creating an environment without landmarks is only possible in a VE, because any real world environment includes numerous landmarks (buildings, signs, marks on the sidewalk) that can be used for wayfinding. Only by using a VE were we able to remove all cues from the environment. VEs are an effective way to study wayfinding because VEs can depict visual and spatial information from a 3D first person perspective (Jansen-Osmann, 2002; Richardson, Montello, & Hegarty, 1999) and successful route learning in VEs can transfer to real environments (Ruddle, Payne, & Jones, 1997). VEs have also been used to assess wayfinding by children (Farran, Courbois, Van Herwegen, & Blades, 2012), and are an efficient tool for improving route learning abilities (Farran, Courbois, Van Herwegen, Cruickshank, & Blades, 2012). Jansen-Osmann and Fuchs (2006) compared how 7- and 11-year-olds learnt their way around a VE with and without landmarks and found that wayfinding performance was poorer without landmarks. However, in Jansen-Osmann and Fuchs children freely explored a whole maze (rather than learnt a specific route), and their knowledge of the maze was then measured by assessing how well they navigated between two places in the maze separated by just two turns. Therefore, Jansen-Osmann and Fuchs's procedure was different from the real world route learning studies described above which always involved participants leaning a specific route with several turns. In a further study, Jansen-Osmann and Wiedenbauer (2004) used a VE to assess children's reliance on landmarks. The children freely explored a VE maze while attempting to reach a goal, and later they had to find the goal again when the landmarks had been removed from the maze. 6- and 8-year-old children found this difficult, but 10-year-olds and older children were successful. In the present study we wanted to extend Jansen-Osmann and Wiedenbauer's (2004) findings. We asked 6-year olds, 8-year olds, 10-year olds and adults to learn a specific route through a VE, rather than asking participants to freely explore the VE. The route had six junctions with a choice of two directions at each junction.","were guided along the correct route once, and were then asked to retrace the route, from the start, on their own. Participants retraced the route until they achieved two consecutive completions without error. We used a between participants design (unlike Jansen-Osmann & Wiedenbauer, 2004) to eliminate the possibility that participants improved their performance due to practice effects. In condition 1, participants were tested along a route that included no landmarks (so that wayfinding was dependent on learning the sequence of turns). In condition 2, participants were tested along a route that included landmarks (which could be used to identify turns). Given the previous research (discussed above) that has indicated children's dependence on landmarks, we expected the children to learn the route better in condition 2 (with landmarks) than in condition 1 (without landmarks), but we expected that adults would be equally proficient at learning both the route with and the route without landmarks. The age when children can learn just a sequence of turns (without landmarks) has never been established before. Hypotheses ~~~~~~~~~~ The ability to use directions such as left and right does not fully develop until the age of 10 years (Blades & Medlicott, 1992; Boone & Prescott, 1968; Ofte & Hugdahl, 2002), and Jansen-Osmann and Fuchs (2006) found that 6- and 8-year-olds had difficulty in an environment without landmarks. Therefore our first prediction was that 6- and 8-year olds would perform poorly in condition 1 without landmarks relative to condition 2. We expected 10-year-olds would learn the route without landmarks, because this age group had been successful in Jansen-Osmann and Fuchs' study, albeit along a route with only 2 turns. We assumed that the adults would be able to learn the route in condition 1 without difficulty. Condition 3 of the present experiment was included to find out if a small amount of training would improve children's route learning. The VE in condition 3 (as in condition 2) had landmarks. When the children in condition 2 were first shown the route by the experimenter none of the landmarks were pointed out, but in condition 3 the landmarks at correct junctions were each named as the child was led by the experimenter. Studies have shown that children benefit from being told to attend to landmarks in real environments (Cornell et al., 1989). Therefore, for our second prediction we expected children would perform better when the landmarks had been named in condition 3 than when they had not been named in condition 2. Participants ~~~~~~~~~~~~ Sixty 6-year olds (M = 6; 3, SD = 0.26), 60 8-year olds (M = 8; 5, SD = 0.31), and 60 10-year olds (M = 10; 4, SD = 0.53), were recruited from a number of primary schools in the UK. Twenty of each age group (10 boys and 10 girls) were randomly allocated to condition 1, condition 2, and condition 3. Forty adult participants (mainly postgraduate students) aged 20–37 years also took part (M = 25 years, 6 months, SD = 3 years, 8 months). Twenty adults (10 male and 10 female) were randomly allocated to condition 1 and condition 2. No adults took part in condition 3, because pilot data suggested that adults' performance was already near ceiling in condition 2 and so no further improvement would have been detected in condition 3. Ethical approval was granted by the University of Sheffield Department of Psychology ethics committee. Virtual environments Five different VEs were created using Vizard, a software program which uses python scripting. VEs were presented to participants on a 17-inch Dell laptop that was placed on a desk. Participants sat in a chair at the desk and were approximately 50 cm from the screen. Participants navigated through the maze using the arrow keys on the keyboard. Practice maze One maze (maze A) was used as a practice maze to familiarise participants with moving in a VE. This maze was a similar but different layout to the test mazes. It did not contain any landmarks. Test mazes Four VE mazes (mazes 1–4) were used to test participants. Each maze was a brick wall maze (see Fig. 1) with six junctions. Each junction was a two-choice junction with a correct path and an incorrect path. The incorrect path ended in a cul-de- sac. From the junction the cul-de-sac looked like a T-junction rather than a dead- end. Therefore, participants could not tell that they had made an error until they had actually committed to walking along a chosen path. Of the six junctions in each maze, there were two right, two left and two straight ahead correct choices that were balanced with the same type and number of incorrect choices. All of the path lengths between junctions were equal. A white duck marked the start of the maze and a grey duck marked the end of the maze. When participants reached the grey duck, the maze disappeared indicating the end of the trial. Maze 1 and maze 2 were used in condition 1. Maze 1 is illustrated in Fig. 2. Maze 2 was exactly the same design as maze 1 except that the start point of maze 2 was the end point of maze 1, and the end point of maze 2 was the start point of maze 1. Maze 1 and maze 2 were therefore equivalent, but each included a different sequence of left, right, and straight ahead correct choices. We included two mazes in condition 1 so that the findings were not specific to a particular route. Half the participants in condition 1 received maze 1 and half received maze 2. Maze 3 and maze 4 were used in condition 2. Maze 3 is illustrated in Fig. 3. Maze 3 was the same as maze 1 but included 12 landmarks. Maze 4 was the same as Maze 2 but included 12 landmarks. In both mazes, the landmarks were all objects that would be familiar to children: ball, bench, bus, bicycle, car, cow, playground slide, street lamp, traffic light, bin, tree and umbrella. There were 6 landmarks placed at path junctions on the correct route and 6 landmarks placed at dead-end junctions on the incorrect route. Maze 3 and maze 4 were also used in condition 3 (the training condition). Like condition 2 half of participants were tested in each maze.","All participants were tested individually. Adults completed the experiment in a quiet office in a University Department. Informed consent was obtained prior to data collection. Children completed the experiment in a quiet room in their school. Informed consent was obtained from all the children's parents, and all the children were asked if they wanted to take part. No children refused to take part. The participant sat at the desk facing a computer and the experimenter sat beside them. The experimenter spent 2 min talking to the participant informally to establish rapport. Participants were asked for their age and birthdate. The experimenter introduced the task by saying, ‘This computer has got some mazes on it that we are going to use. First, we're going to practice using the computer to walk around a maze. I'll go first and show you how, and then you can have a turn.’ The experimenter then demonstrated how to navigate through the practice maze using the arrow keys. Participants were then given time to walk around the maze until they were confident in using the arrow keys to navigate, at which point the experimenter ended the practice phase by saying, ‘Well done, I think you've had enough practice now, do you? Let's have a go at another maze now.’ All participants were given preliminary instructions for the test phase: ‘Now I'm going to show you the way through a new maze. Somewhere in this maze there is a little grey duck to find. I'll show you the way to the grey duck once, and then you can have a go.’ The experimenter then demonstrated the correct route from the start to the end of the maze, giving verbal instructions that differed according to condition. In conditions 1 and 2, the experimenter used generic terms such as ‘You go past here, then you turn this way, and then you turn this way’. In condition 3, the experimenter verbally labelled each landmark. For example, ‘You go past the bench, turn this way at the traffic light, and then you turn this way at the bin’. In all conditions, the experimenter did not use any directional language, such as ‘Turn right’. At the end of the demonstration, the experimenter exclaimed, ‘Hooray, we've found the duck!’, and the screen went blank. The participant was then asked to retrace the route they had been shown from the beginning of the maze and used the arrow keys to walk through the maze. The experimenter sat behind the participant and traced the exact route the participant took on a paper copy of the maze, out of the participant's sight. The experimenter timed how long it took the participant to complete the maze. If, after 5 min, a participant had not reached the end of a maze on a particular trial, the experimenter ended the trial by saying, ‘Oops, it looks like you've got a bit lost. Not to worry, let's start back from the beginning, shall we?’ In practice, this happened infrequently. A note was made that the trial was curtailed, and a new trial commenced. Participants did not receive any help in finding their way after the initial demonstration of the correct route. If a participant asked which way to go, the experimenter said, ‘I want you to show me the way to go. Just try your best.’ If a participant returned to the start position but thought that they had reached the end, they were told, ‘You're back at the beginning of the maze now. Let's turn around and try again to remember the way I showed you to the little grey duck.’ Again, in practice, this happened infrequently. When the participant reached the end of the maze, the experimenter congratulated the participant, and asked them to walk the route again from the start. This procedure was repeated until the participant had walked the route to a criterion of two consecutive completions without error. At the end of the final trial, all participants were thanked, and children, regardless of their performance, received a sticker. If a participant had not walked the route with two consecutive completions after 20 min or after eight attempts, the experiment was stopped and the children were given a sticker.","Successful learning was defined as two consecutive completions of the route without error. To achieve this criterion, participants had to walk the route without walking down any incorrect paths on two consecutive learning trials. Walking down an incorrect path was classed as an error. Looking down an incorrect path was not classed as an error. The total number of learning trials to reach criterion excluded the final two perfect trials. For example, if a participant made an error on trial 1, but then walked the route without error on trials 2 and 3, they would be scored as having required 1 trial to reach criterion. A lower score indicated better performance. If a participant never achieved the criterion, the number of learning trials was calculated as the number of trials that were completed. For example, if a participant completed 6 trials within the 20-min cut-off time, but did not complete 2 consecutive trials without error, they scored 6. Participants received a mark of 1 for every error they made during a trial. On each trial a proportional error score was calculated as the number of errors divided by the number of decisions made. For example, Fig. 4 shows the route taken by one participant. This participant made 5 errors out of a total of 13 decisions, producing a proportional error score of 0.38. This scoring captured participants' wayfinding behaviour every time they made a decision. This scoring method accounted for occasions when participants doubled back and returned to the same junction more than once within a trial. Some participants who got lost did not reach the later junctions, so any junctions not reached were also scored as errors at decision points. A mean proportional error score was calculated for each participant across learning trials. The proportional error score captured all of a participant's behaviour on a trial. We note that alternative coding criteria produced the same patterns of performance. For example, we coded just the decisions made the first time a participant approached a junction in each trial. Participants scored 0 if they chose the correct path or 1 if they chose the incorrect path and any junctions not reached were counted as errors. Therefore 6 indicated the worse performance, and 0 perfect performance. When this scoring was compared to the proportional error score (above), there were no differences in the results. Therefore, in the results section we report only the proportional error scores. Independent sampled t-tests showed that there were no differences between performance on maze 1 and maze 2 or maze 3 and maze 4 for any of the dependent variables (all p-values >0.05). We explored whether there were any differences in performance at path junctions across the six landmarks. The assumption of sphericity could not be met (x2 (14) = 164.20, p < 0.001.) and so a geisser-greenhouse correction is reported instead (ε = 0.72). A one-way repeated measures ANOVA (6 levels: tree, bench, bike, umbrella, traffic light, bin) demonstrated a main effect of landmark type (F (3.957, 492.794) = 5.466, p < 0.001, np2 = 0.03). A Bonferoni corrected post-hoc test revealed that participants made fewer errors at the bench than at the bin (p < 0.01) and at the bike than at the bin (p < 0.001). Performance at all other choice point pairings were equal (all p-values >0.05). Both the bike and bench were straight ahead junctions and therefore at these junctions participants did not have to remember to turn left or right. But at the bin, participants had to remember to turn either left or right. This may explain why there was a difference in performance at certain junctions in the maze. Scoring method ~~~~~~~~~~~~~~ To test our predictions relating to children's and adults' performance across the different conditions, we analysed three dependent variables: – (i) proportion of participants reaching the learning criterion, (ii) number of trials to reach learning criterion (for those who reached criterion) and (iii) proportion of errors. Chi-square analyses were used to explore the proportion of children and adults who reached learning criterion in the three different maze conditions. ANOVA analyses were then conducted to explore number of trials to reach criterion and proportion of errors made by children and adults in the different maze conditions. Proportion of participants reaching the learning criterion ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Chi-square analyses were conducted to explore how many participants reached the learning criterion in the different maze conditions. In condition 1 (no landmarks) 10% of 6-year olds, 40% of 8-year olds, 80% of 10-year olds, and 100% of adults reached the criterion of 2 successive trials without error. In condition 1 there was a relationship between age and reaching the learning criterion (X2 = 39.90, df = 3, p < 0.001). Standardized residuals show that this was accounted for by adults whom were more likely to more likely to reach criterion than 6-year olds. In condition 2 (with landmarks) 90% of 6-year olds, 95% of 8-year olds, 100% of 10-year olds, and 100% of adults reached the criterion. There was no significant relationship between age and reaching the learning criterion in condition 2 (X2 = 3.81, df = 3, p > 0.05). In condition 3100% of 6-year olds, 100% of 8-year olds, 100% of 10-year olds, reached the criterion, so no statistical analyses could be conducted. Adults were not included in condition 3 because, as had been expected they were at 100% in condition 2. Number of learning trials to reach criterion ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In condition 1, for only those participants who reached criterion, 6-year olds required more trials (M = 5.00, SD = 0) than 8-year olds (M = 2.13, SD = 2.36), 10-year olds (M = 2.00, SD = 1.71) and adults (M = 0.30, SD = 0.73). As noted above many children did not reach the learning criterion (notably 6- and 8-year olds), so it was not appropriate to conduct any statistical analyses with such uneven groups. In condition 2 (landmarks) for only those participants who reached criterion, 6-year olds required more trials (M = 2.06, SD = 1.83) than 8-year olds (M = 1.21, SD = 1.90), 10-year olds (M = 0.25, SD = 0.55) and adults (M = 0.30, SD = 0.73). A one-way ANOVA was carried out with age (6 years, 8 years, 10 years, adults) on the number of learning trials to reach criterion. There was an effect of age (F (3, 76) = 7.42, p < 0.001). Tukey post-hoc tests showed that in condition 2 the 6-year olds required more trials to reach criterion than the 10-year olds (p < 0.01) and adults (p < 0.01). All other post-hoc tests were non-significant (p > 0.05). In condition 3 the 6-year olds (M = 0.90, SD = 1.07) required more trials than the 8-year olds (M = 0.15, SD = 0.57) and 10-year olds (M = 0.05, SD = 0.22). Analysis of this condition is considered within the ANOVA below. To consider the data from the child groups across conditions (adults did not complete all condition 3), a 2 (condition 2, condition 3) × 3 (6, 8, 10 year olds) ANOVA was performed on the number of learning trials to reach criterion in each condition. Two 6-year olds and one 8-year old did not reach the learning criterion in condition 2 and therefore were not included. There was an effect of condition (F (1, 111) = 13.73, p < 0.001, np2 = 0.11) because children required fewer trials to reach criterion in condition 3 than condition 2. There was an effect of age (F (2, 111) = 12.56, p < 0.001, np2 = 0.19). As confirmed by Tukey pairwise comparisons, 6-year olds required more trials to learn the route than 8-year olds (p < 0.05) and 10-year olds (p < 0.001). Eight-year olds did not require more trials than 10-year olds (p = 0.13). These significant main effects support our second hypothesis that verbal labelling reduced the number of trials required to reach criterion. There was no interaction between maze condition and age group (F (2, 111) = 1.98, p = 0.14). Proportional errors during the learning phase ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In condition 1 the 6-year olds (M = 0.42, SD = 0.15) made more errors than the 8-year olds (M = 0.23, SD = 0.18), 10-year olds (M = 0.14, SD = 0.13) and adults (M = 0.03, SD = 0.06). In condition 2 the 6-year olds (M = 0.12, SD = 0.14) made more errors than the 8-year olds (M = 0.06, SD = 0.08), 10-year olds (M = 0.01, SD = 0.02) and adults (M = 0.01, SD = 0.03). In condition 3 6-year olds (M = 0.07, SD = 0.08) made more errors than 8-year olds (M = 0.01, SD = 0.03), and 10-year olds (M = 0.01, SD = 0.01). We did not conduct a Maze (condition 1, condition 2, condition 3) × Age (6-, 8-, 10-year old, adults) ANOVA, given that no adults took part in condition 3. Rather, we conducted two separate analyses: A Maze (condition 1, condition 2) × Age (6-, 8-, 10-year olds, adults) ANOVA and to explore the effect of verbal labelling, a Maze (condition 2, condition 3) × Age (6-, 8-, 10-year olds) ANOVA. A 2 (condition 1, condition 2) × 3 (6-, 8-, 10-year olds) ANOVA was performed on the proportion of errors. There was an effect of maze (F (1, 114) = 70.26, p < 0.001, np2 = 0.38) because participants made proportionally more errors in condition 1 than condition 2. There was an effect of age (F (1, 2, 114) = 24.32, p < 0.001, np2 = 0.30). As confirmed by Tukey pairwise comparisons, 6-year olds made a higher proportion of errors than 8-year olds (p < 0.001) and 10-year olds (p < 0.001). The proportional of errors for 8-year olds and 10-year olds did not differ (p = 0.06). There was also an interaction between maze condition and age group (F (2, 114) = 4.43, p < 0.05, np2 = 0.07) as there was a greater age related improvement between 6- and 8-year olds than between 8- and 10-year olds in condition 1, but in contrast there was a greater age related improvement between 8- and 10-year olds than between 6- and 8-year olds in condition 2. This supported our first hypothesis that younger children (6- and 8-year olds) would make more errors when learning the route without landmarks than when learning the route with landmarks. A 2 (condition 2, condition 3) × 3 (6, 8, 10 year olds) ANOVA was performed on the proportion of errors. There was an effect of condition (F (1, 114) = 6.88, p < 0.05, np2 = 0.06) because children made proportionally more errors in condition 2 than condition 3. There was also an effect of age (F (2, 114) = 14.78, p < 0.001, np2 = 0.21). As confirmed by Tukey pairwise comparisons, 6-year olds made more errors than 8-year olds (p < 0.01) and 10-year olds (p < 0.001). The proportional of errors for eight- year olds and 10-year olds did not differ (p = 0.20). These significant main effects supported our second hypothesis that verbal labelling reduced the number of errors made by participants. There was no interaction between maze condition and age group (F (2, 114) = 1.21, p = 0.30).","The aim of the present study was to investigate the development of children's wayfinding strategies. As noted in the introduction there are two predominant strategies that can be used for learning and retracing a route (Kitchin & Blades, 2001). One is a strategy based on landmarks so that a child encodes a turn in relation to a particular landmark, and one is a directional strategy when a child learns a sequence of turns. From previous research we predicted that children would rely heavily on the presence of landmarks and that this would be particularly the case for the younger children (Cohen & Schuepfer, 1980; Jansen- Osmann & Wiedenbauer, 2004). Therefore we expected young children to encode a route much better in a maze with landmarks than they would in a maze without landmarks. In a maze without landmarks children are dependent on learning a sequence of turns and therefore one aim was to establish if any of the age groups of children in the study could use this strategy when wayfinding. Our results showed that only a quarter of the 6- and 8-year olds achieved the learning criterion in condition 1 when there were no landmarks available. In other words, without the presence of landmarks their route learning was poor, and they made a large number of errors. In contrast, four-fifths of the 10-year olds were able to learn the route even when there were no landmarks present. The latter finding indicated that by the age of 10 years children can adopt a directional strategy and consider a route in terms of left, right and straight ahead directions. The 10-year olds were therefore mostly competent wayfinders even when they were unable to rely on landmarks. Though we note that the 10-year olds did not perform as well as the adult group, and therefore children's ability to use just directions for route finding is still developing at the age of 10 years. 6 and 8-year olds found it difficult to find their way in the absence of landmarks. These children may have had a more limited memory capacity, and therefore might have had difficulty encoding six turns. However, this is unlikely because the results from condition 2 showed that nearly all the children could learn six turns when landmarks were present, and other research has shown that learning six turns is easily possible for these age groups (Farran, Courbois, Van Herwegen, & Blades, 2012; Farran, Courbois, Van Herwegen, Cruickshank, et al., 2012). Therefore the younger children's difficulty in the present experiment was probably due to their failure to employ a directional strategy (i.e. using left, right and straight ahead) to find their way in the absence of landmarks. Young children may have difficulty using such a strategy because their use of spatial directions like left and right may not be fully developed until about the age of 10 years (Blades & Medlicott, 1992; Boone & Prescott, 1968; Ofte & Hugdahl, 2002). When landmarks were available in the maze (in condition 2) 90% or more of each group learnt the route successfully. However, even when landmarks were available the 6-year-olds required more trials to reach criterion than the older children and adults. In other words, children can learn a route with landmarks, but how effectively they do so improves with age, and this corresponds to previous findings (Cornell et al., 1989; Siegel & White, 1975). The younger children might have required more trials than older children to learn the maze with landmarks for several reasons. Younger children could have less mature cognitive abilities, and/or their lack of wayfinding experience, because young children, especially 6-year olds rarely have to encode new routes for themselves. We considered whether the children's poorer than optimal performance in condition 2 with the landmarks, was a reflection of difficulty with route learning or because the children who did poorly were less aware of appropriate wayfinding strategies. In condition 3 we pointed out and emphasised the landmarks during the children's first walk through the maze. This in itself led to an improvement in the performance, because all the children reached criterion in condition 3 and all three age groups of children had fewer errors in condition 3 than in condition 2. We can conclude that the wayfinding limitations shown by a few children in condition 2 were most likely due to those children not attending sufficiently to the landmarks during the learning trial. At 6-years of age the executive-frontal functions required for efficient navigation begin to develop (Bullens, Nardini, et al., 2010; Nardini, Jones, Bedford, & Braddick, 2008; Nardini, et al., 2009). Furthermore at 6-years of age children are particularly reliant on landmark cues for successful navigation (Bullens, Iglói et al., 2010). When all the children had the landmarks explicitly pointed out to them in condition 3 they all learnt the route successfully. The effect of verbally emphasising landmarks during learning supports previous research (Cornell et al., 1989; Farran, Blades, Boucher, & Tranter, 2010), possibly by triggering verbal recoding and suggests that this simple training technique can be very effective. We emphasise that in condition 3 we did not train the children by indicating the relationship between a landmark and the corresponding turn (which would have been an explicit wayfinding strategy), all we did was get the children to pay attention to the landmarks the first time they walked the route, and this was sufficient for children's route learning to improve. Therefore the limitations in young children's performance in condition 2 were probably due to lack of experience (in this case not realising the need to attend to all of the landmarks). Therefore, it seems likely that if adults make an effort to emphasise landmarks along new routes then the children's wayfinding will be improved, without the need for any more specific training. Emphasising appropriate landmarks in real world contexts might be particularly important because the real environment can include a great many potential landmarks. Future research could manipulate the number and complexity of landmarks in a VE maze to find out the effects of multiple landmarks at decision points along a route. Our cross sectional design only allowed us to examine a snapshot of how children find their way at a certain age. Future research could also explore how children's wayfinding abilities develop longitudinally over time using controlled VEs which (unlike real environments) do not change over time. Like previous research the present study has demonstrated that VEs can be an effective method for investigating how children learn a route (Farran, Courbois, Van Herwegen, & Blades, 2012; Jansen-Osmann, 2002; Jansen-Osmann & Fuchs, 2006), because a VE allows complete control of the environment and is a safe way to test young children's wayfinding. More importantly a VE allows the creation of routes (like the ones without landmarks in condition 1) that would be impossible in a real environment and this has led to novel findings. However, unlike the real world participants were not immersed in the VE and nor did participants experience the same whole body kinaesthetic information. Therefore we recognise that there are limits to using desktop VEs to study wayfinding. This is the first study to distinguish between route learning dependent on landmarks and route learning dependent on learning directions. Without the presence of landmarks, 6- and 8-year olds were much poorer at learning the route than when landmarks were present. This suggests that landmarks are crucial for children's route learning and that route learning based on directions does not develop until the age of 10. Our findings also showed that when landmarks were labelled this improved wayfinding performance for all children."],["The learnability problem of social life suggests that innate mental representations and motives to navigate adaptive relationships have evolved. Like other species, preverbal human infants form dominance hierarchies where some systematically supplant others in zero-sum conflict, and use the formidability cues of body and coalition size, as well as previous win-lose history, to predict who will prevail. Like other primates, human toddlers also seek to affiliate with allies of high rank, but unlike bonobos they pay unique attention whether others voluntarily defer to their precedence, reflecting the importance of consensual authority in cooperative human society. However, young children appear not to readily infer authority from benevolence, and expectations for inequality correlate with unwillingness to share resources even among infants. -------------------------------------------------------------------------------- PREVERBAL INFANTS INFER THE OUTCOME OF FUTURE ZERO-SUM CONFLICT FROM CUES OF FORMIDABILITY -------------------------------------------------------------------------------- Early ethological, naturalistic observations in daycare groups demonstrated that preschoolers [26–28], and even infants from 8 [29] and 11 [30] months of age, form transitive dominance hierarchies like those of other species such that more formidable and agonistic individuals systematically supplant others in contests over scarce resources (e.g. toys). Consistent with the adaptive importance of accurately perceiving and coordinating along formidability in such dominance hierarchies, 9–13 month-old infants also mentally represent social dominance, and use relative body-size – a cue which not only marks formidability, but also status and authority across cultural practices and conceptual metaphors [2,31–33] – to predict who will prevail in novel zero-sum conflicts before any physical coercion ensues. Although they unlikely have extensive personal experience engaging in right-of-way conflicts, infants looked longest, indicating that their expectations were violated, when a large cartoon figure prostrated and yielded the way to a smaller one in a ‘game-of-chicken’, where the two figures blocked each other from crossing a stage in opposite directions. These effects hinged critically upon the existence of zero-sum conflict so that one agent completed its goal at the expense of the other: When the agents moved alone on the stage or in the same direction behind one another, so that both could have achieved their goals, infants made no prediction whether the large or small figure would manage to cross the stage [33]. Nor did they do so if the animations were stopped immediately after the prostration event, suggesting that in the absence of zero-sum conflict, infants do no associate relative size/formidability with this conventional display of submission and deference across culture and species [34]. Infants also use the formidability cue of previous win-lose history to predict who will dominate whom across conflict domains and dyadic relationships. When a novel agent uses coercive force to push another one out of a contested territory, they expect that the former loser will later yield a contested resource to the winner without physical fight [35] (see also Refs. [36]). Importantly, and as previously demonstrated among even fish [37], recent work now confirms that human infants, too, make transitive inferences and predict relative dominance based on the previous win-lose history that individuals had with (common) third-parties [38•]. Finally, like other species including dolphins, spotted hyenas, lions and several non-human primates [7–10,39,40–45], humans engage in coalitional conflict, which the archeological record indicates also characterized the context in which we evolved [13,19,46,47]. Coalitional conflicts manifest from preschool and primary school playgrounds [48,49] to large-scale politics [13,14,50], resemble dove–hawk dominance dynamics (cf. 14,21]), and have profound psychological and societal implications [4,13,50,51]. And indeed, even 6–9 month-old infants were found to understand coalitional formidability, using the relative number of allies to predict which individuals will yield and prevail [52•] in the right-of-way dominance paradigm [33]. However, representing the combined formidability of a set of individuals by their number is unlikely a psychologically simpler, more salient or reliable indicator of dominance than is individual strength, as indicated by the fact that infants from eight months of age form interpersonal dominance hierarchies [29,30], long before they presumably engage in any joint inter-coalitional conflict. Surely, coalitional formidability relies in part upon the strength of individual members. Consistent with this, several non-human primates selectively recruit and join the more formidable individuals in coalitional conflict [43]. When predicting the outcome of coalitional conflict, three-year-olds, as well as adults, also flexibly weigh the parties’ body sizes against the number of allies, emphasizing the former for smaller, and latter for larger, coalitions [69]. In fact, by primary school, children integrate a full range of asymmetric factors to weigh the total expected costs and benefits of engaging in resource conflict, and make no systematic prediction whether one large or two smaller, allied individuals will win a conflict [48]. Indeed, group- living does not make interpersonal dominance obsolete, but potentiates the long-term fitness implications of rank-based, recurrent outcomes within the group, as they are remembered by individuals who know one another [43,53–55]. Accordingly, a suite of dedicated cognitive architecture translates individual, physical strength into claims for both personal resources, coalitional dominance, and allocated social status [56–60], and also supports inductions about the physical strength and competitive motives of others, even from vocal pitch [61–63]. Further supporting the importance of individual strength, even three-month-olds induce physical size from voice pitch [64•] and pre-schoolers adeptly induce and relate both ‘who is strongest’ and ‘who is in charge’ from non-verbal cues of body posture, muscularity and facial configuration [65–68].","Any adaptive gains from representing fundamental forms of social relations stem not simply from accurately perceiving them, but from using this information to navigate the social world in appropriate, strategic ways [1]. And indeed information about status rank motivates the way infants, toddlers and preschoolers themselves relate to, and coordinate with, others. Whereas dominant individuals are avoided in many species, so as to not provoke aggression, in highly social species they may offer access to resources and influence, so that approaching, ingratiating and affiliating oneself with other individuals of high rank may be adaptive [43,71••]. This relationship between rank and resources is well understood: Infants expect third-party distributions of resources to benefit a physically dominant over subordinate agent [72]; three-year-olds and four-year- olds will themselves favor a dominant individual in third-party distributions of resources [73], and they explicitly expect that those giving or withholding access to use resources will be in charge [74•]. Young children also intuitively enact these dynamics in controlled settings: In a game of chicken — where a child yielding the way to the other lost some of its resources, but both children lost all their resources if nobody yielded — dyads of five-year-olds coordinated such that previously dominant children received higher pay-offs because their partner yielded and deferred to them more often [75•]. Similarly, when 4–5 year-olds dominate a dyad, monopolizing a coveted toy from another child, they later donate less stickers to other, anonymous children than do subordinates, and experimentally induced changes of dyadic rank have similar effects as baseline dominance [76]. Consistent with these expectations that rank begets resources, young toddlers selectively approach those of high rank and so reach for a puppet who prevails in the right-of-way dominance paradigm [33] over one who yields [71••]. However, when one puppet was able to pave the way for both to cross the stage, toddlers reached for the more competent puppet [71••]. This was also the case in another study where two puppets both achieved their goals in the absence of zero-sum conflict, but one did so competently on the first try and the other only after repeated failed attempts [77]. Preschoolers also think that those who receive higher pay-offs are smarter [78], and recent evidence suggests that toddlers may also endorse the preferences and opinions of resource-rich over resource-poor puppet dolls [79], consistent with an ingratiation-for-potential-resources perspective on these findings. If, on the other hand, one puppet prevails in right-of-way conflict by retorting to brute force, knocking the other one over, then toddlers avoid this coercively dominant puppet [71••]. By comparison, infants selectively avoid even a prevailing puppet for whom another voluntarily yields [80••] in the right-of-way paradigm [33]. These findings stand in stark contrast to results that bonobos prefer a dominant individual who monopolizes a territory using brute force [81••] (cf. [35]). Furthermore, a recent study found that toddlers expect that cartoon figures will continue to obey the orders of someone they prostrated before and gave their resources, even when this individual is no longer around, but hold no such expectations for an individual who just beat them up [82••]. Thus, although human toddlers pay close attention to whether someone was able to achieve their goal at the expense of others, compared to our primate relatives including even the generally egalitarian and unaggressive bonobos [83], human toddlers also appear uniquely attuned to whether someone prevails in zero-sum conflict through consensually recognized precedence, rather than through brute coercion, reflecting its importance in cooperative, human societies [2,15–20,23,47], (see also Refs. [84–86]). Similarly, in the political psychology of adults, inequalities between groups are not legitimized by appeal to differences in brute formidability, but by appeal to other, commonly and consensually recognized principles of precedence and merit [4,13,14,50] (e.g. the protestant work ethic). Robust cross-cultural findings, and evolutionary reasoning, supported by evidence from some non-human primates [43] including chimpanzees [9,40], indicate that such coalitional dominance motives are greater among males [13,46,50,87]. Mirroring this, a recent study also found that even three-year-old Norwegian boys predominantly chose a member of a more formidable and larger, rather than smaller, group when they were asked which of two cartoon figures they liked best and wanted to play with and befriend. In contrast, three-year-old girls chose randomly between the two [88].","Leaders in both human and non-human mammalian societies facilitate within-group coordination and conflict resolution [20,89]. Consistent with this, just-linguistic infants expect that a larger puppet, or a puppet whose directions were previously followed, will intervene to rectify resource monopolization between subordinates [90] (cf. also Ref. [82••]). But whereas young children, and even infants, understand that those in power will control resources, permitting or sanctioning their use [74•,90], and contrary to expectations that legitimate authorities take care of their followers and that status may stem from the benefits one can confer upon others, children do not appear to readily infer consensual authority from benevolent help: In fact, although even infants prefer helpers over hinderers [91,92], when North American 4–7 year-olds, as well as adults, view someone refuse a request for help, they take this as evidence that person is in charge [93••]. With the exception of permitting access to scarce resources, it was also easier for young North American children, as well as for adults, to induce such authority from malevolent (imposing costs) rather than benevolent (conferring benefits) ways of using of power [74•]. It is important to find out if this is also the case using more implicit measures and across subsistence systems, culture, languages and less enculturated ages. For instance, future studies might investigate if toddlers also expect novel agents to continue to obey someone towards whom they displayed submissive deference after receiving a resource from them (rather than after offering them a resource, as previously demonstrated [82••]). In terms of learning, North American preschoolers endorsed and imitated the domain-specific preferences and labels of a model whom bystanders had selectively attended over someone whom they ignored [94], supporting the importance of prestige-cues that someone is an expert for culturally transmitted learning (but see also Refs. [78,95]). But other studies reported that French and Mayan preschoolers also selectively endorsed the testimony (which way did the animal go) and word-labels (what is the name of a novel object) of someone who dominated resources with brute force [96,97], although, from an epistemic point of view, such coercive dominance need not imply any superior knowledge. In contrast, however, five recent studies did not find evidence that 249 blind-tested, Norwegian preschoolers selectively endorsed the labels and testimony provided by the prevailing agent in the otherwise robustly dominance-inducing right-of-way paradigm [33,98] (see also Ref. [99] for similar results among a smaller sample of Japanese preschoolers using resource monopolization as dominance cue). Finally, although North American and Chinese 5–7 year-olds distinguish prestige and dominance, neither group was recently found to explicitly predict the outcome of resource-conflict based on dominance rank [100], raising further questions as to how their explicit reasoning link the two.","Finally, although infants, toddlers, and preschoolers understand and strategically respond to hierarchical rank, in fact the default expectations and preferences of most of them appear egalitarian, as is the case for adults cross-nationally [4,13,14,101]. Thus, in the absence of other social information such as relative dominance [72] or effort/merit/deservingness [102,103], most infants expect equal resource distributions between third-parties, prefer people who distribute resources equally, expect others will also prefer them, and expect equal distributors will be more helpful [72,92,102–104,105••,106–110]. Reflecting the appeal of personal resource holding, however, young children themselves react earlier and cross-culturally more reliably when they are personally suffering from disadvantageous, rather than benefitting from advantageous, inequity [109,110]. Nevertheless, the most widely used coordination strategy among five-year-olds when distributing resources in a game of chicken was, in fact, that of turn-taking [75•]. As is the case among adults [4,13,14,50,87,101], individual differences among infants in whether they expect equal versus unequal resource- distribution between others also predict whether infants are themselves willing to share their valued resources [105••,109,110,111]. This is consistent with emergent evidence that dispositions towards hierarchical versus egalitarian expectations and redistribution motives are in part genetically underpinned and related [113], presumably reflecting the balancing selection equilibria of both hierarchical and egalitarian strategies [4,101,113].","Although human societies display unprecedented levels of cooperation, help and egalitarianism compared to our nearest primates relatives, hierarchies within and between groups are ubiquitous across culture. Humans display rich and sophisticated representations to perceive, navigate, and coordinate these dilemmas of resource distribution and conflicting interests that cooperative group living poses: Most of even the youngest humans prefer equal distributors and helpers, but they also prefer not only the most competent, but also those that prevail in zero-sum conflict at the expense of others who defer to them. Expectations that resources will be unequally distributed among third parties relate to unwillingness to share personal resources, and help and benevolence appear not to readily license inferences of authority. Consistent with our evolutionary history, these relational representations and motives reflect, and must optimize, ecological trade-offs between different forms of hierarchical and egalitarian strategies for resource distribution, and are active from earliest childhood and even infancy. It is critical that future work investigate whether, when and how they reliably manifest across culture and subsistence systems to address which relational logic forms part of our innate endowment, undergirding the regulation of social life."],["The effect of mother-infant skin-to-skin contact on Ghanaian infants’ developing social expectations for maternal behavior was investigated. Infants with high and low mother-infant skin-to-skin contact experience in the infants’ first month engaged with their mothers in a Still Face Task at 6 weeks of age. Infants with high skin-to-skin contact experience, but not those with low skin-to-skin contact experience, demonstrated the still face effect with their smiles. Infants with both high and low skin-to-skin contact experience demonstrated the still face effect with their visual attention. The behaviors of the Ghanaian infants and their mothers during the task were compared to archival evidence of Canadian mother-infant dyads’ behaviors in skin-to-skin and control groups who engaged in the Still Face Task at the infant ages of 1 and 2 months. Similarities and differences between the behaviors of the mother-infant dyads in the two cultures were assessed. --------------------------------------------------------------------------------","Developmental psychologists increasingly are exploring how infants’ biological maturation intersects with parenting cultural goals and styles of interaction. Parents from different cultural contexts value different parenting goals that affect parenting interactive practices, which in turn shape infants’ behavior (Kärtner et al., 2008; Keller, 2007; Keller et al., 2004). Distal parenting practices, focusing on face-to-face contexts and object play, are prevalent in Western industrialized societies where parents tend to socialize their infants toward goals of independence and autonomy. Proximal parenting practices, focusing on physical contact and body stimulation, are prevalent in many non- Western societies where parents socialize their infants toward goals of interdependence and relatedness with family and community. Although this duality is likely an oversimplification of parenting goals and practices as complex cultural cross-overs due to globalization as well as individual differences exist in any culture, it is clear that the way of engaging infants prevalent in Western societies is not universal. Yet infants’ biological maturation is universal. Infants are biologically predisposed to engage with others. Their biological predispositions interact with parenting practices early in life and adapt to cultural demands. Mother-infant skin-to-skin contact (SSC) is a method of caring for young infants. The infant is placed between the mother’s breasts, dressed only in a diaper, so that frontal body contact of mother and infant is skin-to-skin. SSC benefits newborns’ neuro-physical adjustment. By engaging in SSC with her infant, the mother provides warmth and stimulation that is believed to simulate the prenatal environment (Whitelaw & Sleath, 1985). Across a number of cultures, SSC for infants early in life has been shown to stabilize infant temperatures, heart rates, respiratory rates, gastrointestinal adaptations (Bauer, Sontheimer, Fischer, & Linderkamp, 1996; Bergman, Linley, & Fawcus, 2004; Bystrova et al., 2003; Charpak, Ruiz-Pelaez, Figueroa de Calume, & Charpak, 1997; Cleary, Spinner, Gibson, & Greenspan, 1997; Conde-Agudelo, Diaz-Rossello, & Belizen, 2000; Feldman, Eidelman, Sirota, & Weller, 2002; Fohe, Kropf, & Avenarius, 2000; Ludington-Hoe, Anderson, Swinth, Thompson, & Hadeed, 2004; Moore, Anderson, & Bergman, 2007; Mori, Khanna, Pledge, & Nakayama, 2010), and reduce pain from routine medical procedures (Gray, Miller, Philipp, & Blass, 2002; Gray, Watt, & Blass, 2000). Studies investigating the effects of SSC beyond the newborn period have shown that the effects extend to infants’ social cognitive development. Feldman and colleagues (Feldman, Eidelman, et al., 2002; Feldman, Weller, Sirota, & Eidelman, 2002) found that, compared to infants who did not receive SSC, infants at 3 months who received SSC as newborns had more interest in novel stimuli, and at 6 months were more advanced in toy exploration and shared attention with mother. Ohgi et al. (2002) found that infants at 6 months who had previous SSC showed better orientation to social and non-social objects. At 12 months, infants with SSC experience performed better on infant development scales than infants without such experience (Feldman, Eidelman, et al., 2002; Ohgi et al., 2002; Tessier et al., 1998). Multiple reasons have been proposed for why SSC early in life would facilitate infants’ subsequent social cognitive development (Feldman, Eidelman, et al., 2002; Feldman, Weller, et al., 2002). The post-birth period constitutes a sensitive period for maternal contact in animal and human models (Field, 1995; Lehmann, Stohr, & Feldon, 2000; Scafidi et al., 1990; Wigger & Neumann, 1999), and tactile and proprioceptive stimulation are especially important during this time (Gottlieb, 1976, 1991). SSC improves infants’ state organization, particularly infants’ sleep/wake cycles, as well as improves stress reactivity and physical maturation (Feldman, Weller, et al., 2002; Michelsson, Christensson, Rothganger, & Winberg, 1996). Such early physiological and behavioral regulation predicts later cognitive development (Beckwith & Parmelee, 1986; Doussard- Roosevelt, Porges, Scanlon, Alemi, & Scanlon, 1997; Feldman, Greenbaum, Yirmiya, & Mayes, 1996; Sigman, Cohen, & Beckwith, 1997). The close proximity between mother and infant in SSC also facilitates infants’ ability to recognize and respond to their mothers and mothers’ ability to recognize and become familiar with their infants’ signals (Feldman, 2004). Mothers who are sensitive to their young infants’ signals engage in more frequent and positive mother-infant interactions. Mothers who provide SSC for their infants report more positive maternal feelings, positive perceptions of their infants, less depression, and more empowerment in their parenting role (Affonso, Bosque, Wahlberg, & Brady, 1993; Johnson, 2007; Neu, 1999; Roller, 2005; Tessier et al., 1998; Whitelaw, 1990). Although most studies of the effects of SSC on maternal feelings and behavior have been conducted only in the newborn period, some have shown increased maternal sensitivity and maternal behaviors of holding, touch, and infant-directed speech throughout the infants’ first year (Bigelow, Littlejohn, Bergman, & McDonald, 2010; DeChateau & Wiberg, 1977, 1984; Feldman, Eidelman, et al., 2002). SSC promotes infants’ social cognitive development by affecting infants’ behavioral and physiological regulation and mothers’ maternal behaviors (Feldman, 2004). Thus in mother-infant interactions, infants with SSC experience may have enhanced understanding of the relation between their own actions and their mothers’ social responsiveness to them. The Still Face Task has been used to examine infants’ awareness of others’ social behavior toward them. This task, first reported in a seminal study by Tronick and colleagues (Tronick, Als, Adamson, Wise, & Brazelton, 1978), has mothers or other social partners engage infants in normal face-to-face interaction, then become suddenly still and expressionless, and then resume normal interaction. Infants tend to exhibit reduced visual attention and decreased positive affect, as demonstrated by changes in smiling and non-distress vocalizations, during the still face phase compared to the interactive phases. Such changes in the infants’ behavior, known as the still face effect, have been shown in numerous studies, typically with infants between 2 and 9 months of age (but see Bigelow & Power, 2012, below for effects in younger infants) and are found regardless of procedural variations, such as length of the episodes (Adamson & Frick, 2003; Mesman, van IJzendoorn, & Bakermans-Kranenburg, 2009). Bigelow and Power (2012) conducted a longitudinal quasi-experiment with Canadian infants in SSC and control groups in which infants engaged with their mothers in the Still Face Task at ages 1 week, 1 month, 2 months, and 3 months. At 1 week, infants in both groups demonstrated the still face effect with their visual attention that continued to be demonstrated at each of the following ages, thus showing that even in the newborn period infants can detect changes in their mothers’ interactive behavior. At 1 month, differences between the groups appeared. Infants in the SSC group began responding to changes in their mothers’ behavior with their affect, suggesting they were responding to violations of their expectations for affect sharing. Infants in the control group did not show affect changes to the phases of the Still Face Task until 2 months. At 3 months, infants in the SSC group increased their non- distress vocalizations in the still face phase, indicative of social bidding to their non- responsive mothers, whereas infants in the control group decreased their non-distress vocalizations in the still face phase. SSC accelerated infants’ social expectations for their mothers’ behavior and enhanced infants’ awareness of themselves as active agents in eliciting social interactions with their mothers. A focus of the present study was to investigate whether SSC would increase infants’ early expectations for the social behaviors of others in a culture with proximal parenting practices. Although SSC studies have been conducted in various cultural contexts, the focus in these contexts has primarily been on the effects of SSC on infants’ neuro-physical adjustments shortly after birth (see Mori et al., 2010, for meta-analysis including SSC studies from 18 different countries). The findings from these studies show similar benefits of SSC to newborns’ physiological adaptations to postnatal life. However, the effect of SSC on infants’ social expectations of others in cultures with different parenting goals and practices has been less well researched. Ghana has a culture with proximal parenting practices. As in many African countries, Ghanaian mothers carry their infants by wrapping them onto their backs as they go about their daily chores or business. Infants sleep in the same bed with their mothers and spend most of their day with their mothers, primarily wrapped on their mothers’ backs. Frontal SSC is not the norm, yet mother-infant tactile contact is high compared to mother-infant dyads in Western societies (Kärtner et al., 2008; Keller et al., 2004). Ghanaian parenting goals and practices resemble those found in other West African societies, of which the Nso of the Cameroon have been the most thoroughly studied. Kärtner and colleagues (Kärtner et al., 2008; Kärtner, Keller, & Yovsi, 2010) compared Nso and German mother-infant interactions. They found that Nso mothers were more directive in their maternal interactions and had more physical contact with their infants. Nso infants experienced less face-to-face interactions and reduced smiling and verbal turn-taking with their mothers. Yet Nso infants experienced enhanced physical closeness with their mothers due to being carried on their mother’s body and sleeping with her. Such practices by Nso mothers facilitate proximal parenting goals of interdependence, such as obedience, deference, and collective responsibility. Maternal responsiveness to their infants was similar across the two cultures, albeit manifested differently. German mothers were more visually responsive with gaze, smiles, and facial expressions; Nso mothers were more tactually responsive with touch and physical stimulation (Kärtner et al., 2008, 2010). These cultural differences emerged between the infants’ second and third month. Mothers in both cultures responded readily to infants’ non-distress vocal signals. For German mothers, maternal verbal responses are the norm; for Nso mothers, physical contact responses are prevalent. Mother-infant verbal turn-taking may be less necessary when physical closeness, such as being carried on the mother’s body, allows for ready physical responsiveness (Mesman et al., 2018). Mother-infant interaction is a bidirectional process in which each partner affects the other. Although this process begins at birth, it increases around 2 months. This 2 month transition is thought to be due to infants’ neural maturational processes and their experience with social partners who reinforce and contingently respond to the infants’ social behaviors (Wörmann, Holodynski, Kärtner, & Keller, 2012). Around 2 months of age, infants’ maturational processes allow them to become more alert, to maintain posture and attention for longer periods, and to systematically explore the internal features of the face (Haith, Bergman, & Moore, 1977; Hopkins et al., 1990; Wolff, 1987). Infants become more aware and interested in social partners at this time (Rochat, 2001). Infants’ visual attention, smiling, and non-distress vocalizations increase (Spitz, 1965; Trevarthen, 1979; Wolff, 1987). They become more responsive (Henning, Striano, & Lieven, 2005) and show effortful patterns of communication (Lavelli & Fogel, 2002, 2005). Kärtner et al. (2010) found that Nso infants’ alertness increased between 6 and 8 weeks, showing evidence of neural maturation associated with the 2 month transition. The culturally specific differences between German and Nso mothers’ responses to their infants occurred around this time. These cultural differences were due primarily to changes in the way German mothers responded to their infants. Between infants’ second and third months, German mothers reduced their tactile responses and increased their gaze and facial responses to their infants, whereas Nso mothers continued to respond to their infants with high levels of tactile responsiveness and did not increase their visual responses. Nso mother-infant dyads engaged in mutual gaze less often than German mother-infant dyads; nevertheless, when in face-to-face interaction, Nso mothers smiled as much as German mothers and Nso infants at 6 weeks gazed and smiled at their mothers as much as German infants at this age (Kärtner et al., 2010; Wörmann et al., 2012). The present study ~~~~~~~~~~~~~~~~~ Infants’ response to the Still Face Task appears to be robust in that it is found in infants of varying ages and in normative and at-risk samples, yet very few studies have conducted the task with infants in non-Western cultures (Mesman et al., 2009). To date only three such studies have been published: Kisilevsky et al. (1998) tested 3 to 6 month old infants in China, Hsu and Jeng (2008) tested 2 month old infants in Taiwan, and Yato et al. (2008) tested 4 and 9 month old infants in Japan. Infants in these studies responded to the task with the still face effect, similar to infants in Western cultures. Yet participants in these studies were from urban, predominantly highly educated, middle- to high-income families, where mother-infant interactive practices may be affected by those of Western cultures. Education, with accompanying changes in income and living standards, influences parenting goals toward independence and autonomy (Kağitçibaşi, 1996; LeVine, Miller, Richman, & LeVine, 1996) as well as increases exposure to Western parenting practices through international acquaintances and travel. To date, no study has used the Still Face Task in an African culture where proximal parenting practices are prevalent. Nevertheless, infants as they approach the 2 month transition may be responsive to the changes in their mothers’ behavior during the task and SSC may accelerate infants’ affective responses to the task, as it did in the Bigelow and Power (2012) study. The present study examined the effect of SSC experience on the behavior of Ghanaian mother- infant dyads in the Still Face Task at the infant age of 6 weeks, when the infants were on the cusp of the 2 month transition. Although at prenatal checkups women were informed about SSC, the amount the Ghanaian mothers provided for their infants varied. The dyads were subsequently divided based on a median split (Bigelow & Power, 2014) into those with high and low SSC experience. Their behavior was compared to archival data from mother- infant dyads in the Bigelow and Power (2012) study at the infant ages of 1 and 2 months. Hypotheses were four: (1) Ghanaian infants with both high and low SSC experience would show the still face effect with their visual attention, like the Canadian infants at both 1 and 2 months; (2) Ghanaian infants with high SSC experience, but not those with low SSC experience, would show the still face effect with their affect, similar to the Canadian infants at 1 month; (3) Ghanaian infants would show overall behaviors of visual attention, smiling, non-distress vocalizations, frowning, and distress vocalizations that are intermediate between those behaviors found in the Canadian infants at 1 and 2 months; and (4) Ghanaian mothers would show more tactile and less verbal behaviors but similar visual attention and smiling behaviors during the interactive phases of the Still Face Task compared to the Canadian mothers. Ghanaian sample Mothers were recruited prior to the birth of their infants during prenatal visits at the Komfo Anokye Teaching Hospital, a large public hospital in Kumasi, the capital city of the Ashanti Region and the second largest city in Ghana. The Antenatal Clinic at Komfo Anokye Teaching Hospital serves pregnant women residing in the urban area and surrounding villages. The Ashanti Region adheres to a matrilineal tradition in which people have an attachment to villages of their mothers’ family. Pregnant women typically spend time in these villages prior to, and after, the birth of their infants. Thus traditional means of child care are prevalent for both urban and village women. Although English is the official language of Ghana, in the Ashanti Region Twi is the common language. All the communications with the mothers were conducted in Twi by native speakers. The participants were 26 infants (14 males) and their mothers. The mean age of the mothers at the infants’ birth was 28.1 years (SD = 4.0 years). The percentage of mothers with a university degree was 15%, 15% had some post-secondary training, 31% had only a high school diploma, and 39% were without a high school diploma. The mothers were predominantly from the ethnic group Akan (92%) with a minority from Ewe (8%). For 40% of the mothers, this was their first child; 28% of the mothers had one previous child, and 32% of the mothers had two or more previous children. All the mothers had telephones, necessary for SSC reporting. The infants’ mean gestation age was 38.3 weeks (SD = 1.9 weeks). Their mean birth weight was 3005.3 g (SD = 489.3 g). The videotaped mother-infant Still Face Task was conducted at the infants’ 6 week checkup (M age = 46 days, SD = 5 days). Canadian sample The participants were 80 infants (38 males) and their mothers. Mothers were recruited prior to the birth of their infants through perinatal clinics at two hospitals with similar demographics in northeastern Canada. One hospital recruited for the SSC group and one hospital recruited for the control group. Approximately halfway through the study, the recruitment sites were switched; the former SSC site became the control site and vice versa. Dyads of the recruited mothers were included if the infants were over 37 weeks gestation age and had no medical problems. Twenty-eight dyads were in the SSC group and 52 dyads were in the control group. Socioeconomic status (SES) of the infants’ families was measured by a Canadian index (Blishen, Carroll, & Moore, 1987) based on education and income. In the index, occupations are divided into 514 groups, ranging in SES scores of 17.81–101.75 (M = 42.74, SD = 13.28). The scores of the higher status parent in the infants’ families yielded a SES mean score of 50.41 (SD = 11.80). The mean age of the mothers at the infants’ birth was 29.7 years (SD = 5.0 years). The percentage of mothers with a university degree was 42%, 41% had some university education, 16% had only a high school diploma, and 1% were without a high school diploma. The racial-ethnic composition of the mothers was 99% non-Hispanic White and 1% Asian. For 47% of the mothers, this was their first child, 29% of the mothers had one previous child, and 24% of the mothers had two or more previous children. The infants’ mean birth weight was 3647.11 g (SD = 530.95 g). The videotaped mother-infant Still Face Tasks used for comparison with the Ghanaian mother-infant dyads were conducted when the infants were 1 month (M = 32 days, SD = 5 days) and 2 months (M = 64 days, SD = 8 days). In Ghana The study received ethical clearance from the first author’s university research ethics board. Women coming for prenatal checkups at Komfo Anokye Teaching Hospital were told about SSC, encouraged to provide SSC for their infants, and asked if they would be willing to participate in the study, which involved keeping daily records of the amount of SSC they provided for their infants through the infants’ first month and engaging in a videotaped Still Face Task with their infants at the 6 week checkup. Women who agreed to participate filled out demographic forms. The first author was notified when the participating mothers gave birth and was given access to the infants’ sex, birth weight, and gestation age. A research assistant telephoned participating mothers weekly through the infants’ first month to gather the daily records of amount of SSC the mothers provided during the previous week. At the infants’ 6 week checkup, the mother and infant engaged in the Still Face Task in a private room in the hospital. In the Still Face Task, the mother and infant sat facing each other approximately 50 to 60 cm apart. The infant sat in an infant car seat that was situated on a table. The mother sat in a chair that allowed her to be at eye level with the infant. Behind and to the side of the infant was an upright mirror (60 cm × 40 cm) in a frame that could be angled to reflect the mother seated opposite the infant. The angle was typically 70–75°. The research assistant videotaping the Still Face Task was behind and to the side of the mother, out of the direct view of the mother and infant. The videotape recorded the infant (full frontal body) and the mother’s reflection (frontal body from the waist up). The Still Face Task consisted of three phases that sequentially followed each other without pause: initial interactive phase, still face phase, reunion phase. For the initial interactive phase, the mother was asked to interact with her infant as she wished for two minutes. For the still face phase, she was asked to become still with a neutral expression, looking at her infant but not talking or touching the infant for one minute. For the reunion phase, the mother was asked to interact with her infant again as she wished for two minutes. The research assistant gave the mother a verbal cue at the beginning of each phase and at the end of the task. The behavior of the mother and infant was scored in the lab of the second author on the Observer Video-Pro 5.0 (Noldus Information Technology, 2003) computer software program by a coder blind to the amount of SSC mothers provided. Infants were scored for duration of visual attention, and positive and negative affect in facial expressions (smiles, frowns) and in vocalizations (non-distress vocalizations, distress vocalizations) during each of the three phases of the Still Face Task. Visual attention was scored as the presence or absence of looking at the mother’s face. Smiles were scored as upward lip movements with or without vocalizations. Frowns were downward lip movements with or without vocalizations. Non-distress vocalizations excluded distress vocalizations (fussing, crying) and digestive sounds (e.g., burps, hiccups). Distress vocalizations excluded non-distress vocalizations and digestive sounds. The duration times were converted to percentage of time within each phase. Mothers were scored for duration of visual attention to the infant’s face, smiles, vocalizations, and physical contact with their infants during the two interactive phases of the task. Mothers’ vocalizations, which were all non-distress vocalizations, were further coded as arousing (energetic, highly stimulating) or neutral (excluding arousing vocalizations). Mothers’ physical contact with their infants was coded as arousing (e.g., pumping the infant’s legs), attention getting (e.g., tapping the infant’s face when not looking at mother), soothing (e.g., gentle stroking), adjusting (e.g., repositioning infant), or passive holding (e.g., hands passively on infant’s body). The duration times were converted to percentage of time within the interactive phases. The coder was blind to the amount of SSC the mothers provided. Coding of the infant and mother behaviors followed that of previous studies (e.g., Bigelow, 1998; Bigelow & Power, 2012; Bigelow & Walden, 2009), with the addition of mothers’ type of vocalizations (arousing, neutral) and physical contact (arousing, attention getting, soothing, adjusting, passive holding). For reliability purposes, a second coder, who was also blind to the amount of SSC mothers provided, independently scored the behaviors of 19% of the dyads. For infant and mother behaviors, the range of intraclass correlations, absolute type with raters random, was between .821 and .997 (all p ≤ .01). In Canada After the study received ethical clearance from the two participating hospitals and the second author’s university research ethics board, the perinatal clinics in the two hospitals distributed Consent to be Contacted Forms to pregnant women in the third trimester of their pregnancy. The women who signed the form were contacted by a research assistant, who explained the study. Mothers who agreed to participate had notices put on their medical charts so that attending nurses would notify the research assistants when the women gave birth. Research assistants (N = 8) visited the mother-infant dyads in their homes when the infants were 1 week, 1 month, 2 months, and 3 months. Mothers were seen by the same research assistant from the contact interview through the data collection visits. Mothers in the SSC group were requested to provide six hours of SSC with their infants cumulative throughout the day during the infants’ first week, and then two hours per day until the infants were one month. No request for mother-infant SSC was made to control group mothers. Mothers in both the SSC and control groups recorded the amount of SSC they provided to their infants each day and records were collected at each visit. The present report presents the results from the Still Face Task conducted in the home when the infants were 1 month and 2 months of age. The setup and procedure of the Still Face Task was identical to that done in Ghana with the exception that the initial interactive phase lasted for 3 min. Coding of the mothers and infants was done in the lab of the second author as described in the Ghana sample by a coder blind to whether the dyads were in the SSC or control groups. For reliability purposes, a second coder, who was also blind to whether the dyads were in the SSC or control group, independently scored the behaviors of 13% of the dyads. For infant and mother behaviors, the range of intraclass correlations, absolute type with raters random, was between .824 and .999 (all p ≤ .01). Plan of analysis ~~~~~~~~~~~~~~~~ In preliminary analyses, analyses of variance (ANOVAs) were conducted on the demographics of the Ghanaian mothers and infants in the high and low SSC groups and on the demographics of the Canadian mothers and infants in the SSC and control groups. ANOVAs also compared the amount of SSC the Ghanaian mothers in the high and low SSC groups did in their infants’ first week and in their infants’ weeks 2 through 4 with the amount of SSC the Canadian mothers in the SSC and control groups did in their infants’ first week and in their infants’ weeks 2 through 4. To test Hypotheses 1 and 2, mixed ANOVAs with the repeated variable phase (initial interactive, still face, reunion) and the between variable SSC group (high, low) were conducted on Ghanaian infants’ visual attention (Hypothesis 1) and smiles, non-distress vocalizations, frowns, and distress vocalizations (Hypothesis 2) during the Still Face Task. Significant group x phase interactions were followed by t-test pairwise comparisons. Similar ANOVAs (between variable group: SSC, control) and follow-up comparisons were conducted on the Canadian infants’ behaviors at 1 month and at 2 months. To test Hypothesis 3, ANOVAs with the between variable country (Ghana, Canada) compared the overall infant behaviors during the Still Face Task of Ghanaian infants (visual attention, smiles, non-distress vocalizations, frowns, distress vocalizations) with those behaviors of Canadian infants at 1 month and at 2 months. To test Hypothesis 4, ANOVAs with the between variable country (Ghana, Canada) compared the Ghanaian mothers’ behaviors toward their infants (visual attention, smiles, vocalizations, physical contact) during the interactive phases of the task with these behaviors of the Canadian mothers when their infants were 1 month and 2 months. Follow-up ANOVAs (between variable country: Ghana, Canada) compared the Ghanaian mothers’ type of vocalizations (arousing, neutral) and type of physical contact with their infants (arousing, attention getting, soothing, adjusting, passive holding) with those of the Canadian mothers when their infants were 1 month and 2 months. Preliminary analyses ~~~~~~~~~~~~~~~~~~~~ Dyads in the Ghanaian sample were divided into high and low SSC groups based on a median split of the amount of SSC reported during the infants’ first month. ANOVAs conducted on the demographics of the mothers and infants indicated there was no significant difference between the mothers in the high and low SSC groups in maternal age, education, or number of previous births; or in their infants’ sex, gestation age, or birth weight. For the dyads in the Canadian sample, ANOVAs conducted on the demographics of the mothers in the SSC and control groups indicated there were no significant differences between the groups in maternal education, number of previous births, or SES, although the mothers in the SSC group were slightly older (M = 32.0 years, SD = 5.5 years) than the mothers in the control group (M = 28.4 years, SD = 4.1 years). Infants in the SSC and control groups did not differ in sex, gestation age, or birth weight. Table 1 shows the amount of SSC the Ghanaian mothers in the high SSC and low SSC groups and the Canadian mothers in the SSC and control groups provided for their infants in the infants’ first week and in the weeks 2 through 4. There was no significant difference in the amount of SSC provided by Ghanaian and Canadian mothers in their infants’ first week either in the Ghanaian high SSC group and the Canadian SSC group or in the Ghanaian low SSC group and the Canadian control group. However, in the infants’ weeks 2 through 4, Ghanaian mothers provided more SSC in the high SSC group than Canadian mothers provided in the SSC group, F (1, 38) = 31.29, p < .001, ηp2 = .452, and Ghanaian mothers in the low SSC group provided more SSC than the Canadian mothers did in the control group, F (1, 62) = 58.54, p < .001, ηp2 = .486, but less SSC than Canadian mothers did in the SSC group, F (1, 38) = 16.44, p < .001, ηp2 = .302. Fig. 1 shows infants’ behaviors (visual attention, smiling, non-distress vocalizations, frowning, distress vocalizations) across the phases of the Still Face Task for the Ghanaian infants at 6 weeks and the Canadian infants at 1 and 2 months. Hypothesis 1: Ghanaian infants with both high and low SSC experience would show the still face effect with their visual attention, like the Canadian infants at 1 and 2 months ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Mixed ANOVAs with the repeated variable phase (initial interactive, still face, reunion) and between variable SSC group (high, low) conducted on the Ghanaian infants’ duration of visual attention showed a significant main effect for phase, F (2, 48) = 5.05, p = .010, ηp2 = .174, indicating the still face effect, with no group main effect or interaction effect. The Canadian infants showed similar results from mixed ANOVAs with the repeated variable phase (initial interactive, still face, reunion) and the between variable group (SSC, control) conducted on the infants’ visual attention at 1 and 2 months. At 1 month, the Canadian infants’ visual attention showed a significant main effect for phase, F (2, 152) = 4.65, p = .011, ηp2 = .058, with no group main effect or interaction effect. These results were replicated at 2 months (phase main effect: F (2, 146) = 18.34, p < .001, ηp2 = .201). As can be seen in Fig. 1a, the Ghanaian infants, like the Canadian infants at 1 and 2 months, discriminated between the phases of the Still Face Task with their visual attention. Hypothesis 2: Ghanaian infants with high SSC experience, but not those with low SSC experience, would show the still face effect with their affect, similar to the Canadian infants at 1 month ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Mixed ANOVAs with the repeated variable phase (initial interactive, still face, reunion) and between variable SSC group (high, low) were conducted on the Ghanaian infants’ duration of smiling, non-distress vocalizations, frowning, and distress vocalizations during the Still Face Task. For smiling, there was a significant main effect for phase, F (2, 48) = 4.31, p = .019, ηp2 = .152, and significant phase x group interaction, F (2, 48) = 3.50, p = .038, ηp2 = .127, with no group main effect. Fig. 2 shows the infants’ smiling across the phases for the Ghanaian infants in the high and low SSC groups. The analysis was redone after adjusting for outliers (n = 2 in initial interactive phase, n = 1 in still face phase, n = 2 in reunion phase) by Winsorizing (replacing the outlier’s data with the next closest participant’s data that is not an outlier). The results were unchanged. Ghanaian infants with high SSC experience, but not those with low SSC experience, showed the still face effect with their smiling. There were no significant effects between the Ghanaian infants in the high and low SSC groups for non-distress vocalizations, frowning, or distress vocalizations. Similar ANOVAs were conducted on the Canadian infants’ duration of smiling, non-distress vocalizations, frowning, and distress vocalizations during the Still Face Task at 1 and 2 months. For smiling, the Canadian infants at 1 month showed a main effect for phase F (2, 152) = 3.36, p = .037, ηp2 = .042, indicating a linear decline over the phases, with no main effect for group or interaction effect (see Fig. 1b). However at 2 months, the Canadian infants’ smiling showed a significant main effect for phase, F (2, 146) = 19.98, p < .001, ηp2 = .215, indicating the still face effect, with no group main effect or interaction effect (see Fig. 1b). Thus Canadian infants in both the SSC and control groups did not discriminate between the phases of the Still Face Task with their smiling at 1 month, but infants in both the SSC and control groups did so at 2 months. However, at 1 month the Canadian infants in the SSC group, but not those in the control group, showed the still face effect with their non- distress vocalizations. A mixed ANOVA yielded a main effect for phase, F (2, 152) = 6.86, p = .001, ηp2 = .083, and a phase x group interaction, F (2, 152) = 3.85, p = .023, ηp2 = .048. Infants in the SSC group showed the still face effect with their non-distress vocalizations at 1 month (initial interactive: M = 9.9, SD = 18.4; still face: M = 2.4, SD = 3.9; reunion: M = 4.6, SD = 6.3), whereas infants in the control group did not (initial interactive: M = 3.9, SD = 4.3; still face: M = 2.8, SD = 4.3; reunion: M = 3.3, SD = 4.7). Thus like the Ghanaian infants in the high SSC group, the Canadian infants in the SSC group prior to 2 months of age demonstrated the still face effect with their affect. The Canadian infants increased their negative affect with their frowns and distress vocalizations through the phases of the task at both 1 and 2 months, as can be seen in Fig. 1d and e. At both ages, infants showed a phase effect for frowns (1 month: F (2, 152) = 4.711, p = .010, ηp2 = .058; 2 months: F (2, 146) = 6.262, p = .002, ηp2 = .079) and for distress vocalizations (1 month: F (2, 152) = 5.686, p = .004, ηp2 = .070; 2 months: F (2, 146) = 6.818, p = .001, ηp2 = .085), indicating an increase in negative affect from the initial interactive phase to the still face phase that remained or increased at the reunion phase. Hypothesis 3: Ghanaian infants would show overall behaviors of visual attention, smiling, non-distress vocalizations, frowning, and distress vocalizations that are intermediate between those behaviors found in the Canadian infants at 1 and 2 months ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 3 shows the duration of visual attention, smiling, non-distress vocalizations, frowning, and distress vocalizations during the Still Face Task for the 6-week-old Ghanaian infants and the Canadian infants at 1 month and 2 months. ANOVAs comparing the Ghanaian infants with the Canadian infants at 1 month indicated there was no significant difference in visual attention or smiling between the Ghanaian infants and the Canadian infants, but the Canadian infants were making more non-distress vocalizations, F (1, 97) = 5.49, p = .021, ηp2 = .054, frowning more, F (1, 97) = 19.61, p < .001, ηp2 = .168, and making more distress vocalizations, F (1, 97) = 14.62, p < .001, ηp2 = .131. ANOVAs comparing the Ghanaian infants with the Canadian infants at 2 months indicated that the Canadian infants were attending more, F (1, 99) = 13.87, p < .001, ηp2 = .123, smiling more, F (1, 99) = 10.25, p = .002, ηp2 = .094, making more non-distress vocalizations, F (1, 99) = 25.28, p < .001, ηp2 = .203, frowning more, F (1, 99) = 11.04, p = .001, ηp2 = .100, and making more distress vocalizations, F (1, 99) = 6.88, p = .010, ηp2 = .065. When analyses were redone adjusting for outliers (n = 2 for smiles; n = 1 for non-distress vocalizations; n = 3 for frowns; n = 3 for distress vocalizations) by Winsorizing, the results were unchanged. Thus, rather than showing overall behaviors that were between the amounts shown by the Canadian infants at 1 and 2 months, Ghanaian infants at 6 weeks showed visual attention and smiling behaviors comparable with the Canadian infants at 1 month, but showed less non-distress vocalizations, frowns, and distress vocalizations than Canadian infants at 1 or 2 months. Hypothesis 4: Ghanaian mothers would show more tactile and less verbal behaviors but similar visual attention and smiling behaviors during the interactive phases of the Still Face Task compared to the Canadian mothers ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 4 shows mothers’ duration of visual attention, smiling, vocalizing, and physical contact with their infants during the interactive phases of the Still Face Task for the Ghanaian mothers and for the Canadian mothers at the 1 and 2 month visits. ANOVAs indicated there was no significant difference in the duration of attention, vocalizations, or physical contact with their infants for the Ghanaian mothers and the Canadian mothers at the 1 or 2 month visits. Ghanaian mothers smiled more to their infants than Canadian mothers did at the 1 month visit, F (1, 104) = 12.10, p < .001, ηp2 = .111, but there was no significant difference in the Ghanaian mothers’ smiling and the Canadian mothers’ smiling at the 2 month visit. Analyses adjusted for outliers (n = 1 for gaze; n = 1 for vocalizations) by Winsorizing did not change the findings. However, the Ghanaian and Canadian mothers differed on the type of vocalizations and physical contact they did with their infants during the interactive phases of the Still Face Task. Fig. 5 shows the mothers’ vocalizations coded as arousing and neutral. Ghanaian mothers’ vocalizations were more arousing than the Canadian mothers’ vocalizations at the 1 month visit, F (1, 97) = 283.62, p < .001, ηp2 = .745, and at the 2 month visit, F (1, 99) = 200.87, p < .001, ηp2 = .670. Canadian mothers’ vocalizations were more neutral at both the 1 month, F (1, 97) = 159.04, p < .001, ηp2 = .621, and 2 month visits, F (1, 99) = 184.39, p < .001, ηp2 = .651 than the Ghanaian mothers’ vocalizations. When adjusted for outliers (n = 3 for neutral vocalizations) by Winsorizing, results did not change. Fig. 6 shows the mothers’ physical contact with their infants coded as arousing, attention getting, soothing, adjusting, and passive holding. Ghanaian mothers provided more arousing physical contact with their infants than Canadian mothers at the 1 month visit, F (1, 97) = 124.22, p < .001, ηp2 = .562, and at the 2 month visit, F (1, 99) = 99.48, p < .001, ηp2 = .501. Ghanaian mothers did less attention getting physical contact than Canadian mothers did at the 1 month visit, F (1, 97) = 10.95, p = .001, ηp2 = .101, but there was no difference at the 2 month visit. Ghanaian mothers did less soothing physical contact than Canadian mothers did at the 1 month visit, F (1, 97) = 9.11, p = .003, ηp2 = .086, and at the 2 month visit, F (1, 99) = 5.53, p = .021, ηp2 = .053. Ghanaian mothers’ amount of adjusting physical contact was similar to that of Canadian mothers at the 1 month visit, but they did more than Canadian mothers at the 2 month visit, F (1, 99) = 7.13, p = .009, ηp2 = .067. Ghanaian mothers did less passive holding than Canadian mothers did at the 1 month visit, F (1, 97) = 41.75, p < .001, ηp2 = .301, and at the 2 month visit, F (1, 99) = 35.21, p < .001, ηp2 = .262. When analyses were redone adjusting for outliers (n = 3 for attention getting; n = 2 for soothing; n = 4 for adjusting; n = 2 for passive holding) by Winsorizing, the only changes in the results were that Ghanaian mothers had less attention getting physical contact than Canadian mothers at the 2 month visit, and Ghanaian mothers and Canadian mothers at the 2 month visit had similar amounts of adjusting physical contact. As a check to determine whether SSC affected the mothers’ behavior during the Still Face Task, ANOVAs with the between variable group (high SSC, low SSC) were conducted on the Ghanaian mothers’ visual attention, smiling, vocalizations, and physical contact with their infants during the task. There were no significant differences between the groups for any of the behaviors. Likewise, ANOVAs with the between variable group (SSC, control) were conducted on the Canadian mothers’ visual attention, smiling, vocalizations, and physical contact with their infants during the task at the 1 month and 2 month visits. Only the ANOVA for maternal physical contact at the 1 month visit showed a significant group effect, F (1, 76) = 6.52, p = .013, ηp2 = .079, indicating mothers in the control group provided more physical contact (M = 81.8, SD = 16.7) during the Still Face Task than mothers in the SSC group (M = 67.3, SD = 31.4). Follow-up ANOVAs were conducted on the type of physical contact mothers provided for their infants at the 1 month visit. Control group mothers provided more passive holding (F (1, 76) = 6.09, p = .016, ηp2 = .074; SSC group M = 31.8, SD = 28.0; control group M = 47.0, SD = 24.6) and SSC group mothers provided more soothing physical contact (F (1, 76) = 4.51, p = .037, ηp2 = .056; SSC group M = 17.2, SD = 24.7; control group M = 8.1, SD = 13.5). Ghanaian mothers and Canadian mothers were similar in the amount of attention, smiling, vocalizing, and physical contact directed toward their infants during the interactive phases of the Still Face Task, with the exception that Ghanaian mothers smiled more than Canadian mothers at the 1 month visit. However, Ghanaian mothers were more arousing in their vocalizations and physical contact with their infants than Canadian mothers. Canadian mothers’ vocalizations were more neutral and their physical contact with their infants was used for more attention getting, soothing, and passive holding. Adjusting physical contact was minimal for both groups.","Ghanaian mother-infant dyads at the infant age of 6 weeks showed both similarities and differences in their interactions in the Still Face Task compared to Canadian mother- infant dyads at the infant ages of 1 and 2 months. Importantly, Ghanaian infants with high SSC experience, but not those with low SSC experience, showed affect responses to their mothers’ changing behavior during the task in ways that were similar to Canadian infants in the SSC group at 1 month. Ghanaian infants in both the high and low SSC groups showed the still face effect with their visual attention, providing support for the first hypothesis. The Canadian infants at 1 and 2 months also showed the still face effect with their visual attention. Young infants from both cultural contexts demonstrated with their attention that they detected changes in the mothers’ behavior. Indeed, Bigelow and Power (2012), who longitudinally followed the Canadian infants from 1 week to 3 months of age, found that even at 1 week the infants showed the still face effect with their attention. Although few other studies have used the Still Face Task with newborns, the ones that have done so indicate that infants reduce their attention from the initial interactive phase to the still face phase (Bertin & Striano, 2006; Ellsworth, 1987, cited in Muir & Hains, 1993; Nagy, 2008). Thus infants show evidence of detecting changes in their mothers’ interactive behavior from the beginning of life. Ghanaian infants with high SSC experience, but not those with low SSC experience, showed the still face effect with their affect. They smiled more in the interactive phases than the still face phase, providing support for the second hypothesis. Researchers speculate that the systems that regulate infants’ attention and affect operate differently during face-to-face interactions (Bigelow & Walden, 2009; Legerstee & Varghese, 2001; Nadel, Soussignan, Canet, Libert, & Gerardin, 2005). They propose that changes in infants’ visual attention during the Still Face Task indicate infants’ detection of changes in their mothers’ behavior, whereas changes in infants’ affect indicate they are responding to violations of their expectations for affect sharing, which develops later than infants’ ability to detect changes in their mothers’ behavior. Ghanaian infants with high amounts of SSC affectively responded to their mothers’ changing behavior during the Still Face Task, suggesting SSC accelerated the infants’ expectations for affect sharing. For Canadian infants, there was no evidence of the still face effect for smiling at 1 month, however by 2 months, infants in both the SSC and control groups showed the still face effect with their smiling, which is evidence of the 2 month transition when infants become more interested and participatory social partners (Henning et al., 2005; Lavelli & Fogel, 2002, 2005; Rochat, 2001). Yet Canadian infants in the SSC group at 1 month showed the still face effect with their affect via their non-distress vocalizations. Although smiling is the most frequently used measure of infants’ positive affect, infants’ non-distress vocalizations convey their emotional reactions and are the primary social signals that elicit maternal communication (Hsu & Fogel, 2001; Keller, Lohaus, Volker, Cappenberg, & Chasiotis, 1999; Papousek, 1989; Van Egeren, Barratt, & Roach, 2001). Mothers use these vocalizations as signals for determining infants’ readiness to interact and for adjusting their own emotional response. Thus both Ghanaian and Canadian infants with high SSC experience responded to the Still Face Task with their affect prior to 2 months of age. Why would Ghanaian infants’ early affect response to the Still Face Task be with smiling and Canadian infants’ early affect response be with non-distress vocalizations? Kärtner et al. (2008, 2010) found that the Nso mothers of the Cameroon were as contingently responsive to their infants as German mothers, but their predominant mode of responsiveness was expressed through physical contact. German mothers responded contingently to their infants primarily visually with gaze, smiles, and facial expressions and with verbal turn-taking. Nso mother-infant dyads engaged less in face-to-face interactions with periods of mutual gaze. Thus Nso infants experienced their mothers’ smiling during maternal interactions less often. Yet when Nso mothers were in mutual gaze with their infants, they smiled as much as German mothers. The Still Face Task involves mother-infant face-to-face interaction with mutual gaze. Sroufe (1996) proposed that the emergence of social smiling is facilitated by infants’ recognition of others’ contingent responsiveness to their behavior, not only in face-to- face contexts but also in contexts of other modalities, such as body contact or stimulation. SSC has been shown to facilitate mothers’ responsiveness to their infants’ signals as well as infants’ sensitivity to their mothers’ behavior toward them (Feldman, 2004; Feldman, Eidelman, et al., 2002). Such increases in mother-infant reciprocity may explain the Ghanaian infants in the high SSC group being more reactive with their smiling to their mothers’ changing behavior during the task. Canadian infants would have more face-to face experience with their mothers in which verbal turn-taking is prevalent. Although mothers also smile responsively in face-to-face interactions with their infants, mothers’ vocal responsiveness may be more easily noticed. Young infants’ smiling is instigated by maternal smiling (Kaye & Fogel, 1980), therefore infants’ smiles tend to overlap with mothers’ smiles. In Western societies where mother-infant vocal turn-taking is prevalent, when infants vocalize, overlap with maternal talking is minimal (Bigelow, MacLean et al., 2010; Bornstein et al., 1992). Mothers tend to stop talking and resume after infant vocalizations have ended, resulting in vocal turn-taking. Infants who experience such turn-taking may recognize their mothers’ responsiveness more readily with vocalizations than with smiles. Interestingly, SSC experience affected infants’ behavior toward their mothers in the Still Face Task, yet SSC did not influence the behavior of the mothers toward their infants during the task. Ghanaian mothers in both high and low SSC groups showed similar behavior toward their infants. Canadian mothers in SSC and control groups also showed similar behavior toward their infants, with the exception that mothers in SSC and control groups differed in the amount of physical contact with their infants during the Still Face Task at 1 month. Mesman et al. (2009) indicate that the history of infants’ experience with their mothers is an important aspect affecting infants’ response to the Still Face Task. The findings suggest that the infants’ history of SSC experience with their mothers was a significant factor influencing the infants’ response to the task. Ghanaian infants’ overall duration of attention and smiling during the Still Face Task paralleled the overall duration of those behaviors by the Canadian infants at 1 month, rather than at 2 months. Kärtner et al. (2010) found that, although Nso infants experience less mutual gaze with their mothers, when in face-to-face interaction with their mothers, Nso infants at 6 weeks gazed and smiled at their mothers as much as German infants at this age (Wörmann et al., 2012). The Ghanaian infants, however, engaged in less non-distress vocalizing than the Canadian infants at 1 month, possibly because the Canadian infants in the SSC group at 1 month were already responding to the Still Face Task with their non- distress vocalizations. By 2 months, the Canadian infants engaged in more attention, smiling, and non-distress vocalizing than the Ghanaian infants. In comparing German and Nso mothers’ interactions, cultural differences in the modes of maternal responsiveness emerged in the infants’ second and third months (Kärtner et al., 2010). The divergence was mostly due to German mothers decreasing their tactile responsiveness and increasing their visual responses of gaze, smiles, and facial expressions; whereas Nso mothers maintained their high level of tactile responsiveness with little increase in visual responses. Mother-infant interactions are transactional in that both partners influence each other. Yet mothers are primarily responsible for establishing and maintaining interactions with their infants, particularly in the infants’ early life (Henning & Striano, 2011; Kaye & Fogel, 1980). Ghanaian infants’ experience of maternal responsiveness may be more similar to that of Canadian infants at 1 month, prior to the divergence of cultural modes of maternal responsiveness, than that of Canadian infants at 2 months. Notably, Ghanaian infants differed from Canadian infants at both 1 and 2 months in distress, as expressed in frowns and distress vocalizations of fussiness and crying. Ghanaian infants were less distressed during the Still Face Task. Proximal parenting cultures value calmness in infants and the suppression of negative affect (Keller & Otto, 2009). Mothers in both Western and non-Western cultures tend to respond readily to infant distress (Mesman et al., 2018). Yet the physical closeness of Ghanaian mother-infant dyads, through mothers carrying their infants and sleeping with them, facilitates the mothers’ prompt response to, and anticipation of, infant distress, which promotes mothers’ nurturance of calmness in their infants. Ghanaian and Canadian mothers were similar in the amount of infant directed behaviors they provided during the interactive phases of the Still Face Task. Nso mothers were found to smile as much as German mothers when in mutual gaze with their infants, despite mutual gaze being less frequent for Nso mother-infant dyads (Kärtner et al., 2010; Wörmann et al., 2012). The Still Face Task involves face-to-face interaction in which mutual mother-infant gaze is prevalent. In this context, the amount of maternal attention, smiling, vocalizing, and physical contact was similar for Ghanaian and Canadian mothers, even though Ghanaian mothers smiled more than Canadian mothers when their infants were 1 month. However, the mothers’ type of vocalizations and physical contact with their infants differed. Ghanaian mothers’ vocal and physical contact behaviors were more arousing. Their vocalizations tended to be boisterous rhythmic sounds, energetic and repetitive infant greetings and singing. These vocal behaviors were accompanied by moving the infants’ limbs in rhythm with the vocalizations. Maternal touch is important in regulating infant emotion. In Western societies, mothers use touch to reduce distress and elicit positive affect (Stack & Arnold, 1998; Stack & Muir, 1992). Canadian mothers’ tactile behavior was most often used for soothing the infant, getting the infant’s attention, or simply passively touching the infant, and their vocalizations were less stimulating. Mothers from proximal caretaking cultures of West Africa have been described as treating their infants as novices who need to learn compliance and subordination (Kärtner et al., 2010). Mothers are directive in this process, which influences the ways in which they engage with their infants. Mothers tend to lead the interaction rather than follow the infant; synchrony rather than reciprocity is prominent, which promotes the cultural goals of interdependency and relatedness (Keller et al., 2004). Mother-infant interaction is characterized by rhythmic structuring with musical chorusing and repeated greetings, synchronized with rhythmic body stimulation. These modes of interaction were seen in the Ghanaian mothers’ behaviors in the interactive phases of the Still Face Task. Ghanaian infants were primarily calmly attentive to these energetic behaviors of their mothers. There are several limitations to the study. The Ghanaian mothers, coming from a proximal parenting culture, were less accustomed to face-to-face interactions with their infants (Kärtner et al., 2008; Keller, 2007; Keller et al., 2004), which may have affected the way in which they vocally and tactually engaged with their infants during the interactive phases of the Still Face Task. Yet the Ghanaian mothers’ behavior was similar to that described by other researchers studying mother-infant engagement during naturalistic face-to-face interactions in West African cultures (Kärtner et al., 2010). Although the study examined Ghanaian mothers’ behaviors toward their infants during the Still Face Task, the study did not assess the mothers’ contingent responsiveness to their infants’ behavior. Mothers from West African societies tend to be predominantly tactually responsive to their infants, but not necessarily while in mutual gaze with their infants (Kärtner et al., 2008, 2010). Thus, to assess Ghanaian mothers’ contingent responsiveness would require measuring their tactile responsiveness during naturalistic mother-infant interactions, which is beyond the scope of the current data. The amount of daily SSC Ghanaian and Canadian mothers provided for their infants during the infants’ first month was based on mothers’ self reports. It is possible that the mothers did not accurately report the amount. The sample size of Ghanaian mother-infant dyads was small, which reduces the power of analyses, particularly in finding significant results. That significant results were obtained with the limited sample size suggest the findings are robust. Yet the size of the sample limits the interpretation of the results. The Ghanaian mother-infant dyads were seen in a cross-sectional study at the infant age of 6 weeks, whereas the Canadian archival data to which they were compared come from a longitudinal study at the infant ages of 1 and 2 months. Both studies examined the effect of SSC on infants’ response to the Still Face Task prior to the 2 month transition; and the Canadian study is the only known SSC study utilizing the Still Face Task in Western culture with infants at ages comparable to the Ghanaian infants. Nevertheless, the differences in the study designs may have affected the results found between the two cultures. The length of the initial interaction differed slightly in the Ghanaian and Canadian studies (2 min and 3 min, respectively), although length of episodes has been shown to have little effect on infants’ response to the task (Adamson & Frick, 2003; Mesman et al., 2009). The infants were seated in a car seat for the Still Face Task, which may have been novel for some of the Ghanaian infants as the use of car seats is not necessarily common practice in Ghana. Yet the Ghanaian infants did not resist the seat and did not seem distracted by it. Ghanaian mothers’ demographics differed from those of the Canadian mothers, which may have affected their maternal behaviors. However, the demographics of the Ghanaian mothers are prototypical of their cultural context, and thus are interwoven with their parenting goals and practices. Despite cultural differences in how mothers engaged with their infants, mother-infant SSC facilitated infants’ early expectations for their mothers’ behavior. Ghanaian infants with high SSC experience, on the cusp of the 2 month transition, responded with their smiling to changes in their mothers’ behavior during the Still Face Task, just as Canadian infants in the SSC group prior to 2 months responded with their affect to their mothers’ behavioral changes during the task. In both cultures, SSC enhanced infants’ ability to emotionally respond to their mothers’ social behavior, suggesting acceleration of the infants’ developing expectations for their mothers’ engagement.","This research was aided by a grant from the Bill and Melinda Gates Foundation to the first (Principal Investigator) and second authors and by a grant from the Nova Scotia Health Research Foundation to the second author. The funders had no involvement in the design, conduction, or writing of the report. Gratitude is expressed to Yasmin Mohammed, Jan Hanifen, Gerry Cameron, Rachel MacFarlane, Mena Enxuga, Jennifer Delaney, Yvonne MacDonald, Cynthia Flannigan, Lynne Lukeman, Dale Fewer, Charlene Kennedy-Chisholm, Chow Shim Pang, Laura Walden, Caitlin Best, and Madison Links, who were research assistants; and to the mothers and infants who participated in the study."],["Previous research has reported mixed findings regarding executive function (EF) abilities in developmental coordination disorder (DCD), which is diagnosed on the basis of significant impairments in motor skills. The current study aimed to assess whether these differences in study outcomes could result from the relative motor loads of the tasks used to assess EF in DCD. Children with DCD had significant difficulties on measures of inhibition and planning compared to a control group, although there were no significant correlations between motor skills and EF task performance in either group. The complexity of the response, as well as the component skills required in EF tasks, should be considered in future research to ensure easier comparison across studies and a better understanding of EF in DCD over development. © 2014 Elsevier Ltd. --------------------------------------------------------------------------------","Executive function (EF) is an umbrella term that includes a range of top-down processes of cognitive control, characterised by Miyake, Friedman, Emerson, Witzki, and Howerter (2000) as comprising three core functions, namely response inhibition, shifting between tasks or mental sets, and updating/monitoring of working memory representations. These core functions provide the basis for higher-order functions such as planning and reasoning (Diamond, 2013). EFs develop over a protracted period, emerging before birth and continuing to develop throughout adolescence and into early adulthood (e.g., Anderson, 2002; Best & Miller, 2010). In a variety of psychological and medical conditions, executive dysfunction has been associated with significant negative consequences for daily life functioning, academic achievement, and employability (Altshuler et al., 2007; Biederman et al., 2004; Garcia-Villamisar & Hughes, 2007; Gilotty, Kenworthy, Sirian, Black, & Wagner, 2002). Various patterns of executive dysfunction have been reported across a number of clinical disorders (see reviews by Hill, 2004; Sergeant, Guerts, & Oosterlaan, 2002). For example, Ozonoff and Jensen (1999) reported poor planning and cognitive flexibility but typical inhibitory skill in children/adolescents with autism spectrum disorder (ASD), and the reverse profile in children/adolescents with Attention Deficit-Hyperactivity Disorder (ADHD). Similarly, Happé, Booth, Charlton, and Hughes (2006) reported specific and differing profiles of executive functioning in children and adolescents diagnosed with ASD vs. ADHD, with individuals with ADHD showing a more widespread and general impairment, particularly in response inhibition, whereas those with ASD showed greater difficulties with response selection and monitoring. The literature regarding EF across neurodevelopmental disorders has paid less attention to developmental coordination disorder (DCD), which is diagnosed on the basis of movement difficulties that interfere with academic achievement or activities of daily living, such as dressing or eating (DSM-5, American Psychiatric Association, 2013). These movement difficulties cannot be the result of any known intellectual disability or medical condition such as cerebral palsy. As in ASD and ADHD, reports suggest that individuals with DCD have difficulties in many aspects of EF (see Wilson, Ruddock, Smits-Engelsman, Polatajko, & Blank, 2012), particularly in the three key components of EF identified by Miyake et al. (2000) of response inhibition (e.g., Mandich, Buckolz, & Polatajko, 2002; Michel, Roethlisberger, Neuenschwander, & Roebers, 2011; Piek et al., 2004; Piek, Dyck, & Francis, 2007; Querne et al., 2008; Wisdom, Dyck, Piek, & Hay, 2007), working memory (e.g., Alloway & Archibald, 2008; Alloway, 2007, 2011; Michel et al., 2011; Piek et al., 2004, 2007; Wisdom et al., 2007), and switching (e.g., Michel et al., 2011; Piek et al., 2004, 2007; Wisdom et al., 2007; Wuang, Chwen-Yng, & Su, 2011). These studies have suggested that children with DCD perform more poorly or with more variability than their typically developing counterparts on a range of tasks, although the patterns of impairments and variability in DCD groups are not always the same, with areas of relative strength in some studies appearing to be relative weaknesses in others. For example, when testing switching, Michel et al. (2011) and Piek et al. (2004) found no differences between children with motor difficulties and controls in terms of the numbers of errors made, while Wuang et al. (2011) and Piek et al. (2007) reported significantly more errors in children with motor difficulties than controls. Differences between studies may be due to the age ranges and tasks used across research groups, or may rely on the recruitment method used (e.g., screening using different percentile cut-offs for motor difficulty vs. recruitment of clinically referred children). The current study includes only children with a clinical diagnosis of DCD in order to better understand this group in terms of their executive functioning profile. In terms of inhibition and working memory, some tasks are used to assess both functions (e.g., the trailmaking/updating task used by Piek et al., 2004, 2007), while Michel et al. (2011) used separate tasks for these two functions. The tasks also differ in the extent to which they rely on motor skills, with tasks such as the trailmaking/updating task requiring button press responses, while the ‘Fruit Stroop’ task used by Michel et al. having no motor demands. Studies within normative samples have reported a significant relationship between motor abilities and response inhibition (Livesey, Keen, Rouse, & White, 2006; Rigoli, Piek, Kane, & Oosterlaan, 2012), and motor skills in DCD have been reported to significantly predict working memory (Michel et al., 2011; Piek et al., 2004). The level of impairment or variability in the DCD group could therefore be affected by the extent to which the EF task relies on complex motor responses. A study by van Swieten et al. (2009) supported this suggestion, demonstrating developmentally inappropriate motor planning in 6–13-year-old children with DCD, but appropriate executive planning (using a Tower of London task) in 7–11-year-olds in this group. The present study aims to address this issue by comparing performance on tasks that require a greater motor load to those that have a reduced motor load. Two EF components were selected for the current investigation, namely planning and inhibition, both of which have previously been tested in DCD with tasks that require greater or reduced motor output. While planning is not one of the core EFs identified by Miyake et al. (2000), it is suggested to build on core functions such as working memory (Diamond, 2013), and deficits in the planning and control of motor actions are likely to be key to the movement difficulties seen in DCD (see Hill, 1998, for a review). Inhibition is often investigated using tasks that involve button presses or other motor responses, and so it is important to assess the extent to which any difficulties or additional processing load associated with producing these responses affects inhibition performance in children with DCD. In the current study, tests of planning and inhibition were taken from different executive functioning measures, and were chosen according to their relative motor loads (i.e., high vs. reduced motor response required). Each executive function was therefore measured using two tasks: Planning was assessed by the NEPSY Tower task (Korkman, Kirk, & Kemp, 1998; reduced motor-load) and the Rotational Bar task (Rosenbaum et al., 1990; high motor-load). Inhibition was assessed by the Stroop task (Stroop, 1935; reduced motor-load) and by the NEPSY Knock-Tap task (Korkman et al., 1998; high motor-load). These tasks are described in more detail in Sections 2.2.1 and 2.2.2, and the high motor-load tasks are presented graphically in Fig. 1. The NEPSY Tower task was used to measure planning with a reduced motor load, in line with van Swieten et al. (2009), and was compared to a motor planning task. Specifically, the Rotational Bar task developed by Rosenbaum et al. (1990) was used, in which participants are required to pick up and rotate a bar so that a coloured end of the bar is placed on a specific coloured disc on a table. This requires participants to plan their grips in order to end in a comfortable position (achieving ‘end-state comfort’). Using this task, Smyth and Mason (1997) found no significant differences between 4- and 8-year- old children screened for movement difficulties and a control group with typical movement skills in the proportion of grips ending in a comfortable state, although van Swieten et al. (2009) found increasing differences with age between children with DCD and controls in grip selection on a related task. Given that the children in the current study were of a similar age range to those tested by van Swieten et al. (6–14 years and 6–13 years, respectively), the hypotheses were based on the latter study. Specifically, it was predicted that children in the DCD group would perform more poorly than the control group on the Rotational Bar task (high motor-load) but not on the NEPSY Tower task (reduced motor-load). In the inhibition tasks, participants with DCD were expected to perform more poorly than the control group in the Knock-Tap task (high motor-load), but not in the Stroop task (reduced motor-load).","Twenty-six children and adolescents diagnosed with DCD, and 24 children and adolescents without a DCD diagnosis (hereafter, ‘typically developing group’) were recruited through schools and DCD support groups. In the DCD group, only those with a clinical diagnosis of DCD made according to the full DSM-IV-TR criteria (American Psychiatric Association, 2000) and without additional diagnoses, such as ADHD, ASD or dyslexia, were included. The following criteria were necessary for a diagnosis of DCD to be given under DSM-IV-TR: (A) performance in daily activities that require motor coordination was substantially below that expected, given the child's chronological age and measured intelligence; (B) the disturbance in Criterion A significantly interfered with academic achievement or activities of daily living; (C) the disturbance was not due to a general medical condition (e.g., cerebral palsy, hemiplegia, or muscular dystrophy) and did not meet criteria for a Pervasive Developmental Disorder; (D) if mental retardation was present, the motor difficulties were in excess of those usually associated with it. All participants completed the Movement Assessment Battery for Children – 2nd edition (Movement ABC-2; Henderson, Sugden, & Barnett, 2007) to further document their level of movement skill and the Wechsler Intelligence Scale for Children-Fourth Edition (WISC-IV; Wechsler, 2004) to measure their IQ (see Section 2.2 for further details of these tests). Children were included in the typically developing (TD) group only if they had not received a diagnosis of any neurodevelopmental disorder prior to participation in the study, and if they performed above the 15th centile on the Movement ABC-2. Children were only included in each group if they had WISC-IV Verbal Comprehension scores above 70 (an IQ below this cut- off suggests intellectual disability). The Verbal Comprehension scores were used rather than a full IQ score, as Full-Scale IQ measures encompass tasks with a high motor load or executive functioning component, and may therefore disadvantage children with DCD by giving a pessimistic estimate of their full IQ. Differences were indeed found between groups in WISC-IV Perceptual Reasoning scores (see Table 1), and so these scores were taken into account in the analyses. Participant characteristics, including Verbal Comprehension and Perceptual Reasoning, are presented in Table 1.","The Movement ABC-2 (Henderson et al., 2007) is a standardised test of motor skills suitable for children aged 3–16. It consists of three subtests: Manual Dexterity, Aiming and Catching, and Static and Dynamic Balance, each of which comprises a series of speeded and non-speeded motor tasks. Scores for each component can be converted to standard scores and percentile ranks, and a Total Standard Score can also be calculated from the components (M = 10, SD = 3). The Total Score percentile can be used as an indicator of motor difficulties, with scores below the 5th percentile suggesting a significant motor difficulty, and between the 6th and 15th percentiles signifying a borderline motor difficulty. The WISC-IV (Wechsler, 2004) is a standardised test of verbal and nonverbal abilities and is suitable for children aged 6–16. There are 10 subtests that contribute to a number of indices of intellectual functioning. For the current report, only the Verbal Comprehension Index (VCI) and Perceptual Reasoning Index (PRI) scores are reported. Both indices have a mean of 100 and a standard deviation of 15. Tests of planning The NEPSY Tower task (Korkman et al., 1998; hereafter, ‘Tower task’) required participants to move a set of three balls on three pegs from a start to a target configuration while following certain rules. Participants were shown a target picture of the apparatus, with the coloured balls in specific positions across the pegs, and were asked to copy the picture using their equipment. Participants first completed a practice trial in which they moved a ball so that the model matched the picture, and were assisted if necessary in order to ensure that they understood the task. Participants were asked to complete the trial in a set number of moves while adhering to several rules, namely: (1) a move was finished once the hand was removed from the ball; (2) only one ball could be moved at a time; (3) balls that were not being moved must remain on their pegs at all times and must not be placed anywhere else. Each time a participant broke one of these rules they scored one mark for a violation and automatically scored zero for the trial in question, but were allowed to continue with the task. A score of one point was awarded for each correct trial. The task had seven stages, with each stage becoming progressively more difficult in terms of planning complexity and the number of moves required to complete the trial. The first stage could be completed in one move, and the last stage could be completed in seven moves. The task was stopped once a participant had obtained four consecutive scores of zero on trials, or if they had completed all 20 trials. Although the standardised task was timed, participants were not automatically stopped at the cut-off point of 45 s for later trials, as it was felt that it was important to assess if the task was too difficult for participants in the DCD group, even once the time limit had been passed. The number of violations committed during the task provided one dependent variable in the analyses. Raw score (out of 20) was the other dependent variable for this task. The Rotational Bar task (Rosenbaum et al., 1990) required participants to move a coloured rod that rested on a tripod on the tabletop and place one end of the rod onto one of two coloured discs (see Fig. 1a). One side of the bar was blue, and the other side was red, and these ends of the bar were placed in front of the matching coloured disc on the table (in Fig. 1a, the blue parts of the apparatus are represented by the darker tones, and the red parts by the lighter tones). Participants were required to pick up the bar and place one of the coloured ends of the bar onto one of the coloured discs. In order to complete this task, participants could use an underhand grip, with their palm facing up, or an overhand grip, with their palm facing down, to pick up the bar, and this choice would affect how comfortable their final arm position would be, i.e., their level of ‘end-state comfort’ (ESC; Rosenbaum et al., 1990). Participants completed four practice trials before the task began. There were 8 trials in which an overhand grip would achieve ESC and 8 in which an underhand grip would be the most comfortable. Two marks were given for each trial in which a participant used the correct grip to pick up the bar and also placed the correct end of the bar onto the coloured disc. If they began with the wrong grip, but adjusted it to place the correct end of the bar onto the coloured disc, one mark was awarded. If participants used the wrong grip throughout the trial, and thus did not achieve ESC, they did not receive any points for that trial. A total of 32 points was therefore available for each participant, and the raw score (out of 32) provided the dependent variable for this task. Tests of inhibition The Stroop task (Stroop, 1935) required participants to name the colour of the ink in which a colour word was printed (e.g., the word ‘blue’ printed in red ink; response ‘red’). Participants had a maximum of 2 min to read out 112 words that were either congruent (set 1) or incongruent (set 2) with the ink in which they were printed. The number of correct responses in congruent and incongruent trials (both out of 112) was recorded and provided the dependent variables for this task. Before the task, participants were asked to read a list of words in order to ensure that any reading difficulties would not affect performance, and also read an example word in an incongruent colour to ensure that they understood the task. The NEPSY Knock-Tap task (Korkman et al., 1998; hereafter, ‘Knock-Tap task’) required participants to lay their non-preferred hands on the table and to use their preferred hands to complete certain actions, which the experimenter explained at the beginning of each set of trials (see Fig. 1b). The actions were: knocking on the table with knuckles, tapping the table with the palm of the hand, placing the side of the fist on the table, or no response. In the first set of 15 trials, only the knock and tap responses were used, and participants were required to do the opposite action to the experimenter (i.e., if the experimenter knocked, the participant should tap). In the second set of 15 trials, the participant response to the experimenter's knock was the side fist, and the response to the side fist was a knock. If the experimenter tapped, the participant should provide no response. Before each set of trials, participants completed a series of practice trials in which the possible action–response pairs were presented twice each. Once the participant understood the task, the main trials began. One point was awarded for each correct response, providing a score out of 30, which was entered as the dependent variable in the analyses. Procedure Participants were part of a larger study into the relationships between movement abilities, cognition and emotional well-being. They were visited in their own homes or at school to complete the testing battery, which could be administered in one session over the course of a day or over several shorter sessions, depending on the needs of the child and the constraints of the testing setting. The executive functioning tasks detailed in the current paper were always completed in one session, and in a quiet room without distractions. Tasks were presented to participants in a randomised order, and the session lasted approximately 30–40 min. Breaks and rewards were given throughout the testing session as necessary. Parents signed consent forms and the participants gave informed verbal consent to take part in the tasks. The experimenter explained the rules to the participant before each task and answered any questions arising from the description.","Two main methods of analyses were conducted on the data. First, hierarchical regressions were carried out, with chronological age and WISC-IV Perceptual Reasoning score entered as predictors in the first step, and Group (DCD vs. TD) in the second step. This meant that any group differences in EF performance revealed in the analyses would be evident even after differences or changes in performance related to age and Perceptual Reasoning had been taken into account. Chronological age was included to account for the improvement of EF ability within the relatively wide age range of the two groups. Perceptual Reasoning scores were included because the DCD group had significantly poorer scores than the TD group, and it was important to account for these differences before assessing the groups on EF performance. Second, within-group correlations were conducted to assess the relationship between motor abilities and EF performance within each group. There were six dependent variables in total: Tower task total raw score, Tower task violations, and Rotational Bar task total score; Stroop correct congruent responses, Stroop correct incongruent responses, and Knock-Tap total raw score. The means, standard deviations and ranges of these six variables are presented in Table 2. As some of these dependent variables were not normally distributed in one or both groups, non-parametric Spearman correlations were conducted on the data. For the regression analyses, bootstrapping procedures were applied, allowing an assessment to be made of the representativeness of the relatively small sample to the population from which it was drawn. Bootstrapping provides estimates of the confidence intervals around the regression coefficients, and relies on fewer assumptions about the distribution of the data and residuals than traditional statistical approaches, which is particularly useful in clinical samples (Wright, London, & Field, 2011). DCD vs. TD group differences ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Six separate hierarchical regressions were conducted on the data, with one planning or inhibition score as the dependent variable in each case. Chronological Age and Perceptual Reasoning scores were entered in Step 1, with Group entered in Step 2. Sections 3.1.1 and 3.1.2 report the summaries of the full models for each of the EF tasks, along with the unstandardised coefficients, standard errors (SEs) and confidence intervals (CIs) for each of the predictors in Step 2 of the regression, based on 1000 bootstrap samples. Planning tasks The final model significantly predicted the Tower task total raw score, F(3,46) = 26.82, p < .001, Adj R2 = .61. The coefficients were significant for Chronological [Age, B = 0.09 (CI = 0.07–0.13), SE B = 0.01, p = .001], Perceptual Reasoning [B = 0.10 (CI = 0.04–0.14), SE B = 0.02, p = .001], and Group [B = 3.26 (CI = 1.76–5.00), SE B = 0.85, p = .002]. For the number of violations in the Tower task, the final model was again significant, F(3,46) = 6.49, p = .001, Adj R2 = .25, although the only significant predictor was Group [B = −1.76 (CI = −2.99 to −0.47), SE B = 0.66, p = .02]. The coefficients for Chronological Age [B = −0.02 (CI = −0.04 to 0.002), SE B = 0.01, p = .11], and Perceptual Reasoning [B = −0.03 (CI = −0.06 to 0.01), SE B = 0.02, p = .12], were not significant. Finally, the final model predicted a significant amount of the variance in scores on the Rotational Bar task, F(3,46) = 15.74, p < .001, Adj R2 = .47, with Group again emerging as the only significant predictor of performance [B = 9.81 (CI = 6.70–12.55), SE B = 1.36, p < .001]. Chronological Age [B = 0.06 (CI = 0.01–0.11), SE B = 0.013, p = .053], and Perceptual Reasoning [B = −0.05 (CI = −0.14 to 0.03), SE B = 0.04, p = .18], were not significant. To summarise, the DCD group performed significantly worse than the TD group on all three measures of planning, even once the effects of Chronological Age and Perceptual Reasoning on planning performance had been taken into account. Inhibition tasks The final model significantly predicted the number of correct responses in the Stroop task on both the congruent trials, F(3,42) = 4.62, p = .01, Adj R2 = .19, and the incongruent trials, F(3,42) = 9.59, p < .001, Adj R2 = .36. For the congruent trials, none of the coefficients of the predictors were significant: Chronological Age [B = 0.21 (CI = 0.05–0.41), SE B = 0.09, p = .08]; Perceptual Reasoning [B = 0.13 (CI = −0.06 to 0.33), SE B = 0.10, p = .23], and Group [B = 7.65 (CI = 0.03–16.38), SE B = 4.48, p = .26]. For the incongruent trials, Perceptual Reasoning was not a significant predictor of correct responses [B = 0.17 (CI = −0.12 to 0.49), SE B = 0.15, p = .26], but both Chronological Age [B = 0.45 (CI = 0.22–0.70), SE B = 0.11, p < .001], and Group [B = 15.73 (CI = 3.03–27.82), SE B = 5.79, p = .01], had significant coefficients. Finally, the final model did not significantly predict performance on the Knock-Tap task, F(3,45) = 2.04, p = .12, Adj R2 = .06, and none of the coefficients were significant: Chronological Age [B = 0.02 (CI = −0.01 to 0.05), SE B = 0.01, p = .17]; Perceptual Reasoning [B = −0.02 (CI = −0.06 to 0.02), SE B = 0.02, p = .36]; Group [B = 1.31 (CI = 0.05–2.52), SE B = 0.65, p = .052]. In summary, the DCD group performed significantly worse than the TD group in the reduced motor-load task (the Stroop incongruent trials) but the difference in performance in the Knock-Tap task (high motor-load) was only a non-significant trend. Within-group relationships between motor and EF abilities ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As some scores were not normally distributed, Spearman's correlations were conducted on all data and were Bonferroni-corrected for multiple correlations (p < .004). Correlation analyses were conducted within each group between Movement ABC-2 Total Standard Score and the six outcome measures of EF ability. There were no significant correlations between MABC-2 Total Standard Scores and any of the EF measures in the DCD group (all rs < 0.43, p > .03), or in the TD group (all rs < 0.29, p > .16). Separate analyses were also conducted for each group using the component standard scores of the Movement ABC-2 (Manual Dexterity, Aiming and Catching, and Balance). None of the components were significantly correlated with any of the EF outcome measures in the DCD group (Manual Dexterity: all rs < 0.42, p > .03; Aiming and Catching: all rs < 0.43, p > .03; Balance: all rs < 0.31, p > .12), or in the TD group (Manual Dexterity: all rs < 0.26, p > .22; Aiming and Catching: all rs < 0.40, p > .05; Balance: all rs < 0.29, p > .16).","The current study aimed to assess the role of motor load on EF task performance in children with DCD. As in previous research, children with DCD had significantly lower scores than a control group on measures of planning, although some measures of inhibition did not differ significantly between groups. Contrary to predictions, the level of motor load did not seem to specifically affect performance on the tasks: the DCD group had significantly lower scores on both the high-motor and reduced-motor planning tasks than the TD group. In the measures of inhibition, the DCD group performed significantly worse than the TD group when there was a reduced motor-load, but any differences in performance in the high motor-load task did not reach significance. Within each group, motor ability did not correlate significantly with any of the EF tasks. The poorer performance in the DCD group across ‘motor planning’ and ‘executive planning’ tasks is at odds with the previous study by van Swieten et al. (2009), which reported no deficits in planning on the Tower task in a group of children with DCD. Although there was no control group comparison in that study, the performance of children with DCD was compared to standardised means, and all children scored within the typical range (between a scaled score of 7 and 13). In the current study, children with DCD not only had lower scores overall than the TD group, but many also scored outside this typical range of standard scores (65% had a scaled score of less than 7, while only 27% of children in the TD group scored below 7). It is difficult to assess how these differences between studies may have arisen, as little information is given about the children who completed the Tower task in the van Swieten et al. study, apart from their chronological age range. It is possible that the current group had lower IQs or more severe movement difficulties than those in the van Swieten et al. sample, resulting in more difficulties in the executive planning task. However, all children in the current sample had a Verbal IQ in what is considered to be the normal range (all were above 70), and differences in Nonverbal IQ (or Perceptual Reasoning in the current study) were taken into account in the analyses, so the children in the present DCD group did not represent a particularly low-functioning group. In addition, motor abilities did not correlate significantly with performance on the Tower task in either the DCD or TD groups, suggesting (as predicted) that any motor skills required to complete the task did not have a significant impact on performance. It may be that the Tower task involves multiple processes over and above planning ability (producing a less ‘pure’ planning measure: Miyake et al., 2000), and it is the complexity of the task that affects performance in the current DCD sample. The fact that Perceptual Reasoning scores were a significant predictor of Tower raw scores in the DCD group may lend some weight to this suggestion. It is important to note that Perceptual Reasoning scores did not significantly predict any of the other measures, suggesting that the other EF measures used in this study are tapping into different underlying processes and constructs to the Perceptual Reasoning tasks. Future studies could compare ‘purer’ measures of executive planning (e.g., a maze or sorting task) with the Tower task in order to examine this hypothesis further in children with DCD. The complexity of the task might also have been a factor in the inhibition performance of the DCD group. While children with DCD performed significantly worse than the TD group on the Stroop task during the incongruent trials, the difference between groups on the Knock-Tap task did not reach significance. It is possible that the high motor-load task actually taps into a simpler, more fundamental skill than the Stroop task, which requires reading a word, then coding (and ignoring) the colour in which the word is presented, and finally verbalising the word. The prepotent response might also be stronger for the Stroop task than the Knock-Tap task, as it involves the highly automatised skill of reading rather than a non-automatic pattern of motor responses that is built up over the task. Previous reports of poorer inhibition in children with DCD could therefore relate not just to the requirement for a motor response per se, but to the complexity of the motor response or the additional requirements of other complex processes (e.g., visuo-spatial processing skills), as well as the strength of the prepotent response to be inhibited. It is also important to remember that while there may be minimal differences between groups on basic motor inhibition tasks in terms of behavioural outcomes, the functional brain responses underlying this performance can be very different between TD and DCD groups (Querne et al., 2008). If even a simple task requires more effortful top-down control in children with DCD than their typically developing counterparts (Querne et al., 2008), this could explain why increasing complexity in a task could have a disproportionate effect on their performance. The complexity of the tasks in the current study is one issue that arises from the adoption of tasks from standardised batteries and experimental measures previously used in the literature. This procedure also meant that it was necessary to manipulate motor load across tasks with different materials and methods, rather than within one task. While a focus of future research might be to manipulate the motor loads within tasks, the current study was important in terms of assessing the motor demands of a range of tasks that are widely used in the literature. In addition, it will be important for future investigations to consider the relationships between age and performance on the different executive functioning measures. In the current data, the relatively wide age range was taken into account in the regressions, as splitting the sample into smaller age bands would have affected the power of the analyses. However, it will be of great interest to assess the developmental trajectories of the different executive functions in children with DCD, as some EFs may develop linearly with age, while others may show different stages of development, and it will be important to know if these patterns could be atypical or delayed in children with DCD. While previous research has reported significant correlations between components of the Movement ABC-2 and measures of response inhibition in normative samples (e.g., Livesey et al., 2006; Rigoli et al., 2012), no significant relationships were found between motor and EF abilities in either the TD or DCD group in the current study. As the development of EF is a protracted process (Anderson, 2002; Best & Miller, 2010), it is possible that its relationship with motor abilities changes over developmental time and that the varied ages of the samples across the studies may explain the different results (i.e., 5–6 years in Livesey et al., 12–16 years in Rigoli et al., 7–14 years in the current study). Interestingly, even the tasks classified as having a relatively high motor-load were not correlated with the Movement ABC-2 Total Score or its components. It is important to note that Rigoli et al. found a significant relationship between motor ability and the time taken to complete an inhibition task, whereas the current study assessed inhibition errors. However, there may be a more fundamental difference between studies in terms of the aspect of response inhibition being assessed. Livesey et al. reported that motor ability was more closely correlated with a modified Stroop task (which they argued measured ‘interference control’) than with a Stop-Signal task, which was regarded as a measure of inhibition of an ongoing response (cf. Nigg, 2000). In the current study, although the Knock-Tap task had a relatively increased motor- load in terms of the response required, differences found between it and the Stroop task in could be due to the fact that they were measuring different aspects of response inhibition. It will be important in future research to take this into account when selecting inhibition tasks to use with children with DCD, as there may be much closer links between neural pathways relating to motor and interference control than between motor control and the inhibition of an ongoing response (Livesey et al., 2006).","The current study aimed to assess the role of motor load on EF performance in children with DCD compared to a control group with typical motor development, adding to the relatively limited literature regarding EF performance in DCD compared to other neurodevelopmental disorders. The DCD group performed significantly worse than the TD group on measures of both planning and inhibition, but the effect of the motor load of the response in the task was not clear-cut. It seems that while this may have an effect on the EF performance of children with DCD, other factors such as the executive ‘purity’ of the tasks, their interaction with age and the different aspects of response inhibition that may be measured could also play an important role. Given the reported negative consequences of executive difficulties on quality of life and achievements (e.g., Altshuler et al., 2007; Biederman et al., 2004; Garcia-Villamisar & Hughes, 2007; Gilotty et al., 2002), it is important that methodological problems are addressed in order to improve our understanding of EF in children and adults with DCD. Investigations considering EF across development in DCD, which take into account performance on different tasks measuring the same EF construct and assessing a range of compound skills, will be vital in this research. The incorporation of parent- and self-report measures of the effects of EF difficulties on functioning outside the laboratory will also help to provide a clearer picture of EF in DCD."],["Spontaneous tool innovation to solve physical problems is difficult for young children. In three studies, we explored the effect of prior experience with tools on tool innovation in children aged 4–7 years (N = 299). We also gave children an experience more consistent with that experienced by corvids in similar studies to enable fairer cross-species comparisons. Children who had the opportunity to use a premade target tool in the task context during a warm-up phase were significantly more likely to innovate a tool to solve the problem on the test trial compared with children who had no such warm-up experience. Older children benefited from either using or merely seeing a premade target tool prior to a test trial requiring innovation. Younger children were helped by using a premade target tool. Seeing the tool helped younger children in some conditions. We conclude that spontaneous innovation of tools to solve physical problems is difficult for children. However, children from 4 years of age can innovate the means to solve the problem when they have had experience with the solution (visual or haptic exploration). Directions for future research are discussed. --------------------------------------------------------------------------------","Tool use is considered to be a hallmark of human cognition, with our substantial technological accomplishments unrivaled by any nonhuman animal species. Observations of infants and children demonstrate their impressive competence in terms of their tool use and knowledge. From around 2 years of age, children show insight into the function of tools (Casler & Kelemen, 2005) and anticipate the target of a tool use action (Paulus, Hunnius, & Bekkering, 2011). In addition, 3-year-olds reliably copy tool use from peers (Hopper, Flynn, Wood, & Whiten, 2010) and are capable of transmitting a tool use action across multiple generations (Flynn & Whiten, 2008; Hopper et al., 2010). It is surprising, then, that despite being proficient tool users, children appear to struggle to innovate tools (i.e., to make a novel tool to solve a task) without prior training (Beck, Apperly, Chappell, Guthrie, & Cutting, 2011; Cutting, Apperly, & Beck, 2011; Cutting, Apperly, Chappell, & Beck, 2014; Hanus, Mendes, Tennie, & Call, 2011; Nielsen, Tomaselli, Mushin, & Whiten, 2014; Sheridan, Konopasky, Kirkwood, & Defeyter, 2016; Tennie, Call, & Tomasello, 2009). Recent research into children’s tool innovation was motivated by studies with corvids. “Betty,” a captive New Caledonian crow, spontaneously manufactured a hook from a straight piece of wire to retrieve a wire-handled bucket from a transparent tube (Weir, Chappell, & Kacelnik, 2002). Researchers were investigating whether crows could choose the correct tool to solve a task and, on a task requiring a hook, had given them a hooked piece of wire and a straight piece of wire. When one crow flew off with the hooked tool, Betty bent the straight piece of wire to make her own hook despite not being shown how to make tools previously. In a later experiment, Betty continued to make functional hooks on the majority of trials when given only straight pieces of wire. New Caledonian crows have been directly observed making and using tools in the wild (Hunt & Gray, 2004; Rutz, Sugasawa, van der Wal, Klump, & St Clair, 2016), although not from wire or wire-like materials. However, similar findings have been replicated by Bird and Emery (2009) with captive rooks, a species that does not use tools in the wild. They were also able to select and manufacture the correct tool to solve the hook-making task used by Weir et al. (2002). Thus, it is even more surprising that children younger than 8 years demonstrate difficulty with similar problems requiring tool innovation. In a paradigm adapted from the corvid literature, children were able to choose the appropriate tool to solve a problem requiring a hook (Beck et al., 2011). From 4 years of age, children were significantly more likely to pick up a hooked pipe cleaner than to pick up a straight one when their goal was to retrieve a handled bucket containing a sticker reward from a tube. However, when other children were given a straight pipe cleaner that required bending into a hook shape to retrieve the bucket, children younger than 5 years rarely solved the task. Performance improved with age, and it was not until 8 years that approximately half of children passed the task. Interestingly, most children found manufacturing a tool (making a tool following adult demonstration) comparatively easy. This finding appears consistent across cultures. Western and Bushman children show a similar pattern of tool innovation performance; innovating a tool independently to solve a physical problem is difficult for children aged 3–5 years, whereas manufacturing a tool following an adult demonstration is significantly easier (Nielsen et al., 2014). Beck et al. (2011) noted that children’s knowledge of tool function and ability to manufacture tools emerges significantly earlier than their ability to innovate tools. Given the findings from tasks involving corvids, children’s difficulty with tool innovation has been met with curiosity. Given that many nonhuman species are known to use tools (Seed & Byrne, 2010), it is the human propensity for tool innovation, and the complex technologies that have arisen because of it, that sets us apart from nonhuman species. Findings such as these raise bigger questions surrounding human cognitive architecture. It is important to better understand those processes that we might share with nonhuman animals and those that may demonstrate human uniqueness (Shettleworth, 2012). Cross-species comparisons between human children and nonhuman animals are often made (e.g., Beck et al., 2011; Cheke, Loissel, & Clayton, 2012; Engelmann, Herrmann, & Tomasello, 2012; Taylor et al., 2014). However, caution is required. To truly understand how human children differ from other species, in this case corvids, it is vital that studies are methodologically sound and fair to both species (Boesch, 2007; Shettleworth, 2012). When experimental procedures systematically differ, the value of cross-species comparisons is compromised. To date, the experiences of children and corvids on versions of the hook-making task have not been consistent. Specifically, the experiences of the corvids and children prior to attempting the hook- making task were inconsistent. Corvids had already seen and had the opportunity to use a premade hook—made either from the same material as was available for tool making (wire; Weir et al., 2002) or from a different material (wood; Bird & Emery, 2009). Some children had not seen a pipe-cleaner hook previously within the task (Beck et al., 2011), and others had seen a hook but did not have the opportunity to use it (Cutting et al., 2014). The effect and potential advantage that such an experience might have on subsequent tool innovation in children is yet to be determined. Cutting et al. (2014) suggested that certain pretest experiences can promote tool innovation on the hook-making task. In their study, 4- to 6-year-olds were shown a premade example of a pipe-cleaner hook (target tool demonstration). Half of the children also had the pliable nature of the test material (pipe cleaners) highlighted to them via “bending practice.” It was found that 5- and 6-year-olds were significantly more likely to solve the hook-making task if they received both a target tool demonstration and bending practice compared with those who saw only a target tool demonstration. This suggests that seeing the correct tool required to solve a problem does not make tool innovation problems trivially easy for children. Unlike the corvids, children in these studies were not permitted to use the premade target tool prior to attempting the hook-making task. The purpose of the current series of experiments was twofold. First, we sought to further explore the effect of prior experience on children’s tool innovation. Second, we aimed to draw fairer cross-species comparisons of tool innovation between corvids and children. In three studies, the role of prior experience in children’s ability to innovate a hook tool to solve the hook-making task was investigated. In Study 1, we explored how children performed on the hook-making task given the same pretest conditions as corvids, that is, the opportunity to use a premade pipe-cleaner hook. Children, but not corvids, have previously been presented with nonfunctional distracter materials during the hook-making task such as string (Beck et al., 2011; Cutting et al., 2014). Therefore, no distracter items were included in these studies and children were presented only with pipe cleaners. Some corvids were also given multiple trials at the hook-making task, although performance was not observed to improve or deteriorate over time. Still, it seems important that the possibility of improved or changed performance over time is explored in children. Therefore, the first study included 3 trials to explore this possibility and to match more closely the experimental methodology of corvids. We did not give children the full 17 trials used in the original corvid study because we judged this as too many for children to cope with. Before attempting the hook-making task, half of the children completed a hook-using phase, where they were given the opportunity to use a premade pipe-cleaner hook on the bucket and tube apparatus. The hook-using phase emulates the condition used by Weir et al. (2002) where corvids were given a hooked piece of garden wire and a straight one.","Participants Participants The participants were 28 children aged 4 or 5 years (M = 4 years 6 months [4;6], range = 4;2–5;1) and 30 children aged 6 or 7 years (M = 6;7, range = 6;3–7;1) recruited from a mainstream school in the United Kingdom. The ethnic composition of the sample was 85% Caucasian, 10% Black, and 5% Asian. An additional 2 children were excluded from analysis due to retrieving the bucket without making a hook or other functional tool.","The apparatus was a transparent plastic tube (22 cm length, opening 5 cm in diameter) attached to a cardboard base (Fig. 1). At the bottom of the tube, there was a small bucket containing a sticker. The bucket had a wire handle that required a hook in order to retrieve it from the tube. Tool-making materials were 30-cm pipe cleaners. Procedure Procedure Procedure The study comprised a hook-using phase and an innovation phase. Children were systematically assigned to either the experimental condition or the control condition by their class list. They were told by their class teacher that they would be playing a game and must not discuss the game in the class because it would spoil the surprise for the other children. Children completed the experiment in a quiet area of the school library with a female experimenter. A second female coder (the second author) was present during 1 or 4 testing days to ensure reliability of coding success/failure on the hook-making task. There was 100% interobserver agreement. Children in the experimental condition first completed the hook-using phase. A 30-cm straight pipe cleaner and a 30-cm pipe cleaner bent at one end to form a hook were placed next to the apparatus in front of the children. They were told, “If you are able to get the sticker out of the tube, then you can keep it.” Children were allowed up to 1 min to complete this phase, and neutral prompts were given by the experimenter where necessary. All children completed the innovation phase. This followed the hook-using phase for those in the experimental condition and was the only phase for those in the control condition. During the innovation phase, children were presented with the same bucket/tube apparatus and a 30-cm straight pipe cleaner only. They were told, “If you can get the sticker out of the tube, then you can keep it.” All children received three trials. In the first trial, children were allowed up to 1 min to attempt the task. The experimenter then reset the task for two further 30-s trials. If children failed to innovate a hook after the third trial, the experimenter provided a demonstration of how to make a hook and kept it herself. Children were then encouraged to make another attempt to retrieve the sticker using their own materials. Only 2 children failed to make a hook following an adult demonstration. In these cases, children were given the experimenter’s premade hook and used this to retrieve the bucket.","Results and discussion Results and discussion Criteria for success on the innovation phase, here and in all further experiments, were making a hook tool and using it to retrieve the bucket from the tube. All children who made a hook were able to retrieve the bucket from the tube. There were no significant effects of gender on performance in Trial 1, 2, or 3, Fisher’s exact test, lowest p = .267. Therefore, data were combined across gender for all analyses. First, we analyzed whether the hook-using phase was successful in changing children’s experience in the experimental condition. In other words, did children use the hook prior to attempting the innovation phase? The 4- and 5-year-olds were significantly more likely than chance to pick up the hooked pipe cleaner first, with 15 of 15 children choosing the hook first, binomial test, p < .001. A binomial test revealed the same pattern for 6- and 7- year- olds, with 12 of 15 children choosing the hook first, p = .035. The 3 children who picked up the straight pipe cleaner first went on to use the hooked pipe cleaner afterward. Therefore, children in the hook-using phase did indeed use the hook before attempting innovation. Second, the effect of age group on performance during the innovation phase was analyzed. The 6- and 7-year-olds in the control condition were significantly more likely to pass the innovation phase than the 4- and 5-year-olds, Fisher’s exact test, p = .018. However, in the experimental condition there was no significant difference in performance related to age, Fisher’s exact test, p = .080. Of most interest was whether condition affected performance during the innovation phase and, therefore, the likelihood to innovate a functional hook tool. Because a significant effect of age was found in the control group, age groups were analyzed separately when comparing performance by condition. The percentage of children who passed each trial is shown in Table 1. Condition had a significant effect in Trial 1: 4- and 5-year-olds, Fisher’s exact test, p = .001; 6- and 7-year-olds, χ2Yates(df = 1, N = 30) = 7.350, p = .007. In Trial 2, condition had a significant effect on performance of 4- and 5-year-olds, Fisher’s exact test, p = .006, but the effect marginally failed to reach significance for 6- and 7-year-olds, χ2Yates(df = 1, N = 30) = 3.750, p = .053. In Trial 3, condition had a significant effect on performance of 6- and 7-year-olds, Fisher’s exact test, p = .002, but the effect marginally failed to reach significance for 4- and 5-year-olds, Fisher’s exact test, p = .055. However, considering the significant results across all other trials, this borderline finding is treated as if it were significant. To determine whether performance changed over time, McNemar tests were used to compare success levels between trials for each age group (Trial 1 vs. Trial 2, 4- and 5-year-olds, p > .999; 6- and 7-year-olds, p > .999; Trial 2 vs. Trial 3, 4- and 5-year-olds, p > .999; 6- and 7-year-olds, p = .500; Trial 1 vs. Trial 3, 4- and 5-year-olds, p > .999; 6- and 7-year-olds, p = .500). Therefore, children’s performance appeared to neither improve nor deteriorate over time. Finally, tool manufacture was easy for children. Of those children who failed to innovate a hook independently in any of the three test trials, 26 of 28 went on to successfully make a hook following an adult demonstration. The results from Study 1 demonstrated that children who had been given the opportunity to use a hook tool prior to being required to make one of their own were at a significant advantage over those children who had no such experience. Therefore, using a hook tool facilitated children’s subsequent tool innovation, so much so that younger children performed as well as older children. These results suggest that, like corvids, children are able to succeed on the hook-making task after having used a hook previously. Next, we aimed to disentangle what aspect of the warm-up phase was facilitating subsequent innovation. There are at least two obvious factors that could be independently or collaboratively facilitating innovation. Children first need to choose the hook as the correct tool and reject the straight pipe cleaner. Second, they use the hook on the hook-making task. We designed a second experiment to investigate whether choosing the hook and rejecting the straight pipe cleaner was an important element of the warm-up phase or whether simply using a premade pipe-cleaner hook was sufficient. It might be that being able to contrast two tools (a straight pipe cleaner and hooked one) helps children to identify what is functional about the target tool. We know that children from 4 years successfully choose a hooked tool over a straight one to solve this task (Beck et al., 2011). If this is the case, we expect children who choose a hook to outperform children who are just given a hook to use on the hook-making task. Alternatively, the additional demand of needing to choose between two tools may be more taxing for children (perhaps through working memory demands), and this may reduce performance on the subsequent hook-making task. Participants In total, 28 children aged 4 or 5 years (M = 4;9, range = 4;5–5;4) and 23 children aged 6 or 7 years (M = 6;10, range = 6;5–7;4) were recruited from the same school as in Study 1 and made up the final sample. None of these children had taken part in the first experiment. The ethnic composition of the sample was 78% Caucasian, 10% Black, 8% Asian, and 4% other/unknown. An additional 9 children were excluded from analysis due to retrieving the bucket without making a hook or other functional tool. Procedure All participants were presented with the plastic tube and bucket apparatus used in Study 1. Children were systematically allocated by their class list to one of two conditions. In the Hook Use condition, a premade pipe-cleaner hook only was placed next to the tube apparatus (Fig. 2). In the Tool Choice condition, a 30-cm straight pipe cleaner and a premade pipe-cleaner hook were placed next to the tube apparatus (Fig. 3). In both conditions, children were then given up to 1 min to try to retrieve the bucket from the tube. Having completed either the Hook Use or Tool Choice phase, all participants then completed the same test phase as used in Study 1. Because no difference in performance across trials was observed in Study 1, children were given only one attempt at the test phase. Results and discussion ~~~~~~~~~~~~~~~~~~~~~~ All children successfully used a premade hook tool in both conditions. There was no significant effect of gender on performance in either condition or age group, Fisher’s exact test, lowest p = .236. Therefore, data were combined across gender for subsequent analysis. Age group did not have a significant effect on performance during the test phase in either the Tool Choice condition, Fisher’s exact test, p > .999, or the Hook Use condition, χ2(df = 1, N = 26) = 1.330, p = .249. Finally, Table 2 shows the percentage of children who passed the test phase in each condition. Analyses were run to investigate whether condition affected innovation. First, an analysis was run for each age group separately. Condition did not affect performance on the test phase for younger children, Fisher’s exact test, p > .999, or for older children, Fisher’s exact test, p = .214. Because age did not affect performance in this sample, a chi-square analysis was also conducted with age groups combined; however, condition still showed no significant effect on innovation, χ2(df = 1, N = 51) = 1.797, p = .249. Study 2 sought to explore what it was about the warm-up phase in Study 1 that facilitated tool innovation. Regardless of whether they chose and used a pipe-cleaner hook or used a pipe-cleaner hook only during their first phase, children were just as likely to innovate a hook during the test phase. This suggests that it is using the hook, rather than the choosing element of the warm-up phase, that improves children’s tool innovation. In other words, rejecting the nonfunctional tool (straight pipe cleaner) during the choice phase does not seem to improve children’s likelihood of innovating a hook during the test phase. Rather, the opportunity to use a functional hook tool on the hook-making task seems to promote children’s tool innovation. The findings from Studies 1 and 2 are especially interesting because Cutting et al. (2014) found that showing children a pipe-cleaner hook was not sufficient to help them innovate one of their own afterward. In their study, Cutting and colleagues showed children a premade pipe-cleaner hook only if they failed to innovate a hook on the hook-making task and then allowed them another go at the task. Half of these children were given pipe- cleaner bending practice, whereas the other half were not. It was found that only older children, who were also given the opportunity to manipulate the materials beforehand, benefited from a target tool demonstration. Cutting et al. (2014) focused on children’s performance after different aspects of the hook-making task were highlighted to children as well as how they could retrieve and coordinate this information to solve the task. As such, the analysis focused on performance between conditions rather than stages. However, it is possible that for some children, seeing a target tool drove their subsequent innovation. This leaves open the question of whether seeing a target tool is as helpful as the opportunity to use a target tool in terms of likelihood of going on to solve the hook- making task. Children aged 3 or 4 years have been shown to demonstrate different strategies of exploration of materials for different tasks. For example, they used visual exploration to determine which spoons were the correct size to transport sweets, but they chose haptic exploration to decide which sticks, of varying rigidity, were suitable to stir sugar and gravel (Klatzky, Lederman, & Mankinen, 2005). It is yet to be made clear whether or not children find visual or haptic exploration of premade tools equally useful in the context of tool innovation or the hook-making task. If the key information that children extrapolate from a hook demonstration is regarding its shape, we would expect visual demonstration to be sufficient. However, if the information children require is regarding some other physical property of the pipe cleaner (e.g., flexibility, rigidity), we might expect that haptic exploration would be more beneficial. In Study 3, we sought to compare the effects of seeing versus using a premade hook on performance on the hook- making task. In Studies 1 and 2, children were given the opportunity to use a hook before attempting the task. Another factor is whether children benefit more from being given solution-relevant information before or after they encounter the problem to be solved. In Studies 1 and 2, children had only used a pipe-cleaner hook before they attempted the hook-making task, which proved to be beneficial to both younger and older children. In a previous study, children (and apes) were shown the location of tools they could choose to use to solve a problem either before or after they had seen the task they were going to solve (Martin-Ordas, Atance, & Call, 2014). Younger children found it easier to locate the tools to solve the problem after they had seen the task they would need to solve. These findings suggest that children are sensitive to the timing of relevant information when engaging in tool use tasks, and the same may be true of tool innovation tasks. Given the findings of Martin-Ordas et al. (2014), we might predict that younger children would benefit more from being given solution-relevant information after already attempting the hook-making task. Hence, in Study 3, we also manipulated the timing of the experience children were given (before or after having attempted the task). Participants The participants were 94 children aged 4 or 5 years (M = 4;7, range = 4;3–5;2) and 96 children aged 6 or 7 years (M = 6;7, range = 6;3–7;2) from two mainstream schools in the United Kingdom. The same proportion of children from each school were present in each age group. The ethnic composition of the sample was 89% Caucasian, 5% Black, and 6% Asian. None of these children had taken part in either of the previous two studies. Procedure Children were presented with the same bucket and tube apparatus from the previous experiments. As per the 2 × 2 design, children were systematically assigned by class list to one of four conditions: See Before, See After, Use Before, or Use After. All children completed the standard test phase (hook-making task) from the previous experiments and were given the same instructions. The experimenter told them, “If you can get the bucket out of the tube, then you can keep the sticker.” However, the experience that children had before and after attempting the test phase varied per condition. In the See Before condition, children were presented with the bucket and tube apparatus. The experimenter then said “Look at this” and showed them a premade pipe-cleaner hook. The experimenter then removed this from sight, said “Here is something that might help you,” and placed a straight pipe cleaner next to the apparatus. Children were then given 1 min to attempt the hook- making task. In the See After condition, children first attempted the hook-making task (pre-demonstration phase). Children who failed to innovate a pipe-cleaner hook and retrieve the bucket were encouraged to put down their materials. The experimenter then showed them a premade pipe-cleaner hook, said “Look at this,” and gave children a new straight pipe cleaner if necessary. Children then attempted the hook-making task for a second time. In the Use Before condition, children were presented with the bucket and tube apparatus and a premade pipe- cleaner hook. They were allowed up to 1 min to attempt the task with the materials given. All children successfully retrieved the bucket using the premade pipe- cleaner hook. Children then attempted the hook-making task. In the Use After condition, children first attempted the hook-making task (pre-demonstration phase). Children who failed to innovate a pipe-cleaner hook and retrieve the bucket were encouraged to put down their materials. Pipe cleaners were removed. The experimenter then gave children a premade pipe-cleaner hook and said “Here is something that might help you.” Children were encouraged to have another go at the task for up to 1 min. Once more, all children successfully retrieved the bucket using the premade pipe-cleaner hook. Children then attempted the hook-making task for a second time. As noted previously, if children failed to make a hook on their second attempt at the hook-making task (After conditions: Use After and See After) or their only attempt at the task (Before conditions: Use Before and See Before), the experimenter demonstrated how to make a pipe-cleaner hook and allowed them to have another go at the task. Results and discussion ~~~~~~~~~~~~~~~~~~~~~~ There was a significant overall effect of gender on performance in 4- and 5-year-olds, with girls performing better than boys, χ2Yates(df = 1, N = 94) = 4.115, p = .042, ϕ = 0.231. However, there was no significant effect of gender on performance in 6- and 7-year- olds, χ2Yates(df = 1, N = 96) = 0.686, p = .408. Because no previous effects of gender were observed on these tasks, and this difference was present in only one age group, the result is not discussed further here. Older children performed significantly better than younger children in the Use After condition, Fisher’s exact test, p = .030, and the See Before condition, χ2(df = 1, N = 47) = 6.139, p = .013, ϕ = 0.006. However, there were no significant differences in performance related to age in either the See After condition, χ2(df = 1, N = 48) = 2.521, p = .112, or the Use Before condition, χ2(df = 1, N = 48) = 1.137, p = .286. Of most interest was whether condition affected performance. Table 3 shows performance across conditions for each age group. Children in the After conditions who passed during the pre-demonstration phase were included in the analysis. This is to match children in the Before conditions who may have passed regardless of the opportunity to see or use a hook and could not be identified. In total, 8 children in the Use After condition (2 4- and 5-year-olds and 6 6- and 7-year-olds) passed during the pre- demonstration phase, and 10 children in the See After condition (2 4- and 5-year-olds and 8 6- and 7-year-olds) passed during the pre-demonstration phase. We first looked at whether the timing of when information relevant to the hook-making task was given affected performance on the hook-making task. There were no significant differences between performance in Before and After conditions in 4- and 5-year-olds, χ2Yates(df = 1, N = 94) = 0.692, p = .405, or in 6- and 7-year-olds, χ2(df = 1, N = 96) = 0.675, p = .411. Second, we looked at whether there was any difference in performance on the hook-making task related to whether children were in a Use or See condition. There were no significant differences in performance of 6- and 7-year-olds between the Use and See conditions, χ2Yates(df = 1, N = 96) = 1.875, p = .171. However, 4- and 5-year-olds performed significantly better in the Use conditions compared with the See conditions, χ2(df = 1, N = 94) = 4.326, p = .038, ϕ = −0.236. We then looked at what was driving the difference between these conditions in 4- and 5-year-olds using a Bonferroni-adjusted alpha level of .025 (.05/2). A continuity-corrected chi-square test indicated no significant difference in success between the Use After and See After conditions, χ2Yates(df = 1, N = 47) = 0.034, p = .853. However, children in the Use Before condition performed significantly better than those in the See Before condition, χ2Yates(df = 1, N = 47) = 6.139, p = .013, ϕ = −0.404. Study 3 revealed no clear indication overall that either using or seeing a target tool is more beneficial than the alternative. Likewise, the timing of clues or information relating to transformation from the start state to the end goal does not appear to be important. In other words, being given tool shape prior to attempting innovation is no more useful than being told the information having already had one failed attempt at the hook-making task. However, younger children performed significantly better in the Use Before condition compared with the See Before condition, being more than twice as likely to succeed in the Use Before condition. This suggests that the Use Before condition was particularly beneficial for 4- and 5-year-olds in terms of facilitating innovation of a hook tool on the hook-making task. This complements the findings from Studies 1 and 2, in which younger children performed at similar levels as older children when they had the opportunity to use a premade hook tool prior to attempting the hook- making task. This can also be related to the findings of Cutting et al. (2014), who concluded that 4- and 5-year-olds experienced less benefit from seeing a target tool than 5- and 6-year-olds. Older children performed well across conditions, with success levels greater than what might be expected given performance in the control condition of Study 1. The results suggest that using and seeing a premade target tool is useful for children of this age regardless of whether this is before or after first attempting innovation. Younger children also performed well across conditions when compared with performance in previous studies and in the control condition of Study 1. This suggests that some younger children were able to coordinate knowledge highlighted to them in order to innovate a hook tool.","Spontaneous tool innovation remains a difficult problem for many children aged below 8 years (e.g., Study 1 control group). Previous studies demonstrate that choosing the correct tool and tool manufacture following an adult demonstration is relatively easy for children from as early as 4 years (Beck et al., 2011). The current studies contribute to our understanding of what children find especially difficult about tool innovation. In Study 1, when children were given the opportunity to use a premade hook tool, making a hook of their own on subsequent trials was significantly more likely. Study 2 suggested that the using element of the pretest experience in Study 1, rather than choosing/rejecting tools, seemed to be what was driving better performance on the hook- making task. Using a premade hook made it easier to solve the innovation problem in the future. This suggests that fundamental to children’s success on the hook-making task was first interacting with the solution. This provides more evidence that the most difficult aspect of the hook-making task for children is bringing to mind the solution to the problem for themselves. However, it also demonstrates that children do not lack the understanding of how to transform the straight pipe cleaner to a functional hook tool without being explicitly shown how to do so in an action demonstration. Finally, Study 3 provided further insight into children’s performance on the hook-making task. Children across age groups and conditions performed better than expected on the hook-making task compared with previous studies (Beck et al., 2011; Cutting et al., 2014) and children in the control group of Study 1. Younger children were most successful when they could use a premade tool before their first innovation trial. For older children, there was no significant difference between trials. In two conditions (Use Before and See After), there was no significant difference in success levels between younger and older children. These results highlight the benefit that younger children experience by being given information about the target solution. The results also suggest that the same information (e.g., seeing a hook before attempting the task) is not equally helpful for all children. The results from Study 3 may reflect individual differences in children’s learning preferences. For example, Flynn, Turner, and Giraldeau (2016) recently suggested that 5-year-olds could choose a learning strategy (social or asocial) for themselves that was effective for them. Children could choose to either attempt a task for themselves or watch an experimenter attempt it first and then had their learning strategy choice either met or violated. Although children showed a strong preference overall to learn socially, unlike 3-year-olds, 5-year-olds were also efficient at solving the task under asocial conditions, when this strategy matched their indicated preference. In the context of Study 3, one might hypothesize that, similarly, children show preferences for the types of information or opportunity for exploration offered (visual or haptic) and the timing of this information (before or after attempting the task). In Study 3, we concluded that using or seeing a premade tool was beneficial for older children in terms of solving the hook- making task later, with no single condition being significantly better than another. For younger children, using a hook was more beneficial than seeing a hook before attempting the hook-making task, but this difference was not present when the information was presented after their first attempt at innovation. This finding provides some support to the claims of Cutting et al. (2014), who concluded that older children benefited from a target tool demonstration (alongside bending practice), whereas younger children did not. However, the focus of Cutting and colleagues was on performance between conditions rather than stages. On closer inspection, younger children in their study did appear to benefit from seeing a target tool demonstration, albeit not to the extent of the older children. Following a target tool demonstration, a further 23% (no bending practice) and 37% (bending practice) of children went on to innovate a hook tool of their own. In the current investigation, seeing a hook increased tool innovation from near floor (Study 1) to 30%–58% (see Table 3). This suggests that, for a significant number of 4- and 5-year- olds, seeing a premade hook facilitated their future hook tool innovation. Cutting et al. (2014) concluded that 4- and 5-year-olds struggle to innovate both the solution (a hook) and the means (bending the pipe cleaner) to solve the hook-making task. It was also concluded that although 5- and 6-year-olds were better at innovating the means to solve the task following a target tool demonstration and bending practice, spontaneous innovation of the hook tool solution remained difficult for the majority. Here, we suggest that Cutting and colleagues’ conclusions may have been slightly pessimistic regarding younger children’s capabilities. The majority of 4- and 5-year-olds in Studies 1 and 2 could innovate the means to solve the hook-making task, having had the chance to use (rather than just see) a premade tool. Younger children were also helped by seeing a target tool where performance was improved from near floor to 30% to 58% (see Table 3). We suggest that it is especially difficult for children aged 4–7 years to independently generate the solution to the problem—that they need a hook. The current series of experiments allows us to comment for the first time on how the performance of children on the hook-making task compares with that of the corvids. When children were faced with similar pretest experiences as that of the New Caledonian crow in the Weir et al. (2002) version of the task, they too were able to innovate a hook to solve the hook-making task. It would now be interesting to see how corvid species perform on tool tasks such as the hook-making task without prior exposure to or experience with the target tool. Rooks had a different pretest experience from that of children and crows (Bird & Emery, 2009). They were exposed to wooden hook tools prior to needing to make a wire hook of their own. In sum, they transferred the solution from the wooden hook tool and transferred this to a new material when they made a wire hook on subsequent trials. Considering recent findings, this seems remarkable. Beck et al. (2014) found that 4- to 7-year-olds do not transfer their knowledge of manufacturing a pipe-cleaner hook when they are later required to make a hook from dowel (or vice versa) to solve a similar task. In line with previous findings (Beck et al., 2011; Nielsen et al., 2014), we noted children’s difficulty with spontaneous innovation of a tool to solve the hook-making task. None of the 4- and 5-year-olds, and only a minority of the 6- and 7-year-olds, innovated a hook tool on their first attempt at the hook-making task without prior experience of a premade tool. Betty the crow demonstrated innovation of the means to solve the task but not innovation of the solution itself, due to her experience with a premade wire hook (Weir et al., 2002). However, spontaneous tool innovation has been reported in “Figaro” the cockatoo (Auersperg, Szabo, von Bayern, & Kacelnik, 2012). Given impressive reports of the abilities of certain nonhuman animal species, it remains intriguing that young children should experience such difficulty in innovating tools to solve problems independently. It is important to bear in mind that reports of tool innovation in nonhuman animal species are often individual cases. For example, although Figaro the cockatoo could innovate a tool, other individuals either failed to use tools at all or showed components of tool-making behavior, which may have been the result of shared social experience with Figaro. Although this does not make the achievements of Figaro any less impressive, this does resonate with the literature examining children’s propensity for tool innovation. It appears, at least in the case of younger children (4- to 7-year-olds), that some children have the capacity for innovation, whereas others do not. A recent study has suggested that individual differences in divergent and creative thinking are not associated with tool innovation (Beck, Williams, Cutting, Apperly, & Chappell, 2016). Future research efforts might explore whether the ability to innovate tools is related to other personal characteristics, personality traits, or a set of cognitive abilities that promote innovative behavior. Within the animal literature, it is noted that to unravel the potential underlying cognitive mechanisms that support innovative tool manufacture, it will be necessary to control the developmental histories and experiences of participants (Auersperg et al., 2012). This presents a challenge to the study of tool innovation in children who come from a wide range of backgrounds with varied experiences. When studying children’s performance on such tasks, it is important to consider the wider sociocultural context. Children at this age are likely to have received varied influence from parents and caregivers, including opportunity for exploration, feedback, and availability of objects. This is especially true of the youngest children, who have spent less time in full-time state education. These factors may greatly influence a child’s likelihood to display innovative behavior, both its onset and its frequency (Tomasello, 1999). Future studies may find ways in which we can manipulate experience with objects and materials to explore its effect on tool innovation. Holyoak, Junn, and Billman (1984) argued that analogical reasoning is an important mechanism for cognitive development because analogy permits the transfer of knowledge and information from domains that are understood to those that are not. We propose that the hook-making task may require children to employ analogical reasoning because they are required to make inferences about novel experiences, identify the relevant and useful information from target tool demonstrations, and transfer what they have learned to the innovation phase (for further discussion, see Beck et al., 2014). It may be that those children with superior analogical reasoning abilities find it easier to use the information highlighted to them to pass the hook-making task. Future studies may seek to investigate whether children’s analogical reasoning is related to their tool innovation ability. Because one of the main aims of these studies was to draw comparisons with the corvid literature, the current findings are restricted to the results from one innovation task (hook making). However, children’s difficulty with spontaneous innovation has been demonstrated on tool innovation tasks such as the Floating Object Task (Hanus et al., 2011; Nielsen, 2013) and the Loop Task (Tennie et al., 2009). Here, children found it especially difficult to innovate the solution to the hook-making task—that they require a hook to retrieve the bucket from the tube. However, when given the opportunity to use the solution to the hook-making task, they could innovate the means to solve the task without prior training; they transformed the straight pipe cleaner into a pipe-cleaner hook. It remains unclear why some children find tool innovation difficult, whereas others do not. A more complete picture of the developmental trajectory of tool innovation should be built, with further investigation into the levels of scaffolding required and conditions that might facilitate children’s independent innovation on a wide range of innovation tasks."],["This experimental study examined the effects of engaging on social media with attractive female peers on young adult women's body image. Participants were 118 female undergraduate students randomly assigned to one of two experimental conditions. Participants first completed a visual analogue scale measure of state body image and then either browsed and left a comment on the social media site of an attractive female peer (n = 56) or did the same with a family member (n = 62) and then completed a post-manipulation visual analogue scale measure of state body image. A 2 × 2 mixed analysis of variance showed a significant interaction between condition and time. Follow-up t-tests revealed that young adult women who engaged with an attractive peer on social media subsequently experienced an increase in negative body image (dependence-corrected d = 0.13), whereas those who engaged with a family member did not (dependence-corrected d = 0.02). The findings suggest that upward appearance comparisons on social media may promote increased body image concerns in young adult women. --------------------------------------------------------------------------------","Social media’s relation to body image is often examined using social comparison theory, which purports people self-evaluate based on comparisons with similar others. In upward social comparisons, people compare themselves to superior individuals (Festinger, 1954). Among women, making upward appearance comparisons is moderately related to negative body image (Myers & Crowther, 2009). On social media, young adult women most frequently make upward appearance comparisons to peers and rarely compare their appearance to family (Fardouly & Vartanian, 2015; Fardouly, Diedrichs, Vartanian, & Halliwell, 2015). Cross- sectional research shows negative associations between body image and active social media engagement (ASME), particularly photo-based ASME (Cohen, Newton-John, & Slater, 2017; Holland & Tiggemann, 2016; Kim & Chock, 2015; Meier & Gray, 2014). We consider ASME behaviours of viewing and commenting on friends’ social media. ASME requires content engagement and may have a greater impact on psychology than passive social media consumption. Among young adult women, ASME has a small, significant positive correlation with drive for thinness (Kim & Chock, 2015). Facebook photo activity has small-to-moderate positive correlations with thin ideal internalization, self-objectification, and drive for thinness, and a small negative correlation with weight satisfaction (Meier & Gray, 2014). Cross-sectional research suggests upward social comparisons to young adult women’s peers on social media weakly mediates the relationship between social media use and drive for thinness and body dissatisfaction, and that these comparisons have a stronger effect on body image concerns than do celebrity and model upward social comparisons (Fardouly & Vartanian, 2015). Frequency of appearance comparison to family is uncorrelated with social media use and body image (Fardouly & Vartanian, 2015). No published studies have shown causal effects of ASME with peers on body image. Women are more likely than men to use social media to view others’ photos, and ASME is how they typically use social media (Smith, 2014). They engage with social media specifically to compare themselves with others (Haferkamp, Eimler, Papadakis, & Kruck, 2012). Men are more likely to use social media to find friends (Haferkamp et al., 2012). Women feel worse about how they look than men (Engeln, 2017). On social media, young adult women present idealized images of themselves. Consequently, women on social media likely see idealized images of their peers and compare themselves with these idealized images (Manago, Graham, Greenfield, & Salimkhan, 2008); men are less likely to use social media like this. Thus, there is a need for research on potential causal effects of photo-based ASME on young adult women’s body image. This study investigated the effect of photo-based ASME with peers (as compared to family) on young adult women’s body image. Because women are more likely to struggle with body dissatisfaction (Mills, Roosen, & Vella-Zarb, 2011) and use image-based social media than are men (Greenwood, Perrin, & Duggan, 2016), we focused on young adult women. We hypothesized ASME with a female peer whom young adult women perceive as more attractive than themselves (upward social comparison) would result in more negative body image, whereas engaging with a woman unlikely to be an appearance comparison target (a non-peer family member not perceived as more attractive than themselves) would not affect body image.","The initial sample consisted of 143 female York University undergraduate students enrolled in an Introduction to Psychology course recruited through an online experiment management system for a study on “social media and relationships.” Eighteen participants were unable to adhere to experimental instructions and were excluded from all analyses, leaving a final sample of 125. Participants ranged in age from 17 to 27 years (M = 19.59, SD = 2.00) and in body mass index (BMI) from 15.43 to 42.26 kg/m2 (M = 24.31, SD = 5.00). The self- reported ethnicity of the sample was: 25.6% South Asian; 16.0% European; 14.4% Middle Eastern; 8.0% Caribbean; 4.0% Pacific Islander; 7.2% African; 4.8% Latin, Central, and South American; 4.8% East Asian; 1.6% African-American; 1.6% Indigenous; and 7.2% other. A few participants (4.8%) did not report their ethnicity. Demographics Participants reported their age, ethnicity, whether English was their first language, and years of post-secondary education in an online questionnaire 1.5 months before the experiment. Level of post-secondary education was self-reported by answering, “How many years have you been a university student?” English as a first language was self-reported by answering, “Is English your first language?” and, if not, by answering, “Do you consider yourself fluent in English?” State body image dissatisfaction Visual analogue scales (VAS) measured state overall appearance (VAS-OAD) and body dissatisfaction (VAS-BD) before and after ASME. Participants rated how dissatisfied they felt about their overall appearance and body by placing a vertical line on a 10-cm horizontal line. Responses were scored to the nearest millimeter, producing a 100-point scale. The range of responses was none to very much. Higher scores indicated greater state body or overall appearance dissatisfaction. Heinberg and Thompson (1995) constructed the VAS-BD and VAS-OAD measures, averaging them to form one body dissatisfaction measure. We refer to this averaged measure as the State Body Image Scale (SBIS). The pre-manipulation SBIS Spearman-Brown coefficient was .86 and the inter-item correlation was .79. The post-manipulation SBIS Spearman-Brown coefficient was .88 and the inter-item correlation was .78. Heinberg and Thompson (1995) demonstrated convergent validity of the VAS-BD and VAS-OAD with the Eating Disorder Inventory–Body Dissatisfaction subscale (Garner, Olmsted, & Polivy, 1983), r = .66, p < .01, r = .76, p < .01, respectively. Height and weight Participants were weighed and measured on a balance beam scale at the study’s end. ASME with a peer manipulation Participants in this condition (the Peer Condition) completed the pre-manipulation SBIS and then identified a female peer within five years of their own age whom she followed on social media and explicitly considered more attractive than herself (see Appendix A for full Peer Condition identification instructions). Participants reported the peer’s relationship to them and her initials. Next, participants viewed and commented on this peer’s social media pages for 5 min on Facebook and then 5 min on Instagram. Participants looked at only their identified peer’s social media, found a photo of her on Facebook, and left an online comment on this photo. Participants could engage with the peer’s other Facebook content during the 5-min task. After participants received instructions, the experimenter loaded www.facebook.com and left. The experimenter came back after 5 min and gave the participant similar instructions, asking her to do the same thing on Instagram. The experimenter then loaded www.instagram.com and left. Participants commented on photos of the same peer across Facebook and Instagram. Once the next 5 min were over, the experimenter came back into the room, confirmed the instructions had been followed, and asked the participant to complete a social media interaction questionnaire and the post-manipulation SBIS. ASME with family manipulation The procedure and instructions in this condition (the Family Condition) were identical to the Peer Condition, except participants engaged with a female family member’s social media. They looked at and commented on a non-peer, not-more- attractive family member’s social media posts for 5 min. per social media platform (i.e., Facebook and Instagram). Non-peer was defined as someone at least five years younger or older than the participant (see Appendix B). Social media interaction questionnaire Participants were asked whether they found a photo of the person they identified earlier on Facebook and Instagram and to write the comments they had left on her social media. Manipulation checks Participants confirmed their and their social media contact’s age, and their relationship to her. Participants were asked again during the debriefing if they followed the instructions. Browser history was checked after each participant completed the experiment. The social media interaction questionnaire also served as a manipulation check, as all participants wrote down the comments they left. As noted above, 18 participants could not identify a suitable contact and were excluded from the analyses.","The authors’ institutional human participants ethics review board approved this study. Testing was conducted from June 6, 2016 to June 13, 2017. Only female students with Facebook and Instagram accounts were eligible and received partial course credit for participating. Participants provided informed consent in the online questionnaire and then again right before the experiment. Participants were randomly assigned to condition using a coin toss. Each participant was seated alone in a private room; the experimenter only entered the room to explain instructions. Questionnaires were anonymous. Participants were weighed and debriefed in writing and verbally after the study. Statistical analyses The data were examined for normality, with no notable violations. Missing data accounted for 0.8% of the total values in the dataset. Both pre- and post-SBIS were missing 0.8% of data. Missing data were missing completely at random (MCAR), as determined by using Little (1988) MCAR analysis, χ2(2) = 1.35, p = .509. We inputted missing values using the multiple imputation expectation maximization technique, generating five imputations. Using G*Power for an a priori power analysis, for an analysis of variance (ANOVA) repeated measures within-between interaction, a Cohen (1988) f of 0.15 (i.e., a small-to-medium effect size), and a power estimate of 90%, the recommended sample size was 120. Additional participants were run to allow for any necessary exclusions. A 2 × 2 mixed ANOVA was conducted to compare the effect of ASME condition on SBIS scores pre- and post- manipulation. The independent variable was condition. Following Finch (2016) approach to address missing data, using the R base package (R Core Team, 2018), we pooled p-values. Results were considered significant at the α = .05 level, unless otherwise specified below. Analyses were performed using SPSS v. 24 and R. Randomization check ~~~~~~~~~~~~~~~~~~~ Preliminary analyses found no between-group differences in the distribution of ethnic groups, Fisher’s exact test p = .120; English as a first language, χ2(1) = 2.30, p = .129; number of years of post-secondary education, t(123) = 0.75, p = .409, d = 0.14; mean age, t(123) = 0.23, p = .819, d = 0.04; and mean BMI, t(123) = 1.06, p = .293, d = 0.19. Of those who reported English was not their first language (n = 31), the number of participants who considered themselves fluent in English did not differ by condition, χ2(1) = 1.35, p = .245. Thus, randomization resulted in equivalent groups on these variables. Manipulation checks ~~~~~~~~~~~~~~~~~~~ If a participant viewed a social media contact’s profile, the history tab showed an address including the contact’s username. History tab checks confirmed participants followed the instructions. After excluding the aforementioned 18 participants, all remaining participants (N = 125) indicated they found a photo of their contact on Facebook and Instagram on which to comment. Generally, social media “friends” are peers (West, Lewis, & Currie, 2009). To maximize ecological validity, we included participants who identified a best, close, or just “friend” (n = 32). The most commonly reported relationship in the Peer Condition was “friend” (n = 22). “Cousin” (n = 38) was the most commonly reported relationship in the Family Condition. Effects of experimental condition on state body image ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Despite randomization to condition and no observed pre-manipulation differences on sociodemographic variables or BMI, a t-test revealed an unexpected difference in pre-SBIS scores between conditions, t(123) = 2.74, p = .006, d = 0.49. We inspected the distribution of pre-SBIS scores to explain the pre-existing group difference and noted some visual outliers. We then omitted cases where the pre-SBIS was greater than or less than 2 SD above or below the mean. That resulted in excluding seven Peer Condition participants with very high pre-SBIS scores, and a sample of 118 for the main analyses. The difference in pre-SBIS scores between conditions was no longer significant, t(116) = 1.83, p = .070, d = 0.30. The time × condition interaction on SBIS was significant, F(1, 116) = 9.98, Wilks’ Λ = 0.92, p = .002, ηp2 = .08. Follow-up t-tests revealed there was no difference between pre-SBIS (M = 35.89, SD = 23.62) and post-SBIS (M = 35.14, SD = 22.99) in the Family Condition, t(61) = 0.77, p = .438, dependence-corrected d = 0.02. However, there was a significant difference between pre-SBIS and post-SBIS in the Peer Condition, t(55) = 3.33, p = .002, dependence-corrected d = 0.13. As predicted, in the Peer Condition only, participants reported worse state body image after the experimental manipulation (M = 48.70, SD = 22.11) than before the manipulation (M = 43.46, SD = 21.06). There was a significant difference in post-SBIS scores between conditions, t(116) = 2.83, p = .005, d = 0.59. The results revealed a significant main effect of time, F(1, 116) = 4.89, Wilks’ Λ = 0.96, p = .029, ηp2 = .04, and condition on SBIS, F(1, 116) = 6.09, p = .015, ηp2 = .05, but they were qualified by the interaction.","We hypothesized that young adult women who actively engaged with the image-based social media of attractive peers (upward social comparison targets) would have more negative body image than before doing so, whereas young adult women who engaged with the image-based social media of family (unlikely social comparison targets) would not. Results showed ASME with attractive peers’ appearance-based social media resulted in worsened body image in young adult women, whereas interacting with that of family had no effect on state body image, supporting our hypothesis. Our findings align with the recommendation that body image media literacy programs should highlight social media use, especially pressures associated with viewing images of others (Holland & Tiggemann, 2016), and peers in particular. However, similar to other social media and body image research findings, our effect size was small, and possibly negligible in real-world terms. Thus, these results should not be overstated. This study adds to literature showing young adult women’s body image is negatively affected by viewing attractive women’s photos on social media (Fardouly, Pinkus, & Vartanian, 2017; Haferkamp & Krämer, 2011; Kim & Park, 2016; Tamplin, McLean, & Paxton, 2018). It extends prior research by showing ASME with known, attractive female peers causes adverse effects on body image, but the same type of interaction with family does not have this effect. Our results were found across a racially heterogeneous sample. Allowing our participants to identify their own known contact may have allowed them to identify personal (e.g., racial) attractiveness concepts. Limitations and future directions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Focusing on short-term effects limited our study. Our female-only sample was drawn from one university. Our results are not generalizable to all young adult women and cannot generalize to men. Despite effective randomization in terms of sociodemographic variables and BMI, participants in the Peer Condition unexpectedly had more negative baseline body image scores than those in the Family Condition. Therefore, a degree of response bias was inadvertently introduced. Cautious interpretation of our findings is warranted. Future work should investigate whether ASME with peers affects young men’s body image and if certain individuals are more affected by ASME with peers.","This is the first study showing actively engaging with attractive peers’ social media causes worsened body image in young adult women. ASME with family not more attractive than oneself does not cause body image changes. It is common for young adult women to engage with peers’ image-based social media; this study shows this activity can cause a small increase in negative state body image.","We thank Anastasia Buchvostov and Dina Nikseresht for their assistance with data collection and entry, and Dr. Rob Cribbie for his assistance with data analysis. This research was supported by a SSHRC Insight Grant awarded to the second author."],["Delusions are defined as irrational beliefs that compromise good functioning. However, in the empirical literature, delusions have been found to have some psychological benefits. One proposal is that some delusions defuse negative emotions and protect one from low self-esteem by allowing motivational influences on belief formation. In this paper I focus on delusions that have been construed as playing a defensive function (. motivated delusions) and argue that some of their psychological benefits can convert into epistemic ones. Notwithstanding their epistemic costs, motivated delusions also have potential epistemic benefits for agents who have faced adversities, undergone physical or psychological trauma, or are subject to negative emotions and low self-esteem. To account for the epistemic status of motivated delusions, costly and beneficial at the same time, I introduce the notion of epistemic innocence. A delusion is epistemically innocent when adopting it delivers a significant epistemic benefit, and the benefit could not be attained if the delusion were not adopted. The analysis leads to a novel account of the status of delusions by inviting a reflection on the relationship between psychological and epistemic benefits. --------------------------------------------------------------------------------","In this paper, I ask whether delusions that have been construed as playing a defensive function have epistemic benefits. Defence mechanisms are “a means of nuancing or processing information such that it is rendered less anxiety-provoking” (McKay, Langdon, & Coltheart, 2005, p. 316). Arguably, delusions can prevent loss of self-esteem and help manage strong negative emotions (Butler, 2000; Ramachandran, 1996; Raskin and Sullivan, 1974; Bentall, 1994). The claim that delusions are psychologically adaptive can be made on these grounds, and it was recently discussed in the psychological literature (McKay & Dennett, 2009; McKay & Kinsbourne, 2010; McKay et al., 2005). Without denying that delusions are typically false and irrational, and that they compromise good functioning to a considerable extent, my goal here is to establish whether the psychological benefits attributed to those delusions that have been construed as playing a defensive function can translate into epistemic benefits. Thinking about delusions in terms of potential epistemic benefits leads to a more balanced view of the role of delusions in a person’s cognitive and affective life and invites a reflection on the relevance of contextual factors in epistemic evaluation. In Section 1, I review the general features of delusions and describe the epistemic features of delusions that can be construed as playing a defensive function (hereafter, motivated delusions) and their adverse effects on functioning. In Section 2, I consider arguments for the psychological benefits of Reverse Othello syndrome, erotomania and anosognosia. I also describe the proposal by McKay and Dennett (2009), according to which some false beliefs (adaptive misbeliefs) are the result of a mechanism that allows motivational factors to influence belief formation. Can motivated delusions be adaptive misbeliefs? In Section 3, I introduce the notion of epistemic innocence. Cognitions are epistemically innocent when, despite their epistemic costs, they carry a significant epistemic benefit (Epistemic Benefit condition) that could not be attained otherwise (No Alternatives condition). I argue that motivated delusions have the potential for satisfying the two conditions for epistemic innocence and I offer an illustration from anosognosia to support this claim. In Section 4, I suggest that the epistemic innocence potential of motivated delusions highlights the need for a more nuanced evaluation of epistemically costly cognitions and invites a new way of understanding the relationship between psychological and epistemic benefits.","Clinical delusions are symptoms of psychiatric disorders such as schizophrenia, dementia, and delusional disorders. Delusions exemplify failures of rationality and are defined on the basis of surface features that have an epistemic character. Here are some popular definitions: A false belief based on incorrect inference about external reality that is firmly held despite what almost everyone else believes and despite what constitutes incontrovertible and obvious proof or evidence to the contrary. The belief is not ordinarily accepted by other members of the person’s culture or subculture (i.e., it is not an article of religious faith). When a false belief involves a value judgment, it is regarded as a delusion only when the judgment is so extreme as to defy credibility. A person is deluded when they have come to hold a particular belief with a degree of firmness that is both utterly unwarranted by the evidence at hand, and that jeopardises their day-to-day functioning. Delusions are generally accepted to be beliefs which (a) are held with great conviction; (b) defy rational counter-argument; (c) and would be dismissed as false or bizarre by members of the same socio-cultural group. The definitions above characterise delusions on the basis of their epistemic features, including lack of warrant, fixity, resistance to counterargument, and implausibility. Not all delusions manifest such features to the same extent, and delusions may differ in the way they interact with the person’s other cognitive or affective states. Types of delusions ~~~~~~~~~~~~~~~~~~ Davies, Coltheart, Langdon, and Breen (2001) helpfully distinguish between circumscribed and elaborated delusions. Circumscribed delusions are not well integrated with the other beliefs the person has, and the epistemic features that characterise these delusions do not necessarily “spread” to the rest of the person’s belief system. Some of these delusions, so-called “deficit” delusions (McKay & Dennett, 2009), are the result of brain damage or cognitive deterioration, and, even if they were given a psychodynamic interpretation in the past, there is now little room for motivational factors in an account of their formation. Examples are the Capgras delusion (the belief that a loved one has been replaced by an impostor) and mirrored-self misidentification (the belief that there is a stranger in the mirror when one looks at one’s own reflection). Other circumscribed delusions have been construed as playing a defensive function, and motivational factors are sometimes advocated in the explanation of their formation. Such delusions often follow trauma. Examples are the Reverse Othello syndrome (the belief that one’s romantic partner is faithful when she is not) and anosognosia (the denial of illness, for instance the denial that one’s limb is paralysed). Delusions emerging in the context of schizophrenia can be systematised and elaborated. They can turn into complex narratives used to explain most of the person’s experience. Examples are the delusion of persecution (the belief that others are threatening and intend to cause harm), the delusion of grandeur (the exaggerated belief in one’s self-worth), and the delusion of reference (the belief that some events are highly significant when they are not). A popular hypothesis is that delusions in schizophrenia are offered as an explanation for the person’s hypersalient experience. Given that hypersalient experiences cause anxiety and distress in the prodromal phase of psychosis, delusions emerge as hypotheses by which the person makes sense of their experiences (Jaspers, 1963; Kapur, 2003; Mishara and Corlett, 2009). Delusions in schizophrenia can put an end to a state of uncertainty that causes anxiety and distress. Sometimes the “need for closure” is discussed in this context (McKay & Kinsbourne, 2010) and it indicates a preference for certainty over uncertainty and for predictability over unpredictability. In addition to satisfying the need to have an explanation as opposed to none, it has been argued that some delusions in schizophrenia can be motivated due to their specific content. Among others, delusions of grandeur and delusions of persecution seem to protect the person from a negative conception of the self and from low self-esteem. In terms of aetiology, it is plausible that a combination of neurobiological and psycho-social factors (including motivational factors) contribute to the formation of delusions (Bentall, Kinderman, & Kaney, 1994; Davies, 2009; McKay & Kinsbourne, 2010; Roberts, 1992). For instance, Aimola Davies and Davies (2009) argue that a two-factor theory of delusion formation can make sense of most delusions, where the first factor explains where the delusion comes from, and the second factor explains why the delusion is not rejected. The first factor usually consists in an anomalous experience or a neuropsychological deficit. The second factor consists in an impairment of belief evaluation, that is, a problem with the assessment of the evidence for and against the delusional belief. Such a problem may be caused by cognitive impairments, but may also be caused by a motivationally-biased handling of the evidence. When no anomalous experience or deficit can be identified, it is possible that motivation constitutes the first factor, in which case the delusion would emerge as a defence mechanism (McKay et al., 2005). Here I do not propose to discuss the role of motivation in delusion formation, rather I am interested in evidence for the claim that delusions have psychological benefits, and such evidence is often reviewed in the discussion of different delusion formation theories. For convenience and due to space limitations, I will concentrate on those delusions that have been explicitly construed as defence mechanisms in the psychological literature, but I make no assumption about such accounts being the best explanations of how the delusions are formed. For some of the delusions I shall describe, it is possible that a defence mechanism explains the origin of the delusion and its content. For all of these delusions, it is plausible that motivational factors play some role in the maintenance of the delusions. But motivational factors need not play any role in delusion formation for delusions to have psychological benefits. What is wrong with motivated delusions? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Let me describe some of the characteristics of motivated delusions that amount to epistemic costs and can also lead to impaired functioning. Motivated delusions can be characterised as irrational beliefs, in that they are implausible, they do not accurately represent reality, they do not respond to evidence, and they may not always be consistently reflected in behaviour (Bortolotti, 2009). For instance, people with motivated delusions may be convinced of the truth of the delusion but at the same time exhibit some “covert recognition” that the content of the delusion is false. When they are circumscribed, motivated delusions may conflict with other beliefs the person has, contributing to an overall inconsistent set of beliefs. When they are elaborated and systematised, they are likely to integrate well with a number of other beliefs that also play a defensive function. I will consider three cases of motivated delusions. An example of a monothematic delusion with a defensive function that emerged as a result of brain damage is the case of Reverse Othello syndrome (Butler, 2000) discussed in some detail by McKay et al. (2005). A man, BX, delusionally believed that he was in a happy relationship, when in fact his partner had left him. Butler’s patient was a talented musician who had sustained severe head injuries in a car accident. The accident left him quadriplegic, unable to speak without reliance on an electronic communicator. One year after his injury, the patient developed a delusional system that revolved around the continuing fidelity of his partner (who had in fact severed all contact with him soon after his accident). The patient became convinced that he and his former partner had recently married, and he was eager to persuade others that he now felt sexually fulfilled. BX’s belief in the fidelity of his previous partner and the continued success of his relationship was very resistant to counterevidence. BX believed that his relationship was going from strength to strength for a few months, even though his former partner did not want to communicate with him and was in a relationship with someone else (Butler, 2000, p. 86). The Reverse Othello syndrome can be seen as a special case of erotomania. In erotomania, a person comes to believe that another person, often of a perceived higher status (e.g., a teacher, an older or more successful person, a celebrity), is in love with her when there is no evidence in support of that belief. Here is the case of a young woman, LT, who started behaving strangely when she became obsessed with the idea that a fellow student was in love with her although the two had never spoken to each other. Her conversation, when unrelated to her delusional process, was rational, coherent, appropriate, and relevant. [ …] When speaking of the delusional process, she went into great detail, explaining the messages she received from her fantasied lover, signs which she received on TV, from the colors of dresses, license plates on cars, and from several other sources. She saw all of this as proof of the fact that the young man was in love with her and was planning to marry her. Numerous attempts to offer LT evidence that her belief was false failed: after two and a half years after the delusion emerged, the alleged lover was asked to talk to LT on the phone, following a suggestion by LT’s mother. He told LT that he did not have any intention to marry her and that he could barely remember who she was. But LT was convinced that her mother had arranged for her to talk to another man and did not abandon her delusion. Motivated delusions are also found in the context of anosognosia, the denial of illness (most commonly, the denial that a limb is paralysed). Delusions take the form of: “I am moving my arm”, when the arm cannot move; or “I can climb stairs but I am a little slow” when a leg is paralysed. Anosognosia has been considered as a pathology of belief: “There is a mismatch between the patient’s estimate of his or her abilities and the reality of the impairment” (Aimola Davies, Davies, Ogden, Smithson, & White, 2009, p. 188). In anosognosia, people deny evidence supporting the fact that they are impaired. Here is the case of a patient with anosognosia for hemiplegia: [A]sked to clap the hands, [she] lifted her right hand and put it in the position of clapping, perfectly aligned with the trunk midline, moving it as if it was clapped against the left hand. She appeared perfectly satisfied with her performance, never admitting that the left arm did not participate in the action. This despite the fact that the patient could see that the left hand did not clap against the right hand and the typical sound of clapping was not heard. It is not clear to what extent people with anosognosia are unaware of their impairment, given that, on occasion, they seem to implicitly acknowledge it. For instance, a person with a paralysed leg might deny paralysis but at the same time acknowledge that she cannot climb stairs properly. She might even provide a confabulatory explanation of her poor performance, saying that it is due to tiredness or to arthritis (Ramachandran, 1995, p. 23). The delusions I have described are not just implausible and irresponsive to evidence but they have an adverse effect on wellbeing and interfere with interpersonal relationships. The person reporting the delusion stops being regarded as a trustworthy source of information about the topic of the delusion and may be socially sanctioned or excluded for that reason. Relationships with family members and with healthcare professionals may be strained as a result of the absence of a “shared reality” (Fotopoulou, 2008, p. 546) and common goals. In many cases of erotomania (Lovett-Doust & Christie, 1978), the person develops an obsession with the delusional theme, loses interest in her family and friends, gives up her daily activities, and becomes isolated as a result. Anosognosia can have negative effects on people’s health. For instance, it can interfere with therapy and rehabilitation (Fotopoulou, 2008, p. 554). If the person does not acknowledge an impairment due to a recent trauma, she will not understand the need to engage in rehabilitation; indeed it has been found that anosognosia is “an inhibitory factor hampering rehabilitation” (Maeshima et al., 1997, p. 691).","Considerations about the epistemic costs of delusions and their adverse effects on functioning seem to rule out the possibility that delusions have any benefits. However, it has been argued that the adoption of a motivated delusion helps manage overwhelmingly negative emotions that would otherwise lead to depression and protects against negative self-conceptions that would otherwise lead to low self-esteem. In the light of these arguments, the relationship between delusions and wellbeing appears more complex than one might have expected. For instance, Lansky (1977, p. 21) writes that delusion “is restitutive, ameliorating anxieties by altering the construction of reality”. Motivated delusions in context ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the case of Reverse Othello syndrome I mentioned earlier, the delusion seemed to protect BX from the undesirable truth that his romantic partner had left him while he was coping with the consequences of permanent disability. Butler, who first reported on this case, describes the period preceding the report of the delusion: [H]is communicative responses initially indicated considerable insight and the beginning of an intense emotional response to a massive disability and a fracturing of his interpersonal relationships. Gradually, in the year following his injury, BX developed the delusion that his former romantic partner was still in a successful relationship with him, and also that they had recently married. While still in hospital, he often asked to go home so that he could see his wife. Butler argues that the delusion relieved the sense of loss that BX was feeling at the time. [A]ppearance [of delusions] may mark an adaptive attempt to regain intrapsychic coherence and to confer meaning on otherwise catastrophic loss or emptiness. As gradually as it had appeared, BX’s delusional system dissolved, and by the end of the process BX realised that his former partner had moved on, was not married to him, and had no intention to go back to him. No other delusions or psychotic symptoms were observed. This happened roughly at the time when BX had completed his physical rehabilitation and was ready to return home. Butler does not advocate a purely motivational explanation for BX’s delusion, but simply argues that different factors might have been at play, and that a psychological defence against depression contributed to the fixity and elaboration of BX’s delusional system. Independent of the role of motivational factors in the formation or maintenance of delusions, Butler makes a case for the presence of psychological benefits. The delusion kept BX’s depression at bay at a very critical time. Acknowledging the end of his romantic relationship might have been disastrous at a time when he was coping with the realisation of his new disability and its effects on his life. Erotomania more generally has also been described as an adaptive response (Raskin & Sullivan, 1974)), and it is significant that it often develops following a physical/psychological trauma or adverse circumstances. In many of the case studies of erotomania reported in the literature, the delusion seems to affect people who have a history of depression and feel lonely or under-appreciated (Hollender & Callahan, 1975). Consider LT, the young woman who developed erotomania and whose case was reported by Jordan and Howe (1980). Her background fits with the general profile described: The patient has a twin sister, who at that time was a junior in college, and a younger sister, who was two years old at the time of the onset of this disorder. In describing L.T.’s twin sister, Mrs. T. indicated that she was outgoing, friendly, and though somewhat reserved, maintained close relationships with people of both sexes. The mother stated that the patient in contrast had always been a quiet and rather inhibited child. She was much more reserved than her popular sister and dated infrequently. She was also described as being studious, an avid reader, highly moralistic, and a loner. Moreover, she tended to be somewhat suspicious and mistrustful. Her limited heterosexual experiences were characterized as being very short-lived. According to the mother, one such relationship had just recently ended abruptly and she related that the patient appeared rather emotionally distraught by this. In many cases of erotomania (Lovett-Doust & Christie, 1978, p. 105), there are identifiable “pharmacological, metabolic, and physiological and structural causes” (including injection of cortisone, alcoholism, meningioma, ingestion of contraceptive pills), but also “psychological and situational triggers”. People experience loneliness and loss prior to adopting the delusion, and in some cases the emergence of the delusion follows a traumatic event (e.g., the discovery of one’s partner’s infidelity, the death of a loved one, or the birth of a child then given up for adoption). This seems to suggest that the delusion plays a defensive function in that it compensates for loss or protects one from low self- esteem. In anosognosia, the connection between trauma and delusion is more explicit. The person refuses to acknowledge a serious impairment as a result of trauma or illness and, also often fails to recognise its implications (although the denial of the impairment and the failure to acknowledge its implications can dissociate, see Aimola Davies et al., 2009). Delusions occurring in anosognosia can be seen as playing a defensive function, but purely motivational accounts of anosognosia have been strongly criticised for failing to account for the fact that anosognosia is much more likely to emerge when the right parietal lobe is damaged, and that the denial seems to be “domain-specific”: the person may deny one impairment and acknowledge another. Ramachandran describes a patient who would go to great lengths to deny the paralysis of her limb but happily admitted to having diabetes (Ramachandran, 1995, pp. 23–24). To explain these phenomena, popular accounts of anosognosia have attempted to combine neuropsychological and motivational factors (Aimola Davies et al., 2009; Ramachandran, 1996). Ramachandran advances the hypothesis that the behaviours that give rise to delusions in this context are an exaggeration of normal defence mechanisms that have an adaptive function. Denying change can sometimes be instrumental to preserving a coherent system of beliefs and behaving in a stable and predictable manner (Ramachandran, 1996). The psychological advantages are not necessarily cashed out in terms of the preservation of the concept of the self as healthy, but in terms of the preservation of the concept of the present self as coherent with that of the past self. Fotopoulou observes the same phenomenon in people with memory impairments and anosognosia who do not seem to update personal information: [P]atients may need to highlight their continuity and coherence with their past selves and may not be able to understand or deal with the loss of their previous family and social role. Aimola Davies and Davies (2009) report positive and negative effects of anosognosia on wellbeing, suggesting that there could be a role for motivational factors in the explanation of anosognosia. As I mentioned in Section 1.2, anosognosia has negative effects, as people who do not acknowledge the illness or impairment may be slow in seeking treatment and unmotivated to engage in rehabilitation. But after the initial stages of illness, anosognosia is associated with fewer negative emotions and reduced anxiety. Delusions as a shear pin ~~~~~~~~~~~~~~~~~~~~~~~~ According to the “shear-pin” account developed by McKay and Dennett (2009), some false beliefs that help manage negative emotions and avoid low self-esteem and depression can count as psychologically adaptive. McKay and Dennett suggest that, in situations of extreme stress, motivational influences are allowed to intervene in the process of belief evaluation, causing a breakage. Although the breakage is bad news epistemically, as the result is that people come to believe what they desire to be true and not what they have evidence for, it is not an evolutionary “mistake”, rather it is designed to avoid breakages that would have worse consequences for the person’s self-esteem and wellbeing. What might count as a doxastic analogue of shear pin breakage? We envision doxastic shear pins as components of belief evaluation machinery that are “designed” to break in situations of extreme psychological stress (analogous to the mechanical overload that breaks a shear pin or the power surge that blows a fuse). Perhaps the normal function (both normatively and statistically construed) of such components would be to constrain the influence of motivational processes on belief formation. Breakage of such components, therefore, might permit the formation and maintenance of comforting misbeliefs – beliefs that would ordinarily be rejected as ungrounded, but that would facilitate the negotiation of overwhelming circumstances (perhaps by enabling the management of powerful negative emotions) and that would thus be adaptive in such extraordinary circumstances. Could motivated delusions be adaptive misbeliefs? The mechanism that inhibits motivational influences on belief evaluation would be compromised, and as a result of this motivated delusions would emerge, making negative emotions easier to manage and depression less likely to ensue. McKay and Dennett consider the possibility that some delusions count as adaptive misbeliefs, but interestingly argue that the extent to which desires are allowed to influence belief formation in the case of delusions is pathological. Delusions are the result of the maladaptive version of a psychologically adaptive mechanism. Delusions may be produced by extreme versions of systems that have evolved in accordance with error management principles, that is, evolved so as to exploit recurrent cost asymmetries. As extreme versions, however, there is every chance that such systems manage errors in a maladaptive fashion. More needs to be said about the precise nature of the advantage that adaptive misbeliefs may have, but shear-pin accounts are helpful in providing a framework for the potential psychological benefits of motivated delusions. First, in the shear-pin account, the situation in which adaptive misbeliefs emerge is already seriously compromised. The premise is that the person is already experiencing high levels of distress, and can come to more serious harm unless her negative emotions are managed. Thus, the benefit here amounts to the prevention of more serious harm than the one the person is already experiencing. In other words, the adaptive misbelief is equivalent to an emergency response. McKay and Dennett talk about the “extraordinary circumstances” in which motivational influences on belief are not just tolerated but desirable, and argue that such influences are not accidental but designed. I am going to suggest that a careful consideration of the circumstances in which beliefs are adopted should play a role in establishing not just whether they have psychological benefits, but also what their epistemic status is. Second, the phrase “adaptive misbelief” and the general description of the shear-pin mechanism may be taken to suggest that there is almost an inverse correlation between psychological and epistemic benefits. The more distant the belief is from a bleak reality, the more psychologically adaptive it is. This is not, however, what McKay and Dennett have in mind. Indeed, a possible explanation for the difference between motivated delusions and non-delusional adaptive misbeliefs is that delusions are ultimately maladaptive because, in the case of delusions, motivational influences affect beliefs to an extent that compromises their overall plausibility and makes them impervious to counterevidence. These epistemic costs are likely to bring also psychological costs, and thus delusions may turn out to have greater psychological costs than benefits. I am going to suggest that we should resist a trade-off view of the relationship between psychological and epistemic benefits. There are reasons to believe that it is psychologically beneficial to have beliefs that are constrained by reality, and that managing negative emotions, relieving anxiety, and protecting self-esteem have epistemically positive consequences.","What is it to be epistemically innocent? Innocence is sometimes cashed out in terms of a person being ‘free from faults’ or ‘free from sins’. This is not the sense of innocence I am advocating here. Ideally, agents would have beliefs that are true and that are supported by, and responsive to, the evidence available to them. But human agents have limited cognitive capacities, and beliefs that are false and badly supported by, or irresponsive to, the evidence are a common occurrence. It is tempting to dismiss epistemically costly cognitions altogether. But sometimes an epistemically costly cognition can also have positive epistemic features. When these epistemic benefits are significant and could not be attained in other ways, then the cognition may gain some sort of innocence. The notion of innocence I have in mind is used in the legal context of justification defence. In general, an innocence defence applies to someone who is not deemed liable for an act that appears to be wrongful. Innocence defence can be due to excuse or justification. An excuse defence applies when there is no criminal intent. A justification defence applies when the act does not constitute an offence in the given circumstances because it prevents greater harm from occurring. Here is an example of a justification defence: Ann injures Ben and by doing so she stops him from detonating a bomb (Greenawalt, 1986, p. 89). It is morally and legally objectionable to cause injury to another, but in the specific circumstances in which Ann injures Ben her action has some significant benefits as it prevents a greater harm from happening. Moreover, other ways of stopping Ben, such as talking him out of detonating the bomb, may be not available at the time. Ann is acquitted because what she did is not wrongful: it is an acceptable response to an emergency. My purpose here is to apply this notion of innocence to the domain of epistemic evaluation. In some contexts, an epistemically costly cognition (say, a false belief) may help avoid bad epistemic consequences, and thus qualifies as an acceptable response to an emergency.1 For instance, a delusion is epistemically innocent if adopting it delivers a significant epistemic benefit that could not be obtained otherwise. I propose the following two conditions for epistemic innocence: Epistemic Benefit: The delusional belief confers a significant epistemic benefit to an agent at the time of its adoption. No Alternatives: Other beliefs that would confer the same benefit are not available to that agent at that time.2 The exact formulation of the conditions for epistemic innocence will vary depending on one’s epistemological commitments. Let us consider the Epistemic Benefit condition first. The epistemic benefits of having a cognition may include maximising the acquisition and retention of true beliefs (for a veritist), promoting intellectual virtues (for a virtue epistemologist), or avoiding epistemic blame (for a deontologist). In line with a broadly consequentialist understanding of epistemic value, I shall argue in Section 3.1 that the adoption of a delusional belief can support the agent’s epistemic functionality that would otherwise be compromised by overwhelming negative emotions and low self-esteem. I take epistemic functionality to be the capacity to perform well epistemically by acquiring true beliefs/knowledge or exercising intellectual virtues. Now let us consider the No Alternatives condition. Different notions and degrees of unavailability can explain the failure to adopt a less epistemically costly cognition. This variety may be due to the nature of the limitations that the agent experiences in the relevant context, ranging from standard reasoning limitations to deficits affecting perception, inference, or memory in clinical settings (see Sullivan-Bissett (2015) for a useful taxonomy of relevant types of unavailability). In short, there may be no alternative to adopting a delusional belief because evidence that would lend support to less epistemically costly beliefs is not available or cannot be weighed up due to a bias or deficit. I shall review these options in Section 3.2. Meeting the Epistemic Benefit condition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The epistemic benefits of motivated delusions are mediated by their psychological benefits. We saw in Section 2 that the adoption of a delusional belief can be psychologically adaptive. In the shear-pin account of the benefits of delusions, the very fact that people believe a more positive version of reality (e.g., “I am now severely disabled, but my girlfriend still loves me”) than the one they have evidence for allows them to manage negative feelings that could become overwhelming, preserve self-esteem, and overcome anxiety and stress. The benefits in question amount to the delusion preventing a serious epistemic harm from occurring. McKay and Dennett focus on the effects of adaptive misbeliefs on wellbeing. The point of allowing motivational factors to influence belief evaluation is to make the person feel better about herself and her situation. But if a belief helps manage negative emotions, protect self-esteem, and relieve anxiety and stress, it will have positive effects not just on the agent’s wellbeing but also on her capacity to function well epistemically (what I called “epistemic functionality”). By having the belief, a person will be more likely to engage with her surrounding physical and social environment in a way that is conducive to epistemic achievements. Consequences of stress and anxiety include lack of concentration, irritability, social isolation, and emotional disturbances. These in turn negatively affect socialisation, making interaction with other people less frequent and less conducive to useful feedback on existing beliefs, and to the fruitful exchange of relevant information. Due to reduced socialisation and engagement, the acquisition and retention of knowledge is compromised and intellectual virtues are not exercised. There is at least one problem with considering relief from stress and anxiety as an indirect source of epistemic advantages for motivated delusions. The delusion may bring relief at the time when it is adopted, due to the person being already in an epistemically compromised situation, but it often increases rather than reduces stress and anxiety when it is maintained in the face of conflicting evidence and challenges from third parties. Stress and anxiety no longer come from the negative emotions associated with trauma or loss (“I’m paralysed”, “My girlfriend left me”, “Nobody loves me”, etc.), but from the fact that the content of the delusion can clash with aspects of the person’s experience, conflict with other things she believes or feels, and alienate other people. For all of these reasons, anxiety and depression do not always lessen after a delusion is adopted, they can also heighten. This is particularly true of delusions with negative content that are correlated with higher depression and lower self- esteem (e.g., Smith et al., 2006). Thus, the adoption of a delusional belief may be beneficial because it prevents the occurrence of a disastrous epistemic breakdown, but its benefits are unlikely to outlive the prevention of the breakdown. It should also be kept in mind that acknowledging that motivated delusions can have some epistemic benefits is not equivalent to claiming that their epistemic benefits outweigh their epistemic costs. My claim is relatively modest: when they are adopted delusions can be epistemically innocent, as opposed to epistemically justified, or epistemically good overall. The obvious question is why the person does not adopt a belief that has the same epistemic benefits as the delusional one but fewer costs. One suggestion emerging from the empirical literature is that, in the “extraordinary circumstances” in which the agent finds herself, no other belief with the relevant characteristics is available. Meeting the No Alternatives condition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ I have reviewed some of the literature suggesting that motivated delusions can help manage negative emotions and avoid low self-esteem because they function as a defence mechanism against the psychological effects of previous adversities or recent trauma. This may be sufficient to claim that, if relief from anxiety and distress benefits agents by supporting their epistemic functionality, then motivated delusions can be thought of as epistemically beneficial in the circumstances. This claim, though, does not amount to an epistemic defence of motivated delusions. In order for delusions to be epistemically innocent, we also need to establish that less epistemically costly beliefs could not deliver the same epistemic benefits. This is analogous to the legal case: for Ann to be found innocent of injuring Ben, the court needs to be convinced that there was no other way for her to stop him from detonating the bomb. How to characterise the availability of alternative beliefs is an issue which deserves greater attention than I can give it here. To make a start, I shall review how the unavailability of alternative beliefs has been assessed with respect to delusions emerging in the context of anosognosia. There are two main obstacles to adopting a non-delusional belief (such as “I am paralysed”): (1) evidence for the truth of the non-delusional belief is not available; and (2) evidence for the truth of the non-delusional belief is available but the agent’s capacity for evaluating competing hypotheses is compromised, and thus the evidence is not taken into account. To start with option (1), is evidence for their impairment available to people with anosognosia who deny their impairment? The possibility that people may be covertly aware of their impairments has been discussed widely in the empirical and philosophical literature (see, for instance, Fotopoulou, Pernigo, Maeda, Rudd, & Kopelman, 2010; Levy, 2008). Ramachandran (1995) found that people with anosognosia for left hemiplegia who are asked to choose between unimanual tasks (tasks that they can perform by using just one hand) and bimanual tasks (tasks that they can perform only if they use both hands) consistently choose bimanual ones because the expected rewards are greater. Controls (people with paralysis but no anosognosia) choose the unimanual tasks instead. This seems to suggest that people with anosognosia genuinely believe that they can succeed in those tasks. Given these results, there is no reason to suppose that evidence of their impairment is available to them. Similar conclusions can be drawn from the use of a virtual reality box, where the person is fooled into seeing her (paralysed) arm moving by following the experimenter’s instructions (in reality, the moving arm belongs to someone else). If people were aware of their paralysis, they would show (verbal and non-verbal) signs of surprise, but this is not the case. [F]ar from being a mere façade-like condition that leaves room for traces of insight to leek through, anosognosia runs deep. With respect to option (2), some cases of anosognosia show that the person can be made temporarily aware of her impairment (via vestibular stimulation), but lacks the capacity to integrate the information in her overall conception of herself (Ramachandran, 1995, p. 36). As we saw earlier, Ramachandran explains the formation of the delusions in terms of the need to preserve a coherent sense of self. Usually, the left hemisphere produces confabulatory explanations aimed at preserving the status quo, but the right hemisphere detects an anomaly between the hypotheses generated by the left hemisphere and reality. So, it forces a revision of the belief system. In people with anosognosia, this discrepancy detector in the right hemisphere no longer works, and the belief system fails to update. For Aimola Davies and Davies (2009), it is both true that the person has no direct evidence of the impairment and that she cannot use the other available evidence to revise her belief that she is healthy. Indeed, the formation of anosognosia is explained by the authors in terms of two factors: (1) the person’s motoric failure does not make itself known to the person via direct experience due to neglect, loss of proprioception, or specific problems of integration and memory; and (2) the person cannot use other available evidence of motoric failure due to problems with working memory and executive function. This analysis strongly suggests that evidence for the belief that there is an impairment is not usually available to people with anosognosia: they cannot learn about their impairment from their own experience (due to a neuropsychological deficit that constitutes factor one), and they cannot use other evidence to come to the conclusion that they are impaired (due to their compromised capacity to evaluate competing hypotheses that constitutes factor two). Even if evidence for non-delusional beliefs were available to people with anosognosia, such beliefs would probably fail to support the agent’s epistemic functionality to the same extent as the delusional beliefs. A more plausible belief (e.g., “I am paralysed”) may not be as well placed as the delusional one to play a defensive function, in terms of preserving a coherent and positive self and defusing the negative emotions caused by trauma and disability. Given that the non-delusional beliefs lack such psychological benefits, they may also lack the epistemic benefits associated with them.","In this paper I have argued that the psychological benefits of motivated delusions can convert into epistemic ones. Motivated delusions have the potential for epistemic innocence, where epistemic innocence is characterised as the epistemic status of those cognitions that have obvious epistemic costs but also have a significant epistemic benefit that would be otherwise unattainable. Those delusions that allow the agent to manage negative emotions and avoid low self-esteem are also likely to support the agent’s epistemic functionality. Moreover, when they are adopted, motivated delusions may be the only beliefs supporting epistemic functionality that are available to agents distressed by the consequences of previous adversities, recent trauma or loss. When we think about epistemically costly cognitions that may have psychological benefits, such as self- deception, positive illusions, confabulatory narratives, and distorted memories, we usually think in terms of there being a trade-off. Believing something false or putting a positive spin on a past event can make us feel better, but it leads us further away from the truth. Thus, it may increase wellbeing, but it is not epistemically good. The case for the potential epistemic innocence of motivated delusions puts some pressure on the trade- off view. It would be misleading to believe that motivated delusions provide anxiety- relief and protect self-esteem by compromising access to the truth. Rather, in the picture I have sketched, delusional beliefs are adopted at a time when access to the truth is already compromised by the effects of trauma or previous adversities, and it would be further compromised unless negative emotions were effectively managed. As a temporary response to an emergency, motivated delusions play a useful epistemic function. These considerations obviously apply to other epistemically costly cognitions. In particular, everyday self-deception and motivated delusions seems to have a very similar shear-pin function in that they are the result of a mechanism that lets desires shape beliefs. In so far as motivational influences on belief formation relieve anxiety and stress, everyday self-deception and motivated delusions can carry some benefits by supporting the agent’s epistemic functionality. But different from self-deception, motivated delusions may invite a radical embellishment of the agent’s reality and create tension both within an agent’s belief system and between the agent and other agents, thereby causing inconsistencies, social isolation, and withdrawal. Thus motivated delusions are more likely to become maladaptive (as opposed to adaptive) instances of misbelief than everyday self-deception. On the other hand, the person with everyday self-deception may have more alternative hypotheses available to her than the person who ends up endorsing a motivated delusion, if we suppose the former is not subject to perceptual abnormalities or reasoning impairments to the same extent as the latter. Thus, it may be harder to argue for the epistemic innocence of non-clinical self-deception. The case of motivated delusions and its analogies and disanalogies with self-deception illustrate perfectly the limitations of the trade-off view: some of the psychological benefits attributed to delusions carry significant epistemic benefits that it would be unwise to neglect. It may seem that wellbeing is safeguarded at the expense of truth when the delusional belief is adopted, but a reflection on the effects of adopting the delusion suggests that safeguarding wellbeing and promoting epistemic functionality go hand in hand. In the case of Reverse Othello syndrome described earlier, the clinical team decided not to challenge the delusion after they realised that there were no other psychotic symptoms and the delusion was playing a defensive function. Persistent attempts […] to challenge B.X.’s delusional beliefs were unsuccessful and usually led him to become tearful and agitated. It was concluded that B.X.’s fantasy system functioned to protect him from the consequences of massive narcissistic injury and attendant depressive overwhelm. All members of the treating team were instructed not to aggressively B.X.’s delusional beliefs but were also cautioned not to become complicit in his elaboration of them. Similarly, Fotopoulou observes that challenging delusions in a person with anosognosia can prove ineffective and psychologically disruptive: [RM’s] engagement in rehabilitation activities was initially very poor as he was not motivated and required constant prompting and supervision. Attempts to contradict his anosognosia and increase his motivation were often ineffective as RM immediately provided a series of confabulations to support his alleged abilities and he was particularly sensitive to poor performance and negative feedback. The excerpts above make a similar point in different contexts (Reverse Othello syndrome and anosognosia): if the delusional belief provides some psychological benefit and the benefit is not available via any other belief that the person would accept at that time, then challenging the delusion is a bad idea. A clinical team might decide not to challenge a delusion if they think that challenging an agent is going to be ineffective or disruptive, or if there is a high risk of depression ensuing from the agent’s insight into her mental illness. My discussion suggests that, in these contexts, challenging the delusion might not be advisable from an epistemic point of view either. At the critical stage, motivated delusions may serve a useful epistemic function, allowing the agent to overcome negative feelings or low self-esteem that would prevent her from exercising her epistemic functionality."],["We examined the effects of the Healthy Body Image (HBI) intervention on positive embodiment and health-related quality of life among Norwegian high school students. The intervention comprised three interactive workshops, with body image, media literacy, and lifestyle as main themes. In total, 2,446 12 th grade boys (43%) and girls (mean age 16.8 years) from 30 high schools participated in a cluster-randomized controlled study with the HBI intervention and a control condition as the study arms. Data were collected at baseline, post-intervention, 3- and 12-months follow-up, and analysed using linear mixed regression models. The HBI intervention caused a favourable immediate change in positive embodiment and health-related quality of life among intervention girls, which was maintained at follow-up. Among intervention boys, however, weak post-intervention effects on embodiment and health-related quality of life vanished at the follow-ups. Future studies should address steps to make the HBI intervention more relevant for boys as well as determine whether the number of workshops or themes may be shortened to ease implementation and to enhance intervention effects. --------------------------------------------------------------------------------","Positive embodiment and body appreciation are important aspects of health and quality of life (Avalos, Tylka, & Wood-Barcalow, 2005; Piran, 2019; Tiggemann, 2011). In previous studies, positive embodiment and body appreciation have been associated with positive self- and body esteem, healthy eating, and performing regular physical activity in boys and girls (Cash & Fleming, 2002; Neumark-Sztainer, Paxton, Hannan, Haines, & Story, 2006; Santos, Tassitano, do Nascimento, Petribú, & Cabral, 2011; Tylka & Homan, 2015). Further, body image has been found to predict health-related quality of life in boys and girls (Griffiths et al., 2017; Haraldstad, Christophersen, Eide, Natvig, & Helseth, 2011). There is however a well-known gender difference, as fewer adolescent boys struggle with body image issues (13–45%) compared to adolescent girls (45–71%) (Martinsen, Bratland-Sanda, Eriksson, & Sundgot-Borgen, 2010; Torstveit, Aagedal-Mortensen, & Stea, 2015). In the same vein, adolescent boys report more satisfaction with their bodies and higher levels of embodiment compared to adolescent girls (Franko, Cousineau, Rodgers, & Roehrig, 2013; Holmqvist, Frisén, & Piran, 2018; Neumark-Sztainer et al., 2006; Santos et al., 2011). From a developmental perspective, changes in the experience of the body during the critical phase of adolescence can have a long-term impact on body image (Wertheim, Paxton, & Blaney, 2009). Promoting positive embodiment in adolescence is therefore vital to establish a good basis for health-related quality of life, as such quality of life has proved stable during the life course (Bisegger, Cloetta, von Rueden, Abel, & Ravens- Sieberer, 2005), and can be viewed as a core issue for public health. Systematic reviews show that universal intervention programs that are successful address the reduction of risk factors, as for example body dissatisfaction, in order to prevent eating disorders among adolescents (Le, Barendregt, Hay, & Mihalopoulos, 2017; Stice, Shaw, & Marti, 2007; Yager, Diedrichs, Ricciardelli, & Halliwell, 2013). Within a health promotion perspective, promoting positive embodiment represents a theoretical and methodological paradigmatic shift from the disease-preventing focus, e.g., by preventing body dissatisfaction, to a health-promotion focus (Le et al., 2017; Stice, Becker, & Yokum, 2013). This shift opens new possibilities to assess health-promotion interventions (Piran, 2015; Tylka & Wood- Barcalow, 2015; for examples, see Alleva et al., 2018; Halliwell, Jarman, Tylka, & Slater, 2018; McCabe, Connaughton, Tatangelo, Mellor, & Busija, 2017). The research-based positive embodiment construct is defined as “positive body connection and comfort, embodied agency and passion, and attuned self-care” (Piran, 2016, p.47). Positive embodiment relates conceptually to body appreciation (Tylka & Piran, 2019), the most commonly used construct in assessing positive body image (Tylka, 2019). Both positive embodiment and body appreciation emphasize positive connection to, and appreciation of, the body, as well as attuned care of the body (Tylka & Piran, 2019). The positive embodiment construct, however, includes in addition, experiences of agency to act in the world and comfort with bodily desires (Piran, 2019). Researchers have called for intervention studies that aim to enhance embodiment and health-related quality of life (Alleva, Sheeran, Webb, Martijn, & Miles, 2015; Tylka & Piran, 2019). Yet, most existing intervention studies lack inclusion of multidimensional instruments of positive embodiment (Webb, Wood-Barcalow, & Tylka, 2015). In particular, no randomized, controlled outcome evaluation studies have been conducted as a universal promoting program aimed at enhancing positive embodiment in both boys and girls in late adolescence (Alleva et al., 2015). Development and implementation of the HBI intervention ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We have developed the universal, multi-component health-promotion intervention “Healthy Body Image” (HBI; Sundgot-Borgen et al., 2018). The HBI intervention focuses on positive embodiment and health-related quality of life among Norwegian high school students, and employs an interactive educational approach, which has been found suitable in school settings (Yager et al., 2013). The HBI intervention comprised three overarching themes related to body image, media literacy, and lifestyle, as these have been found to improve physical self-perception, body satisfaction and appreciation, physical competence, and body esteem, sometimes with large effect sizes (Alleva et al., 2015; Espinoza, Penelo, & Raich, 2013; Franko et al., 2013; Tomyn, Fuller-Tyszkiewicz, Richardson, & Colla, 2016). A more detailed description of the program and its rationale has been published elsewhere (Sundgot-Borgen et al., 2018). The program was constructed to include both boys and girls in late adolescence. This was important because the peer environment is shaped by sociocultural ideals of both genders. Both boys' and girls' attitudes must change if the social environment of the whole school can be changed (Yager et al., 2013). Due to the mixed-gender sample, the intervention contained gender neutralized and gender specific contents (e.g., pictures, videos, communication examples), to make it relevant for both genders. Despite some debate on what age is most appropriate for initiation of body image interventions, evidence suggests that in prevention studies, it might be beneficial to target young adolescents prior to the onset of eating disorders (Espinoza et al., 2018; Rohde, Stice, & Marti, 2015). However, late adolescence involves pubertal, cognitive, and interpersonal changes, which increase adolescents’ ability to reach a more abstract characterization of themselves, the influence of their peers increases (Rohde et al., 2015), and they may become more aware of and vulnerable to pressures to attain sociocultural beauty ideals. They are at an age where the risk for eating disorders peaks (Espinoza et al., 2018; Rohde et al., 2015; Stice et al., 2007), and promotion of positive embodiment is especially crucial, as they are moving towards the independence of young adulthood. Also, their improved ability for abstract reasoning makes them more likely to comprehend the intervention content, relate skills to their own lives, and take advantage of such taught skills. The school context also ensures a relatively comparable participation rate between genders, which is an obvious asset since few existing studies have managed to include a balanced gender sample. Moreover, a mixed-gender approach may offer a more real-life setting in universally implemented health promotion initiatives (Yager et al., 2013). Hypothesis ~~~~~~~~~~ We hypothesized that the HBI intervention would be effective, resulting in more favourable scores on positive embodiment (higher) and health-related quality of life (higher) in intervention students compared to control students. Design and randomization ~~~~~~~~~~~~~~~~~~~~~~~~ A cluster-randomized controlled design was used with schools as the clustering factor at a ratio of 1:1. Schools were randomly allocated to either the HBI intervention or the control group to equalize sample size, and the effect of socioeconomic and demographic variables, notably related to ethnicity and the urban-rural dimension. The sample would be considered representative of the adolescent population of Oslo and Akershus County. The randomization was conducted by a professional not affiliated with the study to minimize contamination biases within schools. During the intervention period, students at the control schools followed their regular school curriculum. Fig. 1 presents a diagram of the inclusion and randomization process of schools and students, respectively. Sample characteristics ~~~~~~~~~~~~~~~~~~~~~~ Thirty schools were randomized and 2,446, 1,254, 1,278, and 1,080 students consented to participate at pre-test, post-intervention, and 3- and 12-months follow-up, respectively (Fig. 1). The mean (range) number of students consenting at each school was 82 (22–184), 42 (5–97), 43 (4–125), and 36 (3–103) at pre-test, post-intervention, and 3- and 12-months follow-up, respectively. The number of students included in the primary outcomes analyses were 1,742, 1,190, 1,172, and 955 for the Experience of Embodiment Scale, and 1,688, 1,173, 1,158, and 925 for the KIDSCREEN-10 and General health across the four measurement occasions. The participants were 16.8 (SD = 0.76) years old, and 11%, and 1% were categorized as overweight and obese, respectively. Among the participants, 13% were categorized as immigrants, 39% had parents with a total income of ≥1 million NOK, and 82% reported one or both parents having a higher education. Ethics approval and consent to participate ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The study met the intent and requirements of the Health Research Act and the Helsinki declaration, and was approved by the Regional Committee for Medical and Health Research Ethics (P-REK 2016/142). It was enrolled in the international database of controlled trials www.clinicaltrials.gov (ID: PRSNCT02901457). Students at consenting schools had the prerogative to decline participation after consent. In such cases, students were allowed to follow the HBI workshops, but without completing the questionnaires. After the final 12- months follow-up, control schools were offered one lecture where the program highlights were compressed. The methods and results are described according to the Consort Statement (Moher, Schulz, & Altman, 2001). Procedure and data collection ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As a result of a subsequent pilot study during March and April 2016 among 120 12th grade high schoolers, a few questionnaire items about body perception and nutrition were deleted to reduce the risk of error variance due to acquiescence bias. In addition, the amount of workshop assignments was reduced to allow for more time allocated to discuss mood and body satisfaction issues. The HBI intervention included all 12th grade high school classes following a general study program, excluding students following a vocational study program. No further exclusion criteria were set. During Spring 2016, principals of all public and private high schools in Oslo and Akershus County in Norway were contacted by e-mail. Oral and written study information was provided to students and staff at the consenting schools. The Norwegian Health Research Act states that adolescents, 16 years or older, can give their informed consent with no parental consent needed. Students were sent an e-mail with study information and a letter of informed consent. If they pressed \"yes\" to the question of consent, they were given access to a link that made the questionnaire package available, and they completed the questionnaire package through the online survey system SurveyXact 8.2. Ethical approval of the study required that the students completed the questionnaires outside regular school hours. Students were informed about their allocation into the intervention or control group after the randomization.","As described in the study protocol (Sundgot-Borgen et al., 2018), participants completed standardized questionnaires related to demographics, positive embodiment, and health- related quality of life at baseline, post-intervention, and at 3- and 12-months follow-up, respectively. All baseline assessments were conducted prior to the randomization. Post- intervention assessment was not available the same day as the last workshop, but within one week (Sundgot-Borgen et al., 2018). Demographic variables The demographic variables were collected at all measurement occasions, including age, gender, and self-reported body weight (kg) and height (cm). BMI was calculated as body weight (kg) divided by the height squared (m2). Categorization of weight status was based on international age- and gender-adjusted cut-off scores (Cole, Bellizzi, Flegal, & Dietz, 2000). Total parental income was measured by asking the students what they believed to be their parents' total income, selecting one of five options (less than NOK 200.000, NOK 200.000 - 400.000, NOK 500.000 - 800.000, NOK 900.000 - 1 million, more than NOK 1 million, respectively). Students also ticked off if their parents had completed 1. Primary school, 2. High school, 3. College/University, or whether they 4. Did not know. Immigration status was measured by asking whether the student or both parents had immigrated (Yes I have, Yes both my parents, No). Positive embodiment Positive embodiment was measured using the Experience of Embodiment Scale (EES) (Teall & Piran, 2012). The Cronbach’s alpha for the current study was .93 for girls and .92 for boys, similar to other studies with the range of .91–.94 (Chmielewski, Bowman, & Tolman, 2019; Holmqvist et al., 2018; Piran, 2019; Teall, 2006, 2014). Test-retest reliability over a 3-week period of the EES was also previously found to be acceptable (r = .93) (Piran, 2019). The 34 items covered positive connection with the body, agency and functionality, experience and expression of desire, body attunement, self-care vs. harm/neglect, and subjective lens vs. self-objectification (e.g., \"I am proud of what my body can do\" and \"I care more about how my body feels than about how it looks\"). The items had a Likert-format ranging from 1 (strongly disagree) to 5 (strongly agree), and the 17 negatively framed items (e.g., \"I ignore the signs my body sends me\" and \"My dissatisfaction with my body/appearance has a negative effect on my social life\") were reversed so that the sum score reflected higher levels of positive embodiment. Adequate construct validity of the EES has been found in previous studies on young adults as reflected by positive correlations with measures of body esteem in women (rs = .76–.79) and men (r = .69), body responsiveness (r = .73), body connection (r = .60), well-being (rs = .55–.80), and life satisfaction in men (r = .68) and women (r = .66). Further, the EES correlated negatively with measures of objectified body consciousness (rs = -0.55, -.73), eating problems (rs = -0.43, -.70), alexithymia (rs = -0.51, -.54), and depression (r = -0.63) (Chmielewski, Tolman, & Bowman, 2018; Holmqvist et al., 2018; Piran, 2019; Teall, 2006, 2014). Young men have reported higher EES scores compared to women (Holmqvist et al., 2018). Since the present investigation included late adolescents, ages 16–17, the study used the adult version of the EES. To date, most validation studies of the EES were conducted in young adult samples, such as Chmielewski et al. (2018) that included 340 women between the ages of 18–26 with an average age of 19.81. Based on a series of confirmatory factor analyses, the global EES score was used as an outcome measure. While its original 6-factor model showed an adequate fit when modeling the method variance related to the positively and negatively worded items, χ2(507) = 3311, p < .001, RMSEA = 0.056, CFI/TLI = .890/.867, SRMR = .066, we used a global score since a general second-order factor, χ2(516) = 3431, p < .001, RMSEA = .057, CFI/TLI = .875/.864, SRMR = .076, accounted adequately for the 6-factor model. Health-related quality of life Health-related quality of life was measured by the KIDSCREEN-10, which is a widely used and validated self-report tool (Ravens-Sieberer, 2006), and has been validated in Norwegian adolescents (Haraldstad & Richter, 2014). The scale consists of 10-items (e.g., \"Have you felt fit and well?\" and \"Have you felt sad?\"). The sum score of the 1–10 provides a general health-related quality of life index. A separate item included in the KIDSCREEN-10 measured perceived General Health (\"In general, how would you say your health is?\"), which has been found to correlate well with measures of physical well-being (r = .63) and psychological well-being (r = .51) (Barthel et al., 2017). All items, 1–11, had a 5-point Likert-type format from 1 (not at all/never) to 5 (extremely/always) for 10 items, and from 1 (excellent) to 5 (poor) for the General Health item. Negatively worded questions were reversed, and hence a higher score indicated higher levels of health-related quality of life. Standardized T-scores were presented at baseline to enable comparison of means across study samples and compare data to health-related quality of life norm data. A score of 50 represents the mean. A T-score < 38 on the KIDSCREEN-10 indicates lower health-related quality of life, while scores ≥ 38 indicate preferable reported health-related quality of life (Ravens-Sieberer, 2006). The internal consistency for this sample was α = .81, and has been found to be satisfactory in other samples of adolescent boys and girls (Haraldstad et al., 2011). The HBI intervention ~~~~~~~~~~~~~~~~~~~~ There is no consensus as to which theoretical orientation may provide the most effective approach when developing a health promotion intervention aiming to promote embodiment and health-related quality of life (Alleva et al., 2015). However, a sociocultural perspective (Thompson, Heinberg, Altabe, & Tantleff-Dunn, 1999) was natural to consider when aiming to change attitudes, beliefs, and knowledge related to idealized lifestyles (involving e.g., extreme exercise and diet regimes) and bodies, to further strengthen the resilience towards unhealthy internalization, and strengthen life-managing skills in a mixed-gender school-based setting. Also, an etiological model of risk and protective factors (Piran, 2015; Smolak & Piran, 2012) as well as the developmental theory of embodiment (Piran, 2017; Teall & Piran, 2012) within the realm of positive psychology (Seligman & Csikszentmihalyi, 2000), were important in its development. Although thoroughly described in the Appendix, some important aspects of the intervention specifically aiming to promote positive embodiment are presented. Through the body image and media literacy workshops, we aimed to improve critical awareness of unhealthy body and lifestyle idealization, critical and constructive use of social media, including consequences of current body ideals for boys and girls. By this, we intended to reduce the risk of internalization of unhealthy ideals, self-harm, and neglect, as well as promote a subjective lens while reducing self- objectification. To improve a positive connection with the body, we aimed to strengthen attitudes towards, and knowledge about, how to promote self-care and experience of body functionality when discussing lifestyle factors, such as nutrition, exercise, and sleep. The intervention was developed to suit the cognitive development among adolescents 16 years of age in terms of their ability for abstract reasoning. The workshop delivery was based on the elaboration likelihood model (Petty & Briño, 2012; Petty & Cacioppo, 1986). According to this model, as well as previous findings (Alleva et al., 2015; Stice et al., 2013, 2007), the program contained three 90-min interactive workshops to facilitate extensive student discussions. All workshops were arranged in classrooms during regular school hours. About 60 boys and girls (i.e., two school classes) participated per workshop. Student attendance was registered at each workshop to calculate program adherence. A 3-week interval between each workshop resulted in a 3-month intervention period. The first and fourth author facilitated the intervention. Both are specialized in physical activity and health, sports nutrition, motivational interviewing, and body image among adolescents. Detailed information about the intervention content and targets can be found in the study protocol (Sundgot-Borgen et al., 2018). Sample size and power analyses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The statistical power estimation was based on two comparison groups (α = .05 and b = .20) with an average within-cluster sample size of 70 students. The expected effect size was .28 according to a meta-analysis (Hausenblas & Fallon, 2006) that included 35 studies examining intervention effects on body images variables. Moreover, we assumed that the within-cluster dependency related to schools accounted for approximately 3% (ICC = .03). This is fair for variables related to psychological or mental health outcomes, as selection factors like socioeconomic status affect these variables less than for example academic performance. These considerations required a minimum of 10 clusters within each group, requiring a total sample size of 10 schools × 2 groups × 70 students ˜ 1400 students. Statistical analysis The software program Mplus, version 8.0, was used to carry out factor analyses, while remaining statistics were analysed using IBM SPSS 24 for Windows. The adequacy of the randomization procedure was examined by comparing group differences at baseline with independent t-tests, chi-square tests, or Kruskal-Wallis tests (Table 2). A case was recorded as dropout if all post-intervention and follow-up data were missing. Due to several layers of dependency in the outcome data, linear mixed regression models were fit, as suggested in comparable studies (Wilksch et al., 2017). Dependency within the school clusters was accounted for by adding school as a random factor, whereas dependency between the repeated measures was accounted for by fitting a compound symmetry matrix to the residual matrices (thus assuming equal-sized correlations between measurement occasions). Students were nested within schools, which also was accounted for. The baseline score was used as a covariate to adjust for imperfections in the randomization procedure and to increase the statistical power. The fixed factors were group (one coefficient for the difference between the intervention and the control group), time (a coefficient for each time point except the last, thus detecting a non-linear change), and group × time (to detect if intervention effects were particularly pronounced at certain time points). In order to examine if the level of participations at workshops influenced the outcomes, workshop attendance (WA-number of workshops) was added as linear covariate, as well as interaction terms examining if WA influenced the outcome particularly at certain time points (WA × time) or additionally within just one of the groups (WA × time × group). The restricted maximum likelihood procedure and Type III F-tests were preferred. The analyses were stratified for gender. Statistically significant effects set to p < .05, were followed-up with planned comparison tests (LSD) examining group differences at each follow-up assessment. Results are expressed as absolute numbers (n) and percentage (%) for categorical data and model estimated means including 95% confidence intervals and standard deviation (SD) for continuous data. Effect sizes are presented as Cohen's d and phi- coefficients. Participant demographics ~~~~~~~~~~~~~~~~~~~~~~~~ Participant demographics for each group are presented in Table 1. At baseline, all participants were 16–17 years of age, with a mean BMI within the normal weight range for youths (Cole et al., 2000). The baseline correlation between EES and KIDSCREEN-10 was r = .60 (p < .001) among both boys and girls. Girls in the intervention had higher scores on positive embodiment, health-related quality of life, and the general health item compared to girls in the control group. No significant difference between groups was found in boys for these outcome measures. Based on parents' total income and education level, girls in the intervention group were more likely to be defined with a higher social economic status compared to girls in the control group. Boys in the intervention group had parents with a higher level of education, and fewer were categorized as immigrants compared to boys in the control group (Table 1). The linear mixed regression models were adjusted for group differences at baseline. Dropout analysis ~~~~~~~~~~~~~~~~ No differences were observed in the outcome variables between dropouts and completers in either boys or girls. More students in the control group (p = .001, φ = 10.61), and more boys (p < .001, φ = 52.48) dropped out. Boys who dropped out had slightly higher BMI (p = .044, d = 0.15) and body weight (p = .010, d = 0.20), while girls who dropped out were slightly older (p = .014, d = 0.17). Effect analyses were therefore adjusted for these variables. Positive embodiment intervention effects ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ For boys, the linear mixed regression model showed that the main effect of group (p = .072), time (p = .756) and the interaction effect of group × time (p = .543) were nonsignificant. The planned comparison analyses showed that boys in the intervention group reported higher positive embodiment at post-intervention compared to boys in the control group, suggesting a short-term favorable small effect. However, this effect was lost at the 3- and 12-month follow-ups (Table 2). For girls, the main effect of group was significant, F(1, 777) = 33.11, p < .001, while time (p = .267) and group × time (p = .133) effects were nonsignificant. The planned comparison analyses showed a significant and favorable effect of the intervention on positive embodiment for girls in the intervention group. This effect was maintained at the 3- and 12-months follow-up, respectively. The effect size increased slightly over time, and with a peak at the last follow-up assessment (see Table 2). Health-related quality of life intervention effects ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ For boys, the linear mixed regression model showed a significant main effect of group for health-related quality of life, F(1, 360) = 4.78, p = .029, while the time (p = .148) and group × time (p = .871) effects were nonsignificant. Although the mean differences between boys in the intervention and control groups increased across the assessment time-points, no planned comparison analyses showed statistical significance (see Table 3). For the general health outcome item, the model showed no effect of group (p = .120), time (p = .953), or group × time (p = .191) for boys. The planned comparison analyses did show a favorable and significant post-intervention effect for boys in the intervention group compared to boys in the control group, which was not maintained at follow-up (see Table 3). For girls, the main effect of group for health-related quality of life was not significant (p = .186), whereas significant time, F(2, 860) = 3.99, p =.019, and group × time, F(2, 860) = 4.47, p = .012, effects were observed. The planned comparison analyses showed no significant difference in health-related quality of life between girls in the intervention and control groups at post-intervention and 3-months follow-up. However, a “sleeping effect” was evident, as girls in the intervention group had a significantly higher health-related quality of life (small effect size) at the 12-months follow-up compared to girls in the control group (see Table 3). The model with the general health variable as outcome showed a significant group effect, F(1, 807) = 10.54, p = .001, while the effect of time (p = .466) and group × time (p = .598) were nonsignificant. The planned comparison analyses showed that girls in the intervention group had significantly more favorable general health at post-intervention compared to girls in the control group (small effect size), which was maintained at follow-up, as well (see Table 3). Dose-response effect related to the number of attended workshops ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Since the degree of attendance was irrelevant for the control group, the group variable was recoded as 0 (control group), and 1–4 (1 = 0 workshops in intervention student, 2 = 1 workshop, 3 = 2 workshops, 4 = 3 workshops). Neither group (p = .290), time (p = .715), nor time × group (p = .750) were significant among boys in the intervention group. However, in girls, the main effect of group was significant, F(4, 756) = 10.96, p < .001. The time (p = .284) and the interaction effects (time × group) (p = .335) were nonsignificant. The follow-up tests, as presented in Table 4, indicate that an increasing attendance yielded a stronger intervention effect. A noteworthy finding was that boys and girls needed to attend at least three and two workshops, respectively, in order to benefit from the HBI intervention. This moderation effect was lost among boys at follow-up, but not among girls. All effect sizes were in the small range (see Table 4). Comparable analyses on health-related quality of life and the general health variable revealed no significant moderation effects.","The HBI intervention promoted a post-intervention effect on positive embodiment and perceived general health for boys, although no sustained effects were observed. However, for girls, the HBI intervention promoted immediate and sustained positive embodiment. Additionally, for girls, there was a consistent pattern of improvement in perceived general health at post-intervention and 12-months follow-up, whereas the effects on health-related quality of life were only demonstrated at 12-months follow-up. These findings seem to converge with other body image programs that include follow-up measures (Espinoza et al., 2013; Neumark-Sztainer et al., 2010). The effect sizes in girls were also strongest at the 12-months follow-up, which is noteworthy. The current study increases the knowledge base of the long-term and delayed effect of body image interventions, which currently is scarce. Our study emphasises the importance of long-term follow-ups as some intervention effects may mature in a slower manner. The intervention was intended to facilitate awareness of how attitudes towards the body and lifestyle choices are transmitted through different learned social channels, and, through that, shape students' attitudes, feelings, and lifestyles. According to a sociocultural perspective (Thompson et al., 1999), an increase in critical awareness could have improved the ability to withstand unhealthy idealization, reducing the risk of internalization of such ideals (Teall & Piran, 2012). Students were also taught to become aware of, and use, factors in everyday life that enhance their embodiment. Further, body functionality and well-being were emphasized, rather than appearance, when discussing lifestyle factors. This could have promoted healthy perspectives on how to engage in lifestyle behaviours, similar to positive embodiment characteristics (Tylka & Wood-Barcalow, 2015). The HBI intervention is to our knowledge, the first one among body image interventions to report on effects on health-related quality of life. In girls, the diffusion of the health- related quality of life effect from improving their embodiment was expected because these variables have been found to be highly correlated (Griffiths et al., 2017; Haraldstad et al., 2011). By strengthening the ability to filter media information, reduce unhealthy comparisons, and promote positive self-talk, it might be easier to improve body acceptance which may transform into better psychological well-being. Moreover, improving self-care and a healthy conscious lifestyle, may ultimately improve physiological health, which may explain the observed improvements in health-related quality of life. The effect sizes were in general small and comparable with previous studies (Franko et al., 2013; Halliwell, Jarman, McNamara, Risdon, & Jankowski, 2015; Lindwall & Lindgren, 2005; Morgan, Saunders, & Lubans, 2012; Sharpe, Schober, Treasure, & Schmidt, 2013). In contrast to clinical studies, the interpretation of small effect sizes may be more favourable. Thus, such small effect sizes are common in universal interventions, and may be expected due to low base rates for clinical symptoms, and a high probability of ceiling effect for positive health indices. Similarly, by definition, study variables in health promotion studies do not pre- select participants having scores within a clinical range (Wilksch, 2014). Attention has been given in the literature (Piran, 2001) to how students perceive the credibility of those who deliver intervention programs. In the present study, students were informed about the facilitators' education and academic position. In addition, the facilitators were attentive to the quality of their verbal and non-verbal communication with the students. Nevertheless, the students’ perceived credibility of the workshop facilitators was not assessed. An explicit rationale for the HBI intervention was to promote the interaction between boys and girls, and to mirror the across-gender sociocultural influences on body experiences that occurs in a realistic real-life setting. Strategies to accomplish this rationale included the use of different interactive components, thus, to enhance the chance of effect in both genders. Our study only found long-term effects in girls. This may support previous suggestions that girls are more receptive to body image interventions (Stice et al., 2007) even when efforts have been made to make the intervention gender neutral. Importantly, our results do not document that a single-gender intervention is preferred. Further, the HBI intervention is a health promotion intervention, where the aim is not only to reduce risk factors, but to promote health- related factors. Based on our findings, a mixed-gender approach might have been important to girls despite the lack of effect in boys. To further investigate whether single- or mixed-gender approaches is most effective, future studies need to include more arms (control, mixed-gender, single-gender- group) into the study design. Similar to the effects of the HBI intervention, weak and transient effects from a body image intervention has been found in other studies on young adult men (Jankowski et al., 2017). Importantly, although undocumented, the presenters observed that the boys found the topics of \"comparison,\" \"self-talk,\" and \"communication\" not as relevant as the girls, which could have made it more difficult to be engaged and receptive to the workshop content. Previous studies have shown that enhancing peer comradery and connection, and including masculine points of reference, helped engage boys and men in an intervention (Seaton et al., 2017). Perhaps the female implementers in the HBI intervention may have had challenges with potentially important factors to engage boys as well as may have under-communicated the masculine aspects. Virtually no effects among boys may also be explained by scores above norm data for health-related quality of life at baseline (Ravens-Sieberer, 2006). Although no norm data for the EES exists for late adolescent boys, one study on young men showed that boys scored significantly higher on the EES compared to girls (Holmqvist et al., 2018). This could reflect that boys at baseline are more accepting of their bodies, and therefore have a lower improvement potential compared to girls. At present, it remains unsettled whether the intervention may work better among boys with lower baseline health- related quality of life and embodiment, and whether it may work equally well in a girls- only group. Our findings contradict the suggestion (Wilksch, 2017) that a single-session (workshop) intervention may suffice. Although a one-session may be more feasible in school settings, our results are in line with the elaboration likelihood model (Petty & Briño, 2012; Petty & Cacioppo, 1986) and previous meta-analyses (Stice & Shaw, 2004; Stice et al., 2007), that at least two workshop sessions were needed for girls to maintain the intervention effects at follow-up. Strengths, limitations, and future directions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Assets of the present study are the theoretical framework, the user involvement through a pilot study, the randomized controlled design and the adequate statistical power. However, a loss of power at the follow-ups may have increased the probability of Type II errors, especially in boys. The fact that boys who dropped out had a slightly higher BMI is consistent with previous observations in health- and body image-related interventions and classroom-based activities (Finn, Faith, & Seo, 2018), that those with higher BMI feel self-conscious when exposed to the intervention content. However, boys who dropped out did not differ in positive embodiment or health-related quality of life, which reduces the reasons to believe that many of those who might especially benefit from our intervention dropped out. Drop-outs seem almost inevitable, yet some steps may be mentioned to counteract them. Although we used measures of positive aspects of body image and not measures of body dissatisfaction, care should be taken when considering the comprehensiveness of the questionnaire, and to decrease the number of included questions, notably those of a sensitive nature. To facilitate improvement potential, one challenge is to select outcome measures where both genders have room for improvement. Before a broader dissemination of the HBI intervention, modifications to the workshops should be tested, with male facilitators, to further investigate whether it might be possible to achieve genuine and sustainable effects for boys. Also, although the credibility of the workshop holders was planned and facilitated for, the students’ perceptions of this credibility were not assessed. This is a limitation, and future studies should include such assessment. In addition, there is a need to study the dismantling potentials. The present findings clearly indicate that among girls, two interactive and multicomponent workshops may suffice. However, future studies need to address the issue of which of the three workshops that may be deleted from the program. This would inform which of the themes (i.e. body image, media literacy, and lifestyle factors) that should be retained.","The HBI intervention promoted a post-intervention effect on positive embodiment and perceived general health in boys. The intervention promoted a sustained effect on positive embodiment and health-related quality of life in girls. Future studies should examine the effect of only two workshops for girls and modifications of the workshops for boys to see if it is possible to obtain sustained effects in boys as well.","The authors declare that they have no competing interests.","This work was supported by The Norwegian Woman`s Public Health Association (H1/2016), the Norwegian Extra Foundation for Health and Rehabilitation (2016/FO76521), and TINE SA. The sponsors came in after the study protocol was developed and did not have any role in development of study design, data collection, analysis or interpretation of data, or manuscript writing and submission."],["Background Developmental coordination disorder (DCD) is a common developmental disorder but its long term impact on health and education are poorly understood. Aim To assess the impact of DCD diagnosed at 7 years, and co-occurring developmental difficulties, on educational achievement at 16 years. Methods A prospective cohort study using data from the Avon Longitudinal Study of Parents and Children (ALSPAC). National General Certificate of Secondary Education (GCSE) exam results and Special Educational Needs provision were compared for adolescents with DCD (n = 284) and controls (n = 5425). Results Adolescents with DCD achieved a median of 2 GCSEs whilst controls achieved a median of 7 GCSEs. Compared to controls, adolescents with DCD were much less likely to achieve 5 or more GCSEs in secondary school (OR 0.27, 95% CI 0.21–0.34), even after adjustment for gender, socio-economic status and IQ (OR 0.6, 95% CI 0.44–0.81). Those with DCD were more likely to have persistent difficulties with reading, social communication and hyperactivity/inattention, which all affected educational achievement. Nearly 40% of adolescents with DCD were not in receipt of additional formal support during school. Conclusions DCD has a significant impact on educational achievement and therefore life chances. Co-occurring problems with reading skills, social communication difficulties and hyperactivity/inattention are common and contribute to educational difficulties. Greater understanding of DCD among educational and medical professionals and policy makers is crucial to improve the support provided for these individuals. --------------------------------------------------------------------------------","Developmental coordination disorder (DCD) is one of the most common developmental conditions of childhood. However, its impact on longer term health and education is not well understood. This paper contributes robust epidemiological evidence of the persisting impact that DCD can have on learning and achievement in secondary school. Using a population-based cohort, individuals with DSM-IV classified DCD at 7 years were 70% less likely to achieve 5 or more qualifications at 16 years than their peers. Co-occurrence of reading difficulties and other developmental traits, such as social communication difficulties and hyperactivity/inattention, were commonly present in adolescence and these difficulties contributed to poor educational achievement. This study also illustrates the extent to which DCD can be a hidden disability – 37% of those with the condition were not in receipt of additional formal teaching support. These results demonstrate the impact DCD can have on educational attainment, and therefore on future life prospects. It is hoped this work will contribute to raising awareness and understanding of the impact of DCD and stimulate discussion about how best to support those with this complex condition.","Developmental coordination disorder (DCD) is a common neurodevelopmental disorder characterised by deficits in both fine and gross motor coordination which have a significant impact on a child’s activities of daily living or school productivity (American Psychiatric Association, 2013). These deficits are present in the absence of severe intellectual or visual impairment, or another motor disability, such as cerebral palsy. It is thought to affect around 5% of school-aged children (American Psychiatric Association, 2013), but despite its high prevalence it remains one of the less well understood and recognised developmental conditions in both educational and medical settings. Schoolwork of children with DCD often does not reflect their true abilities as they struggle with fine motor skills, including handwriting (Missiuna, Rivard, & Pollock, 2004). However, there is also evidence of a wider academic deficit involving reading, working memory and mathematical skills (Alloway, 2007; Dewey et al., 2002; Kaplan, Wilson, Dewey, & Crawford, 1998). Although initially identified on the basis of motor difficulties, the condition may develop into complex psychosocial problems, with difficulties in peer relationships and social participation (Sylvestre, Nadeau, Charron, Larose, & Lepage, 2013), bullying (Campbell, Missiuna, & Vaillancourt, 2012; Skinner & Piek, 2001; Wagner et al., 2012), low self-worth and perceived self-competence (Piek, Baynam, & Barrett, 2006), and internalising disorders, such as anxiety and low mood (Lingam et al., 2012). It may be that these sequelae lead to poor performance in school. As well as secondary psychosocial consequences, those with DCD have a higher risk of displaying other developmental traits, such as hyperactivity and social communication difficulties, and specific learning disabilities, particularly dyslexia. (Kadesjo & Gillberg, 1999; Lingam et al., 2010; Sumner et al., 2016). Overlapping difficulties in two or more developmental and educational domains implies that discrete diagnosis of a single disorder is often not appropriate (Kaplan, Dewey, Crawford, & Wilson, 2001). In some individuals with DCD, it may be that co-occurring difficulties contribute to or explain some of the sequelae of the condition (Conti-Ramsden, Durkin, Simkin, & Knox, 2008; Loe & Feldman, 2007). Previous work has highlighted the importance of identifying co-occurring problems in DCD by demonstrating the mediating effects of social communication difficulties and hyperactivity can have on psychological outcomes (Harrowell, Hollén, Lingam, & Emond, 2017; Lingam et al., 2012). In the evaluation and management of a child with suspected DCD, consideration of other possible co-occurring difficulties is essential. Although the body of literature on DCD demonstrates many reasons why a child with DCD might struggle in school (Zwicker, Harris, & Klassen, 2013), lack of awareness of the condition by medical and educational professionals is widespread, highlighted by parental reports of difficulty accessing support and services for their child (Missiuna, Moll, Law, King, & King, 2005; Novak, Lingam, Coad, & Emond, 2012). One study found that 43% of parents were not offered any practical support (Alonso Soriano, Hill, & Crane, 2015). Even when support is provided, it may not be appropriate for the child, which parents put down to lack of understanding of the condition (Maciver et al., 2011). A study of students in further and higher education found that those with dyslexia were more likely than students with DCD to receive Disability Student Allowance from the government, despite greater self-reported difficulties in the DCD group (Kirby, Sugden, Beveridge, Edwards, & Edwards, 2008). Furthermore, there were no differences between the types of support provided for these two different developmental disorders. This not only emphasises the poor recognition of DCD, but also lack of understanding of the specific needs of those with coordination problems. By definition, children with DCD have motor difficulties that interfere with academic achievement (American Psychiatric Association, 2013). However, longitudinal studies of educational achievement in secondary school for those with DCD are few, and those that have been conducted lack strict diagnostic criteria or have been drawn from clinical samples. (Cantell, Smyth, & Ahonen, 1994; Gillberg & Gillberg, 1989; Losse et al., 1991). Thus, the primary aim of this research was to assess the impact of DCD on educational achievement in secondary school, using prospective data from a large population-based cohort study. Secondly, we aimed to assess the presence of co-occurring difficulties in reading ability, social communication problems and hyperactivity/inattention, and whether these impacted upon educational achievement in DCD. Thirdly, we aimed to determine how many of those meeting the criteria for DCD were identified for formal additional educational support in school, and assess whether provision of support was related to educational achievement. Study participants ~~~~~~~~~~~~~~~~~~ The Avon Longitudinal Study of Parents and Children (ALSPAC) is a population-based birth cohort which invited all pregnant women in the Avon area of southwest England, with expected dates of delivery between 1 April 1991 and 31 December 1992 to take part. The original sample comprised 14062 live-born children, with 13968 surviving to 1 year. ALSPAC has collected data on a large range of socio-economic, environmental and health measures for both parents and children; data were collected using questionnaires, face-to-face assessments and linked health and education data. Recruitment of participants and data collection have been described in detail elsewhere (Boyd et al., 2013). The study website contains details of all the data that are available through a fully searchable data dictionary (http://www.bris.ac.uk/alspac/researchers/data-access/data-dictionary/). Ethical approval for ALSPAC was obtained from the Local Research Ethics Committees, and this study was monitored by the ALSPAC Ethics and Law Committee. Identification of developmental coordination disorder ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Classification of children with DCD in this cohort and derivation of the measures used for the inclusion and exclusion criteria has been described in detail previously (Lingam, Hunt, Golding, Jongmans, & Emond, 2009). Children were defined as having DCD by applying the inclusion (Criteria A and B) and exclusion criteria (Criteria C and D) derived from the Diagnostic and Statistical Manual of Mental Disorders 4th Edition Text Revision (DSM- IV-TR) and adapted for research using the 2006 Leeds Consensus Statement (American Psychiatric Association, 2000; Sugden, Chambers, & Utley, 2006). Children were classified as having DCD at 7–8 years if they met all four criteria: (A) marked impairment of motor coordination, (B) motor coordination impairment significantly interferes with academic achievement or activities of daily living (ADL), (C) absence of another neurological/visual disorder, and (D) absence of severe learning difficulty (IQ > 70). DCD was defined in this cohort prior to the publication of the Diagnostic and Statistical Manual of Mental Disorders 5th Edition (DSM-V; American Psychiatric Association, 2013). However, the criteria for DCD were not significantly altered between the two editions, and so the criteria used in this cohort are compatible with the definition of DCD in the more recent DSM-V. The ALSPAC coordination test, consisting of 3 subtests of the Movement Assessment Battery for Children Test (MABC; Henderson & Sugden, 1992), was applied by 19 trained examiners in a purpose-built research clinic at 7–8 years. Consistency was maintained by standardised training and ongoing supervision and training. The 3 subtests were chosen, following principal component analysis of the original standardization data set of the MABC, as they best represented the three domains of coordination: heel-to-toe walking (balance), placing pegs task (manual dexterity) and throwing a bean bag into a box (ball skills). Age-adjusted scores were derived by comparing the scores of each child to raw scores in the cohort itself. Those scoring under the 15th centile in the coordination test were considered to have a motor impairment, consistent with criterion A of the DSM- IV-TR definition (Schoemaker, Lingam, Jongmans, van Heuvelen, & Emond, 2013). Academic achievement was assessed using linked educational data. As part of the national curriculum in the UK, literacy testing is undertaken at 7 years (Key Stage 1), which involves a writing test with questions on English grammar, punctuation and spelling. The writing test is scored 1–4 (4 being the best, 2 being the expected level at this age): those who scored 1 were considered to have significant difficulties with writing. ADL were assessed using a parent-completed 23-item questionnaire at 6 years 9 months of age, containing items derived from the Schedule of Growing Skills II (Bellman, Lingam, & Aukett, 1996) and the Denver Developmental Screening Test II (Frankenburg & Dodds, 1967). The questions represented skills the child would have been expected to achieve by 81 months. It assessed for difficulties in developmentally age-appropriate skills with which children with DCD struggle, such as self-care, playing, and gross and fine motor skills. Age-adjusted scores were calculated by stratifying the child’s age and those scoring below the 10th centile were considered to have significant impairment. IQ was measured at 8.5 years by trained psychologists in a research clinic using a validated and shortened form of the Wechsler Intelligence Scale for Children-III (WISC-III; Wechsler, Golombok, & Rust, 1992; Connery, Katz, Kaufman, & Kaufman, 1996). Alternate items were used for all subtests, with the exception of the coding subtest which was administered in its full form. Using the look-up tables provided in the WISC-III manual, age-scaled scores were obtained from the raw scores and total scores were calculated for the Performance and Verbal scales. Prorating was performed in accordance with WISC-III instructions. IQ was assumed to be stable over time (Schneider, Niklas, & Schmiedeler, 2014). Those with an IQ < 70 and those with known visual, developmental or neurological conditions were excluded from case definition of DCD (criterion C and D). Children who scored below the 15th centile on the ALSPAC coordination test (criterion A), and had significant difficulties with writing in their Key Stage 1 handwriting test or were below the 10th centile on the ADL scale (criterion B), were defined as having DCD. At 7–8 years, a cohort of 6902 children had all the data required for full assessment, and 329 children met the criteria for DCD. Educational achievement ~~~~~~~~~~~~~~~~~~~~~~~ ALSPAC obtained linked educational data from the National Pupil Database, a central repository for pupil-level educational data, and from the Pupil Level Annual School Census, which captures pupil-level demographic data about special needs support. The linked educational data in ALSPAC covers pupils in England and in state-funded schools; therefore pupils from Wales or in independent schools are excluded from this analysis. Academic achievement at 16 years was assessed using the results from the General Certificate of Secondary Education (GCSE) at Key Stage 4. These are the national achievement exams, covering a range of subjects (English and Maths are mandatory), undertaken at the end of compulsory schooling by all children in state schools in England. They are marked by anonymous external examiners and given grades ranging from A*-G. For the purpose of our analysis, we dichotomised the cohort into those who did and did not achieve 5 or more GCSEs graded A*-C, which is a widely used marker of performance in the UK. Special Educational Needs (SEN) provision data were obtained at 9–10 years (year 5, end of primary school) and 11–12 years (year 7, start of secondary school). At 9–10 years old, SEN status was divided into six groups. ‘No SEN provision’ indicated no formal support was being provided. ‘SEN provision without statement levels 1–4’ indicated extra support was being provided, with increasing level indicating increasing level of severity. ‘Statement of SEN’ indicated that a statutory assessment by the Local Authority had been undertaken and a statement of extra support required had been created. SEN status at 11–12 years was categorised into four main groups: ‘no SEN provision’, ‘School action’ (which indicated the child had been recognised as not progressing satisfactorily and extra support was being provided internally), ‘School action plus’ (which indicated the child had not made adequate progress on ‘school action’ and the school had sought external help from the Local Authority, the National Health Service or Social Services), and ‘Statement of SEN’. For the purpose of our analyses, at both time points, SEN status was dichotomised: those receiving no formal support at all and those receiving some level of support. Confounding variables ~~~~~~~~~~~~~~~~~~~~~ Gender, gestation, birthweight, socioeconomic status and IQ are known to be associated with DCD (Lingam et al., 2009), and have well-established links with academic performance. They were therefore selected as confounders for the multi-variable analysis. Gender, gestation and birthweight were extracted from the birth records in ALSPAC. Family adversity was measured using the ALSPAC Family Adversity Index (FAI). This is derived from responses to a questionnaire about childhood adversity and socio-economic status which mothers completed during pregnancy. The index comprises 18 items which are assigned a score of 1 if adversity is present and 0 if it is absent, giving a total possible score of 18. The FAI includes the following factors: age at first pregnancy, housing adequacy, basic amenities at home, mother’s educational achievement, financial difficulties, partner relationship status, partner aggression, family size, child in care/on risk register, social network, known maternal psychopathology, substances abuse and crime/convictions. Co-occurring difficulties ~~~~~~~~~~~~~~~~~~~~~~~~~ Reading ability (Lingam et al., 2010), social communication difficulties (Conti-Ramsden et al., 2008) and hyperactivity/inattention (Loe & Feldman, 2007) were selected as important conditions to be adjusted for in the multi-variable analysis, as they often co-occur with DCD and can impact on educational achievement. Reading ability was measured in research clinics at 13.5 years using the Test of Word Reading Efficiency (TOWRE), a short test which measures an individual's ability to pronounce printed words and phonemically regular non-words accurately and fluently. Words are used to assess sight word reading efficiency and non-words to assess decoding efficiency. The scoring is based on the number of words/non-words read quickly and correctly during 45 s (Torgesen, Wagner, & Rashotte, 1999). Social communication difficulties were measured at 16.5 years using the Social and Communication Disorders Checklist (SCDC; Skuse, Mandy, & Scourfield, 2005), completed by the main caregiver. It comprises 12 items relating to the child’s social communication ability and cognition, each with 3 responses (not true, quite or sometimes true/very or often true, scoring 0/1/2 respectively). Answers are summed to give a score between 0 and 36. A score of 9 or above was used to predict social communication difficulty trait (Humphreys et al., 2013). Hyperactivity/inattention was measured at 15.5 years using the self-reported Strengths and Difficulties Questionnaire hyperactivity-inattention subscale (SDQ; Goodman & Goodman, 2009). This subscale of the questionnaire comprises 5 items relating to behaviours indicating hyperactivity, each with 3 responses (not true/somewhat true/certainly true, scored 0/1/2 respectively). A score in the top decile of the ALSPAC cohort was taken as indicating significant hyperactivity (Goodman, 2001).","The two-sample test for proportions, Student’s t-test and the Mann-Whitney U test were used, where appropriate, to assess differences between groups. Logistic regressions were used to assess the impact of DCD on educational achievement. Multi-variable models were created to adjust for the effect of the confounding variables and mediating variables sequentially. Covariates significant at the 5% level in the univariate analyses were included in the multivariable models. Model 1 adjusted for gender, birthweight, gestation and socioeconomic status. Model 2 adjusted for IQ. Model 3 adjusted for reading ability. Model 4 adjusted for the developmental traits of social communication difficulties and hyperactivity/inattention. Pearson goodness of fit chi-squared tests were used to ensure satisfactory model fit. Multiple imputation using chained equations (ICE) was used to impute missing data in the covariates only (Appendix A in Supplementary material). Analysis of imputed datasets helps to minimise attrition bias and improve precision of estimates (Sterne et al., 2009). Logistic regressions were used to determine which variables strongly predicted missingness and these were included in the ICE prediction models. Twenty imputations were performed. All analyses were performed using Stata v. 14.1 (StataCorp, College Station, TX, USA). An alpha level of 0.05 was used for all statistical tests. Sample characteristics ~~~~~~~~~~~~~~~~~~~~~~ Of those previously assessed for DCD at age 7 years, educational data at 16 years were available for 5709 adolescents, including 284 (4.9%) of those who met the criteria for DCD. These represented 86% (284/329) of those originally diagnosed with DCD. Characteristics of those with and without educational outcome data available are shown in Appendix B in Supplementary material. Characteristics of the adolescents with DCD and controls with educational data available are compared in Table 1. In this follow-up cohort, when compared to controls, those with DCD were more likely to be male, have been born prematurely, have a low birth weight, have a lower IQ and have experienced greater family adversity. Academic achievement and SEN status ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There were substantial differences in educational achievement and SEN status between the controls and DCD group (Table 2). The control group achieved a median of 7 GCSEs graded A*-C, with 70% overall achieving 5 or more GCSEs at this level. In contrast, those with DCD achieved a median of 2 GCSEs graded A*-C, with 39% overall achieving 5 or more GCSEs. Fig. 1 illustrates the skewed distribution of GCSE attainment in the DCD group. When results were stratified by gender, in the control group girls performed significantly better than boys in achieving 5 or more GCSEs graded A*-C (76% vs. 65%, z = 8.89, p < 0.001). In the DCD group, girls were also more likely than boys to achieve 5 or more GCSEs graded A*-C, but this association did not reach significance (46% vs. 35%, z = 1.84, p = 0.064). The majority of those in the control group were not receiving any SEN provision at both time points; 895/5165 (17%) of controls received some form of SEN provision during their school career. In contrast, of those in the DCD group, 175/273 (63%) received some form of SEN support (z = −18.77, p < 0.001). In the DCD group, 15% were in receipt of support in primary school only and 9% in secondary school only, with 39% being recognised as needing extra support in both primary and secondary school. A sub-group analysis was performed to compare educational achievement for those who received SEN support and those who did not (Table 3). Differences between the groups were assessed using the two-sample test for proportions. Controls who received no SEN support were no more likely to achieve 5 or more GCSEs graded A*-C than those with DCD who received no SEN support (76% vs. 70%, z = 1.37, p = 0.17). Controls who received SEN support at any time were more likely to achieve 5 or more GCSEs graded A*-C than those with DCD who received SEN support at any time (27% vs. 16%, z = 3.00, p < 0.01). In the control group, those who received no SEN support were more likely to achieve 5 or more GCSEs graded A*-C than those who received SEN support at any time (76% vs. 27%, z = 28.10, p < 0.001). This was similar to the DCD group: those who received no SEN support were more likely to achieve 5 or more GCSEs graded A*-C than those who received SEN support at any time (70% vs. 16%, z = 6.85, p < 0.001). Co-occurring difficulties ~~~~~~~~~~~~~~~~~~~~~~~~~ Table 4 details the co-occurring difficulties reported for the control and DCD groups. Those with DCD were more likely to have difficulty with reading on both the word and non- word tests, social communication difficulties and hyperactive-inattentive behaviours. Multi-variable logistic regression ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The results of the multi-variable logistic regressions, using the multiple imputation dataset, are shown in Fig. 2. Results of the analyses of all available data are presented in Appendix C in Supplementary material. In the unadjusted model, compared to their peers, individuals with DCD were much less likely to achieve 5 or more GCSEs graded A*-C (Odds Ratio [OR] 0.27, 95% Confidence Interval [CI] 0.21–0.34). Even after adjusting for confounding variables and IQ, they were still significantly less likely to achieve 5 or more GCSEs (OR 0.60, 95%CI 0.44–0.81). After adjusting for reading ability, this association was attenuated (OR 0.73 95%CI 0.52–1.01) and was attenuated further after adjustment for poor social communication skills and hyperactivity/inattention (OR 0.78, 95% CI 0.55–1.10). Summary of results ~~~~~~~~~~~~~~~~~~ This prospective study using a large population-based UK cohort and strict diagnostic criteria has shown that children with DCD are much less likely to achieve 5 or more GCSEs graded A*-C in secondary school when compared to their peers. Children with DCD were more likely to have difficulties with reading, social communication and hyperactivity/inattention, and these problems contributed to poor achievement in school. Those with DCD were more likely to be identified as requiring formal SEN support than controls, although over a third of those with the condition were not identified as requiring formal SEN support. Adolescents with DCD without SEN support did not perform significantly worse in their GCSEs than controls without SEN support. Conversely, adolescents with DCD who were receiving SEN support performed worse in their exams than controls with SEN support. Discussion of results ~~~~~~~~~~~~~~~~~~~~~ The proportion of pupils achieving 5 or more GCSE qualifications at grades A*-C is not only used as an attainment indicator in the UK (Department for Education, 2016), but at an individual level it is often specified as a requirement to obtain places in further post-16 school education, in college or in certain jobs and apprenticeships. Poor performance in GCSEs can therefore have a large negative impact on an individual’s life prospects, depending on what they wish to do, as it may limit options for the next steps in their educational or vocational development. It is known that median hourly pay rate increases incrementally with increasing level of qualification, and that those with a higher level of qualification are more likely to be in skilled and better paid jobs (Office for National Statistics, 2011). Therefore, those with DCD are at a considerable disadvantage not only educationally, but also economically in their future careers. DCD is one of the more prevalent but less well-understood developmental conditions, and so this impact on life course is concerning. As well as accounting for the effects of important confounders and IQ on the relationship between DCD and educational achievement, we assessed the impact of co-occurring difficulties with reading, social communication and hyperactivity/inattention. As the effect of these traits were consecutively adjusted for in the multi-variable model, the difference between the control and DCD group was attenuated, which is reflected in the sequentially attenuated odds ratios from model 2 to model 4. This suggests that co-occurring developmental traits, which are common in DCD (Lingam et al., 2009), may account for some of the under-achievement at GCSE level in those with DCD compared to controls. In the fully adjusted model, those with DCD were 22% less likely to achieve 5 or more GCSEs compared to controls. However, it should be noted that the confidence intervals are substantially wider in models 3 and 4, which likely reflects the greater attrition in response rates for the variables in these models. These results emphasise the need for a holistic view of a child with DCD – problems with school work may not just be solely due to motor deficits and they should not be viewed as ‘just clumsy’. Co-occurring difficulties in other developmental domains should be actively sought, and if present, addressed in any support provided. Of those who met the criteria for DCD, 37% did not receive any formal SEN provision. This corroborates what has been reported in qualitative work with parents (Missiuna et al., 2005) and illustrates how DCD is often a hidden disability, as motor deficits can be subtle and learning impairments difficult to recognise. This is important to realise because if missed, children with DCD may start to avoid situations which draw attention to their poor coordination as they become more aware of their motor difficulties, to try to minimise embarrassment or bullying. Previous work on this cohort has shown that adolescents with DCD report more bullying and less supportive friendships than their peers (Harrowell et al., 2017). This, along with reducing engagement at school, will contribute to poor perceived efficacy and low self-esteem (Skinner & Piek, 2001). The combined impact of these secondary problems, as well as the primary motor deficit, will ultimately compound poor performance in school. Poor academic performance may further exacerbate engagement at school, leading to a complex vicious circle. Within the DCD group, those with formal SEN provision performed significantly worse in their exams than their counter-parts without SEN provision. The explanation for this is likely two-fold. Firstly, those with the most severe motor deficits, who will have most difficulty in school, are likely to be recognised as requiring extra help. Secondly, it may be that those receiving SEN support were those with co-occurring difficulties, whose multiple problems compound their poor performance, and are more likely to be recognised as needing support. A previous study which found a similar pattern in children with specific language impairments; those who had a ‘Statement of SEN’ performed worse than those without one in their secondary school exams (Durkin, Simkin, Knox, & Conti-Ramsden, 2009). Implications of the findings ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Guidelines for diagnosis, assessment and intervention in DCD were published in 2012 based on the best available evidence (Blank, Smits-Engelsman, Polatajko, & Wilson, 2012), and were more recently adapted for the UK context following a multi-disciplinary consultation (Barnett, Hill, Kirby, & Sugden, 2015). The evidence provided by this study supports the need to raise awareness of the condition and its consequences not only among educational and clinical health professionals, but also policy makers. Clearly defined pathways for children with DCD are required to coordinate services appropriately. The work of organisations like the Dyspraxia Foundation (Dyspraxia Foundation, 2016) and Movement Matters (Movement Matters UK, 2016) provide information about DCD to families, carers, teachers and clinicians, with the aim of improving awareness of and support for individuals with condition. Since March 2014, policy for SEN provision in the UK has been reformed, with the ‘Statement of SEN’ being replaced by an incorporated ‘Educational, Health and Care Plan’ in the Children and Families Act 2014 (Department for Education, 2014). How this will impact upon care of those with DCD remains to be seen, and implementation is likely to vary between regions. It is hoped that the results of the work on DCD in ALSPAC will contribute to improving awareness and understanding of DCD and subsequently improved planning of support. Increased awareness and understanding of the condition needs to be followed by improved intervention. The evidence base for what interventions work best in DCD is limited. Interventions which are task-orientated and individualised appear to have the largest effect on motor skills, based on a recent systematic review of high quality randomised controlled trials (Preston et al., 2016). As highlighted in this study, however, motor skills may only be one domain that requires attention in a child identified as having DCD. More research is needed on the benefit of interventions aimed at addressing other developmental domains concurrently, such as social communication skills (Piek et al., 2015). Further, there is a paucity of research into the needs of adolescents and adults with DCD. Many young people with DCD continue to experience motor and psychosocial difficulties after they have finished school and it is important that they are not forgotten and are supported to gain employment (Kirby, Edwards, & Sugden, 2011; Kirby, Williams, Thomas, & Hill, 2013). Strengths and weaknesses of the study ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The key strength of this study is the use of a large, prospective, population-based sample which is broadly representative of the UK (Boyd et al., 2013), and avoids the biases introduced in clinical samples. Due to the large amount of data collected in ALSPAC, we were able to strictly define our cases of DCD according to DSM-IV-TR criteria, and also to assess the impact of important co-occurring problems on our main outcome. The use of linked educational data was also a strength of the study. The main weakness of this study, as with all prospective cohorts, is attrition. Through the use of linked education data, we were able to minimise loss-to-follow-up in our primary outcome measure. For those who had GCSE data available (n = 5709), 81% had all the covariates available for multi- variable models 1 and 2. Those who did not have GCSE data available at 16 years were more likely to have higher IQ and a higher level of maternal education (Appendix A in Supplementary material), a pattern which was seen in both the control and DCD groups and so should not systematically affect our results and thus our conclusions. The largest attrition occurred in model 4 (social communication deficits and hyperactivity), which relied on questionnaires filled out between 14 and 16 years. In ALSPAC, it is recognised that those who are lost to follow-up in clinics and questionnaires tend to come from lower socio-economic backgrounds, have lower IQ and are more likely to be male (Boyd et al., 2013). These are factors which are also associated with DCD and so differences between the groups may be under-estimated. We have attempted to minimise any bias introduced by missing data in the covariates by the use of multiple imputation, a well-validated statistical technique (Sterne et al., 2009). Another drawback is that no measure of motor competence was performed in adolescence. Motor skill difficulties identified in childhood do not always persist into adolescence (Cantell et al., 1994) and so there may be some in the DCD group who no longer have significant motor deficits. Furthermore, we were only able to comment on whether the child had been identified as requiring formal SEN provision, but not the level or type of support, or whether intervention had any impact on the core deficits in DCD.","DCD is an important but poorly understood developmental condition that has a stark impact upon educational achievement at the end of secondary schooling, which will affect an individual’s future prospects. Co-occurring developmental conditions are common, and may contribute to poor educational outcomes. Addressing these co-occurring difficulties, as well as motor deficits, is necessary to improve outcomes. Increased awareness and understanding of the condition amongst healthcare and educational professionals and policy makers is vital to improve the support provided for those with DCD.","The authors declare that they have no conflict of interests relating to this article.","This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. The UK Medical Research Council and the Wellcome Trust (Grant ref: 102215/Z/13/Z) and the University of Bristol provide core support for ALSPAC."],["The past decades have witnessed a huge interest in uncovering the neural bases of intelligence (e.g., Stelmack, & Houlihan, 1995; Stelmack, Knott, & Beauchamp, 2003). This study investigated the influence of transcranial alternating current stimulation (tACS) on fluid intelligence performance and corresponding brain activation. Previous findings showed that left parietal theta tACS leads to a transient increase in fluid reasoning performance. In an attempt to extend and replicate these findings, we combined theta tACS with fMRI. In a double-blind sham-controlled experiment, N = 20 participants worked on two intelligence tasks (matrices and paper folding) after theta tACS was applied to the left parietal cortex. Stimulation-induced brain activation changes were recorded during task processing using fMRI. Results showed that theta tACS significantly increased fluid intelligence performance when working on difficult items in the matrices test; no effect was observed for the visuo-spatial paper folding test. Whole-brain analyses showed that left parietal brain stimulation was accompanied by lower activation in task-irrelevant brain areas. Complemental ROI analyses revealed a tendency towards lower activation in the left inferior parietal cortex. These findings corroborate the functional role of left parietal theta activity in fluid reasoning and are in line with the neural efficiency hypothesis. --------------------------------------------------------------------------------","Intelligence is associated with diverse relevant real-life outcomes such as educational accomplishment, occupational performance, and even health (Deary, 2012). There is a long- standing research tradition to investigate the neural bases of intelligence (e.g., Ertl & Schafer, 1969; Stelmack & Houlihan, 1995). Neurophysiological models of intelligence emphasize the role of the fronto-parietal-network. The parieto-frontal integration theory of intelligence (P-FIT; Jung & Haier, 2007) postulates a four-phase information processing model which highlights the importance of frontal and parietal brain areas and the associated communication patterns between those. Another important theory is the neural efficiency hypothesis. It posits that more intelligent people use their brain resources more efficiently as compared to less intelligent individuals in terms of lower brain activation (Haier et al., 1988) or faster neural transmission time (Stelmack, Knott, & Beauchamp, 2003). A review of relevant findings by Neubauer and Fink (2009) further suggests that neural efficiency is restricted to tasks of low to moderate task difficulty, whereas highly able individuals may even invest more cortical resources in very difficult tasks (see also, Dunst et al., 2014). More recently, there has been an increasing interest to use and extend our understanding of the brain in attempts to improve intelligence via brain stimulation (Enriquez-Geppert, Huster, & Herrmann, 2013). Stimulating intelligence ~~~~~~~~~~~~~~~~~~~~~~~~ Due to the social relevance of gf, the issue was raised whether there is a way to increase it through training or stimulation. However, recent reviews and meta-analyses come to inconsistent conclusions, either showing that intelligence scores cannot be increased through cognitive training (Haier, 2014; Melby-Lervåg, Monica, & Hulme, 2016), or finding support for the beneficial effect of systematic cognitive training (Au et al., 2015; Karbach & Verhaeghen, 2014). For example, Jaušovec and Jaušovec (2012) reported that working memory training does not only promote performance in a gf task but also changes the associated electrocortical brain activation. One particular issue of the training approach is the missing evidence for transfer effects, which is sometimes challenged by the similarity between assessment and training tasks (Shipstead, Redick, & Engle, 2012). This methodological problem can be circumvented with direct stimulation of brain activation. Available research mostly used transcranial electrical stimulation (TES; Kuo & Nitsche, 2012), which It involves either direct (tDCS) or alternating current (tACS). TDCS manipulates neuronal activity via depolarization or hyperpolarization of cell membranes. TACS influences cortical activity via EEG frequency-specific oscillatory current. tACS leads to EEG-synchronization of neural activity in the respective frequency band, which may facilitate efficient cortical communication patterns (Zaghi, Acar, Hultgren, Boggio, & Fregni, 2010). Much research has focused on stimulation of the dorsolateral prefrontal cortex (DLPFC). For tDCS, DLPFC stimulation was associated with positive effects on cognitive processes such as language, attention, perception, executive functioning, and memory processes (Jacobson, Koslowsky, & Lavidor, 2012; Utz, Dimova, Oppenländer, & Kerkhoff, 2010). For tACS of the DLPFC, Santarnecchi, Polizzotto, Godone, Giovannelli, Feurra et al. (2013) found that gamma stimulation leads to a shortening of (correct) response latencies in a gf task. A growing number of studies examined left-parietal stimulation effects on working memory and intelligence. Left-parietal tDCS was found to enhance performance in a verbal memory task (Jacobson, Goren, Lavidor & Levy, 2012). For tACS, Pahor and Jaušovec showed that left parietal theta tACS leads to an increase in gf performance and accompanying changes in cortical activation (Pahor & Jaušovec, 2014). The same stimulation setting also increased working memory capacity and a decreased P3 component latency, which could reflect a faster allocation of attention-related cognitive resources (Jaušovec & Jaušovec, 2014). Aims of this study ~~~~~~~~~~~~~~~~~~ Previous findings showed that left parietal theta tACS leads to a transient increase in fluid reasoning performance (Pahor & Jaušovec, 2014; Jaušovec & Jaušovec, 2014). In an attempt to replicate and extend these findings, we combined theta tACS with fMRI in a double-blind design. Available evidence indicates that left-parietal theta activity might play a causal role for gf-related demands such as executive functioning (Sauseng, Griesmayr, Freunberger, & Klimesch, 2010), working memory capacity (Postle et al., 2006). Neurophysiological stimulation effects are studied by analyzing concurrent brain activation changes with fMRI. Finally, we test the role of task difficulty as a moderating variable in the relationship between brain stimulation and brain activation (Neubauer & Fink, 2009). We expect that left parietal theta stimulation increases fluid intelligence performance (Pahor & Jaušovec, 2014) and affects brain activation in intelligence-related brain areas according to the P-FIT model (Jung & Haier, 2007). Additionally, the neural efficiency hypothesis predicts that increased intelligence performance is associated with reduced activation of brain areas that are not considered central for intelligence (Basten, Stelzel, & Fiebach, 2013; Neubauer & Fink, 2009).","25 individuals were recruited from a pre-tested pool of participants (Jauk, Benedek, Dunst, & Neubauer, 2013). All participants were right-handed, had normal or corrected-to- normal vision, and no self-reported history of CNS-affecting drugs, mental or neurological diseases. Five participants were excluded due to excessive missing data or failure to complete both test sessions. The final sample hence consisted of 20 participants (11 women; average age = 24.85; SD = 3.30). The sample was kept homogeneous with respect to age (18–30 years), intelligence as measured with the intelligence-structure-battery (INSBAT; for details see Dunst et al., 2014; M = 106.09; SD = 7.34), and educational level (students) to enable a sensitive test of within-subject stimulation effects. Participants did not report any medical treatments or health problems and gave written informed consent. The study was approved by the local ethics committee. Design ~~~~~~ The study was a double-blind sham-controlled experiment. One experimenter only operated the stimulation device and had no interaction with the participant, whereas another one (who was blind to the stimulation condition) instructed participants. Verum (i.e., active) and sham (i.e., placebo) stimulation conditions varied within subjects in a counterbalanced fashion. The two sessions were separated by 28 days to control for potential influences of different phases of the menstrual cycle (Amin et al., 2006). Dependent variables were performance in the two gf tasks as well as BOLD responses in fMRI. Tasks and procedure ~~~~~~~~~~~~~~~~~~~ The experiment was carried out in two sessions. Participants received a standardized instruction and then either sham or verum tACS was applied for 15 min (see below), followed by a questionnaire about intensity and duration of stimulation induced sensations. After the stimulation, the participants were led to the scanning room where they performed two intelligence tasks. The matrices task was based on the Raven's progressive matrices (RPM; Court & Raven, 1995), slightly modified to the requirements of neurophysiological investigations (Pahor & Jaušovec, 2014). The test consisted of 50 items – 18 easy (set B of the CPM, and APM items 1–3), and 32 difficult (APM items 4 to 35). The 50 test items were divided into two parallel forms each consisting of 25 items (10 easy items; 15 difficult items) which were counterbalanced between the sham/verum conditions. As in the original RPM, items were presented in a fixed sequence reflecting increasing task difficulty. Each trial of the RPM started with a jittered fixation cross period (6–10 s). Then, the item was presented for 6 s (easy items) or 10 s (difficult items) together with four response alternatives depicted below, followed by a 3 s response phase (indicated by red interrogation marks presented under the figure). During this, participants had to press one of four buttons corresponding to the four response options. Total task duration of the RPM scanner task was about 9 min. This modified RPM has been shown to correlate substantially with the WAIS-R (r = 0.56) (Jaušovec & Jaušovec, 2012), suggesting that content validity is not impaired by the shorter item presentation time. The cross-form consistency of the modified RPM had been proven in a larger sample (rA,B = 0.73; Pahor & Jaušovec, 2014). As a second intelligence measure, we used the paper folding task (PFT) from the Stanford-Binet test (Cowan et al., 2011). Participants had to judge which of the four presented figures on the right side corresponded to that one on the left side after variable steps of folding and cutting (Jaušovec & Jaušovec, 2012). Again, we used two parallel versions which consisted of 20 items each. Items were presented in a fixed quasi-randomized sequence. For analyses considering task difficulty, these items were divided in to 10 easy and 10 difficult items based on effective task performance. Presentation parameters were the same as for the RPM except that the duration of stimulus presentation was 7 s per item here, and fixation cross period varied randomly from 5 to 9 s. Total duration of the PFT was about 6 min. Again, cross-form consistency of the modified PFT was established in a previous study (rA,B = 0.71, Pahor & Jaušovec, 2014). In both tasks (RPM and PFT) a trial was scored as solved when the correct response alternative was selected before timeout. The total duration of the experiment including instruction, stimulation, and scanning was around 65 min. Electrical stimulation ~~~~~~~~~~~~~~~~~~~~~~ We used a battery-operated stimulator system (DC-stimulator plus, Neuroconn, Ilmenau, Germany). The stimulating electrodes (5 × 7 cm) were attached to the scalp using a rubber band placed over the electrode and attached under the chin. This procedure prevented movement of the electrodes during the experiment. The electrodes were covered by saline soaked sponges, which reduced the electrode impedance. Target electrode was placed over the left parietal location (P3), and the return electrode was placed on Cz. This electrode positioning was chosen because of the key role of the left parietal cortex for gf and gf- related functions (e.g., working memory capacity; Jung & Haier, 2007; Cowan et al., 2011; Chein & Fiez, 2010), and first evidence on effects of left parietal tACS on fluid reasoning (Pahor & Jaušovec, 2014; Jaušovec & Jaušovec, 2014; Postle et al., 2006). Stimulation waveform was sinusoidal without DC offset and a 0° relative phase. The impedance level was kept below 10 kΩ throughout the entire stimulation period. The applied oscillating currents corresponded to mean theta frequency in previous studies (Jacobson, Goren, Lavidor, & Levy, 2012; 5 Hz); current intensity was 1500 μA. In verum condition, tACS was applied for 15 min. The current was ramped up and down over the first and last 15 s of stimulation. In sham condition, the procedure and stimulation parameters were the same as in the verum condition except for the duration of stimulation, which was applied for only 60 s in the beginning and then turned off. Since participants feel stimulation related sensations (e.g. itching) only in the beginning of tACS, this approach prevents individual awareness of the stimulation conditions (Nitsche et al., 2008). MRI data acquisition ~~~~~~~~~~~~~~~~~~~~ Imaging was performed on a 3.0-T Tim Trio system (Siemens Medical Systems, Germany) using a 32-channel head coil. BOLD-sensitive T2*-weighted functional images were acquired using a single shot gradient-echo EPI pulse sequence (TR = 2400 ms, TE = 30 ms, flip angle = 90o, slice thickness = 3.5 mm, matrix size = 68 × 68, FOV = 240 mm, 35 slices per volume). The first two volumes after each scanner pause were discarded to allow for T1 equilibration effects. Field maps were created from a double echo gradient-echo pulse sequence (31 slices, TE1 = 4.92 ms, TE2 = 7.38 ms, TR = 400 ms, slice gap = 0.9 mm, slice thickness = 3.5 mm, matrix size 68 × 68, FOV = 240 mm). Visual stimuli were presented onto a screen using the Software Presentation (Neurobehavioral Systems, Albany, CA) and viewed through a mirror attached to the head coil. MRI data analysis ~~~~~~~~~~~~~~~~~ Functional MRI data analysis was performed using SPM 8 software (Wellcome Department of Imaging Neuroscience, London, UK). Preprocessing steps included field map correction, motion correction, slice time acquisition correction, spatial normalization to an averaged EPI template, and smoothing with a 7-mm full-width at half maximum Gaussian kernel. For each task, scans from the verum and sham fMRI sessions were modeled together in a fixed- effects model including the conditions REST (fixation epochs), VERUM, and SHAM (task performance following verum or sham stimulation, from stimulus onset to onset of response period). Motion parameters were included in the model as regressors of no interest. Task- specific brain activation was modeled with a conjunction of both stimulation and sham conditions (VERUM + SHAM > 0) against implicit baseline (Poline, Kherif, Pallier, & Penny, 2007). Stimulation-specific brain activation was modeled with the contrast of conditions (VERUM vs. SHAM). For the examination of difficulty-specific effects another model was run considering easy and difficult trials separately. At the second level, a random effects analysis was performed computing one-sample t-tests for the subject-specific statistical parametric maps obtained at the first level. Whole-brain results for task-specific effects (VERUM + SHAM) are reported using a conservative criterion of voxel-wise p < 10− 8 (uncorrected) with cluster size of k ≥ 25. For the analysis of stimulation-specific effects (VERUM vs. SHAM), clusters are only reported if they are significant on voxel level (p < 0.0001, uncorrected) and exceed a minimum cluster size of 5 voxels. Finally, a region of interest (ROI) analysis was computed to determine the direction (activation or deactivation) and magnitude of changes in regions showing significant task-specific activation patterns (VERUM + SHAM) using MarsBaR 0.43 (Brett, Anton, Valabregue, & Poline, 2002). The ROIs were functionally defined based on the task-specific activation patterns (Poldrack, 2007). Behavioral results ~~~~~~~~~~~~~~~~~~ The self-reported stimulation-induced sensations during the tACS sessions did not differ significantly between sham and verum tACS settings (Wilcoxon Z19 = − 0.53; ns.). Stimulation effects on intelligence in the RPM task were analyzed with an ANOVA with the within-subject conditions STIM (sham/verum) and DIFFICULTY (easy, difficult). The interaction effect between STIM and DIFFICULTY was significant (F(1, 19) = 4.86, p < 0.05; eta2part = 0.20). Post-hoc paired-sample t-tests showed that participants solved significantly more difficult RPM items (28.44%, M (SD) = 9.10 (1.87)) when verum tACS was applied than after sham stimulation (24.69%, M (SD) = 7.90 (2.53); t(19) = 2.40, p = 0.03, d = 0.53), but there was no stimulation effect for easy items (45.83%, M (SD) = 8.25 (1.07), and 46.39%, M (SD) = 8.35 (0.88) for verum and sham stimulation, respectively; t(19) = − 0.42, p = 0.68; Fig. 1). Supplemental Fig. S1 shows the individual data. Performance increases following verum tACS were evident in nine of 20 individuals; seven displayed no change, four displayed (slight) decreases. Interestingly, especially those individuals with low baseline (sham) performance appear to have benefited from the stimulation. The same GLM procedure was conducted for the PFT, but we observed no effects related to the stimulation condition or task difficulty. Task-specific brain activation Task-specific brain activation in the RPM task included 13 clusters, with seven clusters showing relatively stronger activation and six clusters showing lower activation relative to baseline (Table S1). Task-specific activation was observed bilaterally in the insula, in the left superior/inferior parietal lobe, in the thalamus, the right lingual gyrus, and the postcentral gyrus. Task-specific brain activation in the PFT included 13 clusters with 8 clusters showing stronger activation and 5 clusters showing lower activation relative to implicit baseline (Table S2). Task-specific activation was observed in the middle and the left inferior occipital gyrus as well as in the left superior/inferior parietal lobe, the right insula, the right lingual gyrus, the right supramarginal gyrus, and in the right fusiform gyrus. Stimulation effects on brain activation In both tasks, brain stimulation induced no relative increases but only relative decreases in brain activation (see VERUM < SHAM). In the RPM task, verum stimulation was associated with lower activation in the left middle occipital gyrus as well as in right occipital and frontal lobes (Table 1). In the PFT task, lower activation was observed bilaterally in the precuneus, the left inferior temporal gyrus, and right regions of the cerebellum. In both tasks, stimulation- induced activation changes did not overlap with the task-positive brain regions but partly overlapped with the task-negative brain regions observed for these tasks. We further examined whether stimulation effects are different for easy and difficult items, as it was the case at the behavioral level in the RPM task. The observed activation differences were essentially the same when considering only difficult items, but no stimulation effects were observed for easy items. ROI analyses Finally, we examined the effect of stimulation on brain activity in terms of signal change in the seven clusters showing increased task-specific activation (Table S1), separately for trials classified as easy or difficult. Here we observed a tendency towards a difficulty-depended stimulation effect in the inferior parietal lobe verum stimulation tended to be associated with lower activation as compared to sham stimulation for difficult items (p = 0.09) but not for easy items (p = 0.71). No stimulation effects were observed in PFT task.","We set out to examine the effects of left-parietal theta tACS on intelligence test performance and its respective neurophysiological bases. We found that theta tACS can enhance performance in a gf task (as measured by Raven's Progressive Matrices) for difficult items; no stimulation effect was found for easy RPM items, or the PFT. The stimulation effect was accompanied by distinct brain activation changes. Now, we will discuss reasons and implications of these findings. This study replicates previous research showing that left parietal theta tACS leads to increased task-specific reasoning performance (Pahor & Jaušovec, 2014). It is also in line with findings that left parietal theta tACS increased working memory capacity (Jaušovec & Jaušovec, 2014), as working memory is a central executive function underlying fluid intelligence (Benedek, Jauk, Sommer, Arendasy, & Neubauer, 2014). We presume that theta frequency reflects a general cognitive control mechanism, which might be of general importance for gf performance (Sauseng et al., 2010). Regarding the neurophysiological stimulation effects, left parietal theta stimulation was accompanied by lower brain activation in the right frontal lobe and bilaterally in the occipital lobe for the RPM, and in the precuneus, left inferior temporal gyrus, and right cerebellum in the PFT task. The missing overlap of stimulation-induced brain activation changes across tasks can be seen to corroborate the task-specificity of the stimulation effects. Furthermore, independent of the stimulation condition, there was little overlap in task-specific activation patterns (e.g. inferior parietal lobule, insula; Table S1; Table S2), which supports the notion that each of the tasks addresses distinct facets of intelligence performance. Notably, stimulation-induced decreases in brain activation primarily concerned brain areas that were not part of task- specific brain activation patterns (i.e., task-positive areas), but rather overlapped with task-negative areas (e.g., precuneus). This finding parallels research on neural efficiency: Studies have shown that higher intelligence is related to lower brain activation, particularly in task-negative brain regions (Neubauer & Fink, 2009; Basten et al., 2013). Our findings could thus be tentatively interpreted in terms of induced neural efficiency by means of tACS. A closer look suggests that activation decreases might also be related to task difficulty. Complemental analyses showed that stimulation-related differences in brain activation were only apparent when individuals worked on difficult, but not on easy items (mirroring the behavioral stimulation effects). The stronger deactivation following stimulation during difficult tasks might indicate an efficient downregulation of irrelevant cortical activity during phases of increased cognitive load. Interestingly, we did not observe stimulation-dependent brain activation increases in task-positive brain regions or parietal and frontal areas as predicted by the P-FIT (Jung & Haier, 2007). In contrast, ROI analyses revealed a weak tendency towards a stimulation- induced reduction of brain activation in the left inferior parietal cortex. While the left inferior parietal cortex is part of the task-positive network, it was also the stimulation location. Hence, theta stimulation may not necessarily translate to increased brain activation, but may even induce slight relative deactivation; at least at the stimulation site. However, these stimulation-specific effects did not survive FWE-correction, why they should be seen as first exploratory evidence for a potential impact of theta-tACS on the hemodynamic function measured through fMRI. Our findings add to evidence reported by Jaušovec and Jaušovec (2014), who observed a decreased P3 latency and increased theta power at left parietal brain regions after theta tACS was applied on the left parietal cortex (cf., Beauchamp & Stelmack, 2006). Moreover, Pahor and Jaušovec (2014) showed that parietal theta tACS was associated with a frontal theta power increase. How can these previous EEG findings be reconciled with the fMRI evidence? Scheeringa et al. (2009) showed that, during performance of a working memory task, increases in frontal theta power were correlated with BOLD decreases in regions that together form the default mode network. Thus, our finding of a deactivation of brain regions that are not considered essential for gf performance is generally consistent with the observed frontal theta power increase in a previous study (Pahor & Jaušovec, 2014), and it seems that higher theta power goes along with lower BOLD in independent brain regions. Altogether, our study suggests that the main mechanism underlying theta tACS-stimulation effects can be seen in brain activation decreases of task-irrelevant brain regions rather than increases in task- relevant regions, which is in line with the neural efficiency hypothesis. This finding is potentially consistent with the notion that intelligence is not associated with faster neural transmission at task-relevant regions (Stelmack et al., 2003). It should be acknowledged, however, that the current findings can only be interpreted in terms of a transient increase in the performance on a specific fluid intelligence task, which can be seen as an enhancement of a subfactor (namely fluid reasoning) rather than a general “intelligence boost”. Note also that stimulation effects were restricted to task performance on Raven items of higher difficulty. This includes the behavioral and brain activation effects and is consistent with a previous study (Pahor & Jaušovec, 2014). But why are findings specifically observed for difficult but not for easy Raven items? Easy items had item difficulties ranging between 0.76 and 1 indicating that they were solved by most of the participants. Hence, easy items may have less discriminatory power than more difficult items to discern between differences in reasoning performance. Also, complemental single-subject analyses showed that individuals with low baseline performance benefit the most from tACS, which has direct implications for the differential use of tACS and will hopefully stimulate future research. Finally, stimulation effects were specific to the matrices task in this study. A possible explanation could be that the RPM and the PFT focus on different facets of intelligence. Although the PFT is commonly considered a fluid intelligence task (e.g., Nusbaum & Silvia, 2011), it can also be seen to have a strong visual-spatial focus. Recent research showed that visual-spatial abilities are predominantly represented via right-hemispheric activation (Zacks, 2008), whereas fluid reasoning is mainly represented through left-hemispheric (Barbey et al., 2012) or bilateral activation patterns (Gray, Chabris, & Braver, 2003). A left parietal stimulation hence could have more effect on cognitive processes that are generally left-lateralized. Also, the participants performed the matrices task first, so it is possible that the power of the stimulation effect had already decreased when participants worked on the PFT. In conclusion, this study was the first to explore the neurophysiological basis of fluid intelligence via the combination of tACS and fMRI. Left parietal theta tACS was found to moderately increase performance in a fluid reasoning task when working on difficult items, and this was accompanied by deactivation of task-irrelevant brain regions. For future research, it would be exciting to study tACS effects on additional direct indices of brain functioning like neural transmission time (e.g., Stelmack et al., 2003).","This research was funded by a grant from the Austrian Science Fund (FWF): P23914."],["This study examined the relation between school poverty and educational attainment of adolescents, and tested whether personality trait agreeableness moderated this link. The sample consisted of 4236 adolescents, whose math abilities were assessed twice, at ages around 13/14 and 15/16. Agreeableness was assessed at age 13. School poverty was measured as the proportion of children eligible for free school meals in the school. The results showed a negative relation between school poverty and educational attainment, however, this negative relation was weaker for adolescents with higher levels of agreeableness. Specifically, in low poverty schools, agreeableness did not predict differences in educational attainment. The results were in line with the diathesis-stress model. This suggests that higher levels of agreeableness can contribute to resilience and better coping with contextual stressors in the school environment. --------------------------------------------------------------------------------","Many studies have linked contextual poverty in general and school poverty specifically to educational outcomes of individuals (Lacour & Tissington, 2011; Nieuwenhuis & Hooimeijer, 2016; Nieuwenhuis, Hooimeijer, van Dorsselaer, & Vollebergh, 2013; Portes & MacLeod, 1996). School poverty is negatively related to parental education. Therefore, the social networks within low SES schools consist of lower educated parents, and through these networks, less social capital is available, such as information about after-school programs. Furthermore, parents may not value or understand the benefits of formal education, resulting in students who are less prepared for education (Lacour & Tissington, 2011). Finally, higher poverty schools were found to have less qualified teachers on staff (Peske & Haycock, 2006). This suggests that low SES schools have fewer positive role models showing the benefits and transferring the importance of education. Because positive socialisation mechanisms are not in place in low SES schools, children may become less inclined to perform well. However, educational attainment is not uniform for all children in the same school, some perform better than others. This variation may be induced by differences in resilience, as described by the diathesis-stress model. Low SES schools can be experienced as stressful environments, however, some are better able to cope with environmental stressors than others (Magnusson & Stattin, 2006). Personality traits have been shown to be related to better coping with stressful environments (O'Brien & DeLongis, 1996), such as school poverty. In this case, resilient adolescents are expected to be affected less by school poverty than non-resilient adolescents. In high SES schools they are expected not to differ. Alternatively, the differential susceptibility model predicts that adolescents who are more malleable are more likely to be negatively affected by stressful environments than less malleable adolescents, but also more likely to be positively affected by positive environments (Belsky & Pluess, 2009). In this case, less malleable adolescents are expected not to be affected by the level of school poverty, while malleable adolescents are expected to have better educational outcomes in high SES schools and worse outcomes in low SES schools. Specifically, I have examined the moderating role of agreeableness, a personality trait that is related to being forgiving, patient, warm, considerate, and sympathetic (Goldberg, 1992). Agreeableness has been linked to lower levels of criminal behaviour and higher levels of community involvement (Ozer & Benet-Martinez, 2006; Roberts, Kuncel, Shiner, Caspi, & Goldberg, 2007). This suggests that adolescents with higher levels of agreeableness are less inclined to interact with deviant peers, and more likely to be involved in positive institutions at school, such as clubs or committees. Adolescents with lower levels of agreeableness are more inclined to participate in antisocial behaviour, which may be amplified in a stressful school environment. When adolescents do not function as expected in school, and the school does not foster their positive development, they may focus their attention towards other activities or deviant peer groups, where status attainment is reached through violent behaviour and anti-school attitudes (Ellis et al., 2012; Willis, 1977). For this follows that agreeableness could contribute to adolescents' resilience or malleability in the school environment. Both the diathesis-stress and differential susceptibility models lead to the following hypothesis: The relation between school poverty and educational achievement is moderated by agreeableness such that this relationship will be weaker for adolescents with higher levels of agreeableness (H1). Both models predict that in high poverty schools, adolescents with low agreeableness do worse, however, in low poverty schools they lead to two competing hypotheses. From diathesis- stress follows: In low poverty schools, adolescents with high levels of agreeableness do not differ from adolescents with low levels of agreeableness in their educational attainment (H2a). From differential susceptibility follows: In low poverty schools, adolescents with high levels of agreeableness have lower educational attainment than adolescents with low levels of agreeableness (H2b).","Participants were 4236 adolescents (52% females) from the Avon Longitudinal Study of Parents and Children (ALSPAC; initial recruitment: 14,541 pregnant women with expected delivery dates between 1991/04/01–1992/12/31; total sample: 15,458 fetuses, of which 14,701 were alive at age 1; Boyd et al., 2013). Adolescents' math abilities were assessed twice, at ages around 13/14 and 15/16, resulting in 6813 observations. Students were nested in 336 schools, mostly in the south-west of England. Please note that the study website contains details of all the data that is available through a fully searchable data dictionary (http://www.bris.ac.uk/alspac/researchers/data-access/data-dictionary/). Educational attainment Math scores on the standardised tests Key Stage 3 (age 13/14) and Key Stage 4 (age 15/16) were obtained from the National Pupil Database. The scales of the two scores were different, and were transformed for comparability using the proportion of maximum scaling (POMS; Little, 2013). This transformation retains the rank- order of individuals, while avoiding measuring mean-level changes. The formula used was POMS = (observed − minimum) / (maximum − minimum) (Moeller, 2015). For descriptive statistics and correlations of all variables, see Tables 1 and 2, respectively. Agreeableness Using the International Personality Item Pool (Goldberg, 1992), the Big Five personality trait agreeableness was assessed at 13 years and 6 months. Adolescents were presented with 10 statements, and where asked how well these statements described them on a 5-point answering scale. Cronbach's alpha: 0.72. Agreeableness was centred. School poverty The proportion of children in the school who were eligible for free school meals was used as a proxy for school poverty. This is a heavily studied measure of school poverty, and considered a good indicator (Gorard, 2012). The measure was assessed for both the schools adolescents attended at the two Key Stages. These data were obtained from the Annual School Census. The measure was standardised with 0 as mean and standard deviation 1 to make the interaction term easier to interpret. Control variables. First, sex was measured as female (1) and male (0). Second, parental education was measured as the average of the highest attained education of both parents. Education consisted of five categories: 0) (General) Certificate of Secondary Education ([G]CSE) levels D, E, F, or G; 1) vocational; 2) Ordinary Level (O Level) or GCSE levels A, B, or C; 3) Advanced Level (A Level); and 4) university degree. Third, race was measured as non-white (1) and white (0). Fourth, mother's age at delivery was measured as mother's age in years at the birth of the respondent (ranging from 16 to 44). Analyses ~~~~~~~~ To test the hypotheses, I used multilevel random-effects regression models, with time nested in individuals, nested in schools adolescents attended at the time of Key Stage 3. I created an interaction term composed of the cross-product of school poverty and agreeableness to test for moderation. The quadratic terms of school poverty and agreeableness were included to correct for the non-normality of the response variables. This corrects for spurious interaction effects (Lubinski & Humphreys, 1990). To control for changes in school environment, I ran a sensitivity analysis, restricting the sample to adolescents that did not change schools between Key Stage 3 and 4. This resulted in a reduced sample of 4020 (from 4236), indicating that most adolescents stayed in the same school. The results of the sensitivity analysis were the same, suggesting that changes in school environment did not play a role.","Adolescents attended schools with different levels of poverty: the first quartile of the sample went to schools that ranged from 0% to 4.2% of children eligible for free school meals; the second quartile ranged from 4.2% to 6.5%; the third from 6.5% to 13.3%; and the fourth from 13.3% to 100% school meal eligibility. When examining Model 1 (Table 2), school poverty is indeed negatively related to educational attainment. Higher levels of agreeableness were related to higher educational attainment. The control variables show that parental education was positively related with adolescents' attainment. Model 2 (Table 3) shows a positive interaction between school poverty and agreeableness. The interaction significantly improved the model fit. As predicted, higher levels of agreeableness were related to a weaker relation between school poverty and educational attainment. The interaction plot shows the same result (Fig 1): adolescent with low levels of agreeableness had a steeper slope for the relation between school poverty and educational attainment (b = −0.069; p = 0.000) than adolescents with high levels of agreeableness (b = −0.017; p = 0.019). Calculating the region of significance (with alpha = 0.05; Preacher, Curran, & Bauer, 2006) showed that agreeableness did not relate to educational attainment in schools with a standardised proportion of children eligible for school meals lower than 2.1452 (where proportion of children eligible for school meals ranged from −1.13 to 9.31).","This study showed that higher levels of school poverty were related to lower educational attainment. It is possible that the lack of positive role models and presence of peers from low-educated families results in a bad learning environment, where educational attainment is not valued and stimulated. Next, the relation between school poverty and educational attainment was buffered by personality trait agreeableness, which is in support of hypothesis 1 (H1). This finding is in line with other studies examining the moderating role of personality for other contextual effects on educational attainment (Nieuwenhuis, Hooimeijer, & Meeus, 2015; Nieuwenhuis, Hooimeijer, van Ham, & Meeus, 2017). Next, the diathesis-stress model predicted that adolescents in low poverty schools would not have different educational attainment based on their level of agreeableness (H2a), while the differential susceptibility model predicted adolescents with low agreeableness to be more malleable, and therefore do better in low poverty schools than adolescents with high agreeableness (H2b). The results showed that adolescents with different levels of agreeableness did not have different educational attainment in low poverty schools, which is in line with the diathesis-stress model (H2a). The findings suggest that higher levels of agreeableness contribute to resilience when faced with environmental stressors, such as school poverty. Lower levels of agreeableness contribute to more vulnerability under the same circumstances. It is possible that children from lower educated families end up in schools with lower SES, which could indicate that lower SES schools contain lower attaining adolescents because they sort into these schools based on their socio-economic background. By controlling for parental education, I tried to (partly) overcome this problem. Also, restricting the sample to adolescents who did not move schools between the two measurement points yielded the same results. If the relation between school and attainment would have disappeared, this could suggest that scoring low on Key Stage 3 would have resulted in a move to a worse (lower SES) school, which would then have driven the effect of school poverty for the full sample. This was not the case, which indicates that sorting did not play a major role. The results emphasise the importance of considering individual characteristics such as personality traits when assessing contextual predictors. School effects are not uniform, and have to be studied in tandem with the diversity of children who inhabit the school. This study could be used to identify adolescents who are particularly vulnerable to detrimental effects of school poverty."],["Sensory information is inherently ambiguous. The brain disambiguates this information by anticipating or predicting the sensory environment based on prior knowledge. Pellicano and Burr (2012) proposed that this process may be atypical in autism and that internal assumptions, or “priors,” may be underweighted or less used than in typical individuals. A robust internal assumption used by adults is the “light-from-above” prior, a bias to interpret ambiguous shading patterns as if formed by a light source located above (and slightly to the left) of the scene. We investigated whether autistic children (n = 18) use this prior to the same degree as typical children of similar age and intellectual ability (n = 18). Children were asked to judge the shape (concave or convex) of a shaded hexagon stimulus presented in 24 rotations. We estimated the relation between the proportion of convex judgments and stimulus orientation for each child and calculated the light source location most consistent with those judgments. Children behaved similarly to adults in this task, preferring to assume that the light source was from above left, when other interpretations were compatible with the shading evidence. Autistic and typical children used prior assumptions to the same extent to make sense of shading patterns. Future research should examine whether this prior is as adaptable (i.e., modifiable with training) in autistic children as it is in typical adults. --------------------------------------------------------------------------------","In human vision, the complex dimensions of a visual scene—the shapes of objects, their spatial arrangement, and their material properties—are reduced to flat patterns of excitation of the cones and rods of the retina. Information entering the brain is inherently ambiguous, compatible with a range of interpretations. Visual input, therefore, is “underspecified” for the task of providing the reliable and stable awareness of the environment that we experience. Consequently, perception has long been considered as a process of “unconscious inference” (Helmholtz, 1866/1911), in which existing knowledge is spontaneously and automatically deployed to interpret the meaning of sensory signals. Expectations based on experience of how the material environment works—for example, of faces being convex or of light sources being overhead—feed into the construction of a percept by the brain. The framework provided by Helmholz (1866/1911) was extended to characterize perceptual inferences as “hypotheses” or informed speculations using noisy and limited data (Gregory, 1980). Perceptual decisions are made possible by comparing the probability of the sensory evidence and prior experience. The Bayesian framework has since supplied an established mathematical model for perceptual decision making under conditions of uncertainty. In Bayesian terms, if both sensory signals and knowledge-based hypotheses are represented as probability distributions, techniques of statistical inference can be used to locate the combined point of maximal probability, the “best guess” interpretation (Knill, Kersten, & Yuille, 1996). In Bayesian perceptual inference, the prior probability distribution, or “prior,” represents a “baseline” understanding of the likelihood of particular environmental conditions on the basis of past experience (Gregory, 1980). A number of visual priors have been established and are thought to improve overall perceptual efficiency by weighting perceptual hypotheses in a broadly reliable way. For example, a prior for convexity reflects the predominance of convex, over concave, objects in the world (Langer & Bülthoff, 2001; Sun & Perona, 1996). The statistical inference calculation implies a trade-off between the image data and the prior probability, such that perception will be more prior driven when ambiguity in the sensory input is high. In some circumstances, prior-driven expectations will be misguided, resulting in visual illusions. For example, perceiving a hollow mask as a convex mask implies the operation of overriding expectations of convexity in faces. Hence, when problems of object perception are resolved on a probabilistic basis, the optimal solution may still be “inaccurate” (Kersten & Yuille, 2003). Importantly, in the context of this research, Bayesian priors envisage dynamic connections among perceptual inference, experience of the environment, and behavior (Lee, Yang, Romero, & Mumford, 2002). Atypicalities in sensation and perception are highly characteristic of autism (Baranek, David, Poe, Stone, & Watson, 2006; Leekam, Nieto, Libby, Wing, & Gould, 2007; Mottron, Dawson, Soulieres, Hubert, & Burack, 2006; Simmons et al., 2009). Although difficulties in social communication are considered hallmarks of autism, sensory reactivity, including hypersensitivity (e.g., to light or touch) and hyposensitivity (e.g., to pain), was noted during the first description of autism (Kanner, 1943). Sensory reactivity has since been shown to be present in the majority of autistic children and adults (Ben-Sasson et al., 2009; Leekam et al., 2007; Simmons et al., 2009), to be pervasive and persistent across development (Crane, Goddard, & Pring, 2009; McCormick, Hepburn, Young, & Rogers, 2016), and to have a substantial impact on the lives of autistic people (e.g., Dickie, Baranek, Schultz, Watson, & McComish, 2009; Grandin, 2009; Williams, 1994). In the perceptual domain, the majority of scientific studies have reported atypical processing in aspects of visual perception and visual attention ranging from characteristically nonsocial stimuli and tasks, such as discrimination of chromatic stimuli (e.g., Franklin et al., 2010), cast shadow (Becchio, Mari, & Castiello, 2010), static gratings (e.g., Bertone, Mottron, Jelenic, & Faubert, 2005), moving dots (e.g., Milne et al., 2002; Pellicano, Gibson, Maybery, Durkin, & Badcock, 2005), and complex objects (e.g., “Greebles”) (e.g., Behrmann, Thomas, & Humphreys, 2006; Davies, Bishop, Manstead, & Tantam, 1994), to social stimuli, including faces (e.g., Dalton et al., 2005; Kemner & van Engeland, 2006), eye gaze (e.g., Elsabbagh et al., 2009), and biological motion (e.g., Blake, Turner, Smoski, Pozdol, & Stone, 2003; Klin, Lin, Gorrindo, Ramsay, & Jones, 2009; for a review, see Simmons et al., 2009). Theories of autistic perception have explained these findings in terms of a “detail-focused” perceptual style (Frith & Happé, 1994), generally enhanced perceptual functioning (Mottron et al., 2006), or reduced generalization (Plaisted, 2001). Building on these accounts, Pellicano and Burr (2012) suggested that it is not sensory processing itself that is atypical in autism but rather the interpretation of the sensory input. Specifically, drawing on the tools of Bayesian theory, they proposed that the internal priors of autistic people are underweighted or less used than in typical individuals. Attenuated priors might result in enhanced perception in some contexts and reduced performance in others, depending on whether the task draws more heavily on sensory input itself or on successful perceptual prediction using prior knowledge. This idea was expanded by several related Bayesian and neurobiological accounts that proposed possible atypicalities in predictive processing in autism (Brock, 2012; Friston, Lawson, & Frith, 2013; Lawson, Rees, & Friston, 2014; Sinha et al., 2014; van Boxtel & Lu, 2013; Van de Cruys et al., 2014). Evidence supporting the hypothesis of attenuated priors in autism comes from studies showing that autistic children and adults show reduced adaptation, a form of experience-dependent plasticity in which neural systems fine-tune to the current visual environment according to the previous context. There is now considerable evidence for reduced adaptation in autism for high-level visual attributes, both social stimuli (e.g., faces: Pellicano, Jeffery, Burr, & Rhodes, 2007; biological motion: van Boxtel, Dapretto, & Lu, 2016) and nonsocial stimuli (e.g., numerosity: Turi et al., 2015), as well as for other sensory modalities (e.g., touch: Tommerdahl, Tannan, Holden, & Baranek, 2008; audition: Lawson, Aylward, White, & Rees, 2015); audiovisual calibration: Turi, Karaminis, Pellicano, & Burr, 2016). More recently, Karaminis et al. (2016) demonstrated atypicalities in prior knowledge in autism more formally, in the context of temporal reproduction, using a Bayesian computational model for central tendency (Cicchini, Arrighi, Cecchetti, & Burr, 2012). The computational model proposed that central tendency reflects the integration of noisy temporal estimates with prior knowledge representations of a mean stimulus. This integration serves to reduce overall error and, crucially, is flexible; the noisier the sensory estimates, the greater the reliance on prior knowledge. Karaminis and colleagues contrasted the performance of autistic and typical children completing a time interval reproduction task (measuring central tendency) and a temporal discrimination task (assessing temporal resolution) to the predictions of the Bayesian model. Computational simulations suggested that central tendency in autistic children was much less than that predicted by computational modeling given the poor temporal resolution of these children. Autistic children presented with a much less flexible use of priors across development compared with typically developing children. In the current study, we provided another test of the account of Pellicano and Burr (2012) using a well-established perceptual expectation, namely the “light-from-above” prior. The patterns of light and shade on and around an object provide uncertain cues for the brain to determine its shape and position in space. To resolve this ambiguity, the visual system must assume a light source location (Adams, 2007; Gerardin, de Montalembert, & Mamassian, 2007; Mamassian & Landy, 2001). The light-from-above prior weights inference toward assuming that a light source is located above the scene observed, as is most commonly experienced. The presence of a light-from-above prior is demonstrated when an image lit from above is rotated by 180°, so that shading previously perceived as indicating convexity will indicate concavity (and vice versa). Simple rotation changes interpretation (e.g., Mamassian & Goutcher, 2001). A perceptual bias toward light from above is also demonstrated in visual search; observers can identify shapes compatible with light from above significantly faster than the same shapes reversed (i.e., compatible with light from below), whereas shapes suggesting a light source from the side are recognized both more slowly and less accurately (Kleffner & Ramachandran, 1992). A further characteristic of the light-from- above prior is that it is biased slightly to the left of vertical (Gerardin et al., 2007; Mamassian & Goutcher, 2001; Sun & Perona, 1996; Thomas, Nardini, & Mareschal, 2010), although there is considerable individual variation in the location of the prior (e.g., Adams, 2007; Champion & Adams, 2007; Mazzilli & Schofield, 2013; Morgenstern, Murray, & Harris, 2011). While the preference for an overhead light source is assumed to have an environmental origin, the source of the leftward bias is not well understood (see Mamassian & Goutcher, 2001, for a discussion). The light-from-above prior strongly influences depth perception and assists in reconstructing three dimensions (objects/scenes) from two-dimensional retinal images. Here, we examined the use of this robust and well-characterized light-from-above prior in autistic children. We also investigated age-related differences in the use of the prior in autistic and typical children. To address these aims, we assessed autistic and typical children of similar age and ability on a shape judgment task designed to reveal individual differences in implicit judgments of light source location. Specifically, we adapted the seven-hexagon stimulus developed by Andrews, Aisenberg, d’Avossa, and Sapir (2012), which was used to demonstrate an effect of cultural differences (reading direction) on location of the prior and, therefore, was thought to be sufficiently sensitive to potential group-level variations. Limited shading information in the stimulus delivered a high level of ambiguity about depth, ensuring that perceptual inference was required to resolve shape and a light source location would need to be assumed. We sought to investigate whether autistic children would resolve the ambiguous shape-from-shading information in our stimulus using an assumption of light from above. Specifically, we tested whether autistic children would show the bias in stable conditions designed to encourage reliance on prior experience of lighting conditions. An attenuated light-from-above prior, in line with the hypothesis of Pellicano and Burr (2012), might result in more mixed shape judgments according to rotation of the stimulus by autistic children than by typical children, with fewer compatible with light from above. Because priors are thought to smooth over neural noise, attenuated priors in the autistic group might also lead to less stable perception of convexity or concavity, leading to more varied (or less confident) interpretations. Our test purposefully excluded signals that might compete with the light-from-above prior, such as a visible light source (Morgenstern et al., 2011), to determine whether the prior is fundamentally intact and develops similarly in autistic children compared with typical children.","In total, 18 autistic children (16 boys) and 18 typical children (12 boys), all between 7 and 14 years of age, took part in this study. Children were recruited via community contacts. All autistic children had been previously diagnosed with an autism spectrum condition by independent clinicians and scored above the threshold for an autism spectrum disorder on either the Autism Diagnostic Observation Schedule–Second Edition (ADOS-2) or the Lifetime version of the Social Communication Questionnaire (SCQ) (see Table 1 for scores). All typically developing children scored below the cutoff for autism on the SCQ (score of 15; Rutter, Bailey, & Lord, 2003), reflecting the absence of clinically significant autistic features. All children had normal or corrected-to-normal visual acuity, as reported by their parents. The groups were matched in terms of age, t(34) = 0.06, p = .95, verbal IQ, t(34) = 0.94, p = .35, performance IQ, t(34) = 0.88, p = .39, and full-scale IQ, t(34) = 0.38, p = .70, as measured by the Wechsler Abbreviated Scales of Intelligence–Second Edition (WASI-2; Wechsler, 2011) (see Table 1 for scores). All children obtained full-scale IQ scores of 70 or above and, thus, were considered to be cognitively able. An additional 15 typical adults (14 female; 20–35 years of age), recruited from the university, were tested to establish parameters for adult performance in our task. One additional adult was tested but excluded from the analysis because the estimated light source direction (see “Measurements” section below) lay more than 2 standard deviations to the left of the group average. Removing this outlier did not change the results reported here. Ethics statement ~~~~~~~~~~~~~~~~ The study was conducted in accordance with the principles laid out in the Declaration of Helsinki. Ethical approval was granted by the faculty research ethics committee of the university (FPS456). Parents of all children gave their informed written consent prior to the participation of their children in the project, and children gave their verbal assent. Stimuli Following Andrews et al. (2012), the stimulus comprised seven tessellated gray-scale hexagons on a gray-scale background (RGB = 127, 127, 127), with shading on the inner and outer edges of the shapes (see Fig. 1) providing ambiguous cues to depth. Each shape was 2.5° of visual angle, with an intermediate level of blur. The stimulus was rotated by 360° in 15° increments, compatible with different light source locations. We presented 24 rotations in a randomized order in order to estimate the light source direction most consistent with the judgments of convexity and concavity by children. Each rotation was presented five times, yielding a total of 120 trials. Stimuli were presented on a Dell Precision laptop screen with 1366 × 768 pixel resolution at a refresh rate of 60 Hz and mean luminance of 60 cd/m2. All children viewed the stimuli binocularly at a distance of approximately 57 cm from the screen.","The experiment was written in Matlab using the Psychophysics Toolbox extensions (Brainard, 1997; Kleiner, Brainard, & Pelli, 2007). We measured the implicit judgments of light source location made by children in the context of a child-friendly computer game. During the introduction phase, the cover story was introduced showing colored cartoon honeycomb cells and an animated bee. Children were asked to help the bee decide which “cells” needed to be filled with honey (Fig. 2A and B). Children were told that they would collect points for their decisions and win a small gift with more than 500 points. Dummy points were also provided to maintain children’s interest. To encourage children to distinguish between cells they perceived as empty and those they saw as full, they were also asked not to “waste” honey when cells were “full.” We also showed the bee with a telescope looking at the central hexagon (Fig. 2C) to ensure that answers were given for this cell only. After the introduction to the story, children were taken through the demonstration phase. The cartoon honeycomb was replaced with the hexagon stimulus, which appeared at a rotation of 150° so that the central cell would be most likely to be perceived as concave, according to the known mean prior for adults (Fig. 2D). Children were told that this needed filling with honey, which was demonstrated by the bee in an animation. Two demonstration trials followed, using the exact procedure followed in later test trials. The stimulus appeared with the prompt, “Does this one need filling?” Color-coded “yes” and “no” prompts appeared below the question to the left and right of the screen, respectively, in order to match the location of the response keys (letters “a” and “l” on the keyboard). These keys were also labeled “Y” and “N” using sticky labels matched in color to the font of the on-screen response options (yellow for yes and blue for no). Children were asked to press a key to indicate their response, with left (yes) indicating perceived concavity and right (no) indicating perceived convexity. The first demonstration trial again showed the rotation most likely to be perceived as concave, and children were invited to respond. A “yes” response was required to proceed. The second demonstration trial showed the stimulus rotated 330°, the position most likely to be perceived as convex. This required a “no” response and was followed by the caption “No! That one’s full!” Children could repeat the demonstration trials if required, but there were no such further trials in order to avoid influencing subsequent responses (Stone, Kerrigan, & Porrill, 2009). All children produced correct responses for both trials. The test phase commenced immediately following the demonstration phase. An animated bee hovered on the screen (2 s) as a fixation point. Children saw the honeycomb stimulus in one of the 24 rotations, the order of which was randomized. They answered the question “Does this one need filling?” by pressing the “yes” or “no” key according to their interpretation of the stimulus shape. Trials were self- paced, so that each response triggered the next fixation bee. Children completed three blocks of 40 test trials, yielding a total of 120 trials. Dummy scores and general encouragement were provided at the end of each block. General procedure ~~~~~~~~~~~~~~~~~ Children were tested individually in a quiet room at the university. The room was lit only by the computer screen, so that no environmental cues were available to influence children’s perceptions of lighting direction (Morgenstern et al., 2011). Testing on the experimental task lasted 10–15 min. The WASI-2 (and the ADOS-2 for autistic children) was administered to children in later sessions.","Fig. 3 shows example results from one autistic and one typical child participant. The plots show two-dimensional psychometric functions on polar axes, plotting percentage convex as a function of rendered orientation. In both cases, the judgment tends to be convex for orientations left of vertical and concave for orientations to the right; this is consistent with a light-from-above interpretation, specifically with a light-from-above left interpretation, in line with previous research (Gerardin et al., 2007; Sun & Perona, 1996; Thomas et al., 2010). The average orientation [estimated from Eq. (3)] is indicated by the orange circle: −32° for the autistic child and −17° for the typical child. Fig. 4 shows group averages for the estimated light source direction biases for the three participant groups. As expected (cf. Gerardin et al., 2007; Sun & Perona, 1996; Thomas et al., 2010), average light source biases were to the left of vertical (autistic children: M = −12.67°, SD = 12.51; typical children: M = −12.50°, SD = 7.87; adults: M = −12.06°, SD = 16.72) and were significantly lower than zero for all groups [autistic children: t(17) = −4.44, p < .001; typical children: t(17) = −6.72, p < .001; adults: t(14) = −2.79, p = .01]. A one-way analysis of variance (ANOVA) with group (autistic children, typical children, or adults) as a between-participants factor revealed no significant effect of group on the magnitude of the light bias, F(2, 48) = 0.10, p = .99. We also examined the data by performing a Bayesian one-way ANOVA using JASP software (Version 0.8.0.0; JASP Team, 2016) and estimating a Bayes factor using Bayesian information criteria (Wagenmakers, 2007). The Bayes factor allowed for a comparison of the fit of our data under the null hypothesis and the alternative hypothesis. The estimated Bayes factor (null/alternative) suggested that our results were 6.56:1 in favor of the null hypothesis, that is, 6.56 times more likely to occur under a model without an effect of group on the magnitude of the light-from-above bias rather than a model with an effect of group. Our data, therefore, provided substantial evidence (Wetzels et al., 2011) that adults, typical children, and autistic children interpreted the shape of the stimulus in a similar way in that they used a light-from-above prior to a similar degree. To investigate potential age- related differences in the formation of the prior, we examined the relationship between chronological age and the magnitude of the light bias. There was no significant relationship between age and bias for either group (autistic: r = −.24, p = .33; typical group: r = .03, p = .90). There was also no significant relationship with ability, as measured by full-scale IQ scores on the WASI-2, for either group of children (autistic: r = −.10, p = .95; typical group: r = .23, p = .36), or with autistic symptomatology, as measured by the ADOS-2 severity scores of children (r = −.13, p = .61) (Gotham, Pickles, & Lord, 2009; Hus & Lord, 2014). We also examined whether the judgments of younger children were influenced by a preference for convexity, as suggested by Thomas et al. (2010), by conducting a Pearson correlation analysis between participants’ percentage of convex judgments (across all responses) and age. However, we found no significant relationship between age and preference to interpret the shape as convex for either group of children (typical: r = .05, p = .84; autistic: r = −.02, p = .93).","This study investigated whether autistic and typical children apply a light-from-above prior similarly to interpret shape from shading in conditions where shading cues are ambiguous. The mean light-from-above prior for the children seen in this study was approximately −13° (above and slightly to the left). We found no significant difference in assumed light source location between autistic and typical children. Judgments of depth by autistic children were influenced by a light-from-above prior—that is, they preferentially assumed a light source above and to the left of the stimuli—to a similar extent as those by typical children. Given that all children reported that it was easy to decide whether the cell should be filled, it appears that they used their priors to help resolve noise and ambiguity and achieve stable percepts. This finding is consistent with a recent report suggesting that priors for eye gaze direction—where gaze is more likely to be perceived as direct in conditions of uncertainty—are intact in autistic adults (Pell et al., 2016). In our task, environmental conditions other than rotation were held constant, so that inference relied on participants’ existing priors, summing their long-term experience of lighting conditions. We used a task that had previously shown subtle differences between groups whose long-term (cultural) experience of reading direction differed (Andrews et al., 2012), suggesting that it should have been possible to detect differences based on long-term experience between our groups if they were in fact present. Similar mean levels might nevertheless conceal different group levels of adaptation in the short term to prevailing environmental conditions. Experimental evidence has demonstrated that changes in prevailing environmental conditions can modify light priors in the short term. In the case of typical adults, the light-from-above prior adapts rapidly (Adams, Graf, & Ernst, 2004; Adams, Kerrigan, & Graf, 2010; Champion & Adams, 2007), so that training with haptic feedback can change an individual light prior location by approximately 10° with 1.5 h of training (Adams et al., 2004). Participants’ visual judgments on a separate task after training were recalibrated in line with the new prior (Adams et al., 2010; Champion & Adams, 2007), indicating that the prior had been temporarily updated. Similarly, learning an association between a context and its lighting conditions led to contextually appropriate calibration of the light prior (Kerrigan & Adams, 2013), in line with Bayesian updating of the prior probability distribution by context. Flexible adaptation to prevailing conditions was not tested in the current study but is a worthy avenue for future research and a more direct test of the account of Pellicano and Burr (2012). It is also important to consider precisely how prior knowledge acts on perceptual inference. Although evidence strongly suggests that an assumed light source position is represented early in the visual system and acts on early inference (Champion & Adams, 2007; Lee et al., 2002; Mamassian, Jentzsch, Bacon, & Schweinberger, 2003), perceptual computation also appears to be an interactive process, so that activity in the early visual cortex may nevertheless take into account behavioral experience and higher order perceptual saliency (Lee, 2003; Lee et al., 2002). Champion and Adams (2007) found evidence of both early and late influence of the light-from-above prior on perceptual inference. Although visual haptic training modified the light-from-above prior used in a judgment of shape task, the prior assessed in a visual search task was not affected by the same training (Champion & Adams, 2007). The discrepant effects of the training environment were interpreted as evidence that although training did not touch the “quick and dirty” process of early inference, it influenced later additional stages of visual processing, where recent experience with the world is taken into account. Because our task did not require accounting for recent experience, the mean priors we found may reflect the influence of the prior on the initial stage of visual processing alone, leaving open the possibility of reduced influence of priors during later stages of visual processing for autistic children. Future research should test how far the priors of children adapt to prevailing environmental conditions. Our findings also show that the two groups of children performed similarly to adults. Furthermore, we found no age-related changes in the use of the light- from-above prior, at least in children between 7 and 14 years of age, suggesting that the prior develops early—before 7 years. This would be consistent with recent findings showing that typical children use prior knowledge in magnitude estimations from an early age, in particular for temporal (Karaminis et al., 2016) and spatial (Sciutti, Burr, Saracco, Sandini, & Gori, 2014) interval reproduction. If learning an internal model of the sensory environment is key to the statistical inference process involved in Bayesian perception, and also is critical to learning beyond perception (Fiser, Berkes, Orbán, & Lengyel, 2010), early acquisition of priors, especially robust ones like these, would make sense developmentally. However, Karaminis et al. (2016) found differences between autistic and typical development in time interval reproduction. Autistic children showed significantly less precision in their estimations than matched typical children, with computational simulations suggesting less use of prior knowledge than expected to compensate for this level of imprecision. This result implies differences in the flexible deployment of priors to improve the precision/reliability of estimates. Our task showed no differences between typical and autistic children in the use or development of the light-from-above prior, at least in controlled conditions where competing information is limited (cf. Morgenstern et al., 2011), but the findings of Karaminis and colleagues are another reason to study the flexibility of the light-from-above prior (cf. Adams et al., 2010) in autistic children. Our findings are in contrast to other research reporting developmental increases in responses consistent with light-from-above prior (Stone, 2011; Stone & Pascalis, 2010; Thomas et al., 2010), although the evidence is not clear-cut. Thomas et al. (2010) found an overall effect of age on children’s interpretation of ambiguous stimuli, so that trials answered assuming light from above increased across childhood. Yet this result was complicated by the measurement of two distinct priors—one for convexity and one for light from above. When the two priors biased interpretation in the same direction, there was little change in use of the light-from-above prior over development; when the priors conflicted, younger children (4- and 5-year-olds) preferred a convex interpretation, whereas older children and adults preferred to assume light from above if light could be interpreted as from above and to the left rather than as above right. However, the number of younger children was small (n = 7), with some performing at around chance levels. Stone and Pascalis (2010) showed children aged 4–10 years geometric shapes and photographic stimuli, either upright or rotated 180°, reporting increasing levels of interpretation consistent with light from above by age group. Yet the results for symbolic images—where ambiguity is higher and, therefore, the potential influence of the prior is stronger—did not clearly support this interpretation. In particular, the regression analysis for concave stimuli (if lit from above) against age was nonsignificant, whereas that for convex symbols (if lit from above) indicated greater use of the prior for children aged around 7 or 8 years. Overall, these results point to the difficulty of disentangling the influence of light from above and convexity priors and the difficulty of interpreting which factors influenced the performance of younger children. In sum, we have established that the light-from-above prior is similar in strength in a shape judgment task in school- age autistic and typical children of similar age and ability. Furthermore, contrary to previous reports (Stone & Pascalis, 2010; Thomas et al., 2010), our results suggest a stable level of use of the prior across this age range. Although our methods were developmentally sensitive, with performance not being confounded with linguistic demands (concave and convex), our sample sizes were relatively small and we may have had insufficient power to detect age-related changes in the use or strength of the prior; this is especially important given that priors vary widely between individuals (e.g., Adams, 2007; Champion & Adams, 2007; Mazzilli & Schofield, 2013; Morgenstern et al., 2011). One outstanding question is whether the light-from-above prior is just as adaptable as it is in adults for typically developing children—but especially for children on the autism spectrum."],["Background This study examined the effects of cultivated (i.e. developed through training) and dispositional (trait) mindfulness on smooth pursuit (SPEM) and antisaccade (AS) tasks known to engage the fronto-parietal network implicated in attentional and motion detection processes, and the fronto-striatal network implicated in cognitive control, respectively. Methods Sixty healthy men (19–59 years), of whom 30 were experienced mindfulness practitioners and 30 meditation-naïve, underwent infrared oculographic assessment of SPEM and AS performance. Trait mindfulness was assessed using the self-report Five Facet Mindfulness Questionnaire (FFMQ). Results Meditators, relative to meditation-naïve individuals, made significantly fewer catch-up and anticipatory saccades during the SPEM task, and had significantly lower intra-individual variability in gain and spatial error during the AS task. No SPEM or AS measure correlated significantly with FFMQ scores in meditation-naïve individuals. Conclusions Cultivated, but not dispositional, mindfulness is associated with improved attention and sensorimotor control as indexed by SPEM and AS tasks. --------------------------------------------------------------------------------","Smooth pursuit and saccadic eye movements are two types of eye movements that both human and non-human primates voluntarily employ to allow the image of an object fall and maintain near to or on the fovea. The function of smooth pursuit eye movements (SPEM) is to keep a retinal image within the area of the fovea during the movement of an object. The initiation as well as the maintenance of accurate SPEM requires attentional control (Hutton & Tegally, 2005). The primary measure of pursuit accuracy is the velocity gain which corresponds to the ratio of smooth pursuit velocity over target or object velocity (100% if SPEM velocity matches the target velocity) (Lencer & Trillenberg, 2008). Other indicators of SPEM efficiency are the frequency of compensatory catch-up and intrusive anticipatory saccades made during the smooth pursuit task. Saccades refer to the fast eye movements made to the sudden appearance of a visual target. Prosaccades require the participant to make a saccade to a single-target stimulus as soon as it appears. The antisaccade (AS) paradigm, on the other hand, requires the participant to inhibit a reflex-like saccade towards the target, and instead initiate a saccade in the direction opposite to the target (Hutton & Ettinger, 2006). It examines the conflict between a pre- potent stimulus that produces a strong urge to make a saccade to the target, and the overriding goal to look in the opposite direction. Correct AS performance requires accurate perception, ability to transform location information to a mirror image representation, and suppression of a saccade towards the AS stimulus. Performance is assessed as the error rate, spatial accuracy and latency of (anti)saccades (Hutton & Ettinger, 2006). Given their high test-retest reliability (Ettinger et al., 2003) and the ease of administration, SPEM and AS paradigms have been used to assess cognitive functions in a wide variety of contexts (Hutton & Ettinger, 2006; Lencer & Trillenberg, 2008). However, no published study, to our knowledge, has yet utilised these paradigms to understand the neural and cognitive influence of mindfulness as a trait (‘dispositional’ mindfulness) (Brown & Ryan, 2003), or mindfulness developed through training (‘cultivated’ mindfulness) (Ivanovski & Malhi, 2007). Mindfulness, a translation of the Pali term sati, is operationalized in modern psychology as a quality of awareness that arises from paying attention to the experience on a moment-by-moment basis without judging, elaborating upon, or fixating on this experience in any way (Kabat-Zinn, 1990). Its practice typically begins with the mindfulness of bodily sensations to the awareness of feelings and thoughts, progressing to a present-centred awareness without an explicit focus, in most Buddhist traditions as well as formal intervention-style practices such as Mindfulness- Based Stress Reduction (MBSR; Kabat-Zinn, 1990) and Mindfulness-Based Cognitive Therapy (MBCT; Segal, Williams, & Teasdale, 2002). Dispositional mindfulness refers to the naturally occurring tendency to display this non-judgmental present awareness in everyday life and varies amongst individuals (Brown & Ryan, 2003). There is growing evidence for a positive effect of cultivated mindfulness on a range of cognitive functions (reviews, Chiesa, Calati, & Serretti, 2011; Gallant, 2016). Mindfulness practice enhances stability of attention by reducing cortical noise (Lutz et al., 2009) and appears to increase information processing capacity (Slagter et al., 2007) with the cognitive system more rapidly available to process new targets (Slagter, Lutz, Greischar, Nieuwenhuis, & Davidson, 2008). It is reported to positively influence orienting attention (van den Hurk, Giommi, Gielen, Speckens, & Barendregt, 2010), dual attention (Jensen, Vangkilde, Frokjaer, & Hasselbalch, 2012) and performance on a range of tasks requiring attention and/or cognitive flexibility (Hodgins & Adair, 2010; Jha, Krompinger, & Baime, 2007; Jha, Stanley, Kiyonaga, Wong, & Gelfand, 2010; Kumari, Hamid, Brand, & Antonova, 2015; Moore, Gruber, Derose, & Malinowski, 2012; Semple, 2010; Tang et al., 2007; van den Hurk, Janssen, Giommi, Barendregt, & Gielen, 2010). According to a recent review (Gallant, 2016), mindfulness-led improvements in executive functioning, within Miyake et al.’s model (2000), are more consistently found on specific measures of inhibition (Allen et al., 2012; Heeren, Van Broeck, & Philippot, 2009; Moore & Malinowski, 2009; Sahdra et al., 2011; Teper & Inzlicht, 2013), relative to updating (Jha et al., 2010; Mrazek, Franklin, Tarchin, Baird, & Schooler, 2013) and shifting components (Anderson, Lau, Segal, & Bishop, 2007; Chambers, Lo, & Allen, 2008; Heeren et al., 2009; Moynihan et al., 2013). There are, however, only few data at present examining the association between trait mindfulness, as measured in the general non-meditating population using self-report questionnaires (Brown & Ryan, 2003), and cognitive function, with some studies (Stillman, Feldman, Wambachm, Howard, & Howard, 2014; Stillman et al., 2016; Whitmarsh, Uddén, Barendregt, & Petersson, 2013) indicating a negative relationship between trait mindfulness and implicit learning, i.e. learning without conscious awareness. The aim of the present study was to examine the influence of mindfulness both as a trait and developed through training on visuo-spatial attentional and voluntary inhibition processes indexed by the SPEM and AS paradigms, respectively. Based on previous report of positive effects of mindfulness practice on a range of attention, working memory and inhibition tasks (reviews, Chiesa et al., 2011; Gallant, 2016), we hypothesised that meditators, compared with meditation-naïve individuals, will show more accurate gain and lower frequency of catch-up and anticipatory saccades on the SPEM task, and a lower error rate, higher spatial accuracy and reduced within-subject variability in spatial accuracy (i.e. more consistent performance) on the AS task. Furthermore, we hypothesised tentatively (in absence of direct previous data) that in meditation-naïve individuals, trait mindfulness, as assessed by the Five Factor Mindfulness Questionnaire (FFMQ; Baer, Smith, Hopkins, Krietemeyer, & Toney, 2006), will correlate with better SPEM and AS performance, particularly on measures that significantly differentiate meditators and non-meditators. Participants and design ~~~~~~~~~~~~~~~~~~~~~~~ The study involved two groups. Group 1 consisted of 30 experienced mindfulness practitioners (meditators) and Group 2 of 30 meditation-naïve individuals (all males; age range 19–59 years). All participants were assessed on a single occasion. Meditators were recruited from local and national Buddhist centres via poster advertisement and presentations of the study and its aims at meetings of centre members. They were required to have at least 2 years of consistent meditation practice (minimum 45 min per day at least 6 days a week). Meditation-naive men were recruited from a healthy volunteer database or by circular emails sent to staff and students of King’s College London, UK. They were required to have no experience of mindfulness-related practices, including meditation, yoga, tai-chi, qigong or martial arts. Additional inclusion criteria for all participants were: (i) right-handedness (Oldfield, 1971), (ii) current IQ > 80 as assessed with the two-test version of Wechsler Abbreviated Scale of Intelligence (Wechsler, 1999), (iii) aged 18–60 years, (iv) normal or corrected-to-normal vision, and (v) not drinking more than 28 units of alcohol per week [1 unit = 1/2 pint of beer (285 mls) or 25 ml of spirits or 1 glass of wine], or more than 6 units of caffeinated beverage a day. Those with a positive screen for any neuropsychiatric disorder, a current or past primary diagnosis of substance misuse or on regular medical prescription were not included. The final sample with usable SPEM data consisted of 29 meditators (drawn mainly from Zen, Theravada, Vajrayana and Triratna traditions of Buddhism) and 30 meditation-naïve individuals, and with usable AS data of 27 meditators and 29 meditation-naïve individuals (see Table 1). Study procedures were approved by the King’s College London Research Ethics Committee (PNM/14/15-90). Participants provided written informed consent to their participation and were compensated for their time and travel. Sample characterisation ~~~~~~~~~~~~~~~~~~~~~~~ Meditation history (meditation tradition/style, meditation routine) was obtained from the meditators prior to study participation. In addition, all participants completed the FFMQ (Baer et al., 2006). The FFMQ has been constructed from factor analysis done on five of the most previously popular measures of trait mindfulness, and is currently the most frequently used measure of trait mindfulness. Its five facets are observing (Observe, e.g. “When I’m walking, I deliberately notice the sensations of my body moving.”), describing (Describe, e.g. “I’m good at finding words to describe my feelings.”), acting with awareness (Awareness, e.g. “I find myself doing things without paying attention.” with reverse scoring), non-judging of inner experience (Non-judgment, e.g. “I tell myself I shouldn’t be feeling the way that I am feeling.”) and non-reactivity to inner experience (Non-reactivity, e.g. “I watch my feelings without getting lost in them.”), assessed on a 5-point Likert scale (never or rarely, rarely, sometimes, often, very often or always true) with 39 items (8 items each for Observe, Describe, Awareness and Non-judgement facets and 7 items for Non-reactivity facet). Higher scores indicate higher mindfulness. Eye movements: Paradigms and procedure ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Infrared oculography (IRIS 6500; Skalar Medical BV, Delft, the Netherlands) was used to record eye movements. Participants were seated in a height-adjustable chair with their head resting on a chinrest at a distance of 57 cm from the computer monitor. They were requested to remain as still as possible throughout the experiment. The visual stimulus consisted of a white circular target, with 0.3° diameter, presented on a black background on a 17-inch monitor. A 3-point (+12°, 0°, −12°; stimulus duration = 1000 ms) calibration task was carried out prior to running the SPEM, AS and prosaccade (PS) tasks. During the SPEM task, a target dot moved horizontally across the screen, in a triangular waveform, at three different velocities (12°/s, 24°/s, 36°/s). Participants were requested to keep their gaze on this horizontally moving target as closely as possible. The target dot was initially positioned at the centre of the screen (0°). It then moved either to the left or the right side of the screen (±12°), and then to the opposite side. Each target movement from one side (e.g. −12°) to the other (e.g. +12°) is referred to as half-cycle or ramp. The target dot completed a total of 16.5 half-cycles. The first half-ramp (from 0 to ±12°) was not included in the analysis. The PS and AS tasks used ±6° and ±12° targets, each presented 15 times in a random order (60 PS trials and 60 AS trials in separate blocks). Each PS and AS trial began with the target in the centre of the participant’s visual field (0°) for a random duration of 1000–2000 ms. The target then abruptly stepped to one of the four possible peripheral locations (±6° and ±12°), along the horizontal plane, and remained there for 1000 ms before it stepped back to the central position for the next trial. During the PS task, participants were requested to look at the target when it was in the central position, and then follow it with their eyes (i.e. generate prosaccades) when it stepped to the peripheral positions. During the AS task, participants were requested to keep their gaze at the target when it was in the central position and to generate a saccadic eye movement to the mirror-image projection of the target, in the opposite hemifield, when it moved to any one of the four peripheral positions. Four practice trials, one with each target location, were carried out before the experimental trials, and repeated if necessary. All participants were assessed first on the PS task, followed by the AS and SPEM tasks. The testing took place in a quiet and darkened room. The participants were allowed to have coffee on the day of testing but were provided only with decaffeinated drinks for at least 1 h prior to being assessed on eye movement tasks. The experimental setup, oculomotor tasks and oculographic data recording and scoring procedures were the same as used in a previous study (Schmechtig et al., 2010). Eye movement analysis ~~~~~~~~~~~~~~~~~~~~~ All eye movement recordings were scored blind to group membership. Smooth pursuit The time-weighted average pursuit velocity gain, frequency of catch-up saccades and frequency of anticipatory saccades were calculated for each participant, using LABVIEW 6.0 student version (2000). The time-weighted average pursuit velocity gain was calculated by dividing mean eye velocity by target velocity. This analysis included sections of pursuit which lay in the central half of each ramp (the first and last quarters excluded to avoid effects of pursuit initiation and slowing at target turnarounds). Velocity gain scores for each section of the pursuit, that were free of saccades or blinks, were time-weighted and subsequently averaged across half cycles for each target velocity (score <100% means that the eye is slower, and score >100% that it is moving faster, than the target). Saccadic frequency per second was calculated dividing the total number of anticipatory or catch-up saccades by the duration in seconds of pursuit at each target velocity (N/s). Anticipatory saccades were defined as saccades that began with the eye on or behind the target and ended ahead of it, while catch-up saccades were defined as saccades that began with the eye behind the target and served to bring the eye closer to the target. Saccades that began behind the target and ended ahead of it were defined as anticipatory saccades if more than half of amplitude was spent ahead of the target, and as catch-up saccades if more than half of amplitude was spent behind the target. Back-up saccades and square- wave jerks were not counted as their frequency was too low for a meaningful differentiation of the meditator and meditation-naïve groups. Saccades were automatically detected in EYEMAP 2.1 (AMTech, GmbH, Weinheim, Germany) using minimum amplitude (1°) and velocity (30°/s) criteria. Prosaccade (PS) Spatial accuracy (gain, spatial error) and latency, along with associated SDs, were scored for each participant using EYEMAP 2.1. PS gain was calculated as the percentage of saccade amplitude divided by target amplitude multiplied by 100 (a score of 100% represents a perfectly accurate saccade, <100% a hypometric saccade and >100 a hypermetric saccade). Spatial error (percentage) represented the residual error. It was calculated for each trial by subtracting the target amplitude from saccade amplitude and dividing it by the target amplitude, and then averaged across all trials and multiplied by 100 (higher scores indicate greater spatial error regardless of saccadic over or under-shoot). Saccadic latency represented the time (in ms) from the appearance of the target to saccade initiation. Antisaccade (AS) The AS error rate (% total) was calculated as the percentage of error trials (i.e. trials where the participant’s first saccade is towards the target) over the total number of valid trials (i.e. error trials plus correct trials, excluding eye blink trials). AS gain, spatial error, latency, error rate (% total), as well as the SDs of gain, spatial error and latency were calculated for each participant. In addition, the correction rate (%) was scored to ensure that all included participants knew task requirements (∼100% correction rate). AS gain, spatial error and latency were calculated following the criteria described above (for PS).","Group differences in age, IQ, and FFMQ scores were examined using independent sample t-tests. Each SPEM measure (gain, frequency of catch-up saccades, frequency of anticipatory saccades) was analysed using a 2 (Group: meditators, meditation-naïve) × 3 (Velocity: 12°/s, 24°/s and 36°/s target velocities) analysis of variance (ANOVA) with Group as a between-subjects factor and Velocity as a within-subjects factor, followed by the analysis of simple main effects and lower order ANOVAs as appropriate to test the hypothesised differences between the meditator and meditation-naïve groups. Each AS (error rate, gain, spatial error and latency, as well as the SDs of gain, spatial error, latency) and PS measure (gain, spatial error and latency as well as SDs of these variables) was analysed using a one-way ANOVA. Effect sizes, where reported, are partial eta squared (ηp2; the proportion of variance associated with a factor). Correlational analyses (Pearson’s r) were run to examine the hypothesised association between FFMQ scores and SPEM and AS measures in meditation-naïve individuals; for completeness, similar correlation analyses were conducted in meditators. Possible correlations between age and SPEM and AS variables were also examined. All analyses were performed using the Statistical Package for Social Sciences (for Windows, version 22; IBM, New York, US). Alpha level for testing significance of effects was maintained at p < 0.05 unless stated otherwise. Sample characteristics: Meditators versus Non-meditators ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Meditators scored significantly higher on Observe [t (49) = 3.21, p = 0.002], Non-judgment [t(49) = 3.37, p = 0.001] and Non-reactivity facets [t(49) = 4.05, p < 0.001] of mindfulness as assessed by the FFMQ (Baer et al., 2006) compared to meditation-naïve individuals (Table 1). The two groups did not differ in age or IQ (p > 0.25) (Table 1). Eye movements: Meditators versus non-meditators ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Means (SDs) for eye movement measures, separately for the meditators and meditation-naive groups, are presented in Table 2. SPEM Meditators did not differ from meditation-naïve individuals in gain at any of the three target velocities, as there was neither a main effect of Group [F(1, 57) = 1.96, p = 0.17, ηp2 = 0.03] nor a Group × Velocity interaction [F(2, 114) = 0.06, p = 0.94, ηp2 = 0.001]. There was only a significant main effect of Velocity showing less accurate performance with increasing velocities in both groups [F(2, 114) = 8.74, p < 0.001, ηp2 = 0.13; linear F(1, 57) = 12.66, p = 0.001, ηp2 = 0.18] (Table 2). Meditators made fewer catch-up saccades than meditation-naïve individuals at 12° target velocity (p = 0.01) but did not differ significantly at the other two target velocities (p > 0.13), as revealed by the follow-up analysis of a significant Group × Velocity interaction [F(2, 114) = 3.73, p = 0.03, ηp2 = 0.06]. There was a main effect of Velocity showing a higher frequency of catch-up saccades with increasing velocities in both groups [F(2, 114) = 142.06, p < 0.001, ηp2 = 0.71; linear F(1, 57) = 225.79, p < 0.001, ηp2 = 0.78] (Table 2). Meditators made fewer anticipatory saccades than meditation-naïve individuals at all three velocities as demonstrated by a significant main effect of Group [F(1, 57) = 6.70, p = 0.01, ηp2 = 0.11] and a non-significant Group × Velocity interaction [F(2, 114) = 0.88, p = 0.42, ηp2 = 0.01]. In addition, there was a main effect of Velocity showing a higher frequency of anticipatory saccades with increasing velocities in both groups [F(2, 114) = 48.70, p < 0.001, ηp2 = 0.46; linear F(1, 57) = 69.72, p < 0.001, ηp2 = 0.55] (Table 2). AS Meditator and meditation-naïve groups did not differ in error rate [F(1, 54) = 1.00, p = 0.33, ηp2 = 0.02], latency [F(1, 54) = 0.06, p = 0.80, ηp2 = 0.001], gain [F(1, 54) = 2.52, p = 0.12, ηp2 = 0.04] or spatial error [F(1, 54) = 2.53, p = 0.12, ηp2 = 0.04]. Meditators, however, had significantly lower SDs of gain [F(1, 54) = 5.20, p = 0.026, ηp2 = 0.09] and spatial error [F(1, 54) = 4.24, p = 0.04, ηp2 = 0.07] relative to meditation-naïve individuals. PS There was no difference between the two groups for any PS variables (all p values >0.18). Correlational analyses: Trait mindfulness (FFMQ), age and eye movement measures ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In meditation-naïve individuals, only two, out of 78 in total, correlations reached p < 0.05 (not corrected for multiple correlations), and these two were inconsistent with one showing better (fewer catch-up saccades) and the other showing worse performance (more anticipatory saccades) in association with higher scores trait mindfulness (Describe and Non-judgment facets) (Table 3). Age did not correlate with any eye movement measures. In meditators, higher scores on the FFMQ were associated with better AS performance. Specifically, higher scores on both Observe and Non-reactivity facets were associated with lower spatial error and lower SDs of spatial error and gain, with Non-reactivity facet associating further with a larger gain; the correlations with other mindfulness facets were non-significant but in the same direction. In addition, older age was consistently associated with poorer AS performance (higher error rate, increased spatial error and longer latency).","Supporting our hypothesis in relation to the influence of cultivated mindfulness, the present study revealed superior SPEM and AS performance in meditators relative to meditation-naïve individuals. Specifically, meditators, relative to meditation-naïve individuals, had fewer catch-up (at 12° target velocity) and anticipatory saccades (at all three target velocities) during the SPEM task, and significantly lower SDs of gain and spatial error during the AS task. The two groups did not differ significantly in AS gain and spatial error, though meditators, in line with our a priori hypothesis, were more stable across trials in these measures of spatial accuracy. The findings, however, offered no support for our hypothesis in relation to dispositional (trait) mindfulness. None of the five facets of the FFMQ correlated consistently positively or negatively with either SPEM or AS indices in non-meditators, though significant relationships between higher FFMQ scores (Observe and Non-reactivity facets) and more accurate and more consistent AS performance were present in meditators. In addition, older age was associated with poorer AS performance (higher error rate, increased spatial error and longer latency) in meditators. The finding of fewer catch-up and anticipatory saccades during the SPEM task in meditators, compared to non-meditators, indicating better attentional control in long- term meditators (Hutton & Tegally, 2005), is in line with Lutz et al.’s (2008) focussed attention meditation framework. This superiority most likely developed through regular practice of mindfulness since no significant relationship was found between the SPEM indices and FFMQ scores in non-meditators. SPEM paradigms are well known to elicit activity in frontal and posterior areas that are implicated in attentional and motion detection processes (Lencer & Trillenberg, 2008; Sharpe, 2008) as well as in the neurobiological effects of mindfulness practices (Barnby, Bailey, Chambers, & Fitzgerald, 2015; Hölzel et al., 2011; Marchand, 2014). It would be valuable to further examine the neural basis of the influence of cultivated mindfulness in SPEM performance. Our finding of significantly lower SDs of gain and spatial error during the AS task indicates lower intra-individual variability (i.e. more consistent within-session performance) in meditators, relative to non-meditators. Intra-individual variability, examined mostly in reaction time across a range of tasks, is considered to reflect lapses in attention or cognitive control (Fassbender, Scangos, Lesh, & Carter, 2014; Weissman, Roberts, Visscher, & Woldorff, 2006), sustained attention deficit (Leth-Steensen, Elbaz, & Douglas, 2000), a poor ability to successfully engage cognitive control in demanding situations (Bellgrove, Hester, & Garavan, 2004), and poor regulation of effort (Sergeant, Geurts, Huijbregts, Scheres, & Oosterlaan, 2003). Increased intra-individual variability occurs with aging (e.g. Anstey, 1999; Fozard, Vercruyssen, Reynolds, Hancock, & Quilter, 1994; Shammi, Bosman, & Stuss, 1998) and is associated with many clinical conditions, including frontal- temporal dementia (Murtha, Cismaru, Waechter, & Chertkow, 2002), attention hyperactivity deficit disorder (Leth-Steensen et al., 2000; Vaurio, Simmonds, & Mostofsky, 2009; Zahn, Kruesi, & Rapoport, 1991), schizophrenia (Schwartz et al., 1989) and traumatic brain injury (e.g. Bleiberg, Garmoe, Halpern, Reeves, Nadler, 1997; Segalowitz, Dywan, & Unsal, 1997; Stuss, Murphy, Binns, & Alexander, 2003; Stuss et al., 1989; Zahn & Mirsky, 1999). Intra-individual variability is typically most strongly present on tasks that require executive control (West, Murphy, Armillio, Craik, & Stuss, 2002) and appears particularly sensitive to frontal lobe function, with frontally-deficient clinical groups showing the greatest intra-individual variability (Murtha et al., 2002). Since all our participants were required to be free of a neuropsychiatric condition and known brain injury, the findings of lower SDs of gain and spatial error during the AS task can be taken to indicate superior executive control and frontal lobe functioning developed through training in the meditator group. Interestingly, neuroticism has been associated with greater intra-individual variability, supposedly due to distracting worries about task performance or difficulty in those with high neuroticism (Robinson & Tamir, 2005). If mindfulness improves emotion regulation by exerting a positive influence on executive control processes (Teper & Inzlicht, 2013), this may, at least partly, explain both performance superiority of the meditators observed in our study and the recently demonstrated reduction in neuroticism following MBCT (Armstrong & Rimes, 2016). Furthermore, mindfulness training or practice is known to exert attenuating effects on the Default Mode Network (e.g. Brewer et al., 2011; Farb et al., 2007), associated with mind- wandering or stimulus-independent thought (Mason et al., 2007), which would further enhance the performance on the tasks employed in the current study. This study did not reveal a meaningful pattern of correlations between dispositional mindfulness and eye movement performance indices. It may be that dispositional mindfulness, as measured by self-report questionnaires, indeed exists fairly independently of cultivated mindfulness and is conceptually unique (Rau & Williams, 2016; Wheeler, Arnkoff, & Glass, 2016). It may share some (e.g. negative association with neuroticism) but not all behavioural or neural correlates of cultivated mindfulness (Rau & Williams, 2016). Furthermore, absence of opposite traits, such as neuroticism or mindlessness, might not indicate a presence of mindfulness by necessity (Grossman & Van Dam, 2011). There were, however, meaningful and significant relationships between FFMQ scores and AS measures in meditators. Specifically, both Observe and Non-reactivity facets were correlated with lower spatial error and lower SDs of spatial error and gain, with Non-reactivity facet correlating further with a larger gain. Correlations of the three remaining mindfulness facets with these AS parameters, although in the same direction, were non-significant. Taken together, these observations may suggest that earlier-noted consistent mindfulness training-led improvements found across studies in the inhibition component of executive control (review, Gallant, 2016), may be mediated most strongly by Observe and Non-reactivity aspects of mindfulness training. Our study has some limitations. First, it examined the effects of cultivated mindfulness on eye movement control in a cross-sectional design, without any knowledge of the meditators’ eye movement performance prior to them starting mindfulness practice. Future research could examine the effects of shorter duration mindfulness-based interventions (MBIs) on SPEM and AS performance. If the results show improved SPEM and AS performance following MBIs, they would not only add to our understanding of the neural and cognitive effects of mindfulness but would also provide easily quantifiable and objective markers to index mindfulness training effects and its neurobiological underpinnings. Second, the findings of this study, which involved only men, cannot be generalized to women. Further research is needed to examine the influence of mindfulness in eye movement control in women, preferably controlling for menstrual phases, given menstrual phase- related variability in other psychophysiological measures of attention and inhibitory function (Kumari, 2011). In conclusion, this is the first study, to our knowledge, to have examined and shown superior SPEM and AS performance in established meditators, relative to meditation-naïve individuals. The findings suggest that mindfulness meditation improves attention and the stability of responding on visuo-motor tasks. Future studies are needed to confirm these effects using within-subjects designs (pre- and post-mindfulness training) and firmly establish whether eye movement tasks hold promise as objective measures of mindfulness training.","The authors declare no conflict of interest.","The sponsors had no role in study design; in the collection, analysis and interpretation of data; in the writing of the report; or in the decision to submit the paper for publication."],["Whilst there is evidence for the impact of driving anxiety on behaviour, less exists for the impact of trait anxiety and what does exist is inconclusive. The current study explored the possibility that trait anxiety interacts with driving anxiety to impact the frequency of negative on-road thoughts and behaviours. An online survey was administered to drivers, and the State-Trait Inventory for Cognitive and Somatic Anxiety, the Driving Cognitions Questionnaire, and the Driving Behaviour Survey, were completed. Moderation analyses suggested that in addition to an increase in social concerns and aggressive responses, high trait anxiety reduced positive associations between driving anxiety and exaggerated safety-cautious behaviours, as well as the general use of maladaptive reactions to stressful situations. As scores on these subscales were still higher regardless of the reduced associations, it is argued that both drivers with a generally anxious personality and those with high levels of driving-specific anxiety should be made aware of their potential to violate traffic norms in stressful situations. --------------------------------------------------------------------------------","According to research and national statistics (Department for Transport, 2014; Dula & Geller, 2003), negative emotions increase the likelihood of dangerous driving behaviours and crash involvement. The UK's Department for Transport revealed that over 5000 crashes in 2013 were preceded by negative emotional experiences behind the wheel, with 1900 of these accounted for by nervousness, uncertainly or panic. This suggests that emotions associated with anxiety may be a significant risk factor for traffic crash involvement. Those with an anxious driving style tend to feel distress and anxiety when driving, and express a lack of confidence in their skills (Taubman-Ben-Ari, Mikulincer, & Gillath, 2004). Whilst they often report lower levels of sensation-seeking, suggesting a reduced risk of crash involvement due to avoidance of high-risk situations (Ulleberg, 2001), empirical evidence suggests that this subgroup may be more dangerous on the road. Those associated with this driving style have shown lapses in attention and memory (Lucidi et al., 2010), a greater number of on-road errors (Taylor, Deane, & Podd, 2007), and a greater likelihood of crash involvement (Marengo, Settanni, & Vidotto, 2012). Those with driving anxiety may perceive their abilities as insufficient to deal with the environment they have encountered, resulting in increased stress and use of emotion-focused coping (Lazarus & Folkman, 1984). Transportation research confirms this suggestion by demonstrating that increased levels of stress contribute towards increased errors, lapses, and dangerous driving (Ge et al., 2014; Rowden, Matthews, Watson, & Biggs, 2011). However, recent evidence has also acknowledged the relationship between personality and the appraisal-emotion relationship; neuroticism, a trait associated with anxiety, has demonstrated a moderating and exacerbating relationship between appraisals and negatively valenced emotions (Tong, 2010). Thus when evaluating the ways in which drivers cope with stressful situations, the role of personality should not be ignored. Yet there is a lack of focus in, or consensus on, the relationship between trait anxiety and driving. Whilst self-report evidence associates higher trait anxiety with more violations, errors, and lapses (Pourabdian & Azmoon, 2013; Shahar, 2009), behavioural research looking at areas such as hazard perception (Barnard & Chapman, 2016) and speeding compliance (Stephens & Groeger, 2009) have found it has either no detrimental effects, or actually make drivers safer. It is possible that trait anxiety, rather than consistently affecting driver behaviour, has more of an impact on the negative thoughts associated with driving. Processing theories such as Attentional Control Theory (Eysenck, Derakshan, Santos, & Calvo, 2007), propose that an increased occupation with worrisome thoughts reduces processing efficiency without necessarily impacting behaviours. Research into how this is associated with driving is limited, although it does accord with the observation that trait anxiety is associated with increased reaction times on n-back tasks as well as increased errors and lapses (Wong, Mahar, & Titchener, 2015). Recent questionnaires assessing the effects of anxiety on driver behaviour have referred to some of the principles discussed within Attentional Control Theory. For example, recent research looking at the Driving Cognitions Questionnaire (DCQ- Ehlers et al., 2007) acknowledged the high probability of trait anxiety producing a higher frequency of dysfunctional thoughts in phobic participants (da Costa, de Carvalho, Cantini, da Rocha Freire, & Nardi, 2014). Furthermore, the Driving Behaviour Survey (DBS- Clapp et al., 2011) includes a subscale on anxiety-based performance deficits, defined as behaviours that occur due to an increase in worrisome or anxious thoughts that increase cognitive load. Notably, some of the previously discussed research emphasised an association between an anxious driving style and lapses in attention, suggesting that the principles of the theory could apply to anxious drivers, as well as those with dispositional anxiety. What has received less focus in the literature is the way in which trait and state anxiety may interact with each other. This issue was highlighted in a recent review on anxiety (Wilt, Oehlberg, & Revelle, 2011). It was suggested that the dichotomisation of anxiety into these two dimensions may result in the potential to reduce our understanding of how they exist in a concurrent fashion. The idea of interaction has been supported more recently in research into phonological and set-shifting processing efficiencies (Edwards, Edwards, & Lyvers, 2015, 2016). Research within the field of transportation would benefit from a similar integration of personality and state. Whilst previous research has investigated the relationship between trait and driving anxiety, this has merely confirmed the existence of higher trait anxiety in those with driving anxiety (Taylor et al., 2007), without further exploration of the relationship between the two. The present study explores the potential influences of both driving anxiety, as well as trait anxiety, on the frequency of negative thoughts and behaviours on the road. This was achieved using an online survey. We investigate whether the previous suggestions from Wilt et al. (2011) could be observed within an applied context. Additionally, we provide data to help practitioners understand whether the frequency of specific thoughts and actions is additionally affected by an anxious personality, rather than simply being anxious about driving. Based on the previous literature, it was hypothesised that trait anxiety could have a moderating effect on the frequency of negative thoughts associated with driving, as well as potentially on behaviours associated with worrisome or anxious thoughts.","Participants were approached using social media invitations, advertising on a local newspaper website, advertising on a local study recruitment website, and through a University volunteering database. The study was completed online, and a total of 320 participants with full driver's licences expressed interest in the survey by going to the web link associated with the survey; however, only 227 completed their responses, resulting in a 71% retention rate (149 females, 76 males). Their ages ranged from 17 to 81, with an average age of 35 (sd = 18.44). The majority were in employment or full-time study (80.2%), whilst the remainder of the sample was retired or unemployed (17.2%). Participants had held their full driving licence for 15.19 years (sd = 16.79), and drove 6229.22 miles per year (sd = 5904.19). Most reported driving at least a few times a week (69.46%). Ninety-five reported previous involvement in a crash in which they were the driver (41.85%). Software ~~~~~~~~ The survey was compiled and distributed to participants using LimeSurvey version 2.05, an online open source software tool which can be used to create and publish surveys, as well as compile respondent statistics and collate responses for analysis. Demographic variables The survey initially consisted of a standard information sheet and consent form, after which questions were asked regarding demographics and general driving behaviour. These included questions on licence duration, annual mileage and crash history. Once these had been completed, participants completed the State-Trait Inventory for Cognitive and Somatic Anxiety (STICSA-Grös, Antony, Simms, & McCabe, 2007), the DCQ, and the DBS. State-Trait Anxiety for Cognitive and Somatic Anxiety The STICSA is a 42-item questionnaire, 21 of each measuring state and trait anxiety. Items are distinguished according to whether they measure cognitive (10 items) or somatic symptoms of anxiety (11 items). Items are administered on a 1 to 4 Likert scale, with 1 meaning ‘not at all’ and 4 meaning ‘very much’. Total scores for state and trait range from 21 to 84. In all cases, a higher score indicates higher anxiety levels. However, to measure the effects of driving anxiety, the phrasing of the state STICSA was changed; instead of asking participants how they felt “at this moment”, they were asked how they felt whilst driving. Whilst previous research has used the State-Trait Anxiety Inventory (Spielberger, Gorsuch, Lushene, Vagg, & Jacobs, 1983) to obtain anxiety measures, recent research has suggested that the STICSA is more strongly correlated with anxiety than the STAI (Grös et al., 2007) Average driving anxiety scores were 29.43 (sd = 10.93) and average trait anxiety scores were 31.8 (sd = 11.3) Cronbach's α was 0.946 for driving anxiety, and 0.941 for trait anxiety, indicating good internal consistency. Driving cognitions questionnaire The DCQ is a 20-item questionnaire designed to assess the frequency of concerning thoughts whilst driving. Six items address social concerns, seven address accident concerns, and seven address panic concerns. Items are administered on a 5-point Likert scale, with “0” meaning “Never”, and “4” meaning “Always”. A minimum score of 0 and maximum score of 28 can be obtained for accident and panic concerns, whilst a maximum score of 24 can be obtained for social concerns. An overall maximum score of 80 can be obtained. In the current study, average scores for social, panic and accident concerns were 5.17 (sd = 4.9), 1.98 (sd = 3.48) and 6.34 (sd = 5.26) respectively, whilst the overall average score was 13.49 (sd = 12.26). Overall Cronbach's α was 0.941, and for social, accident and panic concerns were 0.866, 0.897, and 0.888 respectively. Driving behaviour survey The DBS is a 21-item questionnaire measuring the frequency of driving behaviours associated with a hypothetically stressful driving situation on three subscales consisting of seven items each. Anxiety-based performance deficits (ABPD) are related to changes in driving performance due to an anxiety-induced increase in cognitive load and include behaviours such as lane drifting and inappropriate speed adjustments. Exaggerated safety-cautious behaviours (ESCB) increase the perceived safety of a situation by maintaining larger headway distances and unnecessarily slowing down at traffic lights. Finally, the aggression subscale evaluates the ‘fight’ aspect of the fight-or-flight response and includes behaviours such as swearing and pounding on the steering wheel. Items are administered on a 7-point Likert scale, with 1 meaning “Never” and 7 meaning “Always”. Item scores are averaged, meaning that for each subscale participants can obtain a minimum score of 1 and a maximum score of 7; this same principle is applied to the overall DBS score. Average scores for ABPD, ESCB behaviours, and aggression were 2.09 (sd = 0.9), 3.52 (sd = 1.25) and 2.2 (sd = 1.12) respectively, whilst the average overall score was 2.6 (sd = 0.84). Cronbach's α for the total score was 0.885, and for ABPD, ESCB, and aggressive reactions were 0.835, 0.846, and 0.862 respectively.","After expressing interest in the survey, participants took an average of 10–15 min to complete it, after which a chance to win £50 in shopping vouchers was offered as an incentive. The survey was online for approximately 12 months and ethical approval for its administration was obtained from a local Ethics Committee.","Subscale and overall scores for questionnaires were calculated manually; based on previous research (Taylor & Sullman, 2009), where there were no more than two or fewer items missing on a scale, mean item replacement accounted for this. Data from five participants were removed due to missing STICSA scores. Three further cases were also removed from DBS analysis. Scores were subjected to a moderation analysis using PROCESS version 2.16.3 (Hayes, 2013), running in SPSS version 22. Driving anxiety was treated as a predictor variable, whilst trait anxiety was treated as a moderator variable. Demographic variables previously associated with higher levels of anxiety in previous transportation research were entered into moderation models as covariates. These included age, experience (defined as the amount of years since passing a driving test), gender, and cras involvement. The latter two were dummy coded to assume that the participant was female or had been previously involved in an accident. Interactions between driving and trait anxiety were explored using simple slopes analysis. Slopes were generated at ±1 standard deviation of the mean of trait anxiety scores, and the Johnson-Neyman method was used to obtain a zone of significance. Bivariate correlations between variables, as well as output from post-hoc power analyses conducted in G*Power v3.1 (Faul, Erdfelder, Lang, & Buchner, 2007), are available as supplementary material. DCQ ~~~ For all scores, levels of driving anxiety were a significant positive predictor (all ps < 0.001). Trait anxiety also acted as a significant positive predictor of social concerns (p < .01) and total DCQ scores (p < .05); however, it did not moderate the relationship between driving anxiety and DCQ scales. Additionally, those involved in an accident had lower DCQ social scores (see Table 1). DBS ~~~ Levels of driving anxiety positively predicted ABPD scores, ESCB scores, and total DBS scores (all ps < .001, see Table 2). For the aggression subscale, age and gender were significant negative predictors. This suggested that female drivers, as well as older drivers, had lower aggression scores (ps < .05). Trait anxiety was a positive predictor of ABPD scores, aggression scores, and total DBS scores (ps < .001). Trait anxiety acted as a positive predictor of aggression in the absence of a similar effect of driving anxiety (p = .36). Interaction effects were found for ESCB and total DBS scores (ps < .05, see Figs. 1 and 2). Details on unstandardized slopes are available in Table 3. The Johnson-Neyman technique revealed that trait anxiety moderated the relationship between driving anxiety and scores, until trait anxiety scores were 20.91 points above the mean for ESCB, and 19.54 points above the mean for total DBS scores.","The current study aimed to establish whether trait anxiety moderates the relationship between driving anxiety and the frequency of negative thoughts and behaviours on the road, in accordance with previous suggestions and findings (Wilt et al., 2011). We found that whilst trait anxiety did have a moderating effect, this was only for the frequency of reactive behaviours, and not the frequency of negative thoughts, as previously hypothesised. Interactions included an increase in the frequency of ESCB behaviours, as well as an increase in general negative behaviours. However, rather than finding that personality exacerbated this relationship, as previously suggested (Tong, 2010), higher trait anxiety in fact reduced the associative strength between driving anxiety and these behaviours. Furthermore, the frequency of ESCB and total DBS behaviours were consistently higher for those with high trait anxiety than those with high driving anxiety but only low or average levels of trait anxiety (see Figs. 1 and 2). For ESCB behaviours, whilst those with high driving anxiety may be engaging in such behaviours as a means to increase perceptions of control (Baker, Litwack, Clapp, Beck, & Sloan, 2014), those with high trait anxiety may be doing so, and acting dangerously, without necessarily being anxious about driving. This would also apply to the use of maladaptive behaviours, as indicated by total DBS scores. Practically, those with high trait anxiety may need to be made aware that their reactions to stressful situations may be causing danger and violating traffic norms. Whilst it cannot be determined from this study whether those with high trait anxiety are consciously adopting these behaviours, it has been suggested that unrecognised persistence of ESCB behaviours would negatively impact interventions such as exposure-based therapy (Clapp, Baker, Litwack, Sloan, & Beck, 2014). Several interventions could be put in place to increase such awareness, such as continuous feedback from vehicles. Trait anxiety did not moderate concerns whilst driving, but still acted as a predictor of thoughts and behaviours. Some findings are in line with previous research (Deffenbacher, Huff, Lynch, Oetting, & Salvatore, 2000) and indicate that future interventions should seek to reduce feelings of reactive anger. Other findings, such as the relationship between trait anxiety and ABPD, could suggest maladaptive changes in behaviour due to an increase in cognitive load. This would support previous research (Clapp et al., 2011) and suggests that some of the specific behaviours that may need to be targeted by interventions for those with high trait anxiety include attention-based behaviours. To accept this suggestion would imply that these behaviours are due to worrisome thoughts. The data from this study suggest that there are effects of trait anxiety on social and general concerns whilst driving, supporting previous research (Taylor & Deane, 2000), but it is also worth considering how these thoughts may impact behaviour. The DCQ asks participants to rate how frequently they experience certain thoughts whilst they are driving. As these data report thoughts during driving, this could indicate a form of mind wandering. Mind wandering results in dangerous driving behaviours (Yanko & Spalek, 2014); it is possible that those with high trait anxiety may be preoccupied with thoughts of unlikely events, such as others thinking they are a bad driver, resulting in unintentionally dangerous behaviours. Whilst research suggests those reporting anxiety whilst driving are no more likely to show mind wandering (Burdett, Charlton, & Starkey, 2016), this is from a state perspective; those high in neuroticism also report higher levels of mind wandering and poorer attentional control (Robison, Gath, & Unsworth, 2017), thus it is still possible that those high in trait anxiety may show similar behaviours behind the wheel. Whilst these findings could have important implications, there are several limitations to consider. Firstly, the use of self-report data means that these results may not fully reflect on-road behaviour. Whilst there are some positive relationships between self-report and behavioural responses (Taubman-Ben-Ari, Eherenfreund-Hager, & Prato, 2016), to our knowledge such data does not currently exist for the DCQ and DBS. Secondly, the use of retrospective data suggests that the data may need to be interpreted with caution. For example, accident data is often prone to distortions in memory due to forgetting (Maycock, Lockwood, & Lester, 1991) and inflations in intensity, depending on time since recall (Chapman & Underwood, 2000). Additionally, those higher in anxiety are more likely to recall threatening information (Mitte, 2008). For the current study, anxious participants could have used more negative experiences to influence recall, resulting in potential data bias. Finally, there may be issues with temporal stability. The data were collected over 12 months, in which time driver cognitions may have changed. To highlight this, initial DBS construction suggested issues in temporal stability for the aggression subscale (Clapp et al., 2011). Therefore, it is acknowledged that due to the chosen scales, there may be issues with reliability. The current study suggested that whilst trait anxiety independently predicts certain concerns and maladaptive responses to stressful driving situations, it also has the potential to reduce positive associations between driving anxiety and ESCB behaviours. Due to consistently higher ESCB scores in high trait anxiety, this indicates that those high in anxiety could show driving behaviours violating traffic norms and increasing accident likelihood. Practically, this indicates that those with such a personality should not be ignored within the field of transportation, and measures seeking to help those with high driving anxiety may also benefit those with high trait anxiety."],["Background Even though Down syndrome is the most common chromosomal cause of intellectual disability, studies on early development are scarce. Aim To describe movements and postures in 3- to 5-month-old infants with Down syndrome and assess the relation between pre- and perinatal risk factors and the eventual motor performance. Methods and procedures Exploratory study; 47 infants with Down syndrome (26 males, 27 infants born preterm, 22 infants with congenital heart disease) were videoed at 10–19 weeks post-term (median = 14 weeks). We assessed their Motor Optimality Score (MOS) based on postures and movements (including fidgety movements) and compared it to that of 47 infants later diagnosed with cerebral palsy and 47 infants with a normal neurological outcome, matched for gestational and recording ages. Outcomes and results The MOS (median = 13, range 10–28) was significantly lower than in infants with a normal neurological outcome (median = 26), but higher than in infants later diagnosed with cerebral palsy (median = 6). Fourteen infants with Down syndrome showed normal fidgety movements, 13 no fidgety movements, and 20 exaggerated, too fast or too slow fidgety movements. A lack of movements to the midline and several atypical postures were observed. Neither preterm birth nor congenital heart disease was related to aberrant fidgety movements or reduced MOS. Conclusions and implications The heterogeneity in fidgety movements and MOS add to an understanding of the large variability of the early phenotype of Down syndrome. Studies on the predictive values of the early spontaneous motor repertoire, especially for the cognitive outcome, are warranted. What this paper adds The significance of this exploratory study lies in its minute description of the motor repertoire of infants with Down syndrome aged 3–5 months. Thirty percent of infants with Down syndrome showed age-specific normal fidgety movements. The rate of abnormal fidgety movements (large amplitude, high/slow speed) or a lack of fidgety movements was exceedingly high. The motor optimality score of infants with Down syndrome was lower than in infants with normal neurological outcome but higher than in infants who were later diagnosed with cerebral palsy. Neither preterm birth nor congenital heart disease were related to the motor performance at 3–5 months. --------------------------------------------------------------------------------","Even though Down syndrome is the most common chromosomal cause of intellectual disability, with 20–22 individuals per 10,000 births affected (e.g., Kurtovic-Kozaric et al., 2016; Loane et al., 2013), studies on early development are scarce. Infants with Down syndrome are known to be socially competent but show a delay in the acquisition of motor milestones and deficits in early gesture production (Grieco, Pulsifer, Seligsohn, Skotko, & Schwartz, 2015; Özcaliskan, Adamson, Dimitrova, Bailey, & Schmuck, 2016; Saito & Watanabe, 2016). As early as the first months of life they scored lower than typically developing infants on both the Test of Infant Motor Performance (Cardoso, Campos, Santos, Santos, & Rocha, 2015) and the Alberta Infant Motor Scale (Tudella, Pereira, Pedrolongo Basso, & Savelsbergh, 2011). They kicked less often (Ulrich & Ulrich, 1995) and their arm movements were less accurate when reaching for objects of different sizes (de Campos, Cerra, Silva, & Rocha, 2014). Repeated assessments of their spontaneous general movements revealed a heterogeneous movement quality, although the fluency and complexity tended to improve between 1 and 6 months of age (Mazzone, Mugno, & Mazzone, 2004). Initially designed for infants with acquired brain injuries, the Prechtl assessment of general movements (Einspieler & Prechtl, 2005; Prechtl et al., 1997) has recently also been applied to infants with genetic syndromes (Einspieler, Hirota, Yuge, Deijima, & Marschik, 2012; Einspieler, Kerr, & Prechtl, 2005; Einspieler et al., 2014; Marschik, Soloveichick, Windpassinger, & Einspieler, 2015; Mazzone et al., 2004) and infants later diagnosed with autism spectrum disorders (Einspieler et al., 2014; Zappella et al., 2015). The assessment is based on visual Gestalt perception of normal vs. abnormal movements in the entire body (i.e. general movements). It is applied in foetuses, preterm infants, and newborn infants from term to 5 months post-term (Einspieler, Prechtl, Bos, Ferrari, & Cioni, 2004; Prechtl & Einspieler, 1997). The excellent predictive power of general movement assessments (Bosanquet, Copeland, Ware, & Boyd, 2013; Einspieler et al., 2004) is mainly attributable to fidgety general movements, which occur from 3 to 5 months post-term age (Einspieler & Prechtl, 2005; Prechtl et al., 1997). Infants with normal fidgety movements are very likely to develop normally in neurological terms, whereas infants who never develop fidgety movements have a high risk for neurological impairment (Einspieler & Prechtl, 2005; Prechtl et al., 1997). Adding a detailed assessment of concurrent movements and postures to the assessment of fidgety movements, for example, showed a reduced motor optimality score (MOS) to be associated with a limited activity in children who were later diagnosed with cerebral palsy (Yang et al., 2012), or with lower intelligent quotients during school age (Butcher et al., 2009). We therefore assumed that determining the MOS (Einspieler et al., 2004, p. 26) by assessing fidgety movements as well as concurrent movement and postural patterns would enable us to systematically document the motor repertoire of infants with Down syndrome. The MOS makes it possible to quantitatively relate pre- and perinatal data to the motor repertoire of an infant and to data obtained from follow-up studies. The aims of our study were (1) to describe movements and postures in 3- to 5-month-old infants with Down syndrome; (2) to compare their MOS with the MOS of two matched samples, one of which was later diagnosed with cerebral palsy, while the other had a normal neurological outcome; and (3) to analyse to what extent clinical risk factors during pregnancy, at delivery, and during the neonatal period were related to the motor performance in 3- to 5-month-old infants with Down syndrome.","This exploratory study comprised a convenience sample of 47 infants with Down syndrome − 21 females (45%) and 26 males (55%) − who had been admitted to (a) the Darcy Vargas Public Hospital, São Paulo (17 individuals); (b) the Department of Physiotherapy and Rehabilitation at the Hacettepe University in Ankara (ten individuals); (c) the Associação de Pais e Amigos dos Excepcionais at São Paulo University (seven individuals); (d) the Rehabilitation Department of the Children’s Hospital of Fudan University in Shanghai (four individuals); (e) the Children’s Department at the City Hospital of Ostrava (three individuals); (f) the Clinic of Early Intervention at the University Hospital São Paolo (three individuals); and (g) the Medical University of Graz (three individuals) between June 2015 and May 2016. In order for the infants to be included in the study, the infants’ motor performance had to be recorded between 9 and 20 weeks post-term. The infants’ gestational ages at birth ranged from 29 to 41 weeks (median = 37 weeks), with a birth weight range of 1440 g to 3680 g (median = 2585 g). Twenty-seven infants were born preterm (57%), including two monozygotic twin pairs. Other clinical characteristics obtained from the medical histories are presented in Table 1. Three infants were diagnosed with mosaic Down syndrome, one with Robertsonian translocation (14;21). For comparison we used data of our international, MOS-based data bank (N = 365 as of January 15, 2017), picking (i) 47 individuals with a normal neurological outcome at 3–5 years of age, whose gestational age and age of video recording matched our study cases; and (ii) 47 individuals with a comparable gestational age at birth and post-term age at the time of the video recording who were diagnosed with cerebral palsy at 3–5 years of age (Table 2). Since spontaneous (i.e. endogenously generated) movements are not related to ethnicity (Luxwolda et al., 2014), we did not match the ethnic background. All parents gave their written informed consent. The ethical review boards of the various centres approved the study. Recording and evaluation of movements and postures at 3–5 months post-term age ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Within the GenGM network we recorded 5-min videos of the spontaneous motility of each infant at a median post-term age of 14 weeks (P25 = 12 weeks; P75 = 15 weeks; range: 10–19 weeks). The recordings were performed during periods of active wakefulness between feedings, with the infant partly dressed, lying in supine position (Einspieler et al., 2004). The videos were evaluated by at least two certified raters (D.H., C.E., A.M., A.N., J.P., P.B.M.) according to the Prechtl method of global and detailed general movement assessment (Einspieler et al., 2004 Einspieler & Prechtl, 2005). Scorers C.E. and P.B.M. were not familiar with the details of the participants’ clinical histories apart from the fact that they had Down syndrome. In case of disagreement (four recordings; 8.5%), the raters re-evaluated the recordings until consensus was reached on a final score. Fidgety movements and the concurrent repertoire of movements and postures were assessed independently in separate runs of the video recordings. Using the score sheet for the assessment of motor repertoire at 3–5 months (Einspieler et al., 2004, p. 26), we calculated the MOS, with a maximum value of 28 (for the best possible performance) and a minimum value of 5. The score sheet comprises the following five sub-categories: (i) fidgety movements, (ii) age-adequacy of motor repertoire, (iii) quality of movement patterns other than fidgety movements, (iv) posture, and (v) overall quality of the motor repertoire (Einspieler et al., 2004; Yuge et al., 2011). Fjørtoft and colleagues found a high inter-observer reliability for the MOS with intra-class correlation coefficients ranging from 0.80 to 0.94 (Fjørtoft, Einspieler, Adde, & Strand, 2009).","Statistical analysis was performed using the SPSS package for Windows, version 23.0 (SPSS Inc., Chicago, IL). The Pearson Chi-square test was used to evaluate associations between nominal data. To put the medians of non-normally distributed continuous data (e.g. motor optimality score) in relation to nominal data (e.g. preterm birth), we applied the Mann- Whitney-U test or, if there were more than two categories, the Kruskal-Wallis test (e.g. repertoire). To assess the relative strength of the association between variables, we computed the following correlation coefficients: Cramer’s V coefficient was applied when at least one of the two variables was nominal (e.g. preterm birth and age-adequacy of the repertoire). To assess the relation between two continuous variables (e.g. gestational age and motor optimality score), we applied the Pearson product-moment correlation coefficient. Throughout the analyses, p < 0.05 (two-tailed) was considered to be statistically significant. The motor performance of infants with Down syndrome at 3–5 months postterm age (Table 2, first column) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fourteen infants had normal fidgety movements (30%); six infants (12.5%) displayed abnormal fidgety movements (i.e. they look like normal ones though with a greater amplitude, speed and jerkiness); 13 infants (27.5%) displayed no fidgety movements, which were therefore classified as absent; 14 infants (30%) showed fidgety-like movements whose amplitude was too great and whose pace was too slow. Abnormal and absent fidgety movements as well as fidgety-like movements were grouped as aberrant fidgety movements (33 infants; 70%). Twelve infants (25.5%) displayed an age-adequate movement repertoire. The repertoire was found to be reduced in 20 infants (42.5%; mainly due to a lack of movements to the midline), and age-inadequate in 15 infants (32%). The quality of the various movement patterns (other than fidgety movements) was scored as predominantly normal in 39 infants (83%) and predominantly abnormal in three infants (6.5%); five infants (10.5%) showed an equal number of normal and abnormal movements. On average the infants demonstrated three normal movement patterns (range: 0–8) and one abnormal movement pattern (range: 0–3). The most frequent normal movement patterns included visual scanning (32/47; 68%), side-to-side movements of the head (22/47; 47%), foot-to-foot contact (14/47; 30%), hand-to-mouth contact (12/47; 25.5%), and kicking (10/47; 21%). Smiling (9/47; 19%), fiddling (9/47; 19%), hand regards (8/47; 17%), swipes (7/47; 15%), hand-to-hand contact (6/47; 13%), arching (6/47; 13%), and leg lifting (5/47; 11%) were observed in fewer than ten individuals. Other movement patterns such as wiggling-oscillating arm movements, hand-to- knee contact, or rolling to the side were observed in fewer than five individuals (<10%). The most frequent abnormal movement pattern was long-lasting and/or repetitive tongue protrusion (26/47; 55%). In a few infants we observed hand-to-hand contact with no mutual manipulation (4/47; 8.5%), long lasting wiggling-oscillating arm movements (3/47; 6%), repetitive kicking (3/47; 6%), and monotonous side-to-side movements of the head (1/47; 2%). Posture was rated as predominantly normal in 22 infants (47%) and predominantly abnormal in 15 infants (32%); ten infants (21%) showed an equal number of normal and abnormal postures (Table 2). Infants with predominantly normal postural patterns were able to hold their head in midline (31/47; 66%), showed a symmetrical body posture (25/47; 53%) and variable finger postures (23/47; 49%); a persistent asymmetric tonic neck response was absent in all individuals. On average the infants demonstrated three normal postural patterns (range: 1–4) and two abnormal postural patterns (range: 0–7). The most common abnormal pattern was a lack of variable finger postures (24/47; 51%) with just a few monotonous finger postures, finger spreading and/or predominant fisting. Twelve infants (25.5%) kept both arms predominantly extended, while seven infants (15%) kept their legs extended most of the time. Hyperextension of the neck and trunk was seen in three individuals (6.5%). The following two postural atypicalities were observed, but are not captured by the MOS sheet: nine individuals (19%) showed an internal rotation and pronation of one or both wrists, and 22 infants (47%) showed an external rotation and abduction of the hips (which was also the reason why most of them were unable to show foot-to-foot contact). Only three infants (6.5%) exhibited a normal, smooth and fluent overall movement character, while 44 infants (93.5%) displayed a monotonous, stiff, jerky and/or tremulous movement character. The median MOS was 13 (P25 = 12; P75 = 23; range: 10–28). The MOS of infants with Down syndrome compared to infants later diagnosed with cerebral palsy and infants with a normal neurological outcome (Table 2) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The MOS of infants with Down syndrome was significantly lower than that of infants with a normal neurological outcome (p < 0.01) but significantly higher than that of infants later diagnosed with cerebral palsy (p < 0.01). Similar results were obtained for fidgety movements, the age-adequacy of the motor repertoire, and the overall movement character (p-values < 0.01). The quality of movement and postural patterns of infants with Down syndrome were similar to those of infants with a normal neurological outcome (p-values > 0.10), while infants later diagnosed with cerebral palsy scored lower (p-values < 0.01; Table 2). The neonatal period and its relation to the motor performance at 3–5 months post-term age (Table 3 and 4) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ None of the clinical characteristics was related to the MOS or fidgety movements. Preterm birth was not associated with the MOS or its subcategories (Table 3), nor was any other clinical variable. We would particularly like to mention that congenital heart disease (CHD) was not related to the motor performance at 3–5 months. Cranial ultrasound data were only available for a small proportion of the sample (eight infants, 17%). Two infants with abnormal cranial ultrasound findings had abnormal fidgety-like movements. Among the six infants with normal cranial ultrasound findings three showed normal fidgety movements, one infant did not develop fidgety movements, and two infants had slow abnormal fidgety-like movements (). Table 4 only lists significant associations between clinical variables and items of the MOS. A jerky movement character (observed in 26/47 infants; 55%) was associated with caesarean section and hyperbilirubinaemia, although the two clinical variables were not related to each other (Pearson Chi-square test, p = 0.58). Delivery by caesarean section was also related to a higher occurrence of a particular atypical posture at 3–5 months: in many cases external rotation and abduction of the hips led to a lack of (especially lower limb) movements to the midline. No difference was observed between twins and singletons (Fisher sign test, p > 0.05). Neither the three infants with mosaic Down syndrome (two with absent fidgety movements, one with normal fidgety movements; MOS = 10, 12 and 24, respectively) nor the one with Robertsonian translocation (normal fidgety movements, MOS = 23) showed any sort of specific features in their motor performance at 3–5 months post term.","Apart from one individual reported in the context of a larger sample of high-risk infants in Japan (Yuge et al., 2011), this is the first study to use the MOS to describe fidgety movements and the concurrent motor repertoire in infants with Down syndrome. Mazzone et al. (2004) applied the detailed scoring for writhing general movements (observable from birth to 1–2 months post-term) up to the age of 6 months rather than for fidgety and concurrent motor patterns, which occur from 3 months onwards. The motor coordination and performance of children, adolescents and adults with Down syndrome has been found to be highly heterogeneous (Latash, 2007). Interestingly, we already observed this heterogeneity at a much younger age: 30% of our 3- to 5-month-old infants with Down syndrome showed normal fidgety movements, 27.5% had no fidgety movements, and 42.5% showed abnormal fidgety movements. These highly diverse findings are reflected in the MOS, which range from 10 to the highest possible score of 28. Fifty percent of the infants scored between 10 and 13, which is significantly lower than infants with a normal neurological outcome and higher than infants who were later diagnosed with cerebral palsy. So far, two studies carried out in children with cerebral palsy in China, Italy, and the Netherlands demonstrated the lower the MOS at 3–5 months, the more severely limited their gross motor function (Bruggink et al., 2009; Yang et al., 2012). As we intend to monitor our sample of children with Down syndrome for at least 2 more years, we shall also assess the relation between the MOS and the motor, cognitive and language outcomes. The one infant with Down syndrome described by Yuge et al. (2011) displayed abnormal fidgety movements and an MOS of 13. Abnormal fidgety movements, exaggerated in amplitude, speed and jerkiness, were also observed in six infants in the present study. This is an exceptionally high rate of occurrence, as abnormal fidgety movements are usually rare (see also Table 2). It has been a matter of debate whether or not infants with a low muscle tone are more likely to show abnormal fidgety movements (Einspieler, Peharz, & Marschik, 2016; Yuge et al., 2011). Although this may not be consentaneously defined, a low muscle tone is a general feature in infants with Down syndrome (Latash, 2007; Morris, Vaughan, & Vaccaro, 1982). A number of studies have documented an association between abnormal fidgety movements and coordination difficulties and/or disabilities in fine manipulative skills at school age (Einspieler et al., 2007; Einspieler et al., 2016); others describe an exceedingly high rate of abnormal fidgety movements in infants who were later diagnosed with autism spectrum disorders or Rett syndrome (Einspieler et al., 2005; Einspieler et al., 2014). However, most so-called abnormal fidgety movements in infants later diagnosed with autism spectrum disorders or Rett syndrome did not correspond with the category of abnormal fidgety movements described in infants with brain injuries, which were exaggerated in amplitude and speed (Einspieler & Prechtl, 2005; Einspieler et al., 2016; Prechtl et al., 1997). Several infants with a later diagnosis of autism spectrum disorders or Rett syndrome showed continual fidgety activity which was exaggerated in amplitude but too slow (Einspieler et al., 2005; Einspieler et al., 2014; Zappella et al., 2015). We also observed this pattern in 14/47 (30%) infants with Down syndrome (see the abnormal fidgety movements’ category in Table 2). It remains unclear whether this pattern is related to low muscle tone, another early atypicality in autism spectrum disorders (Flanagan, Landa, Bhat, & Bauman, 2012) and Rett syndrome (Nomura & Segawa, 1990), as video analysis does not allow for an assessment of active muscle strength or resistance to passive movements. Data on neurological examinations were not available in 44 of 47 individuals. In any case, the low rate of movements to the midline (i.e., foot-to-foot contact in 30% and hand-to- hand contact in only 13% of our sample) and the lack of kicking (21%) are in line with previous studies, where they have also been discussed as possible consequences of low muscle tone (Tudella et al., 2011). The same is true of the frequent external rotation and abduction of the hip (observed in 47% of infants with Down syndrome) as well as the uni- or bilateral internal rotation and pronation of the wrist(s) (19%). Unfortunately we are unable to confirm this association for lack of information about our infants’ muscle tone (see section 4.1.). Several research groups have reported a significant impact of pre- and perinatal variables on the MOS at 3–5 months. For example, prenatal exposure to environmental pollutants or selective serotonin reuptake inhibitors, and perinatal hypoxic events resulted in a reduced MOS (Berghuis, Soechitram, Hitzert, Sauer, & Bos, 2013; de Vries, van der Veere, Reijneveld, & Bos, 2013; Yuge et al., 2011). No such association was established in our study. Neither gestational age nor birth weight or any other perinatal risk factor was found to be significantly related to the MOS or fidgety movements, notwithstanding the fact that 27 infants (57%) were born preterm. Nor was the percentage of aberrant fidgety movements and/or a lower MOS increased in the 22 infants with CHD, although toddlers with Down syndrome and CHD were reported to have a higher percentage of motor, language and cognitive deficits at 12–14 months than toddlers with Down syndrome and a normal heart structure (Alsaied et al., 2016; Visootsak et al., 2016). In our study, neither aberrant fidgety movements nor a reduced MOS were attributable to preterm birth or CHD. Only two variables were shown to affect movements and postures at 3–5 months: caesarean section and hyperbilirubinaemia were related to a jerky movement character, and external rotation and abduction of hips was more common among infants born by caesarean section. In their study on healthy full-term infants during the first week after birth, Ploegstra, Bos, and de Vries (2014) compared the general movements of low-risk neonates born by vaginal delivery with those of neonates after caesarean section and found no difference. So far, we cannot explain how external rotation and abduction of the hips is related to caesarean section. As for the jerky movement character, there have no doubt been more revealing findings: Groen, de Blécourt, Postema, and Hadders-Algra (2005) reported 11 out of 15 children to have shown a normal neurological outcome in spite of jerky movements at the age of 2–4 months. Limitations to our study ~~~~~~~~~~~~~~~~~~~~~~~~ As various motor patterns such as abnormal fidgety or fidgety-like movements, lack of movements to the midline, and external rotation and abduction of the hips could be attributed to low muscle tone, we would also have needed to assess the infants’ muscle tone. Yet, the assessment of muscle tone is anything but unequivocal. Definitions are not standardised and inter-scorer agreement is prone to be low (Latash, 2007; Prechtl, 2001; Touwen, 1976) except for extremes. Diverging experience and procedures in the tonus assessment would no doubt have been additional challenges in a multicentre approach. Training and application of the Hammersmith Infant Neurological Examination (Haataja et al., 1999) was not available in most of the centres contributing to the study. This being a comparative study, one might consider our participants’ varied ethnic backgrounds as problematic. However, as we are dealing with spontaneous, i.e. endogenously generated movements, the respective care-giving practices are very unlikely to have had an impact on the early motor patterns in question. In fact, general movements have been assessed worldwide for more than 20 years with similar cross-cultural results (e.g., Bruggink et al., 2009; Luxwolda et al., 2014; Yang et al., 2012; Yuge et al., 2011). Nor did the various sensory stimulations affect the infants’ fidgety movements, which demonstrates their robust, environment-independent character (Dibiasi & Einspieler, 2002; Dibiasi & Einspieler, 2004). Of course, the infants’ different cultural backgrounds could be an issue with regard to the subcategory “posture”. Infants raised in a hammock, for example, seem to be faster in acquiring the “head in midline” position and/or a “symmetric body posture” (D.H. and C.E., personal observations), but none of our participants had experienced such exceptional practices, and this goes for all three groups.","The significance of this exploratory study lies in its minute description of the motor repertoire of infants with Down syndrome aged 3–5 months. During this time window, fidgety movements are the predominant spontaneous movement pattern and an excellent marker for the neurological outcome (Bosanquet et al., 2013; Einspieler et al., 2004; Einspieler et al., 2016). Infants with Down syndrome already show motor impairments at this early age, as evidenced by a significantly low MOS. Reassessing the same children as toddlers will show whether the high predictive power of fidgety movements found in infants with and without acquired brain injury (Bosanquet et al., 2013; Einspieler & Prechtl, 2005; Einspieler et al., 2004; Prechtl et al., 1997) also holds true for infants with a genetic disorder. A particularly important aspect will be whether the early spontaneous motor repertoire will also assist to predict the cognitive development of individuals with Down syndrome. A particularly important aspect that needs to be studied is the predictive power of the early spontaneous motor repertoire for the cognitive development of individuals with Down syndrome. But quite apart from this important aspect for clinicians and researchers, atypical postural features such as external rotation and abduction of the hips and/or internal rotation and pronation of the wrists call for earlier intervention."],["Background and Aims: Autistic children often recall fewer details about witnessed events than typically developing children (of comparable age and ability), although the information they recall is generally no less accurate. Previous research has not examined the narrative coherence of such accounts, despite higher quality narratives potentially being perceived more favourably by criminal justice professionals and juries. This study compared the narrative coherence of witness transcripts produced by autistic and typically developing (TD) children (ages 6–11 years, IQs 70+). Methods and Procedures: Secondary analysis was carried out on interview transcripts from a subset of 104 participants (autism = 52, TD = 52) who had taken part in a larger study of eyewitness skills in autistic and TD children. Groups were matched on chronological age, IQ and receptive language ability. Coding frameworks were adopted from existing narrative research, featuring elements of ‘story grammar’. Outcomes and Results: Whilst fewer event details were reported by autistic children, there were no group differences in narrative coherence (number and diversity of ‘story grammar’ elements used), narrative length or semantic diversity. Conclusions and Implications: These findings suggest that the narrative coherence of autistic children's witness accounts is equivalent to TD peers of comparable age and ability. --------------------------------------------------------------------------------","Previous work examining witness skills in autistic children has focused on the volume and accuracy of their recall for a witnessed event. No previous studies have examined whether these accounts are organised logically for the listener in terms of key ‘story grammar’ elements: information about the setting and the initiating event; the intentions of people in the event; the emotions, cognitions and goals of the people in the event; the actions that took place during the event; and the consequences arising from, and resolution of, the event. Such information may help criminal justice professionals to better understand what happened. Using closely matched, relatively large samples of autistic and non- autistic children (all with IQs 70+), the present study found no differences in the inclusion of story grammar elements between these two groups, despite the fact that autistic children reported fewer correct details about the witnessed event than non- autistic children. Similarly, the groups did not differ in terms of how long and semantically diverse their accounts were. The current findings are novel because they attest to the comparability of autistic and non-autistic child witnesses in terms of both narrative coherence and account length/semantic diversity (despite the autistic children reporting fewer details overall). The findings also support a growing body of literature attesting to the overall reliability of autistic child witnesses.","Autistic1 children and adults are more likely to encounter the criminal justice system than non-autistic individuals (Lindblad & Lainpelto, 2011; Turcotte, Shea, & Mandell, 2018; Woodbury-Smith & Dein, 2014). Further, because children and adults with developmental differences are at increased risk of violence, victimisation and abuse (Jones et al., 2012; Petersilia, 2001), it is vital to examine the quality of their evidence and encourage increased rates of reporting, investigation and prosecution. The current study examined the narrative coherence of information remembered by autistic children, comparing them to matched typically developing (TD) children. It extends a growing body of empirical research examining the accuracy and volume of recall for witnessed events in children and adults on the autism spectrum (largely those without intellectual disabilities). Previous findings indicate that children on the autism spectrum often recall a lower volume of information than TD peers of comparable age and ability (IQ) when interviewed about witnessed events (Almeida, Lamb, & Weisblatt, 2019; Bruck, London, Landa, & Goodman, 2007; Henry, Messer et al., 2017; Mattison, Dando, & Ormerod, 2015; McCrory, Henry, & Happé, 2007); although findings for autistic adults are more complex (see review by Maras & Bowler, 2014). There is evidence that group differences in volume of recall (in child and adult samples) are less apparent in more structured interviews (Henry, Crane et al., 2017; Maras & Bowler, 2010), or when additional supports (more specific questioning, physical reinstatement of context, or concrete visual prompts) are provided at recall (Maras & Bowler, 2012a, 2014; Mattison, Dando, & Ormerod, 2018). Importantly, eyewitness information provided by autistic children can be as accurate as that of comparable peers (Almeida et al., 2019; Bruck et al., 2007; Henry, Messer et al., 2017; Henry, Crane et al., 2017; McCrory et al., 2007); although accuracy levels may vary with interview type (Mattison et al., 2018) and the findings are less consistent in autistic adults (Maras & Bowler, 2010, 2011, 2012b; Maras, Memon, Lambrechts, & Bowler, 2013). Critically, autistic people are not more suggestible than non-autistic people (Bruck et al., 2007; Maras & Bowler, 2011, 2012b; McCrory et al., 2007; North, Russell, & Gudjonsson, 2008), despite many legal professionals believing this to be true (see George, Crane, Bingham, Pophale, & Remington, 2018, for a survey on this topic in UK barristers). They may, however, be more compliant (Chandler, Russell, & Maras, 2019; North et al., 2008). To our knowledge, there is no current literature assessing the narrative coherence of witness accounts provided by autistic children. Narrative coherence refers to ‘a global representation of story meaning and connectedness’ (Diehl, Bennetto, & Young, 2006) and is not the same as the amount of information recalled about an event, or its accuracy (Brown, Brown, Lewis, & Lamb, 2018; Feltis, Powell, & Roberts, 2011; Reese et al., 2011). A coherent account can be full or sparse in terms of evidential details, but will show a degree of organisation and structure relating to the context, content and characters associated with an event. Reese et al. (2011) suggest that a coherent narrative is ‘one that makes sense to a naïve listener’ (p.425, emphasis original), and describe developmental changes that occur during childhood in terms of increased complexity and number of narrative features (see also Berman & Slobin, 1994). Narrative coherence is often overlooked in research into witness recall, yet could substantially impact the degree to which a child’s evidence can be easily understood by members of the criminal justice system. For example, coherent accounts may appear more meaningful and credible to jurors (Brown et al., 2018; Feltis et al., 2011; Feltis, Powell, Snow, & Hughes-Scholes, 2010; Gentle, Milne, Powell, & Sharman, 2013; Murfett, Powell, & Snow, 2008). Further, in England and Wales, narrative coherence is used by the police (to establish relevant 'points to prove') and the Crown Prosecution Service (when making decisions about whether or not to authorise charges). Studies looking at narrative coherence in TD children’s witness accounts have adopted a ‘story grammar’ approach (Stein & Glenn, 1979) to capture higher order hierarchical structure, organisation and coherence (often described as ‘macrostructure’ or ‘global structure’). This approach looks for key elements that make accounts clear, organised and understandable for the listener via the inclusion of a number of logically ordered story elements. Seven elements are coded including: setting (contextual details about the event location to orient the listener); initiating event (how the event began); internal response (emotions, cognitions and goals of the people in the event); plan (the intentions of the people affected by the initiating event); action/attempt (the activities that constituted the event); direct consequence (the outcome/s of the event); and resolution (what happened at the end of the event). Many four- to eight-year-old children with TD can provide at least some elements of story grammar when recalling a witnessed event in an open-ended free recall interview (e.g., action/attempt, initiating event, and direct consequence details), although few children in this age range include internal response, plan, setting or resolution elements (Feltis et al., 2011). Using the story grammar framework, Westcott and Kynan (2004) conducted a secondary analysis of investigative interviews with children suspected of being sexually abused, finding that although children included basic story grammar components such as ‘setting’, their narratives were often “incomplete, ambiguous and disordered” (p.37). Several authors have used the story grammar approach successfully to assess narratives about witnessed events in children with intellectual disabilities (ID), finding that some children with ID include proportionately fewer story grammar elements than comparison groups (Brown et al., 2018; Gentle et al., 2013; Murfett et al., 2008). The current study extended the story grammar approach to look at narrative coherence in witness transcripts produced by children on the autism spectrum. An interview comprising only open-ended questions was used, given that such methods are most likely to elicit story grammar elements (Feltis et al., 2010; Snow, Powell, & Murfett, 2009). Children between the ages of 6 and 11 years with and without an autism diagnosis were included, as at least some markers of story grammar are present in witness accounts throughout this age range (Brown et al., 2018; Westcott & Kynan, 2004). We did not look at developmental increases in the inclusion of story grammar elements as these are well-established (e.g., Brown et al., 2018; Feltis et al., 2011); the aim was to compare groups of autistic and TD children matched on age (as well as IQ and receptive language). Previous literature on the broader narrative skills of autistic children and adults presents a conflicting picture, largely noting difficulties in some areas but not others (e.g., Banney, Harper-Hill, & Arnott, 2015; Capps, Losh, & Thurber, 2000; Diehl et al., 2006; King, Dockrell, & Stuart, 2013) or hardly any differences at all (e.g., Capps, Kehres, & Sigman, 1998; Norbury & Bishop, 2003; Tager-Flusberg & Sullivan, 1995; Young, Diehl, Morris, Hyman, & Bennetto, 2005). Studies generally compare autistic children or adults to age and language-matched (or verbal ability-matched) TD children or adults (or in some cases language-matched children with developmental delays). The areas of difficulty or difference identified are varied: less use of complex syntax; reduced use of evaluative devices; reduced use of causal explanations; greater numbers of ambiguous nouns and pronouns; shorter mean length of utterance; fewer different main body words and word roots; less complex ‘high point’ structure; more bizarre or idiosyncratic contributions; less complex episodic structure; and focus on details rather than gist (Banney et al., 2015; Barnes & Baron-Cohen, 2012; Capps et al., 1998, 2000; Goldman, 2008; King et al., 2013; Lee et al., 2018; Losh & Capps, 2003; McCabe, Hillier, & Shapiro, 2013; Norbury & Bishop, 2003; Pearlman-Avnion & Eviatar, 2002). Baixauli, Colomer, Rosello, and Miranda (2016) reflected this variability in their meta-analysis of narrative production tasks in autistic and non-autistic children, reporting small, moderate and large effect sizes over a range of narrative indices. Importantly, this variability in results may depend on the tasks used. Open-ended tasks, such as describing personal narratives, recalling orally presented fairy tales, or making up a story to go with an emotionally ambiguous picture, tend to reveal greater difficulties for autistic children and adults than more structured and guided tasks such as narrating the story to a ‘wordless picture book’ (King et al., 2013; Lee et al., 2018; Losh & Capps, 2003; Losh & Gordon, 2014; Tager-Flusberg & Sullivan, 1995). Wordless picture book tasks often reveal almost no group differences at all (e.g., Losh & Capps, 2003; Norbury & Bishop, 2003), and may reduce the cognitive demands of storytelling by scaffolding memory, attention, story organisation and language production via providing temporally sequenced visual cues. Importantly, although in their meta-analysis Baixauli et al. (2016) reported no effects of narrative task, they could only assess two combined categories of narrative task given the limited number of available studies. Results also differ depending upon the nature of the comparison group. There are relatively few (or no) differences in narrative ability on some tasks between well-matched autistic and non- autistic children (e.g., Capps et al., 1998, 2000; Diehl et al., 2006; Losh & Capps, 2003; Norbury & Bishop, 2003; Tager-Flusberg & Sullivan, 1995; Young et al., 2005). Such matching, most commonly for verbal ability and age, ensures that group differences – if found – are not just a function of language skills or developmental level. When looking at results for the inclusion of the global elements of story grammar, previous research also presents a mixed picture. Norbury and Bishop (2003), for example, compared autistic and non-autistic children matched for age and non-verbal ability, but who showed differences on language measures including receptive vocabulary and grammar. There were no group differences in the inclusion of key story grammar elements. Diehl et al. (2006) included samples of autistic and TD children who all had ability levels in the 80 + IQ range and who were matched on age, IQ and language ability. They found no group differences in the inclusion of gist elements or the proportion of basic story elements recalled (although there were group differences in causal connectivity). Similarly, Lee et al. (2018) found no group differences in the inclusion of story grammar elements between autistic and TD adults (the groups differed on verbal IQ so this variable was controlled in the analyses). Nonetheless, others have described differences in the inclusion of story elements between autistic and non-autistic groups (Banney et al., 2015; Goldman, 2008; Losh & Capps, 2003; see also Pearlman-Avnion & Eviatar, 2002, although groups in this study were not matched for language); found differences in the types of information produced (e.g., reductions in the preponderance or amount of ‘gist’ rather than ‘detail’ information: Barnes & Baron- Cohen, 2012; McCrory et al., 2007); or reported group differences in a combined meta- analysis of relevant studies (Baixauli et al., 2016). In evaluating previous research into narrative coherence in individuals on the autism spectrum, several difficulties emerge. First, studies have used different elicitation methods, ranging from wordless picture books (Losh & Capps, 2003), to describing autobiographical memories (King et al., 2013), to recalling sections from television programmes (Barnes & Baron-Cohen, 2012) or making up a story to an ambiguous picture (Lee et al., 2018). Only one study assessing narrative coherence in autistic individuals has looked at eyewitness memory (Pearlman-Avnion & Eviatar, 2002, using a narrated slide show), which is a more realistic everyday remembering task. In the current study, a secondary analysis of the story grammar elements produced in witness transcripts collected in a previous study of autistic and non-autistic children (Henry, Messer et al., 2017) is presented. Second, the approach to matching has differed. Some studies have matched for chronological age, intelligence (IQ – usually verbal but sometimes non-verbal), and one or more aspects of language ability (e.g., Banney et al., 2015); others have used more than one comparison group (e.g., with matching for chronological age and IQ and language ability and IQ, King et al., 2013); and others have matched on some but not all of these variables (e.g., Capps et al., 1998, 2000; Norbury & Bishop, 2003; Tager-Flusberg & Sullivan, 1995; Young et al., 2005). It may be difficult to match on multiple indices, particularly if language ability is included, given the heterogeneity and variability of language skills found in autistic individuals (Kwok, Brown, Smyth, & Oram Cardy, 2015; Taylor, Maybery, & Whitehouse, 2012). Indeed, some studies do not attempt to do so (Goldman, 2008) or control key variables statistically (Lee et al., 2018). In the present study, all available participants were included from the larger study who could be matched on age, IQ and one aspect of language (receptive vocabulary), although other aspects of language ability varied significantly between groups. This provided a test of narrative skill in autistic children who were similar in age, intellectual ability and receptive vocabulary to a comparison group of non-autistic children. Finally, there are few studies with larger sample sizes. The current sample size (52 autistic children, 52 TD children) was larger than in previous studies, and a conservative significance level was used to account for multiple group comparisons (see Banney et al., 2015). We also measured the length of the transcripts in terms of total number of words, and their semantic diversity in terms of number of different words. This is important, as although many previous studies have found no significant differences between autistic and typical individuals in story length or semantic diversity, albeit using slightly varying measures (Banney et al., 2015; Diehl et al., 2006; Losh & Capps, 2003; McCabe et al., 2013; Norbury & Bishop, 2003; Tager-Flusberg & Sullivan, 1995; Young et al., 2005), there are some reports of differences (Baixauli et al., 2016; Capps et al., 2000; King et al., 2013; Lee et al., 2018). Finally, we included preliminary correlational analyses between story grammar measures and other cognitive and language variables, although few relationships were expected given limited findings in previous research (Banney et al., 2015). Matching TD and autistic groups on several indices including age, IQ and receptive language level may minimise differences in narrative performance. Therefore, it was hypothesised that there would be no significant differences between the groups with regard to the number and diversity of ‘story grammar’ elements included in their accounts. We also tentatively predicted no significant differences in overall length or number of different words.","The data used in this project were taken from interviews carried out as part of a larger study conducted by Henry, Messer et al. (2017), which compared eyewitness memory skills in children with and without autism diagnoses. The original study included 272 participants (162 boys and 110 girls, aged between 76 months and 142 months). Of these children, 71 had a formal diagnosis of autism, obtained independently of the research study by a suitably qualified professional. Children were recruited from mainstream primary schools or special educational needs schools in Greater London or South-East England. Selecting a meaningful comparison group for studies of children on the autism spectrum is not straightforward when discrepant ‘peaks and valleys’ represent a common cognitive profile (Burack, Iarocci, Flanagan, & Bowler, 2004). We tried to avoid ‘over-matching’ where statistical bias is accidentally caused by matching for what is believed to be a confounding variable but is actually a core feature of the autism profile. Burack et al. (2004) advise matching on a subset of features, but caution against matching ‘out’ the key features of autism. In line with previous literature, matching was carried out for age, IQ and one measure of language (receptive vocabulary). Receptive vocabulary was selected following Goldman (2008), who reported non-significant differences in receptive vocabulary scores in their samples, together with substantial differences on two subtests from a broader language battery (the Clinical Evaluation of Language Fundamentals, CELF-4 UK, Semel, Wiig, & Secord, 2006). All children who recalled at least three items of correct information in their interview accounts of the witnessed event were eligible for inclusion (three children were excluded at this stage: two autistic, one TD). Individual matches within +/- 10 points for both IQ and receptive vocabulary and +/- 6 months for age were hand selected. The resulting sub- sample comprised: 104 children (52 in the TD group, and 52 in the autism group) matched on age, IQ and receptive language ability. There were 83 boys and 21 girls (TD group: 38 boys and 14 girls; autism group: 45 boys and 7 girls). It was not possible to match the groups exactly on gender, as the autism group in the original sample included a high proportion of males (62 boys and 9 girls). The composition of the current sample reflects the wider autistic population, where males are estimated to outnumber females (Loomes, Hull, & Mandy, 2017). Table 1 provides mean scores on all background variables. Mervis and Klein- Tasman (2004) suggest that for groups to be well-matched, the group distributions on the control variable in question should overlap, with a p-level of at least .50 on the test of mean differences. An independent samples t-test was used to compare the groups, as these data were normally distributed. On all three indices, age, IQ and receptive vocabulary, the groups were closely matched (see Table 1). Data are also reported on other measures of language, memory and attention (not matched between groups), as well as the total number of correct items of information recalled in the initial interview (see Table 1). Mirroring the original study findings (Henry, Messer et al., 2017), the groups differed in the number of correct details recalled about the witnessed event (Table 1), with autistic children (M = 25.71, SD = 14.37) recalling fewer correct details than TD children (M = 33.42, SD = 14.24), t(102)=-2.75, p = .007. However, the analyses of story grammar elements and length/diversity of narratives did not focus solely on correct details, but included the full transcripts for all children.","Participants’ intellectual abilities were assessed in the original study using the Wechsler Abbreviated Scale of Intelligence (WASI-II; Wechsler & Zhou, 2011). An estimate of full-scale IQ was obtained using ‘Vocabulary’ from the Verbal Comprehension Index and ‘Matrix Reasoning’ from the Perceptual Reasoning Index. As noted, following matching, the groups did not differ on this measure. Language level of participants was measured using several tasks from the original study. Receptive vocabulary was assessed using the British Picture Vocabulary Scale - Third Edition (BPVS-3; Dunn, Dunn, & Styles, 2009), which requires children to select the correct picture from a choice of four upon hearing a word spoken. Two subtests from the Expressive Language Test - 2 (ELT-2; Bowers, Huisingh, Logiudice, & Orman, 2010) included: ‘Sequencing’, to test narrative ability, and ‘Grammar and Syntax’, to test grammatical morphology. A further two subtests were taken from the Clinical Evaluation of Language Fundamentals - 4th edition (CELF-4 UK; Semel et al., 2006). ‘Recalling Sentences’ assesses the ability to accurately repeat sentences using internalised grammatical structure, and ‘Formulated Sentences’ assesses the ability to produce full, grammatically accurate, and meaningful sentences about picture stimuli. When compared on the language measures, the two participant groups were well-matched on receptive vocabulary ability (BPVS), as expected (Table 1), with no significant difference between the autism and TD groups. However, significant group differences (or marginal differences) remained in scores on the other language measures (ELT-2 ‘Sequencing’; ELT-2 ‘Grammar and Syntax’; CELF-4 ‘Recalling Sentences’; CELF-4 ‘Formulated Sentences’). An assessment of general memory ability from the Test of Learning and Memory (2nd Edition, TOMAL-2, Reynolds & Voress, 2007) included: Memory for Stories; Paired Recall; Facial Memory; and Visual Sequential Memory. Table 1 includes composite scores for the Verbal Memory Index, on which groups did not differ, and the Nonverbal Memory Index, on which the autistic children obtained lower scores than the TD children (although no differences were present on the combined Composite Memory Index). Three assessments of focused, sustained and sustained-divided (dual task) attention completed the test battery (Sky Search; Score! ; and Sky Search Dual Task from the Test of Everyday Attention for Children, TEA-Ch, Manly, Robertson, Anderson, & Nimmo-Smith, 1999). There were no group differences on the attention measures. Materials and procedure Ethical approval for the original study and the present secondary data analysis study was granted by the Research Ethics Committees of the universities at which the research took place. All parents or guardians provided informed written consent for their children to take part. The original research study involved data being collected from both TD and autism groups across several phases, in a manner intended to replicate the processes involved in gathering evidence from eyewitnesses in the criminal justice system (evidence gathering statements, investigative interviews, identification line-ups, cross- examinations). Only the first set of interview data (evidence gathering statements, or ‘Brief Interviews’) administered on the same day as the event was witnessed, are relevant to the current paper. These initial questions about the event were included to mimic a response officer’s initial contact with a witness. The procedure used in this phase of the study is summarised briefly below. For a more detailed description, please see Henry, Messer et al. (2017). Experimental procedure Participants were shown a staged event, featuring two actors who gave a short presentation about what school was like in Victorian times. For practical and logistical reasons, this was either viewed live during a school assembly, or a video of the performance was shown to the children. There were no recall differences between live and video presentation for autistic or TD children (Henry, Messer et al., 2017) so these data were combined. The talk consisted mainly of facts about Victorian schools, but involved a staged minor crime, where either a phone or a set of keys was ‘stolen’ by one of the actors (the ‘theft’ was later explained as a misunderstanding). Participants were randomly assigned to one of two versions of the talk, which were virtually identical but used different materials and gave alternative names to the actors. The rationale was to give some indication of the generalisability of the results. As there were no differences between versions, data were collapsed over this variable (Henry, Messer et al., 2017). After viewing the staged event, participants were interviewed individually on the same day. Interviewers used a standard protocol with an initial question designed to elicit a free recall account (‘Tell me what you remember about what you just saw’). This was followed by a series of open-ended questions that could be used to prompt further information depending on the child’s response (e.g. ‘Who was there?’ ; ‘What did they do?’). Finally, participants were asked if they could remember anything else. This final prompt was repeated until the child could no longer provide any new items of information. Interviews were audio-taped and transcribed, then coded by two independent coders according to the narrative frameworks described below. The coders were blind to group status. ‘Story grammar’ narrative analysis The story grammar framework developed by Murfett et al. (2008) was used to analyse participants’ accounts by tallying the number of narrative features included. ‘Correctness’ of information in the transcripts was not taken into account, as this would not be possible in a real case when jurors and legal professionals are unlikely to know what actually happened to a victim/witness. Further, story grammar is a ‘framework’ for organising event details and, as such, is scored independently of accuracy (Feltis et al., 2011). Seven narrative elements were scored: Setting, Initiating event, Internal response, Plan, Action/attempt, Direct consequence, and Resolution (see Table 2). Elements were scored based on the number of times they appeared in the narrative, regardless of their accuracy (e.g., a point would be awarded for setting if the child mentioned the location of the event, even if this was incorrect). The total number of story grammar elements and the number of different story grammar elements (maximum 7) used were noted. Additional structural measures Two structural measures were hand coded: the total number of words in each account, and the total number of different words used (excluding non-word utterances such as ‘um’), to obtain a broad comparison of the overall length and semantic diversity of the narratives produced by each group. Reliability analyses All transcripts were coded independently by two coders, following discussion to agree the coding rules. Intra-class correlation coefficients were calculated to assess the reliability of coding, and these indicated good to excellent reliability for all but one story grammar element (Direct consequence): Setting: ICC = .94; Initiating event: ICC = .83; Internal response: ICC = .81; Plan: ICC = .92; Action/attempt: ICC = .82; Direct consequence: ICC = .08; Resolution: ICC = .80. For the aggregated story grammar element scores, reliability was also good: Total number of grammar elements, ICC = .94; Number of different grammar elements, ICC = .93. Intercoder agreement for the word count measures was excellent: Total number of words, ICC = 1.00; Number of different words, ICC = 1.00. Given the poor agreement for Direct consequence, the first coder looked at the discrepancies in coding (coder 2 noted more instances of this element than coder 1) and recoded data based on commonalities in coding criteria. After recoding, agreement was much higher, ICC = .98. Nevertheless, given the initial discrepancies, we suggest caution in interpreting the results for Direct consequence. First coder ratings were used for all data analyses.","Table 3 provides means, medians and ranges for each story grammar element and the structural measures (total number of words, number of different words). As group comparisons across several areas were conducted, a more stringent significance level (p < .01) was adopted to limit the chances of incorrectly rejecting the null hypothesis (recommended by Banney et al., 2015). Data for the structural measures of total number of words and number of different words were normally distributed, so independent samples t-tests were performed to compare groups. ‘Total number of words’ showed no significant difference between groups [t(102)=-1.17, p = .25, d = .23]: TD group (M = 234.46, SD = 111.08); autism group (M = 207.75, SD = 121.72). There was no significant group difference for ‘Number of different words’ [t(102)=-1.53, p = .13, d = .30]: TD group (M = 97.52, SD = 31.16); autism group (M = 87.15, SD = 37.67). Given the groups did not differ on overall length or diversity of their narratives, data on story grammar elements were analysed using raw scores. Table 3 shows that means and medians for many story grammar elements were low, with medians of zero in one or both groups for initiating event, internal response, plan, direct consequence, and resolution. Kolmogorov-Smirnov normality tests indicated that scores for the story grammar elements and their totals violated the parameters of normality (either for both groups or one group) on every measure, therefore, non-parametric tests were used (Mann-Whitney U). There were no significant differences in the distributions for any of the story grammar elements, or for the total scores [Setting, U = 1651.5, z = 2.02, p = .04, d = .39; Initiating event, U = 1364, z = .09, p = .93, d = .02; Internal response, U = 1235, z = -0.88, p = .38, d = .15; Plan, U = 1591, z = 1.88, p = .06, d = .31; Action/attempt, U = 1340, z=-0.08, p = .94, d = .02; Direct consequence, U = 1382, z = .21, p = .83, d = .04; Resolution, U = 1587.5, z = 2.19, p = .03, d = .30; Total number of story grammar elements, U = 1421, z = .50, p = .65, d = .09, Number of different story grammar elements, U = 1537.5, z = 1.23, p = .22, d = .24]. In many cases, medians across the two groups were identical (see Table 3). To investigate whether any background variables related to Total number of story grammar elements or to Number of different story grammar elements, preliminary Spearman non-parametric correlations were carried out. For the TD and autism groups, Total number of story grammar elements correlated with Total correct details recalled in the interview: Autism group r = .71; TD group r = .79 (ps < .001). Number of different story grammar elements also correlated with Total correct details recalled in the interview: Autism group r = .76; TD group r = .64 (ps < .001). There were correlations between Total number of story grammar elements and age in the TD group, which failed to reach significance in the autism group: Autism group r = .25, p = .07; TD group r = .41, p = .003, and between Number of different story grammar elements and age in the TD group, which again failed to reach significance in the autism group: Autism group r = .27, p = .055; TD group r = .41, p = .003. No other correlations were significant.","This study investigated the narrative coherence of eyewitness accounts in autistic and non-autistic children (6–11 years) matched on age, IQ and receptive language ability. We replicated a previous finding that autistic children recalled fewer correct items of information than non-autistic children in a brief interview about a staged event involving a minor crime, witnessed earlier that day (Henry, Messer et al., 2017). Despite this difference in volume of correct recall, no significant group differences emerged when a ‘story grammar’ framework was used to assess the presence and number of key story elements in participants’ accounts (incorporating all information mentioned, whether correct or not). The findings confirm the utility of previous work using the story grammar framework to evaluate eyewitness accounts of children with developmental differences (Brown et al., 2018; Gentle et al., 2013; Murfett et al., 2008), and provide support for claims that there are relatively few (or no) differences in narrative ability on some tasks between well-matched groups of autistic and non-autistic children (e.g., Capps et al., 1998, 2000; Diehl et al., 2006; Losh & Capps, 2003; Norbury & Bishop, 2003; Tager-Flusberg & Sullivan, 1995; Young et al., 2005). On metrics assessing the length in words and semantic diversity (number of different words) of the accounts, there were also no group differences in performance. This supports much previous literature (e.g., Banney et al., 2015; Diehl et al., 2006; Losh & Capps, 2003; McCabe et al., 2013; Norbury & Bishop, 2003; Tager-Flusberg & Sullivan, 1995; Young et al., 2005). The present findings also suggest that, although children on the autism spectrum may recall fewer items of correct information about an event, their overall narrative accounts are not shorter or less semantically diverse, nor are they less coherent in terms of narrative structure. Taken together with findings that eyewitness accuracy in autistic children is generally as high as in comparable non- autistic children (2017b, Almeida et al., 2019; Bruck et al., 2007; Henry, Messer et al., 2017; McCrory et al., 2007; although accuracy levels may vary with interview type – see Mattison et al., 2018), the present results provide further evidence for the overall reliability of autistic child witnesses. Narrative coherence adds important information, because it may affect the degree to which members of the criminal justice system understand a child’s evidence. This could be via impacting on how meaningful and credible the child appears to jurors (2011, Brown et al., 2018; Feltis et al., 2010; Gentle et al., 2013; Murfett et al., 2008), helping the police to establish relevant 'points to prove', and/or contributing to decisions by the Crown Prosecution Service about whether or not to authorise charges. The most commonly included story grammar element in children’s narratives related to descriptions of what happened in the event (i.e., Action/attempt details, Mdn = 5.5 autism group; Mdn = 6.0 TD group). This confirms previous observations that Action/attempt details are the most commonly elicited story grammar element. Feltis et al. (2011) found about 50 % of story grammar elements in immediate recall were Action/attempt details, close to the 60–65 % recorded in the current study. Nearly all other story grammar elements had median scores at or close to zero, indicating that children in the current age range (6–11 years) often omitted several story grammar elements. For Setting details, the median score of 2 items in both groups indicated that many children included information about where and when the event took place (see also Westcott & Kynan, 2004). However, the median number of types of story grammar elements included was only 3 out of a possible 7, underlining the fact that maturation in narrative coherence was not complete in the current sample of 6–11 year old children. Further research should assess narrative coherence in witness accounts of older children who include more story grammar elements, to negate the possibility that low scores on some measures reduced our ability to detect group differences. Nevertheless, group differences on story grammar measures with higher scores (e.g. total scores, action/attempt) were still absent, increasing confidence in the findings. Although autistic (relative to matched non-autistic) children produced narratives of similar coherence (based on an analysis of story grammar elements), length (based on number of words) and semantic diversity (based on number of different words), the type of narrative task is important. Losh and Capps (2003) noted that experimental findings may not reflect children’s narrative competence “within the less structured and more socially demanding contexts of daily life” (p.248). In the same way that wordless picture book tasks may reduce cognitive demands of storytelling by scaffolding memory, attention, story organisation and language production using temporally sequenced visual cues, adult-led interviews are structured interactions in which specific narrative features may be prompted through direct questions (e.g., ‘Where did this happen?’). This is not necessarily typical of how children’s narratives are produced in real conversational exchanges, although it does reflect how forensic interviews might proceed (although note that we did not use full investigative interviews, which would usually occur later than the ‘same-day’ brief evidence gathering statements used in the current study). Further, since autistic children often experience difficulties with emotional regulation (Samson et al., 2014), they might perform less well when describing a ‘real’ witnessed event, particularly if it induced heightened emotions or distress. Therefore, caution is required when generalising the current findings to more emotionally charged situations such as are likely to be encountered in criminal investigations. Finally, as the current groups were matched on age, general ability and one aspect of language (receptive vocabulary), it is important to note that the findings could be different in less closely matched samples. Preliminary correlational analyses assessed whether background age and cognitive measures related to the number or diversity of story grammar elements included in children’s accounts. Correlations between both of these measures and the number of correct details recalled in the interview suggested that volume of correct recall is a good indicator of the extent to which both autistic and TD children include key story elements. Higher volume of recall, better structure and greater narrative coherence could all have an impact on the credibility of a witness before jurors (see also Henry, Ridley, Perry, & Crane, 2011). The relationships between number and diversity of story grammar elements and age (significant only in the TD group) confirmed the known developmental improvements in the use of story grammar elements (Feltis et al., 2011). Summary ~~~~~~~ This study investigated the narrative skills of 52 autistic children aged 6–11 years (IQs 70+), comparing them to 52 TD peers matched on age, IQ and receptive language ability. Participants had been interviewed using open-ended prompts shortly after viewing a staged mock crime event, and interview transcripts were coded using existing analytical frameworks for narrative features (story grammar elements; narrative length; semantic diversity). There were no significant differences between the autism and TD groups in terms of narrative skills on any of these measures. The findings add to, and extend, previous reports that when children on the autism spectrum are well-matched to peers with TD on cognitive and some language measures, they exhibit a similar level of narrative skill. The current findings also attest to the comparability of autistic and non-autistic child witnesses in terms of accuracy, account length and narrative coherence, if not in volume of correct information recalled."],["Researchers have shown an interest in the aggregated Big Five personality of U.S. states, but typically they have relied on scores from a single sample (Rentfrow, Gosling, & Potter, 2008). We examine the replicability of U.S. state personality scores from two studies (Rentfrow et al., 2008; Rentfrow, Gosling, Jokela, & Stillwell, 2013) across a total of seven samples, two of them new. Same-trait correlations across samples are, on average, positive for all five traits, indicating score agreement. Additionally, three traits (Conscientiousness, Neuroticism, and Openness) show strongly consistent patterns of correlations with sociodemographic variables across samples. We find rank order stability in state personality scores for a 16-year period (1999–2015). --------------------------------------------------------------------------------","The beginning of the twenty-first century has seen an explosion of interest concerning geographical variation in personality within the United States. Before the modern era of the internet, there were a few studies that examined aggregate psychological differences by U.S. cities or regions (e.g., Krug & Kulhavy, 1973; Thorndike, 1939). However, with the widespread adoption of the internet in the U.S., several psychology labs have collected samples of hundreds of thousands of participants across the country via online personality assessments (e.g., Revelle, Wilt, & Rosenthal, 2010; Revelle et al., 2016; Srivastava, John, Gosling, & Potter, 2003). These samples, although not representative of the U.S. population, are more diverse than traditional methods of data collection (Gosling, Vazire, Srivastava, & John, 2004). They also have enough statistical power for analyses to detect small effects between a large number of regional groups. These online assessments typically use self-report personality assessment models based on the Big Five (Goldberg, 1990), a widely-accepted taxonomy that organizes most individual differences into five broad traits: Conscientiousness, Agreeableness, Neuroticism (sometimes referred to by its polar opposite, Emotional Stability), Openness (sometimes called Intellect), and Extraversion. Studies that have used the state scores data from Rentfrow et al. (2008) have assumed that these state scores were representative of the actual personality scores of the states’ residents. For example, these studies assumed that the Extraversion score for Oklahoma accurately represented the mean Extraversion score of all Oklahomans. This assumption could be problematic for at least three reasons. First, these state scores, although based on many participants, are not immune to sampling bias. Idiosyncratic methods of participant selection could lead to a lack of replication in other samples. Second, even if the original scores were accurate, the personality of some states’ residents may have changed since the original study’s data were collected (1999–2005). Third, Rentfrow et al. (2008) measured personality with the Big Five Inventory (BFI; John & Srivastava, 1999). The ranks and standardized scores from Rentfrow et al. (2008) may not generalize across other measures of the Big Five. Therefore, it is useful to examine the extent to which the state-level scores of Rentfrow et al. (2008) will replicate, and to estimate the effect of disagreement attributable to differences in participant recruitment methods, change over time, and Big Five measures. Additionally, because these state scores are often reused in other studies, it is critical to determine the extent to which correlations between state personality and sociodemographics replicate in spite of differences in samples. A study dedicated to replicating Rentfrow et al. (2008) has not yet been reported. Rentfrow, Gosling, Jokela, and Stillwell (2013) compared state scores across five samples, including the sample from Rentfrow et al. (2008). Details from this effort were brief because the focus of the paper concerned the personality profiles of broad regions of the U.S. Findings were not thoroughly discussed, and many of the results were relegated to the supplemental materials. However, their analyses indicated that “there were no clear or consistent statewide differences in any of the scale properties,” and state scores were “reliable and generalizable” (Rentfrow et al., 2013, p. 1003). The study also found that in general, the same traits correlated with the same sociodemographic variables at similar magnitudes across the five samples. The current study used the five samples reported from Rentfrow et al. (2013) and added two new large samples. The goals of the study were as follows: One, for a point of comparison with Rentfrow et al. (2013), determine the extent to which the two replication samples were representative of U.S. states (Section 3.1). Two, across all samples, estimate the reliability of state score differences by evaluating their intraclass correlations (Section 3.2). Three, estimate the effect size of same-trait convergent correlations across all samples for each Big Five trait (Section 3.3). Within this goal, estimate the separate effects of three possible sources of attenuation: differences due to recruitment methods (Section 3.3.1), time of data collection (Section 3.3.2), and personality inventories (Section 3.3.3). Four, from these estimates, determine whether the personality of U.S. states had maintained rank order stability (i.e., relative to each other, states’ personalities did not change) from 1999 to 2015 (Section 3.3.4). And finally, determine the replicability of correlations between state-level personality scores and sociodemographic variables (Section 3.4). Samples 1–5 ~~~~~~~~~~~ Samples 1–5 were originally analyzed in Rentfrow et al. (2013). The samples were collected during different time periods, as part of different research projects, using different personality inventories (Table 1). All five samples had aggregate measures of personality based on Likert-type scales for the 48 contiguous states, as well as Washington, DC. In total, the samples were collected over an 11-year period (1999–2010), with a range in sample size from 18,182 to 612,140. Samples 1–4 were online personality assessments that used self-selecting participant recruitment. Sample 5 was an online assessment that used a recruitment method similar to random digit-dialing to select a representative sample of registered voters. A more detailed summary of Samples 1–5 can be found in Rentfrow et al. (2013), where each sample is referred to by the same name used in this study. Through correspondence with Rentfrow, we received participant counts and unadjusted mean scores for states, for each sample. Other data on the five samples, such as interclass correlations, were collected from Rentfrow et al. (2013) and the supplemental materials. SAPA samples ~~~~~~~~~~~~ The last two samples were from the Synthetic Aperture Personality Assessment (SAPA) project, an online non-commercial personality assessment (https://sapa-project.org; Revelle et al., 2016). For their participation, participants received feedback concerning their personality. Each sample covered an approximate five-year period of time and was named for the last year in which data were collected. The SAPA2010 sample was collected from April 2006 to August 2010. The SAPA2015 sample was collected from August 2010 to December 2015 (Table 1).","Participants were screened to ensure that entries beyond their first were not included in the analysis. Duplicate entries taken in a single internet browser session were removed. Participants who reported having previously taken the assessment also were excluded. Since this study was concerned with state-level analysis, it was also necessary to remove participants who reported not being from one of the 50 U.S. states or Washington, DC. Concerning educational attainment, 40% of the SAPA2010 sample reported being an undergraduate at the time of assessment, while 28% had attained at least a bachelor’s degree. In the SAPA2015 sample, 51% were current undergraduates, while 27% had attained at least a bachelor’s degree. Personality measures Most online personality assessments give every participant the same fixed set of items. Researchers typically analyze complete cases or use mean scores to impute the small amount of missing data. The SAPA project is radically different in this regard; participants receive a random sample of items from a pool of personality, cognitive ability, and interest inventories. Although items given within an inventory are random, sampling rates of inventories differ based on research goals of the SAPA project’s collaborators (Revelle et al., 2010, 2016). In this study, all of the personality items in the SAPA samples were from the International Personality Item Pool (IPIP; http://ipip.ori.org/), an online repository for public domain personality items and inventories (Goldberg, 1999; Goldberg et al., 2006). We assessed participants on these items using a 1–6 Likert-type scale (1 = “Very inaccurate”; 6 = “Very accurate”). The sole Big Five personality measure in the SAPA2010 sample was the Big Five Factor Markers (BFFM), a 100-item IPIP personality inventory based on the Goldberg (1992) conception of the Big Five. On average, participants took 48 BFFM items (48% of the inventory). In the SAPA2015 sample we examined four measures of Big Five personality. The first was the BFFM. On average, participants took 30 BFFM items (30% of the inventory). The second personality measure was the IPIP-NEO, a 300-item IPIP inventory based on the NEO- PI-R conception of the Big Five (Costa & McCrae, 1992). On average, participants took 26 IPIP-NEO items (9% of the inventory). The third personality measure was the SAPA Personality Inventory (SPI), a 75-item inventory of IPIP items in which Condon (2014) determined the “best” items for the Big Five through empirical analyses. On average, participants took 11 SPI items (15% of the inventory). The last measure in the SAPA2015 sample was the IPIP-HEXACO (Ashton, Lee, & Goldberg, 2007), a 240-item IPIP inventory based on the six-factor HEXACO framework, which adds the Honesty-Humility trait to the Big Five (Lee & Ashton, 2004). On average, participants took 26 IPIP-HEXACO items (11% of the inventory). The BFFM was considered to be the primary measure of personality in the SAPA2015 sample because participants took the most items from the inventory, both in terms of raw number and percent of total items in the inventory. However, an analysis in Section 3.3.3 utilized all the above personality measures of the SAPA2015 sample. In Section 3.2, two additional measures in the SAPA2015 sample were examined: the IPIP-NEO 10-item Openness facet of Liberalism, and the 60-item ICAR (International Cognitive Ability Resource) measure of cognitive ability (Condon & Revelle, 2014). On average, participants took 2 Liberalism items (20% of the inventory) and 14 ICAR items (23% of the inventory). U.S. Census Bureau data In order to weight correlations and determine the representativeness of SAPA samples, four measures from the 2010 U.S. Census were used: state populations (United States Census Bureau, 2015), state populations by ethnicity (United States Census Bureau, 2014), state populations by age (United States Census Bureau, 2016) and state populations by adult education (United States Census Bureau, 2016). Sociodemographic measures Thirteen state-level sociodemographic measures were selected to cover a similar breadth of criteria as reported in previous studies (Rentfrow et al., 2008, 2013). All sociodemographic measures analyzed were either per capita or percentage of state residents. There were two measures of physical health in 2008: cancer deaths (American Cancer Society, 2008) and heart disease deaths (Miniño, Murphy, & Xu, 2011); two measures of crime in 2008: violent crime and property crime (Federal Bureau of Investigation, 2009); four measures concerning adult employment by job fields in 2009: “arts, design, entertainment, sports, and media,” “business and financial,” “computer and mathematical science,” and “healthcare practitioner and technical” (Bureau of Labor Statistics, 2009); one measure of innovation: the number of patents issued in 2008 (United States Patent & Trademark Office, 2009); two measures regarding beliefs in 2008: self-identified political liberals and people who responded affirmatively to the question, “Is religion an important part of your daily life?” (Gallup, 2014); and two measures of well-being in 2013: an index of overall well-being and self-reported community recognition in the past 12 months (Gallup, 2016).","Individual personality scores in the SAPA samples were calculated using the simple mean of the observed items. State personality scores were then found by aggregating individual personality scores. State-level scores for the SAPA2010 and SAPA2015 samples, as well as a sample that combines the two, are available in Tables 6–8 of the supplemental materials. Each mean of correlations was calculated by converting correlations to z-scores, determining the mean, and transforming the mean z-score back into a correlation. Analyses were performed in the psych (Revelle, 2016) package and displayed using the psych and corrplot (Wei, 2013) packages in the R statistical system (R Core Team, 2016). Despite large individual sample sizes, aggregating personality decreased the number of “participants” to the number of states (51 in Section 3.1; 49–51 in Section 3.2; 49 in Section 3.3; and 48 in Section 3.4). In the case of 48 participants, a conventional criterion for statistical significance (p < .05) would require rs > ∣.28∣. Multiple comparisons were analyzed, which typically requires one to set an even higher threshold for statistical significance, in order to protect against an increased probability of finding a “significant” correlation by chance. We focused on patterns and means of correlations when possible because the standard error of a mean correlation decreases with the square root of the number of correlations that go into the mean. Reliability of state differences: individual and group variation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Just as small effects can be “statistically significant” in large samples, a small ICC1 and a large average group size will produce a large ICC2. James (1982) suggested using ICC1 as a criterion for aggregation. It would be unrealistic to expect that most of an individual’s personality would be due to state residence. However, studying the personality of states presumes that a non-trivial amount of individual variance in personality is explained by state variance in personality. Same-trait convergent correlations for state-level personality scores ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Each mean same-trait correlation was a composite of 21 correlations between pairs of the seven samples (Fig. 1). Samples differed in terms of their personality inventory, participant recruitment related to underlying research project, and the time period in which they were collected (Table 1). Any one of these differences could have led to attenuation of a same-trait correlation.2 We grouped correlations by these three differences to determine whether any of them were related to attenuation in same-trait correlations (Table 4). Where possible, we estimated the effect of one type of difference by only evaluating correlations whose samples differed in that one way. Replicability of state personality correlations with sociodemographics ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Weighted correlations (as described in Section 3.3) between personality scores of the seven samples and the thirteen sociodemographic variables were partialled, controlling for the other four personality scores within a sample (Tables 1–5 in the supplemental materials). These partial correlations were limited to the 48 contiguous states due to limitations of the secondary data.","State differences in Big Five personality may be used to predict important sociodemographic outcomes. These results replicate for three personality traits (Conscientiousness, Neuroticism, and Openness) across sampling methods, time periods, and personality inventories. In terms of agreement of scores across samples, same-trait correlations were highest for Neuroticism and Openness, but correlations for all five traits were positive, on average, indicating that state scores were somewhat robust. In Section 3.2, we showed that most of the personality variation was within states (small ICC1s), but between-state variation was still very reliable (large ICC2s). By the standard statistical criterion of “variance explained,” states appeared to be a poor candidate for grouping personality. The two replication samples were representative in terms of state populations and ethnicity within a state, were less representative of how educated states were, and were not representative of age groups within states. These representative analyses were similar to findings concerning the five samples from a previous study (Rentfrow et al., 2013). There are at least three reasons why Extraversion and Agreeableness were not found to have replicable correlations with sociodemographics. First, perhaps suboptimal sociodemographic measures were selected. Future studies might find that state-level Agreeableness and Extraversion have large, replicable correlations with sociodemographic measures not included in this study. While this is a possibility, several of the sociodemographic variables were selected based on previous research that showed moderate correlations with Agreeableness (crime and religiosity) and Extraversion (crime and business occupations). Second, the current study found that Agreeableness had a restricted range of correlations with the selected sociodemographic variables, which could indicate that correlations of a large magnitude are required to ensure replication at the state level. A third possible reason for lack of replicability is that lower-level facets within the traits differentially correlated with the sociodemographic variables. Rentfrow (2014) found that when Extraversion was split into two lower-level facets and correlated with six social indicators, the signs of the correlations were consistently opposite for the facets. Combining these kinds of opposing facets into a higher-level trait would result in the trait having smaller correlations and/or an inconsistent pattern of correlations with sociodemographic variables across samples. Future research should examine whether novel sociodemographics correlate more highly with Agreeableness and Extraversion, as well as whether facets within some traits differentially correlate with certain sociodemographic measures. Researchers interested in the links between personality and geography may find new, important insights by examining personality constructs outside or at a lower level than the Big Five traits. Preliminary evidence in this paper indicated that state residence accounted for more total variance in cognitive ability and Liberalism than the Big Five. Additionally, for correlational analyses with sociodemographic variables, the level of a personality characteristic should match the level of a sociodemographic variable (Wittmann, 1988). That is, it may be more theoretically appropriate to correlate certain sociodemographic variables with facets instead of the Big Five. In a similar way that personality constructs like the Big Five may be too broad in the study of geographical personality, geographical constructs may also be too broad. States may not be the optimal level for discovering geographical differences in personality. States share a common government, but they often cover large areas, and state lines can be arbitrary when grouping people by personality. For example, would we expect the personality of Paris, Illinois, a town of less than 10,000, to be more like Versailles, Indiana, or Chicago, Illinois? Both are roughly equidistant from Paris, but one is a town of less than 3000, and the other is the most populous city in the Midwest. Chicago is in the same state as Paris, so a state-level analysis would group both together and contrast them against another group comprised partly of Versailles and Indianapolis. It could be valuable for future research to compare the effect size of state aggregation to other regional groups, such as counties, metropolitan areas, cities, and neighborhoods. Personality scores aggregated by geographical regions could prove to be invaluable to researchers interested in exploring macro-level relationships between patterns of human behavior, environment, and social outcomes. These scores would represent static measures for these regions like any other sociodemographic variable. For example, one could imagine a scenario in which the Extraversion of Illinois in 2020 would be a matter of public record in the same way that its population will be. Self-reported personality items have components of affect, behavior, cognition, and desire (Wilt & Revelle, 2015) and are presumed to reflect observable patterns in the real world. Having aggregate measures of these complex processes could help researchers to better understand mechanisms that drive important outcome patterns, such as rates of obesity, mental illness, crime, and well- being. In order for these scores to be trustworthy, however, they need to represent the population, aggregate individual variance, be reliably different, agree across samples, and have a consistent pattern of correlations with sociodemographic measures."],["Intellectual humility has been identified as a character virtue that allows individuals to recognize their own potential fallibility when forming and revising attitudes. Intellectual humility is therefore essential for avoiding confirmation biases when reasoning about evidence and evaluating beliefs. The present study investigated the cognitive correlates of intellectual humility. The results indicate that cognitive flexibility, measured with objective behavioural assessments, predicted intellectual humility. Intelligence was also predictive of intellectual humility. These relationships were particularly pronounced for the facets of intellectual humility associated with respect for opposing opinions and openness to revising one's attitudes in light of new evidence. The data revealed an interaction: high cognitive flexibility is particularly valuable for intellectual humility in the context of low intelligence, and reciprocally, high intelligence was beneficial for intellectual humility in the context of low flexibility. Notably, there was evidence of a compensatory effect, as participants who scored highly on both flexibility and intelligence did not exhibit superior intellectual humility relative to individuals who scored highly on only one of these cognitive traits. These findings are suggestive of dual psychological pathways to intellectual humility; either cognitive flexibility or intelligence are sufficient for high intellectual humility, but neither is necessary. --------------------------------------------------------------------------------","In an era of polarization, fake news, and the wide spread of misinformation, there is a strong public need for an understanding of how citizens can inoculate themselves against deception and inaccurate information. The capacity to critically evaluate information in nonbiased ways requires intellectual humility – the understanding of one's limitations and biases when making evidence-based decisions. Intellectual humility allows us to avoid psychological tendencies to overlook evidence and confirm prior beliefs. Specifically, intellectual humility has been defined as “recognizing that a particular personal belief may be fallible, accompanied by an appropriate attentiveness to limitations in the evidentiary basis of that belief and to one's own limitations in obtaining and evaluating relevant information” (Leary et al., 2017). Over the last decade, a substantial literature has emerged in philosophy, theology, and psychology, seeking to (a) define intellectual humility (Baehr, 2011; Davis et al., 2016; Gregg, Mahadevan, & Sedikides, 2017; Roberts & Wood, 2003; Samuelson et al., 2015; Whitcomb, Battaly, Baehr, & Howard-Snyder, 2015; Wright et al., 2017), (b) develop measurement tools (Hoyle, Davisson, Diebels, & Leary, 2016; Krumrei-Mancuso & Rouse, 2016; Leary et al., 2017; McElroy et al., 2014; Meagher, Leman, Bias, Latendresse, & Rowatt, 2015), and (c) link intellectual humility to other personality traits such as openness (McElroy et al., 2014; Porter & Schumann, 2018; Leary et al., 2017), prosociality (Krumrei-Mancuso, 2017), dispositional attachment orientation (Jarvinen & Paulus, 2017), and religiosity and religious tolerance (Hopkin, Hoyle, & Toner, 2014; Hook et al., 2017; Krumrei-Mancuso, 2018; Leary et al., 2017; Rodriguez et al., 2017; Van Tongeren et al., 2016; Zhang et al., 2018). So far, research on the psychological roots of intellectual humility has been primarily the concern of social and developmental psychology. In theorising about the cognitive mechanisms that might underlie intellectual humility, Samuelson and Church (2015) proposed that the human tendency to rely on heuristics may lead to intellectually arrogant behaviours. Dual-systems accounts of human cognition suggest that thinking and reasoning are characterized by two distinct systems: System 1 processes, which are fast, automatic, associative, and intuitive, and System 2 processes, which are slow, conscious, deliberate, and analytical (Evans, 2003, 2008; Evans & Stanovich, 2013; Kahneman & Frederick, 2002). The corollary of this dual- systems approach is that in order to reason intelligently and avoid biased thinking, it is necessary to engage System 2 processes which are deliberate and analytical, and to override the automatic biases that are assumed to emerge from System 1 processes (Evans, 2003, 2008). Samuelson and Church (2015) therefore suggest that in order to facilitate intellectual humility, System 2 processes must be engaged and promoted. Interestingly, De keersmaecker and Roets (2017) found that cognitive ability shaped the extent to which individuals adjust their beliefs after learning that their attitudes were based on false information; people with lower levels of cognitive ability adjust their attitudes to a lesser extent than those with higher levels of cognitive ability. Intelligence may therefore be an important cognitive correlate of intellectual humility. Nevertheless, although deliberate, intelligent, analytical thinking may be important for intellectual humility, it might not be sufficient or necessary. For instance, one can persist in believing one's previous ideas and resist changing them in the face of new evidence even with slow and deliberative thinking. Intellectual humility and the capacity to revise one's ideas and be open to the ideas of others may require more than just analytical thinking or cognitive ability. Specifically, in order to be aware of one's cognitive limitations and evaluate evidence appropriately, considerable mental flexibility is required. While the intellectually arrogant or servile individual disregards new information in favour of past beliefs, the intellectually humble individual is able to be flexible in their thinking, overcome biased reasoning, find creative connections between past ideas and new information, and flexibly adjust their attitudes based on new evidence. The aim of this study was therefore to evaluate the hypothesis that the intellectually humble mind is also a flexible mind. The hypothesis that cognitive flexibility and openness to novel ideas may be crucial ingredients for intellectual humility has support in the empirical literature. Indeed, Leary et al. (2017) found that intellectual humility was positively correlated with self-reported openness to alternative ideas and values, and negatively correlated with dogmatism and intolerance of ambiguity. Stanovich and West (1997) found that participants who scored highly on a self-report measure called “Actively Open-minded Thinking”, which the researchers suggested is an indicator of cognitive flexibility and openness to belief change, were more likely to evaluate arguments based on the argument quality rather than relying on prior beliefs, even when controlling for cognitive ability. The study therefore suggests that a flexible thinking disposition may facilitate intellectual humility independently of cognitive ability. Interestingly, cognitive ability, operationalized with SAT scores and a test of verbal ability, was a unique and independent predictor of argument evaluation performance, signifying intelligence may still play a notable role. However, there are methodological problems with relying purely on self-report measures of cognitive flexibility. For instance, effect sizes may be inflated in self-report as compared to behavioural measures of cognition, and at times self-report measures yield opposite effects to theoretically-consistent behavioural assessments (e.g. Van Hiel, Onraet, Crowson, & Roets, 2016; De Keersmaecker et al., 2017; Saunders, Milyavskaya, Etz, Randles, & Inzlicht, 2018). Furthermore, new tools have been developed to accurately measure intellectual humility and its components directly (Krumrei-Mancuso & Rouse, 2016), and so there is a need to empirically investigate the ways in which flexibility of thought can shape intellectual humility. The present study sought to disentangle the relationships between cognitive flexibility, cognitive ability (fluid intelligence), and intellectual humility, using classic tasks from experimental psychology. Notably, cognitive flexibility and intelligence have been theoretically and empirically dissociated (e.g. Friedman et al., 2006; Salthouse, Fristoe, McGuthry, & Hambrick, 1998; Schaie, Dutta, & Willis, 1991), and so it is valuable to examine their relative contributions and interactions. This investigation thus addressed three primary hypotheses: Flexible thinking is positively correlated with intellectual humility (building on Stanovich and West's (1997) work). Cognitive ability is positively correlated with intellectual humility (as suggested by Samuelson & Church, 2015 and De keersmaecker & Roets, 2017). There is an interaction between flexibility and intelligence in shaping intellectual humility. If indeed intellectual humility is associated with high cognitive flexibility (in H1) and high intelligence (in H2), then two plausible, dissociative interaction mechanisms might be at play: H3-A There is an additive or multiplicative interaction, such that the highest intellectual humility would reflect high flexibility and high intelligence, while the lowest intellectual humility would be associated with low flexibility and low intelligence. This hypothesis would predict that individuals who score highly on flexibility, but not intelligence (and vice versa), would have lower intellectual humility than individuals who score highly on both. H3-B There is a compensatory interaction, such that either high flexibility or high intelligence are sufficient for high intellectual humility. Consequently, high flexibility would facilitate intellectual humility particularly for individuals with lower scores on the intelligence test, and vice versa. This hypothesis would predict that individuals who score highly on flexibility, but not intelligence (and vice versa), would have similar levels of intellectual humility as individuals who score highly on both. That is, there is no additive advantage for intellectual humility in scoring highly on both flexibility and intelligence. This would suggest that there are multiple independent psychological pathways to achieving high intellectual humility. The present study sought to investigate the cognitive correlates of intellectual humility and clarify these mechanisms in order to better understand the psychological underpinnings of intellectual humility and its various facets.","In accordance with the guidelines by Simmons, Nelson, and Simonsohn (2012), we report how we determined our sample size, all data exclusions (if any), all manipulations, and all measures in the study. Relevant data and code will be available on the Open Science Framework repository upon publication.","108 participants completed the study in full (see Supplementary Information SI1 for further details). Participants provided their informed consent to participate in the study in accordance with the institution's Department of Psychology Ethics Committee approval. Power analysis was conducted to compute the required sample size (see Supplementary Information SI1 for further details), with the ‘pwr’ package (Champely, 2015) in R (R Core Team, 2017). Intellectual humility – comprehensive intellectual humility scale (CIHS) The CIHS, a 22-item scale developed by Krumrei-Mancuso and Rouse (2016), was used to assess intellectual humility. The CIHS scale measures four distinct factors of intellectual humility: (1) independence of intellect and ego (Cronbach's α = 0.914; e.g. “When someone contradicts my most important beliefs, it feels like a personal attack”), (2) openness to revising one's viewpoint (Cronbach's α = 0.872; e.g. “I am open to revising my important beliefs in the face of new information”), (3) respect for others' viewpoints (Cronbach's α = 0.926; e.g. “I can respect others, even if I disagree with them in important ways”), and (4) lack of intellectual overconfidence (Cronbach's α = 0.822; e.g. “My ideas are usually better than other people's ideas”). The items are rated on a 5-point Likert scale ranging from 1 (strongly disagree) to 5 (strongly agree). Items were summed for the full scale (Cronbach's α = 0.664) and for each of the subscales (factors). Higher scores indicated greater intellectual humility. Cognitive flexibility – alternate uses task (AUT) In this computerized version of the AUT (Guilford, 1967), two common household items (brick and newspaper) were presented each for 1.5 min. Participants were asked to generate as many possible uses for these items. A timed clock was displayed to participants showing them how much time they had left. Flexibility was quantified as the total number of distinct conceptual categories in which the participant's responses belonged, in accordance with convention (e.g. Addis, Pan, Musicaro, & Schacter, 2016; Chermahini & Hommel, 2010; Madore, Addis, & Schacter, 2015). The responses were rated and calculated by two independent raters. Cognitive Flexibility - Verbal Fluency task (VF). In this computerized version of the semantic verbal fluency (Tombaugh, Kozak, & Rees, 1999; Troyer, Moscovitch, & Winocur, 1997), participants are asked to generate words from a given concept (i.e. ‘things on wheels’ or ‘red things’) for 2 min each. Flexibility was computed as the total number of distinct conceptual categories. The responses were rated and calculated by two independent raters. Intelligence - Raven's standard progressive matrices task (Raven's SPM) An abbreviated version of the Raven's SPM (Bilker et al., 2012; Raven, 1938) was used to assess fluid intelligence. The task was composed of nine visual patterns which progressively increased in difficulty. For each matrix pattern, one piece was missing, and participants are asked to select the correct pattern piece from a set of possible solutions.","All analyses were conducted in R (R Core Team, 2017) and SPSS (Version 25.0; IBM Corp., 2017), including the R packages visreg (Breheny & Burchett, 2017), jtools (Long, 2018), and pwr (Champely, 2015). First, we investigated whether the demographic variables of age, gender, and educational attainment, were related to the psychological variables of interest. Age was not significantly correlated with cognitive flexibility measured with the AUT (r = 0.06, p = .521), cognitive flexibility measured with the VF task (r = 0.011, p = .908), fluid intelligence measured with Raven's SPM (r = 0.037, p = .703), or with the comprehensive intellectual humility score (r = 0.056, p = .567). Furthermore, there were no gender differences in AUT cognitive flexibility, t(106) = 0.68, p = .501, VF cognitive flexibility, t(106) = −0.60, p = .552, or in intellectual humility, t(106) = −1.03, p = .306. There was a gender difference in Raven's SPM scores in the current sample, t(106) = 2.14, p = .035, in which males scored higher than females. Educational attainment was significantly correlated with AUT Flexibility (r = 0.23, p = .002) and Raven's SPM (r = 0.31, p = .001), nearly significantly correlated with VF Flexibility (r = 0.18, p = .060), and not correlated with intellectual humility (r = 0.06, p = .552). In all subsequent statistical analyses, age, gender, and educational attainment were included as covariates. H1: Is intellectual humility positively correlated with cognitive flexibility? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Correlational analysis revealed that cognitive flexibility measured with the AUT was significantly positively correlated with general intellectual humility (Fig. 1A). Furthermore, as evident in Fig. 2, decomposing the Comprehensive Intellectual Humility scale into its constituent factors revealed that this association was primarily driven by the correlations of cognitive flexibility with openness to revising one's viewpoint (Factor 2) and respect for others' viewpoints (Factor 3). Given Gignac and Szodorai's (2016) effect size guidelines for individual differences research, these effect sizes can be considered moderate to large. This pattern was corroborated by the correlations of intellectual humility and cognitive flexibility measured with the Verbal Fluency (VF) task. VF Flexibility was positively correlated with the comprehensive intellectual humility scale (r = 0.26, p = .007), and specifically with openness to revising one's viewpoint (Factor 2; r = 0.29, p = .002) and respect for others' viewpoints (Factor 3; r = 0.24, p = .014). There were no significant correlations between VF cognitive flexibility and independence of intellectual ego (Factor 1; r = 0.12, p = .207) or lack of intellectual overconfidence (r = 0.09, p = .339), paralleling the findings for AUT cognitive flexibility. H2: Is intellectual humility positively correlated with intelligence? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As depicted in Fig. 1B and Fig. 2B, intellectual humility was significantly positively correlated with fluid intelligence, such that more intellectually humble individuals tended to score more highly on Raven's SPM. Similarly to the pattern of results revealed for cognitive flexibility, intelligence was specifically positively correlated to the factors of intellectual humility representing openness to revising one's viewpoint (Factor 2) and respect for others' viewpoints (Factor 3; Fig. 2). The correlation effect sizes were generally smaller for the relationship between intellectual humility and intelligence than for intellectual humility and cognitive flexibility. H3: What is the relationship between cognitive flexibility and intelligence in shaping intellectual humility? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In order to investigate whether, and in what way, cognitive flexibility and intelligence interact to produce heightened intellectual humility, hierarchical linear regression analysis predicting general intellectual humility was conducted (Table 1). Note that all independent variables were centred prior to the hierarchical linear regression, as this helps reduce multicollinearity and facilitates testing of simple slopes (Dawson & Richter, 2006). In Step 1, the control variables, including age, gender, and educational attainment, were entered. As shown in Table 1, none of these control variables significantly predicted intellectual humility. In Step 2, the centred cognitive flexibility and centred fluid intelligence scores were entered. These independent variables explained a significant proportion of the variance in intellectual humility (R2 = 0.17). As evident in Table 1, the coefficients of both cognitive flexibility and intelligence were positive and significant, suggesting that both positively predicted heightened intellectual humility and each was a unique predictor. Next, in Step 3, we entered the interaction term for cognitive flexibility and intelligence. As predicted, the interaction of flexibility and intelligence was significant and accounted for an additional 5.5% of the variance in intellectual humility. Simple slope analyses were conducted to examine the relationship between intellectual humility and flexibility at 1 SD above and below mean intelligence, while controlling for age, gender, and educational attainment as covariates (see Fig. 3A). These analyses revealed that flexibility was positively related to intellectual humility in the context of low intelligence (at −1 SD, b = 6.24, SE = 1.51, p < .001) but not high intelligence (at +1 SD, b = −0.39, SE = 1.87, p = .834). Reciprocally, simple slope analyses demonstrated that when flexibility is conceptualized as the moderator, intellectual humility was positively related to intelligence in the context of low flexibility (at −1 SD, b = 2.43, SE = 0.72, p < .001), but not for high flexibility (at +1 SD, b = −0.51, SE = 0.78, p = .516). To validate this finding further, the sample was divided into three equal groups (terciles) rather than according to deviation from the mean. The simple slope analysis results were unchanged following this robustness check; cognitive flexibility was positively related to intellectual humility in the context of low intelligence (at −2.67 SD, b = 6.64, SE = 1.61, p < .001) but not average intelligence (at +0.5 SD, b = 2.23, SE = 1.25, p = .078) and high intelligence (at +2.71 SD, b = −0.85, SE = 2.01, p = .674). Reciprocally, intelligence was positively related to intellectual humility in the context of low flexibility (at −0.87 SD, b = 2.16, SE = 0.66, p = .001), but not for average (at +0.28 SD, b = 0.57, SE = 0.55, p = .300) or high flexibility (at +1.23 SD, b = −0.75, SE = 0.85, p = .381). This interaction effect is visualized in the filled contour plot and the corresponding 3D perspective plot in Fig. 4. This depicts that the relationship between intellectual humility and cognitive flexibility varies depending on intelligence, such that intelligence differentiates between low and high intellectual humility at low levels of cognitive flexibility, but not at high levels of cognitive flexibility. Similarly, cognitive flexibility differentiates between low and high intellectual humility at low intelligence scores, but not high intelligence scores. Moreover, Fig. 4 illustrates that the highest intellectual humility was evident in participants who scored highly on either intelligence or flexibility, and that scoring highly on both is not related to higher intellectual humility. Fig. 4 also highlights a slight bias toward higher intellectual humility scores amongst those with high cognitive flexibility (but low intelligence) relative to those with high intelligence (but low flexibility), which is also reflected in the higher regression coefficients in Table 1 for cognitive flexibility relative to intelligence. To probe the interaction further, we applied the Johnson-Neyman technique (Bauer & Curran, 2005; Hayes & Matthes, 2009; Johnson & Neyman, 1936), which calculates the range of z values of the moderator (in this case, intelligence) in which the predictor (i.e. cognitive flexibility) is a significant versus nonsignificant predictor of the outcome (i.e. intellectual humility). This helps to avoid limitations of traditional simple slopes analysis which require selection of potentially arbitrary values of the moderator at which the relationship between the predictor and outcome variable are assessed (e.g. ±1 SD from the mean). This technique is increasingly used in the psychological and cognitive sciences (e.g. Beach et al., 2012; Bushman, Giancola, Parrott, & Roth, 2012; Salerno & Peter-Hagene, 2013). Furthermore, Esarey and Sumner (2017) pointed out that probing interactions in the traditional way can lead to a multiple comparison problem. To address this, we implemented the method proposed by Esarey and Sumner (2017) to control for multiple comparisons; this leads to a more conservative test in which the false discovery rate in the marginal effects plot is controlled. The findings from the Johnson-Neyman analysis demonstrated that the relationship between intellectual humility and cognitive flexibility was significant when intelligence was less than 0.19 SD above the mean, but not significant with higher values of intelligence (Fig. 3B). This mirrors the finding from the simple slopes interaction analysis (Fig. 3A), in which the relationship between intellectual humility and cognitive flexibility is significant at low intelligence (−1 SD). In accordance with the methodological suggestions of McClelland, Irwin, Disatnik, and Sivan (2017), Spiller, Fitzsimons, Lynch Jr, and McClelland (2013), and Bauer and Curran (2005), this is revealed graphically in Fig. 3B. In Fig. 3B, the transition between significance and non-significance of the conditional effect is indicated by the dashed vertical line, which represents the Johnson-Neyman point at which the 95% confidence band intersects the x-axis. In accordance with Esarey and Sumner's (2017) recommendation, the Johnson-Neyman interval was calculated using the false discovery rate adjusted t = 2.21; note that when not adjusted for multiple comparisons using Esarey and Sumner's (2017) methodology, the Johnson-Neyman point is +0.35 SD. Although we conceptualized intelligence as the moderator of the relationship between intellectual humility and cognitive flexibility, it is important to note that the choice of moderator for analyses is arbitrary – cognitive flexibility could have equally been used as the moderator with paralleling results. We chose to use intelligence as the moderator here because it is largely considered a highly genetically heritable and stable construct, while there is more discussion over the stability and malleability of cognitive flexibility (e.g. Miyake & Friedman, 2012). Nonetheless, as evident in the filled contour plot of Fig. 4A, there is a symmetry in the interaction effect, such that the relationship between intellectual humility and intelligence is most pronounced at low levels of cognitive flexibility, and similarly the relationship between intellectual humility and flexibility is evident at low levels of intelligence.","Intellectual humility has been identified as a character virtue that enables individuals to recognize their own potential fallibility when forming and revising attitudes and beliefs. The present study examined the relationships between intellectual humility and objectively-assessed cognitive flexibility and fluid intelligence. With regards to our first hypothesis (H1), the results indicate that intellectual humility is positively related to heightened cognitive flexibility (Fig. 1A). Secondly, the findings reveal that intellectual humility is also positively correlated with intelligence (Fig. 1B), corroborating our second hypothesis (H2) and Samuelson and Church's (2013) suggestion that System 2 (i.e. analytical and deliberate) thinking styles are important for engaging in intellectually humble behaviour. These effects were driven by the facets of intellectual humility that correspond to openness to revising one's viewpoints and respect for others' viewpoints (Fig. 2). Thirdly, the data revealed an interaction between cognitive flexibility and intelligence in predicting intellectual humility (Table 1). Specifically, there was evidence of a facilitation effect, such that high cognitive flexibility is particularly valuable for intellectual humility in the context of low intelligence, and reciprocally, high intelligence was beneficial for intellectual humility in the context of low flexibility (Figs. 3 & 4). Interestingly, there was no evidence of an additive or multiplicative effect (contrary to hypothesis H3-A), as high flexibility and high intelligence did not produce superior intellectual humility relative to individuals who scored highly on only one of these cognitive traits (corroborating hypothesis H3-B; see Fig. 4). This is suggestive of dual psychological pathways to intellectual humility; either cognitive flexibility or intelligence is sufficient for high intellectual humility, but neither is necessary. The results demonstrate that cognitive flexibility was more strongly implicated in intellectual humility than intelligence, as manifest by the larger effect sizes (in Table 1, Figs. 1, 2, & 4). This may signify that the two pathways may have differential efficacy in producing intellectually humble attitudes and behaviours. Furthermore, the study revealed that not all facets of intellectual humility are equally shaped by cognitive flexibility and intelligence (Fig. 2). While epistemically-oriented features of intellectual humility, such as openness to alternative ideas (captured by Factor 2) and receptivity to attitude change (Factor 3), were positively correlated with both cognitive traits, the aspects of intellectual humility that are more closely associated with intellectual identity, such as the extent to which one feels threatened when contradicted (Factor 1) and one's conviction that one's own beliefs are superior and infallible (Factor 4), were unrelated to cognitive flexibility and intelligence. The specificity of these relationships suggests that future research will need to examine additional psychological and social factors that shape individuals' tendency to be intellectually overconfident. These findings extend research in three key disciplines: (1) cognitive psychology, (2) social psychology, and (3) interventionist and educational approaches. In the realm of cognitive psychology, recent research has provided corroborating evidence for a positive relationship between intelligence and intellectual humility across the lifespan. Danovitch, Fisher, Schroder, Hambrick, and Moser (2017) investigated biopsychological markers of intellectual humility in 6- to 8-year-old children. They found that greater intellectual humility was related to higher intelligence, and this relationship was specific to the epistemic aspect of intellectual humility (i.e. acknowledging the limitations of one's own knowledge) rather than its social component (i.e. representing one's knowledge to other people and being receptive to their ideas). This mirrors the specificity identified in the present study (Fig. 2). Similarly, developmental work by Mills and Elashi (2014) found that intelligence predicted 6- to 9-year-old children's ability to recognize that a source of information may be worthy of doubt and scepticism. Intelligence may therefore also be linked to early forms of intellectual humility. Additionally, research with adults has illustrated that intellectual humility and receptivity to attitude-change are related to cognitive ability (De keersmaecker & Roets, 2017) and higher discriminability in an old/new recognition memory task (Deffler, Leary, & Hoyle, 2016). Furthermore, Lick, Alter, and Freeman (2018) found that cognitive ability (as measured with Raven's Advanced Progressive Matrices) was related to enhanced updating of social stereotypes in light of new information, supporting the present finding that intelligence may be linked to a willingness to revise one's attitudes based on novel evidence. These results are also congruent with research in social and political psychology on the psychological correlates of behaviours that may be conceptualized as the opposite of intellectual humility – dogmatism, prejudice, and rigid adherence to ideological doctrines. Intellectual humility has been linked to lower dogmatism and belief superiority (Leary et al., 2017), fewer negative attitudes toward religious outgroups (Van Tongeren et al., 2016), and a willingness to be exposed to opposing political perspectives (Porter & Schumann, 2018). Frimer, Skitka, and Motyl (2017) have illustrated that liberals and conservatives are similarly motivated to avoid exposure to one another's opinions – a key facet of intellectual humility – suggesting that strong adherence to ideologies is related to a tendency to avoid hearing opposing views. Furthermore, recent empirical work has shown that cognitive ability is negatively related to right-wing ideological attitudes, authoritarianism, and prejudice (e.g. Brandt & Crawford, 2016; De Keersmaecker et al., 2017; Ludeke, Rasmussen, & DeYoung, 2017; Choma & Hanoch, 2017; for meta-analysis: Onraet et al., 2015), and that a cognitive style characterized by rigidity and intolerance of ambiguity is positively related to right-wing attitudes (for meta-analyses: Van Hiel et al., 2016; Jost, 2017). Moreover, a recent set of studies have demonstrated that behaviourally-assessed cognitive inflexibility is related to the extent to which individuals adhere firmly and rigidly to ideologies, in the realm of nationalism (Zmigrod, Rentfrow, & Robbins, 2018), politics (Zmigrod, Rentfrow, and Robbins, under review), and religion (Zmigrod, Rentfrow, Zmigrod, & Robbins, 2018). There is therefore converging evidence that intellectual humility and its opposing interpersonal correlate – rigid ideological thinking – are shaped by cognitive ability and cognitive flexibility. The finding that intellectual humility has multiple distinct psychological underpinnings – an analytical thinking route and a mental flexibility route – provides a fruitful basis on which to expand research into interventions that promote inoculation against misinformation and ideological polarization. Pre-emptively warning individuals about ideologically-motivated efforts to spread misinformation and about the argumentation techniques commonly used in misinformation campaigns has been shown to be effective in neutralizing the effect of misinformation on attitudes (Cook, Lewandowsky, & Ecker, 2017; Van der Linden, Leiserowitz, Rosenthal, & Maibach, 2017). The present findings are complementary to this line of research on inoculating citizens against fake news for several reasons. Firstly, identifying individual differences in cognition that shape individuals' willingness to revise their attitudes may suggest that individuals with certain psychological traits may be more receptive than others to inoculation interventions. Additionally, perhaps interventions that emphasize certain cognitive skills (analytical thinking, flexible thinking, etc.) may be more beneficial for individuals with particular psychological dispositions. Future research that combines the interventionist and individual differences perspective will be fruitful in refining our understanding of these processes. Secondly, these studies have focused on examining the effects of conveying information about expert consensus and potential misinformation campaigns in shaping citizens' attitudes (Van der Linden et al., 2017) and trying to engage individuals' System 2 analytical processing in evaluating evidence. The current findings suggest that fostering mental flexibility and an attentionally-open information processing style may also be a successful focal point for future interventions. Several potential limitations of the present study highlight future avenues for research. Firstly, it will be valuable to replicate these findings in lab settings and not just online samples, as well as in different cultural contexts, and with complementary measures of intelligence and cognitive flexibility. Since the a priori power analysis we conducted based on relevant effect sizes in the literature recommended a sample size of 92 participants, we computed the power actually achieved for the multiple regression models. This revealed that the power was 99.29% (f2 = 0.290), suggesting that the analyses were well powered to detect the present effects. Larger samples in future studies will help to corroborate and generalize these findings. In outlining future directions for the field, Leary et al. (2017) identified that “of particular interest are ways in which people who are high versus low in intellectual humility may differ in how they process information” (p. 810). The present study addressed this question by illustrating that analytical as well as flexible cognitive processing styles predict heightened intellectual humility. Admitting intellectual fallibility helps facilitate more constructive reactions to disagreements and conflict resolution (Porter & Schumann, 2018). Consequently, identifying and cultivating the cognitive factors shaping intellectual humility may be a key endeavour in building more evidence-based, tolerant, and effective discussions about the contested issues that divide and polarize our societies today.","All authors have declared that they have no conflict of interest."],["Tool innovation-designing and making novel tools to solve tasks-is extremely difficult for young children. To discover why this might be, we highlighted different aspects of tool making to children aged 4 to 6. years (N= 110). Older children successfully innovated the means to make a hook after seeing the pre-made target tool only if they had a chance to manipulate the materials during a warm-up. Older children who had not manipulated the materials and all younger children performed at floor. We conclude that children's difficulty is likely to be due to the ill-structured nature of tool innovation problems, in which components of a solution must be retrieved and coordinated. Older children struggled to bring to mind components of the solution but could coordinate them, whereas younger children could not coordinate components even when explicitly provided. © 2013 The Authors. --------------------------------------------------------------------------------","Tools are an essential part of human everyday life (Vaesen, 2012); it is hard to consider how we might get through the day without them. Tool-using capacity is evident from a young age, with children as young as 2 years using simple tools such as spoons (Connolly & Dalgleish, 1989) and rakes (Brown, 1990). Children gain the majority of their tool behaviors by observing others. As such, social learning has been the focus of research into the development of children’s tool use (Flynn & Whiten, 2008, 2010; Lyons, Young, & Keil, 2007; McGuigan & Whiten, 2009; Nielsen, 2006) and also their tool making (Beck, Apperly, Chappell, Guthrie, & Cutting, 2011). However, social learning cannot be a sufficient explanation for the development of all tool making because this would rule out the possibility of children (or anyone else) innovating novel tools (Nielsen, 2012). In contrast to findings when social learning is possible, recent findings suggest that innovation of a novel tool, by which we mean creating a novel tool to solve a problem, is extremely difficult for young children (Beck et al., 2011; Cutting, Apperly, & Beck, 2011). The focus of the current work was to determine what makes innovation so difficult. Our strategy was to highlight different components of the task solution to see whether this improved children’s performance. Children’s tool innovation difficulties have previously been demonstrated in a series of experiments requiring children to innovate a tool in order to retrieve stickers (Beck et al., 2011; Cutting et al., 2011; Chappell, Cutting, Apperly, & Beck, 2013). Children had great difficulty in generating the solution to bend a pipecleaner into a simple hook tool to retrieve a bucket from a narrow vertical tube. Children under 5 years of age rarely innovated a hook tool, and by 8 years of age only around half of children were successful on this task. This difficulty in tool innovation extends to making other tools using pipecleaners (Cutting et al., 2011) and to other materials and methods of tool making (Cutting, Beck, & Apperly, 2013). Children’s difficulty with tool innovation is surprising because children appear to possess all of the relevant knowledge required to solve tool innovation tasks. Children are familiar with the properties of the materials, for example, the pliant nature of pipecleaners. In previous studies, children received manipulation exercises in which they bent pipecleaners prior to being given the tool-making task (Beck et al., 2011, Experiment 3; Cutting et al., 2011, Experiment 1). Practice with bending pipecleaners did not aid children on subsequent tool-making tasks. This suggests that if children did lack knowledge about the properties of pipecleaners (or other materials), this is not sufficient to explain their difficulty. As well as seemingly understanding the properties of pipecleaners and the fact that they are allowed to manipulate them, children also appeared to have the required knowledge about the physics of the problem they faced. In the hook task, children appeared to understand that a hook would be the most functional tool; in a tool selection version of the task, children as young as 4 years chose the hooked tool over the straight tool first when their task was to retrieve a bucket from a vertical tube using pre-made tools (Beck et al., 2011, Experiment 1). Furthermore, children could also recognize a functional tool when shown how to make one: After initial failure on the hook innovation task, children readily manufactured a hook tool and used it correctly when shown a hook-making demonstration (Beck et al., 2011; Cutting et al., 2011). Note that children were only shown how to make the required tool; they were not given a demonstration as to how to use it. Taken together, this evidence suggests that it is not a simple lack of knowledge that limits children’s performance. Children understand the properties of the materials they are given and are aware that they are allowed to manipulate them. Children understand the physics of the task and can recognize a hook as the most functional tool. So, if children possess all of this knowledge, why do they find tool innovation so difficult? One possibility is that children’s difficulty with tool innovation could be due to its ill- structured nature. Although there is no single agreed-on definition of what constitutes an ill-structured problem, a generally agreed-on framework is that an ill-structured problem is one that is missing information from its start state, goal state, or information regarding the transformation required to go between the two (Goel & Grafman, 2000; Wood, 1983). Following this definition, tool innovation is an ill-structured problem; children are given the start state (the apparatus and the materials) and told that the goal is to retrieve the sticker, yet they are given no information regarding how they should go about this task. Compare this with Beck and colleagues’ (2011, Experiment 1) well-structured tool selection task in which young children readily succeed. In this task, children are given the start state (the apparatus and materials) and the goal state (retrieve the sticker) and are given the choice between two possible means for effecting a transformation (use the straight pipecleaner or use the hooked pipecleaner). When information about the start state, goal, and means were provided, children found it trivially easy to retrieve the bucket. Current findings suggest that just having all of the individual items of domain knowledge is not sufficient to be successful in solving ill-structured problems (Chen & Bradshaw, 2007). Domain knowledge must be well integrated into what is termed structural knowledge to enable people to use it effectively (Jonassen, Beissner, & Yacci, 1993). Structural knowledge is knowledge that is well integrated and developed and, as such, allows the person to use this knowledge in a flexible manner. This flexibility enables people to bring to mind the required pieces of knowledge and then successfully coordinate individual pieces of information into a useful solution. Some novices may possess all of the relevant pieces of information, but only in experts is this knowledge integrated into structural knowledge that is flexible enough to solve the problem (Voss, Blais, Means, & Greene, 1986; Wineburg, 1998). Applying this framework to tool innovation, it is possible that although children undertaking these problems appear to possess all of the knowledge required to solve the tasks, if this knowledge is not well integrated, they may still struggle to produce a solution. Children’s difficulty in these tool innovation studies may lie with bringing to mind the required pieces of information from memory, coordinating these different pieces of knowledge, or a combination of both. From previous studies, we know that highlighting the properties of the materials was not sufficient to elicit tool innovation. For example, 4- to 7-year-olds were not aided in making a tool when they were given bending practice that highlighted information about the properties of the pipecleaners (Beck et al., 2011, Experiment 3; Cutting et al., 2011, Experiment 1). We also know that just seeing the target tool that they were required to make, without any information regarding manipulation, was not sufficient to prompt children to make a tool for themselves (Cutting et al., 2013). This is particularly surprising given that children are able to see the utility of the end state tool and select it to use themselves in the context of a tool selection task (Beck et al., 2011, Experiment 1). In the current experiment, we investigated whether children were able to coordinate information and successfully make a tool if we highlighted the properties of the materials and the target tool required. By highlighting property information to half of the children before they attempted the task and then providing all children with a target tool demonstration after initial failure, we can begin to disentangle the minimum amount of information children require to successfully innovate a tool. Given previous findings, we expected children who had experienced bending practice to be no more successful in making a hook tool than children who had not received bending practice. Second, if children failed to innovate during this first stage, we then compared the two groups on their ability to make the tool following the target tool demonstration. Based on findings from Cutting and colleagues (2013), we expected children who had not received bending practice to perform poorly following the target tool demonstration. This would demonstrate children’s difficulty with bringing to mind additional information. Examination of performance following the target tool demonstration by the bending practice group would reveal whether children could successfully coordinate information. If the difficulty is in bringing information to mind, these children who had information about properties and information about hooks highlighted for them should be more likely to solve the task. However, if children’s difficulty is in coordinating information, even children who had the information highlighted for them should still have difficulties with the task. We tested children in the first (ages 4–5) and second (ages 5–6) years of compulsory education (UK) because these children performed near floor on previous tool innovation tasks and, thus, there was room for significant improvement.","The participants were 53 children aged 4 or 5 years (24 boys and 29 girls, mean age = 4 years 7 months [4;7], range = 4;1–5;1) and 57 children aged 5 or 6 years (26 boys and 31 girls, mean age = 5;7, range = 5;2–6;2) from two schools in the West Midlands, UK. Equal proportions of children from each school were present in each age group. The ethnic composition of the sample was 96% Caucasian, 3% Black, and 1% Asian. Participants had not taken part in previous versions of the task.","For the bending practice exercise, we used a pipecleaner (length = 29 cm), a pen, a piece of string (length = 29 cm), and a template of an S shape printed onto card. The apparatus for the main task was a clear plastic tube (length = 22 cm, width of opening = 4 cm) attached vertically to a cardboard base (length = 35 cm, width = 21 cm), a bucket containing a sticker, a pipecleaner (length = 29 cm), and a piece of string (length = 29 cm) that acted as a distracter item (see Fig. 1). The experimenter used an identical pipecleaner (length = 29 cm) for the demonstrations. Procedure Before testing, children were instructed by their class teacher not to tell other children how to play the games they would be playing with the experimenter to ensure that they would be a nice surprise for everyone. All participants were tested by a female experimenter in a quiet area just outside the main classroom. Children and the experimenter sat at right angles to each other at the corner of a table. Children were alternately allocated to either the bending practice group or the no bending practice group based on the teacher’s class list. Bending practice exercise Children in the bending practice group received the exercise prior to being given the main task. The exercise was designed to highlight the properties of the materials to the children and was based on the procedure from Cutting and colleagues (2011). Children watched as the experimenter demonstrated actions with the string and pipecleaner (order counterbalanced), and children then copied these actions. The pipecleaner was wound around a pen and then was removed to demonstrate that it kept its shape. The string was laid over the template to follow the S-shaped pattern. All children were able to perform the bending practice exercise. Main task Children were shown the vertical transparent tube with the bucket containing a sticker already in place in the bottom. They were told that if they could get the bucket out of the tube, they could win the sticker inside it. The experimenter then brought out the string and pipecleaner and told children that these were things that “can help” to get the bucket and sticker out. Children were then given 1 min to try to retrieve the sticker. No feedback was given, but children were given neutral prompts if required. Examples of prompts included “Can you think how you might be able to get the sticker out?” and “Maybe you could use these things to help you.” If, after 1 min, children had not retrieved the bucket, they were encouraged by the experimenter to put down the materials they were using. With the materials remaining in view in front of participants, the experimenter then said “Look at this” and brought out a ready-made pipecleaner hook for children to view (target tool demonstration). Children were again encouraged to retrieve the bucket using their own materials. If after 30 s children still had not retrieved the bucket, they were told to put down their materials. With their materials remaining in view as before, the experimenter said “Watch this” and, taking her own straight pipecleaner held in the middle, bent one end to form a hook (tool creation demonstration). The experimenter did not demonstrate how to retrieve the bucket with the hook because previous studies have shown that such demonstration is not necessary (Beck et al., 2011). Children were again encouraged to use their own materials to retrieve the bucket. If children were still not successful in making a hook tool, they were given verbal prompts such as “Did you see what I did with mine?” and then “Can you do that?” Thus, there were three stages to the main task: Stage 1 after half of the children had experienced the bending practice exercise, Stage 2 after all of the children had seen the target tool, and Stage 3 after children had seen the hook-making action demonstration. Stages 1 and 3 largely replicated previous studies, and so our main interest in the current study was performance in the two conditions at Stage 2. Children were coded as successful if they retrieved the bucket and sticker from the tube using a pipecleaner they had bent into a hook. Having made a hook, children did not require encouragement to use it.","There were no effects of gender on level of success pre-demonstration, χ2(1, N = 110) = 0.42, p = .518, φ = .062, or for success following the first target tool demonstration, χ2(1, N = 97) = 0.64, p = .425, φ = .081, or the second action demonstration, χ2(1, N = 54) = 0.05, p = .821, φ = .031. As such, data were combined across gender for subsequent analyses. The results were first analyzed for all children combined and then for the two age groups separately. Overall, 84 of 110 children were successful in making a hook tool at any of the three stages of the task. Children’s success at innovating a hook during Stage 1 was examined to see whether the bending practice facilitated performance. Overall, children were very poor during their first exposure to the task, with only 13 of 110 children successfully making a hook tool. Of these children, 7 were in the bending practice group and 6 had not received bending practice, demonstrating no effect of condition, χ2(1, N = 110) = 0.05, p = .822, φ = .022. When we break this down into separate age groups, only 2 4- and 5-year-olds were successful, both of whom had not received bending practice, showing no difference between conditions, Fisher’s exact test, p = .236. For the 5- and 6-year-olds, 7 of the successful children were in the bending practice group and 4 were in the no bending practice group, again showing no difference between conditions, χ2(1, N = 57) = 0.89, p = .346, φ = .125. Children who were successful on their first exposure to the task were excluded from subsequent analyses that compared success following the demonstrations. Chi-square analyses were used to compare children’s performance at Stage 2 following the target tool demonstration. For both age groups combined, children were significantly more likely to make a hook tool following the target tool demonstration in the bending practice condition than in the no bending practice condition, χ2(1, N = 97) = 6.59, p = .010, φ = .261. Comparison across age groups shows that older children were significantly more successful than younger children in the bending practice condition, χ2(1, N = 49) = 9.93, p = .002, φ = .450. No difference in success was seen in the no bending practice condition, χ2(1, N = 48) = 0.87, p = .350, φ = .135. When the two age groups were analyzed separately, the difference in success between conditions was found to be driven by the older children, χ2(1, N = 46) = 9.30, p = .002, φ = .450 (see Table 1). This suggests that 5- and 6-year-olds are able to coordinate the information if they received both the bending practice and saw a pipecleaner hook. No such difference was seen for the 4- and 5-year-olds, χ2(1, N = 51) = 0.86, p = .355, φ = .129. Children who were successful following the target tool demonstration were excluded from the following analyses that investigated success at Stage 3, which followed the tool creation demonstration. For children requiring this demonstration, 55% were successful at making the tool needed (see Table 1). Chi-square analysis revealed no difference in the levels of success for each group following the action demonstration for either the 4- and 5-year-olds, χ2(1, N = 35) = 0.02, p = .877, φ = .026, or the 5- and 6-year-olds, Fisher’s exact test, p > .999.","In the current work, we highlighted various aspects of the task solution in order to discover why children have difficulty with tool innovation. Information regarding the properties of the materials and an example of the tool children needed to create were highlighted. The current findings suggest a series of limiting steps in innovation, with children getting stuck at different steps at different ages. Overall, we found that very few 4- to 6-year-olds spontaneously innovated a hook tool with either no additional information or just information about pipecleaner properties highlighted. These results are in line with previous research demonstrating that young children have great difficulty in innovating tools with either no additional information (Beck et al., 2011, Experiment 2) or information highlighting pipecleaner properties (Beck et al., 2011, Experiment 3; Cutting et al., 2011, Experiment 1). It should be noted that success on this task has been shown to improve with age, with children becoming extremely proficient by 9 or 10 years (Beck et al., 2011, Experiment 1). The main aim of the current study was to test children’s ability to make a tool following a target tool demonstration. Children were shown a ready-made pipecleaner hook but were not shown how to make it. This enabled us to discover whether children could bring to mind the means to make the hook for themselves. In comparison with children who had experience of pipecleaner properties, children in both age groups were extremely poor at making the hook tool following the target tool demonstration if they had not had information regarding pipecleaner properties highlighted for them, that is, children who had not received the bending practice. The 5- and 6-year- olds who had information regarding pipecleaner properties highlighted were significantly more successful in making the required hook tool following the target tool demonstration than children who had not received the bending practice. This suggests that if both pieces of information were readily accessible to older children, they were able to coordinate the information successfully into a solution. Conversely, the 4- and 5-year-olds displayed great difficulty in making a hook tool even if they had both pieces of information highlighted for them. This suggests that younger children face a limitation in the domain of tool making in that they are unable to coordinate information even when it is highlighted. The current findings suggest that children’s main difficulty with tool innovation could be due to problems with retrieving information and recognizing it as a useful solution to the problem. Children were unable to bring to mind additional information when given certain aspects of the task. For example, children who received the bending practice that highlighted the pliable property of pipecleaners were unable to bring to mind information about hooks needed to allow them to innovate the task solution. Similarly, following the target tool demonstration, children who did not receive the bending practice were unable to bring to mind information regarding the properties of pipecleaners that would enable them to successfully make their straight pipecleaner into a hook tool. Findings from both age groups fit with the suggestion that tool innovation is an ill-structured problem that requires solvers to both retrieve and coordinate knowledge in order to solve a task. The current study suggests that 4- and 5-year-olds had difficulty with both of these components. Performance improved with age, with 5- and 6-year-olds being able to coordinate information into a useful solution if it had been highlighted, but these older children still displayed great difficulty with bringing to mind this information for themselves. Regardless of how good children’s ability to coordinate knowledge is, they can never succeed in solving the task if they are unable to retrieve the components of knowledge required and recognize their relevance to the solution. As such, we still see poor innovation ability in this older age group under conditions where the required information is not highlighted for them and they must bring it to mind themselves (Beck et al., 2011; Cutting et al., 2011). It is surprising that both age groups had difficulty in bringing to mind the required knowledge needed to innovate the solution because previous evidence suggests that children possess all of the individual pieces of knowledge required to solve this tool innovation task. First, children recognized that a hook was a solution to the task. This is demonstrated in Beck and colleagues’ (2011, Experiment 1) tool selection task and by children readily manufacturing a hook tool and using it correctly when shown a hook-making demonstration (Beck et al., 2011; Cutting et al., 2011). Second, children have knowledge about the properties of pipecleaners (Beck et al., 2011, Experiment 3; Cutting et al., 2011, Experiment 1). So, if children possess all of this information, why can they not retrieve it in the context of a tool innovation task? Having domain knowledge might not be sufficient to solve ill-structured problems (Jonassen et al., 1993). The children in the current study were novices. Although these children may have possessed all of the independent pieces of knowledge the task required, they did not have sufficient experience with the world and the materials to have integrated structural knowledge. We suggest that without this structural knowledge, young children lacked the flexibility needed to retrieve their knowledge from memory and then coordinate it in order to solve these tool innovation tasks. The current study required children to retrieve and coordinate knowledge regarding the transformation they were required to perform. It seems likely that other types of ill-structured problems—that is, those missing information from either their start or goal state—would also require the solver to retrieve and coordinate the relevant pieces of information. Future research is needed to test this. The current findings suggest that the main difficulty for both age groups was retrieving knowledge from memory. Younger children in this study also displayed great difficulty with coordinating their knowledge. As children develop and integrate their knowledge, they first improve in their capacity to coordinate information and can do so readily if all of the information needed is highlighted for them. We suggest that as children develop further, their knowledge will become more integrated. This will allow them to access and retrieve their knowledge more flexibly and, along with their ability to coordinate knowledge, will enable them to solve these ill-structured tool innovation tasks."],["If beliefs and desires affect perception—at least in certain specified ways—then cognitive penetration occurs. Whether it occurs is a matter of controversy. Recently, some proponents of the predictive coding account of perception have claimed that the account entails that cognitive penetrations occurs. I argue that the relationship between the predictive coding account and cognitive penetration is dependent on both the specific form of the predictive coding account and the specific form of cognitive penetration. In so doing, I spell out different forms of each and the relationship that holds between them. Thus, mere acceptance of the predictive coding approach to perception does not determine whether one should think that cognitive penetration exists. Moreover, given that there are such different conceptions of both predictive coding and cognitive penetration, researchers should cease talking of either without making clear which form they refer to, if they aspire to make true generalisations. --------------------------------------------------------------------------------","Some advocates of the predictive coding account of perception have said that if the predictive coding account of perception is correct then cognitive penetration occurs. If true, this would provide a new route to arguing for the existence of cognitive penetration: argue for a predictive coding account of perception. In this paper, I investigate whether the predictive coding account of perception does entail that cognitive penetration occurs. I establish that there are different versions of cognitive penetration and different versions of predictive coding. I map out the relationship between these different versions. I show that some versions of predictive coding entail that there are some forms of cognitive penetration, some versions of predictive coding are compatible with, but do not entail, that there are some forms of cognitive penetration, and some forms of predictive coding entail that some forms of cognitive penetration do not exist. Mapping out these relationships helps us to understand the predictive coding account of perception and cognitive penetration in depth. It also serves to warn us that we should be clearer which version of the predictive coding account of perception and which version of cognitive penetration we are referring to when we make claims about them. And it serves to show that an argument for cognitive penetration that appeals to the existence of predictive coding has to appeal to a specific version of predictive coding, the establishment of which may turn on just the same issues as the establishment of cognitive penetration itself. This paper is divided into seven further sections. Section 2 defines and examines the different forms of cognitive penetration. Section 3 outlines a minimal account of the predictive coding account of perception. In Section 4, I consider possible relationships between predictive coding and cognitive penetration and examine the claims that people have made about this relationship. In Section 5, I look at whether predictive coding is committed to cognitive states playing a role in the predictive coding account, and what this means for the relationship between various versions of predictive coding and cognitive penetration. In Section 6, I examine whether a predictive coding account that is committed to the idea that cognitive states play a role in predictive coding entails one version of cognitive penetration, namely the cognitive penetration of perceptual experience—that is a conscious, personal-level state. In Section 7, I consider whether a predictive coding account that is committed to the idea that cognitive states play a role in predictive coding entails another version of cognitive penetration, namely the cognitive penetration of early vision—that is, an initial level of visual processing that has been defined in different ways by psychologists and neuroscientists. I give three accounts of what early vision might be taken to be and I spell out what predictive coders should say about the penetration of early vision conceived of in each of them. Finally, in Section 8, I summarise what the various accounts of predictive coding and cognitive penetration are, and the relationships between them, in a table, and I comment on the broad picture that emerges of the relationships between them. Readers may wish to consult this table while they are reading the sections below.","There are two related, but nonetheless distinct, alleged phenomena that go under the name ‘cognitive penetration’. That they are clearly distinct emerges in the rest of this paper by consideration of their relation to predictive coding. The first of these phenomena is the penetration of early vision. It is the phenomenon discussed in Pylyshyn’s (1999) much cited paper on the topic. Early vision is defined functionally by Pylyshyn as the system that takes attentionally modulated signals from the eyes (and perhaps some information from other sensory modalities) as inputs, and produces shape, size and colour representations as output. These representations are then categorised and identified by the cognitive system making use of memory, knowledge and judgment. The question of whether cognitive penetration occurs, conceived of as the penetration of early vision, is the question of whether the function that the early visual system computes “is sensitive in a semantically coherent way, to the organisms goals and beliefs, that is, [whether] it can be altered in a way that bears some logical relation to what the person knows” (1999: 343). Pylyshyn explicitly denies that the output of early vision either is, or determines, one’s visual experience, or equivalently, that the representational content of the output of early vision is the representational content of one’s perceptual experience (1999: 362). In contrast to this notion of cognitive penetration, which has mostly been the subject of study by psychologists and neuroscientists, another notion of cognitive penetration, that of the penetration of perceptual experience, has been a major subject of study by philosophers, although the division in what researchers study is far from exceptionless, and increasingly so. ‘Perceptual experience’ refers to the conscious state that we typically go into when we perceive the world—the conscious state of awareness of the world that has a distinctive phenomenal character compared to typical beliefs or judgments about the world. In my (2012) paper on cognitive penetration, I defined the cognitive penetration of experience as follows: in any case of perception, hold fixed what it is that is perceived (the objects properties and relations seen, heard, touched, and so on), the perceiving conditions (the level of light, shadow, mistiness, for example), the state of the sensory organ (perfect human vision, shortsighted human vision, for example) and the location of one’s focus of attention. With those conditions fixed, if it is possible for two subjects (or one subject at different times) to have different perceptual experiences due to the differing content of the states of their cognitive systems, and moreover, there is a semantic or intelligible link between the content of the cognitive states and the content of the perceptual experience, then perceptual experiences are cognitively penetrable. States of the cognitive system include beliefs, judgments and desires, and should likely also be taken to include the concepts that we possess. For further discussion of the terms used in this definition see my (2012) paper. There are several issues with both of these definitions of cognitive penetration. I address the two most pressing ones. The first concerns attention. The second concerns the semantic or intelligible link that is posited as necessary for cognitive penetration to occur. Despite giving the above definition of cognitive penetration in my (2012) paper, later in the paper, I go on to question whether we should include attention in the conditions that we hold fixed. I noted that one rationale that someone might have for thinking that we should hold fixed attention is that a shift of attention is akin to a shift of the location to which one’s eyes point. A shift in the direction that one’s eyes look changes one’s perceptual experience by changing which objects and properties are processed, but if such a shift were driven by one’s cognitive states, such a shift should not count as cognitive penetration occurring, simply a shift in what is perceived. One might think of shifts of attention in a similar manner: a shift of spatial attention, driven by one’s cognitive states, might affect how clearly, or in how much detail, certain objects and properties are experienced, but that is just a shift in what is perceived and shouldn’t count as cognitive penetration. However, as I noted in my (2012) paper, there are forms of attention other than spatial attention, so even if one was persuaded that changes of spatial attention driven by cognition should not count as cognitive penetration, it is not clear that shifts in other forms of attention should also not count. One form of a non- spatial shift in attention is a shift in what properties one attends to. For example, one can shift from attending to the red things in one’s visual field to attending to the yellow things, and that might change one’s experience in certain ways. For example, it may lead one to have an experience in which the salience of the yellow things is greater than that of the red, and so one might be aware of the number yellow things but not be aware the number of red things. If such a shift were driven by cognition in the right way, then it is tempting to consider such a change in experience a case of cognitive penetration because a kind of bias could creep into perception. For example, one might overestimate the proportion of yellow things to red things. A useful discussion of what factors, such as bias, should lead one to classify certain cases as cases of cognitive penetration can be found in Stokes (2015). However, it is even unclear whether shifts in spatial attention caused by cognitive processes should be conceived of as failing to count as cases of cognitive perception. An argument for thinking that some such cases are instances of cognitive penetration has recently been put forward by Cecchi (2014), who discusses an experiment carried out by Schwartz, Maquet, and Frith (2002). In the experiment, subjects underwent a period of monocular training (subjects in one group used their right eye and subjects in another group used their left eye). They looked at a screen on which was displayed a homogenous background of horizontal bars. An “L” or “T” would appear in the centre of the screen randomly rotated for 16 ms. During the same short interval, three lines would appear randomly at some location within the upper-left quadrant in the periphery of their visual field. Subjects had to do two tasks: identify whether the object in the centre was an “L” or a “T”, and identify the orientation of the three lines that appeared in the periphery of their visual field. Subjects’ eye movements were restricted to fixation on the central letter. Detection of the orientation of the lines in the periphery of their visual field therefore required the allocation of voluntary covert spatial attention. The goal of the experiment was to improve subjects’ performance in the detection of the orientation of the three peripheral lines. In the first (training) phase of the experiment, subjects underwent a total of 1760 trials each, and received feedback about their accuracy after each one. During the first phase, performance of the trained and untrained eye was measured. Successful performance was one in which they accurately detected what the central letter was and what the orientation of the peripheral lines was. During the first phase, subjects improved their performance of detecting the orientation of the peripheral lines gradually in the eye that was being trained, but not in the one that was not. Progressive changes in neural behaviour elicited by the trained eye in the part of the V1 cortex (also known as the striate cortex) associated with seeing in the upper-left quadrant were detected using functional magnetic resonance imaging (fMRI). This was compared to activity in that area elicited by the untrained eye and activity in that area elicited by the same eye before training. According to Cecchi (2014: 85), the voluntary covert attention required to detect the orientation of the lines required guidance by the cognitive states of the subjects to select orientation as the feature to be processed, and top-down signals to the cortex were observed (using fMRI), and discovered to be necessary, for the task to be successfully carried out during the first phase. The subjects were investigated again 24 hours later (during the “test” session), after they had slept the night, and this time, again, the part of the V1 cortex corresponding to seeing in the upper-left quadrant was observed to have increased activity when activated by the trained eye. This is in contrast to when it was activated during the same test session by the untrained eye, or the same eye before training. In both cases, no increased activity was observed. However, this time, fMRI indicated that the increased activity observed in the test session when using the trained eye did not require that top- down cognitive or attentional influences occur. The increased responses occurred before any top-down influences were observed. The same task performed by the untrained eye did recruit higher-level cognitive and attentional influences.1 According to Schwartz (2002: 17140), training resulted in neural changes to the V1 area of the cortex, which had learned to detect the orientation of the peripheral lines when seen by the trained eye. during the training session, the constant and synchronic stimulation of striate visual areas by cognitively guided attention resulted in physiological neural changes in the visual system. Therefore, structural modulations occurring in the visual system due to cognitively guided attention are the result of synchronic architectural cognitive penetration.2 The successful detection of the peripheral object during the test session was possible thanks to diachronic architectural cognitive penetration. The consolidation of cognitively induced architectural modulations enabled the visual system to perform the peripheral-detection task without synchronic cognitive intervention. At that point, the cognitive influences that modified the architecture of the visual system were diachronically produced. Thus Cecchi argues that there is cognitive penetration brought about by means of attention. Cecchi argues that the penetration in question is both that of early vision and that of perceptual experience. With regard to the first, changes to the visual cortex occur in early vision, defined temporarily, as activity within the first 100 ms after stimulus presentation, because they occur around 60 ms after stimulus presentation. With regard to the second, there are changes to the content of visual experience, with experience coming to represent the orientation of the lines presented in the periphery of the visual field. In this way, he argues that there can be cases of cognitive penetration of the early visual system and of experience that involve attention. Cecchi certainly presents us with a tough case to adjudicate. An alternative interpretation of this case is that, in the first training phase of the experiment, the attention involved is simply selecting what to process—items in the upper-left quadrant and their property of orientation—and that this mere selection does not amount to cognitive penetration. As mentioned above, one might claim that this is just like selecting what to process as one might do by changing the direction in which one’s eyes point—something that does not count as cognitive penetration. Moreover, one might claim that in the second test phase of the experiment, subjects had simply gained better eyesight, as one might gain from wearing glasses, which one has put on under the guidance of cognition (such as one’s desire to see more clearly and one’s belief that one’s glasses will facilitate that), or as one might gain by voluntarily squinting one’s eyes to see an object in the distance more clearly (caused by one’s desire to see it and one’s belief that squinting will accomplish that). Clearly such activities do not count as cognitive penetration. It is not obvious to me how to adjudicate this case, but it is clear that the question of whether cognitive penetration—of either early vision or of perceptual experience—can involve attention is not a straightforward matter. The second issue with the definition of cognitive penetration that I wish to discuss concerns why it specifies that, in order for cognitive penetration to occur, a semantic or intelligible link is necessary between the content of the cognitive state doing the penetrating and the content of the perceptual experience, or the content of early vision, that is penetrated. Is such a link necessary? For ease of exposition, I will discuss this case with respect to the cognitive penetration of experience, but readers will be able to reconstruct with ease a parallel discussion that could be made about the penetration of the content of early vision. A semantic or intelligible link is posited as a necessary feature of cases of cognitive penetration because the lack of such a link seems to explain our intuitions about why some cases fail to be those of cognitive penetration.3 In other words, there are some cases which, but for the lack of this link, would fall under the definition of cognitive penetration, but which we are rightly disinclined to think are cases of cognitive penetration. Here is one such case:In Migraine, your belief affects your visual experience but, intuitively, this is not a case of cognitive penetration. It is merely an instance of a migraine brought on by belief-triggered stress. This is because there is no semantic or intelligible link between the belief that you have an exam tomorrow and the subsequent visual experience that you have as of flashing lights. If we hold that such a link is necessary for a case of cognitive penetration, its absence explains why the case is not one of cognitive penetration. However, one might question whether the lack of such a link really does explain why such cases should not count as cases of cognitive penetration. For one can gerrymander the above type of case to yield one that is very similar, in which there does seem to be a semantic or intelligible link, and yet the case still does not appear to be one of cognitive penetration. For example, consider the following case:In Migraine’s Revenge, as in Migraine, your belief affects your subsequent visual experience; yet, intuitively, the case is not one of cognitive penetration, merely a case of migraine. However, in Migraine’s Revenge, unlike in Migraine, there is a semantic or intelligible link between your belief that aliens might land on earth and your subsequent perceptual experience of flashing lights in the sky. (However improbable, it is entrenched in popular culture that aliens coming to Earth arrive in spaceships that have flashing lights.) So, it would seem that, in Migraine’s Revenge, it is not the lack of a semantic link that explains why it is not a case of cognitive penetration. Given this, as well as its similarity to Migraine, one might conclude that the lack of such a link also does not explain why Migraine fails to be a case of cognitive penetration. If the lack of a semantic link does not rule out both Migraine and Migraine’s Revenge as cases of cognitive penetration, what other condition would? I believe that we can find a suitable condition, not by abandoning a semantic or intelligible link as a necessary condition for cognitive penetration, but by adding to, and thereby strengthening, the condition. I believe that what goes wrong in Migraine’s Revenge is that the semantic or intelligible link is accidental. After all, I deliberately gerrymandered Migraine to yield a case in which there just so happened to be such a link. I believe that what is required in cases of cognitive penetration is that there be a causal, semantic link between each of the steps in the chain that lead from the belief to the subsequent perceptual experience.4 This does not exist in Migraine’s Revenge. That case is one in which one’s belief that aliens might land causes one to be stressed, which causes a migraine, which in turn causes one’s perceptual experience of flashing lights. While there is a semantic link between the content of the belief and the content of the perceptual experience, the link is not a causal link. Moreover, there is not a causal, semantic link between the content of the belief and the subsequent state of stress. There is a causal link but not a semantic one. A state of stress is a physiological state that does not have content. Even if one thought that states of stress could have content (perhaps content about the source of the stress or the content that something bad was going to happen), the state of having a migraine is clearly not such a state. So there is no causal, semantic link between the state of one being stressed and the state of having a migraine. Moreover, there is no causal, semantic link between the state of having a migraine and the state of experiencing flashing lights in the sky. Note that I proposed (2012) a mechanism by which cognitive penetration could occur. It was one in which a belief caused some imaginative state or imaginative processing to come into existence, which in turn causally interacts with one’s experience yielding an experience that has content influenced both by perception and imagination. This mechanism could be a mechanism whereby cognitive penetration occurs because it allows for the existence of causal, semantic links at each stage of the process. I take it to be a virtue of the proposed mechanism that it could allow for a causal, semantic link between each of the steps, which I take to be necessary for cognitive penetration. In this section, I have looked at what cognitive penetration is and how it should be defined. I noted that there are two distinct phenomena that go under its name: penetration of the early visual system and penetration of perceptual experience. I explained that there are hard questions to answer about whether there can be a role for attention in cases of cognitive penetration, and I explained why I believe that a necessary condition for cognitive penetration is that there be a causal, semantic link between each step in the process whereby a belief comes to affect perceptual experience. In the next section, I consider what predictive coding accounts of perception are by outlining the minimal commitments of such an account.","The minimal commitments of the predictive coding account of perception are as follows: Based on knowledge, assumptions, and expectations at different high-level processing stages, the brain produces top-down generative representations of the world that are predictions of how the world is. Those representations are then modified bottom-up by incoming signals, which are propagated upwards only as a prediction error signal. Those representations are then further modified by top-down high-level processing, and perhaps sideways by other sensory modalities. The aim of these modifications are to reduce global prediction error. The processing is typically thought to be done according to Bayesian rules.5 I will refer to (i)–(v) as the ‘minimal form of predictive coding’. As we will come to see in later sections, there are many, more specific, forms of predictive coding accounts of perception. We will see the ramifications of these further forms for the relationship between predictive coding and cognitive penetration.","There are three possible relationships between predictive coding accounts of perception and cognitive penetration: predictive coding entails cognitive penetration; predictive coding is compatible with, but does not entail, cognitive penetration; predictive coding is incompatible with cognitive penetration. In each case, ‘cognitive penetration’ might mean either penetration of the early visual system or penetration of perceptual experience. One motivation for exploring which of (a), (b) or (c) is true (or, as we will come to see below, whether any one of these positions is true) is because they have been endorsed by different people. Arguably Lupyan (2015a, 2015b) endorses (a). He certainly comes very close to explicitly endorsing it. For example, he states: penetrability should be expected whenever constraining lower-level processes by higher level knowledge minimizes global prediction error, My goal, however is not to simply argue that perception can be cognitively penetrated, but to present a theoretical framework on which cognitive penetrability is expected as a natural part of how perception works.” The primary goal of this paper is to show that the controversy surrounding cognitive penetrability of perception (CPP)—the idea that perceptual processes are influenced by “non-perceptual” states—vanishes when we view perception not as a passive process the goal of which is recreation of a veridical reality, but rather as a flexible and task-dependent process of creating representations for guiding behavior. My goal is to … argue that there is no in- principle limit on the extent to which a given perceptual process can be penetrated by knowledge, expectations, beliefs, etc. Thus Lupyan thinks that if the predictive coding model of perception is correct then we should expect there to be cognitive penetration, and that the model entails that mechanisms exist that allow it to happen. Clark (2013), prima facie at least (but I will explore other interpretations of him below), believes that predictive coding entails cognitive penetration too. He says that predictive coding: makes the lines between perception and cognition fuzzy … In place of any real distinction between perception and belief we now get variable differences in the mixture of top-down and bottom-up influence … To perceive the world just is to use what you know to explain away the sensory signal across multiple spatial and temporal scales. The process of perception is thus inseparable from rational (broadly Bayesian) processes of belief fixation, and context (top-down) effects are felt at every intermediate level of processing. In my (2015) paper, I argued for a version of (b): that the minimal form of predictive coding is consistent with, but does not entail, the cognitive penetration of perceptual experience.6 Drayson (unpublished manuscript) claims this too. With respect to the penetration of early vision, no one has yet commented. I suspect that is because this issue is particularly thorny on account of what predictive coders can say about what early vision is—a topic that I address in Section 7. I know of no one who explicitly endorses (c), however, I think that one reading of Clark (2013) would have him believe (c) with respect to the penetration of perceptual experience. I will explore this in Section 5. What to make of (c) with respect to early vision, is a question that requires detailed consideration that, as previously stated, I will provide in Section 7. The most expeditious way to determine whether (a), (b), or (c) is true is not to address them each in in order but to consider, in turn, the following three questions: Do cognitive states feature in the predictive coding account of perception? If cognitive states do feature, is there reason to think that they penetrate perceptual experience? If cognitive states do feature, is there reason to think that they penetrate early vision? I do so, respectively, in Sections 5,6, and 7.","Should the high-level processes posited by predictive coders be taken to include cognitive states of subjects? By ‘cognitive states’ I mean doxastic states—states that are accessible to consciousness and inferentially integrated—such as beliefs and desires. Or should they be taken to be sub-doxastic information-carrying states of the brain that the subject does not in principle have access to?7 The minimal account of predictive coding, outlined above, simply states that perceptual representations of the world are produced by high-level processes—high-level in relation to the level of perceptual experiences. The account leaves it open whether those high-level processes include cognitive states. This shows that the minimal form of the predictive coding account does not entail cognitive penetration of either the early visual system or experience. Does it show that predictive coding is consistent with cognitive penetration sometimes or always occurring? The answer will depend on whether we hold that cognitive states sometimes or always affect early vision or perceptual experience—an issue that I will address shortly. One could simply stipulate two different versions of the predictive coding account each of which would be a precisification of the minimal account. According to a first, the high-level states that make the predictions sometimes or always involve cognitive states; according to a second, the high-level states in question never involve cognitive states. Whether the first version entails or is consistent with cognitive penetration will depend on whether we think cognitive states sometimes or always affect early vision or perceptual experience—as was the case with the minimal account. The second account would entail that cognitive penetration of either early vision or perceptual experience never occurs, for cognitive states are not involved in the production of perceptual representations. Which version of the theory do predictive coders actually put forward? Predictive coders often talk about knowledge, beliefs, information, expectations, priors, and so on feeding into the production of perceptual representations from high levels. But should we take some of this talk as literally talk of cognitive states? Lupyan (2015a) is very clear that he thinks we should: there is no in-principle limit on the extent to which a given perceptual process can be penetrated by knowledge, expectations, beliefs, etc. The actual extent to which such penetrability happens can be understood in terms of whether it helps to lower system- wide (global) prediction error. In evolving to minimize prediction error neural systems naturally end up incorporating whatever sources of knowledge, at whatever level, to lower global prediction error. And, in addition, it is clear that the specific examples of cognitive penetration that Lupyan (2015a) discusses are ones in which semantic doxastic states of knowledge and belief, as well as other cognitive states, affect perception. This view would entail that there is cognitive penetration if and only if those cognitive states can be seen to affect early vision or perceptual experience (as Lupyan argues that they do). Hohwy (2013) argues for the same position. In Sections 6 and 7, I will investigate this later claim for predictive coding accounts of perception in which cognitive states play a role. For now, however, in the rest of this section, I will explore whether other philosophers accept that cognitive states play a role in the predictive coding account of perception, and what follows from that about cognitive penetration. Consider Clark’s (2013) view. Although I noted above in Section 4 that, prima facie, he holds that cognitive states penetrate perceptual states, there is quite a bit of room to doubt this. He often talks about sub-personal states of the brain having a top- down influence and, when he writes about “belief” having an influence, he often puts the word in quotation marks, signalling an unorthodox usage. Moreover, he never explicitly states that the highest level corresponds to doxastic level cognitive states. And while he clearly recognises that perception can influence doxastic belief, he doesn’t explicitly mention the reverse case (2013: 17). Furthermore, when he mentions the debate concerning whether perception is theory-laden he says that it is “in at least one (rather specific) sense: What we perceive depends heavily upon the set of priors … that the brain brings to bear in its best attempt to predict the current sensory signal” (2013: 7), thereby not explicitly mentioning cognitive states. All of this might suggest that Clark does not think that cognitive states are involved. This is the reading of Clark that Drayson endorses. She says: “I suggest that Clark’s ‘assumptions’ are like Churchland’s assumptions in that they ‘have nothing to do with beliefs, theories, or other doxastic commitments that we may have’ … (Stokes, in press)” (unpublished manuscript: 4). In consequence, she claims: “[Clark’s] model of the brain is entirely consistent with the traditional view of perception and cognition. In particular, the debate over the cognitive penetrability of perception is left untouched” (unpublished manuscript: 1). While this radical reading of Clark is not implausible, and there could be a version of predictive coding that stipulated cognitive states did not play a role, it seems to me that there is an even more radical reading of Clark that is at least as good a reading of Clark as Drayson’s. On this super radical reading, Clark thinks that there is no distinction, or at least no sharp distinction, between perception and cognition. He says: These accounts thus appear to dissolve, at the level of the implementing neural machinery, the superficially clean distinction between perception and knowledge/belief. To perceive the world just is to use what you know to explain away the sensory signal across multiple spatial and temporal scales. The process of perception is thus inseparable from rational (broadly Bayesian) processes of belief fixation, and context (top-down) effects are felt at every intermediate level of processing. As thought, sensing, and movement here unfold, we discover no stable or well-specified interface or interfaces between cognition and perception. Believing and perceiving, although conceptually distinct, emerge as deeply mechanically intertwined. On this interpretation of Clark, one should be eliminativist about the normal categories such as perception and cognition that traditional views of the mind ascribe to. If that is right, then according to such a view, there can be no cognitive penetration for cognitive penetration requires such a distinction. When I discussed (2015: 582) this view as a possible version of predictive coding Lupyan replied that, “I am aware that in the process of challenging these distinctions”, such as, the presence of a clear boundary between (modal) perception and (amodal) cognition/semantics … I am supporting a collapse of perception and cognition that makes the whole question of the penetrability of one by the other, ill-posed. But I would be thrilled if I my arguments contribute to the eventual demise of this question”. Here Lupyan is more cautious with respect to his previous (2015a) commitment to the role of cognitive states in predictive coding and their role in cognitive penetration. Clearly, this is a view of the mind that predictive coders may consider adopting. This version of the predictive coding account of perception entails that there is no cognitive penetration of perception (either of early vision or of perceptual experience), for there is no cognition and there is no perception. A slightly less radical reading of Clark would be that cognition and perception lie on either ends of a spectrum and that cognitive penetration happens when there is the right kind of influence by states that lie on the more cognitive end of the spectrum on states that lie on the more perceptual end. Indeed, Clark sometimes writes as if this were the case, saying that the higher-level something is the “more cognitive” and the lower is the “more perceptual” (2013: 10). One might surmise that the highest levels correspond to doxastic beliefs and desires and the lowest levels to perceptual states. On this account of predictive coding, it entails that there is cognitive penetration. Still, the picture that Clark paints is unclear between the four options just outlined. And Lupyan seems as yet undecided about his final view. Thus, predictive coders should, I think, give further consideration to these questions, specify more precisely which version of predictive coding they wish to endorse, and in that light, reconsider the claims that they make about the relationship between predictive cognitive penetration. To summarise this section, the minimal version of predictive coding leaves open whether cognitive states are some of the high-level states that generate perceptual representations. Given this, the minimal version does not entail, but is consistent with, cognitive penetration. Lupyan’s (2015a) and Hohwy’s (2013) version of predictive coding stipulates that cognitive states are some of the high-level states. If those cognitive states can be seen to affect early vision or perceptual experience, this version of the theory entails cognitive penetration. (I will explore this issue in detail below). Clark is not particularly clear about whether his version of predictive coding stipulates that cognitive states are some of the high-level states. A super radical reading of him would suggest that there is no distinction between cognition and perception and so no cognitive penetration. Interestingly, we saw Lupyan (2015b) indicate that he would consider such a position. A less radical reading of Clark would suggest that perception and cognition lie on a continuum and that cognitive penetration occurs when states that lie far on the cognitive side affect states that lie far on the perceptual side (in the right way). IF PREDICTIVE CODING INVOLVES COGNITIVE STATES, IS THERE REASON TO THINK THAT THEY PENETRATE PERCEPTUAL EXPERIENCE? -------------------------------------------------------------------------------- Recall that in Section 2, I outlined two forms of cognitive penetration. One was the penetration of perceptual experience; the other was the penetration of early vision. In this section, I will focus on the former. Recall that perceptual experience is a conscious state that has a distinctive perceptual phenomenal character. It is the state that we go into when we consciously see, hear, touch, and so on, but also a state that we go into when we consciously hallucinate. The phenomenal character of a perceptual experience of a red rose in front of one (or better as of a red rose in front of one, to include the case of hallucination) is distinct from the phenomenal character of the corresponding belief that there is a red rose in front of one—a belief that one often has when one experiences a red rose (but not always, for one sometimes believes, rightly or wrongly, that one may be suffering from an illusion or hallucination). If cognitive states are involved in the predictive coding that produces perception, is there reason to think that those cognitive states penetrate perceptual experience? In order to answer this question, we must ask: what exactly does the predictive coding account of perception say about perceptual experience? Clark (2013) explains what he and Hohwy, Roepstorff, and Friston (2008) hold is perceptual experience, according to the predictive coding account: [it] is determined by a process of prediction operating across many levels of a (bidirectional) processing hierarchy … and their interactive equilibrium ultimately selects a best overall (multiscale) hypothesis concerning the state of the visually presented world. This is the hypothesis that ‘makes the best predictions and that, taking priors into consideration, is consequently assigned the highest posterior probability’ (Hohwy et al., 2008, p. 690). Other overall hypotheses, at that moment, are simply crowded out: they are effectively inhibited, having lost the competition to best account for the driving signal.” In short, perceptual experience is identified with the best—i.e. surviving—prediction or representation of how the world is. And that is generated by high-level and low-level processing. If one holds that the high-level states are cognitive states (or cognitive enough), then this will entail that perceptual experience is affected by cognition according to predictive coding, and that there is cognitive penetration. Hohwy (2013) holds this position. (Of course, as we have seen, one might desist from this opinion and hold a version of predictive coding that is silent about whether high-level states are cognitive, that denies that cognitive states are involved, or that denies the cognition/perception distinction. These versions will not entail that there is cognitive penetration, with the last two entailing that there is no cognitive penetration.) So far, things have been reasonably straightforward. However, the matter becomes more complex when we consider what predictive coding says about the penetration of early vision, as I explain in the next section. IF PREDICTIVE CODING INVOLVES COGNITIVE STATES, IS THERE REASON TO THINK THAT THEY PENETRATE EARLY VISION? -------------------------------------------------------------------------------- Recall from Section 2, that early vision was defined functionally as the system that takes in information from the eyes (perhaps via attention) and produces shape, size and colour representations. How might predictive coders understand early vision? One option for predictive coders would be to think of early vision, neutrally, as the system—whatever it is—that produces shape, size and colour representations. According to the predictive coding account, high-level systems produce those representations, causally affected by information from the eyes. The information from the eyes therefore feeds into the high- level system that produces those representations; therefore, the whole system, including high-level states will count as early vision. That’s a very odd result—and a very strange conception of early vision! Suppose, for the sake of argument, that we accepted this conception. (I realise that one might object to it.) In that case, a version of the predictive coding account that said that cognitive states were involved would entail that cognitive penetration of early vision occurred. (And if a version of predictive coding was non-committal about whether cognitive states were involved (such as the minimal version) then it would be consistent with but not entail cognitive penetration of early vision. If a version denied that cognitive states were involved, then the view would entail that there was no cognitive penetration of early vision.) As I indicated above, one might object to this account of early vision on the grounds that it is at odds with the conception of it conceived of by Pylyshyn (1999). Therefore, one might reconsider what predictive coders should take early vision to be. Pylyshyn thought that his functional definition would be realized by signals coming from the eyes, perhaps after being filtered by attention, and then a limited portion of bottom-up processing thereafter. Given this, one approach predictive coders might adopt would be to follow Raftopolous and espouse a temporal conception. Raftopoulos (2014: 603) states: “Early vision … includes a feed forward sweep (FFS) in which signals are transmitted bottom-up and which lasts, in visual areas, for about 100 ms.” If we take the first 100 ms of processing to be early vision, what does predictive coding say about that? Does the high-level predictive signal reach down as low as this, or does the first 100 ms contain only an incoming signal that will, in due course, be treated as an error signal? To my knowledge, the current literature on predictive coding does not address this question. Standard evidence concerning whether there is any top-down influence on the first 100 ms of processing would be what we would look to.8 Whether there is penetration of early vision, considered as the first 100 ms of processing, will depend on such evidence (and, of course, as we have seen previously on whether such high-level processes include cognitive states of subjects). There is another conception of early vision that a predictive coder might appeal to. Recall the functional definition given by Pylyshyn: it is the system that takes in information from the eyes (perhaps via attention) and produces shape, size and colour representations. A predictive coder might claim, contra Pylyshyn, that there is no system that takes in information from the eyes (perhaps via attention) and produces shape, size and colour representations. Instead, they could claim that such representations are produced top-down. This would amount to the denial that there is such a thing as early vision! If that is right, then predictive coding is incompatible with the penetration of early vision. Predictive coding entails that there is no cognitive penetration of early vision as there is no such thing as early vision. To summarise this section, I have outlined that there are three conceptions of early vision that someone who endorses the predictive coding account of perception might adopt: that it includes all processing (including top-down processing) that produces shape, size and colour representations; that it is that which occurs during the first 100 ms of processing; that it doesn’t exist. Predictive coding entails penetration of early vision conceived of as in (1) if and only if the high-levels contain cognitive states. Predictive coding entails, or is compatible with, penetration of early vision conceived of as in (2) if and only if the empirical evidence turns out the right way, and the high-levels contain cognitive states. Predictive coding in any form is not compatible with penetration of early vision conceived as in (3).","The results of the above discussion of various forms of predictive coding accounts of perception and various different forms of cognitive penetration are summarized in Table 1. As discussed previously, some proponents of the predictive coding account of perception have claimed that it entails that cognitive penetrations occurs. I have argued that the claim that the predictive coding account entails cognitive penetration is too simplistic. This is because the relationship between the predictive coding account of perception and cognitive penetration is dependent on both the specific form of the predictive coding account and the specific form of cognitive penetration. I have spelled out different forms of each and the relationships that hold between them. While some forms of predictive coding entail some forms of cognitive penetration, some are consistent with but do not entail some accounts of cognitive penetration, and some entail that there is no cognitive penetration. However, many of those forms of predictive coding that entail that cognitive penetration does not occur do so because they give such a radically different account of the mind compared to the standard conception of it. That is, they rule out the existence of cognitive penetration on the grounds that the components of the mind required to play a role in cognitive penetration do not exist. As we have seen, some accounts deny that there is a distinction to be made between perception and cognition and some deny that there is any such thing as early vision. Such models allow for a great deal of top-down processing of the sort that those who have heretofore argued that there is no cognitive penetration would not endorse. In consequence, mere acceptance of the predictive coding approach to perception does not determine whether one should think that cognitive penetration exists. Moreover, as is demonstrated by the preceding sections, establishing as true the version of predictive coding that does entail cognitive penetration occurs will turn on determining many of the same issues as are required for directly establishing cognitive penetration occurs, such as determining that there are there top-down cognitive effects by cognitive states on perceptual experience or early vision. Thus, arguing for cognitive penetration via establishing that predictive coding is the right account of perception does not actually provide a novel means of establishing that cognitive penetration occurs. Nonetheless, it is clear that many forms of predictive coding models of perception either entail that cognitive penetration occurs or they entail that top-down processing of a radical kind occurs—of such a radical kind that they entail a picture of the mind greatly at odds with that accepted by those who believe in cognitive penetration. Finally, a general lesson of this paper is that, given that there are such different conceptions of both predictive coding and cognitive penetration, researchers should cease talking of either without making clear which form they refer to, if they aspire to make true generalisations."],["We examined the relationship between information processing style and information seeking, and its moderation by anxiety and information utility. Information about Salmonella, a potentially commonplace disease, was presented to 2960 adults. Two types of information processing were examined: preferences for analytical or heuristic processing, and preferences for immediate or delayed processing. Information seeking was captured by measuring the number of additional pieces of information sought by participants. Preferences for analytical information processing were associated positively and directly with information seeking. Heuristic information processing was associated negatively and directly with information seeking. The positive relationship between preferences for delayed decision making and information seeking was moderated by anxiety and by information utility. Anxiety reduced the tendency to seek additional information. Information utility increased the likelihood of information seeking. The findings indicate that low levels of anxiety could prompt information seeking. However, information seeking occurred even when information was perceived as useful and sufficient, suggesting that it can be a form of procrastination rather than a useful contribution to effective decision making. --------------------------------------------------------------------------------","Information seeking is a critical component of effective decision making (Griffin, Dunwoody, & Neuwirth, 1999), yet information seeking can be a mechanism for delaying decisions (Jepson & Chaiken, 1990). Hence, a process model must be applied to understand the difference between information seeking as an analytical strategy versus information seeking as procrastination. This study examined the relationship between information processing styles (how decisions are made) and information seeking (the extent to which information is sought), and its moderation by anxiety and information utility. We integrate insights from the risk and information seeking and processing theory (RISP, Griffin et al., 1999), dual process theory (Epstein, 1990; Epstein, Pacini, Denes-Raj, & Heier, 1996), and broaden-and-build theory (Fredrickson, 1998, 2001) to develop and test a model that accounts for individual-level information seeking behaviour, and the contingencies that lead to information seeking as a form of procrastination. Information processing styles ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Information processing styles, typically characterised as tendencies to use analytical or intuitive (heuristic) approaches to choice (Dane & Pratt, 2007) influence decision processes and outcomes. Analytical processes are required for novel, complex problems whereas intuitive or heuristic processes are applied to numerous daily choices (Bargh, Chen, & Burrows, 1996; Epstein, Lipson, Holstein, & Huh, 1992). Theories of analytical and heuristic thinking rest on the dual-process concept which proposes two parallel, interactive systems of thinking (Epstein, 1990; Epstein et al., 1996). System 1 is intuitive, affect-laden and rapid. System 2 is cognitive, resource intense and requires time. Both systems yield positive outcomes. Analytical thinking is associated with effective decision making due to logical reasoning and fewer decision biases (Stanovich & West, 2002), and ability to focus on important aspects of information relevant to decisions rather than non-relevant contextual information (McElroy & Seta, 2003). Intuitive thinking is associated with expertise (Dreyfus & Dreyfus, 2005) and effectiveness in solving everyday problems (Todd & Gigerenzer, 2007). Individual differences in information processing and information seeking ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ While the dual-process model has universal application, the extent to which System 1 and System 2 are applied, and the situational contingencies that influence their use, are subject to individual differences (Epstein et al., 1996). Therefore, theories that rest on dual-process modelling need to take into account individual-level antecedents and moderating factors. Employing this approach, Griffin et al. (1999) developed the risk information seeking and processing (RISP) model. They proposed information seeking is driven by individual differences in perceived information sufficiency, and continues until the point of sufficiency is reached. Griffin et al. (1999) placed information seeking and information processing together as the dependent variables in their model, and proposed that they combine to produce four decisions styles relating to routine/non routine and heuristic/systematic processing. However, recent research into decision processes, also building on dual process models, has added a second information processing style: regulatory processes that influence whether a decision should be made immediately or delayed Dewberry, Juanchich, and Narendran (2013a) proposed both cognitive information processing (rationality vs. intuition) and regulatory information processing have direct effects on decision outcomes. For example, when faced with a decision about whether to eat food that could harbour harmful bacteria, there are choices about whether to go with past experience, i.e. if eating the food has been alright before then it will be alright at this decision point (heuristic processing); or, whether to find out more about the likelihood of bacteria being present in the food product (analytical processing). There are also choices regarding whether to find out the relevant information now (preference for immediate decision making), or whether to put off information seeking until a later date (preference for delayed decision making). Therefore, we propose that perceived information sufficiency, and preferences for analytical and delayed decisions will be associated directly and positively with information seeking. Conversely, we propose that preferences for heuristic and immediate decisions will be associated directly and negatively with information seeking. Individual differences in age and gender also influence decision processes. Older adults are more likely to draw on their history of life experiences when making choices (Finucane, Mertz, Slovic, & Schmidt, 2005), and this increases the likelihood of greater information seeking. Moreover, women tend to be more risk averse when making decisions, and less confident in their choices than men (Graham, Stendardi, Myers, & Graham, 2002), thus increasing tendencies for information seeking. Thus we expect that older adults and women will be more likely to seek information than younger adults and men. Moderators of the relationship between information processing and information seeking ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Dewberry et al. (2013a) and Dewberry, Juanchich, and Narendran (2013b) suggested that anxiety could increase information seeking in order to delay decision making, because the point of choice causes anxiety, so putting off a decision reduces current experiences of anxiety. In a more complete modelling of the relationship between affect and behaviour, Frederickson’s broaden-and-build theory (Fredrickson, 1998, 2001) proposed that positive affect has a broadening and building effect, increasing effectiveness of decisions made. Conversely, anxiety reduces thought-action repertoires and constricts decision processes by limiting access to memory and the cognitive strategies necessary for problem solving. In addition, Fredrickson’s (1998) model suggests that affect moderates the relationship between preferences, perceptions and actions, and this has been confirmed empirically (Soane et al., 2013). Hence, we propose that anxiety moderates the relationships between information processing styles and information seeking because it increases tendencies to search for information that could allay anxiety, and the process delays the pressure of choice. We also propose that information perceptions influence the relationship between information processing style and information seeking. Griffin et al. (1999) suggested that information will be sought when current information is believed to be insufficient. However there will be contingencies that influence this process. Specifically, information utility moderates the relationship between antecedent factors and information seeking (Griffin et al., 1999). Examining this contingency is important to distinguish between information seeking as analytical information processing, and information seeking as a strategy to delay decision making (Bohner, Chaiken, & Hunyadi, 1994; Dewberry et al., 2013a, 2013b). We suggest that perceptions of context-specific information utility will moderate the relationship between information processing style and information seeking. Summary of the current study ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The current study tested a model of information seeking We hypothesised that the relationship between analytical information processing style and information seeking will be positive, and moderated by anxiety, and the information utility. We also hypothesised that the relationship between heuristic information processing style and preference for delaying decisions will be negatively associated with information seeking, and that the negative relationship will be strengthened by anxiety and information utility. Finally, we hypothesised that preferences for delayed decisions will be associated negatively with information seeking, and that the relationship will be moderated by anxiety and information usefulness. Research context ~~~~~~~~~~~~~~~~ To test the research model, we examined a widespread disease, Salmonellosis, that continues to be a threat to human health and a financial burden on society. In Europe, Salmonellosis is the second most common zoonotic disease in humans (after Campylobacter) (European Food Safety Authority, 2010). The most common way of contracting Salmonellosis is through the consumption of raw egg and raw egg products. Although Salmonella bacteria need not cause disease, the incidence of Salmonellosis indicates that changes in domestic behaviour are required to reduce its impact on society. Hence examining decision making in the context of Salmonellosis contributes to practical strategies regarding disease management as well as to understanding decision processes.","An online survey website was used to recruit 3001 participants to complete a questionnaire on food safety. Participants were emailed an invitation to participate in the research and a clickable link to access the survey. Survey responses were stored on the research team’s secure server. Twenty-seven participants were excluded from the analysis because they stated that they had an allergy to either chocolate or eggs and would not eat the chocolate mousse. Fourteen were excluded due to missing data. The final sample was 2960 (96.8% of completions). The mean age was 40.59 (range 18–82, SD = 12.95). There were 1613 men (54.5%) and 1347 women (45.5%). 1102 (37.2%) had a degree or above; 362 (12.2%) had other higher education; 580 (19.6%) had A levels or equivalent; 618 (20.9%) had GSCEs or equivalent (20.8%); 125 (4.2%) had other qualifications; the remaining 111 (3.8%) had no qualifications.","We focused on a food product, home-made chocolate mousse containing eggs, a common source of Salmonellosis and a widely consumed food item. Age was assessed by asking participants to write their age. Gender was assessed by self-rating ‘male’ or ‘female’. Effect of past experience was measured by one item adapted from Miles and Frewer (2001) which asked participants to rate the extent to which past experience has provided them with information about salmonella prevention. There was a 5-point response range from 1 ‘strongly disagree’ to 5 ‘strongly agree’. Anxiety was assessed based on Watson and Tellegen’s (1985) emotion circumplex. Items were ‘I am anxious about being infected with salmonella from eggs’ and ‘I am worried about being infected with salmonella from eggs.’ There was a 5-point response range, 1 ‘strongly disagree’ to 5 ‘strongly agree’. Information sufficiency was measured by two items measuring perceived sufficiency of current information, and adapted from Trumbo and McComas (2003). ‘The information I have at this time meets all of my needs for knowing about how to protect myself from salmonella from eggs’; ‘I have been able to make a decision about how concerned I am about the risk of salmonella in eggs to me by using my existing knowledge’. There was a 5-point response range, 1 ‘strongly disagree’ to 5 ‘strongly agree’. Information utility was assessed using items developed for this study. Participants were presented with four pieces of information: (1) A description of the likelihood of the prevalence of Salmonella in eggs. (2) A description of the reduction of the prevalence of Salmonella in eggs in England between 1995 and 2003. (3) Percentages describing the likelihood of the prevalence of Salmonella in eggs. (4) A graph format showing the reduction of the prevalence of Salmonella in eggs in England between 1995 and 2003. After each piece of information, participants were asked how useful the information was to evaluating whether to eat the mousse or not. There was a 5-point response range from 1 ‘Not at all useful’ to 5 ‘Very useful’. The scale is a mean score of all four items. Information processing styles were measured as four distinct constructs rather than as two bipolar continua following recommendations from Hodgkinson, Sadler-Smith, Sinclair, and Ashkanasy (2009). Four types of information processing style were assessed using scales from Dewberry (2008). All items were in the form of a statement followed by a three-point response range: 1 ‘Disagree’, 2 ‘Uncertain’, 3 ‘Agree’. A sample item from each scale is included with permission from the author (Dewberry, 2008). Analytical information processing was assessed using a three-item scale. Items assessed the extent to which information is sought prior to making a decision, for example, ‘When deciding on something important, I usually stick with the information I already have rather than looking for more’. Heuristic information processing was measured using a three-item scale. Items assessed tendencies to use current knowledge to make a decision rather than information search strategies. A sample item is ‘When making an important decision, I tend to go on the facts I have before me rather than looking around for more information’. Preference for making immediate decisions was assessed using a two-item scale. For example, ‘If I have to make a decision, I start thinking about it straight away.’ Preference for delaying decisions was measured using two items. An example is ‘If I have difficult decision to make, I tend to put it off’. Information seeking behaviour, the dependent variable, was captured by offering participants four extra pieces of information which they could choose to look at. The options were: information concerning health effects of Salmonella; prevalence of Salmonella; national attempts to control Salmonella in eggs; and, individual risk reduction. Items were developed for this study. Participant access to each piece of information was recorded and used to create an index ranging from 0 to 4.","Table 1 shows the means, standard deviations, Cronbach’s alpha where appropriate and inter-scale correlations. Confirmatory factor analysis ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We then examined the model using SEM Confirmatory Factor Analysis in Amos 19. Data indicated that the model fit was acceptable (Hair, Black, Babin, & Anderson, 2009): χ2 = 537.4; df = 114; CFI = .98; NFI = .98; RMSEA = .04; SRMR = .04, apart from the χ2/df value which is 4.7. However, the χ2/df value is sensitive to large sample sizes (Hair et al., 2009) so we proceeded with hypothesis testing. Hierarchical linear regression ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Next we used hierarchical multiple regression for the first stage of hypothesis testing. All continuous variables were standardized using the Z transformation prior to analysis. Model 1 examined direct effects of age, gender, experience, information processing, anxiety, information utility and sufficiency. Model 2 added interaction terms (anxiety, utility and sufficiency × each of the information processing styles). Data are shown in Table 2. Model 1 data showed main effect positive associations between preferences for analytical thinking, tendency to delay decision making, information sufficiency, information utility and information seeking. There were negative associations between heuristic information processing style, anxiety, and information seeking. Thus there was some initial support for our hypotheses concerning information processing style and information seeking. Moreover, women and older adults were more likely to seek information, as expected. Testing for moderation ~~~~~~~~~~~~~~~~~~~~~~ Model 2 data showed six significant interaction terms. The interaction of affect and preferences for making immediate decisions was not examined further because there was no main effect of immediate decision making. The remaining interactions were examined in more detail following procedures discussed in Hayes (2013) and using the ‘process’ syntax. We tested whether the relationship between information processing style and information seeking was different at high and low levels of affect and information utility (1 standard deviation above and below the mean). The conditional effects were calculated using the bootstrapping procedure recommended by Preacher, Rucker, and Hayes (2007) in order to test whether the findings were robust. A T-statistic was computed for the indirect effect. There were two significant interactions: affect × preferences for delaying decision making, and utility × preferences for delaying decision making. Data are shown in Table 3. Fig. 1 shows the interaction between affect and preferences for delaying decision making. There was a positive association between preferences for delaying decisions and information seeking, although there was less information seeking for people experiencing anxiety. As anxiety increased, preferences for putting off decisions reduced the likelihood of information seeking. There was a positive association between information utility and preferences for delaying decision making. Information seeking is most likely for people who perceive the information as useful, yet have a tendency to put off decision making. The relationship is depicted in Fig. 2. Fig. 3 summarises the direct effects and moderation effects.","Integrating dual process theory; (Epstein, 1990; Epstein et al., 1996) with RISP theory (Griffin et al., 1999) and broaden-and-build theory (Fredrickson, 1998, 2001), provides insights into the information seeking process. The current study has demonstrated the importance of individual differences in information processing styles on information seeking, and the susceptibility of information seeking to anxiety and information perceptions in a food-related decision context. In examining these processes, we make two contributions to the literature. First, we proposed that analytical information processing styles would be associated positively with information seeking. Data confirmed this proposal, and showed that there was a direct effect of analytical information processing style on information seeking that was not influenced by anxiety or information utility. Hence, for people with preferences for analytical information processing styles, information seeking is likely to form part of their strategy for finding and evaluating information systematically prior to making a choice. We also hypothesised that preferences for heuristic decision making would be associated negatively with information seeking, and that this relationship would be influenced by anxiety and information utility. Data showed that there was a main effect, but did not support moderation. Thus heuristic preferences were associated directly with low levels of information seeking. These findings show partial fit with Griffin et al.’s (1999) RISP model. We showed that information processing style was associated with information seeking, but there was no evidence for the complex association between the variables proposed in the RISP model. Furthermore, the data indicate that different information processing styles require specific modelling. Our second contribution concerns the application of the regulatory dimension of information processing styles: preferences to make an immediate or delayed decision. Delaying decisions can serve a specific purpose which is to avoid choice, and this is different from the application of heuristics, a process that is concerned with making decisions faster and with less effort. Data showed that preferences for delaying decisions were associated positively with information seeking, and that this relationship was moderated by both anxiety and information utility. Participants sought more information when they experienced lower levels of anxiety. Furthermore, participants sought more information when they perceived what they had read during the study to be useful. Together, these findings suggest that, for people who find it difficult to regulate the decision process, information seeking is a strategy to delay decisions that becomes more likely when information is perceived to be useful, and less likely under conditions of anxiety. The research has several practical implications for policy makers and food safety risk managers. Research into risk communication has moved towards bottom-up development of information that takes lay concerns into account (Bickerstaff, Lorenzoni, Jones, & Pidgeon, 2010; Stern & Fineberg, 1996). This practical strategy could have the benefit of influencing the balance of affect and information perceptions that have a critical influence on information seeking behaviour such that people are motivated enough to read information, e.g. on websites, and educated about how to act on it to change domestic practices and reduce the risk of infection from Salmonella. Such an approach could also avoid raising anxiety levels to the point where people avoid food safety information. Future research could examine further the relationships between information processing styles and information seeking, and the moderating roles of anxiety and information utility. In particular, further examination of the processing that underlies delayed decision making would enable more complete modelling of the relationship, and it is possible that there are other situational moderators that interact with information processing styles. Future research could consider the relationship between information seeking and effective decision making to test for the positive and negative impact of different information processing styles, and do so in different decision contexts. There could also be further examination of the effects of age and gender on decision processes and information seeking. Epstein et al. (1996), for example, found some differences between men and women’s preferences for analytical and heuristic thinking, although the findings were not consistent across studies. It is also possible that decision processes develop with age (Mata, Schooler, & Rieskamp, 2007), thus future research could consider how these demographic factors function in relation to information seeking. The current study has some strengths, notably the assessment of information seeking behaviour rather than information processing (Kahlor, Dunwoody, Griffin, Neuwirth, & Giese, 2003) or intention (ter Huurne, Griffin, & Gutteling, 2009; Yang, 2012). However, while attempts have been made to develop a theory-driven model and test it on a large sample of adults, the current study has acknowledged limitations. We examined information seeking behaviour using online survey technology, however, a laboratory study would enable more complex information seeking behaviour to be assessed. Moreover, an experimental approach could be used to examine whether information processing styles can be influenced by priming or other contextual variables, thus providing more opportunities to examine moderation effects. Finally, different decision contexts, e.g. other kinds of everyday decisions as well as infrequent decision, or decisions with more serious consequences, would add to theoretical and practical developments. In conclusion, this study suggests that individual differences in preferences for analytical and heuristic information processing style have a direct effect on information seeking, and influence the extent to which information is sought. In contrast, regulatory information processing styles have an indirect association with information seeking. Preferences for delaying decisions were exacerbated by information utility and attenuated by anxiety. These findings contribute to a more complete understanding of the decision processes that lead to information seeking. Moreover, the findings suggest that information campaigns could be made effective by providing sufficient information to generate an emotional need to make timely decisions."],["Background: Vocabulary knowledge and speechreading are important for deaf children's reading development but it is unknown whether they are independent predictors of reading ability. Aims: This study investigated the relationships between reading, speechreading and vocabulary in a large cohort of deaf and hearing children aged 5 to 14 years. Methods and procedures: 86 severely and profoundly deaf children and 91 hearing children participated in this study. All children completed assessments of reading comprehension, word reading accuracy, speechreading and vocabulary. Outcomes and results: Regression analyses showed that vocabulary and speechreading accounted for unique variance in both reading accuracy and comprehension for deaf children. For hearing children, vocabulary was an independent predictor of both reading accuracy and comprehension skills but speechreading only accounted for unique variance in reading accuracy. Conclusions and implications: Speechreading and vocabulary are important for reading development in deaf children. The results are interpreted within the Simple View of Reading framework and the theoretical implications for deaf children's reading are discussed. --------------------------------------------------------------------------------","Despite having intelligence scores in the normal range, the majority of deaf children have poorer reading outcomes than their hearing peers (e.g. Conrad, 1979; Kyle & Harris, 2010; Lederberg, Schick, & Spencer, 2013; Wauters, van Bon, & Tellings, 2006). Large scale studies report that deaf school leavers have reading ages far behind their chronological ages (see Qi & Mitchell, 2011 for a review) and reading skills seem to develop at only a third of the rate of hearing children (Allen, 1986; Kyle & Harris, 2010). There is a consistent picture of underachievement in reading skills which can have long-lasting effects upon future employment opportunities. It is therefore imperative to gain a better understanding of which cognitive and language skills are important for reading development in deaf children and the complex relationships between these abilities. Recent research has suggested that speechreading (silent lipreading) and vocabulary are longitudinal predictors of deaf children's reading development (Kyle & Harris, 2010, 2011), and that speechreading is also predictive of reading ability in hearing children (Kyle & Harris, 2011). However, the relative contribution of these two skills to reading is unknown; therefore the main aim of this study is to examine whether speechreading and vocabulary are independent predictors of reading in deaf and in hearing children. The predictors of reading ability, and the often complex relationships between predictors, are well documented in hearing children. One of the most widely-acknowledged predictors of early reading is phonological knowledge and skills (e.g. Adams, 1990; Castles & Coltheart, 2004). Children with better phonological awareness (the ability to detect and manipulate the constituent sounds of words) and greater knowledge about the relationships between letters and sounds tend to make the most progress in reading in the early stages (see Castles & Coltheart, 2004; Goswami & Bryant, 1990). However, it is also well known that different cognitive and language based skills are predictive of different components of the reading process, i.e. letter-sound knowledge and phonological skills are most predictive of word recognition and word reading whereas higher order language skills such as grammar and syntax are most predictive of reading comprehension (see Catts & Weismer, 2006; Muter, Hulme, Snowling, & Stevenson, 2004; Oakhill, Cain, & Bryant, 2003; Storch & Whitehurst, 2002). Vocabulary is generally thought of as being most important in the beginning stages of reading where it predicts initial word recognition (e.g. Dickinson, McCabe, Anastasopoulos, Peisner-Feinberg, & Poe, 2003; Verhoeven, van Leeuwe, & Vermeer, 2011) and emerging comprehension skills (e.g. Anderson & Freebody, 1981; Ricketts, Nation, & Bishop, 2007; Roth, Speece, & Cooper, 2002); however, research suggests that vocabulary knowledge also plays an important role in later reading skills (e.g. Senechal, Ouellette, & Rodney, 2006; Verhoeven et al., 2011). The question of whether phonological skills are important, or even necessary, for deaf children's reading skills is a matter of ongoing debate. In summary, the research evidence is very mixed. Some authors find evidence for the role of phonological skills in deaf children's reading (e.g. Campbell & Wright, 1988; Dyer, MacSweeney, Szczerbinski, Green, & Campbell, 2003; Easterbrooks, Lederberg, Miller, Bergeron, & Connor, 2008) while many others report a very small or non-significant relation (e.g., Hanson & Fowler, 1987; Kyle & Harris, 2006; Leybaert & Alegria, 1993; Mayberry, del Giudice, & Lieberman, 2011; Miller, 1997). These discrepancies hold true even when traditional phonological awareness assessments are adapted to make them more deaf-friendly, i.e. by representing the items pictorially. A recent meta-analysis of the literature concluded that there was little evidence that deaf individuals use phonology in their reading (Mayberry et al., 2011). In their analyses Mayberry et al. (2011) included 25 studies that looked at the relation between reading and phonological coding and awareness in deaf individuals, ranging from young children to adults and from across the spectrum of language and communication preferences (sign/speech). The resulting effect sizes for the relationship between reading and phonological awareness ranged from −.13 to .81 with a mean of .35. This means that on average, across the 25 studies, 11% of the variance in reading skills in deaf participants was explained by phonological abilities. This figure refers to the contribution of spoken phonology to reading for deaf participants. It should be noted that signed languages are also phonologically structured, albeit with different parameters (e.g. Brentari, 1999; Sandler & Lillo-Martin, 2006). Corina, Hafer, and Welch (2014) recently reported a positive correlation between phonological awareness of American Sign Language and phonological awareness of English. Furthermore, McQuarrie and Abbott (2013) reported a correlation of .47 between phonological awareness in American Sign Language (ASL) and reading. However, as yet there is no evidence for a direct causal relationship between the knowledge of sign language phonology and reading. An important caveat regarding the Mayberry et al. meta-analysis is that the study mainly included correlational studies. Due to the low number of longitudinal studies in deaf children only two such studies were included (Harris & Beech, 1998; Ormel, 2008). Correlational relationships between variables do not infer causality. Furthermore, in longitudinal studies of reading development it is important to control for early levels of reading ability on later reading outcomes because of the well documented auto-regressor effect, whereby the strongest predictor of later ability is typically earlier achievement in that particular skill (see Caravolas, Hulme, & Snowling, 2001; Castles & Coltheart, 2004). In the only two longitudinal studies of deaf reading development that have used this approach, the data suggest that deaf children may develop their phonological skills through reading (Kyle & Harris, 2010, 2011). Reading and phonological awareness were found to be related in deaf children, but the direction of the observed relation was from earlier reading ability to later phonological awareness. Therefore this may be different to the typical relation seen in hearing children whereby phonological awareness is a strong initial predictor of reading ability and then the two skills tend to exhibit a reciprocal relation (Burgess & Lonigan, 1998; Castles & Coltheart, 2004). In deaf children, it seems that it is reading ability that initially predicts phonological awareness, but then, similar to hearing children, the two skills become reciprocally related and develop in a common and mutually beneficial manner. This pattern of development also fits in with an emerging pattern in the deaf literature, that of the predictive relation between phonological awareness and reading being stronger, or more apparent, in older deaf children and adults, than in young children. That is, the role of phonological skills in deaf reading becomes stronger once reading skills are more proficient. It is important to note that the studies included in the meta-analysis of reading and phonological awareness studies conducted by Mayberry et al. (2011) covered deaf participants from a very broad age range, from early childhood to adulthood. New insights from the longitudinal studies outlined above suggest that combining data across this very wide age range of deaf participants may not be appropriate. A different way of looking at whether phonological skills have a role in deaf reading is to examine the contribution of visual-based phonological skills such as lipreading or speechreading. There is growing evidence that some deaf individuals do make use of phonology but that their phonological strategy may be slightly different to that of hearing individuals because it is mainly derived from speechreading information rather than auditory input. Speechreading (or lipreading), which is the skill of processing speech from the visible movements of the head, face and mouth, has the potential to be useful in phonological processing when hearing is absent. In addition to providing visual information about vowels (mouth shape), some consonantal phonemes are ‘easy to see’ and ‘hard to hear’. For example, /n/ and /m/ form one pair of consonants that are confusable auditorily, but clear visually (see Summerfield, 1979). Individual differences in speechreading skill have been found to predict reading outcomes in deaf individuals, both cross-sectionally and longitudinally (Arnold & Kopsel, 1996; Geers & Moog, 1989; Kyle & Harris, 2006, 2010, 2011). If the information gleaned through speechreading is used as the input for a phonological code, then the better one is at speechreading, the more specified and distinct the underlying representations are likely to be. These underlying representations can then be used to form the basis for a phonological code. The quality of phonological representations is thought to be related to reading ability (Elbro, 1996; Swan & Goswami, 1997) because the more specified and distinct the underlying representations, the better able the individual is to complete phonological awareness tasks (Elbro, Borstrøm, & Petersen, 1998) and performance on these types of tasks is extremely indicative of reading ability (see Castles & Coltheart, 2004 for a review). The strong predictive relationships found between speechreading and reading in deaf children (e.g. Kyle & Harris, 2010, 2011), the presence of speechread errors in both deaf children's spelling (Burden & Campbell, 1994; Leybaert & Alegria, 1995; Sutcliffe, Dowker, & Campbell, 1999) and in their performance on phonological awareness tasks (Hanson, Shankweiler, & Fischer, 1983; Leybaert & Charlier, 1996) suggest that, when deaf children are learning to read, the better they are at speechreading the more information they will be able to use when making connections between letters and sound (see also Alegria, 1996; Campbell, 1997). Many of the earlier studies only found speechreading was associated with levels of reading ability in orally educated deaf children (e.g. Arnold & Kopsel, 1996; Campbell & Wright, 1988; Craig, 1964; Geers & Moog, 1989); however, more recent research has reported speechreading to be a strong longitudinal predictor of deaf children's reading development, regardless of language preference (Kyle & Harris, 2010, 2011). This shift can be readily explained by the changes in deaf educational practices in the UK over the past 30 years as fewer deaf children are now educated in specialist schools and there is more integration in mainstream schools. This, combined with the introduction of bilingual education through British Sign Language (BSL) and English, has impacted upon the teaching of reading and language to deaf children making it more likely that almost all of them are exposed to both oral speech and sign to some extent. The other skill that is increasingly reported as being important for deaf children's reading ability is vocabulary knowledge (e.g. Geers & Moog, 1989; Kyle & Harris, 2006, 2010, 2011; LaSasso & Davey, 1987; Mayberry et al., 2011; Moores & Sweet, 1990). Vocabulary and language skills seem to be imperative for deaf reading regardless of how either skill is assessed or which component of reading is measured. ‘Language skills’ were found to be the largest contributor to reading ability in the Mayberry et al. (2011) meta-analysis. The category ‘language skills’ included measures of signed and spoken vocabulary (amongst other measures). Vocabulary was also the strongest and most consistent longitudinal predictor of both word reading and reading comprehension in the Kyle and Harris longitudinal studies (2010, 2011). This could be considered relatively unsurprising given the well documented language delays in deaf children (Waters & Doehring, 1990; Musselman, 2000). In hearing children, vocabulary and good language skills have been proposed as providing a possible compensatory mechanism for children who have poor phonological skills (e.g. Nation & Snowling, 1998; Snowling, Gallagher, & Frith, 2003). This explanation is equally plausible, if not more so, for the strong relationship between vocabulary and reading in deaf children. Previous studies looking at the role of vocabulary and speechreading in deaf children's reading have had insufficient sample sizes to determine whether these two skills are independent predictors of reading in deaf children. In Kyle and Harris (2010), speechreading was mainly a longitudinal predictor of early reading skills and thus it is of interest to determine whether this relationship feeds into vocabulary development. It would make sense that deaf children who have better speechreading skills have larger vocabularies yet the converse relationship whereby having a more extensive vocabulary would enable one to be a better speechreader is also likely (Davies, Kidd, & Lander, 2009). What is not known is whether speechreading and vocabulary, which are likely to be related themselves, make independent contributions to reading or whether they are simply reflecting some common underlying language factor or capacity. On the other hand, these two factors could interact in a more complex developmental manner, for example, one skill may ‘jumpstart’ the development of the other skill at one stage, but then become less relevant. Vocabulary and speechreading underpin the model of deaf reading proposed by Kyle (2015), which was based upon the Simple View of Reading (Gough & Tunmer, 1986). The Simple View of Reading postulates that reading is made up of two components: a decoding component and a linguistic component, both of which are necessary for reading. For deaf children, as argued by Kyle (2015) and Kyle and Harris (2011), speechreading contributes to phonological representations and thus forms the basis, along with phonological awareness, for the decoding component, and vocabulary knowledge contributes to the linguistic component. This is not to say that other skills are unimportant for reading in deaf individuals but the role of these two particular skills is explored in the current study. It should also be noted that the relationship between speechreading and reading has mainly been measured at the level of the single word reading and single word speechreading. It is therefore unknown whether speechreading at different linguistic levels also predicts reading and whether the strength of this relationship varies for different reading components. Kyle and Harris (2006, 2010) found that single word speechreading was significantly related to word reading but not reading comprehension and the relationship was stronger with word reading than sentence comprehension. The current study uses a recently developed Test of Child Speechreading (ToCS; Kyle, Campbell, Mohammed, Coleman, & MacSweeney, 2013), which assesses speechreading at three different levels: words, sentences and sort stories. We investigate how performance at these different levels is related to word reading and reading comprehension. This will help to shed light upon the role that speechreading plays in reading, for example, is the relationship simply at the lexical level or does it reflect broader linguistic knowledge? Kyle and Harris (2011) also reported that speechreading of single words was longitudinally predictive of beginning reading development in hearing children. This finding warrants further investigation as although it is easy to understand why speechreading is predictive of reading in deaf individuals, due to impaired auditory access, it may not be immediately obvious why speechreading would also be predictive of reading growth in hearing children. We would argue that a similar explanation also holds for hearing children: speechreading is related to reading because the visual speech information derived through speechreading is likely to be incorporated into phonological representations. Therefore, better speechreading skills may result in more distinct and specified phonological representations which in turn can help children when learning to read. The key difference is the supplementary nature of this information and detail for hearing children. Evidence from research with blind children supports this viewpoint as studies often report delays in discriminating phonological contrasts that are difficult to distinguish in the auditory domain but are visually distinct (Mills, 1987). Lastly, given the suggested role of speechreading in reading, it is important to understand what makes a good child speechreader. Research with adults has shown that better speechreaders tend to be deaf, use oral language to communicate, have higher levels of reading ability and report they can understand the public (Bernstein, Demorest, & Tucker, 1998). No such comparable studies have been conducted with deaf children but the findings from separate studies do help shed light on possible correlates of speechreading. Relations have previously been reported between speechreading and vocabulary in young hearing children (Davies et al., 2009) and between speechreading and working memory (Lyxell & Holmberg, 2000), phonological awareness (Lyxell & Holmberg, 2000; Kyle & Harris, 2010) and NVIQ (Craig, 1964) in hearing-impaired children. Interestingly, recent research has shown no difference between deaf and hearing children in their speechreading ability (Kyle et al., 2013) but speechreading was found to improve with age (Kyle et al., 2013). Moreover, when using an adult speechreading test very similar to the ToCS, deaf adults were found to have superior speechreading skills in contrast to their hearing peers (Mohammed, Campbell, MacSweeney, Barry, & Coleman, 2006; Mohammed, MacSweeney, & Campbell, 2003). The contribution of demographic, background and audiological factors to children's speechreading will be investigated in the current study. This study explores the role of speechreading and vocabulary in reading ability with a large sample of deaf and hearing children. The main aims were to (1) to determine whether speechreading and vocabulary are independent predictors of reading ability in deaf children; (2) to investigate the role of speechreading at different linguistic levels for different components of reading ability (i.e. does speechreading at different linguistic levels exhibit different relationships with reading components); (3) to examine whether speechreading and vocabulary are independent predictors of reading in hearing children; and (4) to explore the effect of demographic and background variables on speechreading.","Eighty-six deaf children and 91 hearing children aged between 5 and 14 years old took part in this study. The mean age of the deaf children was 9 years 6 months (SD = 31.5) and the mean age of the hearing children was 9 years 1 month (SD = 30.2). Thirty-nine deaf children and 53 hearing children were male. Deaf children were recruited from specialist schools for the deaf and resource bases for students with hearing impairments attached to mainstream schools across Southern England. To ensure that deaf and hearing children were similar in terms of demographic backgrounds, the hearing children were recruited from the mainstream schools to which the resource bases were attached. Deaf and hearing children were from a range of different ethnic backgrounds: 59% of the children were White British or White European, 11% were Black British or Black Other, 21% were Asian British or Asian Other and the remaining 9% were mixed race or other. There were no significant differences between the deaf and hearing children in their gender distribution (X2(1) = 2.94, ns), ethnicity (X2(3) = 6.99, ns), chronological age (t(175) = .98, ns) and NVIQ (t(175) = −1.87, ns). All deaf children had a severe or profound bilateral hearing loss of greater than 70 db with a mean loss of 97.7 db. Thirty-five of the deaf children had cochlear implants (CI) and the remaining (apart from two) wore digital hearing aids. The average age at which deafness was diagnosed was 17 months (SD = 12.3). The majority of deaf children were in hearing-impaired resource bases attached to mainstream schools but a third were in specialist schools for the deaf. Children varied in their language and communication preferences: 44 preferred to communicate through speech; 33 preferred to use signing (26 used BSL and 7 used Sign Supported English); six used total communication (a mixture of both signing and speech) and the remaining three were bilingual in spoken English and BSL. Table 1 presents descriptive statistics for background information for the deaf participants separated out for device use (CI, digital hearing aids and no device)","Four tasks were administered to assess reading ability, speechreading skills, expressive vocabulary and NVIQ. Reading ability The Neale Analysis of Reading II (NARA II: Neale, 1997) was used to assess reading accuracy and reading comprehension skills. Children were shown a booklet containing short passages and asked to read them aloud in their preferred communication mode: English, BSL or a combination of the two. They were then asked a series of questions about each passage to test their comprehension, which they were allowed to answer in their preferred communication. Children received an accuracy score for their word reading and for their comprehension skills. The task was administered according to the instruction manual, apart from the instructions being delivered in the child's preferred language or communication method. Speechreading ability Speechreading ability was measured using the Test of Child Speechreading (ToCS: Kyle et al., 2013). The ToCS is a child-friendly, computer-based assessment that measures silent speechreading at three different psycholinguistic levels: words, sentences and short stories. It uses a video-to-picture matching design whereby children are presented with silent video clips of either a man or a woman speaking and they have to choose the picture (from an array of four containing the target and 3 distractors) that matched what was said in the video clip. For example, in the word subtest, for the target item “door”, the pictures were “door”, “duck”, “fork” and “dog”. An example of a sentence trial was the target “The baby is in the bath” and the distractors were pictures depicting a baby reading a book, some pigs on a path and an elephant having a bath, The short story subtest has a slightly different format in which participants see the speaker saying a short story and are then asked two questions about it. They answer each question by choosing the correct picture from an array of four. For example one of the questions is “where is Ben going?” and the correct answer “school” is depicted along with three viable distractors “home”, “cinema” and “library”. Full details about the ToCS design, item selection and development can be found in Kyle et al. (2013). The instructions were specifically designed so that they could be delivered in the child's preferred communication or language, BSL, spoken English or a combination of the two. ToCs has been shown to have high external validity as an assessment of silent speechreading and good internal reliability (α = .80) (Kyle et al., 2013). The task took about 20 min to administer. Expressive vocabulary The Expressive One Word Picture Vocabulary Test II (EOWPVT II: Brownell, 2000) was used to assess children's expressive vocabulary. Children are shown pictures of increasing difficulty and asked to name them. Children were allowed to respond in their preferred communication and therefore this task was providing an indication of their expressive vocabulary regardless of language preference. However, it should be noted that this task was designed to test English vocabulary and not BSL or sign language vocabulary. Following guidelines from Connor and Zwolan (2004), any answer that was not gestural was accepted. Two items in the test were changed to make it more suitable for British children, following Johnson and Goswami (2010) who used this test with deaf children in the UK. The item racoon was changed to badger and the map of USA to a map of the UK. In addition, and following pilot studies, we changed the pictures for two items to pictures more characteristic of British responses: prescription and windmill. Non-verbal skills An estimate of non-verbal intelligence (NVIQ) was derived from the Matrices subtest of the British Abilities Scales II (BAS II: Elliot, Smith, & McCulloch, 1996). This test has been used previously with deaf children of similar age to those in the current study (see Harris & Moreno, 2004; Kyle & Harris, 2010). Procedure Children were all tested individually in a quiet room, normally adjacent to the classroom. Each child was seen over two testing sessions, not lasting more than 20 min each. All standardised tests were administered according to the instruction manuals but the instructions were delivered in the child's preferred communication method. Written parental consent was given for all children and the child's assent was also sought at the beginning of the first testing session. Ethical clearance was granted from the University Research Ethics Committee. Performance on the ToCS, vocabulary and reading tasks ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The means and standard deviations for all tasks are presented in Table 2. Raw scores were used in all analyses as we were unable to obtain standard scores for all participants across all tests as a few of the children were just outside the top age range for the reading assessment. Standard scores are reported for the vocabulary assessment for descriptive purposes only. As a group, the hearing children achieved age-appropriate scores for reading accuracy and reading comprehension (mean chronological age = 9:01; mean accuracy reading age = 10:00; mean comprehension reading age = 9:08). The deaf children exhibited an average reading delay of sixteen months in reading accuracy and 22 months in reading comprehension (mean chronological age = 9:06; mean accuracy reading age = 8:02; mean comprehension reading age = 7:08). The hearing children had significantly higher vocabulary standard scores than the deaf children, t(175) = −10.87, p < .001, 95% CI −28.8 to −19.8. The mean vocabulary standard score for the deaf children was 76 (SD = 15.1) whereas the mean standard score for the hearing children was 100.2 (SD = 15.1). As reported in Kyle et al. (2013), deaf and hearing children did not differ in their speechreading skills (mean 49.0% vs. 50.6%, respectively) and showed an almost identical pattern of performance across the subtests. A two-way mixed design ANOVA (hearing status by ToCS subtest) revealed no statistically significant differences between the deaf and hearing children in their overall performance on ToCS, F(1,172) = .11, ns. There was a main effect of subtest, whereby children achieved higher scores on the single words > sentences > stories, F(2,344) = 294.61, p < .001. There was no significant interaction between group and ToCS subtest F(2,344) = .29, ns. Correlations between reading, speechreading and vocabulary ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Tables 3 and 4 present the partial correlations between reading, speechreading and vocabulary controlling for age and NVIQ. Age was statistically controlled, as performance on ToCS has been found to improve significantly with age over this age-range (see Kyle et al., 2013). NVIQ was also controlled for as although there was no significant association between NVIQ and speechreading, there were small yet significant associations between NVIQ and reading and vocabulary. After statistically controlling for age and NVIQ, performance on ToCS (combined score across three subtests) for both deaf and hearing children was significantly related to reading accuracy (r = .49, p < .001 and r = .31, p = .005, respectively) and reading comprehension (r = .44, p < .001 and r = .28, p = .010). Performance on all three speechreading subtests was related to reading accuracy in deaf children; however only performance on the sentences was related to reading in the hearing children. Speechreading was also significantly associated with vocabulary knowledge (even after controlling for age and NVIQ) in both deaf children (r = .25, p = .021) and hearing children (r = .25, p = .02). Multiple regression analyses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The main research aim was to determine the relative contributions of vocabulary and speechreading to reading ability in deaf children. A set of fixed order multiple regression analyses was conducted (see Table 5) to investigate whether speechreading and vocabulary were independent predictors of reading accuracy and comprehension. After controlling for age and NVIQ (entered in steps 1 and 2 and accounting for 54% of the variance), speechreading (entered in Step 3) accounted for 11% of the variance in the deaf children's reading accuracy scores. When vocabulary was entered in step 4, it accounted for an additional 13%. When the order in which they were entered into the regression analyses was exchanged so that vocabulary was entered in Step 3 before speechreading, it accounted for 15%. Speechreading still accounted for a small yet significant proportion of the variance (8%) in reading accuracy even when entered after vocabulary. Thus, speechreading and vocabulary seem to be relatively independent predictors of reading accuracy in deaf children as the proportion of variance each skill explains is not particularly dependent upon the order in which it is entered into the analysis. The same analysis was conducted with reading comprehension as the dependent variable. Table 5 shows that speechreading and vocabulary were also independent predictors of reading comprehension for deaf children as the proportion of variance that each accounted for was fairly consistent regardless of the order in which they were entered. Age and NVIQ accounted for 59% of the variance in deaf children's reading comprehension scores. When entered in Step 3, speechreading accounted for almost 8% and vocabulary accounted for an additional 15% of the variance (in step 4). When the order in which vocabulary and speechreading was entered was switched, vocabulary accounted for 17% (step 3) and speechreading for 6% (step 4). The same analyses were conducted for the hearing children (see Table 5). After controlling for age and NVIQ (68%), speechreading and vocabulary were both small yet significant predictors of reading accuracy, regardless of the order in which they were entered. Speechreading accounted for 3% (in step 3) and vocabulary accounted for an additional 3% (in step 4). If the order was switched, vocabulary accounted for 4% (in step 3) and speechreading accounted for 2% (in step 4). Therefore, vocabulary and speechreading are accounting for a portion of independent variance in reading accuracy scores in hearing children. A different picture was observed when reading comprehension was the outcome variable for the hearing children. In this instance, speechreading was only a significant predictor if entered before vocabulary. Age and NVIQ accounted for almost 73% of the variance in reading comprehension so there was little variance left that could be accounted for. When entered in step 3, speechreading was a small predictor (2%) and vocabulary (step 4) accounted for 8%. However if the order was changed, vocabulary accounted for 9% but speechreading no longer accounted for any significant variance. Therefore, speechreading and vocabulary were small yet significant independent predictors of reading accuracy for hearing children, but in contrast to the deaf children, speechreading was not a significant independent predictor of reading comprehension. How are background and audiological factors related to speechreading? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As speechreading was a strong predictor of reading in deaf children, the role of background and audiological factors in determining what makes a good child speechreader was examined. The effects of gender and NVIQ were investigated for both deaf and hearing children. A two-way ANOVA (gender × hearing status) revealed no effect of gender, F(1,176) = 3.47, ns, no main effect of hearing status (deaf vs hearing), F(1,176) = 3.47, ns, and no significant interaction, F(1, 176) = 1.44, ns. There was no significant association between NVIQ and performance on ToCS for deaf or hearing children (r = −.01, ns and r = −.06, ns respectively). Effect of hearing/audiological factors (degree of loss, type of hearing aid and age of diagnosis) An additional set of analyses was undertaken for the deaf cohort. There was no significant correlation between degree of hearing loss and overall speechreading ability, r = −.17, ns. However, there was a significant negative correlation between degree of hearing loss and performance on the sentences subtest (r = −.25, p = .024) and stories subtest (r = −.27, p = .015) but not single words (r = .05, ns). Children with less severe levels of hearing loss scored higher on the sentences and the stories section. There was no effect of age of diagnosis on speechreading scores, r = .13, ns and there were no differences between deaf children with hearing aids (n = 49) and those with CIs (n = 35), t(81) = −.87, ns. Effect of communication preference A two-way ANOVA revealed a main effect of child's communication preference, F(1,81) = 27.67, p < .001), whereby those children who preferred to communicate through oral language achieved higher scores on the ToCS than those who preferred to communicate through signing or total communication. There was also a main effect of speechreading subsection F(2, 162) = 178.16, p < .001) but no significant interaction F(2, 162) = 1.67, ns). Those children who preferred to communicate through oral language had lower levels of hearing loss than those who signed or used total communication, t(82) = −3.25, p = .002.","The main aim of the current study was to investigate the relative contributions of vocabulary and speechreading to reading development and determine if these relationships held across different psycholinguistic levels. Speechreading and vocabulary were found to be independent predictors of reading ability in children but the precise strength of the relationship was dependent upon hearing status and the component of reading being investigated. For deaf children, speechreading and vocabulary were independent predictors of both reading accuracy and reading comprehension and accounted for an appreciable portion of the variance in their reading scores. In contrast, for hearing children, while both speechreading and vocabulary were independent predictors of reading accuracy, only vocabulary was an independent predictor of reading comprehension and accounted for a smaller proportion of the variance. These findings provide further evidence of the importance of speechreading and vocabulary for deaf children's reading (e.g. Kyle & Harris, 2010, 2011; Easterbrooks et al., 2008; Mayberry et al., 2011) and for the idea that speechreading skills should not be ruled out as a factor influencing hearing children's reading (see Kyle & Harris, 2011). Due to the large age range in the current study, age understandably explained a lot of the variance in reading accuracy and reading comprehension scores which left little for speechreading and vocabulary to be able to explain. Thus, after age and NVIQ were entered into the analyses, it is remarkable that speechreading and vocabulary were indeed able to account for any further variance, particularly for the hearing children. While it might be reasonable to expect a relationship between reading and speechreading in deaf children (where speechreading can provide access to phonological information in the absence of auditory input), it is not as immediately obvious why a small yet significant relationship would be observed for hearing children. However, the contribution of speechreading and vocabulary for hearing children can also be interpreted through the Simple View of Reading framework whereby vocabulary again provides input for the linguistic component while speechreading feeds into the decoding component. It is well-known that speech processing plays a role in the development of phonological skills and helps form phonological representations. What differs between the role of speechreading for deaf and hearing children's reading in this model is the necessity: for deaf children, speechreading is often one of the main ways of accessing spoken language whereas for hearing children, it is supplementary. The better a child is at speechreading, the more distinct their phonological representations are likely to be and more specified representations are linked with better reading (e.g. Elbro, 1996). This fits in with the viewpoint that a phonological code is not necessarily tied to the auditory domain but is abstract and therefore it can be derived from speech and speechreading (see Alegria, 1996; Campbell, 1997; Dodd, 1987). Information derived through speechreading has been shown to be processed in a similar manner to auditory speech (Campbell & Dodd, 1980; Dodd, Hobson, Brasher, & Campbell, 1983). In prior research, the strong association reported between speechreading and reading in deaf children has been mainly limited to the level of word reading (Kyle & Harris, 2010, 2011) whereas the current study extends this relationship to reading comprehension. However, previous studies only assessed speechreading of single words and the current study measured speechreading of words, sentences and stories; and a composite of these three levels was found to predict reading comprehension. It is also important to note that the age range in the current study was from 5 to 14 years whereas it was only 7–10 year olds in Kyle and Harris (2010) and therefore it is possible that speechreading plays a more important role in deaf reading comprehension as reading skills develop. This is in line with results from the deaf adult literature where speechreading has been found to correlate significantly with reading comprehension (Bernstein et al., 1998; Mohammed et al., 2006). The extension of this association between reading and speechreading beyond single words is important as it suggests that the relationship is not due simply to perceptual matching (i.e. matching a single word token to a single speechread token). It is more likely to have a linguistic basis, especially as the ToCS was shown to be an ecologically valid assessment of speechreading (see Kyle et al., 2013). Speechreading of sentences and short stories requires higher-order linguistic skills such as parsing and grammatical knowledge, which are equally required for comprehension of written texts. It is important to remember that for deaf children, performance on all three psycholinguistic levels was related to reading accuracy and comprehension. For both deaf and hearing children, vocabulary knowledge was the strongest independent predictor of reading accuracy and reading comprehension. This concurs with previous findings that language skills, including vocabulary knowledge, typically exhibit the strongest relationship with reading in deaf children (e.g. Easterbrooks et al., 2008; Kyle & Harris, 2006, 2010, 2011; Mayberry et al., 2011; Moores & Sweet, 1990; Waters & Doehring, 1990). Whilst this also fits with findings with hearing children, the exact strength of the relationship observed with hearing children usually depends upon which components of language and reading are being measured and the age of the children (see Ricketts et al., 2007). Future research should attempt to explore the contribution of broader language skills to reading development in deaf children rather than just vocabulary knowledge. It would also be interesting to determine whether the role of speechreading and vocabulary in deaf children's reading development is constant across different subgroups of deaf children. Although speechreading and vocabulary were independent predictors of reading ability, they were also inter-related to some extent in both deaf and hearing children, as has been reported by Davies et al. (2009) for young hearing children. One interpretation of this relationship is that one cannot speechread a word that is not already in one's vocabulary; however, for deaf children it is equally plausible to suggest that speechreading leads to vocabulary growth and indeed for some it may be the only way that spoken words enter the mental lexicon. For deaf children in particular, it is most likely to be a reciprocal relationship rather than uni-directional. Another explanation can be found in the theories of Metsala and Walley (1998) and Goswami (2001) who argue that the development of phonological awareness and vocabulary are closely linked because as vocabulary knowledge expands, there is increased pressure for the underlying phonological representations to become more distinctive, which in turn leads to improved phonological awareness. It is noteworthy that there were very few relationships observed between speechreading and other background and demographic skills. Similar to findings with deaf adults, those deaf children who preferred to communicate through speech were better speechreaders (see Bernstein et al., 1998). The lack of a significant effect of gender or NVIQ on child speechreading proficiency also concurs with more recent adult investigations (Auer & Bernstein, 2007). Although there was no overall association between degree of hearing loss and speechreading, deaf children with less severe levels of hearing loss scored higher on the sentences and the stories sections, perhaps suggesting a more supplementary functional use. It is reasonable to assume that speechreading combined with higher levels of residual hearing might lead to better speech perception (both audio- visual and visual alone) than lower levels of residual hearing combined with speechreading, although equally, the greater the level of deafness the more reliance one may have to place upon speechreading. This does raise a possible question over the validity of focusing on speechreading if it cannot be determined what makes a good speechreader in children. However, a better way of approaching this issue would be to implement speechreading training to determine whether it is a skill that can be trained in children. Finally it should be noted that the current study is only correlational and therefore causality cannot be inferred. However, there is no reason to assume that the direction of the relationships observed in this study is any different to that reported in recent longitudinal studies (see Kyle & Harris, 2010, 2011) in which speechreading and vocabulary predicted development in reading rather than reading predicting growth of speechreading and vocabulary.","In conclusion, vocabulary and speechreading have been shown to be independent predictors of reading ability in deaf children and to a lesser extent in hearing children. It is likely that better speechreading skills result in more accurate phonological representations and the current results can be understood within reading models that suggest skilled reading necessitates both a decoding (speechreading) and a linguistic component (vocabulary). These findings suggest that focussing on both vocabulary development and speechreading skills in young deaf children may form a fruitful basis for helping support early reading development in young deaf children. There are several educational implications that can be drawn from these findings. Teachers working with deaf children who use speech to communicate are likely to already have an understanding of the importance of visual speech and highlight this as a source of information. However our results show that speechreading is important for reading development in deaf children from other language backgrounds, including signing and possibly those with cochlear implants. Thus teachers working with these cohorts should also draw children's attention to the complementary phonological information that is visible on the face. Teachers and educators working with typically-developing hearing children should be aware that their children are also probably incorporating phonological information derived from visual speech into their representations and that encouraging children to be look at the lips when learning sounds is likely to help them form more distinct phonological representations. Drawing attention to information from visual speech is likely to not only help children initially distinguish between similar sounding phonemes but this information is likely to help create more distinct representations which will help with reading skills. The findings also provide further evidence for teachers working with either deaf or hearing children that reading is not only about decoding words, but that vocabulary, and most likely broader language skills not measured in the current study, also play an essential role in reading development.","This paper investigates the importance of speechreading and vocabulary skills for reading development in a large cohort of deaf and hearing children ranging in age from 5 to 14 years. Previous research reported a relationship between speechreading and reading ability but only between speechreading of single words and word reading. The current study has extended this relationship to include larger units of speechreading (sentences and short stories) and to reading comprehension. Importantly, the current findings suggest that speechreading and vocabulary make independent contributions to reading ability in deaf children and possibly contribute to hearing children's reading. The results are interpreted within current theoretical frameworks of reading development for hearing and deaf children."],["We assessed 3- to 6-year-old children's production of two-clause sentences linked by before or after. In two experiments, children viewed an animated sequence of two actions and were asked to describe the order of events in specific target sentence structures. We manipulated whether the target sentence structure matched the chronological order of events (e.g., “He finished his homework, before he played in the garden” [chronological order]) or not (e.g., “Before he played in the garden, he finished his homework” [reverse order]). Children produced fewer accurate target sentences when the presentation order of the two clauses did not match the chronological order of events, specifically for target sentences linked by after. Independent measures of vocabulary and memory both were related to performance, but vocabulary was the stronger predictor. We conclude that developmental improvements in children's ability to produce two-clause sentences linked by a sequential temporal connective are driven primarily by language ability rather than memory capacity per se. This work also highlights the advantages of using both sentence repetition (Experiment 1) and blocked elicited production (Experiment 2) paradigms to elicit sentence production in young children. --------------------------------------------------------------------------------","We experience events in the world around us in real time as they occur. In the production of speech and text, however, the speaker or writer does not need to relate events in the order in which they occur. Instead, linguistic devices such as the temporal connectives before and after may be used to refer to events in reverse order (e.g., “Before he ate the cookies, he put on his jumper”). Although children produce sentences containing before and after from around 3 years of age (Diessel, 2004), they have difficulties with correct usage up to at least 9 years (Peterson & McCabe, 1987; Winskel, 2003). That is, children’s production of sentences that include these expressions may belie their full competence because they may have better knowledge of one construction over the other. In this study, we focused on 3- to 6-year-old children’s production of two-clause sentences containing the connectives before and after. We demonstrate that language ability has a stronger influence on performance than working memory capacity per se. Successful production of language draws on an integrated and coherent mental representation of the state of affairs being described, also known as a prelinguistic message (Bock, 1987; Levelt, 1989). When a speaker narrates events in reverse order, as in “She put on her gloves, after she had combed her hair,” the language used deviates from the speaker’s mental representation of the actual sequence of events. When adult speakers choose to do this, they draw on greater processing resources than when planning and producing chronological order sentences (Habets, Jansma, & Münte, 2008; Ye, Habets, Jansma, & Münte, 2011). As noted above, children’s understanding and production of before and after continue to develop for several years after these temporal expressions first appear in their speech. What is not known is whether or not other linguistic and structural features of sentences with temporal connectives, such as reverse order narration, contribute to these developmental differences. We present the first systematic study of how sentences expressing different temporal orders of events affect children’s sentence production accuracy. A speaker’s choice to narrate events in their chronological order using either before or after influences whether the temporal connective occurs in the initial position of the sentence (e.g., “After she combed her hair, she put on her gloves”) or in the medial position (e.g., “She combed her hair, before she put on her gloves”). This was our first research question: How do these features—order of events, connective, position of connective—individually or in combination influence young native speakers’ production of sentences containing temporal connectives? Studies of children’s production and comprehension of temporal connectives show that before is acquired earlier than after (Blything & Cain, 2016; Clark, 1971; Pyykkönen & Järvikivi, 2012). Clark (1971) attributed the difference in age of acquisition for before and after to the semantic features of each term; before indicates the prior event, whereas after does not, making the latter more semantically complex. In addition, before is used more consistently as a temporal connective than after, which is commonly used also as a preposition as in “Watch out, he is only after your money” (see the British National Corpus [Leech, Rayson, & Wilson, 2001]). Thus, after has a less consistent form–meaning relationship and is theoretically more complex than before. This literature suggests that the use of after may involve greater planning and processing effort than the use of before, which may influence the accuracy of sentence production. Another feature that might influence children’s sentence production accuracy is the position of the connective. Corpus studies of spoken language have reported that children and adults use connectives in an initial position infrequently (Diessel, 2004, 2008). This finding has been related to the memory load involved in maintaining the information signaled by the connective from the beginning of the sentence while processing the meaning of the first clause (Diessel, 2004, 2008). Conversely, the preference to use a medially placed connective is associated with processing ease because it provides the linguistic information about temporal order at a point close to when the events can be integrated during the incremental processing of language. Our second research question was to identify which framework best explains variation in performance between different sentence structures and across development. A traditional memory capacity-constrained account attributes performance on language processing tasks to the availability of resources within an independent system of working memory, which limits the amount of information that can be maintained during planning and production (e.g., Carpenter, Miyake, & Just, 1994). A more nuanced language-based perspective of working memory shifts emphasis from the “quantity” of information that can be represented to the “quality” (i.e., the content) of the representation of that information in long-term memory (McElree, 2006), which in turn frees up shared processing resources so that they are allocated to the representation of information in active working memory (e.g., MacDonald, 2016). We examined both of these accounts in our study. A classic theory concerning the role of working memory in sentence processing is the memory capacity- constrained account (e.g., Carpenter et al., 1994). According to this viewpoint, an effect or interacting effect of the aforementioned features—order of events, connective, position of connective—is driven by working memory capacity alone. The primary emphasis is that some sentence structures are more difficult to process than others because they require more information to be held within the limited-capacity working memory system. The account builds on a framework that assumes that working memory is a separate system from long-term memory (Baddeley & Hitch, 1974; Baddeley, 2003) to argue that the accurate representation of information is driven by the availability of processing resources specific to the working memory system. The availability of processing resources determines how many individual language units can be accurately represented (but note that the constitution of a “unit” is undefined; see McElree, 2006, for a full review of limitations). It follows that production accuracy is expected to be weaker in individuals with low working memory capacity because they have fewer resources available for maintaining information in working memory. Under such circumstances, the representation of the language form and structure may decay and be forgotten. There is empirical support for the memory capacity- constrained account of sentence processing from a variety of studies. First, there are studies suggesting that difficulties in producing more complex utterances can be attributed to the availability of resources within an independent working memory system. Patients with working memory capacity deficits display a substantially longer speech onset than controls when producing various utterances, and this difference is more pronounced for utterances with more complex structures (e.g., Martin & Freedman, 2001; Martin, Miller, & Vu, 2004). In addition, healthy speakers produce an increased proportion of double object datives (a more complex dative structure; e.g., “The pirate is giving the monk the book”) relative to prepositional datives (a simpler dative structure; e.g., “The pirate is giving a book to the monk”) when they are not required to maintain a verbal memory load (Slevc, 2011). Specifically in relation to the production of two-clause sentences containing temporal connectives, functional magnetic resonance imaging (fMRI) and electroencephalography (EEG) studies with adults have attributed the extra processing effort for reverse order sentences to the maintenance of additional concepts within working memory (Habets et al., 2008; Ye et al., 2011). However, these latter studies did not include an independent measure of working memory. Complementary work from studies of sentence comprehension show that both adults’ (Münte, Schiltz, & Kutas, 1998) and children’s (Blything & Cain, 2016; Blything, Davies, & Cain, 2015) weaker performance for reverse order sentences is related to an independent measure of working memory (but see De Ruiter, Theakston, Brandt, & Lieven, 2018, for a study with children that did not demonstrate a relationship between memory and sentence comprehension). Alternatively, given that the amount of information that can be held in working memory is often far less than the length of a complex sentence, it has been argued that memory capacity alone cannot be an adequate explanation of the pattern of performance seen by children or adults in sentence processing tasks (MacDonald, 2016; McElree, 2006). A language-based account of sentence processing proposes that the effects of working memory are indirect via language knowledge (e.g., MacDonald, 2016). This argument draws on the framework that, rather than being separate systems, working memory and long-term memory are part of a unitary architecture in which working memory is a temporarily active portion of long-term memory (Ericsson & Kintsch, 1995; McElree, 2006). From this viewpoint, language knowledge influences sentence processing because good language skills free up shared resources within the proposed unitary architecture to support the accurate representation of information in active working memory. There is empirical support for this position. The specificity or distinctness of words in the target utterance has been shown to influence adults’ sentence production (Gennari, Mirković, & MacDonald, 2012; Montag & MacDonald, 2014, 2015; Smith & Wheeldon, 2004). In these studies, speakers are less accurate when a task involves the activation of competitors that carry a similar meaning to target items. In a picture description task, Gennari et al. (2012) contrasted conditions in which the pictured agent and patient were highly similar (e.g., builder, miner) or not (e.g., builder, astronaut). Speakers were more likely to avoid more complex structures and omit optional words in the highly similar condition that permitted potential competition between the agent and patient meanings (e.g., producing “The builder who’s being slapped” rather than “The builder who’s being slapped by the miner”). It follows that robust language knowledge will ease the accessibility and retrieval of target items over competitors that are also partially activated in memory. In relation to the production of reverse order sentences containing temporal connectives, this account would posit that an accurate transformation of the mental representation of the order of events is determined by the availability of processing resources that are shared with language retrieval operations. Crucially, a weak lexical representation of a target connective, or other words in the sentence, will lead to less differentiation in activation compared with competitors with similar meaning (i.e., different temporal connectives to the target). This would disrupt sentence planning and production. On this basis, young language users may experience difficulties with complex sentences (i.e., reverse order) because the quality of their lexical representations (i.e., connectives or other words in the sentence) is weaker. We do not yet know precisely how children’s sentence production differs for sentences expressing different temporal orders of events. As noted, studies of adults’ sentence production have reported processing difficulties for reverse order sentences (Habets et al., 2008; Ye et al., 2011). However, these studies have used stimuli in which the connective was presented only in the sentence-initial position. As a result, the effects of connective (before or after) and event order (chronological or reverse) cannot be disentangled. From a developmental perspective, a fully factorial design that includes all permutations of these factors (before–chronological, before–reverse, after–chronological, after–reverse) is essential because children display developmental differences in their understanding of before and after (Clark, 1971). Overview of study aims, methods, and hypotheses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We conducted two experiments designed to determine whether a memory capacity-constrained account (e.g., Carpenter et al., 1994) or a language-based account (e.g., MacDonald, 2016) of sentence processing best explains young children’s production of two-clause sentences containing before and after. Each clause related a single event. We manipulated the connective (before or after) and whether the order of mention of events was chronological or reverse. As a result, the position of the connective was manipulated (medial or initial). Note that a reverse order sentence with after places the connective in the medial position, whereas a reverse order sentence with before places the connective in the initial position. Thus, we manipulated connective and event order in our materials. We also examined the extent to which independent measures of language (receptive vocabulary) and working memory explained variance in performance. If memory capacity is a critical influence on children’s production of complex sentences, we would expect sentences that relate events in reverse order to be produced less accurately than those that relate events in chronological order. Specifically, before–chronological sentences should be produced most accurately because they contain features that would not be expected to increase the amount of information that must be held in working memory (chronological order, medial position, less complex connective). In contrast, the other structures each have two factors that increase the amount of information maintained in working memory: before–reverse (reverse order, initial position), after–chronological (initial position, more complex connective), and after–reverse (reverse order, more complex connective). Furthermore, if memory capacity is the critical influence on accurate production, our independent measure of memory should predict the effect or interacting effect of these features, and also serve as a proxy for age, when both age and memory are included in the model. A language-based account predicts that the influence of working memory is indirect and modulated by language knowledge (MacDonald, 2016). As a result, any difficulties in producing reverse order sentences should be more pronounced when they are linked by the connective after because the planning of a reverse order sentence should be disrupted more easily when it contains after than when it contains before. In addition, the independent measure of vocabulary should modulate performance because it indicates the quality of an individual’s language knowledge. According to this account, the inclusion of an independent measure of vocabulary should improve model fit over and above the inclusion of an independent measure of memory. Furthermore, because vocabulary knowledge grows with age, vocabulary effects should supersede any age effects once included in the model. In summary, there were three aims of this study. The first was to establish which features—order of events, connective, position of connective—influence the accuracy of children’s production of two-clause sentences containing the temporal connectives before and after. Second, we examined which account best explains the pattern of performance—a traditional memory capacity-constrained account or a more nuanced perspective of working memory that argues that language knowledge influences memory processing and storage. Third, we asked whether the same pattern of performance is reproduced across our two paradigms designed to elicit sentence production.","We assessed sentence production using a sentence repetition task in which participants heard a target sentence and were asked to repeat it back to the experimenter. Sentence repetition is a sensitive measure of processing ease because the participants are required to process the syntactic and semantic information and then formulate the sentence themselves using the same sentence production mechanisms as in spontaneous speech (see Boyle, Lindell, & Kidd, 2013; Lust, Flynn, & Foley, 1996). In general, children are less accurate when repeating sentences with more difficult structures. Previous studies of children’s production of sentences containing temporal connectives using sentence repetition have contrasted sequential (e.g., then, before) and simultaneous (e.g., while, when) connectives (Keller-Cohen, 1981; Winskel, 2003), so these do not speak to the issues addressed in this article.","A total of 67 monolingual, typically developing 3- to- 6-year-old children were recruited from schools of mixed socioeconomic status in the North West region of England. Children were in three different school year groups: 20 3- and 4-year- olds (ages 3;5 [years;months] to 4;7; 13 boys), 23 4- and 5-year-olds (ages 4;9 to 5;9; 12 boys), and 24 5- and 6-year-olds (ages 5;9 to 6;8; 11 boys). Written parental consent was obtained, and children provided oral assent before each session.","All children completed a sentence repetition task split between two sessions. In addition, one session included an assessment of receptive vocabulary, and the other session included an assessment of memory. Each session lasted no longer than 20 min. Sentence repetition A total of 32 two-clause sequences containing before and after were constructed. Each of the 32 items conveyed the temporal order of two events that were arbitrarily related (e.g., “He put on the socks, before he ate the burger”). These items were counterbalanced across four lists so that they each represented one of four sentence constructions (shown in Table 1). The four constructions were the product of manipulations of the order of mention of events (chronological or reverse) and the connective (before or after). We also created 32 filler sentences in which the sequence of events in a sentence was typical and supported by world knowledge rather than arbitrary (e.g., “He put on the socks, before he put on the shoes”). Sentences that relate typical sequences (world knowledge present) may reduce the working memory demands of the task by scaffolding the structure of the sentence (Zwaan & Radvansky, 1998). We included these sentences to enhance the likelihood that children would produce full sentences in the task and to maintain their confidence (proportion accuracy reported in Table A1 in Appendix). Each sentence was visually represented by cartoon animations, one for each clause and each lasting 3 s. These were created using Anime Studio Pro 9.1 (Smith Micro Software; http://anime.smithmicro.com/index.html). Animations make children more likely to use the actor, action, and object of the target sentence, thereby increasing accuracy (Ambridge & Lieven, 2011). Each animation segment explicitly showed an object (e.g., shoes) from one of the clauses; the object (e.g., burger) from the other clause was not present. Each animation segment was followed by a freeze-frame judged by the researchers to best represent the action of that clause. Each segment (e.g., Tom eating a hotdog) was 486 pixels in height and did not exceed the left or right half of the presentation (486 × 872 pixels). The experiment was run using the PsyScript 3.2.1 (Slavin, 2013) scripting environment on a Macintosh laptop computer connected to a monitor. Presentation of items was fully randomized. Practice trials emphasized the importance of producing an exact copy of the narrated sentence. Children practiced each of the four sentence constructions used in the experimental items (i.e., two-clause sentences linked by before and after) (see Table 1 for examples). The animation on the left-hand side of the screen was shown first, followed by the animation on the right-hand side of the screen. The instruction to prompt production began with “Can you say …” and was followed immediately by the narration of the target sentence. A response window was signaled by a short beep. The presentation order of the animation segments corresponded to the actual order of events rather than the narrated order. Responses were recorded using a digital voice recorder (Olympus VN-5500; Watford, Hertfordshire, United Kingdom) and were later transcribed and scored. Children who were not able to repeat a sentence after four practice trials completed another set of four practice trials. With this level of practice, each child was able to copy at least one sentence. For each experimental trial, an exact repetition was scored as a target response. Based on recommendations by Lust et al. (1996), a response was also marked as a target response if a change was only minor such as a change to the label for a subject (e.g., Sue, she), a verb (e.g., put on, putted on), or an object (e.g., ketchup, tomato sauce). This lenient criterion was used because marking such changes as non-target responses would create unnecessary noise when the main point of interest was to evaluate the variance that was caused by the factors we had hypothesized to affect children’s ability to accurately communicate the order of events using a temporal connective. The time taken between the beep and the start of children’s response was extracted using Audacity (Mazzoni, 2014). There were no significant differences in response times between age groups or sentence constructions for target responses, so response time data are not reported. Non-target responses were first categorized into three broad types: sense maintained, sense changed, and incomplete. We categorized responses as sense maintained if children inaccurately repeated the target sentence but successfully communicated the order of events by using a temporal connective. The sense maintained responses were counted as non-target responses because at least one critical feature of the target sentence was missing (connective, order of mention, or position) (see Table A2 in Appendix). Responses were categorized as sense changed when a non-target order of events was communicated. Responses were categorized as incomplete when children failed to respond, omitted a clause, failed to use a connective, or used the connective and. Responses that used the connective and (42) were categorized as incomplete because and does not explicitly specify order (Peterson & McCabe, 1987), so we were unable to categorize whether the response maintained or changed the sense or order. Within each of the three broad non-target response categories, we coded the specific change or combination of changes that children had made. The Appendix includes examples and frequency counts of each specific non-target response type (see Table A2). A second coder blind to the hypotheses coded at least 10% of the data (randomly selected) from each year group. Agreement between coders was good for both accuracy (target vs. non-target responses, agreement = 99%, Cohen’s κ = .96) and the categories of non-target responses (agreement = 96%, Cohen’s κ = .80). Memory Working memory was assessed using the digit span task from the Working Memory Test Battery for Children (Pickering & Gathercole, 2001). In this task, children are required to recall the order of a string of digits read aloud by the assessor. The number of digits in a string increases until children cannot successfully recall strings of that length on three separate trials. This assessment of memory was selected because it was most appropriate for our youngest children, who have been reported to perform at floor on more complex measures of working memory (Gathercole, Pickering, Ambridge, & Wearing, 2004). By using this measure, we could capture variance in memory performance across the entire age range. Raw scores were used in the analysis. The test–retest reliability reported in the manual for children aged 5 to 7 years is high at r = .81. Vocabulary Each child completed the British Picture Vocabulary Scale-III (Dunn, Dunn, Styles, & Sewell, 2009). In this task, children hear a word and are asked to point to one of four pictures that best illustrates its meaning. Testing is discontinued when a specified number of errors have been made. Raw scores were used in the analysis. Design A 3 × 2 × 2 mixed design was used. The between-participants independent variable was year group (3- and 4-year-olds, 4- and 5-year-olds, or 5- and 6-year-olds), and the within-participants variables were connective (before or after) and order (chronological or reverse). The position of the connective was manipulated as a function of the manipulations of connective and order. Two analyses were conducted: one with number of target responses as the dependent variable and the other with non-target response types as the dependent variable. Method of analysis The main analysis of the number of target responses was completed using generalized linear mixed-effects models (GLMMs) (Baayen, Davidson, & Bates, 2008; Barr, Levy, Scheepers, & Tilly, 2013). Significant interactions were explored by additional analyses to identify the source of the interaction and are reported below. These were conducted using the lme4 package from the R statistics environment (Bates, Maechler, Bolker, & Walker, 2015; R Development Core Team, 2016). A binomial link function was specified because the outcome variable was binary (i.e., target/non-target). We followed the recommendations of Barr et al. (2013) for obtaining an optimal model. Our maximum random-effects models did not converge, so the decision to incorporate random intercepts and slopes for participants and items was determined by the result of incremental likelihood ratio tests that demonstrated whether each specific random effect significantly improved model fit (Barr et al., 2013). We describe the optimum models for each respective dataset later in Tables 2 and 3, where the first column provides the coefficient estimates of effects (b) due to experimental conditions, the change in the log odds accuracy of responses associated with each fixed effect. A positive coefficient indicates that the effect of a factor is to increase the odds of a target response, whereas a negative coefficient indicates that the factor decreases the odds of a target response. Age in months (continuous), order (chronological or reverse), and connective (before or after) were entered as fixed effects. We used the scale function to scale and center the age, memory, and vocabulary predictors. Memory The raw memory scores [mean (SD)] demonstrated age-related improvements: 3- and 4-year-olds = 21.65 (5.66); 4- and 5-year-olds = 22.65 (3.70); 5- and 6-year-olds = 25.42 (3.45). In addition, the standardized scores of memory were within the normal range of 85 to 115 for each age group: 4- and 5-year-olds = 101.39 (12.43); 5- and 6-year-olds = 105.96 (12.19); standardized scores are not available for 3- and 4-year-olds. Vocabulary The raw vocabulary scores demonstrated age-related improvements: 3- and 4-year- olds = 72.65 (26.16); 4- and 5-year-olds = 78.26 (9.76); 5- and 6-year-olds = 102.30 (8.59). All children had a standardized score above 85, and the mean scores indicated that each age group was performing at an age-appropriate level: 3- and 4-year-olds = 111.35 (13.08); 4- and 5-year-olds = 101.22 (9.14); 5- and 6-year- olds = 101.54 (9.10). Analysis of accuracy data A total of 2144 responses were recorded. Fig. 1 shows the means for each sentence structure by age in years for ease of comparison (note that the analyses were conducted using age in months as a continuous variable). For the two younger groups, 19 responses were removed because they were inaudible, leaving 1357 responses for analysis. Only 13 responses were judged to be inappropriate (nonsense or no response), indicating that children understood the purpose of the task. The initial model included the main predictors of age, order, and connective. The inferential statistics, main effects, and interactions are summarized in Table A3 of the Appendix. Response accuracy was significantly affected by age, indicating that performance improved between 3 and 6 years of age. There was also a significant effect of order, such that children were more likely to repeat chronological order sentences accurately than reverse order sentences. A significant effect of connective was also found; children were more likely to repeat sentences containing before accurately than those containing after. Order and connective were involved in a significant two-way interaction, which was examined by conducting simple interaction analyses of the effects of order for each connective separately. A main effect of order was evident for after sentences but not for before sentences. Children found it more difficult to accurately repeat after–reverse sentences compared with after–chronological sentences, whereas accuracy was equivalent for before–chronological and before–reverse sentences. The final model incorporated memory and vocabulary as additional factors to age, order, and connective (see Table 2). In comparison with the initial model, log-likelihood tests indicated that the fit of the data was significantly improved when we incorporated memory alone, χ2(4) = 20.01, p < .01, when we incorporated vocabulary alone, χ2(4) = 12.67, p = .01, and when memory and vocabulary were incorporated together, χ2(8) = 29.21, p < .01. Memory and vocabulary both significantly influenced performance, such that stronger sets of skills in both domains improved performance. The main effects of order and connective remained significant and were again involved in a significant two-way interaction. There was no three-way interaction with either memory or vocabulary. Analysis of non-target responses The frequency of different types of non-target responses was investigated to determine whether particular types were associated with specific experimental conditions (sentence constructions) and/or age group. This provided an opportunity to examine additional support for either the memory or language account, as outlined in the Introduction. We excluded responses from the oldest age group because their high accuracy scores left too few non-target responses for meaningful analysis (120 of all 702 non-target responses; 17%). The two youngest age groups made 582 non-target responses. The majority of non-target responses involved a change of sense to the meaning of the target sentence (sense changed = 358; 61% of all non-target responses analyzed). Fewer non-target responses maintained the sentence meaning (sense maintained = 131; 23%) or were incomplete (incomplete responses = 93; 16%). These three categories of non-target responses did not vary substantially by experimental condition, although it is worth noting that sense changed errors made up a higher proportion of 3- and 4-year-olds’ non- target responses (66%) compared with those of 4- and 5-year-olds (57%). To further examine non-target responses, we calculated the percentage of sense changed responses that involved a change of connective, order, or position. A change of connective was the most common (252; 70% of all 358 sense changed responses); there were far fewer position (109; 30%) and order (78; 22%) changes. These values add up to more than 100% because the non-target response types are not mutually exclusive; children could include more than one of these changes in their response. We conducted a further analysis to examine the most common type of sense changed response: those involving a change to the target connective (252). Of these, we excluded 34 responses involving a change to the target connective other than before and after (e.g., then, and then, when). This was because children’s use of other connectives to communicate a non-target order of events was not a clear indicator for a weak representation of before or after. The 218 remaining changes to the connective were explicit demonstrations of producing before or after to communicate a non-target event order: before instead of after (e.g., “He put on the socks, before he ate the burger” instead of the target “He put on the socks, after he ate the burger”) or after instead of before (e.g., “He put on the socks, after he ate the burger” instead of the target “He put on the socks, before he ate the burger”). The most obvious reason for why children make these changes is that they have a weak representation of the precise meaning of before or after. We examined the percentage of the total non-target responses (582) in each experimental condition (i.e., age, order, connective) that were caused by a sense changed response involving a change of before instead of after or a change of after instead of before (218: 37% of all non-target responses). The Appendix provides descriptive statistics (Table A4) and a summary of the GLMM (Baayen et al., 2008) for this analysis (Table A5). The changes were involved in a significantly larger percentage of the non-target responses for reverse order sentences (44%) compared with chronological order sentences (30%). The changes were more frequent for the youngest age group than for the middle age group (3- and 4-year-olds = 40%; 4- and 5-year-olds = 35%), but the difference was not significant. Similarly, although these changes were less common for before sentences than for after sentences (before = 33%; after = 41%), the difference was not significant. Finally, although these changes were most common for after–reverse sentences, the interaction between connective and order was not significant (before–chronological = 29%; after–chronological = 31%; before–reverse = 37%; after–reverse = 48%). Note that Table A4 also provides descriptive statistics to show that a similar pattern is present for the percentage of the total non-target responses (582) in each experimental condition that was caused by all changes of connective (i.e., sense maintained and sense changed responses and also inclusive of changes to then, and then, and when). This pattern (described above) was not evident for non-target responses that involved a change to order or a change to position.","The sentence repetition task was successful at eliciting production of complete two-clause sentences linked by an appropriate temporal connective, yielding very few incomplete responses (no more than 8% of all responses in any age group). Our experimental manipulations demonstrated an influence of event order and connective on production accuracy. In addition, performance on independent measures of memory and vocabulary improved the overall fit of the model. These results do not provide unequivocal support for either the memory capacity-constrained account (Carpenter et al., 1994) or the language-based account (e.g., MacDonald, 2016) of sentence processing. When considered together with the analysis of the non-target responses, the results lend greater support to the language-based account for the reasons discussed below. Reverse order sentences are proposed to incur a greater memory load than chronological order sentences because the speaker must produce the first occurring event as the second clause, which requires this information to be maintained in working memory during planning and production (Habets et al., 2008). Our participants were less accurate in producing reverse order sentences linked by after than those linked by before, demonstrating that this effect was specific to the connective. Independent measures of both memory and vocabulary improved the fit of the model. These findings lend greater support to the language-based account of sentence processing, namely that there is an indirect relation between memory and sentence processing that is modulated by language. The explanation for this is that young children’s lexical representations for after are less precise and secure than those for before because after is acquired later and used less consistently as a temporal connective. For that reason, it may be more difficult to accurately plan and maintain in memory multi-clause sentences linked by after during language production, particularly when the event order is reversed. In that way, variation in language knowledge may lead to difficulties with sentence production, particularly for sentence structures that have a high processing load such as those relating events in reverse chronological order. Our findings do not rule out the alternative memory capacity-constrained account because the independent measure of memory made a significant and independent contribution to the fit of our statistical model. However, our analysis of non-target response types provides additional support for the language-based account. Changes to the target connective that used before and after to communicate a non-target event order (sense changed responses) were more likely for reverse order sentences than for chronological order sentences. These made up the majority of sense changed responses that involved a connective change (218 of 252; 87%). The most obvious reason for why children make these changes is that they have a weak representation of the connective itself. This is supported by the main effect of connective type in our main analysis; children were less accurate at producing sentences containing after than those containing before in general. This analysis of non-target responses indicates that an inaccurate representation of the connective itself (as measured by a change in connective) does not provide the support needed for the planning and production of reverse order sentences. Also note that, although not significant, the descriptive statistics by sentence are in line with a language-based account because an inaccurate representation of the connective influenced a greater percentage of the non- target responses to target after–reverse sentences (52%) than to the other target sentence constructions (ranging from 31% to 40%).","A limitation with the sentence repetition paradigm used in Experiment 1 is that it places additional demands on memory compared with speech production because children need to store the just-heard sentence prior to production. For that reason, sentence repetition might not be the most sensitive task to differentiate the memory capacity-constrained and language-based accounts of children’s and adults’ sentence processing. Experiment 2 sought to further test these accounts using a different method to elicit sentence production. We used a blocked design task comprising four blocked sets of items, each assessing the ability to produce one of the four target sentence constructions (e.g., Huttenlocher, Vasilyeva, & Shimpi, 2004). These blocked conditions were designed to complement Experiment 1 by minimizing the contributions of sentence comprehension and memory associated with sentence repetition and maximizing spontaneous production of sentences.","A new sample of participants was recruited (N = 67): 23 3- and 4-year-olds (ages 3;8 to 4;11; 10 boys), 23 4- and 5-year-olds (ages 4;9 to 5;9; 13 boys), and 21 5- to 6-year-olds (ages 5;10 to 6;9; 10 boys). Materials, procedure, and design Children completed the same independent measures of memory and receptive vocabulary as in Experiment 1. Sentence production was assessed using an elicited production task with a blocked design over two separate sessions. Each session lasted no longer than 20 min. One session included the vocabulary assessment, and the other session included the memory assessment. Elicited production: Blocked design The same stimuli from Experiment 1 were used. The 64 items (32 fillers) were split into four testing blocks, each preceded by a training phase in which children were instructed to use a specific target sentence structure. Depending on which block children performed first, the experimenter provided this instruction: “In this game, I am going to ask you to watch two videos and to say what happened using the word before/after. I want you to tell me the order that he/she did these things, and I want you to use before/after in the middle/at the start of your sentence.” Corrective feedback was provided for all four practice items, and training was repeated if children failed to produce a single target sentence. Three 3- or 4-year-olds and one 5- or 6-year-old were excluded from testing after this phase because they each failed to accurately produce any of the target structures. As in Experiment 1, the order in which the animations were presented corresponded to the order of events described by the target sentence. An instruction was narrated: “Can you tell me the order that Tom did these things?” A response window was signaled by a short beep. The four blocked conditions were counterbalanced. Responses were recorded and were later transcribed and scored. We used the same criteria for scoring accuracy of responses and for categorizing non-target responses as in Experiment 1. We did not analyze the time taken to start a response because this measure was found not to be sensitive in Experiment 1. Agreement between the coders was good both for scoring accuracy of responses (target vs. non- target responses agreement = 99%, Cohen’s κ = .97) and for categorizing non- target responses (agreement = 96%; Cohen’s κ = .96). Memory The raw memory scores demonstrated age-related improvements: 3- and 4-year-olds = 21.15 (2.12); 4- and 5-year-olds = 22.57 (2.12); 5- and 6-year-olds = 25.05 (4.95). In addition, the standardized scores of memory were within the normal range of 85 to 115 for each age group: 4- and 5-year-olds = 100.52 (14.85); 5- and 6-year-olds = 105.9 (12.73); standardized scores are not available for 3- to 4-year-olds. Vocabulary The raw memory scores demonstrated age-related improvements: 3- and 4-year-olds = 70.85 (7.78); 4- and 5-year-olds = 82.48 (12.02); 5- and 6-year-olds = 90.95 (9.90). All children had a standardized score above 85, and the mean scores indicated that each age group was performing at an age-appropriate level: 3- and 4-year-olds = 111.75 (7.07); 4- and 5-year-olds = 100.45 (14.85); 5- and 6-year- olds = 102.15 (16.26). Analysis of accuracy data The main analysis of the number of target responses was completed using the same procedures of model fitting described for Experiment 1. A total of 45 responses (2%) were excluded because they were inaudible or interrupted, leaving 1345 responses. Fig. 2 reports the mean accuracy scores for each experimental condition by age group. Of note, performance for each age group was poorer than in Experiment 1, with the most marked difference in scores being for the youngest age group. We report the initial model with age, order, and connective entered as fixed effects in Table A6 of the Appendix. Table 3 shows the final model that incorporates memory and vocabulary as additional factors to age, order, and connective. For the initial model, we found the main effects of age, order, and connective that were reported in Experiment 1. As predicted, older children produced a greater proportion of accurate responses than younger children, chronological order sentences were easier than reverse order sentences in general, and sentences containing before were easier than those containing after. There were two significant two-way interactions. The first, between age and order, was also apparent in Experiment 1; the other, between age and connective, was not found for the sentence repetition task. These effects were qualified by a significant three-way interaction among age, order, and connective. We examined the significant three-way interaction by conducting simple interaction analyses for the effects of age and order for each connective separately. For before sentences, only the main effect of age reached statistical significance (see Table A6 in the Appendix for a full breakdown of results, and see Fig. 2 for graphs of these effects by sentence construction): Accuracy was equivalent for before–chronological and before–reverse sentences. For after sentences, there were main effects of age and order, and these were also involved in a significant two- way interaction. Children found it more difficult to produce after–reverse sentences accurately than after–chronological sentences, and this difficulty with after–reverse sentences was more pronounced for the younger children. We tested three additional models. The addition of memory to the original model significantly improved the fit of the data, χ2(4) = 20.11, p = .01, and resulted in a significant three-way interaction among memory, order, and connective. This suggests that memory modulated the interaction between connective and order. The memory alone model is reported in Table A7 of the Appendix. In another model, we added vocabulary to the original model and also found improved fit compared with the original model, χ2(8) = 33.57, p = .01. In the final reported model (see Table 3), we included both vocabulary and memory. This resulted in improved fit compared with the memory alone model, χ2(4) = 12.08, p = .02, and there was a main effect of vocabulary but not of memory. In addition, the memory by order by connective interaction was not evident when vocabulary was also present. Analysis of non-target responses Responses by 5- and 6-year-olds were excluded because their high accuracy scores resulted in too few non-target responses for meaningful analysis (15% of all non- target responses by the three age groups; 152 of 1019). We analyzed the 867 non- target responses made by 3- and 4-year-olds and 4- and 5-year-olds. The sense maintained responses made up the highest percentage of responses (410; 47%), followed by incomplete responses (305; 35%) and then sense changed responses (152; 18%). These findings contrast with Experiment 1, in which sense changed responses made up the highest percentage of non-target responses. The different types of non-target responses did not vary significantly by experimental conditions, although 3- and 4-year-olds made a substantially greater number of incomplete responses than 4- and 5-year-olds (220; 42% vs. 85; 25% by age group, respectively). To further examine non-target response type, we calculated the percentage of sense maintained responses that involved a change to connective, order, or position. As noted in Experiment 1, these response types do not add up to 100% exactly because more than one non-target change can be used in a single non-target response. Change in connective was the most common type of sense maintained response (313; 76% of sense maintained responses). In addition, position (237; 58%) and order (256; 62%) changes were also evident in more than half of the total responses that maintained the sentence meaning. Of the 313 sense maintained responses involving a change of connective, only 129 were a change to the connective that involved the replacement of before for after or that of after for before. Therefore, unlike Experiment 1, there were too few responses of this type for further analysis.","The elicited production task complements the sentence repetition task used in Experiment 1, yielding complete two-clause sentences linked by an appropriate temporal connective from young children. As in Experiment 1, there was a main effect of connective because after was more difficult than before in general. In addition, there was a main effect of order because reverse order sentences were more difficult than chronological order sentences. Also replicating Experiment 1 was the finding that children were least accurate when instructed to produce reverse order sentences linked by the connective after. That is, we again found that difficulty with reverse order sentences was modulated by connective; the effect was limited to after–reverse sentences. A critical difference between the two experiments was that production of after–reverse sentences was not modulated by children’s working memory capacity in Experiment 2 when vocabulary was entered into our statistical model. Together, these findings suggest that language knowledge, rather than memory, is the stronger determiner of accurate sentence production. Our analysis of non-target responses revealed a lower proportion of these involving a change of sense compared with Experiment 1. It is important to note that there were few incomplete responses made by 5- and 6-year-olds (30; 5% of all responses) and 4- and 5-year-olds (85; 12% of all responses), although a third of 3- and 4-year-olds’ responses were incomplete (220; 34% of all responses). This highlights the utility of the blocked elicitation paradigm to restrict speaker use to target sentence structures (Ambridge & Lieven, 2011).","These two experiments demonstrate that young children have difficulties in producing two- clause sentences containing before and after during the developmental period that follows their emergence in spontaneous speech. In both experiments, children up to 6 years of age had particular difficulties in producing reverse order sentences linked by the connective after. Clear developmental improvements were evident within this age range. Our experiments advance our understanding of the factors that influence young children’s sentence production, demonstrating that memory capacity-constrained accounts of sentence processing need to factor in the influence of language knowledge rather than attribute difficulties to limited working memory capacity per se. We also demonstrated that investigations that include more than a single paradigm are important to yield robust conclusions in the study of children’s language production. We tested two memory-based accounts for why some sentence structures are more difficult than others: a traditional memory capacity-constrained account (e.g., Carpenter et al., 1994) and a more nuanced language-based perspective of working memory (e.g., MacDonald, 2016). Both accounts predicted that reverse order sentences would be produced less accurately than chronological order sentences in general. Our findings support this prediction, with main effects of order evident in both experiments. According to the memory capacity-constrained account (Carpenter et al., 1994), this effect arises because these sentences require more units of information to be held active in working memory. Additional support for this account was evident; higher working memory capacity predicted better overall performance in Experiment 1. However, other findings indicate that such an interpretation cannot fully explain our findings for the reasons discussed below. Our results are in line with the language-based account of sentence processing proposed by MacDonald (2016) and others, in which the effects of working memory are not direct but rather the result of its relation with language knowledge. The ability to represent information accurately in short-term memory is a requirement for good performance on a sentence production task. The language- based account proposes that short-term memory performance is influenced by the quality of language knowledge. We found support for this account in several ways. First, the effect of order was not consistent across connective; in both experiments, there was a significant order by connective two-way interaction, which was further qualified by age in a three-way interaction in Experiment 2. Critically, in both experiments, sentences with the connective after were less likely to be produced accurately in the reverse order condition than in the chronological order condition; this effect was not found for sentences with the connective before. These findings indicate a role for language over and above any memory effects. In addition, an independent measure of language ability explained performance over and above our independent measure of memory. Third, an inferential analysis of non-target responses in Experiment 1 indicated that a weak representation of our target connectives (measured by connective change responses) does not provide the support needed for the planning and production of reverse order sentences, and descriptive statistics for the sentence constructions showed that this influence was most pronounced with target reverse order sentences linked by after. Note that these findings together indicate that the influence of the connective is not explained by features of the language unit placing additional load on working memory per se; rather, the findings are in line with a language-based account proposal that more processing resources are required to retrieve the context-relevant meaning of after compared with before, so there are fewer processing resources available to accurately represent reverse order sentences in active memory. It is also worth noting that if children had displayed low accuracy for before sentences in the reverse order condition relative to the chronological order condition, such a finding would not necessarily have opposed the proposal by a language-based account that accurate production is influenced by the quality of language knowledge (e.g., MacDonald, 2016). Specifically, even if children have a robust representation of before, planning and production of reverse order sentences linked by before can be disrupted by a weak lexical representation for other words in the sentence. Although not significant, descriptive statistics indicated that accuracy levels for chronological and reverse order sentences linked by before were equivalent for Experiment 1 but not for Experiment 2 (lower accuracy for before–reverse). Unlike Experiment 1, Experiment 2 did not provide children with the target words prior to their task to produce a target sentence. Only Experiment 2 reported that vocabulary had a greater influence than memory, which suggests that before–reverse sentences could be more difficult to produce when the task provides less support for words in the sentence. This in line with the proposal that a robust representation of words in the sentence frees up processing resources for the accurate representation of reverse order sentences in active memory. Our sentence constructions were counterbalanced across conditions, and the effects were specific to after, which has a less consistent form–meaning relationship relative to before. Thus, we conclude that the accurate production of reverse sentences linked by after is more likely to be influenced by a weak representation for the target connective rather than representing more general language effects. Although difficulty with after–reverse sentences was replicated across both experiments, there are at least two reasons to remain cautious about accepting a language-based explanation as the sole reason for young children’s difficulties in producing multiple-event sentences. First, we must consider the possibility that order effects are modulated by a confounding variable, connective position, rather than connective. Second, we must address why the stronger influence of vocabulary over memory (determined by examining model fit) was apparent only in Experiment 2. These limitations are considered in turn below. A natural consequence of our design was that the interaction between connective and order was influenced by connective position because this also differs across sentence structures. For example, after is used in a sentence-initial position when the order of events is presented chronologically (“After he put on the socks, he ate the burger”) but is used in a sentence-medial position when events are presented in reverse order (“He ate the burger, after he put on the socks”). The reverse applies to before sentences. Thus, an alternative explanation for a specific difficulty with reverse order sentences is that the position of the connective modulates the effects of order. That is, a reverse order sentence in which the temporal sequence is cued by before may be easier to represent than its after counterpart because the initial position of the connective signals from the beginning that events will be narrated in a reverse order. This viewpoint is supported by evidence that speakers have cognitive biases to highlight certain referents at the beginning of the sentence, in our case the temporal connective, that act as cues to reduce ambiguity for the listener (e.g., Chafe, 1984; Grice, 1975; Myachykov, Garrod, & Scheepers, 2012; Silva, 1991). Conversely, reverse order sentences that contain after may be more difficult to plan and narrate because the critical information about event order is provided midway through the sentence, which may place greater demands on working memory. We believe that this account—that connective position rather than connective itself modulates order effects—does not adequately explain our pattern of findings. If position accounts for our results, the difficulty with after–reverse sentences would arise because the late signaling of reverse order places greater demands on memory capacity than early signaling. However, our independent measure of memory was a weaker predictor of performance than our independent measure of vocabulary. Moreover, as cited in the Introduction, corpus work suggests that speakers have a preference for relating information using the connective in a medial position (Diessel, 2004, 2008). Clearly, more experimental work is needed to investigate the role of connective position in sentence production. The second reason for caution in accepting a language-based account over a memory capacity-constrained account was that both memory and vocabulary improved model fit in Experiment 1. That is, stronger memory and vocabulary both were associated with more accurate performance, and our independent measure of vocabulary did not explain unique variance in children’s specific difficulty with after–reverse sentences (Experiment 1). The greater influence of vocabulary over memory in Experiment 2 compared with Experiment 1 may have arisen due to the task demands. Participants in the sentence repetition task used in Experiment 1 were provided with the language form in their input, whereas participants in the elicited production task used in Experiment 2 needed to use their language knowledge to specify every level of detail of the form themselves (i.e., syntactic, morphological, phonological, articulatory) so that it could be mapped onto the intended meaning (see Garrett, 1980; Gennari & MacDonald, 2009; Vigliocco & Hartsuiker, 2002). Therefore, there may be greater demands on language knowledge retrieval processes in the production task used in Experiment 2, where children were not first provided with the input to repeat. The above explanation may help to understand why the pattern of findings across sentence structures in these production experiments differs from that reported in recent work examining comprehension of the same sentences (Blything & Cain, 2016; Blything et al., 2015). Blything et al. (2015) found that reverse order sentences that contained after were the most difficult to comprehend, the same pattern reported here for production. However, in contrast to the findings of these production experiments, after–reverse sentences were not statistically more difficult to comprehend; instead, an advantage for before–chronological sentences drove the effect. Furthermore, in both of the previous comprehension studies, an independent measure of working memory accounted for significant variance in performance, whereas an independent measure of vocabulary did not. This is in contrast to the current findings; our replication across two production studies of difficulty with after–reverse sentences, in addition to stronger effects of an independent measure of vocabulary than of memory capacity, suggests that a different explanation is required for production than that used for comprehension. The difference between the findings of the current production study and previous comprehension studies can be explained by how comprehension and production draw on memory and language. Comprehension tasks provide participants with the language form in their input in the same way as described earlier for sentence repetition tasks. Therefore, differences across the domains might be explained in the same way as was proposed above for why the current study provides greater support for the language-based account in the more pure production task (Experiment 2) relative to the production task that carried a comprehension component (Experiment 1). Limitations, implications, and future research ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A strength of this work is the replication of the main finding across two different tasks; children up to 6 years of age had difficulties in producing two-clause reverse order sentences linked by the connective after. However, the analysis of non-target responses highlighted differences in the nature of our two experiments, which we believe is informative for researchers considering a marriage of the two paradigms. First, incomplete responses comprised 35% of the non-target responses in the elicited production task (Experiment 2) compared with only 16% in the sentence repetition task (Experiment 1). This may be due to the scaffolding provided by the initial input in a sentence repetition task, which supports children to produce the target response. Another notable difference between the experiments was that sense maintained responses made up the largest percentage of non- target responses in elicited production (Experiment 2), whereas sense changed responses comprised a large percentage of non-target responses in the sentence repetition task (Experiment 1). In sense-maintained responses, children produced a temporal connective as a linguistic device to successfully communicate order but did not use the target structure. This indicates that the elicited production paradigm, which resulted in a high number of sense maintained responses, is more likely to result in children reverting back to a sentence structure with which they are familiar when required to signal temporal order with a connective. Overall, the difference in non-target response types, along with the differences in the nature of the tasks themselves, illustrates that investigations that include more than a single paradigm are important to yield robust conclusions in the study of children’s language production. Age differences in Experiment 2 persisted when memory and vocabulary were incorporated into the model. Given the high accuracy shown by 5- and 6-year-olds in Experiment 1, other experimental methods are required to study and understand better developmental and individual differences in the planning and production of complex sentences such as these. Habets et al. (2008) successfully used event-related potentials (ERPs) to study processing differences in the production of chronological and reverse order temporal sentences in adults. Such techniques might be adapted for use with children. In addition to using different experimental paradigms to assess the time course and difficulty of production, a more comprehensive battery of tasks could be used to measure the constructs of both working memory and language. Ideally, working memory tasks should measure the storage and manipulation of information to tap these two critical functions of working memory (Oberauer & Lewandowsky, 2010). However, as noted, 5-year-olds find such complex span tasks hard to perform (Gathercole et al., 2004). In addition, memory tasks with a low semantic load should be used to determine the relationship between sentence processing and working memory capacity, distinct from language knowledge (Kidd, 2013). Digit-based tasks as used here (forward digit recall) can be advantageous in this respect because they have a lower semantic load and so are less strongly related to independent measures of language (Cain, 2006; Seigneuric, Ehrlich, Oakhill, & Yuill, 2000; but see Jones & Macken, 2015). Similarly, additional measures of vocabulary as well as tests of grammatical knowledge could be included to more fully assess the construct of language and children’s knowledge of cohesive devices such as connectives (Language and Reading Research Consortium, 2015; see also Cain & Nash, 2011, for work with older children demonstrating differences between connective knowledge and use to at least 10 years of age). Note, however, that the measures used in the current study were predictive of performance, and these suggestions do not undermine the current findings; rather they offer ways to develop a more fine-grained picture of the influence of memory and language knowledge on the production of complex sentences. A critical implication is that a memory capacity-constrained account of sentence processing (Carpenter et al., 1994) is likely too simplistic on its own, and we need to factor in the influence of the specificity or distinctness of retrieval cues (i.e., language knowledge). Converging evidence for this viewpoint has been provided in studies of adult language production (Gennari et al., 2012; Montag & MacDonald, 2014, 2015; Smith & Wheeldon, 2004) and comprehension (for a review, see Van Dyke & Shankweiler, 2012). A next question for the language-based account is how language knowledge becomes sufficiently consolidated (precise and robust) to support the comprehension of complex sentences. A straightforward assumption from a developmental perspective is that language representations become stronger through exposure to the language. Thus, differences between vocabulary items, such as before and after, may be due to differences in the frequency of their occurrence in language (Wells, Christiansen, Race, Acheson, & MacDonald, 2009). To explore this possibility, we coded 100 randomly selected occurrences of before and after from the CHILDES Thomas corpus (age range = 2;7.2 [years;months.days] to 4;11.20; Lieven, Salomo, & Tomasello, 2009), which is a corpus of child-directed speech. Only 47 (of the 100) instances of before and after used these terms as a temporal connective within a multi-clause sentence. Of these, there was not a clear bias for either chronological order or before sentences; there were 17 before–chronological, 4 before–reverse, 10 after–chronological, and 16 after–reverse sentences. This does not provide any evidence that children had less exposure to the more difficult (after–reverse) structure in this study. However, the corpus analysis did show that, of the other 53 occurrences of before and after, 47 were of after being used as a non-connective (e.g., “Every now and then we look after the baby next door”). This finding supports an alternative account that the apparent difficulties with after sentences arise because after is used less consistently as a connective than before. This is consistent with the British National Corpus (Leech et al., 2001). Longitudinal work combining corpus and experimental methodologies could test this hypothesis further. A final thought for future research is to what extent production accuracy might be enhanced when the sequence of events can be informed by world knowledge. Theoretical models of mental representations of text and discourse (e.g., Zwaan & Radvansky, 1998) suggest that it should be easier to plan and produce sentences when the events follow a typical sequence because world knowledge can inform the order in which the events should be mentally represented, for example, that socks are typically put on prior to putting on shoes. World knowledge- present sentences (e.g., “He put on the socks, before he put on the shoes”) served as fillers to scaffold the structure of the sentence and so were not part of our experimental design per se. Nevertheless, children’s overall performance was consistent with previous findings in children’s comprehension of two-clause sentences containing before and after (Blything et al., 2015); the filler world knowledge-present sentences were not performed significantly better than the test sentences in which event order was arbitrary (e.g., “He put on the socks, before he ate the burger”). Thus, at least for these very simple two- clause sentences, world knowledge does not appear to play a significant role in language production. For more complex language, such as longer texts that require greater processing resources to integrate information across several sentences, world knowledge may have a more powerful influence on performance (Pratt, Tunmer, & Nesdale, 1989). In conclusion, 3- to 6-year-olds demonstrated an ability to accurately use before and after as temporal connectives in the production of two-clause sentences, but they notably found it difficult to produce reverse order sentences that were linked by after. We did not find unequivocal support for either the memory capacity-constrained account or the language- based account, although our findings lend greater support to the latter. These two apparently contrasting accounts have a common core; they seek to explain why memory limitations affect sentence processing. Further experimental work is needed to understand how memory and language knowledge individually and together influence sentence planning and production and to elucidate the commonalities and differences in their influence on performance in language production and comprehension tasks."],["Annually, nine million people die due to environmental pollution (Landrigan et al., 2017). Unsafe sanitation, and more specifically open defecation, is one of the main causes, leading to fecal contamination of water bodies and the transmission of fecal bacteria (Prüss-Ustün et al., 2014). In 2015, 892 million people still practiced open defecation, with rates being highest in Sub-Saharan Africa (WHO & UNICEF, 2017). In Ghana, where this study is located, 31% of the rural population practiced open defecation in 2015 (WHO & UNICEF, 2017). A recent systematic review found that increasing access to safe sanitation services can reduce diarrheal diseases by 16% (Wolf et al., 2014). However, a single individual or household, by stopping open defecation, can only marginally reduce their diarrheal risk related to a fecal polluted environment (Jung, Hum, Lou, & Cheng, 2017). Research has shown that at least 75% of all households must stop open defecation to achieve a hygienically safe environment that benefits all (Clasen, Boisson et al., 2014; Jung et al., 2017; Wolf, Hunter et al., 2018). Open defecation is thus not only an individual but a collective health hazard (Geruso & Spears, 2018; Vyas, Kov, Smets, & Spears, 2016). This is comparable to other environmental challenges, such as greenhouse gas emissions, which can only be confronted if most of the population show climate- protective behavior such as reduction of individual energy consumption. Activating social norms1 supporting pro-environmental behaviours helps people to act pro-environmentally (Bamberg & Möser, 2007; Steg & Vlek, 2009), such as avoiding littering in public places (Cialdini, Reno, & Kallgren, 1990), conserving household energy (Schultz, Nolan, Cialdini, Goldstein, & Griskevicius, 2007) or using safe water sources sustainably (Contzen & Marks, 2018). Similarly, activating social norms has been used in the context of sanitation (Dooley, Maule, & Gnilo, 2016). It is a key element of the behavior change campaign Community-Led Total Sanitation (CLTS), which has been shown to successfully reduce open defecation by up to 33% (Pickering, Djebbari, Lopez, Coulibaly, & Alzua, 2015; Venkataramanan, Crocker, Karon, & Bartram, 2018). For Ghana, case studies on CLTS report success rates of up to 26% reduction in open defecation (Crocker et al., 2016, 2017) and scientific as well as political interest on CLTS and sanitation outcomes is steadily increasing for the Ghanaian context (Berendes et al., 2018; Nunbogu, Harter, & Mosler, 2019). CLTS consists of a set of community-based, participatory activities, and explicitly focuses on evoking a shift towards a new social norm opposing open defecation. The influence of CLTS on social norms and thus the effect on latrine construction has already been demonstrated in research (Alemu, Kumie, Medhin, & Gasana, 2018; Harter, Mosch, & Mosler, 2018) and the consideration of social norms for the success of CLTS is gaining more attention (Dooley et al., 2016, p. 299; Novotný, Kolomazníková, & Humňalová, 2017; Venkataramanan et al., 2018). This is also true for the Ghanaian context (Osumanu, Kosoe, & Ategeeng, 2019). Because of its success in stopping open defecation, CLTS is the most widely applied sanitation campaign to date (Bongartz, Vernon, & Fox, 2016; USAID, 2018). While randomized trials have shown that CLTS reduces open defecation compared to controls, these effects are highly heterogeneous (Harter, Inauen, & Mosler, n.d.). This means that despite the general success of CLTS, open defecation rates remain high in some communities, and the threshold of 75% households using latrines is often not reached (Crocker et al., 2016; Pickering et al., 2015; Venkataramanan et al., 2018). This indicates that inter-community differences may moderate CLTS effectiveness. A moderator that might be at play here is social identification, defined as an individual's understanding to belong to a social group and to emotionally value the membership (Abrams & Hogg, 1990; Reynolds, Subašić, & Tindall, 2015; Tajfel, 1978). Previous research has shown that social norms particularly affect behavior in individuals strongly identified with the social group in question (e.g. (Terry, Hogg, & White, 1999; White, Smith, Terry, Greenslade, & McKimmie, 2009). One potential explanation for this effect is that strongly identified people want to be accepted and approved by their group, and may thus be eager to conform with the group's expectations, independent of whether they agree with a specific social norm or not (Abrams & Hogg, 1990; Deutsch & Gerard, 1955). Regarding CLTS, households may construct and use a latrine not because they are convinced of it, but simply because they want to be accepted in the community and therefore conform to the newly established social norm. The social identity perspective, however, proposes an alternative explanation (Tajfel & Turner, 1979; Turner, Hogg, Oakes, Reicher, & Wetherell, 1987). Self-categorization as a group member (i.e. the definition of the self in-group terms and in connection to other group members) includes a merging between group and individual; group goals become personal goals and group norms become personal norms. Accordingly, strongly identified members act in line with group norms not only because they want to conform but more so because they perceive the norm (e.g. of constructing and using latrines) as their personal norm, as their right way (Abrams & Hogg, 1990; Deutsch & Gerard, 1955). We therefore expect that CLTS will be especially successful in reducing open defecation in communities with stronger social identification prior to CLTS implementation because people will more readily follow the newly established social norm to stop open defecation. At the individual level, we expect that people, who feel a stronger social identification than other community members, will be more likely to stop open defecation. To test our assumptions, we conducted a cluster-randomized, controlled trial, which is outlined in the following (WHO & UNICEF, 2017).","For this cluster-randomized, controlled trial, CLTS was implemented in four intervention arms and its effects on open defecation reduction were tested and compared to a control arm.2 Social identification prior to the intervention was tested as a moderator of CLTS effectiveness.","We conducted this trial in the Northern Region of Ghana in two rural districts. In both districts we collected baseline data in February to March 2016 (for more information on the baseline survey, refer to Harter et al. (nd)). Afterwards, Global Communities, a local non-governmental organization, implemented CLTS in communities from both districts from July to November 2016.3 This article presents data from the long-term follow-up that was realized 14–16 months after implementation of CLTS, namely in February to March 2018 in both districts. The ethical board of the University of Zurich, Switzerland and the Ethical Review Committee of the Ghana Health Service (GHS-ERC: 05/01/2016) approved this trial. Study site and clusters ~~~~~~~~~~~~~~~~~~~~~~~ The study was realized in collaboration with Global Communities and local government representatives. Global Communities selected the two districts in the Northern Region of Ghana, i.e. Bole and Sawla-Tuna-Kalba, because no CLTS campaign had been implemented there before. The local government representatives selected 132 communities within the two districts according to two eligibility criteria: accessibility (by car or motorbike due to practical reasons) and community size (minimum community size of 25 households). We grouped the communities of both districts into 25 regionally separate clusters to avoid spillover of intervention effects between close communities, and randomly allocated them to the four intervention arms (five clusters per intervention arm) and the control arm (five clusters). Study participants ~~~~~~~~~~~~~~~~~~ Trained data collectors selected study participants in the communities following the random route method (Hoffmeyer-Zlotnik, 2003, pp. 205–217). Data collectors were instructed to start from a central point of the community and interview every third household in an assigned area of the community. If no one or no eligible person was at home or if the household did not want to participate, data collectors selected the next following household. Household members were eligible if aged 18 or older and stable inhabitants of the community. If more than one household member was eligible, the participating member was selected according to their availability. We equally considered men and women, as both might take important decisions for latrine construction. Every participant gave informed written consent to participate in the study. The sample size was calculated a priori for a cluster-randomized trial with repeated measures and a dichotomous primary outcome (Spybrook et al., 2011). Assuming an intra-cluster correlation of ρ = 0.2, 80% power, 5% α-error probability, and 20% dropout, we estimated a required sample size of 3,215 households nested in 132 communities (approx. 25 households in each) to detect a medium effect of the intervention on open defecation. For a detailed description of the sample size calculation, please refer to Harter et al. (nd). Fig. 1 displays the flow of participants through the trial. Interventions ~~~~~~~~~~~~~ Global Communities developed intervention protocols for CLTS based on the Handbook on CLTS (Kar & Chambers, 2008). Local facilitators implemented it in three phases. First was an informative phase, where facilitators visited the community and collected information on the composition of the community and the baseline behavior. A date for a community meeting was agreed and all inhabitants were invited. The community meeting, also called triggering event, formed the second phase of CLTS. During this meeting, the facilitators motivated community members to draw a map of their community on the ground and to indicate their houses as well as the spots they used for open defecation on the map. Through asking questions about possible ways of fecal-oral transmission of pathogens, the inhabitants were expected to recognize the hygienic problems connected to open defecation. The facilitators further identified emerging leaders during the triggering event and invited them to serve as role models and to support others in the process of latrine construction. A community action plan and a date on which the community wanted to be open defecation free (ODF) was agreed. In the end of the triggering event, the facilitators explained the first step of a latrine construction, namely digging the pit and gave further information on the construction process, such as which material to use. No financial support was given to community members (including emerging leaders), however, construction materials were provided at wholesale price instead of retail prices. The third phase of CLTS included follow-up visits in the weeks after the triggering event until the community reached the status ODF, defined as at least 80% latrine coverage. During the follow-up visits, facilitators addressed any arising problems and questions regarding latrine construction. CLTS was implemented in all four intervention arms. For three of the intervention arms additional campaign activities were developed and implemented based on the Risk, Attitudes, Norms, Abilities and Self-regulation (RANAS) approach (for detailed description of implemented interventions and outcomes please refer to the intervention manual4 and Harter et al. (nd)). They included a household action plan and a public commitment for latrine construction. The control arm did not receive any intervention during the research phase but CLTS was implemented after the trial. In intervention communities, 72.8% (n = 1540) of the households attended the CLTS event. Data collection and outcome measures ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A team of 33 local data collectors assessed outcome variables at baseline and both follow- ups. The first author, together with local personnel trained the team in a 1-week training before each of the three data collection phases. The trainings included a detailed discussion of questionnaire items and explained the correct usage of instruments and interview techniques. These were then rehearsed in role-plays. The questionnaire was translated into seven local languages as part of the data collector training of the baseline data collection, and pretested in two days and 66 interviews prior to each data collection in the field. Every interview was supervised (by research managers, interns, master students and local field supervisors), and lasted 50 min on average. Interviews included self-reported behavioral measurements, social identification, and further items on psychosocial determinants of behavior (not relevant to the present paper, for information refer to Harter et al. (nd)). The self-reported open defecation rates at long- term follow-up were assessed with items based on the Safe San Index (Jenkins, Freeman, & Routray, 2014). Six items assessed the self-reported open defecation rate of each individual during the last week. The original Safe San Index includes information about all household members, whereas for this article only individual self-reported behavior at long-term follow-up was considered. Three items asked for the respondent's open defecation frequency in the mornings, middays and evenings/nights of the last week and three items asked the same for latrine use (items displayed in Supporting Information in SI Table 1). The Safe San Index represents the proportion of safely managed feces relative to total defecation instances, resulting in a range of 0–1. However, the data revealed that individuals either exclusively practiced open defecation or used a latrine. This resulted in a binary outcome variable with 0 = no open defecation and 1 = open defecation. Aggregated to community level it accounts for a communities' average open defecation rate, the proportion of people within a community who reported to practice open defecation (0–100%). Social identification at baseline was measured as identification with the community on three dimensions: in-group ties, in-group affect and centrality, following items proposed by Cameron (2004). The selection of two items per dimension for this research was done in accordance with local partners, based on cultural and language considerations. Items were framed as statements with a five-point Likert-type scale for agreement. We used a visual scale with five black dots (in ascending order relative to their size) to help respondents choose one of the answer options. The data collector read out every answer option to the respondent and pointed it out on the visual scale. To test the item factor structure, we conducted an exploratory factor analysis with Principal Components Analysis and Varimax rotation with Kaiser Normalization (Field, 2009) (correlations displayed in SI Table 2 in supporting information). The factor analysis was not able to replicate the dimensions proposed by Cameron (2004), but resulted in one factor for social identification with the items of the two dimensions in-group affect and centrality loading on the factor. Whereas the items of the dimension social ties did not load on it and were therefore excluded. The remaining four items were aggregated to one scale (M = 4.29, SD = 0.30, Cronbach's α = 0.64). Table 1 displays the four items of the scale and according descriptive measures, correlations and intra-class correlation. Aggregated at the community level, it resembles a community's average social identification. Analyses ~~~~~~~~ To test the moderating influence of social identification on the effect of CLTS on open defecation, we fitted a Generalized Estimating Equation (GEE) (see Zeger and Liang (1986, pp. 121–130); Zeger, Liang, and Albert (1988)) using IBM SPSS Statistics for Windows, version 24 (IBM Corp., Armonk, N.Y., USA). The model was set up using binomial distribution with logit link (Homish, Edwards, Eiden, & Leonard, 2010), because the outcome was binary. We used an exchangeable correlation structure, which assumes constant intra-cluster dependency (used for clustered data not assessed in a time-series, see Ballinger (2004)). This model accounted for the nested structure of our data with households nested in communities and further allowed the inclusion of a binary outcome (0 = no open defecation vs. 1 = open defecation). The CLTS intervention (0 = control arm; 1 = intervention arms) was entered together with the community-averaged social identification (grand-mean centered), and the individual's deviation from their community's average social identification (group-mean centering). Thereby, we were able to distinguish between community-level and individual-level effects, which may differ (Hamaker, 2012). We further added the interaction terms of the intervention with social identification at both, individual and community level, to test whether social identification at baseline moderated the intervention effect on reported open defecation at follow-up. As effect size measures, we calculated odds ratios (ORs) with asymptotic Wald 95% confidence intervals (CIs). ORs can be interpreted as increased (OR>1) or decreased (OR<1) odds of practicing open defecation for a unit increase in the predictor. Sample description ~~~~~~~~~~~~~~~~~~ The respondents were on average 44.5 years old (SD = 16.1). Slightly fewer than half were female (42%) and 21% were able to read and write. The households consisted of eight members on average (SD = 5). In terms of religion, 26% named Islam as their religion, 49% Christianity, 19% traditional religions, and 5% mentioned to be atheists. Most of the sample reported to be farmers (80.4%) with an average monthly household income of 202 Ghanaian New Cedi (SD = 380), equivalent to 42 USD. The households of the sample therefore lay on average below the poverty line proposed by the World Bank of 57 USD per individual per month (Atkinson, 2017). Randomization check and dropout analysis ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 2 shows baseline characteristics for intervention and control arms. Chi-Square tests and variance analysis revealed that the groups significantly differed on all characteristics except for age, household size and number of dropouts, which were equally distributed. At baseline, 89.9% of the control and 97.2% of the intervention arm reported to practice open defecation. The main analyses reported in this paper were rerun and characteristics were included that had shown significant differences between intervention and control group at baseline. Even though the effect sizes were small (Cohen, 1992; Ferguson, 2009; Trusty, Thompson, & Petrocelli, 2004), those characteristics were included in sensitivity analyses as they were considered to be potential confounding variables. Furthermore, we compared respondents who participated in both panel surveys (n = 2,607) to respondents only participating in the baseline survey (dropouts, n = 609, 18.9%) on the same characteristics. Chi-square tests and variance analyses showed that the study dropouts were significantly less socially identified with their community, less likely to be farmers, had a higher probability for literacy, were significantly younger and had a higher income compared to analyzed participants. Open defecation rates were not significantly different between study dropouts and participants remaining in the sample (see SI Table 3 in Supporting Information). Intervention effects on open defecation and the influence of social identification ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the CLTS intervention arms, 46.4% (SD = 49.9%) of the individuals reported to practice open defecation at follow-up, compared to 88.4% (SD = 32.0%) in the control arm. As indicated by the GEE model results (see Table 3), the OR for the intervention group indicates that intervention participants were 11 times less likely to practice open defecation at follow-up than controls (Β[SE] = −2.42 [0.33], OR = 0.09, p < 0.001). Fig. 2 shows the community-averaged open defecation rate in control and intervention arms moderated by community's average social identification. In line with our hypothesis, CLTS intervention communities with stronger community-averaged social identification reported less open defecation at follow-up than those with lower community-averaged social identification (Β[SE] = −11.70 [4.15], p = 0.005). The control arm showed opposite effects: communities with stronger community-averaged social identification reported higher open defecation rates at follow-up than those with lower community-averaged social identification (Β[SE] = 7.06 [2.28], p = 0.002). In both, the control and intervention arms, the effects of individuals' social identification pointed in the same direction as the community-averaged social identification, but were not significant (control arm: Β[SE] = 0.25 [0.24], p = 0.305; intervention arm: Β[SE] = −0.65 [1.05], p = 0.534). Sensitivity analyses revealed that including the baseline characteristics, and adjusting for baseline behavior, did not substantively change the findings. Only age and literacy had significant but small reducing effects on open defecation (age: Β[SE] = −0.01 [<0.01], p = 0.002; literacy: Β[SE] = −0.27 [0.10], p = 0.008).","This study corroborated previous findings that CLTS is an effective intervention to reduce open defecation. In our sample, at the long-term follow-up, 53.6% of individuals in the intervention arms did not defecate in the open anymore. While this rate is still behind the threshold of 75% of all community members that would need to stop open defecation to reach an incremental health benefit at community level (Jung et al., 2017; Wolf et al., 2018), it is comparable to most randomized trials of CLTS. A recent review on CLTS reports that the majority of interventions achieve around 50–80% rates of stopping open defecation (USAID, 2018). That the reported rate in our study is at the lower end, can partly be explained by the short time elapsed between intervention and follow-up survey. At the time of the survey, many latrines (61.6%) were still under construction and therefore still not in use. In the intervention arm, the figure of 46.4% of the respondents that reported to practice open defecation, might likely decrease as soon as the construction process of the remaining latrines is completed. We expect this, because in our sample the clear majority of households that owned a completed latrine also used it (93.6%). This is surprising when compared to previous research on latrine ownership and use, for example from India, where it was found that only 47% of the owned latrines were actually used (Barnard et al., 2013). More importantly, this study showed for the first time that social identification within communities moderates the effectiveness of CLTS on open defecation. Specifically, our hypothesis regarding the influence of social identification on the intervention effects was supported: CLTS was more successful in communities with stronger social identification prior to the intervention. In communities with stronger community-averaged social identification, 38.7% of respondents reported to practice open defecation, compared to 54.7% in communities with lower average social identification. Our findings extend previous findings of a randomized trial on CLTS in Indonesia on the importance of communities’ pre-existing social conditions for intervention success (Cameron, Olivia, & Shah, 2015). The researchers were able to show in a randomized trial, that communities with higher initial social capital, i.e., higher trust and cohesion, were more likely to have higher latrine coverages. We suppose that the reported moderating effect of social identification on CLTS effectiveness work through an increase in new social norms that oppose open defecation. In communities in which individuals strongly identify with their community, individuals wish to conform to the new norm, which will lead to better CLTS outcomes (e.g. Terry et al. (1999) and White et al. (2009)). Alternatively, and based on social identity theory, strongly identified individuals might not only conform to the social norms, but internalize them as their own goal and their personal norms (Abrams & Hogg, 1990; Deutsch & Gerard, 1955). Future research can test these assumed mechanisms of CLTS and disentangle whether CLTS truly evokes a shift in social norms and whether these translate, moderated by social identification, into personal norms. Interestingly, our data showed opposite effects in control communities: higher open defecation rates were reported in communities with stronger compared to communities with weaker average social identification. This might be because in communities without CLTS intervention the prevailing social norms were supporting open defecation, as no impulse of change had occurred. This finding supports social identity theory; communities with stronger community-averaged social identification follow, or better said incorporate, the prevailing social norms, whether the social norms suggest stopping open defecation – as in intervention communities – or the opposite – as in control communities (Abrams & Hogg, 1990; Cialdini et al., 1990). Schultz et al. (2007) described this effect of salient norms that lead to an undesired behavior as “the destructive potential of social norms” (p.431). A departure from prevailing social norms, such as stopping open defecation when open defecation is what the rest of the community members are doing, may only be possible for community contexts where social identification is weak, i.e., where community members do not define themselves through their community and are thus less inclined to follow the prevailing social norms (Abrams & Hogg, 1990; Deutsch & Gerard, 1955). Finally, our results did not suggest any additional effects of social identification at the individual level, over and above the community-level effects. It seems that the moderating effect of social identification is a truly community-based phenomenon. To sum up, our results highlight the importance of social identification especially for collective environmental challenges, such as open defecation. This is in line with the increasing interest of the CLTS community to consider the social context in CLTS planning and implementation (Dooley et al., 2016, p. 299; Novotný et al., 2017; USAID, 2018). For the implementation practice, this means that communities with strong social identification provide a fertile ground for CLTS implementation. To improve CLTS planning, we therefore suggest the assessment of social identification in a first step. If social identification is found to be weak, activities should be carried out to foster social identification prior to CLTS implementation, as has been recommended for the field of collective action (Van Zomeren, Postmes, & Spears, 2008). Such activities might include enabling interaction between community members (Jans, Leach, Garcia, & Postmes, 2015) or directing attention to neighboring communities that have already eliminated open defecation, for example forming a competition-like situation and pointing out the differences to an out-group (Jans, Bouman, & Fielding, 2018; Tajfel & Forgas, 2000, pp. 49–63). In cases where social identification cannot be strengthened before a CLTS implementation, by-laws or sanctions for people not following the norms might be enforced, which is proposed by the CLTS Handbook (Kar & Chambers, 2008) and in social psychology literature to solve social dilemma situations (e.g. (De Cremer, Hoogervorst, & Desmet, 2012)). The control group findings imply that communities with strong social identification are potentially at risk of increasing or reinforcing open defecation practices. These communities should therefore be selected with high priority for sanitation interventions to avoid such tendencies, and leverage promising responses to interventions due to strong identification. Strengths and limitations ~~~~~~~~~~~~~~~~~~~~~~~~~ To the best of our knowledge, this study is the first that investigated the influence of social identification on the effect of CLTS on open defecation. It was fully powered with a sample of 2,606 households for the longtime follow-up survey and 3,216 households in the baseline. With 132 communities, it allowed analysis at the community level, and the investigation of the deviation of individuals from community means. CLTS was implemented under real conditions in rural Ghana in a variety of local contexts, such as different community sizes and ethnical compositions. This allows assuming high external validity. The study, however, has the following limitations. The first relates to the causal relationship of social identification moderating CLTS intervention effects. Because we did not experimentally manipulate social identification, the found moderating effect could be attributable to other influencing factors, for example community size or heterogeneity within communities. Future research should manipulate the strength of social identification to provide further evidence for the presented moderating effects in this article. Open defecation was assessed through self-reports. However, strengthening the validity of the self-report measure, the use of latrines was verified by observation of enumerators,5 which correlated strongly with the self-reported behavior (r2 = 0.72, p < 0.001). The scope of this study did not allow for assessment of open defecation rates more than one year, including participants that might have reverted to open defecation. Long- term change should be included in future research. Social identification was measured using six items only and for analysis, only four items were included. However, the scale showed relatively low reliability (Cronbach's α = 0.64). Furthermore, the six items used for the social identification scale applied in this article were not able to replicate the three dimensions of social identification postulated by Cameron (2004). The reason for this may be that only two items per dimension were included to keep the questionnaire as brief as possible to minimize participant burden. Future studies should use more items to allow for a more detailed consideration of social identification dimensions.","This study reports the success of CLTS on reducing open defecation rates and highlights the relevance of including social conditions into planning of sanitation campaigns, such as CLTS. Specifically, the consideration of communities’ social identification is crucial for the success of CLTS on reducing open defecation, as it might be able to intensify the effects of the intervention. We therefore recommend to assess the level of social identification within target communities and plan CLTS interventions accordingly, meaning to strengthen, if needed, the social identification among community members before a sanitation intervention. Further, this is the first time that the concept of social identification was studied in environmental sanitation and points to the potential influence of social identification in other water and sanitation related behaviors in low- and middle-income countries."],["The provision of verbal labels enhances 12-month-old infants’ memory flexibility across a form change in a puppet imitation task (Herbert, 2011), although the mechanisms for this effect remain unclear. Here we investigate whether verbal labels can scaffold flexible memory retrieval when task difficulty increases and consider the mechanism responsible for the effect of language cues on early memory flexibility. Twelve-month-old infants were provided with English, Chinese, or empty language cues during a difficult imitation task, a combined change in the puppet's colour and form at the test (Hayne et al., 1997). Imitation performance by infants in the English language condition only exceeded baseline performance after the 10-min delay. Thus, verbal labels facilitated flexible memory retrieval on this task. There were no correlations between infants’ language comprehension and imitation performance. Thus, it is likely that verbal labels facilitate both attention and categorisation during encoding and retrieval. --------------------------------------------------------------------------------","The ability to flexibly retrieve our memories across a range of situations is an important feature of the declarative memory system. Memory flexibility, or the ability to use pre- existing stores of knowledge to solve new problems, enables us to avoid costly or time consuming re-learning and to benefit from our past experiences. Across the first two years of life, studies using the deferred imitation procedure have demonstrated that the ability to flexibly retrieve memories across changes in the social and physical contexts, or changes in the target stimulus develops gradually (e.g., Hayne, Boniface, & Barr, 2000; Hayne, Barr, & Herbert, 2003; Hayne, MacDonald, & Barr, 1997; Herbert & Hayne, 2000a; Herbert, Gross, & Hayne, 2006; Learmonth, Lamberth, & Rovee-Collier, 2004; Learmonth, Lamberth, & Rovee-Collier, 2005). For example, within the puppet imitation task, 12-month old infants can reproduce the target actions shown on a puppet that differs in colour (Hayne et al., 1997) or form (Jones & Herbert, 2008) from the one present during the original demonstration after a 10 min delay, but not with a puppet that differs in both colour and form, or after a 24 h delay. In contrast, 18-month old infants can reproduce the target actions with a puppet that differs in both colour and form after a 24-h delay but not when the puppets are highly dissimilar (e.g., black/white cow, yellow/orange duck; Hayne et al., 1997). Thus, early memory flexibility appears to be determined in part, by the degree of changes to the central stimulus between learning and retrieval, and the length of the retention interval (also see Herbert & Hayne, 2000a). Language cues have recently been identified as a powerful means by which to enhance flexible memory retrieval at young ages. Being able to spontaneously self-generate a label for a novel stimulus or a hiding location facilitates 2-year-old children’s ability to transfer knowledge to a new situation (Miller & Marcovitch, 2011; Zimmermann et al., 2015), while memory flexibility at younger ages benefits from experimenter-generated labels (Herbert & Hayne, 2000; Herbert, 2011). For example, in Herbert (2011), 12- and 15-month old infants were shown the target actions with the puppet task accompanied by verbal labels for the object and actions (e.g., “Look, a puppet. Off. Shake. On”) or no label during the demonstration. At the test 10 min later, the puppet was labelled again for infants in the verbal label condition only. Infants in both the label and no label conditions reproduced significantly more target actions with a puppet that differed in form (e.g., pale grey mouse at demonstration, pale grey rabbit at test) than baseline performance by infants in the control condition. Thus, infants were already showing some ability to generalise across a change in stimulus form after the short delay. Importantly, however, infants in the label condition reproduced significantly more target actions than infants in the no label condition. In other words, verbal labels enhanced performance above spontaneous flexible memory retrieval. Whether the same effect would be observed when task difficulty increases remains to be determined. Perhaps unsurprisingly, verbal labels do not facilitate infant abilities in all situations, such as when 15-month-olds are asked to transfer knowledge across a complex 2D to 3D action imitation task (Zack, Gerhardstein, Meltzoff, & Barr, 2013). However a unique feature of the puppet task is that task difficulty can be progressively increased by altering the colour, form, or colour and form of the puppet, providing an opportunity to examine the limits of the facilitative effects of language on infant memory abilities. In this study we consider whether verbal labels can facilitate infants’ memory flexibility across a complex change, altering both the form and colour of the stimulus present at retrieval. Although the body of research showing the facilitative effects of language cues on memory flexibility continues to grow, the mechanism for this effect remains unclear. One possibility is that verbal labels may facilitate flexible memory retrieval across changes in the target object by potentially facilitating categorisation of the target and test objects (for similar argument, see Jones & Herbert, 2009). Infant language comprehension begins around 8- to 10-months (Fenson et al., 1994) at which point verbal labels can influence categorisation (Westermann & Mareschal, 2014). Verbal labels can affect perceived object similarity and can help infants categorise perceptually dissimilar exemplars of objects into a single category (Plunkett, Hu, & Cohen, 2008; Waxman & Booth, 2003; Waxman & Braun, 2005). For example, Plunkett et al. (2008) presented 10-month old infants with a series of objects that differed on a number of features where values on one dimension could combine with the full range of values on other dimensions. Infants categorised the objects into two categories (long neck and short neck) shown by a visual preference for the object that averaged across all dimensions. In contrast, when given a single novel label (e.g., “dax”) for the objects during the familiarisation phase, 10-month old infants categorised objects into a single category (Plunkett et al., 2008). Thus, in the imitation task, a label presented at demonstration and test may help infants form a single category for the demonstration and test puppets and facilitate memory retrieval across the stimulus change. Alternatively, verbal labels may function to direct infant’s attention to the relevant aspects of the learning task. Indeed, Taylor and Herbert (2014) found that 6-, 9- and 12-month old infants’ attentional patterns during the puppet demonstration session were related to their ability to reproduce the target actions at test. Using an eye-tracker, Taylor and Herbert (2014) showed that infants distribute their attention more widely than adults when viewing the puppet task. Critically, greater attention to the person and less attention to the background were related to learning outcome on the task (Taylor & Herbert, 2014). Furthermore, studies have shown that 9- to 12-month old infants increase attention to objects that have been labelled compared to objects that have not (Balaban & Waxman, 1997; Baldwin & Markman, 1989). Thus, verbal labels may serve as an attention grabber during learning and at test. The purpose of the present experiment was to determine whether verbal labels facilitate flexible memory retrieval via an attentional or categorisation mechanism. Twelve-month old infants were presented with a puppet that differed in both colour and form during the imitation test session following a live demonstration 10 min earlier. At this age, infants typically fail to reproduce the target actions when presented with a colour and form change puppet (Hayne et al., 1997). Some infants received empty language cues (no verbal label) whilst others received language cues (verbal label for object and actions) in either English or Chinese. The English language labels were used to determine whether language could scaffold learning and push infants into succeeding on the difficult flexible retrieval task. By 12-months of age infants can no longer discriminate between foreign-language phonetics (Best, McRoberts, LaFleur, & Silver-Isenstadt, 1995; Kuhl, Tsao, & Liu, 2003; Maye, Werker, & Gerken, 2002; Narayan, Werker, & Beddor, 2010; Werker & Tees, 1984) and mouth sounds alone are not sufficient to produce categorisation (Fulkerson & Haaf, 2003). Thus, the Chinese language labels were used to determine whether verbal information merely directs infant attention to the relevant aspects of the task. We hypothesised that infants who receive empty language cues will fail to reproduce the target actions with the form and colour change puppet, consistent with prior work (Hayne et al., 1997). In contrast, we predicted that infants who receive English language cues would reproduce the target actions with the form and colour change puppet, consistent with both an attentional and categorisation mechanism. If infants in the Chinese language group do not perform above baseline then the categorisation mechanism will be supported. In contrast, if infants in the Chinese language condition do perform above baseline then the attentional mechanism will be supported. If the effectiveness of verbal labels can be explained by categorisation, then infants’ language comprehension should be related to their subsequent imitation when given English language cues (e.g., Waxman & Booth, 2003).","Participants were 52 12-month old infants (26 males, 26 females) tested within 10 days of their birthday. None of the infants were born more than four weeks premature or experienced birthing difficulties. An additional 15 infants were tested but excluded due to infant fussiness (n = 6), failure to touch the puppet during the test (n = 2), experimenter error (n = 4) and previous exposure to a foreign language (n = 3).","The Oxford Communicative Development Index (Oxford CDI; Hamilton, Plunkett, & Schafer, 2000) was used to measure infant vocabulary comprehension and production. Comprehension scores around 15% and production scores around 0% are considered typical for 12-month old British infants (Hamilton et al., 2000).","Four hand puppets were used in the present study; two resembling a mouse and two resembling a rabbit, both made in either pale grey or pastel pink (see Hayne et al., 1997). The puppets were 30 cm in height and had a removable mitten 8 cm (w) × 9 cm (h) in matching pale grey or pastel pink on the puppet’s right arm. A large jingle bell was attached to the inside of the mitten during the demonstration for the experimental conditions, and attached to the back of the puppet for the control condition (see Hayne et al., 1997).","All infants were tested individually in the Developmental lab at the University of Sheffield. Upon arrival, the purpose of the study was explained to caregivers and informed consent was obtained. Infants were randomly assigned to one of four conditions: an English language condition (n = 14), a Chinese language condition (n = 12), an empty language condition (n = 12) and a control condition (n = 14). Half of the infants in each condition were female and half were male. Following consent, all infants engaged in a warm-up session in the waiting room with the experimenter until a smile was elicited. A native Chinese speaker conducted the Chinese language condition and a native English speaker conducted the English language and empty language conditions, both experimenters conducted the control condition. Infants in the Chinese language condition were exposed to three short Chinese phrases during the warm-up session (see Table 1); the same phrases were also given in English to the infants in the English language, empty language and control conditions. The purpose of exposure to Chinese phrases prior to the experiment was to build up the infant’s familiarity to hearing the experimenter speak in a foreign language. During the warm up session, the parent filled out the Oxford CDI. Parents in the Chinese language condition also answered two questions about the infant’s language exposure (Has your baby ever been exposed to Chinese language before e.g., neighbours, friends? Has your baby ever been exposed to any other languages except English?). After the warm up session, the experimenter then escorted the caregiver and infant to a separate testing room. Demonstration Session Infants were seated on their caregiver’s lap, with the experimenter kneeling on the floor facing the infant. During a warm up phase, the experimenter interacted with the infant until he or she appeared comfortable. Out of view of the infant, the experimenter placed the demonstration puppet on her hand. The puppet was then placed in front of the infant, out of reaching distance. For infants in the English language, Chinese language and empty language conditions, the experimenter performed three actions on the puppet: 1) taking the puppet’s mitten off, 2) deliberately shaking the mitten three times in succession ringing the jingle bell attached inside, and 3) replacing the mitten back on the puppet. This demonstration was accompanied by English language, Chinese language or empty language spoken by the experimenter (see Table 2). Verbal labels in the empty language condition were limited to “Look” and filler phrases in between each repetition in order to limit attention-grabbing cues. This enabled us to consider the role of verbal labels in directing attention during the task in the Chinese language and English language conditions. The actions were repeated three times in succession before the puppet was removed from view. For infants in the control condition, the experimenter shook the puppet side to side three times and repeated the action three times in succession accompanied by empty language cues. The purpose of the control group was to measure infants’ spontaneous production of the target actions. For all conditions, the experimenter said “Shall we do that again” or “Are you watching” in between each repetition. The colour and form of the demonstration puppet was counterbalanced across infants. Test Session The test session was conducted approximately 10 min after the demonstration. During the 10-min delay, the infant, caregiver, and experimenter interacted in the waiting room with unrelated toys. Each infant then returned with their caregiver to the experimental room and was presented with the form and colour change puppet during the test session (e.g., pastel pink rabbit during the demonstration session, pale grey mouse during the test session). The puppets were counterbalanced across condition. The experimenter placed the test puppet on her hand, out of view of the infant. The experimenter then revealed the puppet, which she either labelled (English language, Chinese language) or simply said, “Look” (empty language and control), before placing the puppet within reaching distance of the infant. Infants were then given 90 s to produce the target actions, timed from their first touch. The entire session was videotaped for later analysis.","The videotaped test sessions were coded for the presence or absence of the target actions and infants were given an imitation score based on the number of target actions produced (range 0–3). Approximately 23% (n = 12) of the videos were double coded by an independent experimenter. Inter-observer reliability analysis was 83% (kappa = 0.72). Preliminary analyses revealed no effect of gender on imitation scores so the data was collapsed across gender for subsequent analyses. To determine whether there were differences in infants’ imitation scores as a function of condition, a Kruskal-Wallis test was conducted due to differences in sample size and variance across condition and the control, empty language and Chinese language conditions violating the normality assumption. Overall, there was a significant effect of condition on infant imitation scores, H (3) = 12.93, p = 0.002 (see Fig. 1). In deferred imitation studies, memory is inferred if the imitation score in a demonstration condition exceeds the spontaneous production of target actions produced by infants in the baseline control condition (see Hayne, 2004; Meltzoff, 1985). Given that the imitation scores were not normally distributed, Mann-Whitney tests were used to compare each experimental group (English language, empty language, Chinese language) to the spontaneous production of the target actions in the control group. The English language group reproduced the target actions significantly more than the control group (U = 52.00, p = 0.035, r = −0.43). In contrast, the empty language (U = 60.50, p = 0.231, r = −0.32) and Chinese language (U = 79.00, p = 0.820, r = −0.06) groups did not differ significantly from the control group. Thus, only infants in the English language group reproduced the target actions above spontaneous production by infants in the control group. For the CDIs, data was missing for two infants whose parents did not complete and return the questionnaire. Children were given total scores for the number of words that the child comprehends and the number of words that the child produces. These scores were calculated by summing the number of items that the caregiver had marked as “understands” or “understands and says” for the comprehension score and the number of items that the caregiver had marked as “understands and says” for the production score. Children’s comprehension and production scores were expressed as a percentiles according using the norming data for the Oxford CDI (Hamilton et al., 2000) for analysis. Preliminary analyses revealed a significant effect of gender on vocabulary comprehension scores t (48) = 2.15, p = 0.037 with girls (m = 63.11, sd = 32.12) scoring more highly than boys (m = 46.08, sd = 23.13). There was no significant effect of gender on vocabulary production t (48) = 1.49, p = 0.144 (girls m = 56.00, sd = 32.87; boys m = 44.02, sd = 23.30). Given that there was an even gender split in each condition, the data was collapsed across gender for subsequent analyses. To determine whether vocabulary comprehension or production differed between infants in each condition, Kruskal-Wallis tests were conducted. Overall, there was no significant difference in vocabulary comprehension, H (3) = 2.431, p = 0.488 or vocabulary production, H (3) = 0.485, p = 0.922 between infants in each experimental condition (see Table 3). In addition, Kendall’s tau analyses revealed non-significant correlations between vocabulary comprehension or production scores and imitation scores for any condition.","The present experiment replicates and extends prior work (Hayne et al., 1997) in demonstrating that 12-month old infants fail to retrieve their memories if the stimulus presented during the test session differs in both colour and form from the one present during encoding, even after a short delay. Moreover, consistent with our hypothesis, the addition of experimenter provided language cues did facilitate flexible memory retrieval across the form and colour change stimulus in the present study. While there are limits on their effectiveness (see Zack et al., 2013), verbal cues can scaffold learning and push infants into succeeding on a difficult flexible retrieval task. To start to tease apart the attentional and categorisation mechanisms by which verbal cues facilitate flexible memory retrieval, it is particularly informative to consider the results from the Chinese language condition. The addition of Chinese language labels during the demonstration and test did not facilitate infants’ flexible memory retrieval above the spontaneous production of the target actions by the control group. Given that by 12-months of age infants can no longer discriminate between foreign-language phonetics (Best et al., 1995; Kuhl et al., 2003; Maye et al., 2002; Narayan et al., 2010; Werker & Tees, 1984), our monolingual English infants will not have been able to comprehend the Chinese verbal labels. Furthermore, infants fail to categorise following non-labelling mouth sounds (Fulkerson & Haaf, 2003). Instead, the Chinese verbal labels should serve as an attention- grabber. Thus, the results from Chinese language condition suggest that an attentional mechanism alone is unlikely to explain how verbal labels influence memory retrieval. There was no association between vocabulary comprehension or production at 12-months of age on imitation performance by infants in any condition. It is important to note that the Oxford CDI does not measure children’s comprehension of the words “puppet” or “shake” thus we do not have a validated record of whether infants understood the specific words used in our narration. Anecdotally, after the task, parents frequently stated they did not use these types of words with their infants. However, regardless of whether our infants benefited from the specific words used in the narration, it remains a possibility that categorisation may be the mechanism by which verbal cues facilitate flexible memory retrieval. A considerable body of research has shown that vocabulary comprehension is not essential for categorisation. For example, even before an age at which they can parse individual words (Jusczyk & Aslin, 1995), 3- and 4-month old infants categorise pictures of animals following verbal labels but not tones (Ferry, Hespos, & Waxman, 2010). The null finding for the Chinese language group appear to rule out phonetic discrimination as a potential mechanism for categorisation but not comprehension as a mechanism for categorisation. However, it is likely that verbal labels will facilitate both attention and categorisation during encoding and retrieval. Using eye tracking to monitor infants visual attention during an imitation demonstration session when verbal labels are given will help determine whether attention is one mechanism by which verbal labels can enhance flexible memory retrieval (also see Taylor & Herbert, 2014). In conclusion, the present results, combined with those of Herbert (2011), suggest that verbal labels can facilitate infants’ emerging memory flexibility, even as the task becomes progressively more difficult. The next steps in this research will be to determine the relationship between memory flexibility, language cues, and forgetting. Infants’ ability to flexibly retrieve their memories is influenced by the length of the retention interval between the demonstration and test sessions (e.g., Hayne et al., 1997; Herbert & Hayne, 2000a, 2000b). Moreover, prior work has demonstrated that verbal labels can facilitate the length of time over which a memory can be retained when the test stimuli are the same as those presented during the demonstration (Hayne & Herbert, 2004). Thus, it remains to be determined whether verbal labels can scaffold flexible memory retrieval across longer delays, or whether longer retention intervals are also a task demand too far for the early verbal infant."],["The development of episodic memory in children has been of interest to researchers for more than a century. Current behavioral tests that have been developed to assess episodic memory differ substantially in their surface features. Therefore, it is possible that these tests are assessing different memory processes. In this study, 106 children aged 3 to 6. years were tested on four putative tests of episodic memory. Covariation in performance was investigated in order to address two conflicting hypotheses: (a) that the high level of difference between the tests will result in little covariation in performance despite their being designed to assess the same ability and (b) that the conceptual similarity of these tasks will lead to high levels of covariation despite surface differences. The results indicated a gradual improvement with age on all tests. Performances on many of the tests were related, but not after controlling for age. A principal component analysis found that a single principal component was able to satisfactorily fit the observed data. This principal component produced a marginally stronger correlation with age than any test alone. As such, it might be concluded that different tests of episodic memory are too different to be used in parallel. Nevertheless, if used together, these tests may offer a robust assessment of episodic memory as a complex multifaceted process. --------------------------------------------------------------------------------","Six blind men wanted to discover for themselves the nature of an elephant. Each one went to the elephant and touched it. The first touched the elephant’s leg and said “it is like a tree,” the second touched the elephant’s tail and said “it is like a rope,” the third touched the elephant’s trunk and said “it is like a snake,” the fourth touched the elephant’s tusk and said “it is like a spear,” the fifth touched the elephant’s side and said “it is like a wall,” and the sixth touched the elephant’s ear and said “it is like a fan.” Characterizing healthy episodic memory development in young children is important because it allows problems with memory to be identified and informs appropriate educational strategies. Although the development of memory in children has been studied for nearly a century, to date there is considerable variation in the methodologies used to do so. The fable of the six blind men and the elephant serves to warn us that a single perspective on an intangible phenomenon may provide truth but can also be misleading. As psychologists, we can never directly assess psychological processes but can only measure performance on particular tests that are thought to rely on those processes. Different tests of episodic memory stem from different philosophical, theoretical, and empirical origins, and they differ substantially in the outward behavior they assess. Such eclecticism can be both a strength and a weakness. A range of testing methodologies can allow triangulation on a single common feature. This may allow production of a battery of measures that provides a more complete picture of a psychological process. However, a range of tests that vary largely in their methodologies may merely muddy any possible interpretation. In this study, the same sample of 3- to 6-year-old children was tested on a range of episodic memory tests. These tests are all very different in their surface features, so it might be predicted that they would produce different results. Nevertheless, they all putatively assess the same underlying cognitive ability, and as such it might instead be predicted that there should be a demonstrable association among them, reflecting this latent variable. The tests we chose to investigate are some of those that have been claimed to tap episodic memory or are candidates for such a claim. Therefore, we should expect to see a similar developmental change in all of the tests (Wellman, Cross, & Watson, 2001). In the following section, we briefly review the literature concerning these tasks. Free and cued recall ~~~~~~~~~~~~~~~~~~~~ Free and cued recall paradigms involve learning a series of items (words or pictures) and then later being asked to recall them, either with (cued) or without (free) external cues such as category words to aid recollection. Freely recalled items are more likely to be reported as “remembered” rather than as “known” compared with cued items (Tulving, 1985) and, therefore, are considered to be more reliant on episodic memory. Both free recall and cued recall improve between 3 and 8 years of age, with children of all ages reliably finding cued recall to be the easier of the two (Naito, 2003; Perner & Ruffman, 1995; Sluzenski, Newcombe, & Ottinger, 2004). What–Where–When ~~~~~~~~~~~~~~~ The What–Where–When test requires participants to remember the time and location of a particular event. Clayton and Dickinson (1998) argued that this requires an integrated spatiotemporal representation of the event, which corresponds to Tulving and colleagues’ definition of episodic memory (Tulving, 1972). The What–Where–When test produces cross- sectional developmental patterns similar to those of other tests, with improvements between 2.5 and 5 years of age (Burns, Russell, & Russell, in press; Hayne & Imuta, 2011; Newcombe, Balcomb, Ferrara, Hansen, & Koski, 2014; Russell, Cheke, Clayton, & Meltzoff, 2011). Unexpected source memory ~~~~~~~~~~~~~~~~~~~~~~~~ Source memory tests assess retention of the context of an event rather than its focal targets (Wheeler, Stuss, & Tulving, 1997). Here, participants are required to report not only what was learned but also (and unexpectedly) on the details of the context in which the learning occurred. Zentall and colleagues argued that deliberate encoding reduces the contribution of episodic memory, and thus only a question that is unexpected requires an individual to episodically reexperience the original event (Zentall, Clement, Bhatt, & Allen, 2001; Zentall, Singer, & Stagner, 2008). However, Ornstein, Haden, and Elischberger (2006) argued that it is children’s growing competence with event recall that translates into increasingly deliberate encoding. This would suggest that episodic memory development may facilitate the later emergence of purposeful remembering (Cuvo, 1975). Thus, it is unclear whether younger children may encode “expected” and “unexpected” items differently. Children under 5 years have, however, demonstrated very poor source memory (Drummey & Newcombe, 2002; Gopnik & Graff, 1988; Whitcombe & Robinson, 2000). Absent from our analysis in this article is a consideration of autobiographical memory reports. However, to the extent that they require children to report on previous events, unexpected/source memory tests can be viewed as similar in some ways to lab-based autobiographical/event memory tests. Autobiographical memory has been studied extensively in young children (e.g., Fivush, 2014; Goodman, Ogle, McWilliams, Narr, & Paz‐Alonso, 2014; Reese, 2014). Much of this work has examined children’s verbal recall of events during parent–child conversations (e.g., Burch, Austin, & Bauer, 2004; Farrant & Reese, 2000; Haden, Ornstein, Rudek, & Cameron, 2009) and interviews elicited by experimenters (e.g., Ornstein, Gordon, & Larus, 1992). Fivush (2011) and Fivush and Nelson (2004) argued for a distinction between episodic memory and autobiographical memory, suggesting that the former should be characterized by the content of the memory, whereas the latter requires an additional layer of cognitive sophistication. The current study took as its starting point this more “minimalist” view of episodic memory (see also Clayton & Russell, 2009) and as such does not involve an assessment of autobiographical memory. In summary, the literature using different episodic memory tests suggests that performance on each improves between 3 and 6 years of age. This may mean that these tests are able to produce consistent developmental trajectories despite very different testing methodologies. However, establishing comparable improvement in test performance throughout the preschool years is not sufficient evidence to conclude that a common cognitive process underlies performance on these different tasks. After all, many cognitive (e.g., theory of mind, language) and non- cognitive (e.g., running speed, height) factors improve over this period. What is required is an assessment of the degree to which performances on these different tests are related within the same individuals. This type of investigation has been carried out with respect to different tests of prospection, finding good correlation among most tests (Atance & Jackson, 2009). Here, 3- to 6-year-old children were presented with three putative tests of episodic memory (What–Where–When, Unexpected Source Memory, and Free Recall) and one that is thought to rely less on episodic memory than on semantic memory (Cued Recall) (Tulving, 1985). All of the memory tests were designed to produce continuous data (not pass/fail). This enabled us to investigate whether performances on these tests are correlated and whether they are able to hold together as a “battery” of tests.","At total of 106 children between 36 and 83 months of age were recruited from schools and nurseries in the Cambridge area of England. There were 27 3-year-olds (M = 42.3 months, SD = 3.6), 18 4-year-olds (M = 53.67 months, SD = 3.5), 27 5-year-olds (M = 66.2 months, SD = 3.3), and 34 6-year-olds (M = 76.7 months, SD = 2.7). The sample consisted of 49 girls and 57 boys (3-year-olds: 15 girls and 12 boys; 4-year-olds: 6 girls and 12 boys; 5-year-olds: 15 girls and 12 boys; 6-year-olds: 13 boys and 21 girls). The study was approved by the Cambridge University psychological research ethics committee. Informed written consent was received from parents before any child took part. Testing took place in an empty room or in a quiet corner of a classroom in the school/nursery. The majority of the children were native English speaking, Caucasian, and middle class, representative of the local area.","As shown in Fig. 1, the study had a nested design in which elements of each test described here formed the retention intervals for the other tests. Free Recall and Cued Recall In both recall tasks, children were shown eight photos of familiar objects or animals (e.g., a book, a horse) and were asked to name each one in turn. They were then told to look at the images and try to remember what was in them. Recall occurred after a delay of approximately 5 min. In Free Recall children were asked to tell the experimenter “what pictures had been on the cards,” whereas in Cued Recall they were asked to tell the experimenter what pictures of specific categories (animals or toys) had been on the cards. This methodology followed that of Perner and Ruffman (1995), but the images were not the same. What–Where–When Fewer children took part in this task because children were split between this and another study (not reported). As such, 68 children took part in this task (32 girls and 36 boys; 12 3-year-olds, 13 4-year-olds, 17 5-year-olds, and 26 6-year- olds). As shown in Fig. 2, children were given three pieces of “gold treasure” (plastic £1 coins) and three pieces of “silver treasure” (plastic 20p coins) and asked to hide them in two different trays (the “forest” and the “town”). There were two hiding sessions separated by approximately 5 min. In each session, children could hide in only one tray. During hiding, the experimenter highlighted each coin’s identity by saying, “Where are you going to hide that [gold/silver] treasure?” After a delay (∼5–10 min), a new character (“Mr. Crow”) who had “stolen” a specific subset of the hidden coins was introduced; for example, the gold treasure from the second hiding session had been stolen. Children were explicitly informed which treasure remained (e.g., “the GOLD treasure from BEFORE we looked at cards”). All elements of the treasure that was left (particularly the “when” element) were described to children in a number of ways (e.g., “before,” “earlier,” “first,” “longer ago”) to increase their chances of understanding what was being asked. Children were then asked to indicate the location of the remaining treasure by pointing. Children could then swap the coins for a sticker. Unexpected Source Memory One week after the first stage of the experiment, children were unexpectedly asked about elements of the “games” that had been played. The 11 questions concerned contextual details about the learning episode. Both open-ended questions (e.g., “What animal stole the treasure in the pirate game?”) and cued-choice questions (e.g., “Which video had a teddy bear in it?”) were asked. Some referred to games children had played that day, and others referred to games they had played the previous week.","Data were analyzed using Pearson’s and partial correlations to measure covariation among various test performance. To assess whether performance on the different tests may reflect a single latent variable, principal component analysis was used. Alpha was set at .05. Preliminary analysis revealed no differences in performance between boys and girls, and therefore gender was not considered in the main analyses. Table 1 shows the mean, standard deviation, and range of scores for each age group on each of the tests. As shown in Fig. 3, performance on all tests was positively associated with age. Table 2 indicates the correlations with age as well as the associations among the tests. Performances on many of the memory tests were related. The significant correlation between Unexpected Source Memory and Free Recall remained when Cued Recall was controlled (r = .317, p = .02). The positive correlation between Free Recall and Cued Recall remained after age was controlled. However, all of the other correlations were reduced to non-significance when controlling for age. The central question of this study concerns whether different tests of episodic memory can be said to be assessing the same underlying cognitive process. One way of addressing this is to examine the extent to which an appropriately amalgamated score from all four tests (using the 68 children who took part in all four tests) is able account for more variance than any of the tests on their own. This was done in two stages, a covariance summary and then a prediction of a known relevant variable, as a validation paradigm. A single principal component was produced by this analysis and had an eigenvalue of 1.91, explaining 47.9% of total variance. No other component had an eigenvalue over 1.0, and orthogonal rotation did not increase the eigenvalue of the second principal component over 1.10, indicating a single underlying source of covariance in this dataset. Of the four tests, the What–Where–When test loaded least well into this factor (.340). Finally, the principal component was then correlated with age. This produced a marginally stronger correlation than any of the individual tests (r = .699). This correlation was similar when dropping the weakest loading test (What–Where–When) and re-extracting the first principal component (r = .708).","The aim of this experiment was to investigate the consistency of a number of different tests putatively assessing the same underlying psychological process, namely episodic memory. Performance on each of the four tests was shown to improve gradually between 3 and 6 years of age, in line with previous literature (e.g., Hayne & Imuta, 2011; Naito, 2003; Reese, 2014). These tests differ in their surface features but are conceptually similar in that all aim to assess episodic memory. As would be predicted by this similarity, performances on many of the tests were correlated; however, few correlations remained significant after covariation due to age-related improvement being controlled. This result demonstrates that these tests are not equivalent and should not be treated as so when assessing episodic memory in children. However, the principal component analysis suggests that a single underlying factor may satisfactorily fit the data. Thus, there may be mileage in adapting these tests with the aim of reducing their surface differences and bringing them together into a single battery. Nevertheless, there is still some distance to go before a satisfactory battery of episodic memory tasks can be achieved. The tests used in this study were not designed to be methodologically similar; rather, they were designed to represent the variation present between currently used tests. Therefore, it is difficult to identify what factors contribute to low levels of age-independent correlation. There are many non-mnemonic cognitive and non-cognitive abilities that develop across the age range covered here. Furthermore, the tests differed in the extent to which they required receptive and productive verbal competence, confidence around a strange experimenter, executive functions, and other “extra-target” challenges. However, recent work with adults (Cheke & Clayton, 2013) shows that poor correlations among these tests are present even during adulthood. This implies that development of extra-target factors might not provide a full explanation for low covariation. Each of the tests used in this study explicitly tests different mnemonic skills. The What–Where–When test requires binding of spatiotemporal features. The Unexpected Source Memory task requires the ability to reanalyze previous experiences for new information. The Free Recall task assesses the ability to mentally initiate and guide retrieval in the absence of external cueing, and the Cued Recall task requires the ability to use category words as retrieval cues. These different elements are by no means the only important features of episodic memory: One limitation of this study is the absence of a measure of these children’s autobiographical memory reports. Still, the tests employed cover a range of different perspectives on the “defining features” of episodic memory. Adaptation of these types of test to facilitate the creation of a battery could provide benefits that are greater than the sum of it’s parts. Ultimately, it may allow the assessment of episodic memory as a whole without undue emphasis on any one particular feature. To summarize, this study revealed few associations for performance on different tasks putatively assessing episodic memory among 3- to 6-year-old children when age was controlled. This suggests that these tests are sufficiently different to lead to disparate results in studies across the episodic memory literature. Nevertheless, it was also found that performance across the tasks was well described by a single factor model. This may indicate that although the different tests are too different to use independently as equivalent tests, the conceptual similarities are sufficient to warrant their adaptation to create an episodic memory battery. Like the blind men’s perspective of the elephant, this would have informative powers above and beyond each individual test. Future work, therefore, should focus on development of such a battery in which the tests are more methodologically matched but remain structurally distinct. This battery might further include autobiographical reports and tests of episodic foresight. In this way, researchers will be able to investigate episodic memory not only as a coherent single process but also as a multifaceted one."],["Unlike those with type 1 blindsight, people who have type 2 blindsight have some sort of consciousness of the stimuli in their blind field. What is the nature of that consciousness? Is it visual experience? I address these questions by considering whether we can establish the existence of any structural-necessary-features of visual experience. I argue that it is very difficult to establish the existence of any such features. In particular, I investigate whether it is possible to visually, or more generally perceptually, experience form or movement at a distance from our body, without experiencing colour. The traditional answer, advocated by Aristotle, and some other philosophers, up to and including the present day, is that it is not and hence colour is a structural feature of visual experience. I argue that there is no good reason to think that this is impossible, and provide evidence from four cases-sensory substitution, achomatopsia, phantom contours and amodal completion-in favour of the idea that it is possible. If it is possible then one important reason for rejecting the idea that people with type 2 blindsight do not have visual experiences is undermined. I suggest further experiments that could be done to help settle the matter. --------------------------------------------------------------------------------","Unlike those with type 1 blindsight, people who have type 2 blindsight have some sort of consciousness of the stimuli in their blind field. What is the nature of that consciousness? More specifically, do those people have a visual experience of a stimulus or of some of its features, or do they lack such a visual experience, and have some other type of conscious state, such as a conscious feeling or thought? I address this question by considering whether we can establish the existence of any structural features of visual experience. Structural features of experience are necessary features of experience. I will argue that it is very difficult to establish the existence of any such features. In particular, I investigate whether it is possible to visually, or more generally perceptually, experience form or movement at a distance from our body, without experiencing some colour (chromatic or achromatic colour). The traditional answer, advocated by Aristotle, and some other philosophers, up to and including the present day, is that it is not. I argue that there is no good reason to think that this is impossible, and the evidence, although not conclusive, suggests that it is possible. If this is possible then one important reason for rejecting the idea that people with type 2 blindsight do not have visual experiences is undermined. This result is important for if it can be established that those who have type 2 blindsight are having visual experiences then we have reason to think that area V1 of the visual cortex is not required for visual consciousness. This is because such people suffer lesions to V1.1 (See Zeki and ffytche (1998), Stoerig and Barth (2001), and ffytche and Zeki (2011).) Moreover, there is some evidence to suggest that in fact type 1 blindsight does not exist at all, and that all cases of blindsight are really of type 2 (Overgaard, Fehl, Mouridsen, Bergholt, & Cleeremans, 2008). If that is right then one main source of evidence for thinking that there can be unconscious perception is removed. In section one, I explicate what structural features of experience are. In section two, I outline the nature of type 1 and type 2 blindsight. In particular, I outline the debate about the nature of the conscious state in those said to have type 2 blindsight. In section three, I discuss the difference between different kinds of mental states and argue that those with type 2 blindsight either have visual experiences or conscious thoughts. We should eschew the idea that their awareness or consciousness is a matter of them having feelings. In section four, I examine the evidence about the nature of the awareness or consciousness had in type 2 blindsight. I show that one reason given by Overgaard et al. (2008) and Overgaard and Grünbaum (2011) for thinking that those with type 2 blindsight have visual experiences is not a good one. I then go on to explicate two arguments that Brogaard (2011, 2012) has put forward in favour of thinking that the sort of consciousness in type 2 blindsight is conscious thought. I show that one of these arguments is not suitably backed up by the empirical evidence. So the weight of her position rests on the other argument. That argument relies on the premise that visual experiences have a certain structural feature: they must all be experiences of colour. In section five, I explore whether one should believe that visual experiences must have that feature and conclude that there is no good reason to think that. Indeed, the evidence tells in favour, although not conclusively, of the claim that they do not. That evidence also points towards an account of what the visual experiences of those with type 2 blindsight might be like that has not yet been considered. I show that visual experiences can be like that. I therefore conclude that there are no good arguments for the conclusion that the type of consciousness enjoyed by people with type 2 blindsight cannot be visual experience. And I suggest further experiments that could be done to test whether they do have such experiences.","What are structural features of perceptual experience? Structural features of experience are invariant features of experience. On a weak understanding, they are simply invariant features of human perceptual experience that exist as a matter of nomological necessity given the kind of human brain that we have. Thus, the perceptual experiences of creatures that have other types of brain need not exhibit these invariant features, nor need the perceptual experiences of subjects with human brains in possible worlds with a physics unlike our own. On a strong understanding, structural features of perceptual experience are metaphysically or conceptually necessary invariant features of experience tout court. Such features would be true of any creature with any type of brain in every possible world. Of course, there will be accounts of structural features of experience of strengths intermediate to the strong and the weak kinds just outlined: metaphysically necessary features true of all subjects with humans brains, and nomologically necessary features true of all creatures no matter what kind of brain they have. However, I set these aside in this paper. Here are some examples of propositions that some people have claimed specify structural features of perceptual experience. I offer these up only as candidates for propositions that specify structural features. I do not show that they really are ones.2 I begin with an example that many hold to be true: Necessarily, perceptual experiences are conscious. This is plausibly proposition that specifies a structural feature of perceptual experience—and a structural feature of perceptual experience in the strong sense. Many philosophers hold this to be true a priori. Many candidate structural features of perceptual experience will be features concerning what is represented in perceptual experience. Some may be pertain to all experiences. For example this proposition specifies an alleged such feature: Necessarily, perceptual experiences represent space and time. The representation of space and time is, somewhat plausibly, a weak structural feature of perceptual experience. However, some structural features may pertain only to experiences in a certain modality. Consider these modality specific claims: Necessarily, auditory experiences represent sound. Necessarily, visual experiences represent colour.3 Again these seem, to some degree, to be plausible truths about structural features of experience. For example, Aristotle in De Anima held that each of the sensory modalities had a proper sensible, that is, an object or property that could only be represented by that sensory modality and that was always represented by that sensory modality: sound in the case of hearing, colour in the case of vision, pressure and temperature in the case of touch, smells in the case of olfaction, and tastes in the case of taste. (Aristotle considered there to be only these five senses.) The proper sensibles contrast with common sensibles such as shape, which can be represented in more than one modality: vision and touch. For Aristotle the proper sensibles were the defining features of the sensory modalities. The proper sensibles were that which made each sensory modality the sensory modality it was. They individuated the senses.4 I will come back to discuss claim (iv) later in this paper. Some claims about candidate structural features of perceptual experience concern what is not represented in experience: Necessarily, perceptual experiences do not represent the future. And some such claims are restricted to experiences in specific sensory modalities: Necessarily, visual experience does not represent sound. Necessarily, visual experience does not represent tastes. Some claims about candidate structural features may be ones that pertain to combinations of representational features: Necessarily, auditory experiences represent some volume when they represent some pitch. Necessarily, tactile experiences represent only one object as being at any location. Necessarily, visual experiences represent only one colour to be on a surface at any given time. Necessarily, experiences of red are more similar to experiences of orange than they are to experiences of green.5 As I said above, but let me emphasise, I am not claiming that these are propositions that specific structural features of perceptual experience, only that they are the sorts of claim worthy of consideration for specifying structural features. Indeed claims about structural features of perceptual experience are very often particularly difficult to establish. Let me provide you with one example: whether it is possible for perceptual experiences to represent reddish-green. Wittgenstein discussed the questions of whether there could be a reddish-green colour, whether a reddish-green colour could be perceived, or whether the concept of reddish-green even makes sense. Lugg (2010) makes a careful summary of his remarks suggesting that while the popular interpretation of Wittgenstein has been that at least at some points in his career he claimed that there could be no perceptual experiences of reddish-green, Wittgenstein “is genuinely puzzled, that he is pulled in both directions and cannot commit himself either way.” (Lugg, 2010: 172). Nonetheless, perhaps inspired by Wittgenstein, some philosophers, for example Brenner (1987), have held that there could be no experiences as of a reddish-green. However, this claim looks to be disproved by recent results from Crane and Piantanida (1983) and Billock and Tsou (2004). They assert that they have created experiences of reddish-green in the laboratory. There are interesting questions about whether there is clear proof for the existence of such experiences. (See, for example, Lugg (2010) and Nida-Rümelin and Suarez (2009).) However, if such experiences have been created then a claim that some took to specify a strong structural feature of perceptual experience has been disproved.6 Faced with such a case, one might wonder how one could ever defend any claim about the structure of experience—either strong or weak. One might think that the primary evidence that would support a claim about the structural features of experience would come from the sorts of experience that one has had and the sorts of experience that one has not. However, that evidence concerns what is the case. How does one get from claims about what is the case to modal claims about what could be the case—as claims about the structural features of experience are? One might think that one can extend one’s knowledge from claims about what one has experienced to claims about what one could experience by drawing on one’s sensory imagination. For example, no one has ever seen the national animal of Scotland, the unicorn, for such creatures have never existed. Suppose someone had never visually experienced a unicorn because they had never seen one and had not hallucinated one, or had a visual experience as of one while watching a film or looking at a hologram or the like. Nonetheless, such a person could imagine what it would be like to visually experience a unicorn—as many people have done. One imagines what it would be like to experience a unicorn by conjoining in the imagination, in the appropriate way, what one knows of what it is like to experience a white horse and what it is like to experience a white horn. Based on this, it seems right to say that one can know that it would be possible to have an experience as of a unicorn. Such a combinatorial process allows one to consider numerous experiences that one has not had and whether they are possible. Despite this, however, one might think that the imagination is limited in an important respect. How could one come to have knowledge of whether it is possible to experience things that are not simple conjunctions of what one has experienced? Recall that Hume (1739–40/1975: 6) said that it is possible that one may be able to visually imagine what it would be like to experience that which one has not perceived (and, although Hume did not, we can add here that which one has not had an experience as of) and which is not something that can be visually imagined by a simple conjoining of things that one has seen. I will call such things “novel qualities”. For example, as Hume famously argued, if one had experienced all the shades of colour except one particular shade of blue, one might come to be able to imagine what it would be like to experience that shade if one was presented with all the other shades of blue laid out in order of resemblance. However, cases where one can imagine novel qualities are few and far between—as Hume himself noted. He states of the missing shade of blue, “the instance is so particular and singular, that it is scarce worth our observing” (1739–40/1975: 6). Hume goes on to note that, “We cannot form to ourselves a just idea of the taste of a pine apple [sic], without having actually tasted it” (1739–40/1975: 6) He thus foreshadows Jackson (1982) who claims that if one had only had an experience as of black and white and shades of grey among the colours, one could not come to know what it was like to have an experience as of red.7 Of course, if someone has only had experiences as of black, white and grey then they cannot legitimately conclude that experiences as of other colours, such as red are impossible. A notable feature of the experiments in which it was claimed that people came to have experiences as of reddish-green is that before having the experience as of reddish-green, subjects said that they could not imagine what it would be like to have such an experience, but afterwards they could imagine it. I think that, in general, people who have not had an experience as of reddish-green, and who have just had typical human experience, cannot imagine reddish-green. I certainly cannot. Given this, the inability to imagine reddish-green before having the experience would explain why people might have thought that such experiences were impossible. However, as in the case of red discussed by Jackson in the previous paragraph, we cannot conclude just from the fact that one has not had a certain type of experience, and the fact that one cannot imagine what such an experience might be like, that such an experience is not possible. In arguing about whether a certain sort of experience is possible, there may be further considerations that one can bring into try to establish the matter. If one were trying to establish that experiences as of reddish-green were impossible then one might adduce arguments concerning the opponency of the visual system. According to colour opponent theory, information from the eye is processed in the brain in an antagonistic manner. There are three opponent channels: the red versus green, the blue versus yellow, and light versus dark. The brain cannot signal the presence of red at a location at the same time as it signals that there is green at a location. (It could however, indicate that both red and blue were present at the same locations, in which case, it would be signalling that purple was present.) If all brain processing were subject to these opponent channels, then one might argue that subjects whose brains were so constrained could never experience reddish-green. (And if one made a case that their brains were so constrained as a matter of nomological necessity then, in doing so, one would be putting forward an argument that not being able to be as of reddish-green is a weak structural feature of experience.) However, Crane and Piantanida (1983) speculate that the “filling in” process that produces the experience allows the visual cortex to signal for the presence of red and green at the same location at the same time by allowing it to signal the presence of colours unconstrained from bottom-up opponent channels. What the considerations above show is that in every case where we wish to identify structural features of experience, we will have to look closely at what evidence is available. Evidence from our imagination can play a role in establishing what is a structural feature; however, the inability to imagine an experience should not be taken as conclusive evidence that such an experience is impossible. Furthermore, other evidence should be sought and weighed as best we can. Considering what the structural features of visual experience are will be important in thinking about the nature of the experiences had by those people who have type 2 blindsight. I will consider whether we can visually experience something as being at a distance from our body without at the same time having an experience as of its colour. Can we determine whether or not this is true? Knowing the answer could help settle a debate about the nature of type 2 blindsight. In the next section, I discuss the nature of type 1 and type 2 blindsight, and outline the question of what is the nature of the conscious state that is had in type 2 blindsight. I will distinguish between different candidates for what that conscious state is: a feeling, a perceptual experience, or a thought.","According to the traditional conception of blindsight (Weiskrantz, 1986), a subject has blindsight when he or she does not acknowledge any awareness or consciousness in a portion of his or her visual field (the blind field), yet nonetheless, some visual functioning remains intact. Physiologically, blindsight occurs when there is damage to the part of the primary visual cortex (V1) that corresponds to the blind field. Blindsight subjects report that they are totally blind in that area of their visual field. Yet, it can be determined that some information from the blind field does affect subjects’ behaviour. In particular, in a forced choice paradigm, subjects are able to guess with a high degree of accuracy about some features of a stimulus presented in their blind field. For example, Weiskrantz (1997: 23) says that people with blindsight “have been reported who are able, in their blind hemifields, to detect the presence of stimuli, to locate them in space, to discriminate direction of movement, to discriminate orientation of lines, to be able to judge whether stimuli in the blind field match or mismatch those in the intact hemifield, and to discriminate between different wavelengths of light, that is, to tell colours apart.” People with blindsight are initially unaware that they have blindsight and are not simply blind. It comes as a surprise to them that their guesses are accurate, although they can come to know that their guessing is accurate when experimenters tell them so. It should be noted that subjects cannot generate accurate guesses themselves. One reason is that in the experimental setting in which they display their guessing ability, the experimenter provides two options that they have to choose between, one of which is accurate. And the experimenter uses his or her knowledge of how the world is in the subjects’ blind field to ensure that there is one accurate option. However, subjects on their own cannot reliably generate two options one of which corresponds to the way the world is. It has been found that some patients who have been classified as having blindsight in fact report some limited consciousness in what is typically called their “blind field”.8 In response to such cases, Weiskrantz (1998) introduced a distinction between type 1 and type 2 blindsight. Type 1 blindsight is that which conforms to the traditional definition above in that no consciousness corresponding to the stimuli in the blind field is reported. Type 2 blindsight is defined as occurring when some limited consciousness of the stimulus in the blind field exists. However, how to characterise this consciousness is a tricky business. In attempting to characterise it, some people have focused on the fact that sometimes what is reported in type 2 blindsight is a mere conscious “feeling” or conscious “knowing” of the nature of the stimulus—but a conscious state that does not amount to a visual experience. In his definition of type 2 blindsight, Weizkrantz states that the nature of the consciousness is “acknowledged experience of events in the blind field in the absence of acknowledged ‘seeing’” (1998: xi). So, Weizkrantz, and others who have endorsed the existence of type 2 blindsight, have claimed that the consciousness that is present in these cases does not consist of a visual experience of the stimulus or some of its features. Many of the cases of type 2 blindsight that have been discussed in the literature are type 2 blindsight with respect to movement. Subjects with type 2 blindsight may have type 1 or type 2 blidsight or be completely blind with respect to other features of objects in their blind fields. In contrast to Weizkrantz, other researchers (Zeki and ffytche (1998), Stoerig and Barth (2001), and ffytche and Zeki (2011)) have claimed that the form of consciousness that type 2 blindsight subjects have is visual experience. That is, they claim that these subjects have visual experiences of movement (at least sometimes—typically the more high-contrast the stimulus and the faster it is moving it is the more likely subjects are to report awareness). These researchers often classify the subjects as having “Riddoch syndrome”. Riddoch (1917) examined men injured in war who reported that they were blind in one half of their visual fields, except for the fact that they reported seeing movement in them.9 These researchers noted that type 2 blindsight patients with respect to movement seem to be just like those patients Riddoch studied. Because they classify the awareness as visual experience they tend not to classify the subjects as having “type 2 blindsight”. Why is this? The answer is that if the subjects are having a visual experience, then it is tempting to say that they are just seeing and hence do not have blindsight. If type 2 subjects have visual experiences, should we classify them as just seeing and not having type 2 blindsight? The answer to that question is complicated because a subject who has type 2 blindsight with respect to some feature or features of an object is typically blind and/or has type 1 blindsight with respect to the other features of objects in the blind field. To illustrate, suppose we placed a red X-shaped object on the left of a subject’s blind field and then moved it to the right of that field. A subject might report having some form of consciousness of the left to right movement but deny having any consciousness of anything else. Moreover, at the same time, the subject could, in a forced choice paradigm, be able to guess reliably that the shape was that of an X and not be able to guess reliably that it was red. In such a case, the subject would have type 2 blindsight for the movement of the object, type 1 blindsight for the shape of the object, and simply be blind with respect to the colour of the object. With these considerations in mind, what is the answer to the question of whether, if type 2 blindsight involves having a visual experience, the subject would just be seeing and hence would not have blindsight? The answer is that, at least with respect to the feature in question that they are visually experiencing, the subject would be seeing that feature. Given that, it does seem appropriate to question whether it is right to say that those in whom the consciousness amounts to visual experience have a form of blindsight with respect to the feature they experience. For they just seem to be seeing it. Nonetheless, I will stick to calling people who report any form of consciousness of limited features people with type 2 blindsight—with respect those features that they so report—and I will also speak of their “blind fields”. One reason is that “Riddoch syndrome” is only a term that applies to those who reported blindness except for movement. It is useful to have a term that does not just pertain to cases of reported awareness of movement. Another reason is that it is very useful to have a term that applies to those who report some consciousness of just one, or a limited number, of features in the presence of either blindness or type 1 blindsight for other features, where the terms leaves open what the nature that awareness is. For it is useful to ask of such people whether we can determine whether their awareness is a feeling, a visual experience, or knowledge, or judgment of a feature of an object. To summarise: psychologists have identified a group of subjects who report being, for the most part, blind in a portion of their visual fields. Those subjects do, however, report some form of consciousness of movement—and only movement—in that field. Some researchers think that the awareness is a feeling, or knowing, or judging. Other researchers think that the awareness is a visual experience—a minimal or highly degraded one that represents a limited number of features of a stimulus. In either case, I will classify these subjects as having type 2 blindsight and investigate what we can determine about the nature of their conscious mental state. Do they have visual experiences or do they have some other type of conscious mental state? I will set aside the worry that, if it turns out that these subjects have visual experiences, then the term “type 2 blindsight” may not turn out to be the most appropriate nomenclature for their condition for it would be appropriate to say that the subjects see the feature that they claim awareness of. I said earlier in this section that many cases of type 2 blindsight involved subjects reporting consciousness of movement. A case not involving movement is reported by Overgaard et al. (2008) who investigated a subject GR. When subject to standard blindsight testing, GR displays behaviour which would lead one to classify her as having type 1 blindsight. For example, when asked whether she can detect stimuli in a certain portion of her visual field, given only the options of answering “yes” or “no”, she reports that she does not. When a letter is presented in that portion of her visual field and she is asked to guess in a forced choice paradigm whether the letter is “A”, “B” or “C”, she performs better than chance. However, Overgaard et al. tested GR further. They asked her to rate her experience on a four point scale: (CI) ‘clear image’, (ACI) ‘almost clear image’ (meaning ‘I think I know what was shown’), (WG) ‘weak glimpse’ (meaning ‘something was there but I had no idea what it was’), and (NS) ‘not seen’ (2008: 1). When GR was tested in the damaged portion of her visual field on thirty-three occasions, seven were reported as “clear image”, eleven as “almost clear image”, twelve as “weak glimpse”, and three as “not seen”. Moreover, there was a positive relationship between the accuracy of the “guess” about what the stimulus was and the clarity of the experience. In fact, there was the same relationship between accuracy and clarity in her blind field as there was in her intact field. These results indicate that GR does have some consciousness of the stimulus. It therefore seems right to classify her not as having type 1 blindsight, but type 2.10 As in the case of those who have type 2 blindsight with respect to movement, researchers have different views about what kind of consciousness to ascribe to GR. Overgaard et al. (2008) and Overgaard and Grünbaum (2011) argue that GR has a visual experience, while Brogaard (2011, 2012) argues that GR only has a conscious thought that is about the nature of the stimulus. In this section, I have explained the difference between type 1 and type 2 blindsight and briefly outlined the different answers in the debate about the nature of the awareness had in type 2 blindsight. Some say that it is a feeling or thought. Others say that it is a visual experience—albeit a degraded one. In the next section, I give an account of what these different types of mental state are. I argue that the consciousness of those with type 2 blindsight should not be described as a feeling. The only serious options for what the nature of their awareness or consciousness is, is a visual experience or a conscious thought.","What is the difference between feelings, thoughts, and visual experiences? Thoughts are a type of propositional attitude. Other types of propositional attitudes include beliefs and desires. When one has a propositional attitType 2 blindsight ude one takes an attitude, such as holding it to be true in the case of belief, or wanting it to be true in the case of desire, to some proposition. For example, if one believes that Scotland should be independent, then one takes the attitude of holding it to be true towards the proposition that Scotland should be an independent country. If one desires that Scotland be an independent country, then one takes the attitude of wanting it to be true that Scotland is an independent country. When one has a thought about something one can be merely entertaining a proposition (considering whether a proposition is true), or one can be endorsing a proposition. Propositional attitudes are representational states. They are about something, and what they are about is specified by the proposition to which one takes an attitude. In the case of the thought that Scotland should be an independent country, the proposition, that Scotland should be an independent country, specifies that which is represented. Thoughts, unlike beliefs and desires, are always occurrent states rather than dispositional states. Thoughts can be conscious or unconscious. There is an interesting question as to whether conscious thoughts have phenomenal character.11 One view is that in and of themselves they do not, but that they are usually or always accompanied by states that do. In particular, they might be accompanied by visual imagery. In the case of thinking that Scotland should be an independent country, perhaps one has visual imagery of mountains and lochs, the Saltire, and the face of William Wallace. The thought might also be accompanied by auditory imagery. For example, one might hear the words “Scotland should be an independent country”, or “Freedom”, in one’s own inner voice. On this view, no imagery in particular is essential to thinking the thought (although one might think that some imagery or other usually or always does so). Someone else might have visual imagery of Glasgow and of the faces of Robert the Bruce and the Black Douglas. And they might not have auditory imagery of the words “Scotland should be an independent country”, but the French words “Ecosse devrait être un pays indépendant”. Another view, however, is that thoughts have their own proprietary phenomenal character associated with grasping the meaning of the proposition. Contrasting with the propositional attitudes are feelings, also known as “sensations”, such as pains, itches and tickles. These states clearly have phenomenal character, and indeed plausibly, theses states are individuated by their phenomenal character: what makes a pain a pain is the way that it feels, and what makes an itch an itch is the distinctive itchy feel that such states have. Thus feelings and sensations have been thought of as essentially conscious states. Traditionally feelings have been conceived of in philosophy as states that do not represent. Why is that? Thomas Reid (1785/2002: I. i. 36) said that sensation, “hath no object distinct from the act itself”. The idea is that if I have a sensation of pain, there is no object—a pain—that I am sensing. My sensation is not ‘of’ any object, therefore it is not representational. More recently, the traditional view of feelings as non-representational has been questioned. It is agreed that feeling states, such as pains and itches, do not represent objects called feelings—objects that are pains or itches. Rather, it is claimed that feelings represent different states of the body. (See Armstrong (1962, 1968) and Pitcher (1970, 1971) and, more recently, Tye (2006a, 2006b).) For example, a throbbing pain in one’s big toe might represent that there is an increase and decrease in the volume of damaged and inflamed tissues in one’s toe. A sharp stabbing pain in the chest might represent that something pointed is entering and tearing asunder the flesh in one’s upper torso. A feeling of hunger might represent one’s stomach contractions and low blood glucose. Perceptual experiences have been traditionally thought of as hybrid states: being somewhat like the propositional attitudes and somewhat like feelings or sensations. Like the propositional attitudes, perceptual experiences seem to represent and be about things in the world. My present visual experience of a teapot seems to represent a silver object with a round body with various protrusions, corresponding to the handle, spout and knob of the lid. My auditory experience of the whistle of the kettle, represents a high-pitched loud note.12 At the same time, like feelings and sensations, perceptual experiences are essentially conscious states that have phenomenal character that differentiates one perceptual experience from another. Given the characterisation of these three types of states that are the candidates for what kind of state a person with type 2 blindsight is in, it is clear that feelings are just not good candidates. The reason is that the people with type 2 blindsight say that they have a feeling about some state of affairs in the world. For example, they say that they have a feeling that something in front of them is moving. Such a state would be a state that represented something in the world exterior to their body. As feelings are either not representational at all, or they represent something happening to the body, they are not the type of state that those with type 2 blindsight are reporting. We can explain, nonetheless, why it is the case that people with type 2 blindsight use the word “feeling” to describe their mental state. Sometimes in everyday language we use this word to indicate that we are not quite sure what kind of mental state we are in. One might say, “I have a feeling that our guests are arriving on Tuesday”, when one wants to report that one has some evidence to this effect but one isn’t quite sure where from. One might not know whether one remembers this, or whether one is guessing at this based on what one was previously told about the travel plans (say that some other destination would be reached by Monday), or whether one has worked it out based on the time one knows it takes to travel between certain places, or what have you. However, although this everyday usage is perfectly acceptable, that does not mean that we should take such talk to imply that subjects mean that the person has feelings in the sense outlined above. Therefore, I will limit my investigation of the mental states of people with type 2 blindsight to the investigation of whether such people are having visual experiences or whether they are having thoughts. How does one determine whether someone is having a conscious thought about something or whether he or she is having a perceptual experience? That is a very tricky question indeed. Perceptual experience seems to inform us of the way the world is now. (At least it seems that way to us: even if we are looking at a distant star and our experience is actually informing us of an event that took place many years ago.) However, we may not believe that the world is as our experience presents it to be. For example, one might think that one is suffering from an illusion or a hallucination and so not be inclined to believe what one’s experience seems to tell one. To this extent, experience is different from belief. However, we are considering how to distinguish experience from thought. One could certainly entertain the thought that this is how the world seems, yet be inclined to desist believing that is how it is because one has reason to not to fully trust the thought in question. So we have no reason yet to distinguish experience and thought. One way in which thought and perceptual experience are different is with respect to their phenomenal character. Perceptual experience typically seems to be about the about the world in front of one at a distance from one’s body. It typically tells one, among other things, about colours, shapes, sizes, positions, movements, and so on.13 Clearly thought can be about this too—although it can be about many more things than it seems visual experience, and perceptual experience more generally, can be about. However, what it is like to visually experience colour or shape or size or movement is different to what it is like to think about these things, if indeed there is anything it is like to think about them. But what one can say about this difference is unclear, not least because there is such discord in thinking about the phenomenal character of thought—as I outlined earlier in this section—and not least because it is hard to describe the phenomenal character of visual experience other than to say what it is an experience as of. Nonetheless, although describing differences in phenomenal character is difficult, and although there are different views about the phenomenal character of thought, I take it that visual experience has a distinctive phenomenal character either because thought has a different one or lacks one. Thus, in ascribing a visual experience to a person, one is ascribing to them a distinctive sort of phenomenal character. I discuss this topic in greater length in the section below. There may be other differences between thought and perceptual experience. For example, some people think that thought is conceptual while visual experience is, or is in part, or can be nonconceptual. However, I leave this topic aside for I don’t believe that even if there is this difference, it will help to settle the question of which state is had by people with type 2 blindsight. How one could test for perceptual experience rather than thought—as opposed to merely knowing what the difference is between them—is a difficult matter too. One might think that one could simply scan the brain and see if the visual areas of the brain are active, However, one can only know what parts of the brain correspond to visual experience by correlating reports of visual experience, and its lack, with brain activity. When we are dealing with blindsight, an important and pertinent issue is what areas of the brain should be taken to give rise to visual experience. Stoerig and Barth (2001) and ffytche and Zeki (2011) use their conclusion that people with type 2 blindsight are having visual experiences to deny the claim that has often been made that activation in the V1 areas of the visual cortex is necessary for visual experience. Therefore, we cannot rely on evidence about what areas of the brain are active as a guide to what kind of mental state is being had by the subject. In this section, I have argued that people with type 2 blindsight are either having conscious thoughts about the stimuli in their blind fields or they are having visual experiences. They are not having feelings. Perceptual experiences have different phenomenal character to conscious thoughts, if indeed the latter have phenomenal character. In the next section, I go on to examine the evidence and the arguments that have been put forward by both sides concerning whether the type 2 blindsight subject is consciously thinking about or visually experiencing the world.","Are type 2 blindsight subjects having degraded visual experiences of some sort or are they having conscious thoughts? Let’s look at the evidence from the subjective reports of various subjects. Zeki and ffytche (1998) examined Riddoch’s (1917) accounts of the subjective reports of his patients. In all these cases, the patients “were able to detect the presence of motion within their scotomatous fields, without being able to characterize the other attributes of the stimulus” (1998: 26). Here are the reports: Patient 1: “The ‘moving things’ have no distinct shape, and the nearest approach to colour that can be attributed to them is a shadowy grey”. Patient 2: ‘The ‘moving something’ had neither form nor colour. It gave him the impression of a shadow”. Patient 3: “could detect the movement of feet in the street ‘. . . though they had no shape’” Patient 4: “. . . declared he could distinguish no object . . . but he knew that something had moved through his blind field” Patient 5: ‘They [the moving objects] don’t appear to have any colour or shape. They look like shadows. Sometimes I can tell if the moving things are white.” (1998: 26) In each of these cases it is clear that movement is reported. It is important to distinguish the question of whether subjects could identify the colour of the stimulus from the question of whether they had an experience of some colour or other, and this is not clearly enough done in the reports of the subjects’ experience. Subjects clearly could not detect the colour of the stimulus. (Although one subject says that they could tell if something was white, this is an anomalous report.) However, the subjects report that their experiences are like looking at shadows. And one alludes to the colour grey. One view is that subjects visually experience dark grey indeterminate shapes moving on a slightly lighter grey or slightly darker grey background. Another view is that subjects have visual experiences without experiencing any colour properties. A third view is that the subjects did not have visual experiences, they only had thoughts. The best candidate for the thought that they had is that it was a thought that something moved, perhaps more specifically the thought that something moved in a certain direction, with a certain speed. Which of these views is true is unclear. ffytche and Zeki (2011) found two subjects, GN and FB , who were very similar to the patients of Riddoch (1917) in terms of their reports and responses. In addition to reporting consciousness of the stimulus, GN and FB “could prepare drawings of what they had perceived in their blind field, which compare favourably with the drawings of the same stimuli when presented to their intact fields” (2001: 254). ffytche and Zeki go on to say “their drawings and … descriptions left us in no doubt that the experiences they had were visual in nature and amounted to what might be called ‘visual qualia’” (2011: 254). However, their conclusion goes rather beyond the available evidence because one could draw a picture of how one thought the world was, rather than how one experienced the world to be. Another person with type 2 blindsight for movement who has been studied in modern times is GY. Zeki and ffytche (1998: 29) report that in 1993, GY described his visual experiences as dark and shadowy. However, later, in 1994, he changed his mind. He now said that he had a “’feeling’ of something happening in his blind field and, given the right conditions, that he is absolutely sure of the occurrence” (1998: 29). When Zeki and ffytche pointed out to him that his description had changed he said that he had previously “been using language that he thought a normally sighted person would understand” (1998: 29–30). Two years later, in 1996, he said his experience was “as that of ‘a black shadow moving on a black background’, adding that ‘shadow is the nearest I can get to putting it into words so that people can understand’. (1998: 30). When Zeki and ffytche tested GY’s blindsight for movement, they asked him to indicate the direction of movement of a stimulus and his level of awareness on a four point scale, similar to the way in which Overgaard et al. (2008) tested GR (described in section two above). Like Overgaard et al., Zeki and ffytche found that GY reported consciousness of movement more frequently when he was given the four-point scale to use, compared to the condition in which he had to indicate with a “yes” or “no” whether he was conscious. Moreover, his ability to correctly identify the direction of the movement correlated positively with his reports of awareness.14 The evidence from these varying reports of GY is not enough to settle the question of the nature of his conscious state one way or the other. GY has been investigated by other experimenters concerning the nature of his consciousness. Unfortunately, the evidence points in opposite directions as to the nature of his consciousness. Stoerig and Barth (2001) report that GY denies seeing. And they cite previous descriptions that GY has given of his experience: “He is aware of ‘something moving’ but it appears as ‘black on black,’ like ‘a mouse under a blanket’ (personal communication), or ‘similar to that of a normally sighted man who, with his eyes shut against sunlight, can perceive the direction of motion of a hand waved in front of him’ (Beckers & Zeki, 1995, p. 56).” (2001: 582). They also report that “in other experiments he has, for instance, stressed an absence of color sensation” (2001: 582). Yet, they found that despite denying seeing, when a moving texture of low contrast was presented to GY in his intact visual field, he accepted that it created the same conscious state in him as a high-contrast bar moving in his blind field. Furthermore, they found an even better match when they used an apparent motion stimulus.15 Stoerig and Barth conclude that as a match with a visual experience was made, GY’s awareness or consciousness in his blind field is just the same as the visual experience he has in his intact field. This is a minimal or degraded experience: one that they describe as having a “reduced phenomenal content” (2001: 584). In contrast, Persaud and Lau (2008) gave GY several definitions of “qualia”, a term that I take them to hold is synonymous with “phenomenal character”.16 They then questioned GY as to whether he experienced any qualia in his blind field. His answer was that he denied “having visual qualia of stationary stimuli in his affected field. He was adamant that he never has visual qualia in his affected field in everyday life.” (2008: 1048). And asked whether he had qualia of moving stimuli in the affected field he replied “No, never” (2008: 1047). So the evidence about GY’s conscious state flip-flops over time, and is dependent on how he is tested and what he is asked. Finally, the last piece of evidence comes from the Overgaard et al. (2008) study of GR, outlined in section two above. Recall that previous to their investigations, GR had been classified as having type 1 blindsight. When asked whether or not she can detect stimuli in her blind field and given only the options of answering “yes” or “no”, she reports that she does not. Overgaard et al. then asked GR to rate her experience in the blind field on a four point scale, “(CI) ‘clear image’, (ACI) ‘almost clear image’ (meaning ‘I think I know what was shown’), (WG) ‘weak glimpse’ (meaning ‘something was there but I had no idea what it was’), and (NS) ‘not seen’” (2008: 1). Out of thirty-three occasions, GR reported seven times a “clear image”, eleven times an “almost clear image”, twelve times a “weak glimpse”, and three times “not seen”. In light of this evidence it seems that we should classify GR as having type 2 blindsight, but what we should conclude about her conscious state is that it is not clear what its nature is. In summary, the reports of subjects’ experience and the experiments performed to try to get clearer about the nature of the mental states of those with type 2 blindsight are inconclusive. The evidence is mixed and points in different directions. Those who believe that people with type 2 blindsight are having visual experience point to the evidence which suggests that they are having visual experience; those who deny this draw attention to the evidence which suggests otherwise. Both sides rightly warn of taking introspective reports at face value. I will now look at different arguments that have been made by researchers in favour of one or other of the positions that people with type 2 blindsight are either having (possibly minimal or degraded) visual experiences or that they are having conscious thoughts. These arguments go beyond the citation of the introspective reports in the different conditions mentioned above. One argument in favour of the idea that people with type 2 blindsight are having a visual experience (and hence visual phenomenal character or qualia) is inspired by the causal origin of the reported awareness. For example, Overgaard and Grünbaum suggest that there is reason to hold that if a subject reports a conscious mental state and that it is caused by visual stimuli then the state should count as a visual experience. They state, “a visual process is one in which a subject at some level reacts to something visual. From this … it should follow that if there is any kind of preserved conscious experience in blindsight subjects caused by visual stimuli … those experiences should be conceived of as visual” (2011: 1858).17 This argument is spurious. There are many counterexamples which spring from noting that the nature of the stimulus, the nature of the sensory organ, and the kind of early perceptual processing that takes place does not fix the nature of the conscious state that is subsequently had by a subject. First, there are examples in which the nature of the stimulus, the nature of the sensory organ, and the kind of early perceptual processing that takes place may all be of one modality, while the subsequent experience is in a different modality. For example, in synaesthesia, people have an experience in one modality, say a visual experience of redness, caused by a visual stimulus affecting their eyes, which leads to visual processing. However, at the same time another experience in a different modality is also caused to occur in them, such as an auditory experience of sound. What this shows us is that the criteria for which type of experience is being had is not the same as which type of stimulus, sensory organ, or perceptual processing is taking place. Although often the modality of the experience will match the modality of the stimulus, organ and processing, it need not. These can come apart in interesting ways.18 Second, there are examples in which the nature of an experience had in one modality is affected by processing in another modality—typically in cases labelled as “cross-modal illusions”.19 For example, in the McGurk effect when an auditory stimulus—a /ba/ sound—is heard alone, it is typically reported accurately as a /ba/ sound. But when it is heard whilst looking at lips making movements that would produce a /ga/ sound, then people typically report hearing a /da/ sound instead (McGurk & MacDonald, 1976). Their auditory experience of the /da/ sound has as a causal origin both auditory and visual stimuli, sensory organs, and processing. Another example is the sound- induced illusory flash experience. When one flash is presented together with two tones, subjects frequently reported that they saw two flashes (Shams, Kamitani, & Shimojo, 2000). In this example, at least one of the visual experiences of the flash had both auditory and visual stimuli, sensory organs, and processing. There are many other such cross-modal illusions. Third, there are clear cases in which stimuli, sensory organ activation, and processing all belonging to the same modality cause conscious states other than perceptual experiences. For example, on seeing something or hearing something, or both, one might become incredibly sad, or happy, or angry. Emotions are not perceptual experiences but they are frequently caused by perceptual stimuli, sensory organ activation, and processing. Thoughts beliefs, desires, and volitions can be caused by perceptual stimuli, sensory organ activation, and processing. Likewise, if there are any examples of type 1 blindsight, or any form of unconscious perception such as that apparently caused by masking, then there will be examples of conscious thoughts involved in guessing that are not perceptual experiences. In short, one cannot argue that the modality of the stimulus, the sensory organ, and the existence of perceptual processing, entails that a subsequent conscious state is a perceptual experience of one modality or another. Nor can one even determine that it will be a perceptual experience, rather than some other kind of mental state. Thus, the argument just considered for the conclusion that the conscious mental state of those with type 2 blindsight is not sound. In contrast to Overgaard and Grünbaum (2011), Brogaard (2011, 2012) argues that, based on the evidence to date, there is reason to believe that those with type 2 blindsight are having thoughts about the world and not visual experiences. (She in fact thinks that whether or not in the end she is right about this is an open question that could be settled in the future by some further detailed study of subjects who are more thoroughly instructed with respect to, and asked about their own, phenomenal character. However, from now on, I will present her arguments in favour of the view that those with type 2 blindsight lack visual experience and have merely conscious thoughts without this qualification.) Brogaard holds that people with type 2 blindsight have thoughts about the stimulus in response to their guessing about the way the world is. She says, “Individuals with blindsight can make correct guesses. Guesses come with a phenomenology, just not the kind normal individuals have when a visual stimulus is presented to them… it has not been shown that the phenomenal consciousness blindsight involves is distinctly visual. I suspect that it is not” (2011: 459). Brogaard’s argument for this is that the proximal cause of a visual experience is the visual stimulus, the stimulation of the eyes, and visual processing, but this is not the proximal cause of the conscious state of the type 2 blindsight subject. The cause of the conscious state of those with type 2 blindsight includes the visual stimulus, the stimulation of the eyes, and visual processing, but it has other causes besides. The conscious state of those with type 2 blindsight has guessing as a proximal cause. Moreover, the proximal cause of the guessing is not the same as that of visual experience either. She states, “different mechanisms no doubt underlie guesses and seeings. So guesses and seeings have different proximate causes” (2012: 596).20 Do we have evidence that people with type 2 blindsight only have the conscious state they report when they make guesses, which is what Brogaard’s account requires? One might think that there is such evidence. After all, in some cases outlined above, it is only when subjects were asked to guess what was in front of them that they reported that they had some conscious state (and even then only when they were asked whether they had some conscious state using the four-point scale, rather than when asked simply to state whether or not they were conscious of the stimulus). However, it is unclear whether Riddochs’ patients only responded that they were aware of movement when asked to guess or whether they spontaneously reported it. And the same seems true of GY—although admittedly his case history is complex. However, even if guessing was not required, that does not stop the subjects in question simply having spontaneous thoughts—as opposed to visual experiences— about their situation. Thus, although Brogaard’s account requires guessing, a very similar one that requires only thoughts does not require the evidence about guessing to turn out one way rather than another. Moreover, even if it turned out that subjects did only have the conscious mental state when they guessed, this does not guarantee that their conscious mental state is a though. It could be that they have conscious visual experiences and that guessing is required in order for the people to have those conscious visual experiences. For example, it could be that in order to have a visual experience, people with type 2 blindsight have to focus their attention in a certain way, and that this is facilitated by guessing. This is just one suggestion; there may be others that could explain why the subjects only had a visual experience when they guessed. In short, I don’t see that we have good evidence that guessing is the proximal stimulus of the conscious state that people with type 2 blindsight report, and even if it were, that is not enough to show that those people are not having visual experiences. A second argument is used by Brogaard to suggest that those with type 2 blindsight are having conscious thoughts rather than having visual experiences. Speaking of those who report awareness of movement, she states, “even if a stimulus perhaps does give rise to an experience with a ‘‘clear’’ phenomenology, this does not provide any evidence of visual awareness of color, shape, or location” (2011: 458), and “blindsighters lack the sort of distinctly visual awareness that includes a purely qualitative color phenomenology” (2011: 459). Those attributes, she implies, are necessary for having a visual experience, and because subjects lack them, they are not having visual experience and are only having conscious thoughts about the stimuli.21 As stated briefly above, it is not clear whether those subjects who are aware of movement are aware of colour. One problem is that those who are reporting on the experiences sometimes say that the subjects were not aware of the colour properties of the stimulus. However, while it is clear that subjects can’t tell what colour things are in the world, it is not so obvious whether they experience the world as being some colour or other—even if it does not have that colour. On the one hand, it seems as if the conscious awareness reported is sometimes of a dark shadow moving against a darker background. This would be an experience of: colour: dark grey and black form: an indeterminate shadowy shape, and movement: from one direction to another. We can certainly have visual experiences like those. If we get people reporting awareness of those features, then a crucial premise of Brogaard’s argument is false and she lacks a good reason to deny visual experience to people who have such awareness.22 However, on the other hand, there is reason to think that the experience is even more minimal—at least in some subjects. This would be an experience of: movement: from one direction to another form: an experience of an area lacking clear boundaries Why would one think this? GY and the patients of Riddoch say that their experience is somewhat like looking at a shadow. But they don’t straightforwardly say that it is like looking at a dark shadow moving on a dark background. It is as if subjects are gesturing to something that is not quite like a shadow. When it was pointed out to GY that he had changed his description from having a visual experience of a shadow (1980) to just having a feeling (1994), he said he was using language that he thought the sighted would understand by mentioning shadows. In 1996 “he described his experience as that of ‘a black shadow moving on a black background’, adding that ‘shadow is the nearest I can get to putting it into words so that people can understand” (Zeki & ffytche, 1998: 30). This provides reason to think that some subjects are having very minimal visual experiences: movement without colour and rather indeterminate form. As we have seen, it is very hard to determine what the conscious mental states of people with type 2 blindsight are like. But let us suppose, as Brogaard does—and as we have seen there is some, although not conclusive, evidence to suppose—that those with type 2 blindsight (or at least some people with type 2 blindsight) do have awareness of movement and indeterminate form without colour. Is Brogaard right to think that this is good reason to think that those with this condition lack visual experiences and hence lack visual phenomenology? I will investigate this claim by considering whether there could be other minimal visual experiences of movement or form without colour. In other words, what I want to do is to investigate whether the following claim about an alleged structural feature of experience, mentioned in section one above, is true: Necessarily, visual experiences represent colour. Brogaard’s argument requires that it is. But as we have seen it is exceptionally difficult to establish the existence of any structural features of experience. What evidence is there that this claim about an alleged structural feature of experience is true? As I explained in section one, Aristotle said that sight can be defined as the perception of colour. (See Sorabji (1971) for a clear commentary on this point focusing on Aristotle’s De Anima, Book II, chapters 4 and 6. Sorabji claims that for Aristotle, this claim is not put forward as a truth that holds in all possible worlds, but that people may consider that it does.) Sorabji goes on to say that if one holds this view one should go on to say that there can be no visual experience without experience of colour. Size, shape, and movement are visually experienced in virtue of experiencing coloured things. Other philosophers have echoed this thought. Pete Mandik, for example, considers the question, “Can there be a visual experience devoid of both color phenomenology and black-and-white phenomenology?” (2014: 225) and answers “While I’ve not conducted anything remotely resembling a formal survey, I’m pretty confident that most philosophers of mind will answer ‘no’” (2014: 228).23 John Hyman, who cites both Aristotle and James Clerk Maxwell as inspiration, says: colours, like smells and tastes, are basic properties, relative to the sense with which we perceive them. In other words, whatever else we perceive by the sense of smell–for example, that a fruit is rotten or that a child is ill— we perceive by smelling smells; whatever else we perceive by taste we perceive by tasting tastes; and whatever else we perceive by sight, we perceive by seeing colors—including the achromatic colors, of course. For example, I cannot see the shape of a banana except by seeing its spatial boundaries, however fleeting and uncertain this experience may be, And I cannot see its spatial boundaries except by seeing the differences of color that make it visibly distinct from its surroundings. That is why, as James Clerk Maxwell pointed out, all vision is color vision (Hyman, 2006: 18). But is it? Surprisingly little evidence is garnered by those who state that it is. Usually it is stated in the manner that Hyman and Brogaard state it: as a fact; and any backing is given by citing Aristotle, Maxwell and other luminaries in an argument from authority. I suspect that the alleged fact is taken to be obvious based on repeated introspections of only visual experiences that are of coloured things. However, as I argued in section one, this is not good evidence in support of their being such structural features of experience. In the next section, I consider evidence that we might have against colour being a necessary feature of visual experience—that is against colour being a structural feature of visual experience. I will begin by considering the nature of sensory substitution. I will argue that there are two plausible interpretations of subjects with congential blindness using tactile visual sensory substitution. One interpretation is that they are not having any new perceptual experiences. The other is that they are having new visual experiences of distal form that are not experiences of colour. If the latter interpretation were true, it would show that colour is not a structural feature of visual experience. However, unfortunately, the interpretation that they are not having new visual experiences cannot be completely ruled out. I will then consider the cases of an achromatopsic who can experience distal form created by chromatic boundaries alone, phantom contours, and amodal completion. I argue that these provide evidence of visual experience of distal form and movement with no differences in colour, luminance or texture, but without the complete lack of these. Such experiences provide an alternative explanation of type 2 blindsight and their existence lends support to the idea that there could be visual experiences in the absence of experience of colour. This evidence undermines Brogaard’s second argument in favour of the conclusion that people with type 2 blindsight are not having visual experiences and it should be rejected.","Do we have any evidence of the existence of visual experiences that lack colour? We do. Consider, first, sensory substitution in which one sense is used to try to replace another. In cases of sensory substitution attempts are made to deliver information to a subject via a sense that does not usually deliver that information. Consider a particular kind of sensory substitution: tactile-visual sensory substitution (TVSS).24 A camera produces a black and white picture of the world. This image drives a series of pins arranged in a grid that correspond to the image. The grid is placed against an area of a subject’s skin, such as the back or the stomach. The white areas of the image make the pins in the corresponding area of the grid push forward into and/or vibrate against the skin of the subject. After a few hours of practice in which subjects were able to move the camera and receive feedback on what the camera was detecting—for example by the experimenter telling them and by the subject feeling what the camera was pointed at—subjects could recognise a range of common objects, point accurately to objects in space, and judge their distance and absolute size. After about thirty hours they could make complex pattern discriminations, recognise the faces of members of laboratory staff, and display a looming response when the camera lens was zoomed. Subjects’ reports about their experience indicate that initially they are only aware of the tactile stimulus/tactile experience. Then, after practice, they report experiencing stable objects out in the world in front of them, not the tactile stimulus/experience (although they can pay attention to the tactile stimulation if they want and have the tactile experience). Thus, it is said that subjects report their experience in quasi-visual terms. (Bach-y- Rita, 1972, and Guarniero, 1974). One question to ask about the nature of the subjects’ mental lives after they have practiced using the device is whether they come to have a new sensory experience. When subjects report a quasi-visual experience are they reporting some new kind of experience—visual or tactile or otherwise—or are they merely reporting new (often accurate) judgments that they can now make about the world based on ordinary tactile experiences caused by the pins? This question is, noticeably, rather similar to the one that we are asking with respect to those people with type 2 blindsight: are subjects having a perceptual experience or are they just making judgments? I will briefly review the evidence in the case of sensory substitution. A distinctive new quasi-visual experience is reported in many instances of sensory substitution.25 Sensory substitution subjects more readily and consistently attest to a new experience than do people with type 2 blindsight. But how reliable are these reports? There is, unfortunately no objective test to see if someone is reporting accurately. Moreover, there is reason to think that reports about perceptual experience are sometimes not reliable. Consider the fact that there is disagreement even in ordinary cases of perception as to the nature of experience. For example, there have been centuries of philosophical disagreement about whether visual experience is two-dimensional or three-dimensional. Another contemporary example is the debate about how rich perceptual experience is—that is in how much detail does it represents the world.26 Returning to the case of sensory substitution, the reports about experiences when using sensory substitution vary quite dramatically. Some people report vivid colour experiences, some report two-dimensional experience and some three- dimensional experience. Other people don’t obviously report experience at all. On the one hand, this might lead one to doubt that we can trust such reports. However, on the other hand, the disagreement may occur because the experiences or conscious states had by different subjects are actually different. This would be explained by the fact that different sorts of subjects have trained on sensory substitution devices: sighted, late blind, early blind, and congenitally blind people. Moreover, subjects have undergone different training regimes, most noticeably in the length of the training and in the degree of immersion in everyday life of the use of the device. In addition, the motivation of subjects has been markedly different. Some subjects greatly enjoy the use of the device and want to use it. Others do not like it and do not want to use it. There may be other differences, perhaps innate ones, between different subjects. This might mean that some subjects have new perceptual experiences and some do not. In any case, it is clear that we can’t take introspective reports as straightforward evidence in favour of the existence of new perceptual experiences. As with the case of type 2 blindsight, one cannot look to brain imaging to establish whether people using TVSS are having new perceptual experiences. This is because we would have to be confident that any correlations that had been noted between brain activity and perceptual experience were reflective of all instances of perceptual experience. However, cases such as sensory substitution and type 2 blindsight precisely question that. In the case of sensory substitution in particular, there has been much speculation that practice with the substitution device might change the functional role of different parts of the brain, rendering prior apparent correlations otiose. Besides subjective reports and brain imaging, other evidence has been cited in favour of the proposition that those using TVSS devices are having perceptual experience. One piece of evidence is the “looming” response displayed by subjects. When, unbeknown to subjects, experimenters zoomed the camera lens, subjects displayed the reflex action of backing away from something looming towards them. It is claimed that displaying this fast and automatic response attests to fast and automatic processing of the signal, which it is further claimed, is a sign of perceptual processing. Hence, it is argued, subjects have a quasi-visual experience of something rushing towards them. This evidence certainly tells prima facie in favour of subjects having quasi-visual experiences; however, an opponent could argue that there could be fast and automatic inferences being made. Perhaps subjects can make fast and automatic inferences in light of their training that ordinary subjects cannot, and hence do not have quasi-visual experiences—they just have a thought that something is looming towards them. In addition, an opponent could point out that often recognition of objects by subjects using TVSS takes a relatively long time, effort, and attention. For example, a good object recognition performance might consist in a subject identifying an object in ten or more seconds. This is clearly unlike ordinary vision. Another piece of evidence about whether subjects are having a visual experience that should be considered is the fact that blind people report forming new perceptual concepts such as parallax, shadows and interposition of objects after training with TVSS. Someone might try to argue that, as the blind didn’t form these concepts before, it must be that they can do so now because they are having a new perceptual experience. However, one might think that the blind form these new concepts on account of inferences and judgments that they learn to make on the basis of tactile experiences. So while this evidence may also prima facie tell in favour of the idea that subjects are having quasi-visual experiences, it is not conclusive. Finally, one can re-create persisting illusory visual effects using TVSS devices. One can create the Muller-Lyer illusion and the waterfall illusion in subjects using TVSS. Moreover, one can re-create these illusions when the subject knows the effect is illusory. This is often taken to be a key sign of perceptual experience—that it can persist in the face of conclusive counter-evidence—while belief should disappear. However, one can imagine a thought, judgment or belief that it seems as if the world is a certain way persisting in a subject when the subject knows that it is not that way. Thinking, judging or believing that things seem a certain way is compatible with knowledge that the world is not that way. Thus, one could think that subjects had such thoughts, judgments or beliefs, rather than quasi-visual experiences. So again, while suggestive, this evidence is inconclusive. In summary, showing that subjects do have a new perceptual experience, rather than making fast automatic inferences is a real challenge. At the moment, we don’t have a clear answer; however, in my opinion, the weight of the evidence points towards the conclusion that, at least in some cases, subjects have a new quasi- visual experience. Although I have by no means shown it to be true, let us suppose that some subjects do have a new quasi-visual experience. It is worth doing this to reflect on what the nature of that experience would be. Do such experiences represent colour? It is tempting to think that colour—including black, white and grey—is not represented: for subjects only receive pressure stimuli, not chromatic or light stimuli. And even though the pressure stimuli are driven by a camera detecting light and producing black and white images, subjects need not know whether pressure corresponds to blackness or whether it corresponds to whiteness. So how would they, or their brains, know which colour to assign to an object? In particular, it is tempting to think that a congenitally blind person would not have experiences of colour. One might think that a non-congenitally blind person, or their brain, generates experiences of colour on account of previous chromatic experiences and knowledge of colour that they have. They might assign black or white colours to objects drawing on their colour knowledge or, simply, arbitrarily. Or they might assign chromatic colour. For example, if a subject saw a banana shape, their memory of previous yellow bananas might render the banana that they experience yellow. Ward and Meijer (2010) provide evidence that this does happen in some cases of sensory substitution. However, as congenitally blind people have never experienced colour, this method of colour entering quasi-visual experience cannot be what happens in them. Suppose therefore that congenitally blind people don’t experience colour when using TVSS. Could they be having quasi-visual experiences—quasi-visual because they are experiences of form at a distance from the body—that don’t involve experiencing colour? To assess this further let’s think about bat echo-location. One way that bats perceive without using their eyes is to send out a high frequency ‘chirrup’ and listen for the returning echo. Using this sense, bats can detect 3-D objects at a distance from their body. Doing this allows them to negotiate through their environment in the dark, quickly dodging obstacles such as tree trunks and branches, and skillfully catching moths. Clearly bats don’t detect colour—for they are not making use of wavelengths of light—but they do detect form. In virtue of what property do bats experience form? Perhaps bats experience distal form by experiencing sound-reflectance properties. After all, they are detecting form using sound-reflectance properties. If this were the case, then the experience of colour is not necessary in order to experience distal form. But in experiencing distal form does one need to experience some other quality or other? Or could one have a “pure” experience of distal form alone—without experiencing any other quality? If one must experience distal form by experiencing some quality or other, what quality do congenitally blind people using the TVSS experience? One might suppose that they experience light reflectance in the form of luminance—after all it is light that is driving the camera’s responses and hence ultimately the tactile stimulus on the congenitally blind person’s skin. However, congenitally blind people using TVSS need not know what is driving the TVSS. Indeed, we could have built a device where the pins were not driven by camera, but by an echolocatory device. We could set things up so that a congenitally blind person would not be in a position to know whether it is a light sensitive camera that is gathering the relevant information about distal form or whether is an echo-location device that is doing so. Given this, there is reason to believe that they would not be experiencing distal form in virtue of experiencing light-reflectance or luminance. Another suggestion is that congenitally blind TVSS users might be experiencing distal form in virtue of experiencing the pressure that they feel on their skin. One might think this because to some extent there is reason to say that the proximal stimulus acting on the subject is pressure. (I say, “to some extent” because one could make a case that if light is driving the camera then that is the proximal stimulus. Whether pressure or light should be held to be the proximal stimulus is not obvious.) One reason to resist the thought that congenitally blind TVSS users are experiencing distal form in virtue of experiencing the pressure that they feel on their skin is because experienced users of TVSS say that they no longer have tactile experiences of pressure, or at least not ones that they notice. But they do have experiences—and experiences that they notice—of distal form. To make things difficult for us, however, let us suppose, that the users are experiencing both pressure and distal form. (So let us suppose that when they report the absence of such experience they do so only because they are not attending to that experience, not because they are not having that experience.) In one sense, one could say that the users are experiencing distal form in virtue of experiencing pressure because if one took away the experience of the pressure (say by taking away the pressure on their skin) then the subjects would not have the experiences of distal form. However, this is true because the experiences of pressure cause the experiences of distal form. However, when asking whether TVSS users are experiencing distal form in virtue of experiencing the pressure I do not have a causal reading of “in virtue of” in mind here. When I say “in virtue of” I have a phenomenal relationship in mind. To illustrate what this is, consider the following example. Suppose that whenever one had an experience of a red circle, that experience caused one to have an experience of a blue square. In the phenomenal sense that I intend, one does not experience the rectangle in virtue of experiencing the redness—even though the experience of the redness is a cause of the experience of the rectangle. One experiences the rectangle, in the phenomenal sense, in virtue of experiencing the blueness. The form of the rectangle is experienced to be constituted by the blueness. This is the phenomenal sense of “in virtue of” that I have in mind. Here are some other examples of the in virtue of relation obtaining in the phenomenal sense: one auditorily experiences the pitch of a note in virtue of experiencing the volume of the note, one tactually experiences the roundness of a coin on the palm of one’s hand in virtue of experiencing the pressure against of the coin against one’s skin. In all these cases where one thing is experienced in virtue of (in the phenomenal sense) another, the two things are co-located: the blueness of the rectangle and the form of the rectangle, the pitch of the note and the volume of the note, the roundness of the coin and the pressure of the coin. The apparent co-location of qualities would seem to be required for one to be experienced in virtue of (in the phenomenological sense) another. Therefore, I don’t think that it can be true that the congenitally blind users are experiencing distal form in virtue of experiencing pressure. The pressure that is felt is experienced as located on their skin and not at a location in front of the body where the distal objects with their forms are experienced to be. How could the form that is experienced to be at one location be experienced to be constituted by a quality experienced at another location? It could not. If light- reflectance, luminance and pressure are not good candidates for what a congenitally blind person would experience when they experience distal form, and if there are no other good candidate qualities, this should lead us to think that experiences of distal form, without experience of any other quality, are possible. Of course we should remember that we have not ruled out completely the idea that the congenitally blind are not having perceptual experiences when using the TVSS. However, there is some reason to suppose that they are, and if we do suppose it then we have good reason to think that, as there is no good candidate for the property in virtue of which (in the phenomenal sense) they experience distal form, then there is no such property, and none is required. I will call such experiences experiences of “pure distal form”. If there can be such perceptual experiences, then showing that people with type 2 blindsight lack experiences of colour does not entail that they lack perceptual experiences of distal form. I turn now to consider three other cases: a special case of achromatopsia, the phantom contours created by Rogers-Ramachandran and Ramachandran (1998), and instances of amodal completion. The first of these cases is like the case of sensory substitution because there are two competing accounts of the nature of the conscious state that the subject is having, which we cannot settle definitively. The other two cases are, however, more clear-cut. All three of these cases differ from the case of sensory substitution because, unlike it, they do not provide examples of experiences of pure distal form. What they do provide is examples of experiences of distal form with no difference in colour—including black, white and greys—or texture. These are interesting cases that not only lend weight to the idea that there could be cases of perceptual experiences of pure distal form, they also provide us with another plausible account of the nature of the experiences of people with type 2 blindsight. People with achromatopsia cannot see colour. This is tested for by the Farnsworth–Munsell 100-Hue Test, in which subjects are asked to place 100 patches of different hues in order (Farnsworth, 1943). Moreover achromatopsics can neither name nor match colours. They can detect luminance, and so can perceive many distal forms in virtue of differences in luminance. This condition comes in two forms: cerebral achromatopsia (in which an area of cortex that seems necessary for the experience of colour is lost) and retinal monochromatism (in which pigments (blue-cone monochromatism) or cell-classes (rod monochromatism) fail to be expressed in the retina). There is no wavelength specific input to the visual system in retinal monochromats. In contrast, cerebral achrmoatopsics have a normal set of wavelength selective inputs to the visual system. A particular subject with cerebral achromatopsia, MS, has been studied by Kentridge, Heywood, and Cowey (2004). Despite being achromatopsic, he could detect isoluminant borders of different chromatic composition. When a shape was placed against an isoluminant background, which differed only in chromaticity, MS, could detect it. In other words, he could experience distal form (an edge) but not because he was detecting any changes or difference in luminance. These are very odd results. MS is sensitive to a pure chromatic difference despite, in all other respects, being colour blind. This raises interesting questions about MS’s phenomenology when detecting such edges. Kentridge et al. describe MS’s behaviour and experience when looking at isoluminant edges thus: A world without background-invariant colour constancy for MS is not one of ever-changing unstable colour, it is one in which colour is absent and meaningless. This is not to say, however, that discrimination of local contrast is unconscious. For all but the most difficult discriminations MS either deliberated at length over his decisions or made them unhesitatingly. Only on very rare occasions did he need to be prompted to simply make a guess. The neural response to local chromatic contrast therefore produces a percept which is acted upon consciously (1994: 829).27 Suppose that we thought that because MS is achromatopsic he only sees black and white and shades of grey. This is a common view of what achromatopisic vision is like.28 (However, this view of achromatopsic vision can be challenged as I will discuss in more detail below.) Because the border MS can detect is defined chromatically and not by luminance, we have reason to think that there is not a luminance difference—hence that MS will not experience it as defined by black, white or grey. But, at the same time, MS is an achromat and hence cannot order, discriminate, or name colours. So it is tempting to think that the border cannot be experienced by MS in virtue of his experiencing different chromatic colours. If MS doesn’t experience the border in virtue of differences in shades of black, white or gray, and he doesn’t experience chromatic colours, then one might think that MS must experience distal form without experiencing any difference in colour at all—chromatic or achromatic. In other words, one might postulate that MS experiences a uniform achromatic colour, yet some distal form within that achromatic field. If that is right, then the case of MS would provide us with an example of a visual experience of distal form with no difference in experience of (chromatic or achromatic) colour. Such an experience is not an experience of pure distal form, which is an experience of distal form without any experience of colour. It is an experience of distal form whilst having an entirely uniform experience of colour.29 If this is a correct description of the experience, then this is not exactly what we are looking for: an experience of distal form with no experience of colour. However, it would be something very close: an experience of distal form—an edge or boundary at a distance in front of one—with no difference in experienced colour or texture forming the boundary or existing across the experienced boundary. Nonetheless, their existence tells in favour of the idea that there could be experiences of pure distal form without colour for these experiences show that one needn’t experience distal form in virtue (in the phenomenal sense) of experiencing colour boundaries. This is also a property of experiences of pure distal form. Moreover, the existence of this kind of experience suggests a new account of the nature of the experience of those with type 2 blindsight for moving stimuli. Recall that GY reported that his experience was like “black on black” like “a mouse under a blanket” (Beckers & Zeki, 1995: 56, reported by Stoerig & Barth, 2001: 582). Perhaps GY experienced a uniformly black or uniformly dark grey achromatic surface and yet experienced a moving form or a moving edge across it, in a manner similar to the way in which we are supposing MS experiences distal form. However, one might question this account of MS’s experience. I said above that some people question the traditional view that achromats have experiences of the qualities of black, white and grey that humans with ordinary vision do. Akins (2014) argues that they do not. Similarly, when it comes to people with less severe forms of colour blindness—dichromats—the traditional view is that they lack experiences of red and green, and only experience the world in shades of yellow, blue and the greys. (See Broakes, 2010.) Broakes’ own view is that they experience many more colours than yellows, blues and greys, and perhaps even experience all of the colours that normal observers do just not in as many sitations as those with normal colour vision. However, contrary to the traditional view and Broakes’ view, others hold that dichromats experience none of the colours that those with normal colour vision experience. (See Byrne and Hilbert (2010) for arguments for two different versions of this view.) The reason that people have for holding that achromats and dichromats do not have any of the same sort of colour experiences compared to those with ordinary human colour vision is that the workings of their luminance and/or chromatic colour visual systems is so unlike that of those with ordinary human colour vision that it is reasonable to believe that they just have very different sorts of experience of the qualities of the surfaces of objects. Let us agree to call those qualities—chromatic or achromatic—“alien colours”, as Byrne and Hilbert do.30 If it is right to think that those with different luminance and chromatic visual systems, compared to those of normal humans, experience alien colours (both chromatic and nonchromatic alien colours), then we should think that MS experiences alien colours. If that is right, then it opens up an account of his experience that is different to that considered above. Perhaps MS experiences patches of what those with normal human colour vision would experience as being different chromatic but isoluminant colours as having different alien colours. It would still be possible for MS to be unable to name, discriminate, or order what those with normal human vision would call different colours, in the way that normal humans do, because his visual system cannot in general pull apart luminance from chromaticity in the way that the visual systems of people with normal human vision can. If this interpretation of MS is correct then, in one sense, MS is experiencing distal form without colour—ordinary non-alien colour—but, in another sense, he is experiencing distal form with colour—alien colour. Whichever way one thinks it is best to describe the case, it would not be a case of experience of distal form without a difference in some quality or other.31 Deciding between these two hypothesis—that MS experiences boundaries but not in virtue of differences in colour or that MS experiences different in alien colours that define the boundaries—is very difficult. Recall that MS can accurately detect a shape on a background that has the same isoluminance, and that differs only chromatically. However, in addition, given three such shapes on such a background, one of which has a different colour but the same luminance from the others, he can tell the odd one out. One might think that this means that MS must experience a difference in the surface qualities of the different shapes, and therefore that the alien colour hypothesis must be true. However, there is an alternative explanation. It could be that the boundaries between the shapes and their backgrounds are more or less distinct in the different chromatic cases and so the boundaries appear different to MS in some respects, and that it is this boundary information, rather than difference in surface appearance, that MS is relying on to tell the odd one out. A striking finding is that this ability to tell the chromatic odd one out is thwarted when the shapes are surrounded by a thick black border so that there is no direct contiguity between the colour of the shapes that have to be discriminated and their background (because the background and the shapes are each contiguous to the black border). This might tempt one into thinking that MS cannot be experiencing the shapes as having different alien colours. For, if he was, why would this difference in alien colour not persist in the face of the addition of the black boundaries, and allow him to do the task? However, this fact is not decisive. When MS had to pick out the odd one out among three shapes against a uniform achromatic background, all of which had the same chromaticity (they were gray on a gray background), where one of the shapes varied in luminance from the other two shapes, he could, as one would expect, do so. However, when the shapes were surrounded by a black border he could not do the task. This is rather surprising. If one followed the logic of the reasoning in the above paragraph then one would be forced to say that MS doesn’t experience differences in luminance either. If this is right, then perhaps MS only experiences edges and no surface qualities at all—neither chromatic or achromatic! Whether that tallies with his ability to discriminate luminance on other occasions is unclear. It certainly seems as if MS has difficulties comparing luminance and chromatic values between two areas that are not contiguous. This is compatible with him experiencing alien colours, yet only being able to compare and contrast them when particular boundary conditions obtain. Therefore, it is hard to know how to interpret the results of the experiments where there was the addition of the black border in the chromatic case. Finally, Kentridge et al. (2004) point out that the Farnsworth–Munsell 100-Hue Test consists of colours embedded in, and surrounded by, black casings which resemble the black borders used in their experiments described above. Given that the above experiments show that the addition of black borders impedes discovery of the chromatic differences that MS can detect, further experimentation on MS would be desirable using methods that employ sorting tasks where the colours can be compared in conditions where their borders are contiguous. For all we know, MS might consistently sort the colours consistently and in line with some alien colour scheme in those conditions. To summarise the discussion of MS, I have argued that there is one interpretation of MS according to which he has experiences of distal form without any difference in colour (chromatic, achromatic or alien), and another according to which he does not—he has experiences of form in virtue of alien colours. And the former interpretation has two subvarients. One of these is that MS can experience distal form while perceiving uniform achromatic colour. The other is that MS can experience distal form while perceive no chromatic or achromatic colour at all. Deciding between these accounts is in my opinion impossible at present. It would be interesting to try to probe the nature of MS’s phenomenal character in more detail. One could ask him to compare a chromatic isoluminant boundary to an achromatic non-isoluminant one. One could also describe the two different accounts of his visual experience proposed here to him and ask him which, if any, he would be prepared to endorse. And then one could ask him to compare his experience of a chromatic isoluminant boundary to that of the “phantom contour” experiences and experiences of amodal complation described below. Nevertheless, at present, one cannot conclusively say that MS provides us with an example of experience of distal form perception without a difference in colour, but some plausible accounts of him do. I turn now to consider two other cases: the phantom contours created by Rogers-Ramachandran and Ramachandran (1998), and instances of amodal completion. Unlike the cases of MS and sensory substitution where there were two different accounts of the nature of the conscious state being had which we could not decide between, these cases are ones where we can establish that subjects are having perceptual experiences and what their nature is. Neither case is a case of an experience of pure distal form. But they are examples of experiences of distal form with no difference in colour—including black, white and greys—or texture. Rogers-Ramachandran and Ramachandran (1998) created a stimulus that consisted of a uniformly grey background. On one side of the background were white dots. On the other side were black dots, as shown in Fig. 1. The dots flickered in counterphase, so that when the white dots changed to black, simultaneously the black dots changed to white. When the frequency of the changes of the colour of the spots was low (less than 7 Hz), subjects could tell that the spots on one half of the stimulus were in counterphase with those on the other half. And, unsurprisingly, they could indicate where the boundary was between the out of phase spots. However, when the frequency of the flicker of the spots increased to 15 Hz subjects were no longer able to tell that the spots were flickering in counterphase. Thus, phenomenally one would think that the stimulus would have looked to them to be uniform. However, subjects experienced a “phantom contour” where the border was between the spots that were flickering in counterphase. Subjects could also experience movement of the phantom boundary when the stimulus was changed so that which spots were in counterphase was altered. Rogers-Ramachandran and Ramachandran state, “We were quite surprised, therefore, to observe a distinctly visible horizontal border separating the two fields, i.e., one sees a texture border defined by indistinguishable elements. We call this paradoxical percept a phantom contour” (1998: 71). If the spots in question were red and green, rather than black and white, the boundary could be detected in similar conditions, so long as the spots were not isoluminant. The spots had to have different luminance values in order for the effect to occur. Rogers-Ramachandran and Ramachandran claim: Taken collectively, these findings indicate two different systems exist in human vision. One of these is a fast contour-extracting system that can signal contours but not their polarity and the other is a slow system that signals surface qualities. The contour system signals the presence of a border but cannot tell which side of the border is black and which side is white. That is, it can detect that there is a difference between the two sides and also follow high flicker rates, but it cannot signal the direction or the “sign” of the difference. The surface system, on the other hand, can potentially signal the surface characteristics (color and luminance) but at 15 Hz, the speed of flicker is too high for it to follow. Thus, the phantom contour stimulus seems to isolate or selectively activate the fast boundary extracting system. So what is perceived is the output of the contour system alone, a contour defined by two surfaces which look identical. (1998: 74).32 If Rogers-Ramachandran and Ramachandran’s report is correct then not only does the brain register that there is a boundary without registering which properties lie on either side of it, this information is also reflected in the nature of the experience had when looking at the stimulus. One experiences a uniform surface (albeit it one with apparently uniformly flickering dots on it) yet one experiences some boundary. Rogers-Ramachandran and Ramachandran say the border is experienced in virtue of “indistinguishable elements” (1998: 71) that “look identical” (1998: 74). Given that, it would be arbitrary which side of the boundary was signaled to be light and which dark, and given that the brain cannot discriminate the flickering dots, there is good reason to think that their description of the experience is correct. In addition to one interpretation of the case of MS, phantom contours provide another example of an experience of distal form with no difference in experienced colour or texture forming the boundary or existing across the boundary. As mentioned previously, such cases lend weight to the supposition that there could be experiences of pure distal form without colour, for in these experiences it is not in virtue of experiencing colour boundaries that one experiences form. Thus they lend weight to the supposition that colour is not a structural feature of experience. In turn this backs up the idea that showing that people with type 2 blindsight lack experiences of colour does not show that they lack visual experience. Moreover, as also mentioned previously, the existence of this kind of experience suggests that one account of the experiences of those with type 2 blindsight could be accurate: that they experience a uniformly black or dark grey background and yet form and movement within. This case of phantom contours provides evidence not just of form, as the case of MS does, but of movement perception in such conditions. This is because Rogers- Ramachandran and Ramachandran altered the location of the phantom contour by altering the proportion of black to white dots. Subjects reported experiencing the phantom contour moving across their visual field. Finally, I turn to consider a last case: amodal completion. In order to explain this example, I will first explain what modal completion is. Consider the Kanizsa triangle in Fig. 2.33 Ordinary perceivers report experiencing a bright white equilateral triangle pointing towards the top of the page that is lighter than the background, and is hence defined by lightness boundaries within experience. That triangle is experienced as partially occluding another triangle pointing towards the bottom of the page that is defined by black lines. Three “pacman-like” figures are also experienced as occluded circles. On close inspection of the figure, one can come to realise that part of this experience is illusory. There are no lightness boundaries forming a triangle pointing towards the top of the page. What is important to notice, for our purposes, is that when one experiences that illusory triangle, one has an experience as of edges created by a luminance boundary. We know that there are no such luminance edges and there is not a difference in lightness where there appears to be one; but that is what we experience. Cases such as this are cases of modal completion because although the edges that seem to form the upward pointing triangle do not exist, one has a visual experience that represents such edges in a way that typical visual experiences do: in virtue of a colour, lightness or texture boundary. With the nature of this case of modal completion firmly in mind, consider now a case of amodal completion. Consider Fig. 3. When asked, people say that they have a strong sense that it consists in a square that continues behind an occluding circle—hence, that it seems as if the shapes in Fig. 4 are present. However, Fig. 3 is perfectly compatible with the shapes shown in either Fig. 5 or Fig. 6 being present—and hence with a square with a corner removed or with a square with a jaggy protruding extension to a corner being present—as well as with a square being present (Michott et al., 1964/1991). The visual experience of Fig. 3, however, does not consist in experienced boundaries consisting of colour, lightness or texture corresponding the occluded portion of the square. Yet, nonetheless, it is a square that is reported as being present. This is a case of amodal completion, and it contrasts with modal completion in that it occurs when part of an object is experienced as occluded and is reported as having one of many possible shapes, yet the occluded portion of the object is not experienced as being defined by colour, lightness or texture boundaries. There are two different interpretations of amodal completion. One is that such visual experiences are only of the lines that make up Fig. 3 on the uniform background and that the shape of the occluded figure is inferred, yielding a judgment, thought or belief that an occluded square is present. The second interpretation is that the visual experience is of the lines that make up Fig. 3 on the uniform background and, in addition, the occluded part of the square is represented in the visual experience. A great deal of psychological research has gone into determining which interpretation is correct. Michotte, Thines, and Crabbe (1964/1991) themselves noted that amodal completion occurs despite subjects’ beliefs—indeed knowledge—that the occluded figure is not the way that their experience tells them it is. For example, if one first sees that a square with a protruding jaggy corner, as shown in Fig. 6, is present and then it is occluded by a circle, then what is seen will still be experienced as an occluded square, not an occluded square with a protruding jaggy corner. Moreover, that experience persists even if a lot of jaggy protruding cornered squares are experienced. Looking at Fig. 7 should be illustrative. Likewise, if one draws a broken triangle as in the left-hand side of Fig. 8 and then places a pencil over it in the manner depicted in the right-hand side of Fig. 8 then one experiences a completed occluded triangle even when one knows it not to be such. Using a visual search paradigm, studies have shown conclusively that amodal completion can occur without focused attention and within the time span associated with early visual processing (Enns & Rensink, 1998). And, in a study directly measuring cells’ response in nonhuman primates, neurophysiological data show that “cells as early as V1 have the computational power to make inferences about the nature of partially invisible forms seen behind occluding structures” (Sugita, 1999). These results and others are summarised in Wagemans, Lier and Scholl (2006) and strongly suggest that the occluded parts of objects are experienced visually—and are not inferred. Further evidence of the genuinely experiential nature of amodal occlusion phenomena is given by Briscoe (2011). Consider Fig. 9 that is typically experienced as a number of unconnected two-dimensional forms that lie on the same two-dimensional plane of depth. When these elements appear to be occluded by the insertion of oblongs into the picture, as in Fig. 10, the forms previously seen are no longer experienced as two-dimensional, unconnected, and on the same two-dimensional plane of depth. They are experienced as connected, three-dimensional pieces that form a three- dimensional cube. This change in one’s conscious experience provides as clear a demonstration as one could hope for that changes in one’s experience, rather than thought or propositional attitudes, are brought about by amodal completion. If this is right, then in amodal completion we visually experience occluded distal form at a distance from our bodies but not in virtue of experiencing an edge, border or form that is experienced in virtue of colour, lightness or texture. Such experiences are therefore very similar to the experiences of phantom contours that I discussed above. Again, these experiences are not ones of pure distal form without colour, but experiences of distal form without an experience of a difference in colour, lightness or texture. Nevertheless, as in the case of MS and the phantom contours, the existence of these experiences lend weight to the idea that there could be pure distal experiences of form. This is because, in these experiences, form is experienced but not in virtue of differences in colour, lightness or texture. Moreover, they may themselves be the sorts of experience that those with type 2 blindsight have. It would be interesting, and potentially informative, to ask people to compare and contrast the phenomenal character of their experiences of amodal completion with their experiences of phantom contours. Similarly, it would be interesting to get MS to compare his experience of form defined by differently coloured but equiluminous areas with his experiences of these other phenomena. I have discussed four types of unusual perceptual experience. I argued that if people trained to use TVSS do have new perceptual experiences, then the best account of the experience of a trained congenitally blind person who was using such a device is that they have a perceptual experience of pure distal form without colour. The case of MS has two interpretations. On one he has a perceptual experience of distal form in virtue of alien colours. On another he experiences distal form without experience a difference in colour (chromatic, achromatic, or alien). I have also argued that the case of phantom contours and amodal completion show that there can be experiences of distal form without any differences in colour, lightness or texture. The existence of these experiences lends weight to the suggestion that there could be perceptual experiences of pure distal form with no experience of colour (including luminance) or alien colour. Moreover, they suggest another description of the perceptual experiences of those with type 2 blindsight for moving stimuli: experiences of form and/or movement within a field of uniform colour, lightness and texture. We therefore have two plausible candidates for the nature of the perceptual experiences of those with type 2 blindsight: experiences of pure distal form and experiences of distal form without experience of difference in colour, lightness, or texture (but not a complete lack of colour, lightness or texture). Recall that Brogaard (2011, 2012) held that there cannot be visual experiences of pure distal form and hence that we should reject the idea that those with type 2 blindsight are having perceptual experiences. She holds that we should think that they are having thoughts instead. I have argued that there is no good reason to maintain that people with type 2 blindsight cannot be having perceptual experiences. Are such perceptual experiences visual experiences? One can imagine someone claiming a priori that for an experience to be visual it must be an experience of colour. But I think that such a view would merely be stipulative: a linguistic decision taken on no good grounds. What reason could one have to hold this view, rather than the view that experience of distal form is sufficient for an experience to be visual? I do not believe that there is any. One might think that one could appeal to the fact that these experiences are caused by light, which is the proximal stimulus of vision and caused by the use of the eye, which is the sensory organ of vision. And to the extent that these are relevant, they point towards the experience being visual. However, as I have already argued above, the modality of an experience is not completely determined by facts such as these, as the case of synaesthetic concurrent experiences show. What is represented by the experience—distal form—is that which is represented by only visual experiences among the human senses, and plausibly the phenomenal character of the experience is most like visual experiences. But are those things enough to make the experience visual? I do not think that there is any fact of the matter here. I do not have space to argue it here, but as I have argued in Macpherson (2011a), not only may the distinctions between the sensory modalities be a matter of degree, the distinctions between kinds of experience may be too. Whether or not one withholds the epithet of “visual” from such experiences, they are perceptual experiences and, of all kinds of human experience, most like visual experiences—due to the fact that they are experiences of distal form, and caused by light stimulating the eyes. Showing that there is no good reason to rule out that those with type 2 blindsight have perceptual experiences is the crux of the matter in this paper, not whether they should be classed as visual perceptual experiences or perceptual experiences in some other modality. Thus we can resist the idea that we are forced to conclude that those with type 2 blindsight are only having thoughts about movement or distal form, for resisting depends on showing that they could be having perceptual experiences, not visual experiences per se.","In section four, we saw that people have argued about the nature of the conscious state that is had by people with type 2 blindsight. I argued that the experimental evidence available at present does not settle the matter. I then discussed the argument put forward by Overgaard and Grünbaum (2011) for the conclusion that people with type 2 blindsight must be having visual experiences and argued that it was not sound. This was because it relies on the false premise that the conscious state must be a visual experience because it was caused by a visual stimulus, a visual sensory organ, and early visual processing. This is false because thoughts, beliefs, and perceptual experiences in non-visual modalities, can also have those causes. I then examined the argument given by Brogaard (2011, 2012) for the conclusion that those with type 2 blindsight are having thoughts based on their guessing what is before them. I argued that we lack good evidence that guessing is required to produce the conscious state that people with type 2 blindsight report, and even if it were, that is not enough to show that those people are not having visual experiences. That is because the guessing might cause the occurrence of those visual experiences—perhaps by focusing attention or by some other means. Finally, I turned to address Brogaard’s second argument to the conclusion that those with type 2 blindsight were having thoughts and not visual experiences. Brogaard’s second argument supposed that those who were experiencing type 2 blindsight could not be having visual experiences because they experienced pure distal form, that is form without colour (which I stipulated to include black, white and grey). I showed that there is a long intellectual tradition of supposing that visual experiences must represent colour in order to be visual. I noted that colour was one alleged structural feature of visual experience. I also argued in section two, that it is very difficult to establish the existence of any given alleged structural feature of experience and that, just because one has not had an experience that lacks an alleged structural feature, that is not reason enough to establish that the alleged structural feature is indeed one. The example of experiences of novel colours was illustrative. I set out, in section five, to examine the evidence concerning whether there could be visual experiences in the absence of experience of colour by looking to see if there were any experiences of pure distal form. I argued that there is some, although not conclusive, reason to think that perceptual experiences of the congenitally blind using TVSS might be of this ilk. I also showed that one plausible interpretation of the nature of the experiences of the achromat MS, when he looks at a form delineated by two equiluminant areas that are different in chromatic colour, is that they are of pure distal form. However, there was another interpretation of the nature of the experience of MS that was at least as plausible, and that did not have this consequence. I suggested that further investigation of MS would be instructive, in particular, asking him to compare his experiences of chromatic boundaries with no luminance difference with his experience of phantom contours and amodal completion. I then showed that the evidence concerning phantom boundaries and amodal completion clearly shows that there could be perceptual experiences of distal form and movement with no difference in colour, lightness or texture differences. And I argued that the existence of such experiences was evidence in favour of thinking that there can be experiences of pure distal form with no colour, lightness or texture. This is because, in such experiences, form and movement are not experienced in virtue of experiencing colour, lightness or texture boundaries. Moreover, I showed that experience of form and movement, or just movement, with no difference in colour, lightness or texture, was a second very good candidate for being the sort of experience that people with type 2 blindsight are having, based on their reports of experiencing darkness and black shadows (in addition to the experiences already considered or of pure distal form and/or movement). If I am right and such experiences are possible, then, contra Brogaard, there is no good reason to doubt that the reports of people with type 2 blindsight could be accurate descriptions of perceptual experiences. Thus, Brogaard’s reason for holding that those with type 2 blindsight are not having visual experiences should be rejected. I discussed whether experiences of pure distal form and experiences of distal form with no difference in colour or texture should be thought of as visual. I said that they are more like ordinary visual experiences than experiences in any other modality. Moreover, I claimed that there is no good reason to deny that such experiences are visual. Denying it would merely be a stipulative manoeuvre. In any case, what is important is whether there is any good reason to deny that the conscious states of those with type 2 blindsight are perceptual experiences, and I have argued that there is not. The matter was only originally discussed in terms of visual experiences as that modality seemed the most likely one for the experiences to belong to. Removing Brogaard’s reasons for denying that people with type 2 blindsight are having perceptual experiences does not establish that they are doing so. However, it leaves open the possibility that they are. Further investigation of the phenomenon is clearly called for. In light of the arguments given in this paper, I suggest that it would be good to give those with type 2 blindsight experiences of phantom contours and amodal completion in their non-blind visual fields, and ask them to compare those experiences to the conscious states that they report when their blind field is stimulated, and to see if they are willing to accept that there is a phenomenal match—that those experiences are the same subjectively. In conclusion, I have taken steps forward in the investigation of three alleged structural facts about experience: the necessary experience of colour in visual experience the necessary experience of colour in experience of distal form and movement, and the necessary experience of difference in colour in experience of distal form and movement. I have provided evidence against each being true. Even if some of these steps turn out to be small steps forward, that any steps have been taken makes them significant ones, given the difficulties that attend establishing or dismissing the structural features of experience."],["Preterm birth (<37 weeks' gestation) is sometimes associated with poorer outcomes in adulthood (e.g., poorer health, fewer intimate relationships, and lower income). However, few studies have examined how these adults felt about their lives or how personality affected these associations. 11,592 preterm and 51,460 full term adults completed online surveys measuring their subjective well-being (life, relationship and job satisfaction, and health). Adults born preterm reported similar levels of relationship satisfaction, but poorer health, life satisfaction and job satisfaction. Adults who reported having long hospital stays at birth also reported poorer health, life, relationship and job satisfaction, and this poorer well-being appeared to be accounted for, in part, by factors such as their personality. --------------------------------------------------------------------------------","Outcome studies have demonstrated that preterm birth (before 37 weeks’ gestation) is associated with poorer health (Cooke, 2004; Hack, 2009; Hack, Cartar, Schluchter, Klein, & Forrest, 2007; Lindstrom, Lindblad, & Hjern, 2009), reduced likelihood of forming romantic relationships (Hack, 2009; Moster, Lie, & Markestad, 2008; Wolke, 2011), poorer educational attainment, and lower salaries (despite equal levels of employment; Cooke, 2004; Hack, 2009; Lindstrom, Winbladh, Haglund, & Hjern, 2007; Moster et al., 2008; Wolke, 2011). Small samples tend to involve rich data collected from individuals over time (often starting soon after birth) but do not have sufficient power to explore the role of confounding variables or allow subgroup analyses (Saigal, 2013). In comparison, the large national register samples allow such analyses and include information about objective variables such as educational attainment, employment status, health and living situation (living with a partner, peers or parents; Lindstrom et al., 2009, 2007; Moster et al., 2008) but not subjective evaluations of these circumstances (Saigal, 2013). In order to have a sample comparable in size to that of national register samples, which also included subjective assessments of life, relationship and job satisfaction, we collected data using an online survey. This method allowed us to collect information from a larger number of individuals. The resulting sample closely resembled the general population of UK (Rentfrow, Jokela, & Lamb, 2015), and was large (the subsample studied here included over 60,000 adults) so ensured sufficient power to examine the role of potential covariates and to compute robust effect sizes. Subjective wellbeing reflects individual beliefs and feelings and is related to health and social relationships (Diener, 2012). Therefore, the first aim of this study was to understand the adult sequelae of preterm birth with an emphasis on subjective accounts of their health, relationships, jobs and lives. Despite their importance, little is known about the subjective well-being of adults born preterm (Guyatt & Cook, 1994; Saigal, 2013) although the quality of life experienced by preterm and full term-born individuals tend to be similar when rated by the individuals rather than their parents (Cooke, 2004; Hack, 2009; Hack et al., 2007; Roberts et al., 2013; Saigal & Tyson, 2008; Zwicker & Harris, 2008). Less is known about life satisfaction, which involves comparing current circumstances with subjective standards set by the individuals themselves (Diener, Emmons, Larsen, & Griffin, 1985). Adults born small for gestational age had levels of life satisfaction similar to those of individuals with normal birthweight in one study (Strauss, 2000), but similar data are not available for adults born preterm. In terms of health, adults born preterm tend to report poorer health outcomes than those born at full term. In one study, for example, both male and female adults born preterm had lower levels of physical functioning, female preterm adults reported role limitations due to emotional problems, mental health and lack of energy, and male preterm adults perceived their general health to be poorer (Cooke, 2004). Preterm adults were also more likely to have chronic conditions (primarily, asthma), be hospitalized for psychiatric conditions, have a diagnosis of ADHD or autism spectrum disorder, and be taking prescription medicines (Cooke, 2004; Hack, 2009; Hack et al., 2007; Johnson & Marlow, 2014; Lindstrom et al., 2009). Therefore, prematurity appears to be related to poorer objective and self-reported health. However, a sample of 55 adults born preterm reported slightly better physical health than community norms in a recent study (although mental health appeared worse than community norms in the sample; Natalucci et al., 2013). Research on the socio-economic sequelae of prematurity has generally focused on objective outcomes (for example, relationship or employment status) rather than how individuals feel about their relationships or jobs. Studies focused on objective measures have shown that individuals born prematurely are less likely to form romantic relationships, start co-habiting with partners, find life partners, or become parents (Hack, 2009; Moster et al., 2008; Wolke, 2011). Cooke (2004) found similar proportions of preterm and full term adults in intimate relationships and in sexual relationships, but evidence regarding the effects on relationship satisfaction is less clear-cut. Relationship satisfaction has not been directly measured in adults born preterm although adults with very low birthweight reported less attachment-related anxiety than their normal weight peers (Pyhala et al., 2009). Preterm adolescents reported having fewer social interactions than full term adolescents, although they were rated equivalently adequate (Hallin & Stjernqvist, 2011). Furthermore, studies of objective socio-economic outcomes have demonstrated that preterm adults complete less schooling and leave school earlier than their peers (Cooke, 2004; Hack, 2009; Lindstrom et al., 2007) and despite being equally likely to be employed, preterm individuals had lower salaries (Hack, 2009; Lindstrom et al., 2007; Moster et al., 2008; Wolke, 2011). Because previous work has only focused on these objective measures of employment, this study asked adults born preterm about their job satisfaction. The second aim was to understand how other variables affected the associations between preterm birth and outcomes. Personality may mediate relations between preterm birth and adult outcomes (Hack, 2009; Wolke, 2011), although conclusive evidence is lacking. Such hypotheses are based on findings that adults born preterm tend to score higher on measures of shyness, agreeableness, conscientiousness and neuroticism, and lower on measures of extraversion (Allin et al., 2006; Hertz, Mathiasen, Hansen, Mortensen, & Greisen, 2013; Pesonen et al., 2008; Schmidt, Miskovic, Boyle, & Saigal, 2008). This personality profile may help explain some of the differences in adult outcomes. For example, higher conscientiousness is related with better health and better longevity (e.g., Friedman & Kern, 2014; Friedman, Kern, Hampson, & Duckworth, 2014; Roberts, Walton, & Bogg, 2005; Smith, 2006) while increased neuroticism is related to the increased diagnosis of illness (although results are more mixed for relations between neuroticism and health; Friedman & Kern, 2014; Smith, 2006). In addition, increased extraversion and decreased neuroticism were related to higher levels of subjective well- being (Diener, Suh, Lucas, & Smith, 1999). Therefore, the higher neuroticism and lower extraversion of adults born preterm may, for example, place these individuals at risk of poorer health and subjective well-being; whereas the higher conscientiousness of adults born preterm may help to protect these individuals from poorer health outcomes associated with prematurity. These differences in personality may also mediate relations between prematurity and the later formation of romantic relationships (Hack, 2009; Wolke, 2011) although conclusive evidence is lacking. One goal of the present study was thus to explore the extent to which personality accounts for some of the adult outcomes of prematurity. In addition to differences in personality profiles, Hack et al. (2002) described the behavioral cautiousness of preterm individuals. Risky behavior, or its avoidance, may therefore also account for some adult outcomes (Wolke, 2011), with preterm adults appearing more risk adverse, and less likely to consume alcohol, use drugs, and go to clubs or pubs than full term adults (Cooke, 2004; Hack, 2009; Hack et al., 2007, 2002, 2004; Pyhala et al., 2009; Roberts et al., 2013; Schmidt et al., 2008). Wolke (2011) questioned whether such behavioral cautiousness might reduce opportunities and thus help explain some adult outcomes of prematurity, so we examined the extent to which risky behaviors accounted for adult sequelae of prematurity. Finally, demographic and childhood factors may play important roles in understanding adult outcomes. Premature delivery not only places babies in the world before they are biologically ready, but is also often combined with long periods of hospitalization following birth (Goldberg & DiVitto, 1983). Given long hospital stays at birth may reflect more extreme preterm birth or more medical complications at birth, we examined long hospital stays at birth as a predictor of adult outcomes (as well as birth status). In addition, premature deliveries occur more often among mothers of low socioeconomic status, who are under 15 years old, or who have had many pregnancies close together in time (Behrman & Butler, 2006; Goldberg & DiVitto, 1983). These conditions themselves place children at increased risk. Therefore, preterm delivery may combine biological immaturity with environmental risk. As a result, demographic (family-of-origin socioeconomic status) variables need to be considered when seeking to assess the specific effects of preterm birth on adult outcomes. In addition, preterm adults appear less likely to have children (Moster et al., 2008), which may affect the outcomes of interest (for example, relationship satisfaction) in this study. Therefore, being a parent also needs to be considered as a potential confounding variable for adult outcomes. Study sample ~~~~~~~~~~~~ A sub-sample of cases was drawn from data collected using an online survey advertised (through webpages, and television and radio channels) and hosted by the British Broadcasting Corporation. 556,330 participants, aged between 18 and 80 years, responded between November 2009 and April 2011. Participants were excluded if data were missing (either due to no response or responses such as “rather not say” or “don’t know”) regarding preterm birth or confounding variables, resulting in a subsample of 63,052. Individuals without missing data were younger, t(84429.77) = −43.20, p < .001, d = 0.17, were 1.21 times more likely to be female, χ2 (N = 537,077) = 462.20, p < .001, and were higher on family-of-origin SES, t(90957.62) = 22.92, p < .001, d = 0.10. Models were then run on the subsamples of individuals without missing data for the outcomes of interest. Only individuals in relationships and those who were employed were included in analyses of relationship satisfaction and job satisfaction, respectively.","‘The Big Personality Test’ contained items pertaining to demographic and life histories (childhood, health, education, employment), personality, and well-being, among other topics (see the supplementary materials for the relevant sections of the survey). Before starting the survey, individuals were provided with information about the study and were informed of their right to withdraw at any time and to ignore any questions. All study procedures were reviewed by the Department of Psychology Research Ethics Committee (in the University of Cambridge). All items were answered using a multiple-choice format, producing numeric data. All variables (unless otherwise stated) reflect the average of items making up the respective scale. Preterm birth ~~~~~~~~~~~~~ Respondents were asked if they were born preterm – before 37 weeks (see Section 6, supplementary materials). Respondents answered yes, no or rather not say/do not know. Those who selected the last option were excluded. Seven per cent of respondents were born preterm (6.8% of all respondents and 7.1% of those who answered this question), a rate equivalent to national statistics for the U.K. (Office for National Statistics, 2011). Birth status was related to having missing data, χ2 (1, N = 445,769) = 14357.70, p < .001, with preterm adults 4.11 times more likely to have complete data than full term individuals. Consequently, preterm adults made up 18% of the final sample. Outcomes ~~~~~~~~ Health was measured using only the first item of the RAND SF 36-Item Health Survey (Ware, 2004). Respondents were asked to rate one item about how they perceived their health to be in general on a five-point scale from ‘excellent’ to ‘poor’ (see Section 7, supplementary materials). Life satisfaction was measured using the Satisfaction with Life Scale (SWLS; Diener et al., 1985). Participants selected responses on a seven-point scale that ranged from ‘strongly agree’ to ‘strongly disagree’ in response to 5 items pertaining to their feelings about various aspects of their lives (see Section 8, supplementary materials). For example, So far, I have got the important things I want in life. The internal consistency of the scale was high (α = .90). Relationship satisfaction was measured using the Adapted Triangular Love Scale (Ahmetoglu, Swami, & Chamorro-Premuzic, 2010). Respondents first indicated whether they were in intimate relationships. Those in relationships selected responses on a five-point scale that ranged from ‘disagree strongly’ to ‘agree strongly’ in response to 9 items pertaining to their satisfaction with three aspects of their relationships: intimacy, passion and commitment (see Section 4, supplementary materials). For example, I think my relationship with my partner will last forever or I can tell everything to my partner. The internal consistency of the scale was high (α = .84). Job satisfaction was measured using the General Index of Job satisfaction (Brayfield & Rothe, 1951). Respondents first indicated their occupational status. Those in employment (full time, part time or self-employed) selected responses on a five-point scale that ranged from ‘disagree strongly’ to ‘agree strongly’ in response to 5 items pertaining to their feelings about their job (see Section 2, supplementary materials). For example, I like my job better than the average person does. The internal consistency of the scale was high (α = .91). Control variables ~~~~~~~~~~~~~~~~~ Personality was measured using the Big Five Inventory (John, Donahue, & Kentle, 1991). Participants selected responses on five-point Likert scales that ranged from ‘strongly disagree’ to ‘strongly agree’ to specify the extent that 44 statements describing personality characteristics applied to them (see Section 3, supplementary materials). The internal consistency of subscales was high (α ⩾ .77). Risky behaviors were measured using 3 items from the CDC Youth Risk Behavior Survey (Centers for Disease Control, 2013). Respondents were asked how many cigarettes they smoked per day in the last month. Individuals were categorized as no or yes (any number per day) for smoking. Respondents were asked how many days in the past month they drank 5 or more drinks within a couple of hours. Responses to this item were used to create an ordinal variable of increasing binge drinking (see Table 1 for levels). Finally, respondents were asked how many times in their lifetime they had taken illegal drugs. Responses to this item were used to create an ordinal variable of increasing life-time use of illegal drugs (see Table 1 for levels and see Section 7 of the supplementary materials for items). Child background factors and control variables were measured using single items about the respondents’ age, gender, ethnicity (see Section 1, supplementary materials), whether they thought they had experienced long hospitalizations at birth (see Section 6, supplementary materials) and whether they were parents (see Section 4, supplementary materials). Family-of-origin SES was calculated by standardizing maternal and paternal education, and the occupation of primary breadwinners during the respondents’ childhoods (see Section 5, supplementary materials), and then calculating the mean score for these three standardized variables. Statistical analyses First, a series of multiple regressions separately explored relations among birth status or long hospital stays at birth (independent variables), control variables, and outcomes (dependent variables). Then, for the four outcomes, a series of regression models were run. The first model included birth status and long hospital stays at birth as predictors. Model 2 added personality (all 5 subscales); model 3 added risky behaviors; model 4 added family-of-origin SES; and model 5 added parenthood. All models controlled for age, gender and ethnicity. All analyses were run using R (R Core Team., 2013b) and R packages foreign (R Core Team, 2013a), lsr (Navarro, 2014), and psych (Revelle, 2014). Relations between preterm birth or long hospital stays at birth and other variables ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 1 presents the descriptive statistics for all variables for the preterm and full term samples. Table 2 presents the regression coefficients for models examining differences between the preterm and full term samples, and between individuals who did and did not have long hospital stays at birth. Preterm individuals were lower than full term respondents on openness and extraversion, and higher on conscientiousness, agreeableness and neuroticism. Preterm individuals were less likely to binge drink, use illegal drugs or be parents, but were more likely to have had long hospital stays at birth and have grown up in families with lower SES. In addition, preterm adults reported lower health, life satisfaction and job satisfaction. Effect sizes were small, with Cohen’s d ranging from −0.13 to 0.07 with the exception of long hospitalization (d = 1.10). No differences between preterm and full term adults were found for smoking or relationship satisfaction. Furthermore, individuals who reported long hospital stays at birth showed similar pattern of results to preterm individuals with a few exceptions. Individuals who had long hospital stays at birth showed lower conscientiousness (the reverse of preterm individuals), but showed no difference in their openness and agreeableness to individuals who did not have long hospital stays at birth. There was no difference in family-of-origin SES between individuals who did and did not have long hospital stays at birth, but individuals who did have long hospital stays at birth reported lower relationship satisfaction, and furthermore long hospitalizations at birth appeared to have a stronger effect on all outcome measures than preterm birth. The role of preterm birth, long hospital stays and confounding variables ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 3 presents the effect size – Cohen’s d – for birth status and long hospital stays in the regression models predicting: health, life satisfaction, relationship satisfaction and job satisfaction. Table 1 in the supplementary materials reports correlations between all predictor variables in the regression analyses. Health One per cent of the variance was accounted for by preterm birth and long hospital stays (as well as age, gender and ethnicity), R2 = .01, F(5, 63046) = 87.72, p < .001. Adding personality variables, ΔR2 = .11, F(5, 63041) = 1586.70, p < .001, risky behavior variables, ΔR2 = .02, F(3, 63038) = 377.48, p < .001, childhood variables, ΔR2 = .00, F(1, 63037) = 205.20, p < .001, led to significant increases in the amount of variance accounted for, but adding parenthood, ΔR2 = .00, F(1, 63036) = 2.71, p = .100, did not increase the amount of variance accounted for. Adding interactions between preterm birth or long hospital stays and all other predictor variables added minimally (0.05%) to the variance accounted for. Openness (b = −0.04, 95% CI [−0.05, −0.03], β = −0.03, p < .001) and neuroticism (b = −0.29 [−0.30, −0.28], β = −0.23, p < .001) negatively predicted, and conscientiousness (b = 0.19 [0.18, 0.20], β = 0.13, p < .001) and extraversion (b = 0.10 [0.09, 0.11], β = 0.08, p < .001) positively predicted health, independently of preterm birth and long hospital stays at birth. In addition, smoking (b = −0.30 [−0.32, −0.28], β = −0.12, p < .001) and drug use (b = −0.01 [−0.01, 0.00], β = −0.02, p < .001) negatively predicted, and family-of-origin SES positively predicted (b = 0.09 [0.08, 0.10], β = 0.06, p < .001) health, independently of preterm birth and long hospital stays at birth. Adding control variables into the regression models did not reduce the effect of preterm birth but did reduce the effect of long hospital stays on health. In particular, including personality variables (and risky behaviors to a smaller extent) reduced the unique association between long hospital stays on health. Life satisfaction One per cent of the variance was accounted for by preterm birth and long hospital stays at birth (as well as age, gender and ethnicity), R2 = .01, F(5, 62194) = 184.70, p < .001. Adding personality, ΔR2 = .21, F(5, 62189) = 3348.90, p < .001, risky behavior, ΔR2 = .02, F(3, 62186) = 422.44, p < .001, childhood variables, ΔR2 = .00, F(2, 62185) = 268.59, p < .001, and parenthood, ΔR2 = .01, F(1, 62184) = 821.04, p < .001, led to significant increases in the amount of variance accounted for. Adding interactions between preterm birth and all other predictor variables added minimally (0.04%) to the variance accounted for. Conscientiousness (b = 0.24 [0.23, 0.26], β = 0.13, p < .001), extraversion (b = 0.29 [0.27, 0.30], β = 0.17, p < .001) and agreeableness (b = 0.14 [0.12, 0.16], β = 0.06, p < .001) positively predicted, and openness (b = −0.09 [−0.11, −0.08], β = −0.04, p < .001) and neuroticism (b = −0.49 [−0.50, −0.47], β = −0.29, p < .001) negatively predicted life satisfaction; smoking (b = −0.37 [−0.40, −0.34], β = −0.11, p < .001) and drug use (b = −0.02 [−0.03, −0.02], β = −0.04, p < .001) negatively predicted life satisfaction; and family-of-origin SES (b = 0.15 [0.13, 0.16], β = 0.07, p < .001) and parenthood positively predicted (b = 0.38 [0.35, 0.40], β = 0.13, p < .001) life satisfaction, independently of preterm birth and long hospital stays at birth. Adding control variables into the regression models did not reduce the effect of preterm birth (with the exception of parenthood reducing the effect of birth status from −0.08 to −0.07) but did reduce the effect of long hospital stays on life satisfaction. In particular, including personality variables (and risky behaviors to a smaller extent) reduced the unique association between long hospital stays on life satisfaction. Relationship satisfaction Preterm birth was related relationship status (whether individuals were in an intimate relationship or not), χ2 (1, N = 62,333) = 286.78, p < .001. Full term individuals were 1.43 [1.37, 1.49] times more likely to be in intimate relationships than preterm individuals. For individuals in an intimate relationship, we examined their satisfaction with their current relationship. One per cent of the variance in relationship satisfaction was accounted for by preterm birth and long hospital stays at birth (and age, gender and ethnicity), R2 = .01, F(4, 41635) = 68.33, p < .001. Adding personality variables, ΔR2 = .07, F(5, 41630) = 633.58, p < .001, risky behavior variables, ΔR2 = .01, F(3, 41627) = 212.05, p < .001, led to significant increases in the amount of variance accounted for. However, adding childhood variables, ΔR2 = .00, F(2, 41626) = 1.06, p = .304, and parenthood did not lead to a significant increase in the amount of variance accounted for, ΔR2 = .00, F(1, 41625) = 0.27, p = .601. Adding interactions between preterm birth and all other predictor variables added minimally (0.09%) to the variance accounted for. Conscientiousness (b = 0.13 [0.12, 0.14], β = 0.13, p < .001), extraversion (b = 0.06 [0.05, 0.07], β = 0.07, p < .001) and agreeableness (b = 0.16 [0.15, 0.17], β = 0.14, p < .001) positively predicted, and neuroticism (b = −0.05 [−0.06, −0.04], β = −0.06, p < .001) negatively predicted relationship satisfaction; and smoking (b = −0.12 [−0.14, −0.10], β = −0.07, p < .001), binge drinking (b = −0.03 [−0.04, −0.03], β = −0.07, p < .001) and drug use (b = −0.01 [−0.02, −0.01], β = −0.04, p < .001) negatively predicted relationship satisfaction, independently of preterm birth and long hospital stays at birth. Adding personality variables into the regression models reduce the effect of preterm birth to insignificance and reduced the effect of long hospital stays on relationship satisfaction. In addition, controlling for risky behaviors increased (to a small extent) the unique association between long hospital stays on relationship satisfaction. Job satisfaction Preterm birth was related to employment status, χ2 (2, N = 61,386) = 360.60, p < .001. Full term individuals were 1.41 [1.35, 1.47] times more likely to be employed than preterm individuals. For individuals in employment, we examined their satisfaction with their current job. Three per cent of the variance was accounted for by preterm birth and long hospital stays at birth (and age, gender and ethnicity), R2 = .03, F(4, 38618) = 217.60, p < .001. Adding personality, ΔR2 = .10, F(5, 38613) = 929.10, p < .001, risky behavior, ΔR2 = .00, F(3, 38610) = 25.83, p < .001, childhood variables, ΔR2 = .00, F(2, 38609) = 61.72, p < .001, and parenthood, ΔR2 = .00, F(1, 38608) = 38.04, p < .001, led to significant increases in the amount of variance accounted for. Adding interactions between preterm birth and all other predictor variables added minimally (0.06%) to the variance accounted for. Openness (b = 0.07 [0.06, 0.09], β = 0.05, p < .001), conscientiousness (b = 0.18 [0.17, 0.20], β = 0.13, p < .001), extraversion (b = 0.18 [0.16, 0.19], β = 0.15, p < .001) and agreeableness (b = 0.09 [0.07, 0.11], β = 0.06, p < .001) positively predicted, and neuroticism (b = −0.18 [−0.19, −0.17], β = −0.15, p < .001) negatively predicted job satisfaction; and smoking (b = −0.05 [−0.08, −0.03], β = −0.02, p < .001), binge drinking (b = −0.01 [−0.02, −0.01], β = −0.02, p < .001) and drug use (b = −0.01 [−0.01, 0.00], β = −0.02, p < .001); and family-of-origin SES (b = 0.06 [0.05, 0.08], β = 0.04, p < .001) and parenthood positively predicted (b = 0.07 [0.05, 0.09], β = 0.04, p < .001) job satisfaction, independently of preterm birth and long hospital stays at birth. Controlling for long hospital stays at birth reduced the effect of preterm birth on job satisfaction to insignificance and adding personality variables into the regression models reduce the effect of long hospital stays at birth on job satisfaction to insignificance.","As in previous research, preterm birth was associated with reduced perceived health (Cooke, 2004; Hack et al., 2007; Lindstrom et al., 2009). In addition, preterm birth was negatively related to life and job satisfaction (areas of functioning not previously measured in preterm adults). The effect sizes for the relations between preterm birth and these three outcomes were small, however. As previously reported, preterm individuals were less likely to be in intimate relationships when asking only individuals currently in intimate relationships, relationships were equally satisfying for preterm and full term adults (consistent with previous findings, although the specific outcomes had not been directly measured before; Hack et al., 2007; Moster et al., 2008). By asking preterm adults how they felt about various areas of their life, we extended our understanding of functioning in adulthood by demonstrating that these adults had equally satisfying intimate relationships but less satisfying jobs and lives more generally. The differences that did exist were small though. However, our large sample size allowed us to compute robust effect sizes and thus add to the growing literature suggesting consistent, long- lasting, but very small, long-term correlates of preterm birth (Hack, 2009; Saigal & Doyle, 2008). Similar results with slightly higher effect sizes were found for individuals who had long hospital stays at birth. We also showed that the associations between preterm birth or long hospital stays at birth and outcomes were affected by other variables associated with preterm birth. Individuals who had been born preterm scored higher on measures of conscientiousness (long hospitalizations at birth were associated with lower conscientiousness), agreeableness and neuroticism, lower on binge-drinking and lifetime drug use, had lower family-of-origin SES, were more likely to report long hospital stays at birth, and were less likely to be parents than individuals born at full term (see also Allin et al., 2006; Cooke, 2004; Hack, 2009; Hack et al., 2007, 2004; Hertz et al., 2013; Moster et al., 2008; Pesonen et al., 2008; Schmidt et al., 2008). Personality differences in preterm individuals appeared to account for some of the reduction in relationship satisfaction, while personality differences in individuals who experienced long hospital stays at birth appeared to account for some of the reduction in health, life, relationship and job satisfaction of these individuals. For example, individuals who had long hospital stays at birth were lower on conscientiousness and extraversion, and lower on neuroticism. When personality was controlled for, the unique effect of long hospital stays on all outcomes were reduced, with the effect reduced to insignificance for job satisfaction. While risky behaviors did not appear to account for any of the effects of preterm birth, controlling for these behaviors did have a small effect on the unique association between long hospital stays at birth and health and life satisfaction. Although personality, risky behaviors and having started parenthood were related to both preterm birth and outcome measures, these variables appeared to have limited effect on relations between preterm birth and adult outcomes. However, personality did appear to have a more important role in explaining some of the poorer outcomes for individuals who spent long periods in hospital at birth. Therefore, hypotheses about the role of personality and risky behavior in explaining the adult outcomes of preterm birth (Hack, 2009; Wolke, 2011) were not fully supported. Although personality did not account for reductions in health, job satisfaction or life satisfaction by prematurity (although it did appear to for long hospital stays at birth), adding personality to models resulted in the greatest increase in variance accounted for in subjective well-being. Therefore, individual differences in personality accounted for individual differences in subjective well-being. Furthermore, personality factors appeared to affect well-being in very similar ways for adults born preterm and full term. For example, higher conscientiousness and extraversion, and lower neuroticism, were related to better health, life, relationship and job satisfaction, higher agreeableness was related to higher life, relationship and job satisfaction, and higher openness to experience was related to poorer health but better job satisfaction. Therefore, some aspects of the personality of adults born preterm appears to place them at risk for poorer subjective well-being (for example, higher neuroticism) while other aspects could potentially be protective for individuals born preterm (for example, the higher conscientiousness seen in preterm adults) but not for those who spend long periods in hospital at birth (for example, lower conscientiousness seen in these individuals). All analyses reported controlled for age. Including age in all analyses was important not only due to relations with the outcomes and covariates of interest, but also because the age of preterm individuals reflects aspects of the hospital care they received at birth. Advances in perinatal and neonatal medicine, including the introduction of surfactant therapy in the 1990s, have allowed increasing numbers of preterm infants to survive (Behrman & Butler, 2006; Hintz et al., 2005). In addition to medical advances, social aspects of hospital stays have changed. For example, the level of contact parents are encouraged to have with their newborns during hospitalization has increased dramatically over the last 50 years (Davis, Mohay, & Edwards, 2003; Goldberg & DiVitto, 1983). In addition to age, medical risk at birth also has implications for the care infants receive during the initial hospitalization as well as subsequent re-hospitalizations. That is, infants often experience longer hospitalizations when they are born at younger gestational ages or at higher medical risk. Although we asked individuals whether they experienced long hospitalizations at birth, individual’s perceptions of “long” may vary. The current methodology allowed us to examine self-report data from a very large sample; however, detailed data were not collected about the early medical contexts. Further work should determine whether longer hospitalizations actually accounted for long-term outcomes of prematurity as well as other early medical factors that may predict adult outcomes. Of course, some limitations need to be acknowledged. Data were only collected at one time point and therefore many of the correlations are open to multiple interpretations. However, preterm birth, hospitalization following birth and family-of-origin SES all occurred before adulthood so it seems fair to say that individuals from low SES families who reported long hospitalizations at birth had lower levels of life satisfaction as adults. Because ‘The Big Personality Test’ was an online survey, individuals had to self- report their own prematurity. Preterm individuals were not selectively recruited and therefore rates of preterm birth were as expected based on national rates (around 7% of live births; Office for National Statistics, 2011). Previous studies have demonstrated that mothers accurately report whether their offspring were born preterm (even when the delivery was over 30 years ago) but are more likely to not respond, or to respond less accurately, to questions about their offspring’s exact gestational age at birth (Tomeo et al., 1999; Yawn, Suman, & Jacobsen, 1998). We do not know of any studies examining individuals reporting on their own birth status, however. We therefore asked only whether individuals had been born early; information was not collected about the individual’s exact gestational age. As a result, we were not able to distinguish between extremely preterm and near full term births, or between high- and low-risk preterm individuals. Future studies should attempt to distinguish among these groups. However, it is noteworthy that we observed significant effects despite having a sample presumably dominated by the late preterm individuals who constitute the largest proportion of preterm individuals in the general population. Furthermore, estimates of the effect of preterm birth may be conservative because some of the individuals in the full term sample may actually have been born preterm, as mothers are more likely to report (and thus tell their children) that their offspring were born later rather than earlier (Tomeo et al., 1999; Yawn et al., 1998). Differences between preterm and full term individuals on our control variables (personality, risky behaviors and parenthood) are consistent with previous studies, which gives us further confidence about the representativeness of our preterm sample. Although 7% of the sample reported being preterm, preterm individuals were over 4 times more likely to provide complete data. This higher rate of having complete data for preterm individuals may reflect the higher conscientiousness, younger age (as older individuals were more likely to have missing data) or reduced illegal drug use (and therefore perhaps less likely to skip questions about risky behaviors) of adults born preterm. However, future studies should examine whether and why preterm individuals provide more complete data about their subjective wellbeing. Regardless of the reasons, as a result of the reduced likelihood of missing data, preterm individuals made up 18% of the final subsample. All information was provided by the participants, which may be problematic because preterm individuals reportedly provide more socially acceptable responses (Allin et al., 2006). However, only self-report measures can be used to assess perceptions of life and well- being and only such a large online survey would have allowed us to recruit such a large sample. If adults born preterm indeed provided more socially acceptable responses this would have minimized rather than amplified the associations we found. Furthermore, as all data was collected from only the individual, the higher levels of neuroticism combined with poorer subjective wellbeing may reflect a more general negative and/or pessimistic outlook on life and in turn such an outlook may make individuals more likely to endorse the preterm or long hospital stays at birth items. Therefore, some of the shared variance in such measures could reflect such a negative and/or pessimistic tendency. Future work should attempt to untangle such a possibility.","We asked preterm individuals how they felt about various aspects of their lives. This approach to studying the adult sequelae of preterm birth allowed us to ask, not about relationship or employment status (which have already been explored in various studies), but about the well-being and functioning of preterm individuals in adulthood. We were able to demonstrate poorer health, lower levels of life and job satisfaction but equal levels of relationship satisfaction in preterm adults. Although these differences were small, they could prove important at the general population level because so many individuals are born preterm and the survival rates following preterm deliveries are increasing (Johnson & Wolke, 2013). These findings are also consistent with other evidence of mild effects of prematurity into adulthood (Hack, 2009; Saigal & Doyle, 2008). The large sample size not only allowed us to examine the individuals’ ratings of their lives following preterm birth or long hospital stays, but also allowed us to control for various other variables including personality and risky behaviors. Despite previous suggestions that the personality profile and behavioral cautiousness of preterm individuals may help explain certain adult outcomes (Hack, 2009; Wolke, 2011) and personality accounting for the greatest amount of variability in subjective well-being in our analyses, we found that personality only accounted for some of the reduction in life satisfaction in preterm adults. However, personality did appear to account for some of the poorer well-being of individuals who had long hospital stays at birth. These results help provide a broader understanding of preterm infants’ functioning in a variety of domains well into adulthood."],["Human performance fluctuates over time. Rather than random, the complex time course of variation reflects, among other factors, influences from regular periodic processes operating at multiple time scales. In this review, we consider evidence for how our performance ebbs and flows over fractions of seconds as we engage with sensory objects, over minutes as we perform tasks, and over hours according to homeostatic factors. We propose that rhythms of performance at these multiple tempos arise from the interplay among three sources of influence: intrinsic fluctuations in brain activity, periodicity of external stimulation, and the anticipation of the temporal structure of external stimulation by the brain. --------------------------------------------------------------------------------","Fast periodicity of perceptual systems was proposed on theoretical grounds before it was directly observed [9, see also Ref. 10]. Periodic and discrete sampling of sensory information, or the ongoing ‘parsing’ of continuous sensory input, was proposed to be advantageous for enabling multiplexing of information processing and for providing time stamps for integrating and sequencing information [9]. Behaviourally, rapid performance modulations (>1 Hz) have been observed for making perceptual judgments about visual features, such as motion, depth, or colour [11]. Intrinsic brain fluctuations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Studies combining measurements of behavioural performance and neural activity have revealed that the intrinsic rhythmic fluctuations of brain activity may contribute to the periodicity of behavioural performance. Fluctuations in behavioural performance at timescales above 1 Hz correlate with the ebb and flow of intrinsic neural oscillations2 [12–14]. Supporting the notion of periodic sensory sampling, studies have shown that variability in perceptual performance co-varies with intrinsic brain rhythms, especially in the theta (∼4–8 Hz) and alpha (∼8–12 Hz) frequency bands. For instance, visual target detection of near-threshold stimuli fluctuates in phase with neural oscillations in the alpha band [12,13,15]. So far, most studies investigating perceptual fluctuations in relation to the phase of intrinsic brain oscillations have been in the visual modality. However, a recent paper suggests that oscillations in behavioural performance also exist in the auditory domain in the theta frequency range [16]. In addition, studies using ‘reset events’ provide evidence for periodic sensory sampling in the theta frequency range [14,17–20]. Reset events can be generated by a transient salient event, thought to reset the phase of ongoing intrinsic neural oscillations so that subsequent oscillations become phase-locked to the reset event. Through the presentation of response-relevant stimuli at various intervals after such a reset event, the unfolding of behavioural performance fluctuations can be directly measured. For example, in a spatial attention task with two behaviourally relevant locations, a reset event was used to investigate the periodic sampling of each location [17]. This approach assumed that the reset event captures attention, thus prioritising sampling of its location. Following the reset event, visual target detection accuracy for each location fluctuated at a 4-Hz rate, while performance between the two locations was in antiphase. Periodic sensory sampling of this type has also been found to be triggered by reset events caused by auditory stimuli [20] and movement onsets [21]. Rhythmic performance fluctuations have been predominantly reported for the theta and alpha frequency bands. While we are not aware of studies showing direct evidence for sensory performance fluctuations in the gamma frequency range (>30 Hz), there is evidence that the phase of delta (0.5–4 Hz) and theta rhythms modulates gamma and alpha band activity and neuronal firing [22–25]. The informational content of neural processing is thought to be carried in these high-frequency signals. Therefore, their regulation by the slower delta and theta oscillations supports the early theoretical proposals of periodic sampling and processing of sensory information [9,10]. Such a mechanism could enable the pick-up and relay of information within local neuronal ensembles to be quantised and paced by capitalising on the slower intrinsic oscillations reflecting the circuit-level dynamics of the networks in which they are embedded. Entrainment to periodic stimulation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Intrinsic brain rhythms can contribute to fluctuating patterns of performance even in the absence of temporally structured external stimulation [1,12–15]. However, additionally, performance is also highly sensitive to periodicities in external stimulation [26]. Many natural stimuli that guide behaviour, such as speech, music, or footsteps, follow a regular rhythm, occurring predominantly between 0.5 and 4 Hz. The influence of periodic, rhythmic stimulation on performance was demonstrated in a series of psychophysical experiments in the auditory domain. Perceptual identification and discrimination of auditory tones were most accurate when stimuli occurred in phase with the preceding rhythmic stimulus train [27, but see Ref. 28]. In principle, two different types of mechanisms can account for performance benefits in the context of periodic external stimulation. The simplest is a mechanism of reactive ‘entrainment’ by which intrinsic neural oscillations are reactively and automatically reset and paced by external events [29]. The alternative is a mechanism of proactive anticipation, by which the brain learns about the periodicities in external stimulation and uses top-down signals to prepare sensory systems for relevant upcoming sensory events [30,31]. Entrainment mechanisms were invoked to explain performance benefits in rhythmic stimulation contexts on theoretical and computational grounds in the ‘dynamic attending theory (DAT)’ by Jones and colleagues [27,32]. Neural recordings later confirmed that neural oscillations do entrain to rhythmic stimulation in the delta and theta ranges, thereby providing a plausible physiological basis for DAT [29,33]; see also Ref. [34]. When presented with rhythmic auditory stimulation, fluctuations in behavioural performance are dependent on the phase of the entrained neural oscillations [35], and behavioural modulations can be observed even in the absence of abrupt onsets in the entraining sequence [36,37]. While many studies have focused on the auditory modality, stimulus-driven fluctuations of performance have also been noted in the visual modality, particularly in the alpha range [38,39]. Proactive temporal anticipation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In rhythmic contexts, separating effects attributed to slavish stimulus-driven entrainment versus proactive anticipation is difficult if not impossible [40]. Many different types of anticipation mechanisms have been described [41]. In addition to rhythmic anticipation [27], anticipation can be related to learned temporal associations between individual events [42], sequences of events [31,43], and temporal conditional probability [44,45]. These multiple temporal structures can combine and occasionally interact [46–48]. It is likely that in rhythmic contexts both entrainment and temporal anticipation occur, and that these interact further with other sources of top-down attention-related signals that guide prioritisation and selection of relevant stimuli. For example, when monkeys were presented with interleaved rhythmic auditory and visual stimulation, performance fluctuations and neural oscillatory activity were dependent on which stimulus modality was relevant for performance [29].","Slower oscillations in the range of seconds reflect processes that affect sustained task performance. For example, studies of vigilance examine the capacity to maintain an adequate state of arousal and focus to detect occasional targets within repetitive and non-engaging tasks over minutes or hours [3,49,50]. Whereas traditionally studies of vigilance have investigated the decrement of performance over time [49,51–53], some have emphasised the waxing and waning of performance [50,54]. Potentially, clinical and neurotypical populations can be distinguished based on performance rhythms. For example, studies have shown that children diagnosed with ADHD manifest a unique oscillatory pattern of periodic drops in accuracy every 20–30 s [55]. It was also speculated that children with ADHD exhibit atypical rhythmic fluctuations in the ‘default-mode network’ which normally fluctuates between 0.01 and 0.1 Hz [56]. A different study showed that individuals with ADHD have diminished ability to benefit from rhythmic patterns in a continuous performance task compared to neurotypical individuals [57]. Intrinsic brain fluctuations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Empirical findings concerning occasional disengagements from continuous performance tasks [54] are normally attributed to a gradual inability to sustain attention [58]. However, intrinsic oscillatory properties of performance may also contribute. Task rhythms may partly reflect ebbing and flowing of different brain networks [59,60]. The existence of functionally significant slow oscillations ranging between 0.01 and 0.2 Hz is supported by modelling data [61], local field-potential recordings in monkeys [62], and human electrophysiology [63]. These rhythms are thought to result in periodic changes in psychophysical performance parameters [59,64]. They are associated with the clustering of performance levels in cognitive tasks, for example, when detection rates on consecutive trials are auto correlated for time lags longer than 100 s [64]. At the neural level, they are thought to represent a slow cyclic modulation of gross cortical excitability [65]. Interestingly, these oscillations below 1 Hz (infraslow) can also interact with faster rhythms. Empirical evidence suggests that the phase of infraslow brain oscillations correlates with the amplitude of faster rhythms (1–40 Hz) [64]. Accordingly, it has been proposed that the ongoing intrinsic infraslow fluctuations between 0.01 and 0.1 Hz and the faster oscillations (between 1 and 40 Hz) nested therein may account for the typical correlation in performance among successive trials in behavioural tasks, creating non- random clustering of performance patterns over time, with variability increasing at longer time-scales [64]. When describing the time series of psychophysical performance over minutes, behavioural data exhibit fractal patterns and power-law autocorrelations [66]. Such dynamics that are characterised by patterns of hierarchical self-similarities at multiple time-scales are typical of ‘scale-free dynamics’, or 1/f distributions, which seem to be a common motif of both behavioural performance patterns [2,67] and brain activity [63–65]. Entrainment to periodic stimulation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Some evidence suggests that entrainment to external periodic stimulation is not confined to the timescale of milliseconds. Using intracellular recordings in animals, researchers have identified non-lemniscal auditory neurons in the thalamus with spontaneous up/down transitions at random intervals, which can become entrained to rhythmic stimulation occurring between 3 and 12 s [68]. However, we are not familiar with comparable findings in the human literature showing entrainment at such slow rhythms. Proactive temporal anticipation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ At the slower rhythm of sustained task performance, temporal anticipation of target events may interact with the regulation of arousal. Changes in tonic arousal during task performance have been associated with changes in uncertainty levels [69–71]. The potential interaction between external stimulus rhythms and arousal is often discussed in the sustained-performance literature. To some extent, the very first experimental task manipulating vigilance relied on rhythmic stimulation [49]. Many other task designs that followed the traditional vigilance (and later: sustained attention) research also presented stimuli in fixed or regular rhythmic intervals [50,72,73]. In our lab, we have recently observed that presenting target events within a predictable, rhythmic temporal structure leads to a periodic modulation of pupil size in preparation for stimulus onset, alongside a reduction in the overall arousal as indexed by tonic changes in pupil size [74]. In contrast, temporally unpredictable targets are associated with a continuous state of high arousal. Our results suggest that traditional explanations of changes in arousal caused by habituation of the neural response to repetitive stimulation [75] or by ordinal predictability [76] may be insufficient, as they do not account for the effects specifically attributable to the temporal structure of the task. Intrinsic brain fluctuations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ When moving from minutes to hours, further intrinsic rhythmic components influence performance. For example, researchers have shown that fluctuations in arousal can occur independently of circadian cycles based on a dopaminergic ultradian oscillator [77]. Electrophysiological studies support the notion of such slow rhythmic fluctuations. For example, researchers have identified two separate components of arousal in broadband EEG, one fluctuates in cycles of ∼100 min, and is related to changes in vigilance; the other fluctuates between 3 and 8 h, and represents variations in wakefulness levels [78]. The functional significance of ultradian rhythms of arousal was demonstrated in a study tracking the latency and amplitude of event-related potentials showing reliable rhythmic fluctuations over hours [79]. Similarly, it was shown that the pattern of performance decrement during a prolonged vigilance task is mirrored by a decrease in the trial-by- trial consistency of the neural response in the theta phase (3–7 Hz), providing further evidence for the association between task and event rhythms [80]. Thus, although changes in arousal are often discussed in the context of habituation resulting from repetitive stimulation [75], it is also affected by very slow intrinsic rhythmic components. Entrainment to periodic stimulation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Whereas the literature has not yet highlighted the effect of infraslow periodic stimulation on performance, a clear stimulus-driven component is recognised to affect performance fluctuations related to ultraslow circadian rhythms. In the ‘forced desynchrony’ experimental protocol, participants are placed in artificial light/dark cycles for varying durations [81]. Using this approach, researchers have shown that fluctuations in body temperature and heart rate are determined by both external stimulation and homeostatic rhythmic components [82]. In turn, periodic changes to body temperature are causally related to fluctuations in cognitive performance [83]. Furthermore, models of dynamic changes in human performance within the time scale of days discuss the circadian processes that periodically determine both the sleep drive and waking alertness, with the latter being closely associated with overall performance [84]. A different source of a periodic stimulation which potentially interacts with brain oscillations occurs through the interaction with other physiological processes in the body. A recent study has demonstrated the phase-amplitude coupling between alpha activity in the brain and an infra-slow gastric basal rhythm (∼0.05 Hz) generated by the stomach [85]. The study has shown that approximately 8% of the variance in alpha activity can be explained by the oscillations generated by the gut (for a recent review on how visceral signals shape neural activity see Ref. [86]). However, it is still unclear whether such periodic stimulation is directly associated with changes in behaviour.","Rhythmicity is a hallmark of the brain [1] and its environment, and characterizes many aspects of cognition and behaviour [87]. Observations of the rhythmic facilitation of behaviour appeared in the earliest days of cognitive research [26], and have since been substantiated by neuroscientific evidence [29]. In this review we have considered rhythms affecting behavioural performance at multiple timescales — from single events, to sustained tasks, to throughout the day. We have suggested that multiple factors play a role in structuring performance over time — fluctuations in intrinsic brain and homeostatic mechanisms, reactive entrainment to external period stimulation, and proactive anticipation of the temporal structure of events (see Figure 2). The exciting topic for future research will be to understand whether and how the factors influencing performance at these various tempos interact. Here we noted a few examples available so far, such as putative interactions between task rhythms and event rhythms that could be mediated by the effects of arousal on faster attention-related dynamics through its change in cortical signal-to-noise ratio [69,71], and the interplay between even slower homeostatic functions and faster signatures of brain activity [68]. However, most of the fun work is still ahead, and results are likely to reveal interesting fundamental principles about the coordination of brain activity and behavioural output. Headway will depend on us broadening the temporal focus of our experimental tasks. Rather than just taking performance measures during single trials as isolated events, it will be fruitful to move to dynamic and extended task contexts, to measure brain and homeostatic activity at multiple time scales, to vary the temporal regularity and pace of stimulation, and to manipulate the predictability of temporal structures."],["Drawing on recent work in emotional and cultural geography, the author brings Derrida's concept of hauntology into communication with thinking about atmospheres. The research deployed a mixed-method approach including audio documentation, observation, focus groups and interviews to look at the use of spectrality in the making of atmospheres associated with A Knight's Peril, an interactive game played at Bodiam Castle in the South East of England. The paper argues that the figure of the ghost is a useful heuristic towards understanding how designers conjure and exploit the emotional and affective power of atmospheres. At Bodiam, these techniques are deployed in an attempt to facilitate new understandings of the past. --------------------------------------------------------------------------------","Visitors to Bodiam Castle who play A Knight's Peril are driven by this urgent plea to foil the plot against Sir Dallingridge. Guided by an adventure map, echo horn, and the ghostly voice of twelve-year-old Kate (a fictional 14th century character who drives the narrative and assists in the players' investigation), participants interact with a few of the historic and imagined personalities associated with the castle as they solve the mystery and intervene in the past. In this paper I build on recent work on atmospheres in Emotion, Space and Society (Anderson, 2009; special issue edited by Bille et al., 2015; Urry et al., 2016) and elsewhere to examine the production and staging of A Knight's Peril. Drawing on Jacques Derrida's hauntology (1994) and considerations of spectrality in social phenomena (Gordon, 1997; Edensor, 2005; Wylie, 2007; Cameron, 2008; Maddern and Adey, 2008; Matless, 2008), I focus on how the game's design conjures ghosts through narrative and sound in support of particular atmospheres and experiences at Bodiam Castle. Coined by Derrida in Specters of Marx, hauntology concerns the deconstructive critique of the priority given to concepts such as being and presence (over non-being and absence for example). It is also a philosophical and ethical destabilisation of all manner of dualisms and universalising totalities (Critchley, 2014). In hauntology, the ghost plays a crucial role in this destabilisation via its characteristic uncertainty. As Liz Roberts explains, a hauntological position ‘is one of deliberate indeterminacy, enforced hesitancy or uncertainty over presupposed givens and operations involving visibility and invisibility that constitute our reality’ (2012, 393). Geographical and urban scholarship on spectrality often draws on Derrida's work and has brought attention to the ways in which spaces are always haunted (see Till, 2005, 2012; Wylie, 2007; Edensor, 2005; Cameron, 2008; Maddern and Adey, 2008). Here, ghosts are a pervasive, yet often unnoticed or unaccounted for part of social life. As Jameson writes, spectrality ‘is what makes the present waver’, it is the notion that ‘the living present is scarcely as self-sufficient as it claims to be; that we would do well not to count on its density and solidity … ’ (1999, 38–39). For urban and social research, taking a hauntological position and being mindful of ghosts can serve as a heuristic device towards unsettling commonly-accepted ontological categories and assumptions. ‘The logic of the ghost’, according to Derrida, is that it ‘points toward a thinking of the event that necessarily exceeds a binary or dialectical logic’ (1994, 78), such as its tendency to blur the boundaries between supposedly stable ontological categories (e.g. living/dead, being/non-being, and presence/absence). Atmospheres are arguably the prototypical spatial form of hauntology. Vague, irrational and indeterminate, they haunt the middle ground between subject and object (Böhme, 1993, 2013). ‘We are unsure where they are’ (Bille et al., 2015, 32) yet we feel them all around us. Like ghosts, their ontological status is always insecure. From the perspective of atmospheres, concepts such as ‘presence and absence, materiality and ideality, definite and indefinite, singularity and generality’ are always expressed as ‘relations in tension’ (Anderson, 2009, 80). For hauntology, these relations are not only tense, but are inseparable as each term can be found to contain traces of its opposite (Buse and Stott, 1999). Despite this overlap, theoretical and empirical connections between spectrality, hauntology and atmospheres are relatively underexplored (but see Edensor, 2012). In this paper, I take scholarship forward by bringing these areas into communication and taking a hauntological position in the investigation of the atmospheres associated with Bodiam Castle's A Knight's Peril. I argue that the framework of hauntology brings a fresh perspective to scholarship on atmospheres. The paper demonstrates how the purposeful making and installing of atmospheres (Böhme, 2013) can be a process through which to redress historical absences – in this case, the absence of women and children in medieval record. Research for this paper was conducted over two phases during 2014 and 2015. The first phase centred on general background to Bodiam Castle and interviews with designers and historians (n = 4) who were involved in creating A Knight's Peril. The second phase occurred over six days and encompassed the main research activity at the castle. Methods included the use of visual and audio methods (e.g. photography and recording), participant observation, focus group discussions and interviews with players (15 groups consisting of 46 individuals). Data collection was not focused on reproducing a single or neutral representation of the conditions at Bodiam Castle. Rather, the approach sought to animate some of the atmospheres associated with playing A Knight's Peril by focusing on the affective, emotional and sensuous elements of the game. The approach was explorative and inspired by recent methods discussions within non-representational theory (Vannini, 2015; Anderson and Ash, 2015; McCormack, 2014) where a diversity of methods are often deployed in order to help ‘look at, listen to and feel the space differently’ (Adey, 2008, 303). This information was analysed with particular attention on processes of staging and constructing atmospheres at the castle and the experiences of participants who played A Knight's Peril. Nevertheless, while the paper captures and names discrete atmospheres (Anderson and Ash, 2015), these are circumstantial forms of sense-making (McCormack, 2014) where envelopment (naming) simultaneously gives consistency while remaining open to the contingency and dynamism of social experience. Moreover, as Simpson (2017b) notes, individuals do not arrive at research sites with identical past experiences. Indeed, even among young siblings who participated in the project, experiences and histories will have been diverse and a day out at a National Trust property can stimulate divergent affective and emotional responses. The structure of the paper is as follows. Following this introduction I review A Knight's Peril within the context of contemporary trends in heritage interpretation. I then discuss recent writing on atmosphere, focusing on design and staging as well as the role of spectrality in producing paradoxical temporality surrounding these phenomena. The empirical discussion presents the ways in which narratives and staging techniques such as the introduction of sounds both enable and co-produce atmospheres which inform and mediate heritage experience and understandings of the past. This work is further analysed through a hauntological lens, reflecting on Derrida's non-linear conception of time and the role of spectrality in the production of emotionally resonant social spaces. I conclude with reflections and suggestions for further research. Playing with history in A Knight’s Peril ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A Knight's Peril is a pervasive game (Montola et al., 2009) – an interactive, augmented reality experience – played at Bodiam Castle (see Fig. 1) in the South East of England. Built by Sir Edward Dallingridge in 1383, Bodiam has the look of an archetypal late medieval castle (Saul, 1995). While the structure suffered an extended period of decline, I found that its wide moat, elegant, symmetrical shape, large towers and country setting produced an evocative and picturesque heritage space.1 Its current condition can be attributed to a series of conservation efforts starting in the early 1800s when the castle was repaired and maintained by individuals and families interested in preserving the structure as a ‘romantic ruin’ (National Trust, 2001, 9–10). Today, the castle is owned and managed by the National Trust, a charitable organisation that operates a range of historic houses, properties, landscapes and nature reserves in England, Wales, and Northern Ireland.2 In 2014 an independent design company (Splash & Ripple3) was engaged by the National Trust to create A Knight's Peril. The broad objective of the project was to improve the experience of explorer families – described as families that actively learn and play together (National Trust, 2014). Interest in this demographic is representative of the charity's desire to diversify its visitor base and reach new audiences. A Knight's Peril deploys a choose-your-own-adventure game model to enable player interaction with a few of the personalities associated with the castle. It facilitates this primarily through a fanciful echo horn, constructed specifically for the project. The device is RFID4 enabled to allow digital interaction at the site. However, it has the look of a type of object one might find in the 14th century. According to the game narrative, Kate Dallingridge (Sir Edward's daughter) has left echoes around the castle and where they have seeped into the stonework, a small seal has grown. When placed in contact with a seal, the horn can tap into Kate's echoes. With echo horn in hand, participants interact with her character and work together to foil the plan to assassinate her father (see Fig. 2). A Knight's Peril is illustrative of heritage learning where non-didactic ways of engaging people with past experience are being explored (e.g. beyond the guidebook and audio guide). Such projects typically involve the introduction of digital media or the creation of virtual environments where visitors can engage in learning activities (Mortara et al., 2014). A range of motivations are evident including: to modify and enhance conventional heritage experiences, to decentre heritage experience away from dominant narratives; to facilitate user agency, and to boost attendance through the incorporation of fun and engaging activities and technology (Hertzman et al., 2008; Coenen et al., 2013; Mortara et al., 2013). A Knight's Peril reclaims and reimagines traces of what has vanished over the course of time. It is an example of how heritage interpretation can encompass or contain a part of the past and to bring it to the present (Till, 2005). The game's setting – Bodiam Castle – is itself a preserved piece of the past ‘adrift in a modern sea, an isolated feature that stands out because it alone is old’ (Lowenthal, 2013, 438). Almost by magic, the castle is physically present yet seems to belong to another time. I found Bodiam's evocative, haunting feel was in no small part due to this uncanny quality of being neither wholly past nor present. Part ruin, the castle is particularly conducive for the conjuration of ghosts. Ruins, as DeSilvey and Edensor (2012, 471) note are ‘characterized by multiple temporalities … offer(ing) opportunities for constructing alternative versions of the past, and for recouping untold and marginalized stories’. Bodiam is managed in such a way to allow ghosts to fill its spaces. The structure and grounds are uncluttered and have been restored and maintained with an eye toward simplicity and clarity. The castle has no roof, interior furniture or other medieval artefacts and is surrounded by simple landscaping all of which maintains focus on the massive stone structure. The result is a simple, yet evocative and haunting space. Such an architectural landscape ‘has a particularly important part to play allowing the uncanny nature of the past to become somehow visible to visitors’ (Maddern, 2008, 365; Edensor, 2005). As a site of memorialisation, Bodiam allows visitors to draw on their own memories and expectations of what castle life might have been like. Of course, interpretations of the past always involve decisions about how history is to be represented and understood (Till, 2005). These stories tell us as much about our contemporary selves as any historical moment. Maddern refers to this interpretive work as a conjuration that involves ‘simultaneously … excavating and burying histories and material assemblages’ (2008, 369 italics in original). In other words, the absences – what we hide or do not talk about – are likely to be just as important as that which is memorialised and remembered. Moreover, as I will explore further in this paper, this process of commemoration – evident at Bodiam Castle and countless historic districts, listed buildings, museums, and monuments to past events – involves the purposeful staging of atmospheres through which resonances of the past can be witnessed and experienced.","Much has been written recently about atmospheres and ambiances (Anderson, 2009; Bissell, 2010; Edensor, 2012, 2015; Ash, 2013; Buser, 2014, 2017; Lin, 2015; Sørensen, 2015). Most productively, the concepts have been deployed to express a range of spatio-temporal conditions that challenge static representations of space. To centre on atmospheres is to understand social experience as sensory (Thibaud, 2011), collective, more-than-human (Anderson and Wylie, 2009), and consisting of dynamic and multiple fields of intensities (McCormack, 2008, 414). In social research, atmospheres help capture and characterise the shifting ‘moods, feelings, sensations and dispositions’ (Lin, 2015, 287) associated with being-in-the-world. However, as Bissell (2010) and Simpson (2017b) note, much of what we call atmospheres occurs in the background and is often unrecognised. These are the affects which, while outside cognitive perception, have the potential to modify a body's behaviours, actions and emotions. For Philippopoulos-Mihalopoulos (2016, 151), such ‘affectively directed’ atmospheres are political in that they can seduce and lure, ‘numb(ing) a body … into an affective embrace of stability and permanence’. Indeed, whether or not atmospheres are perceived by (or involve) a human subject does not diminish their existence or power (Ash, 2013; Sørensen, 2015). This is particularly relevant here as much of the design work employed at places such as Bodiam Castle involves manipulating background environments in subtle ways to produce or support particular forms of behaviour (Turner and Peters, 2015). That individual visitors may or may not be cognitively aware of particular atmospheres is less important than the effects they have on the relationships between bodies. At Bodiam, these relations have been carefully staged through design techniques intended to engineer particular atmospheric qualities. In the following, I introduce two areas of research relevant for a hauntological examination of A Knight's Peril. The first centres on critical examinations of the practices and effects of ‘making atmospheres’ (Böhme, 2013, 2). The second foregrounds scholarship that considers the ghostly temporality of atmospheres. Staging atmospheres ~~~~~~~~~~~~~~~~~~~ The practice of purposefully staging and manipulating atmospheres is pervasive, evident amongst the landscape of shopping malls, festival markets, sports and grand events, public spaces and other sites of managed social experience. In these locales, architects and designers manipulate spaces and objects in order to generate ‘imaginative representations’ – generators which influence individual and collective behaviours and emotions (Böhme, 2013, 4). As Simpson notes, ‘through the design of a space and its particular layout/configuration, different sorts of atmospheres might be produced, encouraged and felt’ (2017b, 430). Such generators include the myriad ways in which space can be manipulated through, for example, light and illumination (Edensor, 2015, 2012; Bille, 2015), smell (Hudson, 2015), visual materials (Biehl-Missal, 2012), decay (DeSilvey, 2006; Turner and Peters, 2015) and the wider urban environment (Bissell, 2010; Simpson, 2017b). For Böhme (2013, 5) atmospheres and ambiances are a ‘felt presence … in space’ produced by architects, designers and others (including non-humans) who shape the experiences and affective and emotional connections or engagements to particular places. Designing or staging atmospheres can range from the most extravagant (e.g. Olympic sporting events, music festivals, etc.) to the everyday practice of care and management (e.g. sweeping a pavement, planting flowers) of the built environment (Thibaud, 2015, 43). Yet, any ‘making of atmospheres’ involves setting out the ‘generators’ which make it possible for an atmosphere to materialise (Böhme, 2013, 3–4). Assembled in the material world, atmospheres have social impacts. For example, atmospheres associated with consumption (Healy, 2014), public space (Buser, 2017), securitisation (Adey, 2008; Urry et al., 2016) and other cultural settings manage social experience. In these and other settings, the configuration of material assemblages contributes to the emergence of feelings and emotions (Anderson, 2009) as well as the potential for action (or non-action). Nevertheless, Bille et al. (2015, 36) argue that rather than reinforcing hegemonic positions, the staging of atmospheres can work to ‘create discontinuities’ in understandings and experiences of the world, pointing to the possibility to disrupt dominant views. Of course, such engineering is never certain (Simpson, 2017b) as atmospheres emerge through the coming together of bodies in uncertain ways. For example, at Bodiam Castle, while visitors may share particular demographic characteristics, they can be affected by and experience atmospheres in widely divergent ways reflecting bodily and social contexts, histories and dispositions (Simpson, 2017b). Moreover, atmospheres can co-exist alongside one another without being in conflict or fusing into something new (Anderson and Ash, 2015). Awareness of the potential for multiple atmospheres means any attempt to forge a singular experience – as is common in the design of historic or commemorative sites (Sumartojo, 2015) – is problematic. Yet, while purposeful staging is pervasive, it is still commonly overlooked. Very little scholarship has analysed how such atmospheres are created, or how they might be used to counter dominant discourses (but see Duff, 2010; Thibaud, 2011; Emotion, Space and Society Vol 15, 2015; Urry et al., 2016; Visual Communications Vol 7, 2016). Moreover, little attention has been paid to examining the production of heritage from the perspective of atmospheres (but see Turner and Peters, 2015; Sørensen, 2015). My research seeks to continue the movement towards filling this gap. Temporality and spectrality in atmospheres ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The second area of scholarship on atmospheres of interest here calls attention to their ‘ghostly’, complex and multivalent temporality. The durability of any atmosphere is belied by an incessant potential for transformation through the introduction or realignment of objects, bodies, and affects (Anderson and Ash, 2015; Buser, 2017). Atmospheres are always less solid than our representations. Moreover, as geographers have long argued, the experience of space is socially situated, ‘ … usually conditioned by previous experience, by habit, by familiar emotions and sensations that produce feelings of belongingness or otherwise’ (Edensor, 2012, 1114). Yet, the material artefacts and affective resonances of particular atmospheres and phenomena may linger and ‘circulate as a field of movement’ long after their disappearance or dissipation (McCormack, 2008, 425). In other words, the material and emotional connections to particular places have their own duration which can outlast any unique moment. These understandings reveal a paradoxical non-contemporaneity of atmospheres – a temporal uncertainty that draws on aspects of nostalgia, memory, repetition, expectation and anticipation (Edensor, 2012). Such insights defy efforts to demarcate the ‘pure presence’ or the essential immediacy of any situation (Jameson, 1999, 58). For Bille et al. ‘ … atmospheres emerge as multi-temporal tensions: they are at the same time a product of the past and future’ (2015, 34). This temporal indeterminacy reveals the ways in which the time of atmospheres can seem out of joint5 (Derrida, 1994). This is a hauntological conceit pointing to the ghostly folding of space and time (Maddern and Adey, 2008) where the present, past and future cannot be cleanly divided but rather are co-constitutive, with each always containing traces of each other. Atmospheres, in other words, can express hauntological characteristics, constructed through the ‘persistences, repetitions, [and] prefigurations’ (Fisher, 2014, 29) of social experience. These hauntological qualities are particularly evident in scholarship on the atmospheric qualities of ruins and sites where ghosts have not been exorcised by the need to tell a singular narrative and where the past is less fixed in place (Edensor, 2005, 2011; DeSilvey and Edensor, 2012; Maddern, 2008; Gallagher, 2015). According to Edensor (2005, 834), the cluttered, disorganised and tangled qualities of these spaces produce an ‘excess’ where ‘memory is elusive, dependent upon conjectures about the traces of the overlooked people, places and processes which haunt ruins’. Ruins, can express ghostly qualities of indeterminacy and ambiguity where the past remains strange, not yet eradicated or re-interpreted through dominant forms' memorialisation. Such spaces can draw attention to the ways in which ‘our experience of the world is haunted’ such that the ‘past and future co-exist, and interact, in uncertain and unpredictable ways (Hill, 2013, 381). Within this scholarship, spectrality and the figure of the ghost is crucial to understanding the less-than-settled, non-linear ways people experience the world. In heritage spaces such as Bodiam Castle, ruination is a form of curation (Turner and Peters, 2015) where decay is visible, yet checked in order to facilitate haunting and evocative experiences and to tell stories of the past. Staging atmospheres at bodiam castle ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The making of atmospheres at Bodiam Castle is entangled with issues of conjuration and the deliberate practice of making (certain) traces of the past visible. This section is organised around three analytical concepts. First, I examine how particular spectral traces are made visible through narrative and the effects they have on both the modern-day chronicle of the castle and the production of atmospheres at Bodiam. I then turn to the role of sound in A Knight's Peril and how specific design techniques about how the narration is delivered influences experiences and atmospheres. Finally, I consider the spectral and anachronistic qualities of these elements and how these contribute to the production of new, arguably richer, ways of experiencing and understanding the past. Conjuring ghosts and addressing representational silences: narrative in the atmospheres of A Knight's Peril Geographers are increasingly engaging with narrative and storytelling in research (see Cameron, 2012 for an overview). Here, I focus on the way stories and storytelling might be used to destabilise prevailing discourses (Bulken et al., 2015; Gibson-Graham, 2006; Pratt, 2009). A Knight's Peril is a form of storytelling that engages through ‘everyday’ characters. For the motivated visitor, the ‘official’ story of Bodiam Castle can be easily found in the National Trust guidebook or the DVD. These authoritative materials highlight information about the castle's builder (Sir Edward Dallingridge) and provide detailed reviews of the castle's design including the ubiquitous cross sections and renderings of the structure as it would have been in the 14th century. There is also information about individuals and families involved in the castle's conservation and restoration. These authoritative resources privilege particular medieval histories and serve as ‘affective signals’ (Bissell, 2010, 279) which contribute to the types of atmospheres one might experience (Sumartojo, 2015) at historic sites such as Bodiam Castle. In producing A Knight's Peril project designers and historians did not want to retrace these traditional narratives but sought to approach the history of Bodiam Castle ‘from below’ (Poole, 2017, 10). According to Steve Poole (project historian), A Knight's Peril celebrates non-didactic, personalised ways of engaging people with past experience: ‘ … it requires them [participants] to enter into it and to interpret for themselves how things were. We try to feed in some context without saying “hey, listen, this is what you need to know”’.6 For Poole, interactive and sensory-based games offer heritage sites a means of ‘telling the story’ that is more intimate, engaged and varied than traditional means allow. It is through imagination and immersion into the narrative that this personalised form of heritage interpretation occurs. Central to the approach at Bodiam was the desire to tell a story about lesser-known characters. Unfortunately, records from 14th century England are limited and the ability to tell alternative stories or to speak confidently about everyday characters is constrained by medieval social systems and record keeping procedures. Set within patriarchal 14th century England, well-documented individuals are typically wealthy, male landowners or nobility. As a result, relatively little is known about women, children or the Bodiam household beyond Sir Dallingridge. This absence of evidence continues to influence the National Trust's representation of Bodiam and this period of history. For example, without formal documentation to substantiate their experiences, women and children are largely missing from the guidebook and DVD. A Knight's Peril attempts to overcome these representational silences by mixing known persons and commonly accepted facts with fictional characters and creative storytelling. The experience conjures a female child via fictional, 12-year-old Kate. As one historian noted, Kate was developed to generate visitor interest amongst young girls, but also to address a representational injustice: ‘we wanted a young woman to be the guide partly because this is such a … male dominated environment … it is easy to assume that men are the only doers and shakers in this event. Now, that's an impression created by the evidence because it is men who control the means of recording this and who were allowed any kind of formal role in politics and public life … ’.7 Through A Knight's Peril, project designers and historians introduced a new historical form into the atmosphere of the castle. Kate shapes visitors' experiences and emotional responses to the material environment of Bodiam Castle. She is an ‘emotional opening’ (Gibson-Graham, 2006, 136) towards the construction of a counter-narrative. Several interviewees noted the importance of a female at the centre of the story and their experience of the castle. For some, her character was ‘appealing’8 and provided a sense of personal attachment – ‘I liked it because you could imagine you were Kate’.9 Others found it was the perspective of a child that made absences present. One father noted how he began to think about the nature of childhood in a 14th century castle (e.g. the difficulty getting around, the limited possibilities for play). Whereas previously castle life was about knights and battles, having a child tell the story and listening to his daughter's own reflections altered this frame10. New questions and thoughts about the castle emerged: What was life like for the child of a local nobleman? What kinds of games did she play? Who did she play with? Was she able to run and explore or was her life filled with work and hardship? Did she struggle to make her way though the winding stone staircases? Tracing through Bodiam at the behest of a spectral child coaxes particular forms of movement and bodily empathy (Edensor, 2005) about life at the castle. Conjured by historians and game designers, Kate's ghost possesses and guides modern bodies towards new understandings of history. Kate makes traces of what has been lost – here, the erasure of women and children – visible. Producing real material effects (Gordon, 1997) in the atmospheres of the castle, Kate reminds us of the absences and makes them present in contemporary narratives of medieval history and the way visitors experience Bodiam. Embracing Kate's fictional qualities meant she represented those … ‘ … who just aren't in any records because they didn't know or because the records have been lost, or they're just not important people, so they don't get written down at all’.11 Within A Knight's Peril, ghosts are active agents in the production of Bodiam Castle's atmospheres. Their contribution within the narrative – as part of a new chronicle of the castle – sets in motion the repair of longstanding representational silences, disrupting conventional ideas about medieval castle life and constructing new understandings of the past. The echo horn and the sonic production of atmospheres During A Knight's Peril, participants listen to the voices and sounds of a 14th century castle. This sonic environment is transformative. Sounds, as Michael Gallagher notes, are not only capable of ‘activating feelings and emotions’ but also are ‘a kind of affect – an oscillating difference, an intensity that moves bodies’ (2016, 43; Duffy et al., 2016). Recently, geographers have shown increasing interest in diversifying sensual understandings social experience to include the role of sound (Gallagher and Prior, 2014; Hill, 2015). While there is insufficient space to review all of this literature, of particular relevance is research on the role of sounds and soundscapes in shaping the qualities of place (Anderson, 2004; Simpson, 2017a; Duffy and Waitt, 2013). This attunement to sound ‘calls attention to something that is ordinarily ignored’ (Gallagher and Prior, 2014, 271) but which can greatly influence atmospheres. Moreover, geographers have shown how sounds ‘can produce particular embodied relationships with the past’ (Simpson, 2017a, 91; Gallagher, 2015) Sounds contribute directly to how we come to understand particular spatial settings. As might be expected, sounds – in the form of dialogue and narration – are used in A Knight's Peril to convey meaning. Through the voice of Kate and the characters she encounters the audio communicates specific information for participants. For example, references to war with France and the recent Peasants' Revolt are included as contextual material that historians felt was important for visitors understanding of the time. In addition, at the end of each instalment, the audio provides clues and specific choices about who to follow and where to go next in the castle. As such, playing A Knight's Peril encourages a form of interacting with and moving through Bodiam Castle that requires careful listening and interpretation. In addition to these cognitive elements, the audio also affects participants in less reasoned ways. Most obvious was the hurried way in which players tend to move through the castle. Pleading with visitors to help solve the mystery, Kate's voice incites this urgency, prompting participants to ‘hurry up’. She implores, ‘we don't have much time’ and later, reflecting on the beauty of the castle she laments, ‘we don't have time to see all of it … we need to figure out who might want to hurt my father’. In my experience as a player and as witnessed during observation, this tended to produce repetitions of pausing to focus on dialogue and directions from the echo horn, followed by a quick dash to another room. Other, more ambient, sounds are introduced to facilitate imaginative connection to medieval life. At various points of the game visitors will hear: ‘traditional’ medieval music; banging and clattering of pots and pans in the kitchen; doors creaking open; the drawing and clashing of weapons; and the flight of arrows. These sounds reshape and adapt visitors' relation to the site by ‘amplify(ing) the haunted qualities’ (Gallagher, 2016, 468) of Bodiam Castle. They bring ghosts to the surface of recognition. Particularly powerful is the conjuration of Kate – a remarkable and mysterious sonic moment that occurs when the echo horn comes in contact with the initial seal. At first, a faint voice can be heard, echoing and straining to gain players' attention. Soon, the echoes fade and the clear voice of 12-year-old Kate materialises. This ‘spectrality effect’ (Parkin-Gounelas, 1999, 128) disrupts certainties about the present. Echoes are a well-used storytelling device that often signal the crossing of a threshold and the jumbling or contamination of time and space. This spectral use of sounds is also a well-developed trope with connections to hauntological music (Sexton, 2012) where a range of aesthetic devices such as decay (e.g. a purposeful erosion of sound) and crackle used ‘to effect a ghostly infiltration of the present by the recent past’ (Gallagher, 2015, 481; Fisher, 2014; Fisher, 2012). The sonic environment associated with A Knight's Peril deploys these techniques to facilitate the magical shift away from the 21st century and into the medieval past. Following the (hauntological) appearance of ghostly Kate, players enter the castle immersed in happenings of the 14th century. Indeed, one parent spoke about how his children became thoroughly immersed in the experience and noted how ‘the horn really activated it for them’.12 The audio device helped to bring these visitors into contact with the castle and its history and magnify emotional resonance. This father went on to explain how ‘they were really present and focused’ when participating in A Knight's Peril. Others noted how the horn itself was particularly memorable and helped Bodiam Castle stand out from other heritage sites. ‘we go to a lot of National Trust places and we often can't remember one from the other, they merge into one a bit, this would be quite different’13 ‘Sometimes one castle is like any other castle and this will probably help them to remember this castle in particular’14 These comments are illustrative of how the staging and making of atmospheres through sounds and audio devices can contribute to significant emotional resonance and meaning. This was a purposeful objective of the design team which sought to forge a strong sense of medieval life and attachment to the castle amongst participants through an engaged and immersive experience.15 Yet, this is not a total immersion. Of course, not all players experienced the game in the same way. Indeed, during my observations, some children were clearly bored, others concentrated intensely, while others simply used the echo horn as a piece of medieval fashion. Moreover, adults generally played along but were mostly supporting younger participants by pointing out spots on the map or repeating certain phrases from the narrative. It is evident that bodily capacities and social histories (Simpson, 2017b) played an important role in mediating experiences of A Knight's Peril and the emergence of atmospheres at Bodiam Castle. Moreover, sounds emanating from the game overlap with ‘live’ sounds of the castle as well as other the audio of players in the same room (multiple versions of the audio can be played and heard at once). It is an experience that some non-players I spoke with found frustrating and distracting to the otherwise tranquil atmospheres of Bodiam. Sounds are not universally received, but can variously enrich, ‘plague and pollute’ (Lorimer and Wylie, 2010, 7). Moreover, the overlapping of sounds associated with A Knight's Peril can differently mediate movements through the site (Gallagher, 2015). For some, this occurs in concert with Kate and her quest. For others, it can be a repulsing force. On these occasions, multiple atmospheres come into conflict and result in new relations and new atmospheres (Anderson and Ash, 2015). This affective-materialist perspective points to the ways in which sounds co-produce social space (Simpson, 2017a) and atmospheres (Doughty and Lagerqvist, 2016). The sounds (and associated audio devices) of A Knight's Peril contribute to the production of atmospheres at Bodiam Castle through diverse bodily relations (Simpson, 2017a) and hauntological techniques. Their contribution to conviviality of the site depends upon a range of contextual elements including one's angle of arrival (Ahmed, 2010). As Simpson (2017a, 91) notes, sounds are differently received and how ‘listening bodies’ are ‘disposed towards hearing those sounds’ has significant implications for the types of relations and atmospheres they can produce. Towards a hauntological understanding of atmospheres The time is out of joint. O cursed spite, That ever I was born to set it right! Nay, come let's go together. (Hamlet, Act 1, Scene 5, Page 8) In Specters of Marx, Derrida cites the above passage from Shakespeare as a provocation and challenge to linear understandings of time. Calling on the ghost of Hamlet's father, he draws attention to the possibility that pasts, presents and futures are more likely to be jumbled than linear or clearly compartmentalised. The ghost alludes to how past, present and future mix in our minds and our emotional connection to place. In other words, past histories and future expectations shape how we experience and interpret the present. In this section I apply Derrida's hauntological position to the concept of atmospheres, reflecting on the case of A Knight's Peril. I note the ways in which the atmospheres associated with the game draw on and reinforce disrupted, non-contemporaneous notions of time. The ghosts of A Knight's Peril attempt to destabilise visitors' sense that the past is fixed or inert. Kate, the central figure, arrives from the past, but she is clearly not the same as any individual named Kate from the 14th century. Neither alive nor dead, from the present nor the past – Kate's ghosts reveals the ontological instability associated with both atmospheres and hauntology. Her existence in the 21st century – a time where she does not belong – is anachronistic as she floats between and among various times and spaces of Bodiam Castle. Like atmospheres, she has an uncertain ontological status; never fully present, she remains with us and influences players' emotions and behaviour. Moreover, Kate makes demands of those in the present, recruiting visitors to solve the mystery and save Sir Dallingridge. During my research, this sense of urgency and desire to put right certain events of the 14th century were echoed by many participants. ‘I really liked just trying to help her’, one young visitor noted.16 Kate's ghostly presence pushes us and implores us to act. Of course, there is something more being asked than simply to play a game. Like the ghost of Hamlet's father, Kate is an apparition who provokes. Yet, she does not call for murderous retribution. Rather, her justice rests with our changed sense of the past. This new representational construct involves an expanded recognition of the lives of marginalised persons from medieval history. Moreover, her (re)appearance signals an as-of-yet future, traces of which we might only sense. Indeed, our work in the 14th century salvation of Sir Dallingridge is peripheral to the world it foretells – central to which is the possibility of a more equal and just future. This notion of equality in representation is something A Knight's Peril designers and historians sought to build into visitors' experience of the castle. By introducing the ghost of a young female child into the chronicle of the castle, new atmospheres are facilitated centring on playfulness and whimsy and new questions are asked – what was it like to be a young girl in this place? The making and conjuration of atmospheres at Bodiam Castle is not simply a memorialisation of the past. Rather, it is as Derrida might say, a ‘phantomatic mode of production’ (1994, 120) where visitors are not only encouraged to walk with ghosts, but where their spectral demands are taken seriously. Through simple yet evocative and engaging generators (e.g. hauntological sonic techniques), the atmospheres of A Knight's Peril take on a sense of openness which strengthens emotional connection and allows for ghosts and memories to take hold. This deliberate production of atmospheres is a conjuration that makes (in)visible particular stories (here, the experiences of women and children) through narrative and sonic design techniques. Moreover, by allowing space for ghosts it implies a renegotiation and disruption of the supposed stability between concepts such as presence and absence. At Bodiam Castle, the resulting social spaces make obvious the uncanny, non-linear nature of time (Hill, 2013) and, for many people who play A Knight's Peril, facilitate powerfully emotional and resonant experiences. Such efforts are part of a hauntological strategy of remembrance which does not exorcise ghosts, but conjures them, drawing attention to the ‘discontinuities and irruptions’ which characterise processes of memory (Edensor, 2005, 829).","This paper studied the production and staging of A Knight's Peril at Bodiam Castle. The research examined particular representational and material elements associated with the game including the characters created and deployed as well as the use of hauntological sonic and aesthetic techniques. Building on scholarship in emotional and cultural geography (Wylie, 2007; Maddern, 2008; Sørensen, 2015; Edensor, 2012; Bille et al., 2015) this research detailed the ways in which historians and designers sought to intentionally shape experiences and emotional responses to Bodiam Castle through the production of atmospheres. Endeavouring to address a representational silence of medieval history – the widespread absence of women and children in historical accounts – the design team not only conjured the ghost of Kate Dallingridge, but exploited a suite of hauntological tropes (e.g. echo, decay, anachronism) to disrupt conventional readings of the past. Drawing on these techniques, the atmospheres of A Knight's Peril can be seen as purposefully staged framings which – under certain circumstances – can re-order visitors' understandings of the medieval history and potentially disrupt hegemonic views and assumptions about the world. Bringing Derrida’s (1994) concept of hauntology to the study of atmospheres, the paper explored the role of spectrality in the production of A Knight's Peril. Similar to atmospheres, ghosts belie ontological certainty. To think of the atmospheres at Bodiam Castle as hauntological is to suggest that there is something ‘out-of-joint’, not just right or uncanny about them. I argue that temporal uncertainty, facilitated via the figure of the ghost, is one way in which A Knight's Peril gains its emotional and affective power. This suggests that research which examines spectrality in the making of atmospheres (within or outside heritage contexts) can help understand the effects designers have on the moods, emotions and behaviours of people in social spaces.","This research for was made possible by an Alumni Award from REACT, the South West creative economy hub established by the Arts and Humanities Research Council (AH/J005185/1)."],["How can we explain consciousness? This question has become a vibrant topic of neuroscience research in recent decades. A large body of empirical results has been accumulated, and many theories have been proposed. Certain theories suggest that consciousness should be explained in terms of brain functions, such as accessing information in a global workspace, applying higher order to lower order representations, or predictive coding. These functions could be realized by a variety of patterns of brain connectivity. Other theories, such as Information Integration Theory (IIT)and Recurrent Processing Theory (RPT), identify causal structure with consciousness. For example, according to these theories, feedforward systems are never conscious, and feedback systems always are. Here, using theorems from the theory of computation, we show that causal structure theories are either false or outside the realm of science. --------------------------------------------------------------------------------","We wake up every day and transition from an unconscious to a conscious state. Surely, there is something to explain. In binocular rivalry and visual masking, we can render clearly visible stimuli invisible. Surely, there is something to explain here too. These examples and many others are routinely used by the scientific community as a means to study consciousness and are at the heart of all empirically-minded theories of consciousness (Fig. 1a). Because of the subjectivity of consciousness, the dependent measures in these experiments are subjective reports (or other measurements known to reliably correlate with subjective reports). We cannot use measures of brain activity as a-priori indicators of consciousness because we want to understand how brain activity gives rise to consciousness in the first place. Causal structure theories of consciousness ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Theories of consciousness aim to explain how changes in the proposed mechanism lead to changes from unconscious to conscious states and vice versa. To do so, a number of theories propose that the essential element for understanding consciousness is how parts of a system interact. If a system has the “right” kind of causal structure, in other words, if its elements interact in the “right” way, it is conscious. Otherwise, it is not. We call such theories causal structure theories. For example, in Recurrent Processing Theory (RPT), Lamme proposed that recurrent processing is both necessary and sufficient for consciousness (Lamme, 2006). The first sweep of visual feedforward processing is unconscious. Consciousness kicks in when recurrent, top-down processing interacts with neurons activated during the initial feedforward sweep (Lamme, 2006). According to RPT, what matters is the causal structure because consciousness depends only on how neurons interact with each other: when there is recurrent processing there is consciousness, and there is no consciousness otherwise. Empirical support for RPT was proposed to be provided by neurophysiological experiments (Fig. 1c) in which recurrent processing enhanced neural activity in V1 when visual stimuli were consciously perceived. When the stimuli were not consciously perceived (during anaesthesia or when the stimuli were masked), there was no recurrent processing (Fahrenfort, Scholte, & Lamme, 2007). Information Integration Theory (IIT) is another example of a causal structure theory of consciousness. IIT proposes that an information integration measure called ϕ, which is computed based on the causal structure of a system, quantifies consciousness (Oizumi, Albantakis, & Tononi, 2014). Consciousness is identified with ϕ > 0 systems: if elements of a system interact in the “right” way, the system has ϕ > 0 and is conscious. If ϕ = 0, it is unconscious. For example, ϕ is always greater than zero in recurrent systems (they are always conscious) and always equal to zero in feedforward systems (they are never conscious). Empirical support for IIT was asserted to be provided by studies showing that a practical proxy of ϕ is low in coma, intermediate in minimally conscious states, and maximal during wakefulness (Casali et al., 2013; Tononi, Boly, Massimini, & Koch, 2016). IIT and RPT were amongst the first theories of consciousness to make precise predictions about which systems are conscious. As such, they contributed greatly to the advancement of the science of consciousness. However, we will show that causal structure theories end up in an empirical impasse for principled reasons: they are either false or outside the realm of science.","Recurrent neural networks are universal function approximators (Fig. 2; Schäfer & Zimmermann, 2006). That is, any input-output function can be approximated to any degree of accuracy. Vision is such an input-output function. For example, pictures of animals are presented as inputs on the retina, and the outputs are the elicited percepts of animals (or reports about these percepts). Likewise, the stimuli in a visual masking experiment are inputs, and the outputs may be button presses, verbal reports or any other measure shown to reliably correlate with subjective reports. Importantly, experiments which intervene directly on the brain, for example using implanted electrodes or Transcranial Magnetic Stimulation (TMS) are still input-output functions. The only difference is that part of the input is provided by means of electrodes or TMS rather than through the sensory organs. Feedforward neural networks are also universal function approximators (Fig. 2; Hornik, Stinchcombe, & White, 1989). Hence, for a given input-output function we can find both feedforward and recurrent networks that realize the same function in different ways (LeCun, Bengio, & Hinton, 2015; Oizumi et al., 2014; Werbos, 1988). For instance, if there is a recurrent network that performs image recognition, there is an equivalent feedforward network that does it equally well. If there is a recurrent network that exhibits the characteristics of binocular rivalry, there is an equivalent feedforward network that does so too. If there is a recurrent network that takes a collection of spike trains as input and outputs another collection of spike trains, there is an equivalent feedforward network that does the same thing. Anything that can be done by recurrent networks can also be done in a feedforward manner (Fig. 2). We call this unfolding: any recurrent network can be unfolded into a feedforward network implementing the same function. In particular, any behavioural experiment can be seen as an input-output function, and can thus be implemented by both recurrent and feedforward networks. Any input-output behaviour can be implemented not only by one particular feedforward network, but also by infinitely many equivalent feedforward networks and by infinitely many equivalent recurrent networks, because the universal approximator property does not depend on structural details such as the number of layers or on the precise connectivity. In fact, given an input-output function, we can find infinitely many networks, each with a different ϕ, that all realize the same input-output function (see Appendices A and B, see also Chalmers, 2018). Moreover, the universal function approximator property is not restricted to neural networks but also holds true for Turing machines, cellular automata, cyclic tag systems, and more generally for any universal computing system (Turing, 1937; Wolfram, 2002). These facts are uncontroversial and widely accepted (including by proponents of IIT: see Oizumi et al., 2014). Implication I: causal structure theories are doubly dissociated from empirical data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ An implication of the unfolding argument is that causal structure theories are either false or outside the realm of science. In other words, causal structure theories are doubly dissociated from empirical data: they are neither necessary nor sufficient to explain empirical data (see Fig. 3). For example, according to IIT, the level of consciousness varies with ϕ. A system is conscious if, and only if, ϕ > 0 (oblique red arrows in Fig. 3).1 An experiment can be seen as an input-output function (Fig. 2). Hence, for any recurrent system with ϕ > 0 that reproduces the outcome of an experiment, there are feedforward systems with ϕ = 0 that also reproduce the outcome. According to IIT, one system has consciousness but the other does not. Conversely, for any feedforward system with ϕ = 0, there are recurrent systems with ϕ > 0 that produce the same experimental results. In fact, we show in Appendix A how to implement any function with ϕ = 0 or with arbitrarily high ϕ. That is to say, for each system that provides evidence for IIT, there are other possible systems that falsify it. This argument generalizes to any causal structure theory, not just IIT. All causal structure theories are doubly dissociated from empirical results about consciousness. Hence, it makes no sense to provide experimental evidence for causal structure theories. For example, the figure-ground experiment mentioned earlier cannot support or be explained by RPT because there are feedforward networks that make the same subjective reports as humans when they consciously see the figure, but are unconscious according to RPT. Conversely, there are networks with recurrent activity that make the same subjective reports as humans when they do not consciously perceive the figure, but are conscious according to RPT. Likewise, the finding that awake humans have higher ϕ than sleeping humans cannot be explained by or support IIT because there are feedforward networks with human wakefulness characteristics, and recurrent network with human sleep characteristics. Our arguments are not only of an abstract mathematical nature. In real life, there are many examples where feedforward and recurrent networks realize the same complex functions. For example, deep reinforcement learning has been implemented with purely feedforward convolutional networks to achieve super-human performance in Atari video games (Mnih et al., 2013). Hausknecht and Stone (2015) replicated this superhuman performance using recurrent networks. The unfolding theorems tell us that this is not surprising because we can always find equivalent feedforward and recurrent networks. These systems are empirically identical (to a close approximation). One is conscious but the other is not, according to causal structure theories. Moreover, unfolding provides a recipe to build two small robot systems with exactly identical input-output functions but different causal structure (see Appendices A and B). Experiments on one robot support the theory; experiments on the other falsify it. The unfolding argument shows that there are always systems that empirically falsify causal structure theories. Proponents of IIT try to avoid this problem by claiming that systems with ϕ = 0 are unconscious despite being empirically indistinguishable from conscious systems (Oizumi et al., 2014). We will show next that this claim makes IIT circular, and therefore unfalsifiable. In other words, causal structure theories are falsified by the unfolding argument, unless they decide to become unfalsifiable. For example, proponents of IIT may still insist that the robot with ϕ = 0 is unconscious whereas the one with ϕ > 0 is conscious, owing to their differing causal structure. However, such a proposition quickly ends up in circularity because we have no criteria to settle the matter. In particular, we have no empirical criteria because experimental results about consciousness are all identical for the two robots. The only reason to believe that only the ϕ > 0 robot is conscious is to already believe in IIT, but this is circular. The situation is even worse: there are many causal structure theories, such as IIT and RPT. Which one is the “right” one? Even within IIT, the axioms do not uniquely determine ϕ (Barrett & Mediano, 2019; Bayne, 2018), and different empirical measures of ϕ yield very different results (Mediano, Seth, & Barrett, 2018). Which version of IIT is the “right” one? We can never decide because we have no criteria to test the theories and pit their predictions against each other. We are left with the conclusion that there are different types of “consciousness” (i.e., consciousnessIIT_version_1, … , consciousnessIIT_version_n, consciousnessRPT, and so on), depending on which theory we favour. Insisting that the robot with ϕ = 0 is unconscious whereas the one with ϕ > 0 is conscious even though they are empirically identical leads IIT outside the realm of empirical science. To summarize the unfolding argument, the conclusion follows from four premises. (P1): In science we rely on physical measurements (based on subjective reports about consciousness). (P2): For any recurrent system with a given input-output function, there exist feedforward systems with the same input-output function (and vice-versa). (P3): Two systems that have identical input-output functions cannot be distinguished by any experiment that relies on a physical measurement (other than a measurement of brain activity itself or of other internal workings of the system). (P4): We cannot use measures of brain activity as a-priori indicators of consciousness, because the brain basis of consciousness is what we are trying to understand in the first place. (C): Therefore, EITHER causal structure theories are falsified (if they accept that unfolded, feedforward networks can be conscious), OR causal structure theories are outside the realm of scientific inquiry (if they maintain that unfolded feedforward networks are not conscious despite being empirically indistinguishable from functionally equivalent recurrent networks). Examples ~~~~~~~~ Imagine that one could surgically replace the brain’s native recurrent sound processing system with an equivalent feedforward implant. The implant takes the same collection of spike trains as inputs, and outputs the same collection of spike trains as the native brain areas. We know that such implants exist in principle because of the previously mentioned unfolding theorems. Even though the causal structure in the new implant is completely different, the rest of the brain does not notice any difference.2 The brain can do its normal job. This means that all subjective reports by the person are identical before and after the surgery. The person will claim all the same things about sound as before the implant was placed, such as “I hear the drizzle of the rain, it is music to my ears”, or “I understand what you are saying”, etc. In particular, any experiment about which sounds are consciously perceived will yield exactly the same results as with the native brain area. Therefore, we end up with the dilemma mentioned earlier: either causal structure theories are wrong (if they accept that there is still auditory consciousness with the implant), or they are outside the realm of science (if they claim that consciousness is different with and without the implant even though there are no empirical differences). We can push the example further to entire brains. Since anything that can be done with a recurrent network can also be done with a feedforward network, there could be «feedforward brains» that behave exactly like human brains. Such systems would have all the same functional characteristics as a normal human brain, but completely different causal structure. They behave exactly like a human in all respects, passing the Turing test seamlessly. However, according to causal structure theories, they are not conscious because they do not have the “right” kind of causal structure. Crucially, these systems respond to any empirical experiment exactly like humans. For example, they identically describe what it is like for them to see red, hear sounds, have memories, and so on. They respond to all scientific paradigms (such as masking, binocular rivalry, figure-ground segmentation, etc.) in exactly the same way. They exhibit the same wakefulness characteristics and the same sleep characteristics. In summary, no behavioural experiment can distinguish between human brains and feedforward brains in principle. Therefore, either causal structure theories are wrong or they are outside the realm of science. Implication II: conscious content in IIT is doubly dissociated from experiments ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In causal structure theories the content of consciousness is also doubly dissociated from empirical observations. For example, we can construct systems that behave as having experience X when, according to IIT, they are in fact experiencing Y (see Appendix C). For example, a system participating in a rivalry experiment may report that it is seeing the cat image when, according to IIT, it is experiencing the smell of ham. In principle, as shown in the appendix, it can experience any content of consciousness while reporting that it sees a cat. Of course, it could also experience seeing a cat, but this would just be a coincidence, showing a double dissociation. There is a straightforward reason why causal structure theories are vulnerable to the kind of arguments presented here. All that is required for a system to be conscious is a particular causal structure. At the same time, any function can be implemented by many different systems with different causal structures. Hence, there can be no consistent link between causal structures and experimental results. Network efficiency & evolutionary constraints ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In practice, the brain has to cope with very strong space and energy constraints: processing must be efficient enough to be contained within a small skull, and energy consumption must be limited. Experiments have suggested that, indeed, sufficiently complex tasks can strongly constrain network properties, when the number of neurons is limited (Khaligh-Razavi & Kriegeskorte, 2014; Nayebi et al., 2018; Yamins et al., 2014). In general, feedforward networks require many more neurons to implement a function than equivalent recurrent networks with more efficient causal structure and are therefore impractical (but not always: for instance image recognition is more efficiently implemented in feedforward convolutional networks). In this regard, causal structure theories may turn out to be good markers for consciousness. For example, high ϕ has obvious functional benefits, such as efficiently integrating information. We argue that awake brains have high ϕ for this functional reason. Hence, causal structures may be good correlates for consciousness in humans not because they are identical with consciousness, but because they correlate well with neural information processing in general, which happens to covary with conscious state in humans as a contingent rule. This explains why causal structure theories may provide human consciousness-meters (see for example Casali et al., 2013). However, it is an entirely different thing to identify consciousness with causal structure. In short, brains are recurrent because brain processing must necessarily fit inside a skull, not because consciousness is identical with the brain’s causal structure. The unfolding argument vs. the zombie argument ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The zombie argument is a well-known argument aiming to show that physicalist and functionalist theories cannot account for consciousness. Imagine unconscious “zombies” who are empirically indistinguishable from conscious “non-zombie” equivalents. For example, if a zombie inadvertently places its hand on a hot stove, it yells “ouch” and immediately retracts its hand, just as the non-zombie does. However, the zombie does not feel pain. The 3rd person, observable properties are all identical between zombies and non-zombies, but consciousness differs. Whether or not zombies are in fact possible is heavily debated (e.g., Dennett, 1991). However, if they are, it follows that consciousness cannot be explained in the standard, functionalist framework of science (this is the hard problem of consciousness; Chalmers, 1996). Indeed, if functionally identical systems (the unconscious zombie and its conscious non-zombie equivalent) can have different consciousness, then functional approaches cannot explain consciousness. The unfolding argument is very different from the zombie argument for two reasons. First, the zombie argument aims to dismiss all physicalist accounts of consciousness, including functional ones. In contrast, the unfolding argument only targets causal structure theories of consciousness, and not physicalist or functionalist theories in general. In fact, the unfolding argument favours functionalist theories because (un)folding a network changes only its causal structure but not its function or physical nature. Other major theories of consciousness, such as Global Workspace Theory (GWT; Baars, 1997; Dehaene & Naccache, 2001), Higher-Order Thought Theory (HOTT; Lau & Rosenthal, 2011; Rosenthal, 2004) or Predictive Processing Theory (PPT; Friston, 2013) are not affected by the unfolding argument, as we show in the next subsection. Second, one can choose to dismiss the zombie argument by claiming that zombies are in fact not possible (e.g., Dennett, 1991). In contrast, the existence of unfolded systems is a straightforward mathematical fact, and not a mere thought experiment. In fact, unfolding provides a recipe for creating empirically identical networks with different causal structures (for example, with arbitrarily high ϕ; see Appendix A and B). As mentioned, there even are real-world cases of feedforward and recurrent agents performing the same complex task (see Section 2.1). Hence, for example, the fact that we never have observed an unfolded cortex in practice (and probably never will) is not by itself a sufficient argument to call into question the unfolding argument. Furthermore, even though unfolded brains are impractical, we explicitly showed that the unfolding argument does not rely only on unfolded whole brains (see the previous example with the auditory system). This example can be scaled down again to the smallest part of the brain proposed to be relevant for consciousness by a given causal structure theory. Non-causal structure theories of consciousness are not subject to the unfolding argument ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Global Workspace Theory (GWT; Baars, 1997; Dehaene & Naccache, 2001), Higher-Order Thought Theory (HOTT; Lau & Rosenthal, 2011; Rosenthal, 2004) or Predictive Processing Theory (PPT; Friston, 2013) are examples of functionalist theories of consciousness: they focus on functions proposed to be crucial for consciousness. The unfolding argument does not apply to these theories because they propose that systems are conscious insofar as they implement the right kind of function – independently of the causal structure. Of course, these theories are usually couched in terms of recurrent or top-down processing, or other seemingly causal-structure terminology, but they can be formulated in other kinds of networks too. The unfolding argument only applies to theories in which reccurence per se (or another proposed causal structure) is necessary and sufficient for consciousness. For example, the typical description of GWT is that consciousness occurs when cortical areas, which code for certain contents of consciousness (e.g., sensory areas), “broadcast” their information in a global neuronal workspace that consists of highly recurrent fronto- parietal areas, thus making these contents globally available for widespread use by other areas. The crucial functions here are (a) the creation of contents in sensory areas and (b) making these contents globally available for widespread use by other areas (i.e., broadcasting). GWT is usually explained with recurrent networks. Still, equivalent feedforward networks can maintain the same broadcasting function (see Fig. 4 for a toy model).3 Similar toy models are easily produced for HOTT, PPT and all other functionalist theories. What matters is the function, e.g., broadcasting, but not the (neural) implementation. In summary, functionalist theories differ importantly from causal structure theories in that they propose functions as crucial for consciousness, independently of their implementations. For example, an unfolded global workspace network retains the crucial function of broadcasting. Hence, by their very nature, functionalist theories are not subject to the unfolding argument. Unfolding & the correlation approach ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ For many researchers, consciousness is non-physical and cannot be studied using input- output functions, and therefore might be impossible to explain with standard neuroscience (Chalmers, 1996). In this case, it is impossible to study consciousness directly. However, consciousness may still be linked to neural states by bridging principles based on correlations (Chalmers, 2004; Varela, 1996). For example, we may find correlations between human reports about their conscious experience (first-person data: I am experiencing a face) and observable properties of the brain (third person data: neural activity in the fusiform face area). To quote Chalmers (2004): «In the case of consciousness, we can expect systematic bridging principles that underlie and explain the covariation between third-person data and first-person data.» This approach cannot hold for causal structure theories because of the unfolding argument (Fig. 5). As mentioned, there are infinitely many equivalent systems that produce exactly the same first person reports as humans, but with completely different causal structures. Therefore, linking conscious properties with the brain’s causal structures by relying on first person reports cannot succeed (Fig. 5). The correlation approach cannot work with causal structure theories, although it may (or may not) succeed for other theories of consciousness.","To be considered scientific, IIT and other causal structure theories require empirical support. However, the unfolding argument shows that they are either false or outside the realm of science. For the same reason, different causal structure theories cannot be compared with each other. For example, different mathematical formulations of IIT’s axioms lead to different predictions about which systems are conscious, but we cannot compare them because the predictions are doubly dissociated from empirical data. Proponents of IIT have previously acknowledged that feedforward and recurrent networks can be functionally equivalent but have different consciousness, according to IIT (Oizumi et al., 2014). In other words, they share the same uncontroversial starting point as we do. However, conclusions differ strongly. Proponents of IIT suggest that this should prompt us to focus on the subjectivity of consciousness. In contrast, we conclude that adopting a causal structure theory precludes any experimental approach to consciousness. Indeed, we have shown that all possible experimental results, including the ones focussing on subjectivity, do not depend on causal structure. The unfolding argument rules out a class of explanations of consciousness wherein consciousness supervenes on causal structures. This should prompt us to turn our attention elsewhere in trying to understand consciousness. In this respect, the unfolding argument suggests that consciousness must be explained on a more abstract level than that of neural wiring. Indeed, any proposed framework based on neural connections suffers from the unfolding argument: any network can be replaced by equivalent feedforward networks with different connections that lead to identical empirical observations about consciousness. Only theories that abstract away implementation details and focus on explaining which kinds of functions are important for consciousness can avoid these challenges. To remain within the realm of science, consciousness must be described in terms of what it does, and not how it does it."],["Spatial scaling is the ability to transform distance information between shapes of differing sizes. Research on the developmental trajectories of spatial scaling beyond the pre-school years has been limited by a lack of suitable scaling measures for older children. Here we developed an age-appropriate discrimination scaling task, and demonstrated that children (N = 386) achieve performance gains in spatial scaling skills between 5 and 8-years-of-age, after which no significant improvements were found. Furthermore, the results support the use of relative distance strategies for task completion. These findings contrast to localisation paradigms, where performance reaches a plateau by age 6 and mental transformation strategies are used for scaling. The finding that scaling skills continue to develop until 8 years highlight the potential of scaling interventions in the early primary school years. Such interventions may infer direct benefits on spatial thinking and indirect advantages for science, technology, engineering and maths (STEM) achievement. --------------------------------------------------------------------------------","To successfully navigate, individuals must be capable of representing their own location with reference to their external environment. Hence, navigation within an environment is inherently dependent on spatial thinking. Furthermore, the use of common navigation aids such as maps or GPS systems to assist navigation requires effective spatial scaling, a particular sub-domain of spatial cognition. Spatial scaling is the ability to transform distance information from one representation to another representation of a different size (Frick & Newcombe, 2012). Scaling requires comprehension of both the symbolic and spatial correspondence between a map (model), and an associated referent space, in addition to the ability to mentally manipulate and transform spatial information between spaces of different sizes. Beyond navigation, spatial scaling is also associated with success in aspects of mathematics. For example, spatial scaling skills explain a significant proportion of the variation in proportional reasoning, above that explained by verbal intelligence (Möhring, Newcombe, & Frick, 2015). Across other domains of spatial thinking, significant age based differences in performance have been reported for children up to 8 years of age. For example, children’s performance on tasks such as perspective taking, mental folding and mental rotation improve until at least 8 years with significant individual variation reported at all ages (for example see, Frick, Möhring, & Newcombe, 2014; Newcombe, Uttal, & Suater, 2013). However, no suitable measure is available to explore the development of spatial scaling throughout the primary school years. Hence, this study aims to develop a spatial scaling measure suitable for the investigation of both developmental differences and individual differences in scaling performance in children aged 5–10 years. The development of spatial scaling ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Traditional Piagetian theories suggest that the skills required for spatial scaling are not evident until 10–11-years-of-age when children no longer hold topological views of spatial concepts (Piaget and Inhelder, 1948; Piaget & Inhelder, 1948, 1956, 1967). In contrast, more recent evidence suggests that some prerequisite skills required for spatial scaling including symbolic correspondence and metric encoding emerge in early childhood (e.g. Huttenlocher, Newcombe, & Sandberg, 1994; Newcombe, Sluzenski, & Huttenlocher, 2005; Vasilyeva and Huttenlocher, 2004). First, comprehension of symbolic correspondence, or the correspondence between a model and a referent space has been reported in children as young as 3-years-old (DeLoache, 1987, 1989). At this age children recognise that features on a map or model represent features in the real world. Second, metric encoding, or the ability to encode distances metrically, has been reported in infants as young as 5-months-old, with some infants demonstrating sensitivity to distance differences of just 20 cm (Newcombe, Huttenlocher, & Learmonth, 1999; Newcombe et al., 2005). Similarly, Bushnell, McKenzie, Lawrence, and Connell (1995) reported that 12-month-old infants can locate an object which is hidden in a circular enclosure under one of many randomly placed identical cushions. Given the lack of cues or landmarks and the random arrangement of the cushions, this suggests an ability to use metric encoding relative to the participant, in order to identify the correct cushion. Similar findings from Huttenlocher et al. (1994) propose that metric encoding in children is robust by 16 months of age. Beyond these prerequisite skills, there is also evidence that the ability to successfully map encoded distances between different sized spaces, spatial scaling, develops at a younger age than suggested by Piaget and Inhelder, with significant development in scaling proficiency between the ages of 3 and 6 years (Piaget and Inhelder, 1948; Piaget & Inhelder, 1948, 1956, 1967). Frick and Newcombe (2012) reported that children’s scaling ability, measured using a two- dimensional localisation task, improves with age from 3 to 6-years-of-age, at which time children’s accuracy levels are broadly comparable to adult scores. No significant difference in performance between 5- and 6-year-old children was reported. In a similar computer-based study, Möhring, Newcombe, and Frick (2014) demonstrated improvements in spatial scaling across different scaling factors between the ages of 4–5 years. Similar results have also been reported in studies using more naturalistic environments. For example, Vasilyeva and Huttenlocher (2004) reported that 90% of 5-year-old children tested could successfully place objects on a rectangular rug using a two-dimensional map. In comparison, only 60% of 4-year-olds were successful when presented with the same task. While some studies have reported accurate spatial scaling in children younger than 5 years, these findings may be attributable to the use of simplified tasks in which targets are presented along a single dimension. For example, Huttenlocher, Newcombe, and Vasilyeva (1999) reported accurate spatial scaling for most 3 and 4-year-old participants when tested using a scaling paradigm with a single dimension, the horizontal axis. Similarly, four-year-old children can successfully use a one-dimensional map to locate one of three target bins in a rectangular room (Shusterman, Ah Lee, & Spelke, 2008). In contrast, in more complex studies requiring scaling in two dimensions, 4 year-olds typically struggle. For example, 4-year-old children struggle to correctly place an object in a target location within a room, based on locations learnt from a corresponding map (Uttal, 1996). To conclude, in contrast to findings from Piaget and Inhelder, 1948; Piaget and Inhelder (1948, 1956, 1967), it appears that children as young as three years demonstrate symbolic correspondence and metric coding, prerequisite skills for spatial scaling. Successful performance on one and two-dimensional scaling tasks is typically evident from three years and five years respectively. Importantly, although developmental differences are evident in spatial scaling abilities, individual differences in performance are also reported for scaling performance in children up to 10-years-old (e.g. Liben and Downs, 1993). Cognitive scaling strategies ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Comprehending the strategies used for spatial scaling, and identifying how these strategies may change and adapt throughout development, may explain how and why spatial scaling has been associated with other aspects of cognition, including mathematics. Proposed cognitive strategies for spatial scaling include absolute, relative and mental transformation strategies. Absolute strategies as described by Möhring, Newcombe, and Frick (2016) involve absolute encoding of distance in one space with direct mapping of this distance onto a second space. While effective for similarly sized spaces, this strategy leads to reductions in accuracy when scaling between spaces of differing sizes. Differences in response times between scaled and non-scaled spaces are not expected when using this strategy, as the same cognitive mechanism, the direct mapping of absolute distance, is applied regardless of scaling factor. In contrast, relative scaling strategies involve the use of proportional distances (Huttenlocher et al., 1999). For example, when using a map, a target location may be encoded as being one third of the distance between two landmarks. This relative position (one third) can then be located on a referent map of a different size. Applying this scaling strategy to spaces of differing sizes leads to increased accuracy in comparison to absolute strategies. However, relative scaling strategies do not lead to differences in accuracy or response times for scaled compared to non-scaled trials, as a similar strategy is applied regardless of scaling factor. Thirdly, mental transformation strategies require the mental expansion or contraction of one space to match a second differentially sized space (Vasilyeva & Huttenlocher, 2004). This technique requires distance coding, preservation of metric distances and mental transformation. The use of mental transformation strategies is expected to reduce accuracy with increasing scaling factor. This is attributable to the fact that as scaling factor increases, the mental expansion (or contraction) required is greater, making an accurate transformation more difficult. Furthermore, in line with findings from mental rotation paradigms, response times are expected to increase with increasing scaling factor (Möhring et al., 2016). As summarised in Table 1 each of the aforementioned strategies of spatial scaling are associated with individual patterns of performance, with increasing scaling factor. Recent findings suggest that mental transformation strategies are required for effective spatial scaling (Möhring et al., 2014, 2016). Evidence outlining linear increases in both error rates and response times with increasing scaling factor have been reported in localisation tasks with children aged 4–5 years (Möhring et al., 2014) and both localisation and discrimination tasks in adult populations (Möhring et al., 2016). These performance patterns fit with the proposed mental transformation model of spatial scaling. However, there is also evidence supporting the use of relative scaling strategies in spatial scaling. Frick and Newcombe (2012) reported no linear increase in errors or response time with increasing scaling factor in children aged 3–6 years. Despite reported reductions in performance accuracy for scaled compared to unscaled trials, these results indicate the use of a relative scaling strategy. Taken together, these conflicting findings suggest differences in the use of spatial strategies across contexts and experimental paradigms. Features of spatial task design ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Scaling tasks typically require participants to use a map or model with a labelled feature to identify a corresponding feature on a referent space. However, it appears that many features of task design may influence performance on scaling tasks. Findings from Frick and Newcombe (2012) demonstrate the positive influence of landmarks and other reference points, as aids to effective spatial scaling. In contrast, reduced scaling performance is seen in paradigms with high working memory demands. For example, accuracy appears to be reduced when participants are required to remember the location of multiple targets, or in cases where stimuli and referent spaces are not presented simultaneously (Uttal, 1996). As previously described, dimensionality also influences success in spatial scaling paradigms. Performance accuracy is higher and the age at with children can successfully complete spatial scaling is lower for maps including targets distributed on one dimension (i.e. targets distributed along a single, horizontal axis) compared to targets distributed on two dimensions (i.e. targets distributed using a horizontal and vertical axis) (Vasilyeva & Huttenlocher, 2004). Finally, limitations introduced by using technology for scaling tasks, such as size limitations imposed by screen size, do not appear to influence scaling performance. For example, Möhring et al. (2016) reported no significant differences in scaling accuracy based on the absolute size of the stimuli presented. For each scaling factor tested, participants were presented with a map and corresponding referent space. No significant difference in accuracy was reported for trials in which the referent space was smaller than the map compared to trials in which the referent space was larger than the map. This suggests that the absolute size of the map and referent space used do not significantly influence accuracy at a given scaling factor. However, no known studies investigate the influence of level of acuity, on scaling accuracy at different scaling factors. Here, we propose that the degree of sensitivity (scaling precision) required to identify the correct answer is relatively low when the grid on which targets are presented is less dense (fewer, larger squares within the same surface area). This makes scaling accessible for younger children. However, for denser grids, we propose that the sensitivity required is relatively higher. Hence, 10 × 10 grid trials may highlight variation in the scaling abilities of older children. For example, the inclusion of both 6 × 6 and 10 × 10 grids may enable movement beyond the question of whether participants can scale or not, allowing for the generation of a more sensitive metric for the precision of individual’s scaling abilities. Recent developments in spatial scaling paradigms have been led by Frick, Newcombe, Möhring and colleagues who have produced a series of spatial paradigms aimed at limiting the cognitive load of scaling tasks and reducing the influence of confounding factors (Frick & Newcombe, 2012; Möhring et al., 2014, 2016). For localisation tasks, participants are shown the position of a target and are asked to find the corresponding position on a referent space (for example see Frick and Newcombe, 2012). Responses are coded as absolute deviation from the correct answer. For tasks using localisation paradigms, ceiling effects are reported by 6 years of age (for example see Frick & Newcombe, 2012). In contrast, one recent study of adults used a discrimination rather than a localisation paradigm (Möhring et al., 2016).","were asked to distinguish whether a referent map was a scaled correspondent of a target map, or not. Within discrimination paradigms participant’s responses are encoded as correct or incorrect. This mode of coding is less forgiving and does not afford any marks for responses that are close to being accurate. As such this more stringent scoring method leads to more discrete scores and may allow for the identification of subtler developmental differences in performance in older children. Furthermore, the use of a discrimination paradigm with multiple response options enables the categorisation and analysis of error types. Whilst one might assume similar frequencies of errors along the horizontal and vertical axes, a higher frequency of errors on the horizontal plane for example, might suggest more inaccurate scaling on this, relative to the vertical plane. Analysis of this type offers a novel insight into scaling processes. Current study ~~~~~~~~~~~~~ This study aims to extend previous findings by exploring age-based and individual variation in spatial scaling using a novel, age-appropriate discrimination paradigm. The results will be reported as a developmental profile of scaling ability in childhood from 5 to 10 years. The secondary aim of this study is to explore the cognitive strategies used by children in the completion of discrimination tasks. Previous studies have highlighted both mental transformation and relative encoding as possible cognitive strategies used in localisation scaling paradigms in children under 6 years of age (Frick & Newcombe, 2012; Möhring et al., 2014). However, findings from discrimination paradigms are limited to adult populations (Möhring et al., 2016). It is unknown whether similar developmental patterns in scaling performance and similar trends in cognitive strategy use, are seen for localisation and discrimination tasks in children, in young and middle childhood (age 5–10 years). The findings from this study provide important evidence on the development of spatial scaling by using a novel scaling task in which both scaling factor and required level of visual acuity are manipulated to create a suitable measure of scaling for children aged 5–10 years. The findings presented are vital for informing the design of future spatial scaling interventions aimed at improving both spatial scaling and STEM achievement more generally. Improved information on age-based differences, and indeed individual differences in spatial scaling ability, will enable the identification of age groups for which targeted scaling interventions may lead to greater gains. Furthermore, information pertaining to cognitive strategy use in spatial scaling tasks will guide and inform the approach taken in future spatial scaling interventions. Participants ~~~~~~~~~~~~ This study included 386 participants across 6 age groups. Participants were aged between 5 and 10 years. Approximately equal numbers of males (55.2%) and females (44.8%) participated in the study. All participants had normal or corrected to normal vision. The sample size, mean age and gender ratios of each age group are shown in Table 2. Participants were recruited from a middle-class, suburban London school in the UK.","Participants were tested individually in a quiet room in their school. In this task, participants were required to choose which one of four onscreen referent maps matched a printed model map. Model maps were either the same size as the on-screen referent maps, or were scaled-up versions of the referent maps (further details in materials section). The experimenter sat to the left of the participant while model and referent maps were positioned in front of the participant as shown in Fig. 1. The experimenter introduced the task as a pirate map game explaining that the yellow colouring on the maps represented sand, while the black boxes were targets showing where hidden treasure was buried. Participants were encouraged to respond as quickly and accurately as possible, by manually pressing one of the maps on the screen to indicate their answer. Following each trial a fixation dot appeared on screen, allowing the experimenter time to turn the page on the A3 flip chart and present the next trial. The task was presented as three blocks of six experimental trials preceded by 2 practice trials with a scaling factor of 1. Feedback was given for practice trials. For incorrect practice trials, participants were asked to repeat the trial until the correct referent space was selected. Only participants achieving at least 50% accuracy for practice trials (i.e. correctly answering at least one of the two practice items on their first attempt), continued to the experimental blocks. All participants successfully completed at least one of the practice trials. Between each block the task instructions were repeated. Participants received no feedback on their performance during experimental trials. Materials Task stimuli included paper-based model maps and onscreen referent maps. Model maps were presented on an A3 flipchart. Each map was positioned in the centre of a white A3 page. Onscreen referent maps including both correct and distractor maps were presented on a 13 inch Hewlett Packard touch-screen laptop in a 2 × 2 arrangement. Model maps measured 8 cm × 8 cm, 16 cm × 16 cm and 32 cm × 32 cm, for trials at a scaling factor of 1, 0.5 and 0.25 respectively. These scaling factors equated to trials in which the lengths of the referent maps were, the same size, one half the size, and one quarter the size of the model map, relative to the participant. All referent maps were 8 cm × 8 cm in size. The model and referent maps were positioned equidistantly from the participant. Consequently, the scaling factor in each trial was determined as the difference in the relative length of the referent and model maps with respect to the participant. All maps including both model and referent maps, were coloured yellow. Gridlines (for the model map only) and targets were presented in black ink. The task included three blocks of six trials. Scaling factor varied by block. Within each block, the overall area of the maps, and by extension the scaling factor, did not change. However, the density of the grid on which targets were presented, and hence the size of the grid squares and visual acuity of the maps varied. As shown in Fig. 2, half of the trials in each block were presented using a 6 × 6 square grid (requiring gross-level acuity) while the remaining targets were presented using a 10 × 10 square grid (requiring fine-level acuity). The targets displayed on each map were methodically selected to ensure a balance of left and right side targets. No targets were selected in the outer columns or rows of each grid. In order to counterbalance for any unintended effects of target position, or block order, five versions of the task were generated. First, for the initial set of targets generated, a second mirror-imaged target set was generated. This counter-balancing controlled for any potential left-right bias in target presentation. Second, to ensure that success on particular blocks could not be attributed to the specific targets used, presentation of target sets was counterbalanced between 0.5 and 0.25 scaling blocks. Overall, this created four versions of the task (Versions A-D). For these versions of the task, the order of block presentation was fixed and blocks were presented in order of increasing scaling factor (i.e. scaling factor was set at 1, 0.5 and 0.25 for Block A, B and C respectively). Finally, to confirm that the order of block presentation did not influence task performance, an additional version of the task, Version A2 was added. The targets included in Version A2 were identical to those in Version A. However, blocks B and C were presented in reverse order i.e. trials with a scaling factor of 0.25 were presented before trials with a scaling factor of 0.5. Approximately equal numbers of participants completed each task version. For each trial, four onscreen referent maps were presented including 1 correct map (i.e. the scaled (or unscaled) correspondent of the model map) and 3 distractor maps. As shown in Fig. 3, the distractor maps displayed: a vertical distractor which displayed the target one row directly above or below the correct target (A); a horizontal distractor which displayed the target one column directly to the left or right of the correct target (B) and; a diagonal distractor in which the target was positioned at one of the 4 diagonal positions relative to the correct target (C). The onscreen position of the correct map relative to the three distractor maps was randomised across trials with the correct map appearing in each quadrant of the screen with equal frequency. Analysis strategy ~~~~~~~~~~~~~~~~~ Statistical analyses were completed using IBM SPSS Statistics for windows (version 22). The use of parametric testing was determined by the outcomes of normality tests, the presence of outliers, and the relatively large sample size in this study. Unless otherwise reported, outcomes of parametric tests are reported. For analyses of variance (ANOVA) including task version, block order, gender or age group, where equal variances could not be assumed, the results for unequal variance are reported. Post-hoc Games-Howell or Tukey tests were used appropriately in cases where the assumption of homogeneity of variance was violated or met, respectively (Field, 2009). Performance accuracy, measured as percentage of correct trials, acted as the dependent variable in accuracy analysis. There was no missing performance accuracy data. However, thirty-two participants (Total N = 386) were excluded as their performance accuracy on unscaled trials was below chance (25%), suggesting that they were responding at random. Of those excluded, 15 participants were aged 5 years old, 11 participants were aged 6 years and 6 participants were aged 7 years. Mean response times for correct trials (ms), acted as the dependent variable in response time analysis. In accordance with previous literature, any response times lower than 300 ms, or above 2.5 standard deviations from the median were treated as outliers and coded as missing, prior to the calculation of mean scores (Leys, Ley, Klein, Bernard, & Licata, 2013; Ratcliff & Tuerlinckx, 2002). As mean response times were calculated from correct trials only, it was not possible to calculate mean response times for participants who failed to accurately complete at least one trial out of three for each trial type. As such, we propose that the missing response time data in this study are not random but associated with performance accuracy (for further information see Table A1 in Appendix A). Hence, excluding participants with missing response time values for any trial type (as is the case in complete case analysis) causes bias, with significant under-representation of lower performing participants. Given this consideration, for this study, missing response time data were replaced with plausible values using multiple imputation. Multiple imputation is a statistical technique in which regression modelling is used to calculate plausible values to be used in place of missing data (Rubin, 1987). The use of multiple imputation in this study is supported by the availability of auxiliary variables including performance accuracy and age, that can be included in the multiple imputation model as causes of “missingness” (Collins, Schafer, & Kam, 2001). The number of imputations was calculated using the guidelines set by Graham, Olchowski, and Gilreath (2007) based on both the fraction of missing data (g) and the tolerance for power fall-off. In this study, the fraction of missing response time information is 0.118 and the tolerance for power fall-off was set at the lowest (most conservative) level, <1%. Based on these parameters, it was determined that 20 imputations were required (Table 5, p212, Graham et al., 2007). Hence, for imputed data the pooled results from 20 imputations are reported.","Secondly, to compare the effect of block order (the order of presentation of blocks of each scaling factor) on accuracy and response time, between subject t-tests (comparing Version A and Version A2) were completed. The results indicated no significant effect of block order on accuracy, t (146) = 0.256, p = .798, d = 0.042, or response time, t (146) = 0.653, p = .514, d = 0.108. As no significant effects of task version or block order were reported, these factors were not included in subsequent analyses. Error patterns ~~~~~~~~~~~~~~ As shown in Fig. 7, chi squared analysis was used to investigate differences in the relative proportions of Vertical (V), Horizontal (H) and Diagonal (D) type errors. For all scaling factors, significantly higher proportions of H errors relative to V and D errors were reported: unscaled trials: Χ2 (2, N = 693) = 64.944, p < .001; scaling factor of 0.5: Χ2 (2, N = 948) = 84.994, p < .001; scaling factor of 0.25: Χ2 (2, N = 941) = 167.007, p < .001. Fig. 8 highlights differences in the relative frequencies of V, H and D errors across age groups. Chi squared analysis for each age group indicated significantly more H errors compared to V and D errors for all age groups (p < .001 for all). Furthermore, V errors occurred with significantly higher frequency compared to D errors for all participants aged 8–10 years (p < .001 for all). However, there was no significant difference in the frequency of V and D errors for participants aged 5 years, Χ2 (1, N = 268) = 0.731, p = .392, and participants aged 6 years, Χ2 (1, N = 240) = 1.667, p = .197. For all analysis, N values indicate the number of trials.","The primary aim of this study was to provide a developmental profile of spatial scaling in children between the ages of 5 and 10 years of age. Overall, significant age-based differences in scaling accuracy were reported. The results indicated performance gains in spatial scaling between 5 and 8-years-of-age, after which no significant improvements in task accuracy were found. These findings contrast with results from localisation paradigms, where children’s accuracy on scaling tasks reaches a plateau by age 6 (Frick & Newcombe, 2012; Möhring et al., 2014). Conversely, children’s scaling skills as measured using this discrimination paradigm continue to develop until 8-years-of-age. These contrasting results suggest that the scaling skills for placement, as required in localisation tasks, differ from those of discrimination tasks, enabling younger children to effectively complete tasks of this type (Frick & Newcombe, 2012; Möhring et al., 2014). This may be attributable to the increased scaling precision required for discrimination tasks, in addition to other domain general demands that may be needed for discrimination paradigms, such as working memory or inhibition. The findings of this study also add to previous literature on gender differences in spatial performance in childhood. Significant differences in both accuracy and response time were reported such that females had significantly faster but less accurate performance than males. This may suggest that females are applying scaling strategies less effectively then males. However, these findings should be interpreted in the context of the small effect sizes reported. These results are consistent with previous studies on spatial thinking in which non-significant or small effect sizes for significant gender differences between males and females were found (Alyman & Peters, 1993; Halpern et al., 2007; Lachance & Mazzocco, 2006; LeFevre et al., 2010; Manger & Eikeland, 1998; Neuburger, Jansen, Heil, & Quaiser-Pohl, 2011). The secondary aim of this study was to explore the cognitive strategies used by children in the completion of discrimination type scaling tasks. As previously outlined, specific patterns of performance accuracy and response times in scaling tasks, have been associated with different cognitive strategies including absolute, relative and mental transformation strategies (see Table 1). As this study included a small number of trials at each scaling factor, the findings pertaining to response time should be interpreted cautiously and seen as complementary to the key findings on performance accuracy. Nonetheless, the findings reported in this study indicated no significant increase in response time with increasing scaling factor. Indeed for gross level trials, response times were significantly shorter for trials at a scaling factor of 0.25 compared to unscaled trials. The results reported can be interpreted in the context of the aforementioned cognitive strategies for spatial scaling. Firstly, despite evidence that children use mental transformation strategies for spatial scaling in localisation tasks (for example see Möhring et al. (2014)), the use of mental transformation strategies is not supported in this study. This is particularly interesting given that mental transformation strategies have also been reported for discrimination tasks in adults (Möhring et al., 2016). However, the results reported in this study show neither reduced performance accuracy nor increased response times with increased scaling demands, which would be anticipated for mental transformation strategy use. In contrast, as scaling demands increase, relative scaling strategies are associated with unchanged accuracy and response times, while absolute strategies are expected to generate reductions in performance accuracy only. The unusual pattern of results reported in this study, and mirrored in previous work by Frick and Newcombe (2012), is not entirely consistent with either of these models. Despite showing reduced performance for scaled relative to unscaled trials, no reduction in accuracy with increasing scaling factor is observed. As such the use of absolute strategies in the completion of this task is also deemed unlikely. Although the pattern of results reported in this study is not a perfect fit to the relative scaling model, both mental transformation and absolute strategies provide poor explanations for these findings. Furthermore, despite findings from adult populations where the physical size of the maps used was not found to significantly influence scaling performance (Möhring et al., 2016), future research could investigate whether the physical size of the maps used influences performance in discrimination scaling tasks in children. Overall, the results support the use of relative strategies for discrimination tasks of spatial scaling in children. As such, participants are proposed to encode relative distances to solve scaling problems, for example by encoding that the target is one third of the way between the two sides of the grid. These findings contrast with those for discrimination paradigms in adults, where mental transformation strategies are reported (Möhring et al., 2016). These findings can be viewed in the context of other spatial domains such as mental rotation, for which there is evidence of variation in the strategies used for task completion by different individuals at different developmental stages (Geiser, Lehmann, & Eid, 2008; Glück, Machat, Jirasko, & Rollett, 2002; Janssen & Geiser, 2012). Consequently, it is unsurprising that scaling tasks with differing experimental paradigms may lead to the recruitment of differing scaling strategies. Perhaps therefore, it would not be safe to assume that all children at all ages deploy the same strategies in the completion of scaling tasks. Future studies should explore the conditions under which individuals might be encouraged to use specific cognitive scaling strategies, and whether particular features of task design promote the use of different strategies. For example, in this study, the inclusion of four maps may have encouraged the use of the relative strategy. Participants may have first encoded the target location, using relative distance information, across a single dimension (e.g. the horizontal axis), therefore enabling them to immediately discount some of distractors, before then relatively encoding the remaining maps on the other axis. This is cognitively less demanding than completing mental transformations on four individual maps. Furthermore, the presence of grid lines on the model map may have encouraged the use of relative distance information (i.e. the target is three units from the left of the map). The discrimination task used in this study allowed for the categorisation of incorrect trials into discrete error sub-categories. For all scaling factors and age groups, higher proportions of H errors compared to V and D errors were reported. Furthermore, this pattern remained stable across scaling factors. These findings indicate interesting overall differences in individual’s mapping abilities on the horizontal and vertical axis. This suggests that scaling demands do not induce specific negative effects for mapping on the horizontal or vertical axis respectively. The uniform error patterns may also suggest that participants use similar cognitive strategies to complete both scaled and unscaled trials. The use of relative scaling strategies would fit with this pattern. These findings beg the question as to why individuals appear to be more accurate in distinguishing D and V errors from the correct target, compared to H errors. For D errors, there is less shared contact area with the correct target. As such D errors are further away from the correct target and are understandably the easiest error type to distinguish from the correct target. However, shared contact with the target cannot explain the differences reported in the frequencies of H and V type errors. Both of these error types have the same degree of contact with, and are identical distances from the correct target. Alternatively, differences in the frequency of H and V type errors may be attributable to the horizontal-vertical illusion (Oppel, 1855). This is the illusion that a vertical line appears longer than a horizontal line of the same length. This illusion leads to overestimations in vertical segments relative to horizontal ones (Mamassian & de Montalembert, 2010). One explanation for this phenomenon is that vertical lines may be perceived as “receding into the third dimension” leading to perceptual errors in estimating spatial distance (p. 60, Girgus & Coren, 1975). In the context of this study, over-estimating vertical distances would lead participants to perceive V errors as further away from, and thus less likely to be confused with, the correct answer. While studies assessing the horizontal-vertical illusion typically include lines or dots displayed simultaneously on a single display (McGraw & Whitaker, 1999), the findings of the current study suggest that the horizontal-vertical illusion may extend to mapping tasks. Beyond this study, the implications of the horizontal-vertical illusion in spatial mapping tasks is largely unknown. Future studies could explore differences in horizontal mapping and vertical mapping using both localisation and discrimination paradigms. Alternatively, the observed differences in the frequency in H and V type errors may be attributable to the horizontal layout of the maps and referent spaces used in this task. As participants were required to transfer their attention horizontally from the target map to the referent spaces, focus may have been inadvertently directed to the horizontal axis. Future research could compare performance on this task when maps and referent spaces are presented using a vertical layout (i.e. with the target map presented above or below the referent spaces). The use of a discrimination task in this study offers novel insights into the cognitive processes used by participants in the completion of spatial scaling. As previously outlined, performance on scaling tasks appears to be influenced by features of task design. For example, visual acuity was shown to influence performance in this study. Variation in acuity increased the suitability of this task for a wider age range of children, such that the inclusion of fine-level acuity trials allowed for performance variation for the oldest children, whilst the gross-level acuity trials avoided floor effects in performance for the younger children. The results show that performance accuracy for trials requiring gross level acuity stabilised at 7-years-of-age, in contrast to trials requiring fine level acuity, where performance accuracy stabilised at 8-years-of-age. Future research should further investigate the role of visual acuity in scaling success. As this is the first study to investigate scaling in children using a discrimination paradigm, the exact features of task design that best enable children to complete discrimination based scaling tasks, are largely unknown. Given that this task is the first to require discrimination between four scaled spaces, and uses comparison between digital and paper-based formats, response times are high relative to other scaling tasks. Furthermore, it might be argued that the long-lasting comparison process required between four alternatives may be masking increases in response time across scaling factors. However, given that comparison between four alternatives is required for all trials, we would expect this to elevate response times across all scaling factors uniformly and not to selectively interfere with any effects of scaling or acuity. Overall the findings from this study, in particular findings on age-based performance differences, should also be viewed in light of the high levels of individual variation in task performance across children. However, taken together, the specific features of task design used in this study, have led to the generation of a measure suitable for assessing individual differences in spatial scaling abilities in older children (without ceiling effects) and younger children (without floor effects). In this study, through the use of an age-appropriate discrimination task, it was shown that children achieve performance gains in spatial scaling until 8-years-of-age, some two years older than is typically seen for localisation paradigms. These contrasting results suggest that the spatial skills required for different scaling paradigms (localisation v’s discrimination tasks) may vary. Perhaps, spatial scaling through development may best be understood by combining findings from localisation, discrimination and other paradigms. For example, future work with children of this age could investigate further unexplored paradigms such as the use of scaled maps for navigation or the reconstruction of maps from scaled models. Furthermore, the patterns of errors seen in this study suggest that mapping between differing spaces may be influenced by the horizontal-vertical illusion. Better understanding of this illusion may offer a novel way of teaching spatial mapping and improving scaling skills in children. By achieving a better understanding of the development of spatial scaling skills, children can be encouraged to engage in age appropriate spatial scaling tasks aimed at improving their scaling skills. Given that other aspects of spatial cognition are predictive of success in Science, Technology, Engineering and Mathematics domains (STEM) domains in adults, spatial scaling may also play an important role in STEM achievement (Casey, Nuttall, & Pezaris, 2001; Gunderson, Ramirez, Beilock, & Levine, 2012; Taylor & Hutton, (2013); Wai, Lubinski, & Benbow, 2009). There are many practical applications of spatial scaling in the classroom including map reading, achieving proportionality in drawing, or classroom activities in science for example, such as relating larger scaled diagrams (e.g. on a whiteboard) to smaller printed diagrams. Given their relevance in the science and maths classroom there is a need to better understand the cognitive underpinnings and development of scaling abilities. Through the use of an age-appropriate discrimination paradigm, this study offers novel insights into children’s spatial scaling abilities, highlighting the continuing development of scaling abilities in children up to 8 years of age, and the presence of substantial individual differences in spatial scaling at all developmental ages. These finding suggest room for improvement in children’s spatial scaling skills, and highlight the potential of scaling interventions in the early primary school years. Such interventions may infer both direct benefits on spatial thinking and indirect advantages for STEM achievement more generally."],["Schools are introducing more and more non-evidence-based methods in dyslexia therapy. The aim of the study is to verify whether the novel method – Warnke Method can be regarded as a useful tool in dyslexia therapy in Polish children. The research group consisted of 37 pupils, between 10 and 12 years, diagnosed with developmental dyslexia. Participants were assessed at pretest on literacy and phonological processing and tasks measuring central auditory and visual processing with Warnke Method tools. Subsequently, each child underwent 20 training sessions of Warnke Method. Afterwards, children were assessed with posttest measures. Results showed that phonological processing served as a mediator in relationship between central auditory and visual processing and reading and writing skills. Significant improvement was observed with regard to central auditory and visual processing, phonological processing, as well as reading and writing skills. Furthermore, improvement was seen in students' grades of Polish language and literature classes. --------------------------------------------------------------------------------","DSM-5 (American Psychiatric Association, 2013) defines dyslexia as an alternative term used to refer to a specific learning disorder concerning reading impairment which often coexists with difficulties in other language skills such as spelling and writing. It is primarily characterized by problems with accuracy or fluency of word recognition, poor decoding and poor spelling abilities. Despite decades of study and an ongoing search for the causes and mechanisms of developmental dyslexia, there are still no clear answers. Morton and Frith (1995) proposed considering three levels—biological, cognitive and behavioral—when analyzing and understanding the phenomenon of dyslexia. We will refer to the biological level when considering genetic predispositions and the neurobiological characteristics associated with dyslexia; the cognitive level is related to pathological mechanisms. Finally, we will use the behavioral level for discussing the symptoms of dyslexia (for instance difficulties in reading and writing). The use of the above categories allows the organization of a significant amount of knowledge regarding developmental dyslexia (Frith, 2008; Morton & Frith, 1995). Pathomechanisms for dyslexia ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Many current theories look for the causes of dyslexia at the biological level, particularly in genetic predispositions (Anthoni et al., 2012; Galaburda, LoTurco, Ramus, Fitch, & Rosen, 2006; Kere, 2014; Krasowicz-Kupis, Bogdanowicz, & Wiejak, 2014; Mascheretti et al., 2017; Matsson et al., 2015; Neef et al., 2017; Pennington & Olson, 2008; Wilcke et al., 2009). Some neurofunctional and neuroanatomical differences between individuals with dyslexia and those who do not exhibit any learning difficulties have also been reported (Bloom, Garcia-Barrera, Miller, Miller, & Hynd, 2013; Clark et al., 2014; Démonet, Taylor, & Chaix, 2004; Goswami, 2014; Habib, 2000; Jednoróg, Gawron, Marchewka, Heim, & Grabowska, 2014; Norton, Beach, & Gabrieli, 2015; Płoński et al., 2017; Richlan, 2014; Wajuihian, 2012; Xia, Hoeft, Zhang, & Shu, 2016). Because the diagnostic criteria for developmental dyslexia are various language difficulties—namely, challenges to master accuracy and/or fluency in word recognition, poor spelling, and decoding abilities (Lyon, Shaywitz, & Shaywitz, 2003)—the main approach in studies searching for the pathomechanism of the disorder stresses the role of linguistic processes. The most documented hypotheses which focus on auditory language functioning are: 1) the phonological deficit hypothesis, which examines difficulties with representation, storing, manipulating and retrieving speech sounds (Law, Vandermosten, Ghesquiere, & Wouters, 2014; Peterson, Pennington, Olson, & Wadsworth, 2014; Ramus, 2014; Ramus, Marshall, Rosen, & van der Lely, 2013; Snowling, 2000; Snowling & Hayiou-Thomas, 2006); and 2) the double deficit hypothesis which looks to deficits in both phonological processing and naming speed (Heikkilä, Torppa, Aro, Närhi, & Ahonen, 2016; Norton et al., 2014; Torppa et al., 2013; Wolf & Bowers, 1999). However, in the last 25 years, theories have been put forward which suggest that deficits in dyslexia are more than just phonological in nature and can also include visual deficits. Stein (2001) and other researchers (Gori, Cecchini, Bigoni, Molteni, & Facoetti, 2014; Gori, Seitz, Ronconi, Franceschini, & Facoetti, 2016; Jednoróg, Marchewka, Tacikowski, Heim, & Grabowska, 2011; Stein, 2014) applied magnocellular deficit theory and looked for the causes of dyslexia in anomalies in neural pathways associated with visual analysis. Individuals with dyslexia present deficits in perception organization and in manipulation of visual information (Lipowska, Czaplewska, & Wysocka, 2011; Winner et al., 2001). They experience problems with: the simultaneous processing of multiple pieces of visual information and in visual working memory (Bosse, Tainturier, & Valdois, 2007), visual-motor coordination (Bogdanowicz, 1997; Crispiani, 2015), temporal integration of visual information (Stein, 2014), functional coordination (Lachmann, 2002), and procedural learning (Biotteau et al., 2017; Biotteau, Chaix, & Albaret, 2015; Mariën et al., 2014; Nicolson & Fawcett, 2011; Nicolson, Fawcett, Brookes, & Needle, 2010; Wong & Ho, 2010). There are also theories which point to deficits in temporal processing, which particularly pertain to the processing of short duration auditory and visual elements (Daikhin, Raviv, & Ahissar, 2017; Protopapas, 1634; Szeląg et al., 2014; Tallal, 1980), as well as attention (Borkowska, 2006; Bosse et al., 2007; Dahle & Knivsberg, 2014; Facoetti, Lorusso, Cattaneo, Galli, & Molteni, 2005; Ruffino, Gori, Boccardi, Molteni, & Facoetti, 2014) in the pathomechanism of dyslexia. Language deficits influencing the clinical picture and therapy of dyslexia ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The complex etiology of specific difficulties with reading and writing is mirrored by the variability in clinical pictures of developmental dyslexia exhibited by children and adolescents. Thus, there is a need to employ diverse therapeutic methods, tailored to children's individual needs and initial causes of difficulties (Bogdanowicz & Adryjanek, 2004; Bogdanowicz, Czabaj, & Bućko, 2008; Fletcher, Lyon, Fuchs, & Barnes, 2007; Shaywitz, Morris, & Shaywitz, 2008; Terzi, 2005; Tilanus, Segers, & Verhoeven, 2016). In contrast, the majority of therapeutic methods in many different languages are aimed at directly training phonological awareness and the reading and writing skills of dyslexic pupils. In Morton and Frith's terms (Morton & Frith, 1995), all of these methods are working on the cognitive and behavioral level. English is a language with an opaque alphabetic orthography, i.e., it has many irregular letter-sound mappings. This creates difficulties both in reading and writing acquisition. It is more consistent at the level of morphological units than phonological units. Therefore more global methods of literacy training supported by phonics instruction aimed at both phonemes and larger onset-rhyme particles are the most efficient when the student is an English speaker (Gottardo, Pasquarella, Chen, & Ramirez, 2016). In contrast, Polish is a morphophonemic, inflectional, and consonantal language with semi-transparent correspondence between phonemes and graphemes (more transparent in terms of reading, less transparent in terms of spelling, although more transparent in both aspects than English; Awramiuk & Krasowicz- Kupis, 2014; Łockiewicz & Ciecholewska, 2017; Pawlicka, Lipowska, & Jurek, 2018). Therefore reading acquisition in Polish develops based on analytical and phonological strategies focusing on phonemes and syllables at the initial stage, developing in to global word- and phrase-based reading. When working with a Polish-speaking child affected by developmental dyslexia, training focuses on grapheme-phoneme correspondences, word decoding, phoneme and/or syllable segmentation, and blending skills. In particular, the following methods are commonly used: reading words, sentences, and text excerpts in syllables; alternate reading of syllables, words, and sentences; loud and subvocal reading, quiet selective reading, and syllable reading; reading using an eye-level reading ruler; and group reading (Bogdanowicz, Adryjanek, & Rożyńska, 2014; Skibska, 2016; Trypuć, 2014). Writing acquisition evolves from partial representation of speech units, through a dominant phonetic strategy, to the stage where orthographic and morphological awareness develop and are central to writing (Awramiuk & Krasowicz-Kupis, 2014). Furthermore, when practicing writing skills, therapists employ the following exercises: drawing big letters with one's hand in the air; writing large letters on a blackboard; tracing letter templates; writing letters on sheets of paper of varying sizes; writing letters using a stencil; and copying letters using carbon paper. These exercises are often supplemented by art classes, exercises aimed at improving seeing shapes and backgrounds, as well as tasks related to spatial relationships (Bogdanowicz et al., 2014; Skibska, 2016; Trypuć, 2014). Therapeutic interventions used with children in other countries are similar (Abegg & Gentile, 2016; Alexander & Slinger-Constant, 2004; Denton & Madsen, 2016; Facoetti, Lorusso, Paganoni, Umiltà, & Mascetti, 2003; Snowling & Hulme, 2012; van der Leij, 2013), with some differences reflecting the specifics of the language (e.g. orthography transparency representing the level of grapheme-phoneme correspondence). The Warnke Method in the treatment of dyslexia ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To date, specialists who apply empirical research to clinical practice have been mostly interested in detecting symptoms of developmental dyslexia. The application of research to practice has resulted in the development of an effective system of diagnosing specific learning disorders and the creation of reliable diagnostic tools (Bogdanowicz, Kalka, Karpińska, Sajewicz-Radtke, & Radtke, 2012; Fawcett, Singleton, & Peer, 1998; Feifer, 2015; Flanagan, Ortiz, & Alfonso, 2013; Hook, Macaruso, & Jones, 2001; Jaworowska, Matczak, & Stańczak, 2010; Krasowicz-Kupis, 2009; Nayton, Hettrich, Samar, & Wilkinson, 2017; Nicolson & Fawcett, 1997; Reynolds & Caravolas, 2016; Wagner, Torgesen, Rashotte, & Pearson, 2013; Wiederhold & Bryant, 2012; Wolf & Denckla, 2005). However, few evidence- based effective approaches are available for developmental dyslexia that align with the research-based diagnostic tools. Thus, creating a comprehensive system of intervention as well as diagnostic assessment for children with developmental dyslexia has become a priority (Nayton et al., 2017; Reynolds, Nicolson, & Hambly, 2003). In order to ensure that children with specific learning difficulties can benefit from effective, evidence- based forms of therapy, the current study attempts to provide data on whether the Warnke Method—a fairly new therapeutic approach which has generated significant interest among Polish teachers and therapists—is a valid and efficient therapeutic or supportive method, or at least one that does not lower children's phonological awareness and reading/writing skills when used in the therapy of developmental dyslexia. The method was developed by Fred Warnke (2000) for individuals who experience difficulties with reading, writing, and speaking. It assumes that difficulties in learning the complex skills of reading, writing, and speaking result from deficits in the processing of auditory, visual, and motor stimuli—mainly due to a decreased level of automaticity of these processes (Warnke, 2014). This method requires the use of special equipment. Diagnostic and therapeutic equipment allows the assessment of functioning in eight tasks: visual and auditory order threshold, spatial hearing, pitch discrimination, auditory motor timing, auditory choice reaction time, frequency pattern recognition, and tone duration recognition. The exercises are performed while wearing headphones (Odowska-Szlachcic & Mierzejewska, 2013). According to Warnke (2014), the method is aimed at training central auditory and visual processing with emphasis on auditory motor timing, automatic balance retainment, and auditory choice reaction time as particularly important in reading and writing. The Warnke Method differs from other well-established methods focusing on fluency (e.g. RAVE-O; Wolf et al., 2009). Above all, the main difference is that it does not directly employ typical linguistic exercises—syllable and phoneme manipulation tasks, learning about words in context, etc. The method can be therefore grounded in Tallal's (1980) ‘rapid temporal processing’ theory, which identifies basic perceptual processing as a possible cause of dyslexia, especially the apprehension of temporal auditory and visual patterns. Studies report deficits in perceptual sequence learning in dyslexics, either with use of visual (Bennett, Romano, Howard, & Howard, 2008; Howard, Howard, Japikse, & Eden, 2006) or auditory stimuli (Helenius, Uutela, & Hari, 1999; Tallal, 1980). The lower results of poorer readers on rapid naming tasks (RAN; Wolf & Bowers, 2000) may point to a sequence learning deficit, as sequential processing is also involved in RAN tasks (Bennett et al., 2008). Furthermore, the cerebellar theory (Nicolson, Fawcett, & Dean, 2001) stresses the contribution of the cerebellum to central-auditory functions, speech perception, speech timing, and, hence, phonological awareness (Stoodley & Stein, 2011). However, the question arises as to whether performing compensatory work using this procedural learning system but with the use of non-linguistic material might be beneficial for literacy development in dyslexics. Stoodley (2016) acknowledges the potential benefits. On the other hand, Kearns and Fuchs (2013) question the effectiveness of training of underlying cognitive processes in dyslexia therapy. The present study ~~~~~~~~~~~~~~~~~ Since the Warnke Method may be regarded as being based on scientific theories related to the pathomechanism of dyslexia and is generating interest among practitioners working with dyslexic children, we examined its efficacy in the context of its potential implementation (in terms of the number and frequency of training sessions) in public schools in Poland. The goal of the current study was to investigate the effectiveness of the eight tasks (visual and auditory order threshold, spatial hearing, pitch discrimination, auditory motor timing, auditory choice reaction time, frequency pattern recognition, and tone duration recognition) of the Warnke Method in the therapy of developmental dyslexia in children. Some necessary modifications (regarding the frequency of the sessions) were made so that the Warnke Method could be used in a school setting. In summary, we examined the following three research questions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ First, we examined the relationships between the variables—in particular, we were interested in whether the impact of central auditory and visual processing on reading and writing skills was direct or indirect, through phonological processing. Does phonological processing mediate the relationship between central auditory and visual processing and reading and writing skills? Furthermore, we examined two questions about the efficiency of the Warnke Method in developmental dyslexia therapy in children: Does a course of 20 sessions of training with the Warnke Method improve the functioning of children with developmental dyslexia in terms of the central auditory and visual processing? Does training of central auditory and visual processing with the Warnke Method improve phonological processing in children with developmental dyslexia as well as their reading and writing skills, decreasing the severity of their dyslexic deficits?","Forty students from the fourth and fifth grade of primary schools in northern and central Poland with a formal diagnosis of developmental dyslexia took part in the study (19 children presented severe reading problems, 15 moderate problems, and three children mild problems). The age of the study group is a result of regulations from the Polish Ministry of Education (2010) stating that developmental dyslexia can be diagnosed after the third year of primary school education, when a child is expected to have mastered reading and writing skills. Participants were recruited through contact with teachers and psychologists working in public and non-public schools. The children who took part in the study did not regularly engage in any other therapeutic intervention aimed at difficulties in reading and writing. They also did not exhibit any other types of learning difficulties (e.g. developmental dyscalculia), nor did they have any other clinical diagnoses (e.g. attention deficit hyperactivity disorder, conduct disorder, or autism spectrum disorder). Throughout the course of the intervention, three children were excluded. Two of them began to exhibit neurological symptoms (headaches, nausea) and one was excluded because of long- term hospitalization which resulted in absence from school and lack of regularity in therapeutic training. As a result, the final group comprised 37 children: 17 girls (45.95% of the sample) and 20 boys (54.05% of the sample) aged between 10 and 12 (M = 11.23; SD = 0.63).","Three research tools were used in the study: The Battery of Methods for Diagnosing the Causes of Failure at School 10/12 (Polish: Bateria Metod Diagnozy Przyczyn Niepowodzeń Szkolnych 10/12; Bogdanowicz et al., 2012) was used to assess phonological processing and reading and writing skills; the Brain-Boy Universal Professional device was used to assess central auditory and visual processing at the beginning of the project and after 20 training sessions; and the Brain-Boy Universal device was used during the 20 training sessions. The Battery of Methods for Diagnosing the Causes of Failure at School 10/12 (Polish: Bateria Metod Diagnozy Przyczyn Niepowodzeń Szkolnych 10/12; Bogdanowicz et al., 2012) is a tool used for the initial diagnosis of specific learning disorders in children aged 10–12. The tool is comprised of tests diagnosing visuospatial functions (visuospatial perception and the speed of visual perception when working with visual material) and phonological processing (phoneme differentiation, phonological memory, phoneme analysis skills, phoneme isolation, phoneme synthesis, and attention span). The battery is characterized by good content validity and satisfactory construct validity. The internal consistency of the tool (Cronbach's α) was 0.77. The validity of the tool was determined based on the factor structure and its correspondence with the psychological concept of developmental dyslexia symptoms. A relationship between battery scores and school achievement has also been confirmed (see Bogdanowicz et al., 2012). The current research used the sub-tests investigating phonological processing and reading and writing skills, as the Warnke Method is theoretically addresses children's functioning in these areas. Phonological memory Two of the above subtests (Spoonerisms and Phonological Memory) were used to build a latent variable: phonological processing. The reliability measured by the omega internal consistency coefficient for this variable in the current study was 0.67. The Spoonerisms task is the most complex, requiring a high level of phonological awareness (including phoneme differentiation, phonemic synthesis, and analysis and deletion skills) in order to obtain a high score (from seven to nine points in this task). On the other hand, repeating non-words is considered an important indirect measure of phonological working memory. One of the most important functions of phonological memory is temporary storage of incoming linguistic information. It also supports the subvocal rehearsal process. These functions have significant meaning for vocabulary acquisition, acquiring reading skills as well as reading and oral language comprehension (Baddeley, 2012; Marini, Ruffino, Sali, & Molteni, 2017; Nicolielo-Carrilho, Crenitte, Lopes-Herrera, & Hage, 2018). Reading and writing skills ~~~~~~~~~~~~~~~~~~~~~~~~~~ We used two of the above measured skills (reading words aloud measuring the reading speed and accuracy and dictation—number of errors—measuring the writing skill – writing accuracy) to build a latent variable named reading and writing skills. Reliability measured by the omega internal consistency coefficient for the created variable in the current study was 0.69. Central auditory and visual processing measured using Warnke Method tasks ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The Warnke Method (Warnke, 2000, 2014) requires the use of specialist devices, namely headphones and the following equipment: The Brain-Boy Universal Professional (BUP) device allows the accurate assessment and training of central auditory and visual processing in the tasks listed below. All tasks are performed by the child while wearing headphones. Based on the results of the initial assessment, a training program tailored to the individual needs and abilities of the child is proposed. The Brain-Boy Universal device is used mainly for training the skills listed below. The plan for its use (the sort and number of games played in each session together with the difficulty level) is based on the results of the assessment with the BUP device. The device is equipped with eight games and has options for adjusting the level of difficulty based on the user's performance. The name of the device (Brain-Boy) is intended to associate it with the Game Boy—in order the decrease the anxiety related to undergoing training with a specialized device and make children eager to train. Both the assessment and the training in the Warnke Method are administered in the form of eight different games: The Visual Brain-Boy game – visual order threshold: the child sees one flash of light on the left and one on the right side. The task is to decide on which side it appeared first; pairs of flashes of light appear with progressively shorter time intervals between them (from 400 to 5 ms). The Auditory Brain-Boy game – auditory order threshold: the child hears one click in the left ear and one in the right ear. The child needs to decide in which ear the stimulus appeared first; the pairs of sounds appear with progressively shorter time intervals between them (from 400 to 5 ms). The Klik-Boy game – spatial hearing: the child hears one click and has to decide on which side of their head it appeared. The Sound-Boy game – pitch discrimination: the child hears two sounds of different tones and has to decide in which order they appeared. The Sync-Boy game – auditory motor-timing: the child hears a regular sequence of sounds (clicks) which appear in the left and the right ear, alternately. The child needs to synchronously press buttons to the rhythm of the heard clicks. The Speed-Boy game – auditory choice reaction time: the child hears two sounds of different tones in the left and right ears. The child has to press a button as fast as possible on the side from which the lower tone came. The Trio-Boy game – frequency pattern recognition: the child hears three sounds, one of which differs from the two others in terms of the tone. The child has to identify which sound was different from the other two. The Long-Boy game – tone duration recognition: the child hears three sounds, one of which lasts longer than the other two sounds. The child needs to assess which sound differed in duration from the other two. Finally, five of the eight tasks (visual and auditory order threshold, pitch discrimination, frequency pattern recognition, and tone duration recognition) constituted the latent variable named central auditory and visual processing. Reliability measured by the omega internal consistency coefficient for this variable in the current study was 0.77. Procedure Before conducting this research, consent was obtained from the parents of students recruited to participate in the project. The parents were informed about the anonymity of the collected data and about their use for an empirical study. Research was conducted throughout the 2015/2016 school-year (approximately eight months) in primary schools in northern and central Poland. In the first stage of the research (time one), which took place after the beginning of the school year (September–October), parents filled-in a questionnaire regarding the life and medical history of the child as well as the socio- economic status of the family. In the second stage, the researcher met each child during an initial assessment at the beginning of the school year. The initial assessment included phonological processing and reading/writing skills using the Battery of Methods for Diagnosing the Causes of Failure at School 10/12 (Polish: Bateria Metod Diagnozy Przyczyn Niepowodzeń Szkolnych 10/12;Bogdanowicz et al., 2012), as well as assessment of the child's central auditory and visual processing using the Brain-Boy Universal Professional device. During the next stage, which lasted approximately eight months, the subjects underwent a series of 20 training sessions using the Warnke Method and the Brain-Boy Universal device. Six training sessions were of a combined auditory-visual character; the remaining 14 were based on auditory stimuli. The duration of a single session was approximately 30 min and as the training progressed it was shortened to about 15 min. Each training session took place in the psychologist's or speech therapist's office at school. We cooperated with the specialized group of school psychologists and speech therapists who completed a specialist course in assessment and training with the Warnke Method.1 Training sessions with the children took place weekly; however, during Christmas, Easter, and the Winter school vacation, two-week long breaks between the training sessions took place (resulting in an average of three sessions per month). In the last stage, after completing 20 therapy sessions, the children underwent a final post-test assessment (time two) using the same assessment battery as at first stage of the study. Statistical analyses Statistical analyses were conducted in two stages: in the first stage, a theoretical model describing the relations between the variables was tested; in the second stage, participants' mean scores were compared at baseline and after completing the intervention, in order to assess the effectiveness of the Warnke Method. The model of the relationship between the variables was tested using structural equation modelling (SEM). The model incorporated three latent variables: Central auditory and visual processing (indicators: Other threshold – visual, Order threshold – auditory, Pitch discrimination, Frequency pattern recognition, Tone duration recognition), Phological processing (indicators: Phonological memory and Phonological awareness), and Reading and writing skills (representing by Reading speed and accuracy and Writing accuracy). In order to estimate the parameters of the model, a maximal likelihood estimator (ML) was used. When assessing the fit of the model to data, the following measures were used: χ2, root mean square error of approximation (RMSEA), and comparative fit index (CFI; Kline, 2016). The mediating role of phonological processing was examined by testing the indirect effect of central auditory and visual processing on reading and writing skills. Following the recommendations of Cheung and Lau (2007)—especially in the context of a very small sample—we implemented a bootstrapping procedure in which 1000 bootstrap samples were created at a 95% confidence interval. In order to assess whether the indirect effects are significant, we used the bias-corrected percentile method. Calculations were conducted in an R environment using the lavaan package (Rosseel, 2012). A paired t-test was used to test differences in mean scores of the subjects at baseline and after the completion of 20 training sessions. For the central auditory and visual processing indices, results were normalized in a way such that the subjects' developmental performance change due to age was taken into account (norms for the eight tasks trained with use of the Brain-Boy Universal include ages five to twelve). In the case of phonological processing and reading and writing skills, raw scores were used for comparison.","Table 1 Presents means, standard deviations, and correlation coefficients of the analyzed variables. In line with the hypotheses, phonological memory was highly correlated with producing and recognizing Spoonerisms (r = 0.51), which supports the idea that these skills should be treated as indicators of phonological processing of the first category.2 Importantly, only phonological processing (as a latent variable) was significantly correlated with reading and writing skills. Furthermore, analysis of correlation indicates a significant positive association between central auditory and visual processing with reading and writing skills (as latent variable), as well as with phonological processing (Figs. 1 and 2). Relationship between central auditory and visual processing and reading and writing skills—the mediating role of phonological processing ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the first step of statistical analysis, a theoretical model describing relations between the investigated variables was tested. This model (see Fig. 3) includes three latent variables: 1) central auditory and visual processing, with the results of five of the eight measured tasks in the Warnke Method as indicators; 2) phonological processing, whose indicators are phonological memory (repeating nonwords) and phonological awareness (producing and recognizing Spoonerisms); and 3) reading and writing skills, whose indicators are the speed of reading and accuracy of writing (dictation). It was assumed that the relationship of central auditory and visual processing with reading and writing skills is mediated by phonological processing. The model was tested using data from the baseline measurement. The results of path analysis are presented in Fig. 3. The model fit the data well [χ2 = 29.43 (df = 25; p = .25); CFI = 0.95; RMSEA = 0.069]. Effects of training central auditory and visual processing using the Warnke Method ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In order to analyze the effectiveness of the Warnke Method in the development of central auditory and visual processing, phonological processing, and reading and writing skills, we compared the mean scores for the investigated variables before and after training (see Table 2). In line with expectations, mean scores obtained by the participants were significantly higher after 20 sessions of training, with the exception of the tone duration recognition task results. For this variable, a statistically significant difference was only observed in pupils whose baseline level of central auditory and visual processing was below the median (MBefore training = 25.00, SD = 21.64, MAfter training = 41.16, SD = 25.77, t = −2.10, p = .05). The results of children with developmental dyslexia in terms of phonological processing as well as reading and writing skills were also compared to developmental norms provided with the Battery of Methods for Diagnosing the Causes of Failure at School 10/12 (Polish: Bateria Metod Diagnozy Przyczyn Niepowodzeń Szkolnych 10/12; Bogdanowicz et al., 2012). The norm group was composed of children aged 10–12 years with no dyslexia diagnosis. The standardized scores represented level of skills expected at the specific developmental age and they were set every three months. On this basis, profiles of the severity of dyslexic difficulties exhibited by the children were created. Low scores (0–7 for Phonological memory, 0–3 for Phonological awareness, 0–65 for Reading speed and accuracy and more than 28 mistakes for Writing accuracy) indicate a high severity of difficulties. The aim of this was to verify whether the effects obtained through the Warnke Method training were due to the natural process of development over the course of the ten months of the school year or whether they were the result of the intervention. The analyses indicated a decrease in the dyslexic difficulties in the group of children studied. The percentage of low scores decreased and average and high scores increased in the group, as shown in Fig. 4. The number of children who scored high on Phonological Memory (the first of the phonological processing indicators in our study) increased (from three to eight children), while the number of children who scored low dropped (from 15 to 10). However, the distribution of scores in comparison to developmental norms changed most with regards to the Spoonerisms subtest (the second of the phonological processing indicators in our study). During the final assessment, only seven out of 37 children obtained low scores, which is about 20% of the group. During the initial assessment, 23 children (60% of subjects) scored low in this task. During the primary assessment, only one child scored in the range that qualified as high, while nine children obtained high scores on the final assessment. The children's reading skills (speed and accuracy of reading aloud—the first of the reading and writing skills indicators in our study) also improved in comparison to norms. The number of pupils who scored average on the final measurement did not change, however the number of persons scoring high increased (from one to three) and the number of pupils with low scores decreased (from 23 to 21). One might observe that this change affected two individuals, however the children were about ten months older at this point than during the primary assessment and thus also more was demanded of them. Significant differences in results compared to developmental norms were observed in the Dictation task (the second of the reading and writing skills indicators in our study), aimed at assessing writing skills in terms of accuracy of spelling, punctuation, and writing speed. Average and high scores constituted 60% of all scores during the final assessment, while at baseline they only constituted about 40%. None of the subjects scored high in their initial assessment, and at final measurement three subjects obtained such a result. The majority of scores were average at final measurement, while at baseline a considerable majority of scores were low. This improvement was also visible in the students' grades obtained during the Polish language classes (where a substantial amount of time is spent on reading, reading comprehension, and dictation). At baseline (the previous year's grade from the Polish language course), students' grades were M = 3.1, while grades at the end of the school year (after completion of the 20 training sessions) were significantly higher M = 3.92 (p < .05; with the 1–6 grading system in Poland).","The primary goal of this study was to verify whether implementing the Warnke Method in the therapy of children with developmental dyslexia for a period of approximately eight months (20 training sessions) is an effective form of improving phonological awareness, phonological memory, and, most importantly, reading and writing skills. We were also interested in the role of phonological processing in the relationship between central auditory and visual processing and reading and writing skills. The results showed that after the 20 training sessions, children's scores improved for central auditory and visual processing skills. Both visual and auditory differentiation thresholds improved as well as pitch discrimination of non-speech sounds, most probably as the result of training (since standardized scores were compared). Moreover an improvement in phonological processing was also observed at the final assessment, in terms of both phonological memory (repeating non-words) and phonological awareness (the Spoonerisms task). Comparison of the profiles of the studied pupils, before and after the intervention—in particular on tests from The Battery of Methods for Diagnosing the Causes of Failure at School 10/12 (Polish: Bateria Metod Diagnozy Przyczyn Niepowodzeń Szkolnych 10/12; Bogdanowicz et al., 2012)— indicated a decrease in dyslexic difficulties in children who participated in the intervention. The distribution of the scores with regard to developmental norms changed most on the Spoonerisms subtest. This task requires a high level of phonological awareness, syllable isolation, syllable synthesis, as well as phonological memory; therefore, it is difficult for children. It has proven to be effective at ‘spotting’ children with learning difficulties (Bogdanowicz et al., 2012). The quality of reading non-words aloud also improved in terms of the speed and accuracy in children with dyslexia in comparison to developmental norms, as well as their accuracy of writing. The number of errors with respect to spelling and punctuation was significantly lower from time one to time two both in raw and standardized scores. However the question arises as to what extent the aforementioned improved scores are due to the Warnke Method training rather than other possible factors (e.g. the education process or the home environment). Since our study was not a randomized control trial, conclusions must be drawn very cautiously. There was no control group in our study, which is a significant limitation. However the studied group comprised of children who were diagnosed with developmental dyslexia and were not undergoing other forms of therapy aimed at learning disorders (including dyslexia) during the period of the Warnke Method training. The exposure to only one type of therapeutic intervention and also comparing the obtained results to the standard scores (referring to the specific developmental age) suggest that observed changes might be attributed as an effect of the Warnke Method training. Conclusions regarding the pathomechanisms for dyslexia ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Initially, a theoretical model was tested describing the relations between the variables hypothesized to be underlying deficits for developmental dyslexia. It is worth stressing that this model turned out to appropriately describe the relations between central auditory and visual processing, phonological processing, and reading and writing skills - as previously described in theoretical models of dyslexia (Alt et al., 2017; Christo, Davis, & Brock, 2009; Heikkilä et al., 2016; Nicolson & Fawcett, 2011). The model showed that phonological processing mediated the relationship between the other two variables. We can thus say that central auditory and visual processing predicted the phonological processing of children with developmental dyslexia, which in turn predicted their reading and writing skills. This means that training with the Warnke Method seemed to directly improve phonological processing skills and, through these phonological processing skills (indirectly), reading and writing skills. Specifically, phonological processing included: 1) short term phonological memory and 2) phonological awareness including phonemic and syllable differentiation and manipulation. The obtained results support the theories stressing the importance of basic perceptual processing and automaticity in the pathomechanism for dyslexia. They also support the notions that training of underlying cognitive processes and additionally with use of the non-linguistic material may be beneficial in dyslexia therapy. The role of central auditory and visual processing and its automaticity in the development of reading and writing skills is being emphasized more often in the literature. Nicolson and Fawcett (2011) concluded, based on their own research, that children with dyslexia, apart from difficulties in language functioning, also exhibit deficits in other skills and activities which are ostensibly unrelated to reading or writing (e.g. difficulties in visuomotor coordination). In this context, auditory-motor coordination seems to play the key role (Needle, Nicolson, & Fawcett, 2015; Nicolson & Fawcett, 2011; Thomson, Fryer, Maltby, & Goswami, 2006). The functions of auditory motor timing and auditory choice reaction time seem to be particularly important (Warnke, 2014). The authors believe that problems with automaticity are the common theme in the various difficulties exhibited by dyslexic children. This is in line with the conclusions of this study pointing to the role of the more basic, lower-order processes and the importance of training them for the development of higher order processes such as phonological awareness and indirectly reading and writing. The results of our study pointed to auditory and visual order threshold as well as frequency pattern and tone duration recognition as the aspects of central auditory and visual processing that proved to predict phonological processing and indirectly reading and writing skills. The visual order threshold is proved to be involved in quick visual scanning of written and reading material (Chung et al., 2008; Giovagnoli, Vicari, Tomassetti, & Menghini, 2016) and therefore may be important in reading and writing. The auditory order threshold indicates the smallest time interval which allows one to distinguish between and sequence correctly two auditory stimuli. It is also important for the ability to divide an utterance into segments (Warnke, 2014). Tone differentiation, on the other hand, is crucial for differentiating between similar vowels. Such vowels differ in the structure of their frequency; thus, in order to understand an utterance, the ability to differentiate between tones has to be correctly automatized (Warnke, 2014). Central auditory processing allows a child to receive, memorize, and recognize sounds—in particular, the sounds of speech. Therefore it seems to be significant for phonological awareness and indirectly for reading and writing skills. Our research suggests that central auditory and visual processing training can influence phonological functioning, as well as reading and writing skills, so it may be a valuable way to decrease dyslexic difficulties. It also should be noticed that to our knowledge this is the first, though preliminary attempt to empirically verify the method which is increasingly used in the treatment of developmental dyslexia. Moreover, presented studies were conducted in accordance with the principles of evidence-based psychological practice (EBPP). Qualitative impact of the Warnke Method training ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ It's important to emphasize the emotional and motivational effect of the Warnke Method training. Qualitative observations of children during the intervention training suggested that the children enjoyed the training sessions as they resembled playing a computer game in which certain skills are being improved. Furthermore, children reported a better mood and an increase in their self-esteem, which is of special value to us because the growing body of research indicates that children's mental health is associated with academic achievement (Puskar, Sereika, & Haller, 2003; Schulte-Körne, 2016). Often, before a training session, they would tell us with a smile about their better grades, or praise received from teachers, e.g.: ‘Nobody laughed at me today when I read’, ‘The teacher praised me today during class’, or ‘For the first time, I got a better grade than 3 (equivalent to a C) for reading’. This was also confirmed by teachers we spoke to after the end of the whole project. They observed an improvement in their pupils, both with regards to reading and writing skills and in their attitude towards studying. Implications and conclusion ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our study is the first to systematically measure the functioning of dyslexic children before and after the Warnke Method training in terms of both the trained domain and higher level processes and skills (phonological processing and reading and writing skills), which, ultimately, are the goal of the training. This study addresses gaps in the field for evidence-based approaches that align with the theoretical conceptualization of dyslexia, and that can be implemented by trained practitioners in the field and incorporated into practitioners' therapeutic work. Taking the results of the study into account, we may conclude that the Warnke Method does not hinder the functioning of children in terms of phonological processing and literacy acquisition. After 20 training sessions over approximately eight months, a significant improvement was observed with regards to central auditory and visual processing, phonological processing, as well as reading and writing skills. It is worth stressing that this improvement was visible not only in assessment test scores, but also in the students' grades. More rigorous testing however is needed using randomized control trials to examine the extent to which improvements from the Warnke Method are attributed to the intervention itself. Since the Warnke Method works on non-linguistic material, it can be used regardless of language. We consider this to be an important benefit of the method. It can be used with bi- or multilingual children—both in diagnosis and in training— influencing all of their language skills. Moreover the method's structure and the game-like approach associate the training with entertainment in the children's minds. This decreases the children's anxiety levels, supports a goal-oriented approach, and increases the children's motivation to participate in other methods of dyslexia therapy. Therefore the method can be also perceived as a supportive tool, enriching the therapy of reading and writing difficulties.","Despite the fact that the study was conducted on a relatively large clinical sample of children diagnosed with developmental dyslexia, the sample size somewhat limits the generalization of the results—it would be worthwhile to replicate the study with a larger sample. The research should also be expanded to other countries and languages, which will enable to generalize conclusion about effectiveness of the Warnke Method in dyslexia therapy. Moreover, the quality of the current study would be increased if a randomized control study design were implemented. Another important aspect for further research is assessing the long-term effects of training through a follow-up assessment after the Warnke Method intervention - a year after training and also after the next stage of education (high school)."],["Inner speech is a common experience for many but hard to measure empirically. The Varieties of Inner Speech Questionnaire (VISQ) has been used to link everyday phenomenology of inner speech – such as inner dialogue – to various psychopathological traits. However, positive and supportive aspects of inner speech have not always been captured. This study presents a revised version of the scale – the VISQ-R – based on factor analyses in two large samples: respondents to a survey on inner speech and reading (N = 1412) and a sample of university students (N = 377). Exploratory factor analysis indicated a five-factor structure including three previous subscales (dialogic, condensed, and other people in inner speech), an evaluative/critical factor, and a new positive/regulatory factor. Confirmatory factor analysis then replicated this structure in sample 2. Hierarchical regression analyses also replicated a number of relations between inner speech, hallucination-proneness, anxiety, depression, self-esteem, and dissociation. --------------------------------------------------------------------------------","Inner speech – or the act of talking to yourself in your head – is an experience that is as familiar as it is elusive. While many people will report frequent inner speech (also known as silent speech or verbal thinking), operationalizing inner speech as a measurable process is extremely challenging, with some doubting that it is even possible (Schwitzgebel, 2008). It is, however, necessary for understanding the role of inner speech in psychological processing; whether as the day-to-day narrator of conscious experience, a putative facilitator of abstract thought, or a potential indicator of a developing psychopathology (Alderson-Day & Fernyhough, 2015a). Inner speech has been extensively studied as a cognitive process in research on executive functioning in children and adults, and has been linked to verbal working memory (via rehearsal; Baddeley, 2012), planning (Williams, Bowler, & Jarrold, 2012), inhibition (Tullett & Inzlicht, 2010), and cognitive flexibility (Emerson & Miyake, 2003). Typically, such research has involved blocking the production of inner speech during cognitive tasks, via distractions such as articulatory rehearsal. In contrast, the day-to-day experience of inner speech has been assessed using self-report methods – such as questionnaires and experience-sampling – which have largely focused on the frequency and content of inner speech (Brinthaupt, Hein, & Kramer, 2009; Duncan & Cheyne, 1999; Klinger & Cox, 1987). While the former approach may be thought to lack ecological validity, the latter has been accused of producing unreliable findings with uncertain construct validity (Uttl, Morin, & Hamper, 2011; although see Hurlburt, Heavey, & Kelsey, 2013). A self-report tool that takes an alternative approach – assessing the phenomenological form and quality of inner speech – is the Varieties of Inner Speech Questionnaire (VISQ; McCarthy-Jones & Fernyhough, 2011). Based on Vygotsky’s (1934/1987) developmental theory of inner speech, the 18-item VISQ asks participants to rate the phenomenological properties of their inner speech according to four factors: dialogicality (inner speech that occurs as a back-and-forth conversation), evaluative/motivational inner speech, other people in inner speech, and condensation of inner speech (i.e. abbreviation of sentences in which meaning is retained). In its initial development with university students, the VISQ demonstrated good internal and test-retest reliability, and significant relations with self-reported rates of anxiety, depression, and proneness to auditory but not visual hallucinations (McCarthy- Jones & Fernyhough, 2011). A follow-up study in a similar sample also implicated relations with self-esteem and dissociation, with the latter partly mediating the link between inner speech and hallucination-proneness (Alderson-Day et al., 2014). Since then the VISQ has been used to assess inner speech in people with psychosis (de Sousa, Sellwood, Spray, Fernyhough, & Bentall, 2016) and explore relations with reading imagery (Alderson-Day, Bernini, & Fernyhough, 2017), along with being adapted for use in Spanish (Perona- Garcelán, Bellido-Zanin, Senín-Calderón, López-Jiménez, & Rodríguez-Testal, 2017), Colombian (Tamayo-Agudelo, Vélez-Urrego, Gaviria-Castaño, & Perona-Garcelán, 2016), and Chinese populations (Ren, Wang, & Jarrold, 2016). Self-reported dialogic inner speech on the VISQ has also been observed to correlate with neural activation of areas linked to producing inner dialogue during an fMRI task (Alderson-Day et al., 2016). Notwithstanding its successful use in these contexts, the VISQ also has shortcomings. First, its focus on the key features of a Vygotskian model of inner speech (such as dialogue and condensation) may have neglected other aspects of inner speech phenomenology, such as use of metaphorical language, speaker position, and feelings of passivity. For example, Hurlburt et al. (2013) make a distinction within inner speech between “inner speaking” and “inner hearing“, with the latter being more like an experience of hearing one’s own voice played back on a tape recorder. While subtle, such distinctions are potentially significant given the putative roles of metaphor, positioning and agency in anomalous and psychopathological experiences like hallucinations or thought insertion (Badcock, 2016; Hauser et al., 2011; Mossaheb et al., 2014). Second, the original VISQ did not generally capture positive and regulatory aspects of inner speech phenomenology, which can be a key part of the functional role of dialogue-like verbal thinking. Puchalska-Wasyl (2007, 2016), for example, specifies seven positive functions of dialogue, including support, insight, exploration and self-improvement. Such aspects of inner speech permit more direct comparisons with the way the concept is studied in other fields, such as cognitive research or sports psychology (Hardy, Hall, & Hardy, 2005). In contrast, nominally neutral statements on the original VISQ, such as “I think in inner speech about what I have done, and whether it was right or not” may in some cases have been interpreted in a primarily negative and ruminative sense. This is supported by the fact that more people endorse such statements if they also report lower levels of self-esteem (Alderson-Day et al., 2014). Here we present a revised edition of the VISQ with the aim of addressing some of these concerns and expanding the phenomenological scope of the measure. First, an extended 35-item version of the scale was trialled as part of a larger study on reading imagery (n = 1472; Alderson-Day et al., 2017), and exploratory factor analysis was used to identify a new five-factor model. Second, confirmatory factor analysis was used to assess the fit of the new model in a new sample of 377 university students. Third, we used this sample to attempt to replicate prior findings linking VISQ factors to anxiety, depression, hallucination-proneness, dissociation, and self-esteem (Alderson-Day et al., 2014; McCarthy-Jones & Fernyhough, 2011). We predicted that (i) certain features of inner speech would predict auditory but not visual hallucination-proneness (specifically, dialogic, evaluative, and other people in inner speech); (ii) dissociation would partially mediate the relation between inner speech and hallucination-proneness, and (iii) inner speech would also be related in meaningful patterns to anxiety, depression, and self-esteem. Finally, in order to further explore links to psychopathology, we (iv) compared inner speech characteristics in students with and without a self-reported psychiatric diagnosis. As very little work has been done in on this topic, we made no predictions about group differences: such exploratory information is nevertheless informative for future investigation of inner speech in clinical groups. Sample 1 A total of 1566 participants took part in an online survey on readers’ inner voices (Alderson-Day et al., 2017). The participants were recruited through a series of blog posts for the Books and Science sections of the Guardian website (‘Inner Voices’), social media and publicity at the Edinburgh International Book Festival. The participants that responded to at least 80% of the new VISQ items were included in the analysis (n = 1472, Age M = 38.84, SD = 13.42, Range 18–81). Responses primarily came from English-speaking countries (see Table 1 for demographic information). Sample 2 377 participants (322 females), aged 18–56 (M = 20.02, SD = 3.24) were recruited from university settings. The study was advertised through a departmental participant pool. Ethical approval was given by a university research ethics committee. Participants received course credit or gift vouchers for their participation. Surveys for both sample 1 and sample 2 were delivered on the Bristol Online Survey platform. Varieties of inner speech questionnaire – Revised (VISQ-R) Phenomenological properties of inner speech were evaluated with the 35-item Varieties of Inner Speech Questionnaire – Revised (VISQ-R). In addition to the original 18 items on dialogicality, evaluation, condensation, and the presence of other people in inner speech, the extended version contained new items on literal and metaphorical use of language, speaker positioning and address, and regulation of different moods. The new items were derived in an iterative manner by a working group including the original VISQ authors (SMJ & CF), plus members of the research team with expertise in phenomenology, analytic philosophy, and cognitive science (BAD & SW). A large set of items were initially generated (>50) in response to the aims of the scale, and then these were refined down to a manageable number for exploratory testing. Responses were made on a seven-item frequency scale where respondents evaluated how frequently the inner speech experiences occurred, ranging from “Never” (1) to “All the time” (7). This differed from the original 6-point rating scale based on agreement with statements (e.g. “Certainly applies to me”), in an attempt to be more precise about how often such experiences occur (cf. Hurlburt et al., 2013). Each subscale of the original VISQ has previously shown high internal reliability (Cronbach’s α > .80) and moderate to high test–retest reliability (>.60). See Table 2 for a list of old and new items. Revised Launay-Slade Hallucination scale (LSHS-R; McCarthy-Jones & Fernyhough, 2011; Morrison, Wells, & Nothard, 2000) This scale included nine items used by McCarthy-Jones and Fernyhough (2011), adapted from Morrison et al. (2000)’s Revised Launay-Slade Hallucination scale, including five auditory hallucination and four visual hallucination statements (e.g. “I have had the experience of hearing a person’s voice and then found that there was no one there”). Ratings are made on a four-point Likert scale ranging from “Never” (1) to “Almost always” (4). Scores can range from 9 to 36, where higher scores indicate greater hallucination-proneness. The auditory and visual subscales of the revised LSHS-R have been shown to have adequate internal reliability (Cronbach’s α > .70, e.g. McCarthy-Jones & Fernyhough, 2011). Hospital anxiety and depression scale (HADS; Zigmond & Snaith, 1983) Levels of anxiety and depression were assessed with the 14-item Hospital Anxiety and Depression Scale (HADS; Zigmond and Snaith, 1983). This scale comprises of seven items relating to anxiety (e.g. “I get sudden feelings of panic”) and seven items relating to depression (e.g. “I have lost interest in my appearance”). Responses are made on a four-point Likert scale, with the total scores ranging from 0 to 21. Higher scores indicate higher levels of anxiety and depression. This scale has been used extensively and shown to have satisfactory psychometric properties (Zigmond & Snaith, 1983). Dissociative experiences scale – Second revision (DES-II; Carlson and Putnam, 1993) Frequency of dissociative experiences was measured with the 28-item self-report Dissociative Experiences Scale. Participants are asked to indicate what percentage of the time they experience dissociative states, such as feelings of derealisation or absorption (e.g. “Some people sometimes have the experience of feeling that their body does not belong to them”). Answers can range from 0% to 100%. The original DES has a mean internal reliability of .93 (van IJzendoorn & Schuengel, 1996). Rosenberg self-esteem scale (RSES; Rosenberg, 1965) The 10-item Rosenberg Self-Esteem Scale includes five positive and five negative statements on self-concept and social rank (e.g. “On the whole, I am satisfied with myself. “). Responses are made on a four-point scale ranging from “Strongly agree” to “Strongly disagree”. The scale has been shown to have high test–retest and internal reliability (e.g. Fleming & Courtney, 1984).","All data were analysed in SPSS 20, with the exception of the confirmatory factor analysis (which was conducted using AMOS 22) and the mediation analysis (for which we used the “medmod” package in jamovi (Version 0.9). Relations between variables were assessed using Pearson’s product correlation co-efficient tests and hierarchical linear regression. Alpha values for correlations were Bonferroni corrected to account for multiple comparisons. Sample 1: Exploratory factor analysis of VISQ-R (35 items, n = 1472) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Missing data constituted less than 1% of total responses for the extended VISQ and were replaced by mean values. Exploratory factor analysis (EFA) was performed using Principal Component Analysis (PCA) with oblique rotation (direct oblimin), allowing for inter- correlation of the VISQ factors. Items were removed based on communalities under 0.4 or if they failed to load >0.5 onto a single factor (Costello & Osborne, 2005). For all of the below models, KMO statistics and Bartlett’s test values were within acceptable ranges. The initial solution returned by the analysis produced 8 factors with eigenvalues over 1 (68% of variance explained). However, items 24, 26, 29, and 33 had either low communalities or low factor loadings, and these were subsequently removed from the analysis. The following model identified 7 factors with eigenvalues over 1 and accounted for a greater amount of variance (72%) but this included one factor with only one item (item 23). Inspection of the scree plot suggested five main factors with eigenvalues over 2, with two more factors barely above 1. Forcing a five-factor solution (62% variance explained) led to five items not loading on a single factor (items 21, 22, 23, 30, and 34) of which three also failed the communalities test (Items 22, 23, 30). With a total of nine items removed, a stable model was found with five factors (68.43% of variance explained), with all items loading on at least one factor over 0.5 and with all communalities > 0.4 (KMO statistic = 0.895, Bartlett’s X2 (325) = 23773.10, p < .001). The five factors, displayed in Table 3, broadly mapped on to the original VISQ structure but with an expanded, 6-item evaluative/critical factor (Items 20, 23 and 24 are new) and a new, 4-item positive/regulatory factor (Items 19, 22, 25 and 26 are new). As for the previous VISQ, internal reliability of each factor was excellent (dialogic: 0.87, evaluative/critical: 0.88, positive/regulatory: 0.80, condensed: 0.87, other people: 0.91). Table 4 shows that a number of the factors were inter-correlated: dialogic inner speech was most closely related to each of the factors, followed by evaluative/critical inner speech. Sample 2: Confirmatory factor analysis of VISQ-R (26 items, n = 377) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Confirmatory factor analysis was used to test the five-factor model in sample 2. Less than 0.2% of the data were missing and replaced by the mean response (per item). Frequency of responses for each item (by subscale) are displayed in Table 5. Maximum likelihood estimation was used for model fitting. Following an initial model containing no covariances between the factors, modification indices indicated improved fit if covariances were allowed for each factor to correlate with dialogic and evaluative/critical inner speech, but no pairwise covariances between the other three factors. Covariances between error terms were also added (within factor only) where modification indices suggested improved fit. This resulted in a stable model with a significant χ2 value, χ2 (2 6 0) = 562.53, p < 0.001, but acceptable CMIN/DF ratio of 2.16. Fit statistics for the model were also in a satisfactory/good range: CFI = 0.925, RMSEA = 0.056 (90% C.I = 0.048–0.062), Hoelter index = 200. Replicating McCarthy-Jones and Fernyhough (2011) and Alderson-Day et al. (2014) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Using sample 2, we then attempted to replicate our previous findings using the new VISQ-R. Table 6 displays the descriptive statistics for the five dimensions of inner speech, as well as scores for hallucination-proneness (LSHS-R), anxiety and depression (HADS), self- esteem (RSES), and dissociation (DES). Mean scores were highest for the evaluative inner speech dimension, followed by dialogic, condensed and positive inner speech. The lowest mean scores were recorded for other people in inner speech. Reliability for dialogic, evaluative and other people in inner speech was high (α > 0.8) but lower for positive inner speech (α = 0.6). Cronbach’s α was also found to be adequate for visual hallucination-proneness, HADS anxiety, HADS depression RSES and DES (α > 0.7), but on the boundary of acceptability (0.67) for auditory hallucination-proneness. Table 7 shows the bivariate correlations between VISQ-R, DES, RSES, HADS and LSHS-R scores. To account for testing the five VISQ-R factors against six psychopathology variables simultaneously, a Bonferroni corrected alpha-value of 0.0017 was applied for significance (i.e. 0.05/30). Auditory hallucination-proneness (LSHS-R) was positively associated with evaluative, dialogic and other people in inner speech, and visual hallucination-proneness. A similar pattern was observed for visual hallucination-proneness. Evaluative and other people in inner speech also correlated positively with HADS Anxiety scores, HADS Depression scores, and lower self-esteem. Frequency of dissociative experiences (DES) was positively correlated with dialogic and other people in inner speech. For exploratory purposes, we also compared male and female participants’ scores on the new VISQ-R factors using independent samples t-tests: the only difference observed was for evaluative/critical inner speech (t (372) = −2.84, p = 0.005), with female participants (M = 30.34, SD = 7.13) scoring higher than male participants (M = 27.29, SD = 4.45). Predicting hallucination-proneness controlling for anxiety and depression (McCarthy-Jones & Fernyhough, 2011) A hierarchical multiple linear regression was performed to assess unique contributions of the five dimensions of inner speech (VISQ-R), depression and anxiety (HADS) in predicting auditory hallucination-proneness. Age, gender, depression, anxiety and proneness to visual hallucinations were entered in the first step, and five subscales of the VISQ-R in a second step, with auditory hallucination-proneness (LSHS Auditory) as a dependent variable. Measures of multicollinearity were in an acceptable range. Both Block 1 (F (5, 371) = 21.05, p < 0.001) and Block 2 (F (10, 366) = 12.47, p < 0.001) significantly predicted auditory hallucination-proneness. In Block 1 (R2 = 0.22), only visual hallucination-proneness was found to be a significant predictor of auditory hallucination-proneness (β = 0.46, p < 0.001). In Block 2 (R2 = 0.25; Δ R2 = 0.03, Δ F(5, 366) = 3.26, p = 0.007), visual hallucination-proneness, β = 0.42, p < 0.001, and other people in inner speech, β = 0.13, p = 0.005, significantly predicted auditory hallucination-proneness. To test for the specificity of this relationship, we then reran the above analysis with visual hallucination-proneness as the dependent variable and auditory hallucination-proneness as a predictor variable in Block 1. Both Block 1 (R2 = 0.36, F (5, 371) = 41.55, p < 0.001) and Block 2 (R2 = 0.61, F (10, 366) = 21.51, p < 0.001) were significant, but importantly only auditory hallucination-proneness, β = 0.36, p < 0.001, and depression, β = 0.32, p < 0.001, were observed to significantly predict visual hallucinations in the final model: all of the VISQ-R factors were non-significant (p > 0.11). Therefore inner speech characteristics appeared to be relevant to predicting auditory hallucination-proneness specifically. Predicting hallucination-proneness controlling for self-esteem (RSES) and dissociation (DES) (Alderson-Day et al., 2014) A multiple regression was performed to assess the contribution of different types of inner speech (VISQ-R), self-esteem (RSES) and dissociative tendencies (DES). Age and Gender were entered in the first step, followed by the five subscales of VISQ-R in the second step. Self-esteem (RSES) was entered in the third block, and dissociative tendencies (DES) in the fourth block. The first model with age and gender as predictors of auditory hallucination-proneness was not significant (p > 0.05). The addition of the five dimensions of VISQ-R in Block 2 made a significant change to the model (Δ R2 = 0.09, Δ F (5, 369) = 8.16, p < 0.001), where other people in inner speech, β = 0.21, p < 0.001, and evaluative inner speech, β = 0.13, p = 0.023, both significantly predicted auditory hallucination-proneness (R2 = 0.103, F(7,369) = 6.077, p < 0.001). The addition of the self-esteem measure (RES) to Block 3 did not make a significant change to the model (Δ R2 = 0.03, p > 0.05), and self-esteem did not significantly predict hallucination-proneness (p > 0.05). The addition of dissociation scale (DES) scores to Block 4 resulted in a significant change to the model (Δ R2 = .17, Δ F(1, 367) = 83.88, p < 0.001), with dissociation found to significantly predict auditory hallucination-proneness, β = 0.44, p < 0.001, as well as other people in inner speech, β = 0.17, p < 0.001 (final model: R2 = 0.273, F(9,367) = 15.281, p < 0.001). As in Alderson-Day et al. (2014), we then examined what (if any) mediating role dissociation played in relationship between inner speech and hallucination-proneness. Using the medmod package in jamovi, we tested the direct and indirect effects of other people in inner speech on auditory hallucination-proneness, with DES scores as the mediating variable. Significant effects were apparent for both direct (Z = 4.16, p < 0.001) and indirect (Z = 3.04, p = 0.002) paths, with the latter accounting for 27.9% of the total effect. This suggested that dissociation partially (rather than fully) mediated the effect of other people in inner speech on hallucination-proneness, and that most of this effect was in fact direct. Confounding factors: Item overlap, language status, and psychiatric diagnosis Finally, we reran our main analysis accounting for some potential confounds in the data. One concern with the original VISQ is the inclusion of items that refer to experiencing “actual voices” of other people in their inner speech, which may be thought to index similar experiences as the auditory items of the LSHS. To address this, Alderson-Day et al. (2014) reran their analyses removing items 12 and 16 of the VISQ from the other people subscale. When this was done for the present data, the analysis showed the same results as the regression model 1 (i.e., McCarthy- Jones & Fernyhough 2011), in that visual hallucinations were a significant predictor of auditory hallucinations in the first block (β = 0.46, p < 0.001), and visual hallucinations (β = 0.43, p < 0.001) and other people in inner speech significantly predicted auditory hallucinations in the second block (β = .12, p = 0.013). The analysis also showed the same results as in regression model 2 (i.e. Alderson-Day et al., 2014), where both evaluative (β = 0.14, p = 0.013) and other people in inner speech (β = 0.20, p < 0.001) were found to be significant predictors of auditory hallucination-proneness. The addition of self-esteem scores did not make a significant change to the model (p > 0.05). In Block 4, both dissociation scores (β = 0.44, p < 0.001) and other people in inner speech were found to be significant predictors of auditory hallucination-proneness (β = 0.15, p = 0.001). Other factors which may have affected our data include the presence of non-native speakers and people with a self-declared psychiatric diagnosis. To address these, we reran our main analyses of sample 2 first excluding non-native English speakers, and second excluding those with a psychiatric diagnosis. For the former, the results for the native speakers group (n = 297) showed the same results as reported in regression model 1, in that visual hallucination-proneness (β = 0.48, p < 0.001) and other people in inner speech (β = 0.13, p < 0.05) were significant predictors of auditory hallucination-proneness in the final model. Similarly, for regression model 2, other people in inner speech (β = 0.19, p < 0.001) and dissociation scores in (β = 0.45, p < 0.001) were significant predictors of auditory hallucination-proneness in the final model. Identical results were observed when the analyses were rerun only in people without a psychiatric diagnosis. Finally, we compared those with (n = 43) and without a psychiatric diagnosis (n = 326) on the five subscales of the VISQ-R (8 participants preferred not to report their diagnostic status). The independent- sample t-tests showed that there was a significant difference in the level of reported dialogic inner speech between the two groups (t (367) = 2.08, p = 0.038), with the sample with psychiatric diagnoses reporting higher levels of dialogic inner speech (Dx group M = 23.07, SD = 5.89; no Dx group M = 21.20, SD = 5.47). A similar pattern was evident for evaluative inner speech (t (367) = 5.43, p < 0.001), with lower scores in the sample without a diagnosis (Dx group M = 35.29, SD = 7.47; no Dx group M = 29.12, SD = 6.94). No group differences were found in levels of condensed, positive and other people in inner speech (p > 0.05).","Asking people to report on their inner speech is challenging, but with careful methodological considerations it can produce consistent results. Here we tested and confirmed a five-factor model for an expanded VISQ, the VISQ-R; introduced a new variable of positive/regulatory inner speech; and replicated a number of previous findings of relations between inner speech variables and psychopathology. Compared to the original scale, the VISQ-R captures a broader range of phenomenology associated with the experience of self-directed speech, in line with the primary aim of the study. In particular, new items relating to positive and negative states of inner speech, including the use of inner speech to regulate mood, survived the various stages of scale development. The inclusion of these new items appears to have elaborated the concept of evaluative inner speech captured in the original VISQ (McCarthy-Jones & Fernyhough, 2011). Evaluative states in the present study were clustered with statements about inner speech contributing to feeling anxious or depressed, while motivation and regulation in inner speech clustered around the new positive factor. This is consistent with prior findings of higher evaluative inner speech being associated with a more negative self-concept (Alderson-Day et al., 2014), even though the items included in the original evaluative factor were ostensibly neutral and focused more on deliberative states (e.g., I think in inner speech about what I have done, and whether it was right or not). The inter-correlations of evaluative and positive subscales in sample 1 (r = 0.40) and sample 2 (r = 0.29) also highlight that these are related but ultimately separable factors, rather than positive and negative poles of the same scale. A secondary aim for the new scale was to capture the potential for inner speech to have positive psychological effects, as this is important for drawing together disparate strands of research on self-directed speech. In sports psychology and similar fields, self-talk has often been associated with improved focus and performance (Hardy, Begley, & Blanchfield, 2015; Hardy et al., 2005). However, such research often elides the distinction between overt and covert self-talk. Research on private speech (overt or out-loud self-talk) shows it to be associated with regulatory strategies in childhood (e.g., Fernyhough & Fradley, 2005) and in adulthood (Duncan & Cheyne, 2001). In contrast, the phenomenology of inner speech and its relations to psychopathological states and processes (such as rumination) have rarely been explored. Through our identification of a positive inner speech factor – one with clear parallels with the overt and motivational self-talk deployed in sports research – we have the potential, in future research, to test predictions about cognitive and behavioural performance. For example, one could hypothesise that participants with more positive/regulatory inner speech are likely to perform better on tasks requiring self-talk (or following instructions to explicitly use such a strategy), but those with more evaluative/critical inner speech may not. While some attempts have been made to link different functions of private speech to aspects of task performance, observational studies of private speech are fraught with difficulty (Winsler, Fernyhough, & Montero, 2009). The VISQ-R thus provides us with a new way of empirically linking aspects of self- talk to aspects of performance. The cognitive benefits of positive/regulatory inner speech may also be evidenced in the spheres of creativity and imagination. To give one example, the presence of an imaginary companion in childhood has been linked to greater levels of self-talk in general in childhood (Davis, Meins, & Fernyhough, 2013) and adulthood (Brinthaupt & Dove, 2012). Evaluative/critical inner speech, in contrast, would be expected to relate more strongly to rumination, shame, and perfectionism (Flett, Madorsky, Hewitt, & Heisel, 2002; Orth, Berking, & Burkhardt, 2006). The gender difference observed for evaluative/critical inner speech in the present data – with higher scores in female participants – is consistent with greater rates of rumination in women (Nolen-Hoeksema & Jackson, 2001) and with findings from the previous VISQ (Tamayo-Agudelo et al., 2016). Even with the introduction of the new factor, a number of previous relations to psychopathology were observed in the confirmation sample of university students. In line with the original VISQ study, specific characteristics of inner speech were associated with a greater proneness to auditory hallucinations but not visual hallucinations (hypothesis 1); dissociation appeared to mediate this relationship (hypothesis 2); and a subset of inner speech characteristics also correlated with scores for anxiety, depression, and self-esteem (hypothesis 3; Alderson-Day et al., 2014; McCarthy-Jones & Fernyhough, 2011). One difference from the original study concerns the role of dialogic inner speech. Whereas the dialogic factor predicted auditory hallucination-proneness more strongly than other factors in the original study, in the present study only other people in inner speech did so. It seems unlikely that this is due to the new scale structure: with the original VISQ, pairwise correlations between dialogic inner speech and hallucination-proneness were often evident (e.g., Alderson-Day et al., 2017) but did not always survive exclusion during hierarchical regression analysis (Alderson-Day et al., 2014). The exploratory analysis of those with and without a psychiatric diagnosis suggested here that dialogic inner speech (and evaluative/critical inner speech) may be greater in those with a self-reported diagnosis (cf. the marginally significant reduction in dialogicality in the inner speech reported by Langdon, Jones, Connaughton, & Fernyhough, 2009). In contrast, a recent study with a clinical sample suggested that condensed and other people in inner speech were related to psychopathology (de Sousa et al., 2016). As a potential explanation of these disparate findings, dialogic inner speech is typically the factor that correlates most with all of the other VISQ factors; this is also the case with the findings reported here. It may be that dialogicality acts as a “core” feature of inner speech, and the other VISQ factors mediate its relation to other variables. Alternatively, it may be that the concept of dialogicality is liable to varying interpretations, leading to inconsistent performance across samples: for example, in Spanish translations of the original VISQ (Perona-Garcelán et al., 2017) a further dialogic factor needed to be added that explicitly referred to inner speech in terms of position of self in a dialogue (see also Tamayo-Agudelo et al., 2016). In our experience, English-speaking users of the scale have typically interpreted dialogic items to both include dialogues with oneself and dialogues with another (indeed, the focus of a Vygotskian understanding of dialogicality is arguably in structure, rather than the identity of interlocutors), but it is possible that cultural understandings of dialogue will differ considerably depending on the language used. A further point to note is that the idea of hallucinatory experiences existing on a continuum stretching into the general population has been increasingly challenged (e.g., Garrison et al., 2017), and it may need to be acknowledged that relations between hallucination-proneness and inner speech variables may transpire differently in clinical and non-clinical samples. As in earlier studies (Alderson-Day & Fernyhough, 2015b; Ren et al., 2016), condensed inner speech items were not strongly endorsed, and correlations with other inner speech factors were low. In general, condensed inner speech scores also show few correlations with non-VISQ variables (Ren et al., 2016). As noted above, there is evidence that patients with psychosis endorse this experience more than controls, and that it relates to increased levels of thought disorder (de Sousa et al., 2016). As such, it would seem to be an important factor to retain and may be more informative when used in clinical samples. This may also have been the case for a number of the other new items that did not survive the initial exploratory factor analysis in sample 1. Specifically, questions regarding control, surprise, metaphor, and feelings of passivity in inner speech (such as experiencing it as hearing rather than speaking) did not cluster sufficiently either with themselves or the existing factors to be included in the final scale. While still potentially important experiences to explore, it may be that these characteristics of inner speech – along with condensed inner speech – are too extraordinary for non-clinical respondents and do not become salient until things start to go awry. Such a question bears on contemporary debates in philosophy of mind about how and why we experience our thoughts as our own (Roessler, 2016) – questions that are important to consider when probing the distinction between inner speech, auditory verbal hallucination, and other atypical experiences such as thought insertion (Wilkinson & Alderson-Day, 2016). Some limitations to the present study are important to consider. First, asking people to report on characteristics of their own inner speech via questionnaires raises the perennial concern about how reliably people can report on their own inner experience (see, for example, Alderson-Day & Fernyhough, 2014; Hurlburt et al., 2013; Hurlburt & Schwitzgebel, 2007). Notwithstanding the fact that such self-report methods are relied on extensively in personality and individual differences research, it is important when studying inner speech to employ multiple methods. We therefore recommend that future use of the VISQ-R involve it being deployed alongside other tools such as cognitive tasks (Ren et al., 2016) or neuroimaging (Alderson-Day et al., 2016). Methods of tracking the characteristics of inner speech are developing all the time, with a recent example being provided by Whitford et al. (2017), who used EEG to demonstrate perceptual capture of external sounds during inner speech production. Second, examining the relations between conceptually similar notions (such as inner speech and hallucination-proneness) raises the risk of measurement overlap. We have taken steps to address this issue: for example, by showing that the relation between other people in inner speech and auditory hallucination-proneness remains even when potentially overlapping items are removed, concerns about direct conceptual overlap are to some degree minimised. Moreover, we have endeavoured to design VISQ-R items that focus on the form and phenomenology of inner speech, rather than its contents (to avoid confounding the content of a thought or mood and its structure). Nevertheless, ideally such issues would be explored more extensively in a procedure that would enable clarification and disambiguation, such as a one-to-one interview. This would seem to be particularly important when dealing with patients who may struggle to report on their own inner experience (de Sousa et al., 2016). Finally, both of these samples reflect highly-educated populations: in one case respondents to a survey on reading experiences promoted by the Guardian newspaper (Alderson-Day et al., 2017), and in the other a university sample. Although Alderson-Day et al.’s (2017) sample was an international one, the majority of its respondents were still either from the UK or from other western, English-speaking countries (as was the case for sample 2). We have already discussed some language- specificity issues relating to translations of the VISQ. For these and other reasons, the generalisability of the present findings is limited, and it will be particularly important to test the VISQ-R in more mixed, general population samples and non-English speaking countries. A related issue that will be important for future research on inner speech is to expand its horizons into more diverse populations. For example, research on inner speech (or equivalent experiences of inner language, such as signing) in those who are blind or deaf is still meagre (for exceptions, see Campbell & Wright, 1990; Zimmermann & Brugger, 2013). Similarly, research on neuro-diverse populations – such as autistic people – has been largely confined to examining “deficits” in self-talk rather than exploring qualitative differences in inner speech or more broadly, inner experience. Evidence of differential use of verbal strategies on cognitive tasks in autistic adults (Williams et al., 2012) and vivid accounts of inner experience by adults with Asperger Syndrome (Hurlburt, Happé, & Frith, 1994) suggests that inner speech may be radically different for this group, and deserving of a more systematic exploration. In conclusion, inner speech has been proposed as a key tool for unlocking creative, exploratory, and abstract thought. Our results, with the revision of the VISQ, provide new avenues for probing these relationships while continuing to explore the important connections between the phenomenology of self-talk and psychopathology. By expanding the horizons of inner speech, and picking up a greater range of experiences, a richer narrative will be available to turn Vygotsky’s (1987) ‘cloud of thought’ — to use his evocative phrase — into ‘a shower of words’."],["Despite the coherence and seeming directness of our bodily experience, our perception of the world, including that of our own body, may constitute an inference based on ambiguous sensory data and prior expectations. In this article, I apply a 'psychologised' version of the recently proposed free energy framework to the understanding of certain disorders of neurological unawareness in order to examine how inferential processes may determine our body perception. I specifically consider three facets of body perception in such disorders: namely, the 'external body' as inferred on the basis of exteroceptive signals and related predictions; the 'internal body' as inferred on the basis of proprioceptive and interoceptive signals and related predictions; and lastly the 'impersonalised body' as inferred on the basis of signals from social and third-person perspectives on the body and related predictions. Several conclusions will be drawn from these considerations: (a) there is a deep interdependency of prior beliefs and sensory data; as the brain uses sensory data to update its virtual model of the world, lack or imprecision of sensory prediction errors may lead to aberrant inferences influenced disproportionally by outdated, premorbid predictions; (b) interoception and interoceptive salience have a unique role in our inferences about body awareness and (c) social, 'objectified' prior beliefs about the body may have a silent but potent role in our bodily self-awareness. Finally, the article emphasizes that our learned, virtual model of the body is depended on the nature and thus integrity of the very body that allowed the model to be formed in the first place. --------------------------------------------------------------------------------","Remembering the past, and being able to project oneself in the future, allows the mind to escape the psychological ‘here and now’ of experience. Studies in psychology have long established that we do not only project our current self into the future to build a kind of ‘as if’, imagined future self but we also reconstruct our past self in our memories (Bartlett, 1932). Despite the incredible storage capacity of human memory, what we remember in the now is not always what took place in the past. Instead, the autobiographical incidents that we experience as veridical, coherent and self-defining are frequently unconscious collages of previous recollective attempts, fragments of experienced events, currents thoughts and long term goals (Conway, 2005). In this sense, we have come to understand our autobiographical self as actively, yet unconsciously inferred on the basis of imperfect memory data and current expectations. A similar idea for the nature of our experience of current reality, our embodied perception of the world and ourselves in it, has also being long proposed in psychology (e.g. Gregory, 1966). Despite the coherence and seeming directness of our experience, our perception of the world may constitute an inference based on ambiguous sensory data and prior expectations (von Helmholtz, 1878/1971). However, this idea is less established, perhaps given its counterintuitive nature and complex, philosophical implications. We experience the world via our body and the experience of the latter in the ‘here and now’ is considered as a fundamental aspect of our self-consciousness; our bodily self is the foundation upon which our ‘autobiographical’, ‘narrative’ or ‘extended’ self is built (Gallagher, 2000). If our bodily self is an inference, then our ability to perceive the world and ourselves veridically is called into challenge (see Clark, 2013 for discussion). Leaving aside the majority of the long and complicated philosophical discussions on the nature of reality and our capacity to perceive it, in this paper I will explore similar ideas from the point of view of a recent, influential theory from computational neuroscience. The theory aims to define the idea of perceptual inference using concepts from theoretical physics and mathematics and also aims to ground the same idea on biology and particularly knowledge about the workings of the brain. In the current paper, I will not address the issues of interest in mathematical ways. Instead, I will use a ‘psychologised’ version of the free energy framework in order to examine some ideas regarding neurological unawareness and ultimately bodily self-consciousness. Specifically, I will use clinical observations, behavioural and neural data from a specific neurological aberration of self-awareness, namely anosognosia for hemiplegia, to explore the possibility that our bodily self- awareness is normally imperfect, in the sense that it is based on a set of inferences about the hidden causes of sensorimotor signals. I also hope to demonstrate that the study of the pathologically exaggerated ways in which we may infer the experience of our own body, can provide insights into the mechanisms of normal perceptual and active inference, and particularly the predictive and social nature of motor awareness.","The starting point of the ‘free energy framework’ (Friston, 2005) is that humans are biological, self-organising agents that need to occupy a limited repertoire of sensory states for homeostatic reasons (e.g. humans need certain ranges in environmental temperature in order to survive). However, due to the inherent ambiguity and uncertainty of the signals an organism receives from the world, we risk finding ourselves in dangerous states for longer periods than those we could biologically sustain (e.g. in cold climates). We thus need to be able to predict (infer) the causes of our possible sensory states despite the limited or noisy information available to our sensory organs (von Helmholtz, 1878/1971). The framework proposes that our brain engages in a form of probabilistic representation of the causes (e.g. the weather) of our future states (e.g. our temperature) on the basis of noisy sensory data; in other terms, it maintains hypotheses (‘‘generative models’’) of the hidden causes of sensory input. Furthermore, it uses such input to constantly update its models, so as to reduce its representational errors over time and thus ultimately minimize the risk of surprise (unpredictability, see below for mathematical definition). From a psychological point of view, I will refer to the formation of such models as the ‘mentalisation’ of sensorimotor signals. Although the term mentalisation is traditionally used in psychology to refer to our cognitive ability to infer the mental states of others and our own, I think the two terms are related (see also Kilner, Friston, & Frith, 2007). In fact, the use of the term ‘mentalisation’ in this article is intending to ground this traditional concept in its embodied origin. Returning to the biological level, the free energy framework is biologically constrained by the so- called ‘predictive coding’ models of perception, stemming primarily from biological and behavioural studies in visual perception, with supporting evidence generated in various modalities (e.g. Henson & Gagnepain, 2010; McNally, Johansen, & Blair, 2011). These suggest that a constant filtering of sensations by top-down predictions and a parallel updating of the latter based on prediction errors (signals representing the mismatch between predictions and sensations), with the ultimate goal of minimizing prediction errors, is an imperfect but highly efficient means of perceiving sensations (Rao & Ballard, 1999). Our brain is assumed to achieve the minimisation of prediction errors by recurrent message passing among hierarchical level of cortical systems, so that various neural subsystems at different hierarchical levels minimize uncertainty about incoming information by generating a prediction (or a prior belief, see below) and responding to errors (mismatches) in the accuracy of the prediction, or prediction errors. Such prediction errors are passed forward to drive the units in the level above that encode conditional expectations which optimize top-down predictions to explain away (reduce, inhibit) prediction error in the level below until conditional expectations are optimized. Such message passing is considered neurobiologically plausible on the basis of functional asymmetries in cortical hierarchies; prediction errors are thought to be conveyed via feedforward connections from lower to higher levels in order to optimize representations in the latter. Predictions from higher-levels are transferred via feedback connections that have both driving and modulatory characteristics and can suppress prediction errors in lower levels. This hierarchy is thus reciprocal but asymmetric and models the nonlinear generation of sensory input (Adams, Shipp, & Friston, 2013). Based on such hierarchical, perceptual schemes, the free energy principle, rests upon the idea that the brain as a whole works as an Helmholtzian inference machine that is trying to optimize its own model of the world by actively predicting the causes of its sensory inputs (Friston, 2005). Moreover, this inferential process is mathematically understood in Bayesian terms (Bayes’ theorem describes an optimal procedure for updating the probabilities assigned to a hypothesis in the light of new evidence), in the sense that it relies on a combination of prior beliefs (probability distributions over some unknown cause excluding any sensory data) and new sensory data to update prior beliefs and generate posterior beliefs (probability distributions over some unknown cause after data have been received). Furthermore, in the free energy principle this hierarchical minimization of prediction errors is understood as a minimization of free-energy on the basis of the formal definition of the latter; a quantity from informational theory that bounds (is greater than) the evidence for a model of data (Hinton & von Camp, 1993). In this case the data is sensory and free energy bounds the negative log-evidence (surprise) inherent in sensory data, under a model of how the data were caused (See Friston, 2010 for the mathematical details). Given some mathematical assumptions, free energy can be thought of as the amount of prediction error in any given level of the system. Minimizing free energy then corresponds to explaining away prediction errors following the principles of Bayes (Friston, 2010). However, representing the world in constructive ways (perceptual inference) cannot take us far in terms of our ultimate goal; surviving in an uncertain world. Psychologically speaking, we may become better in predicting (‘mentalising’) the changes in the environment that act to produce sensory impressions on us, but we cannot on this basis change the sensations themselves and hence ultimately their surprise. A highly innovative conceptual move in the free energy principle framework allows us to understand how we do just that. By acting upon the world we can change its states and therefore ‘re- sample’ the world to ensure we satisfy our predictions about the sensory input we expect to receive. By selectively sampling the sensory inputs that we expect we add accuracy to our predictions about sensory states. Thus, action has an intimate relationship with perception, both being governed by the same master principle, namely reduction of free energy; action can reduce free energy by changing sensory input, while perception reduces free-energy by changing predictions. In sum, the framework is consistent with theories of embodied cognition and enactive perception (see Clark, 2013 for discussion) that stress the role of embodiment in shaping cognition and propose a close link between action and perception. The framework further makes strong claims about cognition consisting of predictions (or priors) that do not represent the world and our bodily state directly. Instead, in order to evade the inherent surprise of the world, our cognition serves a constantly-updated, Bayes-optimal, iterative, self-fulfilling prophecy. via a cascade of multilevel processing across the neurocognitive hierarchy, we progressively minimise our own representational errors in perception and maximize the posterior probability of generating the observed sensory states in action. In doing so, we generate a kind of ‘virtual version’ of the sources of our bodily signals. The state of the body and its world is, in this sense, never directly available to perception and always inferred. At this point however, an important clarification needs to be made. The framework does not imply that the mind is (dualistically) divorced from its environment, including its body. The generative models in question are not viewed as mere functions somehow ‘housed’ in the brain, and ‘informed’ about the body by sensory states. Instead, the ‘mentalisation’ of the body implies physical, changes in the structure and function of the body itself, from the periphery to the brain. Indeed, the framework suggests that the structure and physiology of the brain itself are shaped by sensory states in as much as they shape them (in both ontogenetic and phylogenetic development) (e.g. see Adams et al., 2013). In other terms, the inferential, predictive models of possible causes of sensory input are understood as embodied (e.g. reflecting changes in synaptic connectivity) at different levels of the neurobiological hierarchy. It follows that in its totality, the self- organised, agentive virtual model in question is not merely ‘corrected’ by our embodiment (i.e. affected by sensory prediction errors), but rather it is our embodiment. In this sense, perception of the body and the world is both truly indirect (virtual, predictive) in the ‘here and now’ and exact in the long-term: it can ultimately only represent itself.","Anosognosia for hemiplegia (AHP) is defined as the apparent unawareness of one’s paralysis (Babinski, 1914), which occurs typically following stroke-induced right perisylvian lesions, and less often following left perisylvian lesions (Cocchini, Beschin, Cameron, Fotopoulou, & Della Sala, 2009). This prototypical, neurological disorder of body unawareness affects our awareness of action; a composite notion that includes at least two facets, the subjective feeling of moving in the here-and-now, and more general beliefs or judgments about one’s motor abilities, such as being able to perform certain bilateral actions (see also below for the related distinction of on-line versus off-line awareness, as well as the distinction between illusory versus delusional awareness). In a subset of patients with concomitant body delusions (somatoparaphrenias, Gerstmann, 1942), the right- hemisphere damage can also affects the sense of body ownership (the subjective feeling that our body is separate from the world and other bodies). Such patients may reject the ownership of one’s limb (asomatognosia), misattribute it to others, or vice versa (somatoparaphrenia proper), claim they have three or more limbs (supernumerary limbs), or treat the limb as though it was a separate person (personification; Critchley, 1955). The typical duration of AHP ranges from days to weeks (Vocat, Staub, Stroppini, & Vuilleumier, 2010), but in about one third of patients the symptoms may last beyond the acute stage of illness and even years (see Pia, Neppi-Modona, Ricci, & Berti, 2004). AHP can be highly specific in that some patients deny their plegia, while being simultaneously aware of other neurological, or neuropsychological disturbances (Bisiach, Vallar, Perani, Papagno, & Berti, 1986; Berti, Làdavas, & Della Corte, 1996; Marcel, Tegner, & Nimmo-Smith, 2004). In terms of AHP’s ‘extension’ (what kinds, or objects of awareness can be compromised, Marcel et al., 2004), some patients acknowledge their motor deficits but fail to adjust to their functional consequences, while others show the opposite pattern (Marcel et al., 2004; Moro, Pernigo, Zapparoli, Cordioli, & Aglioti, 2011). Furthermore, some patients claim their limbs have moved even upon demonstration of the opposite (illusory movements, Feinberg, Roane, & Ali, 2000; Fotopoulou, Tsakiris, Haggard, Rudd, & Kopelman, 2008), while others admit their on-line failure, but fail to update their long-term or, ‘off- line’ body awareness (Carruthers, 2008; Marcel et al., 2004; Moro et al., 2011; Tsakiris & Fotopoulou, 2008). A related, and at times hard to separate, characteristic of AHP is its ‘partiality’ (whether unawareness of one’s deficit is less than total, Marcel et al., 2004). This property is noted in studies that demonstrate implicit awareness of deficits despite explicit unawareness in verbal (Fotopoulou, Pernigo, Maeda, Rudd, & Kopelman, 2010), or behavioural tasks (Cocchini, Beschin, Fotopoulou, & Della Sala, 2010; Moro et al., 2011; Nardone, Ward, Fotopoulou, & Turnbull, 2007), as well as in studies that show higher awareness of plegia in third-person versus first person tasks. For example, patients have been observed to deny their deficits in direct view but admit them in a video replay (Besharati, Kopelman, Avesani, Moro, & Fotopoulou, 2014; Fotopoulou, Rudd, Holmes, & Kopelman, 2009) or in 3rd-person questions (Fotopoulou et al., 2011; Marcel et al., 2004). Similarly, patients who deny the ownership of their arms (see above) in direct view have been shown to admit them in front of a mirror (Fotopoulou et al., 2011; Jenkinson, Haggard, Ferreira, & Fotopoulou, 2013), and even show improved somatosensation when tested from a 3rd-person perspective (Bottini, Bisiach, Sterzi, & Vallar, 2002), or when they use their ipsilesional hand to actively touch their affected, contralesional arm (Van Stralen, van Zandvoort, & Dijkerman, 2011). More recently, we also showed that anosognosia can be momentarily reduced following affective, social feedback (Besharati et al., 2014). Despite recent rapid progress in the assessment and understanding of AHP (see Fotopoulou, 2014; Jenkinson, Preston, & Ellis, 2011; Orfei, Caltagirone, & Spalletta, 2009 for reviews), little consensus exists regarding its functional and neuroanatomical explanation. Older theories emphasise deficits in afferent (feedback) and bottom-up signals, while more recent hypotheses focus on modular abnormalities in predictive (feedforward) signals and their role in motor awareness (Berti et al., 2005; Frith, Blakemore, & Wolpert, 2000; Heilman, Barret, & Adair, 1998; see also Jenkinson & Fotopoulou, 2010 for review). For example, on the basis of a computational model of motor control (Wolpert, 1997), Frith et al. (2000) have proposed that although patients with AHP are able to predict the expected sensory consequences of intended movements, they fail to register the discrepancy between predicted and actual sensory feedback because of visuospatial neglect or other sensory deficits. Berti and colleagues (see Berti et al., 2007 for review) suggested that this failure may instead relate directly to damage to the lateral premotor cortex (Berti et al., 2005). This group as well as other groups have further produced physiological (Berti et al., 2007; Hildebrandt & Zieger, 1995; but see Gold, Adair, Jacobs, & Heilman, 1994) and behavioural (Garbarini et al., 2012; Jenkinson, Edelstyn, & Ellis, 2009) evidence showing that there are intact motor intentions in AHP. A further study showed for the first time the direct relation between motor intention and awareness (Fotopoulou et al., 2008). The authors were able to show that patients’ illusory awareness of movement reflected an abnormal, selective dominance of motor intentions over visual feedback about the actual effects of movement (elicited by a realistic rubber-hand patients assumed was their own), and this effect could not be explained by neglect. Lastly, while as mentioned above, taking a third-person perspective on the self, verbally (e.g. Marcel et al., 2004) or visually by video-replays (Besharati, Kopelman, et al., 2014; Fotopoulou et al., 2009) may improve awareness, there may also be an alternative, motor explanation of the video-replay results. During video viewing, patients receive visual feedback of their paralysis at a time when they are not intending to move and hence forward signals are rendered irrelevant to motor awareness. Patients’ awareness during video-replay needs to rely exclusively on the visual or auditory feedback they receive via the video clip. Despite the clear value of the ‘feedforward’ hypotheses, it has become apparent to many authors that a strictly modular, motor explanation is not sufficient to account for all the manifestations of AHP. For example, such theories cannot explain why mood induction can temporarily improve AHP (Besharati, Forkel, et al., 2014), nor account for the extension and partiality of AHP (see above). Indeed, recent experimental studies have shown that awareness dissociations between and within patients are linked with different lesion patterns, including limbic areas non-associated with motor functions (Fotopoulou et al., 2010; Moro et al., 2011). Similarly, in a voxel-based lesion-symptom mapping study, Vocat et al. (2010) demonstrated that the neuropsychological and neural profile of AHP patients’ changes in time, and different lesion patterns are associated with AHP at different time points. These studies point to a multi-component disorder occurring due to lesions affecting a distributed set of brain regions, including the insula, premotor and parietal regions but also subcortical areas such as the thalamus, basal ganglia and limbic structures. More broadly, a number of authors have noted that AHP sometimes has delusional features that cannot be explained solely on the basis of sensorimotor deficits (for discussion, see Fotopoulou, 2010; Frith et al., 2000; Ramachandran, 1995; Turnbull & Solms, 2007; Vuilleumier, 2004; Turnbull, Fotopoulou, & Solms, 2014). Feedforward theories are valuable in explaining the illusion of moving (Fotopoulou et al., 2008), but AHP patients do not simply claim that they have the phenomenal experience of moving. In fact, typically patients with AHP do not spontaneously complain of any related, subjectively perceived symptom, whether negative (e.g. I am not moving) or positive (I have the impression that I am moving). On the contrary, AHP is diagnosed on the basis of questioning during which patients are typically asked to report on their current experiences (confrontation questions) and infer their more general motor abilities (see also Marcel et al., 2004). Even patients who report illusory experiences of movement during confrontation and hence presumably base their inferences on such impressions (Fotopoulou et al., 2008), they nevertheless simultaneously ignore the wealth of contrary evidence and medical signs indicating that they are paralysed (e.g. their medical results, disabilities, occasional accidents and others’ feedback). This perceptual ‘selectivity’ is not the same as the one observed in other symptoms such as hemispatial or personal neglect, in the sense that patients with neglect can become aware of the fact that they have neglect after their errors are demonstrated to them. They then continue to do such errors but they are not surprised or in denial when these errors are pointed out to them again. In fact, the subset of patients who cannot become aware of their neglect would be diagnosed as anosognosic for these deficits. Moreover, as aforementioned, there is now also experimental evidence that patients with AHP maintain their denial even after they themselves had admitted their paralysis momentarily (e.g. Besharati, Forkel, et al., 2014). It can thus be said that they adhere to the delusional belief that they have functional limbs. If one accepts that anosognosia has delusional features then theoretical loans from the literature on delusions can be allowed (see Fletcher & Fotopoulou, 2014; Fotopoulou, 2010 for discussions). Of particular interest here is the ongoing debate between one and two-factor theories. According to the former, rational reasoning on the basis of anomalous or unusual experience should be sufficient to ultimately lead to refractory, delusion beliefs (e.g. Maher, 1992). By contrast, two-factor theories claim that delusional beliefs cannot be explained without the role of additional, cognitive dysfunctions such as reasoning biases, or monitoring deficits that are necessary for the generation and maintenance of the false beliefs (e.g. Davies, Coltheart, Langdon, & Breen, 2001; Garety & Freeman, 1999). Indeed, also in the literature on anosognosia, a third set of recent theories emphasise that the explanation of anosognosic beliefs and attitudes requires the postulation of additional dysfunctions that prevents sensorimotor and other failures to be re-represented at a higher level of cognitive and emotional self- representation, beyond the sensorimotor domain. These accounts stress the necessary combination of bottom-up and top-down deficits and corresponding lesioned brain regions (Davies, Davies, & Coltheart 2005; Levine, 1990; Levine, Calvanio, & Rinn, 1991; Vuilleumier, 2004). For example, considering anosognosia in the more general context of delusional beliefs, Davies et al. (2005) proposed that anosognosic beliefs maybe explained by a two-factor account used to explain other delusions; abnormal beliefs arise due to a first impairment in perception that prompts the abnormal belief and a second impairment that interferes with higher-order, monitoring processes thus allowing the abnormal perceptions to become abnormal beliefs. These accounts have undoubtedly being useful in emphasizing the multifaceted nature of AHP, and for attempting to link the understanding of anosognosic phenomena with insights about the cognitive processes that may underlie normal and pathological belief formation (see also Fotopoulou, 2010, 2012). However, these accounts have been criticized for not being falsifiable (Vallar & Ronchi, 2006). Moreover, reflecting the modular epistemology of cognitive neuropsychology (Fotopoulou, 2014 for a critical review), these models treat the described deficits as simply ‘additive’ and as potentially caused by simultaneous damage to functionally independent lesion sites. For example, Vocat et al. (2010) suggested that a combination of lesions to two or more brain areas within the insular, premotor, parietal and temporal cortex, or the white matter connections that link one or more of these areas with subcortical regions, may lead to different combinations of deficits in functions such as proprioception, spatial neglect, and error monitoring, which in turn lead to anosognosia in different patients. While such ‘combinations’ of lesion sites and deficits are consistent with the multifaceted nature of AHP, what these accounts lack is a more precise neurobiological and neuropsychological description of the dynamic and hierarchical relation between the affected areas and their integrated functional role in body awareness. At this point, we turn to the aforementioned free energy framework in order to propose an alternative model of AHP that aims to address precisely this limitation, as well as to describe the unique, virtual nature of perception and the social nature of the bodily self. Finally, this model effectively unifies previous one- and two-factor models of anosognosia, as it does not allow for a distinction between perception (experience) and cognition (inference) at any level. Instead, as explained above, all perception (including all conscious experiences) is always an inference. Thus, according to the model, the difference between the various manifestations and possible subtypes of anosognosia cannot be captured on the basis of this distinction between perception and cognition. Instead, one explanatory factor is indeed sufficient to explain all manifestations of anosognosia, but this factor is not anomalous experience. Rather it is the aforementioned, always embodied and always cognitive form of inference that may become aberrant in anosognosia, as in other delusions (Corlett, Taylor, Wang, Fletcher, & Krystal, 2009). We consider the particular kind of aberration that may lead to AHP in the following section.","On the basis of the free-energy principle, this paper puts forward the idea that AHP can be best explained as aberrant perceptual inference at various levels of the neurocognitive hierarchy. It is specifically proposed that the observed lesions result in weak, absent or unreliable prediction errors about sensorimotor states of the affected body parts, which ultimately lead patients to base their inference on premorbid, non-updated predictions about their motor abilities and their agency in the world. This faulty relationship between premorbid, habitual predictions and imperfect processing of current prediction errors is thought to take place at different levels of the neurocognitive hierarchy, consistently with the varied phenomenology and critical lesion sites of patients with AHP (e.g. Fotopoulou et al., 2010; Vocat et al., 2010). We consider some of the critical types of such Bayes-optimal, yet aberrant, inference in further detail below. Anosognosia for hemiplegia and the ‘External Body’ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ AHP typically occurs in the context of a number of concomitant sensory impairments, including primary exteroceptive (signals relating to the state of world and the external surface of the body) deficits as well as related, higher order impairments such as visuospatial or, personal neglect (see Vallar & Ronchi, 2006; Small & Ellis, 1996 for reviews). During the 1980s and 1990s, studies under the epistemological remit of cognitive neuropsychology, attempted to establish whether any of these deficits or any given combination of deficits could explain the occurrence of one or more of the above anosognosic phenomena (see Fotopoulou, 2014 for review). These studies revealed double dissociations between AHP and such impairments, suggesting that none was necessary for AHP to occur (e.g. Bisiach et al., 1986; Marcel et al., 2004). Nevertheless, under the remit of the free energy principle, these dissociations do not exclude the possibility that some of these deficits may act as predisposing, or contributing factors. Exteroceptive signals about the left side of the body, as represented in the connections of right hemisphere subcortical areas (e.g. the thalamus), or re-represented and organised in cortical functional networks of the right-hemisphere (Berti et al., 2005; Fotopoulou et al., 2010; Moro et al., 2011; Vocat et al., 2010) may be weak, or even absent due to damage to one or more of these areas. Such damage may therefore allow predictive signals at these levels to continue to operate ‘unchecked’ by appropriate prediction errors about the current, exteroceptive state of the body. In other terms, there will not be sufficient or sufficiently precise (see also below) signals to update one’s normally, predictive motor awareness. It is further expected that the more and the greater the deficits in one or more of these domains, the greater the likelihood of faulty (anosognosic) inferences. Indeed, a recent paper revealed a positive correlation between the degree of AHP and the combined quantity of such deficits (Vocat et al., 2010). However, as aforementioned, it is unlikely that these deficits are sufficient to explain the richness of anosognosic phenomena and are best equipped to explain the illusion of moving, rather than the more general delusional beliefs and attitudes s that patients with AHP show in several cognitive and emotional domains. If such deficits were the sole cause of AHP, it would be unclear why other, unaffected bottom-up information about the body (e.g. internal, homeostatic signals, see below) could not provide the necessary prediction errors to update one’s beliefs about the current state of the body. It would also be unclear why other top-down predictions about the body (e.g. the prediction that one will not fall after moving one’s left leg) were not used to update beliefs about the body. Anosognosia for hemiplegia and the ‘Internal Body’ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A further set of important signals about the body arises from within the body’s skin boundary. These include proprioceptive and more generally kinesthetic sensations about the position and dynamic properties of the body in space arising from the vestibular system and from muscles and tendons, as well as interoceptive sensations about the physiological conditions of all internal organs (Craig, 2009). Deficits in prioprioception are not sufficient to cause AHP but they have been shown to be among the most common deficits in AHP patients (Vallar & Ronchi, 2006; Vocat et al., 2010). The vestibular system is also thought to be affected in AHP, given the fact that vestibular stimulation has been shown to temporarily reinstate awareness (Cappa, Sterzi, Vallar, & Bisiach, 1987; Ramachandran, 1995; Ronchi et al., 2013). Such deficits are potentially important when trying to understand motor awareness, as according to the free energy framework, predictive signals during action do not constitute forward signals on the basis of efference copies of motor commands but rather sensorimotor predictions (e.g. proprioceptive predictions) about the effects of movement (for detailed discussions, see Adams et al., 2013; Friston, 2010). This perspective thus implies that some patients may have reduced ability to generate novel predictions about their bodily and spatial effects of their potential left-sided movements (see also Heilman et al., 1998 for a similar proposal based on a previous computational model of motor control and awareness). Although this deficit may contribute to unawareness in some patients, previous studies have shown that at least some patients with AHP have intact ability to generate sensorimotor predictions (see above). Moreover, if patients with AHP are unable to generate such predictions, the fact that some patients insist that they have moved as desired, i.e. fulfilled kinesthetic predictions, is not easy to explain. Importantly, even in patients whose proprioceptive and vestibular systems are intact, there is a degree of paresis due to damage to the motor system. In fact prototypical cases of motor unawareness are considered the ones who show complete paralysis of their left limbs. Thus, an important source of disruption may be the mere fact that patient can no longer fulfill their proprioceptive and other related priors by active sampling of prediction errors (i.e. transmitting such descending somatomotor predictions to the peripheral motor system and moving their affected limbs so as to generate reafference). AHP patients should be able to generate such somatomotor predictions in spared premotor and parietal cortex areas (Berti et al., 2005; Karnath, Baier, & Nägele, 2005). However, damage further down the hierarchy would mean that such predictions are not fulfilled by the motor system. Nevertheless, unlike the aforementioned effects of missing or weak prediction errors that are passed up the cortical hierarchy in order to modify perceptual inference, such disruptions in active inference at the spinal cord or subcortical level should not have an effect of motor awareness. The normal role of such somatomotor (mainly proprioceptive) reafference seems to be to modify descending predictions at spinothalamic and spino-cerebellothalamic circuits, thus allowing, fast, ‘automatic’ correction and control of movement. Indeed, this lack of active inference does not seem sufficient to cause AHP as the syndrome occurs in a minority of patients with stroke-induced hemiplegia. It thus seems that while exteroceptive, proprioceptive and motor deficits may be important contributors to the symptomatology of some patients with AHP, they are unlikely to be its primary or central causes. By contrast, another facet of the internal body may have a more central role in AHP. Recent lesion mapping studies have highlighted that areas such as the insula, limbic structures and subcortical white matter connections may be selectively associated with AHP (Fotopoulou et al., 2010; Karnath et al., 2005; Moro et al., 2011; Vocat et al., 2010). Such areas and their connections are linked with interoception and motivation and are specifically implicated in bodily salience and interoceptive awareness (Craig, 2009; Critchley, Wiens, Rotshtein, Öhman, & Dolan, 2004). Thus, we propose that weak or imprecise (see also below) interoceptive and emotional signals about the current (physiological) state of the body, may lead to difficulties in affectively personalising new sensorimotor information about the affected body parts. This would be coupled with a persistent, necessary adherence to past expectations of how the affected body parts should feel, ultimately leading to the aberrant beliefs about any available contrary information about the body. In support of this hypothesis, a recent study (Romano, Gandola, Bottini, & Maravita, 2014) shown that right hemisphere patients who show somatoparaphrenic beliefs about their affected body parts, also show reduced physiological reactions to the threat of the same body parts, as measured by skin conductance responses. Moreover, given the higher position of such priors in the neurocognitive hierarchy (see Friston, 2013), such faulty inference may also ‘explain away’ contrary exteroceptive signals during instances of multisensory integration. To use the words of one anosognosic patient who also denied the ownership of his paralysed limbs “But my eyes and my feelings don’t agree, and I must believe my feelings. I know they [left arm and leg] look like mine, but I can feel they are not, and I can’t believe my eyes.” (C.W. Olsen, 1937, cited in Feinberg, 1997). The degree to which such interoceptive deficits are linked to delusions of ownership more frequently than delusions of motor awareness remains to be specified in future studies. Furthermore, a related deficit in the processing of salience from the affected body parts needs to be emphasized. As aforementioned, activity in areas such as the insular and the limbic cortex is not only linked with interoception and emotion but more generally with interoceptive salience and motivation. In the free energy framework, these notions are linked to the concepts of ‘precision’ (mathematically inverse dispersion or variance, and hence the inverse of uncertainty) and its neurochemical equivalent, neuromodulation (Friston et al., 2012). Specifically, precision is linked mainly with the neuromodulation of synaptic gain that encodes the uncertainty of random fluctuations about predicted states. It follows that neuromodulators of synaptic gain (such as dopamine and acetylcholine), do not signal (reward or pleasure) prediction errors about sensory data but the context in which such data were encountered. In other words, such neuromodulators report the salience of sensorimotor representations encoded by the activity of the synapses they modulate. This is important, especially in hierarchical schemes, where precision controls the relative influence of bottom-up prediction errors and top-down predictions. In psychological terms, the processing of salience expectancy allows the organism to control the significance it attributes to the sensory data it uses to update its predictions or to explain away prediction errors. The above considerations have added potency in AHP given the above critical lesion cites, as well the recently identified lesions in fronto-striatal circuits (Fotopoulou et al., 2010; Moro et al., 2011; Venneri & Shanks, 2004; Vocat et al., 2010). Such lesions may lead to a more general difficulty in optimizing the precision of prediction errors (Friston et al., 2012), affecting their salience and ultimately both short- and long-term learning (suboptimal synaptic gain and plasticity, Friston, 2010). Indeed, the functional role of the basal ganglia and particularly the striatum has been linked with prediction error-driven learning (O’Doherty et al., 2003) as well as the aberrant salience theories of psychosis (Gray, Feldon, Rawlins, Hemsley, & Smith, 1991; Kapur, 2003). In AHP such deficits can be linked with both specific instances of aberrant motor monitoring in functionally specialised systems (Berti et al., 2005), or more generally in global error monitoring (Davies et al., 2005; Venneri & Shanks, 2004; Vocat, Saj, & Vuilleumier, 2012), mental flexibility (Levine et al., 1991) and ‘surprise detection’ (Ramachandran, 1995) deficits. Indeed, a recent study showed that AHP patients had the tendency to ‘jump to conclusions’ on the basis of limited and rather vague information and then to subsequently get stuck to their former “false” beliefs instead of modifying them based on novel, arguably more salient information (Vocat et al., 2012). Another study (Besharati, Forkel, et al., 2014) further found that anosognosia can be temporarily reduced by the induction of negative mood, presumably because negative emotions prime the organism for defensive action and increase the salience of sensorimotor signals (see Pereira et al., 2010; Gentsch, & Synofzik, 2014). These ‘neuromodulatory’ deficits may explain some of the delusional features of AHP that are harder to explain on the basis of deficits in sensorimotor signals per se, be those exteroceptive or interoceptive. For example, they provide some insight into how patients can remain in denial of their paralysis and/or apathetic towards the normally alarming sight of a paralysed left arm and its related consequences. More generally, this explanation of anosognosia as an inability to update body awareness in a way that takes into account and ‘personalises’ new motor (agentive), interoceptive (emotional) or, just salient information about the affected body parts has the advantage of unifying the hypotheses put forward previously by modular (e.g. Berti et al., 2005; Karnath et al., 2005) and multi- factorial theories (e.g. Davies et al., 2005; Vuilleumier, 2004) on a single, unified and neurobiologically-plausible formulation. Moreover, this formulation integrates both bottom-up and top-down mechanisms of bodily perception, action and belief formation (see also Fotopoulou, 2012; Fotopoulou, 2014). Furthermore taken together, the above considerations on AHP in the light of the free energy framework highlight, not only how our awareness of our own body is based on habitual predictions of its state in the world, but also how this learned, virtual model of the body is depended on the integrity of the very body that allowed the model to be formed in the first place (see introduction). I now turn to a final aspect of body awareness formation that seems less intuitive than the rest and can hopefully be made more evident via the study of AHP. Anosognosia for hemiplegia and the ‘Impersonalised Body’ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As explained above, the proposal regarding the inability to shape unconscious inferences about the body by prediction errors signaling information about the affected body parts as ‘personal’ and relevant to the self can explain a lot of anosognosic features and beliefs. However, a fundamental question remains. Patients seem unable to use other higher-order knowledge, including social feedback, to update their beliefs. This failure is not easy to explain on the basis of any sensorimotor deficits. Of course, their feelings that the paralysed body parts are of limited self-relevance partly explains why they would disregard social feedback for some time. As the aforementioned patient claimed, one has learned to trust one’s feelings about the body over and above other sources of information. This is also consistent with the assumed hierarchy of the proposed model; brain areas assumed to subserve interoception and emotion are thought of as higher in the neurocognitive hierarchy than areas subserving exteroception (Friston, 2013). However, why are aberrant interoceptive inferences not updated by priors at even higher levels? For example, the model includes the possibility of other, already formed generative models about the self with predictions represented at higher levels (e.g. ‘my family and doctors would not lie to me about serious health issues’) that could presumably be used to send predictions down the neurocognitive hierarchy and eventually influence the inferences based on faulty interoception and salience. Given that patients’ beliefs about their motor abilities seem uninfluenced by such knowledge we infer that these processes are not taking place. This assumption is consistent with the aforementioned clinical observation that patients with AHP frequently refer to other people’s opinions on their bodily state, without altering their self-awareness. For example, they say phrases like ‘I know the doctors think I am paralysed, but I do not believe it’, ‘everyone says I cannot move, but I know I can’. If such third-person knowledge about the self is available why is it not affecting their own opinion about themselves? The answer could be that this particular aspect of perception-cognition, particularly as applied to the affected body parts, is also damaged by the right hemisphere damage in question. In this section, I will try to describe the nature of this particular impairment as manifested in AHP and the nature of the presumed corresponding function in the undamaged brain. Borrowing insights from Merleau-Ponty (1945/1962) I will call this aspect of body representation, the ‘impersonalised’ body. Related concepts include the ‘habit’ body (Merleau-Ponty (1945/1962)), the off-line (vs. on-line) body representation (Carruthers, 2008; see also Tsakiris & Fotopoulou, 2008) and of course the various versions of the old and highly problematic distinction between body schema and body image. The proposed aspect of body representation is not proposed as fitting any of these terms exactly, nor being part of a simple contrast between the personal and the impersonal, or social body. As should be obvious in light of what I have written above, no simple dichotomy of body representations would be sufficient to account for the multiple ways in which the body is represented in the mind and our existence is embodied in brain functioning. However, a proper discussion of all these concepts and terms escape the scope of this article. Instead, I use the term ‘impersonalised’ body in order to emphasise with a single, heuristic term the ‘objectified’ aspect of this domain of bodily perception, as well as its personal and interpersonal origins. Perceiving the world, including my own body, also entails the perception of the body’s action possibilities in the same world. Indeed, my perception of the world is rarely confined to the characteristics of the input that reaches my few, sensory organs from my current, unique position and perspective in the world. Instead, my prior experiences of different sensations, possibilities and positions in the world define how I perceive the world in any given time and space. For example, my multisensory expectations and reaching movements adjust to the characteristics and affordances of a glass on a table by the mere sight of it from one particular egocentric perspective and prior to any current movements, or tactile feedback. Such expectations constitute the world as real in relation to my body, even if they remain implicit most of the time. These ideas, present in phenomenology, have been more recently explored in cognitive neuroscience and are accepted by scholars that subscribe to the ideas of ‘multisensory integration’ and ‘enacted perception’. Moreover, such subpersonal processes are also compatible with the reading of AHP in the light of the free energy framework; the perception of the world entails (unconscious) inferences based on the agents prior embodied experiences with the same world. As Merleau-Ponty (1945/1962) first speculated, the affordances and multisensory characteristics of the world continue to appeal to the habitual body of the anosognosic patient, in the same hidden way that they appealed to it prior to the paralysis. Moreover, as I have outlined, the framework explicitly proposes that perception and action as operating in a continuum, being essentially governed by the same operating principles. There is however an aspect of such perceptual and active inference that seems to have received less attention within the framework. The perception of the world, as a coherent canvas of multiple interrelated sensations, action possibilities and perceptual perspectives, entails the implicit possibility that any given object can simultaneously be perceived by different agents. When I look at the front of a chair, my awareness of the chair as one thing entails my tacit perception of its back. In this sense, I may be aware that the chair can be perceived by different positions, by anyone, or by any-body (Taipale, 2014). In more explicit ways, my perception of the world is developmentally shaped by learning opportunities afforded by solitary prior instances of active and perceptual inference, as well as by the active presence of other agents (see Krahe, Springer, Weinman, & Fotopoulou, 2013 for related ideas in the domain of pain perception). A glass on a table is the object that others and myself can manipulate at any given time, from any of our unique perspectives. This rich experience of others interacting with the same world as ourselves is presumably at the basis of our everyday, intuitive sense of ‘veridical’ and ‘shared’ perception. We perceive the world as containing unique, whole-in-themselves objects despite the fact that our actual perception of them from a subjective point of view will always be limited to the constraints of our body in time and space, e.g. the position of our eyes on the head. This human, adult ability to simultaneously perceive the world, from many silent, potential perspectives and the action possibilities they entail, seems to suggest not only that our perception relies on inference but also that such inferences about the world is deeply embedded in the social world. Indeed, other than classic phenomenological views on intersubjectivity (see Gallagher, 2008; Taipale, 2014; Zahavi, 2001 for recent considerations), such ideas have been mostly developed in certain stands of developmental psychology (e.g. Reddy, 2008; Rochat, 2009). This perspective however entails a kind of tension when applied to the body. Signals in the above mentioned domains of the perceived body, including the exteroceptive and interoceptive domains, are ultimately integrated in egocentric coordinates. The 1st-person perspective (spatial and mental) remains fundamental for the perception of the body as mine, as under my volitional control and as separate from other objects in the world (e.g. Damasio, 1994). In other terms, although the body can be perceived (objectified) as a kind of socially perceived, impersonal, unique object, it is also always the subject of all experiences. As aforementioned, this tension has received many names in the history of mind and brain fields, and this paper cannot even begin to address the complexities behind such questions. However, attempting to understand this tension in the context of AHP and the free energy framework, as outlined in this article, may generate some insights, particularly as our everyday conscious perception of the body does not include two, separated experiences of the body. We mostly conceive of the body we see in the mirror as the one who feels itself to be standing in front of the mirror. It turns out that some patients with AHP do not share this experience. When looking at their paralysed body parts directly, they believe they are able body parts, and if they are somatoparaphrenic, they may believe they belong to someone else. However, when confronted with mirror images or video replays of their own body in the third-person perspective, the same patients describe the reflected body parts as being paralysed and as belonging to themselves, respectively (Fotopoulou et al., 2009, 2011; Jenkinson et al., 2013; Besharati, Kopelman, et al., 2014). These findings confirm the distinction between first and third-person perspectives on body perception and they further highlight the primacy of the first person perspective in conscious perception: it is the (delusional) content of the first person perspective that dominates their conscious awareness when mirrors are not made available to them. We can thus confirm Merleau-Ponty’s (1945/1962, p. 82) intuitions: In the case under consideration, the ambiguity of knowledge amounts to this: our body comprise as it were two layers: that of the habit body and that of the body in this moment. In the first appear manipulatory objects that have disappeared from the other and the problem how I can have the sensation of still possessing a limb I no longer have amounts to finding out how the habitual body can act as guarantee for the body at this moment. How can I perceive objects as manipulatable when I can no longer manipulate them? The manipulatable must have ceased to be what I am now manipulating, and become what one can manipulate; it must have ceased to be a thing manipulatable for me and become a thing manipulate in itself. Correspondingly, my body must be apprehended not only in an experience which is instantaneous, peculiar to itself and complete in itself, but also in some general aspect and in the light of an impersonal being. These observations also reveal another intriguing aspect of this distinction. Although these patients seem able to perceive their body ‘correctly’ from a third-person perspective, they do not seem surprised by the difference between the content of the two perceptual instances. Even though they may not have ‘seen’ their own arm for a month or so, they do not scream, ‘oh there it my arm’ when we place a mirror in front of them. Nor do they blame us for suddenly giving them paralysis when we take the mirror away. And they are not even surprised by the fact that they themselves have given a different answer about the ownership and agency of the same body part, just seconds ago. It thus seems that apart from their deficits in updating their body representation in the first-person, they have also lost the ability to perceive their body as an object in the world that needs to have a unique, socially-shared existence. The fact that they can recognize third-person, images of the body as theirs does not seem sufficient for the proper ‘objectification’ of the body (i.e. the perception of the body as a thing in itself). The latter, ‘impersonalised’ sense of the body therefore seems to involve cognitive and perhaps emotional operations that extend the mere self-recognition in third-person perspectives and rather bizarrely seem to also depend on a kind of grounding in the subjective body, or at least a kind of flexible, abstract perception and integration of 1st and 3rd person perspectives on the body. The cognitive and emotional integration between first and third person perspectives on the bodily self is thought to take place progressively in development but developmental psychologists, as well as phenomenological and psychoanalytic thinkers, seem to stress different aspects in these processes of self-objectification and awareness. In fact, the majority of studies and theories focus on how we come to understand or, infer other minds via their bodies, or how we come to regulate our own emotions. Far less attention is paid to the mentalisation (see definition above) of one’s own body via the influence of other people. While this discussion extends the scope of the current paper, I note certain possibilities as regards AHP here. The loss of the ‘impersonalised body’ in AHP can be explained in at least two ways: (a) body awareness from a first person perspective, including the processing of both interoceptive and exteroceptive signals is higher in the neurocognitive hierarchy (Friston, 2013), thus prediction errors relating to the ‘objectified’ body are simply explained away by predictions about the subjectively felt body and its related salience (see also section on the ‘Internal Body’); (b) the very faculty that allows individuals to engage in the act of flexible perspective-taking and integrate first and third person perspectives is impaired. The latter interpretation would be consistent with the observed damage in AHP patients in brain areas such as the temporoparietal junctions and the superior temporal sulcus (e.g. Besharati, Forkel, et al., 2014; Fotopoulou et al., 2010; Moro et al., 2011) and we have preliminary data showing that such lesions in patients with AHP are selectively associated with deficits in perspective taking and theory of mind abilities (Besharati et al., in preparation).","In this paper, I described the counterintuitive syndrome of anosognosia for hemiplegia, the striking, apparent unawareness of paralysis following right hemisphere stroke. I further put forward a large-scale framework from computational neuroscience, namely the free energy framework in order to account for the clinical variability of AHP and unite previous, seemingly divergent hypotheses about its pathogenesis. The framework proposes a view of human perception that relies on inferring the self and the world in both perception and action on the basis of prior expectations and ambiguous sensory signals. Contrary to intuition, our perception of the world is rarely confined to the characteristics of the input that reaches our sensory organs from our current, unique position and perspective in the world. Instead, our prior experiences of different possibilities and positions in the world define how we perceive the world in any given moment and position in space. In this sense, cognition can be thought of as imperfect, yet highly efficient (in a Bayes-optimal sense) strategy for self-organization in an ambiguous world. Anosognosia for hemiplegia, as a prototypical disorder of body unawareness, represents an exaggeration of such ‘imperfect’ body awareness system. I hope that the description of this neuropathology in the light of the free energy framework served to highlight the deep interdependency of prior beliefs and sensory data; as the brain uses sensory data to update its virtual model of the world, lack or imprecision of sensory prediction errors may lead to aberrant inferences influenced disproportionally by outdated, premorbid predictions. Finally, I hope that this consideration of anosognosia stresses that our learned, virtual model of the body is depended on the nature and thus integrity of the very body that allowed the model to be formed in the first instance."],["This study longitudinally investigated the relation between theory of mind (ToM) and verbal language skills in 231 children from preschool to early adolescence. Further, links to reading comprehension of texts at age 13;7 (years;months) were examined. To assess ToM, children completed false belief tasks at 5;6 and the Strange Stories at 12;8. To assess language, children completed a receptive grammar/sentence comprehension test at 3;6 and 5;6, a receptive vocabulary test at 3;6, 5;6 and 12;8, as well as a test of listening comprehension of texts at 13;7. A bidirectional relation between early and advanced measures of children's language skills and ToM was found: Changes in ToM were predicted by language skills, especially by receptive grammar/sentence comprehension; changes in children's receptive vocabulary were predicted by early ToM. However, early ToM had no direct or indirect effect on later listening comprehension or reading comprehension after controlling for early language skills. Only children's advanced ToM had a small indirect effect on reading comprehension, via listening comprehension. The results are discussed in light of ToM stability over time, and theories on how language and ToM development are intertwined. --------------------------------------------------------------------------------","Between 3 and 5 years of age, children gain increasingly greater knowledge of mental states and processes; that is, they have notable advancement in their theory of mind (ToM) development. This is reflected in their mastery of explicit false-belief tasks in which children are asked to predict how a protagonist will act or what a protagonist thinks based on a mistaken belief. This understanding of false beliefs has been shown to be closely connected to children’s language skills (Milligan, Astington, & Dack, 2007). Moreover, longitudinal data and training studies have revealed that in preschool age language skills are more predictive of ToM than vice versa (de Villiers, 2005; Ebert, 2015; Hale & Tager-Flusberg, 2003). However, less is known about how the two domains are related longitudinally beyond the preschool years. For instance, little is known about whether early language skills are related to advanced ToM and whether early explicit false-belief understanding is related to children’s further development in ToM, language, and other language-related or cognitive domains (Apperly, Samson, & Humphreys, 2009; Hughes, 2016; Lockl, Ebert, & Weinert, 2017). Thus, my main aim in this study was to longitudinally investigate the relation between language and ToM from preschool to early adolescence. In addition, I asked how both are connected to children’s later reading comprehension. The role of language in children’s early ToM ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Various theoretical accounts have explained why children’s language skills might be important for the emergence of ToM in preschool years. On the one hand, children’s language skills are a means of communication that enable them to take part in and make sense of verbal communication, especially when people talk about nonvisible mental states and processes (e.g., Harris, 2005; Nelson, 2005; Wellman & Peterson, 2013). Training and longitudinal studies show that verbal communication, especially talking about and elaborating on mental entities, promotes children’s ToM development (e.g., Ebert, Peterson, Slaughter, & Weinert, 2017; Lohmann & Tomasello, 2003). On the other hand, language is an important means for representing mental states and separating them from reality. For example, performance in false-belief tasks was found to improve after the experimenter provided labels for the different locations of hidden objects (Low & Simpson, 2012). Labeling may support children’s representation of nonobservable objects. In addition, the ability to use specific mental terms and mastery of syntax, especially structures with propositions embedded in clauses (e.g., “Hannes thinks that Joshua is climbing outdoors”), may help children to represent mental states and multiple propositions simultaneously (see also Astington & Baird, 2005a). Some language components seem more theoretically relevant for ToM development than others. Whereas pragmatic features of language and ToM are related by definition (Astington & Jenkins, 1999), there has been some deeper discussion of whether semantics or syntax are more important for keeping track of and representing false-beliefs (Astington & Jenkins, 1999; Harris, de Rosnay, & Pons, 2005; Slade & Ruffman, 2005). A meta-analysis by Milligan et al. (2007) reported a larger effect size for the longitudinal correlation between ToM and more general language measures than for receptive vocabulary but reported no other differences in correlations with various language measures. This suggests that in children’s preschool years various facets of language support their understanding of mental states and no single or specific language component is of primary importance (see also Harris et al., 2005; Miller, 2016). However, although this conclusion can be drawn with respect to children’s ToM development in preschool, the situation may be different beyond children’s preschool years. ToM beyond the preschool years ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ At about 5 years of age, when children explicitly understand false beliefs, they are said to have developed a metarepresentational understanding of the mind (Perner, 1991; Wellman, 2014). However, their understanding of mental states and processes continues to develop beyond this point. In particular, children learn that people differ in their opinions and interpretations of the same entity and gain a better understanding of mental states and processes as well as nonliteral meanings of people’s utterances in more complex social situations (e.g., Carpendale & Chandler, 1996; Weimer, Dowds, Fabricius, Schwanenflugel, & Suh, 2017). Accordingly, advanced ToM tests such as the Strange Stories task often incorporate social scenarios (White, Hill, Happe, & Frith, 2009). In these tasks, children need to refer to a character’s mental state to explain her or his actions correctly. Many researchers believe that advanced ToM includes no further changes in children’s conceptual ToM understanding (Apperly et al., 2009; Devine, White, Ensor, & Hughes, 2016; Keenan, 2003; Lecce, Bianco, Devine, & Hughes, 2017). Nevertheless, it is assumed that individual differences in the use of mental concepts in social situations exist. The “genuine variation” account states that although all typically developing children develop a metarepresentational understanding eventually, individual differences in the time point at which children acquire this understanding reflect “differences in the ease or fluency with which children or adults use their theory of mind to attribute mental states to others” (Hughes & Devine, 2015, p. 151). Thus, individual differences in early ToM should be related to individual differences in advanced ToM. The few studies that have measured ToM longitudinally in preschool and middle childhood report medium correlations, supporting this account (Devine et al., 2016; Ensor, Devine, Marks, & Hughes, 2014; Lecce, Caputi, & Pagnin, 2014). These studies also suggest that individual differences in language skills are related to advanced ToM, similarly as for early ToM. Relations between early language and advanced ToM ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Under the assumption that no further conceptual change occurs after having acquired false- belief understanding, language might play a different role for developing an advanced ToM than for developing a metarepresentational ToM in preschool. For example, it is possible that they correlate less or that correlations are rooted in different underlying mechanisms. For example, language might be correlated only with advanced ToM due to shared task demands or because it becomes a necessary component of ToM (Apperly et al., 2009; Miller, 2009). Recent studies show that pragmatic, semantic, and syntactic language measures are significantly related to ToM in middle childhood but more weakly in early adolescence (Banerjee, Watling, & Caputi, 2011; Devine & Hughes, 2013; Im-Bolter, Agostino, & Owens-Jaffray, 2016; Lecce, Ronchi, Del Sette, Bischetti, & Bambini, 2019). However, from a developmental and theoretical point of view, cross-sectional associations say nothing about developmental relations. To my knowledge, three longitudinal studies have assessed language measures as control variables and provide hints on whether early language skills scaffold ToM beyond the preschool years. Lecce et al. (2014) reported medium correlations between early receptive vocabulary at about 6 years of age and advanced ToM (social scenarios) at 10 years in a group of 49 children. Devine et al. (2016) similarly found relations between receptive vocabulary at 6 years of age and ToM at 10 years (various measures) in a group of 137 children even after controlling for earlier ToM, executive functions, and socioeconomic status (SES). Ensor et al. (2014) found an association between more general early verbal abilities at 3 years of age and advanced ToM (social scenarios) at 10 years. Although these three studies suggest a longitudinal relation between early language and advanced ToM, they all included only one language measure at a single time point. In addition, the studies differed in whether they considered early vocabulary or a more general language comprehension measure. Moreover, there were differences in the time point when the language measures were assessed; Lecce et al. (2014) and Devine et al. (2016) measured receptive vocabulary at about 5 or 6 years of age, whereas Ensor et al. (2014) assessed a more comprehensive language measure at 3 years. They also reported that the impact of early language on advanced ToM at 10 years of age was mediated by false-belief understanding at 6 years. Thus, it is not clear whether these various language measures actually were differentially related to advanced ToM. I extended these previous studies while simultaneously including two different language measures, vocabulary and grammar (sentence comprehension), at 3 and 5 years of age and comparing their developmental relations with ToM. Moreover, because early ToM was measured at 5 years of age, I was also able to specify direct and indirect effects of language at 3 years via early ToM on advanced ToM. Relations between early ToM and later language ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The relation between ToM and language skills might change as children get older and start school. Now, more complex language comprehension skills, such as those necessary in text comprehension tasks, become more central. ToM skills might support these language skills. For example, Pelletier and Astington (2004) assumed that children with more advanced understanding of mental entities are better able to connect settings, events, and actions described in a story with the characters’ thoughts, motives, and emotions. Thus, based on Bruner (1986) idea, children with more advanced ToM may integrate the landscape of action with the landscape of consciousness more easily when listening to or reading a story. In addition, an advanced ability to represent mental states and processes also promotes metacognitive knowledge and skills, which may further support text comprehension. Moreover, ToM may support understanding the author’s intentions, that is, understanding the aims of the text, the ability to connect information in the text with earlier knowledge, and the construction of a mental situation model (see Atkinson, Slade, Powell, & Levy, 2017; Guajardo & Cartwright, 2016; Kim, 2015). Indeed, some recent longitudinal studies have reported associations between ToM in preschool and children’s later text comprehension even after controlling for early language (Atkinson et al., 2017; Kim, 2015, 2016). Another study points to a longitudinal association between early ToM and later receptive vocabulary (Lecce et al., 2014). However, to my knowledge, no existing studies have examined the longitudinal association between early ToM in preschool and various advanced language skills systematically. Relations among ToM, language, and reading ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Early ToM may be predictive of not only advanced language skills but also later reading comprehension, which is closely related to children’s language skills. Indeed, according to the simple view of reading, reading comprehension is a product of listening comprehension and decoding skills (Hoover & Gough, 1990). Moreover, there is good evidence that early language skills predict later reading comprehension alongside early phonological information processing and decoding skills (e.g., Dickinson, McCabe, Anastasopoulos, Peisner-Feinberg, & Poe, 2003; Ebert & Weinert, 2013; Storch & Whitehurst, 2002). However, other higher-order cognitive skills, including ToM, are also discussed as further potential contributors. Theoretical explanations for the relation between ToM and reading comprehension of texts are like those for the relation between ToM and listening comprehension of texts (Atkinson et al., 2017; Guajardo & Cartwright, 2016; Kim, 2015, 2016). Thus, ToM may predict reading comprehension indirectly via listening comprehension, i.e. verbal language skills. The results of the few studies that have investigated the impact of early ToM on later reading comprehension are mixed. Whereas some studies found no direct link between ToM and reading comprehension at all or only via language skills (Guajardo & Cartwright, 2016; Kim, 2015, 2016; Lockl et al., 2017), others found direct relations even after controlling for early language skills (Atkinson et al., 2017; Boerma, Mol, & Jolles, 2017). The lack of association between early ToM and later reading comprehension reported in some studies might be explained by the fact that these studies measured reading comprehension at the beginning of learning to read, and tests at this stage may be so easy that higher-order cognitive processes are not relevant. However, Atkinson et al. (2017) found a direct link between early ToM and reading comprehension, even in 6-year-old children, although they did not directly model the role of early language skills in this relation. In contrast, Boerma et al. (2017), who also found a relation between ToM and reading comprehension, investigated only the relations between reading comprehension and advanced measures of ToM and language. Thus, it is not clear how early and advanced ToM as well as language skills together contribute to children’s reading comprehension in older children. The current study ~~~~~~~~~~~~~~~~~ The main aim of the current study was to provide a more systematic investigation of the longitudinal relation between language and ToM from preschool to early adolescence. If language is associated only with advanced ToM due to shared task demands or because language is a necessary component of advanced ToM, only concurrent but no longitudinal associations should be detected. I assumed that language supports children in understanding and keeping track of what is going on in social situations and, thus, in gaining better insights into people’s mental states and processes, as has been found for developing a metarepresentational ToM in preschool. Furthermore, like in preschool, children’s language skills enable them to take part in verbal exchanges and communicate and learn about others’ mental states and processes while inferring the fine-grained meaning of language in more complex social situations. Thus, I expected that early language skills also scaffold the development of advanced ToM both directly and indirectly via early ToM. The current study replicates and extends previous studies in various ways. First, the developmental period up to 13 years of age was investigated, extending the developmental periods observed in previous studies. Second, not only one early indicator for language at each time point was assessed; rather, two were investigated: one for vocabulary and one for grammar (sentence comprehension). This opens the opportunity to analyze whether different language indicators play a differential role in children’s ToM development between preschool and early adolescence. Third, ToM and language measures were assessed in early childhood and early adolescence; thus, the reciprocal relation between ToM and language beyond preschool can be investigated in more detail. Given that knowledge about mental states should support children’s development of metacognitive knowledge and inference-making skills, I assumed that early ToM is a stronger predictor of later listening comprehension of texts as an advanced language measure, than it is for later vocabulary. Furthermore, I investigated how the developmental relation between ToM and language skills is linked to reading comprehension. ToM and language skills are interrelated over time, and both are considered to predict later reading comprehension. However, to my knowledge, how early and advanced ToM and language skills together are related to reading comprehension in early adolescence has never before been investigated. From a theoretical point of view, I expected only indirect effects of early ToM and language measures on reading comprehension via advanced ToM and advanced language measures. A subordinate aim of the study was to investigate whether individual differences in early ToM are connected to children’s advanced ToM longitudinally by covering the age range from 5 to 12 years. In preschool false-belief tasks were used to assess ToM, and in early adolescence social scenario tasks (i.e., a measure of children’s use of mental states in social situations) were administered. According to previous studies and in line with the “genuine variation” account, I expected to find substantial correlations even after controlling for other cognitive and language measures. I assumed that children who understand mental concepts earlier than their peers are those who have learned to pay more attention to mental states and use them more easily and flexibly in early adolescence (see also Devine et al., 2016). Because nonverbal cognitive abilities, working memory, and family SES are related to both language and ToM, I also controlled for these variables in all analyses. Study design ~~~~~~~~~~~~ The current study was based on a subsample of a German longitudinal study that included children from diverse socioeconomic backgrounds as well as rural and urban areas. This subsample was tested for ToM at Wave 5 of the more comprehensive study. According to the study aims, various additional measurement points and measures were included (see Fig. 1 for an overview). For easier comprehension, I renumbered the waves of the more comprehensive longitudinal study and refer to Wave 1 and Wave 5 in preschool as Time 1 and Time 2 and to Wave 11 and Wave 12 in early adolescence as Time 3 and Time 4, respectively. Measures of Wave 1/Time 1 took place in 2005 and were included as baseline measures. Ethical approval for the comprehensive study was given by the university, and compliance with ethical standards was confirmed by the German Research Foundation (DFG), which funded the study. Informed consent to the children’s participation was obtained from their parents, and all information was provided voluntarily. The children were tested individually by trained research assistants in their preschools. In early adolescence, they were tested at home. The children had the opportunity to withdraw from testing at any point and received a small gift (e.g., sticker) after each testing session. The parents were also rewarded with a small gift for their participation in interviews and questionnaires.","The current study’s participants were from a subsample of 267 children from a more comprehensive longitudinal study. These 267 children should have received ToM measures at Time 2. However, 47 children could not be reached at this time point due to dropout or absence during the testing session. Of these 47 children, 11 were reached at Time 3 when advanced ToM was measured. Thus, those 231 children (125 boys) were included in the current sample. At Time 1, the children had a mean age of 3;6 (years;months) (M = 41.62 months, SD = 3.95). At the other measurement points included in this study, the children had mean ages of 5;6 (M = 63.62 months, SD = 3.95), 12;8 (M = 151.70 months, SD = 3.98), and 13;7 (M = 162.86 months, SD = 3.72). All children were born in Germany, and most (n = 213, 92.2%) had at least one parent with German as her or his native tongue. The educational and socioeconomic backgrounds of the sample were diverse. To measure SES, I refer to the family’s Highest International Socioeconomic Index (HISEI; Ganzeboom, De Graaf, Treiman, & De Leeuw, 1992), an international index of occupational status. The mean HISEI in the sample was 52.30 (SD = 15.93) on a scale ranging from 16 (e.g., cleaner, unskilled farm worker) to 90 (e.g., judge in a court of law). An example occupation with an ISEI of 52 is an administrator for electronic data processing. Theory of mind: False belief At Time 2, children completed one first-order unexpected content false-belief task (based on Perner, Leekam, & Wimmer, 1987) and one second-order false-belief task (Sullivan, Zaitchik, & Tager-Flusberg, 1994). Both tasks were acted out with small figures. First-order task After demonstrating that a peanut box unexpectedly contains a ball instead of peanuts, a naive protagonist (P1) arrived and children were asked the false- belief question (“What does P1 think is in the box?”) and a control question (“Did P1 look inside the box?”). Children needed to answer the control question correctly to be given 1 point on the false-belief question. Children were also given a second test question about their own belief (“Before you had a look inside the box, what did you think was inside?”). Total first-order false-belief scores ranged from 0 to 2 (M = 1.26, SD = 0.73). Second-order task Children were told a story about Peter, a boy who found his actual birthday present (a puppy) unbeknownst to his mother (Mum) who had told him that he would receive a different present (a toy). When Peter was absent, Peter’s Mum talked to Grandma about Peter’s present. Children needed to answer two control questions (“What has Mum really got Peter for his birthday?” and “What did Peter’s Mum say to him that he got for his birthday?”) and three test questions. These were one first-order knowledge access question (“Does Mum know that Peter saw the dog?”) and one second-order knowledge access question (“What does Mum answer to Grandma‘s question: Does Peter know what you got him for his birthday?”) as well as one second-order false-belief question (“What does Mum answer to Grandma’s question: What does Peter think you got him for his birthday?”). In total, children could earn up to 3 points for the second- order ToM task (M = 1.71, SD = 1.12). If children responded incorrectly to one of the control questions, they received a score of 0 for the test questions. Scores on the first-order and second-order tasks were correlated, r(220) = .40 and, thus, were summed to form a comprehensive ToM score. Theory of mind: Strange Stories At Time 3, we assessed ToM using two stories of deception, two stories of misunderstanding, and two stories of double bluff. One of the double-bluff stories was rewritten based on a story by White et al. (2009); all other stories were taken from a German translation of the Strange Stories (Rakoczy, Harder-Kasten, & Sturm, 2012). Children listened twice to each story and a subsequent open-ended question about one of the actors’ behavior from an MP3 player via loudspeaker. They answered the questions verbally, and their responses were recorded, transcribed, and coded according to White et al. (2009). Children received 1 point for a partially correct response and 2 points for a fully correct response. About 25% of the transcripts were coded by a second rater. Interrater reliability was good to excellent (intraclass correlation coefficient for absolute agreement between .78 and .89; Cohen’s kappa between .76 and .86). The scores for all items were summed to form a total score. Cronbach’s alpha was .56. Language: Receptive vocabulary At Time 1, Time 2, and Time 3, children’s receptive vocabulary was measured using a German research version of the Peabody Picture Vocabulary Test–Revised (PPVT; Dunn & Dunn, 1981). This test contained 175 items ordered in sets of 12 (with 7 items in the last set). The maximum possible score was 175. Language: Receptive grammar At Time 1, the sentence comprehension subtest of the German Language Development Test for 3- to 5-year-old children was administered (SETK 3–5; Grimm, 2001). In this test, children were given sentences varying in grammatical complexity. In the first section (9 items), they needed to determine which one of four pictures corresponded to the sentence they had just heard. In the second section (10 items), they needed to act out the content of the given sentence (e.g., “Put the blue pen under the bag”). The maximum possible score was 19. At Time 2, children’s sentence comprehension using a research version of the German adaptation of the Test for the Reception of Grammar (TROG; Bishop, 1983/1989; German version: TROG-D; Fox, 2006) was assessed. This adaptation contained 48 items and required children to select which one of four pictures corresponded to a verbally presented sentence. The research version included all the grammatical structures of the original test; the only difference was that the first three sets had 2 items rather than 4 items. The maximum possible score was 48. Language: Text comprehension At Time 4, six stories (each with approximately 100–150 words) from a paper-and- pencil language comprehension test for adolescents adopted from the DELKO project (Marx & Stanat, 2009) were administered. The stories varied in the complexity of vocabulary and syntax and were set in everyday contexts (e.g., conversation in a supermarket) or were more informational (e.g., text about a rare animal). Children listened to each story twice and were then asked three to five questions (multiple-choice and open-ended questions; 25 questions in total). For example, children were asked to recall or compare information or make inferences. Partially correct answers were given 0.5 points. The maximum possible score was 25. About 22% of the answers were coded by a second rater, and interrater reliability was good to excellent (intraclass correlation coefficient [absolute agreement] between .90 and .98; Cohen’s kappa between .76 and .96). The scores for all items were summed to form a total score. Cronbach’s alpha was .64. For some analyses, the two language measures conducted at the same time point were z-standardized, summed, and averaged. Correlations between the vocabulary and grammar measures were r(225) = .70 at Time 1 and r(211) = .57 at Time 2. Reading: Text comprehension At Time 4, a paper-and-pencil test developed as part of the German National Educational Panel Study (NEPS; Gehrer, Zimmermann, Artelt, & Weinert, 2012) was administered. The test version used was originally developed for ninth graders and encompassed five texts of different types (informational, commentary, literary, instructional, and advertising). Each text contained about 230 words and was followed by five to seven questions, mostly multiple-choice questions with one correct option out of four options. The other questions took the form of matching and decision-making tasks. All items required children to complete tasks such as extracting information and making inferences based on the text. Children had 28 min to complete the entire test. For some matching and decision-making tasks, partially correct answers were possible. The maximum possible score was 33. Nonverbal cognitive abilities At Time 1, the Analogies and Categories subtests from the Snijders–Oomen Nonverbal Intelligence Test for 2½- to 7-year-old children (SON-R 2½-7; Tellegen, Winkel, Wijnberg-Williams, & Laros, 2005) were administered as an indicator for children's nonverbal reasoning skills. This test asks children to infer sorting and classification principles and sort, classify, or categorize either abstract materials of various shapes and colors or picture cards. The maximum possible score was 17 for Analogies and 15 for Categories. Scores for the two subtests were standardized, summed, and averaged. The correlation between the subtests was r(224) = .48. Working memory At Time 1 two subtests from the German version of the Kaufman Assessment Battery for Children (K-ABC; Melchers & Preuß, 2003) were administered. In the Digit Span subtest, children answered 12 items clustered into sets of 3 items of identical length (2–5 digits). They earned 1 point for each correctly repeated item. In the Hand Movements subtest, children needed to repeat a sequence of taps on the table performed by the research assistant with their fist, palm, or side of their hand. The test consisted of 12 items differing in length (2–4 hand movements each) grouped into sets of 4 items. Children earned 1 point for each correctly repeated sequence. The maximum possible scores were 12. Test scores for the two subtests were standardized, summed, and averaged. The correlation between the subtests was r(218) = .51. Family background Parents’ education and SES were assessed via a computer-assisted personal interview (CAPI) with the main caregiver in each child’s home at Time 1. Missing data and analytical strategy ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 1 presents descriptive statistics for the cognitive and language measures included in the current study. For various reasons (no further consent to participate was given, the family moved, illness, technical problems, child refusal, etc.), I did not have valid data for all children at all measurement points. Due to the large developmental period under investigation, the dropout rate between the measurement points in early childhood and early adolescence was particularly high. However, there are good statistical solutions for missing data, especially in longitudinal research, that do not lead to biased estimates, although missingness mechanisms need to be considered appropriately (Graham, 2009). I did not assume that the data were missing completely at random (MCAR). Moreover, Little’s MCAR test was significant, χ2(308) = 517.39, p < .00. However, there is no reason to believe that advanced ToM itself predicts whether a participant has missing data on advanced ToM; rather, children from lower SES backgrounds are more likely to drop out of the study, for example, due to more frequent moves or stress. Moreover, it is known that SES is associated with early cognitive abilities. Indeed, the 112 children from whom we had a valid measure in advanced ToM differed in neither age, t(223) = −0.37, p = .71, nor language background from the children who left the study, χ2(2) = 1.85, p = .40, but were advanced in cognitive and language skills, F(9, 185) = 2.26, p = .02, and their family’s HISEI was higher, t(229) = −2.12, p = .04. Thus, I assumed that the data are missing at random. This means that missing values are systematically related to other observed variables (Enders, 2013) and that after controlling for “all the variables one has, any remaining missingness is completely at random” (Graham, 2009, p. 552). I used a full information maximum likelihood (FIML) approach to account for the missing data and included early cognitive and language variables as well as background variables for which missingness was rare as control variables in the model. FIML including control variables is a good means of handling data that are missing at random, especially incomplete outcome variables, and results in less biased parameter estimates than older methods (Enders, 2013; Graham, 2003). Thus, FIML is superior to listwise deletion, pairwise deletion, and similar response imputation, especially in small sample sizes (Enders & Bandalos, 2001). To analyze the longitudinal association between language and ToM as well as between these variables and reading comprehension in more detail, I specified path models using Mplus 7 (Muthén & Muthén, 2012). I refrained from estimating latent variables to keep the structural equation simple and the sample size-to-parameter ratio low. This increases the likelihood that the statistical requirements will be met even though our sample size is not huge and missing data are estimated (Kline, 2016). I controlled for autoregressive effects in all models, which enables determining whether the relations are bidirectional or solely in one direction. Only if a relation between early and later measures is observed after controlling for autoregressive effects can the earlier measure be said to predict developmental change and might be causally related to the later measure (see Ruffman, Slade, & Crowe, 2002). I also considered direct and indirect links between language and ToM over time and allowed the predictor variables assessed at one time point to correlate. Moreover, I included HISEI and the composite scores for nonverbal cognitive abilities and working memory at Wave 1 as control variables in all models. Hence, I specified paths between these control variables and all outcome measures at all measurement points. I also controlled for age by specifying paths between concurrent age and the cognitive and language measures. I report all significant and nonsignificant paths that were specified among ToM, language, and reading measures in the figures. For the other variables, I report only those paths and standardized beta weights that approached significance at p < .10 for easier readability. The criteria for model fit were as follows: a root mean square error of approximation (RMSEA) < .08, a comparative fit index (CFI) > .95, and a nonsignificant chi-square test of model fit (Brown, 2006; Hu & Bentler, 1999). Relations between ToM and language over time ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As shown in Table 2, longitudinal and concurrent correlations between the aggregate language measures and ToM were moderate to high at all measurement points. Table 2 also shows that the association between vocabulary and ToM seems to change across waves and varies for early versus advanced ToM. Thus, PPVT at age 3;6 correlates stronger with ToM at age 5;6 than with ToM at age 12;8, whereas PPVT at age 5;6 correlates stronger with ToM at age 5;6 than with ToM at age 12;8. In contrast, the association between grammar and ToM seems to be more stable over time. Thus, the grammar measure correlates similarly with early ToM and advanced ToM at both measurement points in preschool. To further investigate the reciprocal relations between different facets of language and ToM, I specified three equivalent path models including different language measures for preschool (see Fig. 2) to predict advanced ToM and language. I used the vocabulary measure in Model 1, the grammar measure in Model 2, and the aggregate language score consisting of both vocabulary and grammar in Model 3. All models fit the data very well (see Fig. 2). However, the models revealed both similarities and differences in the longitudinal relation between language and ToM. In all models, there was a significant direct effect of language at Time 1 on ToM in preschool and adolescence. However, the beta weight between vocabulary at age 3;6 and preschool ToM (β = .24) was much smaller than the one between grammar at age 3;6 and preschool ToM (β = .51). In addition, the relation between language measures at age 5;6 and advanced ToM differed across the three language measures; even after controlling for language and various other control variables 2 years earlier, grammar had a direct effect on the change in ToM between preschool and early adolescence, whereas vocabulary and the aggregate score for language did not. Furthermore, ToM significantly predicted changes in vocabulary between ages 5;6 and 12;8, even when controlling for earlier vocabulary and other control variables, but had no direct effect on listening comprehension. Indirect effects of early language on advanced ToM via preschool ToM were found only for vocabulary at age 3;6 (β = .07, p < .05). In addition, there was an indirect effect of early vocabulary on vocabulary at age 12;8 via preschool ToM (β = .06, p < .05). Relations among ToM, language, and reading ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To investigate whether ToM is related to later reading comprehension after controlling for language skills, I specified path models that, besides the models in Fig. 2, included a measure of reading comprehension (see Fig. 3). First, I specified a model that was similar to Model 3 (see Fig. 2) but included only preschool measures and reading comprehension in early adolescence (Model a in Fig. 3). This was done to test for direct effects of early language and ToM on later reading comprehension. In a second model, I also included advanced ToM and language measures (Modell b in Fig. 3). Model a in Fig. 3 shows that early ToM has no direct effect on later reading comprehension. There is only a direct effect of early language measures on later reading comprehension. There were also no significant indirect effects of early ToM and later reading comprehension via language and ToM measures in early adolescence (Model b in Fig. 3). However, a small indirect effect of advanced ToM via listening comprehension on reading comprehension that approaches significance was found (β = .07, p < .10). In contrast, both language measures in early preschool showed indirect effects on later reading comprehension. These effects were mainly mediated by language measures in early adolescence. Relations between early and advanced ToM ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As can be seen in Table 1, early ToM at age 5;6 and advanced ToM at age 12;8 were moderately correlated, r(101) = .39, p < .01. This correlation hardly changed after controlling for age (rp = .38, p < .01) and remained substantial even after additionally accounting for early nonverbal abilities (rp = .32, p < .01). However, Fig. 2 shows that after also controlling for early sentence comprehension (β = .12; see Model 2), the correlation between early and advanced ToM was substantially reduced, whereas it remained significant when controlling for early vocabulary (β = .28; see Model 1). Given that the advanced ToM measure required a large amount of listening comprehension, I tested whether a specific correlation between early ToM and advanced ToM existed beyond the effect of general listening comprehension needed for the Strange Stories. That is, I controlled for children’s general listening comprehension in adolescence. Therefore, I regressed the advanced ToM score on the listening comprehension score (DELKO). The correlation between early ToM and the residual of advanced ToM was substantial, r(93) = .22, p < .05.","The current study replicates and extends former research by showing that individual differences in language measures—especially in early sentence comprehension—predict changes in ToM from preschool until early adolescence. This is a time interval of more than 7 years. The study also revealed a reciprocal relation between language and ToM in this developmental period and showed that early ToM predicts children’s vocabulary. However, early ToM had neither direct nor indirect effects on reading comprehension at 13;7 years of age, and only advanced ToM exerted small indirect effects on reading comprehension via listening comprehension. Moreover, the results confirmed and extended prior studies by demonstrating that individual differences in ToM exhibit a moderate level of stability across a period of 7 years. These results are discussed in more detail in the following sections. Relations between early language and ToM ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Consistent with earlier research, the study clearly demonstrated that children’s language skills are strongly correlated with ToM in their preschool years. In addition, language skills measured as early as age 3;6 predicted the change in ToM between preschool and early adolescence. Even though there were no indirect effects of early language skills via ToM or later language skills on children’s advanced ToM, the findings underscore the importance of early language skills, which provide children with a strong start and continue to promote their social understanding until early adolescence at least. In addition, the study provides evidence that there are differential effects of language measures on ToM development and that these differential effects might change as children grow older. In early preschool, individual differences in vocabulary and grammar both predicted children’s advanced ToM after controlling for early ToM and language; two years later, only children’s grammar predicted their further ToM development. There are at least two possible explanations for why language and especially receptive grammar/sentence comprehension may be important for advanced ToM. First, advanced ToM tasks such as the Strange Stories are linguistic tasks. Children need to comprehend and verbally respond to stories that are not acted out with puppets (in contrast to early ToM tests). One way to investigate whether language is associated only due to linguistic requirements with advanced ToM would be to measure children’s advanced ToM using tasks that are less verbal or even nonverbal. It has been shown in numerous ways that language is associated with early ToM over and above the two constructs’ shared linguistic nature (Astington & Baird, 2005b). With regard to advanced ToM, Devine et al. (2016) study may provide some insight. The authors of that study administered various advanced ToM tasks that differed in their verbal requirements. Unfortunately, they did not report whether differences were found in the associations between the respective tasks and language measures. Thus, it is not known whether the relation between early language and advanced ToM varies with verbal task requirements. However, I assume that the Strange Stories are not simply a test of text comprehension. This assumption is supported by the finding that the advanced ToM measure and the listening comprehension measure were predicted differently by our three language measures. In addition, and as expected, reading comprehension was more strongly correlated with listening comprehension than with advanced ToM. If the advanced ToM measure were simply a measure of listening comprehension of texts, listening comprehension of texts and advanced ToM should show stronger similarities in correlational patterns. Moreover, several prior studies support my assumption by showing that comprehension skills for texts including mental content (like the Strange Stories) are differently correlated with other ToM measures and executive functions than comprehension skills for texts including physical content (Lecce et al., 2019; Rakoczy et al., 2012). Another possible explanation for the relation between language and advanced ToM is that advanced ToM is typically applied in social situations based on language. Misunderstandings, double bluffs, and persuasion are impossible without interpersonal and primarily verbal communication. Thus, the question arises whether advanced ToM occurs only in linguistic situations. I assume that even if advanced ToM means in particular paying more attention to mental states in social situations and, thus, mentally interpreting social situations more easily and flexibly, this does not necessarily mean that the social situation must be a verbal one. Hence, alternative measures of ToM suggest nonverbal forms of advanced ToM. For example, in the “Reading the Mind in the Eyes Test” (Baron-Cohen, Wheelwright, Hill, Raste, & Plumb, 2001), people must infer a mental state from photographs of the eye region. Nevertheless, language might have supported the development of one’s ability to pay more attention to people’s mental states. Indeed, this is exactly what the current study demonstrates, namely that it is particularly children’s early language that predicts children’s advanced ToM. Language and in particular syntactical skills allow children to represent mental phenomena, take part in verbal interactions, and provide opportunities to learn about mental phenomena as well as about people and their mental states. As children grow older, the more “functional properties of language as a vehicle for communication and social exchange with others” (Hughes, 2005, p. 321) may become more important for understanding others. In contrast to syntactical skills, a rich vocabulary was less related to advanced ToM. This finding seems to contradict Lecce et al. (2014) and Devine et al. (2016), both of whom found effects of vocabulary at 5 or 6 years of age on advanced ToM. However, in contrast to the current study, they did not consider language skills at an earlier time point. Nevertheless, even when running a model that includes only vocabulary at age 5;6 and no controls for language and cognition at age 3;6, no effect of vocabulary at age 5;6 on advanced ToM was found. Admittedly, I cannot exclude the possibility that differences in the predictability of various language measures at different measurement points might have emerged by chance, and I agree with Slade and Ruffman (2005) that children’s general language skills are more important than any single component when it comes to taking part in the community of minds. However, the current study showed that language measures that are strongly correlated may nevertheless differ in their predictive utility for ToM. Indeed, receptive vocabulary and receptive grammar/sentence comprehension were very strongly correlated at Time 1, meaning that it was not methodologically possible to include them as separate measures in a single analysis without suppression effects occurring. However, the predictive value of the language measured differed despite the fact that the same sample was considered. Thus, whenever effects of language skills are discussed, it seems important to consider which facets of language one is talking about. The current study’s result that vocabulary also has an indirect effect via early ToM on advanced ToM was consistent with Ensor et al. (2014), who found an indirect effect of general language skills via early ToM on advanced ToM. However, in the current study, no indirect effects of the grammar measure or the aggregate language measure via early ToM on advanced ToM were found. In addition, there were also direct effects of all three language measures on advanced ToM. This suggests that the link between early language and advanced ToM is not mainly due to the fact that language supports a metarepresentational understanding, and this in turn helps to build an advanced mental understanding, but rather that advanced ToM emerges directly due to the support of early language skills. Most important is that language skills predicted even changes in ToM between preschool and early adolescence. Apperly et al. (2009) used a metaphor of cement and scaffolding to describe the roles a variable might play in development. If a variable is like cement, it becomes part of the construct; if it is like a scaffold, it is no longer important after a certain point in development. Based on neuropsychological data, those authors proposed that in adulthood grammar is independent of ToM and, thus, not part of the construct and not cement, although it might have had a scaffolding role earlier in development. The current study suggests that for ToM in early adolescence, language has both roles; the high concurrent correlations between language and ToM in early adolescence propose that language is part of the construct, and the longitudinal relations suggest that language might have scaffolded the development of advanced ToM. Thus, ToM in early adolescence might be an intermediate state before children achieve an adult ToM. In addition, the results support the idea that grammar and vocabulary may play different roles in ToM development and that ToM is particularly related to more complex language skills needed for communication and not simply to a richer vocabulary. Relations between early ToM and later language ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ It was hypothesized that children’s early ToM supports their metacognitive knowledge and inference-making skills and, thus, is particularly predictive for listening comprehension of texts but less for vocabulary. However, the opposite was found to be the case; early ToM was predictive of changes in children’s vocabulary between ages 5;6 and 12;8, but it was not predictive of text comprehension, after controlling for earlier language skills. Three possible explanations for these results are proposed. First, the vocabulary measure used at older ages may have included more abstract and mental words and, thus, may have indirectly tested children’s advanced ToM. However, the more complex terms in the PPVT are not specifically mental. Another possibility is that the effect of ToM on vocabulary actually reflects the effect of a general text comprehension measure and not a specific effect of ToM. However, there is a unique effect of ToM on vocabulary at age 12;8 even after controlling for early sentence comprehension instead of early vocabulary (see Model 2 in Fig. 2). A third explanation is that as children grow older a receptive vocabulary test such as the PPVT not only might measure children’s vocabulary size but also, and more strongly, might reflect children’s verbal intelligence. If this is the case, the current study’s results are even more intriguing. Early ToM would then be an early predictor of children’s later verbal intelligence, which is important for so many aspects of their lives. That vocabulary in early adolescence might be more a measure of verbal knowledge and intelligence may also be reflected in the fact that early receptive grammar/sentence comprehension is more strongly correlated with advanced ToM than early receptive vocabulary. However, because I did not expect a relation between early ToM and later vocabulary, the result will need to be replicated and investigated in more detail in subsequent studies. Relation among ToM, language, and reading comprehension ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the current study, a small indirect effect of advanced ToM on children’s reading comprehension via listening comprehension of texts was observed. This finding is consistent with the theoretical assumption that ToM facilitates the development of higher- order comprehension skills in children (Atkinson et al., 2017; Guajardo & Cartwright, 2016; Kim, 2015). However, the effect was small and found only for advanced ToM, not for early ToM. Moreover, the advanced ToM measure did not mask the effects of early ToM. This might be explained by the fact that, in contrast to other studies, I controlled for early vocabulary and grammar and also included a measure of early and advanced ToM. Indeed, many prior studies that investigated the link between ToM and reading comprehension also found no effect or only indirect effects of early ToM on later reading comprehension via language skills (Guajardo & Cartwright, 2016; Kim, 2015, 2016; Lockl et al., 2017). However, in the current study, reading comprehension was assessed in early adolescence when children are advanced in basic reading skill and, thus, higher-order cognitive skills should be more strongly correlated with reading comprehension than in these previous studies. Nevertheless, our reading comprehension test may have been too demanding for children of this age given that it was originally designed for older children; thus, the demands made on basic reading processes may have masked the effects of higher-order inference-making skills, meaning that even advanced ToM was of little help. Contrary to this assumption, early and advanced language measures both were found to be important predictors for reading comprehension. Thus, there seems to be no specific effect of early ToM on children’s reading comprehension after considering language skills. Stability of individual differences in ToM ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The results demonstrate that young adolescents differ in how they answer the Strange Stories and that these individual differences are related to differences in children’s understanding of basic ToM concepts 7 years earlier. Hence, children who developed a more basic understanding of mental states earlier than their peers during preschool also exhibit better advanced ToM in early adolescence. However, the study’s correlational findings are silent with respect to whether the Strange Stories indeed assess the application of already understood mental concepts or reflect another step in children’s conceptual ToM development (see Peterson & Wellman, 2019). The moderate internal consistency of the Strange Stories suggests that different aspects of advanced ToM may be related to early ToM to differing degrees. Indeed, when differentiating between the mental concepts measured (misunderstanding, double bluff, and deception), the highest correlations between false beliefs in preschool and advanced ToM were found for double bluff stories (r = .50). Because there were only two tasks per concept included, there is a need for further research including more tasks per mental concept in order to examine whether early ToM is associated with better social understanding in general or only with specific mental concepts. There is also recent evidence that ToM shows a diverse structure, at least beyond the preschool years (Warnell & Redcay, 2019). In confirmation of other studies, the current study supports the assumption of stable individual differences in ToM over time (Ensor et al., 2014; Devine et al., 2016), although children were followed for a longer time period, that is, up to 12 years of age. One implication of this is that children’s school environment changed fundamentally during this period: Children left preschool about 6 years of age, went to elementary school for another 4 years, and eventually entered secondary school. In Germany, children enter different school tracks after Grade 4 (about 10 years of age) depending on their performance in elementary school. Hence, children had very different educational environments at 12 years of age (i.e., when advanced ToM was assessed), and they had undergone numerous experiences between assessments. Given that I did not find lower stability in ToM than Ensor et al. (2014) and Devine et al. (2016), an “environmental account of stability” (Devine et al., 2016, p. 769) seems unlikely. However, the factors that support ToM development might not vary greatly across different school environments and, thus, might not affect stability. Moreover, this does not mean that environmental factors have no influence on the development of ToM after preschool. Stability was moderate at best, and there is much room for other factors to affect development as, for example, was shown for language skills. Thus, the study shows that the correlation between early and advanced ToM was dramatically reduced when it was controlled for early grammar, but not when it was controlled for early vocabulary. This indicates that early ToM also reflects language skills.","The current study adds to previous research on ToM with a more thorough analysis of how ToM and language skills are linked over an extended time period of 7 years. Nevertheless, some limitations need consideration. First, internal consistency of advanced ToM was rather low, and scores were close to ceiling. In addition, both reading and listening comprehension tests were newly designed measures, which may implicate reduced reliability. Thus, the internal consistency of the listening comprehension test was rather low, and a number of children had problems in completing the reading comprehension test in time. Thus, the reading comprehension test might cover not only reading comprehension but, to an extent, also reading speed. Another challenge of the current study was the high dropout rate of about 50%. Nevertheless, given that the study covers a period of more than 10 years in childhood including multiple measurement points, the sample size was still substantial and included families of a wide variety of SES backgrounds from urban and rural areas. I am convinced that even if attrition is high, the use of longitudinal data is critical for gathering information about relations between variables over time and development, pending good methods to deal with missing data. In addition, given that previous studies on ToM often refer to SES homogeneous samples, despite our study’s high attrition, the study makes an important contribution. Unfortunately, it was not possible to use the same measures of ToM and language (except for vocabulary) in preschool and adolescence. Thus, it was not possible to deliver evidence of how individual differences in early ToM and language contribute to “real” growth in these concepts. In addition, in the current study, only a verbal measure of children’s advanced ToM was included. This made it somewhat difficult to investigate the “functional” link between early language skills and advanced ToM. However, it must be taken into account that social understanding and language may be inextricably intertwined and that early ToM develops out of language skills (e.g., Astington & Jenkins, 1999). Thus, ToM is interwoven with language. This is also reflected in the results demonstrating that language and ToM are closely reciprocally intertwined over time; language skills drive developmental changes in ToM, and ToM also drives changes in language development."],["Previous studies with preschoolers have reported “East–West” contrasts in children's executive function (East > West) and theory of mind (East < West). This cross-cultural study with two samples of older children from the United Kingdom and Hong Kong aimed to test competing accounts of these contrasts that focus on either global effects of culture or more specific effects of pedagogical experience. Both groups of children in Hong Kong outperformed the British children on executive function tasks. That is, with respect to executive function, general cultural influences appear to be salient. In contrast, compared with their U.K. counterparts, children attending local schools in Hong Kong (but not those attending British-based international schools in Hong Kong) performed poorly on age-appropriate tests of theory of mind. With respect to theory of mind, therefore, pedagogical experiences appear to be more salient than factors related to the broad contrast between individualist and collectivist cultures. Our findings also contribute to the debate surrounding the relationship between theory of mind and executive function; although scores on these two sets of tasks were robustly correlated within each country, the double dissociation between delayed theory of mind but superior executive function for children in local schools in Hong Kong compared with their U.K. peers suggests that variation in executive function may be necessary but is not sufficient to explain variation in theory of mind. --------------------------------------------------------------------------------","Theory of mind is a developmental achievement that emerges early in life and continues to develop during adolescence and adulthood. Developments in cognitive domains such as language and executive function, as well as social factors such as cultural practice, family context, and interactional and pedagogical experience, all support the process of gaining insight into people’s mental world (for a comprehensive review, see Hughes & Devine, 2015). Research in this field has been largely restricted to young children (e.g., Wellman, Cross, & Watson, 2001; Wellman & Liu, 2004), although recent years have seen growing interest in infants’ mental state understanding (e.g., Onishi & Baillargeon, 2005; Sodian, 2011). Moreover, several studies have reported striking individual differences in adults’ perspective taking (Dodell-Feder, Lincoln, Coulson, & Hooker, 2013; Ferguson & Austin, 2010; Keysar, Lin, & Barr, 2003; Royzman, Cassidy, & Baron, 2003). What remains scarce, however, is research on theory of mind during middle childhood that bridges the gap between these two fields. A few studies (e.g., Banerjee, Watling, & Caputi, 2011; Devine & Hughes, 2013; Dumontheil, Apperly, & Blakemore, 2010) have shown age-related gains in theory of mind across middle childhood, but individual differences in theory-of- mind development during this period and their underlying mechanisms remain poorly understood. To address this challenge, the current study examined both internal and external factors that might contribute to individual difference in theory of mind beyond early childhood. With regard to internal factors, we focused on executive function, the higher level cognitive ability that underpins flexible, goal-directed activity using working memory, attention shifting, and inhibitory control (Blair, Zelazo, & Greenberg, 2005; Diamond, 2013). With regard to external factors, we focused on social and pedagogical experiences at school, which increase in frequency and complexity across middle childhood (Eccles, 1999). Theory of mind across cultures ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Although most theory-of-mind research has been conducted within Anglo-Saxon countries (Hughes & Devine, 2015), the past decade has seen a marked increase in cross-cultural research (e.g., Callaghan et al., 2005; Liu, Wellman, Tardif, & Sabbagh, 2008), including several cross-cultural comparisons (Ahn & Miller, 2012; Hughes et al., 2014; Lecce & Hughes, 2010; Lewis et al., 2009; Sabbagh, Xu, Carlson, Moses, & Lee, 2006). This cross- cultural perspective is helpful in unraveling the nature versus nurture riddle in theory- of-mind development: Is theory of mind an innate, culturally universal construct that is merely triggered by environmental factors, or is it cultivated in a context of social interaction, displaying culturally specific developmental routes? Existing findings are mixed. Some cross-cultural studies (e.g., Callaghan et al., 2005; Oberle, 2009) have highlighted synchrony in the onset of false belief understanding, but others have reported dramatic contrasts. For example, Mayer and Träuble (2013) found that under the age of 8 years Samoan children overwhelmingly failed false belief tasks, with one third of 10- to 13-year-old Samoan children also failing. These results echo earlier reports of delays in non-Western children’s understanding of false belief (e.g., Naito & Koyama, 2006; Vinden, 1996). Moreover, a meta-analysis of data from more than 3000 children from mainland China and Hong Kong showed that although mainland Chinese children were more or less in line with North American children in the onset of false belief understanding, children in Hong Kong lagged behind by up to 2 years (Liu et al., 2008). If preschoolers in Hong Kong lag behind their Western peers by up to 2 years in their false belief understanding (Liu et al., 2008), do they catch up eventually? There is some evidence to suggest that Chinese adults might in fact have better perspective taking than their Western counterparts. Wu and Keysar (2007) found that bilingual Chinese American adults outperformed European Americans in perspective taking. However, a later reanalysis of Wu and Keysar’s data using a more time-sensitive approach found that the Chinese participants made the same egocentric mistakes initially in interference as the European American participants but suppressed the interference earlier and more effectively (Wu, Barr, Gann, & Keysar, 2013). That is, rather than having a more advanced perspective taking ability, the bilingual Chinese participants appeared to capitalize on their relatively advanced executive functions (Sabbagh et al., 2006). To date, no study has gone beyond the preschool years in comparing the social understanding of children in Hong Kong and Western children. The first aim of our study was to compare older children from the United Kingdom and Hong Kong on a battery of age-appropriate theory-of-mind tests in order to assess whether children in Hong Kong “catch up” in theory of mind during middle childhood. Explaining cultural contrasts: The role of educational experiences ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ How might potential differences in theory-of-mind use between children from Hong Kong and children from the United Kingdom be explained? According to one hypothesis, our cultural norms (collectivist vs. individualist) about actions and cultural specific epistemologies shape our mental state inferences (Ames et al., 2001). For example, the emphasis on obedience and behavioral inhibitory control in parenting practice within collectivist cultures might lead to reduced conversational exposure to mental state terms (e.g., Mayer & Träuble, 2013). Related to this view, Lu, Su, and Wang (2008) found that, in both a longitudinal study and a training study, Chinese preschoolers’ success on false belief tasks was associated with increased conversational references to other people rather than to talks about mental states in particular. In contrast, Chasiotis, Kiessling, Hofer, and Campos (2006) argued that conflict inhibition, but not delay inhibition, is a culture- independent, universal developmental prerequisite for the development of theory of mind. As Hughes et al. (2014) have noted, attributing the cross-cultural difference in theory of mind to individualist versus collectivist contrasts is an oversimplification. Moreover, this hypothesis fails to account for the discrepancy between Hong Kong children and the mainland Chinese children found by Liu et al. (2008). In similar collectivist cultures, Japanese children passed the false belief task at around 6 to 8 years of age (Naito & Koyama, 2006), considerably later than the Western norm of 4 years (Wellman et al., 2001), whereas Korean children have been found to outperform their U.S. and British counterparts on false belief tasks (Ahn & Miller, 2012). Likewise, contrasts have been reported between children from different parts of Europe: the United Kingdom and Italy (Lecce & Hughes, 2010). In the first cross-cultural study of theory of mind to establish measurement invariance (i.e., to ensure that group contrasts did not simply reflect spurious measurement effects), Hughes et al. (2014) compared means on a latent theory-of-mind factor in school-aged children from the United Kingdom, Italy, and Japan matched on age, gender, and verbal ability. Their findings replicated both the U.K.–Italy contrast (Lecce & Hughes, 2010) and previous reports of a delay in Japanese children (Naito & Koyama, 2006). However, there was no clear difference between Italian and Japanese children, challenging a simple individualist versus collectivist contrast. Instead, Hughes et al. (2014) proposed that the differences reflected contrasts in pedagogical experience because children in the United Kingdom start formal schooling at age 4 or 5 years, at least a year earlier than Italian and Japanese children. The “pedagogical experience” hypothesis rectifies the oversimplification of individualist versus collectivist contrast by shifting the focus from a broad cultural construct to a more practical, socially organized activity (Ratner, 1999), in this case, schooling, as the primary cultural influence on psychology. This pedagogical experience hypothesis is in line with an environmental account that emphasizes the social origins of individual differences in understanding of mind. Evidence for this view comes from several distinctive lines of research, including studies of twins and deaf children born to hearing versus deaf parents, training studies, and longitudinal studies of relations between family discourse and theory of mind (for a comprehensive review, see Hughes & Devine, 2015). Formal schooling immerses children in learning activities that involve interpreting both epistemic mental states (e.g., knowledge, memories, ignorance, false belief) and motivational mental states (e.g., attention, intention). These experiences are likely to benefit children’s theory-of-mind development (Frye & Wang, 2008; Tomasello, Kruger, & Ratner, 1993). Consistent with this view, theory of mind has been related to children’s understanding of the concept of teaching and learning (Frye & Ziv, 2005; Wang, 2010; Ziv & Frye, 2004; Ziv, Solomon, & Frye, 2008) as well as children’s own teaching activity (Davis-Unger & Carlson, 2008a, 2008b). Our study provided a unique opportunity to test the pedagogical hypothesis. As a former British colony in the Far East that is now “Asia’s World City,” Hong Kong includes both international schools, most of which adopt inquiry-based curricula and use English as a mode of instruction, and Chinese-style local schools, which focus heavily on academic learning (Watkins & Biggs, 2001), even in early childhood classrooms. Teachers in Hong Kong local schools emphasize behavioral control and the need to follow instructions instead of inquiry (Cheng, Benson, Lau, & Fung, 2009). Children’s early pedagogical experience is mainly directed to the mastery of language and literacy, which is predominantly taught in local Hong Kong schools via a drill-and-practice approach (Li & Rao, 2005) due to the fact that children live in a trilingual (Cantonese, Mandarin, and English) biliterate (Chinese and English) society. According to the pedagogical hypothesis, children attending international schools in Hong Kong should perform just as well on theory-of-mind tasks as their British counterparts because they share a similar schooling experience. In contrast, children attending local schools in Hong Kong have quite distinct school experiences that are less rich in opportunities for discussing different points of view and so might be expected to lag behind their British counterparts in their theory-of-mind performance. The second goal of the current study was to test this pedagogical hypothesis. By pairing up U.K. pupils and Hong Kong international school pupils, as well as U.K. pupils and Hong Kong local school pupils, we singled out the effect of pedagogical experience that is usually embedded in a broader cultural context. According to the collectivist versus individualist culture divide hypothesis, both groups of children in Hong Kong should differ from children in the United Kingdom in their theory-of-mind performance. According to the pedagogical hypothesis, however, children in the United Kingdom are expected to show better theory of mind than children in Hong Kong attending local schools but not those attending international schools. Executive function in theory of mind during middle childhood ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Executive function is closely associated with the acquisition of false belief during early childhood. The relation between theory of mind and executive function can be summarized by either the “expression” account or the “emergence” account (Moses, 2001). According to the expression account, the executive demands of theory-of-mind tasks (in particular false belief tasks) mask children’s understanding of mind. Specifically, children’s failure on theory-of-mind tasks reflects the executive demands in those tasks rather than a lack of theory of mind. According to the emergence account, however, executive function is a necessary condition for the acquisition of theory-of-mind understanding. Children need to acknowledge different points of view and be able to hold back their own perspective in order to appreciate other perspectives. In a recent meta-analysis based on 102 studies representing close to 10,000 3- to 6-year-old children from 15 countries, Devine and Hughes (2014) reported a consistent and moderate correlation (r = .38) between false belief understanding and executive function that remained significant when effects of age and verbal ability were controlled. Further analysis of 10 longitudinal studies indicated that early individual differences in executive function modestly predicted later theory of mind, even when controlling for previous theory-of-mind scores and verbal ability, but not vice versa. Interestingly, the correlation between executive function and theory of mind appeared to be consistent across different cultures. The findings support an emergence account of the relation between theory of mind and executive function. The expression account posits that children with the greatest levels of executive function performance should outperform their peers on measures of theory of mind, whereas the emergence account is more tolerant to a disassociation between theory-of-mind performance and executive function performance. Sabbagh et al. (2006) compared American and Chinese preschool children on measures of theory of mind and executive function. Despite a marked advantage in executive function relative to their American peers, Chinese children did not outperform their American counterparts on measures of theory of mind. The results seem to support the emergence account. The expression account also predicts that only the theory- of-mind tasks with high executive demands, but not those with low executive demands, would correlate with executive function; on the contrary, the emergence account predicts that executive function should correlate with theory-of-mind tasks with either high executive demands or low executive demands. Testing this possibility, Carlson, Claxton, and Moses (2015) measured preschool children with executive function tasks, theory-of-mind tasks with high executive demands, and theory-of-mind tasks with low executive demands. They found that executive function tasks, specifically conflict executive function tasks, correlated uniformly with theory-of-mind measures that imposed either high or low executive demands when verbal ability was controlled, again supporting an emergence account. During middle childhood and adolescence, children continue to develop both executive function (e.g., Davidson, Amso, Anderson, & Diamond, 2006; Huizinga, Dolan, & van der Molen, 2006) and theory of mind (e.g., Devine & Hughes, 2013; Dumontheil et al., 2010). Moreover, individual differences in theory of mind remain correlated with executive function during middle childhood (e.g., Bock, Gallaway, & Hund, 2014; Lagattuta, Sayfan, & Blattman, 2010; Lagattuta, Sayfan, & Harvey, 2014). To date, however, the links between theory of mind and executive function during middle childhood have not been studied in a cross-cultural context. Examining the relations between theory of mind and executive function in different cultures provides an opportunity to understand the nature of these links. Our third goal, therefore, was to compare executive function and theory-of-mind performance in school-aged children from the United Kingdom and Hong Kong using online theory-of-mind measures and examine the association between individual differences in these two domains in each sample. In doing so, we sought to examine the universality of the relation between theory of mind and executive function during middle childhood. Moreover, our study gave us an opportunity to test the expression account versus emergence account of the executive function and theory-of-mind relation. According to the expression account, if children in Hong Kong outperformed the British children on executive function measures, they should also do better in theory-of-mind measures. According to the emergence account, children in Hong Kong would not necessarily outperform their British counterparts in theory-of-mind measures even if they did so in executive function measures. In contrast to false belief tasks, in which children need to inhibit their own more salient knowledge about a situation in order to infer a protagonist’s false belief, the online measures of theory of mind simply require children to explain the protagonist’s behaviors and so impose a lower executive demand. According to the expression account, these measures would not correlate well with executive function measures. The emergence account predicts that executive function measures should correlate with theory-of-mind measures with either high executive demands or low executive demands (Moses, 2001). In summary, our study had three key aims: (a) to investigate similarities and differences in theory of mind in children from the United Kingdom and Hong Kong, (b) to assess the role of pedagogical experiences in shaping individual differences in theory of mind, and (c) to examine the universality of the relations between theory of mind and executive function during middle childhood. To address these three aims, we drew on two data sets that included aggregate measures of theory of mind during middle childhood that were administered to samples of children in Hong Kong and the United Kingdom.","Sample 1 consisted of 118 children (48% male) recruited from seven English-speaking international schools in Hong Kong and six state primary and secondary schools in the United Kingdom as part of a follow-up study of children’s social development in the United Kingdom and Hong Kong led by the third author (Wong, Freeman, & Hughes, 2014). Follow-up children were matched in age, t(118) = 0.57, p = .57, and self-reported family affluence, t(117) = 0.49, p = .63. Inclusion criteria for both sites were (a) no known developmental delays or disabilities and (b) native speaker of English or spoke English as a second language. The U.K. sample was predominantly White British (77.5%, n = 39; 53% male) with a mean age of 12.42 years (SD = 1.89, range = 9.00–16.07). The Hong Kong sample was more ethnically diverse (n = 78; 45% male) than the U.K. sample (77.5% White British) in that children reported being Chinese (43.6%), White British (17.9%), and mixed race (17.9%) with a mean age of 12.38 years (SD = 2.00, range = 9.10–15.62). However, a t-test confirmed that the two samples did not differ in age, t(116) = 0.10, p = .92. In both samples, English was the primary language spoken at home (United Kingdom = 97.5%; Hong Kong = 61.5%), with Cantonese (16.7%) and Mandarin (9%) being the two other most spoken languages at home. In terms of family background, the children from the United Kingdom had significantly more siblings (M = 1.37, SD = 0.88), t(101) = 2.20, p < .05, d = 0.45, than the children from Hong Kong (M = 1.01, SD = 0.72). Parental education levels were higher in the Hong Kong sample than in the U.K. sample, χ2(1) = 9.89, p < .01, ϕ = .31. Specifically, 68 of 78 mothers in the Hong Kong sample had a university degree or higher, whereas 33 of 39 mothers in the U.K. sample had upper secondary school education or a university degree. In terms of socioeconomic status, 80.3% to 85.4% of U.K. and Hong Kong participants fell in the “affluent” band, respectively, as measured by the Family Affluence Scale (Boyce, Torsheim, Currie, & Zambon, 2006). Sample 2 consisted of 137 children taking part in the fifth wave of an ongoing longitudinal study of social and cognitive development in the United Kingdom (Hughes, 2011). We recruited a comparable sample of 125 children recruited from state primary schools in Hong Kong. Inclusion criteria for both sites were (a) no known developmental delays or disabilities and (b) native speaker of English/Cantonese (as appropriate). From this sample of 262 children, we created two groups of children individually matched on age (in months) and gender from the Hong Kong sample to children in the U.K. sample. In each site, there were 108 children (57% male). The U.K. sample was predominantly White British with a mean age of 10.81 years (SD = 0.39, range = 10.05–11.55). The Hong Kong sample was predominantly ethnic Chinese with a mean age of 10.81 years (SD = 0.35, range = 10.05–11.54); a t-test confirmed that the two samples did not differ in age, t(214) = −0.05, p = .96. In terms of family background and socioeconomic status, the children from the United Kingdom had significantly more siblings (M = 2.02, SD = 1.45), t(214) = 6.34, p < .001, d = 0.87, than the children from Hong Kong (M = 1.01, SD = 0.80). Parental education levels were also higher in the U.K. sample than in the Hong Kong sample, χ2(1) = 16.75, p < .001, ϕ = .28. Specifically, whereas 64 mothers in the U.K. sample had upper secondary school education (e.g., A levels) or a university degree, only 33 of the mothers in the Hong Kong sample had upper secondary school education or a university degree. In terms of self-rated affluence, children in the United Kingdom (M = 6.09, SD = 1.80) gave themselves higher ratings on the Family Affluence Scale than children in Hong Kong (M = 4.46, SD = 1.86), t(212) = 6.53, p < .001, d = 0.90. These between-sample differences in family background and socioeconomic status were controlled statistically in each of our models.","Table 1 summarizes the measures administered to Samples 1 and 2. Theory of mind Each theory-of-mind task involved watching a short film clip or listening to a short audio file (∼30 s in length) and then either explaining a character’s behavior or answering short questions about a character’s thoughts and feelings. The children from both samples completed the Triangles Task (Castelli, Happé, Frith, & Frith, 2000). In this task, the children watched three short silent cartoons depicting two triangles moving about against a white background. Designed to elicit mental state attributions, these cartoons featured instances of coaxing, teasing, and surprising. Using the coding scheme developed by Castelli et al. (2000), the children’s responses were scored for ascription of intentionality (i.e., the proclivity to ascribe mental states when explaining the actions of the triangles) and accuracy (i.e., how closely descriptions matched the intended story). Intentionality scores ranged from 0 (non-deliberate action) to 5 (deliberate actions with the goal of affecting others’ mental states). Accuracy was scored as 0 (incorrect descriptions), 1 (imprecise descriptions), or 2 (correct descriptions of the story presented in the clip). We created total scores by summing together intentionality and accuracy scores for each item (possible range = 0–21). The children in both samples also completed the Silent Film Task (Devine & Hughes, 2013). In this task, the children were required to explain the behavior of characters in five clips (∼30 s in length) from a classic silent comedy played once in a fixed order. The first clip was accompanied by two questions, and the remaining four clips were followed by one question. Following the coding scheme developed by Devine and Hughes (2013), participants received 2 points for responses that correctly described the events shown in the clip with reference to characters’ mental states, 1 point for answers that correctly described the events but did not use mental explanations, and 0 points for irrelevant or factually incorrect responses. Scores for individual items were summed together to create a total score (possible range = 0–12). In addition, the children in Sample 2 completed the Strange Stories Task. For this task, we used five vignettes (taken from Happé, 1994) that depicted social situations involving double bluffs, deception, and misunderstanding. After each vignette, the children were asked to explain a character’s behavior. The vignettes were translated into Cantonese by a panel of three English/Cantonese bilingual developmental psychologists that included the lead author, adopting a collaborative and iterative translation approach to ensure conceptual equivalence (Douglas & Craig, 2007). In each site, children listened to recordings (∼30 s in length) of an adult reading each of the stories. The text of each story remained on-screen after the audio recording stopped. Following the coding scheme developed by White, Hill, Happé, and Frith (2009), fully correct responses received 2 points, partially correct responses received 1 point, and incorrect responses received 0 points; across the five vignettes, therefore, children could score between 0 and 10 points. The validity of these tasks as measures of theory of mind is supported by three sources of evidence. First, these tasks have shown concurrent associations in a large sample of children ages 7 to 13 years even when individual differences in verbal ability and narrative comprehension were taken into account (Devine & Hughes, 2016). Second, performance on these tasks at age 10 years has also been shown to be linked to earlier performance on a battery of false belief understanding tasks at age 6 years (Devine, White, Ensor, & Hughes, submitted for publication). Third, performance on these tasks has been shown to be related to individual differences in children’s self-reported social competence (Devine & Hughes, 2013). For each task, the children’s responses were digitally recorded and later transcribed verbatim for coding. To ensure comparability of ratings, 25% of the transcripts in each site in both studies were double-coded. Individual items exhibited moderate to strong levels of reliability. In Sample 1, the two-way mixed single measure model type with absolute accuracy showed excellent average measure reliability for the Silent Film Task total score (intraclass correlation (ICC)[3, 1] = .96, 95% confidence interval (CI) [.89, .98], p < .001). The kappa values for all items were within acceptable ranges (κ = .73–1.00, SE = .00–.12). Similarly, the total intentionality (TTI) and appropriateness (TTA) ratings on the Triangles Task showed excellent average measures reliability (TTI: ICC[3, 1] = .96, 95% CI [.91, .98]; TTA: ICC[3, 1] = .90, 95% CI [.78, .96]; ps < .001), with item-level ICCs ranging from .76 to .92. In Sample 2, the second author trained the Hong Kong team in the rating procedures using transcripts from the U.K. sample as well as translated transcripts from the Hong Kong sample. For the Strange Stories Task, the mean κ was .92 and total summed scores showed excellent levels of interrater reliability (ICC = .97, 95% CI [.94, .99], p < .001). For the Triangles Task, the mean κ was .77 and the total summed scores showed excellent levels of interrater reliability (ICC = .97, 95% CI [.94, .99], p < .001). Finally, the individual items of the Silent Film Task showed acceptable levels of interrater reliability, the mean κ was .82, and summed total scores showed excellent interrater agreement (ICC = .94, 95% CI [.87, .97], p < .001). Executive function The children in both samples completed the Bead Memory Task from the Stanford–Binet Intelligence Scale (Thorndike, Hagen, & Sattler, 1986). On each trial, the participants were shown (for 5 s) a picture of an arrangement of beads on a stick (across trials, the beads varied in number, shape, color, and position). The participants were then given a box of beads and were asked to reproduce each bead arrangement exactly. The task was discontinued after three failures across four trials. Raw scores were calculated by subtracting the number of failed trials from the highest item attempted. Although initially designed to measure short-term visual memory, the Bead Memory Task also measures multiple executive functions. Specifically, participants must (a) refrain from touching the beads while the image is displayed, (b) hold the image in mind while they select the beads, and (c) plan ahead in order to place the beads on the stick in the correct sequence. Performance on this task is correlated with performance on measures of inhibitory control and set shifting during childhood (e.g., Hongwanishkul, Happaney, Lee, & Zelazo, 2005; Hughes & Ensor, 2011). The children in Sample 1 also completed the Digit Span Backward Test and the Trail Making Test. The Digit Span Backward Test (Wechsler Intelligence Scale for Children–Fourth Edition [WISC-IV]; Wechsler et al., 2004) is a widely used assessment of working memory. The examiner read out a list of numbers at approximately one digit per second, and the participants were asked to repeat this list of digits in reverse order. Each list was read out once and increased by one digit until the participants failed two consecutive lists of the same digit span. The correct responses were summed to create a total score (ranging from 0 to 16) and standardized scores. Scores for the current study ranged from 4 to 16. The Trail Making Test A and B (Corrigan & Hinkeldey, 1987) is a two-part test measuring switching and visual attention. The participants were timed on how quickly they could connect circles without lifting their pen off the page. This was first done sequentially on a page of 25 randomly scattered “numbered” circles on Form A (1-2- … 25) and then using an alternating order of “numbered–lettered” circles on Form B (1-A-2-B, etc.). Each form started with practice items, and mistakes were corrected and factored into total task completion time. The time difference between Form B and Form A provided an index of executive function, where a longer time (in seconds) reflected poorer performance. In this study, the time difference between forms ranged from 0 to 180 s. This was standardized, and a higher value reflected poorer performance. The children in Sample 2 completed the Arrows Task (Davidson et al., 2006), a measure of inhibitory control. In this task, participants viewed one of four images of a purple arrow on a white screen and needed to respond by hitting a key on the left (1) or right (0). The arrow pointed either directly downward or diagonally on either the left- or right-hand side of the screen. During control trials, the arrow appeared on either the left or right and pointed directly downward, and participants needed to press the button on the same side as where the arrow appeared. During test trials, the arrow appeared on either the left or right but pointed diagonally. On these trials, participants needed to press the button on the opposite side of where the arrow appeared. The participants received detailed instructions and practice items prior to completing the task. The arrows were displayed for up to 750 ms and were preceded by a 500-ms interval in which a crosshair was displayed against a white background. The children completed 12 control trials and 12 test trials in a random order. To measure inhibitory control, we calculated efficiency scores based on the total number of correct test trials divided by the total time taken on test trials. The children in Sample 2 also completed the Smiling Faces Task (Huizinga et al., 2006), a measure of cognitive flexibility. During this task, the children needed to respond with a button press to one of four cartoon faces (i.e., a happy boy, a happy girl, a sad boy, and a sad girl) displayed in one of four quadrants on a white screen. In single task trials, the children needed to identify whether the face was a boy or a girl if it appeared in the top two quadrants (by pressing 1 or 2) or to identify whether the face was happy or sad if it appeared in the bottom two quadrants (by pressing 3 or 4). During alternating trials, the children needed to integrate both rules. The trials were administered in four blocks in a fixed order: one set of 16 single task trials (boy or girl), another set of 16 single task trials (happy or sad), and two sets of 16 alternating trials. The trials within each block were presented in a random order. The participants were provided with detailed instructions and were asked to repeat back the rules of the task before each block. Each trial lasted up to 3500 ms, during which time the participants needed to respond with a button press. The trials were preceded by a 500-ms interval in which a black crosshair was displayed in the center of a white screen. We calculated efficiency scores based on the total number of correct alternating trials divided by the total time taken to complete the alternating trials. Verbal ability Unfortunately, there was no measure of verbal ability that had been standardized for use in both Hong Kong and the United Kingdom. The children in Sample 1 completed the Word Reasoning Test from the WISC-IV (Wechsler et al., 2004). In this task, the participants needed to identify the concept being described in a series of 24 clues of increasing difficulty. Each item was scored as either correct (1) or incorrect (0) before five consecutive “blanks” produced a total verbal ability score out of 24. We standardized raw summed scores (with a possible range of 0–24) within each country using T-scores so that each child’s verbal ability was measured against peers in his or her own country. For the children in Sample 2, we adopted the British Picture Vocabulary Scale (BPVS; Dunn & Dunn, 2007) because this is a widely used task that has very simple instructions. On each trial, the examiner read aloud a word and the children needed to point to one of four pictures that provided the best match for that word. We used the stimulus booklets and vocabulary lists from the U.K. version of this task. Adopting a back- translation approach (Brislin, 1970), each word was translated into Cantonese and then back-translated into English by a panel of three English/Cantonese bilingual developmental psychologists that included the lead author. The panel discussed and modified the Cantonese version to ensure that the two versions were equivalent in meaning. The Cantonese version was checked against “Lexical items with English explanations for fundamental Chinese learning in Hong Kong schools” provided by the Hong Kong Education Bureau (2009) to ensure progressive difficulty of the vocabulary. We calculated raw scores by subtracting the number of errors made by each participant from the item number corresponding to the last number in the participant’s ceiling set (i.e., the set containing eight or more errors). We then standardized scores within each country using T-scores so that each child’s verbal ability was measured against peers in his or her own country. Procedures The children in Sample 1 completed a 60-min individual testing session at school with the third author on two theory-of-mind tasks (Silent Film and Triangles), three executive function tasks (Beads, Digit Span Backward, and Trail Making), and a language test, among other tasks. All tasks were administered in English in both the United Kingdom and Hong Kong. At each site, the children in Sample 2 completed individual test sessions in a quiet environment at school. During these sessions (which lasted ∼90 min, including breaks), an experienced graduate researcher administered a battery of tasks presented in a counterbalanced order in either English (in the United Kingdom) or Cantonese (in Hong Kong). The U.K. team provided the Hong Kong team with detailed training and observed pilot sessions in Hong Kong to ensure that the data collection procedures were identical. Analytic approach ~~~~~~~~~~~~~~~~~ In analyzing the data from both samples, we used confirmatory factor analysis (CFA) to test measurement models designed to reduce the number of variables in our data and provide error-free parameter estimates. Next, we used multiple indicators, multiple causes (MIMIC) models (Brown, 2006) to examine cross-cultural differences in both the theory-of-mind and executive function latent factors. Given the presence of non-normally distributed variables in each data set, we estimated each model using robust maximum likelihood estimation (as opposed to maximum likelihood estimation) in Mplus 7 (Muthén & Muthén, 2012). The robust maximum likelihood estimator provides more accurate parameter estimates than standard maximum likelihood estimation (Brown, 2006; Kline, 2011). For each model, we evaluated fit using Brown’s (2006) four recommended criteria: a nonsignificant chi-square (χ2) test, comparative fit index (CFI) ⩾ .90, Tucker–Lewis index (TLI) ⩾ .90, and root mean square error of approximation (RMSEA) ⩽ .08. The effect sizes for the various model parameters were interpreted in accordance with recommendations from Kline (2011); small standardized effect sizes ranged from .10 to .30, moderate effect sizes ranged from .30 to .50, and large effect sizes were greater than .50. Descriptive statistics and data reduction ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 2 shows descriptive statistics (and intercorrelations) for each of the key variables for Sample 1. There was no evidence of ceiling or floor effects on either the Triangles Task or Silent Film Task. There were no differences between boys and girls in age, self- reported family affluence, or performance on any measures of verbal ability, theory of mind, or executive function (−1.36 ⩽ t ⩾ 1.33, all ps > .10). There were modest correlations between the two measures of theory of mind. The three measures of executive function were moderately intercorrelated. With the exception of the Trail Making Task and the Digit Span Backward Task, each of the variables conformed to normality (i.e., Zskew < ±3.29). Table 3 shows descriptive statistics (and intercorrelations) for each of the key variables for Sample 2. Each of the three measures of theory of mind captured a wide range of individual differences, as evidenced by the symmetrical distribution of scores on each task. There was no evidence of ceiling or floor effects on either the Triangles Task or Silent Film Task; no participants scored 0, and only 1% of the participants obtained the maximum score on either task. The Strange Stories Task showed some evidence of negative skew, with 12.5% of the participants achieving the maximum score. With the exception of the Arrows Task efficiency score, Strange Stories Task total score, and BPVS score, all of the key variables conformed to normality (i.e., Zskew < ±3.29). Data reduction ~~~~~~~~~~~~~~ In Sample 1, we tested a two latent factor model using CFA. First, we loaded both theory- of-mind indicators onto a single theory-of-mind latent factor. Next, we loaded each of the three executive function indicators onto a separate (but correlated) executive function latent factor. This measurement model provided a good fit to the data, χ2(4) = 1.60, p = .81, RMSEA = .00, CFI = 1.00, TLI = 1.00, Akaike information criterion (AIC) = 3463.10. The standardized factor loadings for the theory-of-mind latent factor were .75 and .38 (ps < .01). The standardized factor loadings for the executive function latent factor ranged from .63 to .70 (all ps < .001). The theory-of-mind latent factor was moderately correlated with the executive function latent factor (ϕ = .43, p = .01). In Sample 2, we specified a two latent factor model in which each of the three theory-of-mind task indicators loaded onto a theory-of-mind latent factor and each of the three executive function task indicators loaded onto an executive function latent factor. The two latent factor model fit the data well, χ2(8) = 13.64, p = .09, RMSEA = .05, CFI = .95, TLI = .91, AIC = 7146.43. Both latent factors accounted for significant variance in task performance. The three theory-of-mind indicators loaded significantly onto the theory-of-mind latent factor (mean loading = .50, range = .30–.69, all ps < .001). The three executive function indicators loaded significantly onto the executive function latent factor (mean loading = .50, range = .38–.67, all ps < .001). The two latent factors were strongly correlated (ϕ = .77, p < .001). Note that we chose to use a two latent factor solution for two reasons. First, from a conceptual perspective, theory of mind and executive function are related but distinct constructs. Meta-analytic evidence shows that individual differences in executive function and theory-of-mind task performance are only moderately correlated (Devine & Hughes, 2014). Moreover, findings suggest that the “real-life” correlates of executive function and theory of mind are distinct; for example, the relations between executive function and problem behaviors are significantly greater than those between theory of mind and problem behaviors in preschool children (Hughes & Ensor, 2008). Second, from an empirical standpoint, an alternative single latent factor model in both samples did not provide a good fit to the data (Sample 1: χ2(5) = 8.96, p = .11, RMSEA = .08, CFI = .94, TLI = .88, AIC = 3467.98; Sample 2: χ2(8) = 18.24, p = .03, RMSEA = .07, CFI = .92, TLI = .87, AIC = 7149.03). Modeling theory-of-mind task performance ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Using data from Sample 1, we examined whether there were cultural differences in theory- of-mind task performance using a structural equation in which the two latent factors of theory of mind and executive function were regressed onto a binary “nation” variable. To reduce the potential influence of confounding variables, we also regressed each latent factor onto three continuous variables: age, verbal ability, and self-reported family affluence (the latter of which is a socioeconomic status [SES] indicator). This model provided an excellent fit to the data, χ2(16) = 19.23, p = .26, RMSEA = .04, CFI = .97, TLI = .95, AIC = 3423.38. The parameter estimates for this model are presented in Fig. 1. When accounting for individual differences in executive function, age, verbal ability, and family affluence, there were no significant differences between the British children and the Hong Kong international school children in their performance on the theory-of-mind latent factor. In contrast, there was a weak but significant difference in performance between children in Hong Kong and children in the United Kingdom on the executive function latent factor, with children from Hong Kong outperforming their British counterparts by .21 standardized units. This difference was independent of individual differences in age, verbal ability, theory of mind, and family affluence. Next, using data from Sample 2, we specified a similar structural equation model to examine cultural differences in theory- of-mind task performance. To minimize potential confounds, we also regressed each latent factor onto age, gender, number of siblings, parental education, and verbal ability scores. Our initial model did not provide a good fit to the data, χ2(32) = 68.41, p < .001, RMSEA = .07, CFI = .86, TLI = .78, AIC = 6888.13. Inspection of the modification indices revealed evidence of differential item functioning in that nation exerted a direct effect on performance on both the Strange Stories Task and Triangles Task (see below for more details). The inclusion of these regression paths improved the overall model fit, χ2(30) = 37.13, p = .17, RMSEA = .03, CFI = .97, TLI = .95, AIC = 6860.85 (see Fig. 2). Together, these models accounted for 65% of the variance in theory-of-mind latent factor scores and 64% of the variance in executive function latent factor scores. Echoing findings from preschool samples (e.g., Sabbagh et al., 2006), scores on the theory-of-mind latent factor in Sample 2 were significantly lower for children from Hong Kong than for children from the United Kingdom. Even when potential differences in executive function, parental education, number of siblings, and verbal ability were controlled, children from Hong Kong were, on average, .60 standardized units lower than children from the United Kingdom, indicating a medium to large difference in performance (Brown, 2006; Cohen, 1988). The direct effect of nation on the Strange Stories Task and Triangles Task indicators (independent of individual differences in the theory-of-mind latent factor) showed that these indicators were not culturally invariant. That is, performance on these two tasks differed between both sites for reasons other than theory of mind. In sum, there were national differences favoring children in Hong Kong in aspects of these tasks that did not related to mental state reasoning. In direct contrast to the results for the theory-of-mind latent factor, children from Hong Kong outperformed children from the United Kingdom on the executive function latent factor. When effects of theory of mind, parental education, number of siblings, and verbal ability were controlled, children from Hong Kong obtained scores that were, on average, .37 standardized units higher than those for children from the United Kingdom. This is equivalent to a small difference in performance (Brown, 2006). There were no direct effects of nation on the indicators of executive function, suggesting that these items were culturally invariant. The structural equation model revealed moderate to strong correlations between theory of mind and executive function latent factor scores as well as between each of these latent factors and verbal ability. Parental education, SES, age, and gender exerted small but significant effects on executive function (but not on theory of mind). The number of siblings a child had was unrelated to either theory of mind or executive function.","In this article, we have presented data from two cross-cultural comparisons that jointly produced three main findings. First, when effects of general child and family characteristics were taken into account, theory-of-mind latent factor scores were equivalent for children in the United Kingdom and children attending international schools in Hong Kong, but children attending local schools in Hong Kong scored, on average, .60 standardized units lower than children from the United Kingdom, indicating a medium to large effect. Second, the contrast between children in Hong Kong attending local schools and those attending international schools was specific to theory of mind; both groups outperformed their British counterparts on the executive function latent factor with small effect sizes, .23 standardized units in Sample 1 and .37 standardized units in Sample 2. Third, in both samples, the latent factors for theory of mind and executive function were correlated with either a moderate or large effect size. Our motivation in conducting this research was to discover whether children in Hong Kong “catch up” in theory-of-mind performance during middle childhood. The results revealed no evidence for such a catch-up, at least for Hong Kong local school pupils. The findings extend earlier reports of young Hong Kong children’s delay on theory of mind (Liu et al., 2008) and demonstrate a persistent lag in Hong Kong children’s social understanding. Furthermore, this lag could not be explained by factors related to the global contrast between individualist and collectivist cultures because, unlike their peers from Hong Kong local schools, children from Hong Kong international schools performed just as well as their British counterparts on theory of mind. Instead, the results favor the pedagogical experience hypothesis, namely that children who were exposed to inquiry-based pedagogy (children in the United Kingdom and children attending Hong Kong international schools) showed better theory-of- mind performance than children who were exposed to the drill-and-practice pedagogy (Hong Kong local school children). What makes this conclusion more convincing is the fact that children attending international schools in Hong Kong were better on executive function than the U.K. children, just as children attending local schools in Hong Kong outperformed their British counterparts. In other words, Hong Kong international school pupils were culturally distinct from the U.K. children. Hong Kong international school students were more likely to be bilingual, a factor that is believed to facilitate theory-of-mind performance through enhanced attention control and inhibition rather than conceptual mental state understanding per se (Bialystok & Senman, 2004; Kovács, 2009). By including executive function in our models, we also controlled the potential confounding effect of bilingualism. Cantonese was the main mode of instruction in Hong Kong local schools, in contrast to English in Hong Kong international schools. Among other differences, the two languages differ in syntactical complement, which is hypothesized to affect theory-of-mind development (de Villiers & de Villiers, 2000). However, findings from a cross-lingual study (Cheung et al., 2004) suggest that syntax of complement per se does not contribute to theory-of-mind development; rather, it is the general language comprehension that matters. Our study contributes to the ongoing debate about the cultural universality versus specificity of theory-of-mind development by suggesting that it is not prudent to attribute cross-cultural differences in theory-of-mind development to a global individualist versus collectivist cultural distinction. Children’s direct social environments, their micro systems in Bronfenbrenner’s (1977) terms, might play a more important role in shaping the diverse paths of their theory-of-mind development. The pedagogical experience hypothesis offers a more specific mechanism accounting for individual differences in theory of mind and highlights the importance of the quality of education in social understanding development. However, further longitudinal and intervention studies are necessary to establish a causal relation between pedagogical experiences and theory-of-mind development. The strong correlation between theory of mind with low executive demands and executive function in both samples extends Devine and Hughes’s (2014) meta-analytic finding regarding the association between these two constructs during early childhood into middle childhood. The universality of the relation across an extended period in development and across cultures provides a basis for understanding the nature of the relation between these constructs. Hong Kong children’s advantage in executive function and concurrent disadvantage in theory of mind during middle childhood resonate with an earlier report of a similar inconsistency during early childhood (Sabbagh et al., 2006). This finding challenges a simple “expression” account and supports an “emergence” account of the relation between theory of mind and executive function. Viewed alongside the results from a meta-analysis of longitudinal data (Devine & Hughes, 2014), our cross-cultural data indicate that during middle childhood, just as during the preschool years, executive function facilitates (but is distinct from) theory of mind. Limitations ~~~~~~~~~~~ The two samples were originally recruited for different studies. Sample 1 was slightly older than Sample 2, and the children in the two samples completed slightly different tasks. Therefore, it is challenging to compare Hong Kong international school students with the local school students directly. Future studies are needed to explore the individual differences in mental state understanding among different populations within Hong Kong. Another potential limitation in our study is that there were socioeconomic status contrasts between the different samples. Compared with the Hong Kong pupils at local schools in Sample 2, the Hong Kong international school participants in Sample 1 were from more affluent families. Within Sample 2, the Hong Kong participants were from less affluent households with less educated parents when compared with their British counterparts. We statistically controlled for these variables in all of the models to account for these differences. It is worth noting that, with respect to parental education, Sample 2 was representative of the respective populations. In comparison with the 30% to 50% of employed U.K. adults who have higher education qualifications (Higher Education Statistics Agency, 2013), higher education in Hong Kong is available for only approximately 20% of the population (Hong Kong Education Bureau, n.d.). This contrast supports the ecological validity of any conclusions drawn from these samples.","This study is the first East versus West comparison of theory of mind and executive function during middle childhood. By expanding the age range to the much neglected period of middle childhood, our study adds valuable evidence on theory-of-mind use during this stage of development. Building on the well-documented delay in Hong Kong preschoolers’ theory-of-mind acquisition, our study demonstrated, for the first time, the similarities and differences in mental state understanding beyond early childhood in children in Hong Kong when compared with children in the United Kingdom. Specifically, the contrast between children in Hong Kong attending local schools and those attending international schools carries important educational implications. Our study found that children in Hong Kong attending local (but not international) schools showed a delay in theory-of-mind development compared with their British counterparts. In contrast, both groups of children from Hong Kong outperformed the British children on executive function measures. These results highlight the potential cost of drilling and rote learning (the dominant model in local Hong Kong schools) for children’s understanding of others. Our study also extends earlier reports on the association between theory of mind and executive function of preschoolers to middle childhood and demonstrates that, despite a clear advantage on executive function tasks, children in Hong Kong do not outperform their British counterparts on tests of theory of mind. This dissociation suggests an interesting contrast in the salience of social influences on executive function and theory of mind that deserves further examination in future studies."],["Aims To explore the utility of first-person viewpoint cameras at home, for recording mother and infant behaviour, and for reducing problems associated with participant reactivity, which represent a fundamental bias in observational research. Methods We compared footage recording the same play interactions from a traditional third-person point of view (3rd PC) and using cameras worn on headbands (first-person cameras [1st PCs]) to record first-person points of view of mother and infant simultaneously. In addition, we left the dyads alone with the 1st PCs for a number of days to record natural mother–child behaviour at home. Fifteen mothers with infants (3–12 months of age) provided a total of 14 h of footage at home alone with the 1st PCs. Results Codings of maternal behaviour from footage of the same scenario captured from 1st PCs and 3rd PCs showed high concordance (kappa >0.8). Footage captured by the 1st PCs also showed strong inter-rater reliability (kappa = 0.9). Data from 1st PCs during sessions recorded alone at home captured more ‘negative’ maternal behaviours per min than observations using 1st PCs whilst a researcher was present (mean difference = 0.90 (95% CI 0.5–1.2, p < 0.001 representing 1.5 SDs). Conclusion 1st PCs offer a number of practical advantages and can reliably record maternal and infant behaviour. This approach can also record a higher frequency of less socially desirable maternal behaviours. It is unclear whether this difference is due to lack of need of the presence of researcher or the increased duration of recordings. This finding is potentially important for research questions aiming to capture more ecologically valid behaviours and reduce demand characteristics. --------------------------------------------------------------------------------","Variations in mother–infant interactions have a substantial impact on offspring health and functioning in later life. Non-human animal studies have demonstrated stable and enduring changes in the brain as the outcome of variations in maternal behaviour, even in cross- fostering studies which eliminate the influence of genetic transmission (Francis, Diorio, Liu, & Meaney, 1999). A recent human study also demonstrates associations between variation in parenting within the normative ranges and infant brain development (Bernier, Calkins, & Bell, 2016). Experimental manipulations in human mothers further demonstrate the causal role of maternal behaviour on infant and child development. The still-face procedure (Cohn & Tronick, 1983), where the mother is instructed to behave temporarily in a disengaged manner (blank face and non-response) results in immediate infant distress. In addition, manipulation of contingency of maternal verbalisations (either responding to infant vocalisations within an appropriate time frame or not) leads to changes in infant vocalisations (Goldstein, Schwade, & Bornstein, 2009). Longitudinal studies, which have measured maternal behaviour, also highlight associations between variations in maternal behaviour and longer-term emotional, behavioural, and cognitive outcomes in children (Bornstein, Arterberry, & Lamb, 2014). However, many questions regarding the long-term impact of variations in maternal response remain unanswered. The first step is ecologically valid measurement of parental behaviour, and this is the focus of the current paper. Measurement of maternal behaviour, in large longitudinal studies or randomised control trials investigating parenting interventions, is essential to understanding parental behaviour and its effects. The accepted gold standard for measuring mother–child interactions is generally to have a researcher observe or film an interaction between mother and child in a clinical, research, or home setting and film from this third person point of view (3rd PC). There are several limitations to this approach, however: 1. Demand characteristics or reactivity Observation from a third, often unknown party (the researcher) is undeniably intrusive (Heisenberg, 1927). In observational study, the presence of a videographer may represent a kind of novelty that evokes atypical responses from those observed; this phenomenon is termed “reactivity” or ‘demand characteristics’. Observation may promote socially desirable or appropriate behaviours and suppress socially undesirable or inappropriate behaviours (e.g., adults may display higher rates of positive interactions with children; Baum, Forehand, & Zegiob, 1979; Zegiob, Arnold, & Forehand, 1975). This result may be differential according to different maternal characteristics; that is, some mothers may behave more positively, whilst others may become self-conscious and thus behave less positively (Weber & Cook, 1972). 2. False representation of the infant's experiences Generally, if maternal behaviour is coded from the viewpoint of an observer (3rd PC), what is coded is what the observer sees and not necessarily what the infant or mother experiences. From the point of view of developmental research, however, the ideal is to capture the infant's or mother's experience. For example, a mother smiling at her baby while the baby is looking at the floor differs from when the infant actually sees the smile. Whilst in both cases the intent may be the same and the smile is an act of warmth by the mother, from the infant's point of view the maternal behaviour is unlikely to influence the child if the child misses it. 3. Participant and researcher burden Due to demands on participant and researcher time, observations are usually of short duration and therefore only provide a snapshot of the mother–child relationship. Current study: first-person viewpoint ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ First person cameras (1st PC) are small portable cameras worn by the participant facing outward to capture the view point of the individual. A few recent studies have used 1st PCs, worn on head bands in an effort to capture the viewpoint of mothers and infants. Bornstein and Arterberry (2010), for example, used infant and mother worn 1st PC to record and compare the world from infant versus adult perspectives. Yurovsky, Smith, and Yu (2013) also used this technique to assess how infant's co-ordinate attention with a visual partner and highlight the different viewpoints of babies as compared to adults. For example, a child's view largely consists of a single dominating object compared to a mother's wider perspective. Sugden, Mohamed-Ali, and Moulson (2014) also used 1st PCs worn by infants for a number of hours at home, in order to measure the infant's exposure to adult faces Using 1st PCs offers many practical and research advantages that address the three shortcomings listed above. They: (1) eliminate the need for a researcher to be present, reducing potential influences of the researcher on mother and infant behaviour; (2) record the viewpoints of each interactant, so different perspectives are captured; and (3) diminish participant burden by removing the need to attend or host a research visit. So far, however, these studies have only measured the viewpoint of one-half of the dyad and therefore miss the combined footage. In the present paper, we explore and evaluate the gains and limits of using 1st PCs simultaneously worn both mother and infant. We evaluate how well 1st PCs capture relevant information via video and audio data-collection functions, we describe the reliability of existing coding systems that could be used on data captured by 1st PCs, and we explore 1st PCs ability to attain the 3 advantages described above. We first investigated whether two independent raters show reliability when coding behaviours from video footage from 1st PCs alone. This would mean that the recorded footage from 1st PCs alone is of adequate quality that it can be reliably coded as the same event by two different coders. We next explored whether the 1st PC reduced the role of participant reactivity (advantage number 1). We hypothesise that removing the researcher and allowing a longer duration will reduce reactivity and demand characteristics thereby reducing a fundamental bias in observational psychological research. We predict that, for mothers and infants left alone with the 1st PCs without the researcher present over a number of days, there will be a more negative in the types of behaviours recorded as compared to interactions recorded by the 1st PC but with a researcher present. For this comparison we keep the camera view point constant but vary the presence of a researcher. We specifically hypothesise that we will see a greater frequency of less socially desirable maternal behaviours, such as distracted and critical responses. We finally explored the potential advantage of recording from the infants’ view point (advantage number 2) by documenting the concordance of behavioural coding of footage from 1st PCs with footage from 3rd PCs, when recording the same interaction and varying only the camera view point. High concordance would mean that the majority of the same information is picked up by both viewpoints and is interpreted in the same way by coders. Low concordance could indicate that one of the recording methods picks up unique (and potentially important) information. We explored sources of differential concordance to understand information that is potentially lost or gained from using the different viewpoints.","Mothers with infants between 3 and 12 months of age were recruited using email advertisements within the University of Bristol School of Social and Community Medicine social life email list to Staff (academic and admin) and PhD students. Fifteen participants were recruited. Mean maternal age was 32.3 years (SD = 4.7), and mean age of infants was 8 months (SD = 2.4). All mothers were married or cohabiting and had high levels of education (at least one degree); all but one participant was Caucasian.","Infants were placed on a play mat with a selection of the same set of age-appropriate small, soft, and plastic toys, and mothers were asked to play with their infants as they would normally. Mothers and infants were filmed by a researcher, and both wore 1st PCs that recorded play during this time (see materials). The observation lasted 11 min. The researcher then instructed the mother how to record with and charge the 1st PCs, and asked mothers to use them during a variety of play times, meal times, and bed times in the coming days. The researcher left the 1st PCs with the mothers for an average of 1 week depending on the participant's availability. Participants were asked to record at least three 30-min sessions. Mothers were given a packet with instructions on how to use the 1st PCs, a session diary to record when they had filmed sessions, and a questionnaire about how they found using the 1st PCs (see Appendix). Finally, mothers were asked to complete a standard demographic questionnaire to provide information on highest educational level, occupation, age, number of children, and current feeding method. Materials We captured video and audio footage of mother–infant interactions using low-cost, head- worn cameras that have previously been used for recording infant's eye views of their environment (Sugden et al., 2014). The particular 1st PC we used was a ‘Bogdan Digital Spy Hidden Camera DVR Video Recorder’, at a cost of approximately £20 per camera, plus £5 for an SD storage card for each camera. The 1st PCs record video and audio in AVI format. Video has a resolution of 720 × 480 at 30 frames/s. No specification is provided for audio bitrate or sampling rate. 1st PCs store video on an SD card of up to 16 GB capacity or approximately 1 h of video and audio. The 1st PC has an internal Li-Ion battery, which gives approximately 2 h of battery life. 1st PCs are marketed as novelty spy cameras in the form of lapel badges and are yellow in colour with a black smiley face. To reduce the probability of the infant's gaze being unduly attracted to a brightly coloured 1st PC attached to the mother, we coloured them black. To attach the 1st PCs to the head of the mother and infant, they were sewn into elastic headbands. Footage from the two 1st PCs (mother and infant) was synchronised in time according to an early common event in both cameras and then joined using light works video editing software. The combined footage was then coded as below (see Fig. 1 for an example of combined footage). Coding ~~~~~~ Observations were coded using Noldus Observer XT software to categorise each maternal behaviour (event-based coding), using existing, published operational codes of basic maternal behaviour and infant affect (Leerkes, 2010). Mother codes were mutually exclusive, so at any one time point mothers can only be coded in one category. The codes are also exhaustive, so every maternal event is coded, and the duration of this code was automatically recorded. The coding categories were comforting, engagement, encouragement, positive affect, monitoring, routine care, distracted, critical, mismatched affect, persistent ineffective, and intrusive. Full description is available from Leerkes (2010), and we provide a brief description of each code here: Positive maternal behaviours ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Comforting: involves any maternal behaviour that will make the infant feel more at ease, such as hugging. Engagement: is when a mother directs her actions towards the infant (e.g. talking to the infant). Encouragement: consists of any maternal action that is associated with spurring the infant on. Positive affect: is the display of any positive emotion which is not part of one of the other codes (such as comforting), so for example the mother's comforting her child whilst smiling is coded as comforting, as the primary action relates to the objective to comfort; likewise if the positive emotion is primarily used to encourage, then it is coded as encourage. Neutral maternal behaviours ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Monitoring: this involves the mother watching the infant but not being actively involved. Routine care: this category encompasses acts such as cleaning the infant or adjusting clothes. Negative maternal behaviours ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Distracted: the mother is neither engaged with nor monitoring the infant, and may be looking away or involved in a separate activity. Critical: is any maternal behaviour that involves directing negativity towards the infant. Mismatched affect: is when the mother laughs or smiles when the infant is distressed, wary, and nervous, etc. (does not include attempts to distract or reassure the infant while engaging, supporting or calming). May also include contradicting or denying infant's emotional or behavioural reaction (e.g., “you’re not scared” or “that's not scary” or “it's funny” in matter-of-fact, firm tone if infant is distressed). The coding specifies that if mismatched affect occurs it should be coded rather than any other behaviour with which it potentially co-occurs. Persistent ineffective: this behaviour involves the mother repeating a behaviour despite the fact that her behaviour does not have a positive influence on the infant's emotional state. Intrusive behaviour: is when the mother's actions conflict with the infant's desired outcome, so for example preventing the infant from getting a toy reached for. In the coding manual clear guidance is provided regarding which code to prioritise if two codes co-occur and in which situations to score which code. We used event-based, continuous coding. So, the mother's behaviour would be coded into one of the categories described above The duration of this behaviour would then be recorded in the software until a new behaviour is recorded The video was first coded for maternal behaviours and then separately for infant affect (happy, neutral, or distressed) again with operationalised codes for infant emotion based on Leerkes (2010). Therefore, for each period of time, a code for maternal and infant affect is recorded. Coders were trained and supervised by RP, two coders were psychology final year placement students (AC and RL) and a final coder was a psychiatrist (KG). A series of training coding sessions to reach reliability of >80% on standard 3PCs was conducted before coding the headcam videos. Inter-method recording concordance on the same observation Inter-method reliability was calculated using the 1st PC and 3rd PC footage of the same play observation with the researcher present for all 14 dyads (thus 14 pairs of videos). For each the 1st PC and 3rd PC coding were paired, and overall reliability calculated (in The Observer XT 11). As we were interested in the amount of each behaviour as well as the order in which behaviours occurred, the duration/sequence calculation was used with a tolerance window of 2 s (Jansen, Wiertz, Meyer, & Noldus, 2003). For this comparison we used codings from the same coder to ensure that differences between coders did not account for differences between the coding from the different recording methods. As described below, we also ensured that 1st PCs were reliably coded by an independent coder who had not seen the 3rd PC footage. Inter-rater reliability of coding the 1st PCs Given the time intensive nature of coding, it is standard practice to only double code only a proportion (10–20%) of videos for reliability (Nicol-Harper, Harvey, & Stein, 2007; Zosuls et al., 2009). A random sample of 20% (11 of 56 1st PCs observations, as each participant gave multiple 1st PC observations) of all 1st PCs were also coded by an independent trained coder who was blind to the hypotheses of the study and who had not seen the 3rd PCs. Reliability between the same 1st PC codings from the two raters was compared in Observer using the same method as described for inter-method reliability. Comparing frequencies of maternal behaviours within participants but across recording methods Descriptive statistics were evaluated to provide summaries of coding parameters across video recording method. Due to the small sample and proof-of-concept nature of the study, we were not powered to conduct multiple comparisons. We therefore focus on key hypothesised comparisons and use paired t-tests to test specific comparisons between rates of less sensitive behaviours during interactions with and without a researcher present. Number of free sessions recorded ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Across the 15 mother–infant dyads, 56 free session videos were recorded for total of 14 h of these free sessions. The videos were a mean of 20 min and 12 s long, and each dyad conducted 1–5 sessions. All sessions were recorded during meal times or free play. Inter-method concordance ~~~~~~~~~~~~~~~~~~~~~~~~ Overall concordance was measured for 14 pairs of 1st PC and 3rd PC videos filming the same situation. The range for the index of concordance was 0.84–0.98, with an average of 0.90. Kappa values fell between 0.78 and 0.98, the mean being 0.89. Inter-rater reliability ~~~~~~~~~~~~~~~~~~~~~~~ Twenty percent (11) of the 1st PC videos were randomly selected and coded by independent researcher (who did not see the 3rd PC videos) to measure inter-rater reliability of the 1st PCs. The index of concordance from these calculations ranged from 0.78 to 0.98, with a mean of 0.91. The more conservative kappa values ranged from 0.75 to 0.97, with a mean of 0.90. A second analysis was conducted in which behaviour modifiers (intensity of the behaviour: as well as showing infant distress, which the coder qualified as mild, moderate, or intense) were included, resulting in a more granular measure of behaviours. Even in this case, the index of concordance ranged from 0.76 to 0.98, with a mean of 0.90, and the kappa values fell between 0.74 and 0.97, with a mean of 0.89. Descriptive investigation of non-concordant responses between 1st PC and 3rd PC ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Whilst the concordance of the footage from different viewpoints was high, it was not perfect and it is important to establish the nature of prominent discrepancies. Due to the time intensity of this in-depth analysis, a randomly selected sample 9 of the 14 pairs (64%) of 1st PC and 3rd PC videos were selected, and discrepancies between their codes (made by same coder originally) were examined by a third independent rater to understand the nature of information lost or gained from 1st PCs. A total of 178 events that were coded differently from the different view-points were found from pairs (as demonstrated from the high concordance reported above, this is a small proportion of all events), and researchers categorised the nature of differences according to whether the behaviour/action related to the code was visible from both viewpoints (3rd PC, and mother and infant 1st PC combined) and to what extent, as well as more descriptive analysis of the likely cause of different coding. Discrepancies consisted of differences in coding behaviours/actions viewed by both cameras (so likely a coding disagreement), actions/behaviours observed only on 1st PCs, and actions/behaviours observed only by 3rd PCs. We discuss the nature of these different discrepancies below. Approximately one-third (35%) of differences were behaviours/events that were identified on both cameras but coded differently. Thirty-five percent of these discrepancies were due to coding differences despite the action being clear from both viewpoints. However, the majority of miscoded actions potentially resulted from the different points of view offered by the 1st PC and 3rd PC. For example, the behaviours that were most often differently classified were ‘monitoring’ and ‘engagement’, as the level of engagement of mothers was potentially assessed differently when different viewpoints were taken. Monitoring was usually coded from the 3rd PC and not coded from the 1st PC because the whole body of the mother was visible to the 3rd PC and thus her posture and direction of gaze were much clearer from 3rd PC. Approximately one-half (48%) of the differences were actions that were picked up on the 3rd PCs, but missed on the 1st PCs. These actions were typically whole-body movements. For example, mothers picking up their infants, or infants flapping their arms were generally missed by the 1st PC. On one occasion, for example, a mother was silent on the 1st PC recording and no other changes in behaviour were visible. However, the 3rd PC recording showed that the mother was handing the infant a toy and maintaining eye contact. Behaviours that were picked up by the 1st PCs, but missed by the 3rd PCs accounted for about one-fifth (16%) of the discrepancies. Generally, the sound quality of 1st PCs was superior to that of the 3rd PCs as 1st PCs were closer to the mouths of the participants. Sounds such as whispering, which could not be heard on the 3rd PC were recorded by the 1st PC. Also, as 1st PCs were focused directly on participants’ faces, facial expressions that were often missed by 3rd PC were visible on 1st PCs. In the majority of cases in this initial work, 1st PCs were not optimally placed by participants, meaning that many facial expressions and other actions may not have been coded. Thus, the 16% is likely to increase with better positioning of cameras. As technology improves and cameras become smaller and less intrusive, more actions that would be missed by a 3rd PC should be easily identified on 1st PCs. Performance of the cameras ~~~~~~~~~~~~~~~~~~~~~~~~~~ Many aspects of the 1st PCs’ performance were perfectly adequate to the task. The video quality, while not high definition, can discern facial expressions, eye gaze, and general facial and body movements and responses. Likewise, the audio quality is good enough to record speech. Many shortcomings will be rectified with new and developing technologies. The novelty, low-cost nature of the device means that some aspects of their performance and functionality were sub-optimal. The field of view (which we estimate to be about 60°) for one interactant was often too narrow to fully capture the partner interactant. The usability of the device was poor, with operations (e.g., switching on and off, and starting and stopping recording) controlled by unintuitive combinations of buttons presses. Overall behaviours from different recording methods ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As can be seen in Table 1, the 3rd PC picked up more positive and neutral behaviours per min than the 1st PCs when recording the same interaction. The 3rd PC also picked up more maternal behaviours across the categories than the 1st PCs, when viewing the same interaction. However, the 1st PCs picked up more infant distress perhaps because of the superior audio and close facial recording of 1st PC as described above. Given that infant distress is most commonly expressed in crying vocalisations and clear facial expressions, 1st PCs may have greater potential to pick up these subtle displays. During the free sessions, 1st PCs picked up significantly more insensitive behaviours, such as intrusive or distracted maternal behaviour, per min than the 1st PCs while the researcher was present (paired t-test: t14 = 5.6, p < 0.001; mean difference = 0.90, 95% CI = 0.5–1.2). This difference represents an approximately 1.5 SD increase in such behaviours between the sessions with a researcher present and during the free sessions. Summary of key results ~~~~~~~~~~~~~~~~~~~~~~ 1st PCs recorded information that resulted in similar overall coding on a relatively simple mother–infant coding system to that recorded from 3rd PC. As predicted, once the researcher was not present, the videos contained less sensitive maternal behaviour. Although only a small pilot sample, even in this highly educated low-risk group the differences were of a potentially large size (approximately 1.5 SD higher rates per min of less sensitive behaviour). We do not know if this difference was due to the lack of researcher or different scenarios chosen by the mother when using the 1st PC, but either way, following the protocol here we observe a more negative maternal behaviours. We see fewer behaviours from the 1st PC compared to the 3rd PC, as indicated from the rate per min of maternal behaviours (see Table 1). Two possible explanations present themselves: The restricted camera angle means that 1st PCs miss maternal behaviours that are actually experienced by the infant or 3rd PCs record maternal behaviours not actually seen by the infant. Our descriptive analyses of the different codings suggest that the former may be more accurate. However, some behaviours (such as infant distress) were often picked up only by 1st PCs (see later). Advantages of 1st PC ~~~~~~~~~~~~~~~~~~~~ The first advantage relates to reduction in participant reactivity and demand characteristics. Even in this relatively homogenous sample of educated mothers, a substantial difference emerged in the amount of insensitive maternal behaviours over a number of interactions using 1st PCs at home alone than in a one-off researcher present setting, ether with 1st or 3rd PCs. Based on the assumption that mothers are more likely to display more behaviours that are socially considered ‘negative’ parenting when demand characteristics are reduced, the finding of increased negative parenting, provides preliminary evidence for reduced demand characteristics. Demand characteristics may be especially high in educated mothers due to knowledge of research and ‘best parenting’. This potential bias poses a real challenge to measurement of parents’ behaviour, in which consistent associations are made between maternal education levels and observed parenting (see Bornstein, 2015). Our results are consistent with a study demonstrating that a more extreme ‘negative’ maternal behaviour of corporal punishment is recorded more frequently by using passive audio recording in the home than reported frequency from mothers (Holden, Williamson, & Holland, 2014). There are, therefore, a number of ways in which 1st PCs could lead to reduced demand characteristics. It may because no researcher is physically present, or could result from longer duration of recording. In addition, while the mother is of course still aware that she is being recorded it may be easier to ignore or forget about a camera than the physical presence of a researcher. In addition infants are highly unlikely to be aware that their behaviours will be viewed by someone else while wearing the head cams, and this differs from situations when a researcher is present because infants may be aware of the presence of someone other than their mother, and thus ‘react’. In addition the 1st PC allows a longer duration for the mothers to relax and forget that they are being recorded. Briefer observations can be unstable, and observations lasting longer than an hour are likely to include samples of behaviours from highly varied activities or contexts, thus being more varied themselves (Miller, Shim, & Holden, 1998). It is important to point out that the advantage of reduced demand characteristics is not specific to 1st PC but the home alone setting and other methods (such as stationary cameras in houses) may also offer this advantage and result in more negative behaviours being captured. Secondly, 1st PCS were better able to capture subtle facial expressions and vocalisations. Assuming 1st PCs are placed at the optimal position (i.e., at the glabella between the participants’ eyebrows), there was also evidence of advantage 2 (view point of the infant). First, the ranges of behaviours that can be captured using a 1st PC differ to those visible on a 3rd PC. Subtle facial expressions and facial expressions that would otherwise be missed due to the 3rd PC camera angles can be recorded on 1st PCs (e.g., Baby FACS; Oster, 2005). Subtle sounds can also be recorded more intelligibly by 1st PCs because their cameras’ microphones are much closer to the face than are 3rd PCs microphones. The viewpoints offered by 1st PCs are unique and possibly more naturalistic, as the visual angle is that of the first person. Thirdly, the cameras used here have the advantage of reduced costs associated with researcher's time. 1st PCs are less demanding on research time than 3rd PCs as they can be left with participants to record themselves rather than hosting a researcher and associated costs of researcher time and travel. Challenges ~~~~~~~~~~ Although there are a number of advantages associated with the use of 1st PCs, there are also some challenges. One relates to placement. In practice, many participants did not place their 1st PCs completely optimally, meaning that many advantages of 1st PC were not fully achieved. Over the course of our pilot, we found that instructing mothers to check the position of 1st PCs in the mirror was key to optimally positioning the head camera. However, even when the camera is placed correctly, the visual field offered by current 1st PCs is restricted, and many whole-body movements are consequently missed. Also, to record facial expressions, two participants must wear the cameras and face one another. Evaluation of interventions Reducing demand characteristics and reactivity may be particularly important in evaluating parenting interventions. Many evaluations of parenting interventions involve filming mothers during free-play sessions, in the presence of a researcher, using a 3rd PC. However, these interventions explicitly focus on teaching parents to implement certain behaviours in certain situations. There is a possibility that mothers ‘perform’ these learnt behaviours when being filmed by members of the same study team who delivered the intervention (even if they are different researchers they may be seen as connected), and this circumstance is highly likely to introduce bias. Using the 1st PCs, researchers can help reduce this bias through removing the researchers’ presence and increasing duration of recordings which is key in accurate evaluation. Large cohort studies or resource limited studies Given the low cost and reduced researcher time needed, 1st Pcs may be useful for the large-scale recording of mother–infant behaviour in large epidemiological cohort studies which collect multiple measures and aim to reduce participant burden and save costs. Indeed, the 1st PC are currently being piloted in a large UK cohort ALSPAC, see https://proposals.epi.bristol.ac.uk/?q=node/113441. Studies focused on facial expressions Studies interested in facial expressions specifically may benefit from the use of head cameras which capture them well. Studies interested in infant view point Researchers interested in understanding aspects of the environment that infants and mothers focus on. Many of the abilities that are thought of as automatic in adults are learned early in life, meaning that infants are challenged by processes that are simple for more experienced individuals and therefore they will focus on different aspects of their environment. For example, the ability to conceptualise and categorise objects improves with experience (Oakes & Madole, 2003), which infants initially find difficult. This means that infants are more likely to focus on objects in their environment (Bornstein & Arterberry, 2010). Therefore, infants’ viewpoints, and consequently information processing, could differ greatly from that of their adult counterparts. With 1st PCs, researchers will have access to more accurate representations of both infant and adult viewpoints, leading to more informed inferences about infant experiences. Use in video-feedback interventions In the specific mother–infant context, another possible application of 1st PCs is facilitating video feedback interventions, which have previously been shown to be effective in improving maternal sensitivity in a range of populations, including adolescent mothers, mothers with schizophrenia, and mothers with poor attachment styles (Bakermans-Kranenburg, van, & Juffer, 2003; Cassibba et al., 2015 Cassibba, Castoro, Costantino, Sette, & Van Ijzendoorn, 2015; Kalinauskiene et al., 2009; Reddy et al., 2014). This therapeutic approach consists of video feedback intervention sessions in which mothers are recorded interacting with their infants in a naturalistic free-play session. These recordings are then discussed with a mental health professional, and mothers are encouraged to observe their own sensitive and insensitive behaviours, thereby improving their observational skills and empathy. Any sensitive behaviours displayed by the mother are supported and are used as examples to contrast with instances in which insensitive behaviours are displayed. This means that each mother acts as her own model. A key focus of video feedback is ‘mind-mindedness’, referring to a mother's ability to see the world from her baby's view point and thus respond to the baby's emotional needs (Meins, 1997). 1st PCs would be particularly useful in facilitating mind- mindedness as the infant's actual viewpoint would be captured, and mothers may be better able to understand their infant's world and their infant's perceptions of that world. In addition the mother can reflect on her own responses more easily when viewing the situation from her own view point. This process – known as “self- entheaty” or observing one's actions on record – may be particularly beneficial in a therapeutic sense because psychological and behavioural-change programmes rely heavily on introspection. As Lahlou (2011) brings to light, introspection is not easy, and when we do introspect we change the very cognitions we are hoping to describe. Subjective Evidence-Based Ethnography (SEBE; Lahlou, 2011) allows patients to view their 1st PC data after collection and engaging with the task, in an attempt to examine cognitions after the act. In this way, individuals can avoid the problem of modifying cognition with introspection and gain increased insight by experiencing something and learning independently, as opposed to didactically. Studies interested in global environment and whole body If more precise and whole-body movements are of key importance, 1st PCs may not be recommended because these behaviours were often missed by 1st PC. Rather cameras in rooms (but without researchers present) could be used alongside the 1st PC to capture whole body movements such as touching or assessing the positioning and proximity of the mother and child. Furthermore, 1st PCs can provide important information regarding the location of dyads within a room. In addition studies interested in multiple participants, such as those including fathers and siblings, may need to include the wider footage captured from 3rd PCs placed in homes. Future directions ~~~~~~~~~~~~~~~~~ There is considerable interest in on-body camera technology from a variety of domains. Police officers in a number of regions of the United Kingdom and United States are routinely being issued body cameras for recording their interactions with the public. Social care workers in some areas use on-body cameras for similar purposes, providing visual and audio records of meetings with service users. The market sector driving the biggest changes in on-body camera technology is the domestic desire for wearable action cameras. Manufacturers such as GoPro are bringing devices to market that are smaller, lighter, have wider fields of view, higher definition video, longer battery life, and increased flexibility in streaming and storing captured footage to other devices and networks. These improvements mean that a number of devices coming to market will be better suited to on-body video and audio recording of interactions in dyads for psychological research. Future studies of the kind reported here will undoubtedly benefit from these technological improvements, in particular, wider field of view, greater ease of use, longer battery life, greater storage capacity, and ease of video/audio data transfer.","The high level of concordance between 1st PC and 3rd PC videos of the same situation demonstrates that 1st PCs capture a situation reliably. Some elements of these situations, such as whole-body movements, are missed by 1st PCs because of the relatively small field of vision of the 1st PC. Despite the fact that the audio and visual quality of 1st PCs are at this point still sub-optimal, they capture subtle sounds and facial expressions that may be missed by 3rd PCs. As new technology emerges and cameras with better visual and audio quality and wider fields of vision come onto the market, the performance of 1st PCs should improve. Another potentially important (if preliminary) finding, associated uniquely with 1st PCs, is the increase in insensitive behaviours seen in mothers when they were left alone with 1st PCs. This result suggests that researchers are highly likely to capture an unrepresentative picture of maternal behaviour if researchers are present, even in the home setting. The other main advantages of using 1st PCs are their ease of use, low cost, and the relatively small participant and researcher burden. Future directions may include the use of 1st PCs in video feedback interventions and in developmental and psychological research."],["Empirical descriptions of the phenomenology of meditation states rely on practitioners’ ability to provide accurate information on their experience. We present a meditation training protocol that was designed to equip naive participants with a theoretical background and experiential knowledge that would enable them to share their experience. Subsequently, novices carried on with daily practice during several weeks before participating in experiments. Using a neurophenomenological experiment designed to explore two different meditation states (focused attention and open monitoring), we found that self-reported phenomenological ratings (i) were sensitive to meditation states, (ii) reflected meditation dose and fatigue effects, and (iii) correlated with behavioral measures (variability of response time). Each of these effects was better predicted by features of participants’ daily practice than by desirable responding. Our results provide evidence that novice practitioners can reliably report their experience along phenomenological dimensions and warrant the future investigation of this training protocol with a longitudinal design. --------------------------------------------------------------------------------","This article aims to describe a meditation training protocol developed in the context of an empirical brain imaging, cross-sectional study that investigates the mechanisms of mindfulness and compassion meditations. The novelty of this protocol is to obtain a meditation active control group by training healthy, naive participants to verbally express their subjective experience of meditation practice using a multidimensional phenomenological space (Lutz, Jha, Dunne, & Saron, 2015). Phenomenological space refers here to the description of features of the field of experience, as it is lived and verbally expressed in the first person (e.g., Husserl, 1991). This phenomenological matrix has been recently proposed as a framework to map different styles and levels of training in mindfulness, as well as heuristic tool to generate hypotheses for empirical research. The Brain & Mindfulness project attempts to practically apply this theoretical framework (for the study manual, see Abdoun, Zorn, Fucci, Perraud, Aarts, & Lutz, 2018). During the training participants were introduced to various styles of meditation practices and acquainted with phenomenological categories through various experiential exercises. These phenomenological dimensions were then investigated at neural, behavioral and physiological levels during the various cognitive and affective experimental paradigms. Such explicit use of first-person data to guide the analysis of third-person data is inspired by Francisco Varela’s research program of neurophenomenology (Lutz & Thompson, 2003; Varela, 1996). The current training protocol attempts to pragmatically tackle three methodological and conceptual challenges. The first one is concerned with issues regarding the definition of mindfulness meditation in psychology and cognitive neuroscience. The second one pertains to epistemological and methodological issues related to the integration of first- person reports in an experimental protocol. The third one is related to the quality of control groups for cross-sectional studies of meditation expertise. Theoretical context: mindfulness as a dimensional, phenomenological state ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In experimental and clinical psychology, the construct of mindfulness is generally used with three different meanings that refer either to: (a) a mental trait or a dispositional inclination (e.g. the Five Facet of Mindfulness proposed by Baer, Smith, Hopkins, Krietemeyer, & Toney, 2006), (b) a soteriological or spiritual path conceived in therapeutic and health-promotion terms (e.g. in the Mindfulness-Based-Stress-Reduction program; Kabat-Zinn, 1982), and (c) a single cognitive process trained and potentially brought to various human activities (e.g. like in “paying attention in a particular way: on purpose, in the present moment, and non-judgmentally”, Kabat-Zinn, 1994, p.4; or “the optimal interaction between attention and peripheral awareness”, Culadasa et al., 2015, p.30). While these meanings remain useful for many contexts, they are also problematic. Self-report questionnaires to study mindfulness as a trait lack specificity (Goldberg et al., 2016) and may even yield contradictory findings. For instance, Leigh, Bowen, and Marlatt (2005) found that binge drinkers’ mindfulness scores were higher than those of participants in a mindfulness retreat. In addition, findings may be biased by social desirability, consistency effects, or shared language between intervention instructions and scales (see Sauer et al., 2013, Van Dam, Hobkirk, Danoff-Burg, & Earleywine, 2012). Interpreting mindfulness as a soteriological process (meaning [b]) is often too broad to guide empirical research. Up to this point, discussions of mindfulness as a cognitive process (meaning [c]) make it difficult to account for differences in practice styles and levels of expertise, while also lacking the specificity required to formulate mechanistic hypotheses. Because these meanings are too restrictive, with C. Saron, A. Jha and J Dunne, we have argued against formulating a single, universally applicable consensus definition of mindfulness (Lutz et al., 2015). Instead we favor reconceiving mindfulness through a family resemblance approach whereby it can be conceptualized as “a variety of cognitive processes embedded in a complex postural, aspirational, and motivational context that contribute to states that resemble one another along well-defined phenomenological dimensions” (Lutz et al., 2015, p.633). This approach draws on previous efforts to conceptualize mindfulness (Chambers, Gullone, & Allen, 2009; Hölzel et al., 2011; Lutz, Slagter, Dunne, & Davidson, 2008) and the phenomenology of mindfulness practice. It is compatible with multiple explanatory and analytical frameworks from different subdisciplines, including contemplative theories, clinical frameworks and psychological and neuroscientific models. This approach is guided by a pragmatic inquiry: when one is formally practicing mindfulness, what observable and manipulable features of consciousness are most relevant to report in an experimental setting? We identified seven features proposed in a bipartite phenomenological model (detailed in Lutz et al., 2015 and resumed here in Table 1). The model assumes that these dimensions of experience are dynamic and manipulable in that they are affected—directly or indirectly—by different instructions of practice and/or by the level of expertise. This model was used to plot the hypothetical phenomenological characteristics of two styles of mindfulness, Focused attention (FA) and Open monitoring (OM) meditations, for both novice and expert meditators (Lutz et al., 2015). These plots have been created based on various instruction sets and descriptions. They should not be taken as actual plots of any individual’s phenomenology. The same set of mindfulness instructions could be mapped to different points in the phenomenological space. This is due to individual differences between practitioners in the manner in which they interpret and instantiate instructions. One aim of the Brain & Mindfulness project is to implement this heuristic model and to empirically test some of its assumptions. For instance, can we use self-report scales to reliably measure and monitor consistent changes in these features in response to different meditation instructions and training, congruent with the hypothetical plots previously published? Epistemological limitations: Reliability of self-report data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A second methodological aim arises from the first one: are the empirically-obtained plots of the different styles and levels of expertise reliable in this phenomenological space? The reliability of self-reports The perceived demise of early twentieth century introspectionism (Costall, 2006) and the seminal review by Nisbett and Wilson (1977) questioned the ability of the participants to report the real causes of their behavior. Since then, introspective-like methods have been looked upon with distrust by many in the fields of psychology and cognitive science. Contrastingly, others have warned against drawing general conclusions from these failures. For example, Hurlburt and Heavey (2001) have criticized how Nisbett and Wilson’s work has been carelessly taken as “an unconditional refutation of introspection in general, not merely of the attribution of causation”, thus ignoring that “even Nisbett and Wilson recognized the possibility of accurate reports about inner experience” (Hurlburt & Heavey, 2001, p.401). Devising ways to detect and/or limit the diverse types of self-reports distortions is an active field of methodological research. For example, self-administered questionnaires have long included validity scales designed to this effect (Baer, Rinaldo, & Berry, 2003). More recently, there has been a renewed interest for ‘first-person methods’ to study consciousness (see the three special issues of the Journal of Consciousness Studies on this question: Jack and Roepstorff, 2003, 2004; Hasenkamp & Thompson, 2013). First-person methods refer to methods that allow an investigator to bring a participant close to their subjective experience1 (Petitmengin, 2006), as well as to practices that subjects themselves can use to increase their sensitivity to their own experiences (Bitbol & Petitmengin, 2013; Depraz, Varela, & Vermersch, 2003; Petitmengin, Remillieux, Cahour, & Carter-Thomas, 2013; Varela & Shear, 1999). Meditation training has been proposed as a pragmatic response to this challenge due to it's disciplined approach to examining experience. Approaching experience from this perspective allows for the refinement of first person categories' repertoire and strengthen the robustness of the relationship between first and third-person data (Varela, 1996). However, this hypothesis remains to be thoroughly tested. Current available evidence includes the improvement of the congruence between implicit and explicit measures of self-views after brief mindfulness exercises (see Strick & Papies, 2017 for a study on affiliation motives and goals, and Koole, Govorun, Cheng, & Gallucci, 2009 for a study on self-esteem). In contrast, measures of interoceptive awareness based on heartbeat perception in experienced meditators have yielded mixed and contradictory results (Bornemann & Singer, 2017; Khalsa et al., 2008; Melloni et al., 2013). The inconclusiveness of these studies may be due to a lack of methodological validity (Zamariola, Maurage, Luminet, & Corneille, 2018), discrepancies in the experimental designs and/or in the extent of bodily focus in participants’ meditation practice. Demand characteristics and desirable responding In the context of phenomenological research on self-induced mental states (such as in meditation research), demand characteristics is a major source of confound that undermines the credibility of self-reports. Demand characteristics refer to “the totality of cues which convey an experimental hypothesis to the subject[s]” and which consequently “become significant determinants of subjects' behavior” (Orne, 1962, p.779).","volunteering for scientific experiments have various motivations that may, consciously or unconsciously, incite them to play the role of the good participant and try to serve the experiment by producing the data that they think will confirm the (presumed) research hypothesis. To attenuate the confounding effects of demand characteristics, researchers commonly resort to the concealment of – if not the deception about – hypotheses, manipulations, dependent measures and independent variables. Another source of distortion of a participant’s behavior is his/her wish to present herself favorably to the experimenter, who may be perceived as an evaluator. This so-called social desirability bias is related to the effect of demand characteristics, but not identical to it (Weber & Cook, 1972). To eliminate this confound, some researchers advocate the use of scales developed to capture individuals’ inclination to self- enhancement (Crowne & Marlowe, 1960; Paulhus, 1984), as covariates in the models assessing the effects of interest. In the phenomenological study of meditation, demand characteristics lurk in the large overlap between the semantics of meditation instructions taught or familiar to the participants, and the phrasing of self-report scales aimed at measuring the phenomenological dimensions of interest (e.g. terms such as present-centered and nonjudgmental, see Van Dam et al., 2012). Consequently, the magnitude of self-reported phenomenological features of meditation remain overshadowed by doubt, even when shown to be highly specific (see for example Kok & Singer, 2017). Unfortunately, the usual concealment strategies can hardly be applied in this context, considering that participants are necessarily aware of the manipulation in so far as they are asked to implement it through the practice of meditation. This is not to say that all self-report results from meditation studies are inexorably confounded by the effect of demand characteristics. Even when demand characteristics are difficult to attenuate, one can look for evidence that supports an interpretation of the effect beyond their impact. In this study, we adopt a strategy of comparing certain factors (which are unaffected by demand characteristics or desirable responding) with self-reporting effects. We will illustrate this general approach in two ways: (i) within-subject, by testing whether fluctuations of phenomenological self-reports correlate with relevant behavioral measures, and (ii) across subjects, by testing whether participants’ amount and structure of daily practice predict the self-reported effects on phenomenological dimensions. Methodological issue: quality active control group ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A critical effort of the Brain & Mindfulness project was to refine the matching between the control group and expert meditators. This was done primarily by training novices in different styles of meditation practice and by familiarizing them to different phenomenological categories of interest. A prominent issue in the field of neuroscientific studies on mindfulness meditation is the relative paucity of high quality active control (Goldberg et al., 2017). This issue has been repeatedly raised and suggestions of improvement have been discussed in the context of longitudinal studies (Davidson & Kaszniak, 2015; Kuyken et al., 2016). However, cross-sectional studies with long term practitioners have received less methodological attention. In such studies, an active control group is often lacking or too basic when present. Of the nineteen independent cross-sectional studies on the neurofunctional effects of long term meditation in a recent meta-analysis (Fox et al., 2016), only seven included an active control.2 In six of these studies, meditation-naive participants received written and/or oral instructions and were encouraged to sustain a daily practice for 7 to 10 days until the moment of the experiment. In the remaining study, participants received a brief training session by an experienced teacher on the day of the experiment (Kalyani et al., 2011). These approaches, while clearly better than not including an active control group, have several limitations. First, limited possibility for feedback or the lack of guidance by an experienced and qualified instructor induces a high risk of misinterpretations and inadequate implementations of the practices. Second, the short duration of the training limits the opportunities to engage with the practice. Here we addresses some of these issues by: (1) formally training meditation-naive participants in practices from the same meditation background as the long-term practitioners, (2) letting this training be provided by a qualified instructor in a context with attention to sufficient opportunity for guided practice and feedback, and (3) encouraging participants to maintain a daily practice at home for a minimum of 20 min a day for an extended duration (6–22 weeks depending on the availability of participants and experimental resources). We contend that these adaptations make it more likely for participants to reach a refined understanding of the various practices and experiential dimensions at hand. This training has the potential of reducing the risk that any group differences are merely driven by confounding factors (e.g. misunderstanding practices and/or unfamiliarity with meditation terminology for novices but not experts), instead of reflecting the true effect of interest, i.e. meditation expertise. We will first report the specifics of our meditation training protocol. Then we will provide empirical evidence for its effectiveness in teaching meditation-naive participants to use first-person categories to describe their conscious experience and discriminate between phenomenological dimensions. Participants ~~~~~~~~~~~~ The first stage of the research included a meditation training weekend comprising of 42 healthy participants naive to meditation. These individuals were recruited for their interest to learn meditation and their willingness to sustain a regular practice for several months. After a preliminary inclusion procedure (see study manual for details, Abdoun et al., 2018), participants were invited to attend a weekend-long training program in the Lyon Neuroscience Research Center. The program looked to support participants in developing a refined understanding of the states of consciousness involved in the following experimental study. The expert group was comprised of 30 healthy long-term practitioners with more than 10,000 h of formal meditation in their life and trained in the Kagyü and Nyingma schools of Tibetan Buddhism. Both novices and experts participated in up to 8 experimental sessions (see study manual for details, Abdoun et al., 2018). For expert participants, these experimental sessions were gathered in 2 visits of 3 days, or a single visit of 6 days. For novice participants, each visit comprised 1 or 2 experimental sessions, and the visits were spread over a period spanning from 2 to 23 weeks after the training weekend. Visits were scheduled according to participants’ and equipment (MRI, MEG, EEG) availabilities, leading to a large but quasi-random variability across the novice group in time elapsed between the training and the experiment. Among all participants, 25 trained novice practitioners and 25 expert practitioners participated in the MEG experiment described below. The remaining participants included in the larger study were excluded for the MEG experiment because of excessive signal artifacts caused by dental prostheses. The novice and expert groups for the MEG experiment did not differ in gender (16 and 15 males, respectively; χ2(1) = 0.08), age (53.9 ± 7.1 and 51.6 ± 8.0 years, respectively; independent t-tests t = 0.84) and education (3.88 ± 2.15 and 3.20 ± 2.16 years of higher education, respectively; independent t-tests t = 0.46). General outline The training protocol was based on Joy of Living (Rinpoche & Swanson, 2007; Tergar, 2018), a secular meditation program aimed at Western audiences authored by Yongey Mingyur Rinpoche, a renowned master of Karma Kagyü and Nyingma schools of Tibetan Buddhism. This program was selected for its shared background with experts’ training. In its original format, the program is divided into three stages, each lasting two days; in addition, there are minimal practice requirements to attend stages 2 and 3. For our training protocol, we drew from the material of stages 1 and 2, condensing them in a two-day format, and included adaptations to emphasize the specific dimensions of experience of interest to the research program. The training was provided by a qualified instructor with thirteen years of practice under the guidance of Mingyur Rinpoche, and eight years of teaching experience with the Joy of Living program. The training included teachings with the support of instruction videos, guided meditations and experiential exercises, question and answer sessions, as well as sufficient time to reflect and share within the group. The training allowed a basic understanding of a few selected phenomenological dimensions eligible for an active comparison with expert practitioners. In particular, the program introduced participants to the following dimensions: effort, aperture, absorption vs. meditative awareness, foreground vs. background awareness, equanimity, clarity (see Table 1 and supplementary materials). The discernment of these dimensions was implemented by introducing the lived phenomenology of these states and creating occasions for a direct exploration of them. To access both the experience and the meaning of meditation the teacher devised specific exercises with connected theoretical principles. As a sommelier apprentice does in tasting, savoring, comparing and sampling different wines under the guidance of a sommelier, participants were invited to learn, practice and distinguish few states of consciousness under the guidance of a meditation teacher in order to become progressively familiar with some meditative phenomenological dimensions commonly described in contemplative traditions. The training followed a specific day program (Table 2) which will be briefly described here. In day 1, participants were first introduced to the notion of mental ‘effort’ in meditation through an experiential exercise that involved listening to sounds. The rest of the day was spent exploring the concepts of ‘absorption and meditative awareness’. This exercise was done first by using the breath as an anchor for meditative awareness. Participants were asked to restrict their attention to the breath, to notice when their mind had wandered, and to return their attention to the breath when this happened. Later during the day, they were also asked to gradually explore more open forms of awareness, by opening up to sense experiences from the environment (e.g. sounds and vision). While doing so, participants also engaged in two other experiential exercises that introduced the concepts of ‘object orientation and aperture’ and ‘foreground and background awareness’. At the beginning of day 2, participants continued to explore meditative awareness of the environment with various exercises including a walking meditation in open awareness. Then the instructor asked participants to form small groups and share their personal experience of the weekend. After some time, the small groups gathered to share collectively the problems and difficulties which had emerged, so that the teacher could provide adequate feedback. Then the concepts of ‘empathy and compassion’ were introduced and the teacher asked participants to briefly cultivate feelings of self-compassion. After lunch, participants engaged in an exercise that involved switching between focused attention on, and open monitoring of, pain. The rest of the afternoon was spent further elucidating concepts of empathy and compassion, including an experiential exercise that presented images of people’s suffering to participants. Finally, during the closing meditation session, the importance of the intention to practice was discussed and emphasized. Participants were asked to fully engage in their own practice. Experiential exercises Throughout the training weekend, subjects were prompted to familiarize themselves with the dimensions of subjective experience that were going to be of interest in the neuroscientific experiments. This familiarization was carried out by using experiential exercises. During each exercise, a dimension or process was introduced in a more or less explicit form. Some dimensions were experienced and described in the context of guided meditation sessions and teachings, while other exercises were implemented with the specific aim of familiarizing subjects with a phenomenological dimension. We refer the reader to the Supplementary Materials for a full description of these exercises. At the end of the weekend, subjects received a document that briefly described each phenomenological dimension and reminded them how it was introduced by corresponding exercises during the weekend. Compliance and engagement with practice The minimal goal of the training weekend was to give novice participants sufficient understanding and confidence to carry on practice autonomously, thus deepening their familiarity with the practices under study. At the end of the weekend, participants received an explanation of what was expected of them in terms of homework practice. Participants were advised to practice for 20–30 min on a daily basis and to give equal importance to each of the three practices they had learned. In addition, they were asked to report the type and amount of practice in a practice logbook that was provided at the end of the meditation training weekend. In order to ensure truthful reports, participants were assured that non- compliance to these recommendations would not call into question their participation to the study. Participants were provided with three 15 min-long audio recordings of guided meditations by their instructor to aid their practice (one recording for each of the three meditative practices). However, they were strongly encouraged to avoid relying exclusively on them and to get used to meditating unguided. Participants were also given excerpts from the book Joy of Living, summarizing most of the teachings received during the weekend: the physical posture, the mental attitude, as well as various meditative and experiential exercises examined during the weekend (Rinpoche & Swanson, 2007). Practice metrics ~~~~~~~~~~~~~~~~ Four metrics were used to assess participants’ home practice and degree of engagement: the proportion of days involving practice (hereafter referred to as Frequency of practice or simply Frequency); the average amount of practice during days with practice (hereafter, Session length); the daily average amount of practice, all days included (Intensity of practice or simply Intensity); and the total amount of practice (Experience).3 Fig. 1 describes how these four metrics relate to each other and how they were derived from the data contained in the practice logbooks. These metrics were also explored in relation to phenomenological ratings and behavioral measures from the experiments. Protocol of the MEG experiment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ One major purpose of the meditation course undertaken by novice participants was to train them in using phenomenological categories of interest in the study. In order to validate that they understood these categories as intended and used them appropriately, we analyzed the self-report data from a magnetoencephalographic (MEG) experiment with a hierarchical repeated measure design that included several periods of FA and OM meditations, along with a control (resting-state, RS) period (Fig. 2). The experiment started with a staircase visual threshold calibration for a visual task whose detailed procedure and results are irrelevant for the present article but will be described in another publication. Following the calibration, we recorded a 7 min baseline period. Then, participants practiced FA and OM twice each, in sequences of approximately 24 min (Fig. 2a). Each sequence opened with a 7 min-long block of meditation only, during which participants were presented with a 1.5°-wide white dot in the middle of a black screen (Fig. 2b). The instruction was to keep the gaze steady on the white dot; in FA, participants were additionally instructed to use that disk as a support for the attention. This first block was followed by three blocks lasting approximately 5.5 min each, during which participants were instructed to maintain the meditative state while performing a simple visual conscious report task using a threshold stimulus embedded in a passive visual oddball paradigm (Fig. 2c). After each block, participants were invited to rate their experience over 6 different dimensions, using a 7-point Likert-type item (Fig. 2d): Capacity to apply the meditation instructions, Stability of the mind, Clarity of the mind, Aperture of the field of awareness (see table 1 for definitions), Awareness of bodily sensations, and Wakefulness. Here we will limit our analysis to the dimensions featured in the phenomenological matrix: Stability, Clarity and Aperture.5 Rating scales were thus introduced: “Compared to your usual experience, how would you rate the last block in terms of Stability/Clarity/Aperture?”","Statistical modeling and inference. ANOVAs were of type 2. Post hoc tests were performed using one- or two-sample t-tests, adjusted for family-wise multiple comparison using Tukey’s honestly significant difference (HSD) method. One-sample and paired two-sample tests were performed using the non-parametric Wilcoxon signed rank test. Linear mixed models were fitted using maximum likelihood and significance of fixed effects were evaluated using the likelihood ratio test. All linear regressions were ordinary least square (OLS) regressions. Correlation between scale ratings and variability of response times. Outlier trials were defined as trials for which response times were not within 3 standard deviations from the mean value, for each participant, state and response type (yes/no), and excluded. An index of RT variability was defined for each block as the standard deviation of RTs. Finally, we computed for each participant the Pearson correlation coefficient between Stability ratings and RT variability. We did the same with Clarity. Correlation coefficients were z-transformed for the purpose of statistical modeling and are noted z hereafter. The data from one subject was removed because it had no variance in the Clarity scale (the subject responded 6 in all blocks). Model selection for multiple regression analyses. For each of the effects related to phenomenological rating scores, we considered several potential predictors: metrics of home practice (Intensity, Experience and the balance between focus and open styles of practice), and an index of desirable responding (the score to the Balanced Inventory of Desirable Responding, BIDR). Practice data was missing for one participant, who was therefore excluded from subsequent analyses. We used an information-based model selection to determine the most important predictors for our data (Burnham & Anderson, 2002). Model selection is well suited to multiple regression analysis when the number of predictors is high compared to the number of data points; in addition, it virtually guarantees that no potential effect of interest is missed, as long as it is included in the variable set. We performed model selection in 2 steps, using glmulti for R (Calcagno & de Mazancourt, 2010). Firstly, we fitted all possible models that contained a subset of the predictors mentioned above and their two-way interactions, and that satisfied the marginality constraint (i.e. included all interaction terms as main effects). We used the corrected Akaike information criteria (AICc) as a measure of the quality of fit, because it is well adapted to small sample sizes (Hurvich & Tsai, 1989). Secondly, we selected models that were within 2 information criteria (IC) units of the best fitting model (Burnham & Anderson, 2002) for further consideration. Detailed results of the model selection output are presented in the Supplementary Materials. These include the relative evidence weight, a measure of relative importance of each term across the entire model space (Calcagno & de Mazancourt, 2010), comprised between 0 and 1. Structure of home practice ~~~~~~~~~~~~~~~~~~~~~~~~~~ The total duration of participation to the entire study ranged from 41 to 163 days (99.4 ± 31.4 days) after the training weekend. Average daily practice (=Intensity) ranged from 1.3 to 30.5 min (15.9 ± 7.3 min; supplementary Fig. 1a, top left), suggesting that many participants fell short of the recommended amount of practice (20 to 30 min a day). However, when average daily practice was calculated over the number of days with at least some practice (rather than all days), the obtained Session length was found to range from 14.0 to 33.1 min (21.2 ± 5.5 min; supplementary Fig. 1a, bottom right). Participants dedicated 45.2 ± 16.8% of their practice time to OM (Open Monitoring), 33.4 ± 17.5% to FA (Focused Attention) and 21.4 ± 9.8% to CO (Compassion). A one-way repeated measure ANOVA revealed a significant difference between the proportion of time dedicated to the different practices (F(2,80) = 16.95, p < .0001, η2 = 0.30). Post hoc paired t-tests revealed that all comparisons of pairs of practices were significant (supplementary Fig. 1b). Intensity of practice decreased linearly over weeks (supplementary Fig. 2a; R2adj = 0.87, p < .001, β = -0.44, 95% CI [−0.53, −0.35]). A large portion of this drop (80.9%) is imputable to a sharp decline of Frequency of practice over weeks (supplementary Fig. 2b; R2adj = 0.92, p < .001, β = −0.14, 95% CI [−0.16, −0.12]). The remaining 19.1% is due to a slight shortening of practice sessions (supplementary Fig. 2c; R2adj = 0.22, β = −0.10, 95% CI [−0.18, −0.01]). Taken together, these results show that participants managed to follow the recommendation of 20-to-30 min-long practice sessions throughout the study, but failed to practice on a daily basis and were increasingly inclined to skip days. Patterns in self-reports and relationships to practice ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We tested three predictions that should be verified if the phenomenological self-reports are reliable. We examined whether the responses of the novices to the rating scales in the MEG experiment (i) were sensitive to the meditation state, in a way consistent with the known phenomenology of mindfulness practices (Lutz et al., 2015), (ii) exhibited classic temporal dynamics such as dose and fatigue effects, and (iii) were functionally informative, as would be suggested by correlations with behavioral measures. In each instance, we tested whether desirable responding and/or features of participants’ home practice predicted the effects observed on self-reports. The results are summarized in Table 3. Effect of states on self-reports In the phenomenological model introduced in Lutz et al., 2015, Stability and Clarity are described as secondary qualities that are both increased when practicing either FA and OM (compared to mind-wandering), and even more so with expertise. In contrast, Aperture is hypothesized to increase specifically during the practice of OM. Predictors of the state effect on Aperture When participants are asked to rate the broadness of their attentional scope (i.e. Aperture) in two different meditation states referred to as “focus attention” and “open presence”, the fact that the expected response lies in the very names of the state conditions can hardly remain unnoticed. Thus, we cannot exclude the possibility that participants’ responses were influenced by their willingness to please the experimenter, and/or to show that they have correctly understood the meditation instructions. Another, non-trivial hypothesis is that novice participants develop the ability to differentiate a state of focus attention and a state of broader awareness by getting equally familiar with both attentional styles. Said otherwise, we expect participants who have had more practice in OM than in FA to have more ease opening (or more difficulty narrowing) their attentional scope than participants who have developed an equal familiarity of the two styles of practice. As a result, we would expect the latter to better differentiate FA and OM on the Aperture scale than the former. The same reasoning can be straightforwardly applied, mutatis mutandis, to participants who favored FA over OM. We explored the plausibility of this hypothesis by testing whether our data were coherent with the ensuing predictions. We included an index of balance between FA and OM (FOB), along with other practice metrics (Intensity and Experience) and an index of desirable responding (BIDR) in the set of variables tested for model selection (see Section Methods, Statistical analyses). Three models survived the model selection procedure: the best one included a significant FOB-by-Experience interaction (model A1), while the other two contained a significant FOB-by-Intensity interaction (models A2 and A3; see details in supplementary Table 1). BIDR was not present in any of these models, and its evidence across all models was found relatively low (0.31). The FOB-by-Experience interaction had a higher evidence across model space than the FOB-by-Intensity interaction (0.55 and 0.31, respectively). To illustrate how the balance between focus and open styles of practice interacts with the amount of practice, we performed a Johnson-Neyman post-hoc analysis of the interaction in model A1, using FOB as a predictor and Experience as a moderator. We found that for participants who had accumulated more than 23.9 h of practice, the FOB index positively predicts (p < .05) the self-reported difference in Aperture between FA and OM during the MEG experiment (see supplementary Fig. 3). Considering that all our participants except one practiced OM more than FA at home, a higher FOB could have been entirely driven by FA practice in our dataset. Thus, the FOB-by-Experience interaction could actually hide an effect of Experience in FA only, which would lead to a very different interpretation of our results6. To test this alternative explanation, we fitted a model with Experience in FA as a single regressor, as well as a model that also included Intensity as a regressor. None of them was found significant (p > .98 and p > .16 respectively). We repeated this analysis with Experience in OM instead of FA – again, neither model was significant (p > .84 and p > .22, respectively). To summarize, the results of our model selection analysis and post hoc tests are consistent with the hypothesis that equal familiarity with focus and open attentional states, rather than experience in any specific practice, drives the ability to differentiate between states. Temporal dynamics The absence of effect on self-reported Stability and Clarity in novices over the experiment does not necessarily rule out the possibility that novice participants used these categories appropriately and informatively. For example, averaging ratings over the entire experiment could have occluded temporal effects. This is indeed what we have observed in our data (supplementary Fig. 4). We used linear mixed models to account for the nested nature of the MEG experiment structure (blocks within sequences within sessions). The models included all possible interactions in fixed effects, as well as random subjects’ intercepts and by- subject random slopes across blocks, sequences and sessions. We used two models: one for Stability ratings and one for Clarity ratings. In both models, we found an effect of block on ratings only for the first sequence of each session (i.e. sequences 1 and 3; cf. Fig. 2; Stability: χ2(3) = 8.58, p = .035; Clarity: χ2(3) = 7.85, p = .049), suggesting that the middle break had some sort of resetting effect. Post-hoc pairwise t-tests revealed a significant decrease of ratings from blocks 1 to 3 in sequences 3 and 4 (Stability: Δ = 0.62, 95% CI [0.19, 1.05], p < .002; Clarity: Δ = 0.43, 95% CI [0.09, 0.77], p < .007; all other p > .1), but no pairwise differences in other sequences (all p > .5). Predictors of the fatigue effect The decrease of self-reported Stability and Clarity in novices after four blocks (∼24 min) of meditation may reflect fatigue. This is not surprising considering that most novices were not used to meditating for more than approximately 20 min during their daily home practice (see supplementary Fig. 1a). A corollary of this interpretation is that the longer and more frequently the participants were used to meditate, the less likely they should be to experience fatigue in the context of the experiment. We tested this prediction by modeling a fatigue index, defined for each participant as the difference between their ratings in block 3 and block 1 of sequence 1,7 averaged over the dimensions of Stability and Clarity. Five models were selected (supplementary Table 2). Intensity was present in 2 of them as a main effect, and in 2 others in interaction with BIDR. Surprisingly, FOB was present as a main effect in 4 out of the 5 selected models. Across all models, FOB and Intensity had the highest relative weighted evidence (0.74 and 0.66, respectively) closely followed by BIDR (0.60). The fact that FOB was found as important as Intensity suggests that the mitigating effect of practice Intensity on Fatigue is driven by the level of engagement in a specific style of practice. To explore this idea, we performed a second model selection where we replaced Intensity and FOB by the two subcomponents of Intensity: IntensityFA and IntensityOM, corresponding to the two styles of practice. For the sake of parsimony, we also removed Experience, which was already found to be of low importance. The only selected model from this new variable space had a single regressor, IntensityFA (p = .034, β = 0.085, 95% CI [0.007, 0.162]; supplementary Table 3). Correlation with behavioral measures Previous studies have reported intra-individual variability of performance (most notably response times) as a good predictor of whether the participant is on-task at a given moment or not (Bastian & Sackur, 2013; Seli, Cheyne, & Smilek, 2013). Based on this literature, we predicted that self-reported Stability, but not other dimensions, would be significantly correlated to variability of response times at the level of individual participants. We chose Clarity as a control dimension, for its similar pattern of sensitivity to state and group (see Fig. 3). One sample Wilcoxon signed rank tests of z-transformed within-subject Pearson correlation coefficients against zero show that the RT variability correlated negatively with Stability (z = −0.26, 95% CI [−0.40, −0.12], p < .0003) as expected, but also with Clarity (z = −0.14, 95% CI [−0.28, −0.003, p = .046) (Fig. 5a, left). However, a two-sample paired test between zstability and zclarity was found significant (p = 0.041). Thus, even though self-reported ratings of Stability and Clarity were strongly correlated within subjects (Wilcoxon signed rank test on z-transformed correlation coefficients: z = 0.77, 95%CI [0.62, 0.94], p < .0001), the association with RT variability was significantly stronger for Stability than for Clarity. This suggests that although the phenomenological dimensions of Stability and Clarity tend to fluctuate naturally together, novices are able to differentiate them functionally in their reports, just like experts (Fig. 5a, right). In order to further assess the specificity of these findings, we repeated the same analysis using mean RT (instead of RT variability). We found that neither Stability nor Clarity correlated significantly with mean RT (both p > 0.1). Predictors of the phenomenological specificity Based on these results, we further hypothesized that this fine differentiation could have been implicitly trained in our novice group through the practice of meditation. Indeed, the reflexive quality cultivated in contemplative practices is expected to improve one’s familiarity with the specific phenomenal characteristics of different experiential dimensions. To test this prediction, we defined an index of phenomenological specificity as the difference between zstab, the z-transformed Pearson correlation coefficient between Stability ratings and RT variability, and zclar, the analogue measure for Clarity ratings. Only one model survived selection (supplementary Table 4), and it contained a single regressor, Experience (p = .023, β = −0.011, 95% CI [−0.021, −0.002]; Fig. 5b). Across all models, Experience had by far the highest evidence weight (0.84). Following the suggestion of a reviewer, we further explored whether Experience in particular practices drove phenomenological specificity. We performed a second model selection using the amount of Experience in each practice (FA, OM and CO) and the 2-way interactions between these 3 variables (to model for potential synergies between practices). Three models were selected from this new variable space (supplementary Table 5). All of them included ExperienceOM as a regressor but no interactions whatsoever. ExperienceOM had by far the highest evidence weight (0.91), while ExperienceFA and ExperienceCO were on par (0.44 and 0.49 respectively).","We described a meditation training protocol intended for naive candidates with no prior experience of meditation. We designed this protocol out of the need for a high-quality control group for a neurophenomenological study on the effects of meditation state and expertise in meditation on brain, behavior and physiology. The aim of the training was twofold: (1) to provide participants with sufficient background knowledge and direct experience with three types of meditations so that they could sustain a regular practice for an extended period of time, and be comfortable to practice in the laboratory context for the experimental tasks of the study; (2) to establish a common ground of relevant phenomenological categories with the participants, in order to allow them to report their experience during meditation states reliably. Overall we found evidence that the phenomenological training was successful in the sense that participants’ self-reports: (i) were reliable and sensitive to the meditation state manipulation, (ii) exhibited expected temporal dynamics such as dose and fatigue effects, and (iii) were functionally informative; in addition each of these effects was more strongly predicted by the amount and structure of participants’ practice than by desirable responding as indexed by the BIDR questionnaire. Motivation and compliance ~~~~~~~~~~~~~~~~~~~~~~~~~ In meditation research, motivation is often discussed for its potential confounding effect that limits the interpretability of longitudinal studies (Eberth & Sedlmeier, 2012). On the other hand, motivation of novice control participants can be considered a strength for the cross-sectional study of expertise as expert meditators, in so far as they dedicate a large of amount of time and resources to their training and practice, are expected to have strong motivation. Our study was highly demanding for the novices, as they had to engage in daily practice and participate in 6 to 8 experimental sessions over the course of several months. This, along with the multi-step recruitment procedure, acted as a filter for motivation. The high level of motivation of the novice participants was reflected in the satisfactory level of commitment to the practice maintained throughout their participation to the study. Three months after their training, they were still accomplishing more than half of the prescribed amount of practice, in the absence of any reminders or booster sessions. This even though they were assured that dropping the practice would remain without consequences for their participation to the study and financial compensation. This laxity given to participants, while having the effect of revealing their intrinsic motivation, is not without shortcomings. Compassion meditation, for example, was largely neglected. This might be related to the dense set-up of our initiation program, which attempted to train the participants in three different forms of meditation in just two consecutive days. In contrast, the original Joy of living program on which the training was based requires that practitioners first engage in 6 months of regular practice before they can receive teachings on compassion. However, this shortcoming has limited consequences for our study as the goal was primarily to get participants accustomed to the concept and practice of compassion and sensitize them to its difference with empathic resonance. A large majority (75%) of participants favored the practice of OM at home. This may come as a surprise considering that in many Buddhist contemplative traditions, OM practices are considered more advanced and are approached only after some training in FA (Lutz et al., 2008). However, this bias towards OM is consistent with the deliberate stance adopted by Mingyur Rinpoche, author of the Joy of living program, whereby one is invited to enter the field of open awareness from the outset. Phenomenological proficiency ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ During their training, novice participants were introduced to phenomenological categories with the help of practical, experiential exercises. Using rating scales and behavioral data from one of the experiments to which they later participated, we have described three effects that can be interpreted as reflecting an effect of practice, phenomenal training, or both. Based on both a priori considerations and control for desirable responding, we have systematically assessed the potential confounding effect of demand characteristics and found limited support for it. We review the evidence (see also Table 3) and discuss other potential alternate interpretations below. First, novice participants reported greater Aperture in OM compared to FA, just like experts, suggesting that they were able to distinguish the two practices. In addition, their responses on the Aperture scale was not driven by their diligence in any specific meditation style, but rather by the overall structure of their practice: the better they balanced focus and open styles of meditation, the larger the divergence in Aperture they reported (at least for participants who practiced the most: more than 24 h in total). This finding suggests that equal familiarity with different states is important for their optimal dissociation, at least at a beginner level. Future studies should test this hypothesis more rigorously using a longitudinal design with measures collected at baseline. Moreover, although we did not find any evidence for desirable responding in this data, we could not rule out the possibility that basic semantic priming inflated the reported dissociation in Aperture. Ideally, future training programs and experiments should try and avoid semantic overlap between meditation instructions and phenomenological dimensions altogether. Second, self-reported Stability and Clarity had a two-stage dynamic during a series of 4 consecutive six-minute blocks of meditation, with a statistically significant decrease between the second and the last block. We interpret this phenomenon as a fatigue effect, rather than an effect of scale misuse. This is based on the observation that this decrease was negatively and specifically associated to participants’ Intensity of practicing focus attention. Interestingly, this specific association is consistent with the role of concentrative practices in Buddhist contemplative traditions. These practices are used as training to stabilize attention and other basic qualities such as clarity and effortlessness, before applying them to more advanced practices. However, this correlation is not necessarily indicative of an effect of training: it may be mediated by an individual trait (e.g. conscientiousness or stamina), present even before the meditation training, that could predict both sustained diligence in the practice of (the relatively effortful) focus attention, and endurance during meditation sessions in the MEG experiment. Regardless, the index of desirable responding was of lesser importance in comparison. Third, we showed that participants’ ratings of the dimension Stability were functionally relevant, as they correlated with the variability of their response times. Using Clarity as a control dimension, we found that this functional relationship was specific. This finding suggests that participants were able to make fine distinctions between two close dimensions (this effect was true for both novice and expert practitioners). Regarding novices, we found that the more Experience they had at the day of the experiment, the sharper their phenomenological acuity. More specifically, we found that Experience in OM, compared to other practices, was the best predictor of this acuity. This is in line with the insight potential commonly attributed to OM-like practices in Buddhist contexts (Dahl, Lutz, & Davidson, 2015). Here again, in the absence of longitudinal data, the correlation cannot be treated as direct evidence for a causal link involving learning. However, a noteworthy difference with the fatigue effect described above is that the practice metric that predicted phenomenological acuity (accumulated Experience) is not confounded with participants’ assiduity, because it depends as much on Intensity of practice as on the time elapsed between the training weekend and the experiment (which was variable and random across the group). Considering that Intensity of practice did not predict phenomenological acuity, it appears that a likely explanation of these results is that novice participants became more familiar with the phenomenal richness of their experience throughout the regular practice of meditation, and more proficient in reporting it with specificity and subtlety. Taken together, our data support the claim that the novices in our study had some phenomenological literacy, and were able to report about qualities of their experience in an appropriate, meaningful and informative way, even though in some cases we could not conclusively rule out the possibility for an additional effect of demand characteristics. Training and practice ~~~~~~~~~~~~~~~~~~~~~ We have introduced several metrics of practice beyond the oft-used total amount. Although these metrics are derived from each other, they are not entirely collinear. In particular, our study was able to dissociate Experience (=total amount of practice) from Intensity (=daily average of practice) by having a large variability in the time elapsed between the training and the MEG experiment, across the novice group (from 27 to 133 days; M = 79, SD = 29). Moreover, we showed that these two metrics could be functionally dissociated when correlated with subjective ratings or behavioral measures. Such dissociations could point to potentially different mechanisms of trait changes brought about by the practice of meditation. This observation is consistent with the finding that intensive retreat practice, but not routine daily practice is associated with reliable differences in resting respiration rate in experienced meditators (Wielgosz, Schuyler, Lutz, & Davidson, 2016). It is also reminiscent of the work of Antonova, Chadwick, and Kumari (2015) who found that intensity of practice was a better predictor of decreased habituation to the acoustic startle reflex than total hours of practice. Future investigation of the mechanisms of meditation would benefit from a systematic exploration of various practice metrics and their relation to experimental outcomes, for both novice and expert practitioners. What is the minimum amount of practice that should be required from novice participants for the quality of their phenomenological self-reports to match those of experts’, on the dimensions explored here? Based on our experimental results, we can provide tentative, rough estimates. In our samples of participants, 20–40 h were necessary for novices to reach a level of phenomenal specificity comparable to the one of experts; a minimum of 20 min of practice per day on average and a high balance between practices (no more than 40% bias) enabled novice participants to differentiate different styles of meditation as well as expert practitioners. All these criteria are much higher than what most past cross-sectional studies of meditation expertise have required from their control participants (Fox et al., 2016), but are sufficiently low to be practically accessible and implemented in future studies. Self-rating scales ~~~~~~~~~~~~~~~~~~ We used self-rating scales as tools to help participants translate qualities of their lived experience into quantities that can be manipulated, transformed and statistically analyzed just like any other numerical measure. Such tools raise vexed issues: for example, how can we know that participants use the scales as intended? How can we know that participants, or groups of participants, are not construing a given scale in widely different ways? How can we even be sure that a given participant is consistent in the way he/she uses a scale over time or across experimental conditions, for that matter? To take the example of stability of meditation states, we may argue that stability refers to qualitatively different experiences in FA and OM. In FA, stability reflects the sustained focus on a given object and therefore the stability of mental content. In contrast, OM stability reflects the absence of grasping and as such, should not be affected by variations in content. Our approach of phenomenological training pragmatically addresses the issue of interpretation by mapping linguistic definitions of categories onto features of lived experience induced and revealed by simple experiential exercises. Performing these exercises in the context of a group, under the guidance of an instructor and with the possibility to share their understanding and reflect collectively, has the potential to attenuate idiosyncratic apprehensions of the phenomenological categories. In addition, concerns related to the subjectiveness and incommensurability of self-reported ratings were pragmatically addressed using within-subject designs and analyses. Even if our results suggest that our methodological approach was effective in detecting phenomenological fluctuations, it is worth mentioning the low variance of our self-rating data. As an example, 41 out of 50 of our participants used only three values out of the seven available in the Stability and Clarity scales; for Clarity, 18 out of 50 participants used only two values. This suggests a limitation of our experimental design and/or our training program. For instance the relatively short duration of laboratory experiments may not be sufficient to experience large fluctuations in these dimensions. Finer rating scales could be used as a compensation to increase data variance. Another possibility is that our training program was insufficient in developing participants’ fine-grained sensitivity to these scales. Further methodological work is needed to address these limitations. Future directions ~~~~~~~~~~~~~~~~~ The role of meditation practice for cognitive science was extensively discussed by Varela et al. (1991), becoming a part of their ‘enactive’ approach and then of Varela’s neurophenomenological program (Varela, 1996). We have provided preliminary evidence that meditation experience improves the reliability of self-report data by improving the functional specificity of self-reports, and by shielding them from the effect of demand characteristics. However, several questions remain open and should be addressed by future research. First, the impact of training on the quality of first-person data should be more rigorously assessed using high-quality, longitudinal, randomized controlled trials. In particular, future work should tackle the open question of whether specific phenomenological training such as the one we implemented through experiential exercises is required to improve the quality of first-person data, or whether meditation practice is sufficient in itself. In order to evaluate the confounding effect of demand characteristics on our first person-data, we have used an index of desirable responding. Unfortunately, the validity of the questionnaires designed for this purpose, including the one used in this study, has been frequently questioned as they are unable to distinguish between genuine personality traits and self-enhancement (Paulhus, 2002). Special attention should be given to more recent efforts to overcome these limitations using alternative and potentially complementary approaches (Kwan, John, Kenny, Bond, & Robins, 2004; Paulhus, Harms, Bruce, & Lysy, 2003). We have provided evidence for reliable first person reports of the phenomenology of meditation experience. However, the generalizability of the phenomenal insight provided by meditation practice to other, non-meditation-related applications remains disputed (see Khalsa et al., 2008 for an example of negative result) and warrant more research. Rather than provide a standardized, validated, ready for use protocol, our intention was to raise methodological concerns pertaining to the quality of control groups used in cross-sectional studies of meditation, and to argue for the possibility of obtaining informative experiential self-reports from adequately trained participants. Although we described in detail the protocol that we designed to address these issues, including the experiential exercises used for the phenomenological training of the participants, our approach is tailored to the specific context of our study, and to a particular phenomenological model of meditation among others (Bodhi, 2011; Lindahl, Fisher, Cooper, Rosen, & Britton, 2017; Van Dam et al., 2018). Still, we hope that the process, rather than the content, will inspire researchers in the field to further explore these critical issues.","The entirety of the Brain & Mindfulness project, including the meditation training weekend and the MEG experiment, was approved by the local ethics committee (CPP Sud-Est III, authorization number 2015-A01472-47). All participants signed an informed consent prior to their participation to the meditation training."],["This study examined whether trait emotional intelligence (trait EI or emotional self-efficacy) can differentiate between leaders and non-leaders (. N=. 96) employed by a major multinational company in Europe. Available intelligence test scores along with age, gender, and tenure were used as control variables. Trait EI, cognitive ability, and gender were significant predictors in a logistic-regression model. Further, both leaders and non-leaders scored significantly higher on trait EI compared to the standardization sample of the Trait Emotional Intelligence Questionnaire (Petrides, 2009), though the effect size for the former (Cohen's d=. 2.80) was considerably larger than for the latter (Cohen's d=. 1.23). The results support the notion that leadership and management positions require high trait EI. © 2014 Elsevier Ltd. --------------------------------------------------------------------------------","Much has been said about the importance of emotional intelligence (EI) in leadership, which overlaps with the concept of management (Young & Dulewicz, 2008). The extant literature on EI and leadership lacks differentiation among EI conceptualizations and operationalizations. For instance, when combining “emotional intelligence” and “leadership” as search terms in the PsycINFO database, 3,838 entries were returned. Combining the more specific constructs “trait emotional intelligence” or “ability emotional intelligence” with “leadership” led to 345 and 17 results, respectively. It follows from these numbers that the type of EI construct investigated was not specified in most of these studies. This is problematic because self-report measures of EI do not converge with maximum-performance measures; the former correlate substantially with personality and non-significantly with cognitive ability, while the latter show the opposite pattern of results (e.g., Qualter, Gardner, Pope, Hutchinson, & Whiteley, 2012). While both ability EI and trait EI have theoretical relevance to leadership, the focus of this article is on the latter, which is intended to represent the affective aspects of human personality. Trait EI is formally defined as a constellation of emotional self- perceptions located at the lower levels of personality hierarchies (Petrides, Pita, & Kokkinaki, 2007). There is a large literature demonstrating the validity of personality traits in the prediction of leadership-related constructs. Judge, Bono, Ilies, and Gerhardt (2002) conducted a large-scale meta-analysis showing that all Big Five personality traits, with the exception of agreeableness, predicted leadership, independent of industry and the leader’s specific job role, with a multiple correlation of .48. Yet, consistent with Paunonen and Ashton’s (2001) detailed approach to personality assessment, domain-specific traits may well improve the prediction of various leadership criteria, compared to the Big Five. Since leadership draws on many of the attributes assessed with the prevailing typical-performance EI measures (e.g., assertiveness, optimism; emotion expression, perception, and management), trait EI should emerge as a potentially important predictor of leadership-related variables. Many studies have examined and demonstrated associations between trait EI measures and various aspects of leadership. For example, Barling, Slater, and Kelloway (2000) found EQ-i (Bar-On, 1997) scores to predict three aspects of transformational leadership (idealized influence, inspirational motivation, and individual consideration), controlling for attributional style. Similarly, Mandell and Pherwani (2003) observed a predictive effect of EQ-i scores on overall transformational leadership (β = .49), which increased to β = .56 after controlling for gender. Villanueva and Sánchez (2007) found a moderate positive correlation (r = .56) between leadership self-efficacy (belief in one’s ability to lead), and trait EI, as measured with the Assessing Emotions Scale (Schutte et al., 1998). It has yet to be demonstrated whether leaders have higher trait EI than non-leaders, using objective (i.e., naturally- occurring), rather than psychometrically assessed classifications of leaders and non- leaders, and controlling for relevant factors (e.g., industry, gender, and age). Previous studies have focused on leadership attributes assessed with rating scales, often based on self-report. In another article in this issue, managers had higher trait EI scores than the normative sample of the Trait Emotional Intelligence Questionnaire (Petrides, 2009), but many other factors differentiating the two samples could not be controlled. Furthermore, the Trait Emotional Intelligence Questionnaire (TEIQue), which was developed to measure the construct of trait EI comprehensively and has been shown to possess superior psychometric properties over other measures (e.g., Freudenthaler, Neubauer, Gabler, Scherl, & Rindermann, 2008; Martins, Ramalho, & Morin, 2010), has not been used extensively in leadership research. The present study examined the role of trait EI in leadership within an applied context. In particular, leadership assessment was based on the organizational position of participants in a European multinational company. It was, thus, objectively determined and less prone to response biases than in other studies. Logistic regression was used to assess trait EI as a predictor of leader vs. non-leader positions held by the participants, controlling for cognitive ability, age, gender, and tenure. Furthermore, consistent with the management article in this issue (Siegling, Sfeir, & Smyth, 2014), we compared the mean trait EI level of leaders and non-leaders to the TEIQue standardization sample means, as reported in Petrides (2009). Two hypotheses were tested: Trait EI will distinguish leaders from non-leaders, controlling for cognitive ability, age, gender, and tenure. Leaders will have significantly higher trait EI scores than the TEIQue standardization sample.","A major European multinational company consented to participating in this study, which was conducted in Denmark. The company placed an online recruitment invitation on their intranet and those interested completed the questionnaire anonymously. Of a total of 300 contacted employees, 71 men and 25 women participated, yielding a response rate of 32% (N = 96). The mean age of the sample was 37.09 years (SD = 7.73) and the age range was 24–61 years. The majority of participants were of white ethnicity (90.6%) and had attained an undergraduate (40.6%) or Master’s degree (42.8%). Employees were in their present position (tenure) for an average of 7.88 years (SD = 7.59). Participants came from four business units of the company and were involved in various job functions, including technical support, sales and marketing, finance, logistics, and security. A leadership post in the company entails directly supervising three or more employees whom the leader manages and appraises. Leaders have the right to hire and dismiss employees and are expected to drive company values and deliver top quartile results. They are also expected to inspire their team to higher performance through improving engagement and developing their supervisees’ capabilities to perform efficiently. The study sample comprised 23 leaders (22 males) and 73 non-leaders (49 males). The mean age of leaders was 36.62 years (SD = 7.92) and the mean age of non-leaders was 38.61 years (SD = 7.06). Cognitive ability scores and leadership status data were obtained from participants’ records from the human resources department of the company. Additional information, including trait EI scores, demographic background data (age, gender, and educational level) and job tenure (the number of years employees had been in their present position within the organisation), was collected over a time-frame of approximately two months. Trait EI Either the full Trait Emotional Intelligence Questionnaire (TEIQue – v. 1.50; Petrides, 2009) or its short form, the TEIQue–SF, were used as measures of trait EI. Due to time constraints on data collection, the full form was replaced by TEIQue–SF two weeks into the process. In total, 40 participants completed the TEIQue and 56 completed the TEIQue–SF. The full form consists of 153 items, while the short form consists of 30 items. Both yield global and factor scores, although the former also yields scores on the 15 trait EI facets. The internal consistency for global trait EI was .95 for the full form and .89 for the short form. Leadership Leadership was operationalized in accordance with the company’s definition of a leader, as described in the preceding section. Cognitive ability Cognitive ability was measured using an in-house Wonderlic-type test that was developed by a leading global test developer. Scores range from 0 to 50, with a score of 25 corresponding to an IQ of 120.","Trait EI scores based on the full TEIQue were derived from the items also found on the short form. However, one of the 30 short form items (item 10) does not originate from the full form and, thus, it was replaced by a similar item (item 115). Descriptive statistics for leaders and non-leaders are shown in Table 1. Leaders had significantly higher trait EI scores than non-leaders, which was largely an effect of the Well-Being, t(94) = 2.10, p = .04, and Self-Control, t(94) = 2.62, p = .01, factors, which reached significantly higher levels in leaders (Sociability approached significance at p = .05). As shown in Table 1, the average tenure of leaders was significantly higher than that of non-leaders. In contrast, leaders and non-leaders did not differ on age or cognitive ability. Table 2 displays the bivariate correlations between the study variables. Trait EI correlated positively with leadership, which, in turn, correlated positively with tenure and negatively with gender (there was only one female leader). Cognitive ability was negatively associated with tenure and age, which were positively associated between them. Table 3 shows the logistic regression results. A chi-square goodness of fit test showed that the set of predictors included in the model distinguished between leaders and non- leaders, χ2(5) = 19.77, p = .001. The Hosmer–Lemeshow test indicated a good model fit to the data, χ2(8) = 5.47, p = .71. Indices for the usefulness of the model tested here showed that the model explained between 18.6% (Cox & Snell R2) to 27.9% (Nagelkerke R2) of the variance in leadership. Relative to the constant-only model, the model including trait EI and control variables improved the prediction accuracy of leadership cases from 76.0% to 78.1%. Trait EI was a significant predictor of leadership in the presence of the control variables in the equation (cognitive ability, tenure, age, and gender). The predictive effect was such that leaders had higher trait EI scores than non-leaders. Given the levels of the control variables and the proportion of female to male participants, the odds of being a leader are multiplied by 2.77 for each one-unit increase in trait EI. Although preliminary analyses revealed no significant relationship, cognitive ability was a significant incremental predictor in the regression model. Here, the odds of being a leader are multiplied by 1.12 for each on-unit increase in cognitive ability, given the levels of trait EI and the other control variables. Consistent with the zero-order correlations, gender was the only other significant predictor of leadership, with a greater proportion of males among leaders than among non-leaders. Given the levels of the other variables in the model, female employees were 0.10 times as likely to be a leader as male employees. With the single female leader excluded, the mean global trait EI score of leaders was compared to the TEIQue male-normative standardization sample mean. A one- sample t test showed that male leaders in the current sample had significantly higher trait EI scores (M = 5.54, SD = 0.43) than the normative comparison group (M = 4.95, SD = 0.61), t(21) = 6.41, p < .0001. The average trait EI score of male and female non-leaders (M = 5.29, SD = 0.64) was also significantly higher than the combined normative sample mean (M = 4.90, SD = 0.59), t(72) = 5.22, p < .0001; however, the effect size for male leaders (Cohen’s d = 2.80, r =.81) was much larger than for male and female non-leaders (Cohen’s d = 1.23, r = .52).","The results support Hypothesis 1 and are consistent with previous research relating EI scores from self-report measures other than the TEIQue to various leadership attributes (Barling et al., 2000; Judge et al., 2002; Mandell & Pherwani, 2003; Villanueva & Sánchez, 2007). Global trait EI discriminated between leaders and non-leaders within the same company, controlling for pertinent characteristics. Compared to previous research, in which leadership-related variables were assessed through rating scales, leaders in our study were identified based on their actual occupational position within the company. Further, this study used the TEIQue, a measure designed to represent trait EI comprehensively (Petrides et al., 2007). Although Hypothesis 2 was supported, the non- leader group was also significantly above the TEIQue standardization sample mean, though the effect size for the leadership group was more than twice as large. It is important to bear in mind that the standardization sample is not matched on pertinent characteristics to the sample of this study. For example, the average age of the standardization sample (29.65 years, SD = 11.94) is about seven years younger than the mean age of the current sample, which falls within the age interval (34–44) at which trait EI scores were found to peak (Derksen, Kramer, & Katzko, 2002). On the other hand, the overall sample appeared to be generally high in trait EI, as even the non-leaders had global trait EI scores comparable to a managerial sample of similar age described in another study in this special issue (Siegling et al., 2014). The negative relationship between cognitive ability and tenure can explain why even though cognitive ability was not directly related to leadership, it emerged as a significant predictor in the regression model, with the effect of tenure controlled. A plausible explanation is that employees with higher cognitive ability advance to higher positions faster or move onto other jobs sooner. Tenure in this study referred to the number of years employees were in their present position, hence the negative relationship with cognitive ability. Although leadership and management are not interchangeable (Lunenburg, 2011), they are overlapping concepts (Young & Dulewicz, 2008). Therefore, the results reported in this study are consistent with those reported in the parallel article, wherein UK managers also showed higher trait EI than the standardization sample (Siegling et al., accepted). While these findings require replication on other samples and industries, they provide initial evidence to suggest that the range of personality traits linked to emotions is fundamental in occupational roles involving the supervision of, and responsibility for, others."],["Readers often describe vivid experiences of voices and characters in a manner that has been likened to hallucination. Little is known, however, of how common such experiences are, nor the individual differences they may reflect. Here we present the results of a 2014 survey conducted in collaboration with a national UK newspaper and an international book festival. Participants (n = 1566) completed measures of reading imagery, inner speech, and hallucination-proneness, including 413 participants who provided detailed free-text descriptions of their reading experiences. Hierarchical regression analysis indicated that reading imagery was related to phenomenological characteristics of inner speech and proneness to hallucination-like experiences. However, qualitative analysis of reader's accounts suggested that vivid reading experiences were marked not just by auditory phenomenology, but also their tendency to cross over into non-reading contexts. This supports social-cognitive accounts of reading while highlighting a role for involuntary and uncontrolled personality models in the experience of fictional characters. --------------------------------------------------------------------------------","Vivid or immersive experiences are often described in relation to reading fictional narratives (Caracciolo & Hurlburt, 2016; Green, 2004; Ryan, 1999, 2015). In particular, it seems common for readers (and writers) to report “hearing” the voices of fictional characters, in a way that suggests they have a life of their own (Vilhauer, 2016; Waugh, 2015). What, though, does this mean for the psychological processes that may underpin such experiences? Psychological studies on the phenomenological experience of reading have tended to focus on two strands: first, the perceptual and sensory qualities of reading – primarily via notions of ‘voice’ and inner speech (Alexander & Nygaard, 2008; Perrone- Bertolotti, Rapin, Lachaux, Baciu, & Lœvenbruck, 2014); and second, how the reader represents the characters and agents of a text (Kidd & Castano, 2013; Mar & Oatley, 2008). To a certain extent, it is intuitive to understand why a text – even if not read out loud – would need to be voiced in some way to be read. This is sometimes conceptualized either as inner speech – namely, the various ways in which people talk to themselves (Alderson- Day & Fernyhough, 2015) – or more broadly in terms of auditory imagery, i.e. purposefully imagining the qualities of characters’ or narrators’ voices (Hubbard, 2010; Kuzmičová, 2013). Evidence of inner speech involvement comes from psycholinguistic studies on reading: when we read, phonologically longer stimuli take longer to read than shorter stimuli of the same orthographic length (Abramson & Goldinger, 1997; Smith, Reisberg, & Wilson, 1992), while acoustic properties of one’s own voice, such as accent, can affect our expectation of rhyme and prosody (e.g., Filik & Barber, 2011). This suggests that at least some properties of text are sounded out in inner speech during reading (Ehrich, 2006). Readers’ expectations of character and narrator voices can also affect how a text is processed. For instance, readers adjust their reading times for texts written by authors with a slow or fast-paced voice. People reading difficult texts, and those who report more vivid mental imagery, show greater evidence of such “author voice” effects on reading speed (Alexander & Nygaard, 2008). When characters’ words are referred to in direct speech, voice-selective regions of auditory cortex are more active than during indirect reference (e.g., ‘He said, “I hate that cat”’ vs. ‘He said that he hates that cat’), suggesting auditory simulation of character’s voices (Yao, Belin, & Scheepers, 2011). Evidence of inner speech and auditory imagery being involved in reading is consistent with broader theories of reading that place perceptual simulation and embodiment at the heart of textual comprehension (Zwaan, 2004; Zwaan, Madden, Yaxley, & Aveyard, 2004), i.e. the idea that sensorimotor imagery processes are automatically engaged when we read text, as part of understanding the meaning of what is being described. A second strand of research has emphasized the role of social cognition in the reading experience, largely in response to literary fictional texts. Many readers strongly personify characters and narrators by making inferences about their described thoughts and behaviors (Bortolussi & Dixon, 2003) and assigning them intentionality (Herman, 2008). Studies on empathy in literary experiences have focused on how empathetic engagement with characters is triggered by specific discourse strategies (e.g., first-person vs. third- person narratives; Keen, 2006) or on how empathy in the act of reading relies on readers’ previous personal experiences (Kuiken, Miall, & Sikora, 2004; Miall, 2011). Other psychological approaches to reading have investigated the “projection” of knowledge that readers perform – the process by which they assign to each character an individual epistemic view of the narrative world (Gerrig, Brennan, & Ohaeri, 2001), which allows for narrative dynamics such as “suspense” (Gerrig, 1989). Based on such processes, it has been argued that the ways in which readers attribute consciousness, mental states, intentions, and beliefs to characters recruits (Zunshine, 2006, 2012) or even enhances (Kidd & Castano, 2013) readers’ theory-of-mind, i.e. the ability to represent the mental states of others. Indeed, it has been claimed that the “function” of reading fiction may be to simulate social experiences involving other people (Mar & Oatley, 2008). Taken together, the above studies highlight some of the separate perceptual and social-cognitive processes that could explain accounts of ‘hearing’ the voices of characters. But, although the experiences they are based on are intuitively familiar, phenomenological data on the reading experience in the words of readers themselves is surprisingly lacking. Indeed, almost all of the above work has involved either experimental manipulation of texts, or analysis of responses to specific literary texts. Systematic surveys of readers’ experiences of characters in general – that is, as part of their day-to-day experience of reading for pleasure – are largely absent. We know of only one recent exception: Vilhauer (2016) conducted a qualitative analysis of 160 posts that resulted from a search of ‘hearing voices’ and ‘reading’ from a popular message-board website (Yahoo! Answers). Of these, over 80% reported vivid experiences when reading, the majority describing specific auditory properties including volume, pitch, and tone. Qualities that are perhaps more indicative of social representation – such as identity and control – were also reported in some cases, but were often hard to classify or lacking in detail. However, the open-ended structure of the source material used by Vilhauer (2016), the fact that it was gathered based on the specific keywords ‘hearing voices’, and the lack of demographic data from the study participants limit any strong generalizations about the reading experience. As such, it is unclear whether vivid examples of characters’ voices – or indeed other kinds of character representation – are actually a common part of the reading experience. It could be that the act of reading about characters simply involves combining such features in an additive and largely automatic way: if so, phenomenological reports may be expected to consist of vivid perceptual imagery plus some kind of mental state representation – a clear experience of a character’s voice and their emotional state, for example. But such skills also vary considerably in the general population (Isaac & Marks, 1994; Palmer, Manocha, Gignac, & Stough, 2003) and may not be integral for most people, most of the time: for some, experiences of voices, characters, or other features of a text could combine to create something very different entirely, or even nothing at all (a character’s voice without any impression of intentionality, for example). To investigate this question, we collaborated with the Edinburgh International Book Festival and a national UK newspaper (the Guardian), to survey a large sample of readers about their inner experiences. Instead of focusing on the experience of a particular text (e.g., Miall & Kuiken, 1999), or experimentally varying textual properties (Dixon & Bortolussi, 1996), we opted for a general questionnaire about readers’ encounters with voices and characters. This encompassed all kinds of reading (prose vs poetry; crime fiction vs historical novels; or fictional vs non-fictional narratives), although the large majority of eventual responses related to engaging with fiction (82%). The first aim of the survey was to gather quantitative information on the vividness of readers’ experiences, and examine how that related to other individual differences in potentially similar processes. Based on the putative involvement of inner speech in reading, we included a measure of everyday inner speech experiences: the Varieties of Inner Speech Questionnaire (VISQ: McCarthy- Jones & Fernyhough, 2011). Derived from developmental theories of self-talk (Vygotsky, 1987), the VISQ measures a range of phenomenal properties of inner speech, including the extent to which it includes dialogue, if it is experienced in full sentences, whether it is evaluative or motivating, and whether it includes other people’s voices. If readers were more likely to report vivid experiences of voice and character during reading, they might also be expected to have a more vivid experience of their own inner speech in general. In addition, vivid experiences of voices and characters during reading have been likened by some (e.g., Vilhauer, 2016) to be similar to actual experiences of ‘hearing voices’ or auditory verbal hallucinations (AVHs). Although direct parallels with florid and distressing experiences are unlikely, traits towards having unusually vivid and hallucination-like experiences are thought to exist along a continuum in the general population (Johns & van Os, 2001). They may, therefore, relate to reports of particularly vivid reading experiences. To investigate this further, we included a short measure of auditory hallucination-proneness (the Launay-Slade Hallucination Scale – Revised; Bentall & Slade, 1985). We predicted that participants with more vivid experiences of reading in general would also be more prone to hallucination-like experiences. Our second aim was to qualitatively explore readers’ own descriptions of voice and character, via a free-text section of the survey. Such descriptions offer a nuanced and detailed picture of the inner experience of reading that might otherwise be lost in purely quantitative approaches to the topic. Many of the previously mentioned studies have focused on a very specific aspect of reading (activation of inner speech, auditory imagery, mental imagery, empathy for characters, projection of knowledge, etc.) without attempting a unified account of how all these aspects relate to each other and trigger other, richer processes. To address this, readers’ descriptions in our study were coded in terms of representational features of the experience (such as the different sensory modalities involved), but also their dynamics, namely the processes by which the experiences seemed to occur. While this analysis was primarily driven by the main themes apparent in the data, our descriptions in some cases required the creation of new terms or their importing from narratological and linguistic research on fictional narratives. For example, here we have used the term “mindstyle” – a concept borrowed from linguistics (Fowler, 1977; McIntyre & Archer, 2010; Semino, 2007) – to refer to the unique way in which a person or character thinks about and views the world, and an idea that is clearly relevant for social simulation accounts of the reading experience. This and other terms used in the coding are expanded on below.","Participants were invited to take part in an online survey on readers’ inner voices via a series of blogposts for the Books and Science sections of the Guardian website (‘Inner Voices’), publicity at the Edinburgh International Book Festival (EIBF) 2014, social media, and a project website (www.hearingthevoice.org). A total of 1566 participants (75% F/24% M/1% Other; Age M = 38.85y, SD = 13.48y, Range 18–81) took part in the survey, with responses primarily coming from English-speaking countries (UK, USA, Australia, Ireland, and Canada; see Table 1 for demographic details). Participants were also asked to indicate their level of education and their general reading preferences. This indicated that the sample had a high level of educational achievement on average (over 80% possessing a graduate degree or higher), with the most popular reading preferences being for general fiction, literary classics, and historical fiction. The survey was live for 6 weeks, and all procedures were approved by a local university ethics committee.","The survey was divided into two parts. Section 1 – the Readers’ Imagery Questionnaire – specifically asked about participants’ vivid experiences of voices and characters during reading. Section 2 included the questionnaire items on inner speech and auditory hallucination proneness. Reading imagery questionnaire A reading imagery questionnaire was devised specifically for the present study based on commonly used measures of imagery, including Betts’ Questionnaire upon Mental Imagery (Sheehan, 1967) and the Vividness of Visual Imagery Questionnaire (Marks, 1973). It consisted of five items, each answered on a 5-point Likert scale (see Table 2): Do you ever hear characters’ voices when you are reading? Do you have visual or other sensory experiences of characters when reading? How easy do you find it to imagine a character’s voice when reading? How vivid are characters’ voices when you read? Do you ever experience the voices of characters when not reading? To elicit more phenomenological detail about the experience, a further question asked: ‘If you feel that you have had particularly vivid experiences of characters' voices, please describe them in the box below’. Responses up to a 500-word limit were allowed. Qualitative analysis was then conducted on the readers’ responses to this question specifically (see below). Varieties of inner speech questionnaire (VISQ: McCarthy-Jones & Fernyhough, 2011) The VISQ includes 18 items on the phenomenological properties of inner speech. It includes four subscales: dialogic inner speech (e.g., ‘I talk back and forward to myself in my mind about things’); evaluative/motivational inner speech (‘I think in inner speech about what I have done, and whether it was right or not’); other people in inner speech (‘I hear other people’s voices nagging me in my head’); and condensed inner speech (‘I think to myself in words using brief phrases and single words rather than full sentences’). Participants rated their agreement with these statements on a 6-point Likert scale ranging from “Certainly does not apply to me” to “Certainly applies to me”. Each subscale has good internal and test-retest reliability (Alderson-Day et al., 2014; McCarthy-Jones & Fernyhough, 2011). Launay-Slade hallucination scale – revised (LSHS: Bentall & Slade, 1985) A short version of the LSHS was used to assess susceptibility to auditory hallucinations. Five items that specifically related to unusual auditory experiences were selected from the Revised Launay Slade Hallucination Scale used in Morrison, Wells, and Nothard (2000), for example: ‘I hear people call my name and find that nobody has done so’. Participants indicate their agreement on a 4-point scale ranging from ‘Never’ (1) to ‘Almost Always’ (4). Data on the 5-item version reported by McCarthy-Jones and Fernyhough (2011) and Alderson-Day et al. (2014) have shown the scale to have moderate/good internal reliability (Cronbach’s alpha >0.69). As part of a separate study, participants also completed questionnaires on inner speech frequency and imaginary companions. While this will be fully reported elsewhere, here we have included some preliminary data on inner speech frequency that corroborates the other, agreement-based VISQ results (see Footnote 2). Quantitative analysis & qualitative coding ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A mixed methods approach was used to (i) analyse the relations between questionnaire measures collected in Sections 1 and 2, and (ii) qualitatively code free-text responses given at the end of the Readers’ Imagery Questionnaire. Questionnaire answers to Sections 1 and 2 were analysed using Spearman’s Rho correlation co-efficients (due to non-normal distributions in some of the questionnaire outcomes) and hierarchical regression analysis, using total score for reading imagery as the dependent variable. A Bonferroni correction was applied to all pairwise correlations tested to avoid inflated type 1 error incurred from multiple comparisons. Missing questionnaire responses were replaced with the mean per item. Free-text responses from Section 1 were coded using an inductive thematic analysis (Braun & Clarke, 2006). Two raters (BA-D and MB) first independently devised a series of descriptive codes from the entire dataset. Codes were then discussed, refined, and applied to 20% of the dataset for parallel coding by each rater, before independent coding of the remainder. Inter-rater reliability for coding was high (k = 0.82). The coding scheme classified answers in two ways: firstly, for their general features, including presence of specific sensory properties; and secondly in terms of dynamics, or descriptions of the overall experience and imaginative process described by the reader. Although each response could only receive each code once (e.g., descriptions of experiences with multiple visual features nevertheless only received one visual code), any given answer could be classed as having several different features and dynamics. Along with the term “mindstyle”, our novel codes included “blending”, a term used by Fauconnier and Turner (2003) to denote the mixing of concepts from multiple domains to form new combinations (as in, for example, the creation of novel imagery). In contrast, “experiential crossing” is a new term that we use here to refer to instances of voices and characters being experienced beyond the immediate context of reading. A full list of codes is provided in Table 3. Unless otherwise indicated, all italicization in example quotes has been added by the authors to illustrate how specific content relates to specific codes allocated. For clarity, all references to book titles have been placed in inverted commas. Readers’ experiences of voices and characters: summary data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Less than 1% of responses were left blank in Section 1 (the Reading Imagery Questionnaire; RIQ). Table 2 shows the frequency of response for each of the questions on vivid reading experiences. Most participants reported hearing characters’ voices when reading at least some of the time, with over half (51%) hearing them most or all of the time (Q1). Visual and other sensory experiences were endorsed to a similar degree (Q2). Approximately two thirds of participants described it as being fairly or very easy to imagine characters’ voices when reading (Q3), while the vividness of voices varied from being vaguely present to being as vivid as listening to an actual person (Q4). Experience of characters’ voices outside of reading was, in contrast, relatively rare: a fifth of participants described this experience happening sometimes, but less than 4% of participants reported this happening most or all of the time (Q5). Relations between reading experiences, inner speech, and auditory hallucination-proneness ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Within the sample, 1522 participants also completed measures of inner speech (the VISQ) and hallucination-proneness (the LSHS). Less than 3% of responses were left blank. As the internal reliability of the five RIQ items was good (Cronbach’s alpha = 0.80), scores across the five items were summed to provide a total score of readers’ susceptibility to vivid reading experiences. Correlational analysis indicated that total score on the RIQ was positively related to auditory hallucination-proneness (r = 0.25), other people in inner speech (r = 0.38), dialogic inner speech (r = 0.23), and evaluative/motivational inner speech (r = 0.20, all p < 0.001, Bonferroni corrected). Reading scores also negatively correlated with condensed inner speech (r = −0.13), i.e. participants with more expanded than condensed inner speech also had more vivid reading experiences. When these factors were assessed in a hierarchical regression model, only a subset of predictors was retained (see Table 4). Using total score on the reading items as the dependent variable, Age, Gender,2 and Education Level were included as control predictors in block 1, followed by each of the VISQ subscales and total score on the LSHS. The final model significantly predicted total RIQ score (F = 43.06, p < 0.001) accounting for 18% of the variance. Significant predictors retained by the model were: LSHS (p < 0.001, β = 0.11); dialogic inner speech (p < 0.001, β = 0.11); other people in inner speech (p < 0.001, β = 0.30); and condensed inner speech (p < 0.001, β = −0.09). Age, gender, education level, and evaluative inner speech were all non-significant (gender: p = 0.08, all other p > 0.50).3 Reader’s experiences: detailed qualitative descriptions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Further description of reading experiences was provided by 413 participants. Free text responses ranged from 1- to 320-word answers (answer length M = 41.71 words, SD = 35.26 words). Features As may be expected, descriptions of hearing characters’ voices and references to specific characters were very common, occurring in 291 cases (70%). Many participants reported a strong and vivid engagement with the characters of a text. For example, one participant stated [emphasis added throughout]: I become so engrossed in a novel that the characters become real to me. I know that they are not real, but they feel real. It is as vivid as watching the characters in a film on TV where the screen is my mind's eye. In fact, if I can't hear the characters' voices, I find it impossible to carry on reading the novel because it is the vivid experience of characters' voices (or sometimes the author's “voice”) that I want when I read a novel. Characters, however, were not the only voices present, despite the questionnaire focusing on characters specifically. 49 participants (12%) described vivid perceptions of the author and/or narrator’s voice in a text. For example: There are times when reading, that the narrative draws me in such that the words are no longer words just the world they describe. But in a few cases, a narrator’s voice comes through, almost as if one was listening to someone reading it […] Well composed dialogue often creates a sense of 'listening' to a real conversation with varied tones and accents for me. I can also feel as if a strong first person narrative voice is “speaking” fully to me […]. Auditory characteristics of voices were reported in the large majority of cases. In 80 cases (19%) this involved the specific mention of accents without any reference to other auditory properties: I can hear male and female, young and old, accents that I can't speak myself [….] When I read books whose characters are English, Scottish, Irish, or any of the strong American dialects, I hear them speaking with that dialect – even without relying on the author's use of it in writing what they say. No other specific auditory properties were mentioned as regularly (hence they are grouped under the general “auditory” code). Auditory characteristics described in 250 cases (61%) included references to volume, tone, and speed of speech: I can hear the tone and voice inflexions depending on what the character is saying. Usually the tone is distinct, particular words may be over-pronounced or roll together excitedly. Depending on the situation, if the character becomes stressed or scared for e.g., the voice changes pitch and speed… For 76 cases (18%), the experience of a character did not consist in the quasi-sensory perception of their voice, but rather their manner of speaking, thinking, and experiencing the world – i.e. their mindstyle (Fowler, 1977): It's not only the pitch or sound, but the rhythm of the speech, and the character's emotion and movement I get, it's the whole package, as if I'm watching a film or in the same space. With certain favourite books, which I have reread a number of times, I feel I can hear […], the intonation, the way in which they would speak phrases from the text but also their responses to other situations and problems. Visual characteristics were described in 57 cases (14%). In many cases this was described as an immersive, cinematic experience: Every good book I read that I get in depth with, plays like a movie, I don’t really realize it while I’m reading but when I stop I go to find where I’m up to in the book and I don’t recognise reading the words but I know and have “heard” and “ seen” what has happened. Other feelings, such as tactile or olfactory experiences, were also described in a smaller proportion of cases (47 participants; 11%): Lyra whispering to Will in Pullman's “His Dark Materials”. He describes the loud, busy closeness of her whisper, and I could hear it and feel it on my neck. It feels like I'm sharing the surroundings with the characters or simply experience the landscape, weather, smells, touch, sounds etc. Dynamics While almost all responses provided some information of the features of vivid reading experiences, not all answers could be coded with a specific dynamic. In those that could be classified in such a way (281 participants; 68%), the dynamics of vivid reading experiences were very varied. The first dynamic coded was internal blending, or the process of adapting one’s own inner voice to represent a particular character or narrator. Twenty-two participants (5%) referred to “putting on the voices rather than hearing separate people” or that voices were similar to their own in some way, e.g.: I think the voice is a version of mine usually. I probably overlay class, accent etc. Howard Roark in “The Fountainhead” spoke very calmly and cooly but it was not very different from my own. More common was a kind of external blending, occurring in 62 cases (15%). In this, participants drew on their experience of others’ voices to simulate a character (often usually a friend, relative, or actor from an adaptation). The voice of Zooey in Salinger's “Franny and Zooey” always sounds like a friend of mine who shares many of Zooey's character traits, including a resonant voice. If the book has been made into a film then I would tend to hear the actor's voice. Even if I read the book first, the actors would generally override anything I might originally have imagined. Slightly more frequent than both, however, were responses in which the experience of characters’ voices (or thoughts) appeared to break across into new situations, outside of the immediate experience of reading. 77 cases (19%) described this experiential crossing of voices: If I read a book written in first person, my everyday thoughts are often influenced by the style, tone and vocabulary of the written work. It's as if the character has started to narrate my world. Whenever I'm reading a novel I always hear the characters talking even while not reading. They continue a life between bouts of reading. In some cases this was experienced just as a continuation from reading the book; as in one participant’s description of reading Virginia Woolf: Last February and March, when I was reading “Mrs. Dalloway” and writing a paper on it, I was feeling enveloped by Clarissa Dalloway. I heard her voice or imagined what her reactions to different situations. I'd walk into a Starbucks and feel her reaction to it based on what I was writing in my essay on the different selves of this character. In others this crossing was specifically prompted by familiar or new but similar contexts, i.e. the characters would appear when it would be consistent with their own persona or surroundings: The character Hannah Fowler, from the book of the same name was the voice I heard while walking with my family in the area of Kentucky (USA) where the book took place. I loved the book and heard her dialogue as I walked through the woods.[ …] Other than this, participants would often describe their experience of voices and characters in general, non-specific terms that implied an intentional imaginative construction of characters and scenes. This inner simulation occurred in 77 cases (19%), e.g., I see the book as a movie and I am barely aware of the pages. I hear voices, music, and other sounds as described by the author or imagined by me…. I can visualize and imagine what the characters sound and look like. The voices, how they express themselves, verbally and non-verbally. A small number of responses (19 cases; 5%) also described a specific sense of voices “fitting” characters in particular ways, or failing to fit following depiction by an actor or narrator. This dissonance of voices could in some cases noticeably interfere with enjoyment of the text. One participant, for example, stated that their experience of voices was: Vivid enough that when I hear an author speak, I am often surprised how different they sound than the “narrator” in my head. It's the same with films; characters often sound “fake” compared to how I have imagined them. Finally, 24 responses (6%) appeared to describe actual hallucinatory phenomena. Many of these described states in which hallucinations are relatively common, such as in the transition to and from sleep. I have hypnagogic sleep, so will sometimes hear an actual voice whilst falling asleep or awakening, which can be related to what I've been reading before sleep.","The aim of the present study was to survey phenomenological qualities of voices and characters in the experience of readers. Our results indicated that many readers have very vivid experiences of characters’ voices when reading texts, and that this relates to both other vivid everyday experiences (inner speech) and more unusual experiences (auditory hallucination-proneness). However, the features and dynamics of readers’ descriptions of their voices were varied and complex, highlighting more than one way in which readers could be said to “hear” the voice of specific characters. This included quasi-perceptual events across a variety of sensory modalities; personified, intentionally and cognitively rich agents; and characters that both triggered and echoed previous experiences and extra- textual connections. On the RIQ, the large majority of participants often “heard” character’s voices when reading, with visual and other experiences also occurring frequently. Most found it very easy to imagine characters’ voices during reading, but the vividness of this varied considerably: 1 in 7 participants reported voices that were as vivid as hearing an actual person speak, but double that proportion described either no voices being present or only vague experiences of voice. As such, the experiences reported here should be considered particularly vivid examples of auditory mental imagery, in line with an experiential – and not merely propositional – view of imagery (Kosslyn, Thompson, & Ganis, 2006). This is broadly consistent with inner speech and perceptual simulation accounts of reading (Engelen, Bouwmeester, de Bruin, & Zwaan, 2011; Zwaan et al., 2004). However, it would appear to largely fall short of indicating that participants were having literally auditory experiences during reading. This contrasts, for example, with Vilhauer’s (2016) survey, which heavily emphasized such features (e.g.,“An overwhelming majority of [participants] indicated that inner reading voices were audible”, p. 5). We note that the accounts collected in that survey represented a subset of participants who were describing reading experiences after having already referred to “hearing voices”, suggesting that auditory phenomenology (and more specifically, literal hallucinatory experiences or potential psychopathology) may be over-represented in Vilhauer’s (2016) sample. The relevance of inner speech to this topic is emphasized by the observed relations between the overall vividness of reading experiences and individual differences in the ongoing internal self-talk reported by participants. Regression analysis indicated that participants who reported more elaborate inner speech on the VISQ – being expanded in form, containing dialogue, and the voices of others – were those who also had more vivid experiences of voices and characters during reading. This is a unique demonstration (in terms of empirical research) of the more positive and valuable correlates of phenomenologically diverse inner speech, which in prior studies has correlated with anxiety, depression, low self-esteem, and unusual experiences (McCarthy-Jones & Fernyhough, 2011). Notably, evaluative inner speech – which is linked to low self-esteem (Alderson-Day et al., 2014) – was not significantly associated with reading imagery on the RIQ, once other inner speech and demographic factors were controlled for. It could be that evaluative aspects of inner speech, if present, are more closely linked to negatively- valenced moods and processes (such as rumination; Jones & Fernyhough, 2009), while other features of inner speech may reflect more general imaginative capacities, as indexed by the RIQ used here. The other factor relating to reading imagery was auditory hallucination-proneness (i.e. LSHS), despite the fact that relatively few participants reported character voices that were as vivid as hearing another person speak. While it is possible that these experiences may nevertheless be linked on a continuum of quasi- perceptual phenomena, the qualitative analysis of readers’ accounts sheds some light on what may link these experiences. Descriptions of vivid auditory imagery were common, but the ways in which this was experienced were highly varied: for some participants this was an intentional, constructive process, for others an automatic immersion, and for others again an experience that appeared to seep out into other, non-reading contexts. For example, one notable way in which participants talked about their experience of characters was via what we termed “experiential crossing”. We coined this term to refer to instances of characters and voices being experienced outside of the context of reading; a phenomenon that as far as we know has never been studied either in psychological or narratological research. The presence of experiential crossing in nearly a fifth of participants points towards a perfusion of voice- and character-like representations that apparently transgress the boundary between reading and thought. In some cases this was described almost as an echo of prior reading experiences, with auditory imagery re-emerging in a particular context or scenario, but in other accounts it appeared to shape the readers’ style and manner of thinking – as if they themselves had been changed by a character. Indeed, one discontinuity between experiential crossing and other kinds of dynamic observed in readers’ accounts was the preponderance of fictional characters’ thoughts and feeling over specifically perceptual elements – what was coded as their mindstyle (Fowler, 1977; McIntyre & Archer, 2010; Semino, 2007). These readers were engaged with a range of complex cognitive faculties, from beliefs and behavioral patterns to feelings and perceptual biases, with a particular emphasis on the emotional states of characters. All of these become part of the readers’ construction of a fictional consciousness or a “consciousness frame” (Palmer, 2004): a kind of schema for the characters’ worldview and inner experience. This extensive degree of personification that readers process and project into the text can be seen as the counterpart of personifying processes occurring in the writers’ encoding of fictional minds (Taylor, Hodges, & Kohányi, 2003). Thus, one alternative overlap between vivid reading experiences and hallucination-proneness may lie less in auditory phenomenology and more in the uncontrolled experience of another’s point- of-view; the mindstyle of a character shaping or overlaying the everyday thoughts and feelings of the reader; the social rather than the perceptual. On the one hand this is consistent with accounts of auditory verbal hallucinations that emphasize their intrusive and uncontrollable nature (Badcock, Waters, Maybery, & Michie, 2005) along with their social and agentic characteristics (Wilkinson & Bell, 2016). On the other hand, this would also be consistent with psychological approaches to reading fiction that highlight its interrelation with theory-of-mind (Kidd & Castano, 2013; Mar, Oatley, & Peterson, 2009), and in particular empathy (Mar & Oatley, 2008). One caveat, however, is that accounts of both mindstyle and crossing were described less like controlled simulations of others’ minds, and more like a habit of thinking or expectation of what a character would say in a given situation. If so, this is arguably more similar to the generation and maintenance over time of personality models and agents (Hassabis et al., 2014) than a deliberate and focused act of empathizing or reasoning about another’s mental state (Djikic, Oatley, & Moldoveanu, 2013). One could speculate that the creation of such a ‘consciousness frame’ serves to blur the lines between self and other, leading to the cross-activation of fictional experiences in the reader’s actual world. The notion of experiential crossing could also be seen as a counterpart to the more widely studied relationship between mental simulation and reader’s previous experiences. If, in experiential crossing, experiences move from the fictional to the real, simulation during reading is activated and supported by the reader’s previous experiences of real-world scenarios – an experiential baggage that has been referred to previously as “repertoire” (Iser, 1980), “encyclopedia” (Dolezel, 2000; Eco, 1984), or “experiential background” (Caracciolo, 2014; Herman, 2004). The creation of such repositories of real-world experiences bears similarities with simulation theories of social cognition that propose the construction of a biographical database on which judgements about other minds are based (Harris, 1992). For example, Green (2004) found that undergraduate participants with personal experience or knowledge of the key themes of a story (a homosexual man attending a university reunion) reported greater transportation into the story and, correspondingly, tended to have beliefs that were consistent with the story ideas. A recent study by Chow et al. (2015) has also found preliminary evidence that past experience of particular scenes and actions directly modulates functional connectivity of visual and motor brain areas during story comprehension. The role of prior knowledge and experience was most evident in readers’ accounts of internal and external blending of voices, in which participants described drawing on their own voice or others’ to create the voices of characters in the text. In blending, readers seemed to integrate, or “compress” (Turner and Fauconnier & Turner, 2003), their own voice with a textually cued voice (internal blending) or a textual voice with an external source (external blending). The potential differences between internal and external blending remain to be explored: internal blending may depend on personal identification with characters, while external blending, in contrast, could involve an almost unconscious selection of external sources, based on their similarities with the character as described (e.g., in the behavior, beliefs, bodily features, and so on). The fact that – for some participants – the eventual voice could definitely be “right” or “wrong” (as described in cases of dissonance) suggests that representations of characters are quite robust once they are formed. What role character tropes, stereotypes, and prototypes play in this process will be an important avenue of future investigation. Limitations ~~~~~~~~~~~ It is important to acknowledge some limitations of the present study. First, it is dependent on participants’ self-reports about their general reading experience, which contrasts with experimental approaches that use specific texts and more objective methods of reading engagement (such as eye-tracking). What these data can say about the readers’ experiences of voices and characters is therefore limited by the ability of readers to report on their experiences accurately, and what experiences they choose to describe. Some have argued that readers cannot be relied upon to report on their own experiences accurately and reliably (see Caracciolo & Hurlburt, 2016, for a discussion), but it should be noted that a large body of psycho-narratological work on fiction ultimately depends on participants’ self-reported responses to a story, or other individual differences (e.g., Green, 2004; Mar et al., 2009). Moreover, our intention here was to survey the general reading experience for voices and characters, rather than select a specific text. Individual texts may be particularly evocative of those qualities, but they do not necessarily say anything about how common such experiences are in general reading. Second, in addressing the question of how readers may be said to “hear” the voices of characters, our study specifically used those terms (voice and character) to probe participants’ experiences. Similarly, our choice of other measures to include reflected this focus, as inner speech and hallucination-proneness are prima facie factors that were important to consider. The consequences of this are twofold: (a) it may have led participants to describe predominantly auditory and agent-like experiences over other, more vivid factors (such as visual environments), and (b) it means that other potentially relevant factors, such as transportation and absorption (Kuiken et al., 2004) were not measured here. We would argue that these data are nevertheless informative to understanding the reading experience: voices and characters are prominent in psycholinguistic research (Gerrig, 1993) and narratological approaches to reading (Bortolussi & Dixon, 2003; Herman, 2013; Miall, 2011; Palmer, 2004; Vermeule, 2009), and we note that many of our respondents still described their experience in a much broader fashion (including narrators, visual imagery, and tactile sensations). Nevertheless, it will be important for future phenomenological work to both broaden the scope of inquiry and consider a greater range of individual differences. The role of scene construction, for example, would be expected to play an important role in mental simulation during reading. However, with few exceptions (see “Other”, Table 3) specific references to contextual elements did not feature strongly or explicitly in readers’ descriptions of voice and character (Hassabis & Maguire, 2009). Finally, it is important to recognize the nature of the sample surveyed. Having been promoted by an international reading festival and a newspaper with global readership, the sample surveyed is much larger and more diverse than most studies of fiction and narrative, which primarily use undergraduate samples. However the sample is still not necessarily representative of a general population, as it is likely skewed towards a well- educated, highly literate readership who may be passionate about fiction (71% reported enjoying reading classics) but engage very little with other genres (only 4% read books on sport). It was also completed by three times as many women as men. As such, the extent to which these results can be generalized to the wider reading population is limited. Despite these concerns, the present study is to our knowledge one of the only large-sample phenomenological surveys of the reading experience. We argue that the combination of questionnaire items concerning vivid reading experiences in general and more detailed, qualitative analysis of free-text answers offers a broad and varied resource on what it is like to engage with voices and characters in a text. Accounts of the kind collected here pave the way for further analysis of the role played by the reader’s “experiential traces” (Zwaan, 2008) in relation to the simulation of voices and personification. Cases of experiential crossing, on the other hand, suggest a certain kind of rebound, in which fictional agents and worlds are activated in real-life scenarios. And while our data support the existence of vivid voices and characters in the experience of readers, they ultimately highlight a wide range of quasi-sensory qualities and simulatory dynamics in the reading process. As in debates on mental imagery (Kosslyn et al., 2006), the reading processes described here seemed to be highly varied, ‘experiential’ (Fludernik, 1996), and idiosyncratic. Even if readers’ own accounts are taken as mere indicators of the underlying cognitive and perceptual phenomena, they suggest that the processes by which people produce and experience voices and characters could be very different for different individuals. In this light, the endeavor to describe a typical or normative response to a text becomes perilous, risking an attempt to control the “entropy” of the real experience that individual readers are having (Iser, 2001). The voices and characters of a text are many and various; this would seem to also be true for the reader."],["Objectives: To report the theory-based process evaluation of the Bristol Girls' Dance Project, a cluster-randomised controlled trial to increase adolescent girls' physical activity. Design: A mixed-method process evaluation of the intervention's self-determination theory components comprising lesson observations, post-intervention interviews and focus groups. Method: Four intervention dance lessons per dance instructor were observed, audio recorded and rated to estimate the use of need-supportive teaching strategies. Intervention participants (n = 281) reported their dance instructors' provision of autonomy-support. Semi-structured interviews with the dance instructors (n = 10) explored fidelity to the theory and focus groups were conducted with participants (n = 59) in each school to explore their receipt of the intervention and views on the dance instructors' motivating style. Results: Although instructors accepted the theory-based approach, intervention fidelity was variable. Relatedness support was the most commonly observed need-supportive teaching behaviour, provision of structure was moderate and autonomy-support was comparatively low. The qualitative findings identified how instructors supported competence and developed trusting relationships with participants. Fidelity was challenged where autonomy provision was limited to option choices rather than input into the pace or direction of lessons and where controlling teaching styles were adopted, often to manage disruptive behaviour. Conclusion: The successes and challenges to achieving theoretical fidelity in the Bristol Girls' Dance Project may help explain the intervention effects and can more broadly inform the design of theory-based complex interventions aimed at increasing young people's physical activity in after-school settings. --------------------------------------------------------------------------------","Young people become less active during the transition from childhood to adolescence (Nader, Bradley, Houts, McRitchie, & O'Brien, 2008). Girls are less active and experience a steeper decline in activity than boys (Nader et al., 2008). In England, the majority of adolescent girls do not meet the government's recommendations of a minimum of 60 min of moderate-to-vigorous physical activity (MVPA) per day (Joint Health Surveys Unit, 2013). As physical activity is associated with physical and mental health (Janssen & Leblanc, 2010), identifying ways to encourage more girls to be active more often is a national (Department of Health, 2011) and global (World Health Organisation, 2004) health promotion priority. A recent meta-analysis has shown that physical activity interventions for girls are more effective if they exclude boys, are delivered at school and are based on an underlying theory of behaviour change (Pearson, Braithwaite, & Biddle, 2015). Dance is a popular activity amongst girls (O'Donovan & Kay, 2005) and proliferates contemporary culture and media consumed by young people such as music TV, talent shows and singing contests. Dance can be an enjoyable form of cardiovascular exercise in which girls develop their co-ordination, acquire new skills, work independently and in groups and develop friendships and self-expression (Australian Women Sport and Recreation Association, 2010). Dance is an alternative to traditional/competitive sports offered to girls and we have previously highlighted the potential of a dance-based physical activity intervention for adolescent girls: the Bristol Girls Dance Project (BGDP) (Jago et al., 2013, 2012, 2011; Powell, Carroll, Sebire, Haase, & Jago, 2013). The BGDP was a cluster-randomised controlled trial designed to examine the effectiveness and cost-effectiveness of an after- school dance-based intervention in increasing the MVPA of Year 7 girls (aged 11–12 years). Theoretical foundations of BGDP ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Underpinning interventions with behavioural theory is hypothesised to increase their effectiveness (Baranowski, Anderson, & Carmack, 1998; Craig et al., 2008). In addition, theories allow intervention developers to target activities at theoretically-derived mediators (Baranowski et al., 1998). The BGDP intervention was based on self-determination theory (SDT) (Deci & Ryan, 2000; Ryan & Deci, 2007) because its theoretical foundations are concerned with how the psychological and socio-environmental conditions (e.g., created by a dance teacher) can support individuals' motivation (Fortier, Duda, Guerin, & Teixeira, 2012). Motivation quality ~~~~~~~~~~~~~~~~~~ According to SDT, an individual's motivation for a behaviour such as dance, can be more or less self-determined and six different types of motivation are hypothesised to be differently associated with behaviours such as physical activity and related cognitive and affective outcomes (Ryan & Deci, 2007). The more self-determined types of motivation (i.e., intrinsic motivation, integrated & identified behavioural regulation) are broadly grouped as autonomous. Intrinsic motivation is based on the inherent satisfaction or enjoyment that accompanies a given behaviour. The other forms of autonomous motivation are extrinsic in nature and involve undertaking a behaviour for a reason other than its inherent satisfaction. Integrated regulation is where a person aligns their engagement in a behaviour with their broader self (e.g., seeing being active as part of one's identity) and identified regulation represents motivation which is driven by a valued outcome such as health benefits or making new friends. The less self-determined types of motivation (i.e., introjected & external regulation) are broadly grouped as controlled motivations. Introjected regulation refers to motivation based on internalised pressures such as avoiding feelings of guilt, whereas external regulation is characterised by prods and pushes which are external to the person such as complying with demands or avoiding punishments. Previous research suggests that more autonomous physical activity motivation is positively associated with child and adolescent physical activity (Owen, Smith, Lubans, Ng, & Lonsdale, 2014; Sebire, Jago, Fox, Edwards, & Thompson, 2013) and positive psychological outcomes such as quality of life and physical self-concept (Standage, Gillison, Ntoumanis, & Treasure, 2012). On the other hand, adolescents' controlled motivation for exercise has been shown to correlate negatively with health-related quality of life and functioning within physical, social, school and emotional domains (Standage et al., 2012). Fostering high quality motivation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A cornerstone of SDT is that autonomous motivation is developed when people feel that their psychological needs for autonomy (i.e., feelings of volition and free will), competence (i.e., feeling capable to perform challenging tasks) and relatedness (i.e., perceptions of belonging & meaningful connections with others) are fulfilled (Deci & Ryan, 2000). This hypothesis is supported by empirical research among children (Sebire et al., 2013), adolescents (Van den Berghe, Vansteenkiste, Cardon, Kirk, & Haerens, 2014) and adult dancers (Quested & Duda, 2010). Within SDT, people's psychological needs can be supported or undermined by the motivational climate that an authority figure (e.g., dance instructor) creates through their motivating or teaching style (Deci & Ryan, 2008; Su & Reeve, 2011). Need supportive styles are underpinned by the provision of autonomy support, structure and involvement which is reflected in how teachers' (or dance instructors) conduct their classes and interact with pupils (Haerens et al., 2013; Su & Reeve, 2011). When teachers provide autonomy support they give meaningful rationales (especially for tasks which are important but not as enjoyable as others), offer choices which pupils value, seek and acknowledge pupils' perspectives or ideas and nurture pupils' internal motivation, interest and enjoyment. In contrast, controlling teachers aim to motivate pupils by either inducing internal pressures such as guilt, or external pressure such as a deadline and feedback given and language is used to manipulate rather than be informative. Such strategies are likely to frustrate rather than support pupils' psychological needs (Bartholomew, Ntoumanis, & Thogersen-Ntoumani, 2009). Teacher's provision of structure is primarily related to supporting pupil's competence. A well-structured class is where clear expectations are set out before tasks and during tasks, guidance, direction and positive effect-based feedback is given. Without structure, a learning environment can be described as chaotic where students do not know what they should do or what is expected of them (Vansteenkiste et al., 2012). Pupils' relatedness is supported when teachers' are involved by showing the pupils empathy and genuine interest in them (Aelterman, Vansteenkiste, & Van Keer, 2013; Haerens et al., 2013; Su & Reeve, 2011). In contrast, a lack of involvement by teachers will frustrate relatedness. Amongst children, Physical Education (PE) teachers' use of need-supportive styles has been shown to be associated with their pupils' psychological need satisfaction and autonomous motivation for PE (Ntoumanis & Standage, 2009; Van den Berghe et al., 2014). Design of the BGDP ~~~~~~~~~~~~~~~~~~ We have previously reported the study protocol (Jago et al., 2013) and outcome paper (Jago et al., 2015). The study involved 571 Year 7 girls (aged 11–12 years) from 18 schools from the greater Bristol area allocated at the school-level to intervention (n = 9) and control (n = 9) arms. The intervention consisted of 40, 75-min after-school lessons that took place, twice per week for 20 weeks at school and were led by 10 professional dance instructors between January and July 2014. Girls were provided with a dance diary which they could complete and hand in to the dance instructor at the end of each lesson, in which they could record what they had learnt, their feelings and thoughts. Instructors were provided with a manual which provided plans for all 40 lessons in addition to training outlined below. One instructor was unable to complete the full intervention and was replaced at the intervention mid-point with another instructor. One instructor taught in two schools (ID numbers 21 and 51). Embedding SDT within the BGDP intervention ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ BGDP aimed to increase girls' autonomous motivation for both dance and physical activity and this was targeted through the BGDP instructor manual, the lesson design and content, and dance instructor intervention training (Table 1). Dance instructors received a one day training session (5th December 2013) prior to the start of the intervention (13th January 2014) which included 2 h on SDT (delivered by SJS) highlighting the key features of the training manual and how the theory could be applied in dance lessons. The content of the manual and the training focussed on how to provide autonomy, competence and relatedness support and intertwined involvement and structure as ways to achieve this. Comparisons were made between using these need-supportive strategies and more controlling practices. Instructors were given the opportunity to practice using need-supportive techniques by role-playing different dance activities, asking questions and receiving feedback. At the mid-point of the intervention, instructors attended a half-day top-up training session where the SDT components were revisited and instructors shared their experiences of delivery to resolve any problems. Results of BGDP and the need for a theoretical process evaluation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The BGDP intervention was not effective in increasing girls' MVPA at the end of the intervention period when dance lessons were still running or at 12 months after baseline (Jago et al., 2015). Following the intervention, autonomous and controlled motivation and perceptions of competence and relatedness were lower among intervention versus control group participants. This was unexpected and highlighted the importance of using process evaluation to further understand these findings and the motivational processes at play. The part of the BGDP process evaluation which focussed on the dose of the intervention received (i.e. the number of lessons attended) and evaluation of issues pertaining to attendance and retention identified that between 37 and 40 dance lessons were delivered in all schools (Sebire et al.,). Attendance (M = 12.8, SD = 7.0) at dance lessons declined over time and approximately nine girls per school out of a possible 33 received the intervention dose (defined as two thirds of the lessons offered). Participants enjoyed taking part and reported social and physical health benefits. Instructors valued the intervention training and identified that variable attendance challenged delivery but also facilitated the development of a core group of committed participants. Multiple process evaluation papers are justified to make full use of all of the data collected (Moore et al., 2014). Given the large volume of data produced from the detailed process evaluation of BGDP the aim of this paper is to report a process evaluation of the BGDP study with particular focus on theoretical fidelity. Process evaluation can play a crucial role in illuminating how theory-based intervention components are experienced by the participants and importantly how they are received, interpreted and implemented in practice by intervention deliverers. The findings can then be used to interpret findings of intervention effects in greater detail and refine how theory can be best operationalised in practice. Specifically we sought to use both quantitative and qualitative methods to evaluate: (a) the degree to which dance instructors adopted a need-supportive style, (b) participants' perceptions of the dance instructors' teaching practices, (c) participants' qualitative perceptions of satisfaction of autonomy, competence and relatedness and (d) dance instructors' experiences of delivering an SDT-based intervention. Qualitative data collection At the end of the intervention and within three weeks of the final dance lesson (delivered 11th July 2014), the 10 dance instructors participated in individual semi-structured interviews with JMK (mean duration = 67.2 min, range = 41.4 to 91.4). An interview guide was followed which covered instructors' experiences of the intervention training and using the intervention manual (e.g., “Were you able to include any of the motivational ideas that were in the manual and training day?”), delivering the intervention in a need-supportive way (e.g., What strategies did you use to motivate the participants?), their relationship with the participants (e.g., What was your relationship with the girls like? Did it change?) and the challenges they faced (e.g., What didn't work as well during the dance sessions and why?). A focus group was conducted with participants from each intervention school (mean duration = 42.4 min, range = 30.4–50.2 min). Girls were sampled to reflect tertiles of intervention attendance within each school to ensure a range of views were elicited (total n = 59; n = 25 high, 16 moderate & 18 low attenders respectively) and focus groups ranged in size from 3 to 8 participants. Of particular relevance to this paper, the focus group topic guide explored participants' views on and relationship with the dance instructor (e.g., Is there anything you would change about your dance instructor's teaching style?), and perceptions of autonomy, competence and relatedness (e.g., Autonomy: Do you think you had some control over what you did? Competence: How did you find the dance sessions physically? Relatedness: Did your relationships with one another change as the weeks went on?). Importantly, questions aimed at exploring the theoretical elements focused on broad topics through which participants were given the opportunity to express their views. Observed need-supportive teaching style Dance instructors' fidelity to the need-supportive intervention was assessed by rating the delivery of four intervention lessons per instructor which were randomly-selected within four intervention blocks (one lesson from each of weeks 5–12, 13–20, 21–29 & 30–36). Instructors wore an audio recording device on their arm with a microphone attached to their clothing. The observation tool developed by Haerens et al. (Haerens et al., 2013) was used. Although the tool was originally designed to rate videos of PE teachers, due to ethical constraints we adapted the method to combine rating of audio recordings of lessons with real-time observation. The instrument contains 21 items measuring Relatedness support (5 items), Structure before the activity (5 items), Structure during the activity (6 items) and Autonomy support (4 items). One item (“encourages pupils to persist”) was excluded based on low validity within the validation study (Haerens et al., 2013). Fifteen items were rated from the audio recordings as they did not rely on visual observation and five items (e.g., Is physically nearby the pupils) were rated by direct observation. The frequency of each teaching practice occurring was rated using a four-point scale (0 = Never, 1 = Sometimes, 2 = Often, 3 = All the time) for each 5-min segment of the lesson. Observations and ratings were undertaken by a researcher (JMK) who had discussed the definitions and meaning of the higher order constructs and scale items in depth with (SJS) and undertaken a pilot observation/rating. JMK and SJS listened to the pilot recording separately, then together using the observation tool to guide a discussion of the presence of each teaching behaviour in 5 min segments. To check consistency in interpretation, six audio recordings (excluding the items requiring visual observation) were coded by both (JK) and (SJS). Interrater reliabilities (indicated by intraclass correlations; ICC) (Rabe-Hesketh & Skrondal, 2008) were: Relatedness support (ICC = 0.01, poor), Structure before the activity (ICC = 0.31, fair), Structure during the activity (ICC = 0.57, moderate) and Autonomy support (ICC = 0.41, moderate). Internal consistency reliability estimates were as follows: Relatedness support (α = 0.58), Structure before the activity (α = 0.73), Structure during the activity (α = 0.62) and Autonomy support (α = 0.60). Only the coding of one researcher (JK) was used in data analysis. Participants' perceptions of dance instructor autonomy support At the end of the intervention, participants (n = 281, 98.9% of participants randomised to intervention arm) reported their perceptions of the dance instructor's provision of autonomy-support using an adapted version of the Sport Climate Questionnaire (Amorose & Anderson-Butcher, 2007). Six items (e.g., My Active 7 dance instructor listens to how I would like to do things) were scored using a 7 point Likert scale (1 = Strongly disagree to 7 = Strongly agree). Items were averaged to create a perceived autonomy-support score (α = 0.94). Qualitative data analysis Transcripts of interviews and focus groups were analysed by four researchers (JMK, MJE, TM & SJS) using the Framework Method (Gale, Heath, Cameron, Rashid, & Redwood, 2013) using both inductive (i.e., themes arising from the data) and deductive (i.e., a priori SDT-based themes were interrogated within the data) coding strategies. Deductively, the SDT-related components (i.e., motivation types, psychological needs and instructor teaching style) were defined based on extant literature and these definitions were discussed amongst the group of four analysts. Second, text within the transcripts which related to these definitions (including both positive and negative experiences) were collated. Frameworks created from the initial transcripts formed the basis of analysis for the remaining transcripts and flexible to nuances, new examples and refinements which were discussed at regular group meetings. Third, data relating to SDT were interpreted and refined by SJS and discussed and agreed within the team. A framework for dance instructors and girls was created and a convergence coding matrix which allowed comparison of themes between dance instructors and girls was developed in NVivo (Version 10, QSR International Pty Ltd) (Farmer, Robinson, Elliott, & Eyles, 2006). The application of trustworthiness criteria (Shenton, 2004) included credibility, transferability, dependability and confirmability and are reported elsewhere (Sebire et al.,). The qualitative data were organised into six themes which combined deductive (e.g., we asked questions about perceptions of relatedness) and inductive (e.g., discussions emerged about challenges of delivering the SDT components) results from the instructors and the girls. The themes were: (1) Dance instructor training and acceptance of the intervention theory, (2) Autonomy support and perceptions of autonomy need satisfaction, (3) Dance instructors use of controlling strategies, (4) Competence support and perceptions of competence need satisfaction, (5) Relatedness support and perceptions of relatedness need satisfaction and (6) Challenges of delivering an SDT-based physical activity intervention for children. Quotes are reported using linked dance instructor and school ID numbers (i.e., Instructor 21 delivered lessons in school 21) and are the same as those used in the other process evaluation papers from this study to facilitate cross-referencing. Quantitative data analysis For the observed need-support, scores for each item at each 5-min segment were aggregated to give lesson item mean averages which were then combined to form average scores for relatedness support, structure before the activity, structure during the activity, and autonomy support, for each of the four lessons. Means and SDs for each construct over the four observations were calculated and analysed descriptively. Means and standard deviations (SD) were calculated for girls' perceptions of each of the dance instructor's autonomy supportiveness and were analysed descriptively. Quantitative results ~~~~~~~~~~~~~~~~~~~~ Relatedness support was the most highly scored (between “often” & “all the time”) teaching behaviour amongst all instructors (mean = 2.29, SD = 0.47) (Fig. 1). In general, dance instructors provided moderate (between “sometimes” & “often”) structure before and during the dance activities (structure before, mean = 1.73, SD = 0.54; structure after, mean = 1.53, SD = 0.60). Structure was observed less in the lessons led by instructor 42. Autonomy-support was, for all but one instructor, the lowest scoring teaching practice (mean = 1.16, SD = 0.54) reflecting provision of autonomy support only “sometimes”. Pupil- perceived autonomy support was moderate (mean = 4.68, SD = 1.68) and relatively consistent between instructors (Range: Instructor 51 mean = 4.32 SD = 1.65 to Instructor 21 mean = 5.53, SD = 1.20). Dance instructor training and acceptance of the intervention theory Dance instructors reflected positively on the training and believed that the principles of SDT were appropriate to underpin the dance lessons. Most instructors believed that their existing teaching style was aligned with the SDT approach however one instructor felt that the autonomy-supportive style challenged her existing practice, particularly the language she used: A lot of us are quite experienced teaching and you can get into a groove with how you teach and [the introduction of SDT] really made you challenge those sort of key phrases that you say throughout the class. (Dance instructor 23) Autonomy support and perceptions of autonomy need satisfaction Dance instructors reported providing girls with choice within dance lessons, including the music, dance styles, choreography and warm-ups which was corroborated by some girls. [The dance instructor] asked us what types of things we wanted to do. Some people said contemporary, some people said breakdancing [ … ] so that's what we did, which was good. (Focus group 53) Girls choosing the music was an important source of ownership/autonomy within lessons as it made them more engaging and positively influenced activity: If it's music they like then they want the music on all the time … they're going to be more active and more involved. It just makes perfect sense to … let them have that choice in the music and it motivates them more. (Dance instructor 42) Instructors were encouraged to support participants' autonomy within a clear structure which was developed in the early lessons by involving girls in the development of group rules. They also reported responding to feedback from the girls and attempting to include the views of the group not just a vocal minority (although some children argued to the contrary). I would read their [dance] diaries and sometimes they would write things in there, either about the session that would give me clues as to what … you know, 'oh, I loved this game' and you think oh, I didn't realise you loved this game. OK let's do this game more. (Dance instructor 42) Generally girls enjoyed the level of autonomy they were granted. However, some stated a need for dance instructors to balance autonomy with sufficient instruction and supervision to support their engagement: We had to do it by ourselves and I didn't know the counts or anything and I had to like tell them [others in her group] what to do and I didn't like that …. (Focus group 61) Dance instructors' use of controlling teaching strategies While the majority of dance instructors reported using autonomy-supportive styles, participants from several schools (one in particular) described controlling teaching strategies. The following example is where one Dance instructor (DI 53) covered another's (DI 21) lesson: Participant 5: It was like the army. Participant 2: She forced you to do handstands. Participant 6: And if you were talking or something she would make you do ten press-ups. (Focus group 21) Some girls commented that they “had no say in pretty much anything” (Focus group 32) and identified that where choices given they were not perceived as genuine: She was asking us to choose a dance and then she'd choose a dance herself. (Focus group 21) The frequency and length of drink breaks was considered important to participants' autonomy, but was rarely mentioned by instructors, other than as creating an opportunity for disruption which they had to control by being what participants felt was “strict”. Girls rationalised some of their dance instructor's controlling behaviour as being driven by a desire to avoid group arguments or encourage dedication: She was strict because she wanted you to be dedicated and turn up. (Focus group 53) Competence support and perceptions of competence need satisfaction Instructors reported using numerous competence-supportive teaching strategies including affording participants with the required dance skills, using peer role models, differentiation of dance sequences, encouraging self-reflection, giving opportunities for leadership, and providing constructive feedback. I think the easiest way to deal with the [different skill] levels [is to] get those girls who are working really well to help other girls that … are struggling. (Dance instructor 21 & 51) Several approaches were used to encourage girls to reflect on their competence, including using the dance diaries, reflecting on their own progress and ensuring that this reflected genuine progress: [A participant would say] “I can't do it” and I'm like “well, first of all you can do it”, but also … “remember that step that you couldn't do a couple of weeks ago?” and she's like “oh yeah, I can do it well easy now” and I'm like “well, there you go then” … I'd get them to reflect on their own progress and then I didn't have to try and pretend. (Dance instructor 21) Instructors' awareness of girls' abilities allowed them to provide targeted competence support: There was a bigger girl who came quite a lot and she found - or I found when I was teaching the structure [choreography] she would struggle and give up a lot easier, because I think she felt like she couldn't do it. She loved teaching the warm-ups and I found that she worked harder, she got sweatier, she pushed herself more because she was confident doing the things that she already knew how to do. (Dance instructor 53) Girls corroborated the dance instructors' competence support and reported receiving individual and group-level assistance: She helped like if you were stuck on something. But she helps more like as a whole group whereas [a different instructor that the group had] would sort of just help you individually. (Focus group 23) Girls reported increased confidence and competence to dance which was also observed by the dance instructors: After 30 seconds the first time we were like tired and like couldn't do it. Then after a few sessions … well, not a few but like half way through, we could do it for like ten minutes, five minutes. (Focus group 53) In contrast, some participants felt a lack of support when learning more complex skills and thought that the instructor was not aware of the varied competence of the group members. Girls suggested that they could not control the speed at which lessons progressed: If we didn't know how to like do the move, like it was a bit hard to ask [dance instructor] to show us to do the move again because she was already showing the next bit. (Focus group 32) Relatedness support and perceptions of relatedness need satisfaction All dance instructors referred to using strategies to build trusting relationships with and between girls, including asking them about their lives outside the intervention, responding to comments written in dance diaries, asking after girls when they missed lessons, giving regular high-fives, using a ‘head-to-head sharing time’, and discussing non-attendance: I had this one sort of thing where we lie on the floor with all our heads together, and each say one thing about the session that we felt we did really well or it could be one thing that someone else did well … (Dance instructor 32) Dance instructors and girls reported feeling a strong trusting relationship which developed over the course of the intervention: I had a couple of girls really open up to me and talk to me about sort of personal problems that they were having. (Dance instructor 23) There were a lot of “twelve year old teenage dramas” and people would get upset about 'oh no, my friend doesn't like me, oh!' and then [dance instructor] would be like 'right, we're going to dance this out' [… ] or get them to apologise. (Focus group 42) In general, girls considered dance instructors to be enthusiastic, fun and understanding. She was really nice because we came in and she was like 'oh, you're the dancers!' We were like 'oh yeah'. And she was really nice. She came in and like introduced herself and everything. And then she … if one … like some of us is like injured or doesn't really want to do dance then she'll let us sit out and then just like come back when we feel like it so she's really nice. (Focus group 61) However some girls did not feel a genuine connection with the dance instructor: She had to be in charge all the time. If she had kind of like stepped back and been more of a friend than someone like in charge of us then I think we'd have all found it easier. (Focus group 32) Where dance instructors' comments or actions were not perceived as genuine, this undermined the participant's connection with them: Yeah, [dance instructor] always like really clapped for them. She always clapped for us but like, you know, it was like for them it was, it was like a proper clap. (Focus group 61) Girls and dance instructors described the development of existing friendships and the formation of new ones over the course of the intervention: “We bonded together” (Focus group 42). At the start of the intervention, many girls were apprehensive and groups were fractured or consisted of existing cliques. However, throughout the project, these cliques dissolved and participants reported making friends and feeling more socially comfortable. Participant 1: In the first two sessions I was really shy and I always went to the back and I didn't really say [ … ] much. But now in the Active7 sessions I talk quite a lot [… ] because I got to know quite a lot of people … Participant 2: You feel more comfortable around them. (Focus group 72) However, some girls experienced a lack of connection with their peers which seemed rooted in participants' interpersonal comparisons of their dance abilities and divides between the ‘confident’ and ‘shy’ participants. The girls who already do dance are really like strong about it and they always go together in a group, they don't share it. (Focus group 32) Challenges of delivering an SDT-based physical activity intervention for children Two issues that appeared to challenge the dance instructors' theoretical fidelity in the intervention were the management of disruptive behaviour and the use of end-of-project performances. Some dance instructors found it difficult to be autonomy-supportive when faced with disruptive behaviour: There are times, as I said before, when you've got 25 plus of them all going a bit mental [ … ] then you do have to sort of change tactics unfortunately, but generally speaking it [being autonomy-supportive] would be the way I would want the sessions to be. (Dance instructor 42) In one case a dance instructor felt conflicted between maintaining high fidelity (using the SDT-based guidelines) and keeping control of the discipline: They were running wild and I was trying to be, you know, use the ABC [Autonomy, Belonging, Competence] and it was very hard to try and keep to that. Really, really difficult because they were just testing my limits and going crazy [ … ] I think that's because I was so worried about sticking to ‘this is what we had to do’ to then kind of trying to actually respond to the children themselves. (Dance instructor 21 & 51) Some instructors reflected that more role-play based learning in the training would help prepare them for dealing with this more effectively: I would say it would be good to … set up some situations where that skill [referring to need supportive teaching] could be practised because I felt frustrated with myself sometimes that I didn't know [ … ] what to say and I didn't know how to do it. (Dance instructor 23) A further challenge to theoretical fidelity concerned the use of dance performances as a motivational tool and whether or not the girls wanted to perform in front of others. Girls were given choice as to whether they performed a dance in front of an audience or not and groups often chose to perform. Dance instructors considered performances to be an important motivational element of a dance programme: You've almost got to have like an end plan that they're all working towards (Dance instructor 62). Dance instructors also considered performances to be desired by and a positive experience for most girls, although several reflected on performance-related anxiety: Even if they're slightly petrified of it or whatever, you know … it's kind of a good fear [ … ] They did all love doing this as well, even though all moaned a bit but they loved it. (Dance instructor 62) Some girls liked the motivation and focus a performance provided: At least it's [a performance] for something [ … ] Because if you know you're not doing it in front of lots of people you kind of lack a bit, but if you know you're doing it for, in front of people then you know that you've got to try and do your best and try and get the steps right. (Focus group 23) Whereas for others performing was a source of anxiety: When we were doing our performance … [dance instructor] wasn't there like … to support us in a way, she didn't come to it, so like we didn't really know the music and we had to do that ourselves … some people [were] really nervous and [saying] they weren't going to do it. (Focus group 61) The performance element also was a source of pressure for the dance instructors and one identified that it negatively affected her teaching: One session I had to get a bit strict and ended up getting a bit sort of arsy [… ] I think it was the pressure of the performance. (Dance instructor 23)","In this paper we report a theory-based evaluation of the BGDP intervention using a mixed methods approach. The results can be used to better understand theoretical fidelity and shed light on the results of the trial. The majority of instructors believed that the SDT principles of the training aligned with their teaching styles and methods. The training was well received and served as a reminder to instructors to focus on the “how” (i.e., communication practices) in addition to the “what” (i.e., dance content) of their teaching. A previous study among PE teachers (Aelterman et al., 2013) identified similar acceptance of SDT-based intervention training and research has shown that classroom teachers' beliefs that implementing an SDT-based teaching style is likely to be effective, easy and normal/usual are associated with their motivating style (Reeve et al., 2013). We believe that our findings indicate that the dance instructors involved in BGDP did buy-in to the SDT teaching style, and believed that it would be effective however the reality and ease of implementing it in practice with the participant group led to variable fidelity as discussed below. Dance instructors qualitatively reported providing choice (of music, dance styles, warm-up activities, and choreography) within sessions that reflected examples given in the training manual (this therefore also indicates fidelity) and this was perceived by instructors and girls to make lessons more engaging and enjoyable. However, this provision largely reflected option choice, and did not appear to provide action choice, such as having control over the pace of task progression which is a central element of autonomy-support (Reeve, Nix, & Hamm, 2003). Previous research suggests that providing action choice promotes self-determination and intrinsic motivation more than option choice (Reeve et al., 2003). Instructors that mainly provided option choices may have believed that this was sufficient autonomy-support and neglected action choice. The quantitative need-support ratings corroborate this finding and suggest low provision of autonomy-support across all instructors and substantial room for improvement. This finding is consistent with Haerens et al. (Haerens et al., 2013) who reported that autonomy- support was the least common need supportive practice amongst PE teachers relative to other practices. Further, the qualitative results highlighted the importance of combining autonomy-support with structure as some participants felt uncomfortable when they were left to practice on their own with insufficient instructions. Previous work has shown that teaching styles which combine autonomy-support and structure are associated with improved learning, behavioural and motivational outcomes amongst adolescents (Vansteenkiste et al., 2012) and highlight the importance of ensuring that interventions are able to help teachers balance these two teaching dimensions. In addition to instances of low autonomy- supportiveness, the qualitative findings identified that some dance instructors used somewhat controlling motivational practices. Instructors may have used controlling strategies for a number of reasons; first, the intervention training may not have successfully changed their teaching styles to be more need-supportive and they adopted their usual practices which included controlling techniques. However, the instructors reported believing that the philosophy of SDT chimed with their usual teaching practices which suggest that this may not have been the case. The SDT component of the dance instructor training was comparable in duration to previous training for PE teachers (Aelterman et al., 2013), but shorter than others (Aelterman, Vansteenkiste, Van den Berghe, De Meyer, & Haerens, 2014). Instructors reported wanting more time to practice implementing the different motivating techniques, suggesting that the training may have been conceptually clear but not sufficiently practical. Previous studies of SDT-based training have incorporated videos of real teaching scenarios (Aelterman et al., 2013, 2014) which can be used to identify and reflect on real teaching practices and dedicated more time to practicing motivating strategies. Second, some dance instructors may have misinterpreted the theory, confusing autonomy-support with a lack of rules or structure. The need-support ratings suggested that structure was used with moderate frequency which provides some evidence for this hypothesis and may have led to more disruptive behaviour and reversion to the use of controlling strategies. Third, some dance instructors may have adopted more controlling practices in response to the challenges associated with teaching large classes of beginners in a school environment. For example, one dance instructor drifted (Bumbarger & Perkins, 2008) from the need-support foundation by using press ups as punishment for talking. Others used working towards a dance performance as a motivational lever which while seen positively by some, was also a source of pressure for others and their dance instructors. Moving beyond dance, to broader physical activity interventions which rely on trainers leading groups of young people, our findings suggest that future work is needed to identify whether and how physical activity intervention deliverers use controlling strategies. Further, and based on the nature of trainer's existing practices, it is important to ascertain whether they can be equipped with techniques to use in response to challenging behaviour without resorting to controlling techniques. The findings highlight a need for those developing theory-based interventions to identify innovative ways to communicate theoretical nuances (e.g., action vs. option choice) that can be understood and implemented by practitioners. As theory is sometimes viewed as lacking real world validity (Davidoff, Dixon-Woods, Leviton, & Michie, 2015; Rothman, 2004), there is a risk that efforts to ensure practitioners adopt theoretical principles result in theoretical dilution and more room for drift (Bumbarger & Perkins, 2008) from the intended theoretical targets. Embracing technology within theory-based complex physical activity interventions may be an effective way to overcome some of the issues which in our study were commonly derived from having limited time within training to adequately cover detailed and subtle theoretical nuances alongside other content or intervention deliverers facing challenges within lessons. For example, previous work has supported training with self-study websites which teachers are asked to engage with (Reeve, Jang, Carrell, Jeon, & Barch, 2004) which include videos of real teaching scenarios. A possible next step is to develop smartphone apps to support in-person training, extend the possible training time and provide a hub of resources to support intervention fidelity such as multimedia content, tasks which reinforce learning, tips when dealing with challenges & networking with other instructors. Dance instructors reported confidently providing competence support to girls and used a number of techniques which may have utility in future PA interventions. These included a number of examples of providing structure during the dancing activities such as peer–peer teaching, encouraging self-reflection, providing genuine encouragement and provision of structure through clear instructions before the dance activities. The dance instructors' experience in teaching may explain their confidence in using these strategies and supports the importance of identifying well trained intervention deliverers who can bring beneficial innovation to PA interventions (Bumbarger & Perkins, 2008). Girls also had positive experiences of competence support (and some reflected on their own increased levels of perceived competence). However, in the BGDP trial, perceived competence towards dance and physical activity decreased pre-post intervention in the intervention group (Jago et al., 2015). It is possible that both results are correct; that while some participants did feel more competent after the intervention, girls on average did not. Alternatively, the results could point towards changes in the girls' cognitive representations of dance before and after the intervention. Pre-intervention, the majority of girls had no prior experience of formal dance lessons. Dance competence levels were rated as relatively high; which could reflect informal dance experiences (e.g. dancing with friends or alone). However, after experiencing dance in a more structured taught environment, trying more complex choreography and comparing their ability with that of their peers, it is not unreasonable to assume that their cognitive representations of dance and their frame of reference for their perception of competence could have changed post-intervention. This could mask some of the qualitative perceptions of increased competence. Relatedness support was the most commonly observed need-supportive technique used by the dance instructors which in most cases was verified by the qualitative findings. This is consistent with previous research among PE teachers (Haerens et al., 2013), but the dance instructors in our study provided more relatedness support (2.29/3.00) than previously studied PE teachers (≈1.30/3.00) (Haerens et al., 2013). This may reflect the interpersonal style the dance instructors have developed through their teaching of dance to groups of girls in out-of-school settings and be an indicator that they found this particular dimension of need-supportive instruction easy which has been shown to be associated withteacher's motivating style (Reeve et al., 2013). Previous research suggests that relatedness towards PE teachers is associated with girls' engagement in PE (Shen, McCaughtry, Martin, Fahlman, & Garn, 2012). The qualitative findings identified the development of some strong and trusting girl- instructor relationships. The instructors' use of a number of relatedness-supportive techniques which represent effective intervention innovation (Bumbarger & Perkins, 2008) could be adopted in other interventions (e.g., dedicating time at the end of lessons for the instructor and girls to lie on the floor with their heads together and reflect on the lesson, giving regular high-fives, using the dance diary to guide empathically changing lessons in line with what girls enjoy/don't enjoy). However, it was clear from the findings that for a minority of girls, a sense of relatedness was not formed with the instructor which was commonly caused by perceptions that the instructor-participant bond was not genuine. It would be useful in future interventions to identify the use of effective techniques during implementation and share them amongst the network of practitioners who are finding relatedness support difficult. Whilst extra support for instructors was included in BGDP mid-intervention, a more effective dissemination of teaching techniques or more frequent provision of materials/support (e.g., via an intervention resource such as an app as referred to above) could hold promise in challenging low-fidelity during theory-based PA interventions. Strengths & limitations ~~~~~~~~~~~~~~~~~~~~~~~ The theoretical underpinning of BGDP is a strength that has helped to evaluate where the intervention was consistent with, or drifted from, the intended behaviour change strategies. In addition, the combination of quantitative and in-depth qualitative data collected from both dance instructors and participants has facilitated the development of a detailed picture, strengthened by triangulation between participants and across methods. An inherent limitation in theoretical process evaluations is that the intervention deliverers are aware of the theoretical foundations with which they are asked to underpin their delivery. As such, there is the potential for their interview responses and observed lessons to be biased towards good fidelity. However the focus group results largely added credibility to the dance instructor results and we are confident that we heard a diverse range of perspectives including negative experiences which we have reported. Furthermore it is unlikely that the instructors would have been able to change their teaching style significantly during the four observations, particularly given that the instructors were informed on the day of the lesson that they would be observed. A related limitation is that the inter-rater reliability of the rated dance instructor teaching styles was low for relatedness support and low/moderate for the other dimensions but generally lower than previous work with PE teachers (Haerens et al., 2013). A potential reason for this is that we rated audio rather than video recordings of the dance lessons (as the original measure used) and thus underestimated the information given by physical indicators alongside the audio to make the ratings. However, we only used the ratings of one observer whose ratings were consistent as indicated by good internal consistency estimates. An additional limitation is that we did not measure the dance instructors' perceptions of using an autonomy-supportive teaching style pre- and post-training. This would have afforded us a short-term check of training effectiveness and could be incorporated into future intervention designs. Further, we used a general measure of girls' perception of instructor autonomy-support, which prevented us from examining individual psychological need support and comparing this to our more nuanced observation and qualitative data. Finally, research including applications in the sport and PE domains (Bartholomew, Ntoumanis, Ryan, & Thogersen-Ntoumani, 2011; Haerens, Aelterman, Vansteenkiste, Soenens, & Van Petegem, 2015) has separated the concept of need thwarting (e.g., a child perceiving that their teacher/coach is trying to control or manipulate them) from the experience of low need satisfaction (e.g., that a child does not feel that they have much input in lessons/training sessions). In PE, pupils' perceptions of their teacher's controlling teaching was associated with their need frustration whereas perceptions of teacher autonomy support were associated with need satisfaction. In sport coaching, after controlling for need satisfaction, adolescent athletes' perceptions of psychological need thwarting have been positively associated with exhaustion and negatively associated with vitality (Bartholomew et al., 2011). In the present study, we did not quantitatively measure need thwarting nor the controlling practices of dance instructors. Future process evaluations of SDT-based PA interventions which involve teachers or coaches would benefit from considering need thwarting alongside need satisfaction.","It is recommended that complex health behaviour change interventions are based on sound theory and that the theoretical elements are subjected to in-depth process evaluation (Moore et al., 2014). The findings of this theory-based process evaluation indicated that theoretical fidelity within BGDP was variable. We identified a number of instances of high theoretical fidelity and intervention innovation which informs pragmatic techniques that intervention deliverers working with groups of children could use. Illuminating the lack of intervention effectiveness, we also found that there was much room for improvement as we identified examples of low fidelity, some drift from the intended motivational techniques and potential failures to convert theoretical nuances into practice. More broadly, this work has highlighted the value of combining quantitative and qualitative approaches in theory-based process evaluations of physical activity interventions."],["This article reports a novel procedure used to investigate whether ambient light conditions affect the number of people who choose to walk or cycle. Pedestrian and cyclist count data were analysed using the biannual daylight-saving clock changes to compare daylight and after-dark conditions whilst keeping seasonal and time-of-day factors constant. Changes in frequencies during a 1-h case period before and after a clock change, when light conditions varied significantly between daylight and darkness, were compared against control periods when the light condition did not change. Odds ratios indicated the numbers of pedestrians and cyclists during the case period were significantly higher during daylight conditions than after-dark, resulting in a 62% increase in pedestrians and a 38% increase in cyclists. These results show the importance of light conditions on the numbers of pedestrian and cyclists, and highlight the potential of road lighting as a policy measure to encourage active travel after-dark. --------------------------------------------------------------------------------","Encouraging the use of active travel methods such as walking and cycling has a number of benefits. These include improvements in health outcomes such as all-cause mortality (Kelly et al., 2014), obesity (Pucher, Buehler, Bassett, & Dannenberg, 2010) and other health- related measures such as cancer rates and cardiovascular fitness (Oja et al., 2011). Such health improvements can lead to economic benefits (Jarrett et al., 2012). The promotion of active travel can also lead to reductions in the use of motorised transport (Ogilvie, Egan, Hamilton, & Petticrew, 2004), with reductions in CO2 emissions and improvements in air quality as a result (Goodman, Brand, & Ogilvie, 2012; Grabow et al., 2012; Rissel, 2009). Citizens who continue to use their vehicles for transport may also benefit from the promotion of active transport, due to reduced congestion on roads. One of the key purposes of road lighting is to create acceptable conditions for people to walk or cycle after-dark (British Standards Institution, 2012), thus encouraging active travel. For example, Kerr et al. (2016) and Giehl, Hallal, Brownson & d’Orsi (2016) both found that road lighting was positively associated with increased walking. Cervero and Kockelman (1997) also suggested that the presence of road lighting and the distance between lamps were significant aspects of neighbourhood design that contributed to encouraging non-automobile travel. Differences in lighting conditions can lead to changes in behaviour. For example, Painter (1994; 1996) found there was an increase in pedestrian use of a crime blackspot after new lighting was installed. Light conditions can also influence the speed with which pedestrians walk (Donker, Kruisheer & Kooi, 2011). Cyclists, as well as pedestrians, are also likely to be influenced by light conditions. For example, the ability to make a trip during daylight hours was found to be one of the top ten motivations in deciding to cycle, whilst using a route that was not well lit after-dark was one of the top ten deterrents (Winters, Davidson, Kao, & Teschke, 2011). There are several reasons why good light conditions may encourage walking or cycling. First, it allows obstacles and trip hazards to be seen and avoided, and this is a critical task for both pedestrians and cyclists (Fotios, Uttley, Cheal, & Hara, 2015; Vansteenkiste, Cardon, D'Hondt, Philippaerts, & Lenoir, 2013). Lighting characteristics such as illuminance and spectrum can influence the ability of a pedestrian or cyclist to detect an obstacle in the path in front of them (Fotios, Qasem, Cheal, & Uttley, 2016; Uttley, Fotios, & Cheal, 2015) and this may make a person more or less likely to walk or cycle, depending on the light conditions. Second, it may make the pedestrian or cyclist feel safer and less threatened (Boyce, Eklund, Hamilton, & Bruno, 2000; Fotios, Unwin, & Farrall, 2015). Good light conditions are required to allow a pedestrian or cyclist to see far ahead and have an open view. This is one of the three key attributes an area requires to make it feel safe (prospect, refuge and escape, Fisher & Nasar, 1992). The prospect of an area will be at its highest during daylight, but reductions to this after-dark can be mitigated by road lighting. For example Boyce et al. (2000) asked participants to rate how safe they felt at a number of parking lots in the US during daylight and after-dark. Safety ratings were generally lower after- dark than during daylight, but the difference reduced as the illuminance at the parking lot increased. Feeling safe is particularly important for pedestrians as perceptions of neighbourhood safety have been shown to influence walking levels in that neighbourhood (Foster et al., 2016; Mason, Kearns, & Livingston, 2013). The third and final reason why light conditions may influence the decision of a person to walk or cycle is due to their perceived visibility. Daylight or road lighting may make the pedestrian or cyclist feel more visible and less at risk of being hit by a vehicle as the rate and severity of traffic collisions involving pedestrians and cyclists is increased when there is poor or no road lighting (Eluru, Bhat, & Hensher, 2008), and during darkness (Johansson, Wanvik, & Elvik, 2009; Twisk & Reurings, 2013). These three factors of obstacle avoidance, perceived safety and perceived visibility suggest light should influence the decision of potential pedestrians and cyclists to travel or not and there should be a link between frequency of active travellers and light conditions. A causal connection between light and active travel has not been shown however. For example, previous work has linked the presence of road lighting with increased walking but it is not clear whether this is due to the light conditions provided or some other factor. This uncertainty is compounded by the fact that previous research related to this question has tended to use subjective methods for assessing the role of lighting, a good example of this being literature on lighting and perceived safety. A common approach is to ask participants to rate how safe they feel under different light conditions using a category rating response scale (Boomsma & Steg, 2012; Loewen, Steel, & Suedfeld, 1993; Rea, Bullough, & Brons, 2015). This approach has some limitations, if not carried out in a systematic way (Fotios & Castleton, 2016; Fotios, 2016). For example, asking for a rating compels a participant to make an assessment of something they perhaps would not otherwise consider relevant (Fotios, Unwin, et al., 2015). Data collected using subjective rating scales may be prone to range bias (Poulton, 1989) and influenced by the phrasing of the question (Schwarz, 1999). Perhaps most significantly, it is not certain that a subjective response by a participant translates into actual behaviour. For example, if light conditions do influence the subjective assessment of safety this may not necessarily be reflected in actual walking and cycling behaviour. Previous research has linked lighting conditions, perceived safety and physical activity (e.g. Weber, Hallal, Xavier, Jayce, & D'Orsi, 2012), but this has been based on subjective responses and is subject to the limitations outlined previously. Objective measures of behaviour could provide stronger evidence. In the current article we present an alternative procedure to examine whether the amount of ambient light affects the number of pedestrians and cyclists, which is to count the number of pedestrians and cyclists passing a location in the periods immediately before and after daylight savings clock change. This was inspired by the investigation of vehicle collisions reported by Sullivan and Flannagan (Sullivan & Flannagan, 2002). There are a range of factors that influence the volume of pedestrians and cyclists other than the light conditions, two of the most important being the season and the time of day (Aultman-Hall, Lane, & Lambert, 2009). The biannual changes to clock times resulting from daylight saving time provide an opportunity to control these two variables whilst changing the ambient light condition. This is where clock times in Northern hemisphere countries are advanced in Spring and moved back in Autumn by 1 h, changing the time of day at which dawn and dusk occur. This means that, as an example, a walk to or from work could take place during daylight in one week but after-dark the following week, at the same time of day. That is, an abrupt change of light level for the same journey decision. Counting the number of pedestrians and cyclists passing a particular location at this time of day means that the effect of light on the decision to walk is isolated from potential confounds of journey purpose, destination and environment. A similar approach utilising the daylight savings clock changes was used by Sullivan and Flannagan (2002). They analysed vehicle crash statistics in the US between 1987 and 1997. Their aim was to determine the likely effectiveness of adaptive headlamps in different driving situations, by identifying when dark conditions significantly increased the crash risk compared with daylight. They compared crash frequencies in the nine weeks before and after a clock change to see what the effect of the abrupt change in light conditions was. We use a similar before and after clock change method to compare daylight and dark conditions and their effect on active traveller frequencies. We develop this method further by introducing control periods in which light conditions do not change, against which changes between daylight and dark conditions can be compared. Pedestrian and cyclist count data collected over a five year period from the Arlington County area of Virginia state, United States, have been analysed using this daylight saving clock change method. Frequencies during a case hour before and after the Spring and Autumn clock changes are compared relative to changes in control periods in which the light conditions do not change. Arlington pedestrian and cyclist counters ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Automated pedestrian and cyclist counters have been installed in a number of locations within Arlington County, Virginia, in the Washington, DC metropolitan area, since October 2009, on both cycle trails and on-street cycle lanes. Arlington County is a 26 square mile area that was formerly an inner ring suburb of Washington DC. Walking and cycling have been regarded as important complements to rail and bus transit by the Local Authority, and led to the development of a healthy active travel infrastructure, matched by investment in active travel count apparatus to support transport planning. By 2016 there were 10 cyclist-only counters and 19 joint pedestrian and cyclist counters. Examples of these counters are shown in Fig. 1. The counters continuously record pedestrian and cyclist volumes and this data is available down to 15-min aggregations via a web service at the Bike Arlington website (http://www.bikearlington.com/pages/biking-in-arlington/counting- bikes-to-plan-for-bikes/data-for-developers/). Separate data for pedestrians and cyclists are provided. The direction of the traveller is also provided, as ‘inbound’ or ‘outbound’ relative to the centre of the Arlington area: for the analysis presented in this paper, inbound and outbound volumes were combined. Data collation ~~~~~~~~~~~~~~ The dates of Spring and Autumn daylight saving clock changes in the US between November 2011 and March 2016 are given in Table 1. An appropriate 1-h light transition period was identified for each of the clock-change periods, such that it was dark during this hour one side of the clock change date and daylight during the same hour on the other side of the clock change. These times were identified using the sunset times given for the Washington DC area by the Time and Date website (Time and Date, 2016). This period was defined as the case period. In addition, two 1-h control periods were identified, these having the same light condition both before and after the clock change for the same 1-h period. One of these was 1.5 h before the case period, ensuring it was daylight both before and after the clock change. The other was 1.5 h after the case period, ensuring it was dark both before and after the clock change. Two further control periods were also identified, these being 3.5 h before or after the case period, and also having the same light condition either side of the clock change. Multiple control periods were selected because estimates of the effect of the transition in light conditions may depend on the choice of the comparison, control hour. It is possible that any changes in frequencies during these control periods could vary systematically with the time of day (Johansson et al., 2009). As an example, it is possible that the hypothesised effect of the transition in light conditions during the case period could have a spillover effect on nearby times. The decision to walk or cycle may be influenced by the knowledge that there would be more (or less) daylight in the evening after the clock change, even if the person ended up walking or cycling in the control period rather than the case period. One way to test this hypothesis is to compare frequencies in the case period with control periods that are closer or further away in time from the case period. People who are walking or cycling during a 1-h period that is a greater temporal distance from the case period may be less likely to have been influenced in their decision to walk/cycle by the transition in ambient light levels, compared with someone walking/cycling during an hour that is temporally closer to the case period. The hours selected for all case and control periods are given in Table 2. Data for the case and control periods were extracted for the 13 days (Monday of week one to Saturday of week two) before and after the clock change dates given in Table 1, for all available counters. The day of the actual clock change (always a Sunday) was not included. These data were cleaned and checked for anomalous data. This included removing data for one counter that provided combined pedestrian and cyclist data without distinguishing between the two. Outlying data was identified by converting daily counts within each 1-h period into modified z-scores using the median absolute deviation, as recommended by Leys, Ley, Klein, Bernard, and Licata (2013). Daily counts with z-scores greater than ±3.5 were excluded from the final dataset, following recommendations by Iglewicz and Hoaglin (1993). This resulted in the exclusion of 1.6% of daily count data. Data were incomplete or missing entirely for some counters, due to them either not being installed by the dates queried, the counter being offline as a result of damage or routine maintenance, or the removal of outlying data. Counters that had less than 3 days of data in either the weeks before or after the clock change date were excluded. This process resulted in data from 11 counters being included from November 2011 up to 36 counters from March 2016, as new counters were installed during this period. This represented between 67% and 92% of all installed counters in any given season and year. Completeness of data was good, with only 6.0% of included counters for any particular 1-h period before or after the clock change having less than the full 13 days of data. Table 3 shows details of the number and types of counters that were included in the final dataset, and the minimum number of daily counts for pedestrians and cyclists, at each clock change time and year. Overall results ~~~~~~~~~~~~~~~ The mean daily count across all years was calculated for all counters at each of the 1-h case and control periods. Table 4 shows the overall means and standard deviations across all counters for each of the 1-h time periods that data was extracted for. Fig. 2 shows the overall ratio between cyclist and pedestrian frequencies in the case hour and control hours, over the 13 days before and after the biannual clock changes. This shows the direction of change in the frequencies following Spring and Autumn clock changes, highlighting how there is an increase when the case hour is in daylight rather than darkness, relative to changes in the control hours. It is also apparent that this effect appears larger for the Autumn clock change compared with the Spring clock change. The calculated odds ratios and 95% confidence intervals comparing each of the four control periods against the case period are shown in Fig. 3. All odds ratios were significantly greater than one, indicating that the numbers of pedestrians and cyclists were significantly higher during the daylight side of the clock change in the case period compared with the after-dark side and that this increase was more than that seen in all four control periods. The odds ratios were significantly higher for pedestrians than cyclists for three of the four control periods, with the Dark Control period being the exception. This suggests the transition between darkness and daylight during the case period may have had a greater effect on the numbers of pedestrians than on the numbers of cyclists. The overall odds ratio when all control periods are combined is 1.38 (1.37–1.39 95% CI, p < 0.001) for cyclists and 1.62 (1.60–1.63 95% CI, p < 0.001) for pedestrians. Location type ~~~~~~~~~~~~~ For cyclists, frequency data were collected from two types of locations – on-street cycle lanes, and cycle trails that are not associated with a road. The split in number of counters for each of these location types between 2011 and 2016 is shown in Table 3. Fig. 4 compares the odds ratios of cyclist frequencies during daylight compared with dark at on-street cycle lane and cycle trail locations. Pedestrian frequencies are not compared as the on-street cycle lanes only recorded data about cyclists. The odds ratios for both on- street cycle lanes and cycle trails are significantly greater than one for all control period comparisons. However, odds ratios for the cycle trails are also significantly higher than for the on-street cycle lanes for all four control periods. This suggests the effect of the transition between darkness and daylight during the case period had a greater effect on cyclist frequencies at trail locations, compared with cycle lane locations. Temperature The analysis method reported in this paper of comparing active traveller frequencies before and after a daylight saving clock change attempts to isolate the effect of ambient light conditions on the presence of cyclists and pedestrians by comparing the same hour of the day during daylight and after-dark conditions. Although the comparison periods before and after the clock change are contiguous it is possible that weather conditions were not the same, introducing a potential confound. Furthermore, the temperature may have varied systematically, as the daylight period always fell in the part of the year expected to be warmer compared with the after-dark period. For example, in Spring when the clocks move forward 1 h, the after-dark condition in the case period falls before the clock change, in the earlier part of the year when it may be expected to be slightly cooler in temperature. In Autumn, when clocks move backward 1 h, the reverse situation occurs, with the after-dark condition falling after the clock change, as temperatures may be expected to be cooling. Therefore, the after-dark condition may systematically be cooler than the daylight condition. As temperature is a significant factor in whether someone chooses to walk or cycle (e.g. Miranda- Moreno & Nosal, 2011; Saneinejad, Roorda, & Kennedy, 2012) this would provide an alternative explanation for why the daylight condition shows a relative increase in pedestrians and cyclists. To explore this alternative explanation temperature data for the Arlington area of the United States was downloaded from the Weather Underground via the web service provided at the Bike Arlington website (http://www.bikearlington.com/pages/biking-in-arlington/counting-bikes-to-plan- for-bikes/data-for-developers/). Hourly data was extracted for each of the 1-h case and control periods (see Table 2), during the two-week periods before and after each clock change between Autumn 2011 and Spring 2016. The hourly temperature data was averaged to give a mean temperature for each day. Fig. 5 shows the overall mean daily temperature combined across all years, for the day and after-dark periods in Spring and Autumn clock changes. Fig. 5 suggests mean temperatures were higher during the daylight period compared with the dark period, for both Autumn and Spring clock changes. This was confirmed with a 2-way between- subjects ANOVA, with the clock change season and the ambient light condition as independent factors. There was a significant main effect of the light condition with mean temperatures significantly higher during daylight periods (mean = 11.6 °C) than after-dark periods (mean = 8.1 °C, F (1,285) = 40.8, p < 0.001, ηp = 0.13). There was also a significant main effect of the clock change season, with mean temperatures significantly higher at the Autumn clock change (mean = 11.8 °C) than the Spring clock change (mean = 7.8 °C, F (1,285) = 55.5, p < 0.001, ηp = 0.16). These results show that temperature did significantly differ between the before and after clock change periods, with higher temperatures during those periods in the ambient daylight condition than in the ambient darkness condition. However, to confirm whether the change in light conditions can still explain the change in pedestrian and cyclist frequencies demonstrated in section 3.1, differences in temperature were tested before and after every clock change individually, to determine if in some years the temperature did not significantly change. If such years and seasons were found, these could be used to determine whether a change in pedestrian and cyclist numbers was still seen. If this was the case, it would suggest the light condition was an explanatory factor even when temperature remained constant before and after the clock change. A series of independent t-tests were carried out comparing temperatures during the day and after-dark periods for each clock change in each year. Bonferroni correction was applied to account for the multiple testing, giving an alpha of 0.05/10 = 0.005. The results are summarised in Table 5. These tests suggest there were five occasions when temperatures did not differ in the two week periods before and after clock changes (Autumn 2011; Spring 2013; Autumn 2013, Autumn 2015 and Spring 2016). One of these occasions (Autumn 2013) did show a relatively large difference in temperatures and was close to reaching the Bonferroni-corrected alpha level (p = 0.007), and is therefore not included as an occasion when temperatures did not differ, to err on the side of caution. The other four occasions are deemed to not show a change in temperature before and after the clock change. Clock changes that did not show a significant change in temperature before and after the change date were combined and odds ratios were calculated comparing the case period with each of the four control periods. The same was done for clock changes that did show a significant change in temperature. The calculated odds ratios and 95% confidence intervals are shown in Fig. 6. The data shown in Fig. 6 demonstrate that even when changes in temperature are accounted for and only occasions when the temperature did not significantly change before and after the clock change are examined, the odds ratios are still significantly above one. This suggests the effect of the transition in ambient light condition alone can explain the increase in pedestrians and cyclists during the daylight periods independently from any influence of temperature. In fact, the clock changes that showed no change in temperature generally produced larger odds ratios than those clock changes where the temperature did change significantly. This would suggest the effect of the transition in light conditions was larger when there was no temperature change. Precipitation It is possible that precipitation may also influence the decision to walk or cycle (de Montigny, Ling & Zacaharis, 2012; Nosal & Miranda-Moreno, 2014) and thus confound explanation of the change in pedestrian and cyclist numbers. To check this, daily precipitation levels were obtained from the Weather Underground service via the Bike Arlington website, with mean daily precipitation calculated for the periods before and after each clock change. A two-way between-subjects ANOVA was carried out, with the clock change season and ambient light condition as independent factors, to identify any systematic differences in precipitation levels. This suggested there was no difference in precipitation volume between the Spring and Autumn seasons (respective means = 0.09 and 0.11 inches per day, F (1,231) < 0.001, p = 0.98). There was also no difference in precipitation volumes between the daylight and dark periods (respective means = 0.12 and 0.08 inches per day, F (1,231) = 0.77, p = 0.38). There was also no interaction between the season and the ambient light condition (F (1,231) = 2.91, p = 0.09). These results do not suggest systematic variations in precipitation levels between the periods when the case hour was in daylight and in darkness and precipitation can therefore be ruled out as a potential explanation for the differences found in pedestrian and cyclist frequencies when the ambient light condition changed.","The aim of this work was to establish whether ambient light level affects the number of people choosing to walk or cycle. A large number of pedestrian and cyclist counters in Arlington County, Virginia, provided extensive data about the numbers of people walking and cycling. This open-source data provided the opportunity to carry out a novel method of analysis using the daylight-saving clock change to isolate the effect of an abrupt change in ambient light conditions. Data was extracted for two-week periods before and after ten clock-change dates between Autumn 2011 and Spring 2016. Pedestrian and cyclist frequencies during a 1-h ‘case’ time period, in which the ambient light conditions were different before and after the clock change, were compared against four other 1-h ‘control’ periods, in which the ambient light conditions remained the same both before and after the clock change. Two of these control periods were chosen to be close in time to the case period, one during daylight the other after-dark. The other two control periods were chosen to be more distant in time from the case period. When all control periods were combined, the calculated odds ratio suggested daylight conditions resulted in a 62% increase in pedestrians and a 38% increase in cyclists, compared with after-dark conditions. Looking separately at the odds ratios for each of the four control periods, there is a suggestion that the effect of the transition between darkness and daylight in the case period was greater when compared against the control periods that were further away in time, than when compared against the control periods that were nearer in time to the case period's hour of transition. This is confirmed when looking at odds ratios for the near control periods combined (Dark Control and Day Control) and the far control periods combined (Late Dark Control and Early Day Control). For pedestrians, the combined odds ratio for the near control periods was 1.56 (1.54–1.58 95% CI) compared with 1.72 (1.69–1.75 95% CI) for the far control periods. For cyclists, the combined odds ratio for the near control periods was 1.36 (1.35–1.37 95% CI) compared with 1.42 (1.41–1.44 95% CI) for the far control periods. This suggests the odds ratios were significantly greater for the far control periods than the near control periods, for both pedestrians and cyclists. This supports the hypothesis that there is some spillover or displacement effect of the transition in ambient light conditions. This may also partly explain why the OR for pedestrians is smaller than for cyclists when using the dark control hour, but for the other three control periods the pedestrian OR is larger than the cyclist OR (see Fig. 3). This reversal in the size of OR for pedestrians and cyclists is due to a relatively large change in pedestrians during the dark control period when the case hour is in daylight compared with darkness (daily mean count = 24 and 16 respectively; see Table 2). The equivalent change in cyclists is smaller (daily mean count = 18 and 15 for dark control hour when case hour is in daylight and darkness respectively; see Table 2). This may be due to a greater spillover effect for pedestrians compared with cyclists. Possible reasons for this include reduced flexibility in work departure time amongst cyclist compared with pedestrian commuters due to considerations about road traffic volumes or the habitual nature of cycle commuting. There may also be increased opportunity for delays and detours during a pedestrian's journey home (e.g. visit to the shops or to a bar) compared with a cyclist's. Fig. 8 shows standardised hourly frequencies for pedestrians and cyclists during the 13 day periods before and after Spring and Autumn clock changes in 2015, as an illustration of daily patterns in pedestrian and cyclist numbers. There are large peaks in cyclist frequencies at morning and evening commuter times, and whether the case hour is in daylight or darkness does not alter the timing of these peaks. This supports the suggestion that cyclists may be quite rigid in their travel times, producing a relatively limited spillover effect. The travel times of pedestrians is a lot more distributed throughout the day however, with morning and evening peak times much less obvious compared to cyclists. This suggests there may be greater fluidity in travel times of pedestrians, potentially leading to a greater spillover of travelling during the dark control hour. Further data is needed to corroborate this hypothesis though. The Arlington pedestrian and cyclist counters were located in two types of location – on-street cycle lanes, and cycle trails. The calculated odds ratios for cycle trail locations were significantly greater than on-street cycle lane locations for all four control periods (Fig. 3). This suggests the ambient light conditions had a greater effect on the number of cyclists on the cycle trails compared with the cycle lanes. One possible explanation for this is that the cycle lanes may be used more by cyclist commuters travelling to and from work. This journey is likely to be habitual and therefore the decision to cycle or not may be less likely to be influenced by the light conditions. The cycle trails are more likely to be used by recreational cyclists, who can be more selective in what days and times they choose to cycle, and the ambient light condition is likely to have a greater influence on whether such cyclists choose to cycle at a particular time. As a result, there may be less use of the cycle trails when dark, compared with the on-street cycle lanes, which would explain the larger odds ratios for cycle trails. The cycle trails and cycle lanes may also be located in areas of different land use, e.g. residential districts, parks, industrial areas, and this may influence the type of user and their propensity to travel at different times of the day and week. The users of the two types of cycle paths may also differ in their confidence in cycling and perceptions of danger. Cyclists who are more willing to cycle on urban roads and who see themselves as competent may be more likely to see cycling as a safe travel mode. This may result in cyclists who use the on-street cycle lanes being less influenced by the potential safety implications of cycling in darkness rather than daylight, compared with cyclists who use the cycle trails. An alternative explanation for the difference between on-street cycle lanes and cycle trails though could relate to the public lighting that is present in these two types of locations. The on-street cycle lanes are likely to have well-provisioned public road lighting as they are situated on roads used by motor vehicles. This may be less the case on cycle trails however, where the public lighting may be less frequent and dimmer, if present at all. For example, many of the cycle trails transect public parks, and these are frequently not lit after-dark. Greater provision of public lighting after-dark may result in more cyclists travelling after-dark, and this could explain why the effect of the transition between daylight and darkness is greater on cycle trails than on-street cycle lanes. The cycle lane and trail locations can be seen as typologically similar to an urban and rural distinction, with more road lighting at urban than rural locations. Johansson et al. (2009) suggested the reduced road lighting on rural roads may have accounted for their results about vehicle collisions, which showed larger odds ratios related to the effect of dark conditions on rural roads, compared with urban roads. Weather conditions are an important consideration, alongside light levels, in determining whether someone chooses to walk or cycle (e.g. Saneinejad et al., 2012). In particular, temperature, as a relatively predictable and stable variable of climate, is likely to have an influence on active travelling. A limitation of the current approach using clock changes to investigate the effect of light conditions on active travel is that the period around the clock change date that had more daylight was also the period that was likely to have slightly warmer temperatures, all things being equal. The ‘daylight’ side of the clock change had significantly warmer daily mean temperatures than the ‘darkness’ side of the clock change (Fig. 5). However, in some years the mean daily temperature did not change before and after the clock change. These occasions still showed odds ratios significantly greater than one, indicating that the transition in light had a significant effect on pedestrian and cyclist numbers, over and above any effect of temperature (Fig. 6). In fact, the effect was larger when there was no change in temperature, compared with when the temperature also increased during the daylight side of the clock change. This is logical – an increase in temperature may increase numbers in the control periods which will reduce the relative size of the effect of the transition in light when the case period is compared against the control periods. This is why the effect is larger at those clock change times when temperature did not significantly change before and after the clock change date. An increase in temperature serves to partially mask the effect of the transition in light. We also examined precipitation to determine whether this could explain the changes in active traveller counts, but found no difference in precipitation levels before and after the clock changes (see Fig. 7). The 1-h clock change that occurs in the Spring and Autumn of each year is not only marked by an abrupt change in ambient light levels at the same time of the day, but may also be marked by individual behavioural and wider societal changes. For example, although we show that there is a change in the number of active travellers before and after a clock change even when temperature does not change, it is possible that the clock change represents a psychological Rubicon for many people that symbolises the onset of a new season. This may result in changes in behaviour, activity schedules or perceptions about the environment (such as it being warmer or colder) that may not reflect true changes. The transition to and from Daylight Saving Time may also be used by businesses, organisations and local services to change their hours of business. As an example, in our city of Sheffield, UK, household waste sites change between ‘Summer’ and ‘Winter’ opening times in April and October, around the time of the clock changes. Such changes could result in changes to the behaviour of local residents resulting in differences in the numbers of pedestrians and cyclists before and after a clock change. Clock changes can also cause changes to circadian rhythms and waking times (e.g. Kantermann, Juda, Merrow, & Roenneberg, 2007) which may influence behaviour. As an example of this, Daylight Saving Time has been associated with an increase in ‘cyberloafing’ behaviour amongst employees as a result of lost and low-quality sleep (Wagner, Barnes, Lim, & Ferris, 2012). Such behavioural changes may produce variations in the frequency of pedestrians and cyclists during the case hour examined in the current study. Therefore a number of potential behavioural and societal changes could occur as a result of the biannual clock changes but are not overtly linked to changes in ambient light conditions. These may contribute towards changes in active traveller frequencies. Although the use of control hours in the present study attempts to account for such confounding factors that are unrelated to light conditions, further investigation is required to determine the exact influence of these behavioural and societal changes on pedestrian and cyclist numbers.","Active travel, i.e. walking and cycling, has a range of benefits and should be encouraged and facilitated whenever possible. A number of potential barriers to active travel exist, such as physical fitness, habitual behaviour or perceived environmental factors such as personal safety (e.g. Dawson, Hillsdon, Boller & Foster, 2007). One environmental factor that may be important is the light condition. We have shown that ambient light levels significantly influence the numbers of people choosing to walk or cycle. In drawing this conclusion we have accounted for seasonal and time-of-day factors, by using the daylight- saving clock change analysis method. To our knowledge, this is the first time this method has been used to examine active travel behaviour. We also show that light level is a significant determinant of active travel even when temperature is accounted for. The presence or absence of public lighting is also a possible explanation for why bigger reductions in active travellers were seen after-dark on cycle trails compared with on- street cycle lanes, although further work is required to confirm this hypothesis. It is also possible that the types of cyclists using cycle trails and on-street cycle lanes differ. The influence of cyclist and pedestrian characteristics, such as gender and age, on the likelihood of travelling after-dark should therefore also be investigated. In summary, this work shows the significance of ambient light levels on active travel. Although artificial lighting after-dark is not equivalent to daylight, these results highlight a potential role for road lighting in encouraging active travel, to provide adequate light conditions that minimise the transition in light from daylight to darkness.","This work was carried out with support from the Engineering and Physical Sciences Research Council (EPSRC) grant number EP/M02900X/1."],["Wayfinding is the ability to learn and recall a route through an environment. Theories of wayfinding suggest that for children to learn a route successfully, they must have repeated experience of it, but in this experiment we investigated whether children could learn a route after only a single experience of the route. A total of 80 participants from the United Kingdom in four groups of 20 8-year-olds, 10-year-olds, 12-year-olds, and adults were shown a route through a 12-turn maze in a virtual environment. At each junction, there was a unique object that could be used as a landmark. Participants were “walked” along the route just once (without any verbal prompts) and then were asked to retrace the route from the start without any help. Nearly three quarters of the 12-year-olds, half of the 10-year-olds, and a third of the 8-year-olds retraced the route without any errors the first time they traveled it on their own. This finding suggests that many young children can learn routes, even with as many as 12 turns, very quickly and without the need for repeated experience. The implications for theories of wayfinding that emphasize the need for extensive experience are discussed. --------------------------------------------------------------------------------","Researchers have long been interested in how navigation develops (Bullens, Iglói, Berthoz, Postma, & Rondi-Reig, 2010; Karimpur & Hamburger, 2016; Purser et al., 2015). Navigational abilities such as route learning are used by most people every day as they travel from one place to another place. Route learning refers to the ability to encode spatial and other information along a route well enough to retrace that route on future occasions (Merrill, Yang, Roskos, & Steele, 2016; Rissotto & Giuliani, 2006). Adults can often learn routes quickly and effectively after only one or two experiences of the route (Gärling, Böök, Lindberg, & Nilsson, 1981; Montello, 1998), and this is the case even when the routes are 1 or 2 km long and/or include a large number of choice points at junctions (Farran, Blades, Boucher, & Tranter, 2010; Karimpur & Hamburger, 2016). The ease with which adults learn routes suggests that adults have developed appropriate strategies for encoding routes (Montello, 2017). One important strategy is encoding turns in relation to landmarks (e.g., “the left turn after the school”). Adults may be particularly well adapted to focus on landmarks and turns, and there is evidence for distinct brain activation for landmarks and routes (Wegman & Janzen, 2011). Adults show increased activity in the parahippocampal gyrus when attending to landmarks at decision points (Janzen, Wagensveld, & van Turennout, 2007), and the anterior cingulate gyrus and the right caudate nucleus are activated when adults learn the turns along a route (Janzen & Weststeijn, 2007). In contrast to adults’ competence in learning new routes after only brief experience, the evidence about children’s ability to learn new routes after only brief experience is less clear. Siegel and White (1975) argued that children’s route learning requires repeated experience because children first need to learn individual landmarks along a route; only then do they associate those landmarks with particular decisions (e.g., left or right turns) before they can combine a series of landmarks and turns into a fully learned route. There is evidence that children do, like adults, focus on landmarks (Jansen-Osmann & Wiedenbauer, 2004; Lingwood, Blades, Farran, Courbois, & Matthews, 2015a, 2015b). van Ekert, Wegman, and Janzen (2015) found that similar regions of networks involving the hippocampus and the inferior/middle frontal gyrus were activated during a memory test for previously seen landmarks in both children and adults. Despite children’s focus on landmarks, children do not always learn routes as well as adults (Jansen-Osmann & Wiedenbauer, 2004). This may be because in real life or in complex environments, children may be less good at identifying what is an effective landmark. Younger children may be especially dependent on landmarks that are nearby or next to turns, whereas older children tend to use distant landmarks for wayfinding (Cornell, Hadley, Sterling, Chan, & Boechler, 2001; Purser et al., 2012). Older children are also more likely than younger children to use verbal strategies such as counting the number of steps taken or number of buildings passed when retracing a route (Duroisin & Demeuse, 2015). Having better strategies for learning landmarks or estimating distances along routes means that children’s route learning does improve with age and may account for reports of age-related improvements in wayfinding in complex environments (Cornell, Heth, & Alberts, 1994; Cornell, Heth, & Broda, 1989; Heth, Cornell, & Alberts, 1997). In contrast to the studies showing that children do need repeated experience of a route to learn the route, researchers have found that preschoolers can learn a route without error after only a single experience (Spencer & Darvizeh, 1983), and Cornell and Hay (1984) reported that 6- and 8-year-olds who had seen a route only once were able to retrace the route with an average of less than one error. The latter finding implies that a number of children retraced the route without error at all. Cousins, Siegel, and Maxwell (1983) showed that 7-, 10-, and 13-year-olds could retrace a route after one experience of it. The fact that even very young children can retrace novel routes after one experience goes against the suggestion that children need multiple experiences of a route before they can learn it successfully. Rather, it seems that children can encode a route as effectively as adults and do not need to progress through “stages” of route leaning such as learning landmarks, turns, and then completed routes. The latter would support Montello’s (2017) argument that there is no qualitative difference between “landmark” and “route” knowledge. According to Montello, knowledge of landmarks and knowledge of a route develop in unison and are inseparable aspects of route learning rather than sequential stages in learning a route. However, when children have learned routes after one experience in the studies cited above, the routes were short with seven or fewer turns (Cornell & Hay, 1984; Spencer & Darvizeh, 1983), and even though the children may have been unfamiliar with the particular test route, they were very familiar with the environment that the test route ran through (Cousins et al., 1983; Spencer & Darvizeh, 1983). Therefore, these studies indicated the possibility that young children might be able to retrace routes after one experience. However, given the nature of the routes (which were limited) and the children’s advantage in already knowing the general environments of the routes, the results need to be treated cautiously. It may well be the case that young children can learn routes straight away and without multiple experiences, but this needs to be demonstrated in a more demanding context with longer routes in completely unfamiliar environments. Therefore, we investigated whether children could learn a route after just one experience of the route, but we did so in a more rigorous context and with a longer route than in the previous studies. We tested four age groups, including adults, so that we could make developmental comparisons. To maintain a perfectly consistent and safe environment for all participants, including the youngest ones, we tested participants in a virtual environment (VE). VEs are an alternative to studying wayfinding in the real world because they allow children the opportunity to walk a route several times without the physical demands of a real environment (Broadbent, Farran, & Tolmie, 2014). VEs can be used successfully with very young children (Lingwood et al., 2015a, 2015b), and they also allow children to retrace routes without the need to be accompanied by an adult. VEs can depict visual and spatial information from a three- dimensional first-person perspective (Jansen-Osmann, 2002; Richardson, Montello, & Hegarty, 1999), and successful route learning in VEs transfers to real environments (Ruddle, Payne, & Jones, 1997). Therefore, VEs are a very appropriate way to assess young children’s abilities. Our study compared the ability of 8-, 10-, and 12-year-olds and adults to retrace a route with 12 decision points in a VE.","were guided along the correct route just once and then were asked to retrace the route, from the start, on their own. The primary research question was whether the participants could learn the whole route after a single experience of it. A second research question was to consider when children’s performance was equivalent to the performance of adults. Participants ~~~~~~~~~~~~ Child participants were 20 8-year-olds (M = 7;11 [years;months], SD = 3.84 months), 20 10-year-olds (M = 10;2, SD = 6.04 months), and 20 12-year-olds (M = 12;9, SD = 4.91 months) who were recruited from primary and secondary schools in the United Kingdom. There were 10 boys and 10 girls in each age group. Adult participants were 20 students at the University of Sheffield (M = 23;8, SD = 2;2), with an age range of 18;0 to 29;10 (11 women and 9 men). Ethical approval was granted by the University of Sheffield ethics committee. Virtual environments Two different VEs were created using Vizard, a software program that uses Python scripting. VEs were presented to participants on a 17-in. Dell laptop that was placed on a desk. Participants sat in a chair at the desk and were approximately 50 cm from the screen. Participants navigated through the maze using the arrow keys on the keyboard. Practice maze One maze (Maze A) was used as a practice maze to familiarize participants with moving in a VE. This maze was a similar but different layout compared with the test maze. It did not contain any landmarks. Test maze Maze B was used to test the children and adults (see Fig. 1). The test maze was a brick wall maze with 12 junctions. The junctions were “L” shaped. Each junction had two paths: a correct path and an incorrect path. Of 24 landmarks, 12 landmarks were placed on the correct paths (junction landmarks) and 12 landmarks were placed on incorrect paths (off-route landmarks). As in previous studies, all landmarks were placed in the middle of the paths (Jansen-Osmann & Wiedenbauer, 2004; Lingwood et al., 2015a, 2015b). Landmarks were placed in the middle of paths in case participants interpreted a landmark that was placed on the left hand side of the path as one that indicated a left turn or a landmark on the right as one that indicated a right turn. An incorrect path always ended in a cul-de-sac, but from each junction a cul-de-sac looked like a typical path rather than a dead end. Therefore, participants could not tell that they had made an error until they had actually committed to walking down a chosen path. There were four right, four left, and four straight ahead correct choices that were balanced with the same number and types of incorrect choices. All of the path lengths between junctions were equal. A white duck marked the start of the maze, and a gray duck marked the end of the maze. When retracing a route from the start, participants were told to find the route back to the gray duck. The gray duck provided a salient target that was always the end point of the route and did not move. When participants reached the gray duck, the maze disappeared, indicating the end of a trial. All of the landmarks were objects with names that would be familiar to children (and adults) such as ball, playground slide, street lamp, and umbrella. These items were chosen because they had distinctive names, were easily recognizable, and could be distinguished from each other without difficulty. All of the landmarks were static because they did not change position during the experiment. When being shown the route, participants passed close to the landmarks at junctions, and the off-route landmarks were not so close but were easily recognizable from the distance that participants saw them.","All participants were tested individually. Children completed the experiment in a quiet room in their school. Informed consent was obtained from all of the children’s parents, and all children were asked whether they wanted to take part. None of the children refused to take part. Adults completed the experiment in a quiet office in a university department. Participants sat at the desk facing a computer, and the experimenter sat beside them. The experimenter spent 2 min talking to the children informally to establish rapport. Then the experimenter introduced the task by saying, “This computer has got some mazes on it that we are going to use. First, we’re going to practice using the computer to walk around a maze. I’ll go first and show you how, and then you can have a turn.” The experimenter then demonstrated how to navigate through the practice maze using the arrow keys. The practice maze was the same design as the experimental maze, but it was a different layout. Participants were given time to walk around the maze until they were confident about using the arrow keys, at which point the experimenter ended the practice phase by saying, “Well done, I think you’ve had enough practice now. Let’s have a go at another maze now.” Participants were first given a single experience of the correct route, guided by the experimenter. During this initial experience, the experimenter guided participants by moving forward or turning along the correct route (without looking down any of the incorrect paths). All participants were given preliminary instructions for the test phase: “Now I’m going to show you the way through a new maze. Somewhere in this maze, there is a little gray duck to find. I’ll show you the way to the gray duck once, and then you can have a go.” The experimenter demonstrated the correct route from the start to the end of the maze. The experimenter used generic terms such as “You go past here, then you turn this way, and then you turn this way.” The experimenter never used any directional language such as “turn right.” At the end of the demonstration, the experimenter said, “Hooray, we’ve found the duck!” and the screen went blank. Participants were then asked to retrace the route they had been shown from the start of the maze to the gray duck that was always in the same place. No participant ever queried this instruction or asked whether the duck had moved. Participants navigated through the maze using the arrow keys on the keyboard. The experimenter sat behind participants and traced the exact route they took on a paper copy of the maze out of participants’ sight. The experimenter timed how long it took participants to complete the maze. If a participant had not reached the end of a maze after 5 min, the experimenter ended that attempt by saying, “Oops, it looks like you’ve got a bit lost. Not to worry, let’s start back from the beginning, shall we?” Only one 8-year-old did not reach the end of the maze on the first attempt. Participants did not receive any help in finding their way after the initial demonstration of the correct route. If participants asked which way to go, the experimenter said, “I want you to show me the way to go. Just try your best.” If participants returned to the start position but thought that they had reached the end, they were told, “You’re back at the beginning of the maze now. Let’s turn around and try again to remember the way I showed you to the little gray duck.” Participants then made a second attempt to follow the correct route but did not get to see the original route again. When participants reached the end of the maze, this was the end of one trial. The experimenter congratulated the participants and asked them to walk the route again from the start. This procedure was repeated until participants had walked the route to a criterion of two consecutive completions without error. The reason for the criterion of two consecutive routes without error is given in the section on scoring (below). At the end of the final trial, all participants were thanked and the children, regardless of their performance, received a sticker. If participants had not walked the route successfully on two consecutive attempts after 20 min or after eight attempts, the experiment was stopped and the children were given a sticker. This was based on the procedure used by previous researchers (Farran, Courbois, Herwegen, & Blades, 2012; Lingwood et al., 2015a, 2015b). Only two 8-year-olds did not walk the route successfully on two consecutive occasions. Retracing the route ~~~~~~~~~~~~~~~~~~~ The probability of retracing the route correctly, without any mistakes, by guessing at each junction was p < .00025. The number of participants who completed the route without errors on their first attempt is shown in Table 1. The findings from Table 1 show that some children can learn a route consisting of 12 turns immediately, having viewed it only once. Nearly a third of 8-year-olds, half of 10-year-olds and nearly three quarters of 12-year-olds successfully retraced the route on their first attempt without making any errors. Not all participants retraced the route on their first attempt without error, and so in the following analyses we included only those participants who made one or more errors when retracing the route for the first time. This included 14 8-year-olds, 10 10-year-olds, 6 12-year-olds, and 5 adults. Reaching the learning criterion ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To avoid overestimating participants’ performance, we defined a criterion of successful learning as follows: two consecutive completions of the route without error. To achieve this criterion, participants needed to walk the route without walking down any incorrect paths on two consecutive trials. Walking down an incorrect path was classed as an error. Participants needed to fully walk down a path in order for it to be counted as an error. Looking down an incorrect path was not classed as an error. The total number of learning attempts to reach criterion excluded the final two perfect attempts. For example, if participants made an error on Attempt 1, but then walked the route without error on Attempts 2 and 3, they would be scored as having required 1 attempt before reaching criterion. A lower score indicated better performance. If participants never achieved the criterion, we calculated the number of completed attempts. For example, if participants completed 6 attempts within the 20-min cutoff time but did not complete 2 consecutive attempts without error, they scored 6. Three of the 8-year-olds did not complete two consecutive attempts without error within the 20-min cutoff time. Table 2 shows that participants who made one or more errors when retracing the route on their first attempt required a similar number of trials to reach learning criterion irrespective of age. This was confirmed by non-parametric statistics, H(3) = 1.45, p = .69. Errors during route retracing to criterion ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Participants received a mark of 1 for every error they made during an attempt. On each attempt, a proportional error score was calculated as the number of errors divided by the number of decisions made. Fig. 2 shows the route taken by one participant. This participant made 2 errors out of a total of 12 decisions, producing a proportional error score of .17. This scoring captured participants’ wayfinding behavior every time they made a decision. This scoring method accounted for occasions when participants doubled back and returned to the same junction more than once within a trial. Some participants who got lost did not reach the later junctions, so any junctions not reached were also scored as errors at decision points. A mean proportional error score was calculated for each participant across each attempt. We note that alternative coding criteria produced the same patterns of performance. For example, we coded just the decisions made the first time participants approached a junction in each attempt. Participants scored 0 if they chose the correct path or 1 if they chose the incorrect path, and any junctions not reached were counted as errors. Therefore, 6 indicated the worse performance and 0 indicated perfect performance. When this scoring was compared with the proportional error score (above), there were no differences in the results. Therefore, in this Results section we report only the proportional error scores. The mean proportional error scores made by each age group before participants achieved criterion are shown in Table 3. The trend of these results suggests that, irrespective of age, children and adults made a similar number of errors. This was confirmed by non-parametric statistics, F(3) = 3.49, p = .32. We also coded “looking” behaviors when participants retraced the maze. Every time participants used the arrow keys to turn left or right to look down a junction, this was coded as a looking behavior. Table 4 displays the mean number of looking behaviors across age group. This shows that looking behaviors were similar across all age groups, although 10-year- olds produced the fewest number of looks during the learning trials. This was confirmed by non-parametric statistics, H(3) = 2.38, p = .50.","This experiment addressed two research aims. Our primary aim was to find out whether children could successfully retrace a novel route, having experienced that route only once. The findings showed that some of the children, irrespective of age, were able to successfully retrace a 12-turn maze in a VE. Nearly three quarters of the 12-year-olds, half of the 10-year-olds, and a third of the 8-year-olds retraced the route without any errors the first time they traveled it on their own. We emphasize that the route had 12 turns and, therefore, was challenging; nonetheless, some participants in each age group retraced it successfully and without errors the first time they traveled it on their own. We consider the implications of this aspect of children’s performance first. Past evidence for young children’s route learning ability is limited, probably because of the practical difficulties of testing children along safe routes in the real environment. Although past researchers have demonstrated that young children can encode short routes of up to seven turns in real environments (Cornell & Hay, 1984), there has been no previous research into young children’s ability to learn longer routes. The current study is the first to compare several age groups on a long route (with 12 turns). In the current study, the proportion of 12-year-olds who retraced the route after one trial was similar to the proportion of adults who did so (see Table 1). Although the similarity between the 12-year-olds’ performance and the performance of the adults confirms past research (Cornell et al., 1989; Jansen-Osmann & Wiedenbauer, 2004), we note that the past research routes were shorter (e.g., Jansen-Osmann & Wiedenbauer, 2004) or already familiar to children (e.g., Cornell et al., 2001). Therefore, the findings from the current study give a fuller account of children’s wayfinding abilities by demonstrating that 12-year-olds have similar route learning capabilities to adults, with many showing the ability to learn a long route immediately. A second novel finding in the current study was that the 10-year-olds also performed well, with half of the children in that age group retracing the route on their first attempt. The youngest children, the 8-year-olds, were noticeably poorer than the older age groups, and we note that this was the only group in which not all of the children reached criterion (three children did not). Nonetheless, most children in this age group did achieve criterion, and one third of the 8-year-olds retraced the route after seeing it just once. This indicated the potential of even younger children to remember the route successfully. Previous researchers have found that young children can recall routes (Cornell & Hay, 1984; Spencer & Darvizeh, 1983) but tested children only over comparatively short routes or through areas that might have been familiar to the children. Finding that some children as young as 8 years can retrace a completely novel 12-turn route after experiencing it just once was unexpected. Although most participants were able to successfully retrace the route, there were different trends in performance across different age groups. Although not significant at the statistical level, based on the mean scores (see Tables 2 and 3), younger children (8-year-olds) made more errors and required more trials to reach criterion than older children (10- and 12-year-olds), a finding consistent with previous research (Cornell et al., 1989; Farran et al., 2012; Heth et al., 1997; Jansen-Osmann & Wiedenbauer, 2004). These findings were also consistent with a six- turn VE used by Lingwood et al. (2015a) in which older children (10-year-olds) made fewer errors and required fewer trials to reach criterion than younger children (6-year-olds). Younger children may have less mature cognitive abilities required for learning routes. Evidence from previous studies suggests that children with better attention and long-term memory tend to perform better when learning novel routes (Purser et al., 2012, 2015). In addition, younger children may have less experience in retracing novel routes than older children (Kitchin & Blades, 2002; Lingwood et al., 2015a). In the current study, participants needed to avoid making repeated errors on more than one trial (perseverative errors) if they were to succeed in reaching the learning criterion. This may have been more difficult for younger children (Farran et al., 2012; Purser et al., 2012). There were also other demands extraneous to the route following task that children needed to avoid to ensure they sustained attention (such as inhibiting any desire to look away from the screen, engage in conversation with the experimenter, or freely explore the VE). Participants also needed to inhibit the desire to navigate toward irrelevant off-route landmarks. Children who can inhibit inappropriate responses on other tasks (e.g., a Go/No- Go task) do perform better on route learning tasks (Purser et al., 2012). Furthermore, focusing on non-useful landmarks (i.e., the off-route landmarks in this study) during route learning is associated with poorer navigational strategies (Broadbent et al., 2014; Farran et al., 2010). This suggests that a variety of spatial working memory skills may be important for wayfinding, particularly for younger children. The second aim was to investigate at what age children demonstrated route learning abilities similar to adults. Based on the mean scores (see Table 2), it was found that from 8 years of age onward children made a similar number of errors and required a similar number of additional learning trials as adults. Our results extend findings from previous studies (Cornell et al., 1989, 1994; Jansen-Osmann & Wiedenbauer, 2004). Cornell and colleagues found that children did not become adult-like in their route learning abilities until 12 years of age rather than 8 years of age (as our data suggest). However, in Cornell and colleagues’ studies children were either 6 or 12 years of age, and so it was unclear whether children under age 12 but over age 6 would perform at the same level as adults. Jansen-Osmann and Wiedenbauer (2004) found that children as young as 11 years performed similarly to adults when asked to find their way out of a VE maze. However, the children and adults who participated in Jansen-Osmann and Wiedenbauer’s study explored a maze rather than learned a specific route. Our results extend previous research findings and suggest that children are able to learn and remember a route similarly to adults from 8 years of age onward. It was not possible to determine precisely how children and adults retraced the route and made few (if any) mistakes. For example, it is unclear whether participants relied on a particular strategy when retracing the route such as recalling the left–right sequence or focusing on landmarks at particular junctions. Nonetheless, we argue that being able to accurately encode the correct left–right sequence of 12 turns independently of the route would have been too cognitively taxing for participants, especially the younger children (Hayashi, Fujii, & Inui, 1990). Therefore, we suggest that participants’ success depended on recognizing landmarks as they were traveling along the route and that such recognition then led to the recall of the appropriate action at a choice point. Adults may be adapted to recognize landmarks given that neuroimaging studies have shown increased activity in the parahippocampal place area associated with the recognition of landmarks (Epstein & Vass, 2014; Marchette, Vass, Ryan, & Epstein, 2015). We note that some off-route landmarks could also be viewed from the correct route direction heading. Therefore, there was a possibility that participants used off-route landmarks to retrace the route. However, evidence from neuroimaging studies suggests that we respond to on- and off-route landmarks differently. For example, immediately after learning a route, the parahippocampal gyrus shows increased activity only for on-route, as opposed to off-route, landmarks (Janzen & van Turennout, 2004; Janzen et al., 2007). This suggests that the ability to identify previously seen on-route landmarks may be particularly crucial for route learning. Cornell et al. (1994) suggested that individuals used “place recognition” to help find their way. For instance, when approaching a junction, individuals could look down the various paths before deciding which path was most familiar to them. In this sense, the landmark would be a “beacon” (Waller & Lippa, 2007), the assumption being that only one of the paths was likely to contain landmarks or features previously seen when walking the route. However, in the current study, we found that children and adults did not frequently look down the alternative path when deciding which way to go. This may have been because they could always see a landmark ahead of them. By always being able to view this landmark, this may have helped them to differentiate between the correct and incorrect paths, and so in most cases participants would not need to look down the alternative junction. We acknowledge that in our experiments the junctions were made up of two path intersections and that each junction was visually uncluttered. These types of junctions are easy to describe (Asher, Tolhurst, Troscianko, & Gilchrist, 2013; Clarke, Elsner, & Rohde, 2013; Klippel, Tenbrink, & Montello, 2013; Montello, 2005) and, therefore, may be easier to recognize or recall in comparison with the visually complex junctions that are more likely to be found in real environments. VEs have been shown to tap into similar cognitive mechanisms as the same tasks in the real world (Richardson et al., 1999), and route learning in a VE can be transferred to real-world environments (Montello, Waller, Hegarty, & Richardson, 2004; Ruddle et al., 1997). Therefore, we predict that the current findings will generalize to an environment in the “real” world, and future research could investigate this. However, we note that proprioceptive and vestibular information is absent in desktop VEs (Taube, Valerio, & Yoder, 2013). Such idiothetic features require consideration by employing environments in which participants are more active (Chrastil & Warren, 2012). Given the very good performance of some of the youngest children (the 8-year-olds), any further examination of this age group, or indeed younger age groups in real environments, should consider testing young children on routes that are much longer than have been used in previous real-world environments. Previous studies have shown that adults are generally good at learning a route (Karimpur & Hamburger, 2016), and this was confirmed by the performance of the adult group in the current study. Three quarters of the adult participants learned the 12-turn route after a single experience of the route. Adults’ success in the current experiment confirms that most adults can retrace a long route without error after just a single exposure to that route. Because such a large proportion of adults succeeded straightaway, it is unlikely that 12 turns is the maximum length of a route that adults can encode in one experience; therefore, future researchers should consider testing adults over much longer routes. The fact that many of the participants learned the 12-turn route after just one experience of it does not support theories of wayfinding that emphasize the need for route learners to construct a representation of a route only slowly by first learning some or all of the landmarks along a route and then combining these with actions at landmarks to form a route (Siegel & White, 1975). Rather, the rapid learning demonstrated by many participants suggests that some children and most adults can integrate landmark and turn information into a complete route on the basis of a single experience, and this would accord with Montello (2017) that there are not separate stages of landmark and route learning. The current study considered only one route and, therefore, did not examine how multiple routes might be integrated into larger representations of the environment, and so this study does not rule out the possibility of other stages in learning more complex environments. To summarize, the current study showed that most 8-year-olds (85%) and every one of the 10- and 12-year-olds and adults learned a route consisting of 12 turns. Many participants did so without making any mistakes. However, 8-year-olds made more errors and required more attempts to reach criterion than the other groups, whereas these differences among the 10-year-olds, 12-year-olds, and adults were less pronounced. These findings suggest that children from 8 years of age can learn routes, even with as many as 12 turns, very quickly. The perfectly accurate performance of even some of the youngest children, the 8-year-olds, when retracing the route after just one experience demonstrated the potential of very young children to learn a long route straightaway. Such successful performance has not been noted before; therefore, more research with younger children over longer routes would be a useful focus for future research."],["Synesthesia is a neurological condition that gives rise to unusual secondary sensations (e.g., reading letters might trigger the experience of colour). Testing the consistency of these sensations over long time intervals is the behavioural gold standard assessment for detecting synesthesia (e.g., Simner, Mulvenna et al., 2006). In 2007 however, Eagleman and colleagues presented an online 'Synesthesia Battery' of tests aimed at identifying synesthesia by assessing consistency but within a single test session. This battery has been widely used but has never been previously validated against conventional long-term retesting, and with a randomly recruited sample from the general population. We recruited 2847 participants to complete The Synesthesia Battery and found the prevalence of grapheme-colour synesthesia in the general population to be 1.2%. This prevalence was in line with previous conventional prevalence estimates based on conventional long-term testing (e.g., Simner, Mulvenna et al., 2006). This reproduction of similar prevalence rates suggests that the Synesthesia Battery is indeed a valid methodology for assessing synesthesia. --------------------------------------------------------------------------------","Synesthesia is an inherited condition in which everyday stimuli trigger unusual secondary sensations. For example, synesthetes listening to music might see colours in addition to hearing sound (Ward, Huckstep, & Tsakanikos, 2006). One particularly well-studied variant is grapheme-colour synesthesia, in which synesthetes experience colours when reading, hearing or thinking about letters and/or digits (e.g., Simner, Glover, & Mowat, 2006). Despite being first reported over two hundred years ago (by Sachs, 1812; see Jewanski, Day, & Ward, 2009) synesthesia was initially an under-researched and poorly-understood area of human experience until the last decades of the 20th century. A significant factor in the elevation of synesthesia as a tractable topic was the realisation – and subsequent empirical confirmation – that synesthetes’ experiences could be verified behaviourally by the fact that they remain conspicuously stable over time (Baron-Cohen, Wyke, & Binnie, 1987; Jordan, 1917). Specifically, synesthetes tend to be highly consistent when reporting their synesthetic sensations for any given stimulus. For example, if the letter J triggers the colour pale blue for a given synesthete, she will tend to repeat that J is pale blue (not green, not yellow, etc.) when repeatedly tested over days, months and even years. Indeed, one study was able to show that synesthetic sensations had remained consistent over at least three decades (Simner & Logie, 2008). This stability of responses over time is considered one of the central features of synesthesia and is routinely verified in almost every publication on the subject (e.g., Asher, Aitken, Farooqi, Kurmani, & Baron- Cohen, 2006; Baron-Cohen, Burt, Smith-Laittan, Harrison, & Bolton, 1996; Rich, Bradshaw, & Mattingley, 2005; Ward & Simner, 2003; but see Simner, 2012). In other words, while a wide range of behavioural approaches have been employed to assess the nature of the synesthetic experience, experimental methodologies aiming to validate synesthesia have almost exclusively focussed on the feature of consistency. Hence, researchers selecting synesthete participants for study first verify the genuineness of each case by requiring their synesthetes to demonstrate high levels of consistency over time compared to non- synesthete controls (e.g., Asher et al., 2006; Baron-Cohen et al., 1996; Simner et al., 2006). Controls are tested on analogous associations (i.e., they invent colours for the 26 letters, say, and then attempt to recall these colour associations later) and typically perform significantly worse than synesthetes. Although more than a hundred contemporary studies rely on this test of consistency for genuineness, the particular instantiation of the test has varied widely. For example, a wide range of methods have been used to elicit synesthetic colours: participants have indicated these by either giving verbal descriptions (e.g., Ward, Simner, & Auyeung, 2005), written descriptions (Simner, Glover et al., 2006), using Pantone© swatch colour charts (Asher et al., 2006), electronic colour charts (Simner, Harrold, Creed, Monro, & Foulkes, 2009) or even computerised colour pickers offering extensive palettes of >16 million colours (e.g., Simner & Ludwig, 2012). In this way, synesthesia research has used varying methods, which in turn might raise difficulties for researchers when trying to meaningfully compare data. Despite this superficial variability however, the test of genuineness has nonetheless tended to rely on one key shared feature: synesthetes must outperform controls over fairly lengthy re-test intervals. Consider, for example, the most widely cited large-scale screening for synesthesia (Simner, Mulvenna et al., 2006) in which a large sample of participants were opportunistically recruited from the communities of Edinburgh and Glasgow Universities, and individually assessed for synesthesia.","first indicated by questionnaire whether they believed they experienced synesthesia, and those who reported in the affirmative were asked to provide their synesthetic associations (e.g., the colours of letters). These participants were then retraced after considerable time had passed (on average 6.0 months) and were asked in a surprise retest to re-state their associations. A group of controls without synesthesia performed an analogous task but were re-tested after only two weeks. Synesthetes were able to significantly out-perform controls even though much time had passed and the deck was effectively stacked against them. Methodologies such as this allow confident detection of genuine synesthetes because the surprise retest over lengthy intervals places the performance of synesthetes beyond the usual abilities of the average person. The drawback to this methodology, however, is that the task is extremely time-intensive to perform, and risks a high drop-out rate if synesthetes become untraceable at retest. Perhaps for this reason, one of the most important developments in the methodology of synesthesia validation came with the introduction in 2007 of an alternative version of the test of genuineness. Eagleman and colleagues produced the Synesthesia Battery, a toolbox of online tests which provides a standardised set of questions, tests and quantitative scores to assess a range of synesthesias (Eagleman, Kagan, Nelson, Sagaram, & Sarma, 2007). This battery is again based on internal consistency in that synesthetes are validated by high consistency within their own synesthetic associations, stated repeatedly. However, consistency is measured within a single test session lasting only approximately 10 min. Specifically, synesthetes log on to the testing site (www.synesthete.org) and specify which form(s) of synesthesia they experience. The testing platform then presents their triggering stimuli (e.g., the 26 letters) one by one in randomized order, and participants are required to select their synesthetic colour for each trigger. Each stimulus (e.g., letter) is presented three times each, and a score is generated to quantify the consistency of participant’s responses (e.g., did the participant choose the same/similar colours each of the three times she saw a particular letter?) This score represents the geometric distance in RGB (red, green, blue) colour space, where R, G, and B values are all normalised to lie between 0 and 1. If the mean overall score of colour-distance is less than 1, the participant is classified as a synesthete; if the score is 1 or higher, the degree of inconsistency classifies the participant as a non-synesthete. However, it remains an open question whether this limited retest interval is sufficient to truly distinguish synesthetes from non-synesthetes. In the current study we assessed the validity of the Synesthesia Battery by using it to test almost 3000 randomly sampled subjects for grapheme-colour synesthesia. Our aim was to establish the prevalence of grapheme-colour synesthesia by this method. This will allow us to evaluate the Synesthesia Battery by comparing this prevalence – obtained by assessments within in a single test session – to the most widely accepted previous estimate of the prevalence of grapheme-colour synesthesia based on the standard longitudinal test–retest method (Simner, Mulvenna et al., 2006). If the Synesthesia Battery is just as effective a method for detecting synesthesia as the more standard long-term retest method, we anticipate an equivalent prevalence of grapheme-colour synesthesia across both methods. In carrying out our study, we chose to evaluate grapheme-colour synesthesia in particular for several reasons: it is one of the most common forms of synesthesia (Simner, Mulvenna et al., 2006), it is particularly well-understood in behavioural terms, it lends itself readily to online testing, and those who experience it typically demonstrate the high levels of consistency expected from synesthetes (compared to other variants, whose more complex concurrents may make them more difficult to assess via consistency alone; see Simner, Gäartner, & Taylor, 2011 for discussion). It was not our intention to change or try to improve upon the method made available by Eagleman et al. at www.synesthete.org. Rather, we attempted to simply replicate their test and methodology and then evaluate how it performs in comparison to a conventional longitudinal test–retest method. In evaluating the Synesthesia Battery, our data will also provide an independent test of the prevalence of synesthesia. Our baseline study – the widely cited prevalence study of Simner, Mulvenna et al. (2006) – found the prevalence of grapheme-colour synesthesia to be 1.4% (for synesthetes with both coloured letters and numbers) or 2% (for synesthetes with either coloured letters or numbers). This study was based on a sample of 500 individuals, and the prevalence rate it generated was subsequently verified by a secondary method testing a further 1190 individuals (see Section 4 for details of this second method). Two previous studies have also aimed to validate aspects of the Synesthesia Battery (Eagleman et al., 2007; Rothen, Seth, Witzel, & Ward, 2013). Both studies used self-reported synesthetes who had self-referred for study, in comparison to a group of controls declaring they were non- synesthetes. it is important to highlight the difference between self-referred and self- reported synesthetes. Self-reported synesthetes are any individuals who claim they have synesthesia. Self-referred synesthetes are those who have additionally made the effort to contact a university researcher to volunteer to take part in synesthesia studies. All synesthetes tested by Eagleman, Rothen and colleagues were not only self-declared, but also, importantly, self-referred. In comparison, none of the synesthetes tested here are self-referred. Instead, our approach is to screen the general population (some of whom at a certain point during our test, will self-report having synesthesia when asked, but will not be self-referred). There are likely to be significant differences between our own synesthetes, and the self-referred synesthetes of Eagleman, Rothen and colleagues. These latter synesthetes not only know they have coloured letters, but also know this is called synesthesia, and furthermore, they have made the effort to contact a university researcher to volunteer to take part in synesthesia studies. They therefore have an understanding of synesthesia and judge that the extent of their synesthesia is worthy of study by researchers. In other words, it is at least possible that self-referrers have relatively ‘strong’ (or noticeable, or attention-catching) synesthesia in some way and may not be entirely representative of the population of synesthetes at large. In summary, because of the sampling methods of the two previous validations (Eagleman et al., 2007; Rothen et al., 2013) their participant groups (synesthetes versus controls) may have had diametrically opposing synesthesia characteristics, which might have therefore made them relatively easy to distinguish between. Indeed, both Eagleman et al. (2007) and Rothen et al. (2013) obtained a bimodal distribution of scores when assessing the consistency of grapheme-colour associations of their self-referred synesthetes compared to controls. Here however we individually assess a randomly recruited sample of subjects, allowing our own study to extend the previous findings of Eagleman et al. and Rothen et al. and establish how the Synesthete Battery performs when a distinction between synesthetes and non- synesthetes is perhaps more difficult to achieve. Put differently, by testing a random sample of the population, we expect to capture a broader, more representative range of synesthetic experiences, and we are evaluating how the Synesthete Battery performs under these conditions. In addition, our study will provide the largest estimate of the prevalence of grapheme-colour synesthetes to date, with almost 3000 randomly sampled members of the general population.1 Finally, our study also investigates a second aspect of the Synesthesia Battery. After the single session test of consistency of coloured graphemes, participants next immediately perform a second test of synesthesia: a speeded congruency verification task (Eagleman et al., 2007). In this, participants are presented with individual graphemes, but this time the graphemes are coloured either to match the participant’s earlier colour selection (congruent), or to be a different colour (incongruent). Participants must simply answer whether the letter-colour pairing matched their previous choices or not, and their accuracy and speed is measured. Eagleman et al. (2007) report that synesthetes tend to score 90% or higher with a mean RT of 0.64 ± 0.78 s, while non-synesthetes score below 90% with mean RT of 0.91 ± 0.87 s (Eagleman et al., 2007). We will examine the sensitivity of this type of test to determine whether it too has the diagnostic capability to distinguish between participants who score below the consistency threshold score of <1 (i.e., synesthetes) and those who do not. In other words, where Eagleman et al. (2007) compared self-referred synesthetes with non-synesthete controls, we will extend this type of test to assess (a) randomly sampled participants, who are (b) self-declared synesthetes, but who are not self-referred synesthetes, and who are also (c) either verified as genuine versus non-genuine. As such, we are evaluating whether this type of speed-congruency test still holds up in what is likely to be a more sensitive comparison. Participants ~~~~~~~~~~~~ Two thousand eight hundred and forty-seven participants took part in our study (1317 male, 1530 female; mean age 28.6, range 16–90, S.D. 14.3). We had additionally tested 32 further subjects who completed our study but had entered an obviously false date of birth (e.g., 2013). These subjects did not enter our analysis, which was therefore based only on our N = 2847. Participants were recruited as part of a large-scale, centrally co-ordinated undergraduate research project. Every student registered on the 2nd year of the Psychology undergraduate course at the University of Edinburgh acted as a research assistant (RA), and was required to each recruit 8 participants (4 male and 4 female) over 16 years of age. Our student RAs were not allowed to take part in the study themselves. In recruiting our participants, we took a number of steps to ensure as random a sample as possible. First, RAs were instructed not to deliberately seek out, nor to avoid, people they knew to be synesthetes. Furthermore, in order to avoid self-referral biases, RAs were required to pre-select their sample, and then approach participants in a targeted way (rather than send out an advert and accept self-referrals). Indeed, RAs were required to refrain from recruiting participants via any open calls at all, for example, they could not post the testing URL on social media websites or internet forums. Finally, RAs were also instructed not to a priori inform participants that the study involved synesthesia. The instructions given by the RAs to prospective participants were uniform, and clearly stated that participants were only allowed to complete the test once; if they had previously been approached to complete the test by someone else, they were to inform the recruiter and not proceed with the test. Our study was carried out in two waves to maximise participation numbers: 1514 were tested in January 2013, and 1333 were tested in September 2013. Both used identical methods, carried out by two consecutive intakes of 2nd year students. In both rounds, the study was carefully managed and co-ordinated by authors JS and DAC. Data from both rounds are pooled and presented together here. The online test ~~~~~~~~~~~~~~~ The online test consisted of several sections. Participants first provided informed consent via a checkbox and then gave demographic information such as age, sex, handedness and native language. A second section consisted of a health questionnaire not relevant for the current study. (In this, subjects were requested to indicate if they suffered from a range of clinical conditions, and this was for another project to be reported elsewhere.) After this page, our online synesthesia assessment began with our locally stored replica of the Synesthesia Battery. In this replica – as in the original – participants were first asked whether they experienced grapheme-colour synesthesia, with the question “Do numbers or letters cause you to have a colour experience?” This was accompanied by an example, and then an option to accept separately according to whether these colours are triggered automatically by numbers and/or digits. If participants indicated that they saw neither letters nor numbers in colour, they advanced to an early-exit page thanking them for their participation. The rest of the test was completed by participants who answered in the affirmative to having coloured letters/digits. These participants completed two further tests tailored to the particular variant of grapheme-colour synesthesia they had reported (i.e., for either digits, letters, or both). These two tests were a colour consistency test and a speeded congruency task. The colour consistency test was again an identical clone of the consistency test from the Synesthesia Battery (Eagleman et al., 2007). In this, participants were presented with each grapheme (a–z, 0–9) three times in random order (so 30 trials if the subject reported coloured numbers only, 78 trials if letters only, and 108 trials if both letters and numbers). For each trial, participants were required to select the colour that best matched the grapheme presented (see Fig. 1). Selections were made from a palette of 256 × 256 × 256 colours, exactly as in the original Synesthesia Battery. Once their selection was submitted, the screen advanced to show the next grapheme. The colour palette followed an HSL colour model, with colours varying in lightness along the vertical axis and saturation along the horizontal access a separate, horizontal bar allowed hue to be adjusted (see Fig. 1). The colour consistency test was followed by a speeded congruency task. In this section, participants were shown again the graphemes they had just seen in the colour consistency test. This time they saw each relevant grapheme twice, in a random order, each flashed on screen for a maximum of 1 s or until the participant responded (20 trials for just numbers, 52 trials for just letters or 72 trials for both letters and numbers). In 50% of trials, graphemes were coloured congruently with the participant’s earlier specification, and in 50% of trials they were coloured incongruently. Participants were required to indicate by mouse-click on the relevant on-screen button whether the each grapheme they saw either matched or did not match their previous colour pairing (as collected during the consistency test; Fig. 2). Their response mouse-click advanced the test to the next grapheme, and the test continued until all graphemes had been shown.2 In summary, the colour consistency test generated a consistency colour-distance score, and the speeded congruency task generated an accuracy score and a reaction time. For full details of website configuration and how the consistency colour-distance score is calculated, see Eagleman et al. (2007).","In our study, we classified as non-synesthetes all those who were directed to the early- exit page (i.e., those who said they did not experience coloured letters and/or digits) and all those who continued but scored 1 or higher. The remainder were classified as synesthetes (i.e., those who scored <1). From our sample of 2847 participants, 140 subjects (55 male, 85 female, mean age 23.9, range 16–71, SD 9.5) self-reported grapheme- colour synesthesia, giving a self-reported prevalence of 4.9%. Of those 140 self-reported synesthetes, 34 obtained a colour-distance score of <1 on their consistency test (14 male, 20 female, mean age 24.9, range 17–51, SD 5.8), which is the criterion used by Eagleman et al. (2007) to identify genuine synesthesia. This places the prevalence of genuine grapheme-colour synesthesia at 1.2% and we will return to this prevalence value further below. Of the 140 self-reported synesthetes, 55 reported experiencing coloured numbers only, 58 reported experiencing both coloured numbers and letters and 27 subjects reported experiencing coloured letters only (see Fig. 3a). Of the 34 participants that scored <1 on the consistency test, 17 experienced coloured numbers only, 14 experienced both coloured numbers and letters and 3 subjects experienced coloured letters only (see Fig. 3b). As an additional check, the colour choices of the participants obtaining a consistency score of <1 were examined individually to allow us to confirm that none of these achieved their superior consistency by entering the same colour for each grapheme, or by entering an obviously non-synesthetic pattern of colours throughout, e.g. red for ‘R’, green for ‘G’ and blue for ‘B’ (following Simner, Mulvenna et al., 2006). Next we analysed accuracy and RTs in the speeded congruency task. We first divided our sample of self-reported synesthetes into two groups, around to the consistency colour-distance threshold of <1. For clarity, we refer to those who scored <1 as genuine synesthetes, and those who score ⩾1 we refer to as malingerers3 (i.e., non-synesthetes who self-reported synesthesia but failed to achieve what is considered a synesthetic score in the colour consistency test). There were 34 genuine synesthetes, as noted above, and 90 malingerers. There were also 16 subjects who reported too few coloured graphemes to generate a consistency colour-distance score at all. Following Eagleman et al., 2007, participants were required to enter a minimum of two valid graphemes to obtain a consistency score, a valid grapheme being defined as one to which the subject entered a colour for all three presentations. These 16 are omitted from all further analyses below. (Finally, we remind the reader that there are no values for participants who declared themselves to be non-synesthetes from the start because these individuals did not progress to the synesthesia assessment.) Using an independent samples t-test, we calculated the mean accuracy in the speeded congruency task for each group: genuine synesthetes versus malingerers. This mean was 84.4% (SD = 11.8) for synesthetes and 70.5% (SD = 14.9) for malingerers, and this difference was significant (t(122) = 5.45, p < .0001, d = 1.03). There was no significant difference in mean reaction times between groups (Synesthetes M = 1.93 s; SD = 0.77; Malingerers M = 1.76 s, SD = 0.73; t(122) = 1.13, p = .13, d = 0.23) see Fig. 4a and b). Mean reaction time was calculated across all trials, irrespective of whether the participant had answered correctly or incorrectly.4 In order to investigate whether this non-significant result provided evidence for the null hypothesis we calculated a Bayes factor. By comparing the likelihood of two models (in this case, the null and alternative hypotheses) as a ratio, Bayes factors allow the researcher to evaluate to what extent the data supports the null hypothesis (Rouder, Speckman, Sun, Morey, & Iverson, 2009). Following Jeffreys (1961), a Bayes factor of less than 0.33 provides strong support for the null hypothesis, a Bayes factor of greater than 3 provides support for the alternative hypothesis and values inbetween indicate the data are insensitive and no firm conclusions should be drawn. Using the online calculator provided by Rouder et al. (2009), we calculated a Bayes factor of 1.17, indicating that the data are not sensitive enough to enable a conclusion to be drawn. Finally we explored how these two types of test (consistency and speeded congruency) work in tandem in their assessments of synesthesia. If we take self-reported synesthetes as a single group (i.e., collapsing genuine synesthetes and malingerers) there was a significant inverse correlation between the consistency colour-distance score, and the speeded congruency accuracy score (r(122) = −.71, p < .001; see Fig. 5). Remembering that low colour-distance scores and high accuracy scores are both indicative of synesthesia, this inverse correlation shows that those who performed like synesthetes in the first sub-test were also more likely to perform like synesthetes in the second. However, when we calculate this correlation for genuine synesthetes and malingerers separately, it becomes apparent that the effect comes from the malingerer group only. The correlation between consistency colour distance score and speeded-congruency accuracy score for genuine synesthetes is non-significant (r(32) = −.07, p = .35) whereas for the malingerers, the correlation is highly significant (r(88) = −.73, p < .0001). Finally, there was no significant correlation between consistency colour-distance score and reaction time (r(122) = −.05, p = .59), nor for accuracy score and reaction time (r(122) = −.11, p = .22), and this also remained the case when correlations for the synesthete and malingerer groups were calculated separately (synesthete consistency-RT = r(32) = .006, p = .49; malinger consistency-RT = r(88) = .014, p = .45) (synesthete accuracy-RT = r(32) = −.18, p = .14; malinger accuracy-RT = r(88) = −.09, p = .19).5 Prevalence comparison between studies: Comparing short versus long-term testing ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The key aim of this study was to compare the prevalence of grapheme-colour synesthesia generated by the Synesthesia Battery (a single-session test) to a more conventional method based on long term (rather than single session) testing (Simner, Mulvenna et al., 2006). Using the Synesthesia Battery we found the prevalence of grapheme-colour synesthesia in the general population to be 1.2% for those with coloured letters or digits, compared to the previous estimate of 2% in conventional longer-term testing (Simner, Mulvenna et al., 2006). We can now evaluate that these two estimates are not significantly different (chi square = 2.1; df = 1; p = .14). However, we also calculated prevalence for synesthetes with both coloured letters and digits, finding a value here of 0.5%, and this is significantly less than the 1.4% found in conventional longer term retesting (Simner, Mulvenna et al., 2006; chi square = 5.6; df = 1; p = .02).","In our study, we reproduced elements of the Synesthesia Battery (Eagleman et al., 2007) which is a single-session online test for grapheme-colour synesthesia. We used this method to estimate the prevalence of grapheme-colour synesthesia in a very large randomly recruited sample – indeed the largest sample for this purpose to date. We found the prevalence of grapheme-colour synesthesia in the general population to be 1.2% for those with coloured letters or digits, compared to the previous estimate of 2% (Simner, Mulvenna et al., 2006), a difference which is not significantly different. However, we also calculated prevalence of synesthetes with both coloured letters and digits, finding a value here of 0.5%, and this is significantly less than the 1.4% found in conventional longer term retesting (Simner, Mulvenna et al., 2006). Hence, the Synesthesia Battery numerically under-estimated the prevalence of those with coloured letters and digits, compared to longer retesting methods and we explore possible reasons for this below. One explanation for an under-estimation of prevalence in the Synesthesia Battery might stem from the way graphemes are presented to those who report more than one variant: letters and digits are presented together, randomly ordered within the same consistency test. In the longer term retesting used as the baseline here however, letters and digits were always presented in separate blocks (Simner, Mulvenna et al., 2006). It may therefore be that synesthetes are susceptible to interference when selecting colours for graphemes, and we raise this possibility for future studies to consider. A second possible explanation for the lower prevalence found here is that immediate retests might be inherently more conservative, and this might be suggested by the fact that prevalence estimates here were numerically lower across the board. Immediate retesting might be conservative because the scores of non-synaethetes are likely to be better across shorter (versus longer) time intervals. In other words, non-synesthetes in an immediate retest (e.g., over 10 min) should score higher than in a delayed retest (e.g., over several weeks) even though synesthetes’ scores, in comparison, would be likely to remain relatively unchanged. This would raise the control baseline for single session testing, and therefore allow a smaller range of synesthetes to be identified as significantly more consistent. If this is correct, we conclude that the threshold at which synesthetes are identified (currently around the score of 1) might be usefully set marginally higher in the Synesthesia Battery of Eagleman et al. (2007). One recent study has also speculated that the threshold in the Synesthesia Battery perhaps should be raised. Rothen et al. (2013) have recently argued that for the combination of RGB colour space and city block distance used by Eagleman et al. (2007) – and also in this study – a higher cut-off may indeed be more appropriate. In order to maximise sensitivity and specificity for this colour space/distance combination, they proposed a revised threshold within the existing Synesthesia Battery at <1.43 for synesthetes (rather than <1). In our own study, if we recalculate prevalence in our sample using Rothen et al’s suggested cut-off score of 1.43, our prevalence estimates are no longer significantly lower than expected. With this revised threshold, we now find 69 genuine synesthetes with coloured letters OR digits (2.4%; compared to 2% in Simner, Mulvenna et al., 2006; chi square = 0.3; df = 1; p = .57) and 31 synesthetes with coloured letters AND digits (1.1%; compared to 1.4% in Simner, Mulvenna et al., 2006; chi square = 0.367; df = 1; p = .5). In other words, our results are more in line with more conventional longer term evaluations of synesthesia when the threshold is shifted upwards according to the proposal of Rothen et al. (2013).6 Finally, we point out that Rothen and colleagues also suggest that using alternative colour models (CIELUV and CIELAB) and Euclidean distance provide the best combination of sensitivity and specificity in distinguishing between synesthetes and non-synesthetes. However that change would be beyond the scope of this paper, where we are simply evaluating the Eagleman data output as it stands. Our discussion above of where to set the cut-off threshold for consistency in the Synesthesia Battery might be considered part of a more fundamental concern in synesthesia research. In all kinds of consistency tests for synesthesia – and particularly obvious here – participants are divided into two groups by their score on what is in fact an incremental continuum of possible scores. (Using this test, participants can, in theory, score any value between 0 and ∼4, even though the cut-off is conventionally placed at the fixed value of 1). It seems clear that someone who scores 1.05 on the consistency test is “more synesthetic” than someone scoring 2.48, yet according to the cut-off of <1, both would be considered non-synesthetes. Nonetheless, it is a particular strength of the Synesthesia Battery that researchers are free to consider this score in its own right, rather than for categorical groupings alone. The second part of Eagleman et al’s test involved speeded congruency task in which graphemes are presented either in the same colour previously selected (congruent) or a different colour (incongruent). Subjects must indicate whether the colour they saw matched their earlier choice, and Eagleman et al. (2007) report significant differences in speed and accuracy between a pre-selected group of 15 self-referred synesthetes, and a non-synesthete control group. In our current study, when our randomly-sampled respondents were divided into genuine synesthetes versus malingerers (i.e., around the threshold score of 1), genuine synesthetes were again significantly more accurate than malingerers. These data suggest that accuracy scores can not only distinguish between self-referred synesthetes and non-synesthetes (as in Eagleman et al., 2007) but is also subtle enough to distinguish between groups of genuine synesthetes and those who claim to be so, but do not pass a conventional consistency test. Furthermore, considering all self-reported synesthetes irrespective of their consistency, there was a significant inverse correlation between consistency and accuracy, indicating that more consistent (i.e., more ‘genuine’) synesthetes were also more accurate. However, when this correlation was calculated separately for the each group (genuine synesthetes and malingers) it became apparent that the effect was driven by the malingerers only. We suggest this is because genuine synesthetes score highly on the accuracy test, irrespective of what consistency score they achieve. In other words, a synesthete obtaining a consistency score of, say, 0.99 is likely to be highly accurate on the speeded-congruency test – as accurate as a synesthete scoring 0.4 on the consistency test. In contrast, a malingerer obtaining a consistency score of 1.5 is more likely to be more accurate on the speeded-congruency test than a malingerer scoring 3.5 for consistency. The speed accuracy congruency test presented here showed one notable difference in results compared to Eagleman et al. (2007). Our own findings were that genuine synesthetes, although more accurate than malingerers, were no faster. Eagleman and colleagues found genuine synesthetes to be more accurate and faster than controls. This difference to Eagleman et al. (2007) may stem from differences in our control populations: Eagleman et al. (2007) compared synesthetes to self-declared non-synesthetes, while we compared to ‘malingerer’ individuals claiming to have synesthesia. It is possible that some portion of our controls were in fact synesthetic in some way, albeit with lower consistency, and perhaps this is why our genuine synesthetes did not differ from them in their RTs (see above and Simner, 2012 for discussion). Alternatively, our lack of difference in RTs may be the result of the unanticipated variations we introduced in our version of this test. Our method of selecting congruent and incongruent graphemes and the timing of grapheme presentation for this part of the test differed slightly from Eagleman et al.’s original approach (see Section 2 for a full explanation). Our own version may have raised the difficulty of the task (e.g., because graphemes were on-screen for a shorter time on average) and our RTs were certainly longer and hence potentially more noisy. There is no way to distinguish between these two hypotheses in the current study and so we leave this as an open question for future studies to address. We have evaluated whether single session tests of consistency are effective at identifying synesthetes, compared to established longer retesting methods. One previous study has also suggested that single session testing may indeed be valid. Simner, Mulvenna et al. (2006) established the prevalence of synesthesia both with long term testing (which we used here as our key comparison) but also by screening additional 1190 people in a single session. Their method was more basic than that of Eagleman et al. (2007) in that colour choices were made from a palette of just 13 colours, and only absolute matches contributed to consistency scores. Nonetheless, this again produced roughly comparable prevalence estimates as longer term testing (1.1% prevalence for coloured letters and digits). Taken together with the current study, we therefore suggest single session tests of consistency for synesthetic associations do appear to provide an appropriate method by which to identify synesthetes. The widely available nature of the Synesthesia Battery through its open-access online interface makes it a particularly appealing version of this test, as does its comprehensive colour palette, and its ability to give a calibrated estimate (i.e., continuous consistency score) for synesthesia status. Although researchers will want to consider carefully the question of whether consistency testing can reliably identify every type of synesthesia, or indeed every type of synesthete (see Simner, 2012 for discussion), it is clear from our current study that the Synesthesia Battery provides a suitable tool for evaluating synesthetes along this dimension."],["Paired-associate learning (PAL) tasks measure the ability to form a novel association between a stimulus and a response. Performance on such tasks is strongly associated with reading ability, and there is increasing evidence that verbal task demands may be critical in explaining this relationship. The current study investigated the relationships between different forms of PAL and reading ability. A total of 97 children aged 8–10 years completed a battery of reading assessments and six different PAL tasks (phoneme–phoneme, visual–phoneme, nonverbal–nonverbal, visual–nonverbal, nonword–nonword, and visual–nonword) involving both familiar phonemes and unfamiliar nonwords. A latent variable path model showed that PAL ability is captured by two correlated latent variables: auditory–articulatory and visual–articulatory. The auditory–articulatory latent variable was the stronger predictor of reading ability, providing support for a verbal account of the PAL–reading relationship. --------------------------------------------------------------------------------","The ability to create and consolidate associations between letters and corresponding speech sounds is an essential component of learning to read (Melby-Lervåg, Lyster, & Hulme, 2012; Muter, Hulme, Snowling, & Stevenson, 2004). Individual differences in letter–sound knowledge are a powerful predictor of reading success (Lervåg, Bråten, & Hulme, 2009; Muter et al., 2004). Paired-associate learning (PAL) tasks measure the ability to form novel associations between stimuli and responses. Such associations may be unimodal (between either visual or auditory stimuli) or cross-modal (between a visual stimulus and an auditory stimulus). Learning paired associates depends on learning both the individual stimuli and the association between them (Hülse, Egeth, & Deese, 1980). Many studies have shown that performance on PAL tasks predicts children’s word reading ability, and evidence suggests that PAL taps a mechanism, distinct from phonological awareness, that is also important for learning to read (Lervåg et al., 2009; Warmington & Hulme, 2012; Windfuhr & Snowling, 2001). Indeed, it has been suggested that the cognitive processes underlying performance on PAL tasks reflect the very nature of learning to read—the generation of novel associations between letters (and letter strings) and phonological speech output (Ehri, 1992; Hulme & Snowling, 2013a; Snowling, 2000). In previous studies, two different views have been taken about the nature of the relationship between PAL and reading. One view is that this relationship reflects a role for cross- modal learning as a fundamental process underlying reading development (e.g., Hulme, Goetz, Gooch, Adams, & Snowling, 2007). A second view is that the PAL–reading relationship depends specifically on verbal, or phonological, learning mechanisms (Litt, de Jong, van Bergen, & Nation, 2013). There is some evidence that performance on tasks involving cross- modal PAL is a stronger predictor of reading as compared with other unimodal PAL tasks. A study by Hulme et al. (2007) investigated the relationship between reading and three PAL conditions: two unimodal (visual–visual and verbal–verbal) and one cross-modal (visual–verbal). Of the three conditions, visual–verbal PAL was most strongly correlated with reading ability in typically developing children, although verbal–verbal PAL was also correlated, albeit less strongly, with reading. Importantly, performance on visual–verbal PAL was a unique predictor of word reading even after controlling for performance on verbal–verbal PAL and phoneme awareness. Therefore, the authors suggested that the PAL–reading relationship was specific to learning associations between visual (orthographic) and verbal (phonological) representations. This cross-modal hypothesis is consistent with the important role of letter–sound knowledge in predicting early reading ability because acquiring letter knowledge also depends on the formation of cross-modal visual–verbal associations (Hulme & Snowling, 2013b). In addition, the finding that visual–verbal PAL is a unique predictor of reading after controlling for phoneme awareness is in line with previous research (e.g., Windfuhr & Snowling, 2001) and suggests that PAL ability depends on skills that are, at least in part, separable from children’s phonological skills or the quality of stored phonological representations. In addition, there is good evidence that, relative to typically developing controls, children with dyslexia struggle to learn visual–verbal associations (Mayringer & Wimmer, 2000; Vellutino, Scanlon, & Spearing, 1995; Wimmer, Mayringer, & Landerl, 1998). For example, Messbauer and de Jong (2003) reported that children with dyslexia perform worse on measures of visual–verbal PAL compared with a chronological-age-matched control group. Children in this study completed three PAL tasks; two cross-modal (visual–word and visual–nonword) and one unimodal (visual–visual). Children with dyslexia performed worse on both visual–verbal PAL tasks (involving words or nonwords) but did not differ from chronological-age- and reading-age-matched control groups on the visual–visual PAL task. Impaired performance on both visual–verbal PAL tasks might suggest that a cross-modal learning mechanism is important in explaining the PAL–reading relationship. However, performance on such cross-modal PAL tasks also involves verbal learning, whereas the visual–visual task involves only nonverbal stimuli and responses. In addition, Messbauer and de Jong reported that when differences in phonological awareness were taken into account, group differences on visual–verbal PAL tasks disappeared. Therefore, these findings question the notion that cross-modal associative learning drives the PAL–reading relationship. Rather, differences in verbal or phonological processing may be key. Although the cross-modal account clearly has some support, the alternative verbal account arguably has stronger support. The verbal learning account argues that it is individual differences in learning verbal information that differentiates poor readers from good readers. Litt et al. (2013) reported a study in which children learned pairs of stimuli across four experimental conditions (verbal–verbal, visual–visual, visual–verbal, and verbal–visual) in order to dissociate modality and task demands. Verbal stimuli were consonant–vowel–consonant (CVC) nonwords, and visual stimuli were simple letter-like symbols. Correlations with word reading were found only when verbal output was required (verbal–verbal and visual–verbal conditions). Furthermore, performance in the verbal output PAL conditions predicted significant variance in reading accuracy above and beyond known predictors of reading such as phoneme awareness and rapid automatized naming. The unimodal (verbal–verbal) PAL condition did not involve learning any cross-modal associations. Thus, findings from this study provide strong evidence that verbal learning, rather than cross-modal learning, is the most critical component of the PAL–reading relationship. Further evidence in support of this notion comes from the finding that children with dyslexia are impaired on verbal PAL tasks but not on nonverbal PAL tasks (Litt & Nation, 2014; Mayringer & Wimmer, 2000; Vellutino, Steger, Harding, & Phillips, 1975). Across studies, poor readers consistently perform worse on verbal PAL tasks than age-matched typical readers. For example, in one study children were given two cross-modal PAL tasks; visual–verbal and visual–auditory (Vellutino et al., 1975). Children with dyslexia showed deficits only in the visual–verbal task, but not in the visual–auditory task, which involved imitating nonlinguistic sounds (e.g., high hum, cough), suggesting that reading difficulties may be specifically associated with impaired verbal (phonological) learning. Importantly, both conditions required cross-modal learning in addition to oral output. In line with this finding, more recent research indicates that children with dyslexia make more phonological errors, rather than associative errors, in visual–verbal PAL tasks, implying that their poorer performance is driven by difficulties with the verbal demands of the task rather than with associative learning (Litt & Nation, 2014). In summary, there is clear evidence to suggest that verbal learning mechanisms may be important for explaining the PAL–reading relationship. However, to our knowledge no existing studies have combined both cross-modal and unimodal and verbal versus nonverbal PAL tasks. In addition, studies do not consistently address response modality (and therefore response demands), which may be an important determinant of PAL performance. For example, some “nonverbal” PAL tasks have involved learning associations between pairs of visual symbols or pictures, requiring children to point to the correct response item. In other instances, a completely different response, such as drawing the PAL symbol, is required (i.e., Messbauer & de Jong, 2003). Such inconsistencies make it difficult to draw firm conclusions about the mechanisms underlying performance on nonverbal PAL tasks. The current study evaluated whether the PAL–reading relationship is primarily driven by verbal learning demands (e.g., Litt & Nation, 2014; Litt et al., 2013) or cross-modal learning demands (e.g., Hulme et al., 2007). The study included both unimodal and cross-modal PAL conditions: phoneme–phoneme, visual–phoneme, nonverbal–nonverbal, visual–nonverbal, nonword–nonword, and visual–nonword. The use of individual phonemes as stimuli extends previous studies that have typically used nonword stimuli; the visual–phoneme task can be seen as directly analogous to the process of learning letter–sound relationships. As in previous studies, nonword stimuli were three-letter CVC strings (e.g., hib), allowing us to investigate whether learning novel verbal information is a critical predictive component in the PAL–reading relationship. If the PAL–reading relationship is driven by verbal demands, performance in unimodal phoneme and nonword conditions should correlate most strongly with reading measures relative to the nonverbal PAL conditions. On the other hand, if the cross-modal conditions (including nonverbal PAL) correlate most strongly with reading, this would provide support for the cross-modal hypothesis.","A total of 97 children (49 boys and 48 girls) aged 8 years 0 months to 10 years 9 months (M = 9 years 2 months, SD = 11 months) participated in the study. Children were recruited from Years 4 and 5 in two state primary schools serving socially diverse catchment areas in Hertfordshire, England. Reading Children completed the sight word efficiency (SWE) and phonemic decoding efficiency (PDE) subtests from the Test of Word Reading Efficiency (TOWRE-2; Torgesen, Rashotte, & Wagner, 1999). In this task, children were required to read as many words (SWE) or nonwords (PDE) as possible in 45 s. Children also completed the Single Word Reading Test 6–16 (SWRT6-16; Foster, 2007), in which they needed to read aloud a list of words in increasing difficulty. Testing was discontinued after five consecutive incorrect responses. Estimates of reliability for these standardized measures of reading are .98 (Cronbach’s alpha) for the TOWRE-2 and .90 (test–retest) for the SWRT6-16. PAL tasks Children completed six PAL tasks (phoneme–phoneme, visual–phoneme, nonverbal–nonverbal, visual–nonverbal, nonword–nonword, and visual–nonword), each presented as a computerized game. In each task, children were presented with four pairs of items to learn. In the visual–articulatory PAL tasks (visual–phoneme, visual–nonverbal, and visual–nonword), an unfamiliar symbol was presented on the computer screen and children were required to say the corresponding target sound (phoneme, nonword, or nonverbal sound) paired with that symbol. In auditory–articulatory PAL tasks (phoneme–phoneme, nonverbal–nonverbal, and nonword–nonword), the auditory target stimulus was played and children were required to produce the corresponding paired sound. The nonverbal–articulatory sounds included nonspeech sounds (e.g., lip pop, cough). Children were tested on 6 consecutive school days for approximately 15 min and completed one PAL condition on each day as well as a standardized task from the test battery. The sequence of conditions was counterbalanced using a Latin square. The program randomly generated stimulus pairs for each child across the conditions. Each of the six PAL tasks involved children learning to produce the correct sound (a phoneme, nonword, or nonspeech sound) in response to a visual stimulus (a letter-like form) or an auditory stimulus (a phoneme, nonword, or nonspeech sound). In each condition, before teaching children any associations between item pairs, children were presented with each of the auditory stimuli used in that task and asked to reproduce it (they were required to repeat, one at a time, the four auditory stimuli used in each of the visual–articulatory conditions or the eight auditory stimuli used in each of auditory–articulatory conditions). In the rare event that a child had difficulty in articulating one of the auditory stimuli, the experimenter provided a correct demonstration and asked the child to try again. After this, children moved on to the learning trials. These began with a single presentation of each of the four pairs of stimuli the children were to learn. Children then received 24 test study trials. On test study trials, children were presented with each of the four stimuli and were required to produce the corresponding paired response sound. After children responded (irrespective of whether their response was correct or incorrect), the correct pairing was re- presented to reinforce learning. Children’s responses were recorded for each trial (correct, incorrect, or no response). Stimuli Visual stimuli were 12 unfamiliar symbols (800 × 600 pixels) adapted from Taylor, Plunkett, and Nation (2011). These stimuli are listed in the Appendix. All auditory stimuli were recorded by a female native English speaker in a sound-attenuated booth and included 12 phonemes, 12 nonverbal sounds, and 12 nonwords. Nonverbal sounds were adapted from Vellutino et al. (1975) and consisted of sounds that did not involve phonemes and could be easily produced. These sounds were high hum, low hum, smooch, raspberry, cough, blow, pop with lips, gasp, tut, tongue click, sigh, and sucking front teeth. Phonemes consisted of /kə/, /bə/, /pə/, /fə/, /gə/, /nə/, /rə/, /sə/, /wə/, /lə/, /jə/, and /mə/. Nonwords were CVC nonwords taken from the ARC Nonword Database (Rastle, Harrington, & Coltheart, 2002) as used in previous PAL studies (e.g., Litt et al., 2013): /hɪb/, /dʒɒf/, /kæg/, /kæv/, /lɒm/, /mɪb/, /næl/, /pel/, /tʌs/, /vek/, /jɪz/, and /jʌt/.","Stimuli were presented and responses were recorded using a Visual Basic program on a Dell laptop (Latitude E5520) running Windows 7. Auditory stimuli were presented through Beyerdynamic headphones (DT 770).","We first present descriptive statistics and correlations for all measures before presenting the main analyses, which use structural equation models to investigate the relationship between reading ability and different aspects of PAL. Descriptive statistics for all measures are shown in Table 1. Children performed at an age-appropriate level on measures of reading. There were small amounts of missing data due to occasional absences from school across the 6 consecutive days of testing and due to technical difficulties that resulted in the loss of PAL data for 2 children. Correlations between all measures are shown in Table 2 (simple correlations below the diagonal and partial correlations controlling for age above the diagonal). There was a wide range in performance across the PAL conditions. Performance was higher on the visual–articulatory PAL conditions compared with the auditory–articulatory conditions, and performance varied in both sets of conditions according to the type of response (phoneme > nonverbal sound > nonword). To investigate differences in accuracy, we performed a repeated-measures analysis of variance (ANOVA) with modality (2 levels: auditory–articulatory or visual–articulatory) and response type (3 levels: nonword, nonverbal, or phoneme) as within-participant variables. There was a main effect of modality, with performance on the visual–articulatory conditions being better than performance on the auditory–articulatory conditions, F(1, 69) = 250.91, p < .001, partial η2 = .784. There was also a main effect of response type, F(2, 138) = 281.45, p < .001, partial η2 = .168, indicating a significant difference in accuracy across the three stimulus types (with phoneme responses being by far the easiest). This main effect of response type was qualified by a significant interaction between modality and response type, F(2, 138) = 17.79, p < .001, partial η2 = .205, which reflects the fact in the visual–articulatory conditions the ordering of difficulty (phoneme > nonword > nonverbal) differed from that in the auditory–articulatory conditions (phoneme > nonverbal > nonword). This interaction reflects the fact that requiring children to associate two different nonverbal sounds was a particularly difficult learning task. Relationships between PAL measures and reading ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The latent variable path model shown in Fig. 1 was estimated with Mplus 7.4 (Muthén & Muthén, 1998–2016). Little’s MCAR test confirmed that the small number of missing values could be considered missing completely at random, χ2(98) = 115.1759, p = .11, and the small number of missing values (n = 16) was handled with full information maximum likelihood estimation. The theory we wished to test was that an auditory–articulatory PAL could be distinguished from a visual–articulatory PAL factor and that the visual–articulatory factor would show the strongest relationship with reading ability. As a first step to developing the path model shown in Fig. 1, a confirmatory factor analysis model was estimated with the six PAL tasks defining two correlated latent variables: visual–articulatory PAL (visual–phoneme, visual–nonverbal, and visual–nonword) and auditory–articulatory PAL (phoneme–phoneme, nonverbal–nonverbal, and nonword–nonword). Adding a covariance between the two measures that involved a nonverbal response (nonverbal–nonverbal and visual–nonverbal) resulted in a model with a good fit, χ2(11) = 13.907, p = .24, root mean square error of approximation (RSMEA) = .053 (90% confidence interval [CI] = .000–.128), comparative fit index (CFI) = 0.97, standardized root mean residuals (SRMR) = .051. Therefore, this structure was used for the path model shown in Fig. 1. In the final model shown in Fig. 1, all latent variables were regressed on age (although these regressions are not shown in the figure); hence, the model represents relationships between the latent variables that are independent of the shared variance attributable to age. An initial version of this model included paths from both PAL latent variables to reading. However, in this initial model, the path weight from auditory–articulatory PAL to reading was substantial and significant (.536, p = .043), whereas the path weight from visual–articulatory PAL to reading was negligible in size and not significant (−.087, p = .754). Dropping this path resulted in a nonsignificant change in model fit, χ2 difference(1) = .103, p = .75. Therefore, we used this simplified model where the nonsignificant path had been dropped. In this model, after controlling for the effects of age, the two PAL latent variables are quite highly correlated with each other (r = .79), but auditory–articulatory PAL showed a stronger correlation with reading (r = .43) than visual–articulatory PAL (r = .36). The model accounts for 33% of the variance in reading skills and provides an excellent fit to the data, χ2(29) = 29.014, p = .46, RSMEA = .002 (90% CI = .000–.078), CFI = 1.00, SRMR = .057.","This study explored the role of different types of PAL tasks as predictors of reading ability in children. More specifically, we examined the role of different types of associative learning (auditory–articulatory vs. visual–articulatory) and the type of response required (phoneme, nonword, or nonverbal sound) as determinants of the strength of relationship between reading and PAL. The findings from the path model are clear in showing that an auditory–articulatory PAL latent variable is a strong predictor of reading ability (accounting for 33% of the variance). However, after controlling for the effects of auditory–articulatory PAL, the visual–articulatory PAL latent variable accounted for no additional variance. This pattern contradicts earlier claims (Hulme et al., 2007) that cross-modal PAL plays an especially important role in learning to read and supports the view from later research that PAL tasks involving verbal learning are the ones most closely related to learning to read (e.g., Kalashnikova & Burnham, 2016; Litt et al., 2013; Messbauer & de Jong, 2003). The pattern of correlations in Table 2 shows that the PAL tasks with higher auditory–articulatory learning demands show the strongest relationship with reading ability. Specifically, among the auditory–articulatory PAL tasks, the strongest PAL–reading correlation was observed for the nonword–nonword PAL condition, and the lowest correlation was for the phoneme–phoneme condition. Arguably, the amount of phonological information that needs to be retained in memory is far higher in the nonword–nonword condition than in the phoneme–phoneme condition. Phonemes, in contrast to nonwords, are short and highly familiar forms and, therefore, are less demanding to learn. It is interesting that among the auditory–articulatory PAL conditions the nonverbal–nonverbal task was a moderate correlate of reading ability (and stronger than the phoneme–phoneme PAL condition). We selected this stimulus category for being articulatory but nonverbal; however, it seems that the processing demands of learning these nonverbal stimuli share something in common with learning about verbal stimuli (phonemes or nonwords). The fact that nonverbal–nonverbal PAL correlates better with reading than the phoneme–phoneme condition suggests that something akin to the load on memory (where load reflects both stimulus familiarity and complexity) is driving the relationship between PAL and learning to read. This notion of memory load also appears to account for the pattern of relationships in Fig. 1. The auditory–articulatory latent variable, which shows the strongest relationship with reading, involves measures with a greater verbal–articulatory load than the tasks defining the visual–articulatory variable, which relates to reading less strongly. In contrast to visual–articulatory PAL tasks, successful performance on auditory–articulatory PAL depends on children learning both stimulus and response when items are confusable (in the same modality/phonologically similar). Therefore, these conditions involve the highest level of phonological competition and, in turn, place the greatest demands on memory. Although children demonstrated significantly higher accuracy in the visual–articulatory conditions, there was still a reasonable distribution of scores across these conditions (i.e., children were not performing at ceiling); therefore, it is unlikely that differences in task difficulty can account for these results. It is possible that increased memory load is driving the relationship between PAL and reading. However, an alternative theory is that both nonword and nonverbal PAL tasks involve learning the associated articulatory gestures of novel sounds, which may also be implicated in learning to read. According to the motor theory of speech perception (Liberman, 1999), phonemes are encoded as articulatory gestures, and (in line with this) studies have demonstrated improved visual word recognition following training in analyzing articulatory gestures (Boyer & Ehri, 2011; Castiglioni-Spalten & Ehri, 2003). An alternative explanation for this finding is that children were referring to familiar or preexisting verbal labels (e.g., the words “cough” and “tut”) when retrieving the nonverbal sounds in memory rather than encoding and retrieving the actual nonverbal PAL stimuli. Given this possibility, it cannot be argued that this condition performs the function of being entirely nonverbal. However, that is not to say that performance in this condition depends entirely on verbal learning. For example, children may remember a verbal label and its associated meaning and, therefore, may be engaging additional skills rather than simply relying on phonological memory (see Laing & Hulme, 1999). It is clear that there are challenges in creating a nonverbal analogue of PAL while keeping response modality (i.e., articulatory production) consistent, although further research is needed to investigate nonverbal learning mechanisms and the possible role of an articulatory learning mechanism in learning to read. In summary, the results presented here are consistent with recent accounts and provide clear support for the role of verbal learning in explaining the PAL–reading relationship (Litt & Nation, 2014; Litt et al., 2013). We found that an auditory–articulatory latent variable was a stronger predictor of reading ability than the cross-modal visual–articulatory latent variable. However, we also found a strong correlation between reading and nonverbal–nonverbal PAL. This seemingly provides counterevidence for the verbal account and highlights the methodological advantage of the current study in comparing multiple PAL tasks. Thus, in conclusion, the current study provides support for the verbal account of the PAL–reading relationship. However, our results introduce the idea that articulatory learning might be an important demand implicated in both verbal PAL and reading; as such, further research is required to clarify the PAL–reading relationship."],["Background Literacy impairments in dyslexia have been hypothesized to be (partly) due to an implicit learning deficit. However, studies of implicit visual artificial grammar learning (AGL) have often yielded null results. Aims The aim of this study is to weigh the evidence collected thus far by performing a meta-analysis of studies on implicit visual AGL in dyslexia. Methods and procedures Thirteen studies were selected through a systematic literature search, representing data from 255 participants with dyslexia and 292 control participants (mean age range: 8.5–36.8 years old). Results If the 13 selected studies constitute a random sample, individuals with dyslexia perform worse on average than non-dyslexic individuals (average weighted effect size = 0.46, 95% CI [0.14 … 0.77], p = 0.008), with a larger effect in children than in adults (p = 0.041; average weighted effect sizes 0.71 [sig.] versus 0.16 [non-sig.]). However, the presence of a publication bias indicates the existence of missing studies that may well null the effect. Conclusions and implications While the studies under investigation demonstrate that implicit visual AGL is impaired in dyslexia (more so in children than in adults, if in adults at all), the detected publication bias suggests that the effect might in fact be zero. --------------------------------------------------------------------------------","Individuals with dyslexia have severe and persistent difficulties with learning to read and spell. These difficulties occur despite normal intelligence, adequate educational and socio-economic opportunities, and in absence of sensory or neurological impairment1 (DSM- IV; American Psychiatric Association, 2000). A generally accepted hypothesis is that the persistent difficulties with written language result from a core deficit in phonological processing and, specifically, phonological awareness (see Melby-Lervåg, Lyster, & Hulme, 2012 for a meta-analysis). Phonological awareness is the ability to detect and manipulate phonological segments of words (Shankweiler et al., 1995) and is related to the ability to map letters to sounds, which in turn affects the ability to learn to read and spell. Individuals with dyslexia also experience difficulties in other areas of language. Subtle problems have been reported in the area of inflectional morphology (e.g. pluralization and tense marking: Joanisse, Manis, Keating, & Seidenberg, 2000; subject-verb agreement: Rispens & Been, 2007; Rispens, Roeleven, & Koster, 2004) and syntax (relative clauses: Mann, Shankweiler, & Smith, 1984; Stein, Cairns, & Zurif, 1984, passive sentences: Stein et al., 1984; binding: Waltzman & Cairns, 2000). Additionally, dyslexia is associated with a range of non-linguistic cognitive dysfunctions, including impairments in visual and auditory processing (Stein & Walsh, 1997; Tallal, 2004), attention (Facoetti a Paganoni & Lorusso, 2000), motor functioning (Ramus, Pidgeon, & Frith, 2003), and verbal working memory (Gathercole & Alloway, 2006; Gathercole & Baddeley, 1990; Swanson & Jerman, 2007). Several theories have attempted to define the underlying deficit that accounts for the range of problems experienced by individuals with dyslexia. One recent approach is explaining dyslexia as the result of a problem with implicit learning (see Nicolson & Fawcett, 2007; Ullman & Pierpont, 2005). The term implicit learning refers to the process through which humans extract rules and regularities from visual and auditory sequences available in the environment. Importantly, this happens in absence of awareness. Implicit learning and literacy acquisition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Many studies have related implicit learning abilities to different aspects of language acquisition: the ability to segment words from continuous speech (Saffran, Aslin, & Newport, 1996), the acquisition of phonological categories and phonotactics (Nicolson & Fawcett, 2007; Wijnen, 2013), vocabulary acquisition (Evans, Saffran, & Robe-Torres, 2009; Yu, 2008), and more general language processing (e.g. passives: Kidd, 2012; relative clauses: Misyak, Christiansen, & Tomblin, 2010). Most important to the present discussion is the relationship between implicit learning and the acquisition of literacy skills, as these are the skills most affected in individuals with dyslexia. Learning to read and spell involves the mapping between letters and sounds (grapheme-to-phoneme mapping), which requires phonological awareness and knowledge of the orthographic system. This mapping, and the writing system in general, comprises many regularities. For example, a single letter (e.g. ‘c’) can map onto several phonemes (e.g./k/,/s/). Whether the letter ‘c’ is realized as a/k/or an/s/, depends on co-occurring letters (e.g. the letter ‘c’ followed by the letter ‘a’ generally results in the realization of the phoneme/k/as in can’t, but in the phoneme/s/when followed by an ‘e’ as in cent). In other words, the writing system consists of a “set of correlations that determine the possible co-occurrences of letter sequences, which eventually result in establishing orthographic representations” (Frost, Siegelman, Narkiss, & Afek, 2013, p. 2). Although some of these regularities in written language are taught explicitly, it seems plausible that children’s literacy acquisition is aided by implicit learning through exposure to written language. Previous research has suggested a link between implicit learning and literacy skills in the typically developing population (e.g. Apfelbaum, Hazeltine, & McMurray, 2013; Arciuli & Cupples, 2006; Arciuli & Simpson, 2012; Frost et al., 2013; Pacton, Fayol, & Perruchet, 2005; Spencer, Kaschak, Jones, & Lonigan, 2014). For example, typically developing children apply orthographic regularities in pseudo-word spelling (e.g. in French, /εt/ is more often written as <ette> after −v than after −f), which reflects their implicit knowledge of single letters and letter combinations (Pacton et al., 2005). Similarly, Pacton, Perruchet, Fayol, and Cleeremans (2001) show that French-speaking typically developing children are sensitive to the orthographic constraints of the positions of double consonants (e.g. xevvu is more acceptable than xxevu). Additionally, correlational studies have established a link between performance on implicit learning tasks and reading in English (Arciuli & Simpson, 2012), reading in Hebrew as a second language (Frost et al., 2013), and a variety of literacy-related skills including oral language, vocabulary and phonological processing (Spencer et al., 2014). Using a linear regression analysis, Ise, Arnoldi, Bartling, and Schulte-Körne (2012) showed that children’s performance on a visual artificial grammar learning (AGL) task, a measure of implicit learning which will be explained in more detail below, predicts their performance on a spelling task. Together, the abovementioned studies suggest there is a relationship between implicit learning and (the acquisition of) literacy skills in typical populations. Implicit learning in dyslexia ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A number of studies have investigated the hypothesis that individuals with dyslexia have problems with implicit learning, which affect their literacy skills. Several tasks have been deployed to investigate implicit learning skills in dyslexia. Examples include the serial reaction time (SRT) task (e.g. Deroost et al., 2010; Menghini et al., 2010; Vicari et al., 2005), the alternating SRT task (Hedenius et al., 2013), as well as visual AGL tasks (e.g. Ise et al., 2012; Pothos & Kirk, 2004; Rüsseler, Gerth, & Münte, 2006). Although both the SRT and AGL paradigm are methods used to investigate implicit learning, the type of structure learned in each paradigm differs (greatly). Whereas the SRT measures a motoric response to visual sequences and is stimulus-bound (i.e. no generalization rule can be abstracted from the sequence), the visual AGL measures rule learning from visual input. While numerous studies report implicit learning difficulties in individuals with dyslexia (e.g. Du & Kelly, 2013; Ise et al., 2012; Jiménez-Fernández, Vaquero, Jiménez, & Defior, 2011; Vicari et al., 2005), others do not find evidence for such a deficit (e.g. Deroost et al., 2010; Menghini et al., 2010; Pothos & Kirk, 2004; Rüsseler et al., 2006). Because of these mixed results, Lum, Ullman, and Conti-Ramsden (2013) performed a meta- analysis on 14 studies that investigated implicit learning in individuals with dyslexia using the SRT paradigm. Their results show that implicit sequence learning, as measured by the SRT task, is significantly poorer in people with dyslexia than in non-dyslexic controls (average weighted effect size 0.45, p < 0.001). Thus, these results indicate a deficit in implicit visuo-motor learning in dyslexia. In the current study we investigate whether individuals with developmental dyslexia are also affected in visual artificial grammar learning. If individuals with dyslexia have difficulties with implicit learning across the board, group differences should be found using both the SRT and AGL paradigms. However, it could also be the case that poor performance by individuals with dyslexia on the SRT task is due to a specific motor learning deficiency, as dyslexia has previously been associated with motor problems (e.g. Fawcett & Nicolson, 1995; Ramus, 2003; Ramus et al., 2003). In that case, one would not necessarily also expect difficulties in the area of visual AGL learning. Visual AGL in dyslexia ~~~~~~~~~~~~~~~~~~~~~~ Visual AGL refers to an experimental design that investigates participants’ ability to implicitly learn rules from mere exposure to sequences of visual stimuli generated by these rules. First introduced by Reber (1967), the visual AGL paradigm involves structured sequences that can be presented as letters or abstract shapes. In visual AGL tasks, sequences are generated on the basis of a (finite state) grammar that determines which stimuli can and cannot succeed one another (Fig. 1). In the example depicted in Fig. 1, from the node S2, the sequence can proceed either to S4 (a triangle) or S5 (a diamond), but not back to S1. The AGL task typically consists of two phases: a training and a test phase. In the training phase, participants are exposed to a set of structured sequences. Importantly, in the implicit version of the AGL task that is explored in the current meta- analysis, participants are not informed about the presence of the structural rules in the input. The exposure during the training phase can be either passive (i.e. participants are merely exposed to stimuli) or active (i.e. participants are instructed to memorize strings of stimuli and repeat them afterwards). At the beginning of the test phase, participants are often informed that certain rules guided the presentation of stimuli during the training phase. Subsequently, they are tested on their ability to distinguish sequences that adhere to the artificial grammar (grammatical strings) from sequences that do not (ungrammatical strings). Typically, recognition of grammatical strings is tested within a grammaticality judgment task in which participants are requested to specify whether single strings are grammatical or ungrammatical. Other studies adopt a two-alternative forced choice paradigm, where participants are presented with two strings, one grammatical and one ungrammatical, and have to indicate which of the two strings belongs to the grammar. Performance above chance level (50%) during the test phase is taken as evidence that participants have learned the rules of the underlying grammar. Several studies using the visual AGL paradigm have reported learning deficits in dyslexia among adults (Kahta & Schiff, 2016; Laasonen et al., 2014) or children (Ise et al., 2012; Pavlidou, Williams, & Kelly, 2009; Pavlidou & Williams, 2014). In each of these studies, this deficit is reflected by significantly lower accuracy scores in the group of individuals with dyslexia as compared to a control group. Several other studies failed to show a significant effect of dyslexia in children (Nigro, Jiménez-Fernández, Simpson, & Defior, 2016) or adults (Pothos & Kirk, 2004; Rüsseler et al., 2006). These differences in degrees of significance might be due to chance (i.e. sampling error), because no direct statistical comparisons were ever made between the studies. However, differences in group effects might also reflect genuine differences between the studies. Here we will speculate on several factors that may help explain such genuine differences between individual studies. Firstly, the age of the participants may influence the results of individual studies, as several studies have reported that implicit learning improved with age in typical populations (e.g. Arciuli & Simpson, 2011; Maybery, Taylor, & O'Brien-Malone, 1995, but see Jost, Conway, Purdy, & Hendricks, 2011). In a meta-analysis investigating SRT performance, Lum et al. (2013) found smaller differences between participants with and without dyslexia for studies involving adult as opposed to child participants when certain sequences of stimuli were used (second-order sequences) or when the exposure phase was longer. However, no previous studies have examined the developmental trajectory of AGL in individuals with dyslexia. Secondly, the use of either linguistic or non-linguistic stimuli may influence the difficulty of the task, especially for participants with dyslexia. Linguistic stimuli include visually presented letters (e.g. Ise et al., 2012; Nigro et al., 2016; Rüsseler et al., 2006), whereas non-linguistic experiments have used abstract shapes (e.g. Laasonen et al., 2014; Nigro et al., 2016; Pavlidou et al., 2009; Pothos & Kirk, 2004). The results are mixed: several studies report impaired learning within a AGL task involving linguistic stimuli (i.e., letters, e.g. Ise et al., 2012; Samara, 2013), while others do not find evidence for an effect of dyslexia (e.g. Nigro et al., 2016; Rüsseler et al., 2006). Similarly, studies have yielded mixed results in AGL tasks with non-linguistic stimuli such as abstract shapes (evidence for learning deficits: Laasonen et al., 2014; Pavlidou, Kelly, & Williams, 2010; Pavlidou & Williams, 2014, no evidence for learning deficits: Nigro et al., 2016; Pothos & Kirk, 2004). Thirdly, the training method potentially affects participants’ performance. As mentioned, the training phase generally includes one of two possible methods: passive exposure (Du, 2013; Laasonen et al., 2014; Nigro et al., 2016) or active memorization (e.g. Ise et al., 2012; Rüsseler et al., 2006; Samara, 2013). Active training may lead to better learning, as participants are more focused on the stimuli. Whether the observed differences in results between the studies are genuine or due to chance is one of the questions that the present paper tries to address. Thus, mixed results exist for the visual AGL paradigm: whereas several studies report significant differences between participants with and without dyslexia (e.g. Ise et al., 2012; Laasonen et al., 2014), others do not (Nigro et al., 2016; Rüsseler et al., 2006). Schmalz, Altoè, and Mulatti (2016) conducted a meta-analysis on a subset of studies investigating visual AGL in dyslexia. They report significantly poorer performance by participants with dyslexia (average weighted effect size 0.47). However, at the same time they are careful in their interpretation and state “[…] publication bias and questionable research practices result in an inflated effect size” (p. 9). As no meta-regression analysis was performed, the authors could not quantitatively explain the differences in effect size between studies. The primary aim of the present meta-analysis is to extend the findings by Schmalz et al. (2016) to a larger set of (unpublished) studies and determine whether the accumulated evidence indicates a difference in performance on visual AGL between individuals with and without dyslexia. By doing a systematic literature search and by including a number of unpublished studies, we want to provide a more complete update on the strength of the evidence regarding the association between dyslexia and a deficiency in visual artificial grammar learning. Additionally, we aim to investigate the effect of certain methodological variables through a meta-regression analysis. These variables include the age of participants and the nature and complexity of the task used, which potentially help explain heterogeneity in results of individual studies. Factors included in the analysis are (a) age (adult or child participants), (b) stimulus type (letters or abstract shapes), and (c) type of training method (passive exposure or active memorization). Literature search ~~~~~~~~~~~~~~~~~ We identified studies published up until September 2016 through searches in PubMed, PsycInfo, ERIC, MEDLINE, CINAHL, and LLBA databases. Additionally, the OATD database was searched for unpublished work in the form of theses and dissertations. A complete overview of keywords used for each of the databases can be found in Appendix A in Supplementary data. In addition to database searches, references of included articles were reviewed. Finally, the CogDevSoc and LinguistList mailing lists were used to inquire whether subscribers knew of unpublished data (deadline response: September 2016). Study selection ~~~~~~~~~~~~~~~ Fig. 2 depicts the selection of studies according to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines in a flow diagram (Moher, Liberati, Tetzlaff, & Altman, 2009). Out of all 229 records found, 143 duplicates were removed. Subsequently, one researcher examined the abstracts of 86 unique studies. Studies had to fulfill several selection criteria for inclusion in the present meta-analysis. First, only studies that had administered a visual AGL task were considered. The main reason for a focus on visual AGL studies is to eliminate modality as a possible cause of heterogeneity in results. Second, the experiment had to address implicit learning, i.e., participants were not to be informed of the presence of rules in the input. Third, studies had to include two groups of participants, one group of individuals with dyslexia and one group of non-dyslexic controls. Fifty-six records were removed after screening the title and abstract because they did not meet the abovementioned selection criteria. An additional 19 records were removed from the sample on the basis of full-article screening, thus leaving eleven records for inclusion in the present review and meta-analysis. Two out of the eleven records (Ise et al., 2012; Nigro et al., 2016) involved two experiments with distinct participant groups that were included separately in the present meta-analysis, resulting in 13 individual effect size calculations. For the remainder of the present meta-analysis, we will refer to the number of individual effect size calculations as the number of studies included (N = 13). A second researcher performed identical database searches and assessed all abstracts and full texts. For 28 out of 30 full-text studies the reviewers independently came to the same conclusions regarding inclusion in the present meta-analysis (high inter-rater reliability: Cohen’s kappa = 0.851). Consensus on the remaining two records was reached through discussion of the contents. Note that articles did not have to have been published in peer-reviewed journals in order to be included in our meta-analysis. This means that conference papers or posters, unpublished results and dissertations could be included in the final sample (under the category “other” in Fig. 2). This was done to minimize the possibility of a publication bias. 10 out of 12 records in this category were found through the OATD database (Open Access Theses and Dissertations), of which 2 are included in the final sample (Du, 2013; Samara, 2013). The other two were discovered through personal communication with authors or were presented at the Interdisciplinary Advances in Statistical Learning conference (2015, San Sebastian2). At the time of analysis, two out of thirteen individual effect sizes included in the present meta-analysis were unpublished. Data extraction and effect size calculations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The standard measure of learning in an AGL task is the percentage of correct responses (i.e. overall accuracy) during the test phase of the experiment. Therefore, the method for comparing the performance of two groups on an AGL task is to test whether the overall accuracy differs between the study and the control group. In order to calculate a single effect size for each of the included individual studies, the mean, standard deviation (SD) and sample size of each of the study groups were extracted from the article. If these data were not available from studies themselves, we asked the authors to supply these. Authors provided these data in three cases (Ise et al., 2012; Laasonen et al., 2014; Samara, 2013), which allowed us to calculate single effect sizes for each individual study. For the study by Pothos and Kirk (2004), the SDs had to be gleaned from their Fig. 4 (p. 71). Additionally, the mean age of participants was not available, but since participants were (under)graduates, most aged between 18 and 30, this study was classified as a study involving adult participants. Appendix B in Supplementary data presents an overview of the extracted data that was used for effect size calculation for each included study. Tables 1 and 2 summarize characteristics of the sample and experimental design of the 13 studies included in the present meta-analysis. Following data extraction procedures, a single effect size was computed for each individual study, using the “compute.es” package (Del Re, 2014) for R software (R Development Core Team, 2008). In the present meta-analysis, Hedges’ g effect size4 and 95% confidence intervals summarize the results from each individual study. Positive Hedges’ g values indicate that the control group reached higher accuracy levels on the AGL task compared to the group of individuals with dyslexia, whereas negative values indicate the opposite. The 95% confidence interval provides an estimate of the precision of the study’s effect size: the larger the confidence interval, the poorer the precision. A combination of the “metafor” (Viechtbauer, 2010) and “meta” packages (Schwarzer, 2012) for R software was used to convert the computed individual effect sizes and variances to an average weighted effect size5 and variance across studies. AGL in dyslexia ~~~~~~~~~~~~~~~ Our first goal was to elucidate whether, combining the results from 13 previous studies, individuals with dyslexia perform different from their TD peers on visual AGL tasks. To this end, the effect sizes of all 13 individual studies were combined into a single average weighted effect size using a random-effects model (Hedges & Olkin, 1985). Random- effects models, as opposed to fixed-effect models, allow for variation in true effect sizes between independent studies (Borenstein, Higgins, & Rothstein, 2009). The model was run using the rma.uni function in the “metafor” package with the restricted maximum likelihood (REML) method and the adjustment by Knapp and Hartung (2003) for finite numbers of degrees of freedom. Effect sizes for individual studies and the overall average weighted effect size are presented in Fig. 3. Performance was measured as the overall accuracy score in the test phase of the AGL experiment. Effect sizes ranged from −0.68 to 1.37, with only one effect size in the negative direction (Pothos & Kirk, 2004). All other studies report a lower accuracy level for the group of participants with dyslexia than for the control group. Importantly, as mentioned, some of the individual studies report significant differences, whereas others do not. The meta-analysis reveals that, grouping over 13 studies and despite the negative-estimate study, participants with dyslexia performed significantly worse than control participants (average weighted effect size = 0.46, 95% CI [0.14 … 0.77], p = 0.008). Looking at studies involving either child or adult participants separately, we find that the average weighted effect size for child studies is significant (N = 7, average weighted effect size = 0.71, 95% CI [0.36 … 1.07], p < 0.001), whereas it is not for adult studies (N = 6, average weighted effect size = 0.16, 95% CI [−0.36 … 0.69], p = 0.461). Before investigating whether the observed difference between the adult effect (0.16) and the child effect (0.71) reflects a genuine decreasing difference between dyslexic and non-dyslexic people as a function of age, we inspect the possibility of publication bias. Publication bias ~~~~~~~~~~~~~~~~ To verify the interpretability of the abovementioned findings, we examined the possibility of publication bias in our collected sample of studies. This was initially done through examining a standard funnel plot, which plots the standard error (a measure of study precision) against the effect sizes of the individual studies (Fig. 4a). Generally speaking, in the absence of publication bias, studies should be symmetrically distributed around the average weighted effect size. This distribution takes a funnel shape configuration: studies with high precision are closer to the average weighted effect size, whereas lower precision studies are symmetrically scattered around the average weighted effect size. A linear regression analysis (Egger et al., 1997), using the metabias function (Schwarzer, 2012), formally tested the presence of publication bias. It turned out that effect sizes were significantly asymmetrically distributed, skewing to the lower right corner, indicating the presence of a publication bias in our sample (t[11] = 4.014, p = 0.002). To evaluate the effect of the publication bias in our sample we approximated what the effect size might be in absence of this bias, using Duval and Tweedie (2000) trim and fill method (trimfill function in the “metafor” package using the “L0” estimator for the number of missing studies). Importantly, the trim and fill method can be used to investigate how sensitive the observed effect is to the presence of potential missing studies, but is not meant as a way to calculate the actual values of missing studies (Duval & Tweedie, 2000; Duval, 2005). By using small studies on the positive side of the funnel plot to impute missing studies on the negative side, the trim and fill method estimated that five studies reporting negative findings are missing in our present sample (Fig. 4b). When these five imputed missing studies are added to our dataset of 13 studies, the estimated effect size is considerably reduced and is no longer significantly different from zero (average weighted effect size = 0.20, 95% CI [−0.11 … 0.50], p = 0.205). Note, however, that the trim and fill analysis is known to be a sometimes conservative method for adjusting for publication bias (Peters et al., 2007; Schwarzer, Carpenter, & Rücker, 2010) and the creation of imputed studies can be heavily influenced by a single deviant study, such as the study by Pothos and Kirk (2004) in our sample (e.g. Borenstein et al., 2009, p. 286).6 Additionally, this method of adjusting results for publication bias makes the assumption that the asymmetry observed in the funnel plot is caused exclusively by publication bias, while another possible cause for funnel plot asymmetry is heterogeneity between studies (Mavridis & Salanti, 2014). Finally, we cannot be certain that the computed missing studies would indeed have been found in the absence of such a bias (Mavridis & Salanti, 2014). Nonetheless, the results of the present meta-analysis on our selected 13 studies are likely to be overly optimistic in the direction of the existence of the main effect, as the effect can well be nulled by unpublished findings. Heterogeneity in findings and meta-regression ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The second aim of the present study was to explore several factors that may help account for the heterogeneity (between-studies variability) that appears to be present across different studies investigating AGL in participants with dyslexia. Although the main outcome of the present meta-analysis is probably influenced by the observed publication bias, such a bias is less likely to affect meta-regression analyses, which consider secondary effects. Cochran’s Q-test for heterogeneity was significant (Q[12] = 41.07, p < 0.001, I2 = 71% [0.49 … 0.83]). This result allows us to reject the null hypothesis that all the studies share a common true effect size. As can be seen in Fig. 3, it appears that some factors may influence the effect size of individual studies. As mentioned, the average weighted effect size for child studies is larger than for adult studies (0.71 for child studies vs. 0.16 for adult studies). Thus, we decided to explore the effect of several potential moderator variables on the effect sizes of individual studies through meta-regression analysis. In preparation for the meta-regression analysis, all three binary moderator variables were centered, i.e. coded as: (a) age −0.5 (child) versus +0.5 (adult), (b) type of stimulus −0.5 (abstract shapes) versus +0.5 (letters), and (c) type of training −0.5 (passive exposure) versus +0.5 (active memorization). Random-effects model meta-regression was used to explore the potential value of these factors in explaining variance in effect size between studies. Since the three moderators are correlated, we first tested each of the three main effects individually in a separate meta-regression model (Table 3). Additionally, we tested each of the three interaction effects individually, in a separate model that included the two relevant main effects (also in Table 3). None of the interaction effects turned out to significantly affect the effect sizes of individual studies, so we did not attempt to construct any more complicated models. As shown in Table 3, the only model that reaches significance in explaining variance between individual studies is model 1: the main effect of age. This model fits 35% of the heterogeneity (which is greater than 0% with p = 0.041). When studies had adult participants, as opposed to child participants, effect sizes were smaller, reflecting a smaller difference between participants with and without dyslexia. None of the other main or interaction effects were found to significantly fit the heterogeneity between studies. To the extent to which a p-value of 0.041 can be considered statistically significant in this exploration of six possible effects (without correction for multiple testing), we can conclude that the difference between the observed adult and child effect sizes (0.16 and 0.71) indeed reflects a genuine difference between the two ages in the population.","In the present study, we used meta-analysis and meta-regression to quantitatively review previous research on visual AGL in dyslexia. Our first goal was to elucidate whether the combined findings of thirteen previous studies provide evidence for a difference in visual AGL between individuals with and without dyslexia. The average weighted effect size computed from these individual visual AGL studies, reflecting results from 255 participants with dyslexia and 292 control participants, was found to be moderate and statistically significant. If our 13 selected studies were a sample randomly drawn from an imagined infinite set of possible studies, this finding would indicate that, overall, non- dyslexic people outperform people with dyslexia on visual AGL. Our results would then corroborate the earlier analysis in Schmalz et al. (2016) and strengthen these findings by involving a larger sample of studies (13 instead of 9). Taken together with the meta- analysis of SRT studies by Lum et al. (2013), these results would suggest a general implicit learning deficit in individuals with dyslexia. Importantly, however, it seems plausible that these results have been influenced by a publication bias in the field of artificial grammar learning in dyslexia (see Schmalz et al., 2016). After conservatively controlling for publication bias, the computed effect size was no longer significant, and the results of the main effect of the present meta-analysis should therefore be regarded as unreliable. Large-scale future studies are needed to confirm the presence of a difference in performance on visual AGL between participants with and without dyslexia. Extending the previously published meta-analysis by Schmalz et al. (2016) further, the present study aimed to explain the heterogeneity in results of individual studies by investigating the effect of certain methodological variables through a meta-regression analysis. This analysis revealed that the only moderator that (moderately, i.e. without correction for multiple tests) reached significance was the main effect of age: there were smaller differences between dyslexia and control groups for those studies that involved adult participants as opposed to child participants. This is an indication that the implicit learning deficit might be more pronounced in children with dyslexia than in adults with dyslexia, since similar effects of age have been found in the meta-analysis investigating implicit learning in individuals with dyslexia using the SRT task (Lum et al., 2013). In line with the interpretation of their results, a possible explanation is that adults make use of compensatory processes (e.g. visual processing, pattern recognition, attentional resources, declarative memory) that enhance performance on visual AGL tasks. Another potential explanation for the age effect lies in the selection of participants. Whereas most studies with adult participants involved university students, child studies selected their participants from a broader population of primary school children. The performance of university students with dyslexia may not be representative of the whole population of adults with dyslexia, as these high-achieving individuals with dyslexia may have more developed compensatory mechanisms. This in turn may result in a smaller difference between the performance of adults with and without dyslexia. We want to note that this effect of age should be interpreted with caution, as it seems to be largely driven by one study that reports better performance in adults with dyslexia than in controls (Pothos & Kirk, 2004, g = −0.68). Thus, future research should examine the possibility of an age effect in visual AGL in dyslexia in further detail by selecting adult participants with dyslexia from all educational levels and comparing them to children on the same visual AGL task. Although the present meta-analysis suggests that visual artificial grammar learning might be poorer in dyslexia relative to non-dyslexic individuals overall, these results cannot address the issue of causality between implicit statistical learning and literacy skills in this population. Future longitudinal studies are needed to investigate the potential causal link between implicit statistical learning and literacy skills in individuals with and without dyslexia. Additionally, several factors that could influence the effect sizes of individual studies were not included in the present meta-analysis due to the relatively small number of studies. One such factor is the complexity of the underlying grammar. The level of complexity potentially plays a role in whether participants are able to learn the underlying structure. In fact, a recent meta-analysis of AGL studies with typical populations showed that, indeed, there is a significant correlation between grammar complexity and learners’ task performance (Schiff & Katan, 2014). Also related to the difficulty of the task at hand are factors such as the length of the sequences and the amount of exposure to these sequences. Whereas some studies use a fixed sequence length of 4 (Nigro et al., 2016), 5 (Ise et al., 2012), or 7 (Samara, 2013), other studies use sequences of varying lengths (between 2 and 6 (e.g. Pothos & Kirk, 2004; Pavlidou & Williams, 2014), 4 and 7 (Rüsseler et al., 2006) or 6 and 8 (Du, 2013) individual items). Similarly, whereas some studies include only 69 instances of a grammatical string (e.g. Laasonen et al., 2014; Pavlidou et al., 2009), others include as many as 108 instances (Nigro et al., 2016). Another factor worth investigating is the severity of dyslexia in individual participants, as this may be related to the severity of the deficit in implicit statistical learning. Finally, the modality (visual versus auditory) in which the stimuli are presented may affect the learnability of the grammar for individuals with dyslexia. Future research should investigate the potential effect of the abovementioned factors to gain further understanding of what types of methodological characteristics increase or decrease an AGL task’s learnability for individuals with and without dyslexia."],["Background: Previous research suggests that children with cerebral palsy (CP) have impairments in visual-spatial and mathematics abilities, although we know very little about the association between these two domains. Aims: To investigate the extent of visual-spatial and mathematical impairments in children with CP and the associations between these two domains. Method and Procedure: Thirty-two children with predominantly quadriplegic spastic and/or athetoid (dyskinetic) CP (13 years 7 months) and a group of typically developing (TD) children (8 years 6 months) matched by receptive vocabulary were given a battery of visual-spatial and mathematics tasks. Visual-spatial assessments ranged from simple tests of perception to complex reasoning about these stimuli. A standardised test of mathematics ability was administered to both groups. Outcomes and Results: The children with CP had significantly poorer mathematical and visual-spatial abilities than the TD group. For the TD group age was the best predictor of mathematical ability, in the CP group receptive vocabulary and visual perception abilities were the best predictors of mathematical ability. Conclusion and Implications: The CP group had extensive difficulties with visual perception; visual short-term memory; visual reasoning; and mental rotation all of which were associated with their mathematical abilities. These findings have implications for the teaching of visual perception and visual memory skills in young children with CP as these may help the development of mathematical abilities. --------------------------------------------------------------------------------","Results supported and extended previous research. Children with CP, in comparison with a TD group, were found to have impairments to visual-spatial perception including processing information in visual short-term memory, in mental rotation tasks and in matrices reasoning tasks. The CP group also had significantly poorer mathematical abilities than the TD group. We provide additional information about the presence of significant relationships between visual-spatial perceptual abilities and the mathematical abilities of both groups of children. We demonstrated that the CP group showed extensive, significant relationships between visual-spatial tests and their mathematical abilities. These findings suggest the potential importance of teaching and developing visual-spatial abilities in children with CP, particularly those abilities that might have specific roles in the development of mathematics. Further research in the form of interventions, such as those described in a study of the impact of mental rotation training on arithmetic performance (Cheng & Mix, 2014), may greatly benefit children with CP and their progress in mathematical learning.","Cerebral palsy (CP) is typically caused by brain lesion(s) that are usually diagnosed within the first two years of life. It is most often caused by brain trauma in the uterus; at birth; or by other causes in early infancy (Rosenbaum, Nigel, Leviton, Goldstein & Bax, 2007). The condition affects the contraction, tone and function of muscles, and hence motor development, with a prevalence of approximately 2–3 children in every thousand (Odding, Roebroeck, & Stam, 2006). Research into CP has identified that, along with the characteristic motor impairments, there often are cognitive impairments (Odding et al., 2006; Stadskleiv, Jahnsen, Andersen & von Tetzchner, 2017). Related to this, many children with CP have difficulties with developing mathematical concepts (Jenks, val Lieshout & de Moor, 2012; van Rooijen et al., 2012). Previous research into the mathematical abilities in CP have identified associations with mathematics and visual-spatial impairment such as visual short term memory (Jenks et al., 2009; Peeters, Verhoeven, & de Moor, 2009); non- verbal matrices tasks (van Rooijen et al., 2012; Jenks et al., 2007); and mental rotation abilities (Courbois, Coello, & Bouchart, 2004). Thus, our investigation builds on previous research by seeking to determine the precise profile of visual-spatial deficits in a group of children with CP, and by investigating whether the reported deficits in visual-spatial cognition and mathematics are associated. Typical mathematical development and visual-spatial abilities ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Mathematical skills develop on a ‘learning trajectory’ (Purpura & Ganley, 2014), that is, understanding basic knowledge and skills is necessary before more complex mathematical comprehension is achieved. There are a number of cognitive processes which contribute to mathematical development. Research concerning typically developing (TD) children, for example, has demonstrated that visual-spatial abilities are a strong predictor of mathematical competence (see Uttal et al., 2013; Verdine, Irwin, Golinkoff & Hirsh-Pasek, 2014). In addition, visual short-term memory (STM) as assessed by Corsi block recall has been linked to mathematical abilities (Fuchs et al., 2005; Kyttälä & Lehto, 2008). Furthermore, non-verbal intelligence is usually assessed by tasks that involve visual- spatial abilities (e.g. Raven’s Progressive Matrices RPM, Raven, Raven, & Court, 2008) and has also been found to be related to mathematical ability (Deary, Strand, Smith, & Fernandes, 2007; Kyttälä & Lehto, 2008). These findings provide a rationale to investigate the relationship between visual-spatial perception and mathematical abilities in children with CP. Mathematical abilities, visual-spatial abilities, and cerebral palsy ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Researchers comment on the paucity of investigations into numerical and arithmetic abilities of children with CP particularly those involving aspects of visual-spatial perception (Schmetz, Magis, Detraux, Barisnikov, & Rousselle, 2018; van Rooijen et al., 2012). Available research has shown that children with CP are more likely to experience difficulties with mathematical learning than TD children (Frampton, Yude, & Goodman, 1998), and a small number of studies have highlighted that children with CP are often delayed in acquiring numerical skills such as: subitizing (quantifying sets); counting; and basic arithmetic problem-solving (Jenks, van Lieshout, & de Moor, 2012; van Rooijen et al., 2012). In a longitudinal study over two years, van Rooijen et al. (2014) found that those children with CP and higher non-verbal intelligence increased their mathematical skills on a trajectory similar to a comparison group of TD children; however, they did not catch up with their TD peers. Many children with CP have visual-spatial perceptual impairments despite typical or near-typical visual acuity (Akhutina, Foreman, Krichevets, & Vahakuopus, 2003; Ego et al., 2015; Menken, Cermak, & Fisher, 1987; Reed & Drake, 1990; Ortibus, Lagae, Casteels, Demaerel, & Stiers, 2009; Stadskleiv, Jahnsen, Andersen, & von Tetzchner, 2017). Ego et al. (2015), in a systematic review of visual-perceptual impairments in children with CP, identified 15 studies which included one of five standardised tests that assessed visual perception. The review indicated that the proportion of children with visual-perception impairments ranged from 40% to 50%, with no study reporting a significant effect of CP subtype or IQ level, although the severity of neurological lesions appeared to be associated with visual perception abilities. There has been discussion of whether visual-spatial impairments could directly be caused by neurological lesions and their sequalae (Menken et al., 1987), and it is also possible that an inability to explore space because of physical disabilities may limit the development of spatial functioning and visual-spatial skills in these children (see Clearfield, 2004; Stanton, Wilson, & Foreman, 2002). Because of the fine-motor difficulties of many children with CP (Ballester-Plané et al., 2016; Rosenbaum et al., 2007) we decided to focus on visual-spatial perception assessments which do not involve manipulation; the one that was chosen was the Test of Visual Perception ([TVPS-3]; Martin, 2006). We also decided to use three other assessments which not only involve visual- spatial perception, but also additional cognitive processing of visual-spatial information. These were a short term visual-spatial memory task, a mental rotation task, and a matrices task. It has been reported that children with CP have impaired performance on these tasks. In the case of visual-spatial memory, Gagliardi, Tavano, Turconi, and Borgatti (2013) used a computerised version of the Corsi blocks task where children with CP were asked to remember the sequence in which a set of blocks were touched by the experimenter. The researchers found that the children with CP performed significantly worse than TD children of the same chronological age. Different, age-related, errors were observed across the CP group. For example, position errors, e.g. selecting blocks from the wrong place, were more common in the younger children of the CP group, indicating visual spatial memory difficulties, whilst the older children in the CP group were more likely to have problems with sequencing skills (also see Schmetz et al., 2018). Corsi type tasks are of relevance to mathematical ability as the visual spatial sketchpad may be used as a ‘mental blackboard’ (Baddeley, 2007) when distinguishing between different mathematical forms and shape, as well as retaining previous visual information (see Hubber, Gilmore, & Cragg, 2014). To date, there have been few studies on object-based mental rotation abilities in children with CP. In these tasks participants identify whether rotated and upright shapes have the same identity (see Fig. 1) (Schmetz et al., 2018). We included this task for three reasons. Firstly, it measures a relatively pure form of visual-spatial processing ability (Shepard & Metzler, 1971). Booth et al. (1999) have suggested that during this process, mentally rotated stimuli are temporally stored in working memory, and the visual-spatial sketchpad is implicated in the manipulation of the visual images (Gathercole, Pickering, Ambridge & Wearing, (2004); Hyun & Luck, 2007). Thus, reduced working memory capacity (i.e. limited cognitive resources) could be responsible for poor task performances on these rotation tasks. Secondly, children with CP have been reported to have impaired performance on this task (Courbois et al., 2004). Thirdly, mental rotation has consistently been found to be strongly associated with mathematical abilities in TD children and interventions involving mental rotation can help these children’s mathematical performance on some types of problems (e.g. Cheng & Mix, 2014). Matrices tasks are another assessment which have an important visual-spatial component in which children with CP often have significantly low scores. In these tasks, several visual arrangements of a pattern are presented, and the participant chooses an item to complete a sequence which has a different but related visual arrangement. Children with CP have lower scores on Raven’s Progressive Matrices (RPM or Raven’s Progressive Coloured Matrices [RCPM], Raven et al., 2008) than chronologically age matched TD children (Jenks et al., 2012; Peeters, Verhoeven, Van Balkom, & De Moor, 2008). This is usually interpreted as indicating that the children have lower non-verbal intelligence as the task involves non- verbal reasoning, but matrices tasks have an important visual-spatial component and notably Ego et al. (2015) in their systematic review included the RPM as a visual- perceptual assessment. Consequently, we were interested in investigating whether matrices tasks had similar relationships to mathematical abilities as other visual-spatial tasks. As far as we can ascertain there are few studies that have investigated the relationship between matrices and mathematical abilities in children with CP (Jenks et al., 2009; Peeters et al., 2009). van Rooijen et al. (2012) reported that CP children’s RCPM performance was significantly correlated (r = 0.61) with their arithmetic abilities at ages 6–8 years. In other studies, RCPM assessments have had a strong association with basic mathematical abilities in children with CP who attend mainstream schools as well as those who attend special schools (Jenks et al., 2007). Furthermore, this assessment has been found to correlate significantly with addition and subtraction tests in CP children aged between 6 and 11 years (van Rooijen, Verhoeven, & Steenbergen, 2011). Given these previous findings, in the current study we investigated the extent to which visual-spatial cognition and mathematical ability are impaired in children with CP, and the extent to which these visual abilities predict mathematical ability. Thus, our investigation builds on previous research by determining the precise profile of visual-spatial deficits in a group of children with CP, and by investigating whether the reported deficits in visual- spatial cognition predicts mathematical ability. Overview ~~~~~~~~ Although visual-spatial abilities have been identified as relating to mathematical abilities in the typical population, there have been few studies which have focused on the role of a range of visual-spatial assessments and their relationship to the mathematical abilities of children with CP. Accordingly, it was decided to compare children with CP in relation to TD children who had similar receptive vocabulary. Our first question concerned whether there was a difference in the mathematical and visual-spatial abilities of the two groups and the size of any difference. We were also interested in whether visual-spatial impairments occurred when the assessments did not involve an important motor component (i.e. the primary impairment in CP) and whether tasks that involved the processing of additional visual-spatial information (e.g. memory, mental rotation, non-verbal reasoning) made the tasks more difficult for children with CP due to cognitive load, for example, switching attention (Henry, 2011). The second question concerned whether visual-spatial abilities are related to mathematical abilities in CP, and how this compares to children with TD. Because mathematical abilities are affected by age and language (Butterworth, 2005; Fuchs & Fuchs, 2002), we also examined whether visual-spatial abilities still predicted mathematics scores when the effects of these potential confounds were removed. Design ~~~~~~ This is a cross-sectional study in which we are interested in group comparisons on individual tasks between CP and TD children, and the relationships between visual-spatial cognitive abilities and mathematical abilities for both groups. We used IBM: SPSS Statistics software (version 21) and stepwise hierarchical regression analyses were conducted to further examine the data. Given that individuals with CP are able to access assessments involving receptive language (Bishop, Byers Brown, & Robson, 1990; Dahlgren Sandberg, 2001), we decided to match a sample of children with CP to a TD sample using a measure of receptive vocabulary. Given the broad battery of experimental measures that we employed, which spanned across two domains, the choice of a matching measure was not straight forward. Thus, we chose receptive language as a proxy for general IQ. Whilst this does dictate that performance in the domains of interest is likely to be poor relative to the TD group, the use of a battery of tasks enables us to ask questions regarding profiles of performance, i.e. which aspects of visual-spatial cognition are more vs. less impaired, and to detail how cross domain associations compare to those observed in the typical population.","Thirty-two children with a diagnosis of CP (18 males), aged between 7 years and 18 years; 2 months, were selected from two special schools for children with physical disabilities in the UK. The children were selected from special schools as they were an opportunistic sample and the first author was familiar to the children in both schools. Medical information was not available to the researchers, but observations made by experienced teaching staff and the first author indicated that all the children, bar one, had predominantly either quadriplegic spastic or athetoid CP or a mixture of both which had variable effects on their functionality. The remaining child had hypotonic muscle tone and poor posture, possibly ataxic CP (see Rosenbaum et al., 2006). Twenty-six of the children used a wheelchair to move around school, and seventeen of these children used a walking frame for part of the time. Six children moved around school with no aids. Teacher rating of the children’s mobility using the Movement Assessment Battery for Children Checklist (Henderson, Sugden, & Barnett, 2007) was undertaken and the results (higher scores up to a maximum of 90 indicated poorer mobility) were as follows: Mean: 56.94; SD: 16.32; Minimum: 15; Maximum: 87. Most of the children performed typically in the section of the checklist describing fine-motor abilities, e.g. manipulating small objects such as picking up beads and moving small cubes except for one child who used her big toe for pointing and keyboarding, and two others who found moving sheets of paper difficult. However, all the children were able to point. The children were selected using the following criteria by the teaching staff: Sufficient communication and speech abilities to verbally answer questions in test situations; Typical visual acuity (with glasses if necessary); Typical hearing abilities (with a hearing aid if necessary); Sufficient dexterity to respond to the tasks. Ethical approval was obtained from the relevant University Committee prior to the study (BERA, 2011). Parental written consent and the children’s verbal consent were obtained prior to testing. Children were seen on an individual basis in quiet areas or rooms. To avoid fatigue, particularly in the CP group, sessions were kept to between 20 and 30 minutes, over 6–8 occasions. Performance of the CP group was compared to that of 32 TD children (16 males) aged between 6 years 2 months and 11 years 7 months from three mainstream schools in the UK. The two groups of children were individually matched for raw scores on the British Picture Vocabulary Scale ([BPVS-III]; Dunn, Dunn, & NFER, 2009). As can be seen in Table 1, the mean percentile score of children with CP indicates they had poor receptive vocabularies. Procedure and assessments ~~~~~~~~~~~~~~~~~~~~~~~~~ All assessments were administered and scored according to the relevant assessment manual by experienced researchers. The order of the assessments was determined by each researcher according to the length of time available for each session, and the length of time taken by each child to complete a task. Typically, the assessments for the children with CP were undertaken in the following order: BPVS-III, Mathematics Oral test, Matrices from the British Ability Scales 3rd edition ([BAS3] Elliot & Smith, 2011); Mathematics written paper, Mental Rotation Task, Working Memory Test Battery (WMTB-C) Block Recall; TVPS-3 (R). In order to reduce fatigue, a longer test such as the BPVS-3 was followed by a shorter test such as the Mathematics oral paper within one session. The children in the TD group were assessed in sessions that lasted up to an hour. Receptive language Participants completed the British Picture Vocabulary Scale (BPVS III; Dunn et al., 2009), a measure of receptive vocabulary involving the child selecting which of four pictures corresponds to a spoken word. Mathematical ability Mathematics ability was measured using the Mathematics Computation test from the Wide Range Achievement Test ([WRAT-4]; Wilkinson & Robertson, 2006). This has two sections. Part 1 is an oral test comprising fifteen items starting with simple counting, identification of numbers and basic addition and subtraction calculations. Part 2 is a written paper with calculations involving the four rules of number, including fractions and decimals. It starts with simple sums and progresses to complex calculations and has to be completed within fifteen minutes. Scores from both papers are added together to give a total raw score. Visual-spatial measures Participants completed a battery of visual-spatial measures as detailed below. The Test of Visual Perception Skills 3rd Edition (Revised), TVPS-3 (R), (Martin, 2006) This test was used to assess a range of visuospatial abilities. It consists of seven subtests which measure a different aspect of visual or spatial abilities. The test is not timed. A summary visuospatial perception score can be calculated from the results of the seven sub-tests. Visual discrimination is the ability to distinguish between different shapes, patterns, forms or colours. The test requires the child to match a shape at the top of a page with one of the five shapes displayed at the bottom of the page This is the ability to remember the exact shape, form, number of visual stimuli for a short space of time. The test consists of seeing a shape on one page for five seconds. On the next page, the child has to identify the same shape from a display of four shapes. This is the ability to understand or recognise the positioning of shapes or objects. The test involves showing five shapes or drawings on a page of which four are identical and one is slightly different in some way. This is the ability to remember a sequence of visual information such as numbers or patterns. A child has to remember a specific shape or set of shapes which is presented on one page. After five seconds the child is presented with four sets of similar shapes on the next page, only one of which is identical to the first array. This is the ability to recognise or name shapes, forms or objects by their specific attributes. A shape is displayed on the top half of a page. The bottom half displays five shapes one of which is the top shape transformed in some way, e.g. by rotating the image. This is the ability to identify a shape which is inserted or placed against a background. The test requires a child to select which of four displays contains the identical shape to the one at the top of the page. This test requires the ability to make sense of images which are not clear or finished in some way. An incomplete shape is shown at the top of a page, e.g. a rectangle with no corners, and four complete images are shown at the bottom of the page. Only one of these images can be drawn from the incomplete shape. The Block Recall Test This test from The Working Memory Test battery (WMTB-C) (Pickering & Gathercole, 2001) was included as a measure of visuospatial short-term memory (STM). This test is akin to a Corsi block test. A Mental Rotation task This test was included as a measure of the ability to process spatial information. In this task (from Broadbent, Farran, & Tolmie, 2014), participants view two mirror imaged monkeys on the top half of a computer screen (Fig. 1). The children were asked to choose which of the two monkeys on the top half of the screen matched a monkey on the bottom half of the screen. Crucially, the monkey in the bottom half of the screen was displayed rotated at 0°, 45°, 90°, 135° or 180° (6 practice trials; 32 experimental trials). The percentage of correct responses was calculated for each participant. The task was completed on a laptop, but a large-keys keyboard was available for those children who found the laptop keys difficult to access. Two children in the CP group pointed to their chosen monkey and their choices were typed for them. The Matrices subtest from the British Ability Scales 3rd edition (BAS3) (Elliot & Smith, 2011) This test was included as a visual-spatial assessment of non-verbal reasoning. Group differences in mathematics and visual-spatial cognition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The mean raw and standardised scores of the WRAT-4, TVPS-3 (R), Block Recall and BAS-3 matrices are given in Table 2. Independent t-tests revealed that, despite the CP group having a higher chronological age than the TD group, the CP group had significantly lower raw scores than the TD group on WRAT-4 Mathematics (t[62] = −3.49, p < .001, d = −0.87) and the four measures of visual-spatial cognition: the summary score for the TVPS-3 (t[62]= −7.05, p < .001, d = 1.76); WMTB-C Block Recall (t[62] = −4.86, p < .001, d = 1.21); the Mental Rotation task (t[62] = −4.5, p < .001, d = 1.16; and BAS-3 Matrices (t[62] = −5.28, p <.001, d = 1.32). All of these comparisons produced large to extremely large effect sizes (Cohen, 1992). A Bonferroni correction gives the critical significance value for the five tests as p < .003, consequently, all differences can be considered significant. Thus, these comparisons suggest that the CP group had significantly poorer mathematical and visual-spatial abilities relative to the TD group. Because there were group differences in the summary raw scores from the TVPS-3 it was decided to check whether these differences occurred across all of the sub-scales. Table 3 details the mean raw scores on the subtests. The children in the CP group scored lower than the control group and independent t-tests gave the following statistics: Visual Discrimination t[62] = −4.74, p < .001, d = 1.2; Visual Memory t[62] = −6.46, p <.001, d = −1.6; Visual Spatial Relations t[62] = −6.45, p < .001, d = −1.6; Visual Sequential Memory t[62] = 4.93, p < .001, d = 1.25; Form Constancy t[49.53] = 4.62, p < .001, d = −1.31; Figure Ground t[53.65] = −5.74, p < .001, d = 1.57 and Visual Closure t[52.83] = 3.55, p <.010, d = 0.98. These differences involved large to extremely large effect sizes. Form Constancy, Figure Ground and Visual Closure all had unequal homogeneity of variance according to Levene’s test, these were checked with Mann-Whitney U tests (U = 228.50, p <.001; U = 135.00, p <.001; U = 265.50, p <.01 respectively). Applying a Bonferroni correction set the critical significance value for the seven tests to p < .0008, as a result we need to be cautious when interpreting these differences as significant, particularly the findings about Visual Closure. Predictors of mathematical ability ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To investigate the relationships between the four assessments involving visual-spatial abilities (i.e. visual perception; visual STM; rotation; matrices task) to mathematical abilities, stepwise hierarchical regressions were conducted separately for each independent variable. The correlations between these variables are shown in Table 4. For the TD group, the correlation between age, receptive vocabulary and mathematical ability was above r = 0.80 giving rise to concern about multi-collinearity. However, for both groups the relevant statistics for tolerance and VIF were satisfactory. Mahalanobis values were acceptable except for one TD participant for mental rotations and the removal of this case did not change the significance values, all Cook’s distance measures were acceptable. Table 5 provides information about the outcome of the regression analyses. For each multiple regression, age was entered at step 1 and receptive vocabulary was entered at step 2. In this way age and language ability were controlled. At step 3 the independent variable of interest (e.g. visual perception) was entered. For the TD group, the adjusted R2 at step 3 was 0.82. At step 1, the entry of age resulted in a R2 change of 0.82 (p < .001), and at step 2 the entry of receptive vocabulary resulted in a R2 change of 0.02 (p < .070). However, the entry of the visual variables at step 3, failed to produce a significant change in R2. Examination of the standardised beta coefficients at step 3 revealed that age was the only significant predictor, although the beta coefficients for receptive vocabulary had p values between .05 and .10. For the CP group, the regression equation accounted for slightly less of the variance, the R2 at step 3 was 0.75. In contrast to the TD group, at step 1, the entry of age resulted in a non-significant R2 change of 0.04 (p < .001), and at step 2 the entry of receptive vocabulary resulted in a significant R2 change of 0.56 (p < .001). The entry of the visual cognition variable at step 3 always produced a significant increase in R2 (range 0.11–0.17). The beta coefficients at step 3 show that only in one case did age have a significant coefficient (mental rotation), whereas the beta coefficient for both receptive vocabulary and for the visual-spatial variables always had values with p < .05. For some variables (visual discrimination, visual STM and mental rotation), there was a noticeably higher beta coefficient for receptive vocabulary than for the relevant visual ability. Additionally, we investigated the relationship between the visual perception subtests and mathematics performance in each group. For the TD group, the correlation between age and receptive vocabulary was above 0.80 giving rise to concern about multi-collinearity (see Table 6). However, all the relevant statistics for tolerance and VIF were satisfactory, as were Mahalanobis values and Cook’s distance. In the regression analyses, as before, age was entered at step 1, and receptive vocabulary was entered at step 2, with the relevant visual perception subtest being entered at step 3. The findings from these analyses are given in Table 7. For the TD group, the adjusted R2 was 0.83. The changes in R2 at step 1 and 2 were as before (for age at step 1, R2 = 0.82 (p < .001; for receptive vocabulary at step 2 R2 = 0.02 (p < .070). However, at step 3 the entry of the visual perception subtests, failed to significantly increase this variance except for one variable (visual sequential memory). Examination of the standardised beta coefficients at step 3 revealed that age was a highly significant predictor, and that the beta coefficient for visual sequential memory indicated it was the only other significant predictor. For the CP group, the adjusted R2 value was 0.58 and there were similar statistics for step 1 and 2 (step 1, age, R2 change = 0.04, p < .001; and at step 3, receptive vocabulary, R2 change = 0.56, p < .001). The entry of the visual perception subtests at step 3 always produced a significant increase in R2 except for visual discrimination and form constancy. The beta coefficients at step 3 showed that age was never a significant predictor of mathematical ability, whereas receptive vocabulary was always a significant predictor, and five of the visual perception subtests were significant predictors (the exceptions being visual discrimination and form constancy).","This study was carried out in order to determine whether the difficulties that children with CP have with developing mathematical skills are related to their visual and spatial perception impairments. The children in the CP and TD groups were matched on receptive vocabulary, an assessment commonly used with children with CP (Bishop et al., 1990; Dahlgren Sandberg, 2001). The children with CP were selected to have spoken language and manual dexterity which were sufficient to complete the tests, so the results of this study cannot be generalised to all children with CP. The group comparisons revealed that mathematics and visual-spatial abilities in the CP group were at a lower developmental level (based on standardised scores) than in the TD group and that their visual-spatial abilities were generally very poor. The group comparisons showed that the test of visual perception (TVPS-3) was the most problematic of the four assessments for the CP group as shown by Cohen’s d (Tables 2 and 3) and so we will discuss this assessment first. The analyses revealed that the CP group had lower raw scores on all the visual perception subtests than the TD group and for all assessments there was a large to extremely large difference. Given that we carried out analyses on all seven subtests it is possible that some of these findings were due to chance factors, however, all differences were in the same direction and the findings for the different subtests were similar. Across the seven subtests there was no clear evidence of less impaired performance on the subtests without notable memory demands (Visual Discrimination, Spatial Relationships, Form Constancy, Figure Ground, and Figure Closure) compared to those that had memory demand (Visual Memory, Visual Sequential Memory). Thus, our findings suggest that children with CP have impaired performance with all the visual perception subtests in the TVPS-3, irrespective of the form of the task. These figures are comparable to those produced in a study of children with CP who also used the TVPS-3 (Menken et al., 1987). The finding that children with CP had the lowest scores on the visual perception test, in comparison to the other three tests which also assessed cognitive processing (memory, executive functioning, general problem solving), suggests that visual perception impairments rather than more general cognitive impairments might be the primary disability; although there was evidence from the visual perception test that visual memory was both severely impaired and related to mathematical ability. Consequently, there is the possibility that performance on some cognitive and mathematical assessments is impeded by impairments in basic visual perception rather than the cognitive demands of the tasks. We now turn to consider the other three visual-spatial assessments. A computer version of the WMTB-C visual STM task has been used by Gagliardi et al. (2013), who reported significantly lower performance of children with CP compared to a CA comparison group. The children with CP in our investigation also had significantly lower performance than the TD group on the visual STM task. One reason for the low raw scores of the CP group on this test might be because of poor fine motor skills as suggested by Gagliardi et al. (2013), however the children in our study were considered by their teachers to have sufficient manual dexterity to complete simple assessment tasks, and all of the children completed this test. Furthermore, Stadskleiv et al. (2017) found that the physical dexterity of their participants with CP made no difference to their abilities in a recall test (completed on a computer), and informal observations by the first author suggested that the responses on the visual STM task of children in our study were not dependent on precise motor movements other than pointing. All this supports the idea that the CP group’s lower scores are more likely to be due to their difficulties with perceiving the shape and positions of the blocks or remembering the sequences of the blocks rather than their motor skills (see Gagliardi et al., 2013). The Mental Rotation Task relies on the children’s abilities to perceive and store visual information while attention is given to another picture and this process is likely to involve executive working memory processes (Henry, 2011). A previous study using mental rotation tasks which required children with CP to identify shapes that were similar or mirror images of each other, demonstrated that the CP group had significantly longer response times than the TD group, but a similar amount of errors (Courbois et al., 2004). In our study, the children needed to identify markers such as the position of the red and blue gloves on the hands of the monkeys or the position of the arms (see Fig. 1) and remember these orientations while looking at the monkey in the bottom of the picture. We found that the CP group had a significantly lower percentage of correct responses than the TD group, and Cohen’s d indicated that their decrement of performance on mental rotation was similar to the TVPS-3 subtests of Visual Discrimination, Spatial Relationships and Form Constancy, all of which involve the ability to make comparisons between different visual stimuli that are presented on one page. Thus, for the CP children, the Mental Rotation task appeared to have a similar profile of scores as other comparable visual-spatial tasks and it did not appear to be markedly more difficult than other assessments of visual-spatial abilities. The children with CP also had BAS-3 matrices scores that were significantly lower than those of the TD group. The identification of low matrices scores in children with CP is consistent with other studies which have used the RPM/RCPM (Jenks et al., 2012; Peeters et al., 2008). Our findings are also consistent with the evidence that children with CP have better receptive language than non-verbal ability as assessed by matrices tasks (Bishop et al., 1990; Dahlgren Sandberg, 2001). However, our findings differ from those of Ballester-Plané et al. (2016), whose study showed that participants with CP had better RCPM performance than receptive language. The reasons for this discrepancy between the findings from their and our investigation are not clear, although their participants were mainly adults and the types of CP may be different as our participants had mixed types of CP. Research comparing groups of children with CP suggest that children with dyskinetic CP may perform better on receptive language assessments, but results are not conclusive (Ballester-Plané et al., 2018). Further research is needed to understand the reason for any differences between adults and children with CP. One explanation for the impaired visual-spatial abilities in children with CP is that there could be motor or even perceptual difficulties when executing relevant manual tasks that are used in the assessments (Abercrombie, 1964; Akhutina et al., 2003). In addition, van Rooijen et al. (2012) reported that fine and gross motor skills had significant correlations with arithmetic abilities in children with CP, and fine motor skills was a significant predictor in a structural equation model which contained assessments of decoding and non-verbal intelligence. As already mentioned the visual-spatial tasks we administered had minimal motor components which suggest that poorer visual-spatial performance was unlikely to have been directly due to motor impairments. Regression analyses were conducted to investigate whether visual-spatial abilities were related to mathematics performance and whether the relationships were similar in the CP and TD groups. When interpreting these relationships, it is useful to bear in mind that having low scores on a visual-spatial assessment does not necessarily result in the assessments being significantly correlated with mathematical ability. In theory, children could have had low scores on visual-spatial assessments, but these scores might not be significantly related to mathematical ability, however, in their study, van Rooijen et al. (2012) report that the RCPM was significantly correlated with arithmetic ability in children with CP. The first set of regression analyses involved age and receptive vocabulary as covariates, and each of the four visual-spatial perceptual abilities. In the TD group, only age was a significant predictor of mathematical ability. In contrast, in the CP group, receptive vocabulary and the four visual-spatial variables were significant predictors of mathematical ability. This suggests that for the TD group, age and by implication age-related general cognitive development, is more important than receptive vocabulary and visual-spatial perceptual abilities in the development of mathematical competence. Whereas, for children with CP, our findings suggest that receptive vocabulary and visual spatial abilities are more important than age in the development of mathematical competence. We also investigated whether there were significant relationships between the visual perception subtests and mathematics ability (see Table 5). Regression analyses revealed for the TD group that at step 3, mathematical ability was significantly predicted by age, but not by receptive vocabulary and visual perception (except in the case of visual sequential memory). In contrast, for the CP group, receptive vocabulary was the most important predictor of mathematical ability, and all the visual subtests except for visual discrimination and form constancy were significant predictors of the mathematical ability at step 3. These findings suggest that visual-spatial impairments in CP are wide reaching with cascading developmental effects on other abilities. In addition, the presence of the significant relationships between visual-spatial abilities and mathematics, even after accounting for age and receptive vocabulary, suggests that in the CP group impaired visual-spatial abilities limit the development of mathematical competence. However, given the number of regressions conducted caution is needed when interpreting the significance of the findings, although it is reassuring that the different analyses produced remarkably consistent findings. Implications and summary ~~~~~~~~~~~~~~~~~~~~~~~~ This study has confirmed that many children with CP have visual-spatial perception difficulties which are likely to have an effect on their ability to develop mathematical skills. An inability to perceive and identify shapes and patterns can affect how children learn to read and write numbers; learn to calculate sums such as on a number-line; read and understand written problems; and set out sums. These are all basic skills that young children develop before more complex tasks can be learnt. Teachers and clinicians need to be aware of these difficulties and modify their teaching approaches to include tasks and activities that can help to improve the children’s visual and spatial capabilities. Our findings suggest that the largest impairments in these abilities involved visual memory, visual spatial relationships and figure ground, while the strongest predictors of mathematical ability involved visual memory, visual spatial relationships and visual sequential memory. Therefore, activities that target these processes might be of help to the education of children with CP. To summarise, researchers have identified that many children with CP have difficulties with visual and spatial perception and have delayed development in mathematical skills. Our investigation supported and extended these findings. The analyses revealed that children with CP, in comparison to a group of receptive language matched TD children, had significant poorer performance on an assessment of mathematical ability and on a battery of visual-spatial assessments which only required responses involving basic motor abilities. Furthermore, we did not find that tasks which involved additional demands from the active processing of visual-spatial information resulted in markedly worse performance. Moreover, in the CP group virtually all of the visual-spatial assessments were significant predictors of mathematical ability even after controlling for the effects of age and receptive vocabulary; whereas in the TD group, virtually all the visual-spatial assessments had non-significant relationships with mathematical ability. These findings suggest that in children with CP, visual-spatial abilities may be important for carrying our mathematical tasks, and more important than for TD children. Further research, possibly based around experimental interventions, is needed to determine whether these associations involve a causal relationship."],["Participants were asked to assess their own personality (i.e. Big Five scales), the personality of politicians shown in brief silent video clips, and the probability that they would vote for these politicians. Response surface analyses (RSA) revealed noteworthy effects of self-ratings and observer-ratings of openness, agreeableness, and emotional stability on voting probability. Furthermore, the participants perceived themselves as being more open, more agreeable, more emotionally stable, and more extraverted than the average politician. The study supports previous findings that first impressions affect decision making on important issues. Results also indicate that when only nonverbal information is available people prefer political candidates they perceive as having personality traits they value in themselves. © 2014 The Authors. --------------------------------------------------------------------------------","People form first impressions on the basis of appearance and other nonverbal cues. Such cues not only elicit quick attributions of personality traits and emotional states but also appear to provide a sufficiently reliable source of information to support accurate assessment of personality (e.g. Ambady, Bernieri, & Richeson, 2000; Borkenau, Mauer, Riemann, Spinath, & Angleitner, 2004; Kenny, Horner, Kashy, & Chu, 1992). This ability might help to make social interaction smoother but it is undeniable that such snap judgments also affect public decision making. For instance, the verdicts of judges in small-claims courts have been shown to be at least somewhat influenced by the facial features of the defendant (Zebrowitz & McDonald, 1991). Similarly in the political arena: successful self-presenters create social bonds with an audience not only by finding the right words but also by displaying the “right” behavior (Cherulnik, Donley, Wiewel, & Miller, 2001; Stewart, Waller, & Schubert, 2009). Nonverbal cues can be so convincing that attributions of competence and other personality traits to photographs of political candidates have been successfully used to predict electoral outcomes (e.g. Antonakis & Dalgas, 2009; Ballew & Todorov, 2007; Olivola & Todorov, 2010; Poutvaara, Jordahl, & Berggren, 2009). Consequently, in addition to voter and candidate ideology (e.g. Caprara, Barbaranelli, Consiglio, Picconi, & Zimbardo, 2003; Roets & Van Hiel, 2009), politicians’ appearance and their nonverbal behaviors may also affect people’s voting behavior and how they judge candidate personality. Communication, however, is not a one-way process. Although there is often no direct interaction between speakers on stage and their audience, information communicated has to be processed by the intended receivers. This processing depends in part on how the members of an audience perceive and relate to the speaker. Research has shown that people feel more closely connected to others they perceive to be similar to themselves in attitude and personality (e.g. Berscheid, 1966; Byrne, 1961; Tenney, Turkheimer, & Oltmanns, 2009). More pertinent to our research, this connection also appears to play a role in politics, particularly when people do not gather much information about candidates and their positions. For instance, people seem to vote for politicians whose personality traits are similar to their own (Caprara, Vecchione, Barbaranelli, & Fraley, 2007). Other findings even hint that the “similarity creates liking” relationship applies to physiognomic features. If voters are unfamiliar with politicians they seem to prefer candidates in whose faces they recognize themselves (Bailenson, Iyengar, Yee, & Collins, 2008). The current study investigated the relationship between first impressions and people’s tendency to favor others they regard as having a similar personality. We conducted an experiment in which participants rated short video clips of politicians giving a speech. The main focus of the study was first impressions formed by nonverbal information; so to avoid interference from speech content and different degrees of prominence we presented silent video clips showing politicians that were unknown to our participants.","’ impressions were collected using a brief version of a Big Five personality inventory. We also asked participants to report their own personality and give an estimate of the likelihood that they would vote for the speakers they had seen. Judging a stranger’s personality by brief displays of behavior may be fairly accurate (see above); however, people are often misled by first impressions and tend to simplify decision processes by relying on simple rules (e.g. Kahneman, 2003; Olivola & Todorov, 2010). Given that similarity creates liking, we assumed that people sometimes use their own personality as a kind of reference point when expressing preferences for others. This effect may be even more pronounced for personality traits people value highly and in situations of low information such as in our experimental setting. We used polynomial regression analyses with response surface plots to analyze the relationship between perceived personality, self-rated personality and voting probability (e.g. Edwards & Parry, 1993). These statistical procedures provide information about how congruence and incongruence between two independent variables relate to a dependent variable. They also yield coefficients describing the nature of the relationships between variables (i.e. linear and curvilinear), thereby providing a more comprehensive picture of how the different variables are interrelated than other statistical procedures. In summary, on the basis of research providing evidence for “similarity effects” in different domains and the finding that people perceive themselves as being above average in ability and character (e.g. Alicke, Klotz, Breitenbecher, Yurak, & Vredenburg, 1995; Ehrlinger, Johnson, Banner, Dunning, & Kruger, 2008) we made the following predictions. First, we hypothesized that our participants tend to “vote for” politicians in which they perceive personality traits they also preferably ascribe to themselves. We did not have distinct hypotheses to which personality traits this applies but expected to find patterns of congruence between self-ratings and observer-ratings of some traits. To reveal such patterns of congruence we used the aforementioned response surface analyses (RSA), because this procedure provided detailed insights into whether and in what way self-ratings and observer-ratings are related to voting probability. Second, we assumed that potential candidates would be judged according to the self-attributed “high standards” of their voters. We therefore expected the participants of the rating experiment to perceive themselves as being above the average politician in those personality traits they preferably ascribe to themselves. In other words, for personality traits the participants valued highly we expected to find large differences between self-ratings and observer- ratings. Consequently, the aim of this study was not to show that people are able to make quite reliable guesses about someone else’s personality on the basis of brief displays of behavior, but that people integrate self-perceptions into the guesses they make. Participants ~~~~~~~~~~~~ Eighty participants (42 females with age M = 23.1 years, SD = 3.7; 38 males with age M = 23.9 years, SD = 4.4) were approached personally at locations throughout the University of Vienna and asked to take part in a rating experiment. Raters were not reimbursed for participating in the experiment. Stimulus preparation and procedure ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We selected 40 speeches from the German houses of parliament (20 female and 20 male speakers), and randomly extracted brief video segments from each of the speeches to give 40 video clips with an average length of 15 s. The lower portion of the video clips was cut off to remove captions which gave information about the speaker’s name and party. The speakers we chose were ordinary, non-prominent members of the German Bundestag, who were not very well known in Germany and even less well-known in Austria. In addition, we asked participants if they recognized any of the speakers they had judged. None of the politicians were known to any participant. Stimuli were presented using a rating program. The video clips were presented on the left-hand side of the program’s user interface; on the right-hand side 20 bipolar personality terms (e.g. outgoing and reserved) based on a brief German version of the NEO-FFI personality inventory (Borkenau & Ostendorf, 1991) were displayed. This questionnaire assesses the Big Five personality dimensions of extraversion, agreeableness, conscientiousness, neuroticism or emotional stability and openness. Below the list of personality items the user interface displayed an additional bipolar item: “I would vote for this candidate” or “I would not vote for this candidate”. Participants completed their ratings by dragging a “trackbar control” to the right or to the left pole of the bipolar scales using a computer mouse. The position of the bar corresponded to points along a semantic differential, which was divided into 100 subunits with 0 being the maximum value of the item on the left, 50 being the neutral position, and 100 being the maximum value of the item on the right. The same scale was used for the voting choice item, enabling participants to give an estimate of the probability that they would vote for a candidate. In order not to overtax the participants and to reduce the influence of stimulus order, each participant only rated a subset of eight speakers randomly selected from the 40 video clips. Analysis ~~~~~~~~ We aggregated the questionnaire data by summing items belonging to a latent personality dimension (i.e. each personality dimension comprised four items). We then averaged observer-ratings across the subsets of eight stimuli each participant had rated. The resulting dataset − consisting of 80 self-ratings, 80 mean observer-ratings (i.e. for each personality dimension), and 80 ratings on the continuous voting scale − was then analyzed by means of polynomial regressions and response surface analysis (RSA). Polynomial regression yields regression coefficients for two linear terms (i.e. observer-ratings and self-ratings in this study), their interaction (i.e. relationship between observer-ratings and self-ratings), and two quadratic terms (i.e. squares of observer-ratings and self- ratings) and relates them to an independent variable (i.e. voting probability). To guard against multicollinearity and to adjust differences in variances, which often affect the interpretation of multiple regressions coefficients, we applied z-standardization before the polynomial regression analyses. The regression coefficients obtained by polynomial regression are used to calculate the response surface parameters a1–a4. These parameters define the response surface (RS) plane of a three dimensional plot. RS plots have, numerically, a line of congruence (LOC: X = Y) and a line of incongruence (LOIC: X = −Y), which are derived by fully crossing the numeric levels of two continuous predictor variables X and Y. The LOC is defined by a linear slope (a1) and a curvature (a2); similarly the LOIC is defined by a linear slope (a3) and a curvature (a4). Thus, the LOC and the LOIC provide insight into how congruence and incongruence between the independent variables are related to a dependent variable, which is plotted on the Z-axis (see Table 1 and for more information see Edwards & Parry, 1993; Schönbrodt, 2013; Shanock, Baran, Gentry, Pattison, & Heggestad, 2010). Applied to our data this means that self-ratings and observer-ratings were plotted on the X-axis and the Y-axis of the RSA plot, while voting probability was plotted on the Z-axis (see Figs. 1–3). Thus, self-ratings and observer- ratings defined the plane of the three-dimensional plot. The LOC running from the near corner to the back corner of the plot was derived by crossing equal values of self-ratings and observer-ratings (i.e. where self-ratings of −1 corresponded with observer-ratings of −1; self-ratings of 0 corresponded with observer ratings of 0, etc.). The LOIC running from the left to the right corner of the plot was derived by crossing equal pairs of negative and positive values of observer-ratings and self-ratings (e.g. where −1 met +1). Along the LOC the parameter a1 estimated a linear additive effect (i.e. negative or positive slope) of self-ratings and observer-ratings on voting probability, whereas the parameter a2 gave information about the curvature (i.e. negative or positive) of this relationship. In contrast, the parameters a3 and a4 informed about how a linear or a non- linear relationship along the LOIC affected voting probability. To examine which personality traits participants valued in themselves and how they perceived themselves in relation to the politicians they saw, we compared participants’ self-ratings with their average politician rating (i.e. the mean rating of eight politicians each participant rated). The results of these analyses are presented as t-tests with standard effect size measures (i.e. Cohen’s d), which in combination with the means give an estimate of the degree and direction in which self-ratings differed from observer-ratings. All these statistical analyses were done using the statistical software package R (R Core Team, 2013). Power analysis ~~~~~~~~~~~~~~ We had no specific hypotheses about the degree to which the different dimensions of the Big Five traits would be related to voting probability. Other studies have found bivariate and multiple correlation coefficients for certain personality traits with voting behavior that explained more than 16% of total variance, so we expected that there would be a medium effect size for at least some dimensions (Olivola & Todorov, 2010; Poutvaara et al., 2009). We therefore assumed a medium effect (R2 = .15) for polynomial regression analyses with an alpha level of .05, a power level of .8, and five predictors. On the basis of these assumptions an a priori power analyses suggested an optimal sample size of 73 participants. Our study was also exploratory in that regard that the t-tests and their corresponding effect size measures were used to determine in which personality traits self-ratings differed markedly from observer-ratings. We expected that for highly valued personality traits participants’ self-ratings would be at least moderately different from the average politician rating because other findings indicated that the better-than- average effect can be moderate to strong (Alicke et al., 1995; Ehrlinger et al., 2008). We therefore performed an a priori power analyses with a Cohen’s d of .5 (i.e. threshold for medium effect), an alpha level of .05, and a power of .8, which suggested an optimal sample size of 64 individuals for the t-tests we applied. Power analysis was performed using the statistical software package R (Champley, 2012).","On average each politician was rated by 16 different participants with a range of 13–19 ratings. Measures of the questionnaire’s internal consistency with regard to observer- ratings yielded moderate to high reliabilities for extraversion (Cronbach’s α = .80), agreeableness (α = .87), conscientiousness (α = .72), openness (α = .62), and emotional stability (α = .73). The participants’ self-ratings on the Big Five personality dimensions showed a similar pattern. There were moderate to high reliabilities for extraversion (α = .88), agreeableness (α = .84), conscientiousness (α = .82), and emotional stability (α = .70), but low reliability for openness (α = .50). For this reason, the interpretation of the results for openness presented below should be treated with some caution. Results of response surface analyses (RSAs) are presented in Table 1. The proportion of the total variance explained ranged from R2 = .10 for extraversion to R2 = .37 for agreeableness. The parameters a1–a4 are central to RSA and provide estimates of congruence and incongruence. Therefore, interpretation of the results mainly focuses on these parameters. We assumed that similarity would play a role, but we had no specific hypotheses about which personality dimensions would show a self-observer similarity (or congruence) effect on voting behavior. In this regard our analysis was exploratory. Inspection of the correlation coefficients revealed that extraversion yielded the lowest R2, but also that none of the regression coefficients reached statistical significance. This indicates that there was a relatively weak relationship between self-ratings and observer-ratings for extraversion and voting probability. A similar interpretation can be applied to the data on conscientiousness. Although R2 for this personality trait was higher than for extraversion, none of the relevant RSA parameters were significant. In contrast, RSA for all the other Big Five personality dimensions yielded significant results for at least one of the relevant parameters. These results and the corresponding plots are discussed in greater detail in the following subsections. RSA for openness ~~~~~~~~~~~~~~~~ RSA for openness provided a relatively strong positive linear relationship (a1 in Table 1) along the LOC, which runs from the near corner to the far corner of the RSA plot (Fig. 1). At first sight this indicates a congruence effect and that voting probability increased as self- and observer-ratings of openness increased. However, the parameter b1 (i.e. observer-rating in Table 1) and b2 (i.e. self-rating in Table 1) of the polynomial regression revealed that observer-ratings of openness have a prevailing influence on the RSA parameter a1. Apparently the participants’ aptness to vote for somebody they perceived as open was predominantly affected by their first impressions. There was also a strong linear relationship (a3 in Table 1) along the LOIC, which runs from the left corner of the plot to its right corner indicating that self- and observer-ratings for openness showed both congruence and incongruence. Inspection of Fig. 1 reveals more details how the different variables were related and supported interpretations based on the parameters of the polynomial regression. Overall, voting probability increased (i.e. higher values on Z-axis) when observer-ratings for openness increased. Low self-ratings combined with low observer-ratings (i.e. congruence) were not associated with a higher voting probability than high self-ratings combined with low observer-ratings. However, high observer-ratings were associated with higher voting probability regardless of whether they were combined with low self-ratings (i.e. incongruence) or high self-ratings (i.e. congruence). In conclusion, although data points in the RSA plot of openness appear to follow the LOC, we found no clear effect of congruence for openness. Irrespective of their self-rating, participants preferred politicians they rated highly for openness. RSA for agreeableness ~~~~~~~~~~~~~~~~~~~~~ For agreeableness RSA indicated a strong linear relationship (a1 in Table 1) along the LOC. This suggests that congruence between self- and observer-ratings of agreeableness had an impact on voting probability. However, Fig. 2 shows that the relationship between self- and observer-ratings for agreeableness and voting probability were not so simple. When self-ratings and observer-ratings were low, voting probability was also low (i.e. low value on Z-axis of plot); concurrent increases in self- and observer-ratings were associated with higher voting probabilities (i.e. positive slope of a1). There was also a non-significant curvature (a2 in Table 1) along the LOC, indicating that voting probability decreased slightly when ratings for agreeableness were high. In addition, although RSA yielded no significant effect for incongruence (see a3 and a4 in Table 1), high voting probabilities were observed when agreeableness ratings were somewhat incongruent. For instance, some participants who rated themselves high on agreeableness also tended to choose higher voting probabilities even for politicians they rated not agreeable. In summary, conclusions for agreeableness resembled the conclusions we drew for openness. Although congruence effects were more pronounced for agreeableness than for openness, RSA indicates that participants mostly based their judgments on first impressions and preferred politicians they perceived as being highly agreeable. In comparison with this self-ratings of agreeableness seemed to be of minor importance. Consequently, we found no clear similarity effect for this personality trait but a strong effect of perceived agreeableness on voting probability. RSA for emotional stability ~~~~~~~~~~~~~~~~~~~~~~~~~~~ RSA for emotional stability yielded a pronounced positive curvature along the LOC (a2 in Table 1) and a pronounced negative curvature along the LOIC (a4 in Table 1). Consequently, self-ratings and observer-ratings of emotional stability produced non-linear effects of congruence and incongruence. The significant negative slope of parameter a4 (i.e. producing a concave curvature along the LOIC) enhanced the robustness of the congruence effect along the LOC. It shows that voting probability decreased when combinations of self-ratings and observer-ratings were “far from the LOC” and incongruence between these variables was strong. Inspection of Fig. 3 provides more clarity. The path of the LOC in Fig. 3 reveals that voting probability was very high when low ratings were given for both own emotional stability and speakers’ emotional stability. Voting probability decreased slightly at moderate levels of congruence, and rose again at higher levels of congruence between self-ratings and observer-ratings for emotional stability. Slight incongruence was also associated with high voting probability but when self-ratings and observer-ratings for emotional stability differed too much (i.e. low self-ratings combined with high observer-ratings and vice versa), voting probability was lower. Overall, in comparison with the other personality traits results for emotional stability provided the clearest congruence effect. Differences between self-ratings and observer-ratings ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The comparison of observer-ratings of personality with self-ratings revealed that on average participants gave themselves higher ratings on all five personality dimensions (see Table 2). Exploratory analyses of these differences with t-tests and effect size measures revealed strong effects for openness, agreeableness and emotional stability. A medium effect size was found for extraversion and a very low one for conscientiousness. This shows that participants believed themselves to be far more open, agreeable and emotionally stable than the average politician. However, effect size measures (Cohen’s d) should be interpreted with some caution because there were violations of variance homogeneity.","In this study we asked participants to rate themselves and short video clips of politicians on scales measuring the Big Five personality traits and to give an estimate of the probability that they would vote for each politician they evaluated. We used polynomial regression and response surface analyses to examine whether the participants were more likely to vote for politicians they perceived to be similar to themselves in terms of personality. We expected similarity to have an effect, but we had no specific hypotheses about which personality dimensions played a role. It was revealed that regression models for openness, agreeableness, and emotional stability explained a high proportion of the variance in voting probability. Less variance was explained by models for extraversion and conscientiousness. This suggests that our student sample regarded extraversion and conscientiousness as personality traits of minor importance to voting decisions. A tendency to assign a higher voting probability to speakers that were perceived to be agreeable is perhaps not surprising as politicians should appear likeable. However, the propensity shown by our sample to attach great importance to openness and minor importance to conscientiousness when making voting decisions might only hold true for a student sample such as the one used in our experiment. Future studies should use samples from different social groups, because it has been shown that the relationships between political attitude and personality traits depend on social context and ideology (e.g. Gerber, Huber, Doherty, Dowling, & Shang, 2010; Roets & Van Hiel, 2009) and it is reasonable to assume that this also extends to perceptions of politicians’ personalities. Although we found some incongruence between rater-politician similarity and its relationship to voting probability, detailed analyses showed that participants tended to vote for “someone like me” when the candidate was perceived to have similar emotional stability. There were also similarity effects for agreeableness and openness but these were less clear; the most prominent difference from the similarity effect observed for emotional stability was that for openness and agreeableness concurrent low self- and observer-ratings were not associated with higher voting probability. In addition, in both cases observer-ratings were the more dominant contributor to the found similarity effects and had a prominent effect on voting probability. It is conceivable that observer-ratings on these personality traits have to reach a certain level before a similarity effect delivers an increase in voting probability. People refrain from voting for a candidate they rate as low on the personality traits they favor in political candidates. Such reasoning may also be applied to emotional stability, but the surprising result was that for this personality trait voting probability was highest when low self-ratings were paired with low observer-ratings or high self-ratings were paired with high observer- ratings. In between, at moderate levels of self- and observer-ratings voting probability was less high, although the effect was still strong. Explanations for this pattern of results are speculative at present. Our measure of emotional stability contained items such as “aroused” vs. “composed”, which are adjectives that are typical of temporary emotional states occurring during an experiment. It is possible that participants matched their current emotional states with those they ascribed to the speakers they saw. Some participants may have preferred a highly aroused politician because they were also easily aroused; others may have preferred a very composed politician because they were very composed. This kind of matching behavior might have influenced voting decisions more strongly at the extremes of emotional stability. Student’s t-tests revealed that in nearly all cases the participants gave themselves higher ratings for the traits they seemed to value highly in politicians they would vote for (see previous section). Participants regarded themselves as being more open, more agreeable, more emotionally stable and more extraverted than the average politician in our sample. Whilst this might be due to the poor public image of politicians, it is common to find that people regard themselves as “better than average” (e.g. Ehrlinger et al., 2008). Surprisingly, no noteworthy difference between self-ratings and observer-ratings could be found for conscientiousness, which suggests that our student sample did not value conscientiousness as highly as the other personality traits we investigated. This result may be due to the difficulty of evaluating conscientiousness from brief displays of behavior. However, taking into account the results of the regression analyses and findings suggesting that the easy-to-recognize personality trait of extraversion (e.g. Kenny et al., 1992) was not particularly highly valued, it is plausible to suggest that people want their politicians to reach their own “high standards” on valued traits. Such a conclusion might be applied to openness and agreeableness in this study. As decision making is not confined to the political domain, the current findings may also be interpreted in a broader context. It is possible that in first encounters, during which a common problem has to be solved, people favor others they rate as closer to their self-rated “high standard” on certain personality traits. In contrast to other studies which have found that voters preferred politicians with similar personality traits (Caprara et al., 2007) our study was based on appearance cues and first impressions. This may more closely resemble the situation in which politicians find themselves when they are newcomers to the political arena or when they are candidates for a new party about which ideological information is scarce. Furthermore, research suggests that even when more information is available voting decisions are often guided by superficialities and nonverbal cues (e.g. Antonakis & Dalgas, 2009; Olivola & Todorov, 2010). Our experimental set-up has limitations, because we removed speech content and party membership from the video clips. Therefore, caution in extrapolating the results to real life situations is advised and additional work needed to underpin the findings. However, in follow up studies other potentially influential variables can be easily added to the experimental design and tested under controlled conditions. Manipulating the experimental setting this way could bridge the gap between studies that examine the role of nonverbal cues and those that examine the role of ideology in decision making and help to disentangle the role of appearance from that played by other sources of information. Inclusion of variables on party membership, for instance, might reveal whether there is some kind of an “attraction leads to similarity” effect when people make their voting decisions (e.g. Morry, 2007). It is conceivable that people who receive more background information “adapt” their ratings of politicians to make them more or less similar to themselves. In addition, future experiments should enable both randomization of the presented stimuli and assignment of same sets of stimuli to different participants. This would provide insights into inter-judge agreement and to what degree it varies for ratings of different personality dimensions. In summary, our study supports previous findings showing that first impressions and visual appearance cues affect public decision making processes (see also Antonakis and Dalgas, 2009; Ballew & Todorov, 2007; Poutvaara et al., 2009). However, we extended earlier research by investigating the relationship between self-ratings, observer-ratings and voting probability with response surface analyses (RSA). This statistical tool enabled us to uncover relationships among the investigated variables that traditional methods of analyses would not have revealed. For instance, we found similarity effects for agreeableness, emotional stability and openness, but closer examination of the RSA plots showed that these effects had different appearances. RSA might thus help to refine hypotheses about similarity effects for personality traits by providing a more detailed picture of relationships than other analytical tools. All in all, our results suggested that people want a candidate to possess personality traits they believe themselves to possess and which therefore have a very high value for them."],["The purpose of this paper is to determine whether memories can have any benefit for their subjects while being distorted.1 There are a number of things that one may have in mind by ‘benefit’ while referring to memory. For that reason, the formulation of the issue that will occupy us here admits several possible readings. It will therefore make for clarity if we begin our discussion by specifying, in Section 1, the types of benefits with which we will be concerned in this discussion. One may also have different things in mind by ‘distortion,’ depending on one’s views about the function of memory. Thus, in order to formulate the topic of our discussion precisely, I will distinguish, in Section 2, two pictures of what memory is supposed to do, and two associated notions of distortion. Next, I will put forward two types of memories that, I will argue, can qualify as cases of beneficial distortion under very specific circumstances. In Section 3, I will discuss the case of so-called ‘observer memories’ and, in Section 4, I will discuss the case of so- called ‘fabricated memories.’ My contention will be that, in both cases, some of those memories can, on the one hand, be advantageous for the subject to have while, on the other hand, her faculty of memory has failed to perform its proper function by producing them. The significance of this claim for the two pictures of what memory is supposed to do will be explored in Section 5.","There are at least two ways in which having a memory can be beneficial for a subject. One of them is epistemic. Having a memory may provide the subject with knowledge of, or at least justification for a belief about, the past. The memory does this by supplying the subject with evidence, or grounds, for a certain belief; a belief in the content of the memory or, more precisely, in part of that content.2 Thus, the memory allows the subject to be in ‘cognitive contact’ with an event in her past: It puts the subject in a position to think about, and refer to, that event. When the evidence provided by the memory is good evidence, it is beneficial for the subject to be in that position. For the belief that the subject can form on the basis of her memory will, in that case, be justified. Why is that good for the subject? Beliefs are the sort of mental state which can be true or false. They are in some sense normatively governed by, or aimed at, truth. From the point of view of achieving truth, it is good for the subject to have justified beliefs rather than unjustified beliefs because the former, unlike the latter, tend to be true. Furthermore, when justified beliefs are true, they are such that they could not have easily been false. This is because, when justified beliefs are true, they are not accidentally true, or true by luck. This feature of justified beliefs confers a certain stability upon them: Whereas merely true beliefs are fleeting in that they are easy to undermine, justified beliefs are more likely to remain fast in response to conflicting information.3 Notice that, in order for the subject to enjoy this type of benefit from her memories, the subject’s faculty of memory must be trustworthy in the following sense. It must deliver memories that are likely to be accurate when they have been appropriately produced.4 Imagine that, for any memory of a subject, the fact that the memory in question was properly generated makes no difference as to whether the content of that memory is likely to be case or not. Suppose, now, that the subject has a particular memory. It is hard to see why the subject would be justified in believing the content of it. After all, on the scenario that we are considering, any of the memories that the subject is having could easily be misinforming her about her past. Thus, the belief that the subject would form by taking the content of her memory at face value is not likely to achieve truth. Suppose, however, that it does. Still, it does not seem that the belief in question would be justified. For if the fact that the memory was properly generated makes no difference as to whether its content is likely to be the case or not, then it seems that the belief in question could have easily been false. It turned out to be true, but the fact that the memory on the basis of which the belief was formed was properly generated did not contribute to that outcome. It seems, therefore, that in order for the subject to be in a position to form justified beliefs on the basis of her memories (and thus benefit from them epistemically), those memories must have the property of being such that if they have been properly generated, then they are likely to be accurate.5 Another way in which having a memory can be beneficial for the subject is by being adaptive. Having a memory may allow the subject to form a belief about the past which has a certain instrumental value for her. The belief may serve to represent the past in the way in which the subject needs to represent it in order for her to achieve one of her goals.6 Oftentimes, the relevant goal involves experiencing a certain type of emotion. In this scenario, the memory plays the role of supplying the subject with a representation of her past that is conducive to experiencing the emotion that is being sought by the subject; typically a positive emotion. When the memory allows the subject to represent her past in the way in which she is seeking to represent it from an emotional point of view, it is beneficial for the subject to have that memory. (We could call this type of adaptive benefit, an ‘affectively adaptive’ type of benefit.) Other times, the relevant goal involves making sense of one’s own behaviour towards, and one’s own thoughts about, some particular person or situation. In that scenario, the memory plays the role of supplying the subject with a representation of her past which makes her current behaviour towards some person or situation intelligible to herself, and it allows her to explain why she has certain thoughts towards the relevant person or situation. When the memory allows the subject to represent her past in the way in which she needs to represent it for comprehending her own mental life and her own behaviour, it is beneficial for the subject to have that memory as well. (We could call this type of adaptive benefit, an ‘explanatorily adaptive’ type of benefit.) More generally, from an adaptive point of view, it is beneficial for the subject to have a memory if representing the past in the way in which that memory presents it to the subject is an effective means of satisfying one of the subject’s goals. Notice that, in order for the subject to enjoy an explanatorily adaptive type of benefit from her memories, the subject’s faculty of memory must provide a firm structure in the following sense. It must deliver memories that cohere well with the rest of the subject’s mental states when those memories have been appropriately produced. Why is that? Imagine that, for any memory of a subject, the fact that the memory in question was properly generated by the subject’s faculty of memory makes no difference as to whether the content of that memory squares with the things that the subject has learnt about her past through testimony, and the things that the subject has inferred about her past from other things she knows. Suppose, now, that the subject has a particular memory. It is hard to see how the subject could create a representation of her past that made sense of her current behaviour and mental life by believing the content of that memory. For that memory is, on the scenario that we are considering, likely to be in tension with the behaviour and set of mental states which the subject needs to make sense of. In order for the subject to be in a position to form, on the basis of her memories, beliefs about her past which make sense of her current mental life and her current behaviour (and thus benefit from them in an explanatorily adaptive way), those memories must have the property of being such that if they have been properly generated, then they are consistent, and cohere well, with rest of the subject’s mental states. Likewise, in order for the subject to enjoy an affectively adaptive benefit from her memories, the subject’s faculty of memory must provide a firm structure in the same sense. Why is that? Imagine that the fact that, for any memory of a subject, the fact that the memory was properly generated by the subject’s faculty of memory makes no difference as to whether the content of that memory squares with the contents of the rest of the subject’s mental states which concern the seemingly remembered event. Suppose, now, that the subject has a particular memory. It is hard to see how the subject could use that memory to create a representation of a past event which was intended to produce in the subject certain emotions; those emotions which she is seeking to experience towards the event. For the relevant emotions arguably depend on the subject’s current thoughts, needs, expectations, intentions and wishes that involve that event. And the memory which is available to the subject is, on the scenario that we are considering, not likely to be consistent with the contents of those mental states.7 In order for the subject to be in a position to form, on the basis of her memories, representations of a past event which produce in her the emotions that she is seeking to experience towards that event (and thus benefit from them in an affectively adaptive way), those memories must have the property of being such that if they have been properly generated, then they are consistent, and cohere well, with the rest of the subject’s mental states.8 What is the relation between the two types of benefits? Adaptive benefits have been introduced in terms of goal satisfaction. In normal circumstances, a subject has goals of various types. Some of them are practical goals, such as the goal of being at a certain place at a certain time. Others are theoretical goals, such as the goal of having beliefs that are held for good reasons (as opposed to biases or prejudices). The question of how epistemic and adaptive benefits are related concerns, on the one hand, the distinction between these two types of goals and, on the other hand, the distinction between justified belief and true belief. Let me explain. For the purposes of achieving a practical goal, it is enough to have a true belief about the means that will maximise one’s chances of achieving that goal. Suppose, for instance, that I have the goal of attending a meeting in the city at a certain time, and I want to catch the bus to go to the city. I seem to remember that the bus is supposed to come to the relevant stop at 10 am. On the basis of my memory, I form the belief that the bus will be at the stop at 10 am. As it happens, my memory about these matters is highly unreliable, so I am not justified in having that belief. But it turns out that the bus does show up at 10 am. For the purposes of getting to my meeting on time, it was not necessary for my belief to be justified. It turned out to be true, which is all I needed to achieve my goal. Now, justified beliefs are more likely to be true than unjustified beliefs. For that reason, if a memory carries an epistemic benefit for me, then it is likely to carry an adaptive benefit with regards to some practical goals of mine. But there is no guarantee of that, since justified beliefs may not be true. Suppose, for example, that my memory is highly reliable when it comes to public transport schedules, but the bus is late and it shows up at 10.30 am. Then, my memory carried an epistemic benefit for me. It yielded the justified belief that the bus would be at the stop at 10 am. But the belief turned out to be false, which did not help me to achieve my goal of attending my meeting. With regards to that practical goal, therefore, my memory did not carry an adaptive benefit for me. If, however, we take into consideration theoretical goals as well as practical goals, then it seems that any memory that carries an epistemic benefit for me will carry some adaptive benefit. For a memory that carries an epistemic benefit for me allows me to form a justified belief on the basis of it. And if my belief is justified, then it seems that I am satisfying my theoretical goal of believing for good reasons. What about the converse? Do memories which carry an adaptive benefit always provide an epistemic benefit for the subject as well? It depends, once again, on the type of goal that the memory is helping the subject to achieve. If the goal in question is the theoretical goal of believing for good reasons, then it does seem that the memory will provide an epistemic benefit for the subject. But if the memory is helping the subject to satisfy one of her practical goals, then the memory does not need to do this by putting the subject in a position to form a belief that is justified. As we have just seen, a memory may be untrustworthy (and, for that reason, fail to provide an epistemic benefit for the subject), and yet it may help the subject to achieve one of her practical goals (and, for that reason, provide an adaptive benefit for her). As a matter of fact, the discussion in Sections 3 and 4 below is meant to illustrate this possibility with some interesting cases of memory. To sum up, memories can have an epistemic benefit for the subject, and they can also have an adaptive benefit for the subject. Furthermore, those benefits may be related or not, depending on the subject’s goals. However, in order for a memory to carry either of those two benefits, the subject’s faculty of memory must be functioning properly when it delivers it. Interestingly, ‘functioning properly’ seems to mean something different in each case. The two types of benefits that we appreciate in memories reveal two conceptions of what memory is supposed to do; two conceptions of the function of memory. These two conceptions are often assumed to be opposed to each other and, for that reason, they give rise to competing notions of what a distorted memory is. Let us consider the two pictures of memory in order.","There is a popular conception of memory within philosophy according to which the function of memory is content preservation. Memory is supposed to register and store the content of those (typically, perceptual) experiences that we had in the past by producing memories which inherit their contents from those experiences. On this ‘storage’ conception of memory, then, a subject’s faculty of memory functions properly when the contents of the memories delivered by it match the contents of the subject’s past experiences on which those memories originate.9 And, with this preservative notion of proper function, comes an associated notion of distortion: On the storage conception of memory, a subject’s faculty of memory has produced a distorted memory when the content of that memory does not match the content of the subject’s past experience on which the memory originates. To the extent that one appreciates the epistemic benefit of having properly generated memories, one will find the storage conception of memory appealing. After all, if a subject’s faculty of memory adequately carries out the preservative function that this conception of memory attributes to it, then this will allow a memory produced by it to be epistemically beneficial for the subject.10 Would something weaker than the preservation of the whole content of the original experience be enough for the resulting memory to be epistemically beneficial for the subject? Let us recall that the epistemic benefit that a memory may have for a subject depends on the belief that the subject is forming on the basis of that memory. And what is required for the memory to be beneficial relative to the belief that the subject is forming is the reliability of, not the whole content of the memory, but only the parts of it which are relevant for the content of the belief being formed. Suppose, for example, that my faculty of memory is reliable with regards to which actions I witnessed in the past but not with regards to which people were involved. Suppose, furthermore, that I have a memory of my brother tickling my sister mercilessly while we were kids. As it happens, it was me, and not my brother, who ticked my sister mercilessly. Now, imagine that I form, on the basis of that memory, the belief that someone was ticked mercilessly while I was a kid. In that case, my memory is epistemically beneficial with regards to that belief, but the whole content of the experience on which that memory originates has not needed to be preserved for it to carry that epistemic benefit. It was enough that my faculty of memory was trustworthy with regards to certain aspects in the contents of the memories that it delivers, namely, the actions being remembered.11 On the other hand, there is a popular conception of memory within psychology wherein memory is not a passive device for registering and reproducing contents. It is instead a faculty akin to imagination in its creative capacity. The main tenet of this ‘narrative’ conception of memory is that, in memory, we are engaged in an inventive project wherein we build representations of our past by integrating content that we have acquired through our own experience with content from other sources, such as testimony, inference and imagination. Elizabeth Loftus, for example, describes this picture as ‘a new paradigm of memory, shifting our view from the video-recorder model, in which memories are interpreted as the literal truth, to a reconstructionist model, in which memories are understood as creative blendings of fact and fiction.’12 The reference to an element of fiction in the integration process is telling. For memory, on this conception, is not meant to represent the past as we experienced it to be the case. Instead, the function of memory is to reconstruct the past in order to help us build a smooth and robust narrative of our lives. On the narrative conception of memory, then, a subject’s faculty of memory functions properly when the contents of the memories that it delivers have been reconstructed so as to easily fit together with the contents of the subject’s beliefs about her past. And, with this reconstructive notion of proper function, comes an associated notion of distortion: On the narrative conception of memory, a subject’s faculty of memory has produced a distorted memory, when the reconstructive process has not yielded a memory that meshes well with the contents of the subject’s beliefs about herself and her past and, for that reason, it does not fit into the subject’s narrative of her life. To the extent that one appreciates the adaptive benefit of having properly generated memories, one will find the narrative conception of memory appealing. After all, if a subject’s faculty of memory adequately carries out the reconstructive function that this conception of memory attributes to it, then this will allow the memories that the faculty produces to be beneficial for the subject from an adaptive point of view.13 Having distinguished these two conceptions of memory, two questions naturally arise. One question is which of the two conceptions is right in describing what the function of memory is. We will address that question in Section 5. A different (and, at this point, more pressing) question is which of the two conceptions is right in describing what memory does; whether it is the function of memory to carry out the relevant operation or not. The answer to this question seems to be that, to some extent, both conceptions are right. Reconstruction is a matter of degree. The more input from sources of information other than the experience on which the memory originates, the more reconstructed will the content of that memory be. And there may be extreme cases in which no part of the memory’s content has been inherited from the experience on which the memory originates. (Perhaps some pathological cases of confabulatory memory fall into this category.) But the fact of the matter is that most of our memories contain some pieces of information that they have inherited from the contents of the experiences on which they originate. Conversely, preservation is also a matter of degree. The more information is inherited from the content of the original experience, the more preservative will the content of the resulting memory be. And there may be extreme cases in which the whole content of the memory has been inherited from the experience on which the memory originates. (Perhaps cases of so-called ‘eidetic’ memory fall into this category.) But the fact of the matter is that most of our memories contain some pieces of information that they have not inherited from the contents of the experiences on which they originate. It seems, then, that our memories preserve, to some degree, the information that we acquired in past experience and, to some degree, they reconstruct it. Which explains the attraction and popularity of both conceptions of memory. Pulling apart the preservative function of memory from its reconstructive function helps us to appreciate that the original issue of whether memories can have any benefit for the subject while being distorted actually divides into two issues: The first one is whether it is possible for a subject’s faculty of memory not to carry out its reconstructive function while it generates a memory, and yet for that memory to have an epistemic benefit for the subject. The second one is whether it is possible for a subject’s faculty of memory not to carry out its preservative function while it generates a memory, and yet for that memory to have an adaptive benefit for the subject. In what follows, I will concentrate on addressing the latter issue. I will argue that memories which have been unreliably produced can nonetheless have an adaptive benefit for the subject based on two cases of memory distortion; so-called ‘observer memories’ and ‘fabricated memories.’ Let us now turn to these two interesting types of memories.","A memory may present a past event to its subject from two types of visual perspectives. One of them is the type of perspective from which the subject would have experienced the event if the subject had witnessed it (or had went through it) in the past. By having a memory that presents a past event from a perspective of this type, the subject visualises the event, but she does not visualise herself as part of it. Let us call memories that present events from a perspective of this type, ‘first-person’ or ‘field’ memories. A memory may also present a past event from the type of perspective that a different observer would have had to occupy in the past in order to witness the remembered event with the subject as a participant of it. By having a memory that presents the past event from a perspective of this type, the subject visualises not only the event but she also visualises herself, as it were, from the outside. Let us call memories that present events from a perspective of this type, ‘third-person’ or ‘observer’ memories. In this section, I wish to put forward the claim that observer memories can qualify as a case of beneficial memory distortion. In order to argue that observer memories can be beneficial despite being distorted one must, first of all, make the case that they are indeed distorted. From a preservative point of view, it seems quite clear that they are. Suppose that, years ago, I suffered an accident while driving, and I now remember the accident by having an observer memory of it. In virtue of having this memory, I picture the event from the point of view of a nearby pedestrian on the street, thus being able to visualise some details of my own physical appearance while I was at the wheel. Suppose that, on the basis of my memory, I form the belief that, at the time of the accident, I appeared to be unshaven and my hair appeared to be dishevelled. These facts about my appearance are not facts that I perceived at the time of the accident. (Let us stipulate that I was not looking at myself in the mirror while driving.) Thus, the source of this information in the content of my observer memory must be other than the perceptual experience on which my memory originates. It must be testimony, the imagination or perhaps reasoning from some other facts that I remember about myself. In either of those cases, it seems that my observer memory will be distorted with regards to the content of my belief. For the relevant parts of the content of my memory (my having looked unshaven at the time, for instance) do not belong to the content of any of my perceptual experiences during the accident.14 Thus, it seems that my faculty of memory has not carried out its preservative function adequately while delivering that observer memory. But has it, nonetheless, produced a memory that is somehow beneficial for me to have? Let me introduce some terminology that may be useful at this point. By having a mental state of a certain type (such as a perception, a sensation or a memory), a subject may experience some emotions or moods. We could call these the ‘affective properties’ of the mental state. Also, by having a mental state of a certain type, the subject may have experiences that are qualitatively similar to the experiences produced by the subject’s senses during episodes of perception. Let us call these the ‘sensory properties’ of the mental state.15 Let us also use ‘feeling’ as an umbrella term that refers to both the affective and the sensory properties of a mental state. Furthermore, let us refer to a mental state that is rich in affective properties as ‘affectively rich,’ to a mental state that is rich in sensory properties as ‘sensorily rich,’ and to a mental state that is rich in both as ‘phenomenally rich.’ Let us also call mental states that are not sensorily, affectively or phenomenally rich, respectively, ‘sensorily dry,’ ‘affectively dry’ and ‘phenomenally dry.’ Now, there are reasons to think that, whereas field memories tend to be phenomenally rich, observer memories tend to be phenomenally dry. Consider, for example, the following findings in two classical studies on the field/observer distinction (Nigro & Neisser, 1983): Subjects who describe the contents of their field memories often mention their feelings at the time that they witnessed, or went through, the relevant events whereas subjects who describe the contents of their observer memories make significantly fewer references to their emotions and sensory experiences at the time.16 And, conversely, subjects who are trying to describe their feelings at the time that they witnessed, or went through, a remembered event tend to remember that event from the field perspective whereas subjects who are only trying to describe the circumstances surrounding the remembered event tend to remember it from the observer perspective. The narrower point that field memories are affectively richer than observer memories seems to be confirmed by more recent findings. It seems, for example, that a subject who remembers an event from her past is more likely to have a field memory of it when the remembered event has a strong emotional significance for her than when it does not (Talarico, LaBar, & Rubin, 2004). Taken together, these findings suggest that field memories, as opposed to observer memories, have very salient phenomenal properties for the subject. If field memories are phenomenally rich whereas observer memories are phenomenally dry, then one might wonder whether it is possible to change the phenomenology of remembering a past event by switching from remembering it from a field perspective to remembering it from an observer perspective. In particular, one might wonder whether one might be able to diminish, or dampen, the phenomenal properties of the memory by performing that switch. And, interestingly, there does seem to be some evidence suggesting that this effect is possible (Berntsen & Rubin, 2006). Now, if it is possible to change, and in fact diminish, the phenomenal properties of a memory of a past event by switching from remembering the event from the field perspective to remembering it from the observer perspective, then one can imagine a scenario in which it may be advantageous for a subject to perform that switch. This is the scenario in which the event constituted a traumatic experience for the subject in the past. It seems that traumatic events tend to be remembered, by default, from the field perspective (Porter & Birt, 2001). Assuming that field memories are phenomenally rich, a subject who remembers a traumatic event from the field perspective will presumably be forced to relive some of her emotions and sensory experiences during the event, which is likely to result in further trauma for her. It would seem, therefore, that such a subject would benefit from switching to remembering the traumatic event from the field perspective to remembering it from the observer perspective. For if remembering the traumatic event from the observer perspective does indeed dampen the phenomenal properties of remembering the event, then having an observer memory of the traumatic event should alleviate the suffering associated with reliving it in memory. It should allow the subject to achieve some ‘phenomenal distancing’ from the traumatic event.17 Thus, it seems that observer memories of past events may carry an adaptive type of benefit for the subject despite being distorted. Specifically, it seems that they may be affectively adaptive for her. For representing a past event from an observer point of view can, when that event has been traumatic for the subject, be an effective way of satisfying one of the subject’s goals. The goal in question is, in this case, to alleviate the suffering associated with reliving the event in memory. Admittedly, things are not quite that simple. Let us keep in mind that whether or not a memory of a subject is adaptively beneficial for her depends on which of the subject’s goals we are focusing on. A memory may help the subject to achieve one of her goals while not helping with, or perhaps even hindering her prospects of, achieving another. This may be the case with observer memories. Even though a case can be made that observer memories of trauma are adaptively beneficial with regards to the short-term goal of achieving some affective relief, they may not help the subject to achieve the long-term goal of maintaining a healthy self-concept. Picturing oneself from the outside, as it were, might be more conducive to subjecting oneself to evaluation and, thus, it might increase the risk that one finds certain aspects of how one is perceived not to be satisfactory enough.18 Thus, there may be a cost involved in adopting the observer perspective while remembering traumatic events. And yet, observer memories can be adaptively beneficial for the subject. The important point to bear in mind in order to accommodate both of these facts at the same time is that a subject does not draw, from her memories, adaptive benefits per se. Instead, the adaptive benefit of a subject’s memories must be relativised to each of the subject’s goals.","Over approximately the last twenty years, there has been a debate in cognitive science, sociology, psychiatry and the law on whether or not accounts of long-forgotten episodes of childhood trauma elicited by some memory recovery techniques, such as hypnosis and the use of sodium amytal, should be taken at face value. After undergoing treatment as part of certain approaches to psychotherapy, a subject may claim to remember a traumatic event that happened to her as a child, even though she was not able to remember it before her treatment. The issue in this debate has been whether such reports should be trusted as expressions of accurate memories or not.19 Those who believe that these reports should be trusted refer to the mental states being expressed through them as ‘recovered memories’ whereas those who believe that these reports should not be trusted as expressions of accurate memories refer to the mental states being expressed through them as ‘false memories.’ On the false memory camp, theorists such as Richard Ofshe have claimed that memories cannot be completely lost and, later, be recovered (Ofshe & Watters, 1994). In addition, false memory theorists have argued that inaccurate memories can easily be induced under experimental conditions (Loftus & Ketcham, 1994). Theorists on the recovered memory camp, by contrast, have disputed the contention that psychotherapists have the power required to implant fabricated memories of whole events (Harvey & Herman, 1994). Furthermore, they have appealed to independent evidence suggesting that the reactivation of traumatic experiences of other types, such as trauma during war, can occur after periods of time in which individuals experience relatively few symptoms (Schooler, 1994). In this section, I will assume that some of the reports of recovered memories of long- forgotten episodes of childhood trauma which arise during psychotherapy are not expressions of accurate memories. I will assume that they are reports of a kind of memory which, due to the psychotherapist’s intended or unintended acts of suggestion, the subject mistakenly takes to be an accurate memory of some (typically traumatic) past event. I will refer to these memories as ‘fabricated’ memories.20 In this section, I wish to put forward the claim that fabricated memories of traumatic events could, in extremely unusual circumstances, qualify as cases of beneficial memory distortion. In order to argue that fabricated memories of traumatic events can be beneficial despite being distorted one must, first of all, make the case that they are indeed distorted. From a preservative point of view, it should be uncontroversial that they are. Suppose that I have a memory of my childhood in which I represent an uncle who was visiting at the time as having molested me by touching me in a sexual way. It turns out, however, that the uncle in question never molested me. Thus, the source of that piece of information in the content of my memory must be other than my past perceptual experiences of him. Let us stipulate that this memory has arisen during psychotherapy and it is, as a matter of fact, a memory fabricated by me as a result of my therapist’s use of some techniques of suggestion. In that case, it seems that my fabricated memory is certainly distorted. For it fails to present my uncle to me in any way in which I apparently perceived him to be in the past. Thus, it seems that my faculty of memory has not carried out its preservative function adequately while delivering the memory that I am having. But has it, nonetheless, produced a memory that it could be beneficial for me to have? It is hard to imagine how any of the actual cases of fabricated memories of traumatic events and, especially, fabricated memories of abuse could possibly be beneficial for the subject. In actual cases of fabricated memories of abuse, the subjects involved are misled into thinking that they have been abused with terrible consequences. Not only can the subjects themselves be traumatized by those memories, but also their families can be torn apart and reputations can be destroyed by subsequent accusations of abuse. In actual fact, lives are often ruined by fabricated memories of abuse. Nevertheless, one can conceive some highly unlikely sets of circumstances in which, arguably, having a fabricated memory of a past episode of abuse could carry an adaptive benefit for the subject. My contention is that it is in fact possible to imagine two such sets of circumstances; a set of circumstances in which it is affectively adaptive for the subject to have such a memory, and a set of circumstances in which it is explanatorily adaptive for her. Let us consider the two scenarios in order. Let us imagine that, early in my childhood, I once witnessed my uncle giving a terrible beating to my mother; his sister. In fact, let us imagine that it was so early in my childhood that I am no longer able to recover that memory. Many years later, I invariably feel the desire to hate my uncle whenever I need to interact with him. Every time that I am in his presence, I realise that, quite simply, I want to hate the guy. This makes me ashamed of myself since, not being able to remember anything about the violent incident that I once witnessed, I cannot find anything particularly despicable about my uncle. And I strongly disapprove of the type of person who, as I see it, I would become by hating someone unwarrantedly. So I do not allow myself to experience hate towards my uncle. And yet, I wish that I was able to hate him. I cannot deny it. I am fully aware of my desire to hate him. And worse, I am fully aware that, in spite of the fact that all my efforts to find some justifying reason for it have failed, my desire to hate my uncle remains. The resilience of this desire is upsetting for me, so I find myself in a strange dilemma: On the one hand, I have a desire whose satisfaction would have very negative evaluative consequences for my own self-concept. On the other hand, I cannot get rid of it. I experience it as an intrusive desire; a desire that is beyond my rational control. One can picture how the whole situation would be deeply disturbing for me. Consider, now, the fabricated memory wherein I represent my uncle as molesting me while I was a child. Fabricating this memory would provide me with a reason which, in my view, entitles me to hate my uncle. And this, in turn, would allow me to experience hate towards him without any harm to my own self-concept. Thus, it seems that, in this scenario, my fabricated memory of abuse is beneficial despite being distorted. I benefit from having it in the sense that representing a past episode of abuse that never happened turns out to be an effective way of satisfying one of my goals. The goal in question is, in this case, to manage to occupy, consistently with my own set of values, an emotional state that I feel the need to experience. To that extent, my fabricated memory of abuse has an affectively adaptive benefit for me. We may also imagine a set of circumstances in which my fabricated memory of abuse is explanatorily beneficial for me to have. In order to describe it, one only needs to tweak some of the details in the conceivable scenario sketched above. Let us suppose that I did witness my uncle giving a terrible beating to my mother, and that I am no longer able to remember that event. Let us imagine, however, that I do not currently experience the desire to hate him. But I do find that I am inclined to behave negatively towards my uncle whenever I need to interact with him. I avoid giving him a hug or shaking his hand, I often find a reason to leave the room during a family reunion that involves him, I accidentally break his Christmas gifts, and so on. Let us suppose that I have insight into the fact that my behaviour reveals a dislike for him. However, not being able to remember the beating that I once witnessed, I cannot find a reason for that behaviour. I cannot explain why I am behaving in a hurtful way towards my uncle when I need to interact with him, which is puzzling for me. It is also upsetting, in that my disposition to behave in a hurtful way towards him remains despite all my failing efforts to find an explanation for it. That is, the fact that I cannot make that behaviour intelligible to myself has done nothing to change it. Thus, I feel alienated from some of my dispositions to action. I experience them as dispositions that are beyond my rational control. It is easy to picture how this conflict would be equally disturbing for me. Consider, now, the fabricated memory wherein I represent my uncle as molesting me while I was a child. Fabricating this memory would provide me with a reason which, in my view, explains why I am inclined to behave negatively towards my uncle. And this, in turn, would allow me to experience my relevant actions as actions that are rational: It would seem rational for me to perform those actions given that I can find a reason for performing them. Thus, it seems that, in this scenario, my fabricated memory of abuse is beneficial for me despite being distorted. Once again, I benefit from having it in the sense that representing a past episode of abuse that never happened turns out to be an effective way of satisfying one of my goals. The goal in question is, in this case, to make sense of some behavioural dispositions which I am unable to shake off. To that extent, my fabricated memory of abuse has an explanatorily adaptive benefit for me. Once again, though, things are not quite that simple, for reasons that will be reminiscent of our discussion of observer memories. Recall that a memory may help its subject to achieve one of her goals while not helping with, or perhaps even hindering her prospects of, achieving another. And, for that reason, the memory may carry an adaptive benefit for the subject with regards to the former goal, but not with regards to the latter one. This may be the case with fabricated memories. Even though it can be argued that fabricated memories of abuse could be beneficial with regards to the goal of allowing oneself to feel a certain emotion towards a person or event, or the goal of making sense of one’s own mental states and behaviour towards that person or event, they may not help the subject to achieve some of her other goals. As noted above, thinking of oneself as having been abused is likely to result in trauma. It is also likely to damage one’s social relations. (This is obvious when it comes to one’s relations with the person wrongly accused of being the abuser.) From the point of view of the goals of avoiding trauma and maintaining fulfilling social relations, therefore, it is not adaptively beneficial for a subject to fabricate a memory of having been abused. And yet, there are contexts in which it is possible for fabricated memories of abuse to be adaptively beneficial for the subject. Once again, there is no inconsistency here, provided that adaptive benefits are relativised to the subject’s goals.","Let us take stock. In Section 2, we have seen two pictures of what memory is supposed to do. On one of those pictures, memory is supposed to preserve the information that we acquired through perception in the past. On the other one, memory is supposed to build a narrative of our personal past. In Section 1, we have seen the benefits of memory performing each of those two functions appropriately while producing memories; an epistemic benefit and an adaptive benefit for the subject. However, in Sections 3 and 4, we have also seen that memory may fail to perform one of its functions adequately while producing memories which are, in some sense, beneficial for the subject to have. Specifically, we have seen that a subject’s faculty of memory can be unreliable while delivering some memories which have some value for the subject, either from an affectively adaptive point of view or from an explanatorily adaptive point of view. What does this possibility mean for the two conceptions of memory sketched in Section 2? There are two ways of looking at the relation between the storage conception of memory and the narrative conception of memory. If one takes what we may call an ‘exclusive’ approach to them, then one will believe that memory is either a faculty that is meant to perform a preservative function within our cognitive economy, or it is a faculty that is meant to perform a reconstructive function within it; but not both. On this approach, then, either the function of memory is to preserve the information that we acquired in the past through perception, or the function of memory is to build a narrative of our personal past. In the former case, the narrative conception of memory is wrong whereas, in the latter case, it is the storage conception of memory that is wrong. Either way, both of them cannot be right. If one takes an exclusive approach towards the relation between the storage and narrative conceptions of memory, and one endorses the narrative conception of memory, then observer memories of trauma and fabricated memories of abuse do not qualify as cases of beneficial distortion after all. Specifically, if the narrative conception of memory is correct and the storage conception is wrong, then neither observer memories of trauma nor fabricated memories of abuse are distorted. For if the considerations offered in Section 1 are correct, then those memories must have been, in a certain sense, appropriately produced in order for them to be adaptively beneficial for the subject. The relevant sense is that memory must have carried out its reconstructive function appropriately while delivering them. Observer memories of trauma, for example, cannot help the subject to achieve some phenomenal distancing from the remembered traumatic event if they do not cohere well with the rest of things that the subject remembers about the event, and the things that she knows about her own physical appearance in the past. Suppose, for example, that my observer memory of my traffic accident does not represent me as having the physical traits that I believe I had at the time of the accident. Suppose, furthermore, that my observer memory does not represent my car as having the colour, shape and size that I believe it had at the time of the accident. Then, I will not be able to identify myself as the person who suffered the accident by having that observer memory. It is difficult to see, then, how that memory could allow me to stop remembering the traffic accident from a field perspective and start remembering it in a more phenomenally detached way. After all, if I cannot identify myself as the person who suffered the accident, then why would I recognise the mental state that I am occupying as a memory of something that happened to me at all? Similarly, fabricated memories of abuse cannot help the subject to achieve some emotional state that she is seeking to experience, or some understanding of her own current behaviour and mental life, if they do not cohere well with the rest of things that the subject remembers about the circumstances surrounding the alleged episode of abuse, and the things that she knows about the participants in that episode. Suppose, for example, that I have a fabricated memory of abuse involving my uncle, but it is not consistent with some of the things I know about what was going on at the time. Suppose that I know that my uncle was not in town when, according to my fabricated memory, I suffered his sexual abuse. Suppose that I also know that the house where he is supposed to have visited us did not look at all like my memory is presenting it to me. Let us say that my memory does not even represent my uncle as I believe he looked like at the time. It would be surprising if, given these inconsistencies, I still proceeded to trust my fabricated memory as a memory of a genuine event in my past. In fact, I might even be able to suspect that the memory that I am having has been fabricated by me. It is difficult to see, then, how that memory could give me a justifying reason for allowing myself to experience hate towards my uncle, or it could provide me with an explanatory reason of my behaviour towards my uncle. Thus, if one takes an exclusive approach towards the relation between the storage and narrative conceptions of memory, and one assumes that memory can only have a reconstructive function, then one must conclude that memory has performed its function properly while delivering those observer memories of trauma and fabricated memories of abuse which are adaptively beneficial for the subject to have. Otherwise, they could not be adaptive in the first place. Admittedly, this conclusion allows the narrative theorist to capture the intuition that there is something right, and not distorted, about the way in which those memories have been generated. And this is indeed an intuition worth capturing. We do feel its pull. Unfortunately, though, there is a significant cost to adopting this position. As we saw in Section 1, if the fact that memory is carrying out its function appropriately when it delivers a memory makes no difference as to whether that memory is likely to be correct, then one cannot be justified in forming beliefs about one’s personal past on the basis of one’s memories. And, for that reason, one’s memories cannot provide one with knowledge of one’s personal past. Now, if one believes that the function of memory is exclusively reconstructive, then it seems that this is precisely the conclusion that one should draw. For the fact that memory has carried out its reconstructive function appropriately while producing a memory is no indicator of whether that memory is likely to be accurate or not. (As a matter of fact, the types of memories discussed in Sections 3 and 4 illustrate this point.) Thus, an exclusive approach to the relation between the storage and narrative conceptions of memory, combined with an endorsement of the latter, leads us to the conclusion that memory cannot provide us with knowledge of our personal past. This seems too high a price to pay for preserving the view that those memories which are adaptively beneficial for the subject are, in some intuitive sense, not distorted. Things are not better if one adopts an exclusive approach towards the relation between the storage and narrative conceptions of memory, but one endorses the storage conception instead. In that case, one can capture the intuition that there is something distorted about the way in which observer memories and fabricated memories are generated since, as we saw in Sections 3 and 4, they are indeed distorted from a preservative point of view. But there is a significant cost to adopting this position as well. For the fact that some of those memories can be, under certain circumstances, beneficial for the subject becomes, then, a mystery. Let me explain. The storage theorist who takes an exclusive approach towards the two conceptions of memory is committed to the view that the faculty of memory never carries out its function appropriately when it produces observer memories and fabricated memories. However, if memory is not carrying out its function adequately when it produces observer memories and fabricated memories, then it is hard to understand why some of those memories can actually do some good for the subject. After all, in Sections 3 and 4, we have seen that the reason why some of those memories can be beneficial for the subject is that they serve a certain purpose for the subject. They are aimed at providing something for the subject; something that the subject is in need of. (The aim in question may involve either an emotion or an explanation.) But if memory is never doing what it is supposed to do when it generates those memories, then it is hard to see why, in some cases, the generation of those memories happens to serve a purpose for the subject. What explains the fact that those memories are meant to achieve a certain goal, a goal that it is in the subject’s interest to achieve, if they have been accidentally generated? An alternative approach to the relation between the storage and narrative conceptions of memory is what we may call an ‘inclusive’ approach to them. According to it, memory is a faculty that is meant to perform a preservative function within our cognitive economy, and it is also a faculty that is meant to perform a reconstructive function within it. On this approach, then, the function of memory is to preserve the information that we acquired in the past through perception, and the function of memory is to build a smooth narrative of our personal past as well. What is the relevance of the considerations offered in Sections 3 and 4 for this approach? If one takes an inclusive approach towards the relation between the storage and narrative conceptions of memory, then one can capture two important intuitions about beneficial observer memories and beneficial fabricated memories which have been highlighted above; the intuition that there is something wrong, and the intuition that there is something right, about the way in which those memories have been generated. On the one hand, there is something wrong in that memory has not performed its preservative function adequately while delivering those memories. This is why we are inclined to think that there is a sense in which they are distorted. Capturing this intuition by accepting that there is a preservative function of memory which, in those cases, has not been carried out appropriately allows us to hang on to the idea that, in order for our memories to yield knowledge of our personal past, memory must carry out its function appropriately while delivering them. For this is indeed true of the preservative function of memory. On the other hand, there is something right about the way in which observer memories of trauma and fabricated memories of abuse have been generated when those memories are beneficial for the subject. For if the inclusive approach is correct, then memory has a reconstructive function as well. And it seems that memory has performed that function adequately while delivering those memories. After all, the considerations above suggest that, unless memory maintains a certain coherence within the subject’s mental states when it delivers observer memories of trauma and fabricated memories of abuse, those memories cannot be adaptively beneficial for the subject. Assuming that there can be, as argued in Sections 3 and 4, beneficial cases of such memories, it seems that we must accept that memory has performed a certain function appropriately while delivering those memories, namely, a reconstructive function. The outcome of these considerations, therefore, seems to be that the correct approach to take towards the storage and narrative conceptions of memory is the inclusive approach. The view that the function of memory is both to preserve the information that we acquired in the past through perception, and to build a smooth narrative of our personal past, is not new. It resonates, for example, with the so-called ‘Self-Memory System (SMS)’ conceptual framework (Conway, 2005; Conway & Pleydell-Pearce, 2000; Conway, Singer, & Tagini, 2004). One of the central claims within the SMS framework is that memory must negotiate two demands; that of accurately recording ongoing activity (‘correspondence’ in the SMS terminology) and that of maintaining a coherent record of the self’s past activity (‘coherence’ in the SMS terminology). The idea is that a healthy faculty of memory will meet those demands in an appropriately calibrated way. Now, that central idea in SMS is similar to, but different from, the view that has been offered here. For the reasons why, according to the view offered here, the preservative and reconstructive functions of memory are important are different from the reasons why, within the SMS framework, it is important for our memories to meet the demands of correspondence and coherence. I have argued that building a smooth narrative of one’s past is necessary for the purposes of experiencing a certain emotion towards some events in one’s past, and for the purposes of making sense of one’s attitudes towards that event. By contrast, the reason why, within the SMS framework, it is important for our memories to meet the demand of coherence is that our memories must sustain an enduring sense of self. Otherwise, our versions of our past selves will become detached from reality.21 The difference is that, whereas it is necessary for one to achieve a certain emotion towards an event in one’s past (or for one to make sense of one’s attitudes towards that event, for that matter) that one has a stable sense of self, a stable sense of self does not seem to be sufficient for one to achieve those emotional or explanatory goals. After all, we would not want to claim that all subjects who have not emotionally processed, or have not achieved some emotional closure with respect to, some traumatic event in their past no longer have a consistent sense of self.22 I have also argued that possessing a faculty of memory that reliably preserves the information that we acquired in the past through perception is necessary for the memories that such a faculty produces to afford knowledge of our personal past. By contrast, the reason why, within the SMS framework, it is important for our memories to meet the demand of correspondence is that our memories must keep track of where we are in the process of achieving a certain goal. Otherwise, dysfunctional repetitions of action sequences will ensue, since we will not accurately remember having already performed the necessary actions to achieve some of our goals. The difference is that, whereas it is correct that if a memory provides a subject with knowledge of her past, then it will allow her to keep track of the fact that she has just performed an action that needs to be performed in order to achieve one of her goals, the converse is not the case. All the subject needs in order to keep track of the fact that she has just performed an action which she is required to perform in order to achieve one of her goals is the true belief that she has just done so. And, as we saw in Section 1, a true belief which allows us to achieve one of our goals does not need to be justified and, for that reason, it does not need to amount to knowledge.23 What lesson can be drawn, then, from our discussion in this section? If the inclusive approach to the functions of memory is the correct approach to take, then it seems that we can draw an interesting lesson from the fact that there is such a thing as beneficial memory distortion. Instances of beneficial memory distortion teach us that memory has various functions, and they teach us that the adequate performance of each of those functions can, conceivably, come apart from each other. Memory distortion, in other words, reveals that memory is supposed to do various things. At the very least, it is supposed to preserve the information that we acquired through perception in the past, which is why instances of beneficial memory distortion are instances of distortion. And it is supposed to provide us with a narrative of our personal past, which is why instances of beneficial memory distortion are beneficial. Furthermore, cases of beneficial memory distortion illustrate the fact that memory can, in principle, do the latter without doing the former. Ultimately, then, what cases of beneficial memory distortion teach us about the nature of memory is that memory performing its reconstructive function does not necessarily depend on memory performing its preservative function. There is no logical or conceptual link that ties our notions of those two functions together. In that sense, our capacity to reconstruct our personal past in memory is different from our capacity to acquire knowledge of it through memory.24"],["Intelligence, as measured by standardised tests of cognitive function, such as IQ-type tests, is predictive of psychiatric diagnosis and psychological wellbeing. Using genome-wide association study (GWAS) data, a measure of the shared genetic effect across traits, can be quantified; because this can be done across samples, the confounding effects of psychiatric diagnosis do not influence the magnitude of these relationships. It is now known that there are genetic effects that act across intelligence and psychiatric diagnoses, which provide a partial explanation for the phenotypic link between intelligence and mental health. Potential causal effects between intelligence and mental health have been identified, and the regions of the genome responsible for some of these cross trait associations have begun to be characterised. --------------------------------------------------------------------------------","Intelligence, sometimes called general cognitive ability/function, the g factor, or simply g, describes the finding that scores on cognitive tests that each seem to tax disparate aspects of mental ability, positively correlate [1]. This overlap accounts for around 40% of the variance found when administering a broad array of cognitive tests to a group with a range of ability [2,3]. This finding has been known for over a century, and has been replicated in hundreds of data sets [1–3]. Individual differences in intelligence are predictive of mental illness, where a higher level of intelligence in childhood is predictive of a lower level of self-reported psychological distress decades later [4]. This link between intelligence and mental illness also extends to severe psychiatric conditions where individuals who have a level of intelligence one standard deviation below the mean have, on average, a 60% greater chance of being hospitalized for schizophrenia, a 50% increase of being diagnosed with a mood disorder, and a 75% greater risk for having an alcohol-related disorder, over a two-decade follow up period [5]. A higher risk for several psychiatric illnesses has also been associated with a lower level of intelligence, including major depressive disorder (MDD) [6,7], autistic spectrum disorder (ASD), attention/deficit hyperactivity disorder (ADHD) [8], as well as bipolar disorder [9,10], although a higher level of intelligence, particularly as measured by tests of crystallized ability, may also be a risk factor for bipolar disorder [11].","Intelligence, like many other quantitative traits [12•,13], is heritable with twin and family derived estimates of heritability being around 50–80% [14], with genetic factors explaining an increasing proportion of variance as the age of the sample increases [15]. Molecular genetic data can also be used to derive the proportion of phenotypic variation explained by all genome-wide single nucleotide polymorphisms (SNPs), using genomic- relatedness-based restricted maximum-likelihood single component (GREML-SC) [16], implemented in GCTA [17]. Heritability estimates derived using GREML-SC describe the proportion of phenotypic variance that is explained by genetic variants in linkage- disequilibrium (LD), that is to say correlated, with genotyped SNPs. As SNP arrays typically measure common genetic variation, and two events can only be highly correlated if they occur with a similar frequency, GREML-SC estimates of heritability represent a subset of the total heritability in phenotypes where the genetic architecture includes contributions from low frequency variants, and other types of genetic variation that are poorly correlated with common SNPs. Heritability estimates derived using GREML-SC applied to intelligence show that around 22.7% (SE = 2.1%) of phenotypic variance is explained by additive genetic effects that are linked to common SNPs [18••]. Psychiatric disorders have also been shown to be heritable using GREML-SC, where additive common genetic effects explain 23% (SE = 0.8%) of schizophrenia [19], 21% (SE = 2.1%) of MDD [19], 28% (SE = 2.3%) of ADHD [19], 25% (SE = 1.2%) of bipolar disorder [19], 17% (SE = 2.5%) of ASD [19], and 10.8% (SE = 2.0%) of neuroticism [18••], an individual’s propensity to experience psychological distress. A method to capture genetic effects from across the frequency spectrum of causal variants, called GREML-KIN [20], has been applied to intelligence, neuroticism, and MDD. For intelligence 54.0% of phenotypic variation can explained using genome-wide association (GWAS) data [18••]. For neuroticism, 30% of phenotypic variation is captured, with 47.0% of MDD [21] being explained by additive genetic effects when common and rare genetic effects are summed. GWAS for intelligence have recently attained the statistical power required to reliably identify the loci that contribute to these heritability estimates with more than 200 being identified so far [22••,23••,24] (15 novel loci identified in Sniekers et al. 130 novel loci identified in Hill et al., and 58 novel loci identified in Davies et al.). However, the total proportion of variance these loci explain is far lower than the substantial heritability estimates. This `missing heritability’ [25] is also seen when examining psychiatric variables and is indicative of a highly polygenic architecture, where the cumulative effect of all genetic effects may be substantial, but the contribution made by any individual variant is negligible. This substantial heritability, combined with the relative sparsity of loci identified at current sample sizes, is compelling evidence that by increasing sample size, and with it the ability to reliably estimate small effects, will result in an increase in the number of loci identified for both intelligence and psychiatric variables. Figure 1a shows the Manhattan plot from one of the first well powered GWAS on intelligence [22••].","The polygenic signal that drives heritability estimates can also be used to derive genetic correlations to describe the average genetic effect that is attributable to causal variants in LD with common SNPs, and shared across two traits. Using a technique called linkage disequilibrium score (LDSC) regression [26,27••] genetic correlations between two GWAS data sets can be derived. Whilst LDSC regression is less precise than GREML, as indicated by the higher SE even when sample sizes are similar, as well as in instances where there is genetic heterogeneity between the reference panel used to derive LD scores and the sample used to derive the genetic correlations [28], LDSC regression has the advantage that the GWAS data can come from separate samples where individual level data is unavailable. Although an overlap in controls is not uncommon, by performing genetic correlations across data sets neither the symptoms of psychiatric diagnosis, hospitalisation, or drug regimens, can confound the measure of the genetic relationship between intelligence and psychiatric illness. When applied to GWAS on intelligence and psychiatric disorders, both positive and negative genetic correlations are found. Anorexia nervosa, for example, shows a small but statistically significant positive genetic correlation with intelligence (rg = 0.06, SE = 0.03, P = 0.02), as does ASD (rg = 0.21, SE = 0.04, P = 2.46 × 10−8) [22••]. Negative genetic correlations however, are found between intelligence and schizophrenia (rg = −0.14, SE = 0.03, P = 1.49 × 10−9), ADHD (rg = −0.46, SE = 0.03, P = 2.41 × 10−54), neuroticism (rg = −0.29, SE = 0.07, P = 7.01 × 10−6) [22••], and more recently with MDD (rg = −0.30, SE = 0.04, P = 1.28 × 10−13) [23••]. Together this indicates that the genetic variants associated with high levels of intelligence have both protective and facilitative effects on the genetic liability of mental illness, with the genetic variants associated with a decrease in neuroticism, MDD, and schizophrenia being, on average, those linked to higher levels of intelligence, and genetic variants that confer greater risk of ASD, and of anorexia nervosa also being linked to higher levels of intelligence. Bipolar disorder, however, shows a genetic correlation of around zero with intelligence [22••,29••]. Educational attainment (as measured by the number of years in education or by whether an individual attained a University or college level degree) shows a strong genetic correlation with intelligence (rg = 0.70, SE = 0.02, P = 1.28 × 10−285) [22••] and has been used as a proxy phenotype for intelligence [30]. However, in contrast with intelligence the genetic architecture of education shows a positive genetic correlation with schizophrenia (rg = 0.10, SE = 0.02, P = 5.40 × 10−6) [22••] and with bipolar disorder (rg = 0.28, SE = 0.04, P = 4.84 × 10−14) [22••]. This indicates that, on average, the genetic variants associated with an increase in educational attainment are also linked to an increase in the risk of both schizophrenia, and bipolar disorder, despite that the genetic variants associated with an increase in intelligence are also linked to a reduction in the genetic risk for schizophrenia and are not linked to bipolar disorder. This difference between how the genetic aetiology of intelligence and education overlap with schizophrenia and bipolar disorder, can serve as a diagnostic tool to gauge whether a phenotype constructed to measure intelligence, is in fact a better measure of educational attainment [31]. Figure 1b shows the genetic correlations between intelligence, as well as education, with six psychiatric disorders and neuroticism. Whereas the large genetic correlation between intelligence in childhood, and intelligence in older age (rg = 0.71, SE = 0.10, P = 2.26 × 10−12) suggests that many of the same variants are involved in intelligence across the lifespan, the overlap between intelligence and psychiatric variables may be influenced by the age at which intelligence was assessed [29••]. This may be due to the genetic contributions to intelligence, as measured in childhood, being a product of genetic effects involved in the development of intelligence, whereas intelligence in older age will also be a product of the genetic effects involved in the maintenance of intelligence across the life course [29••].","Genetic correlations tell us about the average genetic effect that is shared across traits. As such, they do not identify the variants involved in cross-trait associations, nor do they imply that when such a variant is identified it will have a shared effect on the two traits consistent with the direction of effect of the genetic correlation. One way to identify loci with a shared effect is to examine SNPs that attain genome-wide significance in intelligence, and in psychiatric disorders. However, this method is dependent on the number of genome-wide significant SNPs. In contrast to examining if a SNP is genome-wide significant in two traits, conjunctional false discovery rate [32] (cFDR) can be used to determine if a SNP shows association with two traits simultaneously. cFDR has been used to examine the genetic link between schizophrenia and intelligence by identifying loci that harbour joint effects. A total of 21 loci were identified as acting across schizophrenia and three measures of cognitive ability (two measures of intelligence, and one measure of reaction time). The genes that were implicated using the cFDR approach were expressed across the developing and adult brain consistent with the strong genetic correlations between childhood intelligence and older age intelligence. Again consistent with the genetic correlations between intelligence and schizophrenia, 18 of the 21 loci identified contained risk alleles that were facilitative of intelligence and protective against schizophrenia. The remaining three loci harboured variants that were associated with a higher level of intelligence and a greater risk of schizophrenia, consistent with some reports finding that a number of those diagnosed with schizophrenia have retained their level of intelligence [33].","GWAS exploits the correlation between SNPs (whether genotyped or imputed) and unknown causal genetic variants. By doing so, regions of the genome, defined by the correlation between SNPs, are identified as being linked to potentially causal variants. The presence of genetic correlations, loci associated across traits, and even the same SNP being implicated in multiple traits, can therefore arise in a number of different conditions. Firstly, biological pleiotropy may be in effect and can be the result of a single SNP, that is genome-wide significant in two traits, tagging a single variant that is causal in both of the studied phenotypes [34•]. Alternatively, biological pleiotropy can describe a situation where a single SNP, again genome-wide significant in two GWAS, tags two causal variants (each casual for a different phenotype) that are both in the same gene. Biological pleiotropy can also describe a situation where two SNPs, both within the same gene, each tag independent causal variants that are also located within the same gene. Secondly, mediated pleiotropy can occur where one phenotype is a causal element in a second phenotype. In such instances genetic variants identified for the causal trait will also be associated with the second. For example, mediated pleiotropy is likely to explain the genetic link between intelligence with education and other measures of socio-economic status [22••,35]. Finally, spurious pleiotropy can occur where there is misclassification of the phenotype, this can occur if the low mood observed by those with bipolar disorder is misclassified as MDD or in cases where those with bipolar disorder are misdiagnosed as having schizophrenia. In these instances, a genetic correlation will exist between these traits that is due to contamination of samples, rather than genetic effects that act across traits. Spurious pleiotropy can also occur in cases where a single variant is found to be associated with two traits, but this variant is tagging two, independent causal variants, each of which is found in a different gene. This can occur in regions of the genome where there is strong linkage disequilibrium. As genetic correlations are based on all SNPs within a data set, different forms of pleiotropy may be in effect. Whereas genetic correlations and shared risk loci are informative as to the average genetic effect across traits, as well as the regions of the genome where such effects are localised, they are not informative as to causality. Mendelian Randomisation (MR) can be used to mimic a randomised control trial using observational (GWAS) data under a number of assumptions [36,37]. MR typically uses SNPs that have attained genome-wide significance for a trait, such as intelligence, as proxy variables for the trait. Independent groups can then be made by grouping participants according to genotype. Groups that are created in this way will also be grouped according to the phenotype, in this case intelligence, that is linked to the genotype. As genetic variants are assigned randomly to a child at conception, MR can be seen as a randomised control trial where intelligence, is randomly assigned to a participant at conception. Using MR, causal links have been suggested for the link between intelligence and education [38••], and this relationship appears to be bidirectional with education being a casual factor in intelligence differences. A bi-directional causal relationship has also been seen for schizophrenia where higher levels of intelligence appear to be causally linked to a risk of schizophrenia [38••]. Lower levels of intelligence also appear to exert a causal risk on ADHD [38••], and consistent with the direction of the genetic correlation, and a higher level of intelligence is a causal risk for ASD [38••]. To conclude, intelligence and psychiatric illnesses are highly polygenic traits where common and rare genetic effects appear to be explain a substantial proportion of individual differences. It is the common genetic effects that are known to act on across intelligence and psychiatric disorders and this genetic link between intelligence and psychiatric condition is not confounded by the presence of psychiatric disease, hospitalisation, or treatment regime. The direction of effect is that high intelligence is protective against schizophrenia, MDD, ADHD, and neuroticism, but is a risk factor for ASD, and anorexia nervosa. Loci of shared effect have been identified for intelligence and schizophrenia implicating brain expressed genes that are expressed in developing and adult brain.","IJD and SHE are supported by the University of Edinburgh Centre for Cognitive Ageing and Cognitive Epidemiology which is funded by the UK Medical Research Council and Biotechnology and Biological Sciences Research Council (Grant No. MR/K026992/1). WDH is supported by Age UK (Disconnected Mind Programme)."],["The associations between higher intelligence test scores from early life and later good health, fewer illnesses, and longer life are recent discoveries. Researchers are mapping the extent of these associations and trying to understanding them. Part of the intelligence-health association has genetic origins. Recent advances in molecular genetic technology and statistical analyses have revealed that: intelligence and many health outcomes are highly polygenic; and that modest but widespread genetic correlations exist between intelligence and health, illness and mortality. Causal accounts of intelligence-health associations are still poorly understood. The contribution of education and socio-economic status — both of which are partly genetic in origin — to the intelligence-health associations are being explored. --------------------------------------------------------------------------------","Until recently, an article on DNA-variant commonalities between intelligence and health would have been science fiction. Thirty years ago, we did not know that intelligence test scores were a predictor of mortality. Fifteen years ago, there were no genome-wide association studies. It was less than five years ago that the first molecular genetic correlations were performed between intelligence and health outcomes. These former blanks have been filled in; however, the fast progress and accumulation of findings in the field of genetic cognitive epidemiology have raised more questions. Individual differences in intelligence, as tested by psychometric tests, are quite stable from later childhood through adulthood to older age [1,2]. The diverse cognitive test scores that are used to test mental capabilities form a multi-level hierarchy [1–3]; about 40% or more of the overall variance is captured by a general cognitive factor with which all tests are correlated, and smaller amounts of variance are found in more specific cognitive domains (reasoning, memory, speed, verbal, and so forth). Twin, family and adoption studies indicated that there was moderate to high heritability of general cognitive ability in adulthood (from about 50–70%), with a lower heritability in childhood [4]. It has long been known that intelligence is a predictor of educational attainments and occupational position and success [1]. Relatively recently, the ‘ultimate validity’ of intelligence test scores was discovered, that is, that higher intelligence significantly predicts later death. First, an Australian Vietnam Veterans study found that higher young-adult intelligence predicted lower risk of accidental deaths up to early middle age [5]. Then, a population-representative Scottish study found that intelligence test scores at age 11 years predicted deaths from all causes up to older age (the mid-70 s) [6]. The association between intelligence test scores from early life and mortality from all causes has been widely replicated [7–9]. Intelligence from childhood and adulthood is associated with most of the major causes of death with the exception of non-smoking-related cancers [10••,11]. Broadly speaking, a one-standard-deviation advantage in intelligence in youth lowers the risk of mortality by 20–25% or more up to older age; the effect sizes are hardly attenuated at all by adjusting for childhood socio-economic status, though are partly attenuated after adjusting for education and adult socio-economic status, which are possible mediators of the association [6,7,8,9,10••,11]. In addition to mortality, intelligence test scores are associated with lower risk of many morbidities, such as cardiovascular disease, cerebrovascular disease, hypertension, cancers such as lung cancer, stroke, and many others, as obtained by self-report and objective assessment [12–14]. Higher intelligence in youth is associated at age 24 with fewer hospital admissions, lower general medical practitioner costs, lower hospital costs, and less use of medical services, and intelligence appeared to account for the associations between education and such health outcomes [15,16]. Higher intelligence is related to a higher likelihood of engaging in healthier behaviours, such as not smoking, quitting smoking, not binge drinking, having a more normal body mass index and avoiding obesity, taking more exercise, and eating a healthier diet [16–18]. The flood of intelligence versus mortality/illness/health-behaviours findings was captured by the term ‘cognitive epidemiology’ [19]. From early on until now, there have been speculations about the possible causes of these associations [6,10••,14,20]. Briefly, there is acknowledgement that the causes of the associations are probably multiple, such as there being a constitutional (perhaps partly genetic) association between intelligence and health, and/or that intelligence’s influence might act via more education, higher health literacy, and more affluent social class. Here, we examine evidence for possible genetic links between intelligence and health.","There are at least three reasons to conduct genetic studies of phenotypes. First one wants to understand the genetic architecture of a phenotype, that is, what is the nature of the genetic variants that contribute to variation in the phenotype. For example, a single mutation might have a large effect, as is the case in Mendelian diseases. By contrast, continuous traits might be more likely to be polygenic; that is, to have some of their variance caused by small contributions from many genetic variants. Second, having discovered the genetic architecture, one is interested in the specific genes in which variants have causal effects, that is, one wants to understand the molecular genetic mechanisms of variation. Third, knowing that there is some genetic contribution to a phenotype, one can ask how good a predictor the genotypic information is; that is, how well can one predict some variation in a phenotype from only genotypic information? Much recent progress has been made along these lines for illnesses and for intelligence. Before the mid-2000s, genetic studies were done by three main methods. First, pedigree-based (twins, adoptees, and families) studies of relatives’ phenotypic associations were used to estimate the heritability of phenotypes, and genetic correlations among them. Limitations of pedigree methods include the fact that several assumptions must be made in doing the modelling, and that one does not learn about the specific genes involved. Second, candidate gene studies tested hypotheses concerning whether certain genetic variants were associated with phenotypic differences. For example, the possession of the e4 allele of the gene for Apolipoprotein E (APOE) is associated with an increased risk of developing Alzheimer’s disease. Limitations of the candidate gene method include the fact that most candidate gene findings are not replicated (APOE e4 possession is an exception to this), and that it is difficult to choose a candidate genetic variant from the millions that are known. Third, genetic linkage analysis was used to track genetic markers in families where specific phenotypes were common, to identify regions of the genome that segregate with the phenotype. The main limitations of this method are that large families are required and it identifies relatively large regions of the genome, rather than specific genetic variants or genes. This changed with the advent and rise of genome-wide association studies (GWASs) [21] (See Box 1). Sample sizes for GWASs often began with a few thousand, but, as the polygenic architecture of many traits became clear — , that is, the associations between individual genetic variants and phenotypes typically had very small effect sizes — it was necessary to form consortia so that the Ns of studies rose to the tens and then hundreds of thousands. Some GWAS consortia are now approaching and passing one million participants. The typical finding — there are exceptions — in health and cognitive GWASs is that many genetic variants of small effect contribute to phenotypic variation. In 2017, a survey of the first ten years of GWASs’ discoveries enumerated the SNPs that were associated with, for example, Crohn’s disease, diabetes, blood lipid levels, heart function, height, bone density, red blood cell traits, metabolic traits, blood platelets, breast cancer, rheumatoid arthritis, blood metabolites, menarche, Alzheimer disease, kidney function, lung function, and education [21]. Often, the numbers of genetic loci in which significant SNP associations are found runs to dozens or even hundreds for a single phenotype. In 2011 the first apparently-decently-sized GWAS of intelligence appeared (N approximately 3500), and found no significant SNPs [22]. By the time the sample size was about 100 times greater, the number of independent genomic regions that were associated with intelligence was about or greater than 150 [23•,24•,25••]. Figure 1 shows results from a recent GWAS of intelligence. Many of these SNPs are located in regions of the genome that have previously been associated with physical and mental illnesses. Therefore, we now know many actual DNA variants that have significant associations with intelligence tests’ scores; there are probably thousands in total. Although it found no significant SNPs, the 2011 paper [22] did make a difference; it was the first study to estimate the heritability of intelligence from DNA data alone and in unrelated subjects. This used a then-new method—called GREML, and run in the GCTA framework [26] — which examined people’s overall genetic similarity — based on common SNPs — with their phenotypic similarity (See Box 1). The common-SNP-based heritability of intelligence is estimated to be about 25% [24•]. It is typical for this common-SNP-based heritability to be about half of that estimated from twin studies [27•]. It is thought that this ‘missing heritability’ is because there are types of genetic variants other than causal variants that are in linkage disequilibrium with common SNPs that contribute to heritability. Some new techniques are helping to find these and close the gap between twin-based and SNP-based heritability [28].","Three things are clear. First, higher intelligence in early life is a significant predictor of better health behaviours, fewer and later illnesses, and longer life. Second, many of the relevant health and illness outcomes, as well as health behaviours, have many SNPs associated with them, and have a detectable level of common-SNP-based heritability. Third: the same goes, genetically, for intelligence. Relatively new methods — bivariate extension of GREML run on GCTA [29], and LD regression [30,31] (see Box 1) — have allowed estimates of the genetic correlations between phenotypes. That is, we can test the extent to which the polygenic signature obtained by using the summary results from GWAS contributes to any two phenotypes, including between intelligence and health. Polygenic signatures for many diseases were soon shown to be associated with intelligence [32]. Twin studies had suggested that part of the intelligence-mortality association might be genetic in origin, though there was disagreement about how much genetics contributed [33,34]. However, more recent studies have used genomic data. The list of significant molecular genetic correlations between intelligence and physical health variables is now long [23•,24•,25••,32]. Table 1 gives some examples. With regard to mortality, longevity has been used; parental age at death has also been used, as a proxy, because most relevant studies have not carried on long enough for many participants to have died. There is a positive correlation of 0.36 between intelligence and parental age at death. There are inverse genetic correlations between intelligence and both heart disease and hypertension, with effect sizes between −0.1 and −0.2. There is a small (<0.1) association with cholesterol, with higher ‘good’ cholesterol going with higher intelligence and the reverse for the ‘bad’ cholesterol. There is a moderate-sized inverse genetic association between intelligence and Alzheimer’s disease. There is a positive genetic association, of 0.27, with intracranial volume, which is an indication of maximal brain volume in the life course. There are significant positive genetic correlations between intelligence and birth weight, lung function, happiness, and short-sightedness. There are significant negative genetic correlations between intelligence and body mass index, poor self-rated health, lung cancer, osteoarthritis, insomnia, smoking, waist-hip ratio, and long-sightedness. It must be stressed that these correlations are based on GWASs conducted on different samples; that is the people on whom intelligence was measured were not the people on whom the health-based phenotype was assessed. Associations are interesting, but they do not explain why the correlations exist, or the direction of causation, which require further study and more new GWAS-based methods. UNDERSTANDING THE INTELLIGENCE VERSUS PHYSICAL HEALTH ASSOCIATION, INCLUDING THE PART PLAYED BY GENETICS -------------------------------------------------------------------------------- As described above, genetic correlations have been identified between intelligence and many diseases, and physical health traits; moreover, polygenic scores for diseases and health traits predict intelligence. However, it is not clear if these findings are due to: (1) genetic variants influencing health traits/diseases, and then those health traits/diseases influencing intelligence; (2) genetic variants influencing intelligence, and then intelligence influencing health traits/diseases; or (3) genetic variants influencing general bodily system integrity [20] that influences both intelligence and health traits/diseases. (1) and (2) may be due to mediated pleiotropy which can be tested for using a relatively new technique called Mendelian Randomization (MR) (see Box 1). Using a bi-directional two-sample MR approach we identified no causal association between intelligence or educational attainment (a proxy measure of intelligence), and the physical health traits of body mass index (BMI), systolic blood pressure, height, coronary artery disease and type 2 diabetes [35], using data from the UK Biobank (N approximately 110,000) and large GWAS consortia. However, a larger, more-recent study found MR-based evidence for potentially causal genetic effects of intelligence on larger intracranial volume, lower risk of Alzheimer’s disease, lower body mass index, and greater likelihood of quitting smoking [25••]. A MR study investigating the effect of education on obesity in about 2000 Finns concluded that education could be a protective factor against obesity, as measured using BMI [36]. Another study using education data from the SSGAC consortium and coronary heart disease data from CARDIoGRAMplusC4D (total sample size = 543,733) found that higher education was causally associated with reduced risk of coronary heart disease, lower likelihood of smoking, lower BMI and a more favourable blood lipid profile [37]. Sensitivity tests indicated that the results were unlikely to be driven by biological pleiotropy. A two-step MR study investigated the influence of vitamin B12 intake during pregnancy on cord blood DNA methylation and whether there is a causal influence on offspring’s cognition in the Avon Longitudinal Study of Parents and Children (ALSPAC) [38]. A small causal effect of vitamin B12-responsive DNA methylation changes on children’s cognition was identified. MR analysis has suggested that genetically-predicted intelligence and education both had associations with Alzheimer’s disease [39]. Another part of understanding the genetic contribution to intelligence-health correlations concerns other predictors of health inequalities, and intelligence’s correlations with them. Intelligence, we saw earlier, is related to education and socio-economic status (SES), and those were known to be related to health inequalities before intelligence was known to have health associations. Although education and SES are principally thought of as social-environmental variables, both have been found to be partly heritable, by both twin-based and molecular genetic studies, both have high genetic correlations with intelligence, Mendelian Randomisation results show bidirectional genetic effects between intelligence and education, and both have genetic correlations with health outcomes [25••,40,41,42,43,44,45].","Intelligence has predictive power for many health outcomes. Part of that association is genetic. The genes involved, and the causal pathways of the associations are being explored.","IJD and SEH are supported by the University of Edinburgh Centre for Cognitive Ageing and Cognitive Epidemiology which is funded by the UK Medical Research Council and Biotechnology and Biological Sciences Research Council (Grant no. MR/K026992/1). WDH is supported by Age UK (Disconnected Mind programme)."],["Little is known about why people differ in their levels of academic motivation. This study explored the etiology of individual differences in enjoyment and self-perceived ability for several school subjects in nearly 13,000 twins aged 9-16 from 6 countries. The results showed a striking consistency across ages, school subjects, and cultures. Contrary to common belief, enjoyment of learning and children's perceptions of their competence were no less heritable than cognitive ability. Genetic factors explained approximately 40% of the variance and all of the observed twins' similarity in academic motivation. Shared environmental factors, such as home or classroom, did not contribute to the twin's similarity in academic motivation. Environmental influences stemmed entirely from individual specific experiences. --------------------------------------------------------------------------------","Academic motivation refers to a wide range of traits, such as individuals’ educationally relevant beliefs, perceptions, values, interests, enjoyment, and attitudes (Ryan & Deci, 2000; Urdan & Midgley, 2003; Wigfield & Eccles, 2000) that are associated to school achievement (Elliot & Dweck, 2005). The etiology of individual differences in these traits remains poorly understood. In this paper, we focused on two important motivational constructs: enjoyment of learning (e.g., interest, liking), usually referred to intrinsic motivation; and self-perceived ability, also known as academic self-concept (e.g., children’s perception of how good they are at school subjects). Several recent studies found self-perceived ability to be substantially heritable (Spinath, Spinath, & Plomin, 2008), even when controlling for general cognitive ability (Greven, Harlaar, Kovas, Chamorro-Premuzic, & Plomin, 2009; Luo, Kovas, Haworth, & Plomin, 2011). In terms of environmental contributions, up to 60% of the variance in enjoyment and self-perceived ability is explained by non-shared experiences (Spinath et al., 2008). Despite the absence of significant shared environmental effects shown by recent large twin studies, several educational studies found a link between aspects of academic motivation and family/classroom-wide factors, such as classroom climate, peer influence, and mothers’ motivational practices in child’s education (Church, Elliot, & Gable, 2001; Gottfried, Fleming, & Gottfried, 1994; Marsh, Martin, & Cheng, 2008; Ryan, 2000). One possible explanation for this inconsistency is that environmental influences may be correlated with genetic effects (Plomin, DeFries, Knopik, & Neiderhiser, 2012). For example, parental involvement in child’s education may have a causal effect on motivation or/and reflect partly genetically driven parental levels of education, ability, and motivation. Some observed classroom effects might also stem from intake selection (e.g., ability streaming). Most research into the relevant home environmental influences examines only one child per family, which makes it difficult to establish whether the environmental effects operate in a family-wide or child-specific manner. It is possible that even objectively shared experiences, such as availability of educational resources at home, act as child-specific experiences through gene-environment correlation, a mechanism through which children in the same home modify their shared environment into individual experiences. The role of teachers in shaping children’s academic motivation has been extensively studied (Chirkov & Ryan, 2001; Church et al., 2001; Reeve & Jang, 2006; Urdan & Midgley, 2003). Research suggested that teachers can promote the development of intrinsic motivation (e.g., enjoyment, liking) by encouraging students’ autonomy, providing feedback and optimal challenges, and adopting a caring attitude towards students (Chirkov & Ryan, 2001; Ryan & Deci, 2000). However, teacher effects cannot be easily disentangled from other potential effects of classroom resources, number of children in the class, and teacher unfacilitated classroom-peer interactions (Olson, Keenan, Byrne, & Samuelsson, 2014). Such teacher/classroom effects vary across development, with potentially stronger or persistent effects at the early stages of the formal education when children are facing systematic instruction and academic feedback for the first time (Church et al., 2001; Kovas, Haworth, Dale, & Plomin, 2007; Reeve & Jang, 2006; Urdan & Midgley 2003). If teachers/classrooms have a strong average effect on children’s liking a particular school subject, we should expect twins in different classes to be on average less similar in their enjoyment of the subject than those in same classes. Findings on academic achievement are mixed: several studies have found small teacher/classroom influences (Byrne et al., 2010; Nye, Konstantopoulos, & Hedges, 2004), whereas other studies did not find any (Kovas et al., 2007), with a recent review suggesting that classroom performance differences should not be viewed as indicators of teacher quality (Olson et al., 2014). It could be that teachers and classrooms have a non-shared, child- specific influence, possibly interacting with children’s genetic and unique environmental background - leading to unique perceptions and reactions in different children. The goal of this study was to investigate the relative contribution of genetic and environmental factors to individual differences in enjoyment and self-perceived ability as a function of cultural and educational settings. Twins between 9 and 16 years of age from six different countries were evaluated on their enjoyment of learning and the perception of their competence in several academic disciplines. We also compared twin similarity in same versus different classrooms to evaluate teacher/classroom effects. Finally, we tested whether the first formal teacher/classroom affects later class-wide level of enjoyment and self-perceived ability (Church et al., 2001; Kovas et al., 2007; Reeve & Jang, 2006; Urdan & Midgley, 2003).","Data of nearly 13,000 identical twins (monozygotic, MZ) and non-identical (dizygotic, DZ) same-sex twins came from six different ongoing twin studies conducted in United Kingdom (Twins Early Development Study – TEDS; Haworth, Davis, & Plomin, 2012), Canada (Quebec Newborn Twin Study – QNTS; Boivin et al., 2013), Japan (Keio Twin Project; Ando et al., 2013), Germany (Twin study on Cognitive ability, Self-reported Motivation and School performance – CoSMoS; Spinath & Wolf, 2006), United States (Western Reserve Reading Project – WRRP; Petrill, Deater-Deckard, Thompson, DeThorne, & Schatschneider, 2006); and Russia (Russian School Twin Registry – RSTR; Kovas et al., 2012). Detailed information on each sample is presented in the Appendix A.1.","Across all samples, children reported their level of enjoyment and self-perceived ability of different school subjects by completing questionnaires. Self-reported evaluations of enjoyment and self-perceived ability were collected from the UK twins at ages 9, 12 (Luo et al., 2011; Spinath, Spinath, Harlaar, & Plomin, 2006) and 16 (OECD, 2000, 2003, 2006); Canadian twins at ages 10 and 12 (Guay, Marsh, & Boivin, 2003); Japanese twins at ages 10, 11, 12, 13 and 16 (Pintrich & de Groot, 1990); German twins at ages 9, 11 and 13 (Spinath et al., 2008); US twins at age 12 (Harlaar, Deater-Deckard, Thompson, DeThorne, & Petrill, 2011); and Russian twins at age 16 (OECD, 2000, 2003, 2006). Table 1 summarizes the measures and the overall sample size for each twin study. The table indicates maximum number of children in each sample. Although the measures used across the samples were not identical, they were designed to tap into the same motivational constructs. Convergence of results under these circumstances warrants greater confidence in their generalizability and replicability beyond specific methodological features. Details of each measure are presented in the Appendix A.2. Procedure Analyses were conducted on variables corrected for age and sex within each sample. Where data on opposite-sex DZ twins were available (UK, Canada, Japan, and Germany), we ran the analyses twice, including and excluding opposite sex DZ twins - with very similar results. The information on whether twins and their co-twins were taught in the same or different classes was also available in the UK sample at ages 7 and 9. We tested whether being in different classes for 8 or more months reduces similarity in the level of enjoyment and perceived ability for the two twins by dividing the sample into same versus different class at age 9. The proportions of twins in same vs. different classrooms were very similar for the two zygosity groups: 60% of MZ twins vs. 59% of DZ twins were taught in the same class. In addition, we split the sample at age 9 into the same teacher/class at age 7 to test whether the first teacher or class had a long-lasting class-wide contribution to academic motivation. Twin analyses ~~~~~~~~~~~~~ We examined the etiology of enjoyment and self-perceived ability for every subject at every age and in each sample separately by comparing within-pair similarity for MZ and DZ twins. Heritability (A) can be estimated as twice the difference between the MZ and DZ intra-class correlations (ICCs). Shared environmental influences (C) are suggested if DZ twins’ correlation is more than half of the MZ correlation and are computed by subtracting the heritability from the MZ ICC. Shared environment refers to environmental influences that both members of a twin pair experience and that increases the similarity between them. Factors such as socio-economic status, home environment, and school are often thought to contribute to similarities among family members. Non-shared environmental influences (E) are estimated by subtracting MZ twin correlation from 100% (Plomin et al., 2012). The non-shared environment refers to environmental factors that are experienced differently by each twin of a pair and that increase their dissimilarity. Non-shared environmental influences may include individual specific experiences, such as different peers and classmates, differential treatment by their parents and teachers, and differences in twins’ perceptions of such experiences (Kovas et al., 2007). Non-shared environmental estimates also include measurement error. Classroom heterogeneity model ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The effects of classroom on enjoyment and self-perceived ability were investigated by fitting “the classroom heterogeneity model” to the data available from the UK sample. These model-fitting analyses tested whether the differences in estimates of the ACE parameters for twin pairs studying in the same class and twin pairs studying in different classes were statistically significant. The classroom heterogeneity model is similar to the sex-limitation models used to test for quantitative sex differences (Kovas et al., 2007). To test for the long-lasting (spill-over) effects of the first teacher/classroom experiences on later academic motivation, we performed the same analyses splitting the sample at age 9 into twins who were attending the same versus different classroom when they were 7.","A wide variation in academic motivation scores was observed within each sample. The data for most measures were normally distributed. Prior to all analyses, where data did not meet the criterion of normality, the necessary transformations were applied (e.g., Vander Waerden, reflect and log). Repeated analyses using transformed and untransformed scores yielded similar results. Phenotypic correlations ~~~~~~~~~~~~~~~~~~~~~~~ Pearson correlations between enjoyment and self-perceived ability were performed on one twin randomly selected out of each pair, and conducted on scores adjusted for age. Correlations were moderate to strong in all samples: r = .41–.79, averaged = .65 (see Table B.1 in the Appendix). Effects of sex and zygosity ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Analyses of variance (ANOVA) were performed within each sample in order to assess the effects of sex and zygosity and their interaction on each variable. The results were adjusted for exact age within each sample. Means, standard deviations and the results of ANOVAs are presented in Appendix (see Tables B.2 and B.3). Overall, boys reported higher perceived ability, with 6 out of 16 comparisons reaching significance (p < .05), and higher enjoyment of mathematics and science in all samples, with 5 out of 16 comparisons reaching significance (p < .05). The effect size of these differences was small, ranging from less than 1–6% for self-perceived ability, and ranging from less than 1–9% for enjoyment. On the contrary, girls reported higher perceived ability, with 3 out of 8 significant comparisons (p < .05), and with less than 2% of the variance explained by sex. They also reported higher enjoyment of reading and language academic subjects, with 5 out of 8 significant comparisons (p < .05). Between 1% and 5% of the variance in enjoyment was explained by sex. Overall, MZ and DZ twins showed similar levels of enjoyment and self- perceived ability within each sample. In the UK sample, we were also able to test for a main effect of zygosity, class, and zygosity by class interaction on enjoyment and self- perceived ability. In other words, we tested whether twins showed greater enjoyment and higher self-perceived ability when they were taught in the same (as opposed to different) classroom; and whether this effect was specific (or more prominent) to one of the zygosity groups. We conducted a series of 2 by 2 ANOVAs, for each school subject, with zygosity (MZ vs. DZ) and class (same vs. different) – as two factors. For enjoyment, we found no class or zygosity effect and no interaction. In other words, average levels of enjoyment of the subjects were similar for MZ and DZ twin groups, irrespective of whether they attended the same or different classes. For self-perceived ability, we found no zygosity effect but a significant effect of class on self-perceived ability for English and Maths: children in the same classroom showed a slightly higher level of self-perceived ability. However, this effect was negligible, explaining less than 1% of the variance. No significant interactions were found. These results suggest that twins (both MZ and DZ) have slightly higher self-perceived ability when taught in the same (rather than different) class. However, in this study, this effect was too weak to justify any further interpretation. Heritability of enjoyment and self-perceived ability ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ MZ and DZ ICCs are presented in Tables 2 and 3, separately for enjoyment and self- perceived ability. Striking similarities were observed across the ages, school subjects and samples for both enjoyment (average MZ ICC = .46; average DZ ICC = .16) and self- perceived ability (average MZ ICC = .46; average DZ ICC = .19). Overall, the results showed that individual differences in enjoyment and self-perceived ability are explained to a similar extent by genetic and individual-specific (i.e., non-shared) environmental factors. Genetic contributions ranged from 16% to 69% across the samples; non-shared environmental contributions ranged from 31% to 75%. Some potentially meaningful cultural specific and subject specific effects were observed. For example, modest shared environmental influences were found for enjoyment and self-perceived ability in German- language at age 9, and for self-perceived ability at age 13; modest but significant shared environmental influences (10%) were found for self-perceived ability in science at age 9; and moderate shared environment was found in the Japanese sample for self-perceived ability in mathematics at age 11. Although these four exceptions, no significant shared environmental influences on these constructs were found. Figure 1 presents the results averaged across age, school subject, motivational construct, and country (see Fig. C.1 in the Appendix for the results split by country). Classroom effect on enjoyment and self-perceived ability ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Children in different classes did not rate their enjoyment of the subjects or their self- perceived ability less similar than those in same classes, with equal genetic (A), shared (C) and non-shared environmental (E) estimates for the two groups. We also tested the assumption that the effect of the first formal teacher may have a continuous or delayed effect on later motivational levels by splitting the sample at 9 years of age by whether the children were taught by the same or different teacher at age 7. The ACE parameters could be equated when the UK sample was split by whether the twins attended the same versus different classes at age 7. In other words, no class-wide effect on contemporaneous or later levels of enjoyment and self-perceived ability was found (see Tables B.4–B.9 in the Appendix).","Overall, the pattern of results for enjoyment and self-perceived ability was highly similar, which is not surprising as these constructs were moderately to strongly correlated for each school subject in each sample. The results showed high consistency across ages, school subjects and cultures in the etiology of individual differences in enjoyment and self-perceived ability. This consistency is particularly striking given the vast cross-cultural differences in schooling and the educational systems involved. The familial similarity in levels of academic motivation was only moderate, even for genetically identical children raised in the same home. With few exceptions, neither enjoyment nor self-perceived ability were influenced by shared environment. In other words, similarities in enjoyment and self-perceived ability in twins growing up in the same family and attending the same schools were entirely explained by their genetic, rather than their environmental relatedness. However, genetic effects on enjoyment and self-perceived ability varied substantially across the samples. These differences in heritability could reflect some cultural aspects that lead to differences in amount of observed variation explained by genetic factors. The observed differences could also be explained by differences in sample sizes and associated statistical power. Moreover, attending different classrooms did not increase dissimilarity between twins in their levels of enjoyment and self-perception of competence. Equal similarity between twins attending same and different classrooms cannot be explained with equalising effect of the shared home environment as no such effect was found in this study. These results suggest that similarity in academic motivation for any unrelated individuals stems from their chance genetic similarity, as well as similar individual-specific environmental experiences, rather than similar family/class-wide experiences. Whatever the environmental influences on the levels of enjoyment and self-perceived ability are, they seem to act in a non-shared, individual-specific way, potentially interacting with genetic make-up, experiences and perceptions. Multiple individual-specific life-events, such as birth complications, missing school due to illness, and peer-relations, may contribute to motivation. Effects of family members, teachers, classes, and schools may also be non- shared: parents, siblings, and teachers may actually treat children in the same family/class differently, responding to their individual characteristics (Babad, 1993; Harris & Morgan, 1991; Spengler, Gottschling, & Spinath, 2012). On the other hand, children may perceive their parents, teachers, classmates, and schools differently (Zhou, Lam, & Chang, 2012) – depending on other non-shared environmental and genetic effects. In addition, genetic effects may differ as a function of environment. For example, research suggested that heritability of reading might be moderated by teacher quality or SES status (Taylor, Roehrig, Hensler, Connor, & Schatschneider, 2010).","Considering the striking consistency of these results across different aspects of academic motivation, different subjects, different ages, and different cultures, we believe that it is time to move away from solely environmental explanations, such as “good” or “bad” home, teacher, and school, for differences in enjoyment and self-perceived ability (Olson et al., 2014). The results convincingly show that, contrary to common belief, enjoyment of learning and children’s perceptions of their competence are no less heritable than cognitive ability (Greven et al., 2009). Surprisingly, unlike cognitive ability, for which shared environment makes a small to moderate contribution across the school years (Petrill et al., 2004), no such contribution was found for these motivational constructs. It remains unclear to what extent the genetic and non-shared environmental factors contributing to variation in enjoyment and self-perceived ability also contribute to variation in achievement and intelligence (Gottfried et al., 1994). Academic motivation is not independent of achievement, as it develops partly through feedback on performance and in turn may influence achievement (Guay et al., 2003). For example, some studies found bidirectional effects between motivation and performance (Luo et al., 2011). This and other genetically sensitive studies call for caution when developing interventions aimed at raising academic motivation before more is known about specific mechanisms underlying its variation (Olson et al., 2014). Current educational policies are based on average effects and are designed to operate at the family-wide and class-wide levels. However, the present research suggests that many true effects may be masked within any class or home, and that individual-specific educational approaches are required."],["Three experiments examined the influence of other people's negative emotions on young children's counterfactual thinking. Experiment 1 (N = 48) explored whether 4- to 6-year-olds could think counterfactually about both physical and emotional events using the discriminating counterfactual tasks that children could not respond correctly without thinking counterfactually. Experiment 1 showed that 4- to 6-year-olds could think of counterfactuals associated with emotional events. Experiment 2 (N = 97) and Experiment 3 (N = 48) examined whether a protagonist's emotional state (emotional expression condition) affected 4- to 6-year-olds' ability to think counterfactually about physical events. It was shown that emotional expression conditions enhanced young children's counterfactual thinking about physical events. These findings suggest that 5- and 6-year-olds can think counterfactually and that emotional components trigger such thinking. --------------------------------------------------------------------------------","Counterfactual thinking is a form of thinking that considers alternative possibilities for an event or behavior in the past. Counterfactual thinking has an adaptive significance for humans in that it allows us to learn from past negative experiences and to avoid negative outcomes in the future (Byrne, 2005, 2016; Epstude & Roese, 2008). Moreover, counterfactual thinking has an important role in children’s cognitive development. Counterfactual thinking during early childhood is closely associated with the following key abilities: understanding causal relations (German, 1999; Harris, German, & Mills, 1996), acquiring theory of mind (Guajardo & Turley-Ames, 2004; Rafetseder & Perner, 2018; Riggs, Peterson, Robinson, & Mitchell, 1998), understanding regret and relief (e.g., Beck & Crilly, 2009; Guttentag & Ferrell, 2004), and engaging in pretend play (Buchsbaum, Bridgers, Weisberg, & Gopnik, 2012). Recent years have seen a growing interest in the development of counterfactual thinking in the field of cognitive development (see Beck & Riggs, 2014; Rafetseder & Perner, 2014). One of the main areas explored in the previous studies is the question of when the capacity for counterfactual thinking is acquired. Studies of the development of counterfactual thinking have shown mixed results. In their pioneering work on the development of counterfactual thinking, Harris et al. (1996) presented children with a scenario that involved a causal chain in the following form: initial state -> causal event -> effect state (e.g., the floor is clean -> Susie walks across the floor with muddy shoes -> the floor gets dirty). They then asked each child a counterfactual question (e.g., “What would have happened if Susie had taken off her muddy shoes?” ; correct answer: the floor would have remained clean). More than 70% of 3-year- olds were able to get the right answer to this counterfactual question. Studies based on this approach (Beck, Riggs, & Gorniak, 2010; German, 1999; German & Nichols, 2003; Nakamichi, 2011; Riggs et al., 1998) have demonstrated that children are capable of counterfactual thinking by around 5 or 6 years of age. By contrast, Rafetseder and colleagues (Rafetseder, Cristi-Vargas, & Perner, 2010; Rafetseder, Schwitalla, & Perner, 2013; see Rafetseder & Perner, 2014, for a review) have argued that 5- and 6-year-olds find it difficult to use counterfactual thinking and that the development of counterfactual thinking is a long-term process. In one experiment, Rafetseder et al. (2013, Experiment 2) asked both 5- to 15-year-old children and adults to perform counterfactual tasks referred to as the discriminating tasks. Rafetseder et al. presented each participant with a scenario involving a causal chain in the following form: initial state -> first causal event -> interim effect state -> second causal event -> final effect state (e.g., the floor is clean -> Susie walks across the floor with muddy shoes -> the floor gets dirty -> Max walks across the floor with muddy shoes -> the floor gets dirty”). They then asked each participant a counterfactual question (e.g., “What would happen if Max took off his muddy shoes?” ; correct answer: the floor would still be dirty). This discriminating task looks similar to the simplified task in Harris et al. (1996). Unlike the latter task, however, the discriminating task has the structure that does not allow a child to produce the correct answer without taking into account the premises of the scenario (e.g., even if Max took his shoes off, Susie would still walk across the floor with muddy shoes). The results showed that 5- and 6-year-olds and 7- to 10-year-olds (percentages of correct answers = 18% and 53%, respectively) performed less well than 13- to 15-year-olds and adults (percentages of correct answers = 88% and 95%, respectively). Moreover, 5- and 6-year-olds performed lower than chance, and 7- to 10-year-olds performed higher than chance. In contrast to studies (Beck et al., 2010; German, 1999; German & Nichols, 2003; Nakamichi, 2011; Riggs et al., 1998) based on the approach used by Harris et al. (1996), the results of Rafetseder et al. (2013) suggest that 5- and 6-year-olds find it difficult to use counterfactual thinking and that children do not acquire an adult-level mode of counterfactual thinking until around 12 years of age. Based on these experimental results, Rafetseder and colleagues (Rafetseder et al., 2010, 2013; Rafetseder and Perner, 2014) have argued that the tasks based on Harris et al. (1996) are a simplified structure and that such tasks cannot adequately measure a child’s capacity for counterfactual thinking. For example, in response to the question “If Susie was to take off her muddy shoes, …?”, children might answer “the floor would be clean” based on a generally conceivable premise (e.g., if you take off your shoes, the floor won’t get dirty) even if children ignore the original scenario. Rafetseder and colleagues have called this inference from the general premise basic conditional reasoning and suggested the need to examine the development of counterfactual thinking using tasks that do not allow children to produce a correct answer unless they think of a counterfactual possibility based on the actual event rather than tasks that allow children to produce a correct answer via basic conditional reasoning. By distinguishing between basic conditional reasoning and counterfactual thinking in young children, Rafetseder et al. (2010, 2013) have offered important evidence with respect to the development of counterfactual thinking. However, the fact that 5- and 6-year-olds found it difficult to complete the discriminating tasks in Rafetseder and colleagues’ studies does not necessarily prove that the children lacked all capacity for counterfactual thinking. Several studies have shown success on counterfactual tasks at age 6 years or younger when the causal structure of the events is clear (Nyhout & Ganea, 2019; Nyhout, Henke, & Ganea, 2019) and the task demands on memory and attentional resources are reduced (McCormack, Ho, Gribben, O’Conner, & Hoerl, 2018; Rafetseder & Perner, 2010). For example, Nyhout and Ganea (2019) presented 3- to 5-year-olds with a “blicket detector” machine that had a clear and novel structure for children and then asked counterfactual questions (e.g., “If she did not put the block on the box, would the light still be on?”). As a result, 4- and 5-year-olds’ performances were above chance level, although 3-year-olds’ performances were not. These studies suggest that young children display mature counterfactual thinking under specified conditions. In addition, Rafetseder and Perner (2018) mentioned the possibility that the performance of Rafetseder et al. (2013, Experiment 2) task was affected by materials and procedure (e.g., footprints were differently colored or not, two protagonists entered the scene at the same time or not). Furthermore, it is possible that the tasks developed by Rafetseder and colleagues did not raise young children’s needs to consider counterfactual worlds with better outcomes even among those who did have the ability to think counterfactually. Counterfactual thinking about a specific event is activated particularly by the event’s negative outcome valence (the general value the result can have), which has an adaptive significance because it can preempt negative outcomes that could occur in the future (Byrne, 2005, 2016; Epstude & Roese, 2008). For example, in the course of reviewing previous studies, Byrne (2016) demonstrated that adults tend to think of counterfactual alternatives more frequently for events that are likely to lead to negative outcomes than for events that are likely to lead to positive or neutral outcomes. This characteristic of adult counterfactual thinking is also seen in young children (German, 1999) and school children (Guajardo, McNally, & Wright, 2016). German (1999) told 5-year-olds about an event in which the protagonist’s decision could produce either a positive outcome or a negative outcome (e.g., in a situation where the protagonist could have chosen either a jacket or a cardigan, he or she wore the cardigan and played outside -> the protagonist is able to play outside all day/the protagonist catches a cold and needs to go home early).","Participants Participants were then asked to explain the event outcomes. In this experiment, 5-year-olds produced more counterfactual statements (e.g., the protagonist should have chosen the jacket) for events that could lead to negative outcomes than for events that were likely to lead to positive outcomes. In this way, counterfactual thinking can be invoked by the negative outcome valence of an event even in young children. All of Rafetseder and colleagues’ counterfactual tasks dealt with physical events that were associated with a relatively ambivalent outcome valence. For example, the physical state of Max and Susie’s footprints on the floor in the above scenario does not necessarily have a negative valence for children. This may be the reason why young children found it difficult to perform Rafetseder and colleagues’ counterfactual task. In association with these findings, Epstude and Roese (2008) proposed the functional theory of adult counterfactual thinking. Their theory suggested that the recognition of a problem activates counterfactual thinking and that the recognition process was influenced by human emotion: “Negative affect may also influence the activation of counterfactual thinking” (p. 171). Compared with neutral or positive emotions, negative emotion may signal that there is a potential problem in the outcome of an event or behavior, thereby encouraging the individual to seek an alternative option that could produce a better outcome. In short, human negative emotion plays an essential role in triggering counterfactual thinking by clearly conveying the outcome valence of an event. In line with theoretical postulates of adult counterfactual thinking (Epstude & Roese, 2008), Guajardo, Parker, and Turley-Ames (2009) and Nakamichi (2011) have demonstrated that young children are better able to use counterfactual thinking when considering events that are related to human emotion compared to physical events without human emotion (in particular negative emotion). For example, Nakamichi (2011) presented 3- to 6-year-olds with a scenario about events involving physical change (e.g., there is a glass cup on the desk -> the cup falls from the desk -> the cup breaks) or human emotional change (e.g., Taro looks at the flower, which makes him happy -> a dog steps on the flower -> Taro feels sad) and then asked a counterfactual question about each of these events (e.g., “If the cup had not fallen, what would it look like?”, “If the dog had not stepped on the flower, how would Taro feel?”). Children’s performance in the emotional task (percentage of correct answers = 76%) was better than their performance in the physical task (percentage of correct answers = 55%) and beyond chance for all ages. Furthermore, Guajardo et al. (2009) told 3- to 5-year-olds about events involving either physical or emotional change and asked, “How could the outcome be changed?” —allowing them to freely produce counterfactual possibilities in response to each event. Children produced more upward counterfactual statements (i.e., resulting in better outcomes) for the emotional event than for the physical event. Based on these studies (Guajardo et al., 2009; Nakamichi, 2011) following the approach of Harris et al. (1996), the young children apparently were capable of counterfactual thinking when an event was associated with humans’ negative emotion rather than mere physical change, although this effect on young children’s counterfactual thinking might not be stable; Beck et al. (2010) contrasted physical and emotional contents in counterfactual tasks and did not find consistent differences in 3- and 4-year- olds. These results suggest the possibility that a component of negative human emotion in an event can clarify the event outcome valence and trigger the ability to think counterfactually in young children. The events in the counterfactual tasks developed by Rafetseder et al. (2013) involved physical change only; they had no emotional component such as the protagonist’s emotional state. For this reason, Rafetseder et al. failed to prove that young children genuinely cannot engage in counterfactual thinking. Guajardo et al. (2009) and Nakamichi (2011) based their approaches on Harris et al. (1996); consequently, it is not clear whether children’s better performance in counterfactual tasks associated with an event involving emotional changes was triggered by human emotional components or by basic conditional reasoning (Rafetseder et al., 2010, 2013; Rafetseder & Perner, 2014). The current study explored the impact of negative emotional components on young children’s counterfactual thinking using tasks that are difficult to answer correctly without taking the actual events into account (i.e., discriminating counterfactual tasks) rather than tasks that could be easily answered using basic conditional reasoning. Three experiments were carried out in this study. Experiment 1 compared children’s performance in the physical event task that involved physical change and no negative emotional component with their performance in the emotional event task that involved the change of emotional state. Based on Guajardo et al. (2009) and Nakamichi (2011), young children should perform better on the emotional event task than on the physical event task even using discriminating counterfactual tasks. Moreover, Experiments 2 and 3 used the physical event task from Experiment 1 to discover whether focusing on a protagonist’s negative response to an event outcome would encourage young children to think counterfactually. In these three experiments, according to the results of Rafetseder et al. (2010, 2013), it was predicted that young children would find it difficult to carry out tasks associated with a physical event. However, if a human’s negative emotion did trigger counterfactual thinking in young children, better performance could be expected in tasks involving human emotion across all experiments. Participants The participants were 48 children aged 4 to 6 years (24 boys and 24 girls; Mage = 63.56 months, SD = 7.47, range = 52–75) recruited from three public nursery schools in Shizuoka, Japan. All participants were Japanese native speakers from middle-income families. The children were divided into two age groups: a younger group with 24 children (12 boys and 12 girls; Mage = 57.01 months, SD = 3.38, range = 52–64) and an older group with 24 children (12 boys and 12 girls; Mage = 70.13 months, SD = 3.57, range = 65–75), following the Japanese education system of classification.","Materials Materials Based on earlier studies (Guajardo et al., 2009; Nakamichi, 2011; Rafetseder et al., 2013), six scenarios were created: three involving a physical event (cup, footprint, and sandbox scenarios) and three involving an emotional event (garden, zoo, and birthday scenarios). Original Japanese scripts for the stories are provided in the Appendix. In each scenario, six pictures (approximately 14.5 × 14.5 cm) were used to present the content. Fig. 1 presents the cup and garden scenario pictures used to depict those emotional and physical events. The physical event scenarios did not refer to the emotional state of the protagonist, describing only a physical change. The emotional event scenarios described the protagonist’s emotional state, changing from a positive emotion to a negative emotion. For each scenario, the first picture displayed an introductory scene. This was followed by a series of pictures in which the causal structure took the following form: initial state of a physical object/emotion -> first causal event -> interim state of the physical object/emotion -> second causal event -> final state of the physical object/emotion. In addition, in each scenario three pictures depicted different answers to the counterfactual question (Fig. 2). Procedure Procedure Procedure All of the steps in this experiment were reviewed and approved by the affiliated university’s committee for ethical research involving human participants. The principals in the nursery schools informed the parents about the experiment and received consent from parents for their children to participate. Written consent was obtained from a principal of each nursery school, acting as a surrogate for the parents. Verbal assent was obtained from the children before they participated in the experiment. Each participant was interviewed individually in a quiet room. The interviews lasted for approximately 20 min. All of the participants were asked to perform tasks involving six scenarios: three physical event tasks and three emotional event tasks. In each age group, half of the participants were initially assigned three physical tasks, followed by three emotional tasks; the other half were assigned the same tasks in reverse order. The order of three tasks (both physical and emotional) was counterbalanced. After each scenario, participants were asked to answer the Now control question (“What was the state of object/emotion at the end of the scenario?”), the Before control question (“What was the state of object/emotion at the beginning of the scenario?”), and the counterfactual Test question (“If the second causal event had been different, how would the final state have changed?”). In each task, the scenario was described to participants by using picture story cards. For instance, participants heard the following scenarios: Physical event task, cup scenario: [introductory picture in Fig. 1] Here are Jiro and Satoko. They are playing with blocks on the desk. There is a glass cup on the desk. [ initial state picture] This is a glass cup on the desk. The glass cup has two handles. [ first causal event picture] As they play with the toys, Satoko’s block hits the cup. [ interim state picture] One of the cup handles breaks. [ second causal event picture] Right after that, Jiro’s hand hits the cup and it falls off the desk. [ final state picture] Both cup handles get broken. Emotional event task, garden scenario: [introductory picture in Fig. 1] This is Keiko. Keiko is carrying a piece of candy while looking at a flower in the garden. [ initial state picture] Keiko loves candy and flowers, so she is very happy. [ first causal event picture] After a while, a dog comes along, steps on the flower, and squashes it. [ interim state picture] Keiko feels a bit sad. [ second causal event picture] Right after that, Keiko tries to eat her candy but drops it. [ final state picture] Keiko feels very sad. After the story cards were placed out of children’s sight following the end of the story, participants were asked to answer two control questions. The Now question involved the final situation in the scenario, and the Before question involved the initial situation in the scenario. For example, in the cup scenario, the questions were “At the end of the scenario, did the cup have two handles?” (Now question) and “At the beginning of the scenario, did the cup have two handles?” (Before question). In the garden scenario, the two questions were “At the end of the scenario, did Keiko feel happy or very sad?” (Now question) and “At the beginning of the scenario, did Keiko feel happy or very sad?” (Before question). To avoid response bias, the order of presenting these questions was counterbalanced. Finally, participants were asked to select one of three possible answers to the counterfactual Test question: “If the second causal event had been different, what would the final state have been?” For instance, in the cup scenario, participants were asked to choose an answer to the question “If the cup had not fallen off the desk, what would the cup look like now?” The three choices were a picture of the cup with no handles (Fig. 2A: final state), a picture of the cup with two handles (Fig. 2B: initial state), and a picture of the cup with only one handle (Fig. 2C: correct answer). Likewise, in the garden scenario, participants were asked to select one of three possible answers to the question “If Keiko had not dropped her candy, how would she feel right now?” The three choices were a face looking very sad (Fig. 2A': final state), a face looking happy (Fig. 2B': initial state), and a face looking slightly sad (Fig. 2C': correct answer). For each question, the experimenter presented the final state choice at the beginning, followed by the remaining two choices at random. Two control questions and a test question were prepared for each of the other scenarios, in accordance with storylines.","Results Results The preliminary analyses did not show differences in task performance by gender and task order. So, the following analyses combined data for these factors. Understanding the contents of scenarios The percentage of correct answers was 95.5% overall for the control questions. In the physical event tasks, the percentages of correct answers were 94.7% for the Now questions, 93.2% for the Before questions, and 94.0% overall. In the emotional event tasks, the percentages of correct answers were 98.5% for the Now questions, 95.5% for the Before questions, and 97.0% overall. Many participants demonstrated a good understanding of the scenario events. However, some participants gave only one wrong answer to the control questions, with 9 participants failing in physical event tasks and 6 participants failing in emotional event tasks. For this reason, they were not excluded from the sample. Instead, to reflect participants’ different ways of understanding the scenarios, if a participant answered any of the control questions incorrectly, 0 points were assigned to the corresponding test question regardless of whether the participant gave the correct answer or not. Answers to the counterfactual questions Participants received 1 point if they answered the two control and counterfactual test questions correctly in each scenario. Afterward, the scores for all three scenarios in both physical and emotional event tasks were totaled to provide a counterfactual score (Max = 3 each). Table 1 shows the mean scores and the numbers and percentages of children receiving scores of 0, 1, 2, and 3 by age group and task. First, a generalized estimating equation (GEE) analysis was conducted. This GEE model included age group (younger or older) as a between-participants predictor, task (physical or emotional) as a within-participants variable, and the counterfactual score (out of 3) as the dependent variable. The results showed that the older group had only a marginally higher score compared with the younger group, B = 1.07, SE = 0.59, Wald χ2(1) = 3.28, p = .070, odds ratio = 2.90, 95% confidence interval (CI) [0.92, 9.20]. In addition, the main effect of the task was significant, with the emotional event task gaining a higher counterfactual score than the physical event task, B = 3.51, SE = 0.64, Wald χ2(1) = 29.68, p < .001, odds ratio = 33.57, 95% CI [9.48, 118.82]. The Age Group × Task interaction was not significant (p = .785). Second, a one-sample Wilcoxon signed-rank test using a counterfactual score was carried out by age group and task to analyze the difference between participants’ performances and chance level (chance = 1.00). In the physical event tasks, the younger group’s score was significantly lower than chance (Z = −2.84, p = .005, r = .58), whereas no significant difference was seen between the older group’s score and chance (Z = −0.94, p = .346, r = .19). In the emotional event tasks, the scores were significantly higher than chance for the younger group (Z = 3.84, p < .001, r = .78) and the older group (Z = 4.04, p < .001, r = .83). Moreover, the analyses using children’s actual response (taking no account of answer to control questions) showed similar results (physical event tasks: younger group, Z = −2.52, p = .012; older group, Z = −1.27, p = .206); emotional event tasks: younger group, Z = 4.67, p < .001; older group, Z = 3.84, p < .001).","Discussion Discussion In Experiment 1, the discriminating counterfactual tasks that are difficult to answer correctly without taking the actual events into account were used to assess the counterfactual thinking of 4- to 6-year-olds. First, the older group (percentage of correct answers = 51.3%) performed better than the younger group (percentage of correct answers = 39.7%), meaning that the ability of counterfactual thinking improved from 3 to 6 years of age. Moreover, participants found it more difficult to perform physical event tasks (percentage of correct answers = 21.5%) than to perform emotional event tasks (percentage of correct answers = 69.5%). The performance of the younger group was below chance, whereas the performance of the older group was similar to chance. Some studies that have used simple tasks, solvable using basic conditional reasoning (e.g., German & Nichols, 2003; Harris et al., 1996), have previously suggested that 3- and 4-year-olds are capable of counterfactual thinking. However, the results of Experiment 1 suggest that children under 5 years of age find it extremely difficult to think counterfactually about physical events that do not involve an emotional component. The difficulty with performing physical event tasks in our study was consistent with the findings of Rafetseder et al. (2013) that used discriminating counterfactual tasks (percentage of correct answers in 5- and 6-year-olds = 18%). Participants in this study performed better in emotional event tasks than they did in physical event tasks, with all age groups performing above chance. These results confirm the results of studies that have compared performance in physical events with performance in emotional events (Guajardo et al., 2009; Nakamichi, 2011). In other words, the current study showed that even when using an approach based on more stringent standards (i.e., discriminating tasks), children under 5 years of age can think counterfactually about emotional events but not about physical events without emotional components. This result suggests that young children not only use basic conditional reasoning but also think counterfactually about events with an emotional component. However, the results of Experiment 1 do not prove that negative emotion triggers counterfactual thinking in young children. It might simply be less demanding for young children to think about changes in an emotional state than to think about changes in a physical state. Consequently, a further experiment was needed to examine the effect of negative emotion on young children’s counterfactual thinking. If negative emotion triggers counterfactual thinking, young children should be capable of thinking counterfactually about physical events by presenting them with the protagonist’s negative emotion. Thus, in Experiment 2, we created the new condition (i.e., emotional expression condition) that was designed to add negative emotional components to the physical events used in Experiment 1. Participants The participants were 97 children aged 4 to 6 years (46 boys and 51 girls; Mage = 64.34 months, SD = 6.56, range = 53–76) recruited from three public nursery schools in Shizuoka. All of the participants spoke Japanese as their first language and were members of middle-income families. These children were divided into two groups: a younger group with 47 children (23 boys and 24 girls; Mage = 58.47 months, SD = 3.23, range = 53–64) and an older group with 50 children (23 boys and 27 girls; Mage = 69.86 months, SD = 3.21, range = 65–76), as with the classification of Experiment 1. None of the participants from Experiment 1 participated in Experiment 2. Materials This experiment used the three physical event scenarios from Experiment 1 (the cup, footprint, and sandbox scenarios). The six pictures representing the content of the event and three pictures of possible answers to the counterfactual question were used in each scenario. In the emotional expression conditions, in addition to the six pictures mentioned above, “a picture of the initial emotional state of the protagonist” (positive emotion) and “a picture of the final emotional state of the protagonist” (negative emotion) were used in each scenario. Fig. 3 shows the pictures used for both the control and emotional expression conditions. Procedure All of the steps in this experiment were reviewed and approved by the affiliated university’s committee for ethical research involving human participants. As with Experiment 1, written consent was obtained from a principal of each nursery school, and verbal assent was obtained from the children. Participants were assigned to either the control condition or the emotional expression condition. In total, 24 children in the younger group (12 boys and 12 girls; Mage = 59.04 months) and 25 children in the older group (11 boys and 14 girls; Mage = 70.32 months) were assigned to the control condition, whereas 23 children in the younger group (11 boys and 12 girls; Mage = 57.87 months) and 25 children in the older group (12 boys and 13 girls; Mage = 69.40 months) were assigned to the emotional expression condition. No significant differences in months of age by condition or gender were seen in these age groups. Each participant was interviewed individually in a quiet room. The interview lasted for approximately 15 min. In both conditions, participants were asked to perform three counterfactual tasks. The order of presenting three tasks was counterbalanced among participants. In the control condition, the scenario for each task was described while the picture story cards were presented to participants. For example, the footprint scenario went as follows: [introductory picture in Fig. 3] This is Taro, the older brother, and Hanako, his younger sister. The two are playing outside and have mud on their shoes. [ initial state picture] This is the inside of Taro and Hanako’s house. Today, the floor is very clean. [ first causal event picture] A short while later, Taro walks on the floor with his shoes on. [ interim state picture] The floor has Taro’s footprints. [ second causal event picture] Right after that, Hanako walks on the floor with her shoes on. [ final state picture] Now the floor has both their footprints. After each scenario, participants were asked to answer the three questions (the Now question, the Before question, and the Test question). In the emotional expression condition, two scenes referring to the emotional state of the protagonist were added to each scenario in the control condition: between the introductory scene and the initial state scene and after the final state scene. For instance, in the footprint scenario, two scenes were inserted: a scene where “both Taro and Hanako have fun playing outside” (Fig. 3A) and a scene where “Hanako feels sad after seeing footprints on the floor” (Fig. 3B). Apart from these points, the storyline and task presentation processes were the same as in the control conditions. Results ~~~~~~~ Preliminary analyses showed no gender difference. So, the genders were mixed in the following analyses. Understanding the contents of scenarios Overall, the percentage of correct answers to the control questions was 94.6%. In the control conditions, the percentages of correct answers were 98.0% to the Now questions, 97.3% to the Before questions, and 97.6% overall. In the emotional expression conditions, the percentages of correct answers were 97.9% to the Now questions, 97.2% to the Before questions, and 97.6% overall. Some children gave only one wrong answer (7 participants in the control condition and 7 participants in the emotional expression condition). Therefore, as in Experiment 1, they were not excluded from the sample. Answers to the counterfactual questions The counterfactual scores in each condition were calculated using the same scoring procedure as in Experiment 1. Table 2 shows the mean scores and the numbers and percentages of children receiving scores of 0, 1, 2, and 3 by age and condition. Because a preliminary analysis did not demonstrate any gender difference, the genders were mixed in the following analysis. First, an ordinal regression analysis was conducted to examine the effect of the two predictor variables, age group and condition, on the overall score out of 3. This analysis included age group (younger or older) and condition (control or emotional expression) as categorical predictors and the Age Group × Condition interaction. This model was significant, χ2(3) = 15.87, p = .001. The older group had a significantly higher score compared with the younger group, B = −1.56, SE = 0.55, Wald χ2(1) = 8.20, p = .004, odds ratio = 4.77, 95% CI [−2.63, −0.49]. In addition, the main effect of the condition was significant; participants in the emotional expression condition had a significantly higher score than those in the control condition, B = −1.14, SE = 0.52, Wald χ2(1) = 4.78, p = .029, odds ratio = 3.14, 95% CI [−2.17, −0.12]. The Age Group × Condition interaction was not significant (p = .364). Second, a one-sample Wilcoxon signed-rank test using a counterfactual score was performed by age group and condition (chance = 1.00). In the control condition, the scores of both the younger group (Z = −1.53, p = .127, r = .31) and the older group (Z = 0.87, p = .383, r = .17) were no different from chance. In the emotional expression condition, the score of the younger group was no different from chance (Z = −0.45, p = .655, r = .09), but the score of the older group was higher than chance (Z = 3.20, p = .001, r = .64). Moreover, the analyses using children’s actual response showed similar results (control condition: younger group, Z = −1.58, p = .127; older group, Z = 0.87, p = .383; emotional expression condition: younger group, Z = −0.45, p = .655; older group, Z = 3.29, p = .001). Discussion ~~~~~~~~~~ Experiment 2 compared 4- to 6-year-olds’ counterfactual thinking in control and emotional expression conditions. The results showed that children’s performance was worse under the control condition (percentage of correct answers = 31.3%) than under the emotional expression condition (percentage of correct answers = 46.0%); the performance in the control condition was similar to chance even in the older group. In the control condition in Experiment 2, the same physical event scenarios were used as in Experiment 1. The results, therefore, confirmed the findings of Experiment 1, meaning that young children find it difficult to think counterfactually about physical events that do not involve an emotional component. On the other hand, the task performance was better in the emotional expression condition than in the control condition; in particular, the performance of the older group (percentage of correct answers = 60.0%) was higher than chance. These results show that 5- and 6-year-olds can think of counterfactual alternatives not only about events associated with emotional change but also about events associated with physical change as long as negative emotional responses to the events are presented. However, in the emotional expression condition of Experiment 2, it is possible that children’s better performance was a consequence of paying more attention to one of the two characters in the scenario. In other words, the results of Experiment 2 leave open the possibility that young children’s counterfactual thinking is encouraged when there is one protagonist in a physical scenario. To eliminate this possibility, in Experiment 3 we created a new set of physical scenarios with only one character and examined the influence of emotional expression condition by using these scenarios. Participants The participants were 48 children aged 4 to 6 years (24 boys and 24 girls; Mage = 63.88 months, SD = 6.47, range = 53–75) recruited from two public nursery schools in Shizuoka. All of the participants spoke Japanese as their first language and were members of middle-class households. As in Experiment 2, these children were divided into two groups: a younger group with 24 children (12 boys and 12 girls; Mage = 58.21 months, SD = 2.99, range = 53–64) and an older group with 24 children (12 boys and 12 girls; Mage = 69.54 months, SD = 3.08, range = 65–75). None of the participants from Experiments 1 and 2 participated in Experiment 3. Materials Based on the three physical task scenarios used in Experiments 1 and 2, new scenarios were created in which only one protagonist made an appearance. These scenarios had the same causal structure in Experiment 2. Each scenario used six pictures representing the story content and three pictures as possible answers to the counterfactual question. For example, the footprint scenario went as follows: “This is Taro. He is playing outside and has mud on his shoes. This is the inside of Taro’s house. Today, the floor is very clean. A short while later, Taro walks across the floor with his shoes on to get a drink of water and leaves footprints on the floor. Later, Taro walks back across the floor with his shoes on to go outside, leaving two sets of footprints on the floor.” As in Experiment 2, in the emotional expression condition, “a picture associated with the initial emotional state of the protagonist” and “a picture associated with the final emotional state of the protagonist” were added to the six original pictures for each scenario. Procedure All of the steps in this experiment were reviewed and approved by the affiliated university’s committee for ethical research involving human participants. As with Experiments 1 and 2, written consent was obtained from a principal of each nursery school, and verbal assent was obtained from the children. Participants were allocated to either the control condition or the emotional expression condition. In total, 12 children in a younger group (6 boys and 6 girls; Mage = 58.83 months) and 12 children in an older group (6 boys and 6 girls; Mage = 70.17 months) were allocated to the control conditions, whereas 12 children in a younger group (6 boys and 6 girls; Mage = 57.58 months) and 12 children in an older group (6 boys and 6 girls; Mage = 68.92 months) were allocated to the emotional expression conditions. No significant differences in months of age by condition or gender were seen in either age group. Each participant was interviewed individually in a quiet room. The interview lasted approximately 15 min. In both conditions, participants were asked to perform three counterfactual tasks. The order of presenting the three counterfactual tasks was counterbalanced among participants. In both conditions, task procedures were the same as those in Experiment 2 except for using the new scenarios. Results ~~~~~~~ Preliminary analyses did not demonstrate any gender difference. So, the genders were mixed in the following analyses. Understanding the contents of scenarios Overall, the percentage of correct answers to the control questions was 96.5%. In the control condition, the percentages of correct answers were 97.2% for the Now question, 95.8% for the Before question, and 96.5% overall. In the emotional expression condition, the percentages of correct answers were 94.4% for the Now question, 98.6% for the Before question, and 96.5% overall. Some children gave only one wrong answer (5 participants in the control condition and 5 participants in the emotional expression condition). Therefore, as in Experiments 1 and 2, they were not excluded from the sample. Answers to the counterfactual questions The counterfactual scores in each condition were calculated using the same scoring procedure as in Experiment 1. Table 3 shows the mean scores and the numbers and percentages of children receiving scores of 0, 1, 2, and 3 by age and condition. First, an ordinal regression analysis was conducted. This analysis included age group (younger or older) and condition (control or emotional expression) as categorical predictors and the Age Group × Condition interaction. This model was significant, χ2(3) = 9.57, p = .023. The older group had only a marginally higher score compared with the younger group, B = −1.49, SE = 0.78, Wald χ2(1) = 3.66, p = .056, odds ratio = 4.45, 95% CI [−3.02, 0.04]. In addition, the main effect of the condition was significant; participants in the emotional expression condition had a significantly higher score than those in the control condition, B = −0.58, SE = 0.78, Wald χ2(1) = 4.06, p = .044, odds ratio = 4.85, 95% CI [−3.11, −0.04]. The Age Group × Condition interaction was not significant (p = .466). Second, a one-sample Wilcoxon signed-rank test using a counterfactual score was performed by age group and condition (chance = 1.00). In the control condition, the scores of both the younger group (Z = −0.632, p = .527, r = .18) and the older group (Z = 0.80, p = .426, r = .23) were no different from chance. In the emotional expression condition, the score of the younger group was no different from chance (Z = 1.00, p = .317, r = .29), but the score of the older group was higher than chance (Z = 2.92, p = .004, r = .85). Moreover, there were no differences between the above results and the results of analyses using children’s actual response. Finally, the task scores in Experiments 2 and 3 were compared by age group and condition. Mann–Whitney U tests found no significant differences in the scores between Experiments 2 and 3: the control condition (U = 133, p = .684, r = .07) and emotional expression condition (U = 108.5, p = .283, r = .18) in the middle group and the control condition (U = 148, p = .947, r = .01) and emotional expression condition (U = 125.5, p = .465, r = .14) in the older group. Discussion ~~~~~~~~~~ To resolve the remaining issues in Experiment 2, Experiment 3 examined the difference of young children’s counterfactual thinking between the control and emotional expression conditions using a scenario with only one protagonist. The results showed that participants performed worse under the control condition (percentage of correct answers = 34.7%) than under the emotional expression condition (percentage of correct answers = 55.7%); the level in the control condition was similar to chance even in the older group. On the other hand, the task performance in the emotional expression condition was better than in the control condition; in particular, the performance of the older group (percentage of correct answers = 69.3%) was higher than chance. A comparison of Experiments 2 and 3 showed no differences in young children’s task performance. The Experiment 3 results confirm that the results of Experiment 2 were not caused by an artifact of the task (e.g., a different number of characters) but rather were caused by a protagonist’s emotional response to the event. In sum, the results of Experiment 3 strengthened the claim that 5- and 6-year-olds can think of counterfactuals for events associated with physical change as long as other persons’ negative emotional responses to the event are presented.","One of the main themes in research on the development of counterfactual thinking involves the question of when the capacity for counterfactual thinking is acquired. To address this question, the current study examined young children’s counterfactual thinking about physical and emotional events using discriminating counterfactual tasks that children could not easily answer correctly via basic conditional reasoning. Experiment 1 showed that, although children younger than 5 or 6 years can generate counterfactual alternatives for emotional events, even 5- and 6-year-olds have difficulty in thinking counterfactually about physical events. Moreover, Experiment 2 showed that presenting with the protagonist’s negative feelings about an event outcome (the emotional expression condition) enables 5- and 6-year-olds to think counterfactually about physical events. Finally, Experiment 3 confirmed that the better performance of young children under the emotional expression condition was not the result of an artifact of the task. As mentioned in the Introduction, Rafetseder et al. (2010, 2013) have shown that 5- and 6-year-olds find it difficult to think of counterfactual alternatives using discriminating counterfactual tasks. The current study demonstrates that the performance of tasks associated with physical events does not significantly exceed chance even among children aged 5 and 6 years (percentages of correct answers: physical event task in Experiment 1 = 27.7%, control condition in Experiment 2 = 38.7%, control condition in Experiment 3 = 41.7%). The results of this study support those of Rafetseder et al. (2013). However, the current study also shows that 5- and 6-year-olds can think of counterfactual alternatives as long as there is an explicit reference to the protagonist’s negative feelings about the event (percentages of correct answers: emotion event task in Experiment 1 = 75.0%, emotional expression condition in Experiment 2 = 60.0%, emotional expression condition in Experiment 3 = 69.3%). These results demonstrate that children aged 5 and 6 years are capable of counterfactual thinking in a certain situation such as presenting the other’s emotions. In other words, the results of the current study support the claim that the ability to think counterfactually is acquired by 5 or 6 years of age (Beck et al., 2010; German, 1999; German & Nichols, 2003; Nakamichi, 2011; Riggs et al., 1998) even when discriminating counterfactual tasks are used. Why do children display their capacity for counterfactual thinking only when a scenario protagonist explicitly shows negative emotion? As mentioned in the Introduction, one possible explanation is that human negative emotion triggers the counterfactual thinking in young children by clearly conveying the outcome valence of an event. Humans’ negative emotion plays the role of clearly conveying the outcome valence of an event. Counterfactual thinking in both adults (Byrne, 2005, 2016; Epstude & Roese, 2008) and children (German, 1999; Guajardo et al., 2016) is triggered by the negative outcome valence of an event. Indeed, young children find it easier to think of a counterfactual alternative for an event associated with a change in another person’s emotions than for a physical event that does not have a potentially negative outcome (Guajardo et al., 2009; Nakamichi, 2011). For this reason, when the protagonist’s negative emotions were explicitly presented, the young children in this study should demonstrate their capacity for counterfactual thinking. This explanation may get support from other developmental studies. From their early years of life, children have sensitivity to others’ face and/or facial expressions. Even infants tend to look at face-like images longer than at any other pattern such as colored circles and concentric circles (Fantz, 1961) and react appropriately to distinct facial expressions (Tronick, Als, Adamson, Wise, & Brazelton, 1978). Moreover, children refer to other people’s emotions, and change their mode of thinking and course of action, based on other people’s emotions. For example, even 12-month-old infants go over the visual cliff when their mothers display positive expressions, but not when their mothers display negative expressions (Sorce, Emde, Campos, & Kilinnert, 1985). On the basis of these studies, it seems reasonable to assume that presenting the protagonist’s emotions gets young children’s attention, consequently improving their performances on counterfactual tasks. Of course, the results of the current study do not imply that 5- and 6-year-old children have an adult-like capacity for counterfactual thought. In line with Rafetseder et al. (2013), participants in this study found it difficult to perform physical event tasks. Rafetseder et al. (2013) argued that counterfactual thinking emerged later—at around 12 years of age. By contrast, based on the results of this study, we deem it likely that even preschool children have the ability to think counterfactually. However, we agree that children have a long way to go to be able to think of counterfactual alternatives without any prompting (e.g., the other person’s negative emotion expression) such as adult-like counterfactual thinking. In addition, this study’s participants were native Japanese preschoolers. The current study advances the previous literature by showing the development of counterfactual thinking in non-English speakers. On the other hand, this fact implies that the results might be limited to Japanese. Therefore, they need to be replicated with English speakers’ samples. Despite these limitations, the results of the current study suggest three directions for future research. One direction is to integrate the effect of emotional component into the recent studies focusing on the causal structure of events and the task demands (McCormack et al., 2018; Nyhout & Ganea, 2019; Nyhout et al., 2019). For example, Nyhout et al. (2019) manipulated the causal relation between antecedent events (e.g., two antecedent events were causally connected or disconnected to one another) in Rafetseder et al. (2013) tasks and investigated children’s counterfactual thinking. The results showed that, given events with clear causal structures, children aged 6 to 8 years could reason counterfactually. Based on our finding, children might succeed much earlier on Nyhout and colleagues’ tasks when events include the emotional component. The second direction involves determining the types of event that drive counterfactual thinking in young children. The current study has demonstrated that children’s performance on counterfactual tasks varies, depending on the content of the event (physical or emotional). Previous developmental studies (e.g., German, 1999; Guajardo et al., 2016) have focused on whether an event involves negative consequences; therefore, little is known about the influence of event content on young children’s counterfactual thinking. For instance, Guajardo et al. (2016, Experiment 1) showed that 8- to 11-year-olds’ counterfactual thinking was influenced by not only outcome valence but also outcome expectancy. Moreover, adults tend to think of counterfactual alternatives more frequently for events they can control as opposed to events they cannot control (Byrne, 2005, 2016). Young children may have the same tendency to generate counterfactual alternatives spontaneously. In addition, Beck and Riggs (2014) also suggested that the amount of knowledge of the target domain is one factor that contributes to the development of counterfactual thinking. They argued that having much causal knowledge about a particular domain promotes the thinking about alternative possibilities related to that domain. Young children may spontaneously think of counterfactual alternatives for event domains they understand well or with which they are familiar. Such research would expand further our understanding of the development of counterfactual thinking. Another direction is to clarify the relationship between counterfactual thinking and executive function (EF). EF is a cognitive process that exerts goal-oriented control over thought, behavior, and emotion. It encompasses working memory (WM), inhibitory control (IC), and cognitive flexibility as its components (Garon, Bryson, & Smith, 2008). Counterfactual thinking requires the individual to retain and update two pieces of information: “existing reality (i.e., an actual event)” and an “imagined alternative possibility.” To do this, the individual must inhibit existing real information in order to think of alternative possibilities (Beck & Riggs, 2014; Byrne, 2005, 2016). Hence, this process of counterfactual thinking requires EF. However, relationships between counterfactual thinking and EF during childhood are not stable. Some studies (Drayton, Turley-Ames, & Guajardo, 2011; Guajardo et al., 2009; Müller, Miller, Michalczyk, & Karapinka, 2007) have shown relationships between counterfactual thinking and WM or IC in young children, whereas one study (Beck, Riggs, & Gorniak, 2009) found that counterfactual thinking is unrelated to WM. EF has dual aspects (Zelazo & Carlson, 2012): the process activated in an emotionally neutral context (i.e., cool EF) and the process activated in an emotional context (i.e., hot EF). As shown in the current and previous studies (Epstude & Roese, 2008; Guajardo et al., 2009; Nakamichi, 2011), an emotional element influences counterfactual thinking in young children and adults. Thus, counterfactual thinking may require not only cool EF but also hot EF. So, future studies should examine the relationship between counterfactual thinking and hot EF. In conclusion, our study results reveal that 5- and 6-year-old children demonstrate a capacity for counterfactual thinking even when using discriminating counterfactual tasks that are difficult to answer correctly without taking the actual events into account. The current study also confirms that emotional components trigger their counterfactual thinking. These findings suggest the need to investigate the influence of emotional components and domain-specific knowledge on counterfactual thinking during early childhood."],["Investigating infants’ numerical ability is crucial to identifying the developmental origins of numeracy. Wynn (1992) claimed that 5-month-old infants understand addition and subtraction as indicated by longer looking at outcomes that violate numerical operations (i.e., 1 + 1 = 1 and 2 − 1 = 2). However, Wynn's claim was contentious, with others suggesting that her results might reflect a familiarity preference for the initial array or that they could be explained in terms of object tracking. To cast light on this controversy, Wynn's conditions were replicated with conventional looking time supplemented with eye-tracker data. In the incorrect outcome of 2 in a subtraction event (2 − 1 = 2), infants looked selectively at the incorrectly present object, a finding that is not predicted by an initial array preference account or a symbolic numerical account but that is consistent with a perceptual object tracking account. It appears that young infants can track at least one object over occlusion, and this may form the precursor of numerical ability. --------------------------------------------------------------------------------","Numeracy is a key aspect of adult cognition, and identifying its origins is vital to understanding its development during childhood and thereafter. Thus, a key area of research concerns infants’ ability to understand number. One strong claim is that young infants compute the outcomes of addition and subtraction manipulations. This was first suggested in a study by Wynn (1992). In an addition (1 + 1) condition, 5-month-old infants saw a doll being placed on a stage. A screen then concealed the doll and a hand appeared holding a second doll and placing it behind the screen. In a subtraction (2 − 1) condition, infants saw two dolls being placed on the stage followed by the screen concealing them. A hand then appeared, went behind the screen, and emerged holding one doll. On subsequent test trials, for both conditions the screen was raised, revealing either one doll or two dolls. In both conditions, infants looked longer at the impossible outcome (either 1 + 1 = 1 or 2 − 1 = 2) than at the possible outcome (either 1 + 1 = 2 or 2 − 1 = 1), with longer looking being interpreted as a violation of their expectation regarding the numerical outcome. Replications of Wynn’s findings have used both three- dimensional displays (Clearfield & Westfahl, 2006; Simon, Hespos, & Rochat, 1995; Slater, Bremner, Johnson, & Hayes, 2010; Uller, Carey, Huntley-Fenner, & Klatt, 1999; Walden, Kim, McCoy, & Karrass, 2007) and two-dimensional displays (Berger, Tzur, & Posner, 2006; Moore & Cocas, 2006). These findings are in keeping with at least three types of converging evidence: (a) that infants look more at their caregivers’ faces following unexpected arithmetic outcomes (Walden et al., 2007), (b) that infant event-related potential data show a similar pattern of activity to that of adults when observing correct and incorrect arithmetical outcomes (Berger et al., 2006), and (c) that newly hatched domestic chicks were reported to track small numbers of objects (Rugani, Fontanari, Simoni, Regolin, & Vallortigara, 2009). Collectively, these results are consistent with the larger claim that infants can perceive number (Antell & Keating, 1983; Feigenson & Carey, 2003; Feigenson, Carey, & Hauser, 2002; Lipton & Spelke, 2003; McKrink & Wynn, 2004, 2007; Xu, 2003; Xu & Arriaga, 2007; Xu & Spelke, 2000) and track numerosity of small number sets (Berger et al., 2006; Clearfield & Westfahl, 2006; Moore & Cocas, 2006; Simon et al., 1995; Uller et al., 1999; vanMarle, 2013; Walden et al., 2007). Wynn (1992) concluded that the ability to perform simple arithmetical calculations is innate and may be the foundation on which subsequent arithmetical ability builds. She argued that her results are evidence for a true symbolic number concept, favoring an accumulator mechanism (Meck & Church, 1983) as the basis on which numerical judgments are reached. A key point in this account is that a single symbol represents the number concerned. On the other hand, lower-level interpretations of Wynn’s (1992) findings are possible due to the simple nature of the dependent measure. For example, from a standpoint in which cognitive abilities such as numerical knowledge are constructed progressively during infancy (Cohen, Chaput, & Cashon, 2002), Cohen and Marks (2002) suggested an interpretation in terms of a perceptual process based on two principles: (a) a preference for familiarity (i.e., the display originally seen before the screen occluded it) and (b) a preference for displays containing a larger number of items. This interpretation is clearly important because, if correct, it would indicate that performance on Wynn’s (1992) task indicates little about infants’ ability to keep track of objects across occlusion, let alone whether infants understand operations of addition and subtraction of small numbers. Inevitably, Wynn’s conclusions will continue to attract controversy while the evidence is based on duration of looking anywhere in the display because this measure is open to lower-level interpretations such as that of Cohen and Marks (2002). Thus, it is vital to obtain a measure of infants’ response that allows a choice between low-level accounts and those based on enumerating or keeping track of objects. Even if an interpretation in terms of familiarity preference can be dismissed, it must be recognized that Wynn’s (1992) claim that infants understand the operations 1 + 1 = 2 and 2 − 1 = 1 can be questioned. It is possible that longer looking at violation outcomes is not based on infants’ realization that the specific operations 1 + 1 = 2 and 2 − 1 = 1 have been violated but rather is based on noting that an object added to the scene is not present or that an object removed from the scene is still present. An alternative to Wynn’s symbolic account is based on object file accounts derived from adult research (Kahneman, Treisman, & Gibbs, 1992; Pylyshyn & Storm, 1988; Scholl, 2001) and locates performance in tracking discrete objects and noting violations of continuity for any of these objects. For the current purposes, the key claim of this object file account is that each object is represented separately; there is no symbolic representation of the number of objects. According to Uller et al. (1999), the object file account is numerical in the sense that the system counts objects (there is one object and there is another object) but falls short of a symbolic number concept in the sense that there is no symbol for the collection of objects. In addition, the process is very much perceptual rather than based on reasoning and understanding and, thus, tends to be considered an implicit numerical system (vanMarle, 2013). Our rationale in this investigation was that the precision of eye-tracker data should allow further evaluation of the perceptual preference, object tracking, and symbolic numerical accounts through differential predictions regarding fixation patterns that would appear to arise from each account. The clearest predictions arise in the case of the subtraction violation condition in which two objects remain. If Cohen and Marks’s (2002) familiarity preference is correct, there should be equal looking at each object because both objects would be equally part of the familiar starting array. Wynn’s (1992) symbolic numerical account also predicts longer looking at both objects. The unexpected outcome following the removal of one of the original two dolls is the incorrect numerical outcome (two dolls). Thus, we would expect infants to direct increased looking to both objects because together they constitute the incorrect number. It seems likely, however, that a particularly high proportion of looking would be directed to the object that should not be there because it is at the root of the numerical violation. Still, the important point is that if infants are evaluating the symbolic numerical outcome, the object that was not subject to subtraction should also be a focus of particular attention. In contrast, in the object file account, objects maintain separate files and only one object file is violated (by its continuing presence despite its earlier removal). Thus, infants should devote a high proportion of their looking to the object that should no longer be there, but they should show no increase in looking to the other object because its file remains unviolated. Although in principle the object file or symbolic numerical account might predict longer looking at the empty location in addition violation (one toy outcome), this cannot be a strong prediction because there is nothing to fixate there and so infants’ looking is likely to be drawn to other features, particularly the one toy that is present. But it is possible that infants responding on the basis of object tracking or number violation will look elsewhere, particularly to the place where objects appear from, as if searching for the missing object or that they will look more at the empty location in the addition violation than in the correct outcome of subtraction when the position is correctly empty. Thus, here we followed Wynn’s (1992) procedure for testing infants’ responses to addition and subtraction events. Crucially, however, we gathered precise eye- tracker records of visual fixation. To ensure that we could replicate Wynn’s results, we also measured looking duration to the stage in the conventional way.","A total of 34 4- and 5-month-old infants provided usable data (Mage = 148.06 days, SD = 13.48, range = 119 − 168), 17 in the addition condition (9 boys and 8 girls; Mage = 148.76 days, SD = 14.67) and 17 in the subtraction condition (11 boys and 6 girls; Mage = 147.35 days, SD = 12.98). They were recruited from the local maternity unit with appointments arranged by follow-up phone calls. The majority of infants were White, and all were full term with no known developmental disabilities and from English-speaking families. Data from 28 additional infants could not be used because of experimenter error (n = 8), failure to obtain individual calibration of the point of gaze (n = 4), or excessive movement such that insufficient eye-tracking data were collected (n = 16). In our experience, this attrition rate is typical for the type of eye tracker we used.","The events were presented on a lit stage presenting an aperture 37 cm wide × 27 cm high × 60 cm deep in a dimly lit testing room. A 14-cm-high screen located 30 cm behind the front of the stage could be rotated to conceal the toys, and a blind could be lowered to conceal the whole stage. The objects were two 11-cm-tall toy men that squeaked when pressed. The experimenter presenting the toys wore a long maroon glove. A video camera, placed at the top center of the stage, recorded infants’ head and eyes for live recording of preferential looking and for subsequent reliability testing by a naive observer. A remote optics corneal reflection eye tracker (ASL Model 5000, Applied Science Laboratories, Bedford, MA, USA), located below and at the midpoint of the stage, was used to collect fixation data. A plasma display was mounted immediately behind the stage, and prior to testing each infant’s point of gaze was calibrated in standard fashion by presenting attention-getting videos. Procedure Infants sat either in an infant car seat or on a caregiver’s lap approximately 60 cm from the screen behind which the toys were placed. In the latter case, the caregiver’s eyes were above the stage, and the caregiver could not see the displays. A technician controlled the eye tracker, and another researcher recorded preferential looking during familiarization and test trials. After gaze calibration, the procedure followed Wynn (1992). Infants saw two pretest trials of one toy and two toys, respectively, in counterbalanced order across infants. The blind was raised to reveal either one toy or two toys, and the observer recorded looking at the toy(s). Toys were placed 35 cm behind the front edge of the stage. When one toy was presented, it was placed 7.5 cm to the right of stage midline; when two toys were presented, the other toy was 7.5 cm to left of midline. The trial continued until the infant had accumulated at least 2 s looking time and looked away from the display for 2 s or more. The blind was then lowered, and the procedure was repeated with the other number of toys. Six arithmetic trials were then presented, with each infant being tested in either the addition condition (1 + 1) or the subtraction condition (2 − 1), with trials alternating, in counterbalanced order, between the possible and impossible outcomes in terms of the number of toys revealed. The test trial sequences are illustrated in Fig. 1. In the addition condition, the experimenter’s gloved hand emerged stage left (i.e., to the infant’s left) holding a toy that the experimenter squeaked to capture the infant’s attention. She then moved the toy, still squeaking, and placed it on the right-hand location used during familiarization. She then slowly withdrew her hand, at which time the screen was raised to hide the toy. This event, from appearance of the toy to withdrawal of the hand, took approximately 5 s. The hand then reappeared from stage left, above the screen, clutching another identical squeaking toy. When she had the infant’s attention, the experimenter placed the toy in the left-hand location used during familiarization, raised her hand, clasped and unclasped it to emphasize that it was empty, and then slowly withdrew it, at which time the screen was lowered to reveal the outcome of either one toy or two toys. The period from appearance to disappearance of the experimenter’s hand took approximately 6 s. To replicate Wynn’s (1992) procedure closely so as to be able to interpret our eye-tracking data relative to her original findings, and because we had not detected side preferences in other work with this age group, we did not counterbalance the side from which objects were manipulated, and so the impossible event (presence or absence) always concerned the left-hand object. We were comfortable with this decision for two reasons. First, a looking bias to one side would be revealed in our analyses. Second, our primary analyses concerned comparisons between addition and subtraction conditions in looking to each respective side and, thus, would not be affected by an overall side bias. In the subtraction condition, the experimenter placed two squeaking toys consecutively on the stage, an event that took approximately 9 s. Following the raising of the screen, her empty hand reappeared above the screen, lowered to the left-hand floor of the screen, and reappeared holding one toy that she squeaked above the screen and withdrew screen left, an event that took approximately 6 s, followed by lowering of the screen to reveal the outcome of one toy or two toys. Between each trial in each condition, the roller blind was lowered to obscure the stage. The impossible outcome (either 1 + 1 = 1 or 2 − 1 = 2) was accomplished by silent removal (addition condition) or addition (subtraction condition) of a toy from the left-hand floor of the screen; each of the toys was mounted on a velvet-covered disk to ensure that the addition or removal of the toy was not audible to either the infant or the observer. The observer who recorded infants’ looking was aware of which condition (addition or subtraction) the infant was in but was unaware on each test trial of whether the outcome was possible or impossible. Preferential looking (violation of expectancy) data were recorded on a Mac G4 using Habit software (Cohen, Atkinson, & Chaput, 2004). A total of 27 infants’ data, for both the pretest and test trials, were independently scored by a second observer from the video records, and interobserver reliability was high (r = .986, p < .001). Preferential looking data ~~~~~~~~~~~~~~~~~~~~~~~~~ We replicated Wynn’s (1992) results very clearly. On the test trials, the infants looked at the unexpected/impossible outcome for a mean of 17.82 s (SD = 8.19) and at the expected/possible outcome for a mean of 10.36 s (SD = 5.21). A total of 31 infants looked longer at the impossible outcome, and 3 looked longer at the possible outcome (binomial, p < .0001). Of the 17 infants in the addition condition, 16 looked longest at the impossible outcome (p < .001), and 15 of the 17 infants in the subtraction condition did so as well (p < .001). Analysis of variance (ANOVA) performed on the data confirmed significantly longer looking at the impossible outcomes than at the correct outcomes, F(1, 32) = 55.00, p < .001, ηp2 = .63. This effect was qualified by a 2 (Condition: addition or subtraction) × 2 (Test Trial: expected or unexpected) × 3 (Trial Block) interaction, F(2, 31) = 4.50, p = .019, ηp2 = .23. This effect stems from a slight increase in looking at expected outcomes accompanied by a slight decrease in looking at unexpected outcomes across the first and second trial blocks by infants in the addition condition; these trends were reversed in the subtraction group. There were no other reliable main effects or interactions. Eye-tracker data ~~~~~~~~~~~~~~~~ Using Applied Science Laboratories’ Eyenal software, we reduced the raw data to a list of fixations, and the data analyzed consisted of dwell times in areas of interest (AOIs) that comprised the regions surrounding the two men on the stage. Preliminary analysis of eye- tracker data indicated that the vast majority of looks were to the location of the toy man/men that was/were present on the stage, and there was little looking elsewhere—for instance, at the location from which the hand emerged—and no evidence for different patterns of looks in these regions (top left and right quadrants) depending on condition. Thus, we concentrated on the looking times to AOIs surrounding the locations of the two men. These AOIs measured 18.5 cm horizontally and 13.5 cm vertically and, thus, corresponded to the lower left and right quadrants of the stage aperture. Although infants typically focused on the toys when they were visible, relatively large AOIs were necessary to detect looking toward an empty object position whose exact location might be uncertain to the infant. The raw dwell times for these two AOIs, accumulated for each trial, were converted to proportions of total dwell times recorded for each infant. We performed separate analyses of these data for the two-object outcomes (Fig. 2) and the one-object outcomes (Fig. 3). For the two-object outcome, there was a reliable main effect of condition, F(1, 18) = 5.32, p = .033, ηp2 = .23, the result of longer dwell times overall by infants in the subtraction condition, consistent with the looking time data reported previously, that is, longer looking at the impossible two-object outcome (see Fig. 2). There was also a reliable main effect of position, F(1, 18) = 4.86, p = .041, ηp2 = .21, qualified by a significant Condition × Position interaction, F(1, 18) = 6.31, p = .022, ηp2 = .26. This effect stemmed from longer looking toward the left position than toward the right position by infants in the subtraction condition, F(1, 9) = 7.48, p = .023, ηp2 = .45, but not by infants in the addition condition, F(1, 9) = 0.09, p = .77, ηp2 = .01. In addition, infants looked longer at the left man when it should not be there (subtraction) than when it should be there (addition), F(1, 33) = 9.76, p = .006, ηp2 = .35. In summary, infants in the subtraction condition showed particularly long dwell times to the most recently manipulated left man that should have been absent. This was confirmed by analysis of the number of infants who looked longer at the left man than at the right man. Whereas 10 of 17 infants (binomial p = .63) looked longer at the left man than at the right man in the addition condition, 15 of 17 infants (binomial p = .002) looked longer at the left man than at the right man in the subtraction condition. A 2 × 2 chi-square test on these data confirmed a larger number of infants looking at the left man in the subtraction condition compared with the addition condition, χ2(1) = 3.78, p = .026. For the one-object outcome, there was a reliable main effect of condition, F(1, 18) = 8.78, p = .008, ηp2 = .33, the result of longer dwell times overall by infants in the addition condition, consistent with the looking time data reported previously, that is, greater looking overall at the impossible one-man outcome (see Fig. 3). This effect was qualified by a reliable Condition × Familiarization Order interaction, F(1, 18) = 6.85, p = .017, ηp2 = .28, which stemmed from longer dwell times overall by infants in the addition condition who were first familiarized with the two-man event relative to the one-man event (p = .031) (the reasons for this effect are unclear); the dwell time difference for infants in the subtraction condition was not significant (p = .10). There was also a reliable main effect of position, F(1, 18) = 15.96, p = .001, ηp2 = .47, due to longer dwell times toward the right-man position (M = 33.34 s, SD = 20.33) than toward the (empty) left-man position (M = 7.84 s, SD = 11.11). The interaction between condition and position was not statistically significant, F(1, 18) = 0.96, p = .42, ηp2 = .04; in other words, there was no reliable difference between conditions in looking duration to either the left-man position or the right-man position. Significantly more infants looked longer at the right man than at the left man in both the addition condition (binomial p = .02) and the subtraction condition (binomial p = .013). The lack of longer looking to the left- hand location in the addition violation condition is unsurprising, given that there was no object to look at in that location. However, it is possible that during addition violation trials infants initially looked there on detecting an empty place. To check this, we analyzed data for the first look infants made after the screen lowered sufficiently to reveal the top of the man or men. This analysis revealed broadly the same pattern as the dwell time analysis; totaled across test trials, in the case of one-object outcomes, the majority of first looks were to the right-hand position (addition violation: left = 5, right = 35; subtraction correct: left = 3, right = 35), with no difference in left- position looks between these conditions. It is also possible that infants faced with addition violation looked with high frequency to the empty left-hand location but looked away quickly on seeing nothing there. Thus, we compared mean number of looks across trials to the empty left position and the occupied right position for the same conditions as above (addition violation: left M = 5.88, SD = 5.32, right M = 10.76, SD = 7.58; subtraction correct: left M = 3.94, SD = 3.03, right M = 8.24, SD = 4.87). Although infants looked more often at the empty left location in the addition violation case than at the subtraction correct case, this difference did not reach significance, t(32) = 1.31, p = .20.","Our key finding from use of the eye tracker is that infants looked reliably longer at the left-hand man in the subtraction violation two-toy outcome than in the addition correct two-toy outcome. In the subtraction violation condition, they also looked reliably longer at the left-hand man than at the right-hand man. In other words, infants looked particularly long at an object that they had just seen removed from the occluded scene. This result is not explained by Cohen and Marks’s (2002) interpretation, namely that infants look longer at the outcome that matches the initial array before the numerical manipulation. This account would not predict differential looking in the subtraction two- man outcome because it posits that infants are looking preferentially at a two-object array in which both objects are equally expected. Similarly, our results do not provide strong support for a symbolic numerical account. If infants noted a violation of number, one would expect increased attention to both objects, not just the one most recently manipulated. A third interpretation, which our data support, is that infants can track the existence and locations of objects for brief intervals of occlusion. Object tracking/object file accounts would predict that they “index” both objects in this task but that their attention is directed to the object whose “file” has been violated. This is still a higher-level account than Cohen and Marks presented, and it can be argued that noting a violation of a movement of a single object actually constitutes the fundamental addition or subtraction operation, namely 0 + 1 = 1 or 1 − 1 = 0. Our findings indicate clearly that eye-tracker data can be used to compare alternative interpretations of the findings from violation of expectancy (VoE) studies with infants, and it is our view that eye tracking could also clarify the interpretation of results arising from a number of key VoE experiments investigating different aspects of object knowledge during infancy. For instance, Baillargeon (1986) demonstrated that 6- to 8-month-olds who had seen an obstruction placed in an object’s path behind an occluder looked longer at object reemergence than when there was no obstruction, evidence that they represented the hidden obstruction and understood that the object could not pass through it. Confidence in this conclusion would be strengthened if infants also showed reduced anticipatory tracking (Johnson, Amso, & Slemmer, 2003) when an obstruction was present. In addition, our findings point to a possible perceptual basis for infants’ awareness of number, an account that carries the advantage of avoiding many of the theoretical problems regarding the concept of innateness (Haith, 1998). In the subtraction condition, infants’ attention in the impossible test outcome—two toys on stage—was directed largely at the one toy that “should not be there.” Thus, performance may be accounted for on the basis of tracking presence or absence of one object at one location, namely the location where the toy was added or taken away. Thus, even relatively parsimonious object tracking accounts of small number judgments (e.g., vanMarle, 2013) may be more complex than the findings from this task demand. Although a representation of the other object may be formed (Uller et al., 1999), its presence may act primarily as a referent for the arithmetic operation constituted by the physical manipulation of the added or subtracted object. This possibly forms the initial implicit perceptual basis for later arithmetic abilities. Specifically, prediction of the outcome of tracking a single object across occlusion effectively consists of adherence to the principles that 0 + 1 = 1 and 1 − 1 = 0. In this respect, this account is consistent with the argument that awareness of object permanence develops from perception of object persistence across occlusion (Bremner, Slater, & Johnson, 2015). We know that young infants’ perception of object persistence across occlusion is limited to short spatiotemporal gaps in perception (Bremner et al., 2005; Johnson et al., 2003), and it is likely that this same perceptually constrained process operates in Wynn’s (1992) task, such that young infants form a perceptual expectation about the persistence of an added object when the screen is lowered. The additional process revealed in the current work is that infants apparently track an object off the scene and form a perceptual expectation of its absence behind the screen, an expectation that is violated when it is revealed remaining in its original location. In summary, to our knowledge this is the first study to derive eye-tracking data from a task involving addition and subtraction of objects from a three-dimensional scene. The clearest results were obtained in the subtraction violation condition, where infants directed particular attention specifically to the object that should no longer be there. Selective attention of this sort is not predicted by a low-level account based on familiarity preference (Cohen & Marks, 2002). However, the fact that there was no increase in looking to the object that was not subject to the subtraction operation does not support a symbolic numerical account, according to which detection of a numerical violation should lead to an increase in attention to both objects in the outcome scene. Our results are more closely in keeping with an object file account in which each object is tracked separately, such that attention is directed only to the object whose file is violated. We favor the view that processing at this level forms a precursor of symbolic numerical ability, which may well develop through the constructionist processes advanced by Cohen and Marks (2002)."],["Emergency service workers, military personnel, and journalists working in conflict zones are regularly exposed to trauma as part of their jobs and suffer higher rates of posttraumatic stress compared with the general population. These individuals often know that they will be exposed to trauma and therefore have the opportunity to adopt potentially protective cognitive strategies. One cognitive strategy linked to better mood and recovery from upsetting events is concrete information processing. Conversely, abstract information processing is linked to the development of anxiety and depression. We trained 50 healthy participants to apply an abstract or concrete mode of processing to six traumatic film clips and to apply this mode of processing to a posttraining traumatic film. Intrusive memories of the films were recorded for 1 week and the Impact of Events Scale-Revised (IES-R; Weiss & Marmar, 1997) was completed at 1-week follow-up. As predicted, participants in the concrete condition reported significantly fewer intrusive memories in response to the films and had lower IES-R scores compared with those in the abstract condition. They also showed reduced emotional reactivity to the posttraining film. Self-reported proneness to intrusive memories in everyday life was significantly correlated with intrusive memories of the films, whereas trait rumination, trait dissociation, and sleep difficulties were not. Findings suggest that training individuals to adopt a concrete mode of information processing during analogue trauma may protect against the development of intrusive memories. --------------------------------------------------------------------------------","A power analysis based on the effect sizes from a similar study by Schartau et al. (2009) was used to determine the sample size. Results showed that a sample size of 25 in each condition would have 80% power to detect a significant difference in mean change scores between the abstract and concrete conditions using a two-group t test with a two-tailed .05 significance level. Therefore, we concluded that a total of 50 participants would be needed. Participants over the age of 18 were recruited from King’s College London through an email circular. We excluded participants who self-reported a mental health problem, scored 10 or higher on a standardized measure of depression, or scored 33 or higher on a standardized measure of posttraumatic stress. We excluded these individuals for three reasons: (a) as an ethical measure to reduce distress in already vulnerable individuals; (b) to increase homogeneity of participants across the two conditions; and (c) to replicate the samples of studies that have previously used the trauma film paradigm (e.g., Ehring et al., 2009; Schartau et al., 2009). A total of 2 participants were excluded for scoring above cutoff on the screening measures, and 1 participant was excluded for failure to follow the instructions during the posttraining test film. This left 50 participants: 26 were randomly assigned to the abstract condition (15 female, Mage = 27.15, SD = 9.11), and 24 to the concrete condition (13 female, Mage = 24.71, SD = 5.35). Conditions were comparable on age, t(48) = 1.15, p = .258) and sex, χ2(1) = 0.06, p = .802. Ethical approval for the study was granted by the Psychiatry, Nursing and Midwifery Research Ethics Committee at King’s College London. Demographic Questionnaire (Unpublished) A brief demographic questionnaire was included to collect relevant demographic information to ensure equivalence of conditions on age, gender, and driving frequency (since most of the films contained exposure to RTA footage). Patient Health Questionnaire–9 (Kroenke, Spitzer, & Williams, 2001) Depression was assessed using the Patient Health Questionnaire–9 (PHQ-9; Kroenke, Spitzer, & Williams, 2001), a brief depression severity measure consisting of 9 items. Scores range from 0 to 27, with higher scores indicating greater depression severity. This measure has demonstrated high internal consistency (α = .86–.89), good test–retest reliability (r = .84), and correlates significantly with other standardized measures of depression. The recommended cutoff of ≥ 10 (Kroenke et al., 2001) was used to indicate major depression in this study. Internal consistency for the current study was modest (Cronbach’s α = .50). Impact of Event Scale–Revised (Weiss & Marmar, 1997) Posttraumatic stress symptoms (both prior to participation and at 1-week follow- up) were assessed using the Impact of Event Scale–Revised (IES-R; Weiss & Marmar, 1997), a widely used measure of PTSD severity, consisting of 22 items. Scores range from 0 to 88, with higher scores indicating greater severity of PTSD symptoms. The scale has shown high internal consistency (α = .96; Creamer, Bell, & Failla, 2003) and test–retest reliability coefficients that range from r = .51 to .94 for the individual subscales (Weiss & Marmar, 1997). Concurrent validity has been documented, with the scale being highly correlated (r = .84) with the Posttraumatic Checklist (Blanchard, Jones-Alexander, Buckley, & Forneris, 1996), an established measure of PTSD (Creamer et al., 2003). We used a cutoff of ≥ 33 to indicate clinically significant levels of PTSD, as recommended by Creamer et al. (2003). High internal consistency was found for the current sample both for preexisting symptom levels (α = .87) and at 1-week follow-up (α = .89). Trauma Screener (Unpublished) Participants also completed the Trauma Screener (unpublished), a self-report inventory of prior exposure to traumatic events, which has been used in previous studies (e.g., Ehlers, Mayou, & Bryant, 1998) and is derived from the trauma checklist included in the Clinician-Administered PTSD Scale (Blake et al., 1990). It includes a checklist of 16 events such as serious traffic accidents, serious other accidents (e.g., fire or explosion), sexual or nonsexual assault, and others. It was used to establish an index traumatic event for completing the IES-R and to ensure that the event was one in which the person “experienced, witnessed or was confronted with an event or events that involved actual or threatened death or serious physical injury, or a threat to the physical integrity of self or others” (Diagnostic and Statistical Manual [DSM-IV]; APA, 1994). State–Trait Anxiety Inventory–Trait Version (Spielberger, Gorsuch, Lushene, Vagg, & Jacobs, 1983) The State–Trait Anxiety Inventory–Trait Version (STAI-T) was administered to assess equivalence of conditions in proneness to anxiety. This is a widely used self-report measure that assesses tendency to feel anxious in response to stressful situations (Spielberger et al., 1983). It contains 20 items, with scores ranging from 20 to 80, and higher scores indicating higher levels of trait anxiety. It has demonstrated high internal consistency (median α = .90), acceptable test–retest reliability ranging from r = .73 to r = .86 in nonclinical samples, and it correlates with other trait anxiety measures (Spielberger et al., 1983). Internal consistency for the current study was high (α = .87). Intrusions Diary (Unpublished) Number of intrusive memories experienced in the week following the experiment was measured using an intrusions diary, a standard way of assessing the frequency of intrusions (Holmes & Bourne, 2008). Participants were asked to record daily the number of times they had experienced spontaneously occurring intrusive memories of any of the films. For each intrusion they had, they were asked to record which film it referred to (to check that intrusions corresponded to films viewed in the study). Perseverative Thinking Questionnaire (Ehring et al., 2010) Trait rumination was measured using the Perseverative Thinking Questionnaire (PTQ; Ehring et al., 2010), a content-independent self-report measure of repetitive negative thinking (including rumination). It contains 15 items, with scores ranging from 0 to 60, with higher scores indicating a greater tendency toward repetitive negative thinking. Excellent internal consistency (α = 0.94–0.95) but limited test–retest reliability (r = 0.69) have been reported. It demonstrates convergent validity, correlating significantly with standard measures of rumination and worry (ranging from r = .48 to r = .72; Ehring et al., 2010). It also shows concurrent validity, with significant and moderate correlations found with severity of anxiety (r = .64) and depressive (r = .54–.58) symptoms (Ehring et al., 2010). Internal consistency was high for the current sample (α = .94). Trait Dissociation Questionnaire–Short Version (TDQs; Murray, 1997) Trait dissociation was measured using the Trait Dissociation Questionnaire–Short Version (TDQs; Murray, 1997), a measure of trait disposition to dissociate. It contains 11 items, with scores ranging from 0 to 55, with higher scores reflecting greater trait dissociation. The measure has been validated using an outpatient sample and correlates highly with the original TDQ (r = 0.94), which has been shown to predict intrusive memories in a student population. It evidences high internal consistency (α = 0.86) but limited test–retest reliability (r = 0.56; Murray, 1997). Internal consistency for the current sample was acceptable (α = .74). Proneness to Intrusive Memories Scale (PIMS; Unpublished) To assess individual proneness to intrusive memories, we administered a 5-point scale asking participants to rate the frequency with which they tended to experience intrusive memories of stressful or unpleasant events ranging from 0 (not at all) to 5 (five times a week or more). Insomnia Severity Index (ISI; Morin, Barlow, & Dement, 1993) To assess preexperiment sleep difficulties, we administered the Insomnia Severity Index (ISI; Morin, Barlow, & Dement, 1993), a brief 7-item screening measure for insomnia severity over the previous 2 weeks. Total scores range from 0 to 28, with higher scores indicating higher severity of insomnia. The measure has been validated on younger and older populations attending a sleep clinic or receiving treatment for insomnia, and has demonstrated adequate internal consistency (α = 0.74–0.78). It has also demonstrated concurrent validity, correlating significantly with other measures of insomnia (sleep diaries, polysomnography, and informant and clinician reports), and sensitivity in detecting change in perceived sleep difficulties following treatment (Bastien, Vallières, & Morin, 2001). Internal consistency was high for the current sample (α = .86).","Informed consent was obtained for all participants before taking part in the study. Participants completed the pretraining measures (PTQ, TDQs, PIMS, and ISI) and were randomly allocated to the abstract or concrete training condition to complete the film task. They were told that this involved watching a series of film clips of real-life traumatic footage, which they would be asked to watch in a particular way according to some instructions they would be given. The film task was adapted from Schartau et al. (2009). The two test films (shown pre- and posttraining) contained real-life footage of the aftermath of road traffic accidents (RTAs) and showed emergency service personnel working to extract trapped victims, severely injured individuals in distress, and dead bodies being moved. These were selected from a series of films containing footage of RTAs compiled by Steil (1997), which has been used in previous studies (e.g., Holmes et al., 2004; Murray, 1997; Steil, 1997). The series of films was piloted on a small sample of student volunteers (N = 10) from King’s College London who met the study’s eligibility criteria, to establish which two films evoked the most distress and hence would best serve as a pre- and posttraining test films. The two films selected were rated as the most distressing (> 55 on a 0–100 scale) and were comparable in the level of distress they evoked (pretraining film: M = 55.30, SD = 27.81; posttraining film: M = 59.70, SD = 23.45; t(9) = 1.346, p = .211). Three additional RTA films were selected from this series as practice films. The remaining three practice films were borrowed with permission from Schartau et al. (2009) and included non-RTA scenes of violence involving animals and humans. We selected both RTA and non-RTA films so that participants could practice on a range of stimuli. Films ranged from 21 to 137 seconds (M = 89.00, SD = 35.88). Each film contained a short introduction at the beginning explaining what each depicted. Films were shown on a 15-in. computer screen and the film task followed six steps (see Figure 1). First, participants were asked to rate their current affect on a scale ranging from 0 (extremely negative) to 100 (extremely positive). Second, they were shown the pretraining test film and asked to simply watch it as they would normally, before rating their emotional reactions to it. Emotional reactivity was indexed by two scales ranging from 0 to 100, which measured distress (ranging from no distress to extreme distress) and horror (ranging from no horror to extreme horror). Distress and horror were chosen as they were the target emotions used by Schartau et al. (2009), who demonstrated an effect of appraisal training on these emotions in reaction to distressing films. They also noted that generalized emotion terms (e.g., distress and negative emotion) have been demonstrated as reliable indices of emotion change in response to experimental manipulations (Richards & Gross, 2000) and bear similarity to the Subjective Units of Distress measure, which is often used in clinical research and practice (e.g., Dalgleish & Yiend, 2006), as well as being emotions historically elicited by traumatic experiences. Participants also rated the degree of personal relevance of the pretraining test film. They were asked, “How much personal relevance did this film have for you?” with the scale ranging from 0 (none) to 100 (extreme). This was included to assess any differences between the conditions in the extent to which the films were personally relevant. Third, participants were given instructions for how to watch the subsequent films and shown the six practice films. Instructions were presented verbally at first, as well as being presented on the computer screen before each practice film. After each film, the word “relax” appeared on the screen for 5 seconds, which aimed to minimize any accumulative effect of the training phase on mood. Participants in the abstract condition received the following instructions: “When watching the films, please focus on: (1) Why these sorts of things happen; (2) What it means about the world; (3) What it means for the people involved; (4) What if this were to happen to you, or someone in your family?” For participants in the concrete condition, the instructions were as follows: “When watching the films, please focus on: (1) The specific and objective details of the event, for example, what you can see, what you can hear; (2) The sequence of events as they are unfolding; (3) What needs to happen step by step from here.” Following Schartau et al. (2009), before the first practice film, participants were given examples of how they might apply the assigned processing mode to it, before practicing applying it on their own. For the remaining films, examples were not given beforehand although after each of the first three films, participants were asked to give examples of their thoughts while watching the film so that the investigator could be sure that they understood the instructions and that their thoughts reflected their assigned processing mode. Where participants were clearly not applying the required mode of processing or their feedback reflected thoughts that were inconsistent with that mode of processing, the instructions were repeated with further clarification and examples where necessary. For the remaining three films, participants were simply asked whether they thought they had successfully applied the required processing mode and prompted to continue doing so. Fourth, participants were asked to rate their affect for the second time on the scale ranging from 0 (extremely negative) to 100 (extremely positive).2 Fifth, participants were shown the posttraining test film with instructions presented onscreen telling them to watch it in the way that they had been practicing so far, followed by the assessment of emotional reactivity (i.e., self-report ratings of distress and horror in response to the final film) and personal relevance ratings. Sixth, a manipulation check was carried out whereby participants were asked to rate their level of adherence to the instructions on a scale, which asked: “To what extent did you watch the film according to the instructions given to you?” with responses ranging from 0 (none of the time) to 100 (all of the time). In addition, their level of attention to the film was assessed by asking: “To what extent did you pay attention to the film?” with a scale of responses ranging from 0 (none of the time) to 100 (all of the time). Participants were excluded if their level of adherence or attention was less than 50%, which resulted in one participant being excluded from the analysis. Afterwards, participants were given the intrusions diary and asked to record any spontaneous intrusive memories of any of the scenes they saw during the films for the following week. The experimenter checked that participants were not showing significant signs of distress before leaving the session and provided participants with one of the researcher’s contact details should they wish to discuss their experience of the study. Participants posted the diary back 1 week later following an email prompt, which contained a link to an online version of the IES-R that they were asked to complete. They were also asked to rate to what extent they completed the diary reliably and accurately on a 5-point scale, with responses ranging from 0 (never) to 4 (all of the time). Participants were given a payment of £15 as compensation for their time. Premanipulation Group Differences ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 1 shows descriptive statistics for demographic and baseline variables by condition. Independent samples t tests revealed no differences between the conditions in levels of depression, preexisting PTSD symptoms, number of previous traumas, trait anxiety, trait rumination, trait dissociation, sleep difficulties, proneness to intrusive memories, baseline affect, personal relevance ratings, or emotional reactivity ratings for the pretraining film (all ps > .05). A chi-squared test revealed that conditions were also comparable on driving frequency, χ2(1) = 0.24, p = .877. Instruction and Diary Compliance ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Participants generally reported paying attention to the posttraining test film (M = 92.10, SD = 10.84) and adhering to the instructions (M = 85.90, SD = 10.43), and there were no differences between the conditions on these measures (attention: t[48] = 0.02, p = .983; adherence: t[48] = 0.31, p = .757). Participants generally reported completing the diary reliably and accurately most of the time (M = 3.79, SD = 0.41) and there were no between- group differences, t(48) = 0.63, p = .534. Effect of Processing Mode Training on Affect ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To investigate the effect of training on affect following the practice films, we conducted a 2 (Condition: abstract, concrete) × 2 (Time: pretraining, posttraining) mixed model analysis of variance (ANOVA), with Condition as the between-subjects factor and Time as the repeated measures factor. Results indicated a significant main effect of Time, F(1, 37) = 104.36, p < .001, ηp2 = .74, which was qualified by a significant Condition × Time interaction, F(1, 37) = 5.26, p = .028 ηp2 = .12 (see Figure 2). Paired samples t tests confirmed that participants in both conditions rated their affect as more negative from pre- to posttraining (abstract: t[18] = 8.22, p < .001, d = − 2.25, 95% CI [28.21, 47.58]; concrete: t[19] = 6.06, p < .001, d = − 1.62, 95% CI [15.71, 32.30]). A t test confirmed that the decrease in affect was greater in the abstract condition than in the concrete condition, t(37) = − 2.29, p = .028, d = 0.73, 95% CI [− 26.17, − 1.619]. Effect of Processing Mode Training on the Development of Intrusive Memories and IES-R Scores ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Of the whole sample, 47 participants (94%) reported at least one intrusive memory relating to the films. Figure 3 shows the number of intrusions and IES-R scores by condition. We conducted independent samples t tests to test our prediction that participants in the concrete condition would experience fewer intrusions during the following week and lower IES-R scores at 1-week follow-up. As predicted, participants in the concrete condition reported significantly fewer intrusive memories over the following week than those in the abstract condition, t(48) = 2.07, p = .044, d = 0.59, 95% CI [1.69, 4.53]. Likewise, participants in the concrete condition reported significantly lower IES-R scores than those in the abstract condition at 1-week follow-up, t(48) = 2.78, p = .009, d = 0.78, 95% CI [1.69, 10.95]. We conducted bivariate correlational analyses to assess whether number of intrusions or IES-R scores were related to change in affect over the training session. These showed that there was no significant relationship between change in affect from pre- to posttraining and number of intrusive memories, r = − .20, p = .229, or between affect change and IES-R score, r = − .27, p = .095. Effect of Processing Mode Training on Emotional Reactivity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Figure 4 shows change scores in distress and horror ratings from pre- to posttraining by condition. To investigate the effect of training on emotional reactivity to the posttraining test film, we conducted a 2 (Condition) × 2 (Time) multivariate analysis of variance (MANOVA) on distress and horror ratings. Results indicated a significant multivariate Condition × Time interaction, F(1, 47) = 3.50, p = .038, ηp2 = .13, which allowed us to interpret the univariate ANOVAs. Univariate ANOVAs revealed a significant Condition × Time interaction for distress ratings, F = (1, 48) = 5.95, p = .018, ηp2 = .11, and horror ratings, F = (1, 48) = 5.03, p = .030, ηp2 = .10. This suggests that concrete training led to less emotional reactivity to the posttraining test film relative to abstract training. Paired samples t tests comparing pre- and posttraining distress scores showed a significant increase in subjective distress ratings from pre- to posttraining test films in the abstract condition, t(25) = − 5.51, p < .001, d = 0.68, 95% CI [− 21.93, − 10.00], but no significant increase in the concrete condition t(23) = 0.00, p = 1.000, d = 0.00, 95% CI [− 12.51, 12.51]. Similarly, for horror ratings, a significant increase was found in the abstract condition, t(25) = − 2.87, p = .008, d = 0.44, 95% CI [− 19.50, − 3.20], compared with no significant increase in the concrete condition, t(23) = 0.65, p = .520, d = − 0.12, 95% CI [− 7.67, 14.76]. These findings suggest that abstract training led to both increased distress and horror reactions to the posttraining test film, whereas concrete training did not. There was a significant correlation between change in affect and change in distress ratings, r = − .32, p = .047, with decreased affect from pre-to posttraining being associated with increased distress ratings. Since participants in the abstract condition showed a greater decrease in affect from pre- to posttraining than those in the concrete condition, we conducted a mixed model analysis of covariance (ANCOVA) to investigate whether increases in distress in the abstract condition could be attributed to decreases in affect rather than to mode of processing. However, when change in affect from pre- to posttraining was included as a covariate, the effect of condition on changes in distress ratings remained, F(1, 36) = 6.80, p = .013, ηp2 = .16, indicating a significant increase in distress ratings in the abstract but not the concrete condition (abstract: t[25] = − 5.51, p < .001, d = 0.68, 95% CI [− 21.93, − 10.00], concrete: t[23] = 0.00, p = 1.000, d = 0.00, 95% CI [− 12.51, 12.51]). There was no significant correlation between change in affect and change in horror ratings, r = − .21, p = .191. Relationship Between Predictor Variables and Intrusive Memories ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We conducted correlational analyses to explore relationships between predictor variables (trait rumination and dissociation, proneness to intrusive memories, and sleep difficulties) and primary outcome measures (number of intrusions and IES-R scores). None of these were significant apart from proneness to intrusive memories, which was significantly associated with the development of intrusive memories following the films. Since greater proneness to intrusive memories following negative events in everyday life was associated with a higher number of reported intrusions of the films, r = .32, p = .025, we conducted a univariate ANCOVA to investigate the effect of condition on intrusive memories while controlling for participants’ proneness to intrusive memories. There was a significant effect of condition on intrusions after controlling for the effect of proneness to intrusive memories, F(1, 47) = 4.40, p = .041, ηp2 = .09, with fewer intrusive memories reported in the concrete compared with the abstract condition, t(48) = 2.07, p = .044, d = 0.59, 95% CI [1.69, 4.53].","This study investigated the effect of adopting an abstract or concrete mode of processing during exposure to analogue trauma and the subsequent development of intrusive memories. As predicted, individuals who were trained to adopt concrete information processing while watching the traumatic films reported fewer intrusive memories and lower PTSD symptom severity scores during the following week compared with individuals who were trained to adopt abstract information processing. Concrete information processing also led to less emotional reactivity (measured as distress and horror) to a posttraining film relative to abstract processing. Our results are consistent with the broader literature showing that processing mode influences responses to negative events (Watkins, 2008), and correlational research linking a ruminative, abstract style of processing after trauma to PTSD symptoms (e.g., El-Leithy et al., 2006; Michael et al., 2007). However, this appears to be the first study to show that processing mode (i.e., abstract vs. concrete) may be causally involved in the development of experiences such as intrusive memories, a hallmark feature of PTSD, and that undergoing training in concrete processing could protect against the development of intrusions. Furthermore, it supports models highlighting peritraumatic cognitive processing as a key factor influencing distress and the development of intrusive memories (e.g., Brewin et al., 1996; Ehlers & Clark, 2000). Since the current study focused on analogue trauma in a healthy, nontraumatized population, future research could investigate whether the current findings extend to real-life traumatic events and the subsequent development of PTSD symptoms. There are different possible mechanisms by which concrete processing may be more adaptive than abstract processing. First, in line with existing cognitive models (e.g., Ehlers & Clark, 2000), concrete processing may lead to a more organized memory and more adaptive appraisals of the event. Halligan, Michael, Clark, and Ehlers (2003) found that memory disorganization and negative appraisals of a traumatic event predicted PTSD symptoms in assault survivors. By focusing attention on contextual details and the overall sequence of events, concrete processing may promote the formation of a coherent narrative, whereas abstract processing may disrupt this process. In addition, concrete processing may generate situation-specific appraisals whereas abstract processing may promote unhelpful, overgeneralized appraisals (Watkins, 2008). Future research could measure the coherence or accuracy of memories of the trauma films as well as assessing what kinds of appraisals participants formed of the analogue trauma. Second, in line with hypotheses about the beneficial effect of concrete relative to abstract processing more generally, a concrete style of thinking may facilitate emotional processing of distressing events (e.g., Stöber, 1998; Teasdale & Barnard, 1993), potentially preventing the development of anxious appraisals (e.g., Behar et al., 2012). Since the abstract condition included instructions for participants to think about “Why these sorts of things happen,” “What it means about the world,” and “What if this was to happen to you or someone in your family?” one rival hypothesis to explain the current findings is that abstract training discouraged participants from processing the films in a self-referent way. Self-referent processing refers to organizing information in relation to the self: in relation to one’s present experience as well as to past experiences and future goals (Farb et al., 2007). Self-referent processing thus involves processing what is happening in the context of oneself and distinguishing one’s own experiences from what is being witnessed. Lack of self-referent processing has been associated with the development of PTSD symptoms, such as intrusive memories, in clinical and nonclinical samples (Evans, Ehlers, Mezey, & Clark, 2007; Laposa & Rector, 2012), presumably because when individuals fail to engage in self-referent processing, they are less likely to notice how the experiences they are witnessing are different from their own experiences. Future studies could include a measure of self-referent processing (e.g., Ehlers, 2002). Proneness to intrusive memories of stressful or unpleasant events in everyday life was associated with greater intrusions of the trauma films, which is consistent with Davies and Clark (1998), who found that self-reported tendency to experience intrusive memories predicted intrusive memories of a traumatic film clip. It is unclear why some individuals would have a greater tendency to experience intrusive memories, although this may relate to individual differences in ways of processing negative events (e.g., adopting an abstract-ruminative style) or strategies used to regulate emotions in response to them. Future research could further investigate processing styles or emotion regulation strategies that may be linked to intrusive memories after negative events in everyday life. However, even after controlling for proneness to intrusive memories, a significant effect of condition remained, suggesting that even if one were vulnerable to experiencing intrusions, the likelihood of developing intrusive memories after analogue trauma could be reduced through training in concrete processing. There are limitations that should be considered when drawing conclusions from this study. First, our study adopted an analogue design with a nonclinical sample, therefore we are unable to say whether these findings would generalize to real-life traumatic events. Second, since we did not include a no- training control condition, it is difficult to determine whether concrete training is superior to no training at all. However, because our hypotheses related to the effects of concrete compared with abstract training, our design fit the hypotheses to be tested and is consistent with the design of other studies seeking to compare abstract versus concrete information processing (e.g., Watkins, 2004; Watkins et al., 2008). Future research could include a no-training control condition to ascertain whether the effects of concrete processing are indeed superior to no training at all, as well as to abstract processing. In conclusion, this study provides the first empirical indication that relative to processing analogue traumatic stimuli in abstract ways, processing them in concrete ways leads to fewer intrusive memories over the following week and lower emotional reactivity to a subsequent analogue trauma. Our study also indicates that individuals who self- reported proneness to intrusive memories in everyday life were more likely to develop intrusive memories in response to analogue trauma and that concrete processing may reduce vulnerability to developing intrusions in these individuals. These findings improve our understanding of the role of peritraumatic processing in the development of intrusive memories and PTSD symptoms. It would be important to assess the generalizability of the current findings to at-risk occupational groups who are regularly exposed to trauma, such as emergency service personnel and journalists in conflict zones, since this could inform the development of interventions aimed at protecting against intrusive memories and PTSD in such populations.","The authors declare that there are no conflicts of interest."],["Appearance goals for exercise are consistently associated with negative body image, but research has yet to consider the processes that link these two variables. Self-determination theory offers one such process: introjected (guilt-based) regulation of exercise behavior. Study 1 investigated these relationships within a cross-sectional sample of female UK students (n = 215, 17–30 years). Appearance goals were indirectly, negatively associated with body image due to links with introjected regulation. Study 2 experimentally tested this pathway, manipulating guilt relating to exercise and appearance goals independently and assessing post-test guilt and body anxiety (n = 165, 18–27 years). The guilt manipulation significantly increased post-test feelings of guilt, and these increases were associated with increased post-test body anxiety, but only for participants in the guilt condition. The implications of these findings for self-determination theory and the importance of guilt for the body image literature are discussed. --------------------------------------------------------------------------------","Exercising to lose weight and improve one’s appearance is a prominent goal for physical activity in Western culture, particularly for women: a content analysis of women’s health and fitness magazines found that over 50% of main features were presented in an appearance or weight loss frame (Aubrey, 2010) and women appear to endorse these reasons for exercise more strongly than men (e.g., Furnham, Badmin, & Sneade, 2002). This endorsement of reasons for exercise such as weight loss, improving appearance, and increasing muscle tone is consistently associated with more negative body image (Furnham et al., 2002 Tiggemann & Williamson, 2000). In contrast, health reasons for exercise are associated positively with body image (Strelan, Mehaffrey, & Tiggemann, 2003). Given the potential consequences of poor body image for disordered eating behavior (e.g., Stice, 2002) and physical and mental health more broadly (e.g., Wilson, Latner, & Hayashi, 2013), it is important to understand why appearance reasons for exercise may be linked with negative body image. However, previous research has not directly evaluated the mechanisms underlying these associations. Self-determination theory (SDT) offers a framework within which to contextualize these different associations, with its focus on the motivation underlying human behavior (e.g., Ryan and Deci, 2006). SDT divides individuals’ goals, or reasons for behavior, into extrinsic goals, which focus on externally evaluated attributes or acquisitions, and intrinsic goals, which focus on self-development and supporting others around them. According to Ryan and Deci (2006), the pursuit of intrinsic goals fulfills basic psychological needs, resulting in higher levels of psychological functioning, whereas the pursuit of extrinsic goals does not. This proposition is well supported, with the endorsement of extrinsic goals, such as image and financial success, consistently associated with negative outcomes such as lower subjective well-being and mental health difficulties (e.g., Twenge et al., 2010). Overall life goals have also been shown to predict body image: in a sample of adolescent girls, the intrinsic life goal of health was associated with better body image, whereas the extrinsic goal of image was associated with more negative body image (Thøgersen-Ntoumani, Ntoumanis, & Nikitaras, 2010). Thus, the differential correlations of appearance and health reasons for exercise with body image could be understood to reflect the extrinsic and intrinsic nature of those reasons. Crucially, SDT provides an explanatory mechanism for interpreting these correlations, although it has not been directly tested in the domain of exercise: the regulation underlying the behavior. SDT suggests that the behavior we engage in when pursuing our goals can be regulated in a variety of ways, varying in levels of self-determination (how much the motivation stems from inside the self; Ryan and Deci, 2006). External regulation occurs when we engage in behavior due to external rewards or pressures, such as when someone exercises to please others. Introjected regulation is where the motivation for the behavior has been partially, but not fully, internalized: an individual might exercise to avoid the guilt they experience if they do not attend a session. Identified regulation is associated with valuing the benefits of the behavior, whatever these are believed to be, rather than the behavior itself. Finally, at the most self-determined end of the continuum, intrinsic regulation is experienced by those who engage in a behavior because they enjoy the behavior itself. Ryan and Deci (2006) suggest that more self-determined regulation should be associated with better well-being, due to the feelings of autonomy that it provides, and review a considerable amount of evidence supporting this assertion, across multiple domains. Self-determined regulation of behavior has positive associations with body image, both when considering regulation in general (Pelletier and Dion, 2007) and, in particular, for exercise behaviors (Brunet and Sabiston, 2009; Brunet, Sabiston, Castonguay, Ferguson, & Bessette, 2012; Markland, 2009; Thøgersen-Ntoumani & Ntoumanis, 2007). However, research also suggests that self-determined regulation is more likely to be associated with intrinsic goals, and non-self-determined regulation with extrinsic ones. Within the exercise domain, research has consistently found that extrinsic goals (e.g., weight loss, appearance reasons) are associated with less self-determined regulation, and that intrinsic goals (e.g., health, affiliation) are associated with more self-determined regulation (Gillison, Standage, & Skevington, 2006; Ingledew & Markland, 2008). Introjected regulation, with its foundation in avoiding guilt and shame, may be particularly relevant in this context. Although guilt is often conceived as a potentially positive motivating force, spurring us into action (e.g., Hoffman, 1982), self- determination theory suggests that guilt-based, introjected motivation may be detrimental to individuals’ well-being, especially when related to body-modification behaviors, such as eating regulation and exercise (Verstuyf, Patrick, Vansteenkiste, & Teixeira, 2012). Guilt-based regulation may be particularly relevant for women’s body image, given gender differences in the experience of self-conscious emotions. Women are more prone to experiencing guilt than men, particularly in individualistic cultures, such as the UK and US (Fischer & Manstead, 2000). Roberts and Goldenberg (2007), in fact, explicitly link women’s increased propensity to shame and guilt to the objectification of women’s bodies by society, and suggest that there should be an even greater gender divide in self- conscious emotions when bodies are made salient, such as in the exercise environment. Introjected regulation may therefore be particularly important in linking women’s body image to their reasons or goals for exercise. However, previous research has not considered appearance goals for exercise, introjected regulation, and body image simultaneously, tending to focus on just one of the associations between these three constructs. This approach may obscure the shared variance between these constructs, and a potential pathway between appearance goals and body image: body image may be associated with appearance reasons in part as a result of their shared association with guilt-based regulation.","The current research investigated the proposal that appearance goals for exercise may be associated with body image via their joint association with introjected regulation. As only components of this pathway have been explored previously in the literature, the aim of the first study was to provide initial cross-sectional support for this proposal. Thus, the first study employed a structural equation framework to model the direct and indirect associations between appearance goals, regulation of exercise behavior, and body image. This method allowed the confirmation of the shared variance between these three variables of interest, while controlling for their numerous correlates, such as health goals for exercise and other forms of exercise regulation (e.g., external, identified, and intrinsic). Notwithstanding the importance of cross-sectional evidence, it cannot provide true evidence of mediation: to fully test mediation, the mediator should be manipulated orthogonally from the independent variable (Bullock, Green, & Ha, 2010). Thus, in a second study, guilt in relation to exercise, the proposed mediator, was manipulated orthogonally from appearance goals, the proposed independent variable. Using a 2 × 2 experimental design, appearance goals for exercise and guilt related to not exercising were manipulated separately, allowing a more robust test of this proposed mediation process. By using a combination of correlational and experimental designs, the present research aimed to explore both the direction of causality in these relationships and the naturally occurring relationships between them, allowing for a fuller picture of this process than either method alone. For both studies, a sample of young adult women was used, due to the high frequency of body image issues within this group (Bucchianeri, Arikian, Hannan, Eisenberg, & Neumark-Sztainer, 2013) and research suggesting that exercise has negative associations with body image for this group, but not older women or men (Tiggemann & Williamson, 2000). Furthermore, research suggests that women experience introjected regulation differently than men (Gillison, Osborn, Standage, & Skevington, 2009) and experience greater levels of self-conscious emotions in Western cultures (e.g., Fischer & Manstead, 2000). Exercise is an important behavior in the pursuit of contemporary appearance ideals for both men and women (e.g., Pope et al., 2000; Tiggemann, 2011); however, the studies reported here focus on women’s experiences of exercise, in order to provide a specific examination of the motivational processes involved in linking their appearance reasons for exercise to body image, which may be very different from men’s.","Study 1 was designed to identify the regulations for exercise most strongly associated with body image and appearance goals for exercise and share variance with both. As discussed above, introjected, or guilt-based, regulation may be particularly relevant in linking appearance goals for exercise to body image in women. However, SDT would also predict that external regulation may be associated with appearance goals and with body image, as a non-self-determined form of regulation. In identifying the specific regulations that share most variance with both appearance goals and body image, this study aimed to provide initial evidence for potential pathways between these constructs. Given the lack of previous research demonstrating this degree of shared variance, between the three constructs rather than simply two, this cross-sectional study represents a necessary stage in the development of this research area.","Following institutional ethical approval, 215 female students (17–30 years, M = 19.77 years, SD = 2.0; 86% white) were recruited from a university participant pool to complete an online questionnaire. The ethical procedures of the study complied fully with APA and BPS ethical guidelines, with informed consent given before the study and debriefing for all participants after completion. Goals for exercise The Exercise Motivations Inventory was used to measure participants’ goals for exercise (EMI-2, Markland & Ingledew, 1997). Participants indicated how true (on a 5-point response scale ranging from not at all true for me to very true for me) each of 51 statements was of their reasons for exercising. The appearance goals measure consisted of the Appearance and Weight subscales (8 items; e.g., “I exercise to help me look better”; α = .95). The health goals measure, included to contrast appearance goals, consisted of the Ill Health Avoidance and the Positive Health subscales (6 items; e.g., “I exercise to have a healthy body”; α = .91). Appearance and Health emerged as distinct factors in an exploratory factor analysis of the full inventory, with no substantive cross-loading of items. Regulation of exercise behavior Participants’ regulation of their exercise behavior was measured using the Behavioural Regulation of Exercise Questionnaire 2 (BREQ-2, Markland & Tobin, 2004). This 19-item questionnaire includes measures of four subtypes of regulation: external (e.g., “I exercise because other people say I should”; α = .82), introjected (e.g., “I exercise because I feel guilty when I don’t exercise”; α = .82), identified (e.g., “I exercise because I value the benefits of exercise”; α = .86) and intrinsic (e.g., “I exercise because it’s fun”; α = .95). Participants indicated the extent to which items described their regulation of exercise behavior on a 5-point scale (not at all true for me to very true for me). Body image Three measures of body image were used. Participants completed a trait version of the Physical Appearance State and Trait Anxiety Scale (PASTAS, Reed, Thompson, Brannick, & Sacco, 1991), which presents eight body anxiety items (legs, waist, stomach, muscle tone, buttocks, hips, size, weight) alongside 12 filler items. Participants rated how anxious they had felt over the past six months about each item on a 5-point scale, ranging from not at all to extremely so (α = .91). The Body Appreciation Scale (BAS, Avalos, Tylka, & Wood-Barcalow, 2005) was included as a positive measure of body image. The scale includes 12 items which assess participants’ positive feelings and behaviors toward their body, using a 5-point response scale ranging from not at all true for me to very true for me (e.g., “I take a positive attitude towards my body”; α = .92). Third, participants completed the Self-Discrepancy Index (SDI, Halliwell & Dittmar, 2006). Participants were asked to generate four different things about themselves they would like to change (self- discrepancies) in an open-ended format, and then rated on a scale from 1 to 6 how concerned they were about each of these discrepancies (importance) and how different they were from their ideal (size). Participants’ responses were coded to identify weight, shape or tone (WST) discrepancies (“I am a size 12, but I would like to be a size 8”). These were coded separately from appearance-related discrepancies that could not be affected by exercise. The weight, shape and tone discrepancies were correlated with the PASTAS and BAS scores (r = .44 and −.38, respectively, ps < .05). A second researcher coded a subset of 25% of these discrepancies and inter-rater agreement on the identification of general appearance vs. weight-related discrepancies was high (98.3%). As per the published guidelines, size and importance of discrepancy were multiplied together and summed to provide a composite total score for weight, shape and tone discrepancies. Physical activity and body mass index (BMI) Participants completed the Leisure Time Exercise Questionnaire (LTEQ, Godin & Shephard, 1985). Participants reported how many times within an average week they engaged in mild, moderate, or strenuous physical activity for more than 15 min. A combined moderate-strenuous ‘METs’ score was computed from these figures.1 BMI was calculated using self-reported height and weight.","Mplus 7.4 (Muthén & Muthén, 1998–2015) was used to run a structural equation model, in order to assess the relationships between goals, regulations, and body image (see Table 1 for zero-order correlations and descriptive statistics). Appearance and health goals were modeled to be correlated and to be associated with the four regulations, which, in turn, were associated with body image. Goals and regulations were represented as observed variables using their scale means. Body image was modeled as a latent construct, with the PASTAS scale mean as the reference indicator due to its strong position within the body image literature, and with the BAS scale mean and the WST discrepancies score as the other indicators. Although PASTAS was used as the reference indicator, the weight of the factor loading was fixed to −1 (rather than the traditional +1), in order to keep the latent variable as a positive measure of body image. Residuals did not covary within this latent factor, but the residuals of the regulations (external, introjected, identified, intrinsic) were allowed to covary. BMI and participants’ METs score from the LTEQ were included as covariates, by modeling these as covariates of goals and directly associated with regulations and body image. This model had very good overall fit indices, with CFI above .95, RMSEA below .08 and SRMR below .06 (χ2 = 28.60, df = 16, p = .03; CFI = .99, RMSEA = .06, SRMR = .03; Fig. 1); the local fit of the model was also good with standardized residual covariances suggesting that no relationships in the data were poorly represented by the model (all <2). Thus, no additional paths were inserted. Given the focus of the research on appearance goals specifically, analysis of this model is focused on the associations, both direct and indirect, between appearance goals for exercise and body image.2 Appearance goals were strongly associated with introjected regulation and more weakly with external regulation. There was also a significant but small link between appearance goals and identified regulation. Introjected regulation was negatively associated with body image, whereas intrinsic regulation showed a positive association. External regulation was marginally negatively associated with body image (p = .09). Bootstrapping with 2000 samples was used to assess whether the associations between appearance goals for exercise and body image were due in part to their shared association with regulations. Appearance goals had a strong negative direct association with body image, but also a significant indirect association via introjected regulation (β = −.14, SE = .05, p = .003, 95% bias-corrected CI [−.05, −.23]). The other three indirect pathways (via external, identified, and intrinsic regulation) were non-significant (ps > .05, 95% bias-corrected CIs across zero). The link between appearance goals and body image is therefore partially due to their shared association with introjected regulation. Brief discussion ~~~~~~~~~~~~~~~~ These findings provide novel correlational evidence for the shared variance between appearance goals for exercise, introjected regulation, and body image, and suggest that introjected regulation may be a key link between appearance goals for exercise and body image. The individual importance of introjected regulation for exercise as an associate of body image has been highlighted previously (e.g., Brunet et al., 2012; Thøgersen-Ntoumani, & Ntoumanis, 2007); however, previous research has not identified this type of regulation’s potential importance in linking appearance goals for exercise to body image. These findings provide a framework within which to place previous research relating appearance and health reasons for exercise to body image (Furnham et al., 2002; Strelan et al., 2003; Tiggemann & Williamson, 2000), by considering these as domain-specific extrinsic and intrinsic goals, which are differentially associated with the regulation of exercise behavior and, in turn, body image. These findings, linking appearance goals for exercise, introjected regulation, and body image, provide necessary information for a more causal test of the links between appearance goals and body image. From this cross- sectional work, it is not possible to draw conclusions about the direction of this effect or to truly identify it as a case of mediation (Bullock et al., 2010). For this, we must experimentally manipulate both our proposed independent variable (appearance goals) and the proposed mediator (introjected regulation) to establish causation.","The initial cross-sectional study suggests that introjected regulation shares considerable variance with both appearance goals for exercise and body image. In our second study, to test the causal links between these variables, appearance vs. health frames for exercise were manipulated at the same time as inducing guilt vs. no guilt regarding exercise behavior, using a magazine article style of manipulation, as successfully used in previous research (Aubrey, 2010). It was hypothesized that participants in both of the guilt conditions (health and guilt; appearance and guilt) would experience more post-test guilt than participants in the no guilt conditions, but that post-test guilt would not be influenced by the appearance vs. health manipulation. Furthermore, for the proposed mediation to hold, the guilt manipulation should affect participants’ post-test body anxiety, whereas the appearance vs. health manipulation should not. That is, appearance reasons for exercise should only be problematic when paired with the guilt manipulation, if this is indeed the mediator in this case. By experimentally manipulating the proposed mediator in addition to the independent variable, this study offers a strong test of introjected regulation (guilt-based exercise motivation) as the underlying mechanism through which appearance goals influence body image. In establishing this effect of the guilt manipulation, there is the challenge of individual variation in responses to it: among those in the guilt condition, there is likely to be variation in how susceptible participants are to the manipulation, with some participants feeling guiltier than others as a result. As such, it would be plausible to predict a treatment–mediation interaction effect (Valeri & Vanderweele, 2013), with the guilt manipulation predicting increases in post-test guilt and this, in turn, predicting body anxiety, but only among those in the guilt condition. In other words, the impacts of a guilt manipulation on body anxiety can be expected to the extent that the manipulation succeeds in inducing guilt. Participants and design One hundred and sixty-five female university students (aged 18–27 years, M = 19.44, SD = 1.40) were randomly assigned to a 2 (appearance vs. health frame) × 2 (no guilt vs. guilt) between-subjects design. Participants were recruited through a university participation pool. Participants were predominantly white (77.7%), and within the ‘normal’ range for BMI (75% between 18.5 and 25, M = 21.31, SD = 3.59). Ethical approval for the experiment was granted by the ethics committee of the University, and the research process met APA and BPS ethical standards.","Participants attended group testing sessions, which ranged in size from 1 to 10 participants and took between 20 and 35 min to complete. After reading the information sheet and providing informed consent, participants worked through the pack at their own pace. Participants were informed that the study related to magazine preferences and requested that they read the article carefully. Appearance vs. health manipulation All participants were given a passage of text reportedly written by ‘Helen’, another student at the university. The passage outlined three tips for fitting exercise into a busy schedule. In the “appearance” conditions, the appearance and weight-related benefits of these tips were highlighted, such as toning and calorie burning, whereas in the “health” conditions, the health benefits of these tips were highlighted, such as cardiovascular health and injury prevention. Providing exercise advice with either a health or appearance focus is an effective means of priming health or appearance reasons for exercise respectively (Aubrey, 2010). The texts were closely matched in length and sentence construction, to ensure that the only substantive difference was the framing of the advice. Guilt manipulation The final paragraph of the text differed by guilt condition. In both conditions, the author acknowledged that she did not always do as much exercise as she would like to. In the ‘no guilt’ condition, this was followed by a self-compassionate statement about not feeling guilty for not doing enough. In the ‘guilt’ condition, this statement was adapted to focus on experiencing guilt for not doing enough, by rephrasing key statements to a more self-critical approach to missing a workout.3 Participants were asked to reread the final paragraph and to imagine they were the author. Participants then listed five reasons why they might feel as described in this paragraph. Measures of guilt often use responses to scenarios to assess this emotion (see Robins, Noftle, & Tracy, 2007, for a full review), and thus this was considered an appropriate technique with which to manipulate guilt. The majority of participants provided 5 reasons (84.3%), with only 4 participants providing 2 or fewer. Questions on the article Participants were asked to describe the material to confirm they had read the article; all participants accurately described the content. They also were asked how similar they thought the author was to them and how likeable the author was (rated on a 5-point Likert scale, not at all to extremely). There were no significant main effects or interactions on perceptions of author likeability and similarity to participants (appearance vs. health, guilt vs. no guilt, appearance × guilt; all ps > .05; descriptive statistics in Table 2). Participants were also asked how health- and appearance-focused they thought the author was, as a manipulation check (see","for full details). Post-test guilt and negative emotion Post-test guilt was assessed using a short form of the Positive and Negative Affect Scale (I-PANAS-SF, Thompson, 2007), with one additional item (guilty) included. This item was included as a manipulation check for the guilt conditions. Participants were asked to what extent they were experiencing each of 11 mood adjectives right now and responded on a 7-point Likert scale (not at all to very much). In addition to guilt, the mean of five other negative emotion terms (hostile, upset, nervous, afraid, ashamed) was used to control for a general negative response to the article (α = .79). This scale occurred only after the manipulation had taken place; there was no pre-test of guilt or other emotions. This was a purposeful decision on the part of the research team, to avoid multiple questions relating to guilt sensitizing participants to this emotion and altering their response to the guilt manipulation, a potential pre-test- treatment interaction effect (Shadish, Cook, & Campbell, 2001). Demand characteristics and other manipulation effects may be particularly relevant in the body image domain, where many findings are well-known in popular culture (e.g., the effects of thin ideal media or focusing on appearance) and demand characteristics have been shown to influence repeat measurements (e.g., Fingeret, Gleaves, & Pearson, 2004; Krawczyk, Menzel, & Thompson, 2014). The risk of heightening sensitivity to guilt was also noted in the development of our materials in a pilot study (N = 50): our first guilt-inducing Helen was too obvious in her attempts to manipulate participants’ guilt. Participants in this condition did not feel guilty, and instead disliked the author significantly more than the other conditions. Body anxiety (state) The Physical Appearance State Trait Anxiety Scale (PASTAS, Reed et al., 1991) was used to measure body anxiety. Participants were asked how anxious they were about a range of elements of their lives right now and responded on a 5-point Likert scale (not at all anxious to very anxious). Embedded within the 20-item scale were 7 items relating to appearance issues, such as “my size”, and “the extent to which I look overweight”. These 7 items demonstrated excellent reliability (α = .85). Regulation of exercise behavior (state) An adapted, shortened version4 of the Behavioural Regulation of Exercise Questionnaire 2 (BREQ-2, Markland & Tobin, 2004) was used to measure participants’ immediate motivation for exercise. The introductory text was rephrased to ask participants to consider why they would be exercising today if they did so, to attain a ‘state’ measure. We used the introjected regulation subscale in our analyses (e.g., “I would be exercising today because I feel guilty when I don’t exercise”; α = .85). Demographic information, BMI, and demand characteristics Participants reported their age, ethnicity, height, and weight. Height and weight were used to calculate body mass index (available for 150 participants). Participants were asked what they thought the study was investigating. No participants recognized that they had experienced a guilt manipulation. Trait measures Two weeks after the experimental session, participants were emailed a link to an online survey and provided their trait measures via this portal (n = 130), in order to control for these in later analyses if necessary. As the effects of the exposure manipulation (a 660-word piece of text) were expected to be relatively short-lived, it was considered appropriate to use a two-week follow-up questionnaire to collect trait data, especially as previous research within an exposure paradigm (e.g., Ashikali, Dittmar, & Ayers, 2014) has included trait measures after the exposure and post-test state measures. Participants completed trait measures of body anxiety, goals for exercise, and introjected regulation. For body anxiety, participants completed the PASTAS (Reed et al., 1991) a second time, but this time were asked how anxious they were about a range of elements of their lives in general. The measure once more demonstrated high reliability (α = .92). A 15-item form of the Goal Content for Exercise Questionnaire (GCEQ, Sebire, Standage, & Vansteenkiste, 2008) was used to measure participants’ appearance and health goals for exercise, with three items for each goal (αs = .85 and .82, respectively). Participants rated to what extent various goals for exercise were important to them on a 5-point Likert scale (not at all important to very important). The shortened BREQ-2 (Markland & Tobin, 2004) was used to assess participants’ trait introjected regulation of exercise behavior (α = .84). Missing data There was no missing data on the post-test measures of body anxiety or guilt. The post-test measure of negative emotions comprised five items and on three of these there was missing data for one (but each different) respondent and these were replaced by mean substitution to provide complete data. For the pre-test measure of trait body anxiety, however, there was missing data on just under 10% of cases (16 out of N = 165). For the mediation analyses, this was handled using Full Information Maximum Likelihood (Enders, 2010); for the ANCOVA, we imputed missing values using the EM algorithm in SPSS. Manipulation checks and overall effects of manipulations Manipulation checks were conducted using a 2 × 2 analysis of variance (ANOVA). Tests for the overall effects of the manipulations (appearance vs. health frame; guilt vs. no guilt) were conducted using an analysis of covariance (ANCOVA), with any trait variables that differed significantly between the conditions included as covariates. These analyses were conducted in SPSS (version 23). Mediation analysis As the manipulation was expected to influence post-test body anxiety due to its effect on post-test guilt, this assumption was tested via a structural equation model. However, mediation is complicated in a situation where the treatment or experimental condition may interact with the mediator itself (Muthén & Asparouhov, 2015; Valeri & Vanderweele, 2013), as may be the case in this design: post-test guilt does not represent the same type of guilt in each condition, and may therefore have a different effect on the outcome of body anxiety. ‘Guilty’ participants in the guilt condition should theoretically be feeling this way due to the manipulation; their guilt should be specifically associated with not exercising enough. In contrast, variation in the guilt ratings of participants in the no guilt condition will not necessarily be associated with guilt regarding exercise (which this condition specifically aims to reduce), but rather should represent other, general, sources of guilt. Thus, we would expect variation in guilt associated with not exercising, mostly aroused in the guilt condition, to affect post-test body anxiety, but variation of other kinds of guilt, most aroused in the control condition, not to affect post-test anxiety. Hence, we predicted a mediation by post-test guilt, but also a moderation by the condition of the mediator’s effect, as indicated by a treatment–mediator interaction effect. The counterfactual method detailed by Muthén and Asparouhov (2015; see also Valeri & Vanderweele, 2013) was employed to examine this possibility, allowing the simultaneous consideration of the mediation and treatment–mediator interaction. Briefly put, this method involves the decomposition of the total effect into two components. In classic treatments of mediation (Baron & Kenny, 1986), where there is no treatment–mediator interaction, the total effect comprises a direct effect and an indirect effect, and comprises three path coefficients. The presence of an interaction, however, introduces additional coefficients that contribute to the total effect, and that need to be taken into account when defining direct and indirect effects. Thus, the total effect can be decomposed either into the pure natural direct effect (PNDE) and the total natural indirect effect (TNIE), or into the total natural direct effect (TNDE) and the pure natural indirect effect (PNIE: Muthén & Muthén, 2015; Valeri & Vanderweele, 2013). For our purposes, the first decomposition is the most appropriate since it is the TNIE that represents the change in the outcome when the condition is held constant at the treatment condition (the guilt condition) and the mediator changes from the level of the control (no guilt) to the level of the treatment condition. Thus, it includes the product of the interaction effect and the effect of the treatment on the mediator, and hence here captures the expectation that it is the guilt aroused in the treatment condition that has an effect, but not the guilt found in the control condition. The counterfactual method relies on the assumption that confounding variables of the mediator-outcome relationship are controlled (Valeri & Vanderweele, 2013). We therefore included two variables as covariates of post- test guilt and post-test body anxiety: trait body anxiety and post-test negative emotions (the mean of five negative emotions from the I-PANAS-SF). Thus, we conducted what MacKinnon and Pirlott (2015) refer to as a “comprehensive structural equation model” (p.35), which explicitly models the influence of known confounding variables measured in the study. We further tested the specificity of the mediation via post-test guilt by directly replacing it with post-test negative emotions in a further analysis. Results ~~~~~~~ Descriptive statistics can be seen in Table 2, by condition. Random assignment checks A series of ANOVAs were conducted to assess whether the trait levels of key variables were significantly different between any of the conditions. Only trait levels of body anxiety significantly varied between conditions; specifically, participants in the health conditions had higher trait levels of body anxiety than those in the appearance conditions, F(1, 125) = 7.19, p = .01; health conditions: M = 2.87, SD = 1.09; appearance conditions: M = 2.40, SD = 0.98. As such, trait levels of this variable were controlled for throughout the analyses. No other potential covariates varied significantly between conditions (age, BMI, trait endorsement of health or appearance goals, trait introjected regulation; all ps > .05). Health and appearance focus ANOVAs were conducted to establish whether the articles primed the intended concerns. Participants perceived the author in the appearance conditions as significantly more appearance-focused than the author in the health conditions F(1, 161) = 31.62, p < .001; health conditions: M = 3.33, SD = 0.81; appearance conditions: M = 4.05, SD = 0.82. The two authors were perceived as equally health-focused, F(1, 161) = 1.44, p = .23; health conditions: M = 3.75, SD = 0.79; appearance conditions: M = 3.59, SD = 0.93. This suggests that both articles primed health concerns, rather than only the health condition. However, the clear perception of the appearance author as more appearance-focused suggests that the manipulation was successful in its main purpose of highlighting appearance reasons for exercise. Guilt inducement The success of the guilt manipulation was assessed with two measures: the immediate post-test rating of guilt and the state measure of introjected regulation. In the case of post-test guilt, a 2 × 2 ANOVA indicated that the guilt manipulation had a significant effect on participants’ immediate emotional reports of guilt, F(1, 161) = 13.02, p < .001; guilt conditions: M = 2.95, SD = 1.69; no guilt conditions: M = 2.02, SD = 1.58. There was no main effect of appearance condition, or of the interaction between the two conditions (both ps > .05). In the case of introjected regulation, neither the guilt nor appearance manipulation had a significant effect on this outcome; the interaction between conditions was also non-significant (all ps > .05). Overall effects of manipulations on body anxiety A 2 × 2 ANCOVA was conducted to assess whether the guilt manipulation, the appearance vs. health manipulation, or the interaction between the two predicted post-test state body anxiety (PASTAS), using trait body anxiety as a covariate. There were no main effects of appearance and no interaction effect, but trait body anxiety had a strong effect on state scores, F(1, 160) = 330.63, p < .001. Post-test guilt: mediation and treatment–mediator interaction Mediation analyses were carried out using Mplus 7.4 (Muthén & Muthén, 1998–2015) using bootstrap standard errors with 1000 bootstrap samples. We first carried out an analysis in which post-test guilt mediated the effect of the guilt manipulation (the treatment) on post-test body anxiety, controlling for trait body anxiety and post-test negative emotions. Trait body anxiety significantly affected post-test body anxiety (b = .0.67, p < .001) but not post-test guilt (b = 0.12, p > .05). Negative post-test emotions significantly affected post-test guilt (b = 0.89, p < .001) and post-test body anxiety (b = 0.12, p < .05). The treatment significantly affected post-test guilt (b = 0.73, p < .001), but post-test guilt did not significantly affect post-test body anxiety (b = 0.05, p = .10) and the treatment–mediator interaction was also not significant (b = 0.09, p = .10). The results from the mediation analysis, however show that the TNIE was significant (see Table 3 for direct and indirect effects), indicating that there was a significant indirect effect of experimental condition, via post-test guilt, on post-test body anxiety, but only for those in the guilt condition. Women in the no guilt condition did not demonstrate this mediation effect (PNIE was non- significant), and no other effects in the mediation analysis were significant. When post-test guilt was replaced in the analysis by post-test negative emotions as the mediating variable, TNIE was not significant and there were no other significant effects in the mediation analysis, suggesting the critical role of guilt rather than negative emotion more generally. It seems, then, that the guilt manipulation had an effect on post-test body anxiety via the particular kind of guilt that it aroused (which we assume to be guilt about lack of exercise), guilt that was not aroused in the control condition. Brief discussion ~~~~~~~~~~~~~~~~ The total effect of the guilt condition on body anxiety was not significant and this is due to the fact that the direct effect is negligible and not significant. The effect of the guilt manipulation was fully mediated by the extent to which it aroused guilt; not all women experienced guilt as a result of our manipulation, but those that did felt more anxious about their bodies. Higher levels of post-test guilt for women in the no guilt condition were not associated with higher levels of body anxiety. This finding suggests that guilt related to exercise is a mechanism through which appearance goals may influence body image. The effect of appearance vs. health framing observed by Aubrey (2010) appears to be superseded by the guilt manipulation introduced in this experiment: appearance goal priming was not problematic for body image when combined with the no guilt manipulation. In further support of the importance of guilt, these findings were not replicated when post-test guilt was replaced by post-test negative emotions more generally in our mediation analysis; the negative link to body anxiety appears to be specific to the guilt elicited by our manipulation. The inclusion of negative emotions beyond guilt and their inclusion as controls and replacing guilt in the analysis is a key strength and contribution of this study. Our findings support the importance of guilt in particular in the relationship between appearance reasons for exercise and body anxiety and shows the divergent validity of guilt, compared to negative affect more generally. In considering this study’s contribution, it is important to note that the experimental materials closely imitated the materials that women are regularly exposed to. Guilt was induced not through an artificial cognitive task, such as scrambled sentences (e.g., Zemack-Rugar, Bettman, & Fitzsimons, 2007), but by an active discussion of guilt by the author, an event that regularly occurs in the real-life media exposures that women experience (e.g., ‘true life testimonials’ in magazines). This similarity gives this experiment a much greater degree of ecological validity than might otherwise be expected of a lab-based experiment. Aubrey (2010) argues that this form of exposure represents a single ‘meal’ in women’s ‘media diets’: this is only a single text endorsing appearance goals, but given the cultural prominence of these messages, it is likely that women are exposed repeatedly to these, experiencing these state effects on body image multiple times a day, and that over an extended period these effects may become cumulative, altering trait levels. Future work should consider these relationships longitudinally, to confirm the direction of the relationship between appearance goals for exercise and body image, via introjected regulation, in a naturalistic environment. In spite of the valuable insights from this study, restrictions within the methodology and the results mean that they must be interpreted with caution, particularly with respect to the mediating role of guilt in influencing body anxiety. There was no main effect of the guilt manipulation on post-test body anxiety, and the mediation analysis presented utilizes post-test guilt as part of the indirect effect: the mediator (guilt) is measured at the same time (post-test) as the proposed outcome variable (body anxiety). The study’s findings around mediation and the guilt manipulation’s influence on body anxiety are therefore limited by their cross- sectional nature, in spite of being situated overall within an experimental design. This raises the possibility of either a reverse effect, whereby the manipulation increased body anxiety, which was responsible for an increase in guilt, or of other unmeasured variables being responsible for the association, as is the case in other cross-sectional mediation analyses (e.g., Bullock et al., 2010). Although we included trait body anxiety and post- test negative emotions as potential confounding variables, there is no guarantee that there are not others at work. As such, the study does not provide as strong a test of mediation of appearance goals’ influence on body image via guilt as intended in the initial study design; its findings must be interpreted in this light, and supported by further investigation. A further methodological limitation of the study was that guilt was not measured before the manipulation. Although a deliberate decision, in order to avoid pre-test sensitization, this design leads to two difficulties. First, the study cannot analyze a change in guilt, and therefore must assume that random assignment to conditions has eliminated potential variation between groups or that this is sufficiently controlled for by the associated trait variables, such as body anxiety and introjected regulation. Second, even if the conditions as groups had similar levels of pre-test guilt, the random variation between participants within these conditions may act to introduce additional uncontrolled variation which may serve to obscure the true causal relationships being considered. This may be particularly relevant given the lack of a direct effect found and the specific indirect effect reported: if only particular women respond to the manipulation, pre-test measures of guilt would be vital in future research to identify who these women are and what the consequences are for them.","Across the two studies, there is initial support for the importance of guilt as a key process through which appearance goals for exercise are associated with body image. Study 1 provides cross-sectional evidence for the shared variation in appearance goals, introjected regulation, and body image, assessing all three within a single model, while controlling for other regulations and goals. The experimental manipulation of these variables within Study 2 provides support for the proposition that guilt relating to exercise may result in increased body anxiety, in spite of the limitations discussed previously. These findings support the theoretical proposal that regulation of exercise behavior may mediate the association between women’s goals for exercise and their body image, as predicted by self-determination theory (e.g., Ryan & Deci, 2006), given the consistent association of extrinsic goals with controlled regulations (e.g., Gillison et al., 2006; Ingledew & Markland, 2008) and of controlled regulations with worse body image (e.g., Brunet and Sabiston, 2009; Thøgersen-Ntoumani & Ntoumanis, 2007). However, although the results replicate the broad theoretical predictions of less self-determined regulation being associated with lower well-being (e.g., Sheldon, Ryan, Deci, & Kasser, 2004), these findings also raise a question for self-determination theory: the most controlled form of regulation, external, is not most strongly associated with negative wellbeing outcomes. In our analyses, introjected regulation emerges as the key regulatory pathway linking appearance goals and negative body image, and future theoretical and empirical work should seek to understand why guilt as a motivation for exercise behavior may have more negative associations or consequences than more external pressures. Guilt is often discussed as a positive motivator, driving us to reparatory action to fix a perceived wrong, but the evidence presented here and the growing body of work in the body image domain (e.g., Brunet & Sabiston, 2009; Calogero & Pina, 2011) suggests that this may not be the case. Guilt appears to be an important emotional response and motivational process resulting from exposure to or endorsement of the extrinsic goal of attractiveness. That guilt relating to exercise behavior has such negative associations for body image is an important finding, as it may open up a new avenue of interventions, suggesting that the negative association between appearance goals and body image could be mitigated by decoupling these goals from the guilt associated with not exercising enough. This provides a potential solution for researchers seeking to reduce the negative impact of appearance goals on women’s body image, without appearing to criticize individuals’ reasons for exercise: by introducing interventions aimed at reducing guilt-based motivation for exercise, practitioners can potentially disrupt one of the negative pathways from appearance goals to body image. From a public health perspective, this form of intervention could have a double reward, reducing the associated health issues of negative body image, but also increasing long-term exercise persistence, which has been negatively associated with introjected regulation (Pelletier, Fortier, Vallerand, & Briere, 2001). In addition to methodological issues relating to the individual studies, previously discussed, the nature of the sample limits the extent to which its findings can be generalized beyond female undergraduate students in the UK. Although there is clear justification for selecting the particular samples of young women in the present work, future research should focus on extending such work to other ‘at-risk’ groups, such as young men (Pope et al., 2000). Thus, future research should investigate whether the importance of guilt as motivation for exercise is an issue unique to women, or whether it can be generalized to men as well. This may be particularly important given research suggesting that young men and women experience introjected regulation differently, with women focusing on the avoidance of guilt and men focusing on the attainment of social status and appreciation (Gillison et al., 2009). These results set an agenda for further work to evaluate the unfolding causal relations between appearance motivations for exercise and body image over time. This study provides evidence of the potential importance of guilt in linking appearance goals for exercise and body image; future research should focus on the task of further investigating the causal nature of this relationship, employing longitudinal research to examine this relationship over longer periods of time and in a more naturalistic setting, alongside further experimental manipulations to fully confirm causality."],["Theoretical models of social learning predict that individuals can benefit from using strategies that specify when and whom to copy. Here the interaction of two social learning strategies, model age-based biased copying and copy when uncertain, was investigated. Uncertainty was created via a systematic manipulation of demonstration efficacy (completeness) and efficiency (causal relevance of some actions). The participants, 4- to 6-year-old children (N = 140), viewed both an adult model and a child model, each of whom used a different tool on a novel task. They did so in a complete condition, a near-complete condition, a partial demonstration condition, or a no-demonstration condition. Half of the demonstrations in each condition incorporated causally irrelevant actions by the models. Social transmission was assessed by first responses but also through children's continued fidelity, the hallmark of social traditions. Results revealed a bias to copy the child model both on first response and in continued interactions. Demonstration efficacy and efficiency did not affect choice of model at first response but did influence solution exploration across trials, with demonstrations containing causally irrelevant actions decreasing exploration of alternative methods. These results imply that uncertain environments can result in canalized social learning from specific classes of model. --------------------------------------------------------------------------------","The social learning of behavior, including tool use, language, and cultural norms, is a fundamental aspect of a child’s development. However, social information can be outdated or inappropriate. Thus, children do not socially learn indiscriminately; rather, they implement cognitive decision-making rules and social learning strategies (Boyd & Richerson, 1985; Laland, 2004; Rendell et al., 2011). These biases toward certain information or people dictate who young children copy and under what circumstances (Price, Wood, & Whiten, in press; Wood, Kendal, & Flynn, 2013b). Such social learning strategies have the potential to facilitate the creation of social traditions and the evolution of cumulative culture (Dean, Vale, Laland, Flynn, & Kendal, 2014) where cultural traits are modified over multiple generations, resulting in an increase in the complexity and efficiency of these traits. Accordingly, understanding children’s selective learning can contribute to our knowledge of uniquely human cultural abilities. The investigation of children’s social learning strategies can be achieved through differing experimental paradigms that measure (a) copying choices regarding personal preferences (e.g., Shutts, Kinzler, McKee, & Spelke, 2009), (b) novel object labeling (e.g., Koenig & Harris, 2005), and (c) novel object use (e.g., Wood, Kendal, & Flynn, 2013a). Such empirical methods are used in conjunction with theoretical models predicting that the implementation of social learning strategies dependent on the context of the to-be-learned behavior is advantageous (Laland, 2004). Emerging empirical evidence also suggests that children’s social learning is often too complex to be explained by a single strategy and that strategies may be most beneficial when they can be used flexibly in different contexts. “Copy when uncertain” biases ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A copy when uncertain bias is one social learning strategy dictating when individuals copy others (Rendell et al., 2011). Such a bias has been found in other animal species, including rats (Galef, 2009) and fish (van Bergen, Coolen, & Laland, 2004). The uncertainty in these paradigms may relate to (a) observers’ uncertainty regarding their environment and (b) whether they should use social versus personal information. However, their uncertainty can also relate to (c) the efficacy and efficiency of social information and (d) which of multiple sources of information, or “who,” they should best copy. The different paradigms used to investigate children’s strategies represent different environments of uncertainty relating to efficacy and efficiency. Novel object labeling paradigms involve two or more models labeling a novel object in divergent ways, but the efficacy of either label remains unknown throughout the paradigm. Novel object use paradigms, such as using a tool to extract a reward from a novel container, differ from such word labeling paradigms in that the efficacy of the model(s) is often made evident by the completion of the task. These differing contexts of uncertainty influence who gets copied; when model efficacy is uncertain, as when children are presented with a novel object labeled differently by two models, 3- and 4-year-old children use a label provided by a previously proficient word labeler over a previously inept word labeler (Koenig, Clément, & Harris, 2004). Conversely, when efficacy is known, as when children are presented with a novel puzzle that is successfully solved differently by one previously proficient solver and one previously less proficient solver, 5-year-old children do not show an initial preference for either model’s method and are motivated to try both methods over time (Wood, Kendal, & Flynn, 2015). Wood, Kendal, and Flynn (2015) argued that because the children were certain about the effectiveness of each method, they did not show any model-based bias to either peer. Similarly, Hu, Buchsbaum, Griffiths, and Xu (2013) found that a bias to follow a majority of others was present only when 3- to 5-year-old children did not know whether the socially demonstrated responses were effective. If children could see that all socially demonstrated responses were effective, the bias was lost. Thus, task-naive children may implement a model-based bias only when there is some uncertainty as to the efficacy of the social information the models are providing. Another form of social information uncertainty corresponds to the efficiency of the social information. Models may produce a plethora of behaviors toward a novel task, and understanding which of those actions are necessary or unnecessary for completing a goal with the object may prove to be important. Likewise, models who perform numerous unnecessary actions may be viewed and copied differently than those who do not. In the domain of social learning research, such causally unnecessary actions have been labeled “irrelevant actions,” and the copying of such actions is commonplace among children and adults (McGuigan, Makinson, & Whiten, 2011). This copying is intriguing and has been argued to enable the development of unique aspects of human culture such as complex cultural practices (Boyd & Richerson, 1996). Although 3- to 5-year-old children often faithfully reproduce such actions, they may nevertheless identify them as “silly” (Wood et al., 2013a) and unnecessary (Lyons, Damrosch, Lin, Macris, & Keil, 2011) and omit them if produced by certain models such as children (Wood, Kendal, & Flynn, 2012). Therefore, there is good reason to think that irrelevant actions might add some ambiguity to the social learning context and children’s perception of model efficiency. A major aim of the current study was to discover whether similar-aged children’s use of a social learning strategy would be affected by the observer’s degree of uncertainty by manipulating the effectiveness and efficiency of the social information provided. The social learning strategy we investigated was a model age-based bias. “Model age-based” biases ~~~~~~~~~~~~~~~~~~~~~~~~ There has been a wave of research demonstrating that young (2- to 6-year-old) children apply model-based social learning strategies (Henrich & McElreath, 2003), where there is selective copying dependent on the identity of the model providing the information (Birch, Vauthier, & Bloom, 2008; Corriveau & Harris, 2009a, 2009b; Corriveau et al., 2009; Koenig & Harris, 2005; Lane, Wellman, & Gelman, 2013; McGuigan, 2013; Zmyj, Buttelmann, Carpenter, & Daum, 2010). We chose to investigate a bias for model age because it has been shown to be salient in a number of contexts (Brody & Stoneman, 1981; Jaswal & Neely, 2006; Ryalls, Gul, & Ryalls, 2000; Seehagen & Herbert, 2011; Zmyj, Aschersleben, Prinz, & Daum, 2012), although not always in the same direction. “Vertical” or “oblique” intergenerational transmission where adults are copied rather than children, has been found with videotaped target acts (Seehagen & Herbert, 2011), and novel object labeling (Jaswal & Neely, 2006). Conversely, infants have shown higher fidelity copying of a 3-year-old child versus an adult (Ryalls et al., 2000) and of peers in preference to older children and adults (Zmyj et al., 2012) when the context was play (but see Rakoczy, Hamann, Warneken, & Tomasello, 2010, for children protesting over a puppet that copies a child over an adult). Similarly, 3- to 5-year-olds selected an adult as a model when answering questions within an adult domain, such as the nutritional value of food, but deferred to a child model when the domain was toys (VanderBorght & Jaswal, 2009). Establishing traditions ~~~~~~~~~~~~~~~~~~~~~~~ Social learning can be assessed by children’s first responses to a novel object, but sustained social learning, the hallmark of social traditions, measured through continued interaction with that object can give a more detailed picture of social transmission. Wood et al. (2013a) found that 5-year-olds who were previously naive to a puzzle box that could be operated by two different methods became canalized to using just one demonstrated method significantly more than children who had previously explored the box and successfully innovated solutions. Likewise, the presentation of social information before interaction with an object can canalize interactions and limit exploratory play in 4-year- olds (Bonawitz et al., 2011). These results suggest that a copy when uncertain strategy could inhibit innovation and lead to a canalized social tradition. The current study aimed to investigate whether this canalization would happen in relation to a model age-based bias; would uncertainty increase conservatism toward a particular model? The current study ~~~~~~~~~~~~~~~~~ In this study, we set out to extend our knowledge of 4- to 6-year-old children’s flexible use of social learning strategies by manipulating observer certainty in the efficacy and efficiency of differently aged models. The study employed a puzzle box, the “Slotbox,” which was designed so that each of two functionally different tools could be used to extract a soft toy. Given mixed findings regarding the direction of model biases for age, we did not make a prediction in this respect. Rather, we aimed to explore the effects of uncertainty on such a bias occurring in either direction. Uncertainty about the efficacy of the social information was created by varying the completeness of the demonstrations such that whereas some children saw both models complete all of the necessary series of actions involved and have a token extracted, others saw a degraded, less complete series of these actions that did not reveal eventual successful removal of the token from a puzzle box task. Uncertainty about the efficiency of the social information was created by varying whether models incorporated visibly causally irrelevant, and thus inefficient, actions into their demonstrations. A final group of children did not witness any social information. We investigated children’s success with a solution method and conservatism to this solution over five response trials. We predicted, first, that children who received social information as compared with the control group would socially learn as indicated by increased success, but that a degraded (vs. full) demonstration would reduce success as measured by latency to success. Second, any model bias would be most pronounced when social information lacked evidence of efficacy and efficiency because the demonstrations would create the most uncertainty. Third, once children have achieved a successful solution, conservatism to this solution would be greatest when they are uncertain of the alternative method. Thus, children who witness two complete and efficient solutions would be predicted to be motivated to explore both demonstrated methods, whereas those with the least complete and inefficient demonstration would be predicted to show more canalization to a particular method.","In total, 151 4- to 6-year-old children completed the study. Of these, 11 children were excluded from analysis (English not first language [n = 2], technical problems during experiment [n = 6], or assistance offered by caregiver [n = 3]). The remaining 140 children (86 girls) ranged from 4 years (48 months) to 6 years (83 months) of age (M = 64.1 months, SD = 9.9). Children were recruited while visiting Edinburgh Zoo through a poster that read, “Win stickers. We are interested in children’s learning and would like to see how you play with toys we have made.” Another poster showed a picture of the Slotbox with “Our toy” written above it. Consent for participation was obtained from children’s caregivers provided that they were parents or grandparents. There was no significant difference in the distribution of boys and girls, χ2(6, N = 140) = 0.30, p > .99, across the seven conditions. There was some difference in the distribution of age, F(6, 133) = 2.49, p < .05, across these conditions, but post hoc pairwise comparisons failed to reveal any statistically significant differences (ps > .05). Design ~~~~~~ Within each condition, there was a within-group variable of model age (adult or child), with each model demonstrating one of two different methods for reward retrieval (arrow or rake, counterbalanced and explained further below) sequentially (order of model demonstration was counterbalanced). There were two between-participant variables: (a) demonstration efficacy shown through completeness (three levels: complete, near-complete, or partial) and (b) demonstration efficiency shown through irrelevant actions (two levels: present or absent). Table 1 summarizes the seven six experimental conditions. There was an additional “no-demonstration” control condition that offered no demonstration of any kind; here each model was paired only with a photograph of one of the two tools.","Video recordings of the models’ actions were embedded in PowerPoint presentations shown on two 19-inch monitors on either end of a table facing the child participant. Children were tested in a pop-up gazebo within an indoor area of Edinburgh Zoo. Fig. 1 depicts the location of the apparatus within the gazebo. The monitors were each attached to computers out of the child’s view. In the middle of the table lay the Slotbox with the rake and arrow tools. The Slotbox is a largely transparent plastic puzzle box (length = 25 cm, width = 8 cm, maximum height = 14 cm, minimum height = 7 cm) that contained a small soft toy token (a gibbon: height = 10 cm). This token was put in the box through a hole in the top (Fig. 2A) to rest at the back of the box (Fig. 2B). There were two other openings to the box: one at the front (height = 5 cm, width = 6 cm), which was covered by a top-hinged door (with crossbar height = 2 cm, width = 8 cm) that could be lifted up, and a slit (height = 0.5 cm, length = 23 cm) along the side of the box. To the right of the Slotbox on the table were two tools: the rake and the arrow. The rake tool was a long rectangular piece of brown opaque plastic (width = 4 cm, length = 25 cm, diameter = 0.2 cm) with three prongs at the end. The arrow tool was a narrow-shaped piece of black plastic (width of arrow base = 10 cm, width of handle = 4 cm, length = 15 cm, diameter = 0.2 cm). Critically, these tools provided two different ways to extract the token. Either the rake could be inserted through the front opening and used to pull the token out of the front opening (Fig. 2C) or, alternatively, the arrow tool could be inserted into the side slit of the box and used to push the token out of the front opening. In addition, a hollow tin with a lid (height = 15 cm, diameter = 10 cm) was used in a warm-up task as well as 2-cm stickers used for rewards. Models and demonstrations ~~~~~~~~~~~~~~~~~~~~~~~~~ Three female adults and three female children acted as the models in different videos, so that biases would not be due to characteristics, beyond age, of individual models. These adult and child models were paired for presentations in each of the nine possible combinations. The three adult models were aged 20 to 22 years, and the three child models were aged 4 to 6 years. All models were unfamiliar to the participants and were recorded first looking at the camera and waving. Initially, an attempt was made to train the child models to perform the demonstration. However, these demonstrations differed significantly from the adult demonstrations in terms of duration, precision, and clarity. Although this may reflect a naturally occurring difference in competence, it was important to avoid confounding an age bias with a competence bias. Thus, as in Wood et al. (2012), the video demonstrations focused on the task, showing just the model’s hands and unclothed arms, and were performed by adults. The “child” demonstration was modeled by a 20-year-old who had small, child-like hands, and the adult demonstration was modeled by a 21-year-old who had average-sized hands, with only the hands and lower arms being visible in the video presentations. No hands had jewelry, nail varnish, or overtly manicured nails. Copies of these video demonstrations are available in the online Supplementary material. In Wood et al. (2012), 20 adults, blind to the study, did not notice that adult and child demonstrations were both performed by adults. In the current study, no child said that the demonstrations were not performed by the model. Procedure The research assistant greeted each interested child and caregiver, saying to the child, “Would you like to play a game with some toys and see if you can win some stickers?” The interested party was shown into the gazebo and introduced to the experimenter (E), who was sitting behind the table with the Slotbox (see Fig. 1). The research assistant stood with the parent at the entranceway to the tent. Photos of the child and the adult models were displayed, one on each monitor. E said, “All of our games today involve getting the Gibbon out of things. It is my turn first.” The experimenter used a simple warm-up task to help explain that when the child gets the token out, the child is rewarded with a sticker. E placed the token in a tin and closed the lid, reopened the lid, and took the token out, saying, “That’s a sticker for me.” She then placed the token back into the tin, closed the lid, and put the tin down on the child’s side of the table, saying, “Now it’s your turn.” Once the child removed the token from the tin, E said, “Well done, that’s a sticker for you. Let’s start you a pile.” Next, E placed the token into the Slotbox in sight of the child via the hole in the top of the box, saying, “Now we put him in here, and before we start I would like to introduce you to two of my friends. Here is my friend Tina [E points to the first monitor]. Tina is an adult. Can you see her waving?” As this was said, E played a 5-s clip of the model smiling and waving. The same was then done with the second monitor, “Here we have my friend Sophie [E points to the second monitor]. Sophie is a child, the same age as you. Can you see her waving?” Declaring that the child model was “the same age” was done to avoid children assuming that the model was either older or younger. Model introduction order (first or second) and monitor position (left or right) were counterbalanced. Two monitors were used so that there would be a clear distinction between the two models and the two tools. The following content was dependent on the experimental condition. Children in the no-demonstration condition were told, “Tina and Sophie both played with this, and now it’s your turn and you can do anything you like.” All other children were told, “Tina and Sophie both played with this, and we are going to watch what they did. Let’s watch what Tina, the adult [same order as they were introduced], did when she played with the toy.” E then played the video clip twice. Half of the children saw clips in which both of the models performed irrelevant actions. For the rake tool, the irrelevant action was to tap the rake end at the front opening of the slot box four times. For the arrow tool, the flat surface of the tool was slid down the back of the box four times. These actions were performed after the token was inserted and before the relevant action. E then said, “So that is what Tina the adult did. Now let’s watch what Sophie, the child, did when she played with the toy.” E then played the child clip twice on the other monitor. The end of both clips showed a picture of the model paired with a picture of the tool for the remainder of the experiment. Video clip duration ranged from 12 to 22 s depending on content. The two video clips shown to the same child (one from the adult and one from the child) never differed by more than 4 s. E recapped by saying, “Now, do you remember that Tina used this tool and Sophie used this tool [E points to appropriate monitors and tools]? Well, now it’s your turn and you can do anything you like.” The participant was given up to 3 min to interact with the Slotbox. There were a number of set prompts in place if 60 s had elapsed and the child had (a) not yet touched any part of the apparatus (“Can you pick up a tool?”), (b) picked up a tool but not made contact with the Slotbox (“Can you play with the toy?”), or (c) moved the token to the front of the task but not opened door (“You can open the door”). If 2 min had passed and there was no success, the child was asked, “Can you get the toy out?” If the child had not touched the box after 3 min or if there was no success after 4 min, the child was told that he or she had done very well and the experiment ended. If the child was successful, E said, “Well done, that’s a sticker for your pile. It’s your turn again. You can do whatever you like.” E took a sticker from a pile and added it to the child’s pile and put the token back in the Slotbox. The child was allowed up to five successes. All children were rewarded with six stickers irrespective of success. Coding and analysis ~~~~~~~~~~~~~~~~~~~ The second author (R.A.H.) coded 100% of the sample. A research assistant, blind to the study’s aims, coded 60 participants (43%) by watching each video from the moment the experimenter first said, “Now it’s your turn and you can do anything you like.” Each participant’s interaction was coded for successful retrieval of the token (yes or no) and method of success (child demonstrated, adult demonstrated, or alternative) on each of the five trials as well as latency to first success (from “Now it’s your turn and you can do anything you like” to full extraction of the token). There was almost perfect agreement (Viera & Garrett, 2005) on all nominal variables (kappa scores > .97). Latency to first success had a moderate intraclass correlation of .54 (p < .003). For within-participant first response and comparisons with the control group, non-parametric tests were used and were two-tailed. Between-participants effects were analyzed using multiple regressions. Levels of success ~~~~~~~~~~~~~~~~~ Of the 140 children, 122 (87%) were successful at retrieving the token within 3 min. Of these, 60 used the arrow tool and 54 used the rake tool (see below for the other 8), indicating no bias toward either tool (p = .64, binomial test). Children who saw a demonstration were significantly more likely to obtain the token than children who did not (p < .001, Fisher’s exact test [FET]); fully 116 of 120 children from experimental conditions were successful, whereas 6 of 20 no-demonstration children were successful. From the no-demonstration condition, 1 child used the arrow tool paired with the adult model for five response trials, and the other 5 children put their hand in the front opening and pulled the token out. Thus, they did not use a tool. Only 3 (2.5%) of the other 116 successful children in demonstration conditions initially used their hands rather than a tool to obtain the token; thus, they were significantly less likely to use their hands on the first trial (p < .001, FET). These 3 children also used demonstrated tool methods in their other response trials. No children from the no-demonstration condition spontaneously produced one of the two irrelevant actions. Because the no- demonstration group had only 1 child who was successful with a tool, this group was removed from subsequent analysis, as were the four unsuccessful children. The 3 children who were successful with hands were excluded from first trial analysis. Latency to first success (in seconds) for those who used either demonstrated method was entered into a linear regression with the same predictor variables. Here the prediction model was statistically significant, F(4, 108) = 7.52, p < .001, and accounted for approximately 22% of the variance of latency to success. Table 2 shows a summary of the predictor variables. Latency to first success was predicted by participant age and demonstration completeness. Older children, and those who saw a more complete demonstration, were more likely to be faster to success. Efficiency and participant sex were not significant predictors. Model age-based bias ~~~~~~~~~~~~~~~~~~~~ The 113 children who were successful with a tool were more likely to use the tool demonstrated by the child model (n = 75, 66%) than the tool demonstrated by the adult model (n = 38, 34%) (p < .001, binomial test) in their first trial. Of the 116 successful children, 114 completed five trials. Across the five trials, the child method was used more often (median = 3, interquartile range [IQR] = 1.5) than the adult method (median = 2, IQR = 2), n = 114, Wilcoxon Z = –2.53, p < .05. Five children used both methods simultaneously at some point, all of them after they had used each of the two methods separately. Irrelevant action reproduction ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Although irrelevant action demonstration from the model was a major focus of the current study, irrelevant action reproduction of the participants was not. Thus, here we give a concise description of children’s behavior following demonstrations of irrelevant actions. No children from the no-irrelevant-action conditions produced an irrelevant action, whereas 37 of 60 children (62%) who watched irrelevant actions from both models produced an irrelevant action of some sort on their first trial. Of these, 4 used a hybrid (e.g., tool of the adult, action of the child), 21 performed the irrelevant actions modeled by the child, and 12 performed those modeled by the adult, which was not a difference (binomial p = .163). Efficacy and efficiency ~~~~~~~~~~~~~~~~~~~~~~~ Fig. 3 shows a summary of each child’s first solution method. A logistic regression was run with model copied (adult or child) entered as the dependent variable and demonstration efficacy (completeness), demonstration efficiency (irrelevant actions), and participant sex and age entered as predictor variables. A test of the full model against a constant- only model was not statistically significant, indicating that the predictors did not reliably distinguish between those who used the child model and those who used the adult model, χ2(4, N = 113) = 3.25, p = 51. None of the independent variables were significant predictors. Conservatism ~~~~~~~~~~~~ The final analysis investigated children’s conservatism to an initial solution. There were an unanticipated 11 children who used their hands to remove the token on at least one trial (including 3 children who did this on their first trial). These 11 children and the 4 who were initially unsuccessful were removed, leaving a total of 105 children (see Fig. 4). A logistic regression was run with method conservatism (yes or no) entered as the dependent variable and efficacy, efficiency, model choice on first trial, and participant sex and age entered as predictor variables. A test of the full model against a constant- only model was statistically significant, indicating that the predictors as a set reliably distinguished between children who copied just one model and children who copied both models (see Table 3). Efficacy, participant sex, and first method were not significant predictors. Age was a significant predictor, with increasing age predicting decreased likelihood of conservatism. Efficiency was a significant predictor, with inefficient demonstrations increasing the likelihood of conservatism. Whereas only 29 of 57 children (51%) who saw an irrelevant action displayed both methods, 45 of 59 children (76%) who did not see an irrelevant action showed both methods (p < .01, FET).","We tested for a model age-based bias in differing contexts by examining children’s first solution choice and also their conservatism in relation to this first solution over five response trials. Our first prediction that children would socially learn from the information provided was supported; children with no demonstration were generally unsuccessful. Context had some effect; demonstration completeness significantly predicted latency to success, with children in the partial condition taking longer to solve the task than children who received more complete information. This would indicate that these children found retrieving the token to be harder yet possible, even with limited information within the demonstrations. We did not make a specific prediction regarding the direction of model biases for age. Our results showed that children tended to preferentially copy the child model both on their first response trial and over time. The use of similar two-action reward retrieval tasks has elicited some variant biases in previous studies; when similar-aged children were shown either an adult or a child demonstrating relevant and irrelevant actions, they were more faithful in their copying of the adult (McGuigan et al., 2011; Wood et al., 2012). However, a bias toward copying peer models has been found when the context was less goal directed and more overtly playful (Ryalls et al., 2000; Zmyj et al., 2012). We presented the Slotbox overtly as a playful game-like activity, and children appeared correspondingly to treat it as such by copying the child over the adult. Furthermore, the device used in Wood et al. (2012) and McGuigan et al. (2011) was much less transparently a toy for which a child would have privileged knowledge, whereas the Slotbox was called a toy, and given that 3- to 5-year-olds are known to select children over adults when the domain is toys (VanderBorght & Jaswal, 2009), this labeling could have created the preference for a child model. Our second prediction, that the model age-based bias would be most pronounced when social information lacked evidence of efficacy and efficiency because the demonstrations would create the most uncertainty, was not supported. Neither efficacy (demonstration completeness) nor efficiency (irrelevant actions) affected the strength of the initial model-based bias toward child peers. The presence of the model age-based bias, when both models gave demonstrations, varies from Wood et al. (2015), who found that children did not distinguish between two peers who differed in previous proficiency when two equally valid solutions were demonstrated. One explanation for this could be that the age contrast in the current study may be considerably more salient than the peer proficiency contrast in Wood et al. (2015) study. For the latter, the models differed in their general levels of proficiency, as indexed by their behavior toward a novel apparatus as well as teacher ratings. This subtler difference between models may have diluted children’s attention to proficiency in the context of the test task. Indeed, the peer ratings of proficiency in Wood et al. (2015) were dominated by age, such that children rated older peers as more proficient than younger peers irrespective of their actual proficiency. Furthermore, previous work has shown that even though children can identify models who “know” versus “don’t know” how to do a task, their imitation of the model is driven more by their age than by their professed knowledge state (Wood et al., 2012). An alternative explanation for the current finding is that the preference for matching the child model, irrespective of demonstration content, was driven by affiliative reasons rather than learning reasons (Uzgiris, 1981). If children copy to affiliate, then it is of no consequence which method is more effective. Affiliative versus learning goals may explain the discrepancy between the current study and Hu et al. (2013), who found that a bias toward copying the majority was lost when children could see that both methods worked. This “second” function of imitation is receiving increasing attention (Over & Carpenter, 2012; Wood et al., 2013b) and may be an important factor in children’s model-based social learning choices. Our third prediction, that conservatism to the initial solution would be greatest when children are uncertain of the alternative method, was partly supported. Children’s continued interaction with the task revealed behavior differences across conditions. Specifically, children who viewed both models using irrelevant actions were significantly less likely to use both methods than children who viewed both models using relevant-only actions. One possible explanation for why demonstrated irrelevant actions discouraged use of the alternate method is that participants may have had more uncertainty about the quality and competence of the demonstrations offered and so continued to use their initial previous solution and were reluctant to try the alternative method. Conversely, when there was a lack of irrelevant actions (and potentially more information regarding success), this led to confidence in both models and, thus, exploration beyond children’s initial bias and personal success. The current results support the hypothesis that a copy when uncertain bias could promote conservatism to an original method and, thus, inhibit exploration of an alternative method, as found with Wood et al. (2013a) and Bonawitz et al. (2011). There was an additional result that was not specifically predicted; the select few children with no demonstration who were successful tended not to use a tool, instead—more efficiently perhaps—using their hand. This result again highlights both the advantage and cost of using social information; demonstration children were much more likely to be successful but also more likely to use a tool that was actually unnecessary (cf. Nielsen & Tomaselli, 2010). We take this as further evidence that when children have no prior information, social information can promote high-fidelity copying of demonstrated actions (Wood et al., 2013a) to the point of copying inefficient components (McGuigan et al., 2007). It would be interesting to investigate what the children with social demonstrations would have done if the tools were removed during their response; would they have been even less likely to succeed than children with no social demonstrations because they were reliant on the social information of using a tool for success? Children’s biases in social learning identified in this study are likely to have wider implications for our understanding of social traditions. In the current study, the most likely context in which a solution was socially learned and persisted over time was when it was performed by a child and when the information was constrained (incomplete) and inefficient. Thus, incipient traditions in the kind of context created in our experiment may be more likely to form when there is limited scope for confidence in the approaches seen in models and there are biases toward learning from individuals (in this case children) with certain characteristics. Conversely, confidence in the quality of competing socially demonstrated solutions may promote motivation for further exploration of alternative approaches witnessed. A combination of social learning and exploration is thought to be the bedrock of cumulative culture; thus, seeing multiple effective solutions from multiple models may be an influential component in humans’ cumulative cultural capacities."],["“Selfies” (self-taken photos) are a common self-presentation strategy on social media. This study experimentally tested whether taking and posting selfies, with and without photo-retouching, elicits changes to mood and body image among young women. Female undergraduate students (N = 110) were randomly assigned to one of three experimental conditions: taking and uploading either an untouched selfie, taking and posting a preferred and retouched selfie to social media, or a control group. State mood and body image were measured pre- and post-manipulation. As predicted, there was a main effect of experimental condition on changes to mood and feelings of physical attractiveness. Women who took and posted selfies to social media reported feeling more anxious, less confident, and less physically attractive afterwards compared to those in the control group. Harmful effects of selfies were found even when participants could retake and retouch their selfies. This is the first experimental study showing that taking and posting selfies on social media causes adverse psychological effects for women. --------------------------------------------------------------------------------","Within the past decade, social networking has become a hugely popular form of online communication, especially among young people (Perloff, 2014). Facebook, Instagram, and Snapchat are among some of the most widely used social media platforms available and can be accessed via computer, smartphone, computer tablet, and through other forms of technology (Perloff, 2014). In comparison to conventional mass media, social media are interactive, allowing individuals to create their own personal profiles and share information and photos with users on their social network (Stefanone, Lackaff, & Rosen, 2011). A national survey by the Pew Research Center found that in the U.S., 18- to 29-year-olds who access the Internet are the most likely of any demographic group to use a social networking (i.e., social media) site, and that women are more likely than men to use these sites (Duggan & Brenner, 2013). Over 95% of college students regularly maintain and manage their social networking profiles (Perloff, 2014; Stefanone et al., 2011). Women, in particular, have been found to upload photos to social media more frequently than do men, and tend to spend more time updating, managing, and maintaining their personal profiles (Stefanone et al., 2011). Emerging evidence provides insight into the effects that social media behaviours may have on users. On one hand, social media use may be beneficial as it allows greater connectedness with others, leading to an increased sense of well-being (Tiggemann & Miller, 2010). On the other hand, social media use may lead to a preoccupation and focus on physical appearance, such as engagement in appearance-related photo activities (Cohen, Newton-John, & Slater, 2017), which could cause appearance concerns and lowered body image and self-esteem (de Vries, Peter, Nikken, & de Graaf, 2014). As users are frequently exposed to a variety of other profiles, they can compare their own appearance to friends, relatives, and strangers (Haferkamp & Kramer, 2011). Hancock and Toma (2009) found that people select their own online dating profile photos in an attempt to look as attractive as possible without being judged to be deceptive. Cross-sectional data have revealed that for both women and men, Facebook use is associated with greater (upward) social comparison and self-objectification, which are both related to lower self-esteem, poorer mental health, and body image concerns (Hanna et al., 2017). Social media and body image ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Various studies have documented widespread body and weight dissatisfaction among girls and women, and social media has been found to be a significant catalyst for these appearance concerns (Brown & Tiggemann, 2016; Holland & Tiggemann, 2016; Tiggemann & Miller, 2010). Given that social media provide the opportunity for social comparison, as well as exposure to unrealistic beauty expectations, body dissatisfaction is likely to result from frequent use (Fardouly, Pinkus, & Vartanian, 2017; Tiggemann & Slater, 2013; Want & Saiphoo, 2017). Social media present innumerable idealized images of thin, lean/tone, beautiful, photo- shopped women, and the “thin ideal” and “athletic ideal” are displayed as a normal, desirable, and attainable body type for every woman (Kim & Chock, 2015; Meier & Gray, 2014; Robinson et al., 2017). Furthermore, the Internet and social media have been found to promote thinness, dieting behavior, and weight loss through idealized images of “perfect” women (Perloff, 2014). Women who use social media often internalize the “thin ideal,” causing them to strive for an unrealistic, unnatural standard of beauty and to feel ashamed when they are unable to achieve it (Kim & Chock, 2015; Meier & Gray, 2014; Tiggemann & Slater, 2013). Studies have found that frequent exposure to the Internet and social networking websites results in high levels of weight dissatisfaction, drive for thinness, and body surveillance in young women (Tiggemann & Miller, 2010; Tiggemann & Slater, 2013), regardless of race (Howard, Heron, MacIntyre, Myers, & Everhart, 2017). Additionally, Perloff (2014) suggests that women who have relatively higher levels of thin ideal internalization, perfectionism, and/or low self-esteem would be especially likely to spend time on appearance-focused online comparisons and that they probably do not use ‘self-protective’ downward appearance comparisons (i.e., comparing their appearance to less attractive friends). These predictions are concerning, since high body dissatisfaction among women is a primary risk factor for the development of eating disorders and is correlated with low self-esteem and depression (Meier & Gray, 2014; Tiggemann & Miller, 2010). Therefore, it is important for researchers to understand the causal effects that social media and self-presentation strategies have on young women by using experimental research methods. Self-presentation and impression management ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Past research on the psychological effects of social media has mainly focused on the implications of social media use for body satisfaction in general. However, there is a lack of empirical research that evaluates the effects of the specific self-presentation strategies that social media users rely on. According to Toma and Hancock (2010), self- presentation involves “adjusting and editing the self during social interactions to create a desired impression on the audience.” The motivation to selectively self-present also relates to impression management, whereby individuals carefully present themselves in order to make specific impressions on their viewers (Pounders, Kowalczyk, & Stowers, 2016). As a result, social media users are driven to present the most attractive versions of themselves to others in order to make a favorable impression (Toma & Hancock, 2010). These photos, however, often do not portray an accurate depiction of one’s true physical appearance (Toma & Hancock, 2010). The most common way that users selectively self-present on social media is through the taking and uploading of “selfies” (photos taken by and of oneself). Users tend to capture selfies from flattering angles and using bright lighting, and may also edit their photos using colour correction, skin-retouching, and even photo- shopping to make body parts appear thinner (Anderson, Fagan, Woodnutt, & Chamorro- Premuzic, 2012). In this way, social media users are able to manage the impressions they have on others by presenting only the most flattering images of themselves and minimizing perceived flaws or imperfections (Anderson et al., 2012; Bell, Cassarly, & Dunbar, 2018; Pounders et al., 2016). It has also been found that individuals who desire to boost their self-esteem upload selfies more frequently, and that women of 16–25 years of age spend up to 5 h per week taking selfies and sharing them on social media (Pounders et al., 2016). Research on gender differences in Internet activities has found that, compared to men, women tend to be more motivated to create a positive self-presentation on their social media profiles, and as a result, they engage in more photo-enhancement behaviours (Haferkamp, Eimler, Papadakis, & Kruck, 2012; Toma & Hancock, 2010). Overall, research has suggested that the taking and retouching of selfies may be a particularly risky behaviour in terms of its potential to negatively impact the body image and self-esteem of young girls and women. The current study ~~~~~~~~~~~~~~~~~ In summary, previous research demonstrates that social media use is positively correlated with appearance concern. Furthermore, the literature suggests that selfie-taking and photo-retouching, which are very common social media behaviours, are associated with poorer self-esteem and body image among young women. It has been suggested that editing and uploading selfies may worsen appearance concerns (de Vries et al., 2014), but it is not yet known whether a causal relationship exists. To fill this gap in the literature, the current study tested the effects of selfie taking on body image and mood in women. It was hypothesized that updating one’s social media profile with a selfie photo would result in lowered mood and increased body concerns as compared to a control group. To answer a secondary research question, we also tested the effects of having control over self- presentation on social media, by retaking and retouching a selfie photo, on women’s body image and mood. It was hypothesized that participants who were allowed to retake and retouch their selfie would experience better mood and body image compared to women who were not allowed to modify their selfie before posting it on social media. This is because women typically react to seeing a photo of themselves by feeling dissatisfied with their appearance (Mills, Shikatani, Tiggemann, & Hollitt, 2014) and photo modification allows a person to present an idealized version of themselves to others (Tiggemann & Miller, 2010).","Participants were 113 psychology undergraduate students recruited through an online experiment management system at York University in Toronto, Canada. Inclusion criteria included being female, being between 16 and 29 years old (M = 19.00, SD = 1.66), and having an active account on Facebook or Instagram. In exchange for their participation in a single, hour-long lab session, participants received partial course credit toward their Introduction to Psychology course. The self-reported ethnic distribution of the sample was 24.8% South Asian, 20.2% European/Caucasian, 12.8% Black/African-American, 10.1% Middle Eastern, 9.2 Caribbean, 6.4% Pacific Islands American, 5.5% East Asian, 2.8% Latino/ Hispanic, and 8.2% other ethnic identification. Body mass index (BMI = kg/m2) scores ranged from 15.84 to 36.23 (M = 23.71, SD = 4.03) across the sample, with the mode, median, and mean all falling within the “normal” weight range (18.5 < BMI <24.9) (Centers for Disease Control & Prevention, 2015). One participant who mistakenly signed up for the study was excluded because he self-identified as male. Two participants declined to participate after reading the informed consent form because they were uncomfortable taking a photo of themselves for religious reasons. iPad Participants used the Internet browser, camera, and photo modification app (“You- Cam Now”), if applicable, installed on an iPad. Mood and body image A series of visual analogue scales (VAS) was used to measure mood and body image at baseline as well as after the experimental manipulation (described below). This commonly used set of scales was designed to assess pre-post fluctuations in psychological states, typically in experimental research designs (Heinberg & Thompson, 1995). The measure consisted of six VAS, each with a 10-centimeter horizontal line labeled with a specific attitude or emotional state. Participants are asked to place an X on the point on the line that most accurately depicts the degree to which they were experiencing that feeling at the moment, from Not at all to Very much. The mood items included anxiety, depression, and confidence. The body image items included feelings of fatness, physical attractiveness, and body size satisfaction. Rather than collapsing scores into global affect or appearance concerns, we separated the items so that we could examine specific affective changes among participants. VAS format is recommended over Likert scales for pre- post research designs since it reduces recall bias (i.e., participants cannot recall their previous response), can be completed quickly, is sensitive to emotional changes (Hargreaves & Tiggemann, 2003). The measure used in the current study is the same one used in other published studies. Demographics Age and race/ethnicity demographics were collected from each participant. Filler items not of interest to the study were included on the questionnaire (e.g., living arrangements, year of study, university program, and media consumption).","Ethics approval was received from York University’s Human Participants Review Committee. Female undergraduate students volunteered for an advertised study examining “the relationship between personality and social media use.” Participants were tested individually behind a partition wall from the experimenter and were asked to leave their bags and any personal electronic devices (including phones) outside of the testing area. Participants were randomly assigned to one of three experimental conditions prior to arriving at the lab. Upon arrival to the lab, participants read and signed a written informed consent form, were given a baseline VAS, and then the demographics questionnaire with additional filler items to distract from the purpose of the study. For ethical reasons, the informed consent form contained the information that participants may be asked to post a selfie to their own social media profile. For the experimental task, participants in the Untouched Selfie condition were asked to take a single photo (a headshot) on the lab’s iPad and upload it to their preferred social media profile (Facebook or Instagram). Participants in the Retouched Selfie condition were asked to take one or more photos of themselves on the lab’s iPad and were told that they could use the photo editing app installed on the iPad to retouch the photo to their satisfaction before uploading it to their social media profile. Participants in the Control condition were also given the lab’s iPad but were asked to read a short article from a social media news website chosen for neutral, non-appearance related content (i.e., popular travel ideas for university students) and to answer questions about the article. This task was chosen to maintain the cover study of social media use and to control for using an iPad, and for the amount of time elapsed between pre-post measures. It was intentional that Control condition participants not engage on Facebook or Instagram (theirs or other people’s profiles, since we could not be certain that they were not exposed to appearance-related content, which could affect mood and/or body image). The assigned tasks in the Untouched Selfie and Control conditions were timed (5 min each). The Retouched Selfie condition was not timed so that participants could retake and retouch their selfie to their satisfaction. However, time to completion was recorded by the experimenter and participants in the Retouched Selfie condition took a similar amount of time to complete their task (mean time to completion = 4.5 min). Instructions and set up in all three conditions took approximately 1–2 min. As manipulation checks, Control condition participants were asked to answer written questions about their article to ensure that they read the article. Selfie condition participants were asked verbally by the experimenter whether they completed the tasks as instructed. In addition, at the end of the study the experimenter checked the photo and browser histories, and any deleted files on the iPad to ensure that participants in all conditions adhered to the instructions and did not open any other websites or social media profiles. All participants confirmed that they followed the instructions and there was no evidence of non-adherence. Upon completion of the experimental tasks, all participants completed the post-manipulation VAS. Participants were asked to complete the scales based on how they were feeling at that particular moment. The elapsed time between the baseline and post-manipulation VAS measure was approximately 10 min. Furthermore, the format of the VAS scale is such that participants cannot recall their previous answer; thus, recall bias is minimized. Participants were then debriefed and probed as to what they believed to be the purpose of the study. Lastly, height and weight were measured by the experimenter on a balance beam scale. Data analysis Statistical analyses were conducted using SPSS version 24. An alpha level of .05 was used for significance testing. A power analysis was conducted using G*Power (Faul, Erdfelder, Lang, & Buchner, 2007); an alpha of .05, medium effect size, and power estimate of .80 resulted in a recommended sample size of 110, which was obtained. Repeated measures analysis (Time 1 – Time 2) was chosen to analyze the effects of experimental condition instead of VAS change scores to maximize power and use within-subject error estimates. To control for Type I error, an initial repeated measures multivariate analysis of variance (RM-MANOVA) was performed with time (Time 1 – Time 2) and test (VAS item) as the within- subject factors, and experimental condition (Untouched Selfie, Retouched Selfie, and Control) as the between-subjects factor. Any significant multivariate 3-way interaction (time × test × condition) on the combined dependent measures was followed by univariate repeated measures ANOVAs, with time (Time 1 – Time 2) as the within-subject factor and experimental condition as the between-subjects factor. Any significant within-subjects contrasts (time × condition) were followed by post-hoc t-tests to examine which conditions differed. For ease of interpretation, change scores (Time 1 – Time 2) were used only for these post hoc t-tests to examine the direction and magnitude of change to psychological states as a function of condition. Preliminary analyses ~~~~~~~~~~~~~~~~~~~~ Inspection of histograms, skewness, and kurtosis suggested that all of the variables were normally distributed. There were no statistical outliers (± 3.0 SD) among the dependent variables; therefore, no adjustments were made. Groups did not differ significantly on baseline levels of any variable, suggesting that randomization resulted in equivalent groups. Multivariate effects of experimental condition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Means and standard deviations for all dependent variables of interest (pre- and post- manipulation) as a function of the experimental condition are shown in Table 1. For ease of interpretation, Table 1 also shows the change in participants’ self-ratings across the psychological states. A significant 3-way (test × time × condition) multivariate effect on the combined dependent variables was found, Hotelling’s Trace = .21, F(10, 201) = 2.14, p = .02, partial η2 = .10, meaning that the experimental groups differed with respect to how mood and body image ratings changed between Time 1 and Time 2. Significant 2-way (time × condition) interactions were found for anxiety, Hotelling’s Trace = .06, F(2, 107) = 3.32, p = .04, partial η2 = .06, confidence, Hotelling’s Trace = .07, F(2, 107) = 3.69, p = .03, partial η2 = .07, and physical attractiveness, Hotelling’s Trace = .07, F(2, 107) = 3.59, p = .03, partial η2 = .06, meaning that the experimental groups were not equal with respect to changes on those items from Time 1 to Time 2. Interactions were not significant for depression, Hotelling’s Trace = .01, F(2, 107) = 0.48, p = .62, feelings of fatness, Hotelling’s Trace = .02, F(2, 107) = 0.97, p = .38, or satisfaction with body size, Hotelling’s Trace = .01, F(2, 107) = 0.75, p = .47. Changes to psychological states as a function of condition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The significant 2-way interactions reported above were followed up with t-tests to compare changes to psychological states across experimental groups. As can be seen in Fig. 1, participants in the Untouched Selfie condition experienced an increase in anxiety and this was significantly greater than the Control condition t(71) = 2.35, p = .02. The Retouched Selfie condition also experienced an increase in anxiety but was not significantly different from the Control condition, t(71) = 1.80, p = .08. The Untouched and the Retouched Selfie conditions did not differ with respect to changes in anxiety, t(71) = 0.79, p = .43. Fig. 2 shows that participants in the Untouched Selfie condition experienced a decrease in confidence and this was significantly greater than the Control condition, t(71) = 2.48, p = .01, and marginally greater than that experienced in the Retouched Selfie conditions t(72) = 1.92, p = .06. There was no difference in changes to feelings of confidence between the Retouched Selfie and Control conditions, t(71) = 0.60, p = .55. Fig. 3 shows that participants experienced decreases in feelings of physical attractiveness that were significantly greater than in the Control condition in both the Untouched Selfie condition, t(71) = 2.43, p = .02, and the Retouched Selfie condition, t(71) = 2.32, p = .02. These decreases were equivalent between the two selfie conditions t(72) = 0.12, p = .90.","This is the first experimental study of the causal effects of posting selfies to social media on young women. The findings generally supported our hypothesis that taking and posting a selfie on social media would result in lowered mood and worsened self-image. We also found that women who had the opportunity to retake and modify their selfie before posting it to social media still experienced decreases to mood and anxiety that were similar to the reactions of those who could not retouch their photo. Participants who took and uploaded a selfie onto social media, without the option to retouch or take multiple photos, felt more anxious, less confident, and less physically attractive afterward, and these differences were significantly greater than the control condition (i.e., reading a neutral news article online). These results all yielded medium effect sizes. These findings are consistent with the previous suggestion that appearance concerns are heightened when women interact with and construct their social media profiles, manifesting in poorer body image and mood (e.g., de Vries et al., 2014). However, we did not find significant effects of selfie-taking on all of the dependent variables of interest in the current study; we found null effects on state feelings of fatness, satisfaction with one’s body, and depression. We interpret these findings to suggest that the psychological states affected by taking and posting selfies to social media are specifically related to feelings of self-consciousness and/or fear of negative evaluation by others. This interpretation seems likely given that participants in the study were sharing their selfie photos on their own social media profiles and for other people they know to see. It is interesting that feelings of physical attractiveness were negatively affected by selfie taking and posting, but not feelings of fatness or satisfaction with one’s body size. However, it is important to note that the current study involved taking a photo only of one’s head and face. In other words, it may not be surprising that effects of taking a selfie on body-related constructs were not found, since the current study looked only at the effects of taking selfies of one’s face. If the current study had examined the effect of taking and posting photos that showed the participant’s body the results might have been different. Celebrities, but probably many social media users, often post body- conscious selfies on their social media (e.g., wearing bathing suits, lingerie, or no clothing at all). Posting selfies of one’s body (and not just the face), even when clothed, could trigger body-specific appearance concerns but we did not capture those effects in the current study. This is an area for future research. We had a secondary research question related to whether being able to retake, select, and modify one’s selfie (as is commonly done by many social media users) might, in fact, improve subsequent mood or body image. As suggested by Kim and Chock (2015), women are motivated to present perfected images and idealized versions of themselves on their social media profiles in order to make a favorable impression on their viewers. Photo-retouching behaviours allow women to present the most attractive versions of themselves and minimize perceived imperfections (Toma & Hancock, 2010). In the current study, women in the retouched selfie condition were able to take multiple photos, delete unwanted photos, and could retouch their photos to their satisfaction using a photo editing application. However, we found little evidence of any psychological benefit of being able to modify the photo women posted to their social media. In terms of state anxiety, women who posted an untouched selfie to social media felt significantly more anxious than those who did not post a selfie at all. But women who were able to retouch their selfie before posting it also felt marginally more anxious than those in the control condition and equally anxious to those in the untouched selfie group. In other words, having the ability to retake and retouch their selfie to their satisfaction before posting it did not mitigate women’s anxiety significantly. This lack of difference between the effects of the two experimental selfie tasks on anxiety was unexpected. A similar result was found regarding feelings of physical attractiveness. Participants who could retouch their selfie felt significantly less attractive after posting it online (as did those who were asked to post an untouched selfie), and there was no significant difference between the retouched and untouched selfie groups on changes to feelings of physical attractiveness. In terms of feelings of confidence, women who could retouch their selfie did feel more confident afterward than those in the untouched selfie group, but they felt just as confident as those who did not post a selfie at all. In other words, posting a retouched selfie did not improve women’s confidence, as compared to engaging in an appearance-neutral task. To explain these findings, it could be that scrutinizing and modifying images of themselves makes women think more about their flaws or imperfections. Retouching could activate feelings of self- objectification. Even though self-presentation strategies like photo-editing provide a sense of control over physical appearance (Tiggemann & Miller, 2010), they do not actually appear to improve mood or self-image. The current study found no evidence that posting retouched photos to social media makes women feel better than usual and found some evidence that it makes them feel worse than usual. Although women might feel less anxious about posting a selfie if they have the chance to retouch it and make it more flattering, the process of taking and editing the photo still draws their attention to feeling dissatisfied about aspects of their appearance. Clinical implications ~~~~~~~~~~~~~~~~~~~~~ These findings have clinical implications for the prevention and treatment of mental health difficulties. Women who took a selfie and posted it to their social media profile had increased levels of anxiety, decreased confidence, and lowered perceived physical attractiveness compared to those who did not take a selfie. Given that women between 16–25 years of age spend up to 5 h per week taking selfies and uploading them to their personal profiles (Pounders et al., 2016), these findings raise significant concern about social media use and well-being. Posting selfies to one’s social media has adverse causal effects on the self-image and mood of young women, and could make them more vulnerable to clinical eating, mood, and/or anxiety disorders. Frequently taking selfies could be considered a body checking behavior, such as repeated weighing and recurrent checking of one’s reflection in mirrors (Mills et al., 2014). As a result, frequently taking and posting selfies should be considered a risky online health-related behavior for young women in terms of mental health, especially if they trigger weight and shape dissatisfaction. High body dissatisfaction is the primary risk factor for the development of eating disorders and is correlated with low self-esteem and depression (Meier & Gray, 2014; Tiggemann & Miller, 2010). Interventions that aim to diminish or eliminate the harmful effects of social media engagement on psychological functioning should be validated and implemented. Limitations and future directions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A unidirectional causal relationship between posting selfies to social media and worsened mood and body image was demonstrated in the current study. The reverse relationship – the effect of low mood or body dissatisfaction on posting selfies – is a future research question of importance. There could be a bidirectional and self-perpetuating cycle between appearance-based social media engagement and negative mood and/or body image. Because the current sample included only young women who regularly use social media, these results may not generalize to older women or women who do not use social media. We did not include men since the existing literature on social media and body image has focused on women. Future studies should include men and relevant appearance-related psychological constructs (e.g., drive for muscularity; Mills & D’Alfonso, 2007). For ethical reasons, participants were informed on the consent form that they may be asked to take a photograph of themselves and post it to social media. We attempted to minimize demand characteristics by including filler questions between repeated measures, by using the visual analogue scale format, by stating the purpose of the study in only vague terms, and by probing what participants thought to be the true purpose of the study. There was no evidence that demand characteristics were a threat to the validity of the study. Nevertheless, it is possible that participants might have had implicit assumptions about the effects of the experimental tasks on how they felt. We did not examine personality moderators in the current study; future research should investigate individual differences and whether certain types of women (e.g., those who are high on perfectionism, those who frequently post selfies in their everyday lives) are more or less vulnerable to the adverse effects of posting selfies to social media than others. Future research should also study the specific modifications that participants make to their photos using retouching. We did not include this outcome variable in the current study, but future studies could explore ways of assessing selfie modification behaviours surreptitiously. Participants in the control condition of the current study did not interact on social media to avoid any possible exposure to appearance-related online content and to make the control condition entirely appearance neutral. However, a different control task (e.g., uploading a neutral, non- selfie photo to social media) might produce different results. Therefore, an important next step is to dismantle what aspects of posting selfies to social media produce the observed effects (e.g., taking selfies without posting them on social media). Finally, future research should examine the longer term and/or cumulative effects of posting selfies to social media using prospective, longitudinal research designs.","This is the first study to show experimentally that selfie posting on social media is harmful in terms of young women’s mood and self-image. Being able to retouch or modify their photo did not result in women feeling better about themselves after posting a selfie to social media. Future research should look at the longer-term effects of posting photos of oneself on social media, which is an increasingly common aspect of contemporary media use."]]